Sunday, 30 January 2011

Alignment of NMR spectra – Part II: Binning / Bucketing

In my last post, I wrote that spectra of biological samples are usually poorly aligned due to wide changes in chemical shift arising from small variations in pH or other sample conditions such as ionic strength or temperature.
The most widely used method of addressing this chemical shift variability across spectra is by means of the so-called binning (or bucketing), procedure that consists in segmenting a spectrum into small areas (bins / buckets) and taking the area under the spectrum for each segment. Preferably, the size of the bins should be large enough so that a given peak remains in its bin despite small spectral shifts across the spectra, but not so large as to include peaks belonging to multiple compounds within a single bin.
As a simple example to illustrate how binning works, let’s consider the spectrum of Taurine (Fig. 1)


Fig. 1: 1H-NMR spectrum of Taurine synthesized with Mnova NMRPredict. Only the spectral region corresponding to the methylene protons is shown.


Taking the spectrum shown in Fig. 1 which has been predicted using Mnova NMRPredict, seven additional spectra were created by changing the chemical shift of the CH2 protons randomly in an effort to simulate the chemical shift variability observed in real life biofluid NMR spectra.

Fig. 2: Synthesized data set comprised by 8 simulated spectra of Taurine with random chemical shifts for the CH2 protons and displayed in superimposed mode in Mnova.

These spectra have been synthesized using 32768 data points and a spectral width of 6001.6 Hz with a spectrometer frequency of 500.13 MHz. If the size of each bin is set to 0.02 ppm (represented by the vertical grid lines in Fig. 2), this will result in the generation of 6001.6 / (0.02 x 500.13) = 600 bins.

When the binning command is issued in Mnova, a new spectrum with 600 data points in which every point is the sum of all the points within each bin is produced. The result of this binning or bucketing operation applied to one single spectrum of the synthetic Taurine data set is depicted in fig. 3, where the circles correspond to the area of each bucket in the original spectrum. Fig. 4 shows the result applied to all spectra in superimposed mode. Digital resolution of the resulting binned spectrum is 10 Hz/point

Fig. 3: Methylene region of one synthetic 1H-NMR spectrum of Taurine after data reduction by uniform binning


Fig. 4: Result of applying data reduction by uniform binning to the 8 1H-NMR spectra of Taurine

Once the spectra have been binned, they are ready to be exported in a convenient format (e.g. ASCII) for further statistical analysis (e.g. PCA).
It can be noticed that binning greatly minimizes the effects from variations in peak positions (in this case, all peaks get perfectly aligned). Additionally, binning reduces the data size for multivariate statistical analyses, although today’s computers and optimized linear algebra algorithms are able to handle large data volumes very efficiently.

The major drawback of this procedure is the loss of a considerable amount of information enclosed in the original spectra. In this particular case, the fine structure of the two triplets is totally lost (the coupling constant is 6.6 Hz whilst the digital resolution is 10 Hz), precluding the direct interpretation of multivariate models. In addition, peaks moving on borders between bins might cause artifacts. Another source of loss of information occurs, for example,when peaks belonging to several compouns are included within a single bin.

There exist several better alternatives to binning, typically involving some form of peak alignment without data reduction. But this will be the subject of my next post …

Thursday, 27 January 2011

Alignment of NMR spectra – The problem: Part I

The chemical shift is of great importance for NMR spectroscopy because it reflects the chemical environment of the nuclides under observation providing detailed information about the structure of a molecule.
Although the chemical shift of a nucleus in a molecule is generally assumed to be fairly stable, there are a number of experimental factors (pH, ionic strength, solvent, field inhomogeneity –bad shimming, temperature, etc) which might produce slight or even quite significant variations in chemical shifts.
This is particularly important in metabonomics/metabolomics where shifts of NMR peaks due to differences in pH and other physico-chemical interactions are quite common in NMR spectra of biological samples. For example, some important metabolites, such as citrate or taurine, have peaks whose chemical shifts fluctuate in an uncontrolled way from sample to sample. These variations can cause spurious grouping of samples in chemometric models.

Example of peak position variation in the citrate region (simulated data)

Whilst it is critical to setup the experimental conditions in the best way to minimize these chemical shift fluctuations (for example by using an appropriate buffer; BTW, there exists a standard protocol for biofluid [urine, serum/plasma] and tissue sample collection and preparation as described by Beckonert et al. [1]), spectral misalignments may still occur and special post-processing methods have to be employed.
Another example in which variation in the chemical shift is important occurs in the context of kinetics or reaction monitoring experiments by NMR. For example, consider the following reaction monitoring example [2]:

Reaction monitoring data set for the solution of phenylethylamine and 2-methoxyphenyl acetate in D2O, with every 35th spectrum from the first (bottom) to the last (top) shown (see [2])

It can be appreciated that during the course of the reaction, the chemical shifts of several signals change as a result of the change in pH (in this case, as a hydrolysis proceeds)
Although characterizing these chemical shifts fluctuations can be sometimes important (pH or drug binding-induced chemical shifts, for example) in general they obscure the process of pattern recognition (metabonomics) and impede the performance of data analysis (e.g. selection of the peaks whose intensities/heights need to be monitored becomes more difficult).

In my next posts, I will cover different ways to deal with the peak misalignment problem, first in the field of metabonomics and then in reaction monitoring.

References:
[1] O. Beckonert, H.C. Keun, T.M. Ebbels, J. Bundy, E. Holmes, J.C. Lindon, J.K. Nicholson, Metabolic profiling, metabolomic and metabonomic procedures for NMR spectroscopy of urine, plasma, serum and tissue extracts, Nat. Protoc. 2 (2007) 2692–2703

[2] M. Khajeh, M. A. Bernsteinb, G. A. Morrisa, Magn. Reson. Chem. 2010, 48, 516–522




Wednesday, 26 January 2011

Here I am again!

As you've undoubtedly noticed there has been little activity on my blog lately, which contrasts with the high activity I'm having in my real life, with a plethora of exciting (and challenging) new projects going on in my company, Mestrelab Research.
One particular area in which we have been working on quite intensively for the last few months belongs to the broad subject of the alignment of NMR spectra. This appears to be a very important topic for those scientists working, amongst others, on fields like metabonomics/metabolomics and reaction monitoring by NMR. Starting from today, I will start blogging about this issue, firstly covering some very basic concepts and then moving on to some more advanced techniques for the efficient alignment of NMR spectra.

So stay tuned!

Saturday, 23 October 2010

Conformational analysis of cyclic compounds using Mspin and RDCs

On the occasion of the release of a new version of Mspin (BTW, this is the very first multiplatform version of Mspin: it works now in Windows, Mac OS X and Linux), I would like to bring into your attention one of the many applications where this software plays an instrumental role: The application of Mspin to the study of seven-membered rings compounds by NMR.

The NMR study of seven-membered ring compounds is a classical problem in conformational analysis. They are commonly studied by means of NOE-based experiments or 3J coupling analysis using Karplus-Altona relationships. In a recent work, recently published in Chemical Communications, [ Chem. Commun. , 2010, 46, 5879–5881 ]from the groups of Roberto Gil (Carnegie Mellon University) and Navarro-Vázquez ( Universidade de Vigo ) have demonstrated that the conformation of a 3-benzazepine compound can be completely determined by using 1DCH residual dipolar couplings (RDCs). These RDCs were easily obtained by performing HSQC experiments coupled in the direct dimension using a polydimethylsiloxane gel as oriented medium


Conformational search indicated the presence of 11 possible conformations for this molecule. In agreement with computed DFT energies for these conformers, as well as observed 3J couplings, chemical shifts, and NOE's RDC analysis shows the preference of the system for a crown-chair conformation with equatorial disposition of the substituents.
Here you can download the rdc data file ( click here ) in the MSpin format and a multiconformer XYZ file with the DFT optimized conformations (click here) . Load them into the RDC module of Mspin, select the singular value decomposition method (SVD), just click the calculate button and see how convenient is to perform the RDC analysis with Mspin

Friday, 30 July 2010

New Fast NMR technique

Instrument time is precious and a plethora of different fast NMR experiments are continuously being proposed in order to reduce the time required to record an NMR spectrum. Actually, money is not the only reason, there are many other factors which motivate the development of techniques to increase the speed of data collection. For example, if one wants to make real-time studies of kinetic processes or protein folding, it’s pivotal to speed up the acquisition of NMR data, in particular multidimensional spectra.
On this issue, we have just put our bit into this field and published an article which describes the use of localized spectroscopy for parallel multidimensional NMR data acquisition. The key idea is to interleave the data acquisition at a variety of localized bands within a given interscan repetition time.



In other words, the method is based on MRI-type slice selection techniques (e.g. spinecho multislice and gradient echo multislice) where nuclear spins in different parts of the tube are excited and detected during subsequent transients while the previously used spins have time to relax towards equilibrium before being excited again, hence achieving a considerable timesaving in the overall acquisition.
We believe that this method; named PALSY, is a very powerful yet simple and general technique to reduce experimental time. Of course, there is a sensitivity penalty approximately proportional to the number of slices chosen, but the good thing is that the achievable resolution in any dimension is not compromised in any way.
Another point of interest is that it does not require any fancy data processing, just a simple data shuffling operation needed to extract the different sub-spectra contained into the acquired raw data matrix. This operation has been implemented into the Mnova Alpha version and will be available in the next official release. Meanwhile, as always, should anyone be interested in evaluating this alpha version, just drop a comment here and I will get in touch to provide an executable.
The article can be accessed here:

Fast multidimensional localized parallel NMR spectroscopy for the analysis of samples

I would like to take this opportunity to acknowledge and congratulate my friend Manolo for this work. He is the intellectual father of this pulse sequence and is currently extending this idea further to cover other NMR experiments.

Tuesday, 27 July 2010

Riding up the peaks …

… around the Pyrenees. NMR peaks are not the only ones that interest us in Mestrelab, we also love the Cols in the Tour de France, and I have some proof :-) :



Quoting a friend of mine: I had never seen the Yellow Jersey before. It is a bit surprising that the competitors would fight so hard for the right to wear it :-)

Even though I’m not a great fan of Lance Armstrong, I reckon he has incredibly popularized cycling in the US. The guy in the photo has been following him for several years already in the Tour de France, running by the riders. It looks easier than it actually is, as the riders go faster than 20 Km/h (i.e. 3 min/Km) in fairly hefty slopes, so you have to be in a pretty good shape to keep up with them for 200 m (that is approximately the distance he is doing with them).



Well, enough for this off topic post. I will follow up later on tonight with some real NMR stuff.

Friday, 2 July 2010

Fermentanomics


We are now in an era in which a plethora of domains of sciencific investigations are combined with the suffix ‘omics’. These include neologisms such as genomics, proteomics, metabonomics, pharmacogenomics, nutrigenomics and many others.
Scientists at Lilly have now established the basis of a new –omics related discipline which they have dubbed Fermentanomics and consists in a new rapid and robust NMR method for monitoring mammalian cell cultures.
This work has been published as a JACS communication and I’m delighted to see that they have used our Global Spectral Deconvolution (GSD) technique available in Mnova for the extraction of the concentrations of the components from the NMR spectra of the spent media of mammalian cell culture

Friday, 21 May 2010

Bruker Smiles

The above title is not aimed to mislead and it does not refer to Bruker’s state of mind: Quite simply, the purpose of this post is merely technical and the relation between Bruker and its smiles will be apparent in a moment… Just keep reading…

Because it wouldn't make sense otherwise, NMR instruments use receiver systems equipped with digital filters since a relatively long time ago. The advantages of such digital filters (generally designed as low-pass filters and applied together with oversampling and decimation methods) are many fold, ranging from higher quality spectral baselines to SNR and effective dynamic range improvements, enhanced reduction of potential sources of folded signals, etc

It´s not all about advantages though … I’m sure most of you are already very well aware of the pesky problem that is infamously known as group-delay artifact in Bruker (and Jeol) data which has plagued the NMR community since these companies switched to digital receivers. In short, the FID resulting from the digital filter does not start at time = 0 but only after a long and slowly rising oscillation of length G (G = Group Delay).



Some empirical procedures to correct it were presented on the internet but they are palliative and do not resolve the problem completely, particularly when apodization is applied.
Typically and depending on how the FID is processed, the spectrum might exhibit smiles (baseline artifacts pointing up) or frowns (baseline artifacts pointing down) at the outer regions of the spectrum as depicted below:




The ultimate solution

These small artifacts are in general not a big problem as one could use a spectral width large enough so that the peaks of interest in the spectrum will not be affected by these artifacts (although some processing algorithms such as backward Linear Prediction could be somewhat problematic with the Group Delay). In any case, we did not feel very comfortable with present solutions to this problem. A few months ago, I went for dinner with Stan and right after it, the power of the red wine and above all, the Galician octopus inspired Stan in such a way that he managed to understand the engineering drawback and proposed a new correction algorithm which we implemented together in Mnova just a few minutes later (whilst still under the influence of the wine :-) ).
Basically we have now a new pre-processing algorithm that corrects in a totally automatic way any Bruker FID corrupted by the group-delay artifact, producing a normal and physically correct FID so that the smiles will not be seen in the f-domain spectrum. The performance of the new algorithm is illustrated in the figure below:




This enhanced correction is available in Mnova since version 6.1.1 onwards, although it is not the default processing method for the moment. In order to activate it, it is necessary to select it via Processing/Group Delay menu command.

I guess the take home from the story is never underestimate the power of red wine and Galician octopus :-)

Thursday, 20 May 2010

Back in the blogosphere

After a 3 month-long hiatus, I'm back in the blogosphere. I’ve been travelling and working very hard on several exciting projects, but my entries have consequently suffered.

Although my workload has not decreased, I am not planning to travel for the next few weeks. Hopefully I will manage to blog on a more regular basis from now on, especially now that there are many things we have been working on lately that I hope will be of interest to the NMR community.

For now, I’d just like to point out a very interesting blog entry written by my friend Stan on his well-known NMR blog. One of the tools in Mnova for which we are most proud of is GSD (Global Spectral Deconvolution) which, as any other fitting process, requires the definition of a line shape model. GSD uses a Lorentzian model and some of the reasons for this choice have been elegantly exposed on his blog.

Why spectral lines are Lorentzian

Thursday, 4 February 2010

Learning NMR with Deep Purple

It’s not that I’ve gone crazy (well, I hope not :-) ) or that Deep Purple has moved from Hard Rock to Science (or at least I‘m not aware of this), but after my last post about the acoustic reproduction of NMR FIDS, I thought that it would be fun to compose a simple song and, at the same time, create some stuff which can serve as an educational tool for some very basic NMR.

I won’t actually compose anything original, but rather make an NMR cover of the famous Smoke on the Water riff by Deep Purple. It’s very simple with a central theme consisting in a four-note "blues scale" melody. If you don’t know this song, you can watch and listen here:




These are the notes for the central theme I learnt by ear. I know it’s not 100% accurate, but for the purpose of the exercise this should do just fine:

DFG DFAG DFG FD

You can play this riff with the virtual piano below, just click on the picture and click the notes above. It’s fun!


NMRing Smoke on the water

Before any further ado, these are the instructions to listen to the NMR version of Smoke on the Water:

  1. First you will need Mnova. If you don’t own a license, you can download a free, fully functional demo from our web site.
  2. Download this NMR document (SmokeOnTheWater.mnova)
  3. Download this script (playArray.qs)
  4. Finally, put on your earphones or amplify the volume on your computer speakers, open SmokeOnTheWater.mnova file with Mnova and run the script.
If everything works as I hope, you should be listening to 12 different pings pretending to be the melody of Smoke on the Water riff. I hope you won’t get too disappointed and more importantly, I hope that Deep Purple won’t detest me for ruining their song :)

Behind the scenes

As Dr. Walter Bauer explained on his web site, the most convenient way to create melodies with NMR is by means of pulse programmers to define the appropriate frequencies and delays. I’m taking a much simpler approach which, of course, cannot be used to create complex harmonies.

The basic idea is to synthesize NMR FIDs in such a way that each FID will represent a note of the song. For example, in this particular case I have simulated 4 FIDS with frequencies corresponding to A, D,F, and G. Next I stacked these 4 FIDs to create a new stacked item with 12 FIDs formed by the combination of the 4 main FIDs (main tones) and sorted to yield DFG DFAG DFG FD. The result is illustrated in the figure below.

As you can see, not all the FIDs have the same length. This is, of course, because some notes have to last longer than others. For example, we can consider the first FID (D) as the quarter note (crotchet), the third FID a half note (minim) and the sixth FID an eighth note (quaver). Next, I will comment on how the duration of the FIDs are controlled.

Some points of interest

At first sight, it may seem that in order to create the NMR version of Smoke on the Water one simply has to create the 4 FIDs with the ‘real’ frequencies corresponding to A, D, F and G notes and then organize them accordingly. For example, using the A220 pitch, the frequencies for the different notes should be:

A = 220 Hz
D = 146.83
F = 174.61
G = 196 Hz

However, if these notes are used in this way, the resulting song will be totally different than expected! The reason for that is very simple: NMR FIDs are expected to be in the so-called Quadrature Detection mode, that is, zero frequency in the center of the spectral width. Thus, it becomes necessary to translate the original frequencies into the NMR quadrature frame. The equation for such transformation is very simple:

NMR Note Frequency = Note Frequency + SpectralWidth/2

For example, in this case, as I have used a spectral width of 3000 Hz, the NMR frequencies for the 4 notes will be:


A = 220 + 3000/2 = 1720 Hz
D = 146.83
+ 3000/2 = 1646.83 Hz
F = 174.61
+ 3000/2 = 1674.61 Hz
G = 196
+ 3000/2 = 1696 Hz
Another point worth mentioning concerns the way to modulate the duration of the different FIDs to create crotchets, quavers, etc. In NMR terms, this is equivalent to the acquisition time (AT) which is defined as:

AT = N / SW

Where N is the number of points and SW is the spectral width. We can modify any of these two values to change the duration of the note (= FID acquisition time). In this example, I have kept the spectral width constant and modified the number of points.


Finally …

In this example I have created pure tone notes (e.g. one single frequency for each note) but it could also be possible to create some kind of guitar chords (e.g. power chords) to make the song more realistic by simply combining the root tone and a fifth. This can be easily done by using the Spin Simulation toolkit available in Mnova.