Saturday, 16 August 2008

SPECTRa

During the flight to Stockholm (where I’m having a great time btw), I read with great interest this article:

SPECTRa: The Deposition and Validation of Primary Chemistry Research Data in Digital Repositories
DOI: 10.1021/ci7004737

In my opinion, the project presented in this article is a great initiative which I can only applaud and support. There is, however, a point which I would like to comment on because it’s not fully clear to me which regards to the way in which NMR spectra are stored in the repository.

The authors have decided to use JCAMP as the format for file input to their repositories. They do not specify which actual data is being saved in these JCAMP files, the processed spectrum or the FID (or both). I hope that they are saving the original FID and not only the processed spectrum, otherwise data preservation will be broken. I think this is a very important point which deserves some further clarification. This is how I see it:

The most important piece of information in an NMR experiment is, for sure, the acquired FID, not the processed spectrum. A chemist could have processed an FID to produce a spectrum in such a way that some spectral features are lost. For example, he/she could have applied a very large line broadening function which will make the analysis of the finer structure of some multiplets impossible. If the original FID has not been kept, the option to re-process the spectrum to calculate those lost couplings will not be viable (in some cases, Inverse Fourier Transform and/or some resolution enhancements procedures could help, but only in a very limited extent). In fact, there are many processing operations which could alter, irreversibly, either the qualitative or quantitative information present in the NMR experiment.

In short, I’m strongly convinced that any system aimed at preserving all information contained in NMR experiments should keep the original acquired data points, the FID. This is something I have learnt during the many years of development of both MestReC and Mnova: Mnova keeps, in addition to the processed spectrum, all the original files as they were acquired in the spectrometer. And I know that iNMR does the same thing, though in a different way (Mnova packs all the files within a single binary file, whereas iNMR keeps all the files separately and a processing log file so that the processed spectrum does not need to be saved. From my point of view, both approaches are equivalent and perfectly valid).

Tuesday, 22 July 2008

Bayesian DOSY: a New Approach to Diffusion Data Processing

Yesterday I blogged about basic concepts on DOSY NMR. From an experimental stand point, the pulse sequences are not very complex, being the Pulse Field Gradient Stimulated Echo (PFGSE) experiment proposed by Tanner one of the basic pulse sequences used to determine diffusion by NMR. This pulse sequence can be considered as a building block for a number of extended pulse sequences designed to minimize some sources of artefacts caused by thermal convection currents in the sample, background gradients, radiation damping, zero order coherences in strongly coupled spin systems, etc. Other effects such as J-modulation and Cross-Relaxation have also been considered.

The tricky part comes in the mathematical DOSY transformation. As I mentioned earlier, there exists different approaches, each of them with their own strengths and weaknesses. Today I would like to introduce a new method for DOSY transformation based on Bayesian Theory which has been implemented in our software Mnova as a result of our collaboration with Stan Sykora. We call this new algorithm BDT (Bayesian DOSY Transform)

A formal description of the Bayesian theory and its particular implementation for DOSY transformation is beyond this blog entry, but we are currently writing an article which will give all the details of the method as implemented in Mnova.
In short, this Bayesian approach assigns an a-priori probability (this is the key word in Bayesian context) to the elements in a space defined by the entities to be estimated. In this context, an entity is the pair (f,d) where f is frequency and d is a mono-diffusion coefficient. Actually, it should be (f,d,w) where w is a weight (intensity), but this weight is factored out by finding its optimal value (since the dependence is linear, this can be done explicitly). So we are actually finding the a-posteriori probability for the sentence there is a component - no matter how intense - at f that has such and such d. Next, a final normalization process modifies the probability still further (composite probability) by looking at the intensity in the spectrum at location f (this would be real probability if the spectrum were normalized to 1).

Stan Sykora will be presenting the mathematical background and physical insights necessary to understand this new method as well as some real life examples processed with Mnova in the GIDRM conference to be held in Bressanone (Brixen) next September.

So how does the algorithm perform? We first tested the algorithm by simulating with Matlab a diffusion experiment with 2 synthetic peaks with frequencies of 100 and 200 Hz and diffusion coefficients of 1 and 0.01 (dimensionless).


Application of BDT yields the following DOSY spectrum:



The synthetic DOSY spectrum was changed in order to include artificial noise and different degrees of peak overlap and diffusion distances. We’ve found the performance of the algorithm to be excellent in all the cases analyzed.

Next we implemented the algorithm in Mnova and we applied it with a sample of an aqueous solution of potassium N-methyl-N-oleoltaurate (a surfactant) with TSP at 23 C (The original Varian FID file has been obtained from the VARIAN NMR USER GROUP LIBRARY which was submitted by Brian Antalek as a sample for this DECRA algorithm). This is the original raw data after automatic Fourier Transform, phase and baseline correction in Mnova:


After applying BDT, we obtain a DOSY spectrum in which the three components are clearly well resolved in the diffusion dimension and the absolute values are in accordance with the expected values


Below you can see the DOSY spectrum of a mixture of Caffeine + 2-EthoxyEthanol + Water acquired in a Bruker instrument by my friend Andy Soper.

Main properties of the algorithm

Compared to other approaches, we believe that this Bayesian method implemented in Mnova appears extremely promising. It automatically avoids having exact, unnatural zeros anywhere in the resulting DOSY map since every point of the 2D map has a well defined value of statistical congruence with the data. Moreover, the BDT maps show ‘normal’ line widths in the f-direction, correctly positioned and resolved peaks in the d-direction and quantitatively correct horizontal and vertical projections – a combination difficult to achieve by any other means.
The approach can be easily extended to non-exponential cases arising from overlapping lines and, like all Bayesian methods, incorporate additional information available from other sources (a-priori knowledge). Likewise, it is possible to place a statistical premium on alignment of spectral peaks along horizontal lines in the [f,d] plot.

Do you want to try it out yourself?

The algorithm is readily available in the alpha stage in Mnova. From this post, I would like to offer this special version to anyone interested in trying the BDT algorithm with his own data sets. I will very much appreciate any feedback from you.
Just write to me at the email address below and I will give you the instructions on how to get the software.



Sunday, 20 July 2008

DOSY NMR

NMR diffusion experiments provide a way to separate the different compounds in a mixture based on the differing translation diffusion coefficients (and therefore differences in the size and shape of the molecule, as well as physical properties of the surrounding environment such as viscosity, temperature, etc) of each chemical species in solution. In a certain way, it can be regarded as a special chromatographic method for physical component separation, but unlike those techniques, it does not require any particular sample preparation or chromatographic method optimization and maintains the innate chemical environment of the sample during analysis.
The measurement of diffusion is carried out by observing the attenuation of the NMR signals during a pulsed field gradient experiment. The degree of attenuation is a function of the magnetic gradient pulse amplitude (G) and occurs at a rate proportional to the diffusion coefficient (D) of the molecule. Assuming that a line at a given (fixed) chemical shift f belongs to a single sample component A with a diffusion constant DA, we have

(1) S(f,z) = SA(f) exp(-DAZ)


where SA(f) is the spectral intensity of component A in zero gradient (‘normal’ spectrum of A), DA is its diffusion coefficient and Z encodes de different gradient amplitudes used in the experiment.


Depending on the type of experiment, there are various formulae for Z in terms of the amplitude G of the applied gradient and one or more timing parameters such as
Δ (time between two pulse gradients, related to echo time) and δ (gradient pulse width). In the original Tanner-Stejskal method using two rectangular gradient pulses, for example,

(2) Z =γ2G2δ2(Δ-δ/3)


Eq. (2) holds strictly for simple PFG-NMR experiments, and is modified slightly to accommodate more complicated pulse sequences.

In practice, a series of NMR diffusion spectra are acquired as a function of the gradient strength. For example, the figure below shows the results of a series of 1H-NMR diffusion experiments for a mixture containing caffeine, 2-Ethoxyethanol and water.



It can be observed that the intensities of the resonances follow an exponential decay. The slope of this decay is proportional to the diffusion coefficient according to equation (1). All signals corresponding to the same molecular species will decay at the same rate. For example, peaks corresponding to water decay faster than the peaks of caffeine and 2-Ethoxyethanol

The DOSY transformation


As far as data processing of raw PFG-NMR spectra is concerned, the goal is to transform the NxM data matrix S into an NxR matrix (2D DOSY spectrum) as follows:



The horizontal axis of the DOSY map D is identical to that of S and encodes the chemical shift of the nucleus observed (general 1H). The vertical dimension, however, encodes the diffusion constant D. This is termed Diffusion Ordered Spectroscopy (DOSY) NMR.
In the ideal case of non-overlapping component lines and no chemical exchange, the 2D peaks align themselves along horizontal lines, each corresponding to one sample component (molecule).

The horizontal cut along such a line should show that component’s ‘normal’ spectrum. Vertical cuts show the diffusion peaks at positions defining the corresponding diffusion constants.


The mapping S=>D will be henceforth called the DOSY transformation. This transformation is, unfortunately, far from straightforward. Practical implementations include mono and biexponential fitting, Maximum Entropy, and multivariate methods such as DECRA.

Recently, Mathias Nilsson and Gareth A. Morris have proposed the so-called ‘Speedy Component Resolution’ (Anal. Chem., 2008, 80, 3777–3782) as an improved variation of the Component Resolved (CORE) method (J. Phys. Chem, 1996, 100, 8180). This is a multivariate-based method and the examples used in the article show an excellent performance of the algorithm.

Following a different approach, we have recently developed a brand new method for DOSY processing which has been included in Mnova. I will leave the details (and how to get the program to try it out) for the next post. In the meantime, should you be interested, just drop me a line.

Tuesday, 1 July 2008

Sensitivity Enhancement for free?

In my previous post I discussed about NMR FID/spectra having two channels, the Real and Imaginary parts and that in general, only the Real part is displayed. In fact, we could use the term Stereo-FID in the same fashion as we use Stereo-Sound (remember that NMR spectra span audio frequencies).
If we think about these stereo signals as coming from two receiver coils in quadrature and assuming simultaneous detection, we could formulate the following scenario:

It’s a common practise to acquire several (hundred or thousand) FIDs which are then added together (signal averaging). Because NMR responses build up in proportion to the number of signals recorded, N, whilst the noise varies randomly from one measurement to the next and thus, adds up more slowly, as sqrt(N), there is an overall improvement in sensitivity of sqrt(N).
If the noise in the Real and Imaginary channels were totally independent, it should be possible to add to the Real channel the Imaginary channel (after a 90º phase shift to make it phase coherent with regard to the Real part) so that we could achieve a further improvement of sensitive of sqrt(2)!. Would this be possible? This will be the subject of this post.

Of course, this ‘trick’ would only work if the noise in the two channels were totally independent. In the NMR field it is assumed that the experimental noise is stationary and white (i.e. the correlation between two consecutive points is zero). But what about the correlation between the noise in the two different channels? Are they correlated? This is very easy to analyze experimentally with Mnova by writing a very simple script which calculates the Correlation Coefficient (Pearson). For example, consider the spectrum below. We could use this script to calculate the Pearson Correlation Coefficient using the points between 5000 & 10000 which is a signal free region


As shown in the figure below, the correlation coefficient between the noise in both channels is almost zero. You can repeat this operation with different spectra and you will arrive to similar results

So it appears as if the noise in both channels is statistically uncorrelated, something which should be intuitively expected as the 2 acquisition channels are orthogonal. Does this mean that the noise in both channels is independent?
We can make a very simple experiment: We could phase a spectrum in order to make the real part perfectly in phase and then change the phase of the imaginary part in such a way that the peaks in both channels become perfectly phase-coherent. Next, both channels can be added so that the signals will increase 2 fold. On the other hand, if the noise in both channels were phase-incoherent, it will increase more slowly and therefore the overall S/N should increase by sqrt(2).

We can carry out this experiment very easily by applying a 90º phase shift into the imaginary channel and then summing up both channels. This is done with the following script:


Before applying this script, we need to calculate the SNR of the original spectrum. We can use the central peak of the Chloroform multiplet as a reference peak to estimate the SNR as depicted in the figure below:

Next, we can apply the script above (sumReImPhased) to apply a 90º phase shift to the imaginary part followed by the addition of the imaginary channel to the real one. If the SNR of this new spectrum is calculated we get the following result:

We can appreciate that the Chloroform peak is now twice the height, but the standard deviation has also increased two-fold, so the SNR remains constant. What a disappointment!

Conclusions:

We have first found out that the noise in the Real and Imaginary channels is statistically uncorrelated. However, this does not mean that they are independent. Just as sin(x) and cos(x) functions are orthogonal "uncorrelated", but not independent (sin^2(x) = 1 – cos^2(x)).
And we have shown that noise in the Re and Im channels is certainly not independent: the SNR does not improve at all.
The essence is that if the spectrometer has just one coil (which is, to the best of my knowledge, always the case in spectroscopy), then the noise in the two orthogonal channels would not be independent and thus the sqrt(2) sensitivity enhancement is not possible. It will be necessary to have two separate coils (and two receivers) in quadrature to have uncorrelated real and imaginary noise. Why is this not possible? I don’t really know, most likely because of lack of space …

References:

Experimental Noise in Data Acquisition and Evaluation III. Exponential Multiplication, Discrete Sampling, and Truncation Effects in FT Spectroscopy. Dr. S. Sýkora.



Friday, 20 June 2008

Real NMR

We often forget that NMR spectra are complex entities. By complex I don’t mean complicated or difficult to understand (which is, unfortunately, a very common thing too), but by the numbers in the complex space, that is, formed by real and imaginary numbers. One reason that can justify why NMR spectra are barely considered as formed by complex numbers is that NMR software, in general, only displays the real part. However, the imaginary part is present in the background (unless discarded e.g. to save memory) and it’s actually very important. The imaginary part makes it possible to correct the phase of the spectra in such a way that we can get nice absorptive lines (e.g. Lorentzian / Gaussian lines) which provide higher resolution than out-of-phase spectra (e.g. combination of absorptive and dispersive lines. See figure below).




NMR spectra, being of complex nature, can also be displayed in magnitude mode which makes the spectra phase insensitive. This can be advantageous in those cases in which phase correction is difficult, but at the expense of poorer resolution (though in the case of 13C NMR, magnitude mode can facilitate automatic processing, as I have blogged before). In addition, phase information can be very useful in some cases (e.g. DEPT, NOESY, edited NMR experiments, etc).


Acquisition of NMR spectra as complex numbers is also important (in addition to making phase correction possible) to distinguish positive from negative frequencies (quadrature detection. See this link for more information).

Mnova does not have a built-in function to show the imaginary part of the spectrum, but in the case you really need to see it, I will offer here two solutions:

1. One obvious way is by realizing that that the Re and Im channels differ by a phase shift of 90º. Thus, if we apply a (zero order) phase correction of +/-90º, the real part displayed in the software becomes analogous or equivalent to the imaginary counterpart.

2. There is a second alternative: it’s possible to write a very simple script which swaps the real with the imaginary part of the spectrum. This is the script:


  1. function swapRI()

  2. {

  3. var w = new DocumentWindow(mainWindow.activeWindow());

  4. //get active spectrum

  5. var spectrum = nmr.activeSpectrum();

  6. if (!spectrum.isValid() || spectrum.isReal)

  7. return;

  8. if (spectrum.dimCount > 1)

  9. return; //1D only

  10. var dCount = spectrum.dimCount;

  11. var npts = spectrum.count(1);

  12. //To print the value of one of the spectral points

  13. for (var i=0; i<npts; i++)

  14. {

  15. var tmp = spectrum.real(i);

  16. spectrum.setReal(i, spectrum.imag(i));

  17. spectrum.setImag(i, tmp);

  18. }

  19. spectrum.update();

  20. w.update();

  21. }

  22. }

Having said that NMR spectra are of complex nature and why this is important, then, how is it possible, from a hardware point of view, to get real and imaginary numbers? After all, NMR detectors are just measuring ‘real’ values (i.e. induced oscillating voltage in a resonant RF coil). When we say Re and Im numbers in this context, we could have said X and Y values, or u and v values, etc. What matters here is that we are using 2 (orthogonal) detectors to measure (generally simultaneously) the magnetization along X and Y axis. We know that the Re and Im components in the complex space are related by a phase angle of 90º, analogous to the X and Y axis. Thus, the values recorded by the detector placed along the X axis can be considered as the real part of the spectrum (or the cosine component), whereas the values measured in the detector placed in the Y axis will correspond to the imaginary part (or sine component) of the spectrum.

What I would like to highlight here is that the Re and Im, in principle, could appear as 2 INDEPENDENT physical measures, even though they are (generally) recorded simultaneously. In other words, the Re and Im channels are UNCORRELATED. However, this is not completely true. I have already mentioned that the dispersive and absorptive lines (corresponding to Real and Im components of a phased spectrum) are related by a phase shift of 90º. More strictly speaking, the Re and Im components are connected by the well known Kramers–Kronig relation being the Hilbert Transform the mathematical bridge connecting both components. This transform makes it possible to obtain the imaginary part from the real one and the other way round. How about the noise? Is the noise in the imaginary and real parts uncorrelated? In that case, theoretically, it should be possible to achieve a further sqrt(2) gain in sensitivity when quadrature detection is used. This issue will be the subject of my next post.

References
Quadrature versus linear detection, and where do complex MR data come from
Nothing is better

Friday, 16 May 2008

Automatic chemical shift calibration of 2D NMR spectra

Accurate calibration of the chemical shift scales of NMR spectra is very important for both reproducibility and the correlation of the chemical shift with structural properties. Usually, this task is very easy in 1D NMR and in general, it is carried out graphically by simply selecting with the mouse cursor a reference signal (e.g. TMS or a residual solvent peak).

In the case of 2D NMR, this operation is more tedious as it has to be applied to both dimensions. In addition, the low digital resolution typically obtained in 2D NMR (nD in general) spectra makes calibration more sensitive to the actual position of the reference peak selected for referencing.

Wouldn’t it be possible to use the calibration information from the 1D spectra (where the digital resolution is usually >20 times higher than in a 2D spectrum) to automatically reference 2D spectra? In theory, this should be possible (at least with several 2D experiments), but I was surprised when I found that none of the NMR software applications I’m aware of incorporates such capability (if anyone knows a software with this feature, please let me know).

Most of the NMR software packages I know include facilities to manually align the 2D and 1D spectra, that is, to correlate signals of 2D and 1D spectra. However, I was interested in a fully automatic procedure to take advantage of the higher digital resolution of the 1D spectrum to calibrate the 2D spectrum.

After studying this issue in depth, we came out with a new algorithm which works exactly as I wanted: Once the 1D spectrum is properly referenced, the 2D spectrum is automatically calibrated with the information contained in the 1D spectrum.

As always, a picture is worth a thousand words. Consider the following 2D COSY spectrum:

And compare it with the corresponding high resolution 1H-NMR counterpart:

It can be observed that the chemical shift scales of both spectra differ by about 1.5 ppm. If the 1D spectrum is attached to the 2D COSY as an external projection, the result would look something like this:


Application of the newly developed autoalignment algorithm will yield the following result (please note that there is not need to enter any user-defined parameter: it's fully automatic):


The scales of the 2D spectrum have been calibrated by using the information contained in the high resolution 1H spectrum.

It’s important to mention that at present, this algorithm does not work with all 2D spectra. Basically, it works fine in those spectra in which internal projections are compatible, in terms of number of resonances, with the external 1D traces. This would be the case of 2D homonuclear experiments and the proton dimension of HSQC and related spectra, but it will not work in the 13C dimension in these experiments.

This new algorithm has been incorporated into the latest version of Mnova (Mnova 5.2.2). Should you be interested in testing it, just download Mnova and play with it. Any comments, suggestions, bug reports, etc, will be very welcome.

Wednesday, 30 April 2008

Metabonomics databases

Having access to NMR databases of most common metabolites is very important if you’re interested or involved in NMR-based metabonomomics/metabolomics studies. Here are two good sites worth checking out:

For a more comprehensive review, see this article

Wednesday, 26 March 2008

13C Chemical Shift prediction: Typos and NMR misinterpretations

It is well known that 13C Chemical Shifts are essential for structure verification and elucidation of organic molecules by NMR. For example, in the field of structure verification, one can compare the observed 13C chemical shifts with the calculated (predicted) values in such a way that the structure hypotheses can be assessed by means on some numerical matching factor.
What is less evident is the great potential that 13C chemical shift prediction has in order to reveal typos and misinterpretations of NMR data that often (much more than expected, I’m afraid) appear in the experimental sections of scientific literature.

Wolfgan Robien has recently created a very interesting Web page in which he shows how his famous CSEARCH program for 13C –NMR prediction can be used for automatic data-checking. According to his personal opinion, there are at least 3 scenarios in which such checks should be done:

  • Daily routine during generation and interpretation of NMR-data
  • Check again during preparation of a manuscript
  • Check again during the peer-reviewing process ('robot-referee')

Mnova is honoured to include Wolfgang Robien’s CSEARCH algorithm within its NMRPredict Desktop plugin and includes simple yet very useful verification tools which can be used to easily identify errors in 13C-NMR data assignments.

Basic Misinterpretations, Typos and other Sad Events in NMR-Spectroscopy


Friday, 22 February 2008

Why aren’t Bruker FIDs time corrected?

Most of NMR has finally gone totally digital and is now routinely resorting to high-frequency ADC sampling at rates of several tens of MHz, with a tendency to ever higher sampling frequencies. Such a drastic oversampling has clear advantages in simplifying the front-end electronics. It has also many side benefits due to the fact that the subsequent down-conversion to low frequency range is done digitally and therefore theoretically artefact free. This includes perfect, calibration-free quadrature detection and drastic reduction of quantization noise. By the way, such techniques are routine in other areas of electronics (military, astronomy, audio, …); NMR is just a latecomer to this world.

High-frequency (HF) oversampling, however, has also some problems of its own. The digital decimation from the HF range to the audio range and the contextual digital filtering (a combination of CIC and FIR filters) need to be properly implemented in the hardware in order to be completely transparent to the User.

However, in the case of Bruker spectra a death time or group delay can be observed in the FID: it starts with very small values and then, after some points (usually between 60-80 points) the normal FID starts.

If a plain FT is applied to this FID, we will get a spectrum with a lot of wiggles in the baseline analogous to the convolution with a sinc function centred in the middle of the spectral window. This can be explained by recalling the time shift theorem of the Fourier Transform which says that if the time domain signal is shifted by n points, the frequency domain spectrum corresponds to the standard spectrum (when the FID has not been shifted) multiplied by exp(-i2*pi*w*n). In other words, we have introduced a very large first order phase correction in the spectrum. For example, if the FID is right shifted by 60 points (death time = 60 points), f-spectrum will exhibit a first order phase distortion of 60 * 360 = 21600 degrees.

In order to work around this problem, most NMR software packages read the decimation factor (and the DSP firmware version) from Bruker files and calculate the required phase correction. So far, so good.
However, the fact remains that Bruker FID’s are not time corrected. Evidently, Varian also uses oversampling and digital filtering and their FIDs are time corrected, that is, they start at time = 0. If the digital filter is known in advance, which is always the case, the group delay should be compensated in the spectrometer, therefore, in my opinion, this death time or group delay is a bug in the spectrometer. For this reason, and going back to the title of this article, any input given as to why Bruker FID’s are not time corrected, would be greatly appreciated.