NoFold: RNA structure clustering without folding or alignment

(Downloading may take up to 30 seconds. If the slide opens in your browser, select File -> Save As to save it.)

Click on image to view larger version.

FIGURE 1.
FIGURE 1.

Normalization of the empirical feature space. Examples of CM score characteristics before (A,B) and after (C,D) normalization, for sequences and CMs of length ≤500 nt. (A) A representative example of the scores given to sequences of various lengths against a single CM, in this case tRNA. We consistently observe a relationship between sequence length and score that is most pronounced for sequences that are smaller than the size of the CM (73 nt in this case, indicated by the dashed line). Gray lines show separate linear regression fits to the scores of sequences shorter or longer than 73 nt, with slopes (m) indicated. (B) We additionally observed a relationship between the length of a CM and the average score that it produces. Average score was calculated based only on sequences with a length longer than the CM. (C) The length- and CM-specific procedure to calculate Z-scores greatly reduced the relationship between sequence length and score on an independent data set. Linear regression fit lines and slopes are indicated as in A. (D) Using Z-scores greatly reduced the relationship between CM length and the average score produced by the CM, and the average score for all CMs was close to zero.

This Article

  1. RNA 20: 1671-1683