Skip to main content

dataset - ID Swapping: Efficient use of a reference table to convert ID values



The rise of various bio-companies for high-throughput sequencing technology and open-source bioinformatic software suites for analysis thereof, has resulted in a equivalent rise of reference conventions (often species specific) to referring to genes, exons, transcripts, etc. Most useful of all of these conventions - to the biologist - is the common gene name (which itself likely has several aliases... an issue for another day).



For example, consider the gene KRAS: enter image description here


Thus it proves useful to convert genes from their ID (e.g. in this case P01116, or ENSG00000133703) to its common name (KRAS). To further drive this point home, consider IDs just within the same system - ENSEMBL. The human ID for KRAS is ENSG00000133703, while for mice the KRAS ID is ENSMUSG00000030265. Note that even stripping away the species specific element, yields different numerals. This type of situation might arise when one tries to compare two gene lists. My previous question, related to merging sets via ID, aims at finding efficient ways to combine different gene lists, once both have the same ID convention.





The goal of this question is to find the most efficient way of scanning a reference file for a gene's give ID, acquiring its associated gene name, and then replacing the original ID entry with the associated gene name. If this is unclear, an example comes shortly below.





For this question we will be using an ENSEMBL reference file. They are not large (ranging from ~30MB for yeast to ~200MB for humans). You can download such a file here:ENSEMBL BioMart.


To get the file do the following:




  1. On the drop down that states "- CHOOSE DATABASE -", select "Ensembl Genes 86"

  2. On the next drop down that states "- CHOOSE DATASET -" select whatever species you want. You could try "Mus musculus genes (GRCm38.p4)"

  3. On the left hand column, click "Attributes"

  4. On the new screen, click the box next to "GENE:"

  5. Add at least "Associated Gene Name" (you can also add other useful information such as "GENCODE basic annotation").

  6. At the top of the screen click results

  7. I prefer .CSV files, so you can change the file extension if you wish.

  8. Click Go to download the file. This link may or may not maintain the information (maybe good link)


Naturally if one works with gene lists often, it may be practical to get several different species.






Using SemanticImport we can see our reference file: enter image description here


For example purposes one could use RandomChoice to get a few IDs. I am just pasting the output here:


demoIDs={
"ENSMUSG00000028334", "ENSMUSG00000054079", "ENSMUSG00000027575",
"ENSMUSG00000041528", "ENSMUSG00000030231", "ENSMUSG00000078606",
"ENSMUSG00000028559", "ENSMUSG00000028431", "ENSMUSG00000020253",
"ENSMUSG00000019818"}


The corresponding gene names are:


{"Nans", "Utp18", "Arfgap1", "Rnf123", "Plekha5", 
"Gm4070", "Osbpl9", "Ikbkap", "Ppm1m", "Cd164"}

Naturally the actual list may be quite large.


Thus if this is our starting file: enter image description here


We want to end up with: enter image description here


Note, these were printed using TableForm, but they are actually Datasets.






I am providing my own answer, as I know of a way to do it. However it is surely not the most efficient or elegant way to do so. Thus I am looking for other answers that given the following:



  • the file name

  • the starting ID column header (Ensembl ID)

  • the final ID column header (gene)

  • a gene list data set

  • and the column header of that set contain the column of IDs (see below)


enter image description here


replaces the IDs in the gene list with the converted gene name (or N/A if not found).





Comments

Popular posts from this blog

plotting - How to draw lines between specified dots on ListPlot?

I would like to create a plot where I have unconnected dots and some connected. So far, I have figured out how to draw the dots. My code is the following: ListPlot[{{1, 1}, {2, 2}, {3, 3}, {4, 4}, {1, 4}, {2, 5}, {3, 6}, {4, 7}, {1, 7}, {2, 8}, {3, 9}, {4, 10}, {1, 10}, {2, 11}, {3, 12}, {4,13}, {2.5, 7}}, Ticks -> {{1, 2, 3, 4}, None}, AxesStyle -> Thin, TicksStyle -> Directive[Black, Bold, 12], Mesh -> Full] I have thought using ListLinePlot command, but I don't know how to specify to the command to draw only selected lines between the dots. Do have any suggestions/hints on how to do that? Thank you. Answer One possibility would be to use Epilog with Line : ListPlot[ {{1, 1}, {2, 2}, {3, 3}, {4, 4}, {1, 4}, {2, 5}, {3, 6}, {4, 7}, {1, 7}, {2, 8}, {3, 9}, {4, 10}, {1, 10}, {2, 11}, {3, 12}, {4, 13}, {2.5, 7}}, Ticks -> {{1, 2, 3, 4}, None}, AxesStyle -> Thin, TicksStyle -> Directive[Black, Bold, 12], Mesh -> Full, Epilog -> { Line[ ...

dynamic - How can I make a clickable ArrayPlot that returns input?

I would like to create a dynamic ArrayPlot so that the rectangles, when clicked, provide the input. Can I use ArrayPlot for this? Or is there something else I should have to use? Answer ArrayPlot is much more than just a simple array like Grid : it represents a ranged 2D dataset, and its visualization can be finetuned by options like DataReversed and DataRange . These features make it quite complicated to reproduce the same layout and order with Grid . Here I offer AnnotatedArrayPlot which comes in handy when your dataset is more than just a flat 2D array. The dynamic interface allows highlighting individual cells and possibly interacting with them. AnnotatedArrayPlot works the same way as ArrayPlot and accepts the same options plus Enabled , HighlightCoordinates , HighlightStyle and HighlightElementFunction . data = {{Missing["HasSomeMoreData"], GrayLevel[ 1], {RGBColor[0, 1, 1], RGBColor[0, 0, 1], GrayLevel[1]}, RGBColor[0, 1, 0]}, {GrayLevel[0], GrayLevel...

Is there a way to do conditional matrix loop using 'continue'

I have the following: n = 3; m = 5; ww = RandomReal[{0, 0.1}, {n, n}]; uu = RandomReal[{0, 1}, {m, n}]; pp = RandomReal[{0, 1}, {n, n}]; ss = RandomInteger[{0, 5}, {m, n}]; Grid[{{"ww", "uu", "pp", "ss"}, {ww // TableForm, uu // TableForm, pp // TableForm, ss // TableForm}}, Spacings -> {5, 2}, Dividers -> All] where I would like to look at every element of matrix ss and produce a matrix tt , with zeroes at the locations in ss which have zeroes, and in all other positions do the following: tt = (-1/Subscript[ww, m]) Log[(1 - uu)/(Subscript[pp, m - 1])], where Subscript[ww, m] is the value at index of ww matrix and where Subscript[pp, m - 1] is the value at index-1 of pp matrix. So for example if the first value ever read from matrix ss happens to be 2, then value taken from matrix ww would be from the row 2, but from pp would be from row 1. Also how to tell difference between a 0 as a valid value from within the matrix elemen...