Thursday, 11 September 2014

IntPath: comprehensive annotation database and novel algorithms to better extract biology from large gene lists

Background

Pathway data are important for understanding the relationship between genes, proteins and many other molecules in living organisms. Pathway gene relationships are crucial information for guidance, prediction, reference and assessment in biochemistry, computational biology, and medicine. Many well-established databases--e.g., KEGG, WikiPathways, and BioCyc--are dedicated to collecting pathway data for public access. However, the effectiveness of these databases is hindered by issues such as incompatible data formats, inconsistent molecular representations, inconsistent molecular relationship representations, inconsistent referrals to pathway names, and incomprehensive data from different databases.

Results

In this paper, we overcome these issues through extraction, normalization and integration of pathway data from several major public databases (KEGG, WikiPathways, BioCyc, etc). We build a database that not only hosts our integrated pathway gene relationship data for public access but also maintains the necessary updates in the long run. This public repository is named IntPath (Integrated Pathway gene relationship database for model organisms and important pathogens). Four organisms--S. cerevisiae, M. tuberculosis H37Rv, H. Sapiens and M. musculus--are included in this version (V2.0) of IntPath. IntPath uses the "full unification" approach to ensure no deletion and no introduced noise in this process. Therefore, IntPath contains much richer pathway-gene and pathway-gene pair relationships and much larger number of non-redundant genes and gene pairs than any of the single-source databases. The gene relationships of each gene (measured by average node degree) per pathway are significantly richer. The gene relationships in each pathway (measured by average number of gene pairs per pathway) are also considerably richer in the integrated pathways. Moderate manual curation are involved to get rid of errors and noises from source data (e.g., the gene ID errors in WikiPathways and relationship errors in KEGG). We turn complicated and incompatible xml data formats and inconsistent gene and gene relationship representations from different source databases into normalized and unified pathway-gene and pathway-gene pair relationships neatly recorded in simple tab-delimited text format and MySQL tables, which facilitates convenient automatic computation and large-scale referencing in many related studies. IntPath data can be downloaded in text format or MySQL dump. IntPath data can also be retrieved and analyzed conveniently through web service by local programs or through web interface by mouse clicks. Several useful analysis tools are also provided in IntPath.

Conclusions

We have overcome in IntPath the issues of compatibility, consistency, and comprehensiveness that often hamper effective use of pathway databases. We have included four organisms in the current release of IntPath. Our methodology and programs described in this work can be easily applied to other organisms; and we will include more model organisms and important pathogens in future releases of IntPath. IntPath maintains regular updates and is freely available at http://compbio.ddns.comp.nus.edu.sg:8080/IntPath

Saturday, 6 September 2014

电脑开机报警

开机报警下面是报警的声音全部解释
一、1短声 —— 系统正常启动。
二、2短声 —— 常规错误。应进入CMOS重设错误选项。
三、1长1短声 —— RAM或主板出错。
四、1长2短声 —— 显示器或显卡错误。
五、1长3短声 —— 键盘控制器出错。
六、1长9短声 —— 主板FlashRAM或EPROM错误。
七、不间断长声 —— 内存未插好或有芯片损坏。
八、不停响声 —— 显示器未与显卡连接好。
九、重复短声 —— 电源故障。
十、无声且无显示 —— 无电源,或CMOS电池耗尽,或CPU出错。
1、如果声音长声不断响,有可能是内存条未插紧
2、如果声音快速短声响两次(例如滴、滴):有可能是CMOS设置错误,需要重设,这种情况出现最多,屏幕画面的底部(一般在左下)会出现英文提示“CMOS CHECKSUM FAILURE ”,然后要求按F?(F?可能是F1-F12,视乎主板)载入默认设置或者进入。另外,刷新/更新BIOS后也有可能出现这个问题,载入默认设置就好了。
3、如果声音是一长声 一短声的,有可能是内存或主板错误。
4、如果声音是一长声 两短声的,有可能是显示器或显卡错误。
5、如果声音是一长声 三短声,有可能是键盘控制器错误。
6、如果声音是一长声 九短声,有可能是主板BIOS的FLASH RAM 或EPROM错误。
1、如果声音只是一短声响,有可能是内存刷新故障。
2、如果声音是快速短声响两次,有可能是内存ECC校验错误。
3、如果声音是快速短声响两次,有可能是系统基本内存检查失败。
4、如果声音是快速短声响四次,有可能是系统时钟出错。
5、如果声音是快速短声响五次,有可能是CPU出现错误。
6、如果声音是快速短声响六次,有可能是键盘控制器错误。
7、如果声音是快速短声响七次,有可能是系统实模式错误。
8、如果声音是快速短声响八次,有可能是显示内存错误。
9、如果声音是快速短声响九次,有可能是BIOS芯片检验错误。
10、如果声音是一长声 三短声,有可能是内存错误。祝你好运!

Wednesday, 30 July 2014

How to parse MiRProf Results.

 MiRProf Results contains may unused information which might affect the effective analysis.

Here shows how to parse the MiRProf exported .csv files.

First get all the names of micro RNA from the file:

grep - All-Expression-hg19-MirBase.csv > mir.txt

 The get the corresponding Total number of reads:

grep Total All-Expression-hg19-MirBase.csv > Allnumber.txt

Then combine the two files on excel, then you are done.

Convert all fastq files in folder to fasta

for fq in *; do fn=$(echo $fq|sed s'/.fq//'); echo $fn;awk 'NR % 4 == 1 {print ">" $0 } NR % 4 == 2 {print $0}' $fq > ../MergedFasta/$fn.fa; done;

Merge two Replicates into one

for fq in 890/*; do fn=$( echo $fq | sed 's/_R1_001.fastq//'| sed 's/_R2_001.fastq//'); echo $fn; cat $fq >>$fn.fq; done;

Thursday, 24 July 2014

How to release the /boot space.

If you've a lot unused kernels. Remove all but the last kernels with:
sudo apt-get purge linux-image-{3.0.0-12,2.6.3{1-21,2-25,8-{1[012],8}}}
 
This is shorthand for:
sudo apt-get purge linux-image-3.0.0-12 
linux-image-2.6.31-21 linux-image-2.6.32-25 
linux-image-2.6.38-10 linux-image-2.6.38-11 
linux-image-2.6.38-12 linux-image-2.6.38-8