faculty

Publications

Sequence Databases

Groups and Associations Vivek Kumar Chaturvedi, Divya Mishra, Aprajita Tiwari, V. P. Snijesh, Noor Ahmad Shaik, and M. P. Singh
Essentials of Bioinformatics, Volume I: Understanding Bioinformatics: Genes to Proteins 2018

Bioinformatics involves the use of information technology to collect, store, retrieve, and analyze the enormous amount of biological data that are available in the form of sequences and structures of proteins, nucleic acids, and other biomolecules (Toomula et al. 2011). Biological databases are mainly classified into sequence and structure databases. The first database was created soon after the sequencing of insulin protein in 1956. Insulin was the first protein to be sequenced; it contains 51 amino acid residues (Altschul et al. 1990). Over the past few decades, there has been a high demand for powerful computational methods which can improve the analysis of exponentially increasing biological information, finally giving rise to a new era of “bioinformatics.” Development in the field of molecular biology and high-throughput sequencing approaches has resulted in the dramatic increase in genomic and proteomic data such as sequences and their corresponding molecular structures. Submission of such facts into public information has led to the development of several biological databases which can be accessed for querying and retrieving of stored information through the research community. During the mid-1960s, the first nucleic acid sequence of yeast tRNA was found out. Around this period 3D structure of the protein was explored, and the well-known Protein Data Bank (PDB) was developed as the primary protein structure database with approximately ten initial entries (Ragunath et al. 2009). The ultimate goal of designing a database is to collect the data in the suitable form which may be easily accessed through researches (Toomula et al. 2011). In this chapter, we are cataloguing the various biological databases and also provide a short review of the classification of databases according to their data types.