;
WELCOME!!!

Monday, October 12, 2009

MICROARRAY TECHNOLOGY

Microarrays exploit the preferential binding of complementary single-stranded nucleic acid sequences. A microarray is typically a glass slide, on to which DNA molecules are attached at fixed locations (spots). There may be tens of thousands of spots on an array, each containing a huge number of identical DNA molecules (or fragments of identical molecules), of lengths from twenty to hundreds of nucleotides. (According to quick napkin calculations by Wilhelm Ansorge and John Quackenbush in Schnookeloch in Heidelberg on 4 October 2001, the number of DNA molecules in a microarray spot is 107-108). For gene expression studies, each of these molecules ideally should identify one gene or one exon in the genome, however, in practice this is not always so simple and may not even be generally possible due to families of similar genes in a genome.

Microarrays that contain all of the approximate 6000 genes of the yeast genome have been available since 1997. The spots are either printed on the microarrays by a robot, or synthesised by photolithography (similarly as in computer chip productions) or by ink-jet printing. The spot diameter is of the order of 0.1 mm, for some microarray types can be even smaller.

There are different ways how microarrays can be used to measure the gene expression levels. One of the most popular microarray applications allows the comparison of gene expression levels in two different samples, e.g., the same cell type in a healthy and diseased state. This is called cDNA microarray. An other array technique is oligo arrays.

  • cDNA Microarray

In the preparation of a cDNA microarray, the total mRNA from the cells in two different conditions is extracted and reverse transcription PCR (RT-PCR) is used to convert the RNA transcripts into cDNA. The cDNAs are usually composed of 500 -2000 basepairs long. The complete pool of cDNA is representative of transcriptional events in the tissue source of the RNA. The genes that were being actively transcribed in the sample will have mRNA copies that should have been first purified and then copied into cDNA during the RT-PCR step. The reverse transcription event for the control and experimental mRNA are identical in every step except one, and it is this step that enables differential gene expression to be determined. Nucleotides labelled with a green fluorescent dye Cy3 are incorporated into the control cDNA, while nucleotides labelled with a red fluorescent dye Cy5 are incorporated into the experimental DNA. After preparation, both probes are mixed and allowed to hybridise to the glass slide. Excess hybridisation buffer is washed off following an overnight incubation, and the slides are then ready to be scanned. Labelled gene products from the extracts hybridise to their complementary sequences in the spots due to the preferential binding - complementary single stranded nucleic acid sequences tend to attract to each other and the longer the complementary parts, the stronger the attraction.

  • Oligonucleotide Microarray

The physical chemistry of hybridisation is oligonucleotide microarrays is clearly different from that of cDNA microarrays. Oligonucleotides range in size from 10-25 bases. So, the DNA fragments in the spots are much smaller than cDNA fragments are. Oligonucleotide microarrays are used to detect point mutations (the missing, adding or changing of a single base) in a known DNA sequence. Single base mismatches do have much more influence on binding to an oligonucleotide sequence compared to cDNA. For example, a small genome can be synthesized on a chip as a set of thousands of 20 bp long fragments. When a single basepair match exists, the fluorescence intensity decreases significant. This technique gives possibilities to find most of the point mutations in a known DNA sequence.

Data quantification

The dyes enable the amount of sample bound to a spot to be measured by the level of fluorescence emitted when a laser excites it. If the RNA from the sample in condition 1 is in abundance, the spot will be green, if the RNA from the sample in condition 2 is in abundance, it will be red. If both are equal, the spot will be yellow, while if neither are present it will not fluoresce and appear black. Thus, from the fluorescence intensities and colours for each spot, the relative expression levels of the genes in both samples can be estimated.

The raw data that are produced from microarray experiments are the hybridised microarray images. To obtain information about gene expression levels, these images should be analysed, each spot on the array identified, its intensity measured and compared to the background. This is called image quantification and is done by image analysis software. To obtain the final gene expression matrix from spot quantification's, all the quantities related to some gene (either on the same array or on arrays measuring the same conditions in repeated experiments) have to be combined and the entire matrix has to be scaled to make different arrays comparable.

Microarrays are already producing massive amounts of data. These data, like genome sequence data, can help us to gain insights into underlying biological processes only if they are carefully recorded and stored in databases, where they can be queried, compared and analysed by different computer software programs. The EBI as well as the NCBI are establishing a public repository for microarray gene expression data analogous to banks for DNA sequence data.

Microarray is fundamentally a technique to identify complete gene expression profiles in selected tissues. Microarray experiments can give false positive and false negative results. Additional means of analysing gene expression (Northern blotting or RNAse protection assays) must be used to control microarray conclusion.


CLICK ON THE LINK BELOW FOR ANIMATED EXPLANATION:

http://www.youtube.com/watch?v=ePFE7yg7LvM&feature=related

GENE EXPRESSION ANALYSIS


Northern blotting

Northern blotting is a laboratorium technique to analyse RNA expression. The experiment takes several steps. First the total amount of RNA is isolated. Then it is separated on fragment length by gel electroforesis. Then it is transferred to nitrocelloluse or nylon filter paper. The filter can then used to search for a particular RNA by several probing techniques, for instance radioactive labeling of the probe. The probe should be complement to the RNA you are looking for. This simple procedure can indicate in which tissues or cell types a particular gene is expressed. In this way a Northern blot is often used for diagnostic purposes. It can also be used to confirm results from other experimental technques like microarray .
Micro Array Analysis: (see Pictorial representation):


ESTs,Physical maps,Cytogenic maps

ESTs : ESTs are small pieces of DNA sequence (usually 200 to 500 nucleotides long) that are generated by sequencing either one or both ends of an expressed gene. The 3' ESTs serve as a common source of STSs because of their likelihood of being unique to a particular species and provide the additional feature of pointing directly to an expressed gene.

ESTs as Gene Discovery Resource:As observed ESTs represent a copy of just the interesting part of a genome, that which is expressed, they have proven themselves again and again as powerful tools in the hunt for genes involved in hereditary diseases. ESTs also have a number of practical advantages in that their sequences can be generated rapidly and inexpensively, only one sequencing experiment is needed per each cDNA generated. ESTs are powerful tools in the hunt for known genes because they greatly reduce the time required to locate a gene. Using this method, scientists have already isolated genes involved in Alzheimer's disease, colon cancer, and many other diseases.

Cytogenetic Map:

A cytogenetic map is the visual appearance of a chromosome when stained and examined under a microscope. Particularly important are visually distinct regions, called light and dark bands, which give each of the chromosomes a unique appearance. This feature allows a person's chromosomes to be studied in a clinical test known as a karyotype, which allows scientists to look for chromosomal alterations.

Physical map:

A physical map is a collection of overlapping clones that have been arranged into a tiling path based on either fingerprinting (digestion of clones with restriction enzymes and comparison of the fragment sizes) or hybridisation.

The genetic markers can help to integrate these three maps mentioned above.

For Arabidopsis thaliana, TAIR's comprehensive MapViewer is an integrated graphic display of each Arabidopsis chromosome. TAIR is the internet site where all information and data about Arabidopsis is combined. MapViewer shows genetic, physical, and sequence maps in one site and allows users to search, browse, and align different maps in a region of interest. In the future, all the maps will be fully integrated into a genome map for the organism.

An Important concept of "GENE MAPPING"


Genetic map

Well lets use some imagination: Like interstate maps having cities and towns that serve as landmarks, the genetic maps have landmarks known as genetic markers, or "markers" for short. The term "marker" is used very broadly to describe any observable variation that results from an alteration, or mutation, at a single genetic locus. A marker may be used as one landmark in a map if, in most cases, that stretch of DNA is inherited from parent to child according to the standard rules of inheritance. Markers can be within genes that code for a noticeable physical characteristic such as leaf colour, or a not so noticeable trait such as a disease. The greater the distance between two linked genes, the greater the chance that two nonsister chromatids would cross over in the region between the genes and the greater the proportion of recombinants that would be produced. Thus, by determining the frequency of recombinants, we can obtain a measure of map distance between the genes. Today, several other genetic markers are used to detect linkage. There are several genetic markers:

  • RFLPs/ (Restriction Fragment Length Polymorphism's): They were among the first developed DNA markers. RFLPs are defined by the presence or absence of a specific site, called a restriction site, for a bacterial restriction enzyme. This enzyme breaks apart strands of DNA wherever they contain a certain nucleotide sequence...
  • VNTRs/ ( Variable Number of Tandem Repeat Polymorphisms): They mainly occur in non-coding regions of DNA. This type of marker is defined by the presence of a nucleotide sequence that is repeated several times. In each case, the number of times a sequence is repeated may vary..
  • Microsatellite polymorphism's: Defined by a variable number of repeats of a very small number of base pairs. Oftentimes, these repeats consist of the nucleotides, or bases, cytosine and adenosine. The number of repeats for a given microsatellite may differ between individuals, hence the term polymorphism--the existence of different forms within a population;
  • SNPs/ Single Nucleotide Polymorphism's: They are individual point mutations, or substitutions of a single nucleotide, that do not change the overall length of the DNA sequence in that region. SNPs occur throughout an individual's genome;
  • AFLP/ Amplified Fragment Length Polymorphism: They mainly involve a DNA fingerprinting technique which detects DNA restriction fragments by means of PCR amplification.
Currently, the most powerful mapping technique, and one that has been used to generate many genome maps, relies on Sequence Tagged Site (STS) mapping: A STS is a short DNA sequence that is easily recognisable and occurs only once in a genome (or chromosome).

What are Databases..??

Databases

At the beginning of the "genomic revolution," a bioinformatics concern was the creation and maintenance of a database to store biological information, such as nucleotide and amino acid sequences. Development of this type of database involved not only design issues, but also the development of complex interfaces whereby researchers could both access existing data as well as submit new or revised data.

Ultimately, however, all of this information must be combined to form a comprehensive picture of normal cellular activities so that researchers may study how these activities are altered in different disease states. Therefore, the field of bioinformatics has evolved such that the most pressing task now involves the analysis and interpretation of various types of data, including nucleotide and amino acid sequences, protein domains, and protein structures. The actual process of analysing and interpreting data is referred to as computational biology. Important sub-disciplines within bioinformatics and computational biology include:

  • The development and implementation of tools that enable efficient access to, and use and management of, various types of information;

  • The development of new algorithms (mathematical formulas) and statistics with which to assess relationships among members of large data sets, such as methods to locate a gene within a sequence, predict protein structure and/or function, and cluster protein sequences into families of related sequences.

Biological databases

A biological database is a large, organised body of persistent data, usually associated with computerised software designed to update, query, and retrieve components of the data stored within the system. A simple database might be a single file containing many records, each of which includes the same set of information. For example, a record associated with a nucleotide sequence database typically contains information such as contact name; the input sequence with a description of the type of molecule; the scientific name of the source organism from which it was isolated; and, often, literature citations associated with the sequence. For researchers to benefit from the data stored in a database, two additional requirements must be met:

  • Easy access to the information;
  • A method for extracting only that information needed to answer a specific biological question.

Entrez

At the site of the NCBI, many of the databases are linked through a unique search and retrieval system, called Entrez. Entrez allows a user to not only access and retrieve specific information from a single database, but to access integrated information from many NCBI databases. For example, the Entrez protein database is cross-linked to the Entrez taxonomy database. This allows a researcher to find taxonomic information of the protein of interest. An overview of the most important databases is given in the part Databases on this site.

Why Bioinformatics is important..??

The genome sequencing projects has produced large amounts of nucleotide and protein sequence data. Traditionally, molecular biology research was carried out entirely at the experimental laboratory bench but the huge increase in the scale of data being produced in this genomic era has seen a need to incorporate computers into this research process.


Sequence generation, and its subsequent storage, interpretation and analysis are entirely computer dependent tasks. However, the molecular biology of an organism is a very complex issue with research being carried out at different levels including the genome, proteome, transcriptome and metabalome levels. Following on from the explosion in volume of genomic data, similar increases in data have been observed in the fields of proteomics, transcriptomics and metabolomics.


The first challenge facing the bioinformatics community today is the intelligent and efficient storage of this mass of data. It is important to provide easy and reliable access to this data. The data itself is meaningless before analysis and the sheer volume present makes it impossible for even a trained biologist to begin to interpret it manually. Therefore, incisive computer tools must be developed to allow the extraction of meaningful biological information.


There are three central biological processes around which bioinformatics tools must be developed:

DNA sequence determines protein sequence
Protein sequence determines protein structure
Protein structure determines protein function

what is GENOMICS..??

Must tell you most of the topics in Bioinformatics that u will be coming across are simple in meaning...i mean GENOMICS u can again use ur grey cells passively and without scratching ur head can tell "its basically study of particular organism's genome" off course ur right...!!!
But again my blog would be a failure if I cant get u into understanding what exactly is going on...RIGHT!!!
so then Wat exactly is constituting our/any organism's genome..??
the Chromosomes.DNA sequences with genes embedded onto the chromosomes/DNA sequence(both are possible) Ref-GENE 8...
Primarily a genome study (GEnomiCS) involves the study of these sequences...
Though Higher Genomics study involves sequencing, comparative genomics, genome annotation, microarray technology....
Some of these techniques i have kept as video tutorials(Micro array technique) keeping in view the easy understanding in relatively less time...

GENOMICS

Now then once we are done with the "what's,the who's and where's"...i feel i can start up the topics on BIOINFORMATICS....how bout GENOMICS.....??

So then where are we using our Bioinformatics skills..??

once we are clear with the basic answers of what and how (u have to plz give a feedback on dat)
I am sure we can pass onto the next level so as to know where exactly do we utilize The Great BIOINFORMATICS:
Post Genomic analysis: The journey is not over yet..and our great fraternity of scientists and researchers are still hungry for more knowledge on genomes & proteomes and bout wats unknown the greatest allay of our research team is a team comprising of people called (bioinformaticians) who constantly thrive on making their works easier for genomics study and the study of the unknown..
phylogentic analyses: This is best explained in the forms of phylograms and cladograms..convenient way of observing the evolution..
visualization, modelling and simulation of metabolic pathways and regulatory networks
software tools: part of this is dealt in form of systems biology as treating the tissue system or organism as a complex circuit or machinery..
protein structure prediction,
data processing, data management,
database searches,
gene expression, expression data analysis,
recognition of genes and regulatory elements..

Why Bioinformatics...??

I am sure that this must be the question you might be thinking,that after cramming these many topics and syllabis wat exactly are we planning to achhieve in Bioinformatics...well then here's the answer for that:
provide new insights into biological function
understand the functioning of living things
to apply the knowledge for "improving the quality of life" (in order to understand and fight against diseases)
databases (design, handling, ...)
gene and genome mapping
analysis of gene expression data
search for gene functions
identification of genetic risk factors
pathways, metabolic and regulatoric networks (modelling, simulation, ...)
structure prediction (protein folding, sequence alignments, homology search, ...)
target identification
drug design (docking studies, HTS, screening tests, ...)
gene therapy....