Unit-1: Protein Engineering and Bioinformatics
1. Protein Based Products and Designing
Protein engineering involves the modification of existing protein structures or the synthesis of novel proteins to produce molecules with desirable functional properties. This branch of biotechnology integrates molecular biology, computational modeling, and protein chemistry to develop high-value biological products.
Major Categories of Protein-Based Products
- Therapeutic Proteins: Biopharmaceuticals such as human insulin, monoclonal antibodies (e.g., Trastuzumab, Rituximab), growth factors (e.g., Erythropoietin), and tissue plasminogen activator (tPA).
- Industrial Enzymes: Enzymes engineered for harsh industrial environments, including bacterial amylases, lipases, and proteases used in laundry detergents, textiles, and biofuel processing.
- Diagnostic Proteins: Monoclonal and polyclonal antibodies, engineered fluorescent proteins (e.g., GFP variants), and diagnostic enzymes used in ELISA kits and biosensors.
- Agricultural and Food Proteins: Modified proteins like Bacillus thuringiensis (Bt) delta-endotoxins for pest resistance, and chymosin used in cheese manufacturing.
Protein Design Strategies
Designing effective proteins requires precise approaches based on structural availability and targeted functional goals.
Definition: Protein Engineering
The purposeful alteration of amino acid sequences in a protein using recombinant DNA techniques or computer-aided design to create proteins with improved or novel functions.
- Rational Design:
This approach relies on detailed knowledge of the protein's three-dimensional structure and catalytic mechanism. Specific amino acid residues are mutated using site-directed mutagenesis to achieve desired properties like altered substrate specificity or enhanced thermostability.
- Directed Evolution:
This strategy mimics natural evolutionary processes in a laboratory setting. It does not require prior knowledge of the protein's 3D structure. Random mutations are created using techniques like error-prone PCR or DNA shuffling, generating a diverse genetic library, followed by high-throughput screening to select variants with enhanced traits.
Comparison: Rational Design vs. Directed Evolution
| Property | Rational Design | Directed Evolution |
|---|---|---|
| Structural Knowledge | Required (3D structure must be known) | Not required |
| Mutagenesis Method | Site-directed mutagenesis | Error-prone PCR, DNA shuffling |
| Library Size | Small, highly targeted | Large, diverse variant library |
| Screening Strategy | Low throughput needed | High-throughput screening/selection needed |
| Design Basis | Hypothesis-driven computational modeling | Random variation and selective pressure |
Exam Notes and Important Observations
Key Note: Rational design is limited by our understanding of protein folding and sequence-structure relationships, whereas directed evolution is limited primarily by the efficiency of the screening system.
Common Mistake: Confusing site-directed mutagenesis with random mutagenesis. Site-directed targets specific nucleotide codons, whereas random mutagenesis introduces stochastic changes across the whole sequence.
2. Proteins
Proteins are biological macromolecules composed of linear chains of amino acid residues joined by peptide bonds. They execute virtually all biological processes, acting as catalysts, structural scaffolds, transport molecules, and signal transducers.
Definition: Peptide Bond
A covalent amide link formed between the alpha-carboxyl group (-COOH) of one amino acid and the alpha-amino group (-NH2) of an adjacent amino acid with the release of a water molecule.
Four Levels of Protein Structure
- Primary Structure: The unique linear sequence of amino acids linked covalently by peptide bonds along the polypeptide chain backbone.
- Secondary Structure: Highly localized, regular spatial arrangements of the polypeptide backbone stabilized by hydrogen bonds between amide nitrogen and carbonyl oxygen atoms.
- Alpha-helix: A right-handed coiled conformation where hydrogen bonds form between the carbonyl oxygen of residue i and the amide hydrogen of residue i+4.
- Beta-pleated sheet: Strands of polypeptide chains aligned side-by-side (parallel or antiparallel) stabilized by inter-strand hydrogen bonds.
- Turns and Loops: Non-repetitive structure elements that reverse the direction of the chain.
- Tertiary Structure: The complete three-dimensional folding of a single polypeptide chain. Stabilized by non-covalent interactions (hydrophobic interactions, electrostatic salt bridges, hydrogen bonds) and covalent disulfide bridges between cysteine residues.
- Quaternary Structure: The structural assembly formed by the interaction of two or more independent polypeptide chains (subunits), as observed in multimeric proteins like Hemoglobin (a tetramer).
Summary of Protein Structural Hierarchies
| Structural Level | Description | Primary Stabilizing Interactions |
|---|---|---|
| Primary | Linear amino acid sequence | Covalent peptide bonds |
| Secondary | Local folding (Helices, Sheets, Turns) | Backbone hydrogen bonding |
| Tertiary | 3D monomeric conformation | Hydrophobic effect, Disulfide bonds, Salt bridges |
| Quaternary | Multi-subunit association | Hydrophobic interactions, Electrostatic forces |
3. Proteomics: An Introduction
Proteomics is the large-scale, systematic study of the proteome—the entire complement of proteins expressed by a genome, tissue, or organism under specific environmental and physiological conditions.
Definition: Proteome
The complete set of proteins expressed by a given cell type, tissue, or organism at a specific point in time under designated physiological or stress conditions.
Comparison: Genome vs. Proteome
| Feature | Genome | Proteome |
|---|---|---|
| Definition | Entire set of genes in an organism | Entire set of expressed proteins |
| State | Static (constant across somatic cells) | Dynamic (varies with cell state and time) |
| Complexity | Fixed size (~20,000 genes in human) | Highly complex (>1,000,000 protein forms due to alternative splicing and PTMs) |
| Analytical Methods | DNA sequencing, PCR, Microarrays | 2D-PAGE, Mass Spectrometry, Protein Arrays |
Major Branches of Proteomics
- Expression Proteomics: Quantitative identification and comparative analysis of protein expression levels between different biological conditions (e.g., healthy vs. diseased tissue).
- Structural Proteomics: High-throughput identification and mapping of 3D protein structures and protein-protein complexes.
- Functional Proteomics: Determination of protein biological functions, cellular interactions, molecular networks, and post-translational modifications (PTMs).
Key Proteomic Analytical Techniques
- Two-Dimensional Polyacrylamide Gel Electrophoresis (2D-PAGE): A technique that separates complex protein mixtures using two properties:
- First Dimension (Isoelectric Focusing - IEF): Proteins separate according to their Isoelectric Point (pI) in a pH gradient.
- Second Dimension (SDS-PAGE): Separates proteins according to their Molecular Weight (MW) perpendicular to the first axis.
- Mass Spectrometry (MS): Measures the mass-to-charge ratio (m/z) of gas-phase protein or peptide ions. Ionization methods include MALDI (Matrix-Assisted Laser Desorption/Ionization) and ESI (Electrospray Ionization).
Mass-to-charge ratio formula:
m / z = (Mass of Ion) / (Charge of Ion)
4. Introduction to Bioinformatics
Bioinformatics is an interdisciplinary domain combining biology, computer science, mathematics, statistics, and information technology to process, store, analyze, and interpret large volumes of biological data.
Definition: Bioinformatics
The application of computational tools and statistical methods to collect, store, manage, analyze, and visualize biological biological sequence, structural, and functional data.
Primary Goals and Applications
- To organize biological data into structured, easily accessible databases.
- To develop algorithms and statistical software for sequence alignment, pattern recognition, and evolutionary tree construction.
- To analyze biological system behavior, facilitate computer-aided drug design (CADD), and support personalized medicine.
5. Sequences and Nomenclature
Standardized sequence representations and biological nomenclature allow researchers worldwide to seamlessly exchange biological data without ambiguity.
Common Biological Sequence Formats
1. FASTA Format: The standard text format for representing nucleotide or amino acid sequences. It begins with a single description header line starting with a greater-than character (>), followed by sequence lines.
Example FASTA Format:
>sp|P68871|HBB_HUMAN Hemoglobin subunit beta
MVHLTPEEKSAVTALWGKVNVDEVGGEALGRLLVVYPWTQRFFESFGDLSTPDAVMGNPK
VKAHGKKVLGAFSDGLAHLDNLKGTFATLSELHCDKLHVDPENFRLLGNVLVCVLAHHFG
KEFTPPVQAAYQKVVAGVANALAHKYH
IUPAC Amino Acid Nomenclature
| Amino Acid | Three-Letter Code | One-Letter Code |
|---|---|---|
| Alanine | Ala | A |
| Arginine | Arg | R |
| Asparagine | Asn | N |
| Aspartic Acid | Asp | D |
| Cysteine | Cys | C |
| Glutamic Acid | Glu | E |
| Glutamine | Gln | Q |
| Glycine | Gly | G |
| Histidine | His | H |
| Isoleucine | Ile | I |
| Leucine | Leu | L |
| Lysine | Lys | K |
| Methionine | Met | M |
| Phenylalanine | Phe | F |
| Proline | Pro | P |
| Serine | Ser | S |
| Threonine | Thr | T |
| Tryptophan | Trp | W |
| Tyrosine | Tyr | Y |
| Valine | Val | V |
6. Information Sources
Biological databases act as centralized repositories for experimental biological data. They are classified as primary databases (containing raw data direct from experiments) or secondary/curated databases (containing annotated data derived from primary sources).
Primary Biological Databases
- GenBank (NCBI): Comprehensive, publicly accessible database of nucleotide sequences maintained by the National Center for Biotechnology Information.
- EMBL-EBI (ENA): European Nucleotide Archive hosting nucleotide raw reads, assemblies, and functional annotations.
- DDBJ: DNA Data Bank of Japan, which collaborates with GenBank and EMBL through the International Nucleotide Sequence Database Collaboration (INSDC).
Secondary and Specialized Protein Databases
- UniProtKB (Universal Protein Resource Knowledgebase):
- UniProtKB/Swiss-Prot: Manually annotated, peer-reviewed, high-quality protein sequence database.
- UniProtKB/TrEMBL: Automatically annotated, unreviewed protein sequence records.
- PDB (Protein Data Bank): Primary worldwide archive of atomic structural data for biological macromolecules solved via X-ray crystallography, NMR, and Cryo-EM.
- Pfam: Large database of protein domain families, classified using Hidden Markov Models (HMMs).
- KEGG (Kyoto Encyclopedia of Genes and Genomes): Database resource for understanding high-level biological systems and metabolic pathways.
Major Databases Overview
| Database Name | Category | Primary Data Type | Maintained By |
|---|---|---|---|
| GenBank | Primary | Nucleotide Sequences | NCBI (USA) |
| UniProtKB/Swiss-Prot | Secondary / Curated | Protein Sequences & Function | SIB / EBI / PIR |
| PDB | Primary | 3D Macromolecular Structures | wwPDB |
| Pfam | Secondary | Protein Domains & Families | EBI (UK) |
7. Analysis Using Bioinformatics Tools
Computational tools enable sequence alignment, identity scoring, functional prediction, and biological structure modeling.
1. Sequence Alignment Methods
Sequence alignment compares two or more sequences by looking for series of matching character patterns to infer functional, structural, or evolutionary relationships.
- Global Alignment: Aligns sequences across their entire length. Best used for highly similar sequences of equal length. Algorithm: Needleman-Wunsch Algorithm.
- Local Alignment: Identifies regions of highest local similarity within longer, variable sequences. Algorithm: Smith-Waterman Algorithm.
2. Basic Local Alignment Search Tool (BLAST)
BLAST is a fast heuristic sequence comparison tool used to search database libraries for matching sequence regions based on statistical significance.
| BLAST Variant | Query Type | Database Search Type | Primary Application |
|---|---|---|---|
| BLASTN | Nucleotide | Nucleotide | Mapping genomic DNA / RNA sequences |
| BLASTP | Protein | Protein | Identifying protein homologs and function |
| BLASTX | Translated Nucleotide (6-frame) | Protein | Annotating coding regions in novel transcripts |
| TBLASTN | Protein | Translated Nucleotide (6-frame) | Finding protein-encoding genes in genomic DNA |
| TBLASTX | Translated Nucleotide (6-frame) | Translated Nucleotide (6-frame) | Cross-species gene comparison without protein annotation |
3. Protein Structure Prediction Methods
- Homology Modeling (Comparative Modeling): Constructs 3D models of an unknown target protein based on solved structural templates that share significant sequence identity (>30%).
- Fold Recognition (Threading): Compares a sequence against structural template libraries to find compatible folding patterns when sequence identity is low (<25%).
- Ab Initio (De Novo) Prediction: Predicts three-dimensional structural conformations purely from the primary sequence using physical principles and energetic calculations without template structures.
Sequence Alignment Score Equation:
Score = Sum of Match Scores - Sum of Mismatch Penalties - Sum of Gap Penalties