Proteins, biological macromolecules composed of amino acid units, have a complex structure that can be categorized at four levels of organization: (i) primary structure, which refers to the sequence of amino acids; (ii) secondary structure, which includes highly regular substructural motifs such as alpha helix and beta sheet; (iii) tertiary structure, which consists of three-dimensional conformations essential for the biological activity of the protein; and (iv) quaternary structure, which involves the interaction of several units forming a supramolecular arrangement [1].
The chemical and physical properties of proteins are influenced not only by their amino acid sequences, but also by their three-dimensional structure, which in turn can undergo modifications. Proteins interact via their interfaces, i.e., specific areas of protein surfaces that are geometrically and physico-chemically complementary, allowing for shape-wise and energetically favorable interactions to occur [2]. For drug design purposes, it is crucial to investigate the protein–protein and protein–small-molecule interactions in order to uncover potential binding or interactions, which are often related to the geometrical features of the proteins involved. To this end, this study focuses specifically on the key variants of the spike protein of SARS-CoV-2. This is an important case study of current interest, but applications to other cases can be made in a similar manner.
SARS-CoV-2, the strain of the Coronavirus responsible for COVID-19, contains a viral particle that is studded with 24–40 spike proteins [3]. These spike proteins are randomly positioned on the surface and play a crucial role in enabling the virus to fuse with the host cell. In contrast to the mostly rigid fusion proteins found in influenza viruses, SARS-CoV-2 fusion proteins are flexible, facilitating the bond with host cells [4]. The spike protein is a trimeric glycoprotein composed of three identical associated chains. It consists of a region resembling a flower's stem, with the Receptor Binding Domain (RBD) forming the corolla. The RBD is responsible for contacting the host cells and initiating infection. This part of the molecule is flexible and undergoes large motions to interact with the receptor [5].
The spike protein consists of two functional subunits: the N-terminal S1 subunit, which is responsible for binding to the host cell receptor, and the C-terminal S2 subunit, which is responsible for fusion between the host cell and the virus. The initial and crucial step in the fusion process is the binding between the S1 domain and the host cell. This binding occurs due to the RBD on the N-terminal S1 domain of the Spike protein, which binds to the ACE2 receptor of the host cell. The fusion process is then facilitated by the Transmembrane Protease Serine 2 (TMPRSS2), which is present on the surface of the host cell, following binding to the ACE2 receptor. TMPRSS2 cleaves a site on the S2 subunit of the spike protein and exposes a set of hydrophobic amino acids that facilitate the fusion of viral and cell membranes, allowing the viral genome to enter the host cell directly. This action leads to the transformation of the rough endoplasmic reticulum's conformation, creating vesicles that enable viral RNA replication and translation. As a result, viral proteins are produced at the cost of native ones, causing a suppression of the cell's gene expression in favor of the viral one. The newly formed viral proteins combine to form a complete viral particle following viral RNA translation. This particle exits the cell by exocytosis through lysosomes and the Golgi complex, ready to infect another human host cell [6].
Below, we describe the background to the geometry-based methodology used in the analysis of biomolecular interactions. In the early 1980s, Kuntz et al. [7] introduced a geometric approach to study interactions between macromolecules and ligands. They proposed a method to explore geometrically feasible alignments of ligands and receptors with known structures, assuming that the protein can be approximated as a rigid body (a sphere) with binding sites represented by pockets or grooves. In their study, they used a graphing program to generate a set of spheres that could fill all the pockets and grooves on the surface of the protein. Meanwhile, the ligand molecule can be represented by a set of spheres that approximate the space it occupies. If the ligand and protein exhibit a favorable affinity, their respective sets of spheres should be compatible. This compatibility can be used to identify viable protein–ligand interactions and to classify them based on the molecular and geometric properties of the protein. Using a relative combined score based on putative interface sites on the surface of protomers, Jones and Thornton [8] analyzed residue patches on the surfaces of protein structures in order to predict the locations of protein–protein interactions (PPIs). Later, Lee et al. [9] developed a method for classifying binding sites based on geomorphic features of the Earth (cave, crater, canyon, plain, and valley). A ligand, or probe, was used to classify the binding sites, which were approximated with a specific number of unitary spheres in an approach similar to that employed by Kuntz et al. The compatibility between the number of spheres and their positioning was used to determine the type of associated binding site.
A recent strategy for classifying proteins has been to employ machine-learning models trained on a set of features derived from proteins' sequences and/or structures. Weston et al. [10] used semi-supervised classification in conjunction with kernel clustering, achieving 94.3 % accuracy. (Their approach is considered semi-supervised because it uses both labeled and unlabeled data for protein classification.) Two kernel cluster methods were tested: the neighborhood kernel and the bagged kernel. The accuracy was measured relative to existing cluster kernel methods, with the authors reporting that their methods outperformed these alternatives. Tsuda et al. [11] addressed the computational challenges of protein classification using semidefinite programming-based Support Vector Machines (SDP/SVMs). Their aim was to introduce an efficient method for classifying proteins using multiple protein networks in order to overcome the high computational costs associated with large datasets and protein networks. Their method allows for the direct incorporation of various protein networks and vectorial data, with combination weights determined through convex optimization. Notably, this approach reduces computation time significantly while maintaining a degree of accuracy comparable to that of SDP/SVM, making it a promising solution for protein classification. Another notable study explored the use of supervised machine learning to automate the classification of protein structures [12]. It evaluated 15 algorithms on a dataset of 11,360 protein domain pairs and found that the boosted random forest algorithm performed the best, classifying protein structures at different levels within the Structural Classification of Proteins hierarchy with 97.0 % accuracy. This research demonstrated the potential of machine learning for accurate protein structure classification, especially in less common classes. Yu et al. [13] developed a PPI predictor using Relaxed Variable Kernel Density Estimator (RVKDE) as a classification method. The proposed predictor utilizes a robust classification algorithm designed to effectively manage datasets with significant imbalances, i.e., situations where one class outnumbers the others in terms of observations. In the realm of binary protein-protein interaction (PPI) classification, this algorithm adeptly addresses the imbalance in number of samples between the positive class (interacting protein pairs) and the negative class (non-interacting pairs). Their method was evaluated using several unbalanced datasets with varying positive-to-negative ratios (ranging from 1:1 to 1:5).
Managing data imbalance is crucial in PPI prediction, since the number of interacting protein pairs is much smaller than that of non-interacting ones. Northey et al. [2] presented IntPred, a machine-learning-based tool for predicting protein–protein interfaces. It was found to outperform several established methods in terms of specificity, making it a valuable tool for predicting these interfaces when crystal structures are unavailable. The choice between IntPred and other methods is application-dependent due to varying sensitivity-specificity trade-offs. In contrast, Paladin et al. [14] presented RepeatsDB 3.0, a database that classifies and annotates protein tandem repeat structures from the PDB. It introduces a hierarchical classification system that combines structural and sequence similarity, including levels such as 'Class,' 'Topology,' 'Fold,' 'Clan,' and 'Family.' This update aims to enhance data organization and user experience, providing a unified view of structures from similar sequences and serving as a valuable resource for researchers exploring tandem repeat protein structures. Das and Chakrabarti [15], meanwhile, introduced a computational method that uses protein–protein interface properties and SVMs to distinguish reliable protein complexes from non-native ones generated by docking. Their approach was found to perform well for both homo- and hetero-complexes and was also applied to predict interfaces for experimentally verified protein interactions without structural data. They also developed a web server, PCPIP, for assessing PPIs and therapeutic potential by comparing interacting interfaces in protein–protein dimer complexes to known ones.
Deep learning techniques have also yielded significant results in PPI research, as outlined by Gao et al. [16]; a notable breakthrough occurred in recent years with the development of DeepMind's AlphaFold [17,18], which uses neural networks to accurately and efficiently predict protein structures and protein–protein complex models. AlphaFold employs a three-track network to integrate the proteins' one-dimensional sequence information, 2D map distance, and 3D coordinate data. This approach has enabled the creation of the most extensive archive of human protein structures available, with over 200 million protein structures, thereby revolutionizing applications of biological research ranging from drug design, to enzyme design, genetically modified organisms, and virus research. As another example of the use of deep learning in this domain, Gainza et al. [19] developed a method to predict protein interactions based on their shared fingerprints. They utilized geometric deep learning to extract fingerprints for three specific tasks: protein pocket-ligand prediction, PPI location, and prediction of protein–protein complexes. Abdin et al. [20], meanwhile, presented two innovative methods for predicting protein–peptide binding sites, PepNN-Struct and PepNN-Seq. It should be noted in this regard that protein–peptide binding is a vital interaction in biology but a challenging one to model due to the flexibility of peptides. PepNN uses reciprocal attention and graph neural networks, achieving good performance on various datasets. It also identifies potential peptide binding proteins, addressing the lack of structural data for disease and infection research. As another notable contribution in this area, Dai et al. [21] developed PInet, a unified Geometric Deep Neural Network, to improve the precision and recall of PPI interface prediction. This represents a crucial step toward better understanding molecular processes, as experimental methods for determining such interfaces are limited in scalability. PInet combines data-driven and physics-based approaches by analyzing the structures of interacting proteins, achieving impressive performance in identifying interaction regions and providing valuable insights into the underlying physical complementarity driving molecular recognition. Gligorijevic et al. [22] presented DeepFRI, a Graph Convolutional Network (GCN) that predicts protein functions by combining sequence data from a protein language model and protein structures. DeepFRI was found to outperform existing methods and to be capable of effectively handling large databases. It was also found to be capable of predicting functions using protein models (as an alternative to experimental structures) with minimal loss in accuracy. The model enables precise site-specific annotations and predicts protein functions with a high degree of confidence. DeepFRI is available as a web server for practical use, addressing the challenge of automated function prediction in the context of the growing body of protein sequence data and diverse functions. As yet another notable application of deep learning, Liu et al. [23] introduced GeoPPI, a deep learning model designed to predict changes in binding affinity based on the 3D geometrical representation of proteins. GeoPPI was found to be capable of differentiating between the binding affinity of various SARS-CoV-2 antibodies and the RBD of the spike protein. Mallet et al. [24], meanwhile, created InDeep, a tool that predicts functional binding sites within proteins, with a focus on PPIs. These interactions are crucial in drug discovery, but designing drugs for them is challenging. InDeep, using deep learning and a curated dataset, was found to outperform existing predictors and to be capable of identifying potential binding sites for proteins or future drugs, aiding drug design for PPIs by pinpointing relevant binding pockets near PPI interfaces. In another notable study, Orasch et al. [25] used graph representation learning and transfer learning to predict interaction sites and interactions of proteins based on their geometrical representation. Their approach represented the relationships between local molecular interactions as edges in a graph representation, which in turn was used by a neural model to learn how to distinguish between different PPIs.
In a recent application of emerging technologies, Di Grazia et al. [26] applied features commonly used for three-dimensional face analysis to characterize proteins. Both human faces and proteins are considered free forms in terms of their geometry, enabling the use of features from differential geometry to describe them. As an example of the application of this method, tubulin was investigated in detail. Following the geometry-based feature extraction phase, two classifiers, an SVM and a k-means method, were employed to classify the isotypes of the human β tubulin protein as well as to differentiate them from tubulin's bacterial counterpart, FtsZ. Their findings suggest that geometry could represent a viable new pathway for protein classification, albeit one that has yet to be thoroughly explored.
The present work investigates the geometrical properties of SARS-Cov-2 spike proteins in order to understand why the Omicron variant is more contagious than the Alpha and Delta variants [27,28], with a focus on the binding interface of spike with the ACE2 protein of the host cell. To achieve this goal, descriptors from the Differential Geometry background [29,30] were mapped point-by-point onto the protein surfaces. Moreover, an SVM classifier was adopted in order to automatically categorize the surfaces based on shape affinity, while, in parallel, Root Mean Square Error was used to compare them.
Comments (0)