Peptide Libraries for Undruggable Protein Interaction Screening
Find Innovative SolutionsGenerate Solutions
Solution Overview
Problem
Current genomics-based technologies are limited in identifying and targeting 'undruggable' protein-protein interactions and novel drug sites at the proteome level due to their genetic focus rather than protein-level screening, which restricts the identification of druggable candidates and understanding how to inhibit them effectively.
Innovation Solution
Development of a library of nucleic acids encoding peptides from naturally occurring proteins, allowing for high-throughput phenotypic screening and targeting of protein-protein interactions by encoding a diverse set of peptides that can interact with mammalian proteins, overcoming limitations of bacterial-derived libraries which are underpowered in interacting with human proteins.
Engineering Contradictions & Design Principles
Engineering Contradiction Analysis
1Productivity
If genomics-based technologies (CRISPR, RNAi) are used for high-throughput screening, then screening capacity and throughput are improved, but the ability to identify druggable protein targets and understand inhibition mechanisms deteriorates because these methods screen at the genetic level rather than protein level
Solution Approach 1:
The patent introduces a peptide library as an intermediary between genetic screening and small molecule drug discovery. The peptides serve as mediators that can physically interact with protein targets, allowing the screening process to capture both target identification and druggability information simultaneously. This intermediary layer enables the system to maintain high throughput while recovering the lost druggability information.
Solution Approach 2:
The patent replaces the genetic-level mechanical system (CRISPR/RNAi) with a protein-level system (peptide library screening). By substituting the level of biological organization being screened, the system can maintain throughput while gaining the ability to identify druggable targets and understand inhibition mechanisms at the protein level.
2Ease of manufacture
If bacterial-derived protein-fragment libraries are used for Protein-i screening, then library generation is simplified and throughput is improved, but the functional interaction capability with mammalian proteins deteriorates because bacterial protein fragments are underpowered in interacting with human proteins
Solution Approach 1:
The patent changes the source parameter of the protein fragments from bacterial to human/mammalian origin. This parameter change fundamentally alters the interaction capability of the library members with human proteins, while the high-throughput screening methodology remains unchanged, thus maintaining ease of manufacture while improving reliability.
Solution Approach 2:
The patent creates a copy of the human proteome in peptide form, rather than using bacterial protein fragments. This copying approach allows the library to naturally possess the correct interaction interfaces for human proteins, solving the compatibility issue while maintaining the simplicity of library generation through synthetic biology approaches.
3Reliability
If mammalian genome is used to create protein-fragment libraries, then functional interaction with human proteins is improved, but the complexity of library construction increases significantly due to large number of coding sequences and manual cloning requirements
Solution Approach 1:
The patent replaces the manual cloning mechanical system with a synthetic biology approach. By using in silico design and automated synthesis methods, the system can handle the complexity of mammalian genome-derived peptide libraries without requiring manual intervention for each clone, thus maintaining reliability while reducing construction complexity.
Solution Approach 2:
The patent changes the construction methodology parameter from manual cloning to automated synthetic biology approaches. This parameter change enables the system to process the large number of coding sequences in mammalian genomes efficiently, reducing the perceived complexity while maintaining the ability to generate functionally relevant peptide libraries.
4Manufacturing precision
If peptide libraries with longer peptides are used to represent native protein structures, then structural accuracy is improved, but the ability to describe discrete spatial sites for small molecule docking deteriorates
Solution Approach 1:
The patent segments the peptide library into multiple length categories (15-25, 25-50, and 50-100 amino acids). This segmentation allows different peptide lengths to serve different functions: longer peptides maintain structural accuracy while shorter peptides provide precise spatial site definitions. The segmented approach resolves the contradiction by distributing the functional requirements across different library subsets.
Solution Approach 2:
The patent applies local quality by having different regions of the peptide library exhibit different properties. Some peptides are optimized for structural representation (longer), while others are optimized for spatial site definition (shorter). This local differentiation allows the overall library to satisfy both requirements simultaneously through compositional diversity.
Data Source
Figure 1
Figure 2
Figure 3
AI summary
The present invention relates to nucleic acid libraries, peptide libraries and uses thereof. The invention relates to libraries of nucleic acids that encode a plurality of peptides that represent fragments of naturally occurring proteins. In particular, the invention relates to a library of nucleic acids, each nucleic acid comprising a coding region of defined nucleic acid sequence encoding for a peptide having a length of between 25 and 110 amino acids, and having an amino acid sequence being a region of a sequence selected from the amino acid sequence of a naturally occurring protein of one or more organisms; wherein the library comprises nucleic acids that encode for a plurality of at least 10,000 different such peptides, and wherein the amino acid sequence of each of at least 50 of such peptides is a sequence region of the amino acid sequence of a different protein of a plurality of different such naturally occurring proteins.