Protein Sequence Generation via Structural Adapter and Language Model

Resolve Bottlenecks,
Find Innovative Solutions
Generate Solutions

Solution Overview

Problem

Current neural structure-based protein design approaches face challenges due to limited available protein structure data and the difficulty in handling structurally non-deterministic regions, leading to sub-optimal sequence predictions and functionally invalid sequences.

Innovation Solution

The proposed solution leverages protein language models to generate sequences prompted by desired structures, utilizing a machine learning model configured with a structural adapter and a pretrained sequence decoder to overcome data limitations and handle non-deterministic regions.

Engineering Contradictions & Design Principles

VSEngineering Contradiction Analysis

1Quantity of substance

If current neural structure-based protein design approaches are used, then protein sequences can be generated from structures, but the available protein structure data is limited and structurally non-deterministic regions cannot be handled effectively

Engineering Contradiction:
Improveavailable protein structure dataVSAvoidsequence prediction accuracy
Core Design Contradiction:
Quantity of substanceVSReliability

Solution Approach 1:

The patent introduces protein language models as an intermediary component between the structure encoder and sequence decoder. These language models, pre-trained on large corpora of protein sequences, serve as mediators that provide evolutionary and functional knowledge to compensate for limited structure data, thereby improving sequence prediction reliability without requiring additional experimental structure data

Inventive Principle:
Principle #24Intermediary (Mediator)

Solution Approach 2:

The patent applies preliminary action by pre-training the language model components on extensive protein sequence data before the actual structure-to-sequence generation task. This pre-training phase allows the model to learn evolutionary patterns, functional constraints, and sequence-structure relationships in advance, enabling it to handle cases with limited structure data and non-deterministic regions more effectively

Inventive Principle:
Principle #10Preliminary action

2Device complexity

If purely structure-based approaches are used, then sequence generation is simplified, but structurally non-deterministic regions produce functionally invalid sequences

Engineering Contradiction:
Improvedesign approach complexityVSAvoidfunctional validity of sequences
Core Design Contradiction:
Device complexityVSReliability

Solution Approach 1:

The patent employs a composite approach by combining structure-based encoding with language model-based sequence generation. The system integrates structural information from the structure encoder with evolutionary and functional knowledge from the pre-trained language model, creating a composite system that maintains the simplicity of structure-based input while incorporating additional knowledge sources to ensure functional validity in non-deterministic regions

Inventive Principle:
Principle #40Composite materials

Solution Approach 2:

The patent applies local quality by allowing different parts of the sequence generation process to utilize different information sources. Structurally deterministic regions are primarily guided by the structure encoder, while structurally non-deterministic regions receive additional guidance from the language model's learned representations, ensuring each region is handled with appropriate information quality

Inventive Principle:
Principle #3Local quality

3Measurement precision

If more protein structure data is collected to improve model performance, then sequence recovery rates may improve, but the cost and time required for data collection increases

Engineering Contradiction:
Improvesequence recovery rateVSAvoiddata collection time
Core Design Contradiction:
Measurement precisionVSLoss of time

Solution Approach 1:

The patent enables the system to serve itself by using pre-trained language models that have already learned from large protein sequence databases. Instead of requiring additional experimental structure data collection, the system leverages the pre-acquired knowledge in the language model to achieve high sequence recovery rates, making the system self-sufficient regarding additional data collection

Inventive Principle:
Principle #25Self-service

Solution Approach 2:

The patent performs the data-intensive learning phase in advance during the pre-training of the language model. By completing the heavy data processing and pattern learning beforehand, the system avoids the need for time-consuming data collection during the actual protein design task, achieving high sequence recovery rates without additional data collection time

Inventive Principle:
Principle #10Preliminary action

Data Source

PatentUS20250173602A1Generating protein sequences using machine learning models
Publication Date: 2025.05.29 BEIJING YOUZHUJU NETWORK TECH CO LTD
  • US20250173602A1 patent drawing
  • US20250173602A1 patent drawing
  • US20250173602A1 patent drawing

AI summary

The present disclosure describes techniques for generating protein sequences using machine learning models. A machine learning model is configured by implanting a structural adapter into a sequence decoder. The machine learning model is configured to generate a protein sequence from a specified structure. The machine learning model is endowed with protein structural awareness by the structural adapter. The machine learning model is equipped with protein sequential evolutionary knowledge by the sequence decoder. The machine learning model comprises the structural adapter, the sequence decoder, and a structure encoder. An initial sequence is generated based on the specified structure by the structure encoder. The protein sequence is optimized through an iterative process. The iterative process comprises progressively refining the protein sequence by iterative decoding. The structural adapter non-linearly imposes representations of the specified structure on a sequence predicted in the iterative process.