Nucleic acid sequence iterative optimization method and device based on multiple agents

By iteratively optimizing nucleic acid sequences using a multi-agent system, the problem of multi-objective optimization conflict in existing technologies is resolved, and the synergistic improvement of nucleic acid sequences in terms of expression efficiency, stability, immunogenicity, and degradation risk is achieved, resulting in a better nucleic acid sequence design.

CN121789780APending Publication Date: 2026-04-03INST OF COMPUTING TECH CHINESE ACAD OF SCI
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-25
Publication Date
2026-04-03

AI Technical Summary

Technical Problem

Existing nucleic acid sequence design technologies suffer from conflicting and difficult-to-coordinate objectives in multi-objective optimization, and lack iterative optimization capabilities, which makes the optimization process prone to getting trapped in local optima and unable to achieve comprehensive improvement in expression efficiency, stability, immunogenicity, and degradation risk.

Method used

A multi-agent system is adopted, in which multiple agents evaluate nucleic acid sequences separately, generate multiple candidate sequences, and integrate optimization strategies through hypothesis generation agents to perform comprehensive evaluation and iterative optimization, thereby achieving multi-objective collaborative optimization.

Benefits of technology

Under conditions of conflicting objectives, the system automatically adjusts the exploration path, continuously proposes improved hypotheses, breaks out of local optima, significantly improves the global optimization efficiency of nucleic acid sequences, and obtains nucleic acid sequences with more balanced performance and better overall indicators.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121789780A_ABST
    Figure CN121789780A_ABST
Patent Text Reader

Abstract

The invention provides a nucleic acid sequence iterative optimization method based on multiple agents. The method comprises the following steps: acquiring a nucleic acid sequence to be optimized and setting a plurality of targets; calling a multi-single-target agent to carry out multi-single-target evaluation on the nucleic acid sequence to give a single-target optimization strategy; a hypothesis generation agent is called to integrate multiple single target optimization strategies to form a unified optimization strategy, and a sequence generation agent is called to generate multiple candidate sequences; calling a multi-single-target agent to evaluate the multiple candidate sequences to obtain a multi-single-target score, and calling a comprehensive evaluation agent to perform comprehensive quantitative evaluation according to the multiple single-target scores of the candidate sequences to obtain a comprehensive score; and selecting the candidate sequence with the highest comprehensive score to enter a next round of iteration until a final optimized sequence is generated. The invention further provides a multi-agent-based nucleic acid sequence iterative optimization device, a storage medium and electronic equipment. Therefore, the sequence global optimization efficiency can be remarkably improved, and the nucleic acid sequence with more balanced performance and better comprehensive indexes can be obtained.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the fields of biomedicine and bioinformatics analysis technology, and in particular to a method, apparatus, storage medium and electronic device for nucleic acid sequence iterative optimization based on multiple agents. Background Technology

[0002] With the rapid development of nucleic acid drugs (including mRNA, siRNA, ASO, etc.) in the fields of vaccines, tumor immunotherapy, genetic disease treatment, and gene regulation, the optimized design of nucleic acid sequences has become a key step in improving drug expression efficiency, stability, half-life, and safety. Existing nucleic acid sequence design technologies mainly include artificial design, biochemical rule-based methods, statistical learning models, and deep learning-based sequence optimization methods.

[0003] In early technologies, nucleic acid sequence optimization largely relied on empirical rules or sequence characteristic statistics. For example, mRNA expression efficiency was improved by adjusting codon usage bias and increasing the codon adaptation index (CAI); stability was improved by reducing GC (guanine, cytosine) content and decreasing unstable motifs (such as AU-rich elements). Immunogenicity risk was reduced by compressing CpG content. Replacing all codons in a foreign mRNA sequence with the host cell's optimal codons might actually prevent protein expression, as some proteins require rare codons for expression, slowing ribosome movement and providing sufficient time for proper protein folding. The codon adaptation index (CAI) refers to the degree of similarity between the frequency of codon usage in a foreign mRNA sequence and the host cell's optimal codons. The closer this value is to 1, theoretically, the higher the protein expression of the foreign mRNA in the host cell. Therefore, the most basic principle of codon optimization is to replace codons in the exogenous mRNA sequence with synonyms that are frequently used in the host cell, ensuring a better match between the codon usage bias in the exogenous mRNA sequence and the host cell, and avoiding the use of rare codons. However, we need to know that codons are not the only factor affecting protein expression; other factors also exist, such as rare codons, GC content, and secondary structure (free energy).

[0004] In 2006, Grzegorz Kudla et al. published an article titled "High Guanine and Cytosine Content Increases mRNA Levels in Mammalian Cells," which found that in mammalian cells, genes rich in GC are expressed several to a hundred times more efficiently than genes with low GC content. This phenomenon is due to the more efficient mRNA transcription or processing of genes rich in GC, resulting in more stable mRNA.

[0005] In 2009, Grzegorz Kudla et al., in their paper "Coding-sequence determinants of gene expression in *Escherichia coli*", constructed 154 GFP mRNAs with randomly mutated synonymous codons and placed them under the same promoter to study the effect of synonymous codon mutations on protein expression. The results showed that the correlation between the fluorescence signal characterizing GFP protein expression and the codon fitness index (CAI) was not very strong; some GFP mRNAs showed high fluorescence signals but low CAI. A strong correlation was found between the folding free energy at the 5' end of GFP mRNA and the fluorescence signal of GFP protein expression. Highly expressed GFP mRNAs had many unpaired nucleotides near the 5' start codon, resulting in a low folding free energy; low-expressed GFP mRNAs formed a typical long hairpin structure at the 5' end, resulting in a high folding free energy, which limited protein translation initiation (the rate-limiting step in protein expression during translation initiation).

[0006] In 2019, David M. Mauger et al. published a paper titled "mRNA structure regulates protein expression through changes in functional half-life," which demonstrated that the fewer secondary structures formed by the first ten codons of the 5' UTR+CDS region of mRNA, the higher the expression level of the encoded protein; and the more secondary structures formed by the remaining CDS region + 3' UTR, the higher the expression level of the encoded protein (higher SHAPE activity indicates fewer RNA secondary structures, i.e., more relaxed).

[0007] In October 2021, the Baidu team published an article titled "LinearDesign: Efficient Algorithms for Optimized mRNA Sequence Design," which developed a new algorithm to more effectively optimize mRNA sequences. The website is http: / / rna.baidu.com / . The algorithm takes into account both the codon fitness index (CAI) and folding free energy (MFE) of mRNA, and can obtain more structurally stable mRNA, prolong the mRNA half-life and protein expression time, thereby increasing the final intracellular mRNA yield.

[0008] Internationally, companies like Moderna and Amazon are collaborating to address sequence optimization challenges. Abogen Biosciences, founded in 2019, has advanced its mRNA COVID-19 vaccine to Phase II clinical trials, while Stemirna Therapeutics' mRNA vaccine has entered Phase I clinical trials. In addition, companies such as Deepin Biotech, Bluemage Biotech, Livanda, Jiachen Haixi, and Zhiyuan Messenger all have investments in mRNA drugs. Domestically, Stemirna Therapeutics, through collaboration with global companies and leveraging AI + cloud computing technologies, has developed a unique algorithmic technology (proprietary) to correct digital models with experimental data, improving the accuracy and efficiency of prediction and sequence design. Compared to traditional sequence optimization platforms, this technology can increase mRNA expression efficiency by 3-4 times and reduce the rate of mRNA degradation.

[0009] Although existing nucleic acid sequence design methods have made some progress in expression optimization, structure prediction, and stability assessment, they still have significant shortcomings when dealing with complex nucleic acid drug design scenarios. The main problems are as follows:

[0010] First, existing technologies are mostly unidirectional optimizations and lack cross-target collaborative capabilities.

[0011] Current nucleic acid sequence optimization techniques typically target a specific area (e.g., expression efficiency, CpG reduction, structural stability) independently, and these different optimization objectives often have inherent conflicts. For example, increasing expression efficiency is often accompanied by increased GC content, which may lead to decreased structural stability or increased immunogenicity; conversely, reducing immunogenicity may lead to decreased translation efficiency. Existing technologies lack mechanisms for effective coordination and comprehensive decision-making among multiple objectives, making it difficult to balance overall performance across multiple dimensions.

[0012] Second, it gets stuck in a fixed process and lacks an iterative optimization mechanism.

[0013] Many existing technologies employ static scoring methods or one-time generation methods, which cannot automatically adjust strategies based on new information emerging during the optimization process, nor can they generate directions for "how to optimize in the next round" based on historical results. Without hypothesis-driven and stochastic exploration space for candidate sequence generation, sequence optimization often presents a linear or fixed process, unable to proactively summarize the reasons for failure, infer the most critical features, or adaptively propose new optimization directions.

[0014] In conclusion, the existing technology obviously has inconveniences and defects in practical use, so it is necessary to improve it. Summary of the Invention

[0015] To address the aforementioned shortcomings, the present invention aims to provide a multi-agent nucleic acid sequence iterative optimization method, apparatus, storage medium, and electronic device, which can automatically adjust the exploration path under conditions of target conflict, continuously propose improved hypotheses and escape local optima, thereby significantly improving the global optimization efficiency of the sequence and obtaining nucleic acid sequences with more balanced performance and better overall indicators.

[0016] To solve the above-mentioned technical problems, the present invention is implemented as follows:

[0017] In a first aspect, embodiments of the present invention provide a nucleic acid sequence iterative optimization method based on multiple agents, comprising:

[0018] The steps for obtaining the set parameters are as follows: obtain the nucleic acid sequence to be optimized, and set multiple targets that need to be optimized in a coordinated manner;

[0019] The single-objective evaluation step involves calling multiple single-objective agents to perform multiple single-objective evaluations on the nucleic acid sequence and providing corresponding single-objective optimization strategies.

[0020] The integration and optimization steps involve calling a hypothesis generation agent to integrate multiple single-objective optimization strategies into a unified optimization strategy, and then calling a sequence generation agent to generate multiple candidate sequences based on the unified optimization strategy and the nucleic acid sequence.

[0021] The comprehensive evaluation step involves calling multiple single-target agents to evaluate multiple candidate sequences separately, obtaining multiple single-target scores for each candidate sequence, and then calling a comprehensive evaluation agent to perform a comprehensive quantitative evaluation based on the multiple single-target scores of each candidate sequence to obtain the corresponding comprehensive score.

[0022] The iterative optimization step involves selecting the candidate sequence with the highest comprehensive score to enter the next round of iteration, and repeating the above steps until the final optimized sequence is generated.

[0023] According to the multi-agent-based nucleic acid sequence iterative optimization method of the present invention, the step of obtaining the setting further includes:

[0024] Obtain the nucleic acid sequence to be optimized;

[0025] Several objectives that need to be optimized collaboratively are defined, including stability, expression efficiency, immunogenicity and / or degradation risk;

[0026] Set up a corresponding scoring tool, which is called by multiple intelligent agents, and set constraints.

[0027] According to the multi-agent-based nucleic acid sequence iterative optimization method of the present invention, the single-target agent in the single-item evaluation step further includes:

[0028] An expression efficiency agent is used to provide a first single-objective optimization strategy to improve expression efficiency.

[0029] A stability agent is used to provide a second single-objective optimization strategy to improve stability.

[0030] An immunogenicity agent is used to provide a third single-objective optimization strategy to enhance immunogenicity;

[0031] A degradation risk agent is used to provide a fourth single-objective optimization strategy to reduce degradation risk.

[0032] According to the multi-agent-based nucleic acid sequence iterative optimization method of the present invention, the integration optimization step further includes:

[0033] The hypothesis-generating agent is invoked to integrate multiple single-objective optimization strategies into a unified optimization strategy;

[0034] The sequence generation agent is invoked to generate multiple candidate sequences based on the unified optimization strategy and the nucleic acid sequence, the candidate sequences including:

[0035] An expression efficiency improvement sequence that focuses on expression efficiency;

[0036] An improved stability sequence with an emphasis on stability.

[0037] An immunogenicity-enhanced sequence with a focus on immunogenicity;

[0038] An improved degradation risk sequence that focuses on degradation risk;

[0039] Multiple integrated improvement sequences combining the above directions;

[0040] An improved sequence that combines the above approaches and incorporates an exploratory sequence with random mutations;

[0041] The expression efficiency improvement sequence, stability improvement sequence, immunogenicity improvement sequence, degradation risk improvement sequence, comprehensive improvement sequence, and exploration sequence are all generated through codon substitution, local fine-tuning, or random exploration while keeping the amino acid sequence unchanged, and serve as the candidate sequences for this round.

[0042] According to the multi-agent-based nucleic acid sequence iterative optimization method of the present invention, the comprehensive evaluation step further includes:

[0043] Multiple single-target agents are invoked to evaluate multiple candidate sequences respectively, resulting in multiple single-target scores for each candidate sequence;

[0044] Then, the comprehensive evaluation agent is invoked to perform a comprehensive quantitative evaluation based on the multiple single-objective scores and their respective weights of each candidate sequence to obtain the corresponding comprehensive score.

[0045] According to the multi-agent-based nucleic acid sequence iterative optimization method of the present invention, the iterative optimization step further includes:

[0046] The candidate sequence with the highest comprehensive score is selected to proceed to the next iteration.

[0047] The process of obtaining the setting step, the single evaluation step, the integration optimization step, the comprehensive evaluation step, and the iterative optimization step is executed repeatedly until the preset termination condition is met, and the final optimization sequence is output.

[0048] According to the multi-agent-based nucleic acid sequence iterative optimization method of the present invention, the termination condition includes:

[0049] The overall score of the best sequence in a certain round reaches a preset threshold;

[0050] The overall score improvement in multiple consecutive rounds is less than the set increment;

[0051] Reaching the maximum number of iterations;

[0052] When execution is terminated, the candidate sequence with the highest comprehensive score in the last round is taken as the final optimized sequence.

[0053] Secondly, embodiments of the present invention provide a nucleic acid sequence iterative optimization device based on multiple agents, comprising:

[0054] The acquisition and setting module is used to acquire the nucleic acid sequence to be optimized and set multiple targets that need to be optimized in a coordinated manner;

[0055] The single-objective evaluation module is used to call multiple single-objective agents to perform multiple single-objective evaluations on the nucleic acid sequence and provide corresponding single-objective optimization strategies.

[0056] The integration and optimization module is used to call the hypothesis generation agent to integrate multiple single-objective optimization strategies into a unified optimization strategy, and then call the sequence generation agent to generate multiple candidate sequences based on the unified optimization strategy and the nucleic acid sequence.

[0057] The comprehensive evaluation module is used to call multiple single-target agents to evaluate multiple candidate sequences respectively, obtain multiple single-target scores for each candidate sequence, and then call the comprehensive evaluation agent to perform a comprehensive quantitative evaluation based on the multiple single-target scores of each candidate sequence to obtain the corresponding comprehensive score.

[0058] The iterative optimization module is used to select the candidate sequence with the highest comprehensive score to enter the next round of iteration, and repeat the above steps until the final optimized sequence is generated.

[0059] Thirdly, embodiments of the present invention provide a storage medium for storing a computer program for performing any of the methods described herein.

[0060] Fourthly, embodiments of the present invention provide an electronic device, including a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, wherein the processor executes the computer program to implement any of the methods described above.

[0061] Therefore, this invention provides a multi-agent-based nucleic acid sequence iterative optimization technique, comprising: acquiring the nucleic acid sequence to be optimized and setting multiple objectives requiring collaborative optimization; invoking multiple single-objective agents to perform multiple single-objective evaluations on the nucleic acid sequence and provide single-objective optimization strategies; invoking a hypothesis generation agent to integrate the multiple single-objective optimization strategies into a unified optimization strategy, and then invoking a sequence generation agent to generate multiple candidate sequences; invoking multiple single-objective agents to evaluate the multiple candidate sequences separately, obtaining multiple single-objective scores for each candidate sequence, and then invoking a comprehensive evaluation agent to perform a comprehensive quantitative evaluation based on each candidate sequence to obtain the corresponding comprehensive score; selecting the candidate sequence with the highest comprehensive score to enter the next iteration, until the final optimized sequence is generated. This invention, by introducing a multi-agent hypothesis-driven and collaborative exploration mechanism, enables coordinated joint exploration of multiple optimization objectives such as stability, expression efficiency, immunogenicity, and degradation risk of nucleic acid sequences, no longer relying on local fine-tuning in a single direction. Therefore, this invention can automatically adjust the exploration path under conditions of conflicting objectives, continuously propose improved hypotheses and escape local optima, thereby significantly improving the global optimization efficiency of the sequence and obtaining nucleic acid sequences with more balanced performance and better overall indicators. Attached Figure Description

[0062] Figure 1This is a flowchart illustrating the nucleic acid sequence iterative optimization method based on multiple agents provided in Embodiment 1 of the present invention;

[0063] Figure 2 This is a schematic diagram of the structure of the nucleic acid sequence iterative optimization device based on multiple agents provided in Embodiment 2 of the present invention;

[0064] Figure 3 This is a schematic diagram of the structure of the electronic device provided in Embodiment 3 of the present invention. Detailed Implementation

[0065] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the invention.

[0066] It should be noted that references to "an embodiment," "embodiment," "example embodiment," etc., in this specification refer to the described embodiment including specific features, structures, or characteristics, but not every embodiment must include these specific features, structures, or characteristics. Furthermore, such expressions do not refer to the same embodiment. Moreover, when describing specific features, structures, or characteristics in conjunction with embodiments, whether or not explicitly described, it is indicated that incorporating such features, structures, or characteristics into other embodiments is within the knowledge of those skilled in the art.

[0067] Furthermore, certain terms are used in the specification and subsequent claims to refer to specific components or parts. Those skilled in the art will understand that manufacturers may use different names or terms to refer to the same component or part. This specification and subsequent claims do not distinguish components or parts by differences in name, but rather by differences in function. The terms "comprising" and "including" used throughout the specification and subsequent claims are open-ended and should be interpreted as "including but not limited to." Additionally, the term "connection" here includes any direct and indirect electrical connection means. Indirect electrical connection means include connections made through other means.

[0068] The following description, in conjunction with the accompanying drawings, details the multi-agent-based nucleic acid sequence iterative optimization method provided by the present invention through specific embodiments and application scenarios.

[0069] This invention belongs to the field of biomedicine and bioinformatics analysis technology, specifically relating to intelligent design and optimization methods for nucleic acid sequences. More particularly, this invention relates to a method for iterative optimization of nucleic acid sequences using a multi-agent hypothesis generation and collaborative exploration mechanism, encompassing the automated generation of multi-objective comprehensive evaluation and improvement strategies for sequence stability, expression efficiency, immunogenicity, degradation risk, and other factors.

[0070] In their research on nucleic acid drug sequence design and performance optimization, the inventors analyzed the structure, stability, expression efficiency, and immunogenicity of a large number of mRNA, siRNA, and ASO sequences. They discovered that existing technologies generally suffer from key problems such as difficulty in co-optimizing multiple objectives, susceptibility to local optima, and a lack of iterative optimization capabilities. Further research revealed that the root cause of these problems lies in the fact that most existing methods rely on fixed rules or single-model outputs, failing to achieve interactive decision-making among multiple objectives, feedback-based generation of new candidate sequences, and dynamic exploration of the search space.

[0071] To address the inherent shortcomings of existing nucleic acid sequence design technologies in multi-objective collaboration, iterative reasoning, and global exploration, this invention fundamentally improves upon existing technologies by introducing a multi-agent division of labor mechanism, a hypothesis-driven strategy generation method, and a collaborative global search framework. This improves upon existing technologies by shifting from a "single-module static optimization" model to a "multi-role collaborative, self-reasoning, and backtracking and exploration-enabled dynamic iterative optimization" model.

[0072] This invention allows different agents to independently evaluate the same sequence from the perspectives of expression efficiency, structural stability, immunogenicity, and degradation risk. Then, an agent that generates hypotheses about the candidate sequence proposes the next round of optimization directions based on historical feedback. In the multi-agent discussion process, the output is not limited by rigid rules, but rather the optimization suggestions from different dimensions are flexibly considered in conjunction with the overall optimization goal, thereby continuously improving the sequence. This effectively solves the problems of existing technologies, such as the difficulty in coordinating multi-objective conflicts, the tendency to get trapped in local optima during the optimization process, and the lack of adaptive iteration capabilities.

[0073] Figure 1 This is a flowchart illustrating the nucleic acid sequence iterative optimization method based on multi-agent technology provided in Embodiment 1 of the present invention. The method includes the following steps:

[0074] Step S101: Obtain the set steps, obtain the nucleic acid sequence to be optimized, and set multiple targets that need to be co-optimized.

[0075] Preferably, the step of obtaining the setting further includes:

[0076] Obtain the nucleic acid sequence to be optimized.

[0077] Multiple objectives that need to be optimized in a coordinated manner are defined, including stability, expression efficiency, immunogenicity and / or degradation risk.

[0078] Set up a corresponding scoring tool, which is called by multiple intelligent agents, and set constraints.

[0079] Step S102, Single-objective evaluation step, calls multiple single-objective agents to perform multiple single-objective evaluations on the nucleic acid sequence and provides corresponding single-objective optimization strategies.

[0080] Preferably, the single-target intelligent agent further comprises:

[0081] An expression efficiency agent is used to provide a first single-objective optimization strategy to improve expression efficiency.

[0082] A stability agent is used to provide a second single-objective optimization strategy to improve stability.

[0083] An immunogenicity agent is used to provide a third single-objective optimization strategy to enhance immunogenicity.

[0084] A degradation risk agent is used to provide a fourth single-objective optimization strategy to reduce degradation risk.

[0085] Step S103, integration and optimization steps: call the hypothesis generation agent to integrate multiple single-objective optimization strategies into a unified optimization strategy, and then call the sequence generation agent to generate multiple candidate sequences based on the unified optimization strategy and nucleic acid sequences.

[0086] Preferably, the integration and optimization step further includes:

[0087] Hypothesis-generating agents are invoked to integrate multiple single-objective optimization strategies into a unified optimization strategy.

[0088] The sequence generation agent is invoked to generate multiple candidate sequences based on a unified optimization strategy and nucleic acid sequences. The candidate sequences include:

[0089] An expression efficiency improvement sequence that focuses on expression efficiency.

[0090] An improved stability sequence that focuses on stability.

[0091] An immunogenicity-enhanced sequence with a focus on immunogenicity.

[0092] An improved degradation risk sequence that focuses on degradation risk.

[0093] Multiple integrated improvement sequences combining the above directions.

[0094] An improved sequence that combines the above directions and incorporates an exploratory sequence with random mutations.

[0095] The expression efficiency improvement sequence, stability improvement sequence, immunogenicity improvement sequence, degradation risk improvement sequence, comprehensive improvement sequence, and exploration sequence are all generated through codon substitution, local fine-tuning, or random exploration while keeping the amino acid sequence unchanged, and serve as candidate sequences for this round.

[0096] Step S104, comprehensive evaluation step: call multiple single-objective agents to evaluate multiple candidate sequences respectively, obtain multiple single-objective scores for each candidate sequence, and then call the comprehensive evaluation agent to perform comprehensive quantitative evaluation based on the multiple single-objective scores of each candidate sequence to obtain the corresponding comprehensive score.

[0097] Preferably, the comprehensive evaluation step further includes:

[0098] Multiple single-objective agents are invoked to evaluate multiple candidate sequences separately, resulting in multiple single-objective scores for each candidate sequence.

[0099] Then, the comprehensive evaluation agent is invoked to perform a comprehensive quantitative evaluation based on the multiple single-objective scores and their respective weights for each candidate sequence to obtain the corresponding comprehensive score.

[0100] Step S105, iterative optimization step: select the candidate sequence with the highest comprehensive score to enter the next round of iteration, and repeat the above steps until the final optimized sequence is generated.

[0101] Preferably, the iterative optimization step further includes:

[0102] The candidate sequence with the highest overall score is selected to proceed to the next iteration.

[0103] The algorithm iteratively executes the following steps: setting the parameters, evaluating individual parameters, integrating and optimizing the results, evaluating the overall results, and iterating and optimizing the results, until the preset termination conditions are met, and then outputs the final optimized sequence.

[0104] Preferably, the termination condition includes:

[0105] The overall score of the best sequence in a certain round reaches a preset threshold.

[0106] The overall score improvement over multiple consecutive rounds has fallen short of the set increment.

[0107] The maximum number of iterations has been reached.

[0108] Execution terminates when any of the termination conditions are met. Upon termination, the candidate sequence with the highest overall score in the last round is selected as the final optimized sequence.

[0109] This invention is applicable to the sequence design, optimization and evaluation of various types of nucleic acid drugs, such as mRNA, siRNA, and ASO (antisense oligonucleotides), and can be widely used in vaccine development, nucleic acid therapeutic drug research and development, functional nucleic acid construction and artificial intelligence-based nucleic acid sequence design platforms.

[0110] This invention proposes an iterative optimization mechanism for nucleic acid sequences based on "multi-agent hypothesis generation and collaborative exploration." Multiple agents independently evaluate the same sequence and generate improvement suggestions. Then, one agent summarizes the suggestions and proposes several feasible exploration directions. Finally, a holistic evaluation of the target sequence is performed to determine the initial sequence for the next round. This process is repeated cyclically to achieve multiple rounds of adaptive optimization. Through this mechanism, this invention can continuously propose feasible sequence improvement strategies under complex multi-objective constraints, effectively avoiding local optima and achieving continuous iterative optimization of nucleic acid sequences.

[0111] This invention introduces a multi-agent hypothesis-driven and collaborative exploration mechanism, enabling coordinated joint exploration of nucleic acid sequences across multiple optimization objectives, such as stability, expression efficiency, immunogenicity, and degradation risk, eliminating reliance on local fine-tuning in a single direction. This mechanism automatically adjusts the exploration path under conflicting objectives, continuously proposing improved hypotheses and escaping local optima, thereby significantly improving the efficiency of global sequence optimization and obtaining nucleic acid sequences with more balanced performance and superior overall indicators.

[0112] To better understand the technical solution of the present invention, specific embodiments are described below, with 5 candidate sequences fixed for each round.

[0113] For example, consider a sequence of 60 nucleotides in length:

[0114] Initial sequence (60 nt):

[0115] ATG GAG GAG CCG CAG TCA GAT CCT AGC GTC GAT CCC CCC CTG AGT CAG GAAACA TTT TCA

[0116] The present invention provides an iterative optimization method for nucleic acid sequences based on the multi-agent hypothesis and cooperative exploration, which specifically includes the following 5 steps:

[0117] Step 1: Sequence acquisition and multi-target setting.

[0118] The above 60 nt initial sequence was obtained as the target to be optimized. Multiple objectives that need to be optimized in a coordinated manner were set, including stability, expression efficiency, immunogenicity, and degradation risk. Corresponding scoring tools were prepared and each agent was responsible for calling them. Constraints were set (such as keeping the amino acid sequence unchanged, keeping the sequence length unchanged, etc.).

[0119] Step 2: The multi-agent system performs a single-objective evaluation of the current sequence and provides optimization directions.

[0120] For the "baseline sequence" of the current round (the first round uses the initial sequence, and subsequent rounds use the optimal sequence selected in the previous round), multiple single-objective agents are invoked to evaluate the sequence, as follows:

[0121] The expression efficiency agent provides specific directions for "improving expression efficiency," such as: if the calculated CAI is 0.95, the current expression efficiency is already relatively high, and there is still room for reduction; or if the CAI is 0.73, the current expression efficiency is relatively low, and several low-frequency codons are selected for replacement, such as 10-12 bit codon CCG->CCC, 19-21 bit codon GAT-GAC).

[0122] The stability agent provides directions for "reducing overly strong structures or improving overall stability": for example, if the CG frequency in the codon is 0.53, it can be reduced as appropriate.

[0123] The immunogenicity agent suggests a direction of "reducing CpG or high-risk motifs": no CpG sequence risk was found, but a long sequence CCCCCCC exists at positions 34-40, and modification is recommended.

[0124] The degradation risk agent indicates that the 38-bit region is at risk of degradation and suggests correction.

[0125] Each agent provides a clear optimization direction and modification suggestions for its own objective in this round, forming a set of directional information under multiple objectives.

[0126] Step 3: Integrate multi-agent suggestions to generate 5 improved sequences for exploration.

[0127] Suppose that the generating agent integrates the optimization directions of each single-objective agent in step 2 to form a unified optimization strategy for this round (e.g., prioritizing the improvement of expression efficiency without excessively sacrificing stability, while reducing one CpG and weakening AU-rich motifs), and this strategy is handed over to the sequence generating agent for execution.

[0128] In this step, the sequence generation agent generates 10 improved sequences, the specific methods of which include, but are not limited to:

[0129] One improved sequence focusing on expression efficiency; one improved sequence focusing on stability; and one improved sequence focusing on immunogenicity.

[0130] One improved sequence that focuses on degradation risk; five improved sequences that combine the above directions;

[0131] One exploratory sequence that combines the above directions and incorporates certain random mutations (to escape local optima).

[0132] These 10 sequences were generated through codon substitution, local fine-tuning, or random exploration while keeping the amino acid sequence unchanged, and were selected as the 10 candidate sequences for this round.

[0133] Step 4: Perform a comprehensive quantitative evaluation of the 10 candidate sequences.

[0134] For the 10 candidate sequences generated in step 3, each single-target agent is called again for parallel evaluation to obtain quantitative scores for each sequence in terms of stability, expression efficiency, immunogenicity, degradation risk, etc.

[0135] The comprehensive evaluation agent calculates a comprehensive score for these 10 sequences based on the pre-set weights of each objective. This can be done by constructing a weighted total score or a multi-objective comprehensive index. At the same time, the performance of each sequence on each individual objective is recorded for subsequent analysis and comparison.

[0136] Step 5: Select the one with the best overall performance to proceed to the next iteration.

[0137] In step 4, the comprehensive scores of the 10 candidate sequences are compared, and the sequence with the highest comprehensive score is selected as the "optimal sequence" for this round. This sequence is then used as the "benchmark sequence" for the next iteration, and the process returns to step 2 for a new round of multi-agent evaluation and collaborative exploration.

[0138] The iteration terminates when any of the following conditions are met:

[0139] The overall score of the best sequence in a certain round reaches a preset threshold;

[0140] The overall score improvement over multiple consecutive rounds has fallen short of the set increment.

[0141] The predetermined maximum number of iterations has been reached.

[0142] At the end, the sequence with the highest overall score in the last round is output as the final optimization result.

[0143] It should be noted that the nucleic acid sequence iterative optimization method based on multi-agent technology provided in this embodiment of the invention can be executed by an electronic device, apparatus, or a control module within that apparatus for executing the method. This embodiment of the invention uses an apparatus executing the method as an example to illustrate the nucleic acid sequence iterative optimization apparatus based on multi-agent technology provided in this embodiment of the invention.

[0144] Figure 2This is a schematic diagram of the structure of the nucleic acid sequence iterative optimization device based on multi-agent intelligence provided in Embodiment 2 of the present invention. The nucleic acid sequence iterative optimization device 100 based on multi-agent intelligence includes an acquisition and setting module 10, a single evaluation module 20, an integrated optimization module 30, a comprehensive evaluation module 40, and an iterative optimization module 50, wherein:

[0145] The acquisition and setting module 10 is used to acquire the nucleic acid sequence to be optimized and set multiple targets that need to be optimized in a coordinated manner.

[0146] The single-objective evaluation module 20 is used to call multiple single-objective agents to perform multiple single-objective evaluations on the nucleic acid sequence and provide corresponding single-objective optimization strategies.

[0147] The integration and optimization module 30 is used to call the hypothesis generation agent to integrate multiple single-objective optimization strategies into a unified optimization strategy, and then call the sequence generation agent to generate multiple candidate sequences based on the unified optimization strategy and the nucleic acid sequence.

[0148] The comprehensive evaluation module 40 is used to call multiple single-objective agents to evaluate multiple candidate sequences respectively, obtain multiple single-objective scores for each candidate sequence, and then call the comprehensive evaluation agent to perform a comprehensive quantitative evaluation based on the multiple single-objective scores of each candidate sequence to obtain the corresponding comprehensive score.

[0149] The iterative optimization module 50 is used to select the candidate sequence with the highest comprehensive score to enter the next round of iteration until the final optimized sequence is generated.

[0150] Preferably, the acquisition setting module 10 further performs the following:

[0151] (11) Obtain the nucleic acid sequence to be optimized;

[0152] (12) Set multiple objectives that need to be optimized in a coordinated manner, including stability, expression efficiency, immunogenicity and / or degradation risk;

[0153] (13) Set up corresponding scoring tools, which are called by multiple agents, and set constraints.

[0154] Preferably, the single-target intelligent agent further comprises:

[0155] An expression efficiency agent is used to provide a first single-objective optimization strategy to improve expression efficiency.

[0156] A stability agent is used to provide a second single-objective optimization strategy to improve stability.

[0157] An immunogenicity agent is used to provide a third single-objective optimization strategy to enhance immunogenicity.

[0158] A degradation risk agent is used to provide a fourth single-objective optimization strategy to reduce degradation risk.

[0159] Preferably, the integration and optimization module 30 further performs the following:

[0160] (31) Call the hypothesis-generating agent to integrate multiple single-objective optimization strategies into a unified optimization strategy.

[0161] (32) Invoke the sequence generation agent to generate multiple candidate sequences based on the unified optimization strategy and nucleic acid sequences. The candidate sequences include:

[0162] An expression efficiency improvement sequence that focuses on expression efficiency.

[0163] An improved stability sequence that focuses on stability.

[0164] An immunogenicity-enhanced sequence with a focus on immunogenicity.

[0165] An improved degradation risk sequence that focuses on degradation risk.

[0166] Multiple integrated improvement sequences combining the above directions.

[0167] An improved sequence that combines the above directions and incorporates an exploratory sequence with random mutations.

[0168] The expression efficiency improvement sequence, stability improvement sequence, immunogenicity improvement sequence, degradation risk improvement sequence, comprehensive improvement sequence, and exploration sequence were all generated as candidate sequences in this round by codon substitution, local fine-tuning, or random exploration while keeping the amino acid sequence unchanged.

[0169] Preferably, the comprehensive evaluation module 40 further performs the following:

[0170] (41) Call multiple single-target agents to evaluate multiple candidate sequences respectively, and obtain multiple single-target scores for each candidate sequence.

[0171] (42) Then call the comprehensive evaluation agent to perform a comprehensive quantitative evaluation based on the multiple single-objective scores and their respective weights of each candidate sequence to obtain the corresponding comprehensive score.

[0172] Preferably, the iterative optimization module 50 further performs:

[0173] (51) Select the candidate sequence with the highest comprehensive score to enter the next round of iteration.

[0174] (52) Repeatedly execute the acquisition setting step, single evaluation step, integration optimization step, comprehensive evaluation step and iterative optimization step until the preset termination condition is met, and output the final optimization sequence.

[0175] Preferably, the termination condition includes:

[0176] First, the overall score of the best sequence in a certain round reaches a preset threshold.

[0177] Second, the overall score improvement in multiple consecutive rounds was less than the set increment.

[0178] Third, reach the maximum number of iteration rounds.

[0179] Execution terminates when any of the termination conditions are met. Upon termination, the candidate sequence with the highest overall score in the last round is selected as the final optimized sequence.

[0180] The nucleic acid sequence iterative optimization device based on multi-agent provided in this embodiment of the invention can achieve Figure 1 The various processes implemented in the embodiment of the nucleic acid sequence iterative optimization method based on multi-agent shown are not described in detail here to avoid repetition.

[0181] The purpose of this invention is to address the technical shortcomings of existing nucleic acid sequence design methods, such as difficulty in coordinating multiple objectives, lack of feedback-based hypothesis-driven optimization mechanisms, insufficient exploration capabilities, and susceptibility to local optima. This invention proposes a nucleic acid sequence iterative optimization method and system based on multi-agent hypothesis and collaborative exploration. This method achieves intelligent coordination among multiple evaluation directions, including stability, expression efficiency, immunogenicity, and degradation risk. Through multi-agent collaboration and hypothesis-driven generation of candidate sequences, multiple agents interact and discuss to determine optimization suggestions for the current sequence. Finally, through a collaborative exploration strategy, continuous iterative optimization of the nucleic acid sequence is achieved, overcoming the limitations of existing technologies, such as difficulties in multi-objective optimization, limited search space, and lack of iterative reasoning capabilities.

[0182] This invention provides an iterative optimization mechanism for nucleic acid sequences based on "multi-agent hypothesis generation and collaborative exploration." Multiple agents independently evaluate the same sequence and generate improvement suggestions. One agent then summarizes these suggestions and proposes several feasible exploration directions. Finally, a holistic evaluation of the target sequence is performed to determine the initial sequence for the next round. This process is repeated cyclically to achieve multi-round adaptive optimization. Through this mechanism, this invention can continuously propose feasible sequence improvement strategies under complex multi-objective constraints, effectively avoiding local optima and achieving continuous iterative optimization of nucleic acid sequences.

[0183] The present invention also provides a storage medium for storing, for example, Figure 1The computer program for any of the multi-agent-based nucleic acid sequence iterative optimization methods is described above. For example, computer program instructions, when executed by a computer, can invoke or provide the methods and / or technical solutions according to the present invention through the operation of the computer, achieving the same technical effect. To avoid repetition, further details are omitted here. The program instructions for invoking the methods of the present invention may be stored in a fixed or removable storage medium, and / or transmitted via data streams in broadcast or other signal carrying media, and / or stored in the storage medium of a computer device operating according to the program instructions.

[0184] According to one embodiment of the present invention, the present invention also provides such a Figure 3 The illustrated electronic device 400 may optionally include a storage medium 200 for storing a computer program and a processor 300 for executing the computer program. When the computer program is executed by the processor 300, it implements any of the aforementioned multi-agent-based nucleic acid sequence iterative optimization methods, triggering the electronic device 400 to execute methods and / or technical solutions based on the foregoing embodiments, achieving the same technical effect. To avoid repetition, these methods will not be elaborated upon here. It should be noted that the electronic devices in this embodiment include mobile electronic devices and non-mobile electronic devices. For example, mobile electronic devices may be mobile phones, tablets, laptops, handheld computers, in-vehicle electronic devices, wearable devices, super mobile personal computers, netbooks, or personal digital assistants, etc., while non-mobile electronic devices may be servers, network-attached storage (NAS), personal computers (PCs), televisions (TVs), ATMs, or self-service machines, etc. This embodiment does not specifically limit the scope of the invention.

[0185] It should be noted that the present invention can be implemented in software and / or a combination of software and hardware, for example, using an application-specific integrated circuit (ASIC), a general-purpose computer, or any other similar hardware device. In one embodiment, the software program of the present invention can be executed by a processor to implement the steps or functions described above. Similarly, the software program of the present invention (including associated data structures) can be stored in a computer-readable recording medium, such as RAM memory, a magnetic or optical drive, a floppy disk, or similar devices. Furthermore, some steps or functions of the present invention can be implemented in hardware, for example, as circuitry that works with a processor to perform the various steps or functions.

[0186] This invention can be implemented on a computer as a computer-based method, or in dedicated hardware, or a combination of both. Executable code or portions thereof for the method according to the invention can be stored on a computer program product. Examples of computer program products include memory devices, optical storage devices, integrated circuits, servers, online software, etc. Optionally, the computer program product includes non-transitory program code components stored on a computer-readable medium so as to execute the method according to the invention when the program product is executed on a computer.

[0187] In an optional embodiment, the computer program includes computer program code components adapted to perform all the steps of the method according to the invention when the computer program is run on a computer. Optionally, the computer program is embodied on a computer-readable medium.

[0188] It should be noted that, in this document, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes that element. Furthermore, it should be noted that the scope of the methods and apparatuses in the embodiments of the present invention is not limited to performing functions in the order shown or discussed, but may also include performing functions substantially simultaneously or in the reverse order, depending on the functions involved. For example, the described methods may be performed in a different order than described, and various steps may be added, omitted, or combined. Additionally, features described with reference to certain examples may be combined in other examples.

[0189] Of course, the present invention may have other various embodiments. Without departing from the spirit and essence of the present invention, those skilled in the art can make various corresponding changes and modifications according to the present invention, but these corresponding changes and modifications should all fall within the protection scope of the appended claims.

Claims

1. A multi-agent-based iterative optimization method for nucleic acid sequences, characterized in that, include: The steps for obtaining the set parameters are as follows: obtain the nucleic acid sequence to be optimized, and set multiple targets that need to be optimized in a coordinated manner; The single-objective evaluation step involves calling multiple single-objective agents to perform multiple single-objective evaluations on the nucleic acid sequence and providing corresponding single-objective optimization strategies. The integration and optimization steps involve calling a hypothesis generation agent to integrate multiple single-objective optimization strategies into a unified optimization strategy, and then calling a sequence generation agent to generate multiple candidate sequences based on the unified optimization strategy and the nucleic acid sequence. The comprehensive evaluation step involves calling multiple single-target agents to evaluate multiple candidate sequences separately, obtaining multiple single-target scores for each candidate sequence, and then calling a comprehensive evaluation agent to perform a comprehensive quantitative evaluation based on the multiple single-target scores of each candidate sequence to obtain the corresponding comprehensive score. The iterative optimization step involves selecting the candidate sequence with the highest comprehensive score to enter the next round of iteration, and repeating the above steps until the final optimized sequence is generated.

2. The nucleic acid sequence iterative optimization method based on multi-agent systems according to claim 1, characterized in that, The step of obtaining the setting further includes: Obtain the nucleic acid sequence to be optimized; Several objectives that need to be optimized collaboratively are defined, including stability, expression efficiency, immunogenicity and / or degradation risk; Set up a corresponding scoring tool, which is called by multiple intelligent agents, and set constraints.

3. The nucleic acid sequence iterative optimization method based on multi-agent technology according to claim 1, characterized in that, The single-target agent in the single-item evaluation step further includes: An expression efficiency agent is used to provide a first single-objective optimization strategy to improve expression efficiency. A stability agent is used to provide a second single-objective optimization strategy to improve stability. An immunogenicity agent is used to provide a third single-objective optimization strategy to enhance immunogenicity; A degradation risk agent is used to provide a fourth single-objective optimization strategy to reduce degradation risk.

4. The nucleic acid sequence iterative optimization method based on multi-agent systems according to claim 1, characterized in that, The integration and optimization steps further include: The hypothesis-generating agent is invoked to integrate multiple single-objective optimization strategies into a unified optimization strategy; The sequence generation agent is invoked to generate multiple candidate sequences based on the unified optimization strategy and the nucleic acid sequence, the candidate sequences including: An expression efficiency improvement sequence that focuses on expression efficiency; An improved stability sequence with an emphasis on stability. An immunogenicity-enhanced sequence with a focus on immunogenicity; An improved degradation risk sequence that focuses on degradation risk; Multiple integrated improvement sequences combining the above directions; An improved sequence that combines the above approaches and incorporates an exploratory sequence with random mutations; The expression efficiency improvement sequence, stability improvement sequence, immunogenicity improvement sequence, degradation risk improvement sequence, comprehensive improvement sequence, and exploration sequence are all generated through codon substitution, local fine-tuning, or random exploration while keeping the amino acid sequence unchanged, and serve as the candidate sequences for this round.

5. The nucleic acid sequence iterative optimization method based on multi-agent technology according to claim 1, characterized in that, The comprehensive evaluation steps further include: Multiple single-target agents are invoked to evaluate multiple candidate sequences respectively, resulting in multiple single-target scores for each candidate sequence; Then, the comprehensive evaluation agent is invoked to perform a comprehensive quantitative evaluation based on the multiple single-objective scores and their respective weights of each candidate sequence to obtain the corresponding comprehensive score.

6. The nucleic acid sequence iterative optimization method based on multi-agent technology according to claim 1, characterized in that, The iterative optimization step further includes: The candidate sequence with the highest comprehensive score is selected to proceed to the next iteration. The process of obtaining the setting step, the single evaluation step, the integration optimization step, the comprehensive evaluation step, and the iterative optimization step is executed repeatedly until the preset termination condition is met, and the final optimization sequence is output.

7. The nucleic acid sequence iterative optimization method based on multi-agent systems according to claim 1, characterized in that, The termination conditions include: The overall score of the best sequence in a certain round reaches a preset threshold; The overall score improvement in multiple consecutive rounds is less than the set increment; Reaching the maximum number of iterations; When execution is terminated, the candidate sequence with the highest comprehensive score in the last round is taken as the final optimized sequence.

8. A nucleic acid sequence iterative optimization device based on multiple agents constructed according to the method described in any one of claims 1 to 7, characterized in that, The device includes: The acquisition and setting module is used to acquire the nucleic acid sequence to be optimized and set multiple targets that need to be optimized in a coordinated manner; The single-objective evaluation module is used to call multiple single-objective agents to perform multiple single-objective evaluations on the nucleic acid sequence and provide corresponding single-objective optimization strategies. The integration and optimization module is used to call the hypothesis generation agent to integrate multiple single-objective optimization strategies into a unified optimization strategy, and then call the sequence generation agent to generate multiple candidate sequences based on the unified optimization strategy and the nucleic acid sequence. The comprehensive evaluation module is used to call multiple single-target agents to evaluate multiple candidate sequences respectively, obtain multiple single-target scores for each candidate sequence, and then call the comprehensive evaluation agent to perform a comprehensive quantitative evaluation based on the multiple single-target scores of each candidate sequence to obtain the corresponding comprehensive score. The iterative optimization module is used to select the candidate sequence with the highest comprehensive score to enter the next round of iteration until the final optimized sequence is generated.

9. A storage medium, characterized in that, Used to store a computer program for performing the method according to any one of claims 1 to 7.

10. An electronic device comprising a storage medium, a processor, and a computer program stored on the storage medium and executable on the processor, characterized in that, When the processor executes the computer program, it implements the method of any one of claims 1 to 7.