An engineered bacteria intelligent design system and method based on a large language model

CN122762013APending Publication Date: 2026-09-15上海创智学院
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202611200419.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-08-10
Publication Date
2026-09-15

AI Technical Summary

Technical Problem

1)传统基因线路设计通常依靠人工阅读大量文献以理解生物元件的功能特性、调控关系与组装规则,并据此手动构建方案,该过程经验依赖性强、设计周期长且优化效率不高;虽然近年来出现的Cello、SBOLDesigner等计算机辅助工具在一定程度上能够辅助设计,但其依赖结构化输入或预定义逻辑规则,难以直接理解研究者以自然语言形式提出的设计目标,导致设计需求与计算工具之间存在语义鸿沟;

Benefits of technology

[0018]Beneficial effects: This invention solves the problems of scattered and inconsistent standards of multi-source biological element data by constructing a vectorized index library with a unified format. It also directly understands the design requirements put forward by users in natural language by using a semantic parsing and reasoning module, eliminating the semantic gap between design requirements and computational logic in traditional tools. At the same time, by automatically evaluating the compatibility and orthogonality of elements, generating standard gene circuit logic diagrams, and mapping them to engineered bacterial functional templates, it significantly reduces the reliance on human experience and multidisciplinary knowledge in the design process, greatly shortens the design cycle, improves the efficiency of interdisciplinary integration, and realizes the automated and intelligent transformation from high-level design goals to computable element combination and regulation logic.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122762013A_ABST
    Figure CN122762013A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of synthetic biology and artificial intelligence, in particular to an engineered bacterium intelligent design system and method based on a large language model, comprising a multi-source biological element database module, collecting multi-source biological element data and constructing a unified format vector index library; a semantic analysis and reasoning module, analyzing natural language design requirements, retrieving candidate elements and generating an optimal design scheme; a logic circuit generation module, evaluating element compatibility and orthogonality, and automatically generating a standard gene circuit logic diagram; a template mapping and output module, mapping the gene circuit diagram to an engineered bacterium standard function template, and generating a visual report. The present application solves the problem of data dispersion through a unified vector index library, eliminates the semantic gap by using a large language model to understand natural language requirements, reduces manual dependence through automatic evaluation and generation, shortens the design cycle, improves interdisciplinary integration efficiency, and realizes intelligent conversion from design goals to element combination and control logic.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the interdisciplinary field of synthetic biology and artificial intelligence, specifically to an engineered bacterial intelligent design system and method based on a large language model. Background Technology

[0002] With the rapid development of precision medicine and intelligent medical technologies, the application prospects of micro- and nano-medical robots in targeted drug delivery, tumor treatment, and tissue repair are attracting increasing attention. Currently, these micro-robots mainly rely on external magnetic fields, sound fields, or chemical gradients and other physicochemical signals for motion control. Their functions are mostly concentrated on drug loading and mechanical motion execution, while their ability to autonomously perceive and intelligently respond to complex physiological environments remains relatively limited.

[0003] The rise of synthetic biology has provided a new technological path for constructing engineered cells with complex functions. By designing gene regulatory networks, cells can acquire abilities such as signal sensing, logical operations, and environmental responses, thereby triggering target functions under specific conditions.

[0004] However, designing such gene circuits requires integrating a large number of biological elements from diverse sources. Existing databases such as iGEMParts Registry, BioPartsDB, and RDBSB contain a wealth of elements, but their data formats are inconsistent and their annotation standards differ. Researchers often need to manually search and screen across databases, which is inefficient.

[0005] It is evident that the design process for existing engineered bacteria or biorobots still heavily relies on researchers' personal experience. This manifests in the following significant shortcomings: 1) Traditional gene circuit design usually relies on manually reading a large amount of literature to understand the functional characteristics, regulatory relationships and assembly rules of biological elements, and then manually constructing schemes accordingly. This process is highly dependent on experience, has a long design cycle and low optimization efficiency. Although computer-aided tools such as Cello and SBOLDesigner have emerged in recent years, they can assist in design to some extent. However, they rely on structured input or predefined logical rules and have difficulty directly understanding the design goals proposed by researchers in natural language, resulting in a semantic gap between design requirements and computational tools. 2) Existing biological component data is scattered and inconsistent in standards, and there is a lack of a unified platform that can integrate multi-source data and support natural language-driven design. At the same time, the design of engineered bacteria often involves knowledge from multiple disciplines such as synthetic biology, materials science, and control theory. Existing tools are difficult to effectively link component functions with the overall design requirements of robots. Designers need to have a multidisciplinary background, and cross-disciplinary integration is quite difficult.

[0006] The above reasons together make it difficult for existing technologies to achieve efficient, intelligent, and interdisciplinary engineered bacterial or bio-robot design, and there is a lack of a comprehensive platform that can automatically transform high-level design goals into computable component combinations and control logic. Summary of the Invention

[0007] To address the above technical problems, this invention provides a technical solution for an engineered bacterial intelligent design system and method based on a large language model.

[0008] The technical problem solved by this invention can be achieved using the following technical solution: an intelligent design system for engineered bacteria based on a large language model, comprising: a multi-source biological element database module, used to collect biological element data from multiple public databases and preprocess the biological element data to construct a vectorized index library with a unified format; a semantic parsing and reasoning module, connected to the multi-source biological element database module, used to receive and parse the natural language design requirements input by the user, retrieve a set of candidate biological elements from the vectorized index library based on the parsing results, and reason and generate the optimal design scheme; a logic circuit generation module, connected to the multi-source biological element database module and the semantic parsing and reasoning module, used to evaluate the compatibility and orthogonality of the biological element combinations in the optimal design scheme, and automatically generate a gene circuit logic diagram that conforms to synthetic biology standards; and a template mapping and output module, connected to the logic circuit generation module, used to map the gene circuit logic diagram to a preset standard functional template for engineered bacteria and generate a visualization report.

[0009] Preferably, the multi-source biological element database module includes: a data acquisition unit, used to automatically acquire biological element data from at least one publicly available biological element database to obtain multi-source biological element data; wherein the types of biological element data include catalytic elements, regulatory elements, coding sequences, signal transduction elements, memory elements, logic gate elements, and robot function-related elements; a data cleaning and deduplication unit, connected to the data acquisition unit, used to clean the multi-source biological element data and deduplicate it according to the element name, sequence consistency identifier, and database source identifier; a standardized annotation unit, connected to the data cleaning and deduplication unit, used to standardize and classify the deduplicated biological element data and uniformly convert it into a standardized synthetic biology open language format to establish a local biological element database that can be called by large language models; and a vectorized index construction unit, connected to the standardized annotation unit, used to vectorize the element information in the local biological element database to construct the vectorized index library.

[0010] Preferably, the semantic parsing and reasoning module includes: a natural language receiving unit, used to receive natural language design requirements input by a user through a user interface; a requirement parsing unit, connected to the natural language receiving unit, used to perform structured parsing of the natural language design requirements using a trained large language model, generating a structured semantic feature vector; wherein the structured semantic feature vector includes host type, perceptual signal type, signal threshold, logical characteristics, and output substance; and a retrieval enhancement generation unit, connected to the requirement parsing unit, used to perform semantic retrieval in the vectorized index library using the structured semantic feature vector as a query vector to obtain retrieval results; wherein the retrieval results include a set of candidate biological elements and the information of each candidate biological element. The system includes: an associated annotation information, comprising at least one of the following: component name, component type, functional description, control direction, applicable host, response signal type, sequence information, performance parameters, source database, and retrieval relevance score; a model input prompt construction unit, connected to the retrieval enhancement generation unit, for performing hierarchical contextual concatenation of the retrieval results, the structured semantic feature vector, system instructions, and preset output format constraints to generate model input prompts; and a strategy recommendation unit, connected to the model input prompt construction unit, for inputting the model input prompts into the trained large language model, reasoning and outputting the optimal design scheme, wherein the optimal design scheme includes combinations of biological components and logical relationships between the various biological components.

[0011] Preferably, the training method of the large language model is as follows: using a general large language model as the base model, using synthetic biology corpus as the training dataset, and performing domain adaptation training on the base model through a low-rank adaptation fine-tuning method; wherein, the synthetic biology corpus is derived from a publicly available biological element database.

[0012] Preferably, the logic circuit generation module includes: a performance evaluation unit, used to perform compatibility and orthogonality evaluation on candidate element combinations in the optimal design scheme based on domain knowledge; the compatibility evaluation includes restriction enzyme site conflict detection, repetitive sequence and homologous recombination risk assessment, and biological vector matching judgment; the orthogonality evaluation includes protein interaction interference assessment and input signal independence assessment; an element mapping unit, connected to the performance evaluation unit, used to map each biological element in the evaluated biological element combination as a standard biological element node, and map the logical relationship between each biological element as a directed regulatory edge; a loop generation unit, connected to the element mapping unit, used to automatically generate an initial gene loop logic diagram according to synthetic biology design rules and a preset gene loop template; and a constraint optimization unit, connected to the loop generation unit, used to perform constraint checks and optimization on the initial gene loop logic diagram according to preset assembly rules, to obtain a gene loop logic diagram that conforms to synthetic biology standards and a standardized description file, the standardized description file including the regulatory relationship between each biological element, the connection order, and the recommended assembly strategy.

[0013] Preferably, the gene circuit logic diagram includes a cis expression module and a complex logic structure. The cis expression module is constructed in the order of promoter, ribosome binding site, coding sequence and terminator. The complex logic structure is generated according to a predefined regulatory topology.

[0014] Preferably, the constraint optimization unit generates the standardized description file using the Open Language for Synthetic Biology, version 3 (OLLP) standard. Specifically, it creates an Interaction object for each regulatory relationship in the gene circuit logic diagram, whereby the Interaction object defines a stimulus or inhibitory relationship; it creates a Constraint object for the cis-expression module in the gene circuit logic diagram, whereby the Constraint object defines the linear physical order constraint between the promoter, ribosome binding site, coding sequence, and terminator; and it integrates the Interaction object and the Constraint object to obtain the standardized description file.

[0015] Preferably, the template mapping and output module includes: a template storage unit for storing the preset engineered bacterial standard functional template, wherein the engineered bacterial standard functional template includes an input sensing submodule, a logic processing submodule, a memory storage submodule, an output execution submodule, and a drive coupling submodule; and a function mapping unit connected to the template storage unit for allocating each biological element in the gene circuit logic diagram to the corresponding submodule in the engineered bacterial standard functional template according to the element function type, and generating a visualization report; wherein the visualization report includes the sequence information, assembly strategy, verification primer sequence, and experimental operation suggestions for each biological element.

[0016] A method for intelligent design of engineered bacteria based on a large language model is applied to an intelligent design system for engineered bacteria based on a large language model as described above. The method includes: Step S1, receiving natural language design requirements input by a user, and performing structured parsing of the natural language design requirements using a trained large language model to generate structured semantic feature vectors; Step S2, based on the structured semantic feature vectors, retrieving a set of candidate biological elements from the vectorized index library using retrieval-enhanced generation technology, and inferring the optimal design scheme, which includes combinations of biological elements and logical relationships between them; Step S3, mapping the biological elements in the optimal design scheme to standard biological element nodes, and mapping the logical relationships to directed regulatory edges, generating a gene circuit logic graph conforming to synthetic biology standards; Step S4, assigning each biological element in the gene circuit logic graph to a corresponding sub-module of a preset engineered bacteria standard functional template according to its functional type, generating a visualization report, which includes sequence information of each biological element, assembly strategy, verification primer sequences, and experimental operation suggestions.

[0017] Preferably, in step S1, when performing structured parsing of the natural language design requirements using the trained large language model, JSON format is used to constrain the output and extract the host type, perceptual signal type, signal threshold, logical characteristics, and output substance from the natural language design requirements.

[0018] Beneficial effects: This invention solves the problems of scattered and inconsistent standards of multi-source biological element data by constructing a vectorized index library with a unified format. It also directly understands the design requirements put forward by users in natural language by using a semantic parsing and reasoning module, eliminating the semantic gap between design requirements and computational logic in traditional tools. At the same time, by automatically evaluating the compatibility and orthogonality of elements, generating standard gene circuit logic diagrams, and mapping them to engineered bacterial functional templates, it significantly reduces the reliance on human experience and multidisciplinary knowledge in the design process, greatly shortens the design cycle, improves the efficiency of interdisciplinary integration, and realizes the automated and intelligent transformation from high-level design goals to computable element combination and regulation logic. Attached Figure Description

[0019] Figure 1 This is a schematic diagram of the system framework of the present invention; Figure 2 This is a schematic diagram illustrating the system working principle of the present invention; Figure 3 This is a schematic diagram of the multi-source biological element database module of the present invention; Figure 4 This is a schematic diagram of the semantic parsing and reasoning module of the present invention; Figure 5 This is a schematic diagram of the logic circuit generation module of the present invention; Figure 6 This is a schematic diagram of the template mapping and output module of the present invention; Figure 7 This is a schematic diagram of the engineered bacterial standard functional template architecture and functional module division of the present invention; Figure 8 This is a schematic diagram of the method flow of the present invention. Detailed Implementation

[0020] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0021] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other.

[0022] It should also be noted that the engineered bacterial intelligent design, construction and application described in this invention must be carried out in accordance with relevant national laws and regulations, ethical norms and biosafety requirements, and shall only be implemented within the scope of legality and compliance.

[0023] The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but this is not intended to limit the scope of the invention.

[0024] Reference Figure 1 This invention provides an intelligent design system for engineered bacteria based on a large language model, comprising: a multi-source biological element database module 100, used to collect biological element data from multiple public databases and preprocess the biological element data to construct a vectorized index library in a unified format; a semantic parsing and reasoning module 200, connected to the multi-source biological element database module 100, used to receive and parse natural language design requirements input by the user, retrieve a set of candidate biological elements from the vectorized index library based on the parsing results, and reason and generate the optimal design scheme; a logic circuit generation module 300, connected to the multi-source biological element database module 100 and the semantic parsing and reasoning module 200, used to evaluate the compatibility and orthogonality of the biological element combinations in the optimal design scheme, and automatically generate a gene circuit logic diagram that conforms to synthetic biology standards; and a template mapping and output module 400, connected to the logic circuit generation module 300, used to map the gene circuit logic diagram to a preset standard functional template for engineered bacteria and generate a visualization report.

[0025] Specifically, in this embodiment of the invention, to address the problems in the existing engineered bacterial design process, such as the scattered and inconsistent standards of multi-source biological element data, the low efficiency of manual cross-database retrieval, the semantic gap between natural language design goals and computational tools, and the difficulty in integrating interdisciplinary knowledge leading to long design cycles and high dependence on expert experience, a unified format vectorized index library is constructed through a multi-source biological element database module 100. Combined with the semantic parsing and reasoning module 200's application of a large language model, it directly understands the user's natural language requirements and automatically retrieves candidate elements, infers and generates the optimal design scheme, and then performs compatibility and orthogonality evaluation through a logic circuit generation module 300 to output a standard gene circuit logic diagram. Finally, the template mapping and output module 400 maps the logic diagram to the standard functional template of engineered bacteria and generates a visual report. This avoids the inefficient mode of manually consulting literature, repeatedly screening across databases, and relying on predefined logical rules for manual combination optimization in traditional design, eliminates the semantic gap between design requirements and automated tools, and realizes end-to-end intelligent generation from high-level design goals to computable biological element combination and regulation logic, significantly improving the efficiency, repeatability, and interdisciplinary collaborative capabilities of engineered bacterial design.

[0026] As a preferred embodiment of the present invention, refer to Figure 3The multi-source biological element database module 100 includes: a data acquisition unit 110, used to automatically acquire biological element data from at least one publicly available biological element database to obtain multi-source biological element data; wherein the types of biological element data include catalytic elements, regulatory elements, coding sequences, signal transduction elements, memory elements, logic gate elements, and robotic function-related elements; a data cleaning and deduplication unit 120, connected to the data acquisition unit 110, used to clean the multi-source biological element data and perform deduplication based on element name, sequence consistency identifier, and database source identifier; a standardized annotation unit 130, connected to the data cleaning and deduplication unit 120, used to standardize and classify the deduplicated biological element data and uniformly convert it into a standardized synthetic biology open language format to establish a local biological element database that can be called by large language models; and a vectorized index construction unit 140, connected to the standardized annotation unit 130, used to vectorize the element information in the local biological element database to construct the vectorized index library.

[0027] Specifically, due to the inconsistencies in data formats and annotation standards among publicly available biological element databases, directly using raw data can lead to low retrieval efficiency and semantic comprehension bias. Therefore, in this embodiment of the invention, reference is made to... Figure 2 We construct a local biological element database in a unified format and vectorize it, thereby providing an efficient and accurate semantic retrieval foundation for subsequent retrieval enhancement generation.

[0028] Specifically, firstly, in the data acquisition unit 110, raw biological element data is automatically collected from multiple publicly available biological element databases using methods such as web page retrieval or application programming interface (API) calls. Among them, the publicly available biological element databases include iGEM Parts Registry and BioPartsDB, etc. The types of biological elements collected include, but are not limited to, catalytic elements (such as transferases and hydrolases), regulatory elements, coding sequences, signal transduction elements, memory elements, logic gate elements, and robot function-related elements.

[0029] Next, in the data cleaning and deduplication unit 120, the collected multi-source biological element data is cleaned. The specific cleaning process can be as follows: missing values ​​are handled using deletion or KNN interpolation algorithms, i.e., when the missing value ratio is less than 5%, the corresponding row is directly deleted, and when it is 5% to 30%, the mean or median is used to fill the missing value; outliers are identified and handled using the 3σ principle or interquartile range (IQR) method (direct deletion or boundary value replacement), and the date, numerical value, character encoding and other formats are standardized; then, deduplication is performed according to the element name, sequence consistency identifier and database source identifier. Specifically, completely duplicate records are directly deduplicated, and partially duplicate records are retained based on timestamp or data source priority, thereby eliminating redundant data.

[0030] Then, in the standardized annotation unit 130, the deduplicated biological element data is standardized and classified, and uniformly converted into the Synthetic Biology Open Language (SBOL) format to establish a local biological element database that can be called by a large language model.

[0031] Finally, the FAISS vector retrieval library is preferably used to perform vectorized embedding processing on the functional description and sequence feature information of each element in the local biological element database to construct the vectorized index library, which supports fast retrieval based on semantic similarity.

[0032] Thus, through the above construction process, the present invention can achieve unified integration and vectorized representation of multi-source heterogeneous biological element data, providing a standardized data access interface and efficient semantic retrieval capability for the subsequent semantic parsing and reasoning module 200.

[0033] As a preferred embodiment of the present invention, refer to Figure 4The semantic parsing and reasoning module 200 includes: a natural language receiving unit 210, used to receive natural language design requirements input by a user through a user interface; a requirement parsing unit 220, connected to the natural language receiving unit 210, used to perform structured parsing of the natural language design requirements using a trained large language model, generating a structured semantic feature vector; wherein the structured semantic feature vector includes host type, perceptual signal type, signal threshold, logical characteristics, and output substance; and a retrieval enhancement generation unit 230, connected to the requirement parsing unit 220, used to perform semantic retrieval in the vectorized index library using the structured semantic feature vector as a query vector to obtain retrieval results; wherein the retrieval results include a set of candidate biological elements and each candidate biological element. The associated annotation information includes at least one of the following: component name, component type, functional description, regulation direction, applicable host, response signal type, sequence information, performance parameters, source database, and retrieval relevance score; a model input prompt construction unit 240, connected to the retrieval enhancement generation unit 230, is used to perform hierarchical context splicing of the retrieval results, the structured semantic feature vector, system instructions, and preset output format constraints to generate model input prompts; a strategy recommendation unit 250, connected to the model input prompt construction unit 240, is used to input the model input prompts into the trained large language model, infer and output the optimal design scheme, wherein the optimal design scheme includes the combination of biological components and the logical relationship between each biological component.

[0034] Specifically, in order to address the semantic gap between natural language design requirements and computational tools, in this embodiment of the invention, referring to... Figure 2 It adopts a dual-path collaborative architecture that combines domain fine-tuning of a large language model with retrieval-augmented generation (RAG). On the one hand, fine-tuning enables the large language model to acquire knowledge of synthetic biology. On the other hand, RAG retrieves the latest component information in real time, thereby achieving high-precision parsing of users' natural language needs and component combination reasoning, and outputting structured design schemes that can be directly used for logic circuit generation.

[0035] Specifically, firstly, the natural language receiving unit 210 receives the natural language design requirements input by the user in the graphical user interface. These natural language design requirements describe the perception and response output of target environmental signals. The natural language receiving process is described in detail below with a specific embodiment: Example 1 (Natural Language Input): The user's natural language design requirement, entered in the graphical user interface, is: "Design a probiotic robot that can sense reactive oxygen species (ROS) levels in the intestinal environment. When ROS concentration increases, it should initiate the expression of antioxidant proteins, and maintain a low expression state under normal ROS levels. The host bacterium is E. coli Nissle 1917." The specific execution code is shown below: Design a probiotic bio-robot that can sense high ROS levels in the gut and express an antioxidant protein only when oxidative stress occurs. Requirements: Host: E. coli Nissle 1917 Input signal: ROS Logic: ROS-responsive activation Output: antioxidant protein expression Avoid constitutive expression to reduce metabolic burden. Next, in the requirement parsing unit 220, the trained large language model is used to perform semantic parsing on the above natural language design requirements, extract key functional elements, and use JSON formatted output constraints to convert unstructured text into structured semantic feature vectors. The extracted key functional elements include: perception signal type: temperature, light, or specific small molecules, etc.; signal threshold: specific temperature, light intensity, or concentration value, which the user can subsequently specify; memory characteristics: permanent memory (continuous activation after a single stimulus); output substance type: target protein; host type: probiotics.

[0036] The output of the demand parsing unit 220 will be described in detail below with a specific embodiment: Example 2 (Requirements Analysis Output): The key functional elements in the natural language design requirements are extracted by the requirements parsing unit 220, and a structured semantic feature vector is output. The specific execution code is shown below: "host": "E. coli Nissle 1917", "signal": "ROS", "logic":"ROS activation", "output":"antioxidant protein", "constraints":"low metabolic burden". The key functional elements extracted by the requirement analysis unit 220 are: the sensing signal type is ROS; the signal threshold is high concentration; the logical characteristic is conditional activation (started when high concentration, turned off when normal); the output substance type is antioxidant protein; and the host type is probiotic E. coli Nissle 1917.

[0037] Secondly, in the retrieval enhancement generation unit 230, the structured semantic feature vector is used as the query vector to perform semantic similarity retrieval in the vectorized index library, and a set of candidate biological elements related to the functions of "ROS perception" and "antioxidant protein expression" is returned.

[0038] The search results of the search enhancement generation unit 230 are described in detail below using a specific embodiment: Example 3 (RAG search results): Regarding the "ROS Sensing" functionality requirement, the following candidate components were retrieved: Element 1: Transcription factor OxyR Component type: Control element Functional description: An oxidative stress regulatory protein in E. coli that undergoes a conformational change and activates downstream promoters in the presence of ROS. Regulation direction: Activation Suitable host: E. coli Response signal type: ROS (H2O2) Sequence information: NCBI Gene ID: 948577 Performance parameters: Response threshold approximately 10 µM H2O2 Source database: iGEM Parts Registry Search relevance score: 0.96 Component 2: ROS-responsive promoter pOxyS Component type: Control element (promoter) Functional description: A promoter activated by the OxyR protein, driving the expression of downstream genes under oxidative stress. Regulation direction: positive Suitable host: E. coli Response signal type: OxyR combination Sequence information: 5′-TTGACA...TATAAT-3′ Source database: BioPartsDB Search relevance score: 0.94 Then, in the model input prompt building unit 240, the search results (candidate component set), structured semantic feature vector, system instructions and preset JSON format constraints are spliced ​​in a hierarchical context according to the splicing order of "system instructions → user requirements → structured requirement parameters → search context → output format constraints" to generate model input prompts.

[0039] In particular, in addition to using the above-mentioned hierarchical context splicing method to construct model input prompts, other design schemes can be adopted, such as placing the retrieval context before the user's requirements and the structured requirement parameters after the system instructions, or adding a few example outputs to the context to guide the model to generate a format that meets expectations.

[0040] Finally, in the strategy recommendation unit 250, the constructed model input prompts are input into the trained large language model, and the model infers and outputs the optimal design scheme. This design scheme includes recommended combinations of biological elements and the logical relationships between the elements.

[0041] Taking the candidate element in Example 3 as an example, the logical relationship between the elements is explained as follows: When the intracellular ROS level is low, OxyR is in an inactive state, the pOxyS promoter maintains a low expression level, and therefore the expression of downstream antioxidant proteins is low; when the ROS concentration increases, OxyR is oxidized and undergoes a conformational change, thereby activating the pOxyS promoter, initiating the expression of downstream antioxidant proteins, thereby helping the cell to clear ROS.

[0042] Through the aforementioned dual-path collaborative architecture, this invention achieves end-to-end intelligent conversion from natural language requirements to structured design solutions, providing accurate and executable input for the subsequent logic circuit generation module 300.

[0043] As a preferred embodiment of the present invention, the training method of the large language model is as follows: using a general large language model as the base model, using synthetic biology corpus as the training dataset, and performing domain adaptation training on the base model through a low-rank adaptation fine-tuning method; wherein, the synthetic biology corpus is derived from a publicly available biological element database.

[0044] Specifically, to enhance the ability of the large language model to understand synthetic biology terminology, component functional descriptions, and regulatory logic, this embodiment of the invention performs domain-specific fine-tuning on the general large language model. The specific training process is as follows: Based on the general large language model base, a low-rank adaptation (LoRA) fine-tuning method is adopted. While keeping most parameters of the base model fixed, only the query and value weight matrices of the self-attention layer are injected with low-rank decomposition matrices for adaptation training. The rank of LoRA fine-tuning is set to 16, the scaling factor is set to 32, the training epochs are 3, and the learning rate is 2e-4.

[0045] The synthetic biology corpus used for training was derived from publicly available biological component databases, such as iGEM PartsRegistry and BioPartsDB, totaling approximately 2GB in CSV format, and was deduplicated and standardized in terms of attribute format.

[0046] The fine-tuned model significantly improved accuracy in natural language parsing, component recognition, and combinatorial reasoning tasks.

[0047] More specifically, in this embodiment of the invention, the open-source large language model Qwen2.5 (released by AlibabaDAMO Academy) is preferably used as the base model. This model adopts the Transformer decoder architecture, is pre-trained on a massive multilingual corpus, and has powerful natural language understanding and generation capabilities.

[0048] Accordingly, Qwen2.5 has a model size of 7B parameters; the model version is Qwen2.5-7B-Instruct, which optimizes instruction following and dialogue interaction, enabling it to better understand the design requirements described by users in natural language and generate structured design solutions according to preset output formats (such as JSON).

[0049] In terms of technical framework implementation, PyTorch is adopted as the deep learning framework, LoRA training is implemented using the PEFT (Parameter-Efficient Fine-Tuning) library, and FAISS is integrated as a vector retrieval library to support retrieval enhancement generation.

[0050] Through the above training method, the large language model obtained by this invention retains both general semantic understanding ability and professional knowledge in the field of synthetic biology, providing a reliable model foundation for subsequent natural language parsing and intelligent reasoning.

[0051] As a preferred embodiment of the present invention, refer to Figure 5The logic circuit generation module 300 includes: a performance evaluation unit 310, used to perform compatibility and orthogonality evaluations on candidate element combinations in the optimal design scheme based on domain knowledge. The compatibility evaluation includes restriction enzyme site conflict detection, repetitive sequence and homologous recombination risk assessment, and biological vector matching judgment. The orthogonality evaluation includes protein interaction interference assessment and input signal independence assessment. An element mapping unit 320, connected to the performance evaluation unit 310, is used to map each biological element in the evaluated biological element combination as a standard biological element node and map the logical relationships between each biological element as directed regulatory edges. A loop generation unit 330, connected to the element mapping unit 320. The system is used to automatically generate an initial gene circuit logic diagram based on synthetic biology design rules and a preset gene circuit template. A constraint optimization unit 340, connected to the circuit generation unit 330, is used to perform constraint checks and optimizations on the initial gene circuit logic diagram according to preset assembly rules, obtaining a gene circuit logic diagram conforming to synthetic biology standards and a standardized description file. The gene circuit logic diagram includes a cis-expression module and a complex logical structure. The cis-expression module is constructed in the order of promoter, ribosome binding site, coding sequence, and terminator. The complex logical structure is generated based on a predefined regulatory topology. The standardized description file includes the regulatory relationships between various biological elements, the connection order, and recommended assembly strategies.

[0052] Specifically, in this embodiment of the invention, a structured graph generation method is used as the core. Through compatibility and orthogonality evaluation, element standardization mapping, templated circuit generation and assembly rule constraint optimization, a gene circuit logic graph that conforms to synthetic biology standards is automatically output. This transforms the combination of elements and logical relationships into experimentally constructable gene circuits, thereby eliminating the experience bias and error risk in manual design.

[0053] Specifically, firstly, to ensure that the candidate element combinations are physically and functionally assemblable and functional, the performance evaluation unit 310 performs compatibility and orthogonality evaluations on the candidate element combinations in the optimal design scheme. The compatibility evaluation of the candidate element combinations is conducted in the following ways: restriction enzyme conflict detection: checking whether there are identical restriction endonuclease recognition sites in the DNA sequence of the candidate elements; if a conflict exists, it suggests replacing the element or modifying the assembly strategy; repetitive sequence and homologous recombination risk assessment: scanning long repetitive fragments (≥20 bp homologous) in the element sequence to predict regions where homologous recombination may occur and excluding high-risk combinations; biological vector matching judgment: verifying the compatibility of the candidate element's replication origin, resistance marker, and target host bacteria to ensure that the element can stably replicate and express in the selected vector.

[0054] The specific methods for orthogonality evaluation of candidate element combinations are as follows: Protein-protein interaction interference evaluation: using protein-protein interaction databases (such as STRING) or domain alignment to determine whether there are unexpected bindings or signal crosstalk between proteins expressed by different elements; Input signal independence evaluation: verifying whether there is cross-activation or inhibition between regulatory pathways corresponding to multiple input signals (such as different small molecules, temperature, light) to ensure the orthogonality of logic gates.

[0055] Only components that pass the evaluation are allowed to proceed to the next stage.

[0056] Next, through the component mapping unit 320, each biological element in the component combination that has passed the evaluation output by the performance evaluation unit 310 is mapped to a standardized biological element node (such as the ComponentDefinition in the SBOL standard), and the logical relationships between elements (such as "activation" and "inhibition") are mapped to directed regulatory edges to construct the initial directed graph structure.

[0057] Then, the loop generation unit 330 automatically generates an initial gene loop logic diagram according to synthetic biology design rules and a preset gene loop template. Specifically: for cis expression modules, the basic structure of "promoter-ribosome binding site (RBS)-coding sequence (CDS)-terminator" is arranged; for complex logical structures (such as AND gates and OR gates), the abstract logical relationship is mapped to gene regulatory topology (such as by combining promoters, transcription factors and repressors) with reference to the gate-level mapping idea.

[0058] Finally, the constraint optimization unit 340 performs constraint checks and optimizations on the initial gene circuit logic diagram according to preset assembly rules such as Golden Gate, BioBrick, or GibsonAssembly. Specifically, this includes: checking whether the connection order conforms to the assembly rules, such as BioBrick requiring standard prefix or suffix sequences on both sides of the element; verifying the consistency of element orientation, such as the promoter direction must face the CDS; optimizing restriction enzyme site conflicts, adjusting the connection order, or recommending alternative assembly methods.

[0059] Finally, the output includes a gene circuit logic diagram conforming to synthetic biology standards, as well as a standardized description file containing regulatory relationships, connection sequences, and recommended assembly strategies.

[0060] The specific execution code of the logic circuit generation module 300 mentioned above is shown below: "name": "OxyR", "type": "transcription factor", "function": "ROS sensor regulator". "name": "pOxyS", "type": "promoter", "function": "activated by OxyR under oxidative stress". "name": "katG", "type": "enzyme", "function": "catalase-peroxidase". As can be seen from the above code, the logic circuit generation module 300 maps the OxyR transcription factor to an upstream regulatory node, the pOxyS promoter to a control node of the cis-expression module, and the katG coding sequence to an output execution node. Among them, OxyR drives katG expression by activating the pOxyS promoter, forming a complete regulatory chain of "ROS signal sensing → promoter activation → antioxidant enzyme expression", providing a standardized gene circuit description for the subsequent template mapping and output module 400.

[0061] The logic circuit generation process will be described in detail below with a specific example: Example 4 (Logic Circuit Generation): Based on the optimal design scheme output by the semantic parsing and reasoning module 200 in the aforementioned embodiments, the recommended component combination is: OxyR transcription factor, ROS-responsive promoter pOxyS, ribosome binding site RBS, antioxidant protein coding sequence ahpC, and transcription terminator. First, the performance evaluation unit 310 performs a compatibility evaluation on this combination. The test shows that there are no restriction enzyme site conflicts between pOxyS and ahpC, no repeat fragments ≥20 bp in the sequence, and all components are applicable to the E. coli vector pSB1C3, thus the compatibility evaluation is passed. Simultaneously, the orthogonality evaluation shows that OxyR only responds to ROS signals, has no protein-protein interaction with the antioxidant protein expressed by ahpC, and the input signals are independent, thus the evaluation is passed.

[0062] Next, the element mapping unit 320 maps OxyR to a control node, pOxyS, RBS, ahpC, and the terminator to standard element nodes, and maps "ROS→OxyR activation→pOxyS startup→ahpC expression" to a directed control edge.

[0063] Subsequently, the initial gene circuit logic diagram is generated by the circuit generation unit 330 according to the cis-expression module structure. The initial gene circuit logic diagram is: pOxyS—RBS—ahpC—terminator; and OxyR is linked to pOxyS as an upstream regulatory factor.

[0064] Then, by checking the constraint optimization unit 340 according to the Golden Gate assembly rules, it was found that there were no additional restriction site conflicts between pOxyS and RBS, the orientation was consistent, the connection order was reasonable, and no adjustment was needed.

[0065] The final output gene circuit logic diagram clarifies the following regulatory relationships: when ROS concentration increases, OxyR is activated and binds to the pOxyS promoter, driving ahpC to transcribe and express antioxidant proteins; when ROS level is normal, OxyR is in an inactive state, and pOxyS maintains a low expression level.

[0066] As can be seen, through the automatic processing of the logic circuit generation module 300, the present invention can transform natural language design requirements into standardized gene circuits that can be directly used for experimental construction, avoiding sequence conflicts and assembly errors in manual design, and significantly improving design efficiency and success rate.

[0067] In a preferred embodiment of the present invention, the constraint optimization unit generates the standardized description file using the Open Language for Synthetic Biology, version 3. Specifically, it creates an Interaction object for each regulatory relationship in the gene circuit logic diagram, wherein the Interaction object defines a stimulus or inhibition relationship; it creates a Constraint object for the cis-expression module in the gene circuit logic diagram, wherein the Constraint object defines the linear physical arrangement order constraint between the promoter, ribosome binding site, coding sequence, and terminator; and it integrates the Interaction object and the Constraint object to obtain the standardized description file.

[0068] Specifically, considering that the differences in the format of gene circuit descriptions between different synthetic biology design software and experimental platforms can lead to difficulties in data exchange, and that manually writing standardized description files is prone to errors, in this embodiment of the invention, the Synthetic Biology Open Language 3 (SBOL3) standard is used to automatically generate standardized description files, thereby achieving a machine-readable, platform-independent, and standardized expression of gene circuit logic diagrams.

[0069] Specifically, first, for each regulatory relationship in the gene circuit logic diagram (e.g., "OxyR activates pOxyS"), an Interaction object is created; the Interaction object contains the following attributes: a list of participants, listing the source element (e.g., OxyR) and target element (e.g., pOxyS) involved in the regulation; a regulation type, set to "stimulation" or "inhibition"; and optional regulation parameters (e.g., response threshold, cooperation coefficient).

[0070] Next, for each cis-expression module (e.g., "pOxyS—RBS—ahpC—terminator") in the gene circuit logic diagram, a Constraint object is created; this Constraint object uses the SBOL3 predefined constraint type "SBOL_PRECEDES" to explicitly constrain the promoter to precede the ribosome binding site, the ribosome binding site to precede the coding sequence, and the coding sequence to precede the terminator, thereby defining the linear physical arrangement order constraint.

[0071] Then, all the created Interaction and Constraint objects are encapsulated together with the ComponentDefinition into an SBOL3 document. This SBOL3 document contains the complete gene circuit topology, regulatory logic, and physical assembly constraints, and is output as a standardized description file in XML or JSON format.

[0072] Through the automated SBOL3 standard description file generation process described above, this invention can effectively ensure the standardization, interchangeability, and reproducibility of gene circuit design results, making it easy for downstream researchers to directly import them into other synthetic biology design or simulation tools for further analysis.

[0073] As a preferred embodiment of the present invention, refer to Figure 6 The template mapping and output module 400 includes: a template storage unit 410 for storing the preset engineered bacterial standard functional template, the engineered bacterial standard functional template including an input sensing submodule 510, a logic processing submodule 520, a memory storage submodule 530, an output execution submodule 540, and a drive coupling submodule 550; and a function mapping unit 420 connected to the template storage unit 410 for allocating each biological element in the gene circuit logic diagram to the corresponding submodule in the engineered bacterial standard functional template according to the element function type, and generating a visualization report; wherein, the visualization report includes the sequence information, assembly strategy, verification primer sequence, and experimental operation suggestions of each biological element.

[0074] Specifically, in order to transform the abstract gene circuit logic diagram into a modular design scheme that conforms to the functional architecture of a robot, making it easier for experimental personnel to understand and use, in this embodiment of the invention, a template mapping method is adopted to automatically allocate each biological element in the standardized gene circuit output by the logic circuit generation module 300 to the corresponding sub-module of the predefined engineered bacterial standard functional template according to its biological function.

[0075] Specifically, firstly, refer to Figure 7 The diagram shown illustrates the architecture and functional module division of the engineered bacteria standard functional template. The predefined engineered bacteria standard functional template in template storage unit 410 includes the following sub-modules: The input sensing submodule 510 includes small molecule responsive signaling pathway elements, light / heat / magnetic field responsive elements, and guiding ribonucleic acid (gRNA) elements, which are used to sense external environmental signals (such as small molecule concentration, temperature changes, light intensity, magnetic field direction, or specific nucleic acid sequences) and convert them into intracellularly processable biochemical signals.

[0076] The logic processing submodule 520, connected to the input sensing submodule 510, includes transcription factors, clustered regular-interval short palindromic repeat (CRISPR) related elements, promoters, and replication origin elements. It is used to perform logical operations (such as AND gates, OR gates, and NOT gates) on the signals transmitted by the input sensing submodule 510 and generate corresponding regulatory instructions.

[0077] The memory storage submodule 530 and the connection logic processing submodule 520 include recombinant elements such as recombinases or recombinant sites, which are used to flip or lock the gene expression state when a specific event occurs, so as to realize long-term storage of information or permanent function switching.

[0078] The output execution submodule 540 is connected to the memory storage submodule 530 and includes coding sequence (CDS), terminator, ribosome regulator, degradation tag, separator intron, ribosome binding region (RBS) and insulator region elements, which are used to perform specific biological functions according to the instructions of the logic processing submodule 520, such as expressing effector proteins, synthesizing metabolites or modifying cell structures.

[0079] The drive coupling submodule 550 is connected to the output execution submodule 540 and includes motion module-related elements (such as chemokines, flagella assembly factors, etc.) to couple the signals generated by the output execution submodule 540 to the bacterial motion behavior (such as chemotaxis, obstacle avoidance, or directional migration) to realize the active motion control of engineered bacteria.

[0080] Next, the function mapping unit 420 automatically assigns the element to the corresponding sub-module of the aforementioned standard function template based on the functional annotation of each biological element in the gene circuit logic diagram output by the logic circuit generation module 300. For example, in the ROS-responsive probiotic design of the aforementioned embodiment, the gene circuit includes the OxyR transcription factor, pOxyS promoter, RBS, ahpC coding sequence, and terminator. The function mapping unit 420 assigns OxyR and pOxyS to the input sensing sub-module 510 (because they work together to sense and convert ROS signals), and assigns ahpC to the output execution sub-module 540 (because it expresses an antioxidant protein and performs a protective function), while the logic processing sub-module 520, memory storage sub-module 530, and drive coupling sub-module 550 do not need to be assigned elements in this embodiment and remain empty.

[0081] Then, the function mapping unit 420 integrates all the assigned sub-modules and their component information to generate a visual report.

[0082] This visualization report includes: standardized nucleotide sequences of each element (e.g., OxyR, pOxyS, RBS, ahpC, Terminator), a schematic diagram of the gene circuit structure (marking the regulatory relationships and physical connection order between elements), recommended assembly strategies (e.g., using the Golden Gate assembly method, with specific restriction enzyme sites and connection order provided), verification primer sequences (automatically designed upstream and downstream primers for PCR verification of correct gene circuit assembly), and experimental operation suggestions (e.g., hydrogen peroxide induction conditions, protein expression detection methods, etc.).

[0083] In this way, through template mapping and output module 400, the present invention transforms the complex gene circuit logic diagram into a modular construction scheme that can be directly used by experimenters, significantly reducing the implementation threshold of engineered bacterial design.

[0084] In addition, in this embodiment, the system recommends using the Golden Gate assembly method to construct expression plasmids and automatically generates PCR primer sequences for verifying the correct assembly of gene circuits. Simultaneously, the system provides suggestions for functional validation experiments, such as: adding hydrogen peroxide to a final concentration of 50 µM to the culture system to induce oxidative stress, collecting bacterial cells after 2 hours of culture, and detecting the expression level of antioxidant proteins using Western blotting or fluorescent reporter protein assays to verify the system's functional response under ROS signal triggering.

[0085] Reference Figure 8This invention also provides an intelligent design method for engineered bacteria based on a large language model, applied to an intelligent design system for engineered bacteria based on a large language model as described above, comprising: Step S1, receiving natural language design requirements input by a user, and performing structured parsing of the natural language design requirements through a trained large language model to generate structured semantic feature vectors; Step S2, based on the structured semantic feature vectors, retrieving a set of candidate biological elements in the vectorized index library through retrieval enhancement generation technology, and inferring the optimal design scheme, wherein the optimal design scheme includes combinations of biological elements and logical relationships between various biological elements; Step S3, mapping the biological elements in the optimal design scheme to standard biological element nodes, mapping the logical relationships to directed regulatory edges, and generating a gene circuit logic graph conforming to synthetic biology standards; Step S4, assigning each biological element in the gene circuit logic graph to the corresponding sub-module of a preset engineered bacteria standard functional template according to the element functional type, and generating a visualization report, wherein the visualization report includes the sequence information of each biological element, assembly strategy, verification primer sequence, and experimental operation suggestions.

[0086] In step S1, when the natural language design requirements are structured and parsed using the trained large language model, JSON format is used to constrain the output and extract the host type, perceptual signal type, signal threshold, logical characteristics, and output substance from the natural language design requirements.

[0087] Specifically, in this embodiment of the invention, through the coordinated execution of steps S1 to S4, the user's colloquial requirements are first transformed into structured semantic feature vectors using a large language model, eliminating the semantic gap; then, combined with retrieval enhancement generation technology, the optimal combination of components and logical relationships are intelligently retrieved from a standardized component library; next, a gene circuit logic diagram conforming to the SBOL3 standard is automatically generated according to synthetic biology design rules; finally, each functional component is mapped to a predefined robot standard template, outputting a visual report containing sequences, assembly strategies, verification primers, and experimental suggestions. This method can automate the entire process from natural language design requirements to experimentally constructable engineered bacterial robot solutions, significantly reducing the dependence of engineered bacterial design on expert experience, avoiding the inefficiency and errors caused by manual cross-library retrieval and manual design of gene circuits, and achieving high efficiency, standardization, and reproducibility in the design process.

[0088] In summary, this invention proposes an engineered bacterial intelligent design system and method based on a large language model. By constructing a unified format multi-source biological element database and vectorized index library, and combining a domain-fine-tuned large language model with retrieval enhancement generation technology, it achieves fully automated end-to-end generation from natural language design requirements to structured element combinations, gene circuit logic diagrams, and visualized construction scheme reports.

[0089] Compared with the prior art, the present invention has the following significant advantages: Natural language-driven approach eliminates the semantic gap: Directly understands the design goals described by researchers in natural language (such as "building probiotics with ROS response"), without the need for manual conversion into structured parameters or predefined logical rules, significantly reducing the design threshold.

[0090] Unified integration of multi-source data improves retrieval efficiency: Bio-element data scattered across multiple public databases (iGEM, BioPartsDB, RDBSB, etc.) are cleaned, deduplicated, and standardized into SBOL format, and a vectorized index library is built to support fast retrieval based on semantic similarity, avoiding the inefficiency of manual cross-database screening.

[0091] Intelligent component recommendation and circuit generation: Utilizing a LoRA-tuned synthetic biology big language model and RAG technology, the system automatically infers the optimal component combination and regulatory logic; it generates gene circuit logic diagrams through compatibility / orthogonality evaluation, cis-expression, and gate-level mapping templates, and performs constraint optimization by combining assembly rules to reduce human design errors.

[0092] Templated output reduces the difficulty of experimental implementation: Gene circuits are automatically mapped to predefined engineered bacterial standard functional templates (input sensing, logic processing, memory storage, output execution, and driving coupling), generating a visual report containing sequence information, assembly strategies, verification primers, and experimental suggestions, which researchers can directly use to conduct subsequent experiments.

[0093] Significantly improves design efficiency and success rate: It achieves minute-level intelligent generation from "high-level design goals" to "experimentable construction solutions", avoiding the long cycle mode of traditional design that relies on expert experience, repeated literature review and manual combination optimization. At the same time, it improves the first-time success rate of design solutions through constraint checks and orthogonality evaluation.

[0094] The above description is merely a preferred embodiment of the present invention and does not limit the implementation and protection scope of the present invention. Those skilled in the art should realize that any equivalent substitutions and obvious changes made based on the description and illustrations of the present invention should be included within the protection scope of the present invention.

Claims

1. An engineered intelligent bacterial design system based on a large language model, characterized in that, include: A multi-source biological element database module is used to collect biological element data from multiple public databases and perform preprocessing operations on the biological element data to construct a vectorized index library with a unified format. The semantic parsing and reasoning module is connected to the multi-source biological element database module. It is used to receive and parse the natural language design requirements input by the user, retrieve the candidate biological element set from the vectorized index library based on the parsing results, and reason and generate the optimal design scheme. The logic circuit generation module, connected to the multi-source biological element database module and the semantic parsing and reasoning module, is used to evaluate the compatibility and orthogonality of the combination of biological elements in the optimal design scheme and automatically generate a gene circuit logic diagram that conforms to synthetic biology standards. The template mapping and output module, connected to the logic circuit generation module, is used to map the gene circuit logic diagram to a preset engineered bacterial standard functional template and generate a visualization report.

2. The engineered bacterial intelligent design system based on a large language model according to claim 1, characterized in that, The multi-source biological element database module includes: a data acquisition unit, used to automatically acquire biological element data from at least one publicly available biological element database to obtain multi-source biological element data; wherein the types of biological element data include catalytic elements, regulatory elements, coding sequences, signal transduction elements, memory elements, logic gate elements, and robotic function-related elements; a data cleaning and deduplication unit, connected to the data acquisition unit, used to clean the multi-source biological element data and deduplicate it according to element name, sequence consistency identifier, and database source identifier; a standardized annotation unit, connected to the data cleaning and deduplication unit, used to standardize and classify the deduplicated biological element data and uniformly convert it into a standardized synthetic biology open language format to establish a local biological element database that can be called by large language models; and a vectorized index construction unit, connected to the standardized annotation unit, used to vectorize the element information in the local biological element database to construct the vectorized index library.

3. The engineered bacterial intelligent design system based on a large language model according to claim 1, characterized in that, The semantic parsing and reasoning module includes: a natural language receiving unit, used to receive natural language design requirements input by the user through a user interface; a requirement parsing unit, connected to the natural language receiving unit, used to perform structured parsing of the natural language design requirements using a trained large language model, generating a structured semantic feature vector; wherein the structured semantic feature vector includes host type, perceptual signal type, signal threshold, logical characteristics, and output substance; and a retrieval enhancement generation unit, connected to the requirement parsing unit, used to perform semantic retrieval in the vectorized index library using the structured semantic feature vector as a query vector to obtain retrieval results; wherein the retrieval results include a set of candidate biological elements and the association of each candidate biological element. The annotation information includes at least one of the following: component name, component type, functional description, regulation direction, applicable host, response signal type, sequence information, performance parameters, source database, and retrieval relevance score; a model input prompt construction unit, connected to the retrieval enhancement generation unit, is used to perform hierarchical contextual concatenation of the retrieval results, the structured semantic feature vector, system instructions, and preset output format constraints to generate model input prompts; a strategy recommendation unit, connected to the model input prompt construction unit, is used to input the model input prompts into the trained large language model, infer and output the optimal design scheme, wherein the optimal design scheme includes the combination of biological components and the logical relationships between the various biological components.

4. The engineered bacterial intelligent design system based on a large language model according to claim 3, characterized in that, The training method of the large language model is as follows: a general large language model is used as the base model, and synthetic biology corpus is used as the training dataset. The base model is trained for domain adaptation using a low-rank adaptation fine-tuning method. The synthetic biology corpus is derived from a publicly available biological element database.

5. The engineered bacterial intelligent design system based on a large language model according to claim 1, characterized in that, The logic circuit generation module includes: a performance evaluation unit, used to perform compatibility and orthogonality evaluations on candidate element combinations in the optimal design scheme based on domain knowledge. The compatibility evaluation includes restriction enzyme site conflict detection, repetitive sequence and homologous recombination risk assessment, and biological vector matching judgment. The orthogonality evaluation includes protein interaction interference assessment and input signal independence assessment. An element mapping unit, connected to the performance evaluation unit, is used to map each biological element in the evaluated biological element combination as a standard biological element node and map the logical relationships between each biological element as directed regulatory edges. A loop generation unit, connected to the element mapping unit, is used to automatically generate an initial gene loop logic diagram according to synthetic biology design rules and a preset gene loop template. A constraint optimization unit, connected to the loop generation unit, is used to perform constraint checks and optimizations on the initial gene loop logic diagram according to preset assembly rules to obtain a gene loop logic diagram that conforms to synthetic biology standards and a standardized description file. The standardized description file includes the regulatory relationships between each biological element, the connection order, and recommended assembly strategies.

6. The engineered bacterial intelligent design system based on a large language model according to claim 5, characterized in that, The gene circuit logic diagram includes a cis expression module and a complex logic structure. The cis expression module is constructed in the order of promoter, ribosome binding site, coding sequence and terminator. The complex logic structure is generated according to a predefined regulatory topology.

7. The engineered bacterial intelligent design system based on a large language model according to claim 5, characterized in that, The constraint optimization unit generates the standardized description file using the Open Language for Synthetic Biology, version 3 (OLLP) standard. Specifically, it creates an Interaction object for each regulatory relationship in the gene circuit logic diagram, whereby the Interaction object defines a stimulus or inhibitory relationship; it creates a Constraint object for the cis-expression module in the gene circuit logic diagram, whereby the Constraint object defines the linear physical order constraint between the promoter, ribosome binding site, coding sequence, and terminator; and it integrates the Interaction object and the Constraint object to obtain the standardized description file.

8. The engineered bacterial intelligent design system based on a large language model according to claim 1, characterized in that, The template mapping and output module includes: a template storage unit for storing the preset engineered bacterial standard functional template, wherein the engineered bacterial standard functional template includes an input sensing submodule, a logic processing submodule, a memory storage submodule, an output execution submodule, and a drive coupling submodule; and a function mapping unit connected to the template storage unit for allocating each biological element in the gene circuit logic diagram to the corresponding submodule in the engineered bacterial standard functional template according to the element function type, and generating a visualization report; wherein the visualization report includes the sequence information, assembly strategy, verification primer sequence, and experimental operation suggestions for each biological element.

9. A method for intelligent design of engineered bacteria based on a large language model, characterized in that, An intelligent design system for engineered bacteria based on a large language model, as described in any one of claims 1-8, comprises: Step S1, receiving a natural language design requirement input by a user, and performing structured parsing of the natural language design requirement using a trained large language model to generate a structured semantic feature vector; Step S2, based on the structured semantic feature vector, retrieving a set of candidate biological elements in the vectorized index library using retrieval enhancement generation technology, and inferring the optimal design scheme, wherein the optimal design scheme includes combinations of biological elements and logical relationships between various biological elements; Step S3, mapping the biological elements in the optimal design scheme to standard biological element nodes, mapping the logical relationships to directed regulatory edges, and generating a gene circuit logic graph conforming to synthetic biology standards; Step S4, assigning each biological element in the gene circuit logic graph to a corresponding sub-module of a preset engineered bacteria standard functional template according to the element functional type, and generating a visualization report, wherein the visualization report includes sequence information, assembly strategy, verification primer sequence, and experimental operation suggestions for each biological element.

10. The engineered bacterial intelligent design method based on a large language model according to claim 9, characterized in that, In step S1, when the natural language design requirements are structured and parsed using the trained large language model, JSON format is used to constrain the output and extract the host type, perceptual signal type, signal threshold, logical characteristics, and output substance from the natural language design requirements.