Biochemical experiment automation script training generation method based on RAG and DPO

By using locally deployed RAG and DPO technologies, combined with hybrid searchers and simulation verification, we can optimize the generation of biochemical experiment scripts, solve the problems of data security and high costs, and improve generation efficiency and reliability.

CN120633819APending Publication Date: 2025-09-12SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510516552.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-23
Publication Date
2025-09-12

AI Technical Summary

Technical Problem

Existing automated script generation technologies for biochemical experiments have data security risks, high operating costs, lack of consistency and end-to-end optimization capabilities, and low generation pass rates, especially in resource-constrained laboratories.

Method used

A locally deployed large-scale language model is combined with retrieval-augmented generation (RAG) and direct preference optimization (DPO). The external knowledge base is retrieved through a hybrid retriever, and simulation verification and iterative optimization are performed to generate biochemical experiment scripts.

Benefits of technology

It significantly improves the script generation pass rate of small-scale large language models, ensures data security, reduces operating costs, and builds end-to-end optimization capabilities to improve the efficiency and reliability of biochemical experiments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120633819A_ABST
    Figure CN120633819A_ABST
Patent Text Reader

Abstract

The invention discloses a biochemical experiment automation script training generation method based on RAG and DPO. The method comprises the steps that biochemical experiment technical documents are collected to serve as an external knowledge base; the BM25 and the Fiss are fused to construct a hybrid retriever; generating experimental process description by using a large language model, determining equipment, protocols and materials, and retrieving related documents from a knowledge base on the basis of the experimental process description, the protocols and the materials; generating an experiment script based on a retrieval result, and performing platform simulation verification, and performing iterative optimization on the script which fails in verification; marking successful and failed scripts as preference data pairs, and constructing a training set; carrying out LoRA fine tuning on a local large language model by adopting direct preference optimization; and according to a target and appliance prompt input by a user, generating a verified experiment script by using the large language model after direct preference optimization training. According to the method, the script generation automation level and the passing rate are remarkably improved, the manual intervention cost is reduced, and an efficient solution is provided for automation of biochemical experiments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and in particular to a method for generating automated script training for biochemical experiments based on RAG and DPO. Background Art

[0002] Biochemical laboratory automation technology is an important tool for improving the efficiency and reproducibility of biochemical experiments. Researchers using laboratory automation equipment can accelerate scientific research progress. However, operating the computer programs of the robots requires advanced programming skills, which is a high barrier to entry for ordinary researchers.

[0003] Over the past few years, the rapid development of large language models (LLMs) such as GPT-4 has provided new solutions to this problem. Their powerful reasoning capabilities have reduced the difficulty for humans to learn and master laboratory automation technologies. LLMs can generate script programming through simple prompts combined with inherent knowledge, which reduces the learning cost for biological researchers. However, large language models typically use autoregressive methods to generate text. This word-by-word generation method relies on probability distributions and is prone to hallucinations, which may lead to inaccurate or irrelevant answers. In addition, large language models are pre-trained on data from general neighborhoods. When faced with specific professional neighborhoods, large language models may lack the necessary expertise.

[0004] To alleviate the hallucination problem of large language models, Retrieval-Augmented Generation (RAG) technology has emerged. Retrieval-Augmented Generation combines information retrieval and generative models. By retrieving relevant context from external knowledge bases and injecting it into the language model input, it improves the accuracy and factual basis of generated content. RAG effectively reduces the hallucination problem of large language models and is particularly suitable for tasks that require specialized knowledge or up-to-date information. However, the information retrieved by RAG may contain noise (such as redundant or irrelevant details), which can affect the quality of the generated results.

[0005] To further optimize the output quality of LLM, Direct Preference Optimization (DPO) was proposed as a language model optimization method based on human preference optimization. It directly adjusts model parameters through preference data, making it more inclined to generate high-quality output without relying on explicit reward modeling.

[0006] In this context, Inagaki et al. (Takashi Inagaki, Akari Kato, Koichi Takahashi, Haruka Ozaki, Genki N. Kanda, "LLMs can generate robotic scripts from goal-oriented instructions in biological laboratory automation") proposed a large language model-based technology in 2023, using GPT-4 to generate Python scripts for the Opentrons OT-2 liquid handling robot from natural language instructions through the OpenAI API. However, this work relies on the API provided by GPT-4 and cannot fully guarantee data security. At the same time, the cost of calling the API is high, and the cost of actual application is not cheap. In addition, it only relies on the built-in knowledge of GPT-4 and does not use external knowledge bases or retrieval technologies to optimize the context, resulting in the script lacking the accuracy of specific experiments.

[0007] The existing biochemical experiment automation script generation technology has four main problems in its application:

[0008] (1) Relying on the external API interface of large language models for script generation may lead to the risk of experimental data leakage and fail to fully guarantee data security; (2) Using paid large language model services such as GPT-4 has high call and operation costs; (3) Existing biochemical experimental methods are mostly scattered functional modules, failing to form a systematic process framework from input to verification, resulting in a lack of coherence and end-to-end optimization capabilities in the script generation process, affecting the efficiency and reliability of biochemical experiments. (4) In resource-constrained laboratories, the pass rate of biochemical experiment scripts generated by smaller-scale (referring to parameter scales below 8B) large language models is low. Summary of the Invention

[0009] The purpose of the present invention is to overcome the shortcomings and deficiencies of the existing technology and provide a method for training and generating automated biochemical experiment scripts based on RAG and DPO to solve the problems that large-scale language models with smaller parameter sizes have a low pass rate in generating biochemical experiment scripts, calling external large-scale language model APIs cannot fully guarantee data security and has high operating costs, as well as lack of consistency and end-to-end optimization capabilities in the script generation process.

[0010] The present invention is achieved through at least one of the following technical solutions.

[0011] The method for generating automated scripts for biochemical experiments based on RAG and DPO includes the following steps:

[0012] The user inputs the biochemical experiment's operational objectives and experimental equipment as prompts into a local large-scale language model trained with Direct Preference Optimization (DPO). Combined with the hybrid retriever's retrieval context, the experimental process is iteratively generated and simulated for verification, ultimately resulting in a script that is successfully executed.

[0013] Training a local large language model consists of the following steps:

[0014] S1. Collect biochemical experimental technical documents and use the text blocks obtained after preprocessing the biochemical experimental technical documents as an external knowledge base;

[0015] S2, integrates the BM25 algorithm and the Faiss library to build a hybrid search engine, which is used to extract the content and metadata of text blocks and output hybrid search results;

[0016] S3. Generate an experimental process using a large language model, and use a hybrid retriever to retrieve relevant documents from the external knowledge base according to the experimental process description generated by the large language model;

[0017] S4. Generate experimental scripts based on the hybrid search results of the retriever, verify them through platform simulation, and iteratively optimize the scripts that fail the verification;

[0018] S5. Mark the successful and failed scripts as preferred data pairs and construct a training dataset;

[0019] S6. Fine-tune LoRA on a local large language model using training datasets and direct preference optimization.

[0020] Furthermore, the preprocessing of step S1 specifically includes:

[0021] Remove interference information from biochemical experimental technical documents and unify their formats;

[0022] Use the olmOCR tool to convert technical documents in PDF format into Markdown format;

[0023] Create a text splitter, set the size of each text block to N characters, and overlap n characters between adjacent blocks;

[0024] Use the text splitter to store Markdown documents in chunks.

[0025] Furthermore, step S2 specifically includes:

[0026] Extract the content and metadata of each text block and form a text list and a metadata list respectively;

[0027] Building a first search engine based on the text list, wherein the first search engine uses the BM25 algorithm to sort and return the top k relevant documents according to the keyword relevance scores of the query;

[0028] The content in the text list is converted into a semantic vector representation using the Conan-embedding-v1 embedding model. Based on the semantic vector, a second retriever is constructed using the Faiss library's L2 distance-based exact nearest neighbor search retrieval method to retrieve the top k relevant documents.

[0029] Assigning a first weight to the relevant documents returned by the first search engine, and assigning a second weight to the relevant documents returned by the second search engine;

[0030] For the same document that appears in the results of two search engines, the final score is obtained by adding the weighted score of the first search engine and the weighted score of the second search engine;

[0031] For documents that only appear in a single retriever's results, their weighted scores are used directly as the final scores;

[0032] All relevant documents are sorted according to the final scores, and top k documents with the highest scores are returned as hybrid retrieval results.

[0033] Furthermore, in step S3, the large language model searches for a prompt template according to an external knowledge base to generate an experimental process description corresponding to the operation target;

[0034] The external knowledge base search prompt template is a structured natural language instruction template, including the experimental equipment and materials, laboratory instruments and reagents, experimental conditions and parameters required for the operation target, the steps required for the experimental protocol of the operation target, and the corresponding Opentrons API function.

[0035] Further, experimental equipment and materials include laboratory appliances and reagents;

[0036] The experimental conditions and parameters include volume, time and number of samples.

[0037] Furthermore, in step S4, the large language model is based on the retrieval results of the hybrid retriever and the experimental script generated by the experimental process, which is simulated and verified by the opentrons_simulate tool of the Opentrons software platform. The scripts that fail the verification are iteratively optimized. Step S4 is specifically as follows:

[0038] Filling the operation target input by the user and the relevant documents retrieved from the external knowledge base using the hybrid retriever into the placeholder of the experiment script generation prompt template;

[0039] Use the populated experimental script to generate a prompt template, construct a conversation history message, input the conversation history message into the large language model, and extract the Python experimental script from the reply message in the large language model language by parsing the Markdown code block tags;

[0040] Use the opentrons_simulate tool of the Opentrons software platform to simulate and verify the extracted Python experimental script to generate verification results;

[0041] If the verification result indicates that the script execution fails, the error message generated during the verification process is extracted and formatted as a prompt instruction, and the prompt instruction is added to the end of the conversation history message between the user and the large language model for update;

[0042] Based on the updated conversation history messages, the large language model combines the operation goals in the initial prompt instructions and the mixed retrieval results to analyze the error information, infer the cause of the problem, and generate a new experimental script. If the new experimental script still fails the simulation verification, the error information generated during the verification process is extracted again and added to the end of the conversation history messages between the user and the large language model. The next round of experimental script generation continues until the simulation verification passes or the number of simulation verifications reaches the upper limit.

[0043] Furthermore, the experiment script generation prompt template is a structured natural language instruction template, which incorporates the operation purpose input by the user into the experiment script generation prompt template and incorporates the search document as context into the experiment script generation prompt template, providing a guidance framework for generating Python experiment scripts for large language models; the experiment script generation prompt template includes:

[0044] Role of the model: Biological, computer and engineering experts;

[0045] Placeholders are dynamically filled with the operation purpose entered by the user and the retrieved documents when the program is running.

[0046] Furthermore, in step S5, the successful scripts and the failed scripts are marked as preferred data pairs to construct a training data set, including:

[0047] Obtain successful and failed scripts related to laboratory experiments. Successful scripts are Python experimental scripts that have passed platform verification, and failed scripts are Python experimental scripts that have not passed platform verification.

[0048] Based on the success script and the rejection script, construct the preference data set D prefer , where each preference data pair includes the experimental task description, success script y w and reject script y l。

[0049] Further, in step S6, a low-rank matrix A and a low-rank matrix B are added in parallel beside the linear module of the attention layer of the large language model;

[0050] During training, only the low-rank matrix A and the low-rank matrix B are updated. The formula is expressed as: W0 + ΔW = W0 + BA, W0 ∈ R d×k , B ∈ R d×r , A ∈ R r×k , where ΔW represents the incremental parameter, d and k represent the input dimension and output dimension of the original attention layer in the large language model, R d×k represents the feature space of the parameter W0, the number of parameters is d × k, r represents the rank in the low-rank decomposition, r << min(d, k), making d × k << d × r + r × k;

[0051] Direct preference loss function is:

[0052]

[0053] where, π θ , π ref respectively represent the local large language model with LoRA adapter and the local large language model without LoRA adapter. During the optimization process, only the local large language model π θ with LoRA adapter is updated for LoRA parameters; y w and y l respectively represent the successful script and the rejected script in the preference dataset D prefer ; σ represents the sigmoid function, which is used to convert the output value into a probability value between (0, 1); β represents the hyperparameter; π θ (y w |x) represents the cumulative probability that the local large language model with LoRA adapter generates the successful script y w given the input x; π θ (y l |x) represents the cumulative probability that the local large language model with LoRA adapter generates the rejected script y l given the input x; π ref (y l |x) represents the cumulative probability that the local large language model without LoRA adapter generates the rejected script y l given the input x; π ref (y w |x) represents the cumulative probability that the local large language model without LoRA adapter generates the successful script y w given the input x.

[0054] A computer device of the present invention comprises: a memory and a processor, and a computer program stored in the memory. When the computer program is executed on the processor, it implements the method for generating automated biochemical experiment script training based on RAG and DPO.

[0055] Compared with the prior art, the present invention has the following beneficial effects:

[0056] 1. This method uses a locally deployed large-scale language model, injects professional knowledge base context through retrieval enhancement generation, and combines direct preference optimization to fine-tune the model based on preference data. It then performs platform simulation verification and iterative optimization, significantly improving the script generation pass rate of locally deployed large-scale language models with smaller parameter scales. This method can effectively avoid the risk of experimental data leakage and ensure the data security of laboratory research, while significantly reducing long-term operating costs, providing a cost-effective solution for laboratories with limited resources.

[0057] 2. This method constructs an end-to-end pipeline framework from user input to context retrieval, script generation, and then to experimental platform verification, replacing scattered functional modules, ensuring the consistency and optimization capabilities of the generation process, thereby improving the efficiency and reliability of biochemical experiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0058] Figure 1 Schematic diagram of the process of the method for generating automated biochemical experiment script training based on RAG and DPO according to an embodiment of the present invention;

[0059] Figure 2 Comparison chart of the baseline and the present invention for the pass rate of biochemical experiment scripts generated for three locally deployed large-scale language models.

[0060] Figure 3 Comparison of the baseline and our method for generating biochemical experiment scripts for three locally deployed large-scale language models using average iterative optimization rounds. DETAILED DESCRIPTION

[0061] In order to enable those skilled in the art to better understand the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative work shall fall within the scope of protection of the present invention.

[0062] This embodiment discloses a method for generating automated script training for biochemical experiments based on RAG and DPO, which significantly improves the level of script generation automation and the pass rate, reduces the cost of manual intervention, and provides an efficient solution for biochemical experiment automation. Figure 1 As shown, the specific steps include:

[0063] S1. Collect technical documents related to biochemical experiments, clean the data (remove noise information, format text), convert them into Markdown format, divide the documents into blocks, and build an external knowledge base to support the search function. The specific process includes:

[0064] S101. Remove interference information from biochemical experimental technical documents and unify their formats;

[0065] S102. Use the olmOCR tool to convert technical documents in PDF format into Markdown format;

[0066] S103, create a text segmenter, set the size of each text block to 800 characters, and have an overlap of 200 characters between adjacent blocks;

[0067] S104: Use a text splitter to divide the document in Markdown format into blocks and store them.

[0068] S2. Use the BM25 algorithm (Best Match 25) to build a keyword-based search engine from the text block, and use Faiss (Fast Similarity Search Library) to construct a vector search engine. Combine these two search engines to create a hybrid search engine that integrates keyword and semantic search results. The specific process includes:

[0069] S201, extracting the content and metadata of each text block, forming a text list and a metadata list respectively;

[0070] S202: constructing a first search engine based on the text list, wherein the first search engine uses the BM25 algorithm (Best Match 25) to sort and return top k relevant documents according to the query keyword relevance score;

[0071] S203, using the embedding model Conan-embedding-v1 to convert the content in the text list into a semantic vector representation, and based on the semantic vector, using the Faiss library (Fast Similarity Search Library) based on the L2 distance accurate nearest neighbor search retrieval method to construct a second retriever to retrieve the top k relevant documents;

[0072] S204: Assign a first weight to the relevant documents returned by the first search engine, and assign a second weight to the relevant documents returned by the second search engine;

[0073] S205: For the same document that appears in the results of the two search engines, a final score is obtained by adding the weighted score of the first search engine and the weighted score of the second search engine;

[0074] S206. For documents that only appear in a single search engine result, directly use their weighted scores as the final scores.

[0075] S207 , sorting all relevant documents according to the final scores, and returning the top k documents with the highest scores as mixed retrieval results.

[0076] S3. Generate experimental process descriptions using large language models, identify equipment, protocols, and materials, and retrieve relevant documents from the knowledge base based on these descriptions. The specific process includes:

[0077] S301: Receive an operation goal input by a user, where the operation goal is related to a specific laboratory experiment.

[0078] S302. Based on the operation goal, use a large language model (LLM) to search for a prompt template in an external knowledge base to generate an experimental process description corresponding to the operation goal. The large language model is an existing natural language processing technology that is based on the Transformer architecture and can generate natural language and code through large-scale text data pre-training (Vaswani, Ashish, et al. "Attention is all you need." Advances inneural information processing systems 30 (2017).);

[0079] The external knowledge base search prompt template is a structured natural language instruction template, which provides a guidance framework for the large language model to generate the experimental process description corresponding to the operation target, where the experimental process description is based on the Opentrons API manual and Opentrons protocol library.

[0080] The external knowledge base search prompt template specifically includes the experimental equipment and materials required for the operation target, the key steps required for the experimental protocol of the operation target and its corresponding Opentrons API function (such as protocol.load_labware, pipette.transfer), and the experimental conditions and parameters of the operation target.

[0081] As an example, experimental equipment and materials include laboratory equipment and reagents (eg, 96-well plates, buffer solutions).

[0082] Experimental conditions and parameters, including volume, time, and sample number (e.g., 100 μL volume).

[0083] S303: Retrieve relevant documents from the external knowledge base using a hybrid retriever according to the experimental process description generated by the large language model.

[0084] S4. The experimental scripts generated by the large-scale language model based on the retrieval results of the hybrid retriever and the experimental process are simulated and verified by the opentrons_simulate tool of the Opentrons software platform. The scripts that fail the verification are iteratively optimized. The specific process includes:

[0085] S401, filling the operation target input by the user and the relevant documents retrieved from the external knowledge base using the hybrid retriever into the placeholder of the experiment script generation prompt template;

[0086] S402. Generate a prompt template using the filled experimental script, construct a conversation history message (i.e., messages), input it into the large language model, and extract the Python experimental script from the reply message of the large language model language by parsing the Markdown code block tag (marked with three backticks).

[0087] S403, using the opentrons_simulate tool of the Opentrons software platform to simulate and verify the experimental script and generate a verification result;

[0088] S404: If the verification result indicates that the script execution fails, the error information generated during the verification process is extracted and added to the end of the conversation history message (i.e., messages) between the user and the large language model;

[0089] S405. The large language model analyzes the error information based on the updated conversation history messages, infers the cause of the problem, and generates a new experimental script. If the new experimental script still fails the simulation verification, the error information generated during the verification process is extracted again and added to the end of the conversation history messages between the user and the large language model. The next round of experimental script generation continues until the simulation verification passes or the number of simulation verifications reaches an upper limit.

[0090] As an embodiment, the experimental script generation prompt template is a structured natural language instruction template, which provides a guidance framework for generating Python experimental scripts for large language models. The main contents of the experimental script generation prompt template are: defining the role of the model (defining the model as a biological, computer and engineering expert), integrating the operation purpose input by the user into the experimental script generation prompt template, and integrating the retrieved document as context into the experimental script generation prompt template; the experimental script generation prompt template contains placeholders, and the operation purpose input by the user and the retrieved document can be dynamically filled into the placeholders when the program is running.

[0091] S5. Obtain successful scripts and failed scripts related to laboratory experiments, where successful scripts are Python experimental scripts that have been verified by the experimental platform, and failed scripts are Python experimental scripts that have not been verified by the experimental platform;

[0092] Based on the success script and failure script, construct the preference dataset D prefer , where each preference data pair includes the experimental task description, success script y w and reject script y l .

[0093] S6. Use direct preference optimization (DPO) to fine-tune the local large language model using LoRA. The specific process includes: adding low-order rank matrix A and low-order rank matrix B in parallel next to the linear module of the attention layer of the large language model. During training, only low-order rank matrix A and low-order rank matrix B are updated. The formula is expressed as:

[0094] W0+ΔW=W0+BA,W0∈R d×k , B∈R d×r , A∈R r×k ;

[0095] Where ΔW is the incremental parameter, d and k represent the input and output dimensions of the original attention layer in the large language model, and R d×k Represents the feature space of parameter W0, the parameter quantity is d×k, r represents the rank in the low-rank decomposition, r<<min(d,k), so that d×k<<d×r+r×k;

[0096] The direct preference loss function is:

[0097]

[0098] Among them, π θ Represents the local large language model with LoRa adapter, π ref Represents the local large language model without the LoRA adapter. During the optimization process, only the local large language model π with the LoRA adapter is used. θUpdate LoRA parameters; w and y l Represents the preference data set D prefer The success script and rejection script in ; σ represents the sigmoid function, which is used to convert the output value into a probability value between (0,1); β represents the hyperparameter; π θ (y w |x) indicates that given input x, the local large language model with LoRA adapter successfully generates script y w The cumulative probability of π ref (y l |x) means that given input x, the local large language model without LoRA adapter generates rejection script y l The cumulative probability of π ref (y w |x) indicates that given input x, the local large language model without the LoRA adapter successfully generates script y w The cumulative probability of π θ (y l |x) indicates that given input x, the local large language model with LoRA adapter generates the rejection script y l The cumulative probability of .

[0099] The user inputs the biochemical experiment operation objectives and experimental equipment as prompts into the local large-scale language model trained by direct preference optimization (DPO). Combined with the hybrid retriever retrieval context, the experimental process script is iteratively generated and simulated for verification, and finally a script that is run through is obtained.

[0100] Figure 2 and Figure 3 Comparisons of the biochemistry experiment script pass rates and average iterative optimization rounds for three locally deployed models are presented, comparing the baseline model with the proposed method. The baseline model refers to the original model without the proposed method, while the proposed method refers to the optimized large-scale language model that combines Retrieval Enhanced Generation (RAG) and Direct Preference Optimization (DPO) from steps S2 to S4. The three models are: llama3-8b-instruct, qwen2.5-1.5b-instruct, and llama3.2-1b-instruct.

[0101] Figure 2 It shows that the passing rate of Llama3-8B-Instruct is improved from 63% of the unmodified model to 81% of the present invention; Qwen2.5-1.5B-Instruct is improved from 45% to 72% of the present invention; and Llama3.2-1B-Instruct is improved from 33% to 63% of the present invention.

[0102] Figure 3 It shows that the number of iterations of Llama3-8B-Instruct dropped from 5.5 of the unimproved model to 2.4 of the present invention, Qwen2.5-1.5B-Instruct decreased from 5.6 to 3.8, and Llama3.2-1B-Instruct increased from 2 to 3.4. As for the abnormal increase in the number of iterations of Llama3.2-1B-Instruct, this may be because its parameter scale is the smallest (1B), and under the baseline state, due to limited capabilities, it only generates simple but low-quality scripts (pass rate 33%), and the iteration requirement is relatively low (2 times). After the introduction of RAG and DPO in the present invention, the pass rate is greatly improved to 63%, but because the model capacity is not enough to quickly adapt to complex contexts and preference optimization, more iterations are required to converge to high-quality output. However, its pass rate increased from 33% to 63%, which still verifies the effectiveness of this method.

[0103] Table 1 shows the performance comparison of four models for generating automated scripts for biochemical experiments. The four models are: GPT-4, llama3-8b-instruct, qwen2.5-1.5b-instruct and llama3.2-1b-instruct. The evaluation indicators include script pass rate, average iteration optimization rounds and data security. The GPT-4 API achieved a 100% script pass rate, but it relies on a cloud interface and has an average iteration optimization round of 3.2. The optimized Llama3-8B-Instruct model of the present invention has a pass rate of 81% in a local deployment environment, and the average iteration rounds are only 2.4, which is 0.8 less than the average iteration rounds of the GPT-4 API. The pass rates for Qwen2.5-1.5B-Instruct and Llama3.2-1B-Instruct were 72% and 63%, respectively, with 3.8 and 3.4 iterations, respectively. While these pass rates are lower than those of GPT-4, the differences in average iterations are only 0.6 and 0.2, respectively. Furthermore, this approach offers the advantages of high security for local deployment and minimal requirements for environmental resources. Therefore, this method effectively improves the generation performance of small-scale models in resource-constrained environments and is particularly suitable for automated script generation in biochemical experiments with high security requirements and limited resource costs.

[0104] Table 1. Comprehensive analysis of experimental results

[0105]

[0106]

[0107] In summary, this paper proposes a method for training and generating automated biochemical experiment scripts based on RAG and DPO. This method constructs an external knowledge base and a hybrid search engine, uses a large language model to retrieve relevant documents, and then generates experimental scripts through platform validation and iterative optimization. Furthermore, Direct Preference Optimization (DPO) is used to fine-tune the local large language model using LoRA to improve script generation performance.

[0108] It is worth noting that in the above-mentioned device embodiment, the modules included are only divided according to functional logic, but are not limited to the above-mentioned division, as long as the corresponding functions can be achieved; in addition, the specific names of the modules are only for the convenience of distinguishing each other, and are not used to limit the scope of protection of the present invention.

[0109] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present invention. Various modifications to the embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present invention. Therefore, the present invention is not limited to the embodiments shown herein, but is intended to be embodied in the widest possible scope consistent with the principles and novel features disclosed herein.

Claims

1. A biochemical experiment automation script training and generation method based on RAG and DPO, characterized by: The following steps are involved: The user inputs the biochemical experiment's operational objectives and experimental equipment as prompts into a local large-scale language model trained with Direct Preference Optimization (DPO). Combined with the hybrid retriever's retrieval context, the experimental process is iteratively generated and simulated for verification, ultimately resulting in a script that is successfully executed. Training a local large language model consists of the following steps: S1. Collect biochemical experimental technical documents and use the text blocks obtained after preprocessing the biochemical experimental technical documents as an external knowledge base; S2, integrates the BM25 algorithm and the Faiss library to build a hybrid search engine, which is used to extract the content and metadata of text blocks and output hybrid search results; S3. Generate an experimental process using a large language model, and use a hybrid retriever to retrieve relevant documents from the external knowledge base according to the experimental process description generated by the large language model; S4. Generate experimental scripts based on the hybrid search results of the retriever, verify them through platform simulation, and iteratively optimize the scripts that fail the verification; S5. Mark the successful and failed scripts as preferred data pairs and construct a training dataset; S6. Fine-tune LoRA on a local large language model using training dataset and direct preference optimization.

2. The method for generating automated biochemical experiment script training based on RAG and DPO according to claim 1, characterized in that: The preprocessing of step S1 specifically includes: Remove interference information from biochemical experimental technical documents and unify their formats; Use the olmOCR tool to convert technical documents in PDF format into Markdown format; Create a text splitter, set the size of each text block to N characters, and there is an overlap of n characters between adjacent blocks; Use the text splitter to store Markdown documents in chunks.

3. The method for generating automated biochemical experiment script training based on RAG and DPO according to claim 1, characterized in that: Step S2 specifically includes: Extract the content and metadata of each text block and form a text list and a metadata list respectively; Building a first search engine based on the text list, wherein the first search engine uses the BM25 algorithm to sort and return the top k relevant documents according to the keyword relevance scores of the query; The content in the text list is converted into a semantic vector representation using the Conan-embedding-v1 embedding model. Based on the semantic vector, a second retriever is constructed using the Faiss library's L2 distance-based exact nearest neighbor search retrieval method to retrieve the top k relevant documents. Assigning a first weight to the relevant documents returned by the first search engine, and assigning a second weight to the relevant documents returned by the second search engine; For the same document that appears in the results of two search engines, the final score is obtained by adding the weighted score of the first search engine and the weighted score of the second search engine; For documents that only appear in a single retriever's results, their weighted scores are used directly as the final scores; All relevant documents are sorted according to the final scores, and top k documents with the highest scores are returned as hybrid retrieval results.

4. The method for generating automated biochemical experiment scripts based on RAG and DPO according to claim 1, characterized in that: In step S3, the large language model searches for a prompt template based on an external knowledge base and generates an experimental process description corresponding to the operation target; The external knowledge base search prompt template is a structured natural language instruction template, including the experimental equipment and materials, laboratory instruments and reagents, experimental conditions and parameters required for the operation target, the steps required for the experimental protocol of the operation target, and the corresponding Opentrons API function.

5. The method for generating automated biochemical experiment script training based on RAG and DPO according to claim 1, characterized in that: Experimental equipment and materials include laboratory appliances and reagents; The experimental conditions and parameters include volume, time and number of samples.

6. The method for generating automated biochemical experiment script training based on RAG and DPO according to claim 1, characterized in that: In step S4, the large language model is based on the retrieval results of the hybrid retriever and the experimental script generated by the experimental process. It is simulated and verified by the opentrons_simulate tool of the Opentrons software platform. The scripts that fail the verification are iteratively optimized. Step S4 is specifically as follows: Filling the operation target input by the user and the relevant documents retrieved from the external knowledge base using the hybrid retriever into the placeholder of the experiment script generation prompt template; Use the populated experimental script to generate a prompt template, construct a conversation history message, input the conversation history message into the large language model, and extract the Python experimental script from the reply message in the large language model language by parsing the Markdown code block tags; Use the opentrons_simulate tool of the Opentrons software platform to simulate and verify the extracted Python experimental script to generate verification results; If the verification result indicates that the script execution fails, the error message generated during the verification process is extracted and formatted as a prompt instruction, and the prompt instruction is added to the end of the conversation history message between the user and the large language model for update; Based on the updated conversation history messages, the large language model combines the operation goals in the initial prompt instructions and the mixed retrieval results to analyze the error information, infer the cause of the problem, and generate a new experimental script. If the new experimental script still fails the simulation verification, the error information generated during the verification process is extracted again and added to the end of the conversation history messages between the user and the large language model. The next round of experimental script generation continues until the simulation verification passes or the number of simulation verifications reaches the upper limit.

7. The method for generating automated biochemical experiment script training based on RAG and DPO according to claim 6, characterized in that: The experiment script generation prompt template is a structured natural language instruction template. It incorporates the user's input operation purpose into the experiment script generation prompt template and incorporates the search document as context into the experiment script generation prompt template, providing a guidance framework for generating Python experiment scripts for large language models. The experimental script generates a prompt template including: Role of the model: Biological, computer and engineering experts; Placeholders are dynamically filled with the operation purpose entered by the user and the retrieved documents when the program is running.

8. The method for generating automated biochemical experiment script training based on RAG and DPO according to claim 1, characterized in that: In step S5, the successful scripts and failed scripts are marked as preferred data pairs to construct a training data set, including: Obtain successful and failed scripts related to laboratory experiments. Successful scripts are Python experimental scripts that have passed platform verification, and failed scripts are Python experimental scripts that have not passed platform verification. Based on the success script and the rejection script, construct the preference data set D prefer , where each preference data pair includes the experimental task description, success script y w and reject script y l .

9. The method for generating biochemical experiment automation script training based on RAG and DPO according to claim 1, characterized in that: In step S6, the low-rank matrix A and the low-rank matrix B are added in parallel next to the linear module of the attention layer of the large language model; During training, only the low-rank matrix A and the low-rank matrix B are updated. The formula is: W0+ΔW=W0+BA,W0∈R d×k , B∈R d×r , A∈R r×k , where ΔW is the incremental parameter, d and k represent the input and output dimensions of the original attention layer in the large language model, and R d×k Represents the feature space of parameter W0, the parameter quantity is d×k, r represents the rank in the low-rank decomposition, r<<min(d,k), so that d×k<<d×r+r×k; Direct preference loss function for: Among them, π θ , π ref Represent the local large language model with LoRA adapter and the local large language model without LoRA adapter respectively. During the optimization process, only the local large language model π with LoRA adapter is optimized. θ Update LoRA parameters; w and y l Represents the preference data set D prefer The success script and rejection script in ; σ represents the sigmoid function, which is used to convert the output value into a probability value between (0,1); β represents the hyperparameter; π θ (y w |x) indicates that given input x, the local large language model with LoRA adapter successfully generates script y w The cumulative probability of π θ (y l |x) indicates that given input x, the local large language model with LoRA adapter generates the rejection script y l The cumulative probability of π ref (y l |x) means that given input x, the local large language model without LoRA adapter generates rejection script y l The cumulative probability of π ref (y w |x) indicates that given input x, the local large language model without the LoRA adapter successfully generates script y w The cumulative probability of .

10. A computer device, characterized in that: The invention comprises: a memory and a processor and a computer program stored in the memory, and when the computer program is executed on the processor, it implements the method for generating automated script training for biochemical experiments based on RAG and DPO as described in any one of claims 1 to 9.

Citation Information

Cited By

  • Water conservancy inspection retrieval execution method and system based on LoRA fine tuning

    CN121350230A

  • Intelligent teaching assisting system base LLM training method for well drilling simulator

    CN121414554A

  • Drilling simulator-oriented intelligent teaching assistant system base LLM training method

    CN121414554B