A handwritten experiment record digitization processing method, system, terminal and medium based on multi-agent cooperation

CN122548006APending Publication Date: 2026-08-11粤港澳大湾区(广东)量子科学中心
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-20
Publication Date
2026-08-11

AI Technical Summary

Technical Problem

[0005]有鉴于此,本发明的目的在于提供一种基于多智能体协同的手写实验记录数字化处理方法、系统、终端及介质,以至少解决现有技术中在手写实验记录数字化过程中容易生成不存在的化学反应或错误的计算结果的问题及缺乏自修正能力的问题

Benefits of technology

[0016]本发明提供的一种基于多智能体协同的手写实验记录数字化处理方法、系统、终端及介质,所述基于多智能体协同的手写实验记录数字化处理方法,应用于包含状态流转控制与反馈回路的多智能体系统,包括:接收实验影像记录,并利用多模态大语言模型和预定义的实验数据结构化描述文件对所述实验影像记录中的手写实验数据进行识别与推理,得到原始结构化数据;对所述原始结构化数据执行双层校验操作得到相应的审核结果;所述双层校验操作包括调用外部文献向量数据库进行方法学合理性验证的外部检索增强生成校验和查询内部历史实验数据库进行参数异常检测的内部检索增强生成校验;在所述审核结果表明所述原始结构化数据未满足预设的校验收敛条件时,则触发反馈修正机制,直至所述审核结果表明所述原始结构化数据满足所述预设的校验收敛条件,并基于通过审核的所述原始结构化数据执行化学计量计算,并将计算结果作为硬性约束先验输入所述多模态大语言模型,生成目标结构化数据报告。由此可知,本发明针对多模态大语言模型基于预定义的实验数据结构化描述文件识别与推理实验影像记录中的手写实验数据得到的原始结构化数据,通过双层检索增强生成校验操作对该原始结构化数据进行校验审核,能够抑制事实幻觉,避免生成不存在的化学反应,进而在审核结果表明原始结构化数据满足预设的校验收敛条件时,通过单独执行化学计量计算任务,将计算结果作为硬性约束先验输入多模态大语言模型生成目标结构化数据报告,能够彻底解决大语言模型在数值运算上的不稳定性,确保实验数据的化学计量比准确度,彻底消除数值计算错误,并在双层检索增强生成校验的过程中发现原始结构化数据未满足预设的校验收敛条件时触发反馈修正机制,实现自修正闭环。本申请的技术方案通过包含状态流转控制与反馈回路的多智能体系统模拟提取、审核、修正的人类专家工作流,能够减少人工校验成本,解决复杂手写场景下单一模型识别率低、稳定性差的问题,并能够具备自修正闭环能力,实现从手写实验记录到高可信度结构化知识的可靠转化,也即将这些沉淀在实验记录本上的非结构化数据转化为高精度、符合化学逻辑的结构化数据,显著提升手写实验数据提取的准确性。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122548006A_ABST
    Figure CN122548006A_ABST
Patent Text Reader

Abstract

The application provides a handwriting experiment record digitization processing method, system, terminal and medium based on multi-agent cooperation, which is applied to a multi-agent system containing state flow control and feedback loop, and the method comprises the following steps: handwritten experiment data in experiment image records are recognized and inferred by using a multimodal large language model and a pre-defined experiment data structured description file, and original structured data is obtained; a double-layer checking operation is performed on the original structured data, when the audit result shows that the original structured data does not satisfy a preset checking convergence condition, a feedback correction mechanism is triggered, until the condition is satisfied, and based on the original structured data that passes the audit, chemical metrology calculation is performed, the calculation result is taken as a hard constraint prior to inputting the multimodal large language model, and target structured data report is generated. Through the workflow of simulation extraction, audit and correction of the multi-agent system, a self-correcting closed loop is realized, and the accuracy of handwriting experiment data extraction is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of text recognition technology, and in particular to a method, system, terminal, and medium for digitizing handwritten experimental records based on multi-agent collaboration. Background Technology

[0002] In materials science research, especially in crystal growth experiments, researchers typically record key information by hand in lab notebooks. These notebooks contain a wealth of core data, including formulation components, heating programs, temperature settings, and crystal morphology characterization. This data usually exists in the form of unstructured handwritten text, sketches, chemical formulas, and abbreviations. To advance AI for Science and reverse engineering of materials, there is an urgent need to transform this unstructured data stored in paper notebooks into a high-precision, chemically logically consistent structured database.

[0003] While existing digitization techniques for extracting handwritten experimental records can achieve this, several problems remain. For instance, traditional OCR-based methods can only transcribe text at the pixel level, lacking the ability to understand chemical formulas (e.g., I2 is easily misinterpreted as 12), the arrow relationships in process flow diagrams, and the logical structure of tables. This results in a high error rate and unusable extraction results. Furthermore, end-to-end extraction methods based on general multimodal large language models, while possessing some text-image understanding capabilities, are prone to generating non-existent chemical reactions or incorrect calculation results, and lack self-correction mechanisms.

[0004] Therefore, existing technologies have shortcomings and need to be improved and developed. Summary of the Invention

[0005] In view of this, the purpose of the present invention is to provide a method, system, terminal and medium for digitizing handwritten experimental records based on multi-agent collaboration, so as to at least solve the problems of generating non-existent chemical reactions or erroneous calculation results and lacking self-correction ability in the prior art during the digitization of handwritten experimental records.

[0006] The technical solution adopted by this invention to solve the technical problem is as follows: In a first aspect, the present invention discloses a method for digitizing handwritten experimental records based on multi-agent collaboration, applicable to a multi-agent system including state transition control and feedback loops, wherein the method includes: The system receives experimental image recordings and uses a multimodal large language model and a predefined experimental data structure description file to identify and infer the handwritten experimental data in the experimental image recordings, thereby obtaining the original structured data. The original structured data is subjected to a two-layer verification operation to obtain the corresponding audit results; the two-layer verification operation includes external search enhancement generation verification by calling an external literature vector database for methodological rationality verification and internal search enhancement generation verification by querying an internal historical experimental database for parameter anomaly detection; If the audit result indicates that the original structured data does not meet the preset verification convergence condition, a feedback correction mechanism is triggered until the audit result indicates that the original structured data meets the preset verification convergence condition. Based on the audited original structured data, chemometric calculations are performed, and the calculation results are used as hard constraint priors input into the multimodal large language model to generate a target structured data report.

[0007] Optionally, before using a multimodal large language model and a predefined experimental data structured description file to identify and infer the handwritten experimental data in the experimental image recording, the method further includes: The experimental image recordings were denoised, resolution adjusted, and Base64 encoded.

[0008] Optionally, the step of performing a two-layer verification operation on the original structured data to obtain the corresponding audit result includes: The main compound name and experimental method fields in the original structured data are vectorized and embedded to generate a query vector; Using the query vector, an approximate nearest neighbor search is performed in a pre-constructed external document vector database to obtain a preset number of the most relevant document fragments; Based on the domain methodological common sense in the most relevant literature fragments, the main compound and experimental method are verified to conform to the domain methodological common sense, and misreading errors of similar-looking characters are identified to obtain the corresponding verification results; wherein, the misreading errors of similar-looking characters include the situation of misreading chemical element symbols or molecular formulas as numerical values ​​or irrelevant characters; Using the name of the main compound or the category of similar compounds as the search key, retrieve historical similar experimental records in the historical experimental database, and extract key process parameters from the historical similar experimental records. Calculate the statistical distribution characteristics of the key process parameters, and determine the mean value of the deviation of the key process parameters from the statistical distribution characteristics; Determine whether the mean exceeds a preset threshold; When the average value exceeds the preset threshold, a parameter anomaly warning sign is generated; Based on the verification results, the parameter anomaly warning indicators, and the original structured data, a structured audit result containing error type descriptions, severity levels, and correction suggestions is generated.

[0009] Optionally, the triggering feedback correction mechanism includes: Based on the review results, a correction prompt is generated and input into the multimodal large language model. The steps of recognizing and reasoning the handwritten experimental data in the experimental image record using the multimodal large language model and the predefined experimental data structured description file are then re-executed.

[0010] Optionally, the step of performing stoichiometric calculations based on the approved original structured data and using the calculation results as hard constraint priors input into the multimodal large language model to generate a target structured data report includes: A deterministic symbolic reasoning engine, independent of the multimodal large language model, is invoked to perform numerical calculations on the chemical component information in the original structured data that has passed the review, and to obtain a set of calculation results containing the molar ratio, mass fraction and stoichiometry of each component. The computation result set is constructed as a reference data object with the highest execution priority, and the reference data object is used as a hard constraint prior input into the multimodal large language model so that the multimodal large language model, under the guidance of the hard constraint prior, converts the original structured data into a standardized natural language experiment report format to obtain the target structured data report.

[0011] Optionally, after using the calculation results as hard constraint prior inputs to the multimodal large language model to generate the target structured data report, the method further includes: Receive user review instructions for the target structured data report; When the review instruction indicates that the review has been approved, the feature hash values ​​of the structured source data and multimodal content corresponding to the target structured data report are extracted, and the feature hash values ​​of the structured source data and multimodal content are stored as new trusted samples in the system knowledge base. The system's internal state memory is then updated based on the new trusted samples. If the review instruction indicates that the application has not been approved, the error information provided by the user is obtained and used as a strong correction prompt to be input into the multimodal large language model. The steps of recognizing and reasoning the handwritten experimental data in the experimental image record using the multimodal large language model and the predefined experimental data structured description file are then re-executed until the review instruction indicates that the application has been approved.

[0012] Optionally, the preset verification convergence conditions include any one or more combinations of the following: the audit result shows no errors, the audit result shows a confidence level higher than a preset value, or the number of iterations reaches a preset maximum threshold.

[0013] Secondly, this invention also discloses a digital processing system for handwritten experimental records based on multi-agent collaboration, applicable to a multi-agent system including state transition control and feedback loops, wherein the system includes: The recording and receiving module is used to receive experimental image recordings; The visual perception processing module is used to identify and reason about the handwritten experimental data in the experimental image recording using a multimodal large language model and a predefined experimental data structured description file to obtain the original structured data. The audit and verification module is used to perform a two-layer verification operation on the original structured data to obtain the corresponding audit results. The two-layer verification operation includes external search enhancement generation verification by calling an external literature vector database for methodological rationality verification and internal search enhancement generation verification by querying an internal historical experimental database for parameter anomaly detection. The correction trigger module is used to trigger a feedback correction mechanism when the audit result shows that the original structured data does not meet the preset verification convergence condition, until the audit result shows that the original structured data meets the preset verification convergence condition. The data report generation module is used to perform chemometric calculations based on the approved original structured data, and input the calculation results as hard constraint priors into the multimodal large language model to generate a target structured data report.

[0014] Thirdly, the present invention discloses a terminal, comprising: a memory, a processor, and a multi-agent collaborative handwritten experiment record digitization processing program stored in the memory and executable on the processor, wherein the multi-agent collaborative handwritten experiment record digitization processing program, when executed by the processor, implements the steps of the multi-agent collaborative handwritten experiment record digitization processing method described above.

[0015] Fourthly, the present invention discloses a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program that can be executed to implement the steps of the multi-agent collaborative handwritten experimental record digitization processing method described above.

[0016] This invention provides a method, system, terminal, and medium for digitizing handwritten experimental records based on multi-agent collaboration. The method, applied to a multi-agent system including state transition control and feedback loops, includes: receiving experimental image records and using a multimodal large language model and a predefined experimental data structured description file to identify and infer the handwritten experimental data in the image records to obtain raw structured data; performing a two-layer verification operation on the raw structured data to obtain corresponding review results; the two-layer verification operation includes external retrieval enhancement generation verification by calling an external literature vector database for methodological rationality verification and internal retrieval enhancement generation verification by querying an internal historical experimental database for parameter anomaly detection; when the review result indicates that the raw structured data does not meet the preset verification convergence conditions, a feedback correction mechanism is triggered until the review result indicates that the raw structured data meets the preset verification convergence conditions; and based on the approved raw structured data, chemometric calculations are performed, and the calculation results are used as hard constraint prior inputs to the multimodal large language model to generate a target structured data report. Therefore, this invention addresses the original structured data obtained from handwritten experimental data in the experimental video recordings of a multimodal large language model based on predefined experimental data structured description documents for recognition and reasoning. It verifies and reviews this original structured data through a two-layer retrieval enhancement generation verification operation, suppressing factual illusions and preventing the generation of non-existent chemical reactions. Furthermore, when the verification results indicate that the original structured data meets the preset verification convergence conditions, a separate stoichiometric calculation task is executed. The calculation results are used as hard constraint prior inputs to the multimodal large language model to generate a target structured data report. This completely solves the instability of large language models in numerical computation, ensures the accuracy of the stoichiometric ratios of experimental data, and completely eliminates numerical calculation errors. Moreover, when the original structured data fails to meet the preset verification convergence conditions during the two-layer retrieval enhancement generation verification process, a feedback correction mechanism is triggered, achieving a self-correcting closed loop. The technical solution of this application simulates the human expert workflow of extraction, review, and correction through a multi-agent system that includes state transition control and feedback loop. This reduces the cost of manual verification, solves the problems of low recognition rate and poor stability of a single model in complex handwritten scenarios, and has self-correcting closed-loop capability. It enables reliable transformation from handwritten experimental records to highly reliable structured knowledge, that is, transforming the unstructured data accumulated in the experimental record into high-precision, chemically logical structured data, which significantly improves the accuracy of handwritten experimental data extraction. Attached Figure Description

[0017] Figure 1 This is a flowchart of a preferred embodiment of the digital processing method for handwritten experimental records based on multi-agent collaboration in this invention; Figure 2 This is a flowchart of a specific method for digitizing handwritten experimental records based on multi-agent collaboration disclosed in this invention; Figure 3 This is a functional principle block diagram of a preferred embodiment of the handwritten experimental record digitization processing system based on multi-agent collaboration in this invention; Figure 4 This is a functional principle block diagram of a preferred embodiment of the terminal in this invention. Detailed Implementation

[0018] To make the objectives, technical solutions, and advantages of this invention clearer and more explicit, the invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative of the invention and are not intended to limit the invention.

[0019] In materials science research, especially in crystal growth experiments, researchers typically record key information by hand in lab notebooks. These notebooks contain a wealth of core data, including formulation components, heating programs, temperature settings, and crystal morphology characterization. This data usually exists in the form of unstructured handwritten text, sketches, chemical formulas, and abbreviations. To promote "AI for Science" and reverse engineering of materials, there is an urgent need to transform this unstructured data stored in paper notebooks into a high-precision, chemically logical, structured database.

[0020] Existing digital extraction technologies for such handwritten experimental records, such as extraction methods based on traditional OCR (Optical Character Recognition) technology and end-to-end extraction methods based on general multimodal large language models, can extract handwritten experimental records, but still have shortcomings such as lack of semantic understanding, factual illusion, numerical calculation errors, and lack of self-correction capabilities.

[0021] Specifically, extraction methods based on traditional OCR technology utilize computer vision to recognize text characters in images, but can only perform literal character recognition, lacking semantic understanding and domain logic. They cannot correctly interpret structural relationships in complex experimental tables or process flow diagrams with arrows (such as heating and cooling curves); furthermore, they cannot distinguish between characters with similar shapes (e.g., easily misidentifying the chemical transport agent I2 as the number 12), resulting in a high error rate in the extracted data, rendering it completely unusable and unsuitable for subsequent data mining. In other words, extraction methods based on traditional OCR technology can only achieve pixel-level text transcription, lacking the ability to understand chemical formulas (e.g., I2 is easily misidentified as 12), the arrow relationships in process flow diagrams, and the logical structure of tables, thus failing to interpret experimental semantics and leading to high error rates and unusable extraction results.

[0022] However, end-to-end extraction methods based on general multimodal large language models directly utilize general visual large language models (such as GPT-4V, Qwen-VL, etc.) to understand experimental images and output text. Although they possess a certain level of image-text comprehension ability, they still suffer from serious problems, such as severe "fact illusions" and logical deficiencies. Specifically, the general large language model lacks domain-specific knowledge constraints, such as those related to crystal growth, often resulting in the fabrication of non-existent chemical formulas or the generation of experimental parameters that violate common physicochemical principles. For example, it may incorrectly identify gaseous transport agents in solid-state sintering and cannot utilize external knowledge bases to verify the rationality of the formulation. Furthermore, their numerical calculation and logical reasoning abilities are weak. When processing ingredient lists, the large language model struggles to perform precise mathematical calculations, such as calculating molar ratios from mass, frequently outputting incorrect stoichiometric ratios, leading to distortion of key experimental parameters. The lack of self-correction capabilities in "open-loop" workflows, where existing methods typically employ a one-off, single-step reasoning approach, means that if the visual perception stage malfunctions, the system cannot reflect, verify, and correct itself by incorporating contextual logic like human experts. This results in the final data failing to meet the high standards required for constructing research-grade databases. In other words, end-to-end extraction methods based on general multimodal large language models suffer from problems such as "fact illusion," unreliable numerical calculations, and the absence of self-correction mechanisms.

[0023] To this end, this application provides a digital processing scheme for handwritten experimental records based on multi-agent collaboration, which can suppress factual illusions, completely eliminate numerical calculation errors, and has self-correcting closed-loop capabilities, thereby significantly improving the accuracy of handwritten experimental data extraction.

[0024] Please see Figure 1 , Figure 1 This is a flowchart of the digital processing method for handwritten experimental records based on multi-agent collaboration in this invention. (For example...) Figure 1 As shown in the embodiment of the present invention, the digital processing method for handwritten experimental records based on multi-agent collaboration is applied to a multi-agent system including state transition control and feedback loops, and includes: Step S11: Receive experimental image recordings and use a multimodal large language model and a predefined experimental data structure description file to identify and infer the handwritten experimental data in the experimental image recordings to obtain the original structured data.

[0025] The experimental image recording can be images of handwritten experimental records, PDF scans exported from experimental instruments, screenshots of electronic experimental notebooks, or photos of whiteboard notes containing experimental procedures. The experimental image recording can also record handwritten text characters, chemical structural diagrams, and experimental process flow diagrams. The multimodal large language model can be Qwen-VL, and it loads a predefined experimental data structure description file (JSON Schema) defining a multi-level data architecture including metadata fields, ingredient component fields, and process flow fields, and setting data type constraints, value range constraints, and logical association constraints for each field.

[0026] In this embodiment, after receiving the experimental image recordings, the recordings can be denoised, have their resolution adjusted, and be Base64 encoded before initial structured extraction. For example, a multimodal large language model combined with a predefined JSON Schema can be used to recognize and understand handwritten text, chemical formulas, and process diagrams in the images to generate raw structured data. The predefined JSON Schema includes constraints such as metadata, ingredients, and process fields.

[0027] Specifically, experimental image recordings and predefined experimental data structured description files are input into a multimodal large language model. The model's visual encoding capabilities are used to extract visual features from the image recordings, identify handwritten text semantics, parse the topological connections of chemical structures, and reconstruct the sequence of steps and operating conditions in the process flow diagram. The model's reasoning capabilities are then used to semantically align the extracted visual features with the field definitions in the structured description file. Guided by logical association constraints, unstructured visual information is mapped into intermediate data objects that conform to data type constraints. Finally, based on the value range constraints and logical association constraints in the structured description file, the intermediate data objects are validated and corrected for compliance, generating the original structured data.

[0028] Step S12: Perform a two-layer verification operation on the original structured data to obtain the corresponding audit results; the two-layer verification operation includes external search enhancement generation verification by calling an external literature vector database for methodological rationality verification and internal search enhancement generation verification by querying an internal historical experimental database for parameter anomaly detection.

[0029] In this embodiment, a multimodal large language model and a predefined experimental data structured description file are used to identify and infer handwritten experimental data from experimental image recordings. After obtaining the original structured data, a two-layer verification operation can be performed on the original structured data. Specifically, through a two-layer RAG (Retrieval-Augmented Generation) architecture, an external literature knowledge base is introduced to perform fact-checking on the original structured data identified by the multimodal large language model. This effectively intercepts methodological illusions. That is, in this embodiment, it is not just simple digitization (OCR), but through the RAG mechanism, external literature knowledge and historical experimental memory are introduced to automatically associate isolated experimental records with external physical properties, such as melting point and lattice parameters. This makes the generated database not only contain the experimental process but also the physical properties, effectively suppressing the factual illusions of general large models in professional fields, significantly improving the accuracy and scientific nature of data extraction, and providing a high-dimensional, high-value structured data foundation for subsequent AI-based reverse engineering of materials.

[0030] Specifically, the main compound name and experimental method fields in the original structured data are vectorized and embedded to generate query vectors. These query vectors are then used to perform an approximate nearest neighbor search in a pre-constructed external literature vector database to obtain a preset number of most relevant literature fragments. Based on the domain methodological common sense in these most relevant literature fragments, the main compound and experimental methods are verified to conform to the domain methodological common sense, and misreading errors involving similar-looking characters are identified, yielding corresponding verification results. Misreading errors involving similar-looking characters include misidentifying chemical element symbols or molecular formulas as numerical values ​​or irrelevant characters. Using the main compound name or similar compound category as the search key, historical experimental records of similar experiments are retrieved from the historical experimental database, and key process parameters are extracted from these records. The statistical distribution characteristics of the key process parameters are calculated, and the mean value of the deviation from the statistical distribution characteristics is determined. It is then determined whether the mean value exceeds a preset threshold. If the mean value exceeds the preset threshold, a parameter anomaly warning is generated. Based on the verification results, the parameter anomaly warning, and the original structured data, a structured audit result containing error type descriptions, severity levels, and correction suggestions is generated.

[0031] Understandably, the first layer of verification is external knowledge semantic verification, which is essentially external search-enhanced generation verification (External RAG) that calls external literature vector databases to verify the methodological rationality. This involves vectorizing and embedding the extracted main compound name and experimental methods, and then searching for Top-K relevant literature fragments (i.e., the most relevant fragments) in external literature vector databases (i.e., vector search sources), such as ChromaDB, Milvus, Faiss, or Elasticsearch. Based on the search results, the system verifies whether the extracted formulation conforms to the methodological common sense in the field, such as verifying whether a certain precursor is suitable for a specific crystal growth method, and identifies potential character misreadings, such as misinterpreting the transport agent I2 as the number 12. The second layer of verification is internal memory anomaly detection, which involves querying the internal historical experimental database for parameter anomaly detection using Internal Retrieval Enhanced Generative Validation (Internal RAG). This involves retrieving historical experimental records of similar compounds from a historical experimental database, such as SQLite, extracting key process parameters, and statistically analyzing their distribution (e.g., the distribution of maximum growth temperature and cooling rate). If the currently extracted key process parameter deviates from the historical average of its distribution by more than a preset threshold (±100°C), it is marked as an anomaly warning. Finally, the search results from External RAG and Internal RAG are combined with the original structured data to output the audit results, which include a detailed description of the error, its severity, and suggested corrections.

[0032] Among them, the external document vector database can be unstructured PDF documents, structured knowledge graphs such as MatBase and OQMD, and can also perform deeper logical verification through SPARQL or subgraph matching.

[0033] Step S13: When the audit result indicates that the original structured data does not meet the preset verification convergence condition, a feedback correction mechanism is triggered until the audit result indicates that the original structured data meets the preset verification convergence condition. Based on the audited original structured data, chemometric calculations are performed, and the calculation results are used as hard constraint priors input into the multimodal large language model to generate a target structured data report.

[0034] The preset verification convergence conditions include any one or more combinations of the following: the audit result shows no errors, the audit result shows a confidence level higher than a preset value, or the number of iterations reaches a preset maximum threshold.

[0035] Understandably, in the initial structuring stage, the handwritten experimental data in the experimental image records is extracted under the constraints of the experimental data structure description file predefined by the multimodal large language model to obtain the original structured data. If the original structured data has serious errors, or the confidence level is not higher than the preset value, or the number of iterations in the current self-correction loop has not reached the preset maximum threshold, then the feedback correction mechanism is triggered to achieve fully automated error detection and self-repair.

[0036] Specifically, correction prompts are generated based on the review results and input into a multimodal large language model. The process then re-executes the steps of recognizing and inferring handwritten experimental data from experimental image records using the multimodal large language model and a predefined experimental data structured description file. In other words, the review results are converted into correction prompts in natural language form and input into the multimodal large language model. This allows the model to perform secondary focusing on specific error areas, outputting corrected original structured data. A two-layer verification operation is then performed on the corrected original structured data until it meets preset verification convergence conditions. For example, if logical contradictions exist in the original structured data, such as abnormal temperature or missing components, specific correction instructions are automatically generated, triggering the multimodal large language model to re-examine the handwritten experimental record images for the errors indicated by the correction instructions. This feedback correction mechanism mimics the process of human experts checking, reflecting, and correcting, improving the first-pass yield of various complex handwritten records and reducing the cost of manual intervention.

[0037] Furthermore, if the audit results indicate that the original structured data meets the preset verification convergence conditions, that is, the original structured data is error-free, or the confidence level is higher than the preset value, or the number of iterations in the current self-correction loop reaches the preset maximum threshold, then a neural symbolic collaborative strategy is adopted for data processing. That is, chemometric calculations are performed based on the audited original structured data, and the calculation results are used as hard constraint prior inputs to the multimodal large language model to generate a target structured data report.

[0038] Specifically, a deterministic symbolic inference engine, independent of the multimodal large language model, is invoked to perform numerical calculations on the chemical component information in the approved raw structured data, resulting in a set of calculation results including the molar ratio, mass fraction, and stoichiometric ratio of each component. This set of calculation results is then constructed as a reference data object with the highest execution priority, and this reference data object is used as a hard constraint prior input into the context prompt window of the multimodal large language model. This locks the key numerical fields in the subsequent generation process, preventing numerical deviations caused by probabilistic illusions. Under the guidance of the hard constraint prior, the multimodal large language model performs the style transfer task from structured data to unstructured documents, converting the raw structured data into a standardized natural language experimental report format, resulting in the target structured data report.

[0039] Specifically, through a neural symbolic collaboration strategy, deterministic Python code (i.e., a deterministic symbolic inference engine independent of the multimodal large language model) can take over numerical calculation tasks such as molar ratios, and force the calculation results into the multimodal large language model to generate the target structured data report via a reference data table. This division of labor, where the multimodal large language model is responsible for understanding semantics and the code is responsible for precise calculations, completely solves the instability of the multimodal large language model in numerical operations and ensures the accuracy of stoichiometry in experimental data.

[0040] It should be noted that, in addition to calculating molar ratios, the deterministic symbolic reasoning engine also possesses symbolic computation capabilities for other deterministic logic, such as calculating solution dilution factors in biological experiments and verifying the rationality of voltage and current data based on Ohm's law in physics experiments. Furthermore, during report generation, material physical property data associated with chemical components are simultaneously acquired and integrated with the calculation result set into a standardized natural language experimental report, resulting in the target structured data report. The target structured data report file is a machine-readable data format, such as JSON / XML, that strictly adheres to structured description file format specifications, and fully preserves the key experimental parameters and logical flow from the handwritten experimental data.

[0041] For example, by stripping away the computational functions of the large model and calling the Python-based pymatgen library to accurately calculate the molar ratios of each component based on chemical formulas and molecular weights, the calculation results are constructed into an immutable reference data table with the highest execution priority. This table is then injected into the context window of the multimodal large language model. Based on this reference data table, the multimodal large language model converts the JSON data into a standardized Markdown format experimental report and automatically supplements it with material physical properties such as melting point and space group retrieved from RAG. In other words, code enhancement technology completely solves the numerical error problem in the stoichiometric ratio calculation of the large model, ensuring the chemical logical consistency of the extracted data. This involves delegating mathematical calculations to a deterministic symbolic reasoning engine independent of the multimodal large language model, such as a Python script, a generalized rule engine, an Excel formula engine, or the WolframAlpha API, or using a knowledge graph inference engine for logical deduction, or using a traditional rule engine for hard-coded logical verification. After calculating the precise results, a reference data table is constructed and used as a hard constraint prior input into the context prompt window of the multimodal large language model. This can lock the key numerical fields in the subsequent generation process and prevent numerical deviations caused by probabilistic illusions. This approach completely eliminates numerical illusions.

[0042] Furthermore, in this embodiment, after using the calculation results as hard constraint prior input to the multimodal large language model to generate the target structured data report, the process may further include: receiving a user's review instruction for the target structured data report; if the review instruction indicates approval, extracting the feature hash values ​​of the structured source data and multimodal content corresponding to the target structured data report, storing the feature hash values ​​of the structured source data and multimodal content as new trusted samples in the system knowledge base, and updating the system's internal state memory based on the new trusted samples; if the review instruction indicates non-approval, obtaining the error information from the user feedback, and inputting the error information as a strong correction prompt into the multimodal large language model, re-executing the steps of recognizing and reasoning the handwritten experimental data in the experimental image record using the multimodal large language model and a predefined experimental data structured description file, until the review instruction indicates approval. It is understood that if the user approves, the structured data and image hash values ​​are stored in the database, and the internal memory is updated. If the user feedback is incorrect, the user feedback is obtained as a strong correction prompt, forcibly triggering a rollback mechanism until the review instruction indicates approval, thus achieving reinforcement optimization based on human feedback.

[0043] For example, in the review interaction interface, the generated target structured data report and associated anomaly warning information are displayed, and the user's review instruction for the target structured data report is received. In response to the review instruction being a confirmation instruction, the feature hash values ​​of the structured source data and multimodal content corresponding to the target structured data report are extracted, and these are persistently stored in the system knowledge base as new trusted samples. The system's internal state memory is updated based on the new trusted samples to enhance the consistency of subsequent generation. In response to the review instruction being an error feedback instruction, user feedback information is captured and a strong correction constraint signal is constructed. Based on the strong correction constraint signal, a process rollback mechanism is triggered to force the current processing flow to roll back to the initial structure extraction step, and the strong correction constraint signal is injected as prior knowledge into the subsequent execution flow to achieve closed-loop reinforcement optimization based on human feedback.

[0044] As can be seen, in this embodiment of the invention, the original structured data obtained by the multimodal large language model based on the handwritten experimental data in the experimental image recording of the recognition and reasoning of predefined experimental data structured description documents is verified and reviewed through a two-layer retrieval enhancement generation verification operation. This can suppress factual illusions and avoid generating non-existent chemical reactions. Furthermore, when the verification results show that the original structured data meets the preset verification convergence conditions, a separate stoichiometric calculation task is performed, and the calculation results are used as hard constraint prior inputs to the multimodal large language model to generate a target structured data report. This can completely solve the instability of the large language model in numerical calculation, ensure the accuracy of the stoichiometric ratio of the experimental data, completely eliminate numerical calculation errors, and trigger a feedback correction mechanism when the original structured data does not meet the preset verification convergence conditions during the two-layer retrieval enhancement generation verification process, thus realizing a self-correcting closed loop. The technical solution of this application simulates the human expert workflow of extraction, review, and correction through a multi-agent system that includes state transition control and feedback loop. This reduces the cost of manual verification, solves the problems of low recognition rate and poor stability of a single model in complex handwritten scenarios, and has self-correcting closed-loop capability. It enables reliable transformation from handwritten experimental records to highly reliable structured knowledge, that is, transforming the unstructured data accumulated in the experimental record into high-precision, chemically logical structured data, which significantly improves the accuracy of handwritten experimental data extraction.

[0045] In one specific implementation, a self-correcting closed-loop workflow is constructed based on the LangGraph state machine. See also Figure 2As shown, in a multi-agent system containing state transition control and feedback loops, the multi-agents can be a visual perception agent (Role A), a domain review agent (Role B), a data engineering agent (Role C), and a loop state machine containing a feedback mechanism. First, the visual perception agent extracts raw experimental data from handwritten images, i.e., it performs denoising, resolution adjustment, and Base64 encoding on the experimental image records. The Qwen-VL in the visual perception agent, combined with a predefined JSON Schema, recognizes and understands the handwritten text, chemical formulas, and flowcharts in the handwritten images to generate raw structured data (Raw JSON). The VLM in the visual perception agent infers and extracts the JSON, determining if the JSON format is valid. If invalid, the visual perception agent receives a correction prompt and re-infers the JSON format. If the JSON format is valid, it enters the review stage. When the review agent finds a logical error, it does not directly output the error result but generates a correction prompt containing the error context and drives the visual perception agent to perform secondary focusing recognition on the image with this prompt. In this stage, a two-layer retrieval enhancement generation mechanism is introduced. The domain audit agent (Role B) receives the raw structured data and performs two-layer RAG hybrid verification and logical audit. The domain audit agent calls the external literature vector database to verify the methodological rationality of the formulation and calls the internal historical experimental database to detect anomalies in parameter distribution. If the RAG data layer includes an external literature vector database and a historical experimental SQL database, firstly, in the external knowledge semantic verification of the first layer, the extracted main compound name and experimental method are vectorized and embedded. Top-K relevant literature fragments are retrieved in the ChromaDB vector search engine. Based on the search results, it is verified whether the extracted formulation conforms to the methodological common sense in the field, such as physical properties and formulation verification. Then, in the internal memory anomaly detection of the second layer, parameter statistics and anomaly detection are implemented. That is, historical experimental records of similar compounds are retrieved in a historical experimental database, such as SQLite. Key process parameters are extracted from the historical experimental records of similar compounds, and the distribution of key process parameters is statistically analyzed. If the current extracted key process parameter deviates from the historical average of the distribution and exceeds a preset threshold, it is marked as an anomaly warning. Finally, the search results of external knowledge semantic verification and internal memory anomaly detection of the second layer are combined with the original structured data to output the audit result. The audit result includes a specific error description, severity, and correction suggestions.The LLM intelligent audit determines whether the audit passes based on the audit results. If the audit fails, it indicates that the original structured data does not meet the preset verification convergence conditions, meaning the original structured data contains serious errors, or the confidence level is not higher than the preset value, or the number of iterations in the current self-correction loop has not reached the preset maximum threshold. This triggers a feedback correction mechanism, converting the audit result into a correction prompt in natural language. This prompt is then input into the multimodal large language model in the visual perception agent, allowing the model to perform secondary focusing on specific error areas and output corrected original structured data. A double-layer verification operation is then performed on the corrected original structured data until it meets the preset verification convergence conditions. If the audit passes, it indicates that the original structured data does not meet the preset verification convergence conditions, meaning the original structured data is error-free, or the confidence level is higher than the preset value, or the number of iterations in the current self-correction loop has reached the preset maximum threshold. The data engineering agent (Role) then triggers a feedback correction mechanism, converting the audit result into a correction prompt in natural language. This prompt is then input into the multimodal large language model in the visual perception agent, allowing the model to perform secondary focusing on specific error areas and output corrected original structured data. The corrected original structured data is then subjected to a double-layer verification operation until it meets the preset verification convergence conditions. C) Upon receiving approved raw structured data, deterministic Python code is used to replace the large language model for numerical calculations such as molar ratios, achieving accurate molar ratio calculations. The calculation results are then forcibly injected into the multimodal large language model through a reference data table to generate the target structured data report. Based on the reference data table, the multimodal large language model converts the JSON data into a standardized Markdown format experimental report, ultimately outputting a structured data experimental report that has undergone logical verification and physical property expansion. This not only digitizes handwritten records but also automatically supplements the physical properties of materials (such as melting point and space group) through knowledge base association, transforming fragmented experimental records into a high-value-added multidimensional structured knowledge base. This achieves a deep transformation of unstructured data into a high-value knowledge base, directly empowering reverse engineering and screening of materials. Through collaboration and state machine transitions among multiple agents, the workflow of human experts in extraction, review, and correction is simulated, significantly reducing the cost of manual verification and solving the problems of low recognition rate and poor stability of single models in complex handwritten scenarios.

[0046] Furthermore, after using the calculation results as hard constraint prior input to the multimodal large language model to generate the target structured data report, the process can also include a human-computer feedback loop and knowledge accumulation stage. Specifically, the generated target structured data report and associated anomaly warning information are displayed on the review interface, and the user's review instructions for the target structured data report are received. If the user confirms approval, the structured data and image hash values ​​are stored in the database, updating the internal memory. If the user reports an error, the user feedback is used as a strong correction prompt, forcibly triggering a rollback mechanism until the review instruction indicates approval, achieving reinforcement optimization based on human feedback.

[0047] It should be noted that the visual perception agent can be GPT-4o, Claude 3.5 Sonnet, or Gemini 1.5 Pro, while the domain auditing agent can be a domain-specific fine-tuned (SFT) Llama 3 or Mistral model. The domain auditing agent and the data engineering agent can be merged into a single backend processing agent, or the visual perception agent can comprise two sub-agents: a text recognition agent and an image segmentation agent.

[0048] In this embodiment, a closed-loop flow of perception, review, and correction is achieved through a multi-agent collaborative architecture, and a deterministic symbolic reasoning engine independent of the multimodal large language model is invoked to solve the numerical illusion problem of the large model. This linear loop of perception, review, and correction can be extended to a parallel integrated architecture, that is, perception, review, and correction can be achieved through serial loops or parallel voting mechanisms. For example, three perception agents can be activated simultaneously for identification, and a referee agent can vote to determine the final result. If the votes are inconsistent, a rework is triggered. Alternatively, multiple review agents with different roles can be set up, such as one checking chemical formulas and another checking process parameters. They work in parallel and vote to decide whether to pass or fail. Another approach is to set up a commander agent to distribute tasks, and multiple engineer agents to execute them, with the commander determining whether a rework is necessary, thereby further improving the rigor of the review.

[0049] In one embodiment, such as Figure 3 As shown, based on the above-mentioned method for digitizing handwritten experimental records based on multi-agent collaboration, this invention also provides a system for digitizing handwritten experimental records based on multi-agent collaboration, applicable to a multi-agent system including state transition control and feedback loops, comprising: Record receiving module 11 is used to receive experimental image recordings; The visual perception processing module 12 is used to identify and reason about the handwritten experimental data in the experimental image recording using a multimodal large language model and a predefined experimental data structure description file to obtain the original structured data. The audit and verification module 13 is used to perform a two-layer verification operation on the original structured data to obtain the corresponding audit results; the two-layer verification operation includes external search enhancement generation verification by calling an external literature vector database for methodological rationality verification and internal search enhancement generation verification by querying an internal historical experimental database for parameter anomaly detection; The correction trigger module 14 is used to trigger a feedback correction mechanism when the audit result shows that the original structured data does not meet the preset verification convergence condition, until the audit result shows that the original structured data meets the preset verification convergence condition. The data report generation module 15 is used to perform chemometric calculations based on the approved original structured data, and input the calculation results as hard constraint priors into the multimodal large language model to generate a target structured data report.

[0050] Furthermore, it is worth noting that the working process of the handwritten experimental record digitization processing system based on multi-agent collaboration provided in this embodiment is the same as the working process of the handwritten experimental record digitization processing method based on multi-agent collaboration described above. Therefore, it will not be repeated here. For details, please refer to the working process of the handwritten experimental record digitization processing method based on multi-agent collaboration described above.

[0051] It should be noted that the handwritten experimental record digitization system based on multi-agent collaboration can be deployed on a local workstation to ensure the privacy and security of experimental data (e.g., using a locally deployed Qwen model and local vector library), or it can adopt a cloud-edge architecture, in which the front end (e.g., the Streamlit interface) runs on the user terminal, while multi-agent reasoning and knowledge base retrieval run on a private cloud server cluster.

[0052] Figure 4 A schematic diagram of the structure of a terminal provided in an embodiment of this application. The terminal may include: The memory 501, the processor 502, and the computer program stored on the memory 501 and capable of running on the processor 502.

[0053] When the processor 502 executes the program, it implements the digital processing method for handwritten experimental records based on multi-agent collaboration provided in the above embodiments.

[0054] Furthermore, the terminal also includes: Communication interface 503 is used for communication between memory 501 and processor 502.

[0055] The memory 501 is used to store computer programs that can run on the processor 502.

[0056] Memory 501 may include high-speed RAM memory, and may also include non-volatile memory, such as at least one disk storage device.

[0057] If the memory 501, processor 502, and communication interface 503 are implemented independently, they can be interconnected via a bus to communicate with each other. The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of representation, only one line is used in the diagram, but this does not imply that there is only one bus or one type of bus.

[0058] Optionally, in a specific implementation, if the memory 501, processor 502, and communication interface 503 are integrated on a single chip, then the memory 501, processor 502, and communication interface 503 can communicate with each other through an internal interface.

[0059] Processor 502 may be a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits configured to implement the embodiments of this application.

[0060] This embodiment also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described method for digitizing handwritten experimental records based on multi-agent collaboration.

[0061] Other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein.

[0062] In the description of this specification, the references to terms such as "one embodiment," "some embodiments," "example," "specific example," or "some examples," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of this application. In this specification, the illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples. Moreover, without contradiction, those skilled in the art can combine and integrate the different embodiments or examples described in this specification, as well as the features of different embodiments or examples.

[0063] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a sequenced list of executable instructions for implementing logical functions, and can be embodied in any computer-readable medium for use by, or in conjunction with, an instruction execution system, apparatus or device (such as a computer-based system, a processor-included system or other system that can read and execute instructions from and from an instruction execution system, apparatus or device).

[0064] It should be understood that the various parts of this application can be implemented using hardware, software, firmware, or a combination thereof. In the above embodiments, the N steps or methods can be implemented using software or firmware stored in memory and executed by a suitable instruction execution system. If implemented in hardware, as in another embodiment, it can be implemented using any one or a combination of the following techniques known in the art: discrete logic circuits having logic gates for implementing logical functions on data signals, application-specific integrated circuits (ASICs) having suitable combinational logic gates, programmable gate arrays (PGAs), field-programmable gate arrays (FPGAs), etc.

[0065] It should be understood that the application of the present invention is not limited to the examples above. Those skilled in the art can make improvements or modifications based on the above description, and all such improvements and modifications should fall within the protection scope of the appended claims.

Claims

1. A handwriting experiment record digitization processing method based on multi-agent cooperation, applied to a multi-agent system containing state flow control and feedback loop, characterized in that, The method includes: The system receives experimental image recordings and uses a multimodal large language model and a predefined experimental data structure description file to identify and infer the handwritten experimental data in the experimental image recordings, thereby obtaining the original structured data. The original structured data is subjected to a two-layer verification operation to obtain the corresponding audit results; the two-layer verification operation includes external search enhancement generation verification by calling an external literature vector database for methodological rationality verification and internal search enhancement generation verification by querying an internal historical experimental database for parameter anomaly detection; If the audit result indicates that the original structured data does not meet the preset verification convergence condition, a feedback correction mechanism is triggered until the audit result indicates that the original structured data meets the preset verification convergence condition. Based on the audited original structured data, chemometric calculations are performed, and the calculation results are used as hard constraint priors input into the multimodal large language model to generate a target structured data report. 2.The method of claim 1, wherein, Before using a multimodal large language model and a predefined experimental data structured description file to identify and infer the handwritten experimental data in the experimental image recording, the method further includes: The experimental image recordings were denoised, resolution adjusted, and Base64 encoded.

3. The multi-agent collaboration based handwritten lab record of claim 1 The digital processing method is characterized by, The process of performing a two-layer validation operation on the original structured data to obtain the corresponding audit results includes: The main compound name and experimental method fields in the original structured data are vectorized and embedded to generate a query vector; Using the query vector, an approximate nearest neighbor search is performed in a pre-constructed external document vector database to obtain a preset number of the most relevant document fragments; Based on the domain methodological common sense in the most relevant literature fragments, the main compound and experimental method are verified to conform to the domain methodological common sense, and misreading errors of similar-looking characters are identified to obtain the corresponding verification results; wherein, the misreading errors of similar-looking characters include the situation of misreading chemical element symbols or molecular formulas as numerical values ​​or irrelevant characters; Using the name of the main compound or the category of similar compounds as the search key, retrieve historical similar experimental records in the historical experimental database, and extract key process parameters from the historical similar experimental records. Calculate the statistical distribution characteristics of the key process parameters, and determine the mean value of the deviation of the key process parameters from the statistical distribution characteristics; Determine whether the mean exceeds a preset threshold; When the average value exceeds the preset threshold, a parameter anomaly warning sign is generated; Based on the verification results, the parameter anomaly warning indicators, and the original structured data, a structured audit result containing error type descriptions, severity levels, and correction suggestions is generated.

4. The method for digitizing handwritten lab records based on multi-agent collaboration according to claim 1, wherein, The trigger feedback correction mechanism includes: Based on the review results, a correction prompt is generated and input into the multimodal large language model. The steps of recognizing and reasoning the handwritten experimental data in the experimental image record using the multimodal large language model and the predefined experimental data structured description file are then re-executed.

5. The multi-agent collaboration based handwritten lab record of claim 1 The digital processing method is characterized by, The process involves performing chemometric calculations based on the approved original structured data, and using the calculation results as hard constraint priors input into the multimodal large language model to generate a target structured data report, including: A deterministic symbolic reasoning engine, independent of the multimodal large language model, is invoked to perform numerical calculations on the chemical component information in the original structured data that has passed the review, and to obtain a set of calculation results containing the molar ratio, mass fraction and stoichiometry of each component. The computation result set is constructed as a reference data object with the highest execution priority, and the reference data object is used as a hard constraint prior input into the multimodal large language model so that the multimodal large language model, under the guidance of the hard constraint prior, converts the original structured data into a standardized natural language experiment report format to obtain the target structured data report.

6. The multi-agent collaboration based handwritten lab record of claim 1 The digital processing method is characterized by, After using the calculation results as hard constraint prior inputs to the multimodal large language model to generate the target structured data report, the method further includes: Receive user review instructions for the target structured data report; When the review instruction indicates that the review has been approved, the feature hash values ​​of the structured source data and multimodal content corresponding to the target structured data report are extracted, and the feature hash values ​​of the structured source data and multimodal content are stored as new trusted samples in the system knowledge base. The system's internal state memory is then updated based on the new trusted samples. If the review instruction indicates that the application has not been approved, the error information provided by the user is obtained and used as a strong correction prompt to be input into the multimodal large language model. The steps of recognizing and reasoning the handwritten experimental data in the experimental image record using the multimodal large language model and the predefined experimental data structured description file are then re-executed until the review instruction indicates that the application has been approved.

7. The method according to any one of claims 1 to 6, wherein, The preset verification convergence conditions include any one or more combinations of the following: the audit result shows no errors, the audit result shows a confidence level higher than a preset value, or the number of iterations reaches a preset maximum threshold.

8. A digital processing system for handwritten experimental records based on multi-agent collaboration, applied to a multi-agent system including state transition control and feedback loops, characterized in that, The system includes: The recording and receiving module is used to receive experimental image recordings; The visual perception processing module is used to identify and reason about the handwritten experimental data in the experimental image recording using a multimodal large language model and a predefined experimental data structured description file to obtain the original structured data. The audit and verification module is used to perform a two-layer verification operation on the original structured data to obtain the corresponding audit results. The two-layer verification operation includes external search enhancement generation verification by calling an external literature vector database for methodological rationality verification and internal search enhancement generation verification by querying an internal historical experimental database for parameter anomaly detection. The correction trigger module is used to trigger a feedback correction mechanism when the audit result shows that the original structured data does not meet the preset verification convergence condition, until the audit result shows that the original structured data meets the preset verification convergence condition. The data report generation module is used to perform chemometric calculations based on the approved original structured data, and input the calculation results as hard constraint priors into the multimodal large language model to generate a target structured data report.

9. A terminal, characterized by comprising: include: The system includes a memory, a processor, and a multi-agent collaborative handwritten experimental record digitization processing program stored in the memory and executable on the processor. When the processor executes the multi-agent collaborative handwritten experimental record digitization processing program, it implements the steps of the multi-agent collaborative handwritten experimental record digitization processing method as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program that can be executed to implement the steps of the method for digitizing handwritten experimental records based on multi-agent collaboration as described in any one of claims 1 to 7.