A method and related apparatus for experimental verification and evaluation based on AI intelligent agents
By using an AI-based experimental verification and evaluation method, and employing a parent-child segmentation model and a large language model to perform structured processing and data extraction of experimental documents, the problem of low efficiency, poor accuracy, and high cost in experimental verification and evaluation during equipment development is solved. This method achieves rapid and accurate evaluation results, supporting equipment development progress and quality control.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- BEIJING RAINFE TECH
- Filing Date
- 2025-07-11
- Publication Date
- 2026-05-26
AI Technical Summary
In existing technologies, the testing, verification and evaluation processes during equipment development suffer from low efficiency, difficulty in ensuring accuracy, and high labor costs. Traditional evaluation methods are difficult to automate and process data in real time, which affects the progress and quality of equipment development.
An AI agent-based experimental verification and evaluation method is adopted. The document is segmented in a structured manner through a parent-child segmentation pattern to generate a vectorized knowledge base. The experimental data is extracted and analyzed using a large language model, and multiple agents are combined for evaluation to provide fast and accurate evaluation results.
It has achieved automation and intelligence in test verification and evaluation, significantly improving evaluation efficiency and accuracy, reducing labor costs, ensuring equipment development progress and quality, and supporting leadership decision-making.
Smart Images

Figure CN120873522B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular to an experimental verification and evaluation method and related apparatus based on AI intelligent agents. Background Technology
[0002] In the complex and critical field of equipment development, testing and verification are the core steps to ensure equipment performance and quality. With the rapid development of modern technology, the functions and structures of equipment are becoming increasingly complex, and the requirements for the accuracy and comprehensiveness of testing and verification are also becoming higher and higher.
[0003] However, current information exchange between design and testing departments suffers from numerous problems. Most results are transmitted in Word and PDF formats, which are unstructured and have limitations in data sharing and processing. After testing, evaluating the completeness, sufficiency, consistency, progress, and cost of the test verification becomes extremely difficult. Traditional evaluation methods are mostly manual, requiring evaluators to sift through numerous Word and PDF documents provided by the design and testing departments to extract relevant data. This lack of automation makes it difficult to meet the need for real-time access to test verification evaluation results for timely decision-making by test management personnel, resulting in low efficiency and impacting the overall progress of equipment development. Summary of the Invention
[0004] The purpose of this application is to provide an experimental verification and evaluation method and related apparatus based on AI intelligent agents, which can quickly obtain evaluation results and support leadership decision-making.
[0005] To achieve the above objectives, this application provides the following solution:
[0006] Firstly, this application provides an experimental verification and evaluation method based on AI intelligent agents, including:
[0007] Obtain the test evaluation documents; these are the documents used in the test evaluation.
[0008] Based on the parent-child segmentation pattern, the test evaluation document is structured into segments to obtain a document with a two-level segmentation structure; the document with the two-level segmentation structure includes a parent block and a child block; the parent block is used to provide context, and the child block is used for precise retrieval;
[0009] The content in the sub-blocks is vectorized to generate a vectorized knowledge base for the test evaluation documents;
[0010] Based on data extraction rules, a data extraction agent is invoked, using the database of the experimental data management software as the data source, and a large language model is used to extract experimental plan information, experimental implementation information, and experimental result information from the experimental evaluation documents.
[0011] Based on the extracted test plan information, test implementation information and test result information, the intelligent agent for test verification and evaluation task planning is used to vectorize user evaluation requirements. Based on the vectorized knowledge base, the semantic similarity of the vectorized user evaluation requirements is calculated to obtain the evaluation algorithm description and test cost accounting standard related to the user evaluation requirements.
[0012] Based on the evaluation algorithm description and the test cost accounting standard, the test verification sufficiency evaluation agent, the test verification progress evaluation agent, and the test cost evaluation agent are used to evaluate the sufficiency of the test verification process, the test progress, and the cost of the test project, and obtain the test verification evaluation results.
[0013] Optionally, in the parent-child segmentation mode, the parent block is a text unit containing complete semantics, and the child block is a subdivided paragraph of the parent block; in the parent-child segmentation mode, the child block is matched first during retrieval, and then the parent block is associated to supplement the context.
[0014] Optionally, the data extraction rules of the data extraction agent include:
[0015] Periodically extract the project names, verification indicator requirements, and project duration from the test plan table;
[0016] Real-time sampling of equipment usage time and consumable consumption in the test implementation table;
[0017] Dynamically capture the item status and completion time in the test results table.
[0018] Optionally, the test verification and evaluation results are presented in the form of bar charts or line graphs.
[0019] Optionally, the experimental verification adequacy assessment agent is used to perform experimental verification adequacy assessment based on the input provided by the experimental verification assessment task planning agent and a large language model, and to feed back the adequacy assessment results to the experimental verification assessment task planning agent.
[0020] Optionally, the experimental verification progress evaluation agent is used to evaluate the experimental verification progress based on the input provided by the experimental verification evaluation task planning agent and a large language model, and to feed back the experimental verification progress evaluation results to the experimental verification evaluation task planning agent.
[0021] Optionally, the test cost assessment agent is used to assess the test cost based on the input provided by the test verification assessment task planning agent and a large language model, and to feed back the cost assessment results of the test project to the test verification assessment task planning agent.
[0022] Secondly, this application provides an experimental verification and evaluation device based on an AI intelligent agent, comprising:
[0023] The data acquisition module is used to acquire test evaluation documents; the test evaluation documents are the documents used for test evaluation.
[0024] The segmentation module is used to structurally segment the test evaluation document based on a parent-child segmentation pattern, resulting in a document with a two-level segmentation structure. The document with the two-level segmentation structure includes a parent block and a child block. The parent block is used to provide context, and the child block is used for precise retrieval.
[0025] The vectorization processing module is used to vectorize the content in the sub-blocks and generate a vectorized knowledge base for the test evaluation documents.
[0026] The data extraction module is used to call the data extraction agent based on the data extraction rules, using the database of the experimental data management software as the data source, and using a large language model to extract experimental plan information, experimental implementation information and experimental result information from the experimental evaluation documents;
[0027] The similarity calculation module is used to plan an intelligent agent based on the extracted test plan information, test implementation information and test result information, vectorize the user evaluation requirements based on the test verification and evaluation task planning agent, and perform semantic similarity calculation on the vectorized user evaluation requirements based on the vectorized knowledge base, so as to obtain the evaluation algorithm description and test cost accounting standard related to the user evaluation requirements.
[0028] The evaluation module is used to evaluate the adequacy of the test verification process, the test progress, and the cost of the test project based on the evaluation algorithm description and test cost accounting standards, using the test verification adequacy evaluation agent, test verification progress evaluation agent, and test cost evaluation agent, and obtain the test verification evaluation results.
[0029] Thirdly, this application provides a computer device, including: a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the experimental verification and evaluation method based on an AI agent as described above.
[0030] Fourthly, this application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the experimental verification and evaluation method based on an AI agent as described above.
[0031] According to the specific embodiments provided in this application, the following technical effects are disclosed:
[0032] This application provides an AI-based experimental verification and evaluation method and related apparatus. First, the experimental evaluation document is structurally segmented using a parent-child segmentation model, forming a two-layer segmentation structure including parent blocks and child blocks. This structured processing makes the document content clearer and more orderly, facilitating subsequent information extraction and processing. The parent block provides contextual information, helping to understand the background and purpose of the entire experiment; while the child blocks are used for precise retrieval, enabling rapid location of key information and improving processing efficiency. Second, the method vectorizes the content in the child blocks, generating a vectorized knowledge base for the experimental evaluation document. Vectorization converts textual information into numerical form, facilitating efficient computation and storage by the computer. Simultaneously, the establishment of the vectorized knowledge base provides a foundation for subsequent semantic similarity calculation, enabling the system to more accurately understand user evaluation needs and quickly find relevant evaluation algorithm descriptions and experimental cost accounting standards. Third, a large language model is used to extract experimental plan information, experimental implementation information, and experimental result information from the experimental evaluation document. The large language model has powerful natural language processing capabilities, accurately identifying and understanding key information in the text. By invoking a data extraction agent and using the database of the experimental data management software as the data source, the system can quickly acquire the necessary experimental data, providing strong support for subsequent evaluation. Finally, evaluation is conducted using multiple agents, including those for experimental verification and evaluation task planning, experimental verification adequacy evaluation, experimental verification progress evaluation, and experimental cost evaluation. These agents are responsible for different evaluation tasks and can fully utilize their respective expertise and algorithmic models to comprehensively and accurately evaluate the experimental verification process. By integrating the evaluation results of multiple agents, the system can quickly obtain the experimental verification evaluation results, providing timely and reliable data support for leadership decision-making. Attached Figure Description
[0033] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0034] Figure 1This is an application environment diagram of an AI-based experimental verification and evaluation method according to an embodiment of this application;
[0035] Figure 2 A flowchart illustrating an experimental verification and evaluation method based on an AI agent, provided as an embodiment of this application;
[0036] Figure 3 A functional module diagram of an AI agent-based experimental verification and evaluation method provided in an embodiment of this application;
[0037] Figure 4 This is a schematic diagram of the structure of a computer device provided in an embodiment of this application. Detailed Implementation
[0038] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0039] Currently, the most common method for evaluating test verification is manual assessment. Evaluators need to meticulously review numerous Word and PDF documents provided by the design and testing departments, filtering out data relevant to the test verification evaluation. For example, when evaluating the test verification of a new type of aviation equipment, evaluators must sift through hundreds or even thousands of pages of design documents and test reports, manually extracting key information such as equipment performance indicators, test process records, and test data. Then, based on their professional knowledge and experience, they analyze and judge this data to arrive at the evaluation conclusion of the test verification.
[0040] However, the disadvantages of manual evaluation are:
[0041] 1) Extremely low efficiency: Manually reviewing and screening large amounts of document data requires a significant amount of time and effort. Faced with the ever-increasing volume of experimental data, the evaluation cycle is greatly extended. Taking large-scale equipment testing as an example, manual evaluation may take several months or even more than half a year, seriously affecting the progress of equipment development and preventing the equipment from being put into use in a timely manner or undergoing subsequent improvements.
[0042] 2) Difficulty in ensuring assessment accuracy: Assessment results are highly dependent on the professional competence and experience of the assessors. Different assessors may have different understandings and judgments of the data, which leads to a lack of consistency and accuracy in the assessment results. For example, when assessing a certain performance indicator of the same equipment, different assessors may give completely different assessment conclusions due to differences in their understanding of relevant standards or differences in personal experience, thus affecting the judgment of the overall performance of the equipment.
[0043] 3) High labor costs: Manual evaluation requires a large number of professional evaluators, which involves not only recruitment and training costs, but also significant time and manpower costs during the evaluation process. As the scale of equipment development continues to expand, the cost of manual evaluation becomes increasingly prominent, placing a heavy economic burden on enterprises and research institutions.
[0044] Currently, some enterprises and research institutions are beginning to experiment with simple data processing software to assist in experimental verification and evaluation. This software typically possesses basic data import, organization, and simple analysis functions. For example, it can initially extract data from Word or PDF documents and convert it into tabular format, making it easier for evaluators to view and conduct preliminary analysis. Evaluators can utilize the software's filtering and sorting functions to process the data to a certain extent, thereby assisting in the evaluation process.
[0045] However, the drawbacks of using simple data processing software to assist in experimental verification and evaluation are as follows:
[0046] 1) Significant Functional Limitations: Most of these software programs can only perform simple data processing and analysis, failing to delve into the intrinsic connections and potential value between data. They struggle to provide effective support for complex assessments of completeness, sufficiency, consistency, and factors such as test progress and cost during experimental verification and evaluation. For example, when evaluating the experimental verification of multi-system collaborative operation of equipment, the software cannot comprehensively analyze and evaluate data from multiple systems, nor can it identify interaction problems between systems.
[0047] 2) Weak data fusion capabilities: Faced with large amounts of data from different departments and in different formats, these software programs perform poorly in data fusion. They struggle to efficiently integrate design data, experimental data, etc., resulting in the inability to effectively reflect the correlation between data, thus affecting the comprehensiveness and accuracy of the evaluation.
[0048] 3) Reliance on manual intervention: During software use, a significant amount of manual work is still required for data preprocessing, interpretation of analysis results, and other tasks. This does not fundamentally solve the problems of low efficiency and high cost of manual assessment; it only reduces the workload to a certain extent and cannot achieve automation and intelligence in the assessment process.
[0049] The purpose of this application is to provide an experimental verification and evaluation method and related apparatus based on AI intelligent agents, which can quickly obtain evaluation results and support leadership decision-making.
[0050] To make the above-mentioned objectives, features and advantages of this application more apparent and understandable, the application will be further described in detail below with reference to the accompanying drawings and specific embodiments.
[0051] The experimental verification and evaluation method based on AI intelligent agents provided in this application can be applied to, for example... Figure 1 In the application environment shown, terminal 102 communicates with server 104 via a network. A data storage system can store the data that server 104 needs to process. The data storage system can be set up independently, integrated into server 104, or placed in the cloud or on other servers. Terminal 102 can send the test evaluation document to be processed to server 104. After receiving the test evaluation document, server 104 performs structured segmentation of the test evaluation document based on a parent-child segmentation pattern, resulting in a document with a two-layer segmentation structure. The document includes a parent block and a child block; the parent block provides context, and the child block is used for precise retrieval. The content in the child block is vectorized to generate a vectorized knowledge base for the test evaluation document. Based on data extraction rules, a data extraction agent is invoked, using the database of the test data management software as the data source, and a large language model is used to extract data from the test evaluation document. The system retrieves test plan information, test implementation information, and test result information. Based on the extracted test plan information, test implementation information, and test result information, an intelligent agent is planned based on the test verification and evaluation task. User evaluation requirements are vectorized, and semantic similarity is calculated on the vectorized user evaluation requirements based on a vectorized knowledge base to obtain evaluation algorithm descriptions and test cost accounting standards related to user evaluation requirements. Based on the evaluation algorithm descriptions and test cost accounting standards, an intelligent agent for test verification sufficiency evaluation, test verification progress evaluation, and test cost evaluation is used to evaluate the sufficiency of the test verification process, the test progress, and the cost incurred by the test project, thus obtaining the test verification evaluation results. The server 104 can feed back the obtained test verification evaluation results to the terminal 102. In addition, in some embodiments, the test verification and evaluation method based on AI intelligent agents can also be implemented by the server 104 or the terminal 102 alone. For example, the terminal 102 can directly process the test evaluation document to be processed, or the server 104 can obtain the test evaluation document to be processed from the data storage system and process it.
[0052] The terminal 102 can be, but is not limited to, various desktop computers, laptops, smartphones, tablets, IoT devices, and portable wearable devices. IoT devices can include smart speakers, smart TVs, smart air conditioners, and smart in-vehicle devices. Portable wearable devices can include smartwatches, smart bracelets, and head-mounted devices. The server 104 can be implemented using a standalone server or a server cluster composed of multiple servers, or it can be a cloud server.
[0053] In one exemplary embodiment, such as Figure 2 As shown, an experimental verification and evaluation method based on AI intelligent agents is provided. This method is executed by computer devices, specifically by a terminal or server alone, or by both a terminal and a server. In this embodiment, the method is applied to... Figure 1 Taking server 104 as an example, the explanation includes the following steps 201 to 206. Wherein:
[0054] Step 201: Obtain the test evaluation document; the test evaluation document is the document used for the test evaluation.
[0055] Step 202: Based on the parent-child segmentation pattern, the test evaluation document is structurally segmented to obtain a document with a two-level segmentation structure; the document with the two-level segmentation structure includes a parent block and a child block; the parent block is used to provide context, and the child block is used for precise retrieval;
[0056] Step 203: Vectorize the content in the sub-blocks to generate a vectorized knowledge base for the test evaluation document;
[0057] Step 204: Based on the data extraction rules, call the data extraction agent, using the database of the experimental data management software as the data source, and use the large language model to extract experimental plan information, experimental implementation information and experimental result information from the experimental evaluation document;
[0058] Step 205: Based on the extracted test plan information, test implementation information and test result information, the intelligent agent for test verification and evaluation task planning is used to vectorize the user evaluation requirements, and based on the vectorized knowledge base, the semantic similarity of the vectorized user evaluation requirements is calculated to obtain the evaluation algorithm description and test cost accounting standard related to the user evaluation requirements.
[0059] Step 206: Based on the evaluation algorithm description and the test cost accounting standard, and using the test verification sufficiency evaluation agent, test verification progress evaluation agent, and test cost evaluation agent, evaluate the sufficiency of the test verification process, the test progress, and the cost of the test project to obtain the test verification evaluation result.
[0060] In one exemplary embodiment, when performing steps 201-206, as follows: Figure 3 As shown, the specific details are as follows:
[0061] Specifically, in the knowledge preparation and vectorization stage: the documents used for experimental evaluation, such as the Word format evaluation algorithm description and experimental cost accounting standards, are segmented, and embedded models are used to vectorize these documents, facilitating accurate searching and rapid processing by the large language model. This can be done as follows:
[0062] 1) Document segmentation:
[0063] Each document is segmented using a two-tiered segmentation structure, known as a "parent-child segmentation pattern," to balance retrieval accuracy and contextual information. First, the document is divided into multiple larger paragraphs (parent blocks) to provide rich contextual information. Then, each larger text unit is divided into smaller paragraphs (child blocks) for precise retrieval. The system first performs precise retrieval through the child blocks to ensure relevance, and then retrieves the corresponding parent blocks to supplement the contextual information. This ensures both accuracy and complete background information in the generated output.
[0064] 2) Document vectorization:
[0065] After dividing the document into words or phrases, removing irrelevant words, extracting stems, and restoring part-of-speech tags, a vector representation model is selected to transform the content of the document sub-blocks into vectors, which can be represented as K = [K1, K2, ..., K]. n ], where n represents the dimension of the feature space, and each K i (i = 1, 2, ..., n) are values in one dimension of the feature space, representing the specific value of the feature at that vector or point.
[0066] Specifically, in the data extraction phase based on the AI agent, the timing of data extraction and the rules for the content to be extracted are determined, and data for experimental verification and evaluation is extracted from the database periodically based on these rules. This can be done as follows:
[0067] 1) Rule settings:
[0068] Rules are set in the data extraction agent, including rules for scheduled extraction time and extraction content. The extraction content rules mainly include: extracting the planned test project name, verification indicator requirements, and test duration from the test plan table; extracting the implemented test project name, verification indicators, test start time, test equipment usage time, and consumable consumption from the test implementation table; and extracting test project status and completion time from the test results table.
[0069] 2) Data extraction:
[0070] Based on data extraction rules, a data extraction agent is invoked, using the database of the experimental data management software as the data source, and a large language model is used to extract experimental plan information, experimental implementation information, and experimental result information from structured or unstructured data.
[0071] Specifically, in the experimental verification and evaluation phase based on AI agents, the experimental verification and evaluation task planning agent first plans the task according to user needs. Then, it calls upon specific functional agents to evaluate the sufficiency, progress, and cost rationality of the experimental verification based on data from the experimental verification and evaluation database and knowledge matched from the knowledge base. The evaluation results are then fed back to the experimental verification and evaluation task planning agent. Specifically, this can be done as follows:
[0072] 1) Knowledge matching:
[0073] The user evaluation requirements are vectorized, and then semantic similarity is calculated with paragraphs in the vectorized knowledge base. The top_K paragraphs with the highest semantic similarity are matched to obtain knowledge such as evaluation algorithm descriptions and test cost accounting standards related to user evaluation requirements.
[0074] 2) Assessment of the adequacy of experimental verification:
[0075] The experiment verification adequacy assessment agent is invoked. Based on the relevant assessment algorithm descriptions matched from the knowledge base, all the actual implemented experiment items and verification indicators are compared with the planned experiment items and verification indicators. The adequacy of the experiment verification process is assessed using a large language model.
[0076] 3) Assessment of experimental verification progress:
[0077] The test verification progress assessment agent is invoked. Based on the assessment algorithm description matched by the knowledge base, the status and completion time of each test project are compared with the planned completion time of each test project. The large language model is used to assess whether the test progress is lagging behind or overdue.
[0078] 4) Trial cost assessment:
[0079] The test cost assessment agent is invoked to assess whether the cost of each test project exceeds the budget based on the completion time of each test project, the usage time of the test equipment, the consumption of consumables, and the budget data of the test project, and the assessment algorithm description and test cost accounting standards matched by the knowledge base.
[0080] In the results presentation and report generation stage, the experimental verification and evaluation results data are displayed in the form of bar charts, line charts, etc. A report template is then used to integrate the experimental verification and evaluation results and graphical data to generate corresponding experimental verification and evaluation reports, including an experimental verification adequacy evaluation report, an experimental verification progress evaluation report, and an experimental cost evaluation report. Specifically, this can be done as follows:
[0081] 1) Results Display:
[0082] Based on intuitive and easy-to-understand visualization controls, the evaluation results are presented to users in the form of bar charts, line charts, etc. For example, a bar chart can show the completion status of various test indicators, and a line chart can show the reasonableness of the test verification cost.
[0083] 2) Report generation:
[0084] Based on the test evaluation results and graphical data, the system calls the corresponding test verification evaluation report template to generate a test verification evaluation report containing text, data, tables, and images, providing users with real-time feedback on the current test verification evaluation results.
[0085] This application also provides an application scenario in which the above-described AI-based intelligent agent-based test verification and evaluation method is applied. Specifically, the AI-based intelligent agent-based test verification and evaluation method provided in this embodiment can be applied in the quality control process of various industrial products. For example, in the automotive manufacturing field, this method can be used to test and verify key indicators such as the durability and safety of automotive parts, quickly generating detailed evaluation reports to help manufacturers promptly identify and resolve potential quality problems. In the electronics manufacturing field, this method can also be used to comprehensively evaluate the performance and stability of electronic products, ensuring that the products meet high-standard quality requirements. In addition, this method can also be applied to multiple industries such as food processing and pharmaceutical manufacturing, providing strong technical support for the quality control of industrial products.
[0086] Based on the same inventive concept, this application also provides an experimental verification and evaluation apparatus for implementing the above-described experimental verification and evaluation method based on AI intelligent agents. The solution provided by this apparatus is similar to the implementation scheme described in the above method; therefore, the specific limitations in one or more apparatus embodiments provided below can be found in the limitations of the experimental verification and evaluation method based on AI intelligent agents described above, and will not be repeated here.
[0087] In one exemplary embodiment, an experimental verification and evaluation device based on an AI agent is provided, comprising:
[0088] The data acquisition module is used to acquire test evaluation documents; the test evaluation documents are the documents used for test evaluation.
[0089] The segmentation module is used to structurally segment the test evaluation document based on a parent-child segmentation pattern, resulting in a document with a two-level segmentation structure. The document with the two-level segmentation structure includes a parent block and a child block. The parent block is used to provide context, and the child block is used for precise retrieval.
[0090] The vectorization processing module is used to vectorize the content in the sub-blocks and generate a vectorized knowledge base for the test evaluation documents.
[0091] The data extraction module is used to call the data extraction agent based on the data extraction rules, using the database of the experimental data management software as the data source, and using a large language model to extract experimental plan information, experimental implementation information and experimental result information from the experimental evaluation documents;
[0092] The similarity calculation module is used to plan an intelligent agent based on the extracted test plan information, test implementation information and test result information, vectorize the user evaluation requirements based on the test verification and evaluation task planning agent, and perform semantic similarity calculation on the vectorized user evaluation requirements based on the vectorized knowledge base, so as to obtain the evaluation algorithm description and test cost accounting standard related to the user evaluation requirements.
[0093] The evaluation module is used to evaluate the adequacy of the test verification process, the test progress, and the cost of the test project based on the evaluation algorithm description and test cost accounting standards, using the test verification adequacy evaluation agent, test verification progress evaluation agent, and test cost evaluation agent, and obtain the test verification evaluation results.
[0094] In one exemplary embodiment, a computer device is provided, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 4As shown, this computer device includes a processor, memory, input / output interfaces (I / O), and communication interfaces. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. It has sufficient AI computing power to stably run large language models such as Deepseek-r1-32b and qwen2.5:32b. It possesses sufficient data processing capabilities to stably run the evaluation system, ensuring its efficiency and reliability. It has a large-capacity storage device for storing experimental verification data and evaluation results, facilitating data management and retrieval. A network interface is configured to enable data interaction with external devices, ensuring timely data acquisition and rapid result transmission. The computer device's memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides the environment for the operating system and computer programs in the non-volatile storage media to run. The database is used to store data. The I / O interfaces are used for information exchange between the processor and external devices. The communication interface of this computer device is used to communicate with external terminals via a network. When the computer program is executed by the processor, it implements an experimental verification and evaluation method based on an AI agent.
[0095] Those skilled in the art will understand that Figure 4 The structure shown is merely a block diagram of a portion of the structure related to the present application and does not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than those shown in the figure, or combine certain components, or have different component arrangements.
[0096] In one exemplary embodiment, a computer device is also provided, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.
[0097] In one exemplary embodiment, a computer-readable storage medium is provided storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.
[0098] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.
[0099] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).
[0100] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic devices, quantum computing-based data processing logic devices, etc., and are not limited to these.
[0101] In summary, this application has the following technical effects:
[0102] 1) Significantly improved efficiency: Compared with traditional manual evaluation and simple software-assisted evaluation, this invention utilizes AI intelligent agents and automated processes to process a large amount of data in a short time, shortening the evaluation cycle by several times or even dozens of times, greatly accelerating the progress of equipment development, and enabling the equipment to be put into use or undergo subsequent optimization more quickly.
[0103] 2) Significantly improved evaluation quality: By leveraging large language models and scientific evaluation models, this invention achieves more objective and accurate evaluations. It avoids the subjective biases of manual evaluations and the limitations of simple software evaluations, enabling in-depth data analysis to uncover potential problems and provide more targeted suggestions for improving equipment testing and verification designs.
[0104] 3) Effective cost reduction: It reduces reliance on a large amount of manual labor, thus lowering labor costs. At the same time, by improving evaluation efficiency, it shortens the evaluation cycle, reduces time costs, and improves resource utilization efficiency, saving enterprises and research institutions a significant amount of money.
[0105] 4) Strong decision support: The rapid and accurate evaluation results provide leaders with clear and comprehensive information on equipment testing and verification, helping them to make timely and scientific decisions, avoid decision-making errors caused by insufficient or inaccurate information, and promote the smooth progress of equipment development.
[0106] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.
[0107] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.
Claims
1. A method for experimental verification and evaluation based on AI intelligent agents, characterized in that, include: Obtain the test evaluation documents; these are the documents used in the test evaluation. Based on the parent-child segmentation pattern, the experimental evaluation document is structurally segmented to obtain a document with a two-level segmentation structure. The document with the two-level segmentation structure includes a parent block and a child block. The parent block is used to provide context, and the child block is used for precise retrieval. In the parent-child segmentation pattern, the parent block is a text unit containing complete semantics, and the child block is a subdivision of the parent block. In the parent-child segmentation pattern, the child block is matched first during retrieval, and then the parent block is associated to supplement the context. The content in the sub-blocks is vectorized to generate a vectorized knowledge base for the test evaluation documents; Based on data extraction rules, a data extraction agent is invoked, using the database of the experimental data management software as the data source, and a large language model is used to extract experimental plan information, experimental implementation information, and experimental result information from the experimental evaluation documents. Based on the extracted test plan information, test implementation information and test result information, the intelligent agent for test verification and evaluation task planning is used to vectorize user evaluation requirements. Based on the vectorized knowledge base, the semantic similarity of the vectorized user evaluation requirements is calculated to obtain the evaluation algorithm description and test cost accounting standard related to the user evaluation requirements. Based on the evaluation algorithm description and the test cost accounting standard, the test verification sufficiency evaluation agent, the test verification progress evaluation agent, and the test cost evaluation agent are used to evaluate the sufficiency of the test verification process, the test progress, and the cost of the test project, and obtain the test verification evaluation results. The experimental verification adequacy assessment agent is used to perform experimental verification adequacy assessment based on the input provided by the experimental verification assessment task planning agent and the large language model, and to feed back the adequacy assessment results to the experimental verification assessment task planning agent. The experimental verification progress evaluation agent is used to evaluate the experimental verification progress based on the input provided by the experimental verification evaluation task planning agent and the large language model, and to feed back the experimental verification progress evaluation result to the experimental verification evaluation task planning agent. The test cost assessment agent is used to assess the test cost based on the input provided by the test verification and assessment task planning agent and the large language model, and to feed back the cost assessment results of the test project to the test verification and assessment task planning agent.
2. The experimental verification and evaluation method based on AI intelligent agents according to claim 1, characterized in that, The data extraction rules of the data extraction agent include: Periodically extract the project names, verification indicator requirements, and project duration from the test plan table; Real-time sampling of equipment usage time and consumable consumption in the test implementation table; Dynamically capture the item status and completion time in the test results table.
3. The experimental verification and evaluation method based on AI intelligent agents according to claim 1, characterized in that, The experimental verification and evaluation results are presented in the form of bar charts or line graphs.
4. An experimental verification and evaluation device based on an AI agent, used to implement the experimental verification and evaluation method based on an AI agent as described in claim 1, characterized in that, include: The data acquisition module is used to acquire test evaluation documents; the test evaluation documents are the documents used for test evaluation. The segmentation module is used to structurally segment the test evaluation document based on a parent-child segmentation pattern, resulting in a document with a two-level segmentation structure. The document with the two-level segmentation structure includes a parent block and a child block. The parent block is used to provide context, and the child block is used for precise retrieval. The vectorization processing module is used to vectorize the content in the sub-blocks and generate a vectorized knowledge base for the test evaluation documents. The data extraction module is used to call the data extraction agent based on the data extraction rules, using the database of the experimental data management software as the data source, and using a large language model to extract experimental plan information, experimental implementation information and experimental result information from the experimental evaluation documents; The similarity calculation module is used to plan an intelligent agent based on the extracted test plan information, test implementation information and test result information, vectorize the user evaluation requirements based on the test verification and evaluation task planning agent, and perform semantic similarity calculation on the vectorized user evaluation requirements based on the vectorized knowledge base, so as to obtain the evaluation algorithm description and test cost accounting standard related to the user evaluation requirements. The evaluation module is used to evaluate the adequacy of the test verification process, the test progress, and the cost of the test project based on the evaluation algorithm description and test cost accounting standards, using the test verification adequacy evaluation agent, test verification progress evaluation agent, and test cost evaluation agent, and obtain the test verification evaluation results.
5. A computer device, comprising: A memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that the processor executes the computer program to implement the experimental verification and evaluation method based on an AI agent according to any one of claims 1-3.
6. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the experimental verification and evaluation method based on an AI agent as described in any one of claims 1-3.