Test analysis and report generation method and system based on large model retrieval enhancement

By introducing large-scale model retrieval enhancement technology, the problem of insufficient flexibility of traditional test data analysis methods is solved, efficient and intelligent test report generation in the field of electronic test measurement is realized, and the accuracy and efficiency of data analysis is improved. It is suitable for technical scenarios with high accuracy and professionalism.

CN120579528APending Publication Date: 2025-09-02CHINA ELECTRONIS TECH INSTR CO LTD
View PDF 0 Cites 17 Cited by

Patent Information

Application Number
CN202510719864.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-09-02

AI Technical Summary

Technical Problem

In the field of electronic testing and measurement, traditional test data analysis and report generation methods are insufficiently flexible and cannot meet the diverse and personalized test needs, resulting in poor template reusability, high labor costs, long development cycles, and difficulty in effectively processing multi-source heterogeneous and real-time dynamic test data, affecting data utilization and analysis efficiency.

Method used

Using a method based on large-scale retrieval enhancement, the query optimization module, natural language generation SQL module, database query module and output result fusion module are used, and the multi-strategy retrieval enhancement generation mechanism of RAG, Self-RAG and Graph-RAG is combined to realize the supplementary knowledge of user queries and the intelligence of data analysis, and generate accurate test reports.

Benefits of technology

It improves the accuracy and intelligence level of data analysis, reduces manual participation, reduces enterprise operation costs, supports multimodal knowledge fusion and the construction of complex knowledge graphs, and has good scalability and domain migration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120579528A_ABST
    Figure CN120579528A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of automatic report generation, and provides a test analysis and report generation method and system based on large model retrieval enhancement, and the method comprises the steps: firstly carrying out the semantic task decomposition of user query, and meanwhile, achieving the multi-dimensional information extraction through the butt joint of a knowledge vector library and a structured image-text knowledge system established by a private knowledge base module. And three strategies of RAG, Self-RAG and Graph-RAG are combined to enhance the generation capability of the large model so as to accurately obtain background knowledge. And constructing a standard SQL statement according to a structured query requirement, and performing data filling according to template prompt by an output result fusion module in combination with a query background and an SQL execution result to form a test report. By optimizing a retrieval enhancement generation method, professional knowledge supplement related to query is realized, and the reasoning ability of a large model in a professional scene is effectively enhanced, so that the accuracy and the intelligent level of data analysis are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field related to automatic report generation, and more specifically, to a test analysis and report generation method and system based on large model retrieval enhancement. Background Art

[0002] The statements in this section merely provide background information related to the present disclosure and do not necessarily constitute prior art.

[0003] Another key future development direction for electronic measuring instrumentation is the intelligent analysis of test and measurement data. Test data analysis and report generation are key components for visualizing test results, evaluating performance metrics, and supporting decision-making. Traditional test data analysis and report generation technologies generally rely on fixed, manually maintained templates and rules. This approach lacks flexibility when faced with ever-changing test tasks and DUTs. Especially with the growing demand for diverse and personalized testing, existing methods struggle to meet the demands for automated and intelligent upgrades in electronic equipment testing.

[0004] The mainstream approach to intelligent analysis of test measurement data relies primarily on testers designing report templates based on test tasks. Program developers then implement data collection, preprocessing, statistical analysis, and other processes based on these templates to ultimately generate test reports. This human-driven model presents the following major issues: First, template design is complex and has a limited scope of application, resulting in poor reusability and versatility. Second, the implementation of analysis tasks relies heavily on specialized personnel, resulting in high labor costs, long development cycles, and low efficiency. Furthermore, faced with multi-source, heterogeneous, and real-time dynamic test data, traditional static analysis methods suffer from significant shortcomings in processing power and response efficiency, resulting in low data utilization and inability to deeply mine and unlock the value of test data.

[0005] With the continuous innovation of electronic technology and the integration, networking, and digitalization of automatic test equipment, the scale and complexity of test data have increased significantly. How to effectively utilize these massive data resources and improve the automation and intelligence of data analysis has become one of the core challenges in the current development of electronic measurement equipment. In recent years, artificial intelligence-generated content (AIGC) technology has developed rapidly, thanks to the increase in hardware computing power and the evolution of deep learning, especially the Transformer model architecture. Transformer-based large language models (LLMs) have powerful language understanding, information extraction, reasoning analysis, and content generation capabilities, and are becoming an important tool for promoting the intelligent upgrade of traditional analysis processes. In various fields such as finance, manufacturing, and retail, LLMs have shown broad application prospects, such as providing efficient solutions in scenarios such as intelligent recommendation, automatic report generation, and decision support.

[0006] During their research, the inventors discovered that existing LLMS still exhibit significant limitations when dealing with data close to daily life, such as finance, accounting, and user analysis, and when processing specific, highly specialized problems in the field of electronic component testing and measurement. Furthermore, in intelligent report generation methods based on large models in other fields, general models are unable to understand the meaning and context of data in specific fields, resulting in an inability to accurately process and interpret data, affecting the performance and reliability of the model. Summary of the Invention

[0007] To address the aforementioned issues, this paper proposes a test analysis and report generation method and system based on large-scale model retrieval enhancement, which significantly improves data analysis efficiency, reduces labor costs, and exhibits good scenario adaptability. By optimizing the retrieval enhancement generation method, this method supplements query-related professional knowledge in the field of electronic test and measurement, effectively enhancing the reasoning capabilities of large models in professional scenarios, thereby improving the accuracy and intelligence of data analysis.

[0008] In order to achieve the above objectives, the present disclosure adopts the following technical solutions:

[0009] One or more embodiments provide a test analysis and report generation system based on large model retrieval enhancement, including:

[0010] The query optimization module is used to split tasks based on the acquired user queries, extract test domain expertise from the constructed knowledge vector library, and use a multi-strategy retrieval enhancement generation mechanism that integrates RAG, Self-RAG, and Graph-RAG to process the extracted relevant knowledge and user queries to obtain enhanced sub-query tasks.

[0011] The natural language generation SQL module is used to generate SQL query statements using the large model based on the enhanced sub-query tasks;

[0012] Database query module, used to execute SQL query based on SQL query statement and extract target test data;

[0013] The output result fusion module is used to combine the target test data and the template background information according to the relevance, fill the template, and obtain the test report.

[0014] One or more embodiments provide a test analysis and report generation method based on large model retrieval enhancement, comprising the following steps:

[0015] Tasks are split based on the acquired user queries. Test domain expertise is extracted from the constructed knowledge vector library. A multi-strategy retrieval enhancement generation mechanism that integrates RAG, Self-RAG, and Graph-RAG is used to process the extracted relevant knowledge and user queries to obtain enhanced sub-query tasks.

[0016] Based on the enhanced sub-query task, the large model is used to generate SQL query statements;

[0017] Execute SQL queries based on SQL query statements to extract target test data;

[0018] Combine the target test data and template background information according to the relevance, fill in the template, and obtain a test report.

[0019] An electronic device includes a memory and a processor, and computer instructions stored in the memory and executed on the processor. When the computer instructions are executed by the processor, the steps in the above-mentioned test analysis and report generation method based on large model retrieval enhancement are completed.

[0020] Compared with the prior art, the present invention has the following beneficial effects:

[0021] This disclosure introduces a multi-strategy retrieval enhancement mechanism to effectively supplement and clarify user query intent, improving the accuracy and intelligence of the report generation process. This system can reduce manual intervention, improve analysis efficiency, and lower enterprise operating costs, making it particularly suitable for technical scenarios requiring high accuracy and professionalism. Furthermore, this method supports multimodal knowledge fusion and the establishment of graphic and text structures, enhancing the system's ability to construct complex knowledge graphs and demonstrating good scalability and domain transferability.

[0022] The advantages of the present disclosure and additional advantages will be described in detail in the following specific embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0023] The accompanying drawings, which constitute a part of the present disclosure, are used to provide a further understanding of the present disclosure. The exemplary embodiments of the present disclosure and their descriptions are used to explain the present disclosure but do not constitute a limitation of the present disclosure.

[0024] Figure 1 is a block diagram of a test analysis and report generation system based on large model retrieval enhancement according to embodiment 1 of the present disclosure;

[0025] Figure 2 is a flow chart of a test analysis and report generation method based on large model retrieval enhancement according to embodiment 1 of the present disclosure; DETAILED DESCRIPTION

[0026] The present disclosure will be further described below with reference to the accompanying drawings and embodiments.

[0027] It should be noted that the following detailed descriptions are exemplary and intended to provide further explanation of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those skilled in the art to which the present disclosure belongs.

[0028] It should be noted that the terms used herein are only for the purpose of describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "including" are used in this specification, they indicate the presence of features, steps, operations, devices, components and / or combinations thereof. It should be noted that, in the absence of conflict, the various embodiments in the present disclosure and the features in the embodiments can be combined with each other. The embodiments will be described in detail below with reference to the accompanying drawings.

[0029] Example 1

[0030] In the technical solutions disclosed in one or more embodiments, Figure 1 As shown, a test analysis and report generation system based on large model retrieval enhancement includes: a query parsing module, a private knowledge base module, an external knowledge retrieval enhancement module, a query optimization module, a natural language SQL generation module, a database query execution module, an output result fusion module and a chart rendering module.

[0031] The query parsing module is used to parse and semantically analyze the acquired user queries, extract entities and keywords in the queries, and determine the keywords and entity information used to represent the user's intent;

[0032] A private knowledge base construction module is used to integrate OCR image and text recognition with natural language processing methods to extract text information from test and measurement-related knowledge document resources and construct image and text structural relationships;

[0033] The query optimization module is used to split tasks based on acquired user queries, extract relevant knowledge from the constructed knowledge vector library, and process the extracted relevant knowledge and user queries using a multi-strategy retrieval augmentation generation (RAG), self-retrieval augmentation generation (Self-RAG), and graph-retrieval augmentation generation (Graph-RAG) to generate enhanced sub-query tasks.

[0034] The natural language generation SQL module is used to generate SQL query statements using the large model based on the enhanced sub-query tasks;

[0035] Database query module, used to execute SQL query based on SQL query statement and extract target test data;

[0036] The output result fusion module is used to combine the target test data and the template background information according to the correlation, fill the template, and obtain the test report;

[0037] In this embodiment, after obtaining the user's natural language input through the query parsing module, semantic analysis is performed to identify the key entities and keywords in the query, and then the user's intention is extracted. On this basis, the query optimization module decomposes the user query into semantic tasks, and at the same time connects the structured graphic knowledge system established by the private knowledge base module through the knowledge vector library to realize multi-dimensional information extraction. The retrieval enhancement module combines the three strategies of RAG, Self-RAG and Graph-RAG to enhance the generation capability of large models, ensuring the accurate acquisition of background knowledge in the highly specialized test and measurement field. Subsequently, the natural language generation SQL module constructs a standard SQL statement based on the structured query requirements, and the database query module executes it in the preset test data set to extract the required data. The output result fusion module combines the query background and the SQL execution results, fills in the data according to the template prompts, and forms a structured, semantically consistent professional test report. The chart rendering module generates visual charts based on the report data to achieve user-friendly output display.

[0038] This embodiment introduces a multi-strategy retrieval enhancement mechanism to effectively supplement and clarify user query intent, improving the accuracy and intelligence of the report generation process. This system can reduce manual intervention, improve analysis efficiency, and lower enterprise operating costs, making it particularly suitable for technical scenarios requiring high accuracy and professionalism. Furthermore, this method supports multimodal knowledge fusion and the establishment of graphic and text structures, enhancing the system's ability to construct complex knowledge graphs and demonstrating good scalability and domain transferability.

[0039] In some embodiments, the query parsing module is configured to perform the following process:

[0040] Step 11: Text segmentation: Perform character-level segmentation on the obtained user query;

[0041] In this embodiment, the user query can be text, which can be the text information of the data analysis requirements edited and input by the user;

[0042] Specifically, we use the pre-trained BertTokenizer provided in Transformers to perform character-level Byte Pair Encoding (BPE) segmentation to achieve fine-grained character-level token decomposition processing;

[0043] In this way, the contextual details in natural language can be effectively retained, and the model's semantic understanding ability of long and short words, compound words and professional terms (such as physical quantities, parameter symbols, etc.) can be enhanced.

[0044] Step 12: Embed the text into a vector space: Use the BGE (Bidirectional Global Embedding) model to vectorize the segmented text and map the natural language into a high-dimensional semantic vector for subsequent semantic understanding and similarity calculation.

[0045] Step 13: Entity recognition: Identify and extract entities related to the test measurement domain from the vectorized data;

[0046] Specifically, the entity may include the instrument name, test task type, geographic location information, and measurement date;

[0047] Furthermore, entity recognition can be based on the Named Entity Recognition (NER) algorithm and combined with a dictionary of professional terminology in the test and measurement field to form a NER recognition model with domain knowledge perception capabilities, so that professional entities with abbreviations such as "I_VEE" (current at the negative power supply voltage node) and "I_VDD" (current at the positive power supply voltage node) can be accurately identified.

[0048] Step 14: Entity classification: Map the identified entities to a predefined entity category system to achieve entity classification;

[0049] In this embodiment, the identified entities are mapped to a predefined entity category system and aligned with the structure of the knowledge vector library to achieve semantic unification and classification, thereby improving the module's generalization ability and semantic standard processing ability for vertical domain languages;

[0050] The query parsing module of this embodiment adopts a hybrid approach of entity-aware NER model combined with professional domain dictionaries. In vertical professional contexts such as testing and measurement, it greatly improves the recognition accuracy of complex domain terms and key entities, providing high-quality semantic input support for subsequent knowledge retrieval and report generation.

[0051] In some embodiments, a private knowledge base construction module is used to build efficient knowledge semantic retrieval capabilities for the test and measurement field, providing context-enhanced support for large models. The private knowledge base construction module is used to integrate OCR image and text recognition with natural language processing methods to extract text information from test and measurement-related knowledge document resources and construct image and text structural relationships. It is configured to perform the following steps:

[0052] Step 21: Knowledge Collection: Collect knowledge document resources related to the test and measurement field, including but not limited to user manuals, development guides, and historical test data analysis reports of test instruments;

[0053] Step 22: Optical Character Recognition: Perform character recognition on the collected knowledge document resources, and perform semantic classification and integration;

[0054] For unstructured data such as PDF documents and scanned images, we use advanced OCR recognition technology, combined with graphic layout analysis algorithms, to not only extract pure text content but also identify document structure information (such as titles, tables, and legends). Natural language processing technology is then used to perform semantic classification and content integration, unifying the data into a vectorizable knowledge content format.

[0055] This step integrates OCR and natural language understanding, not only achieving text recognition, but also preserving the structural semantics of the document and improving the depth of document parsing.

[0056] Step 23: Data cleaning: Detect and remove duplicates from the integrated characters, and then process them into a unified data format;

[0057] Specifically, the knowledge content is detected and deduplicated using MD5 hash comparison, cosine similarity algorithm and Levenshtein edit distance algorithm in turn; then, the knowledge data is formatted and processed using a combination of regular expressions and a lexical analyzer, unified into a structurally standardized data format (such as JSON, XML, etc.) to improve data consistency and operability.

[0058] Step 24: Knowledge content segmentation: Use a hierarchical recursive semantic segmentation strategy to split the long text. Specifically, divide it into natural paragraphs and then refine it into semantically complete small blocks based on preset thresholds to ensure that each knowledge block is self-consistent and contextually coherent.

[0059] This method effectively avoids semantic fragmentation and input overflow problems, and helps improve the accuracy of subsequent model reasoning.

[0060] Step 25: Knowledge semantic vectorization and index construction: Input the segmented knowledge text into the BGE model for high-dimensional semantic embedding;

[0061] Specifically, semantic vector representation is obtained, and the vector and its metadata information are stored in the high-performance Faiss vector database, an index system that supports fast semantic retrieval is constructed, and millisecond-level query capabilities for massive knowledge blocks are achieved.

[0062] Step 26: Build a domain vocabulary and perform terminology normalization mapping on the semantic vectors of high-dimensional semantic embedding;

[0063] To address the uneven distribution and diverse forms of terminology in the test and measurement field, a multi-level domain vocabulary was constructed based on corpus statistics and expert knowledge, covering common equipment names, test parameters, measurement instructions, and more. Combining morphological restoration and entity disambiguation techniques, this standardized mapping of terms in user queries improves semantic consistency and vector search hit rates.

[0064] Step 27: Knowledge graph construction: Extract entities and their semantic relationships from the mapped text vectors and construct a knowledge graph.

[0065] Based on the constructed knowledge base, named entity recognition (NER), relation extraction (RE) and dependency parsing technologies are further introduced to automatically extract entities and their semantic relationships from structured and unstructured texts, and construct a test and measurement domain knowledge graph based on "entity-relationship-entity" triples to support subsequent semantic query, causal analysis and reasoning tasks.

[0066] The private knowledge base construction module of this embodiment deeply integrates OCR image and text recognition with natural language processing technology, not only extracting text information, but also identifying image and text structural relationships (such as tables, legends, and titles), realizing the integrated conversion of document structure and semantics.

[0067] In some embodiments, a private knowledge base construction module is used to construct a knowledge vector library containing professional knowledge in the test and measurement field; a query optimization module extracts relevant knowledge from the constructed knowledge vector library, retrieves knowledge closely related to the query content through keywords and entity information in the user query, and provides contextual support information for the large model, so that the general large model can more accurately understand and handle complex problems in vertical fields, enhancing its professional reasoning ability and task response capabilities;

[0068] In some embodiments, after receiving a user's query, the query optimization module performs operations such as parsing, rewriting, and routing on the query to help the big model understand the user's intentions and needs, and provide guidance to the database query module and output result fusion module;

[0069] The query optimization module is configured to perform the following process:

[0070] Step 31: Semantic parsing and task splitting of user queries: Perform semantic analysis on the query input by the user to identify syntactic components, core intent, and hypernym and hypernym concept structures, and then divide the query into sub-query tasks;

[0071] The query statement is the data analysis requirement edited and input by the user;

[0072] In this implementation, we abandon the traditional rule-based NLP pipeline and adopt a unified, differentiable, large-scale pre-trained language model (LLM) for end-to-end semantic parsing and query task splitting. Specifically, it includes the following key components:

[0073] Optionally, a lightweight intent classification head is connected after the Transformer encoder to identify the semantics of the query statement; roles are annotated for entities in the text information of the user query;

[0074] After step 1, the encoded user query is obtained. A lightweight intent classification head is connected after the Transformer encoder to map the entire sentence to high-level intent categories such as operational queries and analytical queries.

[0075] Furthermore, a fine-tuned entity extraction and classification head can be deployed in parallel to process the entity recognition results output by steps 13 and 14, and to label the semantic roles (SRL) of entities, attributes, and relationships to obtain the "object-operation-dimension-indicator" quadruple.

[0076] (3) Perform sub-query task boundary detection and divide sub-query tasks. Specifically: For long sentences or compound expressions, a token-level delimiter predictor is used to directly output segmentation tags in the vector space to achieve automatic segmentation of atomic-level sub-user queries and obtain multiple sub-query tasks.

[0077] Furthermore, after the first round of optimized parsing, the initially generated list of subquery tasks is fed back to the same large model for secondary correction through prompting or few-shot examples. Subsequently, based on user or system execution feedback, the model's decoding strategy is fine-tuned to ensure that each subquery task matches the processing target.

[0078] According to the above key components, a subquery task and a subquery description list matching the subquery task target are obtained.

[0079] Step 32: Retrieval judgment: Use the Self-RAG (Self-reflective Retrieval-Augmented Generation) strategy to perform context completeness assessment on each subquery using small sample reasoning;

[0080] In-Context Few-Shot Prompting (IFSP) guides the large model to determine whether the current query requires external knowledge supplementation. This phase, combined with the large model's built-in LLM-based self-evaluation mechanism, automatically determines whether the current knowledge background is sufficient, effectively avoiding invalid RAG calls, reducing resource waste, and improving reasoning efficiency.

[0081] Step 33: Based on the integrity assessment results, the RAG process is used to extract relevant knowledge blocks from the constructed knowledge vector library for supplementation;

[0082] If the subquery is judged as "needing external knowledge supplementation", the standard RAG process is initiated: based on the user query and the cosine similarity calculation between the embedded vectors in the constructed knowledge vector library, candidate knowledge blocks are recalled for supplementation;

[0083] Optionally, during semantic retrieval, cosine similarity is used as the primary metric to calculate the angle between the user query vector and the knowledge chunk vector to measure their semantic relevance. Compared to Euclidean distance, cosine similarity is more robust to text length and corpus sparsity, and can more accurately match semantics, improving retrieval performance and the ability to interpret semantic context.

[0084] Step 34: Reranking: The supplemented knowledge blocks are input into the BGE-Reranker reranking model. The semantic relevance of each candidate knowledge block to the user query is evaluated through the cross-attention mechanism. The knowledge blocks with semantic relevance greater than the set threshold are retained as the reranked text blocks.

[0085] Specifically, the query-document pairs are input into the BGE-Reranker model in batches to obtain the relevance scores (cross attention scores), and the five most relevant text blocks are retained to reduce the noise input in the generation stage.

[0086] Step 35: Based on the knowledge graph in the knowledge vector library, we use Graph-RAG to extract core entities from the rearranged knowledge blocks and user query statements, and construct a structured graph path of "entity-relationship-entity" as the enhanced subquery task.

[0087] Based on the rearrangement, the Graph-RAG mechanism is introduced and combined with the pre-built test and measurement domain knowledge graph to complete the following operations:

[0088] Step 351: Identify key domain entities involved in the query;

[0089] Step 352: construct a structured graph path based on the "entity-relationship-entity" triple and generate a template prompt word;

[0090] The query optimization module of this embodiment integrates the multi-strategy retrieval enhancement generation mechanism of RAG, Self-RAG, and Graph-RAG, achieving the first RAG tri-modal collaborative optimization in the test and measurement automatic analysis scenario. Combining the reflection scoring mechanism with the semantic redundancy suppression strategy, it improves the reliability and consistency of results in multiple rounds of large language model generation. A task-granular query decomposition mechanism sequentially executes a modular processing process of deconstruction, optimization, and completion, enabling the automated decomposition and resolution of complex problems in the template construction process.

[0091] In some embodiments, the natural language SQL language generation module is configured to perform the following process:

[0092] Step 41, SQL prompt word template extraction: For each enhanced sub-query task obtained by task decomposition, semantic analysis is performed to extract keywords and entities; based on the keywords and entities, the structural information of the target table in the test database is extracted and organized into an SQL prompt word template for input to the large language model;

[0093] Specifically, the test database is a database that stores test data of the test equipment and may be a relational database.

[0094] Optionally, keywords extracted from each subquery task can include metrics, time ranges, and object attributes. This is used to locate the corresponding entity table or logical view of the entity in the test database. The database metadata is then queried to obtain the target table's structured metadata, which can include table names, field names, data types, primary and foreign key constraints, and logical relationships between fields. This structured metadata is organized into standardized prompt templates, which serve as key contextual support for SQL generation in the large model, achieving a "semantic-structural" integrated input expression.

[0095] Step 42, SQL generation: Based on the SQL prompt word template input into the large language model, generate SQL statements that conform to the database structure and query semantics;

[0096] Input the SQL prompt word template (including natural language semantics and database meta-structure information) into a large language model (such as GPT) to form a prompt word with structural context. The large language model combines natural language intent with database structure information to output a query statement that conforms to SQL syntax.

[0097] In a specific example, the prompt word may include: indicator field, filter condition, time range, logical table structure description, etc.

[0098] Furthermore, a double check is performed on the obtained SQL statement as follows:

[0099] 1) Syntax verification: Use language models or regular expressions to check whether the SQL statement conforms to the standard syntax;

[0100] 2) Semantic verification: Verify the field names, table names, and connection relationships of SQL statements to verify their authenticity and ensure their existence;

[0101] 3) Automatic repair: Check SQL statements and automatically correct them according to the schema if spelling errors, missing fields, or logical conflicts are found;

[0102] Schema correction refers to the process of automatically identifying and correcting errors made by users or language models in table names, field names, field paths, and association logic during the automatic SQL generation process, so that the SQL statement is structurally consistent with the actual database.

[0103] 4) Trial run mechanism: Trial run SQL statements in the database to determine semantic correctness and legality, and avoid execution failures in advance.

[0104] The above solution integrates database metadata and natural language context, and improves the accuracy and controllability of SQL generation through structure-aware prompt construction.

[0105] In some embodiments, the database query module is configured to perform the following process:

[0106] Step 51: Establish a communication connection with the test database;

[0107] Specifically, use the database connection driver or API to establish a connection with the test database based on the configuration information of the target test database (such as host name, port number, user name, password, etc.);

[0108] Step 52: pass the SQL query statement generated by the natural language SQL language generation module to the test database, and execute the query operation, including sending the query statement to the database server, waiting for the query result to be returned, etc.;

[0109] Step 53: Receive the query result data returned by the database and convert it into a set data format;

[0110] The query result data is test data, including the test statistics and specific test values ​​of the device;

[0111] Specifically, the query result data is converted into an appropriate data format (such as JSON, CSV, etc.) for subsequent analysis and display. If needed, the query results may also need to be further processed, such as filtering, sorting, etc.

[0112] Step 54: Disconnect from the test database.

[0113] Close the database connection, release occupied resources, ensure data privacy and security, and record logs for subsequent monitoring and troubleshooting.

[0114] In some embodiments, the output result fusion module is configured to perform the following process:

[0115] Step 61: Construct a chart to generate prompt words: The template background information output by the query optimization module is structured and encapsulated into prompt words (Prompt), and input into the large language model;

[0116] This prompt integrates user intent, contextual data structure, and analysis requirements: the user-initiated task description and retrieval-enhanced data, including relevant term explanations and data table metadata, serve as semantically driven input prompts, effectively improving the large model's understanding depth of chart requirements and output stability.

[0117] Step 62: Obtain the analysis results returned by the large language model and extract semantic components related to chart generation, including structural information such as chart type, data sequence, label, unit, and dimension indicator;

[0118] The large language model can independently determine which chart type is suitable for the data results of a query and output the specified request parameters for drawing;

[0119] Among them, the chart type may include bar chart, line chart, pie chart, etc.;

[0120] Step 63: Render the chart and generate graphics: Based on the extracted semantic components, including the chart type and populated data sequence, the system calls the chart rendering module (supporting ECharts or Matplotlib) to generate code graphics. The chart rendering module integrates chart style beautification, color strategy selection, label optimization, and other capabilities to achieve professional chart output and enhance interactivity.

[0121] Step 64: Summarize charts and automatically arrange report structure: Based on the subqueries and timing logic after user task division, extract the information of rendered icons and plan the layout structure of the report;

[0122] The layout structure can include forced linear arrangement of time-series tasks and parallel layout of multi-dimensional comparison. Each layout includes chart order, block logic, cross-references, etc., to achieve a structured arrangement with clear content hierarchy and logical closure.

[0123] The final report can be exported to multiple formats, including HTML (for embedding in system platforms), PDF (standard archiving format) and image files (such as PNG, SVG, etc.), and retain key semantic information metadata to support subsequent indexing, retrieval, version comparison and other functions.

[0124] Through the above steps, the output result fusion report generation module can convert the analysis result data of the large model into intuitive and clear charts and reports, helping users to better understand and use the analysis results.

[0125] This embodiment fully automates the process of semantic analysis, SQL generation, data summarization, and icon rendering in sequence according to the needs of users such as testers. It overcomes the shortcomings of traditional customized development templates for test result analysis, such as poor flexibility, high manual maintenance costs, and low data utilization. It shortens the process and duration of data analysis, and can improve the comprehensiveness of data analysis, increase data utilization, timely discover problems in the test process, and promote the optimization of test process solutions. It will play an important role in the future field of automatic test and measurement.

[0126] Example 2

[0127] Based on Example 1, this embodiment provides a test analysis and report generation method based on large model retrieval enhancement, such as Figure 2 As shown, the following steps are included:

[0128] Step 1: Split tasks based on acquired user queries, extract test domain expertise from the constructed knowledge vector library, and use a multi-strategy retrieval enhancement generation mechanism that integrates RAG, Self-RAG, and Graph-RAG to process the extracted relevant knowledge and user queries to obtain enhanced sub-query tasks. This step can be implemented in the query optimization module.

[0129] Step 2: Generate SQL query statements using the large model based on the enhanced sub-query task; this can be implemented in the natural language SQL generation module;

[0130] Step 3: Execute SQL query based on SQL query statement to extract target test data; this can be implemented in the database query module;

[0131] Step 4: Combine the target test data and template background information according to their relevance, fill the template, and obtain a test report. This can be achieved in the output result fusion module.

[0132] Furthermore, a multi-strategy retrieval enhancement generation mechanism integrating RAG, Self-RAG and Graph-RAG is used to extract relevant knowledge and user queries, extract template prompt words and generate template background information, including the following steps:

[0133] Perform semantic analysis on the query statement entered by the user to identify syntactic components, core intent, and hyponym / hypernym concept structures, and then divide the query statement into sub-query tasks;

[0134] Adopting the Self-RAG strategy, a small sample reasoning is used to evaluate the context completeness of each subquery;

[0135] Based on the completeness assessment results, the RAG process is used to extract relevant knowledge blocks from the constructed knowledge vector library for supplementation;

[0136] The supplemented knowledge blocks are input into the BGE-Reranker reranking model, and the semantic relevance between each candidate knowledge block and the query sentence is evaluated through the cross-attention mechanism. The knowledge blocks with semantic relevance greater than the set threshold are retained as the reranked text blocks.

[0137] For the rearranged knowledge blocks and query statements, Graph-RAG is used to extract core entities based on the knowledge graph in the knowledge vector library, and a structured graph path of "entity-relationship-entity" is constructed to obtain the enhanced sub-query task.

[0138] Example 3

[0139] This embodiment provides an electronic device, including a memory and a processor, and computer instructions stored in the memory and executed on the processor. When the computer instructions are executed by the processor, the steps of the test analysis and report generation method based on large model retrieval enhancement in Example 2 are completed.

[0140] The foregoing description is merely a preferred embodiment of the present disclosure and is not intended to limit the present disclosure. Those skilled in the art will readily appreciate that various modifications and variations are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present disclosure shall be included within the scope of protection of the present disclosure.

[0141] Although the above describes the specific implementation methods of the present disclosure in conjunction with the accompanying drawings, it is not intended to limit the scope of protection of the present disclosure. Those skilled in the art should understand that on the basis of the technical solution of the present disclosure, various modifications or variations that can be made by those skilled in the art without creative work are still within the scope of protection of the present disclosure.

Claims

1. A test analysis and report generation system based on large model retrieval enhancement, characterized by: include The query optimization module is used to split tasks based on the acquired user queries, extract test domain expertise from the constructed knowledge vector library, and use a multi-strategy retrieval enhancement generation mechanism that integrates RAG, Self-RAG, and Graph-RAG to process the extracted relevant knowledge and user queries to obtain enhanced sub-query tasks. The natural language generation SQL module is used to generate SQL query statements using the large model based on the enhanced sub-query tasks; Database query module, used to execute SQL query based on SQL query statement and extract target test data; The output result fusion module is used to combine the target test data and the template background information according to the relevance, fill the template, and obtain the test report.

2. The test analysis and report generation system based on large model retrieval enhancement according to claim 1, characterized in that: It also includes a query parsing module for parsing and semantically analyzing the acquired user queries, extracting entities and keywords in the queries, and determining keywords and entity information used to represent user intent.

3. The test analysis and report generation system based on large model retrieval enhancement according to claim 1, characterized in that: The private knowledge base construction module is used to integrate OCR image and text recognition and natural language processing methods, extract text information from test and measurement related knowledge document resources, and construct image and text structural relationships to obtain a constructed knowledge vector library.

4. The test analysis and report generation system based on large model retrieval enhancement according to claim 3, characterized in that: The method integrates OCR image and text recognition with natural language processing methods to extract text information from test and measurement-related knowledge document resources and construct image and text structural relationships, including the following steps: Collect knowledge document resources related to the test and measurement field; Perform character recognition, semantic classification and integration on the collected knowledge document resources; Detect, remove duplicates and process the integrated characters into a unified data format; A hierarchical recursive semantic chunking strategy is used to split long texts, first by natural paragraphs, and then by small, semantically complete chunks based on a preset threshold. The knowledge text after segmentation is input into the BGE model for high-dimensional semantic embedding; Build a domain vocabulary and perform terminology standardization mapping on the semantic vectors of high-dimensional semantic embedding; For the text vectors after mapping, entities and their semantic relationships are extracted to construct a knowledge graph.

5. The test analysis and report generation system based on large model retrieval enhancement according to claim 1, characterized in that: The query optimization module is configured to perform the following process: Perform semantic analysis on the query statement entered by the user to identify syntactic components, core intent, and hyponym / hypernym concept structures, and then divide the query statement into sub-query tasks; Adopting the Self-RAG strategy, a small sample reasoning is used to evaluate the context completeness of each subquery; Based on the completeness assessment results, the RAG process is used to extract relevant knowledge blocks from the constructed knowledge vector library for supplementation; The supplemented knowledge blocks are input into the BGE-Reranker reranking model, and the semantic relevance between each candidate knowledge block and the query sentence is evaluated through the cross-attention mechanism. The knowledge blocks with semantic relevance greater than the set threshold are retained as the reranked text blocks. For the rearranged knowledge blocks and query statements, based on the knowledge graph in the knowledge vector library, knowledge graph-assisted reasoning is used to extract core entities and construct a structured graph path of "entity-relationship-entity" as the enhanced sub-query task.

6. The test analysis and report generation system based on large model retrieval enhancement according to claim 1, characterized in that: The natural language SQL language generation module is configured to perform the following process: For each sub-query task obtained by task splitting, semantic analysis is performed to extract keywords and entities; based on the keywords and entities, the structural information of the target table in the test database is extracted and organized into SQL prompt word templates; Based on the SQL prompt word template input into the large language model, SQL statements that conform to the database structure and query semantics are generated.

7. The test analysis and report generation system based on large model retrieval enhancement according to claim 1, characterized in that: The database query module is configured to perform the following process: Establish a communication connection with the test database; Pass the SQL query statement generated by the natural language SQL language generation module to the database and execute the query operation; Receive the query result data returned by the database and convert it into the set data format.

8. A test analysis and report generation method based on large model retrieval enhancement is characterized in that: The steps include: Tasks are split based on the acquired user queries. Test domain expertise is extracted from the constructed knowledge vector library. A multi-strategy retrieval enhancement generation mechanism that integrates RAG, Self-RAG, and Graph-RAG is used to process the extracted relevant knowledge and user queries to obtain enhanced sub-query tasks. Based on the enhanced sub-query task, the large model is used to generate SQL query statements; Execute SQL queries based on SQL query statements to extract target test data; Combine the target test data and template background information according to the relevance, fill in the template, and obtain a test report.

9. The test analysis and report generation method based on large model retrieval enhancement according to claim 8, characterized in that: The method integrates the multi-strategy retrieval enhancement generation mechanism of RAG, Self-RAG and Graph-RAG, extracts relevant knowledge and user queries, extracts template prompt words and generates template background information, including the following steps: Perform semantic analysis on the query statement entered by the user to identify syntactic components, core intent, and hyponym / hypernym concept structures, and then divide the query statement into sub-query tasks; Adopting the Self-RAG strategy, a small sample reasoning is used to evaluate the context completeness of each subquery; Based on the completeness assessment results, the RAG process is used to extract relevant knowledge blocks from the constructed knowledge vector library for supplementation; The supplemented knowledge blocks are input into the BGE-Reranker reranking model, and the semantic relevance between each candidate knowledge block and the query sentence is evaluated through the cross-attention mechanism. The knowledge blocks with semantic relevance greater than the set threshold are retained as the reranked text blocks. For the rearranged knowledge blocks and query statements, Graph-RAG is used to extract core entities based on the knowledge graph in the knowledge vector library, and a structured graph path of "entity-relationship-entity" is constructed as the enhanced sub-query task.

10. An electronic device, characterized in that: It includes a memory and a processor, and computer instructions stored in the memory and running on the processor. When the computer instructions are run by the processor, the steps in the test analysis and report generation method based on large model retrieval enhancement as described in any one of claims 8 to 9 are completed.

Citation Information

Cited By

  • Gas pipe network maintenance method and system based on large language model and electronic equipment

    CN120780816A

  • Image-text report generation method fusing multi-mode large language model and RAG mechanism

    CN120995994A

  • Heat supply industry advanced report generation method and system based on large language model

    CN121052227A

  • Multi-scene intelligent financial analysis system based on large model and composite agent

    CN121073689A

  • NL2SQL method and system based on large language model and retrieval enhancement

    CN121326959A