Design basis text intelligent auditing system and use method thereof

By designing an intelligent text-based audit system, using the Milvus database and large language model for text preprocessing, vector search and intelligent judgment, the problem of inefficient and error recognition based on the consistency of design in the document is solved, and efficient and accurate audit results are achieved.

CN120374041APending Publication Date: 2025-07-25GUANGZHOU TIANYUE ELECTRONICS TECH
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510445981.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-10
Publication Date
2025-07-25

AI Technical Summary

Technical Problem

The prior art is inefficient and prone to errors when checking the consistency of design based content and design based on standard database in documents, especially when facing massive data, it is difficult to deal with complex similarity judgments and error type recognition.

Method used

The intelligent text review system based on design is adopted, including the standard library construction module, text preprocessing module, vector search module and large model intelligent judgment module. The Milvus database storage design is used to store the standard library, and accurately match and error type judgment are made through text preprocessing, vector search and large language models.

Benefits of technology

It improves document review efficiency, reduces manual intervention, can accurately identify differences in design basis under massive data and give modification tips, solving the problem of difficult to deal with complex similarity judgments in the prior art.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120374041A_ABST
    Figure CN120374041A_ABST
Patent Text Reader

Abstract

The invention provides a design basis text intelligent auditing system and a use method thereof. The design basis text intelligent auditing system comprises a standard library construction module, a text preprocessing module, a vector retrieval module and a large model intelligent judgment module. The output end of the standard library construction module is connected with the input end of the text preprocessing module, the output end of the text preprocessing module is connected with the input end of the vector retrieval module, and the output end of the vector retrieval module is connected with the input end of the large model intelligent judgment module. By means of the design basis text intelligent auditing system, the consistency of the design basis content and the design basis standard library in the document can be efficiently checked, prompts are given according to the problem basis, then corresponding modification and replacement content is given, and compared with an existing manual mode, the efficiency is improved, errors are not prone to occurring, and the effect is obvious especially under mass data; compared with a simple keyword matching pair, the problem that it is difficult to accurately recognize the situation that the numbers are slightly different or the name expressions are slightly different can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of engineering design, and particularly to an intelligent review system for design basis text and its usage method. Background Art

[0002] In the practice of engineering design, standardized design documents are crucial for ensuring engineering quality and system security. In engineering fields such as architecture, machinery, and electronics, accurately citing industry standards, specification documents, and technical manuals is a key factor in achieving high-quality design and smooth project progress. These standards not only provide a scientific basis for engineering design but also ensure product quality and compliance with relevant regulations, thus avoiding potential legal penalties and market access issues.

[0003] Currently, there are many challenges in checking whether the design basis content in the document is consistent with the design basis standard library. From the industry status quo, many enterprises or institutions rely on manual checking methods. This method faces various problems in the product application environment. First, manual checking is inefficient. Especially when faced with a large amount of design basis, staff need to spend a lot of time comparing item by item. Second, manual checking is prone to errors. Due to limited human energy, some subtle but crucial differences may be overlooked, such as numbering errors, name errors, or non-standard writing, etc. In addition, with the development of technology and the update of standards, the design basis library is constantly changing, and it is difficult for the manual method to keep up with this dynamic change in a timely manner, which may lead to the use of invalid design basis without awareness. In the existing product application environment, although there are also some automatic checking tools based on traditional databases, these tools often have a single function and can only perform simple keyword matching, unable to effectively handle complex similarity judgments and intelligent identification of error types. This makes a large amount of manual intervention still required to make up for the deficiencies of the tools in actual applications.

[0004] Therefore, there is an urgent need for an intelligent review system for design basis text to solve this problem. Summary of the Invention

[0005] To solve the above problems, the present invention adopts the following technical solutions: An intelligent review system for design basis text, comprising: a standard library construction module, a text preprocessing module, a vector retrieval module, and a large model intelligent judgment module;

[0006] The output end of the standard library construction module is connected to the input end of the text preprocessing module, the output end of the text preprocessing module is connected to the input end of the vector retrieval module, and the output end of the vector retrieval module is connected to the input end of the large model intelligent judgment module.

[0007] Further, the standard library construction module is used to implement the construction and maintenance of the design basis standard library, and the database created by the standard library construction module is a Milvus database, which is used to store the design basis standard library as a knowledge base for review.

[0008] Further, the text preprocessing module is used to obtain the design basis document to be reviewed and perform text preprocessing on the design basis document to standardize and preprocess the design basis content.

[0009] Further, the vector retrieval module is used to obtain the content that best matches the design basis content and the standard library content.

[0010] Further, the large model intelligent judgment module is used to judge whether there are problems with the design basis, and use the large language model to understand the retrieval results and the content of the design basis and judge the error type, and give specific error type prompts.

[0011] A usage method of a design basis text intelligent review system, using any one of the above-mentioned design basis text intelligent review systems, includes the following steps:

[0012] Step S1. Upload the design basis standard to the standard library construction module;

[0013] Step S2. Obtain the design basis document to be reviewed and perform preprocessing through the text preprocessing module;

[0014] Step S3. Retrieve the content preprocessed by the text processing module in the vector retrieval module, and then discriminate the retrieved results from the design basis standard to obtain the different content;

[0015] Step S4. Send the different content to the large model intelligent judgment module, and analyze it through the large model intelligent judgment module to obtain modification prompts and optimization strategies.

[0016] Further, in step S1, build a vector database, connect to the vector database, define the data structure, read the standard library file and extract text data, perform text vectorization processing, and insert vector data into the vector database in the standard library construction module.

[0017] Further, in step S2, perform content acquisition and line-by-line processing, duplicate detection processing, non-review item classification and annotation, text unified format processing, text classification processing, structured data encapsulation, and loop control and processing termination in the text preprocessing module.

[0018] Further, in the step S3, the vector retrieval module first connects to the vector database, then obtains the top-k contents with the closest semantics, and then performs screening and sorting by designing a method combining BLEU and ROUGE to achieve obtaining the content that best matches the design basis content and the standard inventory content.

[0019] Further, in the step S4, the large model intelligent judgment module further judges whether there are problems with the design basis, and uses the large language model to understand and judge the error type of the retrieval result and the content of the design basis, and gives specific error type prompts. The specific steps are as follows: (1) Obtain the review items: Traverse the list to be judged, and obtain the data in the list item by item for independent judgment; (2) Judge whether there is a mark: Process the previously preprocessed mark. If there is a non-review mark, the result of this item to be reviewed is non-reviewed content, and this judgment ends; If there is a duplicate mark, the result of this item to be reviewed is duplicate content, and this judgment ends; If there is no mark, proceed to the subsequent process; (3) Vector retrieval: Obtain the retrieval result for the review item through the vector retrieval module. If the result is empty, the result of this item to be reviewed is not in the library, and this judgment ends; If the result exists, proceed to the subsequent process; (4) Judgment on the validity of the retrieval result: Judge by obtaining the metadata content of the retrieval result. If it is invalid, the result of this item to be reviewed is that the review item is invalidated, and the updated content is the updated content in the metadata, and this judgment ends; If it is valid, proceed to the subsequent process; (5) Compare the text content page_content of the review item and the retrieval result: If the content is the same, the result of this item to be reviewed is passed, and this judgment ends; If the content is inconsistent, proceed to the subsequent process; (6) Judge the previous difference between the review item and the text content page_content of the retrieval result in the large model intelligent judgment module, and give an error correction explanation through the designed agent. Then the result of this item to be reviewed is error correction and error correction explanation, and this judgment ends; The error correction explanations are enumerated as: missing space between standard category and standard number, name error, incorrect writing of year or edition, standard category error; (7) Structured data encapsulation: Process the result after the above judgment and encapsulate it into json data with "review item", "result", "prompt", and "replacement content" as keywords. The specific content form and description are as follows:

[0020] {

[0021] "content": "content of the review item", # Input original text content

[0022] "result": 0 / 1, # The result for passed and non-reviewed content is 1, and others are 0

[0023] "instruction": "Output prompt", # The result of this content to be reviewed

[0024] "replace": "Retrieval content (top-1)" # The main content page_content of the retrieval result

[0025] };

[0026] (8) Loop control and process termination: After processing each piece of content in the list to be judged each time, append the finally encapsulated structured data to the result list; determine whether the current list subscript has exceeded the list to be judged; if not, continue to process the next element for judgment; otherwise, end the intelligent judgment process.

[0027] The beneficial effects of the present invention are as follows: Using this design-based text intelligent review system, the consistency between the design basis content in the document and the design basis standard library can be efficiently checked, and prompts can be given according to the problems, and then corresponding modified and replaced content can be given. Compared with the existing manual method, the efficiency is improved and it is not easy to make mistakes, especially in the case of a large amount of data; compared with simple keyword matching pairs, it can solve the problem that it is difficult to accurately identify the situation where the numbers are slightly different or the name expressions are slightly different. Brief Description of the Drawings

[0028] The drawings further illustrate the present invention, but the embodiments in the drawings do not constitute any limitation to the present invention.

[0029] Figure 1 It is a schematic diagram of the module connection of a design basis text intelligent review system provided for an embodiment;

[0030] Figure 2 It is a usage flow chart of the standard library construction module;

[0031] Figure 3 It is a usage flow chart of the text preprocessing module;

[0032] Figure 4 It is a usage flow chart of the vector retrieval module;

[0033] Figure 5 It is a usage flow chart of the large model intelligent judgment module. Detailed Embodiments

[0034] The following will further describe the technical solutions of the present invention in conjunction with the drawings of the embodiments of the present invention. The present invention is not limited to the following specific embodiments. It should be noted that, without conflict, the embodiments in the present invention and the features in the embodiments can be combined with each other.

[0035] Such as Figure 1As shown in the figure, a text intelligent review system based on design basis includes: a standard library construction module 100, a text preprocessing module 200, a vector retrieval module 300, and a large model intelligent judgment module 400; the output end of the standard library construction module 100 is connected to the input end of the text preprocessing module 200, the output end of the text preprocessing module 200 is connected to the input end of the vector retrieval module 300, and the output end of the vector retrieval module 300 is connected to the input end of the large model intelligent judgment module 400.

[0036] Specifically, the standard library construction module 100 is used to implement the construction and maintenance of the design basis standard library, and the database created by the standard library construction module 100 is a Milvus database, and the Milvus database is used to store the design basis standard library as a knowledge base for review. The text preprocessing module 200 is used to obtain the design basis document to be reviewed and perform text preprocessing on the design basis document to achieve the standardization and preprocessing of the design basis content. The vector retrieval module 300 is used to obtain the content that best matches the design basis content and the standard library content. The large model intelligent judgment module 400 is used to judge whether there are problems with the design basis, and use the large language model to understand and judge the error type of the retrieval result and the content of the design basis, and give specific error type prompts.

[0037] That is to say, using this text intelligent review system based on design basis, the consistency between the design basis content in the document and the design basis standard library can be efficiently checked, and prompts are given for the problematic basis, such as not being stored in the library, being invalid, or having writing errors such as number / name errors, non-standard writing, etc., and then the corresponding modified and replaced content is given. Compared with the existing manual method, the efficiency is improved and errors are not easy to occur, especially in the case of massive data; compared with simple keyword matching pairs, it can solve the problem that it is difficult to accurately identify the situation where the numbers are slightly different or the name expressions are slightly different.

[0038] A usage method of a text intelligent review system based on design basis, using any of the above-mentioned text intelligent review systems based on design basis, includes the following steps:

[0039] Step S1. Upload the design basis standard to the standard library construction module 100.

[0040] Specifically, in the standard library construction module 100, a vector database is built, the vector database is connected, the data structure is defined, the standard library file is read and the text data is extracted, text vectorization processing is performed, and the vector data is inserted into the vector database.

[0041] That is, as Figure 2As shown in the figure, by using the standard library construction module 100, a vector database is first built, and a Milvus database is created on the server to store the knowledge base for design basis standard library review. Then, the langchain library is used to connect to the Milvus service, and the connection parameters of the target database are configured. Then, according to the field structure of the data table in the design basis standard library, each row is generated into a Document object: the standard is used as the page_content of the text content, and other information such as validity and updated content is used as metadata. Then, by reading the files of the design basis standard library, the text column content to be processed is extracted, and the pre-trained vectorization model (bge-large-zh) is loaded to convert the text data into a vector representation of a fixed dimension. Finally, the text vectorized data is inserted into the specified knowledge base in the Milvus database.

[0042] Step S2. Obtain the design basis document to be reviewed and preprocess it through the text preprocessing module 200.

[0043] Specifically, in the text preprocessing module 200, content acquisition and line-by-line processing, duplicate detection processing, non-review entry classification and annotation, text unified format processing, text classification processing, structured data encapsulation, and loop control and processing termination are performed.

[0044] That is to say, such as Figure 3As shown, by using the text preprocessing module 200, first, the content is obtained and processed line by line. The design basis text content is obtained from the input source, and each line of data is independently processed in an iterative manner line by line. Then, duplicate detection processing is performed. The currently processed content is compared one by one with the currently generated list of items to be processed. If the same content is found, a duplicate flag is set. Next, the large model agent classifies the input content into categories such as "project files", "standards", and "others". That is to say, by customizing the large model agent, it is possible to intelligently identify project file entries that do not require review. In addition, in practical applications, non-review entries only exist in the first few lines of the text. Therefore, when processing content related to review, the classification judgment process is closed, thereby optimizing the calculation efficiency. Then, the text is uniformly formatted. The characters and symbols of the text are uniformly regularized, specifically including: full-width and half-width character conversion: all characters are uniformly converted to half-width characters to avoid matching errors caused by mixed character encodings. Space processing: Only the necessary spaces between the standard category and the standard number are retained, and unnecessary redundant spaces are removed. Symbol and format unification: The relevant symbols appearing in the text are normalized, such as deleting irrelevant symbols or presenting multi-form symbol content uniformly to ensure the consistency of text expression. After processing, the text content is classified based on the feature of whether it contains a number composed of a combination of uppercase letters and numbers. Standard content processing: For file content that conforms to the standard format, its serial number, standard prefix, and relevant ending punctuation marks are removed, and only the core content is retained for subsequent analysis. Non-standard content processing: For non-standard text content such as notices, only its serial number and ending punctuation marks are removed, and the rest of the content is retained in its complete form. The text content after preprocessing is encapsulated into a structured JSON format, which includes the original text, the content after preprocessing, and relevant annotation information. Such structured data is added to the list of items to be processed for subsequent module calls and analysis. It is worth mentioning that after processing each line of content, the system will automatically determine whether the current line number has exceeded the upper limit of the total line number. If not completed, the next line of content will continue to be processed; if all have been processed, the preprocessing process will end to ensure that the module processing efficiency is consistent with the user's expectations.

[0045] Step S3. Retrieve the content preprocessed by the text processing module in the vector retrieval module 300, and then discriminate the retrieved result from the design basis standard to obtain the content with differences.

[0046] Specifically, the vector retrieval module 300 first connects to the vector database, then obtains the top-k content with the closest semantics, and then screens and sorts by designing a method combining BLEU and ROUGE to achieve obtaining the content that best matches the content in the standard inventory.

[0047] That is, as Figure 4 shown, by using the vector retrieval module 300, first connect to the vector database Milvus, perform a similarity search on the input query text, obtain the top k candidate contents and their similarity scores, then calculate the Rouge-1 score of each candidate content and the query text, filter out the candidate contents with scores lower than the preset threshold, tokenize the text of the candidate contents passed through the Rouge-1 filter, and calculate its BLEU-2 score with the query text. Then, sort the candidate contents according to the BLEU-2 score, select the candidate content with the highest score as the optimal matching result and return it. If the content is empty after the previous filtering, return empty.

[0048] Step S4. Send the different content to the large model intelligent judgment module 400, and analyze it through the large model intelligent judgment module 400 to obtain modification prompts and optimization strategies.

[0049] Specifically, as Figure 5As shown in the figure, by using the large model intelligent judgment module 400, further judge whether there are problems with the design basis, and use the large language model to understand the retrieval results and the content of the design basis and judge the error type, and give specific error type prompts. The specific steps are as follows: (1) Obtain the audit item: Traverse the list to be judged, and obtain the data in the list item by item for independent judgment; (2) Judge whether there is a mark: Process the marks preprocessed before. If there is a non-audit mark, the result of this item to be audited is non-audit content, and this judgment ends; If there is a duplicate mark, the result of this item to be audited is duplicate content, and this judgment ends; If there is no mark, proceed to the subsequent process; (3) Vector retrieval: Obtain the retrieval result for the audit item through the vector retrieval module 300. If the result is empty, the result of this item to be audited is not in the library, and this judgment ends; If the result exists, proceed to the subsequent process; (4) Judge the validity of the retrieval result: Judge by obtaining the metadata content of the retrieval result. If it is invalid, the result of this item to be audited is that the audit item is invalid, and the updated content is the updated content in the metadata, and this judgment ends; If it is valid, proceed to the subsequent process; (5) Compare the body content page_content of the audit item and the retrieval result: If the content is the same, the result of this item to be audited is passed, and this judgment ends; If the content is inconsistent, proceed to the subsequent process; (6) Judge the previous differences between the audit item and the body content page_content of the retrieval result in the large model intelligent judgment module 400, and give an error correction description through the design agent, then the result of this item to be audited is error correction and error correction description, and this judgment ends; The error correction description enumeration is: missing space between standard category and standard number, name error, incorrect writing of year or edition, standard category error; (7) Structured data encapsulation: Process the result after the above judgment, and encapsulate it into json data with "audit item", "result", "prompt", "replacement content" as keywords. The specific content form and description are as follows:

[0050] {

[0051] "content": "The content of the audit item", # The original content entered

[0052] "result": 0 / 1, # The result for passed and non-audit content is 1, and others are 0

[0053] "instruction": "Output prompt", # The result of this item to be audited

[0054] "replace": "The body content of the retrieval result (top-1)" # The body content page_content of the retrieval result

[0055] };

[0056] (8) Loop control and processing termination: After each item in the list to be judged is processed, append the last encapsulated structured data to the result list; check whether the current list index has exceeded the list to be judged; if not, continue to process the next element for judgment; otherwise, end the intelligent judgment process.

[0057] That is to say, by using a text retrieval and matching method designed based on this design basis text intelligent review system, when conducting design basis review, design a standard library built with Milvus, and use vector retrieval combined with BLEU and ROUGE metrics for matching. Its data management and retrieval capabilities far exceed those of traditional manual databases or spreadsheet storage, and different from traditional character matching or word frequency statistics, it can more accurately capture the similarity of design basis at the semantic level, and is suitable for review requirements in cases such as expression variations or synonymous substitutions. Then, introduce a large model to intelligently judge the error type. The large model has strong learning ability and can adapt to new error types. Different from traditional rule-based methods that need to pre-set each error rule and are difficult to handle new situations or out-of-rule cases, such as the situation of non-standard writing without clear rule definition, the large model can learn and identify from a large number of samples, enabling more comprehensive and accurate discovery of various problems when reviewing design basis, and improving the reliability and effectiveness of the review.

[0058] In summary, the above embodiments are not restrictive embodiments of the present invention. Any modification or equivalent transformation made by those skilled in the art on the basis of the substantial content of the present invention falls within the technical scope of the present invention.

Claims

1. An intelligent text-based design review system, characterized in that, Including: A standard library construction module, a text preprocessing module, a vector retrieval module, and a large model intelligent judgment module; The output end of the standard library construction module is connected to the input end of the text preprocessing module, the output end of the text preprocessing module is connected to the input end of the vector retrieval module, and the output end of the vector retrieval module is connected to the input end of the large model intelligent judgment module.

2. The intelligent text review system for design basis according to claim 1, wherein: The standard library construction module is used to implement the construction and maintenance of the design basis standard library, and the database created by the standard library construction module is a Milvus database, and the Milvus database is used to store the design basis standard library as an audit knowledge base.

3. The intelligent text review system for design basis according to claim 2, wherein: The text preprocessing module is used to obtain the design basis document to be audited and perform text preprocessing on the design basis document to realize the standardization and preprocessing of the design basis content.

4. The intelligent text review system for design basis according to claim 3, wherein: The vector retrieval module is used to obtain the content that best matches the design basis content and the standard library content.

5. The intelligent text review system for design basis according to claim 4, wherein: The large model intelligent judgment module is used to judge whether there are problems with the design basis, and use the large language model to understand the retrieval results and the content of the design basis and judge the error type, and give specific error type prompts.

6. A method of using a text intelligent review system based on design, using the text intelligent review system based on design according to any one of claims 1-5, characterized in that, Including the following steps: Step S1. Upload the design basis standard to the standard library construction module; Step S2. Obtain the design basis document to be audited and perform preprocessing through the text preprocessing module; Step S3. Retrieve the content preprocessed by the text processing module in the vector retrieval module, and then discriminate the retrieved results from the design basis standard to obtain the different content; Step S4. Send the different content to the large model intelligent judgment module, and analyze it through the large model intelligent judgment module to obtain modification prompts and optimization strategies.

7. The method for using the text intelligent review system based on the design basis according to claim 6, characterized in that: In step S1, a vector database is built in the standard library construction module, the vector database is connected, the data structure is defined, the standard library file is read and the text data is extracted, text vectorization processing is performed, and the vector data is inserted into the vector database.

8. The method for using the text intelligent review system based on the design basis according to claim 7, characterized in that: In step S2, content acquisition and line-by-line processing, duplicate detection processing, non-audit item classification and annotation, text unified format processing, text classification processing, structured data encapsulation, and loop control and processing termination are performed in the text preprocessing module.

9. The method for using the text intelligent review system based on the design basis according to claim 8, characterized in that: In step S3, the vector retrieval module first connects to the vector database, then obtains the top-k content with the closest semantics, and then screens and sorts by designing a method combining BLEU and ROUGE to realize obtaining the content that best matches the design basis content and the standard library content.

10. The method for using the text intelligent review system based on the design basis according to claim 9, characterized in that: In the step S4, the large model intelligent judgment module further determines whether there are problems with the design basis, and uses the large language model to understand the retrieval results and the content of the design basis and judge the error type, and gives specific error type prompts. The specific steps are as follows: (1) Obtain the review items: Traverse the list to be judged, and obtain the data in the list item by item for independent judgment; (2) Judge whether there is a mark: Process the previously preprocessed mark. If there is a non-review mark, the result of this item to be reviewed is non-reviewed content, and this judgment ends; If there is a duplicate mark, the result of this item to be reviewed is duplicate content, and this judgment ends; If there is no mark, proceed to the subsequent process; (3) Vector retrieval: Obtain the retrieval result for the review item through the vector retrieval module. If the result is empty, the result of this item to be reviewed is not in the library, and this judgment ends; If the result exists, proceed to the subsequent process; (4) Judgment on the validity of the retrieval result: Judge by obtaining the metadata content of the retrieval result. If it is invalid, the result of this item to be reviewed is that the review item is invalidated, and the updated content is the updated content in the metadata, and this judgment ends; If it is valid, proceed to the subsequent process; (5) Compare the text content page_content of the review item and the retrieval result: If the content is the same, the result of this item to be reviewed is passed, and this judgment ends; If the content is inconsistent, proceed to the subsequent process; (6) Judge the previous difference between the review item and the text content page_content of the retrieval result in the large model intelligent judgment module, and give an error correction explanation through the design agent. Then the result of this item to be reviewed is error correction and error correction explanation, and this judgment ends; The error correction explanations are enumerated as: missing space between the standard category and the standard number, wrong name, wrong writing of year or edition, wrong standard category; (7) Structured data encapsulation: Process the result after the above judgment, and encapsulate it into json data with "review item", "result", "prompt", and "replacement content" as keywords. The specific content form and description are as follows: { "content": "content of the review item", # original input content "result": 0 / 1, # The result for passed and non-reviewed content is 1, and others are 0 "instruction": "output prompt", # result of this item to be reviewed "replace": "retrieved content (top-1)" # text content page_content of the retrieval result }; (8) Loop control and process termination: After each processing of one item in the list to be judged, append the finally encapsulated structured data to the result list; Judge whether the current list subscript has exceeded the list to be judged; If not exceeded, continue to process the next element for judgment; Otherwise, end the intelligent judgment process.

Citation Information

Patent Citations

  • Document auditing method, device and equipment and storage medium

    CN116663525A

  • Legal contract document generation method and system based on large model and vector retrieval

    CN118364056A

  • Construction scheme compliance auditing system and method based on knowledge graph and large model

    CN118643168A

  • Question answering system construction method based on large language model and question answering system

    CN118964587A

  • Method and apparatus for analyzing medical data using large language model

    KR102747558B1