Document auditing method and device, product, equipment and storage medium

By combining robot process automation, natural language processing and large language model technology, the problem of artificial dependence in scientific research project plagiarism checks and literature aggregation is solved, efficient and accurate document review and information aggregation are achieved, and the user experience and the quality of audit results are improved.

CN120068841APending Publication Date: 2025-05-30CHINA THREE GORGES CORPORATION
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510019715.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-07
Publication Date
2025-05-30

AI Technical Summary

Technical Problem

In the prior art, scientific research project plagiarism checking and literature aggregation mainly rely on manual methods, and there are problems such as low efficiency, easy errors, difficulty in retrieving information, incomplete coverage, and cumbersome information aggregation.

Method used

Robot process automation technology, natural language processing technology and large language model technology are used to obtain key data information in the documents to be reviewed, detect duplicate documents, conduct budget review and information aggregation analysis, and generate quantitative analysis reports and modification suggestions.

Benefits of technology

It improves the efficiency and accuracy of document review, expands the scope and depth of review, generates specific modification suggestions and natural language reports, and improves the clarity of user experience and audit results.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120068841A_ABST
    Figure CN120068841A_ABST
Patent Text Reader

Abstract

The invention discloses a document auditing method and device, a product, equipment and a storage medium, and relates to the technical field of document management, and the method comprises the following steps: obtaining key data information in a to-be-audited document; based on the key data information, detecting whether the to-be-audited document is a duplicate document, and obtaining a corresponding duplicate checking report based on a detection result; checking budget information in the to-be-checked document based on the key data information to obtain a checking report; performing information aggregation and analysis according to the key data information to obtain an information aggregation and analysis report; and performing quantitative analysis on the to-be-audited document based on the review report, the information aggregation and analysis report and the duplicate checking report to obtain a score and a modification suggestion of the to-be-audited document. By means of the method and device, the technical problems that in the prior art, RPA and AI technologies cannot be fully utilized, the quality and reliability of document auditing are low, and specific modification suggestions and reasons cannot be given are solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of document management, and particularly to a document review method, device, product, equipment and storage medium. Background Art

[0002] In the process of scientific research project establishment, it is necessary to conduct formal review and content review on the scientific research project application to ensure that the application meets the specification requirements and avoid duplication or plagiarism of existing work. At the same time, it is also necessary to conduct information retrieval and aggregation on external patents and literature to obtain relevant market dynamics and research results. Currently, scientific research project duplicate checking and literature aggregation mainly rely on manual methods, but there are the following problems: (1) Manual review is inefficient and error-prone; (2) Manual information retrieval is difficult and it is hard to cover comprehensively; (3) Manual information aggregation is cumbersome and it is difficult to form a report.

[0003] To solve the technical problems encountered in the process of mainly relying on manual methods for scientific research project duplicate checking and literature aggregation, a document review method based on the combination of RPA (Robotic Process Automation) and AI (Artificial Intelligence) technologies has been proposed in the prior art. Although this method realizes the automation and intelligence of document review to a certain extent, there are still the following deficiencies:

[0004] (1) There is a lack of in-depth interaction and collaboration between RPA and AI technologies. RPA mainly realizes mechanical data operations, and the data processing by AI is also relatively simple, unable to make full use of the advantages and characteristics of both;

[0005] (2) The scope and depth of document review are not comprehensive enough. Only the content and format of the document are considered, without considering the semantics and meanings of the document, as well as the relevance between the document and business matters and work processes. The accuracy and reliability of the review need to be improved;

[0006] (3) The result and feedback method of document review are not clear and friendly enough. Specific modification suggestions and reasons are not given, nor are they presented to the user in the form of natural language, resulting in poor user experience. Summary of the Invention

[0007] To solve the above-mentioned technical problems existing in the prior art, the present invention provides a document review method, device, product, equipment and storage medium.

[0008] To achieve the above object, the present invention provides the following technical solutions:

[0009] In a first aspect, an embodiment of the present invention provides a document review method, the method comprising:

[0010] Obtaining key data information in the document to be reviewed through robotic process automation technology, natural language processing technology and large language model technology;

[0011] Based on the key data information, detect whether the document to be reviewed is a duplicate document, obtain a detection result, and obtain a corresponding duplicate check report based on the detection result;

[0012] Review the budget information in the document to be reviewed based on the key data information, and obtain a review report;

[0013] Perform information aggregation and analysis according to the key data information, and obtain an information aggregation and analysis report;

[0014] Perform quantitative analysis on the document to be reviewed based on the review report, the information aggregation and analysis report, and the duplicate check report, and obtain a score and modification suggestions for the document to be reviewed.

[0015] Optionally, the obtaining of the key data information in the document to be reviewed through robotic process automation technology, natural language processing technology, and large language model technology includes:

[0016] Obtain the data information in the document to be reviewed through robotic process automation technology and natural language processing technology;

[0017] Perform information extraction on the data information through large language model technology to obtain key data information.

[0018] Optionally, the detecting whether the document to be reviewed is a duplicate document based on the key data information to obtain a detection result includes:

[0019] Retrieve the documents in the first database based on the keywords included in the key data information, and select the top N documents in the relevance ranking as similar documents;

[0020] Calculate the similarity between the key data information and the data information in each similar document;

[0021] Compare each similarity with a similarity threshold;

[0022] If each similarity is less than the similarity threshold, obtain a detection result that the document to be reviewed is not a duplicate document;

[0023] If there is any similarity greater than or equal to the similarity threshold, obtain a detection result that the document to be reviewed is a duplicate document.

[0024] Optionally, the step of calculating the similarity between the key data information and the data information in each similar document includes:

[0025] Encode the key data information using large language model technology to obtain a first numerical vector;

[0026] Encode each of the similar documents using large language model technology to obtain multiple second numerical vectors;

[0027] Based on the first numerical vector and each of the second numerical vectors, calculate the similarity between the key data information and each similar document through a multi-dimensional cosine similarity formula.

[0028] Optionally, obtaining a corresponding duplicate check report based on the detection result includes:

[0029] When the detection result indicates that the document to be reviewed is a duplicate document, use large language model technology to analyze the sub-topic correlation between the document to be reviewed and each of the similar documents, and obtain the duplicate sub-topics, similar sub-topics, and different sub-topics between the document to be reviewed and each of the similar documents;

[0030] Use large language model technology to analyze and summarize the duplicate sub-topics, the similar sub-topics, and the different sub-topics to obtain a duplicate check report.

[0031] Optionally, reviewing the budget information in the document to be reviewed based on the key data information to obtain a review report includes:

[0032] Based on the calculation basis included in the key data information, review the budget information in the document to be reviewed through large language model technology to obtain a review result;

[0033] If the review result indicates that there is at least one preset error type in the budget table, then based on the analysis result, use robotic process automation technology to obtain a review report;

[0034] Wherein, the preset error types include budget calculation errors, information omissions, format errors, logical errors, and other errors.

[0035] Optionally, aggregating and analyzing information according to the key data information to obtain an information aggregation and analysis report includes:

[0036] Based on the title, abstract, and keywords included in the key data information, use robotic process automation technology to retrieve documents in the second database, and select the top N documents in the relevance ranking as relevant documents;

[0037] Through large language model technology, extract the title, abstract, and keywords of each of the relevant documents, and encode the extracted information to obtain multiple third numerical vectors;

[0038] Encode the title, abstract, and keywords included in the key data information through large language model technology to obtain a fourth numerical vector;

[0039] Based on the fourth numerical vector and each of the third numerical vectors, calculate the similarity between the key data information and each relevant document through the multi-dimensional cosine similarity formula;

[0040] Based on the similarity between the key data information and each relevant document, perform analysis and summary through large language model technology to obtain an information aggregation and analysis report.

[0041] Optionally, the quantitative analysis of the document to be reviewed based on the review report, the information aggregation and analysis report, and the duplicate check report to obtain the score and modification suggestions for the document to be reviewed includes:

[0042] Perform data cleaning on the review report, the information aggregation and analysis report, and the duplicate check report;

[0043] Based on the cleaned review report, the information aggregation and analysis report, and the duplicate check report, obtain the innovation evaluation score and feasibility evaluation score of the document to be reviewed through large language model technology;

[0044] Calculate based on the innovation evaluation score and the feasibility evaluation score to obtain the score of the document to be reviewed;

[0045] Based on the innovation evaluation score and the feasibility evaluation score, perform analysis and summary through large language model technology to obtain the modification suggestions for the document to be reviewed.

[0046] In a second aspect, an embodiment of the present invention provides a document review device, and the device includes:

[0047] An information acquisition module, configured to obtain key data information in the document to be reviewed through robotic process automation technology, natural language processing technology, and large language model technology;

[0048] A report generation module, configured to detect whether the document to be reviewed is a duplicate document based on the key data information to obtain a detection result, and obtain a corresponding duplicate check report based on the detection result; review the budget information in the document to be reviewed based on the key data information to obtain a review report; perform information aggregation and analysis based on the key data information to obtain an information aggregation and analysis report;

[0049] A comprehensive analysis module, configured to perform quantitative analysis on the document to be reviewed based on the review report, the information aggregation and analysis report, and the duplicate check report to obtain the score and modification suggestions for the document to be reviewed.

[0050] Optionally, the report generation module is specifically configured to:

[0051] When the detection result indicates that the document to be reviewed is a duplicate document, use large language model technology to analyze the sub-topic correlation between the document to be reviewed and the similar document, and obtain the duplicate sub-topics, similar sub-topics, and different sub-topics between the document to be reviewed and the similar document;

[0052] Use large language model technology to analyze and summarize the duplicate sub-topics, the similar sub-topics, and the different sub-topics to obtain a duplicate check report.

[0053] Optionally, the report generation module is specifically configured to:

[0054] Based on the calculation basis included in the key data information, review the budget information in the document to be reviewed through large language model technology to obtain a review result;

[0055] If the review result indicates that there is at least one preset error type in the budget table, based on the analysis result, use robotic process automation technology to obtain a review report;

[0056] Wherein, the preset error types include budget calculation errors, information omissions, format errors, logical errors, and other errors.

[0057] Optionally, the report generation module is specifically configured to:

[0058] Based on the title, abstract, and keywords included in the key data information, use robotic process automation technology to retrieve documents in the second database, and select the top N documents in the relevance ranking as relevant documents;

[0059] Through large language model technology, extract the title, abstract, and keywords of each relevant document, and encode the extracted information to obtain multiple third numerical vectors;

[0060] Through large language model technology, encode the title, abstract, and keywords included in the key data information to obtain a fourth numerical vector;

[0061] Based on the fourth numerical vector and each third numerical vector, calculate the similarity between the key data information and each relevant document through the multi-dimensional cosine similarity formula;

[0062] Based on the similarity between the key data information and each relevant document, conduct analysis and summary through large language model technology to obtain an information aggregation and analysis report.

[0063] In a third aspect, an embodiment of the present invention further provides an electronic device, including: a memory and a processor; the processor is configured to read and execute a computer program stored in the memory to implement the steps of the foregoing document review method.

[0064] In a fourth aspect, an embodiment of the present invention further provides a computer storage medium, in which computer-executable instructions are stored, and when the computer-executable instructions are executed, the steps of the foregoing document review method are implemented.

[0065] In a fifth aspect, an embodiment of the present invention further provides a computer program product, including a computer program / instructions, and when the computer program / instructions are executed by a processor, the steps of the foregoing document review method are implemented.

[0066] The beneficial effects brought by the technical solutions provided by the embodiments of the present invention include:

[0067] 1. Improve the efficiency and accuracy of document review: Utilize robotic process automation technology to automatically obtain document content, avoiding the time and errors of manual operations. At the same time, utilize large language model technology to deeply understand the document content, and through automation and in-depth analysis, perform functions such as text summarization, information extraction, and similarity calculation, improving the quality and reliability of document review, and enhancing the efficiency and accuracy of document review;

[0068] 2. Expand the scope and depth of document review: Utilize large language model technology to review documents from multiple dimensions, considering the content and format of the documents, and at the same time paying attention to the semantics, meanings, and relevance to business matters and work processes of the documents. The comprehensive and in-depth review expands the scope of document review, making the review results more comprehensive and accurate;

[0069] 3. Use large language model technology to generate specific modification suggestions and reasons based on the review results and present them to users in the form of natural language, improving the user experience. At the same time, through robotic process automation technology, automate the generation and sending of document review reports, facilitating users and managers to consult and make decisions. This optimization makes the results and feedback of document review clearer and more valuable for reference, enhancing user satisfaction.

[0070] Through the present invention, the technical problems in the prior art that RPA and AI technologies cannot be fully utilized, the quality and reliability of document review are low, and specific modification suggestions and reasons cannot be given are solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0071] To more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, other drawings can be obtained based on these drawings without creative efforts.

[0072] Figure 1 It is a schematic flowchart of an embodiment of the document review method of the present invention;

[0073] Figure 2 For Figure 1 It is a detailed flowchart of step S10 in

[0074] Figure 3 For Figure 1 It is a detailed flowchart of step S20 in

[0075] Figure 4 For Figure 1 It is a detailed flowchart of step S40 in

[0076] Figure 5 For Figure 1 It is a detailed flowchart of step S50 in

[0077] Figure 6 It is a schematic diagram of the functional modules of an embodiment of the document review device of the present invention;

[0078] Figure 7 It is a schematic diagram of the structure of an electronic device according to an embodiment of the present invention. Detailed implementation manners

[0079] The following will clearly and completely describe the technical solutions in the embodiments of the present invention with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are only some embodiments of the present invention, rather than all embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts belong to the scope of protection of the present invention.

[0080] To make the purpose, technical solutions and advantages of the present invention clearer, the following will describe the embodiments of the present invention in detail with reference to the drawings.

[0081] In a first aspect, an embodiment of the present invention provides a document review method.

[0082] In one embodiment, referring to Figure 1 , Figure 1 It is a schematic flowchart of an embodiment of the document review method of the present invention. As Figure 1 shown, the document review method includes:

[0083] Step S10, obtain the key data information in the document to be reviewed through robotic process automation technology, natural language processing technology, and large language model technology;

[0084] In some specific embodiments, such as Figure 2 shown, step S10 includes:

[0085] Step S101, obtain the data information in the document to be reviewed through robotic process automation technology and natural language processing technology;

[0086] Step S102, perform information extraction on the data information through large language model technology to obtain the key data information.

[0087] In this embodiment, the robotic process automation technology (Robotic process automation, abbreviated as RPA) is used to obtain the data information of the document to be reviewed.

[0088] If the document to be reviewed is a picture, technologies such as OCR technology, table recognition technology, layout analysis technology, and formula recognition technology are used to restore the picture-type PDF to a DOCX document. According to the structured position information, determine the data information in the document to be reviewed.

[0089] Combined with the structured information, analyze the data information in the document to be reviewed, extract the data information in the document to be reviewed, and use large language model technology to summarize the extracted data information in the document to be reviewed, and extract key data information such as title, keywords, research content, budget calculation basis, abstract, research direction, research goal, research method, expected results, and application.

[0090] Specifically, first determine the categories and formats of the key information to be extracted, such as research direction, research goal, research content, research method, expected results, etc., as well as the corresponding tags or keywords; then use the structured text position to extract the text information from the text content of the project application form, initially obtain the data information in the document to be reviewed, and then use large language model technology to summarize the initially obtained data information in the document to be reviewed to obtain key data information such as title, keywords, research content, budget calculation basis, abstract, research direction, research goal, research method, expected results, and application.

[0091] Furthermore, based on the obtained key data information such as title, keywords, research content, budget calculation basis, abstract, research direction, research goal, research method, expected results, and application, use large language model technology to generate the corresponding table and output the result.

[0092] Among them, OCR (Optical Character Recognition) is a technology that converts text in images into editable text. It is widely used in fields such as document digitization, information extraction, and data processing. OCR can recognize printed text, handwritten text, and even certain types of fonts and symbols.

[0093] Table recognition technology is a technology that automatically recognizes and extracts the content and structure of tables from documents or images, and is widely used in fields such as data entry, information retrieval, and document analysis. By using computer vision and machine learning algorithms, table recognition can convert complex table information into an editable format, facilitating users to further process and analyze data. The general table recognition production line includes a table structure recognition module, a layout area analysis module, a text detection module, and a text recognition module.

[0094] Layout analysis technology is a technology that extracts structured information from document images, mainly used to convert complex document layouts into machine-readable data formats. This technology has a wide range of applications in fields such as document management, information extraction, and data digitization. Layout analysis technology can recognize and extract text blocks, headings, paragraphs, pictures, tables, and other layout elements in documents by combining optical character recognition (OCR), image processing, and machine learning algorithms. This process usually includes three main steps: layout analysis, element analysis, and data formatting, and finally generates structured document data, improving the efficiency and accuracy of data processing. The general layout analysis production line includes a table structure recognition module, a layout area analysis module, a text detection module, a text recognition module, a formula recognition module, a seal text detection module, a text image correction module, and a document image orientation classification module.

[0095] Formula recognition technology is a technology that automatically recognizes and extracts the content and structure of LaTeX formulas from documents or images, and is widely used in document editing and data analysis in fields such as mathematics, physics, and computer science. By using computer vision and machine learning algorithms, formula recognition can convert complex mathematical formula information into an editable LaTeX format, facilitating users to further process and analyze data.

[0096] Step S20, based on the key data information, detect whether the document to be reviewed is a duplicate document, obtain a detection result, and obtain a corresponding duplicate check report based on the detection result;

[0097] In some specific embodiments, as Figure 3 shown, the detecting whether the document to be reviewed is a duplicate document based on the key data information and obtaining a detection result includes:

[0098] Step S201: Retrieve the documents in the first database based on the keywords included in the key data information, and select the top N documents in the relevance ranking as similar documents.

[0099] In this embodiment, the first database is an internal database such as a scientific research database. According to the keywords included in the key data information, use the large language model technology to retrieve the documents in the first database, select the top N documents in the relevance ranking as similar documents, and calculate the similarity between the key data information and each of the similar documents. It is easy to understand that the documents with a higher proportion of keyword repetition rank higher in the relevance ranking.

[0100] Step S202: Calculate the similarity between the key data information and each of the similar documents.

[0101] In some specific embodiments, step S202 includes:

[0102] Use the large language model technology to encode the key data information to obtain a first numerical vector;

[0103] Use the large language model technology to encode each of the similar documents to obtain multiple second numerical vectors;

[0104] Based on the first numerical vector and each of the second numerical vectors, calculate the similarity between the key data information and each similar document through the multi-dimensional cosine similarity formula.

[0105] In this embodiment, use the large language model technology to numerically encode the key data information to obtain a first numerical vector. Use the large language model technology to encode each similar document to obtain multiple second numerical vectors.

[0106] Exemplarily, take a simple three-dimensional vector as an example for illustration. In actual applications, numerical vectors with higher dimensions (such as 768 dimensions or 1024 dimensions) may be used.

[0107] Vector representation: Let the vector representation of the document to be reviewed be: V_document_to_be_audited. Assume its encoding result is: V_document_to_be_audited = (v1, v2, v3) = (0.5, 0.2, 0.1). Let the vector representation of the first similar document be V_first_similar_document. Assume its encoding result is: V_first_similar_document = (s1,1, s1,2, s1,3) = (0.2, 0.1, 0.05).

[0108] Multi-dimensional cosine similarity formula:

[0109] In a general N-dimensional space (N ≥ 3), two vectors a = (a 1 ,..., a n ) and b = (b1 ,..., b n ) The cosine similarity is defined as:

[0110]

[0111] Calculation process:

[0112] Step 1: Vector dot product. The dot product formula for two vectors is: a×b = a 1 ×b 1 +a 2 ×b 2 +a 3 ×b 3 . Calculate for the vectors in the example:

[0113] V document to be reviewed × V first similar document = 0.5×0.2 + 0.2×0.1 + 0.1×0.05 = 0.125.

[0114] Step 2: Vector magnitude. The magnitude formula for each vector is:

[0115]

[0116] The calculation process for the magnitude of V document to be reviewed is:

[0117]

[0118] The calculation process for the magnitude of V first similar document is:

[0119]

[0120] Step 3: Calculate the cosine similarity. Substitute the dot product and magnitude into the multi-dimensional cosine similarity formula to calculate the similarity between the document to be reviewed and the first similar document. The specific calculation process is:

[0121]

[0122] In step S203, compare each of the similarities with the similarity threshold;

[0123] In step S204, if each of the similarities is less than the similarity threshold, the detection result is that the document to be reviewed is not a duplicate document;

[0124] In step S205, if there is any one of the similarities greater than or equal to the similarity threshold, the detection result is that the document to be reviewed is a duplicate document.

[0125] In this embodiment, after obtaining the similarity between the key data information and each of the similar documents, each similarity is compared with a set similarity threshold to check the duplicate of the document to be reviewed. If there is only 1 similar document, a similarity threshold θ is set, for example, θ = 0.80. If the calculated similarity Sim≥θ, the document is determined to be a duplicate document; if all similarities are <θ, the document is determined not to be a duplicate document. It is easy to understand that the value of θ is set according to needs and is only for illustration here without limitation.

[0126] Specifically, when there are multiple similar documents, if each similarity is less than the set similarity threshold, the detection result is that the document to be reviewed is not a duplicate document. If there is any similarity greater than or equal to the set similarity threshold, it is considered that there is duplicate or highly similar content, and the detection result is that the document to be reviewed is a duplicate document.

[0127] Further, the document to be reviewed corresponding to the similarity greater than or equal to the set similarity threshold is sent to the terminal for relevant personnel to conduct a re-detection.

[0128] In some specific embodiments, obtaining a corresponding duplicate check report based on the detection result includes:

[0129] When the detection result is that the document to be reviewed is a duplicate document, using large language model technology, analyze the sub-topic correlation between the document to be reviewed and each of the similar documents to obtain the duplicate sub-topics, similar sub-topics, and different sub-topics between the document to be reviewed and each of the similar documents;

[0130] Using large language model technology to analyze and summarize the duplicate sub-topics, the similar sub-topics, and the different sub-topics to obtain a duplicate check report.

[0131] In this embodiment, when the detection result is that the document to be reviewed is a duplicate document, using large language model technology, analyze the sub-topic correlation between the document to be reviewed and each similar document to obtain the duplicate sub-topics, similar sub-topics, and different sub-topics between the document to be reviewed and each similar document. Among them, the duplicate sub-topics are the sub-topics with the same technical operation process, technical research direction, research effect, and research purpose; the similar sub-topics are the sub-topics with different research themes, similar technical operation processes, and similar technical research directions and research effects; the different sub-topics are the sub-topics with different research themes, technical research directions, research effects, and research purposes.

[0132] Using large language model technology to analyze and summarize duplicate sub - topics, similar sub - topics, and different sub - topics to obtain a duplicate check report. Specifically, taking the analysis and summary of similar sub - topics using large language model technology as an example, there is a certain overlap between two sub - topics in edge computing, data fusion, and intelligent analysis, but they are different in specific research content and application fields. Project number: WWKY - 2020 - 0263 pays more attention to the overall data collection and processing architecture of large bases, while project number: Key Special Project - Research and Demonstration Application of Multi - source Data Fusion and Intelligent Monitoring Technology for Typical Wind Turbines in Xinghua Bay, Fujian focuses more on the specific monitoring requirements of offshore wind turbines and the development of intelligent analysis algorithms. Among them, during the process of using large language model technology to analyze and summarize duplicate sub - topics, similar sub - topics, and different sub - topics, data will be desensitized.

[0133] And so on, by using large language model technology to analyze and summarize the duplicate sub - topics, similar sub - topics, and different sub - topics between the document to be reviewed and each similar document, a duplicate check report between the document to be reviewed and each similar document can be obtained.

[0134] Step S30, based on the key data information, review the budget information in the document to be reviewed to obtain a review report;

[0135] In some specific embodiments, step S30 includes:

[0136] Based on the calculation basis included in the key data information, use large language model technology to review the budget information in the document to be reviewed to obtain a review result;

[0137] If the review result is that there is at least one preset error type in the budget table, then based on the analysis result, use robotic process automation technology to obtain a review report;

[0138] Among them, the preset error types include budget calculation errors, information omissions, format errors, logical errors, and other errors.

[0139] In this embodiment, based on the key data information, review the budget information in the document to be reviewed to obtain a review report. During the review process, special attention is paid to the accuracy of the budget information and the calculation basis to ensure the standardization and integrity of the document to be reviewed.

[0140] Specifically, based on the calculation basis included in the key data information, use large language model technology to analyze the budget information in the document to be reviewed to identify whether there is at least one preset error type in the budget information of the document to be reviewed. Among them, the preset error types include budget calculation errors, information omissions, format errors, logical errors, and other errors. The preset error types and their error factors are shown in Table 1.

[0141] Table 1

[0142]

[0143]

[0144] It should be noted that in the document to be reviewed, errors other than budget calculation errors, information omissions, format errors, and logical errors are other errors.

[0145] If there is at least one preset error type in the budget table in the review result, then based on the analysis result, count the number of various errors, and use robotic process automation technology to obtain a review report to facilitate understanding of the overall quality of the document to be reviewed.

[0146] Step S40: Perform information aggregation and analysis based on the key data information to obtain an information aggregation and analysis report;

[0147] In some specific embodiments, as Figure 4 shown, step S40 includes:

[0148] Step S401: Based on the title, abstract, and keywords included in the key data information, use robotic process automation technology to retrieve documents in the second database, and select the top N documents in the relevance ranking as relevant documents;

[0149] Step S402: Through large language model technology, extract the title, abstract, and keywords of each relevant document, and encode the extracted information to obtain multiple third numerical vectors;

[0150] Step S403: Through large language model technology, encode the title, abstract, and keywords included in the key data information to obtain a fourth numerical vector;

[0151] Step S404: Based on the fourth numerical vector and each third numerical vector, calculate the similarity between the key data information and each relevant document through the multi-dimensional cosine similarity formula;

[0152] Step S405: Based on the similarity between the key data information and each relevant document, perform analysis and summary through large language model technology to obtain an information aggregation and analysis report.

[0153] In this embodiment, the second database is a patent database and various academic databases. Based on the title, abstract, and keywords included in the key data information, the documents in the second database are retrieved using robotic process automation technology, that is, relevant literature and patents are retrieved from the patent database and various academic databases (such as CNKI, Wanfang, Google Scholar, etc.), and the top N documents in the relevance ranking are used as relevant documents and saved in text or table format.

[0154] Through large language model technology, the title, abstract, and keywords of each of the said relevant documents are extracted, and the extracted information is encoded to obtain multiple third numerical vectors; through large language model technology, the title, abstract, and keywords included in the key data information are encoded to obtain a fourth numerical vector;

[0155] Based on the fourth numerical vector and each third numerical vector, the similarity between the key data information and each relevant document is calculated through the multi-dimensional cosine similarity formula; based on the similarity between the key data information and each relevant document, analysis and summary are carried out through large language model technology to obtain an information aggregation and analysis report.

[0156] Using large language model technology, relevant literature and patents related to the project establishment content are retrieved from the patent database and academic databases, and information extraction and aggregation are carried out to provide the project establishment personnel and management with the latest market dynamics and research results. The comprehensiveness and timeliness of the information enable decision-makers to better understand the current scientific research project situation and make more informed decisions.

[0157] Furthermore, using robotic process automation technology, the format and content of the information aggregation and analysis report are automatically generated and sent to relevant personnel via email or other means.

[0158] Step S50: Based on the review report, the information aggregation and analysis report, and the duplicate check report, a quantitative analysis is performed on the document to be reviewed to obtain the score and modification suggestions for the document to be reviewed.

[0159] In some specific embodiments, as Figure 5 shown, step S50 includes:

[0160] Step S501: Perform data cleaning on the review report, the information aggregation and analysis report, and the duplicate check report;

[0161] Step S502: Based on the cleaned review report, the information aggregation and analysis report, and the duplicate check report, through large language model technology, obtain the innovation evaluation score and feasibility evaluation score of the document to be reviewed.

[0162] Step S503: Calculate based on the innovative evaluation score and the feasibility evaluation score to obtain the score of the document to be reviewed.

[0163] Step S504: Based on the innovative evaluation score and the feasibility evaluation score, analyze and summarize through large language model technology to obtain modification suggestions for the document to be reviewed.

[0164] In this embodiment, large language model technology and robotic process automation technology are used to evaluate the innovation and feasibility of the project proposal, and a modification suggestion report for the document to be reviewed is generated. Specifically, data from review reports, information aggregation and analysis reports, and duplicate check reports are collected using data pipelines or scripts, and the collected data is cleaned. Errors, incomplete, inconsistent, or duplicate information in the data is identified, corrected, and removed to improve the accuracy, consistency, and integrity of the data, ensure that the data can truly reflect the business situation, and provide a reliable basis for data analysis and decision-making.

[0165] Based on the data in the cleaned review reports, information aggregation and analysis reports, and duplicate check reports, through large language model technology, the innovative evaluation score and the feasibility evaluation score of the document to be reviewed are obtained.

[0166] Specifically, the number of innovation points in the document to be reviewed is counted through large language model technology, and based on the detection result of whether the document to be reviewed is a duplicate document, it is determined whether the content in the document to be reviewed is mentioned a large number of times. It is easy to think that if the detection result is that the document to be reviewed is a duplicate document, it is determined that the content in the document to be reviewed is mentioned a large number of times; if the detection result is that the document to be reviewed is not a duplicate document, it is determined that the content in the document to be reviewed is not mentioned a large number of times.

[0167] Based on the number of innovation points in the document to be reviewed and the judgment result of whether the content in the document to be reviewed is mentioned a large number of times, through a preset calculation method for the innovative evaluation score, the innovative evaluation score of the document to be reviewed is obtained.

[0168] Exemplarily, integrate the form review content, content review content, comparison of external and internal projects, similarity of plan names, similarity of main research content and objectives, similarity of expected results and applications, comprehensive similarity, relevance, content of historical related projects, reasons for relevance. Return the relevant results in the form of a table or system, and give the final report and suggestions.

[0169] 1. Send the prompt words in batches.

[0170] The system needs to understand the innovation points, technical route feasibility, and market prospects in the document to be reviewed more deeply. For this reason, the script of robotic process automation technology will perform multiple rounds of large language model calls on each document to be reviewed, and the common prompt words are as follows:

[0171] Prompt for innovation evaluation:

[0172] "You are a scientific research review expert. The following is a description of the core research content and technical advantages of this project. Please judge whether there are obvious innovation points. If so, please indicate the approximate number (0 or 1 - N), and explain whether there is a disruptive breakthrough. The answer should be in JSON format, with fields including: innovation_count and tech_breakthrough (true / false)".

[0173] Prompt for similarity impact analysis:

[0174] "The following is the core technology description of this project approval document, as well as known similar literature paragraphs and similarity scores. Please read and judge: Does this similarity only focus on the background review? Or does it involve the key technology implementation? Please write a paragraph of about 150 words to explain the possible negative impact of similarity on the project innovation".

[0175] Prompt for feasibility and budget rationality:

[0176] "The following is the budget details and a brief description of the technical route of the project. Please indicate: 1) Whether there are major budget errors (such as a difference exceeding 30%); 2) Whether the technical route has a complete implementation path; 3) Whether the team and resources basically meet the needs. Please return JSON, with fields including:

[0177] budget_defect (none / small / major), route_integrity (good / partial / lack) and team_fit (high / medium / low)".

[0178] To reduce the randomness of the single - model output, the robotic process automation technology calls the large - language model (LLM) three times continuously for the same prompt each time, and averages or votes on the three results obtained.

[0179] For example, the median can be taken for the "innovation_count" field; for "team_fit", if two outputs are "medium" and one is "low", then the final result is set to "medium".

[0180] 2. The evaluation of the document to be reviewed is divided into three major dimensions, with a total of 100 points: Innovation: 30 points, Feasibility: 30 points, and Comprehensive Evaluation: 40 points.

[0181] Innovation: The basic score is 20 points.

[0182] If the cosine similarity (similarity_score) ≤ 50%, add 1 point; if the cosine similarity is between 50% and 70%, deduct 3 points; if the cosine similarity ≥ 70%, deduct 5 points;

[0183] If patent_conflict = TRUE (external patent conflict is found), deduct 3 points;

[0184] If the large language model identifies ≥ 2 innovation points in the document to be reviewed, add X points for X innovation points, and no points are added if no innovation points are identified. The total upper limit is 30 points.

[0185] In this way, the database has fields such as similarity_score, patent_conflict, innovation_count, and whether there is a disruptive breakthrough. At this time, start calculating "innovation_score" (initial value 20, maximum not exceeding 30).

[0186] Influence of similarity:

[0187] If similarity_score ≤ 50%, then innovation_score += 1;

[0188] If 50% < similarity_score < 70%, then innovation_score -= 1;

[0189] If similarity_score ≥ 70%, then innovation_score -= 2.

[0190] Document conflict:

[0191] If patent_conflict = TRUE:

[0192] If there is only partial conflict, innovation_score -= 3;

[0193] If the conflict is severe or the core idea has been covered by other documents, innovation_score -= 5.

[0194] No conflict does not affect the score.

[0195] Number of innovation points:

[0196] If innovation_count = 0, no points are added;

[0197] If innovation_count = 1, innovation_score += 2;

[0198] If innovation_count >= 2, innovation_score += 5.

[0199] If tech_breakthrough = true is detected simultaneously, additionally innovation_score += 2.

[0200] Score truncation:

[0201] Finally, if innovation_score < 0, the innovation evaluation score is set to 0;

[0202] If innovation_score > 30, the innovation evaluation score is set to 30. After completion, write it back to the database, and the innovation evaluation score scoring_result.innovation_score = [final value].

[0203] Statistically match the resources required for the document to be reviewed with the existing resources through large language model technology; obtain the maturity and risk level of the technology in the document to be reviewed based on the external information report. Based on the matching degree, maturity, and risk level, obtain the feasibility evaluation score of the document to be reviewed through a preset feasibility evaluation score calculation method.

[0204] Exemplarily, the basic feasibility score is 20 points.

[0205] Deduct 1 point for each "general budget error", and deduct 3 points for major budget defects (such as a difference of more than 30%). If the large language model determines that "the technical route is complete and has preliminary verification", 2 points can be added; if a key route is detected to be missing, 5 points will be deducted; if the team resources do not match the project establishment requirements (determined by the large language model), 2 - 5 points will be deducted.

[0206] In the feasibility dimension, the system will refer to elements such as "budget rationality", "completeness of the technical route", and "team and resources". The initial feasibility score feasibility_score = 20, with a maximum of 30 points.

[0207] Budget error:

[0208] In the form review report, for each "general budget error" found, feasibility_score -= 1. For "major budget defects" such as a difference exceeding 30% or a key item being blank, each feasibility_score -= 3.

[0209] If the budget is completely normal, feasibility_score += 2 can be given as an encouragement (but the maximum cannot exceed 30 points).

[0210] Completeness of the technical route:

[0211] If the large language model returns route_integrity = good, then feasibility_score += 2;

[0212] If route_integrity = partial, then no addition or subtraction;

[0213] If route_integrity = lack, then feasibility_score -= 5.

[0214] Team resources:

[0215] If team_fit = high, then feasibility_score += 3;

[0216] If team_fit = medium, then no addition or subtraction;

[0217] If team_fit = low, then feasibility_score -= 5.

[0218] Feasibility score truncation:

[0219] After calculation, if feasibility_score < 0, then the feasibility evaluation score is set to 0;

[0220] If feasibility_score > 30, then the feasibility evaluation score is set to 30; and write it back to the database.

[0221] Calculate based on the innovation evaluation score and the feasibility evaluation score to obtain the score of the document to be reviewed. Exemplarily, the comprehensive evaluation: the basic score is 30.

[0222] Market / application prospect: If the large language model identifies that the market / application prospect is "highly compatible", then add 3 points; if the market / application prospect is "highly saturated", then subtract 3 points; if the quantification degree of the research goal is insufficient, then subtract 2 - 3 points; if the goal is clear and measurable, then add 2 points; if the risk control is "no countermeasure", then subtract 3 points; if the "countermeasures are complete", then add 2 points.

[0223] The initial comprehensive evaluation score comprehensive_score = 30, and the maximum is 40 points. The system mainly focuses on the market / application prospect, the quantification degree of the research goal, and the risk control situation.

[0224] Market / application prospect:

[0225] If the large language model identifies market_need = high, then comprehensive_score += 3;

[0226] If the large language model identifies market_need = low or "highly saturated", then comprehensive_score -= 3.

[0227] Quantification of research objectives:

[0228] If no clear quantification indicators are found, such as "the plan is expected to achieve several results", then comprehensive_score -= 2 uniformly;

[0229] If it is completely blank, then comprehensive_score -= 4;

[0230] If the number of patents / stage achievements is clearly stated, then comprehensive_score += 2.

[0231] It should be noted that the logic of adding and deducting points cannot be superimposed. It is necessary to first evaluate whether the quantification is met, and then decide whether to add or deduct points.

[0232] Risk control:

[0233] If the large language model determines that risk_control = none, then comprehensive_score -= 3;

[0234] If risk_control = adequate, then comprehensive_score += 2.

[0235] Score truncation and storage:

[0236] If comprehensive_score < 0, then the score of the document to be reviewed is set to 0; if comprehensive_score > 40, then the score of the document to be reviewed is set to 40; then the score scoring_result.comprehensive_score of the document to be reviewed is written back to the database.

[0237] Final project approval score and subsequent decision-making support. At this time, the system already has innovation_score, feasibility_score, and comprehensive_score. Add them up to get final_score.

[0238] Determine the automatically output text suggestions according to the score range:

[0239] If final_score ≥ 80, then "Suggest project approval";

[0240] If 60 ≤ final_score < 80, then "Review after revision or offline defense is possible";

[0241] If final_score < 60, then "Project establishment is not recommended for the time being."

[0242] Exemplarily, use the following prompt words to let the model make a natural language description of the scoring result:

[0243] "You are an assistant for summarizing the review of scientific research project establishment. Given that the innovation score of this project is XX, the feasibility score is XX, the comprehensive evaluation score is XX, and the total score is XX, the recommended option for the given range is 'Recommended for project establishment / Review after revision / Not recommended for project establishment for the time being'. Please use 100-150 words to elaborate on the core advantages, risk points of this project, and how to proceed next."

[0244] The system combines the text output by the model with the score details and sends it to the management and the project proposer in the form of a final PDF or Word report for viewing.

[0245] RPA (Robotic Process Automation technology) automatically retrieves the email addresses of relevant personnel corresponding to the document to be reviewed according to the document ID to be reviewed, completes the distribution of PDF / Word, and records the report generation time and sending record.

[0246] Based on the innovation evaluation score and the feasibility evaluation score, analysis and summary are carried out through large language model technology to obtain modification suggestions for the document to be reviewed. Specifically, after obtaining the innovation evaluation score and the feasibility evaluation score, the following information is usually combined to further generate modification suggestions:

[0247] 1. Comprehensive dimension analysis;

[0248] 2. Innovation evaluation score and feasibility evaluation score;

[0249] If the innovation score is on the low side: the large language model will give modification suggestions such as "There are no obvious innovation points in some technical points, and new algorithms, theories, and experimental verifications can be added". If the feasibility score is on the low side: suggestions such as "Resources are insufficient or the technology is not yet mature, and experimental verification needs to be supplemented, cooperation units need to be increased, and equipment investment needs to be improved" will be prompted.

[0250] 3. Suggestions generated by the large language model;

[0251] 4. Information aggregation and analysis reports, review reports, and duplicate check reports.

[0252] Combining information such as the review report, duplicate check report (if the duplication degree is too high, it is recommended to make literature citation marks or modify the research idea), and information aggregation and analysis report (similar research has been done externally, and differential positioning is required), generate targeted modification items.

[0253] 5. Example;

[0254] 6. Suggestions;

[0255] If the feasibility score of a certain project is 60 points and the innovation score is 80 points, the modification suggestions given by the system may include:

[0256] (1) It is recommended to strengthen the cooperation between the experimental link and the actual application unit to reduce potential risks;

[0257] (2) It is recommended to increase the funding budget or external resource support to ensure the implementation of key technical links;

[0258] (3) In terms of the team composition, experts in a certain field can be added to make up for the weak technical links.

[0259] 7. Generate an evaluation report and send it;

[0260] RPA or the script automatically fills in the score, advantages / disadvantages analysis, and modification suggestions into the template, and pushes the evaluation report to the applicant / manager corresponding to the document to be reviewed via email or the system; the evaluation report can include recommended opinions: such as "recommended for project establishment", "postpone project establishment", and "needs to be re-evaluated after modification", etc., for management decision-making.

[0261] To sum up, combining large language models and robotic process automation technology can seamlessly connect the processes from document review, scoring to automatically generating modification suggestions and reports, greatly improving the review efficiency and user satisfaction.

[0262] Furthermore, after obtaining the modification suggestions for the document to be reviewed, through the formulated standardized evaluation report template, using the script or RPA robot, fill in the scoring results and analysis content into the report template to obtain the evaluation report of the document to be reviewed, and display the evaluation report of the document to be reviewed in the form of charts, and send the evaluation report to relevant personnel via email or the system, and clear suggestions are provided in the evaluation report of the document to be reviewed, such as "recommended for project establishment", "needs to be re-evaluated after modification" or "not recommended for project establishment". Among them, the evaluation report template includes at least one of the scoring summary, innovation evaluation score, feasibility evaluation score, advantages and disadvantages analysis results of the technical solution in the document to be reviewed, and suggestions. Using large language model technology to generate specific modification suggestions and reasons according to the review results and presenting them to users in natural language form improves the user experience. At the same time, through robotic process automation technology, it realizes the automatic generation and sending of document review reports, which is convenient for users and managers to consult and make decisions. This optimization makes the results and feedback of document review clearer and more valuable for reference, and improves user satisfaction.

[0263] In this embodiment, the key data information in the document to be reviewed is obtained through robotic process automation technology, natural language processing technology and large language model technology; based on the key data information, it is detected whether the document to be reviewed is a duplicate document, and the detection result is obtained, and the corresponding duplicate detection report is obtained based on the detection result; based on the key data information, the budget information in the document to be reviewed is reviewed to obtain a review report; information aggregation and analysis are performed based on the key data information to obtain an information aggregation and analysis report; based on the review report, the information aggregation and analysis report and the duplicate detection report, the document to be reviewed is quantitatively analyzed to obtain the score and modification suggestions for the document to be reviewed. The following technical effects are achieved:

[0264] 1. Improve the efficiency and accuracy of document review: Use robotic process automation technology to automatically obtain document content, avoiding the time and errors of manual operation. At the same time, use big language model technology to deeply understand the document content, and perform text summarization, information extraction, and similarity calculation through automation and in-depth analysis. The interaction and collaboration between robotic process automation technology and big language model technology improves the quality and reliability of document review, and the efficiency and accuracy of document review are also enhanced;

[0265] 2. Expand the scope and depth of document review: Use large language model technology to review documents from multiple dimensions, taking into account the content and format of the document, while paying attention to the semantics, meaning, and relevance of the document to business matters and workflows. Comprehensive and in-depth review expands the scope of document review, making the review results more comprehensive and accurate;

[0266] 3. Use large language model technology to generate specific modification suggestions and reasons based on the review results, and present them to users in natural language, which improves the user experience. At the same time, use robotic process automation technology to automatically generate and send document review reports, which is convenient for users and managers to review and make decisions. This optimization makes the results and feedback of document reviews clearer and more valuable for reference, thereby improving user satisfaction.

[0267] This embodiment solves the technical problems in the prior art that RPA and AI technologies cannot be fully utilized, the quality and reliability of document review are low, and specific modification suggestions and reasons cannot be given.

[0268] In a second aspect, an embodiment of the present invention further provides a document review device.

[0269] In one embodiment, referring to Figure 6 , Figure 6 Schematic diagram of the functional modules of an embodiment of the document review device of the present invention. Figure 6 As shown, the document review device includes:

[0270] An information acquisition module 10, configured to obtain key data information in a document to be reviewed through robotic process automation technology, natural language processing technology, and large language model technology;

[0271] A report generation module 20, configured to detect whether the document to be reviewed is a duplicate document based on the key data information, obtain a detection result, and obtain a corresponding duplicate check report based on the detection result; review the budget information in the document to be reviewed based on the key data information, and obtain a review report; perform information aggregation and analysis based on the key data information to obtain an information aggregation and analysis report;

[0272] A comprehensive analysis module 30, configured to perform quantitative analysis on the document to be reviewed based on the review report, the information aggregation and analysis report, and the duplicate check report, and obtain a score and modification suggestions for the document to be reviewed.

[0273] Optionally, in one embodiment, the information acquisition module 10 is configured to:

[0274] Obtain data information in the document to be reviewed through robotic process automation technology and natural language processing technology;

[0275] Perform information extraction on the data information through large language model technology to obtain key data information.

[0276] Optionally, in one embodiment, the report generation module 20 is configured to:

[0277] Retrieve documents in the first database based on the keywords included in the key data information, and select the top N documents in the relevance ranking as similar documents;

[0278] Calculate the similarity between the key data information and each of the similar documents;

[0279] Compare each of the similarities with a similarity threshold;

[0280] If each of the similarities is less than the similarity threshold, obtain a detection result that the document to be reviewed is not a duplicate document;

[0281] If there is any similarity greater than or equal to the similarity threshold, obtain a detection result that the document to be reviewed is a duplicate document.

[0282] Optionally, in one embodiment, the report generation module 20 is configured to:

[0283] Encode the key data information using large language model technology to obtain a first numerical vector;

[0284] Encode each of the similar documents using large language model technology to obtain multiple second numerical vectors;

[0285] Based on the first numerical vector and each of the second numerical vectors, calculate the similarity between the key data information and each similar document through the multi-dimensional cosine similarity formula.

[0286] Optionally, in one embodiment, the report generation module 20 is configured to:

[0287] When the detection result is that the document to be reviewed is a duplicate document, use large language model technology to analyze the sub-topic correlation between the document to be reviewed and each of the similar documents, and obtain the duplicate sub-topics, similar sub-topics, and different sub-topics between the document to be reviewed and each of the similar documents;

[0288] Use large language model technology to analyze and summarize the duplicate sub-topics, the similar sub-topics, and the different sub-topics to obtain a duplicate check report.

[0289] Optionally, in one embodiment, the report generation module 20 is configured to:

[0290] Based on the calculation basis included in the key data information, review the budget information in the document to be reviewed through large language model technology to obtain a review result;

[0291] If the review result is that there is at least one preset error type in the budget table, then based on the analysis result, use robotic process automation technology to obtain a review report;

[0292] Wherein, the preset error types include budget calculation errors, information omissions, format errors, logical errors, and other errors.

[0293] Optionally, in one embodiment, the report generation module 20 is configured to:

[0294] Based on the title, abstract, and keywords included in the key data information, use robotic process automation technology to retrieve the documents in the second database, and select the top N documents in the relevance ranking as relevant documents;

[0295] Through large language model technology, extract the title, abstract, and keywords of each of the relevant documents, and encode the extracted information to obtain multiple third numerical vectors;

[0296] Through large language model technology, encode the title, abstract, and keywords included in the key data information to obtain a fourth numerical vector;

[0297] Based on the fourth numerical vector and each of the third numerical vectors, the similarity between the key data information and each relevant document is calculated through the multi-dimensional cosine similarity formula;

[0298] Based on the similarity between the key data information and each relevant document, analysis and summary are performed through large language model technology to obtain an information aggregation and analysis report.

[0299] Optionally, in one embodiment, the comprehensive analysis module 30 is configured to:

[0300] Perform data cleaning on the review report, the information aggregation and analysis report, and the duplicate check report;

[0301] Based on the review report, the information aggregation and analysis report, and the duplicate check report after cleaning, through large language model technology, the innovation evaluation score and feasibility evaluation score of the document to be reviewed are obtained;

[0302] Based on the innovation evaluation score and feasibility evaluation score, the score of the document to be reviewed is calculated;

[0303] Based on the innovation evaluation score and feasibility evaluation score, analysis and summary are performed through large language model technology to obtain modification suggestions for the document to be reviewed.

[0304] Wherein, the function implementation of each module in the above document review device corresponds to each step in the above document review method embodiment, and its function and implementation process will not be elaborated here one by one.

[0305] In a third aspect, an embodiment of the present invention further provides an electronic device, the structure of which is as Figure 7 shown, including: a memory and a processor, and the processor is used to read and execute the computer program stored in the memory to implement the foregoing document review method.

[0306] In a fourth aspect, an embodiment of the present invention further provides a computer storage medium, in which computer executable instructions are stored, and when the computer executable instructions are executed, the foregoing document review method is implemented.

[0307] In a fifth aspect, an embodiment of the present invention provides a computer program product, which is stored in a storage medium, and the program product is executed by at least one processor to implement each process of the above document review method embodiment, and can achieve the same technical effect. To avoid repetition, it will not be elaborated here.

[0308] Finally, it should be noted that in some processes described in the embodiments of the present invention, multiple operations or steps appear in a specific order. However, it should be understood that these operations or steps may not be executed in the order in which they appear in the embodiments of the present invention or may be executed in parallel. The serial numbers of the operations are only used to distinguish different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations or steps may be executed in sequence or in parallel, and these operations or steps may be combined.

[0309] The foregoing are only preferred embodiments of the present invention and are not intended to limit the present invention. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art may still modify the technical solutions described in the foregoing embodiments or perform equivalent replacements for some of the technical features. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included in the protection scope of the present invention.

Claims

1. A document review method, characterized in that: The method comprises: Obtain key data information from documents to be reviewed through robotic process automation technology, natural language processing technology and big language model technology; Based on the key data information, detect whether the document to be reviewed is a duplicate document, obtain a detection result, and obtain a corresponding duplicate check report based on the detection result; Reviewing the budget information in the document to be reviewed based on the key data information to obtain a review report; Performing information aggregation and analysis based on the key data information to obtain an information aggregation and analysis report; Based on the review report, the information aggregation and analysis report and the duplicate checking report, the document to be reviewed is quantitatively analyzed to obtain a score and modification suggestions for the document to be reviewed.

2. The document review method according to claim 1, characterized in that: The key data information in the document to be reviewed is obtained by using robotic process automation technology, natural language processing technology and large language model technology, including: Obtain data information from documents to be reviewed through robotic process automation and natural language processing technology; The data information is extracted through large language model technology to obtain key data information.

3. The document review method according to claim 1, characterized in that: The detecting, based on the key data information, whether the document to be reviewed is a duplicate document and obtaining a detection result includes: Based on the keywords included in the key data information, documents in the first database are searched, and the top N documents in the relevance ranking are selected as similar documents; Calculating the similarity between the key data information and each of the similar documents; comparing each of the similarities with a similarity threshold; If each of the similarities is less than the similarity threshold, the detection result is that the document to be reviewed is not a duplicate document; If any of the similarities is greater than or equal to the similarity threshold, the detection result is that the document to be reviewed is a duplicate document.

4. The document review method according to claim 3, characterized in that: The step of calculating the similarity between the key data information and the data information in each similar document comprises: Encoding the key data information using a large language model technology to obtain a first numerical vector; Encoding each of the similar documents using a large language model technology to obtain a plurality of second numerical vectors; Based on the first numerical vector and each of the second numerical vectors, the similarity between the key data information and each similar document is calculated using a multi-dimensional cosine similarity formula.

5. The document review method according to claim 3, characterized in that: The obtaining of a corresponding duplicate checking report based on the detection result includes: When the detection result is that the document to be reviewed is a duplicate document, the large language model technology is used to analyze the sub-topic association between the document to be reviewed and each of the similar documents, and the duplicate sub-topics, similar sub-topics and different sub-topics between the document to be reviewed and each of the similar documents are obtained; The large language model technology is used to analyze and summarize the repeated sub-topics, the similar sub-topics and the different sub-topics to obtain a duplicate checking report.

6. The document review method according to claim 1, characterized in that: The step of reviewing the budget information in the document to be reviewed based on the key data information to obtain a review report includes: Based on the calculation basis included in the key data information, the budget information in the document to be reviewed is reviewed by using a large language model technology to obtain a review result; If the review result is that at least one preset error type exists in the budget table, a review report is obtained using robotic process automation technology based on the analysis result; The preset error types include budget calculation errors, information omissions, format errors, logical errors and other errors.

7. The document review method according to claim 1, characterized in that: The information aggregation and analysis is performed according to the key data information to obtain an information aggregation and analysis report, including: Based on the title, abstract and keywords included in the key data information, the documents in the second database are retrieved using robotic process automation technology, and the top N documents in the relevance ranking are selected as relevant documents; Extracting the title, abstract and keywords of each of the relevant documents by using a large language model technology, and encoding the extracted information to obtain a plurality of third numerical vectors; By using a large language model technology, the title, abstract and keywords included in the key data information are encoded to obtain a fourth numerical vector; Based on the fourth numerical vector and each of the third numerical vectors, the similarity between the key data information and each related document is calculated by a multi-dimensional cosine similarity formula; Based on the similarity between the key data information and each related document, analysis and summary are performed using large language model technology to obtain information aggregation and analysis reports.

8. The document review method according to claim 1, characterized in that: The quantitative analysis of the document to be reviewed based on the review report, the information aggregation and analysis report, and the duplicate checking report to obtain the score and modification suggestions for the document to be reviewed includes: Performing data cleaning on the review report, the information aggregation and analysis report, and the duplicate checking report; Based on the cleaned review report, the information aggregation and analysis report, and the duplicate check report, the innovation evaluation score and the feasibility evaluation score of the document to be reviewed are obtained through the large language model technology; Calculate based on the innovation assessment score and the feasibility assessment score to obtain a score for the document to be reviewed; Based on the innovation evaluation score and the feasibility evaluation score, analysis and summary are performed through large language model technology to obtain modification suggestions for the document to be reviewed.

9. A document review device, characterized in that: The device comprises: An information acquisition module is configured to acquire key data information in the document to be reviewed through robotic process automation technology, natural language processing technology and large language model technology; The report generation module is configured to detect whether the document to be reviewed is a duplicate document based on the key data information, obtain a detection result, and obtain a corresponding duplicate check report based on the detection result; review the budget information in the document to be reviewed based on the key data information to obtain a review report; and aggregate and analyze information based on the key data information to obtain an information aggregation and analysis report; The comprehensive analysis module is configured to perform quantitative analysis on the document to be reviewed based on the review report, the information aggregation and analysis report, and the duplicate checking report to obtain a score and modification suggestions for the document to be reviewed.

10. The document review device according to claim 9, characterized in that: The report generation module is specifically configured to: When the detection result is that the document to be reviewed is a duplicate document, the large language model technology is used to analyze the sub-topic association between the document to be reviewed and the similar documents to obtain the duplicate sub-topics, similar sub-topics and different sub-topics between the document to be reviewed and the similar documents; The large language model technology is used to analyze and summarize the repeated sub-topics, the similar sub-topics and the different sub-topics to obtain a duplicate checking report.

11. The document review device according to claim 9, characterized in that: The report generation module is specifically configured to: Based on the calculation basis included in the key data information, the budget information in the document to be reviewed is reviewed by using a large language model technology to obtain a review result; If the review result is that at least one preset error type exists in the budget table, a review report is obtained using robotic process automation technology based on the analysis result; The preset error types include budget calculation errors, information omissions, format errors, logical errors and other errors.

12. The document review device according to claim 9, characterized in that: The report generation module is specifically configured to: Based on the title, abstract and keywords included in the key data information, the documents in the second database are retrieved using robotic process automation technology, and the top N documents in the relevance ranking are selected as relevant documents; Extracting the title, abstract and keywords of each of the relevant documents by using a large language model technology, and encoding the extracted information to obtain a plurality of third numerical vectors; By using a large language model technology, the title, abstract and keywords included in the key data information are encoded to obtain a fourth numerical vector; Based on the fourth numerical vector and each of the third numerical vectors, the similarity between the key data information and each related document is calculated by a multi-dimensional cosine similarity formula; Based on the similarity between the key data information and each related document, analysis and summary are performed using large language model technology to obtain information aggregation and analysis reports.

13. A computer program product comprising a computer program / instructions, characterized in that When the computer program / instruction is executed by a processor, the steps of the document review method as described in any one of claims 1-8 are implemented.

14. An electronic device, characterized in that: include: Memory and processor; The processor is used to read and execute the computer program stored in the memory to implement the steps of the document review method as described in any one of claims 1-8.

15. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed, the steps of the document review method as described in any one of claims 1 to 8 are implemented.