Project file risk analysis method and system
By combining the advantages of large language model, thinking chain reasoning model and dialogue pre-training model, we identify the semantic vectors of project files, filter risk points and generate modification suggestions, and solve the problems of inefficiency of traditional methods and lack of targeted suggestions, and achieve efficient and accurate project file risk analysis and processing.
Patent Information
- Application Number
- CN202510486492.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2045-04-18
AI Technical Summary
Traditional project file processing methods rely on manual review, are inefficient and prone to information omission or misjudgment due to human factors, and cannot effectively identify potential risk points in project files, and the generated modification suggestions are not targeted and practical.
A large language model is used to identify the semantic vectors of project files, and the difference semantic vectors are screened by comparing with pre-generated standard semantic vectors. The thinking chain reasoning model is used for logical reasoning, and the dialogue pre-trained model is used to generate risk analysis text and modify suggestions text.
It realizes accurate identification of project file semantics, comprehensive analysis of risk points and accurate generation of modification suggestions, improves the intelligence and automation level of project file processing, reduces labor costs, and improves project quality.
Smart Images

Figure CN120012784A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of artificial intelligence technology, and specifically relates to a project file risk analysis method and system. Background Art
[0002] With the rapid development of information technology, the processing and analysis of project documents plays a vital role in many fields such as enterprise management, software development, and engineering construction. Project documents usually contain key information such as detailed description, planning, progress, and resource allocation of the project, and are the basis for project execution and monitoring. However, traditional project document processing methods mainly rely on manual review and analysis, which is not only inefficient, but also prone to information omissions or misjudgments due to human factors, thus bringing potential risks to the smooth implementation of the project. Although some text recognition and induction technologies have achieved certain results in the field of text processing, they still face many challenges in practical applications. Especially when processing text data such as project documents with complex structures and professional terms, the risk analysis ability is limited, and it is impossible to fully identify potential risk points in the document; the generation of modification suggestions lacks pertinence and practicality, and it is difficult to meet the actual needs of users.
[0003] These problems not only affect the efficiency and accuracy of project document processing, but also restrict the application and promotion of related technologies in a wider range of fields. Summary of the invention
[0004] In view of the above-mentioned deficiencies in the prior art, the present invention provides a project document risk analysis method and system to solve the above-mentioned technical problems.
[0005] In a first aspect, the present invention provides a project document risk analysis method, comprising: Obtaining a project file, and identifying a semantic vector of the project file using a large language model; Comparing the semantic vector with a pre-generated standard semantic vector, and screening out a difference semantic vector that is inconsistent with the standard semantic vector; Use the thinking chain reasoning model to perform logical reasoning on the semantic vector of the project file based on user questions to obtain the abnormal semantic vector; The difference semantic vector and the abnormal semantic vector are sequentially input into the dialogue pre-training model to obtain corresponding risk analysis text and modification suggestion text; The difference semantic vector and the abnormal semantic vector are converted into sentences, and the sentences, corresponding risk analysis texts and modification suggestion texts are written into a preset report template to obtain a risk analysis report.
[0006] In an optional implementation, obtaining a project file and identifying a semantic vector of the project file using a large language model includes: Preprocessing the project files, including parsing and converting the file formats; Use natural language processing technology to perform word segmentation and syntactic analysis on the pre-processed project files to obtain structured information; Performing deep semantic analysis on the structured information using a large language model to identify key entities and relationships between key entities in the structured information; Semantic embedding technology is used to map key entities and their relationships into high-dimensional vector space to obtain semantic vectors.
[0007] In an optional implementation, the semantic vector is compared with a pre-generated standard semantic vector to screen out a difference semantic vector that is inconsistent with the standard semantic vector, including: Pre-building a standard library, wherein the standard library stores standard semantic vectors, wherein the standard semantic vectors are semantic vectors of legal entities and relationships mapped to high-dimensional vector spaces; Calculating the cosine similarity between the semantic vector and a standard semantic vector in a standard library; Determine whether there is a standard semantic vector in the standard library whose cosine similarity with the semantic vector reaches a set similarity threshold: If yes, it is determined that the semantic vector is normal; If not, it is determined that the semantic vector is a difference semantic vector.
[0008] In an optional implementation, the thought chain reasoning model is used to perform logical reasoning on the semantic vector of the project file based on the user's question to obtain an abnormal semantic vector, including: Use machine learning or deep learning techniques to identify project scenarios and user needs; Retrieving a matching target reasoning path from predefined reasoning paths based on the project scenario and user requirements; Identify the hierarchical structure of the semantic vector of the project file, and decompose the user question into multiple sub-questions based on the hierarchical structure; The target reasoning path is used to verify the sub-problems corresponding to each level and identify abnormal semantic vectors.
[0009] In an optional embodiment, the method further comprises: Use contrastive learning technology to automatically label unlabeled project data; After confirming that laws and regulations have been updated, the updated content is automatically extracted, and the new regulatory information is compressed and migrated to the existing model through knowledge distillation technology based on the Teacher-Student Model, so that it can be seamlessly integrated into the existing system while maintaining the model's reasoning ability and stability; Through the active learning mechanism, based on the confidence of the large model in the query, sample data for further annotation and learning is automatically selected to improve the learning efficiency of the model; Expand the annotated sample library, extract representative samples from unlabeled data, improve the accuracy of comparative learning, and enhance the understanding of complex regulations by analyzing the applicability of different regulations and policies.
[0010] In an optional implementation, the difference semantic vector and the abnormal semantic vector are sequentially input into the dialogue pre-training model to obtain corresponding risk analysis text and modification suggestion text, including: Generate prompt engineering based on difference semantic vectors and abnormal semantic vectors; The prompt project is input into the dialogue pre-training model to obtain corresponding risk analysis text and modification suggestion text.
[0011] In an optional embodiment, the method further comprises: The vocabulary selection, structure and language expression of the prompting project are adjusted based on the answer instructions fed back by the user.
[0012] In a second aspect, the present invention provides a project file risk analysis system, comprising: A semantic recognition module, used for acquiring a project file and identifying a semantic vector of the project file using a large language model; A vector comparison module, used to compare the semantic vector with a pre-generated standard semantic vector, and screen out a difference semantic vector that is inconsistent with the standard semantic vector; The logical reasoning module is used to use the thinking chain reasoning model to perform logical reasoning on the semantic vector of the project file based on the user's question to obtain the abnormal semantic vector; A risk analysis module, used to input the difference semantic vector and the abnormal semantic vector into the dialogue pre-training model in sequence to obtain corresponding risk analysis text and modification suggestion text; The report generation module is used to convert the difference semantic vector and the abnormal semantic vector into a statement, write the statement and the corresponding risk analysis text and modification suggestion text into a preset report template, and obtain a risk analysis report.
[0013] The beneficial effect of the present invention is that the project file risk analysis method and system provided by the present invention, by combining the advantages of a large language model, a thinking chain reasoning model and a dialogue pre-training model, can achieve accurate recognition of project file semantics, comprehensive analysis of risk points and accurate generation of modification suggestions. Through the implementation of the present invention, not only can the intelligent and automated level of project file processing be improved, but also more efficient and convenient file processing tools can be provided for enterprise managers and project team members, thereby promoting the sustainable development and innovation of related industries.
[0014] The beneficial effects specifically include the following aspects: 1. Improve the efficiency of risk identification: Traditional project file risk analysis mainly relies on manual review, which is not only inefficient but also easy to miss. The present invention greatly improves the efficiency of risk identification by automatically acquiring project files and using a large language model to quickly identify semantic vectors. 2. Enhance the accuracy of risk identification: The present invention can accurately screen out differential semantic vectors that are inconsistent with the standard by comparing the pre-generated standard semantic vectors. At the same time, by using the thinking chain reasoning model for logical reasoning, it can further discover potential abnormal semantic vectors, thereby enhancing the accuracy of risk identification. 3. Provide comprehensive risk analysis: The present invention not only identifies differential semantic vectors and abnormal semantic vectors, but also generates corresponding risk analysis texts and modification suggestion texts through the dialogue pre-training model. This provides users with comprehensive risk analysis information, which helps users better understand the risk points in project files and their impacts. 4. Reduce labor costs: Since the present invention realizes automated risk analysis, the workload of manual review is greatly reduced, thereby reducing labor costs. At the same time, the automated analysis process also avoids misjudgments or missed detections caused by human factors. 5. Improve project quality: By timely identifying and processing risk points in project files, the present invention helps to improve the overall quality of the project. During the project execution process, early discovery and resolution of potential problems significantly reduces the cost of later modifications and rework, ensuring the smooth progress of the project. 6. Promote technological innovation and application: The implementation of the present invention will promote the innovation and application of related technologies in the field of project risk management. By combining a variety of advanced models and technical means, the present invention provides a new solution for project risk management, which is expected to lead the technological progress of the industry. In summary, the present invention realizes comprehensive, efficient and accurate risk analysis of project documents through intelligent methods, bringing significant technological progress and economic benefits to the field of project management. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] In order to more clearly illustrate the embodiments of the present invention or the technical solutions in the prior art, the drawings used in the embodiments or the description of the prior art will be briefly introduced below. Obviously, for ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.
[0016] Figure 1 is a schematic flow chart of a method according to an embodiment of the present invention.
[0017] Figure 2 is a schematic block diagram of a system according to an embodiment of the present invention. DETAILED DESCRIPTION
[0018] In order to enable those skilled in the art to better understand the technical solutions in the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work should fall within the scope of protection of the present invention.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as those commonly understood by those skilled in the art of the present invention. The terms used in the specification of the present invention herein are only for the purpose of describing specific embodiments and are not intended to limit the present invention.
[0020] The key terms appearing in the present invention are explained below.
[0021] Large Language Models (LLMs) are artificial intelligence models trained on large amounts of text data, designed to understand and generate human language. The following is a detailed introduction to large language models: 1. Features Large Parameters: Large language models typically contain billions or even tens of billions of parameters, which enables them to learn rich language features and patterns.
[0022] Deep learning architectures: They are usually based on deep neural networks, such as the Transformer architecture, which includes self-attention mechanisms and is able to handle long-range dependencies.
[0023] Pre-training capability: Pre-training on large amounts of text data to learn universal representations of language, which enables the model to generalize to a variety of different tasks.
[0024] Fine-tuning flexibility: Fine-tuning on specific tasks to adapt to different application scenarios, such as translation, summarization, question answering, etc.
[0025] Contextual Understanding: Ability to understand the context of input text and generate coherent and relevant output.
[0026] Multi-task learning: Some large models can handle multiple language tasks and show a certain degree of versatility.
[0027] Generative capabilities: In addition to understanding language, many large models are also able to generate coherent and grammatically correct text.
[0028] 2. Basic composition Word embeddings: Convert words into continuous vectors so that neural networks can process them. The words represented by the vectors contain semantic information, making similar words closer in the vector space. Typical methods include Word2Vec, GloVe, BERT, etc.
[0029] Encoder-Decoder Architecture: Many large models use the Transformer architecture, which consists of an encoder and a decoder. The encoder converts the input text into an internal representation, and the decoder converts the internal representation into output text. A typical architecture such as the Transformer model contains multiple layers of encoders and decoders, each with a self-attention mechanism and a feed-forward neural network.
[0030] Self-Attention Mechanism: When processing an input sequence, the model focuses on different parts of the sequence and understands the dependencies between words. It processes all words in the sequence in parallel to improve computational efficiency.
[0031] Feedforward Neural Networks: In each layer of the transformer, a feedforward neural network is used to further process and transform the encoded representation. The structure is usually a fully connected layer with an activation function (such as ReLU).
[0032] Positional encoding: Since the transformer architecture has no order information, positional encoding is added to the word embedding to provide the position information of each word in the sequence. Implementations include fixed positional encoding or trainable positional encoding generated by sine and cosine functions.
[0033] Loss function: measures the gap between the model output and the actual target, and is used to guide the update of model parameters. Common types such as the cross-entropy loss function are commonly used in language models.
[0034] Optimization algorithm: According to the feedback of the loss function, adjust the model parameters to minimize the loss. Common methods include Adam, SGD (stochastic gradient descent), etc.
[0035] The Chain of Thought (CoT) is an advanced prompting technology that aims to improve the reasoning ability of large language models (LLMs) by simulating the human thinking process. The following is a detailed explanation of the Chain of Thought: The Thinking Chain Reasoning Model simulates the way humans think by breaking down complex problems into simpler sub-problems and inserting intermediate reasoning steps. Its core principle can be summarized as "simplifying the complex and solving them one by one." Specifically, when faced with a complex problem, the model will gradually break down the problem, then solve these simple problems one by one, and finally get the final answer.
[0036] The implementation of the mind chain reasoning model is usually combined with specific AI technologies and frameworks. For example, the LangChain framework has implemented the Plan-and-Executor mechanism and the ReAct mechanism, both of which are used to build the mind chain reasoning model. In specific applications, developers guide the model to show its reasoning process by designing precise prompts.
[0037] The conversation pre-training model is pre-trained through a large-scale corpus to enable the model to have intelligent and natural conversation capabilities. The basic principle is to enable the model to understand the grammar, semantics and contextual information of natural language by learning from a large corpus, thereby achieving efficient understanding and generation of natural language. This type of model usually adopts a Transformer-based deep learning architecture, such as BERT, GPT, etc. Transformer achieves a high-level abstract representation of natural language through a self-attention mechanism and a multi-layer superimposed network structure, which can capture long-distance dependencies in language, thereby more accurately understanding the context of the conversation.
[0038] The project file risk analysis method provided in the embodiment of the present invention is executed by a computer device, and accordingly, the project file risk analysis system runs in the computer device.
[0039] Figure 1 is a schematic flow chart of a method according to an embodiment of the present invention. Figure 1 The execution subject is a project document risk analysis system. According to different requirements, the order of the steps in the flowchart is changed and some are omitted.
[0040] like Figure 1 As shown, the method includes: S1. Obtain a project file and use a large language model to identify a semantic vector of the project file.
[0041] Get project files from the specified source. These files include documents, codes, data tables and other formats. Then, use large language models (such as BERT, GPT, etc.) to perform semantic analysis on project files. Large language models can deeply understand the text content and convert it into high-dimensional semantic vectors. These semantic vectors are mathematical representations of text content, which can capture key information and contextual relationships in the text, providing a basis for subsequent comparison and reasoning.
[0042] S2. Compare the semantic vector with a pre-generated standard semantic vector, and filter out the difference semantic vector that is inconsistent with the standard semantic vector.
[0043] After obtaining the semantic vector of the project file, it is then compared with the pre-defined standard semantic vector. The standard semantic vector is generated based on industry specifications, project requirements, or historical data, and represents the correct or expected semantic features that the project file should have. By calculating the similarity or distance between the semantic vector of the project file and the standard semantic vector, the difference semantic vectors that are inconsistent with the standard are screened out. These difference semantic vectors represent problems or parts that do not meet the requirements in the project file.
[0044] S3. Use the thinking chain reasoning model to perform logical reasoning on the semantic vector of the project file based on user questions to obtain the abnormal semantic vector.
[0045] After obtaining the difference semantic vector, in order to have a deeper understanding of the problem, the thinking chain reasoning model is used for logical reasoning. The thinking chain reasoning model can simulate the human thinking process, gradually dismantle the problem and solve it one by one, and finally obtain the abnormal semantic vector. These abnormal semantic vectors are obtained by logical reasoning based on the questions raised by the user and the semantic vectors of the project file, and represent the logical errors, inconsistencies or other problems in the project file.
[0046] S4. Input the difference semantic vector and the abnormal semantic vector into the dialogue pre-training model in sequence to obtain the corresponding risk analysis text and modification suggestion text.
[0047] After obtaining the difference semantic vector and the anomaly semantic vector, they are then input into the dialogue pre-training model. The dialogue pre-training model can generate corresponding risk analysis text and modification suggestion text based on these semantic vectors. The risk analysis text will elaborate on the risks and problems caused by the differences and anomalies, while the modification suggestion text will provide targeted solutions or modification suggestions. These texts will help users better understand the problem and take appropriate measures to solve it.
[0048] S5. Convert the difference semantic vector and the abnormal semantic vector into a statement, write the statement and the corresponding risk analysis text and modification suggestion text into a preset report template, and obtain a risk analysis report.
[0049] Convert the difference semantic vectors and abnormal semantic vectors into easy-to-understand statements, and write these statements into the pre-set report template together with the risk analysis text and modification suggestion text. The report template is customized according to the specific needs and format requirements of the project to ensure that the generated risk analysis report has a clear structure and easy-to-understand content. After completing this step, users will get a detailed risk analysis report that will help them better understand the problems and risks in the project files and take appropriate measures to solve them.
[0050] The present invention mainly solves the problem of how to efficiently and accurately identify the difference semantic vectors and abnormal semantic vectors that are inconsistent with the standard semantics during the project file analysis process, and generate corresponding risk analysis reports. By combining the large language model, the thinking chain reasoning model and the dialogue pre-training model, the present invention provides a new and intelligent project file risk analysis method.
[0051] In an embodiment of the present invention, based on step S1, an example is given below to illustrate its specific implementation scheme in a non-limiting manner.
[0052] S101. Preprocessing the project file, including parsing and converting the file format In this subdivision step, the project files obtained from the specified source are first processed. These files exist in various formats, such as Word documents, PDFs, Excel spreadsheets, code files, etc. For the convenience and consistency of subsequent processing, these files are pre-processed.
[0053] The main work of preprocessing includes parsing and converting file formats. For document files, such as Word and PDF, extract the text content and keep its original format and structure information, such as title, paragraph, list, etc. For data table files, such as Excel, read the data in the table and understand its row and column relationship and data type. For code files, parse its syntax structure and extract key code segments and comment information.
[0054] In the process of parsing and converting file formats, some professional tools or libraries are used, such as PDF parsing library, Excel reading library, code parser, etc. These tools can help to extract and convert file content more accurately and provide high-quality data input for subsequent processing steps.
[0055] S102. Use natural language processing technology to perform word segmentation and syntactic analysis on the pre-processed project files to obtain structured information After obtaining the preprocessed project file, natural language processing technology (NLP) is used to perform word segmentation and syntactic analysis on the text content.
[0056] Word segmentation is one of the basic tasks in NLP. It can segment continuous text into independent words or phrases. For Chinese text, word segmentation is a relatively complex process because there is no clear space between Chinese words. With the help of word segmentation tools or algorithms, such as rule-based word segmentation, statistical-based word segmentation or deep learning word segmentation models, text can be segmented accurately.
[0057] Syntactic analysis is an important step in understanding the grammatical structure of a text. Through syntactic analysis, sentence components, phrase types, and grammatical relationships in a text can be identified. This information is crucial for the subsequent extraction of structured information and understanding of text semantics.
[0058] On the basis of word segmentation and syntactic analysis, we further extract structured information from the text, such as named entities (names of people, places, institutions, etc.), time, numbers, keywords, etc. This information will be represented in a structured form to facilitate subsequent processing steps.
[0059] S103. Performing deep semantic analysis on the structured information using a large language model to identify key entities and relationships between key entities in the structured information After obtaining the structured information, the large language model is further used to perform deep semantic analysis. The large language model has powerful language understanding and generation capabilities and can capture the deep semantic information in the text.
[0060] In this step, the structured information is input into the large language model, and the model is guided to identify the key entities and the relationships between key entities. Key entities are core nouns or phrases in the text, such as names of people, places, institutions, concepts, etc. The relationships between key entities describe the interactions or connections between these entities, such as the relationships between characters, the causal relationships between events, etc.
[0061] Through deep semantic analysis, we can understand the text content more accurately and extract valuable information for subsequent processing. This information will provide support for subsequent comparison, reasoning and decision-making.
[0062] S104. Using semantic embedding technology, map key entities and their relationships into high-dimensional vector space to obtain semantic vectors After obtaining the key entities and their relationships, semantic embedding technology is used to map this information into a high-dimensional vector space. Semantic embedding technology can convert text or text fragments into high-dimensional vector representations that can capture the semantic information and contextual relationships in the text.
[0063] In this step, key entities and their relationships are used as input, and semantic embedding algorithms (such as Word2Vec, BERT, etc.) are used to generate corresponding semantic vectors. These vectors will serve as mathematical representations of text content and provide a basis for subsequent tasks such as comparison, classification, and clustering.
[0064] Through semantic embedding technology, text information can be converted into a computable numerical form, which is convenient for subsequent processing and analysis. At the same time, these semantic vectors can also retain the key information and contextual relationships in the text, providing strong support for subsequent comparison and reasoning.
[0065] In summary, through the steps of preprocessing, word segmentation, syntactic analysis, deep semantic parsing and semantic embedding of project files, we can gradually extract and transform text content and finally obtain high-quality semantic vector representation. These semantic vectors will serve as the basis for subsequent comparison, reasoning and decision-making, providing strong support for the smooth progress of the project.
[0066] In an embodiment of the present invention, based on step S2, an example is given below to illustrate its specific implementation scheme in a non-limiting manner.
[0067] Pre-built standard libraries A key step in the project processing flow is to pre-build a standard library. This standard library is a knowledge base that stores standard semantic vectors, which are obtained by mapping legal entities and relationships into high-dimensional vector space. This process is similar to the semantic embedding technology mentioned earlier, but here, it focuses on text content related to laws and regulations.
[0068] When building the standard library, we first collect a large number of legal texts, such as laws and regulations, judicial interpretations, and cases. Then, we use natural language processing technology to perform word segmentation, syntactic analysis, and structured information extraction on these texts. Next, we use a large language model to perform deep semantic analysis on these structured information to identify legal entities (such as legal names, institution names, personal names, etc.) and legal relationships (such as rights and obligations, causal relationships, etc.). Finally, we use semantic embedding technology to map these legal entities and relationships into a high-dimensional vector space to obtain standard semantic vectors, which are then stored in the standard library.
[0069] Calculate the cosine similarity between the semantic vector and the standard semantic vector in the standard library After obtaining the semantic vectors of the project files, the cosine similarity between these semantic vectors and the standard semantic vectors in the standard library is calculated. Cosine similarity is a commonly used indicator to measure the similarity between two vectors. Its value range is [-1,1], where 1 means that the two vectors are completely similar, -1 means that the two vectors are completely opposite, and 0 means that the two vectors have no correlation.
[0070] When calculating cosine similarity, the dot product of the vectors is divided by the product of their moduli. Specifically, for two vectors A and B, their cosine similarity is calculated by the following formula: cos(θ) = (A·B) / (|A| * |B|) Where A·B represents the dot product of vectors A and B, and |A| and |B| represent the modulus of vectors A and B, respectively.
[0071] Determine whether there is a standard semantic vector in the standard library whose cosine similarity with the semantic vector reaches a set similarity threshold After calculating the cosine similarity between the semantic vector and all standard semantic vectors in the standard library, it is determined whether there are one or more standard semantic vectors whose cosine similarity with the semantic vector reaches a set similarity threshold. This similarity threshold is a value set according to actual needs, which determines the standard for considering whether two vectors are sufficiently similar.
[0072] In actual operation, the calculated cosine similarity is compared with a set similarity threshold. If there are one or more standard semantic vectors whose cosine similarity with the semantic vector is greater than or equal to the similarity threshold, it is considered that there are standard semantic vectors similar to the semantic vector in the standard library.
[0073] Determine whether the semantic vector is normal According to the judgment result of the previous step, the semantic vector is judged. If there is a standard semantic vector in the standard library whose cosine similarity with the semantic vector reaches the set similarity threshold, the semantic vector is judged to be normal, that is, it meets the expected legal semantic characteristics. This means that the relevant text content in the project file is semantically similar to the legal text in the standard library, and there is no significant difference.
[0074] However, if there is no standard semantic vector in the standard library whose cosine similarity with the semantic vector reaches the set similarity threshold, the semantic vector is determined to be a difference semantic vector. This means that there is a significant difference in semantics between the relevant text content in the project file and the legal text in the standard library, and further review and analysis is required.
[0075] In summary, by pre-building a standard library, calculating the cosine similarity between the semantic vector and the standard semantic vector, judging whether the similarity reaches the set threshold, and determining whether the semantic vector is normal, the legal text in the project document can be semantically analyzed and compared effectively, thereby discovering the differences and problems therein. This is of great significance for ensuring the legality and compliance of the project document.
[0076] In an embodiment of the present invention, based on step S3, an example is given below to illustrate its specific implementation scheme in a non-limiting manner.
[0077] S301. Use machine learning or deep learning techniques to identify project scenarios and user needs In the project processing process, a crucial step is to accurately identify the project scenario and user needs. This is not only the basis for subsequent reasoning and verification, but also the key to ensuring that the final project results meet user expectations. To achieve this goal, machine learning or deep learning technology is used to assist in identification.
[0078] Specifically, a neural network model is trained to identify different project scenarios, such as contract review, legal consultation, compliance check, etc. This neural network model makes judgments by analyzing the text content, keywords, format and other information of the project file. At the same time, it also uses the questions or descriptions entered by the user to further refine the scenario recognition, so as to more accurately grasp the user's needs.
[0079] In order to train such a model, a large amount of project files and user problem data are collected, annotated and preprocessed. During the training process, model parameters and feature selection are continuously adjusted to improve the recognition accuracy and generalization ability of the model.
[0080] S302. Retrieving a matching target reasoning path from a predefined reasoning path based on the project scenario and user needs After identifying the project scenario and user needs, the matching target reasoning path is retrieved from the predefined reasoning paths. These reasoning paths are predefined based on industry knowledge, laws and regulations, historical experience, etc., and are used to guide the subsequent problem disassembly and reasoning verification process.
[0081] To achieve this goal, a reasoning path library is built, which contains multiple reasoning paths for different project scenarios and user needs. After identifying the project scenario and user needs, the reasoning path library is searched and matched based on this information to find the target reasoning path that best suits the current situation.
[0082] When building the inference path library, we ensure the accuracy and completeness of the inference path. This requires in-depth research and understanding of industry knowledge, laws and regulations, etc., and continuous optimization and improvement based on historical experience. At the same time, we also regularly update the inference path library to adapt to changes in industry development and user needs.
[0083] S303. Identify the hierarchical structure of the semantic vector of the project file, and decompose the user question into multiple sub-questions based on the hierarchical structure After obtaining the semantic vector of the project file, we further identify its hierarchical structure, which helps us to have a deeper understanding of the content and organizational structure of the project file, and provides strong support for subsequent problem analysis and reasoning verification.
[0084] Specifically, natural language processing technology is used to hierarchically divide the semantic vectors of the project file, such as paragraph division, sentence division, phrase division, etc. Then, based on this hierarchical information, the user question is broken down into multiple sub-questions. These sub-questions are inquiries or verifications for a specific paragraph, sentence or phrase in the project file.
[0085] When breaking down user problems, ensure the accuracy and independence of sub-problems. This allows for in-depth analysis and understanding of user problems, and for reasonable division in combination with the hierarchical structure of project files. At the same time, it also ensures that the logical relationship between sub-problems is clear and explicit, so that subsequent reasoning verification and integration can be performed.
[0086] S304. Using the target reasoning path, perform reasoning verification on the sub-problems corresponding to each level and identify abnormal semantic vectors After obtaining the decomposed sub-questions, they are reasoned and verified using the target reasoning path. This process aims to check whether the sub-questions meet the expected legal semantic features or industry standards and identify abnormal semantic vectors.
[0087] Specifically, the sub-questions are verified one by one according to the steps and rules in the target reasoning path. This includes checking the semantic consistency, legality, and compliance of the sub-questions. During the verification process, relevant laws and regulations, industry specifications, historical cases, and other information are referenced to assist in judgment.
[0088] If an abnormal semantic vector is found in a sub-problem during the verification process, that is, its semantic features do not match expectations or there are obvious errors, it will be marked and recorded in time. These abnormal semantic vectors represent potential problems or risk points in the project file for further review and analysis. At the same time, necessary adjustments and optimizations are made to the target reasoning path based on the verification results to improve its accuracy and applicability.
[0089] In summary, by using machine learning or deep learning technology to identify project scenarios and user needs, retrieving matching target reasoning paths from predefined reasoning paths, identifying the hierarchical structure of the semantic vectors of project files and breaking down user questions into multiple sub-questions, and using the target reasoning path to perform reasoning verification on sub-questions and identify abnormal semantic vectors, the legal texts in project files can be more effectively semantically analyzed and compared, thereby ensuring the legality and compliance of the project.
[0090] In an embodiment of the present invention, based on step S4, an example is given below to illustrate its specific implementation scheme in a non-limiting manner.
[0091] A prompt project is generated based on the difference semantic vector and the abnormal semantic vector; the prompt project is input into the dialogue pre-training model to obtain a corresponding risk analysis text and a modification suggestion text.
[0092] During the compliance review process, a series of difference semantic vectors that are inconsistent with the legal semantic vectors and abnormal semantic vectors identified by thought chain reasoning are obtained through semantic vector comparison and difference screening. These vectors represent the compliance issues existing in the project files. In order to present these issues more specifically and generate corresponding risk analysis texts and modification suggestion texts, a prompt project is built based on these difference and abnormal semantic vectors and input into the dialogue pre-training model.
[0093] Sample content 1. Differential semantic vector example and its prompt engineering construction Example of differential semantic vector: Vector 1: The “Environmental Assessment Report” mentioned in the project documents does not explicitly include a “Noise Pollution Assessment” section.
[0094] Vector 2: There are subtle differences between the "liability for breach of contract" clause in the contract and the description of liability for breach of contract in relevant laws and regulations.
[0095] Prompt project construction: Tip 1: "Please analyze whether the 'Environmental Assessment Report' in the project documents is complete, especially whether the 'Noise Pollution Assessment' section is missing, and provide a risk analysis." Tip 2: "Please compare the 'liability for breach of contract' clause in the contract with relevant laws and regulations, point out the differences, and assess the risks." 2. Abnormal semantic vector example and its prompt engineering construction Example of abnormal semantic vector: Vector 1: In the "Fund Usage Report" in the project file, the description of the use of a certain fund does not match the actual use.
[0096] Vector 2: There is a logical contradiction in the "project schedule" in the project plan, such as the completion time of a task in a certain stage is earlier than the start time of the predecessor task.
[0097] Prompt project construction: Tip 1: "Please review the description of the use of a certain fund in the 'Fund Use Report' to confirm whether it is consistent with the actual use and provide modification suggestions." Tip 2: "Please check the 'Project Timeline' in the project plan, identify any logical inconsistencies, and provide suggestions for adjustments." 3. Application and output of dialogue pre-training model Input the above-constructed prompt projects into the dialogue pre-training model (such as ChatGLM3-6B-32k, etc.) in sequence, and the model will generate corresponding risk analysis text and modification suggestion text.
[0098] Example of risk analysis text: For the missing section of the “Environmental Assessment Report,” the model generates the following risk analysis: “Environmental assessment reports that do not include the ‘Noise Pollution Assessment’ section cause projects to be blocked at the environmental approval stage, increasing the risk of project delays or cancellation.” Regarding the differences in the "liability for breach of contract" clauses in the contract, the model points out: "There are subtle differences between the 'liability for breach of contract' clauses in the contract and the laws and regulations, which lead to adverse consequences in dispute resolution." Example of modification suggestion text: Regarding the problem of inconsistent description of fund use, the model suggests: "It is recommended to update the 'Fund Use Report' to ensure that the description of fund use is consistent with the actual use, so as to avoid problems in subsequent audits or inspections." Regarding the problem of logical contradictions in the project timetable, the model suggests: "It is recommended to adjust the 'project timetable' in the project plan to ensure that the time sequence of tasks in each stage is reasonable and avoid logical contradictions." By generating prompts based on difference semantic vectors and abnormal semantic vectors and inputting them into the dialogue pre-training model, specific risk analysis texts and modification suggestion texts are obtained. These texts provide clear guidance to the project team, helping them to quickly identify and resolve compliance issues and reduce project risks.
[0099] In an embodiment of the present invention, based on step S5, an example is given below to illustrate its specific implementation scheme in a non-limiting manner.
[0100] Difference and exception semantic vector conversion: The difference semantic vector and exception semantic vector are converted into easy-to-understand sentence forms for presentation in reports.
[0101] Report template setting: Pre-set the report template, which contains the basic structure, format and style requirements of the report.
[0102] Filling and generating report content: Write the converted statements, risk analysis text and modification suggestion text into the pre-set report template to generate a complete risk analysis report.
[0103] Report output and sharing: Output the generated risk analysis report as PDF, Word or other formats, and share it with the project team, management or relevant stakeholders for review and decision-making.
[0104] On the basis of the above embodiments, in order to further improve the performance of the model provided by the above embodiments, an implementable method is to use contrastive learning technology to automatically annotate data, and use continuous learning and knowledge distillation technology to dynamically update model parameters, so as to ensure that the system can adapt to the latest laws, regulations and policy changes and maintain efficient compliance review capabilities.
[0105] Application of contrastive learning technology: The system uses contrastive learning technology to automatically annotate unannotated project data. By comparing the similarities and differences in project documents, key content is identified and the corresponding legal clauses or policy requirements are annotated. Not only can obvious compliance risk points be identified, but also project content with hidden risks can be extracted from potential semantic differences.
[0106] Continuous learning mechanism: The system dynamically updates the large model to absorb changes in new regulations and policies through a continuous learning mechanism. It automatically extracts the updated parts and integrates them into the existing model through small batch online learning technology. It is suitable for new models and new requirements in legal and regulatory updates and project management.
[0107] Application of knowledge distillation technology: During the update process, new regulations and policies are integrated into the existing model structure using knowledge distillation technology. This ensures that old knowledge is not forgotten while learning new knowledge, and keeps the model performance stable. Through knowledge distillation technology based on the Teacher-Student Model, the new policy content is compressed into a lightweight model to reduce computing resource consumption.
[0108] Active learning and sample expansion: The system automatically selects sample data for further annotation and learning through an active learning mechanism. It prioritizes annotation and learning of legal issues with high uncertainty, reduces annotation workload and improves compliance judgment capabilities. It automatically expands the annotated sample library to improve the accuracy of comparative learning and understanding of complex regulations.
[0109] Regular model updates and performance monitoring: The system sets up a regular update mechanism to retrieve the latest regulations and policy content from the legal database and retrain the large model. The model performance monitoring module monitors the performance of the large model in different compliance review tasks in real time. Dynamically adjust model parameters based on reasoning accuracy, response time and user feedback.
[0110] Real-time regulatory updates and adaptation: The system can automatically retrieve and update the corresponding legal terms according to real-time regulatory changes. Automatically learn new regulations, update the knowledge base and adjust the model reasoning path. Real-time connection with the legal and regulatory database ensures that the model can cope with the challenges of the latest policies.
[0111] Improved compliance review efficiency: The efficiency of compliance review is significantly improved through the combination of comparative learning and continuous learning. Quickly respond to policy changes and provide immediate compliance review suggestions. Quickly adapt to new regulations with fewer labeled samples, maintain efficient and accurate reasoning capabilities, and reduce manual intervention.
[0112] Together, these technologies form the core framework for system compliance review, ensuring that the system can maintain efficient, accurate and real-time compliance review capabilities in the face of an ever-changing regulatory and policy environment.
[0113] In addition, in order to further improve the matching degree between the answers of the dialogue pre-training model and user needs, the prompt engineering is constantly adjusted. The specific adjustment methods include: 1. Automatic Prompt Generation Basics Based on user questions or needs, the system automatically analyzes the task context, combines project documents with legal requirements, and generates the most suitable prompt. Through natural language understanding technology, the system extracts the core information input by the user, such as project stages and review questions, to ensure that the generated prompt is highly targeted and can guide the large model to accurately understand user needs.
[0114] 2. Prompt Engineering and Application Prompt engineering, or prompt learning, guides the model to correctly understand and process input data by providing text information (prompt) in a specific format. Prompts come in various types, including descriptions, examples, guiding questions, etc., suitable for different tasks and data sets. In application scenarios such as text classification, question-answering systems, and machine translation, prompts can significantly improve the performance and interpretability of the model.
[0115] 3. Dynamic Prompt Optimization and Customization Dynamic optimization: The system dynamically adjusts the prompt structure based on user feedback and the quality of the model's answers, optimizing language expression, vocabulary selection, and structure to ensure that the prompt can better guide the model to provide high-quality answers.
[0116] Scenario customization: The system generates customized prompts based on project scenarios and user-specific needs or roles. For example, the project review focuses on funding sources, approval process, etc., while the acceptance stage focuses on contract performance, delivery standards, etc.
[0117] 4. Semantic Context The system tracks user contextual input to ensure that the generated prompt is consistent with previous conversations or operations. This helps the model maintain an understanding of the context of the question in complex multi-round conversations and provide accurate compliance analysis or legal opinions.
[0118] 5. User problem optimization and conversion The system automatically analyzes the complexity and expression of user questions, and transforms complex, ambiguous or incomplete questions into clear questions. Through question optimization and transformation, it ensures that the big model can accurately understand the user's intention and generate high-quality answers when dealing with complex questions.
[0119] 6. Continuous optimization and feedback mechanism The system continuously optimizes the prompt generation strategy through feedback loops. Every time a user interacts with the model, it provides new data to the system to help it improve the quality of future prompt generation. Based on the feedback from the model output, the system adjusts the keywords, structure, and context information in the prompt to ensure that the answer meets the user's needs.
[0120] 7. Prompt learning based on feedback The system automatically learns the optimization direction of Prompt based on the user's rating and correction of the big model's answers. Through machine learning technology, the system accumulates a large amount of user feedback data, dynamically adjusts the Prompt generation algorithm, and improves the big model's question-answering capabilities.
[0121] 8. Combining Prompt with Model Reasoning Path When the system generates a prompt, it combines the current reasoning path to ensure that the reasoning logic of the prompt and the large model remain consistent. This combination reduces reasoning bias and improves the reliability and accuracy of the answer.
[0122] 9. Improve the relevance and accuracy of model responses Through intelligent prompt generation and optimization, the system has greatly improved the accuracy and pertinence of the big model's answers in different project scenarios. The system can automatically generate prompts for specific compliance reviews to ensure the professionalism of the big model's answers to detailed questions.
[0123] In summary, the automated prompt generation mechanism significantly improves the quality and pertinence of the large model's answers in different project scenarios through comprehensive analysis of user questions, dynamic optimization of prompts, customized generation, associating semantic context, optimizing and transforming user questions, continuous feedback and learning, and combining reasoning paths.
[0124] The following is a specific embodiment, which provides a project file risk analysis method, including: 1. Get project files and identify semantic vectors The project files are preprocessed, word segmented and syntactically analyzed, deep semantic analysis is performed using a large language model, and key entities and relationships are mapped to a high-dimensional vector space through semantic embedding technology.
[0125] Preprocessing: Let the project file be F, and get F′ after file format parsing and conversion.
[0126] Word segmentation and syntactic analysis: Use natural language processing technology to process F′ and obtain structured information S.
[0127] Deep semantic parsing: Use a large language model to parse S and identify the key entity set E={e1,e2,⋯,e n} and the key entity relationship set R = {r1,r2,⋯,r m}.
[0128] Semantic embedding: Use the semantic embedding function fembed to map E and R into a high-dimensional vector space to obtain a semantic vector v: v=fembed(E,R) 2. Filtering differential semantic vectors Calculate the cosine similarity between the project file semantic vector and the standard semantic vector, and determine whether it is a difference semantic vector based on the similarity threshold.
[0129] Suppose the set of standard semantic vectors in the standard library is {v std 1 ,v std 2 ,⋯,v std k}.
[0130] Calculate the project file semantic vector v and the standard semantic vector v std i The cosine similarity sim i :
[0131] Set the similarity threshold θ, if all sim i <θ, then v is judged to be a differential semantic vector.
[0132] 3. Inference to obtain abnormal semantic vector Retrieve the reasoning path based on the project scenario and user needs, break down the user problem into sub-problems, and perform reasoning and verification on the sub-problems corresponding to each level.
[0133] Use machine learning or deep learning technology to identify project scenes Sscene and user needs Uneed. In this embodiment, the LSTM model is used for identification: Data preprocessing: Perform word segmentation on the input text data and convert the text into a sequence of words. Build a vocabulary table and map each word to a corresponding integer index. Pad or truncate the input sequence to make it of the same length.
[0134] Model training: Input the preprocessed data into the LSTM model. Use the cross entropy loss function to calculate the loss between the predicted result and the true label; use an optimization algorithm (such as the Adam optimizer) to update the model parameters and minimize the loss function.
[0135] Prediction stage: Perform the same preprocessing on the new input text. Input the preprocessed text into the trained model. The model outputs the predicted category probability distribution, and selects the category with the highest probability as the final prediction result.
[0136] From the predefined reasoning path set P = {p1,p2,⋯,p l} to retrieve the matching target reasoning path p target .
[0137] Identify the hierarchical structure H of the semantic vector v and decompose the user question Q into a set of sub-questions Q sub ={q1,q2,⋯,q s}.
[0138] Use the target reasoning path ptarget to verify the reasoning of each sub-problem and identify abnormal semantic vectors: Rule-based reasoning is to match and deduce input facts according to a predefined set of rules to draw new conclusions. Rules are usually expressed in the form of "IF - THEN". When the condition part (IF) is met, the conclusion part (THEN) is executed.
[0139] Assume that the rule set R = {r1,r2,⋯,r n}, each rule r i Represented as a two-tuple (C i ,A i ), where C i is the conditional part, A i This is the conclusion part.
[0140] The inference function infer is defined as follows: def infer(facts, rules): new_facts = set(facts) while True: old_facts= set(new_facts) for rule in rules: condition, action = rule if all(fact innew_facts for fact in condition): new_facts.add(action) if new_facts == old_facts: break return new_facts.
[0141] The processing flow includes: Initialization: Take the initial fact set F as known information.
[0142] Rule matching: traverse the rule set R, and for each rule ri, check whether its condition part Ci is included in the current fact set.
[0143] Conclusion derivation: If the condition part of the rule is met, the conclusion part Ai is added to the fact set.
[0144] Termination condition: When the set of facts no longer changes, the reasoning process ends.
[0145] For example: # Define rule set rules = [ ({"The project has legal risks", "Legal risks have not been resolved"}, "The project has compliance issues"), ({"The project has compliance issues"}, "Rectification is needed") ] # Define initial fact set facts = {"The project has legal risks", "Legal risks have not been resolved"} # Perform reasoning result = infer(facts,rules) print("Inference result:", result).
[0146] 4. Generate risk analysis text and modification suggestion text Generate prompts based on the difference semantic vector and abnormal semantic vector, and input them into the dialogue pre-training model to obtain the results.
[0147] Let the difference semantic vector be v diff , the abnormal semantic vector is v anomaly , generate the prompt project Pprompt.
[0148] Input Pprompt into the dialogue pre-training model M dialogue , get the risk analysis text T risk and modify the suggested text Tsuggest:(T risk ,T suggest )=M dialogue (Pprompt).
[0149] 5. Generate risk analysis report The difference semantic vector and the abnormal semantic vector are converted into sentences, and written into the report template together with the risk analysis text and the modification suggestion text.
[0150] v diff and v anomaly Convert to statement S diff and S anomaly .
[0151] S diff , S anomaly , T risk and T suggest Writing report template T template , get the risk analysis report R report .
[0152] Automatic labeling: using contrastive learning technology contrast For unlabeled project data D unlabeled Perform automatic labeling to obtain labeled data Dlabeled: D labeled =f contrast (D unlabeled ).
[0153] Knowledge distillation: When laws and regulations are updated, extract the updated content C update , through the knowledge distillation technique of the teacher-student model distill Migrate new regulatory information to existing Model M old , and get the updated model M new :M new =f distill (M old ,C update ).
[0154] Active learning: Based on the confidence of the large model in the query, the active learning mechanism f active Select sample data D for further annotation and learning selected : D selected =f active (confidence).
[0155] Expand the labeled sample library: extract representative samples D from unlabeled data representative , expand the labeled sample library D labeled :D labeled =D labeled ∪Drepresentative .
[0156] Through the above embodiments, the data processing principle and process of the project file risk analysis method are explained in detail.
[0157] In some embodiments, the project document risk analysis system includes a plurality of functional modules composed of computer program segments. The computer program of each program segment in the project document risk analysis system is stored in the memory of the computer device and executed by at least one processor to perform (see Figure 1 Description) Function of project document risk analysis.
[0158] In this embodiment, the project file risk analysis system is divided into multiple functional modules according to the functions it performs, such as Figure 2 As shown. The functional modules of the system 200 include: a semantic recognition module 210, a vector comparison module 220, a logical reasoning module 230, a risk analysis module 240, and a report generation module 250. The module referred to in the present invention refers to a series of computer program segments that can be executed by at least one processor and can complete fixed functions, which are stored in a memory. In this embodiment, the functions of each module will be described in detail in subsequent embodiments.
[0159] A semantic recognition module, used for acquiring a project file and identifying a semantic vector of the project file using a large language model; A vector comparison module, used to compare the semantic vector with a pre-generated standard semantic vector, and screen out a difference semantic vector that is inconsistent with the standard semantic vector; The logical reasoning module is used to use the thinking chain reasoning model to perform logical reasoning on the semantic vector of the project file based on the user's question to obtain the abnormal semantic vector; A risk analysis module, used to input the difference semantic vector and the abnormal semantic vector into the dialogue pre-training model in sequence to obtain corresponding risk analysis text and modification suggestion text; The report generation module is used to convert the difference semantic vector and the abnormal semantic vector into a statement, write the statement and the corresponding risk analysis text and modification suggestion text into a preset report template, and obtain a risk analysis report.
[0160] Optionally, as an embodiment of the present invention, the semantic recognition module includes: A preprocessing unit, used for preprocessing the project file, including parsing and converting the file format; The first processing unit is used to perform word segmentation and syntactic analysis on the pre-processed project file using natural language processing technology to obtain structured information; A second processing unit is used to perform deep semantic analysis on the structured information using a large language model to identify key entities and relationships between key entities in the structured information; The semantic embedding unit is used to map key entities and the relationships between key entities into a high-dimensional vector space by using semantic embedding technology to obtain a semantic vector.
[0161] Optionally, as an embodiment of the present invention, the vector comparison module includes: A standard library construction unit, used to pre-construct a standard library, wherein the standard library stores standard semantic vectors, wherein the standard semantic vectors are semantic vectors of legal entities and relationships mapped to a high-dimensional vector space; A similarity calculation unit, used for calculating the cosine similarity between the semantic vector and a standard semantic vector in a standard library; A judgment module, used for judging whether there is a standard semantic vector in the standard library whose cosine similarity with the semantic vector reaches a set similarity threshold; A first determination unit, configured to determine that the semantic vector is normal if there is a standard semantic vector in the standard library whose cosine similarity with the semantic vector reaches a set similarity threshold; The second determination unit is configured to determine that the semantic vector is a difference semantic vector if there is no standard semantic vector in the standard library whose cosine similarity with the semantic vector reaches a set similarity threshold.
[0162] Optionally, as an embodiment of the present invention, a thought chain reasoning model is used to perform logical reasoning on the semantic vector of the project file based on the user question to obtain an abnormal semantic vector, including: Use machine learning or deep learning techniques to identify project scenarios and user needs; Retrieving a matching target reasoning path from predefined reasoning paths based on the project scenario and user requirements; Identify the hierarchical structure of the semantic vector of the project file, and decompose the user question into multiple sub-questions based on the hierarchical structure; The target reasoning path is used to verify the sub-problems corresponding to each level and identify abnormal semantic vectors.
[0163] Optionally, as an embodiment of the present invention, it further includes: Use contrastive learning technology to automatically label unlabeled project data; After confirming that laws and regulations have been updated, the updated content is automatically extracted, and the new regulatory information is compressed and migrated to the existing model through knowledge distillation technology based on the Teacher-Student Model, so that it can be seamlessly integrated into the existing system while maintaining the model's reasoning ability and stability; Through the active learning mechanism, based on the confidence of the large model in the query, sample data for further annotation and learning is automatically selected to improve the learning efficiency of the model; Expand the annotated sample library, extract representative samples from unlabeled data, improve the accuracy of comparative learning, and enhance the understanding of complex regulations by analyzing the applicability of different regulations and policies.
[0164] Optionally, as an embodiment of the present invention, the difference semantic vector and the abnormal semantic vector are sequentially input into the dialogue pre-training model to obtain corresponding risk analysis text and modification suggestion text, including: Generate prompt engineering based on difference semantic vectors and abnormal semantic vectors; The prompt project is input into the dialogue pre-training model to obtain corresponding risk analysis text and modification suggestion text.
[0165] Optionally, as an embodiment of the present invention, it further includes: The vocabulary selection, structure and language expression of the prompting project are adjusted based on the answer instructions fed back by the user.
[0166] Although the present invention has been described in detail by referring to the accompanying drawings and in combination with the preferred embodiments, the present invention is not limited thereto. Without departing from the spirit and essence of the present invention, a person skilled in the art may make various equivalent modifications or substitutions to the embodiments of the present invention, and these modifications or substitutions shall be within the scope of the present invention. Any person skilled in the art who is familiar with the present invention may easily conceive of changes or substitutions within the technical scope disclosed by the present invention, and these shall be within the scope of protection of the present invention.
Claims
1. A project document risk analysis method, characterized in that: include: Obtaining a project file, and identifying a semantic vector of the project file using a large language model; Comparing the semantic vector with a pre-generated standard semantic vector, and screening out a difference semantic vector that is inconsistent with the standard semantic vector; Use the thinking chain reasoning model to perform logical reasoning on the semantic vector of the project file based on user questions to obtain the abnormal semantic vector; The difference semantic vector and the abnormal semantic vector are sequentially input into the dialogue pre-training model to obtain corresponding risk analysis text and modification suggestion text; The difference semantic vector and the abnormal semantic vector are converted into sentences, and the sentences, corresponding risk analysis texts and modification suggestion texts are written into a preset report template to obtain a risk analysis report.
2. The method according to claim 1, characterized in that: Obtaining a project file and identifying a semantic vector of the project file using a large language model, including: Preprocessing the project files, including parsing and converting the file formats; Use natural language processing technology to perform word segmentation and syntactic analysis on the pre-processed project files to obtain structured information; Performing deep semantic analysis on the structured information using a large language model to identify key entities and relationships between key entities in the structured information; Semantic embedding technology is used to map key entities and their relationships into high-dimensional vector space to obtain semantic vectors.
3. The method according to claim 2, characterized in that Comparing the semantic vector with a pre-generated standard semantic vector, and screening out a difference semantic vector that is inconsistent with the standard semantic vector, including: Pre-building a standard library, wherein the standard library stores standard semantic vectors, wherein the standard semantic vectors are semantic vectors of legal entities and relationships mapped to high-dimensional vector spaces; Calculating the cosine similarity between the semantic vector and a standard semantic vector in a standard library; Determine whether there is a standard semantic vector in the standard library whose cosine similarity with the semantic vector reaches a set similarity threshold: If yes, it is determined that the semantic vector is normal; If not, it is determined that the semantic vector is a difference semantic vector.
4. The method according to claim 1, characterized in that: The thinking chain reasoning model is used to perform logical reasoning on the semantic vector of the project file based on user questions to obtain abnormal semantic vectors, including: Use machine learning or deep learning techniques to identify project scenarios and user needs; Retrieving a matching target reasoning path from predefined reasoning paths based on the project scenario and user requirements; Identify the hierarchical structure of the semantic vector of the project file, and decompose the user question into multiple sub-questions based on the hierarchical structure; The target reasoning path is used to verify the sub-problems corresponding to each level and identify abnormal semantic vectors.
5. The method according to claim 4, characterized in that The method further comprises: Use contrastive learning technology to automatically label unlabeled project data; After confirming that laws and regulations have been updated, extract the updated content and compress and migrate the new regulatory information to the existing model through knowledge distillation technology based on the teacher-student model, so that it can be seamlessly integrated into the existing system while maintaining the model's reasoning ability and stability; Through the active learning mechanism, based on the confidence of the large model in the query, sample data for further annotation and learning is automatically selected to improve the learning efficiency of the model; Expand the annotated sample library, extract representative samples from unlabeled data, improve the accuracy of comparative learning, and enhance the understanding of complex regulations by analyzing the applicability of different regulations and policies.
6. The method according to claim 1, characterized in that The difference semantic vector and the abnormal semantic vector are sequentially input into the dialogue pre-training model to obtain corresponding risk analysis text and modification suggestion text, including: Generate prompt engineering based on difference semantic vectors and abnormal semantic vectors; The prompt project is input into the dialogue pre-training model to obtain corresponding risk analysis text and modification suggestion text.
7. The method according to claim 6, characterized in that The method further comprises: The vocabulary selection, structure and language expression of the prompting project are adjusted based on the answer instructions fed back by the user.
8. A project document risk analysis system, characterized in that: include: A semantic recognition module, used for acquiring a project file and identifying a semantic vector of the project file using a large language model; A vector comparison module, used to compare the semantic vector with a pre-generated standard semantic vector, and screen out a difference semantic vector that is inconsistent with the standard semantic vector; The logical reasoning module is used to use the thinking chain reasoning model to perform logical reasoning on the semantic vector of the project file based on the user's question to obtain the abnormal semantic vector; A risk analysis module, used to input the difference semantic vector and the abnormal semantic vector into the dialogue pre-training model in sequence to obtain corresponding risk analysis text and modification suggestion text; The report generation module is used to convert the difference semantic vector and the abnormal semantic vector into a statement, write the statement and the corresponding risk analysis text and modification suggestion text into a preset report template, and obtain a risk analysis report.
9. The system according to claim 8, characterized in that The semantic recognition module comprises: A preprocessing unit, used for preprocessing the project file, including parsing and converting the file format; The first processing unit is used to perform word segmentation and syntactic analysis on the pre-processed project file using natural language processing technology to obtain structured information; A second processing unit, configured to perform deep semantic analysis on the structured information using a large language model, and identify key entities and relationships between key entities in the structured information; The semantic embedding unit is used to map key entities and their relationships into a high-dimensional vector space by using semantic embedding technology to obtain a semantic vector.
10. The system according to claim 9, characterized in that The vector comparison module comprises: A standard library construction unit, used to pre-construct a standard library, wherein the standard library stores standard semantic vectors, wherein the standard semantic vectors are semantic vectors of legal entities and relationships mapped to a high-dimensional vector space; A similarity calculation unit, used for calculating the cosine similarity between the semantic vector and a standard semantic vector in a standard library; A judgment module, used for judging whether there is a standard semantic vector in the standard library whose cosine similarity with the semantic vector reaches a set similarity threshold; A first determination unit, configured to determine that the semantic vector is normal if there is a standard semantic vector in the standard library whose cosine similarity with the semantic vector reaches a set similarity threshold; The second determination unit is configured to determine that the semantic vector is a difference semantic vector if there is no standard semantic vector in the standard library whose cosine similarity with the semantic vector reaches a set similarity threshold.
Citation Information
Patent Citations
Multi-hop inference knowledge editing method based on model cognitive verification
CN118364912A
Business contract risk intelligent review method and system based on large-scale language model
CN118747904A
Intelligent question and answer method based on cooperation of large language model and knowledge graph
CN118797017A
Intelligent financial question-answering system realized based on large language model
CN119441404A
Intelligent analysis method for junior middle school English reading understanding test questions based on large model
CN119514553A
Cited By
Method for automatically identifying difference between domestic and overseas standard files
CN120197607A
Project intelligent management analysis platform based on security risk standardization
CN120746499A
Project intelligent management and analysis device based on security risk standardization
CN120746499B
Power grid project comprehensive risk knowledge base construction method and device, equipment and storage medium
CN121683966A