Intelligent file review and result pre-judgment method based on machine learning and application
Through intelligent file review and result prediction methods based on machine learning, the limitations of the existing technology in understanding context and semantics are solved, the accuracy and efficiency of file review are improved, and the accuracy in handling complex or fuzzy legal issues is ensured.
Patent Information
- Application Number
- CN202510465541.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-15
- Publication Date
- 2025-05-13
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing intelligent document review methods have limitations in understanding context and semantics, especially when dealing with complex sentence structures and metaphors, it is easy to cause misjudgment, and the inability to fully understand the intention of the document, resulting in limited accuracy when dealing with complex or vague legal issues.
Using intelligent file review and result prediction methods based on machine learning, we use natural language processing and regular matching algorithms to extract specification file features, combine machine learning algorithms to establish a file review and judgment model, and deploy an intelligent prediction model to predict results for non-compliant content.
It improves the identification and processing capabilities of the document review judgment model, ensures the accuracy of document review, can efficiently and accurately determine whether the document meets usage requirements, and provides modification suggestions to avoid potential risks.
Smart Images

Figure CN119990104A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of document review technology, and specifically, relates to an intelligent document review and result prediction method and application based on machine learning. Background Art
[0002] Intelligent document review refers to the analysis and review of files of text, images or other data types through machine learning, natural language processing and other artificial intelligence technologies to identify key information, potential risks, compliance issues or sensitive data.
[0003] The combination of intelligent document review methods and equipment provides enterprises and institutions with an efficient and accurate solution to meet compliance and risk management needs. This method not only improves the efficiency of document processing, but also reduces the error rate of manual review, providing strong support for document management in all walks of life.
[0004] However, existing intelligent document review methods still have limitations in understanding context and semantics, especially when dealing with complex sentence structures and metaphors, which can easily lead to misjudgments. When dealing with language features and industry terms in certain specific fields, they are unable to fully understand the intent of the document. For example, when dealing with complex or ambiguous legal issues, their accuracy is limited, which may cause contract disputes or adverse consequences for the company. Summary of the invention
[0005] In order to solve the above problems and technical defects, the present application adopts the following technical solution, a method for intelligent file review and result prediction based on machine learning, comprising the following steps: Obtain historical standard template files, use natural language processing algorithms to analyze the content in historical standard template files, use regular matching algorithms to extract features from the analysis results, obtain standard file features, combine standard file features with machine learning algorithms for training, and establish a file review and judgment model; Obtain the file to be identified, use the file review judgment model to review the file to be identified, calculate the features of the file to be identified, and determine whether the file to be identified is compliant based on the calculation results; Deploy an intelligent prediction model and use it to predict the results of non-compliant content in the file, obtain the adverse results caused by the non-compliant content, and provide modification suggestions based on the adverse results.
[0006] Preferably, the historical specification template files need to be classified before analyzing the contents in the historical specification template files, and the process is as follows: Use the maximum entropy classifier to type-label historical specification template files; Use named entity recognition technology to count the type nouns of the files that have been type-annotated; Calculate the NER ratio based on the statistical results and use the NER ratio as the basis for file type identification; Perform type matching on the files that have completed type marking, and use the matching result as the file type of the historical specification template file.
[0007] Furthermore, the analysis of the content in the historical specification template file is to segment each part of the content in the file, mark and identify each part of the content, and obtain the key sentences of each part of the content by analyzing the content structure and semantics of each part of the content.
[0008] Furthermore, the process of obtaining the specification file features is as follows: Perform regular expression matching on the key sentences of each part of the content, and obtain the content purpose of each key sentence according to the type of the target file; Document type clues and document type identification are based on feature extraction of the content purpose of each key sentence; Calculate the weight of each feature in each part of the content, and select the feature with the highest weight as the canonical file feature of the target part of the content.
[0009] Furthermore, the process of establishing the document review judgment model is as follows: Design a machine learning architecture using a machine learning algorithm, use the specification file features as labels according to the type of each specification file, and annotate and classify the labels; Treat different tags as different subtasks, obtain the subtask set, and determine whether each subtask can be further decomposed. If it cannot be decomposed, clarify the logical relationship between each subtask and other subtasks; In the machine learning architecture, a logical process framework is generated based on logical relationships, and repeated training is performed. The trained machine learning architecture is used as a document review and judgment model.
[0010] Preferably, the process of calculating the features of the file to be identified using the file review judgment model is as follows: Segment the file to be identified into multiple content parts; The document review judgment model identifies and calculates the logical connection relationship between the target content part and other content parts, and obtains the logical correlation score between the target content part and other content parts.
[0011] Furthermore, the calculation process of the logic correlation score is as follows: Perform preliminary sentence division and keyword marking on the two content parts to be compared to obtain the content structure; Filter keywords based on semantics, retain the key parts related to semantics, and remove irrelevant parts; Perform semantic similarity matching on sentences according to content structure and obtain matching scores of content parts; The matching score and the number and proportion of keywords are combined to obtain the logical correlation score between the target content part and other content parts.
[0012] Furthermore, the determination of whether the to-be-identified file is compliant is performed by presetting two review judgment thresholds, namely a first review judgment threshold and a second review judgment threshold, wherein the first review judgment threshold is less than the second review judgment threshold; If the logical relevance score is less than the first review judgment threshold, it is determined that the document to be identified does not meet the review requirements; If the logical relevance score is greater than or equal to the first review judgment threshold and less than the second review judgment threshold, the file to be identified will be manually judged by the administrator; If the logical relevance score is greater than or equal to the second review judgment threshold, it is determined that the file to be identified meets the review requirements.
[0013] An electronic device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the program, the content of the intelligent file review and result prediction method based on machine learning as described above is implemented.
[0014] A computer-readable storage medium stores a computer program, which, when executed by a processor, implements the contents of the intelligent file review and result prediction method based on machine learning as described above.
[0015] Compared with the prior art, the beneficial effects of this application are: (1) This application obtains historical standard template files, analyzes the content in the historical standard template files, obtains key sentences of each part of the content by analyzing the content structure and semantics of each part of the content, obtains standard file features based on the key sentences, and improves the recognition and processing capabilities of the file review judgment model; (2) This application performs regular expression matching on the key sentences of each part of the content, obtains the content purpose of each key sentence according to the type of the target file, calculates the weight of each feature in each part of the content, and selects the feature with the highest weight as the standard file feature of the target part of the content to ensure the accuracy of the file review; (3) This application calculates the logical connection relationship between the target content part and other content parts to obtain the logical correlation score between the target content part and other content parts. After multiple threshold judgments, it can efficiently and accurately determine whether the file meets the usage requirements. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] In the attached picture: Figure 1A schematic diagram of the method steps of an embodiment of the present application; Figure 2 This is a schematic diagram of the device structure of an embodiment of the present application. DETAILED DESCRIPTION
[0017] In order to make the purpose, technical solutions and advantages of the embodiments of the present application clearer, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application. Obviously, the described embodiments are part of the embodiments of the present application, not all of the embodiments. Generally, the components of the embodiments of the present application described and shown in the drawings here can be arranged and designed in various different configurations.
[0018] Example 1: Figure 1 As shown, a method for intelligent document review and result prediction based on machine learning includes the following contents: Obtain a historical specification template file, and use a natural language processing algorithm to analyze the content in the historical specification template file; Before analyzing the content in the historical specification template file, you need to classify the historical specification template file. The process is as follows: Use the maximum entropy classifier to label the historical specification template files, such as the labels: procurement, confidentiality, and service; Use named entity recognition technology to count the type nouns of the files that have been type-annotated; The BERT-BiLSTM-CRF model can be used to perform entity recognition on the annotated content and extract entity nouns such as "Party A", "Party B", and "Confidentiality Period".
[0019] Calculate the NER ratio based on the statistical results and use the NER ratio as the basis for file type identification; Perform type matching on the files that have completed type marking, and use the matching result as the file type of the historical specification template file.
[0020] Analyzing the content in the historical specification template file is to segment each part of the content in the file, mark and identify each part of the content, and obtain the key sentences of each part of the content by analyzing the content structure and semantics of each part of the content.
[0021] For example, a text segmentation algorithm can be used to segment a contract into "definition terms", "payment terms", "liability for breach of contract" and other parts according to the terms. Key sentences include "Party A must pay 70% of the contract amount within 30 days."
[0022] Use regular matching algorithm to extract features from analysis results and obtain standard file features; The process of obtaining the characteristics of the specification file is as follows: Perform regular expression matching on the key sentences of each part of the content, and obtain the content purpose of each key sentence according to the type of the target file; Document type clues and document type identification are based on feature extraction of the content purpose of each key sentence; Calculate the weight of each feature in each part of the content, and select the feature with the highest weight as the canonical file feature of the target part of the content.
[0023] For example, use regular expressions to design the payment amount, ratio, and term for the "payment terms", and generate feature tags after matching, such as payment time constraints.
[0024] Use TF-IDF combined with mutual information to screen high-weight features, such as the penalty ratio weight.
[0025] Combine the standardized document features with machine learning algorithms for training and establish a document review and judgment model; The process of establishing a document review judgment model is as follows: Design a machine learning architecture using a machine learning algorithm, use the specification file features as labels according to the type of each specification file, and annotate and classify the labels; Treat different tags as different subtasks, obtain the subtask set, and determine whether each subtask can be further decomposed. If it cannot be decomposed, clarify the logical relationship between each subtask and other subtasks; In the machine learning architecture, a logical process framework is generated based on logical relationships, and repeated training is performed. The trained machine learning architecture is used as a document review and judgment model.
[0026] Obtain the file to be identified, and use the file review judgment model to calculate the features of the file to be identified; The process of calculating the features of the files to be identified using the file review judgment model is as follows: Segment the file to be identified into multiple content parts; The document review judgment model identifies and calculates the logical connection relationship between the target content part and other content parts, and obtains the logical correlation score between the target content part and other content parts.
[0027] Logical relevance calculation can be used to perform semantic similarity matching on confidentiality clauses and breach of contract liability clauses in the contract to be reviewed. Sentence segmentation and keyword annotation are used to extract keywords such as confidentiality obligations and compensation for leaks. The cosine similarity is calculated using the Sentence-BERT algorithm, and finally a logical score is performed through weighted calculation.
[0028] The calculation process of logical relevance score is as follows: Perform preliminary sentence division and keyword marking on the two content parts to be compared to obtain the content structure; Filter keywords based on semantics, retain the key parts related to semantics, and remove irrelevant parts; Perform semantic similarity matching on sentences according to content structure and obtain matching scores of content parts; The matching score and the number and proportion of keywords are combined to obtain the logical correlation score between the target content part and other content parts.
[0029] According to the calculation result, whether the file to be identified is compliant is judged. The judgment of whether the file to be identified is compliant is performed by presetting two review judgment thresholds, namely, a first review judgment threshold and a second review judgment threshold, and the first review judgment threshold is less than the second review judgment threshold; If the logical relevance score is less than the first review judgment threshold, it is determined that the document to be identified does not meet the review requirements; If the logical relevance score is greater than or equal to the first review judgment threshold and less than the second review judgment threshold, the file to be identified will be manually judged by the administrator; If the logical relevance score is greater than or equal to the second review judgment threshold, it is determined that the file to be identified meets the review requirements.
[0030] Deploy an intelligent prediction model and use it to predict the results of non-compliant content in the file, obtain the adverse results caused by the non-compliant content, and provide modification suggestions based on the adverse results; The intelligent prediction model is trained by deploying the local deepseek intelligent model and then downloading historical non-compliant files and their corresponding dispute consequences online.
[0031] Example 2: Figure 2 As shown, from a hardware level, the present application provides an embodiment of an electronic device that implements all or part of the content of an intelligent file review and result prediction method based on machine learning, wherein the electronic device includes a service processor and a distributed memory, wherein the service processor is connected to the memory, wherein the distributed memory stores a service self-management program configured to store machine-readable instructions, and wherein the service processor executes the service self-management program, and when the instructions are executed by the processor, a public data storage management system based on artificial intelligence as described above is implemented.
[0032] From the hardware level, in order to effectively improve the flexibility, versatility and efficiency of data collection, the present application provides an embodiment of an electronic device that includes all or part of the contents of the intelligent file review and result prediction method based on machine learning, and the electronic device specifically includes the following contents: Processor, memory, communication interface and bus; wherein the processor, memory and communication interface communicate with each other through the bus; the communication interface is used to realize information transmission between the data acquisition device based on the distributed model and the core business system, user terminal and related database and other related equipment; the logic controller can be a desktop computer, a tablet computer and a mobile terminal, etc., but the present embodiment is not limited thereto. In the present embodiment, the logic controller can be implemented with reference to an embodiment of a public data storage management system based on artificial intelligence and an embodiment of a data acquisition device based on a distributed model, and the contents thereof are incorporated herein and repeated parts are not repeated.
[0033] It is understandable that the user terminal may include a smart phone, a tablet electronic device, a network set-top box, a portable computer, a desktop computer, a personal digital assistant (PDA), a vehicle-mounted device, a smart wearable device, etc. Among them, the smart wearable device may include smart glasses, a smart watch, a smart bracelet, etc.
[0034] In practical applications, the academic opinion annotation and analysis method can be performed on the electronic device side as described above, or all operations can be completed in the client device. The specific selection can be based on the processing capability of the client device and the limitations of the user's usage scenario, etc. This application does not limit this. If all operations are completed in the client device, the client device may also include a processor.
[0035] The above-mentioned client device may have a communication module (i.e., a communication unit), which can communicate with a remote server to realize data transmission with the server. The server may include a server on the task scheduling center side, and other implementation scenarios may also include a server on an intermediate platform, such as a server on a third-party server platform that has a communication link with the task scheduling center server. The server may include a single computer device, or a server cluster consisting of multiple servers, or a server structure of a distributed device.
[0036] Embodiment 3: The embodiments of the present application also provide a computer-readable storage medium capable of implementing the entire contents of an intelligent file review and result prediction method based on machine learning in the above embodiments, where the execution subject is a server or a client. A computer program is stored on the computer-readable storage medium, and when the computer program is executed by a processor, the entire contents of an artificial intelligence-based public data storage and management system in the above embodiments are implemented in which the execution subject is a server or a client.
[0037] The embodiments of the present application may be provided as methods, devices, or computer program products, and therefore, the present application may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware. Moreover, the present application may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program codes.
[0038] The present application is described with reference to the flowcharts and / or block diagrams of the methods, devices (apparatus), and computer program products according to the embodiments of the present application. It should be understood that each process and / or block in the flowchart and / or block diagram, as well as the combination of the processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing device to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate instructions for implementing the processes in the flowchart and / or block diagram. Figure 1 A process or multiple processes and / or boxes Figure 1 A device that provides the functions specified in a block or multiple blocks.
[0039] These computer program instructions may also be stored in a computer-readable memory capable of directing a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce an article of manufacture comprising an instruction device, which implements the process Figure 1 A process or multiple processes and / or boxes Figure 1 A function specified in one or more boxes.
[0040] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, thereby providing instructions for implementing the process. Figure 1 A process or multiple processes and / or boxes Figure 1 The steps for the functions specified in one or more boxes.
[0041] The above-mentioned embodiments only express the preferred implementation methods of the present application, and the descriptions thereof are relatively specific and detailed, but they cannot be understood as limiting the scope of the present application. It should be pointed out that, for ordinary technicians in this field, several modifications, improvements and substitutions can be made without departing from the concept of the present application, and these all belong to the protection scope of the present application.
Claims
1. An intelligent file review and result prediction method based on machine learning, characterized in that: The following steps are involved: Obtain historical standard template files, use natural language processing algorithms to analyze the content in historical standard template files, use regular matching algorithms to extract features from the analysis results, obtain standard file features, combine standard file features with machine learning algorithms for training, and establish a file review and judgment model; Obtain the file to be identified, use the file review judgment model to review the file to be identified, calculate the features of the file to be identified, and determine whether the file to be identified is compliant based on the calculation results; Deploy an intelligent prediction model and use it to predict the results of non-compliant content in the file, obtain the adverse results caused by the non-compliant content, and provide modification suggestions based on the adverse results.
2. According to claim 1, a method for intelligent file review and result prediction based on machine learning is characterized in that: Before analyzing the content in the historical specification template file, it is necessary to classify the historical specification template file first, and the process is as follows: Use the maximum entropy classifier to type-label historical specification template files; Use named entity recognition technology to count the type nouns of the files that have been type-annotated; Calculate the NER ratio based on the statistical results and use the NER ratio as the basis for file type identification; Perform type matching on the files that have completed type marking, and use the matching result as the file type of the historical specification template file.
3. The intelligent file review and result prediction method based on machine learning according to claim 2 is characterized in that: The analysis of the content in the historical specification template file is to segment each part of the content in the file, mark and identify each part of the content, and obtain the key sentences of each part of the content by analyzing the content structure and semantics of each part of the content.
4. The intelligent file review and result prediction method based on machine learning according to claim 3 is characterized in that: The process of obtaining the specification file features is as follows: Perform regular expression matching on the key sentences of each part of the content, and obtain the content purpose of each key sentence according to the type of the target file; Document type clues and document type identification are based on feature extraction of the content purpose of each key sentence; Calculate the weight of each feature in each part of the content, and select the feature with the highest weight as the canonical file feature of the target part of the content.
5. The intelligent file review and result prediction method based on machine learning according to claim 4 is characterized in that: The process of establishing the document review judgment model is as follows: Design a machine learning architecture using a machine learning algorithm, use the specification file features as labels according to the type of each specification file, and annotate and classify the labels; Treat different tags as different subtasks, obtain the subtask set, and determine whether each subtask can be further decomposed. If it cannot be decomposed, clarify the logical relationship between each subtask and other subtasks; In the machine learning architecture, a logical process framework is generated based on logical relationships, and repeated training is performed. The trained machine learning architecture is used as a document review and judgment model.
6. The intelligent file review and result prediction method based on machine learning according to claim 1 is characterized in that: The process of calculating the features of the file to be identified using the file review judgment model is as follows: Segment the file to be identified into multiple content parts; The document review judgment model identifies and calculates the logical connection relationship between the target content part and other content parts, and obtains the logical correlation score between the target content part and other content parts.
7. The intelligent file review and result prediction method based on machine learning according to claim 6 is characterized in that: The calculation process of the logical correlation score is as follows: Perform preliminary sentence division and keyword marking on the two content parts to be compared to obtain the content structure; Filter keywords based on semantics, retain the key parts related to semantics, and remove irrelevant parts; Perform semantic similarity matching on sentences according to content structure and obtain matching scores of content parts; The matching score and the number and proportion of keywords are combined to obtain the logical correlation score between the target content part and other content parts.
8. The intelligent file review and result prediction method based on machine learning according to claim 7 is characterized in that: The determination of whether the to-be-identified file is compliant is performed by presetting two review judgment thresholds, namely a first review judgment threshold and a second review judgment threshold, wherein the first review judgment threshold is less than the second review judgment threshold; If the logical relevance score is less than the first review judgment threshold, it is determined that the document to be identified does not meet the review requirements; If the logical relevance score is greater than or equal to the first review judgment threshold and less than the second review judgment threshold, the file to be identified will be manually judged by the administrator; If the logical relevance score is greater than or equal to the second review judgment threshold, it is determined that the file to be identified meets the review requirements.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, the content of the intelligent file review and result prediction method based on machine learning described in claim 1 is implemented.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, it implements the content of the intelligent file review and result prediction method based on machine learning described in claim 1.
Citation Information
Patent Citations
Text auditing method and device
CN110675269A
Contract review method, device and system and computer readable storage medium
CN114549241A
Contract review method and device, electronic equipment and storage medium
CN117909499A
Enterprise compliance examination method, apparatus and device, and storage medium
CN118709678A
Electronic contract auditing method and device, computer equipment and storage medium
CN119090681A