A fuzzy determination method and device for error information in contract text
Automatic text comparison by calculating text cosine similarity and difference comparison algorithm, the problems of inefficiency and frequent errors in the contract review process are solved, and more efficient, accurate and intuitive audit results are achieved.
Patent Information
- Application Number
- CN202310672681.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-06-07
- Publication Date
- 2025-05-13
- Estimated Expiration
- 2043-06-07
AI Technical Summary
The existing contract review process consumes a lot of time and labor costs, is inefficient, is prone to errors and missed judgments, and the results are not intuitive enough.
By calculating the cosine similarity between texts, automatically comparing texts with the difference comparison algorithm, and designing programs to automatically annotate the review results to provide more intuitive and accurate review results.
It reduces the possibility of misunderstanding and misjudgment of audit results, improves audit efficiency and accuracy, provides intuitive audit results, and reduces labor costs.
Smart Images

Figure CN116702739B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of natural language processing, and in particular to a method and device for fuzzy determination of erroneous information in a contract text. Background Art
[0002] In the prior art, the process of contract review is that the relevant departments submit contracts, manually review contracts, manually annotate contracts, and return contracts. This process often has many disadvantages. For example, first, it takes a lot of time and manpower to review, especially for a large number of contract reviews, which requires a large amount of human resources to process, increasing the cost and management difficulty of the enterprise. Second, the efficiency of manual review is low. Manual review requires a lot of reading and comparison work, and these tasks may involve multiple departments and positions, and require constant communication and coordination, resulting in a slow review process and low efficiency. Third, since the manual review process is based on the experience and subjective judgment of professionals, there may be problems of errors and missed judgments. Manual review requires a lot of comparison and judgment work, which may cause errors or missed judgments due to personal opinions, knowledge level and other factors. Fourth, the results of manual review are not intuitive enough, and often require a large amount of review data to be summarized and analyzed before the conclusion of the contract review can be drawn, which not only increases the uncertainty of the review results, but also increases the possibility of misunderstanding and misjudgment of the review results.
[0003] Therefore, the existing contract review process requires a lot of time and manpower costs, and is also very prone to errors and omissions, which leads to some unnecessary disputes and conflicts, and the final results are not intuitive and concise enough. How to solve the problems of waste of human resources, low efficiency, possible errors and less intuitive results in the contract review process has become one of the many problems that technicians in this field need to solve. Summary of the invention
[0004] The purpose of this application is to overcome the existing technical defects and provide a fuzzy judgment method and device for erroneous information in contract texts. By calculating the cosine similarity between texts, automatically comparing texts with the diff algorithm, and designing a program to automatically annotate the audit results, it can provide more intuitive and accurate audit results and reduce the possibility of misunderstanding and misjudgment of the audit results.
[0005] The purpose of this application is achieved through the following technical solutions:
[0006] In a first aspect, the present application proposes a fuzzy determination method for erroneous information in a contract text, the method comprising:
[0007] Based on the secondary encapsulation of Python-docx library, the contract text is read and written to obtain multiple paragraphs;
[0008] Comparing the multiple paragraphs with the paragraphs in the contract text template, and calculating the paragraph cosine similarity;
[0009] Using a difference comparison algorithm to perform a difference comparison process on the multiple paragraphs and the paragraphs in the contract text template according to the paragraph cosine similarity to obtain text difference content and text difference position;
[0010] The contract text is modified according to the text difference content and text difference position.
[0011] In a possible implementation, the reading and writing schemes include: reading and writing by paragraphs, reading and writing by runs, reading and writing by text, reading and writing by specified text content, and reading and writing by specified text index.
[0012] In a possible implementation, the method further includes:
[0013] The highest value of the paragraph cosine similarity is taken as a matching result, and the plurality of paragraphs and the paragraphs in the contract text template are processed according to the matching result using a difference comparison algorithm.
[0014] In a possible implementation, the step of using a difference comparison algorithm to perform a difference comparison process on the multiple paragraphs and the paragraphs in the contract text template according to the paragraph cosine similarity to obtain text difference contents and text difference positions includes:
[0015] Using a difference comparison algorithm to insert, delete and match the multiple paragraphs with the paragraphs in the contract text template according to the paragraph cosine similarity to obtain text difference content and text difference position;
[0016] The insert operation is to insert new characters or lines into the contract text;
[0017] The deletion operation is to delete characters or lines in the contract text;
[0018] The matching operation is to match characters of the contract text with characters of the contract text template, or lines of the contract text with lines of the contract text template.
[0019] In a possible implementation manner, the difference comparison algorithm obtains the text difference content and the text difference position by finding the maximum value of the common subsequence in the matching operation.
[0020] In a possible implementation manner, the step of modifying the contract text according to the text difference content and text difference position includes:
[0021] S1. Pass in the start index and end index, and calculate the run list and index of the text difference position;
[0022] S2, split the run in the run list into three sections: the text before the mark, the target text to be marked, and the text after the mark. The run represents a formatted text block;
[0023] S3, add a new run after the original run and set the specified color and target text;
[0024] S4. Modify the text content of the original run to the text before the mark.
[0025] S5. If the target text to be marked is not the entire text of the original run, you need to create a new run for the text after the mark and set it to the color of the original run.
[0026] S6. Repeat S3-S5 to traverse all runs.
[0027] In a possible implementation, the contract text template includes the project contractor, project subcontractor, signing location and signing date.
[0028] In a second aspect, the present application further proposes a device for determining fuzzy information of a contract text, the device comprising:
[0029] The reading and writing module is used to read and write the contract text to obtain multiple paragraphs based on the secondary encapsulation of the Python-docx library;
[0030] A comparison module, used to compare the multiple paragraphs with the paragraphs in the contract text template and calculate the paragraph cosine similarity;
[0031] A processing module, configured to use a difference comparison algorithm to perform a difference comparison process on the multiple paragraphs and the paragraphs in the contract text template according to the paragraph cosine similarity to obtain text difference contents and text difference positions;
[0032] A modification module is used to modify the contract text according to the text difference content and text difference position.
[0033] In a third aspect, the present application further proposes a computer device, comprising a processor and a memory, wherein the memory stores a computer program, and the computer program is loaded and executed by the processor to implement the fuzzy judgment method for erroneous information in a contract text as described in any one of the first aspects.
[0034] In a fourth aspect, the present application further proposes a computer-readable storage medium, in which a computer program is stored. The computer program is loaded and executed by a processor to implement the fuzzy judgment method for erroneous information in a contract text as described in any one of the first aspects.
[0035] The above-mentioned main scheme of the present application and its further options can be freely combined to form multiple schemes, all of which are schemes that can be adopted and claimed for protection in the present application; and in the present application, (non-conflicting options) options and other options can also be freely combined. After understanding the scheme of the present application, those skilled in the art can understand that there are multiple combinations based on the prior art and common knowledge, all of which are technical schemes to be protected by the present application, and they are not exhaustively listed here.
[0036] The present application discloses a fuzzy judgment method and device for error information in a contract text. First, the contract text is read and written based on the secondary encapsulation of the Python-docx library to obtain multiple paragraphs, and then the multiple paragraphs processed by difference comparison are compared with the paragraphs in the contract text template, and the cosine similarity of the paragraphs is calculated. Then, a difference comparison algorithm is used to perform difference comparison processing on the multiple paragraphs processed by difference comparison and the paragraphs in the contract text template according to the cosine similarity of the difference comparison processing paragraphs to obtain the text difference content and text difference position, and finally, the difference comparison processing contract text is modified according to the text difference content and text difference position of the difference comparison processing. By calculating the cosine similarity between texts, automatically comparing texts with the difference comparison algorithm, and automatically annotating the audit results, it has the advantages of accurate recognition, fast speed, and high efficiency, and can also provide intuitive and accurate audit results, reducing the possibility of misunderstanding and misjudgment of audit results.
[0037] The technical effects achieved by the present invention are:
[0038] First, based on the secondary encapsulation of the Python-docx library, the reading and writing of contract documents are realized, and more convenient calls are achieved through the functions provided by the encapsulation, and new functions are added.
[0039] Second, we traverse each paragraph of the two docx files and calculate the cosine similarity between them. By calculating the similarity, we can find the matching paragraphs in the two files to avoid comparing the entire document in a relatively large document, thus saving time and improving the accuracy of the comparison.
[0040] Third, after finding the matching paragraphs, the difference comparison algorithm is used to compare the text between the two paragraphs, and the difference content and the location of the text can more intuitively show the differences in the document instead of just showing different paragraphs.
[0041] Fourth, mark the content that is different, and use the relevant algorithm to find the text content that needs to be modified between the runs of the paragraph, and modify the format of the text. The content in the document can be modified more accurately, thereby improving the quality and accuracy of the document.
[0042] Fifth, introduce natural language processing technology to realize automatic extraction and analysis of important information in the contract, and improve the accuracy and efficiency of the review. BRIEF DESCRIPTION OF THE DRAWINGS
[0043] Figure 1 A flow chart of a method for fuzzy determination of erroneous information in a contract text provided in an embodiment of the present application is shown. DETAILED DESCRIPTION
[0044] The following describes the embodiments of the present application through specific examples, and those skilled in the art can easily understand other advantages and effects of the present application from the contents disclosed in this specification. The present application can also be implemented or applied through other different specific embodiments, and the details in this specification can also be modified or changed in various ways based on different viewpoints and applications without departing from the spirit of the present application. It should be noted that the following embodiments and features in the embodiments can be combined with each other without conflict.
[0045] Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making any creative work shall fall within the scope of protection of this application.
[0046] In the prior art, the existing contract review process has the following shortcomings:
[0047] First, there are accuracy and efficiency issues in the contract review process in the background technology.
[0048] Second, manual review causes time-consuming and wasteful work.
[0049] Third, errors and omissions caused by subjective judgment and lack of experience of professionals in contract review.
[0050] Fourth, the complexity and uncertainty of contract documents lead to difficulties in auditing.
[0051] Fifth, the audit results presented in the audit process are not intuitive enough, and it is difficult to summarize and analyze data when the amount of data is large.
[0052] Therefore, how to solve the problems of waste of human resources, low efficiency, possible errors and lack of intuitive results in the contract review process has become one of the many problems that technical personnel in this field need to solve.
[0053] In order to solve the above-mentioned problems, the present application proposes a fuzzy judgment method and device for erroneous information in a contract text, which can not only use computer-assisted manual review to effectively solve the problems of waste of human resources, low efficiency, possible errors and lack of intuitive results in the contract review process, but also can calculate the cosine similarity between texts, cooperate with the difference comparison algorithm to perform automatic text comparison, and automatically annotate the review results. It has the advantages of accurate recognition, fast speed and high efficiency, provides intuitive and accurate review results, and reduces the possibility of misunderstanding and misjudgment of the review results. The method is described in detail below.
[0054] Please refer to Figure 1 , Figure 1 A flow chart of a method for determining fuzzy contract text error information provided by an embodiment of the present application is shown, wherein the steps are as follows:
[0055] S100. Based on the secondary encapsulation of Python-docx library, the contract text is read and written to obtain multiple paragraphs.
[0056] S200: compare the multiple paragraphs with the paragraphs in the contract text template, and calculate the paragraph cosine similarity.
[0057] S300: using a difference comparison algorithm to perform a difference comparison process on multiple paragraphs and paragraphs in the contract text template according to paragraph cosine similarity to obtain text difference content and text difference position.
[0058] S400. Mark and modify the contract text according to the text difference content and text difference position.
[0059] The contract text template includes the project contractor, project subcontractor, signing location and signing date. The contract document can be read and written by repackaging the Python-docx library. The repackaging means functional packaging of the Python-docx library. The functions provided by the packaging can be more conveniently called and new functions are added, such as reading and writing according to text content or index.
[0060] In an optional embodiment, the reading and writing schemes include: reading and writing by paragraphs, reading and writing by runs, reading and writing by text, reading and writing by specified text content, and reading and writing by specified text index.
[0061] Since the code after secondary encapsulation can read and write by paragraphs, by runs, by text, by specified text content, and by specified text index, etc. Among them, reading and writing by paragraph means splitting the contract document into multiple paragraphs and reading and writing each paragraph as a whole, reading and writing by runs means reading and writing paragraphs according to runs (continuous areas of text), reading and writing by text means directly reading and writing the plain text content of the entire document, reading and writing by specified text content means only reading or writing paragraphs or runs containing specified content, and reading and writing by specified text index means reading or writing paragraphs or runs at specified index positions in the document. The specific implementation of these reading and writing schemes can be detailed in the encapsulation code as needed.
[0062] It is worth mentioning that the reading and writing of contract documents is realized based on the secondary encapsulation of Python-docx library. The reading and writing schemes can be selected or combined according to the actual situation. For example, when the contract document needs to be processed as a whole, the contract document is classified and analyzed according to different paragraphs. At this time, the read and write part can be the entire paragraph or a part of the paragraph.
[0063] The above reading and writing scheme can be used when the format of the contract document needs to be modified, when specific text content in the contract document needs to be found and processed, and when processing needs to be performed based on the location information of the text. Using the above reading and writing scheme, the reading and writing part is the paragraph or runs where the specified text or specified index is located, and the text content in the contract document can be processed more finely, such as adjusting the font size, color and other formats.
[0064] In another possible real-time method, the highest value of paragraph cosine similarity is taken as a matching result, and a difference comparison algorithm is used to process multiple paragraphs and paragraphs in the contract text template according to the matching results.
[0065] The cosine similarity algorithm is an indicator used to measure the similarity of two vectors in direction, and is commonly used in text classification, information retrieval and other fields. This application uses the contract document as the file to be reviewed and the contract document template as the benchmark file. The two documents come from different places. Assuming that both documents are docx files, the matching process is explained. First, traverse each paragraph of the two docx files and calculate the cosine similarity between paragraphs in different files. The cosine similarity is obtained by calculating the cosine value of the angle between the two vectors. Each paragraph is first regarded as a vector, and the cosine similarity between paragraphs is used to measure their similarity in meaning.
[0066] Each paragraph can be represented as a vector. Two paragraphs are represented as vectors a and b. Then the cosine similarity between the two paragraphs is calculated using the formula. Each paragraph in the two documents is traversed. During the traversal, the cosine similarity value is compared as the matching result. The cosine similarity value is used as a score. The scores between all paragraphs in the two documents are compared, and the group of paragraphs with the highest score is selected as the matching result.
[0067] When comparing the advantages and disadvantages of different algorithms, several aspects need to be considered. First, the cosine similarity algorithm is simple, intuitive, and fast in calculation, making it suitable for similarity calculations on long texts. Second, the cosine similarity algorithm can highlight important words that appear less frequently in the text. However, the cosine similarity algorithm also has some problems. For example, when the length of the text varies greatly, it is easily affected by the longer text and the result is not accurate enough. At the same time, the cosine similarity algorithm cannot handle the differences in the position of the text well, such as the situation where the same word appears in different positions in two paragraphs but is still semantically similar.
[0068] In a possible implementation, the step of using a difference comparison algorithm to perform a difference comparison process on multiple paragraphs and paragraphs in the contract text template according to paragraph cosine similarity to obtain text difference content and text difference position includes:
[0069] Use a difference comparison algorithm to insert, delete and match multiple paragraphs with the paragraphs in the contract text template according to the paragraph cosine similarity to obtain the text difference content and text difference position;
[0070] The insert operation is to insert new characters or lines into the contract text;
[0071] The deletion operation is to delete characters or lines in the contract text;
[0072] The matching operation is to match characters of the contract text with characters of the contract text template, or lines of the contract text with lines of the contract text template.
[0073] The difference comparison algorithm obtains the text difference content and text difference position by finding the maximum value of the common subsequence in the matching operation.
[0074] The diff algorithm is a classic text comparison algorithm that can find the differences between two texts. The core idea is to determine the differences between texts by finding the longest common subsequence. Compared with other text comparison algorithms, the diff algorithm has the following advantages: First, the results can intuitively show the differences between texts, including the content and location of the differences, which is convenient for users to revise and modify. Second, the diff algorithm does not require preprocessing of the text and can directly compare the differences between two texts. Third, the diff algorithm can process larger texts, and the time complexity is O(N^2), so it runs faster.
[0075] The result data structure of the difference comparison algorithm of the present application is usually represented in the form of a difference block. A difference block is a data structure consisting of a group of rows, including insert, delete and match operations. Each difference block represents a difference between two texts. Specifically, the difference block includes: difference block type (Insert, Delete, Match), difference block starting position (position in the source text and target text), difference block length and difference block content. Using this data structure, the content and position of the difference block can be output to the document, which is convenient for users to revise and modify.
[0076] This application uses the algorithm to first convert the text into lines, and then compare the differences between the lines. The difference comparison algorithm performs insertion, deletion and matching to obtain the text difference content and text difference position, where the insertion operation is to insert a new character or line in the first text, and the deletion operation is to delete a character or line in the first text. The matching operation is to match a character or line in the first text with a character or line in the second text.
[0077] In the matching operation, the algorithm tries to find the longest possible common subsequence. If a matching line cannot be found, the algorithm tries to match the line through insertion or deletion operations. By constantly repeating these operations, the algorithm can find all the differences between the two texts.
[0078] In a possible implementation manner, step S400 of marking and modifying the contract text according to the text difference content and text difference position specifically includes:
[0079] S1. Pass in the start index and end index, and calculate the run list and index of the text difference position;
[0080] S2, split the run in the run list into three sections: the text before the mark, the target text to be marked, and the text after the mark. The run represents a formatted text block;
[0081] S3, add a new run after the original run and set the specified color and target text;
[0082] S4. Modify the text content of the original run to the text before the mark.
[0083] S5. If the target text to be marked is not the entire text of the original run, you need to create a new run for the text after the mark and set it to the color of the original run.
[0084] S6. Repeat S3-S5 to traverse all runs.
[0085] This application marks the content that is different in the comparison results, finds the text content that needs to be modified between the runs of the paragraph through the three modules of modify_text_color_by_index, search_run_info and add_run, and modifies the text format. The specific implementation process is: first, the list of runs that need to be modified and the corresponding indexes are calculated through the start index and end index passed in by the search_run_info() function, and then the modify_text_color_by_index() function uses the result returned by search_run_info() to obtain the list of runs that need to be modified. For each run that needs to be modified, it is divided into three sections: the text before the mark, the target text that needs to be marked, and the text after the mark. Then use the add_run() function to add a new run after the original run, set the color to the specified color, and set the text to the target text that needs to be marked. At the same time, the text content of the original run is modified to the text before the mark. If the target text that needs to be marked is not the entire text of the original run, it is necessary to create a new run for the text after the mark and set it to the color of the original run. Repeat the above steps to mark each run that needs to be marked.
[0086] By finding run objects and analyzing them, the text content that needs to be annotated is determined, and the annotation is added by creating a new run object. At the same time, the code also retains the original format of the word file to ensure that the annotated text is consistent with the format of the surrounding text. By parsing the text paragraphs and splitting and merging the run objects, the function of marking the specified content is realized.
[0087] Compared with the prior art, the embodiments of the present application have the following beneficial effects:
[0088] First, improve the efficiency and accuracy of contract processing by using automated methods to process contracts, complete tasks faster and maintain high accuracy. Compared with traditional manual processing methods, automated processing can save time and reduce errors, improve efficiency and accuracy. Automated processing can quickly read and analyze information in documents and automatically enter data into spreadsheets, improving accuracy and efficiency.
[0089] Second, reduce the labor cost of contract processing. When a large number of documents need to be processed, the traditional manual processing method requires hiring more employees to complete the task, while the method of the present application can complete a large amount of work through machines, thereby saving labor costs. For example, if a company needs to process thousands of contracts, using the traditional manual processing method may require dozens of employees to complete the task for several months, while using the automated processing method of the present invention can shorten the processing time and reduce the number of employees required.
[0090] Third, by using the same processing method, it is ensured that the same terms and expressions are accurately processed in each contract during the contract processing process. This can prevent errors and omissions caused by different processing methods. For example, when a company processes a contract, it may use different processing methods and templates to process different contracts, resulting in some important information being omitted or improperly processed. The automated processing method of the present invention can ensure that all contracts are processed using the same processing method and template, thereby ensuring consistency.
[0091] Fourth, the modular design and encapsulation method makes it easy to modify and maintain the code, which enhances the maintainability of contract processing. For example, if a new contract type appears or the legal provisions change, the code can be modified to adapt to these changes without rewriting the entire program. This can save time and cost and improve the flexibility of contract processing.
[0092] In a possible implementation, the present application further proposes a fuzzy determination device for erroneous information in a contract text, the device comprising:
[0093] The reading and writing module is used to read and write the contract text to obtain multiple paragraphs based on the secondary encapsulation of the Python-docx library;
[0094] A comparison module is used to compare multiple paragraphs with paragraphs in the contract text template and calculate the paragraph cosine similarity;
[0095] A processing module, used for performing a difference comparison process on the multiple paragraphs and the paragraphs in the contract text template according to the paragraph cosine similarity using a difference comparison algorithm to obtain text difference content and text difference position;
[0096] The modification module is used to modify the contract text according to the text difference content and text difference position.
[0097] An embodiment of the present application provides a computer device, which can implement the steps in any embodiment of the fuzzy judgment method for contract text error information provided in the embodiment of the present application. Therefore, the beneficial effects of the fuzzy judgment method for contract text error information provided in the embodiment of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0098] Those skilled in the art will appreciate that all or part of the steps in the various methods of the above embodiments can be completed by instructions, or by controlling related hardware through instructions, and the instructions can be stored in a computer-readable storage medium and loaded and executed by a processor. To this end, an embodiment of the present application provides a storage medium, in which a plurality of instructions are stored, and the instructions can be loaded by a processor to execute the steps of any embodiment of the fuzzy determination method for contract text error information provided in the embodiment of the present application.
[0099] The storage medium may include: a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk, etc.
[0100] Since the instructions stored in the storage medium can execute the steps in the embodiment of the fuzzy judgment method for error information in a contract text provided in the embodiments of the present application, the beneficial effects that can be achieved by the fuzzy judgment method for error information in a contract text provided in the embodiments of the present application can be achieved. Please refer to the previous embodiments for details and will not be repeated here.
[0101] The above description is only a preferred embodiment of the present application and is not intended to limit the present application. Any modifications, equivalent substitutions and improvements made within the spirit and principles of the present application should be included in the protection scope of the present application.
Claims
1. A fuzzy determination method for error information in a contract text, characterized in that: The method comprises: Based on the secondary encapsulation of Python-docx library, the contract text is read and written to obtain multiple paragraphs. The reading and writing schemes include: reading and writing by paragraph, reading and writing by runs, reading and writing by text, reading and writing by specified text content, and reading and writing by specified text index; Comparing the multiple paragraphs with the paragraphs in the contract text template, and calculating the paragraph cosine similarity; Using a difference comparison algorithm to perform a difference comparison process on the multiple paragraphs and the paragraphs in the contract text template according to the paragraph cosine similarity to obtain text difference content and text difference position; S1. Pass in the start index and end index, and calculate the run list and index of the text difference position; S2, split the run in the run list into three sections: the text before the mark, the target text to be marked, and the text after the mark. The run represents a formatted text block; S3, add a new run after the original run and set the specified color and target text; S4, modify the text content of the original run to the text before the mark; S5. If the target text to be marked is not the entire text of the original run, a new run needs to be created for the text after the mark and set to the color of the original run; S6. Repeat S3-S5 to traverse all runs.
2. The fuzzy determination method for error information in a contract text according to claim 1, characterized in that: The method further comprises: The highest value of the paragraph cosine similarity is taken as a matching result, and the plurality of paragraphs and the paragraphs in the contract text template are processed according to the matching result using a difference comparison algorithm.
3. The fuzzy determination method for error information in a contract text according to claim 1, characterized in that: The step of using a difference comparison algorithm to perform a difference comparison process on the multiple paragraphs and the paragraphs in the contract text template according to the paragraph cosine similarity to obtain text difference content and text difference position includes: Using a difference comparison algorithm to insert, delete and match the multiple paragraphs with the paragraphs in the contract text template according to the paragraph cosine similarity to obtain text difference content and text difference position; The insert operation is to insert new characters or lines into the contract text; The deletion operation is to delete characters or lines in the contract text; The matching operation is to match characters of the contract text with characters of the contract text template, or lines of the contract text with lines of the contract text template.
4. The fuzzy determination method for error information in a contract text according to claim 3, characterized in that: The difference comparison algorithm obtains the text difference content and text difference position by finding the maximum value of the common subsequence in the matching operation.
5. The fuzzy determination method for error information in a contract text according to claim 1, characterized in that: The contract text template includes the project contractor, project subcontractor, signing location and signing date.
6. A fuzzy determination device for error information in a contract text, characterized in that: The device comprises: The reading and writing module is used to read and write the contract text based on the secondary encapsulation of the Python-docx library to obtain multiple paragraphs. The reading and writing schemes include: reading and writing by paragraph, reading and writing by runs, reading and writing by text, reading and writing by specified text content, and reading and writing by specified text index; A comparison module, used to compare the multiple paragraphs with the paragraphs in the contract text template and calculate the paragraph cosine similarity; A processing module, configured to use a difference comparison algorithm to perform a difference comparison process on the multiple paragraphs and the paragraphs in the contract text template according to the paragraph cosine similarity to obtain text difference contents and text difference positions; Modify the module to pass in the start index and end index, and calculate the run list and index of the text difference position; Split the runs in the run list into three sections: the text before the mark, the target text to be marked, and the text after the mark. A run represents a formatted text block. Add a new run after the original run and set the specified color and target text; Modify the text content of the original run to the text before the mark; If the target text to be marked is not the entire text of the original run, you need to create a new run with the text after the mark and set it to the color of the original run; Iterate over all runs.
7. A computer device, characterized in that: The computer device includes a processor and a memory, wherein a computer program is stored in the memory, and the computer program is loaded and executed by the processor to implement the fuzzy determination method for erroneous information in a contract text as described in any one of claims 1-5.
8. A computer-readable storage medium, characterized in that: The storage medium stores a computer program, which is loaded and executed by a processor to implement the fuzzy determination method for erroneous information in a contract text as described in any one of claims 1 to 5.
Citation Information
Patent Citations
Contract auditing method, apparatus and device, and medium
CN110837998A
Automatic contract version comparison tool and method in field of bad asset operation
CN110852054A
Text similarity analysis method and device and storage medium
CN113255369A