Contract text comparison method and device and nonvolatile storage medium
The contract documents are analyzed through the document layout segmentation model and the similarity calculation model, which solves the problem of low accuracy in contract documents comparison and achieves efficient and accurate comparison of contract documents of different versions.
Patent Information
- Application Number
- CN202510413932.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-07-11
AI Technical Summary
In the prior art, when comparing contract documents with different versions of contract documents, the accuracy rate is low, and it is impossible to quickly find out the differences between the same contract in different versions.
The document layout segmentation model is used to analyze the contract documents, and feature extraction and fusion is performed through the channel attention module of the encoder and decoder, and the text data is vectorized and similarity calculation model is combined to determine the differences and locations of the contract documents.
It realizes accurate comparison of contract documents of different versions, improves the accuracy of contract documents comparison, and can quickly find out the differences and locations of contract documents.
Smart Images

Figure CN120299059A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of natural language processing. Specifically, it relates to a method, apparatus, and non-volatile storage medium for comparing contract texts. Background Art
[0002] In modern business activities, contracts, as legal documents, are crucial for ensuring the legality and transparency of transactions. They detail the rights, obligations, and responsibilities of all parties, such as product specifications, amounts in sales contracts, and research contents and goals in technology project contracts. The number of contract texts is huge, and contracts are frequently added, deleted, and revised. Therefore, the management, retrieval, review, and comparison of contracts are daily tasks that enterprises must face.
[0003] In the process of contract management, it is often necessary to compare different versions of contracts to accurately grasp the changes in the contracts and ensure that the rights and interests of enterprises are not violated. In contract comparison, the focus is mainly on the changes in information such as contract texts, tables, and pictures. By comparing the differences between different versions of contracts, the changed contents of the contracts are analyzed. However, in related technologies, the management of contracts relies on manual comparison of paper contracts and electronic contracts, which is time-consuming, inefficient, and error-prone. The accuracy of document layout analysis on complex documents is insufficient, especially when dealing with areas of different sizes and uneven sample quantities. Secondly, when dealing with complex contract texts, it performs poorly in key indicators and is difficult to capture long-distance syntactic structures. Therefore, when comparing different versions of contract documents, the accuracy of document comparison is relatively low, and it is impossible to quickly find the differences between different versions of the same contract.
[0004] No effective solution has been proposed for the above problems yet. Summary of the Invention
[0005] Embodiments of this application provide a method, apparatus, and non-volatile storage medium for comparing contract texts to at least solve the technical problem that the accuracy of document comparison is relatively low when comparing different versions of contract documents in related technologies.
[0006] According to one aspect of the embodiments of the present application, a method for comparing contract texts is provided, including: obtaining a first contract document and a second contract document, where the first contract document and the second contract document are different versions of the same contract; performing layout analysis on the first contract document through a document layout segmentation model to obtain first document data corresponding to the first contract document; performing layout analysis on the second contract document through the document layout segmentation model to obtain second document data corresponding to the second contract document; comparing the first document data and the second document data to obtain a first comparison result between the first contract document and the second contract document, where the first comparison result is used to indicate the difference points and the positions of the difference points in the contract documents of the first contract document and the second contract document.
[0007] In some embodiments of the present application, the document layout segmentation model includes an encoder and a decoder. Each network block of the encoder includes a first channel attention module, and the decoder includes a second channel attention module; performing layout analysis on the first contract document through the document layout segmentation model to obtain first document data corresponding to the first contract document includes: using the first channel attention module to determine the first attention weights of each channel when the encoder performs layout analysis; and weighting each channel when the encoder performs layout analysis based on the first attention weights to obtain each weighted channel of the encoder; using each weighted channel of the encoder to perform per-channel feature extraction on the first contract document to obtain multi-level features of the first contract document; using the second channel attention module to determine the second attention weights of each channel of the decoder, and weighting each channel of the decoder based on the second attention weights to obtain each weighted channel of the decoder; performing feature fusion on the multi-level features of the first contract document based on each weighted channel of the decoder to obtain a fused feature map; determining first document data corresponding to the first contract document based on the fused feature map.
[0008] In some embodiments of the present application, determining a first comparison result between a first contract document and a second contract document based on first document data and second document data includes: performing character recognition on the tables and text areas in the first document data and the second document data to obtain a first contract text corresponding to the first document data and a second contract text corresponding to the second document data, where the first document data includes at least the tables, pictures, and text areas in the first contract document, and the second document data includes at least the tables, pictures, and text areas in the second contract document; vectorizing each paragraph in the first contract text to obtain a first vector set, where the first vector set includes vectors corresponding to each paragraph in the first contract text; vectorizing each paragraph in the second contract text to obtain a second vector set, where the second vector set includes vectors corresponding to each paragraph in the second contract text; determining, based on the first vector set and the second vector set, a comparison paragraph in the second contract text corresponding to each paragraph in the first contract text, and a comparison sentence in the corresponding comparison paragraph for each sentence in each paragraph in the first contract text; determining a second comparison result between each sentence in the first contract text and the corresponding comparison sentence through a similarity calculation model, and determining a set of second comparison results between all sentences in the first contract text and the corresponding comparison sentences as the first comparison result.
[0009] In some embodiments of the present application, determining, based on the first vector set and the second vector set, a comparison paragraph in the second contract text corresponding to each paragraph in the first contract text, and a comparison sentence in the corresponding comparison paragraph for each sentence in each paragraph in the first contract text includes: Step 1: Calculate the first cosine similarity between each vector in the first vector and the second vector set respectively, and determine the vector in the second vector set corresponding to the first cosine similarity with the largest value as the matching vector corresponding to the first vector, and determine the paragraph corresponding to the matching vector as the comparison paragraph of the paragraph corresponding to the first vector, where the first vector is any vector in the first vector set; Step 2: Determine the paragraph corresponding to the first vector as the target paragraph, calculate the second cosine similarity between each sentence in the target paragraph and each reference sentence in the comparison paragraph respectively, and determine the reference sentence corresponding to the second cosine similarity with the largest value as the comparison sentence of the sentence in the corresponding target paragraph; Iteratively execute Step 1 to Step 2 until all vectors in the first vector set are traversed, to obtain the comparison paragraphs corresponding to each paragraph in the first vector set, and the comparison sentences corresponding to each sentence in each paragraph.
[0010] In some embodiments of the present application, determining a second comparison result between each sentence in the first contract text and the corresponding comparison sentence through a similarity calculation model includes: determining the syntactic similarity between each sentence in the first contract text and the corresponding comparison sentence through the syntactic similarity calculation module of the similarity calculation model; determining the semantic similarity between each sentence in the first contract text and the corresponding comparison sentence through the semantic similarity calculation module of the similarity calculation model; determining the text similarity between each sentence in the first contract text and the corresponding comparison sentence based on the syntactic similarity, semantic similarity, preset syntactic similarity weight, and preset semantic similarity weight; and determining the second comparison result between each sentence in the first contract text and the corresponding comparison sentence based on the text similarity.
[0011] In some embodiments of the present application, determining the syntactic similarity between each sentence in the first contract text and the corresponding comparison sentence through the syntactic similarity calculation module of the similarity calculation model includes: performing graph-based dependency syntactic analysis on each sentence in the first contract text and the corresponding comparison sentence to obtain the first syntactic tree of each sentence in the first contract text and the second syntactic tree of the corresponding comparison sentence; extracting the first syntactic relationship of the first syntactic tree and the second syntactic relationship of the second syntactic tree; vectorizing the first syntactic relationship of each sentence to obtain the first syntactic relationship vector set of each sentence, and vectorizing the second syntactic relationship of the corresponding comparison sentence of each sentence to obtain the second syntactic relationship vector set; determining the similarity matrix between each sentence and the corresponding comparison sentence based on the first syntactic relationship vector set and the second syntactic relationship vector set; and determining the syntactic similarity between each sentence and the corresponding comparison sentence as the sum of the maximum values of each row in the similarity matrix.
[0012] In some embodiments of the present application, determining the semantic similarity between each sentence in the first contract text and the corresponding comparison sentence through the semantic similarity calculation module of the similarity calculation model includes: determining the word vector of each word in each sentence in the first contract text to obtain the first word vector set corresponding to each sentence in the first contract text; determining the word vector of each word in each comparison sentence to obtain the second word vector set corresponding to the comparison sentence of each sentence in the first contract text; aggregating the first word vector set into a first sentence vector through a pooling operation, and aggregating the second word vector set into a second sentence vector; calculating the third cosine similarity between the first sentence vector of each sentence and the second sentence vector of the corresponding comparison sentence of each sentence, and determining the third cosine similarity as the semantic similarity between each sentence in the first contract text and the corresponding comparison sentence.
[0013] In some embodiments of the present application, determining a second comparison result between each sentence in the first contract text and the corresponding comparison sentence based on text similarity includes: when the text similarity is lower than a preset threshold, marking the positions of the sentences in the first contract text with text similarity lower than the preset threshold and the corresponding comparison sentences to obtain marking information, and determining the marking information as the second comparison result.
[0014] According to another aspect of the embodiments of the present application, there is also provided a comparison device for contract texts, including: an acquisition module for acquiring a first contract document and a second contract document, where the first contract document and the second contract document are different versions of the same contract; a first analysis module for performing layout analysis on the first contract document through a document layout segmentation model to obtain first document data corresponding to the first contract document; a second analysis module for performing layout analysis on the second contract document through the document layout segmentation model to obtain second document data corresponding to the second contract document; a comparison module for comparing the first document data and the second document data to obtain a first comparison result between the first contract document and the second contract document, where the first comparison result is used to indicate the difference points and the positions of the difference points existing in the contract documents of the first contract document and the second contract document.
[0015] According to another aspect of the embodiments of the present application, there is also provided a non-volatile storage medium in which a program is stored. When the program runs, it controls the device where the non-volatile storage medium is located to execute the contract text comparison method of any one of the above.
[0016] According to another aspect of the embodiments of the present application, there is also provided an electronic device, including: a memory and a processor, where the processor is used to run the program stored in the memory. When the program runs, it executes the contract text comparison method of any one of the above.
[0017] According to another aspect of the embodiments of the present application, there is also provided a computer program product, including computer instructions, which implement the contract text comparison method of any one of the above when executed by a processor.
[0018] In the embodiments of the present application, the method includes obtaining a first contract document and a second contract document, where the first contract document and the second contract document are different versions of the same contract; performing layout analysis on the first contract document through a document layout segmentation model to obtain first document data corresponding to the first contract document; performing layout analysis on the second contract document through the document layout segmentation model to obtain second document data corresponding to the second contract document; comparing the first document data and the second document data to obtain a first comparison result between the first contract document and the second contract document, where the first comparison result is used to indicate the difference points between the first contract document and the second contract document and the positions of the difference points. By performing layout analysis on the documents to be compared through the document layout segmentation model respectively, accurate layout analysis results, that is, the first document data and the second document data, are obtained, and then the first document data and the second document data are compared to obtain the first comparison result between the first contract document and the second contract document. The purpose of accurately comparing different versions of contract documents is achieved, and further, the technical problem in the related art that the accuracy rate of document comparison is relatively low when comparing different versions of contract documents is solved. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The drawings described herein are used to provide a further understanding of the present application and form a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation of the present application. In the drawings:
[0020] Figure 1 is a hardware structure block diagram of a computer terminal for implementing a method for comparing contract texts according to an embodiment of the present application;
[0021] Figure 2 is a flowchart of a method for comparing contract texts according to an embodiment of the present application;
[0022] Figure 3 is a network structure diagram of an improved Segformer model according to an embodiment of the present application;
[0023] Figure 4 is a network structure diagram of a similarity calculation model according to an embodiment of the present application;
[0024] Figure 5 is a structural schematic diagram of a device for comparing contract texts according to an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] To enable those skilled in the art to better understand the solution of this application, the technical solution in the embodiments of this application will be clearly and completely described below in conjunction with the accompanying drawings in the embodiments of this application. Obviously, the described embodiments are only a part of the embodiments of this application, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in this application without creative efforts shall fall within the protection scope of this application.
[0026] The information collected in the embodiments of this application is information and data authorized by the user or fully authorized by all parties. Moreover, for the processing of relevant data such as collection, storage, use, processing, transmission, provision, disclosure, and application, all comply with the relevant laws, regulations, and standards in the relevant regions, necessary confidentiality measures are taken, it does not violate public order and good customs, and a corresponding operation entry is provided for the user to choose to authorize or reject the automated decision result; if the user chooses to reject, the expert decision-making process will be entered.
[0027] It should be noted that the terms "first", "second", etc. in the description and claims of this application and the above-mentioned accompanying drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such data can be interchanged under appropriate circumstances so that the embodiments of this application described here can be implemented in an order other than those illustrated or described here. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0028] To better understand the embodiments of this application, the technical terms involved in the embodiments of this application are explained as follows:
[0029] Optical Character Recognition (OCR for short): It is a technology that converts the text in an image into computer-readable text information through image processing and pattern recognition algorithms.
[0030] In the related art, the management of contracts relies on manual comparison between paper contracts and electronic contracts, which is time-consuming, inefficient, and error-prone. The accuracy of document layout analysis on complex documents is insufficient, especially when dealing with areas of different sizes and uneven sample quantities. Secondly, when processing complex contract texts, it performs poorly on key indicators and is difficult to capture long-distance syntactic structures. Therefore, when comparing different versions of contract documents, the accuracy of document comparison is low, and it is impossible to quickly identify the differences in the same contract in different versions. Therefore, there is a technical problem that the accuracy of document comparison in the contract document comparison method in the related art is low when comparing different versions of contract documents. To solve this problem, relevant solutions are provided in the embodiments of the present application, which are described in detail below.
[0031] According to the embodiments of the present application, an embodiment of a method for comparing contract texts is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0032] The method embodiments provided by the embodiments of the present application can be executed on a computer terminal or a similar computing device. Figure 1 A hardware structure block diagram of a computer terminal for implementing the method for comparing contract texts is shown. As Figure 1 shown, the computer terminal 10 may include one or more (shown as 102a, 102b,..., 102n in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA), a memory 104 for storing data, and a transmission device 106 for communication functions. In addition, it may further include: a display, an input / output interface (I / O interface), a universal serial bus (USB) port (which may be included as one of the ports of the BUS bus), a network interface, a power supply, and / or a camera. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above electronic device. For example, the computer terminal 10 may further include more or fewer components than Figure 1 shown, or have a different configuration from Figure 1 shown.
[0033] It should be noted that the above-mentioned one or more processors 102 and / or other data processing circuits can generally be referred to as "data processing circuits" herein. The data processing circuit can be embodied in software, hardware, firmware, or any combination thereof, in whole or in part. In addition, the data processing circuit can be a single independent processing module, or be incorporated in whole or in part into any one of the other elements in the computer terminal 10. As involved in the embodiments of the present application, the data processing circuit is a kind of processor control (such as the selection of a variable resistance terminal path connected to an interface).
[0034] The memory 104 can be used to store software programs and modules of application software, such as the program instructions / data storage devices corresponding to the contract text comparison method in the embodiments of the present application. The processor 102 executes various functional applications and data processing by running the software programs and modules stored in the memory 104, that is, implements the above-mentioned contract text comparison method. The memory 104 can include high-speed random access memory, and can also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memories. In some instances, the memory 104 can further include a memory remotely set relative to the processor 102, and these remote memories can be connected to the computer terminal 10 through a network. Examples of the above-mentioned network include but are not limited to the Internet, intranet, local area network, mobile communication network, and combinations thereof.
[0035] The transmission device 106 is used to receive or send data via a network. Specific examples of the above-mentioned network can include the wireless network provided by the communication provider of the computer terminal 10. In one instance, the transmission device 106 includes a network adapter (Network Interface Controller, NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 can be a radio frequency (RF) module, which is used to communicate with the Internet wirelessly.
[0036] The display can be, for example, a touch-screen liquid crystal display (LCD), which enables the user to interact with the user interface of the computer terminal 10.
[0037] Under the above operating environment, the embodiments of the present application provide an embodiment of a contract text comparison method. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0038] As Figure 2As shown in the figure, it is a flowchart of a method for comparing contract texts provided according to an embodiment of the present application, including:
[0039] Step S202, obtain a first contract document and a second contract document, where the first contract document and the second contract document are different versions of the same contract.
[0040] In the technical solution provided in step S202, although the obtained first contract document and second contract document essentially belong to the same contract, they represent different versions of the contract and have the following differences: changes in text content (including modification, deletion, or addition of new clauses), format and layout adjustments: even if the text content remains unchanged, the layout of the contract may be different, such as margins, font sizes, line spacings, paragraph indents, etc. The presence or absence of signatures or seals, addition or deletion of attachments, etc. The contract document can be in various formats, such as PDF, word, pictures, etc. If it is in the form of a picture or a scanned copy, these contracts are classified according to their respective versions and forms and stored in the database.
[0041] Step S204, perform layout analysis on the first contract document through a document layout segmentation model (for example, an improved Segformer model. The SegFormer model is an advanced image segmentation architecture, full name "Segmentation Transformer", and the SegFormer model has demonstrated excellent performance in various image segmentation tasks), to obtain first document data corresponding to the first contract document.
[0042] In the technical solution provided in step S204, the document layout segmentation model includes an encoder and a decoder. Each network block of the encoder includes a first channel attention module, and the decoder includes a second channel attention module; there are various ways to perform layout analysis on the first contract document through the document layout segmentation model to obtain the first document data corresponding to the first contract document. For example: use the first channel attention module to determine the first attention weights of each channel when the encoder performs layout analysis; and weight each channel when the encoder performs layout analysis based on the first attention weights to obtain each weighted channel of the encoder; use each weighted channel of the encoder to perform per-channel feature extraction on the first contract document to obtain multi-level features of the first contract document; use the second channel attention module to determine the second attention weights of each channel of the decoder, and weight each channel of the decoder based on the second attention weights to obtain each weighted channel of the decoder; perform feature fusion on the multi-level features of the first contract document based on each weighted channel of the decoder to obtain a fused feature map; determine the first document data corresponding to the first contract document based on the fused feature map.
[0043] The following are specific embodiments:
[0044] For the first contract document in a non-word format, first use appropriate tools or libraries to convert these files into standard image formats such as JPEG or PNG. Then, preprocess the images, including but not limited to unifying the size of the input images to meet the input requirements of the model. Normalize the pixel values of the images, usually scaling the pixel value range to between [0,1], and perform necessary data augmentation (e.g., increasing the diversity of the dataset and the generalization ability of the model through operations such as random flipping, rotation, and scaling) to ensure the consistency of the model input and improve the generalization ability of the model.
[0045] Improvements to the improved Segformer model include adding a first channel attention module (e.g., the Efficient Channel Attention mechanism, simply referred to as the ECA module) to each network block (Transformer Block) in the encoder, and adding a second channel attention module (e.g., the Squeeze-and-Excitation mechanism, simply referred to as the SE module) to the decoder. The improved SegFormer encoder integrates the first channel attention module in each Transformer block to dynamically adjust the weights of each channel. The specific steps are as follows: After inputting the preprocessed image, the SegFormer encoder first extracts low-level features through a convolutional layer. In each network block of the encoder, the ECA module dynamically calculates and assigns attention weights to different channels, strengthens important features, and suppresses irrelevant information, thereby obtaining a more focused and accurate feature representation. Finally, multi-level features of the first contract document are obtained under multiple weighted channels. The decoder part is responsible for fusing and upsampling the multi-level features extracted by the encoder to restore the spatial information of the image. In this process, the second channel attention module (SE module) integrated in the decoder further optimizes the representation of the feature map. The specific operation is as follows: Starting from the output of the encoder, the decoder gradually upsamples the feature map, and at the same time uses the SE module to re-weight the channels in the feature map of the decoder, highlighting the features crucial for layout analysis. Especially when dealing with complex elements such as tables, pictures, and text paragraph boundaries, the segmentation accuracy is improved. This feature fusion method can effectively retain and enhance multi-scale information, providing a more detailed and clear feature representation for subsequent layout prediction. The fused feature map output by the decoder is processed by a multi-classification head, and a class is assigned to each pixel point. The specific steps are as follows: In the fused feature map, each pixel point is assigned a probability vector, and each element of the vector corresponds to the probability of a different layout element class (referring to text regions, images, or tables). The probability vector is converted into a normalized class probability distribution, and the class with the highest probability is selected as the prediction result, thereby obtaining a preliminary layout segmentation map. The preliminary layout segmentation map may contain some noise or discontinuous regions and needs to be post-processed and optimized: Remove small noise points in the segmentation map and optimize the boundary coherence. Identify continuous pixels of the same class in the segmentation map and cluster them into separate layout element regions to obtain clear and complete layout element boundary information. Convert the optimized layout segmentation information into structured data: Each layout element (table, picture, text content) is assigned a unique ID, and information such as its type, position coordinates, and size is recorded to form a standardized layout analysis result, that is, the first document data. These structured data can be used to guide subsequent OCR recognition and text processing processes to ensure the correct interpretation and analysis of all layout elements. The structured information is stored in a database as input data for subsequent processing processes.For the recognized text content (or text paragraph), use the Conditional Text Proposal Network (CTPN) model (or a model that can perform the same function) to complete the text detection in the text paragraph area and predict the location information of the text area; use the Convolutional Recurrent Neural Network (CRNN) model (or a model that can perform the same function) to complete the text recognition and obtain the final text information.
[0046] As Figure 3 shown, it is a network structure diagram of an improved Segformer model provided according to an embodiment of the present application. First, the Overlap Patch Embeddings processes the first contract document and converts it into a suitable feature representation. Then, the input is fed into multiple network blocks (SegFormer Block 1 to SegFormer Block 4 in the figure) to determine the attention weights of each channel. Among them, in the ECA module, the Efficient Self-Attn captures the dependencies between channels and adaptively adjusts the channel weights. The Mix-FFN is used to perform non-linear transformation and information fusion on the features, and feature extraction and transformation are carried out under the weighted multiple channels. Finally, multi-level features of the first contract document are obtained (H in the figure represents the height of the feature map corresponding to the feature during feature processing, W represents the width of the feature map corresponding to the feature during feature processing, C represents the number of channels, and C at different positions (i.e., C1 - C4) represents the number of channels at different stages of feature processing). Then, the multi-level features enter the Multi-Layer Perceptron Layer (MLP Layer) for further processing. The Multi-Layer Perceptron (MLP) consists of multiple fully connected layers and is used to further map and process the features. UpSample is used to increase the size of the feature map and restore the image resolution. The second attention module (SE module) compresses and excites the channels, adjusts the channel weights, and enhances the weights of important channel features in the decoder. Finally, a fused feature map is obtained (the height of the fused feature map in the figure is H, the width is W, and N cls is the number of channels).
[0047] The improved Segformer model is trained as follows: Obtain the dataset for training, divide the dataset into a training set and a validation set, train the initial improved Segformer model based on the training set, and verify it through the validation set until the preset number of iterations is reached to obtain the improved Segformer model. Specifically:
[0048] Initialize the parameters of the initial improved Segformer model, unify the image size of the dataset, and set the batch size (batch_size) of the model, that is, the size of the training set data processed by the model in each iteration; Initial learning rate (init_lr) size: The learning rate determines the step size of model parameter update. For example, setting the initial learning rate to 0.0001 means that at the beginning of training, the update step size of model parameters is small, which helps for more refined weight adjustment and avoids violent oscillations of the loss function. Momentum parameter (momentum) size: The momentum parameter is used to accelerate gradient descent. By adding the direction of past gradient updates to the current gradient update, it can reduce oscillations during training and help the model converge faster. Weight decay (weight_decay) size: Weight decay is a regularization strategy used to prevent model overfitting. It adds a penalty term of the sum of squares of model parameters to the loss function, prompting the model parameters to converge towards smaller values, thus obtaining a more generalized model; Number of iterations (for example, set to 100 times). At the same time, set the learning rate decay method to gradually decrease according to the cosine annealing strategy. In the initial stage of training, the learning rate is relatively high, which helps the model quickly explore the terrain of the loss function; as training progresses, the learning rate gradually decreases, and the model can more carefully adjust the weights to reach a better solution.
[0049] The loss function for training the improved Segformer model is Dice loss. Dice loss is a loss function used for medical image segmentation tasks, especially in the case of sample imbalance. Dice loss is based on the Dice coefficient, which is a method for measuring set similarity and is commonly used to calculate the similarity between two samples. For example, it is expressed by the following formula:
[0050]
[0051] Among them, the Dice coefficient is used to measure the similarity between the segmentation result predicted by the model and the true label. N represents the number of pixels, i is an integer greater than 0 and less than N, P represents the predicted result map (obtained by the model predicting the training set during training), G represents the true label map (that is, the correct segmentation label provided for each image in the training dataset), p i ∈P, g i ∈G, p i, g i They respectively represent the predicted value and the true value of each pixel in the predicted result map and the true label map. The Dice coefficient can describe the similarity between the predicted result map and the true label map. The value range of the Dice coefficient is from 0 to 1, where 1 indicates perfect overlap and 0 indicates no overlap. In practical applications, the Dice loss used is the complement of the Dice coefficient, and the specific formula is as follows.
[0052]
[0053] This formula is directly applied to the output of the model and the true label to guide the model training. The Dice loss is particularly effective for small objects and boundary regions because it directly optimizes the overlap between the predicted segmentation region and the true segmentation region.
[0054] Step S206: Perform page layout analysis on the second contract document through the document page layout segmentation model to obtain the second document data corresponding to the second contract document.
[0055] In the technical solution provided in step S206, the way that the document page layout segmentation model performs page layout analysis on the second contract document to obtain the second document data corresponding to the second contract document is the same as the way that the document page layout segmentation model performs page layout analysis on the first contract document to obtain the first document data corresponding to the first contract document.
[0056] Step S208: Compare the first document data and the second document data to obtain the first comparison result of the first contract document and the second contract document, where the first comparison result is used to indicate the difference points and the positions of the difference points existing in the contract documents of the first contract document and the second contract document.
[0057] In the technical solution provided in step S208, there are multiple ways to determine the first comparison result of the first contract document and the second contract document based on the first document data and the second document data. For example: perform character recognition on the tables and text areas in the first document data and the second document data to obtain the first contract text corresponding to the first document data and the second contract text corresponding to the second document data, where the first document data includes at least the tables, pictures, and text areas in the first contract document, and the second document data includes at least the tables, pictures, and text areas in the second contract document; vectorize each paragraph in the first contract text to obtain a first vector set, where the first vector set includes vectors corresponding to each paragraph in the first contract text; vectorize each paragraph in the second contract text to obtain a second vector set, where the second vector set includes vectors corresponding to each paragraph in the second contract text; determine the comparison paragraphs in the second contract text corresponding to each paragraph in the first contract text based on the first vector set and the second vector set, and the comparison sentences in the corresponding comparison paragraphs for each sentence in each paragraph of the first contract text; determine the second comparison result of each sentence in the first contract text and the corresponding comparison sentence through a similarity calculation model, and determine the set of the second comparison results of all sentences in the first contract text and the corresponding comparison sentences as the first comparison result.
[0058] There are multiple ways to determine the comparison paragraphs in the second contract text corresponding to each paragraph in the first contract text based on the first vector set and the second vector set, and the comparison sentences in the corresponding comparison paragraphs for each sentence in each paragraph of the first contract text. For example: Step 1: Calculate the first cosine similarity between the first vector and each vector in the second vector set respectively, and determine the vector in the second vector set corresponding to the first cosine similarity with the largest value as the matching vector corresponding to the first vector, and determine the paragraph corresponding to the matching vector as the comparison paragraph of the paragraph corresponding to the first vector, where the first vector is any vector in the first vector set; Step 2: Determine the paragraph corresponding to the first vector as the target paragraph, and for each sentence in the target paragraph, calculate the second cosine similarity with each reference sentence in the comparison paragraph respectively, and determine the reference sentence corresponding to the second cosine similarity with the largest value as the comparison sentence of the sentence in the corresponding target paragraph; Iteratively execute Step 1 to Step 2 until all vectors in the first vector set are traversed to obtain the comparison paragraphs corresponding to each vector in the first vector set, and the comparison sentences corresponding to each sentence in each paragraph.
[0059] The following are specific embodiments:
[0060] Vectorize each paragraph in the first contract text to obtain a first vector set, vectorize each paragraph in the second contract text to obtain a second vector set, calculate the first cosine similarity sim between the vectors corresponding to each paragraph in the first vector set and the vectors corresponding to each paragraph in the second vector set, set the threshold of the cosine value to 0.3, and finally take the paragraph corresponding to the maximum sim as the comparison paragraph of the corresponding paragraph in the first vector set. Confirm the comparison paragraphs in the second contract text corresponding to each paragraph in the first contract text and the comparison sentences in the second contract text corresponding to each sentence in each paragraph of the first contract text in the same way.
[0061] For the above steps, the second comparison result between each sentence in the first contract text and the corresponding comparison sentence can be determined through a similarity calculation model in various ways. For example: determine the syntactic similarity between each sentence in the first contract text and the corresponding comparison sentence through the syntactic similarity calculation module of the similarity calculation model; determine the semantic similarity between each sentence in the first contract text and the corresponding comparison sentence through the semantic similarity calculation module of the similarity calculation model; determine the text similarity between each sentence in the first contract text and the corresponding comparison sentence based on the syntactic similarity, semantic similarity, preset syntactic similarity weight, and preset semantic similarity weight; determine the second comparison result between each sentence in the first contract text and the corresponding comparison sentence based on the text similarity.
[0062] In the above steps, there are various implementation methods for determining the syntactic similarity between each sentence in the first contract text and the corresponding comparison sentence through the syntactic similarity calculation module of the similarity calculation model. For example: perform graph-based dependency syntactic analysis on each sentence in the first contract text and the corresponding comparison sentence to obtain the first syntactic tree of each sentence in the first contract text and the second syntactic tree of the corresponding comparison sentence; extract the first syntactic relationship of the first syntactic tree and the second syntactic relationship of the second syntactic tree; vectorize the first syntactic relationship of each sentence to obtain the first syntactic relationship vector set of each sentence, and vectorize the second syntactic relationship of the corresponding comparison sentence of each sentence to obtain the second syntactic relationship vector set; determine the similarity matrix between each sentence and the corresponding comparison sentence based on the first syntactic relationship vector set and the second syntactic relationship vector set; determine the syntactic similarity between each sentence and the corresponding comparison sentence by summing the maximum values of each row in the similarity matrix.
[0063] The following are specific embodiments:
[0064] For each sentence in the first contract text, a graph-based dependency syntax analysis joint model is applied by the syntax similarity calculation module of a similarity calculation model (e.g., the Contextualized Similarity with Enhanced Numeric-aware Token Embedding and Sentence Transformer, simply referred to as the CoSENT similarity calculation model). Such models can effectively capture the dependency relationships in the sentence structure and generate the first syntax tree for each sentence. Similarly, for each comparison sentence corresponding to each sentence, the same operation is performed to generate the second syntax tree. The obtained first syntax tree and second syntax tree contain the complex syntactic relationships between the words in the sentence, laying the foundation for the calculation of syntax similarity. Extract the first syntactic relationship from the first syntax tree, which involves identifying the dependency relationships between words, such as subject-predicate relationships, object relationships, etc. Similarly, extract the second syntactic relationship from the second syntax tree. This process aims to convert the sentence structure into a quantifiable set of syntactic relationships for subsequent vectorization and similarity calculation. Vectorize the extracted first syntactic relationship and second syntactic relationship. Specifically, convert each syntactic relationship into a fixed-length vector representation. Vectorization can be achieved using pre-trained word embeddings (e.g., the output of a specific layer of Global Vectors for Word Representation, simply referred to as GloVe, or Bidirectional Encoder Representations from Transformers, simply referred to as BERT), or learned through a custom neural network layer. The set of syntactic relationships for each sentence is thus converted into a first set of syntactic relationship vectors and a second set of syntactic relationship vectors. Based on the first set of syntactic relationship vectors and the second set of syntactic relationship vectors, calculate the cosine similarity between each pair of syntactic relationship vectors (i.e., the above-mentioned second cosine similarity) to obtain a similarity matrix A. Cosine similarity measures the consistency of the directions of two vectors, and the closer its value is to 1, the more similar the syntactic relationships are. Take the maximum value of each row of numerical values in the similarity matrix A, which represents the most similar match between the first syntactic relationship of each sentence and all the second syntactic relationships. Subsequently, sum these maximum values and divide by the total number of syntactic relationships in the sentence. The obtained average value is the syntactic similarity GSim of the sentence. GSim reflects the degree of similarity of the sentence at the syntactic structure level, and its value can be used as an important part of the overall sentence similarity calculation.
[0065] Suppose there are the first paragraphs of two contract texts, denoted as C1 and C2 respectively. C1 has 5 sentences and C2 has 6 sentences. First, perform graph-based dependency syntactic analysis on each sentence in C1 to generate 5 syntactic trees, and extract syntactic relations from each tree. Then, perform the same operations on each sentence in C2 to generate 6 syntactic trees and 6 sets of syntactic relations. Next, vectorize these 5 + 6 = 11 sets of syntactic relations to obtain two vector sets. For the first sentence S1 in C1, we calculate its similarity with the set of syntactic relation vectors of all sentences in C2 to obtain a 5×6 similarity matrix. Then, take the maximum value of each row of the matrix to obtain a 5-dimensional vector, where each element represents the maximum syntactic similarity between S1 and a certain sentence in C2. Sum the elements of this vector and divide by the number of syntactic relations in S1 to obtain the average syntactic similarity GSim between S1 and the sentences in C2. Repeat the above steps to calculate the GSim between the remaining sentences in C1 and the sentences in C2, and finally obtain the syntactic similarity of the entire contract paragraph. This method not only considers the lexical-level similarity but also deeply analyzes the sentence structure, thus providing a more comprehensive text similarity evaluation scheme.
[0066] In the above steps, there are various implementation methods for determining the semantic similarity between each sentence in the first contract text and the corresponding comparison sentence through the semantic similarity calculation module of the similarity calculation model. For example: determine the word vectors of each word in each sentence in the first contract text to obtain the first set of word vectors corresponding to each sentence in the first contract text; determine the word vectors of each word in each comparison sentence to obtain the second set of word vectors of the comparison sentences corresponding to each sentence in the first contract text; aggregate the first set of word vectors into a first sentence vector through a pooling operation, and aggregate the second set of word vectors into a second sentence vector; calculate the third cosine similarity between the first sentence vector of each sentence and the second sentence vector of the corresponding comparison sentence of each sentence, and determine the third cosine similarity as the semantic similarity between each sentence in the first contract text and the corresponding comparison sentence.
[0067] The following are specific embodiments:
[0068] For each sentence in the first contract text, the semantic similarity calculation module of the similarity calculation model uses a pre-trained bidirectional Transformer pre-trained model (Bidirectional Encoder Representations from Transformers, abbreviated as BERT) to perform lexical embedding representation. The model can generate a word vector containing its contextual semantics for each word in the sentence. Similarly, for each comparison sentence, the same operation is performed to generate a set of word vectors. The pooling operation is used to aggregate the set of word vectors of each sentence into a sentence vector of a fixed length. The pooling operation can be average pooling, max pooling, or more advanced pooling (taking the first word vector of the BERT model output, which is usually used for sentence representation), depending on specific application requirements and performance considerations. Calculate the cosine similarity between each sentence vector and the corresponding comparison sentence vector (i.e., the above-mentioned third cosine similarity). Cosine similarity is a metric for measuring the similarity of vector directions in two vector spaces, and its value ranges from -1 to 1. The closer the value is to 1, the more similar the two vectors are. By converting the sentence vector into a unit vector and calculating the dot product between the two unit vectors, the cosine similarity can be obtained. The calculated cosine similarity is used as the semantic similarity value MSim between each sentence in the first contract text and the corresponding comparison sentence. MSim measures the similarity degree of the two sentences at the semantic level and is crucial for text comparison and analysis tasks.
[0069] The similarity calculation model determines the text similarity between each sentence in the first contract text and the corresponding comparison sentence based on syntactic similarity, semantic similarity, a preset syntactic similarity weight, and a preset semantic similarity weight. Specifically: Sim(P,Q) = a × MSim(P,Q) + (1 - a)GSim(P,Q),
[0070] where P represents any sentence in the first contract text, Q represents the comparison sentence corresponding to P in the second contract text, Sim(P,Q) represents the text similarity between P and Q, a represents the preset semantic similarity weight (a preset constant), MSim(P,Q) represents the semantic similarity value between P and Q, GSim(P,Q) represents the syntactic similarity value between P and Q, and (1 - a) represents the preset syntactic similarity weight (a preset constant). As Figure 4As shown in the figure, it is a network structure diagram of a similarity calculation model provided according to an embodiment of the present application. First, for a sentence pair (P, Q), where P represents any sentence in the first contract text and Q represents the corresponding comparison sentence in the second contract text. After being embedded and represented by the BERT model, it then undergoes graph-based dependency syntactic analysis through a syntactic similarity calculation module to obtain the syntactic similarity GSim, and the semantic similarity value MSim is obtained through multiple pooling operations of the semantic similarity calculation module. Finally, the text similarity is determined based on the semantic similarity value MSim and the syntactic similarity GSim.
[0071] In the above steps, there are various implementation manners for determining the second comparison result of each sentence in the first contract text and the corresponding comparison sentence based on the text similarity. For example: when the text similarity is lower than a preset threshold, the positions of the sentences in the first contract text with text similarity lower than the preset threshold and the corresponding comparison sentences are marked to obtain marking information, and the marking information is determined as the second comparison result.
[0072] The following are specific embodiments:
[0073] When the text similarity is lower than a preset threshold (for example, 0.7), the positions of the sentences in the first contract text with text similarity lower than the preset threshold and the corresponding comparison sentences are marked (for example, marked in the form of highlighting) to obtain marking information, and it is determined that there is a large change in the text position of this sentence before and after the contract, and the sentence with the changed semantics is output.
[0074] The embodiment of the present application also provides a structural schematic diagram of a contract text comparison device, as Figure 5 shown, including:
[0075] An acquisition module 502, configured to acquire a first contract document and a second contract document, where the first contract document and the second contract document are different versions of the same contract.
[0076] A first analysis module 504, configured to perform layout analysis on the first contract document through a document layout segmentation model to obtain first document data corresponding to the first contract document.
[0077] A second analysis module 506, configured to perform layout analysis on the second contract document through a document layout segmentation model to obtain second document data corresponding to the second contract document.
[0078] A comparison module 508, configured to compare the first document data and the second document data to obtain a first comparison result of the first contract document and the second contract document, where the first comparison result is used to indicate the difference points and the positions of the difference points existing in the contract documents of the first contract document and the second contract document.
[0079] It should be noted that Figure 5 the contract text comparison device shown is used to execute Figure 2 the contract text comparison method shown. Therefore Figure 2 the relevant explanations in the contract text comparison method in also apply to this contract text comparison device, and will not be elaborated here.
[0080] It should be noted that each module in the above contract text comparison device can be a program module (for example, a set of program instructions that implements a specific function), or a hardware module. For the latter, it can be presented in the following forms, but not limited to: the presentation form of each of the above modules is a processor, or the functions of each of the above modules are implemented by a processor.
[0081] The embodiment of the present application also provides a non-volatile storage medium. The non-volatile storage medium includes a stored program. Among them, when the program runs, it controls the device where the non-volatile storage medium is located to execute the above contract text comparison method. For example, obtain a first contract document and a second contract document, where the first contract document and the second contract document are different versions of the same contract; perform layout analysis on the first contract document through a document layout segmentation model to obtain first document data corresponding to the first contract document; perform layout analysis on the second contract document through the document layout segmentation model to obtain second document data corresponding to the second contract document; compare the first document data and the second document data to obtain a first comparison result of the first contract document and the second contract document, where the first comparison result is used to indicate the difference points and the positions of the difference points in the contract documents of the first contract document and the second contract document.
[0082] The embodiment of the present application also provides an electronic device. The electronic device includes a processor, and the processor is used to run a program. Among them, when the program runs, it executes the above contract text comparison method. For example, obtain a first contract document and a second contract document, where the first contract document and the second contract document are different versions of the same contract; perform layout analysis on the first contract document through a document layout segmentation model to obtain first document data corresponding to the first contract document; perform layout analysis on the second contract document through the document layout segmentation model to obtain second document data corresponding to the second contract document; compare the first document data and the second document data to obtain a first comparison result of the first contract document and the second contract document, where the first comparison result is used to indicate the difference points and the positions of the difference points in the contract documents of the first contract document and the second contract document.
[0083] According to another aspect of the embodiments of the present application, a computer program product is further provided, including a computer program, which when executed by a processor, implements the above contract text comparison method. For example, obtain a first contract document and a second contract document, where the first contract document and the second contract document are different versions of the same contract; perform layout analysis on the first contract document through a document layout segmentation model to obtain first document data corresponding to the first contract document; perform layout analysis on the second contract document through the document layout segmentation model to obtain second document data corresponding to the second contract document; compare the first document data and the second document data to obtain a first comparison result between the first contract document and the second contract document, where the first comparison result is used to indicate the difference points and the positions of the difference points existing in the contract documents of the first contract document and the second contract document.
[0084] In the above embodiments of the present application, the descriptions of each embodiment have their own emphases. For parts not detailed in a certain embodiment, reference may be made to the relevant descriptions of other embodiments.
[0085] In several embodiments provided by the present application, it should be understood that the disclosed technical content can be implemented in other ways. Among them, the device embodiments described above are only illustrative. For example, the division of the units can be a logical function division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces. The indirect coupling or communication connection of units or modules can be in an electrical or other form.
[0086] The units described as separate components may or may not be physically separated. The components displayed as units may or may not be physical units, that is, they can be located in one place, or they can be distributed to multiple units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0087] In addition, each functional unit in the various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.
[0088] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the related technology, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of this application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, read-only memories (ROM, Read-Only Memory), random access memories (RAM, Random Access Memory), mobile hard disks, magnetic disks, or optical discs.
[0089] The above are only the preferred embodiments of this application. It should be noted that for those of ordinary skill in the art, without departing from the principle of this application, several improvements and refinements can still be made, and these improvements and refinements should also be regarded as the protection scope of this application.
Claims
1. A method for comparing contract texts, characterized in that Including: Obtain a first contract document and a second contract document, where the first contract document and the second contract document are different versions of the same contract; Perform layout analysis on the first contract document through a document layout segmentation model to obtain first document data corresponding to the first contract document; Perform layout analysis on the second contract document through the document layout segmentation model to obtain second document data corresponding to the second contract document; Compare the first document data and the second document data to obtain a first comparison result between the first contract document and the second contract document, where the first comparison result is used to indicate the difference points and the positions of the difference points in the contract documents of the first contract document and the second contract document.
2. The method according to claim 1, characterized in that The document layout segmentation model includes an encoder and a decoder. Each network block of the encoder includes a first channel attention module, and the decoder includes a second channel attention module; The performing layout analysis on the first contract document through the document layout segmentation model to obtain the first document data corresponding to the first contract document includes: Use the first channel attention module to determine the first attention weights of each channel when the encoder performs layout analysis; and weight each channel when the encoder performs layout analysis based on the first attention weights to obtain each weighted channel of the encoder; Use each weighted channel of the encoder to perform channel-by-channel feature extraction on the first contract document to obtain multi-level features of the first contract document; Use the second channel attention module to determine the second attention weights of each channel of the decoder, and weight each channel of the decoder based on the second attention weights to obtain each weighted channel of the decoder; Perform feature fusion on the multi-level features of the first contract document based on each weighted channel of the decoder to obtain a fused feature map; Determine the first document data corresponding to the first contract document based on the fused feature map.
3. The method according to claim 1, characterized in that Determining the first comparison result between the first contract document and the second contract document according to the first document data and the second document data includes: Perform character recognition on the tables and text areas in the first document data and the second document data to obtain a first contract text corresponding to the first document data and a second contract text corresponding to the second document data, where the first document data at least includes the tables, pictures, and text areas in the first contract document, and the second document data at least includes the tables, pictures, and text areas in the second contract document; Vectorize each paragraph in the first contract text to obtain a first vector set, where the first vector set includes vectors corresponding to each paragraph in the first contract text; Vectorize each paragraph in the second contract text to obtain a second vector set, where the second vector set includes vectors corresponding to each paragraph in the second contract text; Determine the corresponding comparison paragraphs in the second contract text for each paragraph in the first contract text based on the first vector set and the second vector set, and the comparison sentences in the corresponding comparison paragraphs for each sentence in each paragraph of the first contract text; Determine the second comparison results between each sentence in the first contract text and the corresponding comparison sentences through a similarity calculation model, and determine the set of the second comparison results between all sentences in the first contract text and the corresponding comparison sentences as the first comparison result.
4. The method according to claim 3, wherein The determining the corresponding comparison paragraphs in the second contract text for each paragraph in the first contract text based on the first vector set and the second vector set, and the comparison sentences in the corresponding comparison paragraphs for each sentence in each paragraph of the first contract text includes: Step 1: Calculate the first cosine similarity between the first vector and each vector in the second vector set respectively, and determine the vector in the second vector set corresponding to the first cosine similarity with the maximum value as the matching vector for the first vector, and determine the paragraph corresponding to the matching vector as the comparison paragraph for the paragraph corresponding to the first vector, where the first vector is any vector in the first vector set; Step 2: Determine the paragraph corresponding to the first vector as the target paragraph, for each sentence in the target paragraph, calculate the second cosine similarity with each reference sentence in the comparison paragraph respectively, and determine the reference sentence corresponding to the second cosine similarity with the maximum value as the comparison sentence for the sentence in the corresponding target paragraph; Iteratively execute Step 1 to Step 2 until all vectors in the first vector set are traversed, to obtain the comparison paragraphs corresponding to each vector in the first vector set, and the comparison sentences corresponding to each sentence in each paragraph.
5. The method according to claim 3, characterized in that, The determining the second comparison results between each sentence in the first contract text and the corresponding comparison sentences through the similarity calculation model includes: Determine the syntactic similarity between each sentence in the first contract text and the corresponding comparison sentence through the syntactic similarity calculation module of the similarity calculation model; Determine the semantic similarity between each sentence in the first contract text and the corresponding comparison sentence through the semantic similarity calculation module of the similarity calculation model; Determine the text similarity between each sentence in the first contract text and the corresponding comparison sentence based on the syntactic similarity, the semantic similarity, the preset syntactic similarity weight, and the preset semantic similarity weight; Determine the second comparison results between each sentence in the first contract text and the corresponding comparison sentence based on the text similarity.
6. The method according to claim 5, characterized in that, The determining the syntactic similarity between each sentence in the first contract text and the corresponding comparison sentence through the syntactic similarity calculation module of the similarity calculation model includes: Perform graph-based dependency syntactic analysis on each sentence in the first contract text and the corresponding comparison sentence to obtain the first syntactic tree of each sentence in the first contract text and the second syntactic tree of the corresponding comparison sentence; Extract the first syntactic relationship of the first syntactic tree and the second syntactic relationship of the second syntactic tree; Vectorize the first syntactic relationship of each sentence to obtain a set of first syntactic relationship vectors for each sentence, and vectorize the second syntactic relationship of the corresponding comparison sentence for each sentence to obtain a set of second syntactic relationship vectors; Determine a similarity matrix for each sentence and the corresponding comparison sentence based on the set of first syntactic relationship vectors and the set of second syntactic relationship vectors; Determine the sum of the maximum values of each row in the similarity matrix as the syntactic similarity between each sentence and the corresponding comparison sentence.
7. The method according to claim 5, characterized in that, The determination of the semantic similarity between each sentence in the first contract text and the corresponding comparison sentence by the semantic similarity calculation module of the similarity calculation model includes: Determine the word vectors of each word in each sentence in the first contract text to obtain a set of first word vectors corresponding to each sentence in the first contract text; Determine the word vectors of each word in each comparison sentence to obtain a set of second word vectors corresponding to the comparison sentence of each sentence in the first contract text; Aggregate the set of first word vectors into a first sentence vector through a pooling operation, and aggregate the set of second word vectors into a second sentence vector; Calculate the third cosine similarity between the first sentence vector of each sentence and the second sentence vector of the corresponding comparison sentence of each sentence, and determine the third cosine similarity as the semantic similarity between each sentence in the first contract text and the corresponding comparison sentence.
8. The method according to claim 5, characterized in that, The determination of the second comparison result between each sentence in the first contract text and the corresponding comparison sentence based on the text similarity includes: In the case where the text similarity is lower than a preset threshold, mark the positions of the sentences in the first contract text with text similarity lower than the preset threshold and the corresponding comparison sentences to obtain marking information, and determine the marking information as the second comparison result.
9. A comparison device for contract texts, characterized in that, including: An acquisition module for acquiring a first contract document and a second contract document, where the first contract document and the second contract document are different versions of the same contract; A first analysis module for performing layout analysis on the first contract document through a document layout segmentation model to obtain first document data corresponding to the first contract document; A second analysis module for performing layout analysis on the second contract document through the document layout segmentation model to obtain second document data corresponding to the second contract document; A comparison module for comparing the first document data and the second document data to obtain a first comparison result between the first contract document and the second contract document, where the first comparison result is used to indicate the difference points and the positions of the difference points in the contract documents of the first contract document and the second contract document.
10. A non-volatile storage medium, characterized in that, A program is stored in the non-volatile storage medium, and when the program runs, it controls the device where the non-volatile storage medium is located to execute the contract text comparison method according to any one of claims 1 to 8.
11. An electronic device, characterized in that, including: A memory and a processor, the processor being configured to run a program stored in the memory, wherein, when the program runs, it executes the method for comparing contract texts according to any one of claims 1 to 8.
12. A computer program product, comprising computer instructions, characterized in that, When the computer instructions are executed by the processor, the method for comparing contract texts according to any one of claims 1 to 8 is implemented.