Text comparison method, computer device and computer storage medium
By combining the pre-trained language model and the Transformer bidirectional encoder representation model, the semantic and event consistency verification problems between documents are solved, efficient and reliable document matching is achieved, and the accuracy and efficiency of document content matching is improved.
Patent Information
- Application Number
- CN202210591024.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-27
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2042-05-27
AI Technical Summary
Existing document comparison methods cannot effectively implement event-based semantics and consistency verification, especially when the content is long, it is difficult to scientifically and effectively realize semantics and event consistency verification between documents. The existing technology lacks effective solutions.
The text representation vector model is trained using a pre-trained language model. Combined with unsupervised and supervised learning, the unitized vector is extracted through the text representation vector model, and the text pair matching relationship data set is constructed. The Transformer bidirectional encoder representation model is used for semantic matching and event feature extraction to realize semantic and event consistency verification between documents.
It improves the efficiency and reliability of document matching, realizes semantic and event consistency verification between multiple documents, and improves the accuracy and efficiency of document content matching.
Smart Images

Figure CN115017879B_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of text processing, and specifically to a text comparison method, a computer device, and a computer storage medium. Background Art
[0002] Most of the existing document comparison methods adopt an unsupervised method to calculate the literal coincidence / similarity of the content of specific text paragraphs of two documents, directly determine the candidate paragraph with the highest score, and implement the process of content comparison and information matching, so as to realize the function of prompting the differences between multiple texts.
[0003] The past methods mainly solved the corresponding relationship of document paragraphs, but could not achieve further verification from the perspective of events. In the financial industry, there are generally scenarios that require attention to the event consistency between documents, such as the consistency verification of numerical values between documents, the consistency comparison of reference events between report files and material files, etc. Different people refine, modify, and finally form a summary report for the same data file. Although there are differences in text organization and language expression methods and techniques, the event basis contained in it is unchanged and objectively exists. Further, when the content lengths of two documents are long, there are great challenges in scientifically and effectively realizing the process of semantic and event consistency verification. There is no ready-made solution for this problem in existing papers, patents, and commercial software. Summary of the Invention
[0004] The embodiments of the present application provide a text comparison method, a computer device, and a computer storage medium, which are used to realize the semantic and event consistency verification between multiple documents, and improve the efficiency and reliability of document matching.
[0005] In the first aspect of the embodiments of the present application, a text comparison method is provided, and the method includes:
[0006] Obtain a target document and a comparison document, obtain a pre-trained language model, and train the pre-trained language model according to the target document and the comparison document until the training stops when the convergence condition is met, so as to obtain a text representation vector model;
[0007] Extract the unit vector of the target document and the unit vector of the comparison document according to the text representation vector model, and determine the candidate paragraph of the comparison document from the comparison document according to the unit vector of the target document and the unit vector of the comparison document;
[0008] Construct a text pair matching relationship data set according to the matching relationship between the manually annotated target document and the comparison document, and train the pre-trained language model according to the text pair matching relationship data set to obtain a text pair semantic matching model;
[0009] Calculate the matching relationship probability of each paragraph of the target document with each paragraph in the candidate paragraphs according to the text for the semantic matching model, and respectively determine the maximum matching relationship probability from the multiple matching relationship probabilities corresponding to each paragraph of the target document;
[0010] Prompt that the paragraphs in the target document with the maximum matching relationship probability less than the preset probability do not match any paragraph of the comparison document.
[0011] The second aspect of the embodiments of the present application provides a computer device, including:
[0012] A training unit, configured to obtain a target document and a comparison document, obtain a pre-trained language model, and train the pre-trained language model according to the target document and the comparison document until the training stops when the convergence condition is met, to obtain a text representation vector model;
[0013] A determination unit, configured to extract the unit vector of the target document and the unit vector of the comparison document according to the text representation vector model, and determine the candidate paragraphs of the comparison document from the comparison document according to the unit vector of the target document and the unit vector of the comparison document;
[0014] The training unit is further configured to construct a text pair matching relationship data set according to the matching relationship between the manually annotated target document and the comparison document, and train the pre-trained language model according to the text pair matching relationship data set to obtain a text pair semantic matching model;
[0015] A calculation unit, configured to calculate the matching relationship probability of each paragraph of the target document with each paragraph in the candidate paragraphs according to the text pair semantic matching model, and respectively determine the maximum matching relationship probability from the multiple matching relationship probabilities corresponding to each paragraph of the target document;
[0016] A prompting unit, configured to prompt that the paragraphs in the target document with the maximum matching relationship probability less than the preset probability do not match any paragraph of the comparison document.
[0017] The third aspect of the embodiments of the present application provides a computer device, including a memory and a processor, where the memory stores a computer program, and when the processor executes the computer program, it implements the method in the foregoing first aspect.
[0018] The fourth aspect of the embodiments of the present application provides a computer storage medium, in which instructions are stored, and when the instructions are executed on a computer, the computer is caused to execute the method in the foregoing first aspect.
[0019] From the above technical solutions, it can be seen that the embodiments of the present application have the following advantages:
[0020] In this embodiment, an innovative document comparison method for realizing semantic and event consistency verification is proposed. Starting from the semantic comparison level at the paragraph granularity, NLP is innovatively combined to process the two-stage text matching semantic consistency comparison and event element joint consistency judgment. Through this text comparison method, the process of content matching between documents is solved, and unsupervised learning and supervised learning are combined to jointly improve the efficiency and reliability of matching. At the same time, starting from the fact comparison level at the sentence / phrase granularity, this embodiment innovatively proposes a class of method frameworks based on event element extraction combined with content consistency discrimination to solve the task of event consistency verification. BRIEF DESCRIPTION OF THE DRAWINGS
[0021] Figure 1 It is a schematic flowchart of a text comparison method in an embodiment of the present application;
[0022] Figure 2 It is another schematic flowchart of a text comparison method in an embodiment of the present application;
[0023] Figure 3 It is a schematic structural diagram of a computer device in an embodiment of the present application;
[0024] Figure 4 It is another schematic structural diagram of a computer device in an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0025] The embodiments of the present application provide a text comparison method, a computer device, and a computer storage medium for realizing semantic and event consistency verification between multiple documents and improving the efficiency and reliability of document matching.
[0026] The text comparison method in the embodiments of the present application is described below:
[0027] Please refer to Figure 1 , an embodiment of the text comparison method in the embodiments of the present application includes:
[0028] 101. Obtain a target document and a comparison document, obtain a pre-trained language model, and train the pre-trained language model according to the target document and the comparison document until the training stops when the convergence condition is met, so as to obtain a text representation vector model;
[0029] The method of this embodiment can be applied to a computer device, which can exist in the form of a terminal device or a server device, etc., and is used to provide services and functions of tag calculation and marking for users. When the computer device is a terminal, it can be a terminal device such as a personal computer (PC), a desktop computer, etc.; when the computer device is a server, it can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud databases, cloud computing, and big data and artificial intelligence platforms.
[0030] In this embodiment, a large amount of text paragraph data in a specific technical field can be obtained, and parameter learning can be performed based on pre-trained language models such as the Transformer bidirectional encoder representation model, such as BERT, Roberta, XLNET, etc., and then a pre-trained language model corresponding to each specific technical field is constructed, denoted as ModelA.
[0031] Given multiple documents, including the target document A and the comparison document B, after parsing the file content of each, all text paragraph sets are obtained, denoted as {a1, a2,..., an} and {b1, b2,..., bm} respectively, where n and m represent the number of paragraphs in the target document A and the comparison document B respectively. Therefore, the pre-trained language model can be trained according to the target document and the comparison document until the training of the pre-trained language model meets the convergence condition and stops training, and a text representation vector model is obtained.
[0032] Specifically, the specific implementation of training the pre-trained language model to obtain a text representation vector model may include the following steps:
[0033] Input the target document and the comparison document into the pre-trained language model so that the pre-trained language model performs model training according to the self-supervised learning algorithm and outputs the representation vector of the target document and the representation vector of the comparison document;
[0034] Construct an InfoNCE Loss function, calculate the InfoNCE Loss value according to the representation vector of the target document and the representation vector of the comparison document, and determine that the model training of the pre-trained language model meets the convergence condition when the InfoNCE Loss value meets the preset numerical range, and stop the model training of the pre-trained language model to obtain a text representation vector model.
[0035] For example, let i be any paragraph in the target document A and j be any paragraph in the comparison document B. Using the method of contrastive learning, different data augmentation methods are applied to the texts i and j within a batch (assumed to be 2), such as synonym replacement, word addition and deletion, back translation, dropout and other data augmentation methods. Based on the above ModelA, fixed-dimensional representation vectors of the target document A and the comparison document B are extracted, obtaining the representation vectors vi’ and vi” of the target document A, as well as the representation vectors vj’ and vj” of the comparison document. And an InfoNCE Loss function is constructed, and the InfoNCE Loss value is calculated according to the representation vectors of the target document and the representation vectors of the comparison document. When the InfoNCE Loss value meets the preset numerical range, it is determined that the model training of the pre-trained language model meets the convergence condition, and the model training of the pre-trained language model is stopped, obtaining a text representation vector model, which can be denoted as ModelB. Therefore, in this embodiment, the representation vectors of the target document and the representation vectors of the comparison document are used to calculate the InfoNCE Loss function and the model training is performed according to the calculated InfoNCE Loss value, which can realize the self-supervised learning of the model training.
[0036] 102. Extract the unit vectors of the target document and the unit vectors of the comparison document according to the text representation vector model, and determine the candidate paragraphs of the comparison document from the comparison document according to the unit vectors of the target document and the unit vectors of the comparison document;
[0037] In this embodiment, to extract the unit vectors of the target document and the unit vectors of the comparison document according to the text representation vector model, the specific implementation manner may be to input the paragraph set of the target document and the paragraph set of the comparison document into the text representation vector model, so that the text representation vector model extracts the semantic vectors of each paragraph of the target document and the semantic vectors of each paragraph of the comparison document respectively, and unitize the semantic vectors of each paragraph of the target document and the semantic vectors of each paragraph of the comparison document respectively, obtaining the unit vectors of each paragraph of the target document and the unit vectors of each paragraph of the comparison document.
[0038] For example, continuing with the previous example, input the paragraph set {a1, a2, …, an} of the target document A and the paragraph set {b1, b2, …, bm} of the comparison document B into the text representation vector model ModelB, and extract the semantic vectors of each paragraph of the target document, which can be denoted as {Va1, Va2, …, Van}, and extract the semantic vectors of each paragraph of the comparison document, which can be denoted as {Vb1, Vb2, …, Vbm}.
[0039] In order to maintain the consistency of the training process of the text representation vector model ModelB, distance calculation can be performed using measurement methods such as vector inner product / cosine similarity. The larger the distance value, the closer the vectors are and the more similar the semantics. Also, since the cosine similarity of the unit vectors i and j (with a modulus of 1) is equivalent to the vector inner product, in order to improve the efficiency of candidate text paragraph recall, all text paragraph sets can be unitized respectively, that is, the semantic vectors {Va1, Va2, …, Van} of each paragraph of the target document and the semantic vectors {Vb1, Vb2, …, Vbm} of each paragraph of the comparison document are unitized respectively to obtain the unitized vectors of each paragraph of the target document, which can be denoted as {Va1’, Va2’, …, Van’}, and the unitized vectors of each paragraph of the comparison document, which can be denoted as {Vb1’, Vb2’, …, Vbm’}.
[0040] After obtaining the unitized vectors of the target document and the unitized vectors of the comparison document, candidate paragraphs of the comparison document can be determined from the comparison document according to the unitized vectors of the target document and the unitized vectors of the comparison document. The specific implementation method is to perform matrix calculations on each unitized vector of the target document with the set of unitized vectors of the comparison document respectively to obtain multiple scores corresponding to each unitized vector of the target document, determine the largest K scores from the multiple scores corresponding to each unitized vector of the target document, and determine the paragraphs of the comparison document corresponding to the largest K scores as candidate paragraphs, where K is a positive integer.
[0041] For example, continuing with the previous example, for any text i in the target document A, its corresponding unitized vector is Vai’. Perform matrix calculations on it with the unitized vectors {Vb1’, Vb2’, …, Vbm’} of each paragraph of the comparison document B to obtain the corresponding scores, that is, {Vai’^T*Vb1’, Vai’^T*Vb2’, …, Vai’^T*Vbm’}, and denote the matrix calculation result as {si1, si2, …, sim}. At the same time, the recall number K can be set, obtain the largest K matrix calculation results, and obtain the text paragraphs corresponding to the largest K scores of the matrix calculation results from the comparison document, and determine them as candidate paragraphs, denoted as {b(1), b(2), …, b(k)}. Assume that the value of K is 6, then obtain the text paragraphs corresponding to the largest 6 scores of the matrix calculation results from the comparison document, thus forming candidate paragraphs. By analogy, the candidate paragraphs corresponding to each paragraph in the target document A can be obtained.
[0042] 103. Construct a text pair matching relationship dataset according to the matching relationship between the manually annotated target document and the comparison document, and train a pre-trained language model according to the text pair matching relationship dataset to obtain a text pair semantic matching model;
[0043] In this embodiment, the pre-trained language model may include a bidirectional encoder representation model of Transformer. Training the pre-trained language model to obtain a text pair semantic matching model, and its specific implementation may include the following steps:
[0044] Construct a text pair matching relationship dataset corresponding to each paragraph of the target document. The text pair matching relationship dataset is a set of manually annotated information between any paragraph of the target document and each paragraph in the paragraph set of the comparison document;
[0045] Based on the text pair matching relationship dataset, splice the paragraphs of the target document and the paragraphs of the comparison document to obtain spliced paragraphs, and add the CLS flag bit and the SEP flag bit to the spliced paragraphs;
[0046] Characterize the spliced paragraphs with the CLS flag bit and the SEP flag bit added and input them into the Transformer bidirectional encoder representation model, so that the classification layer of the Transformer bidirectional encoder representation model processes the CLS flag bit of the spliced paragraphs, obtains the predicted probability of the label output by the Transformer bidirectional encoder representation model, and calculates the binary cross-entropy loss function LOSS value according to the predicted probability. When the LOSS value meets the convergence condition, the text pair semantic matching model is obtained.
[0047] Among them, the pre-trained language model trained in this step and the pre-trained language model trained in step 101 may be the same pre-trained language model or different pre-trained language models, which is not limited here.
[0048] For example, continuing with the previous example, for the text paragraph set {a1, a2,..., an} in the target document A, for any paragraph ai, manually screen in the paragraph set {b1, b2,..., bm} of the comparison document B, select the paragraph bj with the closest semantic information, label the paragraph bj as 1, and the rest as 0. This specific method means constructing a text pair matching relationship dataset, denoted as {aibj = 1 if ai is semantically equivalent to bj else 0}, then the manually annotated information (i.e., 1 or 0) between the paragraph ai and each paragraph in the paragraph set of the comparison document B can be determined.
[0049] Next, based on the text pair matching relationship dataset, paragraphs ai and bj are concatenated to obtain a concatenated paragraph, and flag bits [CLS] and [SEP] are added to the concatenated paragraph. After being characterized, it is input into the Transformer bidirectional encoder representation model, so that the classification layer of the Transformer bidirectional encoder representation model processes the CLS flag bit of the concatenated paragraph to obtain the predicted probability of the label output by the Transformer bidirectional encoder representation model. The binary cross-entropy loss function LOSS value is calculated according to the predicted probability. When the LOSS value meets the convergence condition, a text pair semantic matching model is obtained. Specifically, a vector matrix of the text is obtained, and self-attention is used to realize the interaction between text features, that is, for the vector of each word, its Query, Key, and Value vectors are obtained. The Query vector of each word is respectively inner-producted with the key vectors of other words to calculate the attention coefficient between words. Finally, it is multiplied by the Value matrix through softmax to obtain the output vector of each word. In order to realize the classification of sentence pairs, the output vector of the flag bit [CLS] is taken, and after being processed by the classification layer, the predicted probability of each label is obtained. Based on the constructed binary cross-entropy loss function LOSS, the modeling of text pair semantic consistency classification is realized, and a text pair semantic matching model is obtained when the model establishment and training are completed, which can be denoted as ModelC.
[0050] 104. Calculate the matching relationship probability between each paragraph of the target document and each paragraph in the candidate paragraphs according to the text pair semantic matching model, and respectively determine the maximum matching relationship probability from the multiple matching relationship probabilities corresponding to each paragraph of the target document.
[0051] 105. Prompt that the paragraph in the target document with the maximum matching relationship probability less than the preset probability does not match any paragraph in the comparison document.
[0052] In this embodiment, since there will be a large number of negative samples in the relationship dataset, to solve this problem, on the one hand, the downsampling method can be adopted, that is, for any text i in the target document A, its corresponding unit vector is Vai’, and it is matrix-calculated with the unit vectors {Vb1’, Vb2’, …, Vbm’} of each paragraph in the comparison document B to obtain the score corresponding to each paragraph in the target document A, and the negative samples with scores lower than the threshold threshold are screened out, and the strong negative samples are retained; on the other hand, the weighted focal loss method can be adopted to optimize the objective function and reduce the weight of simple samples in the optimization objective.
[0053] After obtaining the candidate paragraphs {b(1), b(2), …, b(k)} corresponding to any paragraph ai of the target document A, establish the input form of {(ai, b(1)), (ai, b(2)), …, (ai, b(k))}, and use ModelC to obtain the matching relationship probabilities {prob(1), prob(2), …, prob(k)} of each text pair (the text pair is the paragraph ai and one paragraph in the candidate paragraphs) respectively, and screen out the matching paragraph bj' with the highest score. At the same time, determine the minimum confidence level alpha (0 < alpha < 1). If max{prob(1), prob(2), …, prob(k)} < alpha, it is prompted that paragraph ai does not match any paragraph of the comparison document. The prompting method here is not limited. For example, paragraph ai can be highlighted, or specific prompting text can be displayed.
[0054] The following will be based on the Figure 1 embodiment shown above, and further describe the embodiments of the present application in detail. Please refer to Figure 2 , another embodiment of the text comparison method in the embodiments of the present application includes:
[0055] 201. Obtain a target document and a comparison document, obtain a pre-trained language model, and train the pre-trained language model according to the target document and the comparison document until the training stops when the convergence condition is met, and obtain a text representation vector model;
[0056] 202. Extract the unit vector of the target document and the unit vector of the comparison document according to the text representation vector model, and determine the candidate paragraphs of the comparison document from the comparison document according to the unit vector of the target document and the unit vector of the comparison document;
[0057] 203. Construct a text pair matching relationship data set according to the matching relationship between the target document and the comparison document marked manually, and train the pre-trained language model according to the text pair matching relationship data set to obtain a text pair semantic matching model;
[0058] 204. Calculate the matching relationship probabilities of each paragraph of the target document with each paragraph in the candidate paragraphs respectively according to the text pair semantic matching model, and determine the maximum matching relationship probability from the multiple matching relationship probabilities corresponding to each paragraph of the target document;
[0059] 205. Prompt that the paragraphs in the target document with the maximum matching relationship probability less than the preset probability do not match any paragraph of the comparison document;
[0060] The operations performed in steps 201 to 205 are similar to the operations performed in steps 101 to 105 in the Figure 1 embodiment shown above, and will not be elaborated here.
[0061] 206. Determine whether the paragraphs of the target document match the paragraphs of the comparison document in terms of event consistency in the paragraphs of the target document that match the comparison document;
[0062] If there is a target paragraph in the target document with a maximum matching relationship probability greater than the preset probability, it is necessary to determine whether the paragraphs of the target document match the paragraphs of the comparison document in terms of event consistency. The specific method is as follows:
[0063] Determine the comparison paragraphs in the comparison document that match the target paragraphs, and perform word segmentation on the target paragraphs and the comparison paragraphs respectively to obtain the input sequence of the target paragraphs and the input sequence of the comparison paragraphs;
[0064] Semantically represent the input sequence of the target paragraphs and the input sequence of the comparison paragraphs respectively according to the siamese network architecture to obtain the context representation corresponding to each word in the input sequence of the target paragraphs, and the context representation corresponding to each word in the input sequence of the comparison paragraphs;
[0065] Establish the event element label categories of the target paragraphs and establish the event element label categories of the comparison paragraphs;
[0066] Perform element extraction modeling on the event element label categories of the target paragraphs and the event element label categories of the comparison paragraphs respectively to obtain the element labels at the corresponding token positions of the target paragraphs and the element labels at the corresponding token positions of the comparison paragraphs.
[0067] For example, assume that for any paragraph ai in the target document A, the paragraph it matches in the comparison document B is denoted as bj. Then, information extraction modeling can be performed on the events contained in these two text paragraphs, and content consistency judgment can be made at the same time. The specific method is to first perform word segmentation on the text pair {ai, bj} respectively to obtain the input sequence of paragraph ai {(ai1, ai2,..., ailm)} and the input sequence of paragraph bj (bj1, bj2,..., bjlm), where lm represents the longest sequence length;
[0068] Next, use the siamese network architecture to semantically represent the above two input sequences respectively. Here, bidirectional encoders such as CNN / RNN / Transformer can be selected to obtain the context representation corresponding to each word in the input sequence of the target paragraphs, and the context representation corresponding to each word in the input sequence of the comparison paragraphs, denoted as {Vai1, Vai2,..., Vailm}, {Vbj1, Vbj2,..., Vbjlm} respectively, where Vai and Vbj represent fixed-length vectors at the corresponding token positions;
[0069] After that, establish the event element label category label_ent for the target paragraph and the event element label category label_ent for the comparison paragraph. Then, perform element extraction modeling on the event element label categories of paragraph ai and the event element label categories of paragraph bj respectively. Here, the element extraction modeling can be achieved by combining the decoding structure. Softmax / CRF / pointer network / Biaffine, etc. can be selected for the element extraction modeling to obtain the element labels {le_ai1, le_ai2, …, le_ailm} at the corresponding token positions in paragraph ai and the element labels {le_bj1, le_bj2, …, le_bjlm} at the corresponding token positions in paragraph bj.
[0070] In this embodiment, based on the true labels, the LOSSner-a loss function corresponding to paragraph ai and the LOSSner-b loss function corresponding to paragraph bj can be constructed, representing the error of the event element extraction of the text paragraph itself.
[0071] After obtaining the element labels at the corresponding token positions in the target paragraph and the element labels at the corresponding token positions in the comparison paragraph, establish the target matrix for the event element label category of the target paragraph and the comparison matrix for the event element label category of the comparison paragraph. Map the output result of each token in the target paragraph to the corresponding vector according to the target matrix to obtain the element label vector at the corresponding token position in the target paragraph. And, map the output result of each token in the comparison paragraph to the corresponding vector according to the comparison matrix to obtain the element label vector at the corresponding token position in the comparison paragraph. Fuse the context representation and the element label vector at the corresponding token position in the target paragraph to obtain the label-fused context vector at the corresponding token position in the target paragraph. And, fuse the context representation and the element label vector at the corresponding token position in the comparison paragraph to obtain the label-fused context vector at the corresponding token position in the comparison paragraph. Fuse the label-fused context vector at the corresponding token position in the target paragraph and the label-fused context vector at the corresponding token position in the comparison paragraph to obtain the interactive attention weighted vector at the corresponding token position in the target paragraph and the interactive attention weighted vector at the corresponding token position in the comparison paragraph.
[0072] For example, continuing with the previous example, construct the target embedding matrix for the event element label categories of the target paragraph, and construct the comparison embedding matrix for the event element label categories of the comparison paragraph. Map the output result of each token in paragraph ai to the corresponding vector according to the target embedding matrix, obtaining the element label vectors at the corresponding token positions in paragraph ai, denoted as {Emble_ai1, Emble_ai2, …, Emble_ailm}, and, map the output result of each token in paragraph bj to the corresponding vector according to the comparison embedding matrix, obtaining the element label vectors at the corresponding token positions in paragraph bj, denoted as {Emble_bj1, Emble_bj2, …, Emble_bjlm}.
[0073] After that, fuse the context representations {Vai1, Vai2, …, Vailm} at the corresponding token positions in paragraph ai with the element label vectors. The fusion method can be addition, obtaining the label-fused context vectors {S_ai1, S_ai2, …, S_ailm} at the corresponding token positions in paragraph ai, and, fuse the context representations {Vbj1, Vbj2, …, Vbjlm} at the corresponding token positions in paragraph bj with the element label vectors, obtaining the label-fused context vectors {S_bj1, S_bj2, …, S_bjlm} at the corresponding token positions in paragraph bj.
[0074] Since the goal of this process is to compare the consistency of the element content between two texts, an element segment flag element entity_mask can be added to the output result of each token. If the token belongs to the segment of any valid entity, the value of entity_mask is 1, otherwise it is 0.
[0075] After that, fuse the label-fused context vectors {S_ai1, S_ai2, …, S_ailm} at the corresponding token positions in paragraph ai with the label-fused context vectors {S_bj1, S_bj2, …, S_bjlm} at the corresponding token positions in paragraph bj. Here, a self-attention mechanism can be introduced and combined with the information of entity_mask to achieve the interaction between elements, obtaining the interaction attention weighted vectors at the corresponding token positions in paragraph ai and the interaction attention weighted vectors at the corresponding token positions in paragraph bj, denoted as {O_ai1, O_ai2, …, O_ailm, O_bj1, O_bj2, …, O_bjlm}.
[0076] After obtaining the interactive attention weighted vectors corresponding to the token positions of the target paragraph and the interactive attention weighted vectors corresponding to the token positions of the comparison paragraph, according to the flag elements of the element segments, obtain the pooling vectors of the interactive attention weighted vectors corresponding to the token positions of the target paragraph, and, according to the flag elements of the element segments, obtain the pooling vectors of the interactive attention weighted vectors corresponding to the token positions of the comparison paragraph. Concatenate the pooling vector of the target paragraph with the pooling vector of the comparison paragraph to obtain a concatenated pooling vector. In the fully connected interaction layer, map the concatenated pooling vector to a value within a preset numerical range according to the sigmoid non-linear function to obtain the target concatenated pooling vector. Construct a binary cross-entropy loss function, construct an optimization objective function according to the target concatenated pooling vector, the adjustment coefficient, and the binary cross-entropy loss function, and perform parameter update according to the gradient descent optimization method to obtain an event matching relationship model.
[0077] For example, continuing with the previous example, divide {O_ai1, O_ai2, …, O_ailm, O_bj1, O_bj2, …, O_bjlm} into two equal parts to obtain the interactive attention weighted vector {O_ai1, O_ai2, …, O_ailm} corresponding to the token positions of paragraph ai, and the interactive attention weighted vector {O_bj1, O_bj2, …, O_bjlm} corresponding to the token positions of paragraph bj. Then, use the entity_mask value to obtain the pooling vectors of each part. Here, the addition method can be used to obtain the pooling vector P_ai corresponding to {O_ai1, O_ai2, …, O_ailm}, and the pooling vector P_bj corresponding to {O_bj1, O_bj2, …, O_bjlm}. Next, concatenate the pooling vector P_ai of paragraph ai with the pooling vector P_bj of paragraph bj to obtain a concatenated pooling vector. In the fully connected interaction layer, map the concatenated pooling vector to a value within a preset numerical range according to the sigmoid non-linear function to obtain the target concatenated pooling vector, where the preset numerical range can be (0, 1). At the same time, construct a binary cross-entropy loss function LOSScls, construct an optimization objective function beta * LOSScls + (1 - beta) * (LOSSner-a + LOSSner-b) according to the target concatenated pooling vector, the adjustment coefficient beta (0 < beta < 1), and the binary cross-entropy loss function LOSScls, and perform parameter update according to the gradient descent optimization method to obtain an event matching relationship model, denoted as ModelD.
[0078] Therefore, after obtaining the event matching relationship model, it is possible to determine whether there is event consistency between any paragraph in the target document and the paragraph matched in the comparison document according to this model. Specifically, the paragraphs matched between the target document and the comparison document are input into the event matching relationship model, so that the event matching relationship model processes the matched paragraphs, and outputs the event element results of the paragraphs in the target document and the event element results of the paragraphs in the comparison document among the matched paragraphs, as well as outputs the event similarity probability between the paragraphs in the target document and the paragraphs in the comparison document among the matched paragraphs. If the event similarity probability is greater than the preset threshold, it is determined that the paragraphs in the target document and the paragraphs in the comparison document among the matched paragraphs conform to event consistency; if the event similarity probability is less than the preset threshold, it is determined that the paragraphs in the target document and the paragraphs in the comparison document among the matched paragraphs do not conform to event consistency.
[0079] For example, continuing with the previous example, for any paragraph ai in the target document A, the paragraph semantically matched in the comparison document is denoted as bj. Through ModelD, the respective event element results can be obtained, denoted as {le_ai1, le_ai2,..., le_ailm} and {le_bj1, le_bj2,..., le_bjlm} respectively. At the same time, the event similarity probability Prob_final between paragraph ai and paragraph bj is output. When Prob_final > classification threshold theta (0 < theta < 1), it is determined that paragraph ai and paragraph bj conform to event consistency; if Prob_final < theta, then paragraph ai and paragraph bj do not conform to event consistency.
[0080] In this embodiment, an innovative document comparison method for realizing semantic and event consistency verification is proposed. Starting from the semantic comparison level at the paragraph granularity, it innovatively combines NLP to process the two-stage text matching semantic consistency comparison and event element joint consistency judgment. Through this text comparison method, the process of content matching between documents is solved, and unsupervised learning and supervised learning are combined with each other to jointly improve the efficiency and reliability of matching. At the same time, starting from the fact comparison level at the sentence / phrase granularity, a class of method frameworks based on event element extraction combined with content consistency discrimination is innovatively proposed to solve the task of event consistency verification.
[0081] The text comparison method in the embodiments of the present application has been described above. Next, the computer device in the embodiments of the present application will be described. Please refer to Figure 3 , one embodiment of the computer device in the embodiments of the present application includes:
[0082] A training unit 301, configured to obtain a target document and a comparison document, obtain a pre-trained language model, and train the pre-trained language model according to the target document and the comparison document until the training stops when a convergence condition is met, so as to obtain a text representation vector model;
[0083] A determination unit 302, configured to extract a unit vector of the target document and a unit vector of the comparison document according to the text representation vector model, and determine candidate paragraphs of the comparison document from the comparison document according to the unit vector of the target document and the unit vector of the comparison document;
[0084] The training unit 301 is further configured to construct a text pair matching relationship dataset according to a matching relationship between the manually annotated target document and the comparison document, and train the pre-trained language model according to the text pair matching relationship dataset to obtain a text pair semantic matching model;
[0085] A calculation unit 303, configured to calculate a matching relationship probability between each paragraph of the target document and each paragraph in the candidate paragraphs according to the text pair semantic matching model, and respectively determine a maximum matching relationship probability from multiple matching relationship probabilities corresponding to each paragraph of the target document;
[0086] A prompting unit 304, configured to prompt that a paragraph in the target document with a maximum matching relationship probability less than a preset probability does not match any paragraph in the comparison document.
[0087] In a preferred implementation manner of this embodiment, the determination unit 302 is specifically configured to:
[0088] Input the paragraph set of the target document and the paragraph set of the comparison document into the text representation vector model, so that the text representation vector model extracts semantic vectors of each paragraph of the target document and semantic vectors of each paragraph of the comparison document respectively;
[0089] Normalize the semantic vectors of each paragraph of the target document and the semantic vectors of each paragraph of the comparison document respectively to obtain unit vectors of each paragraph of the target document and unit vectors of each paragraph of the comparison document;
[0090] The determining the candidate paragraphs of the comparison document from the comparison document according to the unit vector of the target document and the unit vector of the comparison document includes:
[0091] Perform matrix calculations on each unit vector of the target document and the set of unit vectors of the comparison document respectively to obtain multiple scores corresponding to each unit vector of the target document;
[0092] Determine the largest K scores from the multiple scores corresponding to each unit vector of the target document, and determine the paragraphs of the comparison document corresponding to the largest K scores as the candidate paragraphs, where K is a positive integer.
[0093] In a preferred embodiment of this embodiment, the training unit 301 is specifically configured to:
[0094] Input the target document and the comparison document into the pre-trained language model so that the pre-trained language model performs model training according to the self-supervised learning algorithm, and outputs the representation vector of the target document and the representation vector of the comparison document;
[0095] Construct an InfoNCE Loss function, calculate the InfoNCE Loss value according to the representation vector of the target document and the representation vector of the comparison document, and determine that the model training of the pre-trained language model meets the convergence condition when the InfoNCE Loss value satisfies the preset numerical range, and stop the model training of the pre-trained language model to obtain the text representation vector model.
[0096] In a preferred embodiment of this embodiment, the pre-trained language model includes a bidirectional encoder representation model of Transformer;
[0097] The training unit 301 is specifically configured to:
[0098] Construct a text pair matching relationship data set corresponding to each paragraph of the target document, where the text pair matching relationship data set is a set of manually annotated information between any paragraph of the target document and each paragraph in the paragraph set of the comparison document;
[0099] Based on the text pair matching relationship data set, splice the paragraphs of the target document and the paragraphs of the comparison document to obtain spliced paragraphs, and add a CLS flag bit and a SEP flag bit to the spliced paragraphs;
[0100] Characterize the spliced paragraphs with the CLS flag bit and the SEP flag bit added and input them into the Transformer bidirectional encoder representation model, so that the classification layer of the Transformer bidirectional encoder representation model processes the CLS flag bit of the spliced paragraphs to obtain the predicted probability of the label output by the Transformer bidirectional encoder representation model, and calculate the binary cross-entropy loss function LOSS value according to the predicted probability. When the LOSS value meets the convergence condition, obtain the text pair semantic matching model.
[0101] In a preferred implementation of this embodiment, if there is a target paragraph in the target document whose maximum matching relationship probability is greater than the preset probability, the determining unit 302 is further configured to:
[0102] Determine a comparison paragraph in the comparison document that matches the target paragraph, and perform word segmentation on the target paragraph and the comparison paragraph respectively to obtain the input sequence of the target paragraph and the input sequence of the comparison paragraph;
[0103] Semantically represent the input sequence of the target paragraph and the input sequence of the comparison paragraph respectively according to the siamese network architecture to obtain the context representation corresponding to each word in the input sequence of the target paragraph and the context representation corresponding to each word in the input sequence of the comparison paragraph;
[0104] Establish the event element label category of the target paragraph and establish the event element label category of the comparison paragraph;
[0105] Perform element extraction modeling on the event element label category of the target paragraph and the event element label category of the comparison paragraph respectively to obtain the element label at the corresponding token position of the target paragraph and the element label at the corresponding token position of the comparison paragraph.
[0106] In a preferred implementation of this embodiment, the determining unit 302 is further configured to:
[0107] Establish a target matrix for the event element label category of the target paragraph and establish a comparison matrix for the event element label category of the comparison paragraph;
[0108] Map the output result of each token of the target paragraph to a corresponding vector according to the target matrix to obtain the element label vector at the corresponding token position of the target paragraph, and map the output result of each token of the comparison paragraph to a corresponding vector according to the comparison matrix to obtain the element label vector at the corresponding token position of the comparison paragraph;
[0109] Fuse the context representation and the element label vector at the corresponding token position of the target paragraph to obtain the label-fused context vector at the corresponding token position of the target paragraph, and fuse the context representation and the element label vector at the corresponding token position of the comparison paragraph to obtain the label-fused context vector at the corresponding token position of the comparison paragraph;
[0110] Fuse the label fusion context vectors at the token positions corresponding to the target paragraph with the label fusion context vectors at the token positions corresponding to the comparison paragraph to obtain the interactive attention weighted vectors at the token positions corresponding to the target paragraph and the interactive attention weighted vectors at the token positions corresponding to the comparison paragraph.
[0111] In a preferred implementation manner of this embodiment, the determination unit 302 is further configured to:
[0112] Obtain the pooling vector of the interactive attention weighted vector at the token position corresponding to the target paragraph according to the flag element of the element fragment, and obtain the pooling vector of the interactive attention weighted vector at the token position corresponding to the comparison paragraph according to the flag element of the element fragment. Concatenate the pooling vector of the target paragraph with the pooling vector of the comparison paragraph to obtain a concatenated pooling vector;
[0113] Map the concatenated pooling vector to a value within a preset numerical range according to the sigmoid non-linear function in the fully connected interaction layer to obtain a target concatenated pooling vector;
[0114] Construct a binary cross-entropy loss function, construct an optimization objective function according to the target concatenated pooling vector, the adjustment coefficient, and the binary cross-entropy loss function, and update the parameters according to the gradient descent optimization method to obtain an event matching relationship model.
[0115] In a preferred implementation manner of this embodiment, the determination unit 302 is further configured to:
[0116] Input the paragraphs that match between the target document and the comparison document into the event matching relationship model, so that the event matching relationship model processes the matching paragraphs, and outputs the event element results of the paragraphs of the target document and the event element results of the paragraphs of the comparison document in the matching paragraphs, and outputs the event similarity probability between the paragraphs of the target document and the paragraphs of the comparison document in the matching paragraphs;
[0117] If the event similarity probability is greater than a preset threshold, it is determined that the paragraphs of the target document and the paragraphs of the comparison document in the matching paragraphs conform to event consistency;
[0118] If the event similarity probability is less than a preset threshold, it is determined that the paragraphs of the target document and the paragraphs of the comparison document in the matching paragraphs do not conform to event consistency.
[0119] In this embodiment, the operations performed by each unit in the computer device are similar to those described in the foregoing Figures 1 to 2 illustrated embodiment, and will not be described in detail here.
[0120] In this embodiment, an innovative method for document comparison to achieve semantic and event consistency verification is proposed. Starting from the semantic comparison level at the paragraph granularity, it innovatively combines NLP to process the two-stage text matching semantic consistency comparison and event element joint consistency judgment. Through this text comparison method, the process of content matching between documents is solved, realizing the combination of unsupervised learning and supervised learning to jointly improve the efficiency and reliability of matching. At the same time, starting from the fact comparison level at the sentence / phrase granularity, this embodiment innovatively proposes a framework for a class of event element extraction combined with content consistency discrimination methods to solve the task of event consistency verification.
[0121] The computer device in the embodiment of the present application will be described below. Please refer to Figure 4 One embodiment of the computer device in the embodiment of the present application includes:
[0122] The computer device 400 may include one or more central processing units (CPUs) 401 and a memory 405, and one or more applications or data are stored in the memory 405.
[0123] Among them, the memory 405 may be volatile storage or persistent storage. The program stored in the memory 405 may include one or more modules, and each module may include a series of instruction operations on the computer device. Further, the central processing unit 401 may be set to communicate with the memory 405 and execute a series of instruction operations in the memory 405 on the computer device 400.
[0124] The computer device 400 may further include one or more power supplies 402, one or more wired or wireless network interfaces 403, one or more input / output interfaces 404, and / or one or more operating systems, such as Windows ServerTM, Mac OS XTM, UnixTM, LinuxTM, FreeBSDTM, etc.
[0125] The central processing unit 401 may execute the operations performed by the computer device in the foregoing Figures 1 to 2 illustrated embodiment, and details are not described herein again.
[0126] The embodiment of the present application also provides a computer storage medium. One embodiment includes: Instructions are stored in the computer storage medium, and when the instructions are executed on the computer, the computer is caused to execute the operations performed by the computer device in the foregoing Figures 1 to 2 illustrated embodiment.
[0127] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments and will not be elaborated herein.
[0128] In several embodiments provided in the present application, it should be understood that the disclosed systems, devices, and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed couplings or direct couplings or communication connections to each other can be through some interfaces, and the indirect couplings or communication connections of the devices or units can be in electrical, mechanical, or other forms.
[0129] The units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they can be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0130] In addition, in each embodiment of the present application, the functional units can be integrated into one processing unit, or each unit can exist physically alone, or two or more units can be integrated into one unit. The above-mentioned integrated units can be implemented in the form of hardware or in the form of software functional units.
[0131] If the above-mentioned integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in each embodiment of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memories (ROM, read-only memory), random access memories (RAM, random access memory), magnetic disks, or optical discs that can store program codes.
Claims
1. A text comparison method, characterized in that, The method includes: Obtain a target document and a comparison document, and calculate the matching relationship probability between each paragraph of the target document and each paragraph in multiple candidate paragraphs of the comparison document; Determine the maximum matching relationship probability from the multiple matching relationship probabilities corresponding to each paragraph of the target document respectively; Prompt that the paragraphs in the target document with the maximum matching relationship probability less than the preset probability do not match any paragraph in the comparison document; If there is a target paragraph in the target document with the maximum matching relationship probability greater than the preset probability, then the method further includes: Determine the comparison paragraph in the comparison document that matches the target paragraph, tokenize the target paragraph and the comparison paragraph respectively to obtain the input sequence of the target paragraph and the input sequence of the comparison paragraph; Semantically represent the input sequence of the target paragraph and the input sequence of the comparison paragraph respectively according to the siamese network architecture to obtain the context representation corresponding to each word in the input sequence of the target paragraph and the context representation corresponding to each word in the input sequence of the comparison paragraph; Establish the event element label categories of the target paragraph and establish the event element label categories of the comparison paragraph; Perform element extraction modeling on the event element label categories of the target paragraph and the event element label categories of the comparison paragraph respectively to obtain the element labels at the corresponding token positions of the target paragraph and the element labels at the corresponding token positions of the comparison paragraph; Train an event matching relationship model according to the context representation and element labels corresponding to the target paragraph, and the context representation and element labels corresponding to the comparison paragraph, and input the paragraphs that match between the target document and the comparison document into the event matching relationship model to determine whether the paragraphs that match between the target document and the comparison document meet event consistency.
2. The method according to claim 1, wherein The calculation of the matching relationship probability between each paragraph of the target document and each paragraph in the candidate paragraphs includes: Obtain a pre-trained language model, train the pre-trained language model according to the target document and the comparison document, and stop training until the convergence condition is met to obtain a text representation vector model; Extract the unit vector of the target document and the unit vector of the comparison document according to the text representation vector model, and determine the candidate paragraphs of the comparison document from the comparison document according to the unit vector of the target document and the unit vector of the comparison document; Construct a text pair matching relationship data set according to the matching relationship between the target document and the comparison document, and train the pre-trained language model according to the text pair matching relationship data set to obtain a text pair semantic matching model; the matching relationship between the target document and the comparison document is obtained by manual annotation; Calculate the matching relationship probability between each paragraph of the target document and each paragraph in the candidate paragraphs according to the text pair semantic matching model.
3. The method according to claim 2, wherein The extraction of the unit vector of the target document and the unit vector of the comparison document according to the text representation vector model includes: Input the paragraph set of the target document and the paragraph set of the comparison document into the text representation vector model, so that the text representation vector model extracts the semantic vectors of each paragraph of the target document and the semantic vectors of each paragraph of the comparison document respectively; Normalize the semantic vectors of each paragraph of the target document and the semantic vectors of each paragraph of the comparison document respectively to obtain the normalized vectors of each paragraph of the target document and the normalized vectors of each paragraph of the comparison document; The determining of the candidate paragraphs of the comparison document from the comparison document according to the normalized vectors of the target document and the normalized vectors of the comparison document includes: Perform matrix calculations on each normalized vector of the target document and the set of normalized vectors of the comparison document respectively to obtain multiple scores corresponding to each normalized vector of the target document; Determine the largest K scores from the multiple scores corresponding to each normalized vector of the target document, and determine the paragraphs of the comparison document corresponding to the largest K scores as the candidate paragraphs, where K is a positive integer.
4. The method according to claim 2, characterized in that, The training of the pre-trained language model according to the target document and the comparison document until the training stops when the convergence condition is met to obtain the text representation vector model includes: Input the target document and the comparison document into the pre-trained language model so that the pre-trained language model performs model training according to the self-supervised learning algorithm and outputs the representation vector of the target document and the representation vector of the comparison document; Construct an InfoNCE Loss function, calculate the InfoNCE Loss value according to the representation vector of the target document and the representation vector of the comparison document, determine that the model training of the pre-trained language model meets the convergence condition when the InfoNCE Loss value meets the preset numerical range, and stop the model training of the pre-trained language model to obtain the text representation vector model.
5. The method according to claim 2, wherein The pre-trained language model includes a bidirectional encoder representation model of Transformer; The constructing of a text pair matching relationship data set according to the matching relationship between the target document and the comparison document, and training the pre-trained language model according to the text pair matching relationship data set to obtain a text pair semantic matching model includes: Construct a text pair matching relationship data set corresponding to each paragraph of the target document, and the text pair matching relationship data set is a set of manually annotated information between any paragraph of the target document and each paragraph in the paragraph set of the comparison document; Based on the text pair matching relationship data set, splice the paragraphs of the target document and the paragraphs of the comparison document to obtain spliced paragraphs, and add a CLS flag bit and a SEP flag bit to the spliced paragraphs; Characterize the spliced paragraph with the CLS flag bit and SEP flag bit added and input it into the Transformer bidirectional encoder representation model, so that the classification layer of the Transformer bidirectional encoder representation model processes the CLS flag bit of the spliced paragraph to obtain the predicted probability of the label output by the Transformer bidirectional encoder representation model. Calculate the binary cross-entropy loss function LOSS value according to the predicted probability. When the LOSS value meets the convergence condition, obtain the text pair semantic matching model.
6. The method according to claim 1, wherein The method further includes: Establishing a target matrix for the event element label categories of the target paragraph and a comparison matrix for the event element label categories of the comparison paragraph; Mapping the output result of each token of the target paragraph to a corresponding vector according to the target matrix to obtain the element label vector at the corresponding token position of the target paragraph, and mapping the output result of each token of the comparison paragraph to a corresponding vector according to the comparison matrix to obtain the element label vector at the corresponding token position of the comparison paragraph; Fusing the context representation at the corresponding token position of the target paragraph with the element label vector to obtain the label-fused context vector at the corresponding token position of the target paragraph, and fusing the context representation at the corresponding token position of the comparison paragraph with the element label vector to obtain the label-fused context vector at the corresponding token position of the comparison paragraph; Fusing the label-fused context vector at the corresponding token position of the target paragraph with the label-fused context vector at the corresponding token position of the comparison paragraph to obtain the interactive attention weighted vector at the corresponding token position of the target paragraph and the interactive attention weighted vector at the corresponding token position of the comparison paragraph.
7. The method according to claim 6, characterized in that, The method further includes: Obtaining the pooling vector of the interactive attention weighted vector at the corresponding token position of the target paragraph according to the flag element of the element segment, and obtaining the pooling vector of the interactive attention weighted vector at the corresponding token position of the comparison paragraph according to the flag element of the element segment. Concatenating the pooling vector of the target paragraph with the pooling vector of the comparison paragraph to obtain the concatenated pooling vector; Mapping the concatenated pooling vector to a value within a preset numerical range according to the sigmoid non-linear function in the fully connected interaction layer to obtain the target concatenated pooling vector; Constructing a binary cross-entropy loss function, constructing an optimization objective function according to the target concatenated pooling vector, the adjustment coefficient, and the binary cross-entropy loss function, and performing parameter update according to the gradient descent optimization method to obtain the event matching relationship model.
8. The method according to claim 7, wherein The method further includes: Input the paragraphs that match between the target document and the comparison document into the event matching relationship model, so that the event matching relationship model processes the matching paragraphs and outputs the event element results of the paragraphs of the target document and the event element results of the paragraphs of the comparison document in the matching paragraphs, and outputs the event similarity probability between the paragraphs of the target document and the paragraphs of the comparison document in the matching paragraphs; If the event similarity probability is greater than the preset threshold, it is determined that the paragraphs of the target document and the paragraphs of the comparison document in the matching paragraphs meet event consistency; If the event similarity probability is less than the preset threshold, it is determined that the paragraphs of the target document and the paragraphs of the comparison document in the matching paragraphs do not meet event consistency.
9. A computer device, comprising a memory and a processor, the memory storing a computer program, characterized in that, When the processor executes the computer program, it implements the method according to any one of claims 1 to 8.
10. A computer storage medium, characterized in that, Instructions are stored in the computer storage medium, and when the instructions are executed on the computer, the computer is caused to execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Document analysis method and device, intelligent terminal and storage medium
CN112651222A