Electronic document difference analysis method, system and equipment and storage medium

Through preset classification model and matching algorithm, electronic documents are classified, annotated and differentially analyzed, which solves the problem that the existing technology cannot efficiently analyze electronic documents, and achieves the effect of efficient identification and strong readability of differential analysis.

CN120087354APending Publication Date: 2025-06-03BEIJING ANZHENGTONG INFORMATION TECH HLDG CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510578849.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-06-03

AI Technical Summary

Technical Problem

The existing technology cannot efficiently perform differential analysis of electronic documents, the character-level comparison particle size is too fine, and deep learning methods are highly dependent on labeled data, which cannot meet the needs of efficient differential analysis.

Method used

Classify and annotate the document to be analyzed through a preset classification model, extract the text to be analyzed, and use the matching algorithm to obtain the differential characters between the document to be analyzed and the target document, locate the difference field according to the difference characters, and generate field difference annotations on the document.

Benefits of technology

Efficient difference analysis of electronic documents is realized, differential characters between the document to be analyzed and the target document are identified, and the readability of the difference is improved through field difference annotation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120087354A_ABST
    Figure CN120087354A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of natural language processing, and discloses an electronic document difference analysis method, system and device and a storage medium. The method comprises the following steps: classifying and labeling a to-be-analyzed document through a preset classification model to obtain to-be-analyzed information, extracting a to-be-analyzed text from the to-be-analyzed information, performing character matching on the to-be-analyzed text and a target text through a matching algorithm to obtain a matching result, determining difference characters of the to-be-analyzed text and the target text according to the matching result, and analyzing the to-be-analyzed text according to the difference characters. Obtaining target position information of a difference field where the difference character is located; the difference characters comprise a first difference character corresponding to the to-be-analyzed text and a second difference character corresponding to the target text; and according to the target position information and the difference field, generating corresponding field difference annotations on the to-be-analyzed text and the target text respectively. According to the method, the difference characters between the to-be-analyzed document and the target document can be efficiently recognized, the difference is labeled in a field mode, and the readability of the difference is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of natural language processing, and particularly relates to a method, system, device and storage medium for analyzing differences in electronic documents. Background Art

[0002] Currently, the comparison of electronic documents mostly uses preset text processing models and algorithms to achieve automated comparison, such as character-level comparison, deep learning-based comparison, etc. Among them, the difference granularity of the character-level comparison method is too fine and the readability is poor, while the method of comparison based on deep learning is highly dependent on labeled data, and existing methods cannot efficiently analyze the differences in electronic documents. Summary of the Invention

[0003] In order to solve the problem that existing technologies cannot efficiently analyze the differences in electronic documents, the present invention provides a method, system, device and storage medium for analyzing differences in electronic documents.

[0004] In a first aspect, the present invention provides a method for analyzing differences in electronic documents, including: Classifying and annotating a document to be analyzed through a preset classification model to obtain the information to be analyzed in the document to be analyzed, and extracting the text to be analyzed from the information to be analyzed; Performing character matching on the text to be analyzed and a target text through a matching algorithm to obtain a matching result, determining the different characters between the text to be analyzed and the target text according to the matching result, and obtaining the target position information of the different fields where the different characters are located; the different characters include the first different characters corresponding to the text to be analyzed and the second different characters corresponding to the target text; Generating corresponding field difference annotations on the text to be analyzed and the target text respectively according to the target position information and the different fields.

[0005] In an optional implementation manner, the classifying and annotating a document to be analyzed through a preset classification model to obtain the information to be analyzed in the document to be analyzed includes: Classifying the document to be analyzed through the preset classification model to obtain a classification result, and annotating the document to be analyzed according to the classification result to obtain annotation information; the classification result includes header and footer, title, text, table and picture; Obtaining the information type corresponding to the annotation information, and using the target annotation information belonging to the target information type in the annotation information as the information to be analyzed, where the target information type includes the text, the table and the picture.

[0006] In an optional implementation manner, the extracting the text to be analyzed from the information to be analyzed includes: If the information to be analyzed is readable information, directly extract the text to be analyzed from the information to be analyzed; the readable information includes readable text and readable tables; If the information to be analyzed is unreadable information, generate a text detection frame corresponding to the information to be analyzed; the unreadable information includes unreadable text, unreadable tables, and pictures; Use the text detection frame to select the text area in the information to be analyzed, and perform text extraction on the text area to obtain the text to be analyzed.

[0007] In an alternative embodiment, the method further includes: Generate corresponding text positioning coordinates according to the number and positions of the detection frames of the text detection frame; Perform word segmentation on the text to be analyzed by a preset word segmentation method to obtain multiple fields, and generate position information for each field according to the text positioning coordinates.

[0008] In an alternative embodiment, the obtaining the target position information of the difference field where the difference character is located includes: Perform character-by-character matching on each field in the text to be analyzed and the target field in the target text to obtain the difference character; Determine the difference field corresponding to each difference character from the fields, and generate the target position information according to the text positioning coordinates and the difference field.

[0009] In an alternative embodiment, the separately generating corresponding difference annotations on the text to be analyzed and the target text includes: Obtain the number of difference characters in each difference field; If the number of difference characters is greater than a preset number threshold, generate a difference label corresponding to the difference character; the difference label includes a set of multiple difference annotations; If the number of difference characters is less than or equal to the preset number threshold, generate the difference annotation corresponding to the difference character.

[0010] In an alternative embodiment, the extracting the text to be analyzed from the information to be analyzed includes: If the object to be analyzed is not a text object, select a corresponding information extraction model according to the type of the object to be analyzed, and use the information extraction model to extract the object to be analyzed from the information to be analyzed, where the type of the object to be analyzed includes electronic signatures and images; If the object to be analyzed is the text object, extract the text to be analyzed from the information to be analyzed.

[0011] In a second aspect, the present invention provides an electronic document difference analysis system, including: A classification module, configured to classify and label a document to be analyzed through a preset classification model, obtain information to be analyzed in the document to be analyzed, and extract text to be analyzed from the information to be analyzed; A matching module, configured to perform character matching on the text to be analyzed and a target text through a matching algorithm to obtain a matching result, determine different characters between the text to be analyzed and the target text according to the matching result, and obtain target position information of a target field where the different characters are located; the different characters include first different characters corresponding to the text to be analyzed and second different characters corresponding to the target text; A comment module, configured to generate corresponding field difference comments on the text to be analyzed and the target text respectively according to the target position information and the different fields.

[0012] In a third aspect, the present invention provides a computer device, which includes a processor and a memory. The memory stores a computer program, and the processor is configured to execute the computer program to implement the electronic document difference analysis method described in the first aspect.

[0013] In a fourth aspect, the present invention provides a computer storage medium, which stores a computer program. When the computer program is executed on a processor, it implements the electronic document difference analysis method described in the first aspect.

[0014] The embodiments of the present invention have the following beneficial effects: The electronic document difference analysis method provided by the present invention classifies and labels a document to be analyzed, thereby extracting text to be analyzed, then uses a matching algorithm to obtain different characters between the document to be analyzed and a target document, locates the different fields according to the different characters, and performs difference marking on the document to be analyzed and the target document respectively according to the different fields. This application can efficiently identify different characters between the document to be analyzed and the target document, and marks the differences in the form of fields, improving the readability of the differences. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the present invention, the drawings required for the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention, and thus should not be regarded as limiting the protection scope of the present invention. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.

[0016] Figure 1 The flowchart showing a method for analyzing differences between electronic documents provided by an embodiment of the present application; Figure 2 shows a schematic flowchart of a method for obtaining information to be analyzed provided by an embodiment of the present application; Figure 3 shows a schematic flowchart of a method for extracting text to be analyzed provided by an embodiment of the present application; Figure 4 shows a schematic framework diagram of an electronic document difference analysis system provided by an embodiment of the present application. Detailed implementation manners

[0017] Next, the technical solutions in the embodiments of the present invention will be clearly and completely described with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments.

[0018] Generally, the components of the embodiments of the present invention described and illustrated herein can be arranged and designed in various different configurations. Therefore, the following detailed description of the embodiments of the present invention provided in the drawings is not intended to limit the scope of the claimed present invention, but merely represents selected embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative efforts fall within the scope of protection of the present invention.

[0019] In the following, the terms "including", "having" and their cognates that can be used in various embodiments of the present invention are only intended to represent specific features, numbers, steps, operations, elements, components or combinations of the foregoing items, and should not be construed as first excluding the existence of one or more other features, numbers, steps, operations, elements, components or combinations of the foregoing items or increasing the possibility of one or more features, numbers, steps, operations, elements, components or combinations of the foregoing items.

[0020] In addition, the terms "first", "second", "third", etc. are only used for distinguishing descriptions and cannot be construed as indicating or implying relative importance.

[0021] Unless otherwise defined, all terms (including technical terms and scientific terms) used herein have the same meaning as commonly understood by those of ordinary skill in the art to which various embodiments of the present invention belong. The terms (such as those defined in a general-use dictionary) will be interpreted as having the same meaning as the contextual meaning in the relevant technical field and will not be interpreted as having an idealized meaning or an overly formal meaning unless clearly defined in various embodiments of the present invention.

[0022] In the case of no conflict, the following embodiments and the features in the embodiments can be combined with each other.

[0023] Refer to Figure 1 ,Figure 1 The flowchart of an electronic document difference analysis method provided in this embodiment. The method includes: S101. Classify and annotate the document to be analyzed through a preset classification model to obtain the information to be analyzed in the document to be analyzed, and extract the text to be analyzed from the information to be analyzed.

[0024] Electronic documents are very common in daily life. For some ordinary electronic documents, the requirements for their accuracy may not be high. However, for electronic documents involved in industries such as finance and law, such as electronic contracts, the requirements are extremely strict. Small errors and differences may bring large economic losses and disputes. Therefore, strict review and difference analysis of electronic contracts are required.

[0025] An electronic contract is usually formulated by one party and then reviewed and modified by the other party until both parties reach an agreement. For example, Party B first formulates an electronic contract A and sends it to Party A. Party A makes changes, signs, or seals on the basis of the electronic contract A and then returns it to Party B. The modified electronic contract is defined as electronic contract B. After receiving electronic contract B, Party B needs to determine the specific modification information on electronic contract B. At this time, it is necessary to perform difference analysis and comparison between electronic contract A and electronic contract B.

[0026] Therefore, the document to be analyzed can be either electronic contract A or electronic contract B, or perform comparison and analysis on both electronic contract A and electronic contract B at the same time. There are usually various types of information in an electronic contract, such as contract title, header and footer, contract text, tables, pictures, contract seals, signatures, etc. Some of this information is unimportant. Therefore, before performing difference analysis on an electronic contract, a preset classification model can be used to classify and annotate the electronic document, distinguish which information is the contract title, which information is the header and footer, which information is the contract text, etc., and annotate this information. According to the annotation, extract the text to be analyzed from the information to be analyzed. The text to be analyzed usually includes contract text, text in tables, and text in pictures, etc.

[0027] S102. Perform character matching between the text to be analyzed and the target text through a matching algorithm to obtain a matching result. Determine the different characters between the text to be analyzed and the target text according to the matching result, and obtain the target position information of the different fields where the different characters are located; the different characters include the first different characters corresponding to the text to be analyzed and the second different characters corresponding to the target text.

[0028] The text to be analyzed and the target text can be segmented simultaneously by using a predefined dictionary or rules, in combination with rules such as the forward maximum matching method and the bidirectional maximum matching method, to obtain the fields to be analyzed of the text object to be analyzed and the corresponding target fields of the target text. Then, character matching is performed on the fields to be analyzed and the target fields to determine the different characters. Then, the different fields where the different characters are located are reversely searched according to the different characters, and the target position information of the different fields is determined.

[0029] S103. Generate corresponding field difference annotations on the text to be analyzed and the target text respectively according to the target position information and the different fields.

[0030] According to the different fields and the target position information, corresponding field difference annotations are made on the text to be analyzed and the target text respectively, so as to realize the difference comparison between the text to be analyzed and the target text.

[0031] In this embodiment, by classifying and annotating the document to be analyzed, the text to be analyzed is extracted, and then the matching algorithm is used to obtain the different characters between the document to be analyzed and the target document. Then, the different fields are located according to the different characters, and difference annotations are made on the document to be analyzed and the target document respectively according to the different fields. The present application can efficiently identify the different characters between the document to be analyzed and the target document, and mark the differences in the form of fields, improving the readability of the differences.

[0032] Refer to Figure 2 , step S101 includes: steps S1011 - S1012.

[0033] S1011. Classify the document to be analyzed through the preset classification model to obtain a classification result, and annotate the document to be analyzed according to the classification result to obtain annotation information; the classification result includes header and footer, title, text, table and picture.

[0034] S1012. Obtain the information type corresponding to the annotation information, and use the target annotation information belonging to the target information type in the annotation information as the information to be analyzed, where the target information type includes the text, the table and the picture.

[0035] The preset classification model can be a neural network model. After obtaining the document to be analyzed, the annotation information can be classified first, and the parts containing text in the electronic document can be divided to obtain the header and footer, title, text, table, picture, etc. Then, corresponding annotations are given to each part. For example, the header and footer, title and other information can be marked as the first information type, and the text, table and picture can be marked as the second information type. Then, the information of the first information type is removed, and the information of the second type is determined as the information to be analyzed.

[0036] In this embodiment, by classifying and annotating the text to be analyzed, and removing unimportant information according to the annotation results, only the information to be analyzed that needs to be analyzed for differences is retained, thereby improving the efficiency and accuracy of the difference analysis.

[0037] In one implementation, with reference to Figure 3 , step S101 further includes: steps S1014 - S1016.

[0038] S1014. If the information to be analyzed is readable information, directly extract the text to be analyzed from the information to be analyzed; the readable information includes readable text and readable tables.

[0039] S1015. If the information to be analyzed is unreadable information, generate a text detection box corresponding to the information to be analyzed; the unreadable information includes unreadable text, unreadable tables, and pictures.

[0040] S1016. Use the text detection box to select the text area in the information to be analyzed, and perform text extraction on the text area to obtain the text to be analyzed.

[0041] An electronic document may contain various types of text. For example, if the electronic document is a word document, the text in the word document is basically editable and readable. Or, if the electronic document is a PDF document, the text in the document generally cannot be directly edited and read. Even in some electronic documents, such as a word document that contains both editable and readable text, as well as various types of files such as tables and pictures, when the tables and pictures also contain text that needs to be analyzed for differences, text extraction needs to be performed first.

[0042] Therefore, before extracting the text to be analyzed, it is also necessary to first determine whether the information to be analyzed is readable information. If it is readable information, direct text extraction can be performed. If it is unreadable information, such as tables and pictures, it is necessary to first determine the position of the text in the tables and pictures, then generate a corresponding text detection box, use the text detection box to determine the corresponding text area to be analyzed from the unreadable tables and pictures, and then use methods such as a text extraction model to extract the readable text from the text area to be analyzed.

[0043] In this embodiment, by determining the type of the information to be analyzed, and then selecting the corresponding text extraction method according to the type of the information to be analyzed to extract the corresponding text to be analyzed from the information to be analyzed, the all-round extraction of the text to be analyzed is realized, and the extraction efficiency of the text to be analyzed is improved.

[0044] In one implementation, the method further includes: Generate corresponding text positioning coordinates according to the number and position of the detection boxes of the text detection box; Tokenize the text to be analyzed using a preset tokenization method to obtain multiple fields, and generate position information for each field according to the text positioning coordinates.

[0045] For example, an electronic document includes multiple pages. Taking one page as an example, if the current page is all text, the current page can be framed by a text detection box, and then a corresponding text positioning coordinate system is generated according to the boundary of the text detection box, so as to determine the coordinates of each field and character in the current page.

[0046] If the current page includes both readable text, table text, image text, etc., multiple text detection boxes need to be generated. The size and position of each text detection box may be different. Then, from multiple text detection boxes, the current page can be regarded as a rectangle, and the target text detection boxes at the four corners of the current page are selected. A text positioning coordinate system is constructed according to the target text detection boxes, and the text positioning coordinates corresponding to each field and character in other text detection boxes are generated according to the text positioning coordinate system.

[0047] In this embodiment, a text positioning coordinate system is constructed through the number and position of text detection boxes. According to the text positioning coordinate system, the coordinates of each field and character in the text can be determined, thus providing accurate position information for subsequent document difference analysis.

[0048] In one implementation manner, obtaining the target position information of the difference field where the difference character is located includes: Perform a character-by-character match between each field in the text to be analyzed and the target field in the target text to obtain the difference characters; Determine the difference field corresponding to each difference character from the fields, and generate the target position information according to the text positioning coordinates and the difference field.

[0049] Specifically, a method of character-by-character matching comparison can be used to find the differences between the text to be analyzed and the target text, and determine all the characters with differences. Then, reverse search for all the fields corresponding to the difference characters according to the difference characters. The field can be the sentence or paragraph where the difference character is located. Then count the number of difference characters in each field, and perform different difference displays according to the number of difference characters.

[0050] For example, a preset quantity threshold for difference characters can be set. For a field, if the number of difference characters is greater than the preset quantity threshold, a difference label corresponding to the difference characters is generated; the difference label includes a set of multiple difference annotations; if the number of difference characters is less than or equal to the preset quantity threshold, the difference annotation corresponding to the difference characters is generated.

[0051] When there are multiple different characters in a field, if the method of character differences is used for annotation and display, it will lead to cumbersome and chaotic annotations, which is not convenient for users to view and understand. Therefore, when there are many different characters in a field, the form of difference tags can be used. The content of the difference tags can be how many different characters exist in the field. When the user clicks on the difference tag, the specific different character information can be expanded and viewed.

[0052] This embodiment reversely searches for the corresponding field through different characters, and then generates different difference tags or difference annotations according to the number of different characters, improving the readability of document differences.

[0053] In one implementation manner, the extracting the text to be analyzed from the information to be analyzed includes: If the object to be analyzed is not a text object, an information extraction model corresponding to the type of the object to be analyzed is selected according to the type of the object to be analyzed, and the object to be analyzed is extracted from the information to be analyzed by using the information extraction model. The types of the object to be analyzed include electronic signatures and images; If the object to be analyzed is the text object, the text to be analyzed is extracted from the information to be analyzed.

[0054] For electronic documents of the electronic contract type, in addition to comparing the differences in the text itself, for some relatively important information, such as signatures, electronic signatures, images, etc., difference comparisons may also be required to avoid document tampering or signature and seal forgery, etc. Therefore, before performing difference analysis on the electronic document, it is also necessary to determine the type of the object to be analyzed. If only the text object needs to be analyzed, the text to be analyzed can be directly extracted from the information to be analyzed for analysis.

[0055] If the object to be analyzed also includes information such as signatures, electronic signatures, images, etc., the information such as signatures, electronic signatures, images, etc. in the electronic document and the target document can be respectively extracted through corresponding information extraction models, such as neural network models, and then compared respectively by using the comparison model, and finally the corresponding comparison results are output.

[0056] This embodiment extracts information in different ways for different objects to be analyzed in the electronic document, which can not only perform difference comparison on the text in the electronic document, but also perform corresponding difference analysis on information such as images, signatures, and electronic signatures, making the difference analysis of the electronic document more comprehensive and more applicable.

[0057] Refer to Figure 4 , Figure 4 FIG. A classification module 401 is configured to classify and label a document to be analyzed through a preset classification model, obtain the information to be analyzed in the document to be analyzed, and extract the text to be analyzed from the information to be analyzed.

[0058] A matching module 402 is configured to perform character matching on the text to be analyzed and a target text through a matching algorithm to obtain a matching result, determine the different characters between the text to be analyzed and the target text according to the matching result, and obtain the target position information of the different fields where the different characters are located; the different characters include the first different characters corresponding to the text to be analyzed and the second different characters corresponding to the target text.

[0059] A comment module 403 is configured to generate corresponding field difference comments on the text to be analyzed and the target text respectively according to the target position information and the different fields.

[0060] It can be understood that the electronic document difference analysis system in this embodiment corresponds to the electronic document difference analysis method in the above embodiment. The optional items in the above embodiment are also applicable to this embodiment, so they will not be described repeatedly here.

[0061] The present invention also provides a computer device. Exemplarily, the computer device includes a processor and a memory. The memory stores a computer program, and the processor runs the computer program to enable the computer device to execute the above electronic document difference analysis method or the functions of each module in the above electronic document difference analysis system.

[0062] Among them, the processor may be an integrated circuit chip with signal processing capabilities. The processor may be a general-purpose processor, including at least one of a central processing unit (CPU), a graphics processing unit (GPU), a network processor (NP), a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, and discrete hardware components. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc., which can implement or execute the various methods, steps, and logic block diagrams disclosed in the embodiments of the present invention.

[0063] The memory can be, but is not limited to, Random Access Memory (RAM), Read Only Memory (ROM), Programmable Read-Only Memory (PROM), Erasable Programmable Read-Only Memory (EPROM), Electric Erasable Programmable Read-Only Memory (EEPROM), etc. Among them, the memory is used to store computer programs, and after receiving the execution instruction, the processor can execute the computer program accordingly.

[0064] The present invention also provides a computer storage medium for storing the computer program used in the above computer device. Among them, the computer storage medium can be a readable storage medium, a non-volatile storage medium or a volatile storage medium. For example, the computer storage medium can include, but is not limited to: USB flash drives, mobile hard disks, Read Only Memory (ROM), Random Access Memory (RAM), magnetic disks or optical discs and other various media that can store program codes.

[0065] In several embodiments provided by the present invention, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device embodiments described above are only illustrative. For example, the flowcharts and structure diagrams in the drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to multiple embodiments of the present invention. In this regard, each block in the flowchart or block diagram can represent a module, a program segment or a part of code, and the module, program segment or part of code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in an alternative implementation, the functions marked in the blocks can occur in a different order from that marked in the drawings. For example, two consecutive blocks can actually be executed substantially in parallel, and they can sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the structure diagram and / or flowchart, and the combination of blocks in the structure diagram and / or flowchart, can be implemented by a dedicated hardware-based system for performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.

[0066] In addition, each functional module or unit in various embodiments of the present invention may be integrated together to form an independent part, or each module may exist alone, or two or more modules may be integrated to form an independent part.

[0067] If the above-mentioned function is implemented in the form of a software functional module and sold or used as an independent product, it may be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present invention, in essence, or the part that contributes to the prior art, or a part of this technical solution, may be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a smart phone, a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present invention.

[0068] The above is only the specific implementation manner of the present invention, but the protection scope of the present invention is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present invention can easily think of changes or substitutions, which should all be covered by the protection scope of the present invention.

Claims

1. A method for analyzing differences in electronic documents, characterized in that: include: Classify and annotate the document to be analyzed by using a preset classification model to obtain information to be analyzed in the document to be analyzed, and extract the text to be analyzed from the information to be analyzed; Performing character matching on the text to be analyzed and the target text by a matching algorithm to obtain a matching result, determining the difference characters between the text to be analyzed and the target text according to the matching result, and obtaining target position information of the difference field where the difference characters are located; the difference characters include a first difference character corresponding to the text to be analyzed and a second difference character corresponding to the target text; According to the target position information and the difference fields, corresponding field difference annotations are generated on the text to be analyzed and the target text respectively.

2. The electronic document difference analysis method according to claim 1, characterized in that: The method of classifying and labeling the document to be analyzed by using a preset classification model to obtain information to be analyzed in the document to be analyzed includes: The document to be analyzed is classified by the preset classification model to obtain a classification result, and the document to be analyzed is annotated according to the classification result to obtain annotation information; the classification result includes headers and footers, titles, texts, tables and pictures; The information type corresponding to the annotation information is acquired, and target annotation information belonging to a target information type in the annotation information is used as the information to be analyzed, where the target information type includes the text, the table, and the picture.

3. The electronic document difference analysis method according to claim 2, characterized in that: The step of extracting the text to be analyzed from the information to be analyzed includes: If the information to be analyzed is readable information, the text to be analyzed is directly extracted from the information to be analyzed; the readable information includes readable text and readable table; If the information to be analyzed is unreadable information, a text detection box corresponding to the information to be analyzed is generated; the unreadable information includes unreadable text, unreadable tables and pictures; The text detection box is used to select a text area in the information to be analyzed, and text is extracted from the text area to obtain the text to be analyzed.

4. The electronic document difference analysis method according to claim 3, characterized in that: The method further comprises: Generate corresponding text positioning coordinates according to the number of detection frames and the positions of the detection frames of the text detection frames; The text to be analyzed is segmented by a preset segmentation method to obtain a plurality of fields, and the position information of each field is generated according to the text positioning coordinates.

5. The electronic document difference analysis method according to claim 4, characterized in that: The step of obtaining target position information of the difference field where the difference character is located includes: Match each of the fields in the text to be analyzed with the target field in the target text word by word to obtain the difference characters; The difference field corresponding to each of the difference characters is determined from the fields, and the target position information is generated according to the text positioning coordinates and the difference fields.

6. The electronic document difference analysis method according to claim 5, characterized in that: The generating corresponding difference annotations on the to-be-analyzed text and the target text respectively includes: Get the number of difference characters in each of the difference fields; If the number of the difference characters is greater than a preset number threshold, a difference label corresponding to the difference characters is generated; the difference label includes a set of multiple difference annotations; If the number of the difference characters is less than or equal to the preset number threshold, the difference annotation corresponding to the difference characters is generated.

7. The electronic document difference analysis method according to claim 1, characterized in that: The step of extracting the text to be analyzed from the information to be analyzed includes: If the object to be analyzed is not a text object, a corresponding information extraction model is selected according to the type of the object to be analyzed, and the object to be analyzed is extracted from the information to be analyzed using the information extraction model, where the type of the object to be analyzed includes an electronic signature and an image; If the object to be analyzed is the text object, the text to be analyzed is extracted from the information to be analyzed.

8. An electronic document difference analysis system, characterized in that: include: A classification module is used to classify and annotate the document to be analyzed by using a preset classification model, obtain information to be analyzed in the document to be analyzed, and extract text to be analyzed from the information to be analyzed; A matching module, used for performing character matching between the text to be analyzed and the target text through a matching algorithm to obtain a matching result, determining the difference characters between the text to be analyzed and the target text according to the matching result, and obtaining target position information of the difference field where the difference characters are located; the difference characters include a first difference character corresponding to the text to be analyzed and a second difference character corresponding to the target text; The annotation module is used to generate corresponding field difference annotations on the text to be analyzed and the target text respectively according to the target position information and the difference fields.

9. A computer device, characterized in that: The computer device comprises a processor and a memory, wherein the memory stores a computer program, and the processor is configured to execute the computer program to implement the electronic document difference analysis method according to any one of claims 1 to 7.

10. A computer storage medium, characterized in that: The computer program is stored therein, and when the computer program is executed on a processor, the electronic document difference analysis method according to any one of claims 1 to 7 is implemented.

Citation Information

Patent Citations

  • Document information comparison method and device, equipment and medium

    CN119476251A

  • Multi-dimensional comparison method and device of PDF document and electronic equipment

    CN119514518A

  • Method for comparative analysis of document and apparatus for executing the method

    KR102009901B1