Multi-source electronic document verification method based on large model and OCR technology

By using a multi-source electronic document verification method based on large models and OCR technology, scanning information and identifiers of portable documents are generated, relationship chains are established, and scanning information is output to verify the accuracy of text documents. This solves the problem of conversion errors between text documents and portable documents and improves the accuracy and efficiency of the retrieval system.

CN121835666APending Publication Date: 2026-04-10STATE GRID XINJIANG ELECTRIC POWER CO LTD CHANGJI POWER SUPPLY CO
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
STATE GRID XINJIANG ELECTRIC POWER CO LTD CHANGJI POWER SUPPLY CO
Filing Date
2025-12-02
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In existing information document retrieval systems, conversion errors occur during the conversion between text documents and portable documents, making it difficult for users to verify the accuracy of search results and reducing the efficiency of the retrieval process.

Method used

A multi-source electronic document verification method based on large model and OCR technology is adopted to generate scan information and document identifiers of portable documents, establish relationship chains, match the identifier set and relationship chains with the search information input by the user, output scan information to verify the accuracy of text documents, and sort multiple scan information by calculating the interval distance.

Benefits of technology

It improves the accuracy and efficiency of users' document retrieval, ensures the accuracy of search results, and enhances users' trust in search results and work efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121835666A_ABST
    Figure CN121835666A_ABST
Patent Text Reader

Abstract

The invention discloses a multi-source electronic document verification method based on a large model and an OCR technology, and relates to the technical field of information retrieval, and the method comprises the steps: generating scanning information of a portable document based on the portable document; generating a plurality of document identifiers corresponding to the name of the portable document based on the name of the portable document; generating an identification set according to a plurality of document identifiers corresponding to the portable document name; establishing a relation chain among the portable document, the character document corresponding to the portable document, the scanning information corresponding to the portable document and the identification set corresponding to the portable document; obtaining retrieval information input by a user; based on the retrieval information, matching an identification set containing document identifiers corresponding to the retrieval information; and outputting scanning information corresponding to the identification set based on the identification set and the relation chain. The method has the effect that the user can conveniently check the text document popped up by the retrieval system.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of information retrieval technology, and in particular to a method for verifying multi-source electronic documents based on large models and OCR technology. Background Technology

[0002] In today's digital age, information is exploding, and documents of all kinds exist in massive quantities across different fields and scenarios. Information document retrieval systems have become a key tool for people to accurately and quickly obtain the knowledge they need from vast information databases. They break down information barriers, enabling people from different industries and regions to conveniently share and utilize information, greatly improving work efficiency and the speed of knowledge dissemination, and providing solid information support for decision-making, innovative research, and other fields.

[0003] The types of files entered into the information document retrieval system include portable documents (i.e., PDF documents) and text documents (i.e., Word documents). In the process of entering portable documents into the retrieval system, it is often necessary to translate and convert the portable documents into text documents so that the retrieval system can retrieve the corresponding files based on the keywords and terms entered by the user.

[0004] However, there are often conversion errors between the text documents entered into the system and the original portable documents, such as garbled text, incorrect characters, and recognition errors. These conversion errors are difficult for the system to fully identify and correct. When a user enters keywords into the search system and the retrieved text documents appear, the user may have doubts about the accuracy of the retrieved documents. This makes it difficult for the user to verify the accuracy of the retrieved documents, thus reducing the efficiency of the user's file retrieval work. Summary of the Invention

[0005] To facilitate users in verifying text documents displayed by the retrieval system, this application provides a multi-source electronic document verification method based on large-scale models and OCR technology.

[0006] The multi-source electronic document verification method based on large model and OCR technology provided in this application adopts the following technical solution: A multi-source electronic document verification method based on large-scale models and OCR technology includes the following processing steps: Based on the portable document, generate scan information of the portable document; Based on the name of the portable document, generate multiple document identifiers corresponding to the portable document name; Generate an identifier set based on multiple document identifiers corresponding to the portable document name; Establish a relationship chain between portable documents, their corresponding text documents, their corresponding scan information, and their corresponding identifier sets; Obtain the search information input by the user; Based on the search information, a set of identifiers containing the document identifiers corresponding to the search information is matched; Based on the identifier set and relationship chain, the scanning information corresponding to the identifier set is output.

[0007] The aforementioned technical solution allows for the creation of portable documents containing text, images, formulas, and other information on each page. Generating scanned information of the portable document based on this information means scanning and saving each page of the portable document (e.g., in image format or other readable formats) to preserve its detailed information and prevent issues such as document loss or inaccessible paths.

[0008] Furthermore, based on the name of a portable document, multiple document identifiers corresponding to the portable document name are generated. For example, the name of a portable document may contain multiple keywords or key words, and a keyword or key word may correspond to a document identifier, document identifier A (A may be a word or phrase). Therefore, the name of a portable document may correspond to multiple document identifiers.

[0009] Since the name of a portable document may include multiple document identifiers, for example, if the name of a portable document includes three document identifiers, namely A, B, and C, then these three document identifiers are included in the identifier set (that is, the set of multiple document identifiers corresponding to the name of this portable document) so that users can search by entering keywords.

[0010] Next, establish a relationship chain between the portable document, its corresponding text document, the scanned information, and the identifier set. The text document refers to the text document converted from the portable document; the scanned information refers to the scanned information formed from the specific content of the portable document; and the identifier set refers to the set of multiple document identifiers contained in the portable document name. This relationship chain is generated so that when users verify and retrieve information, they can find the original information based on the one-to-one correspondence, complete information verification, and improve the accuracy of document retrieval.

[0011] After a user enters search information (keywords) into the search system, the system will convert the search information into corresponding document identifiers and match the identifier set containing the document identifiers corresponding to the search information in the system using this method.

[0012] After the system matches the corresponding identifier set based on the search information, it can output the scan information corresponding to this identifier set based on the identifier set and the relationship chain corresponding to the identifier set, and display it on the screen. At this time, the user can click to view the scan information to verify whether the text document that popped up in the previous search is correct, thereby improving the accuracy of the user's file search and the efficiency of the search work.

[0013] In a preferred embodiment, this application can be further configured such that, after matching the identifier set containing the document identifiers corresponding to the search information, the following processing steps are included: If multiple different identifier sets are output, then based on each identifier set and the corresponding portable document name, as well as the multiple document identifiers corresponding to the retrieval information, the position of different document identifiers in the same portable document name is obtained; Based on the positions of different document identifiers within the same portable document name, obtain the spacing between adjacent document identifiers within the portable document name; Calculate the total interval distance based on multiple interval distances corresponding to the portable document; Based on the total interval distance corresponding to multiple different identifier sets, the portable documents corresponding to different identifier sets are sorted to generate an audit sequence list.

[0014] If the system matches multiple corresponding identifier sets based on the search information entered by the user, it means that multiple scan information will be displayed on the screen. At this time, it is necessary to sort the multiple scan information.

[0015] Through the above technical solution, if the system matches multiple corresponding identifier sets, it obtains the position of different document identifiers in the portable document name corresponding to each identifier set. For example, after the search information entered by a user is converted into document identifiers, it includes three document identifiers, namely A, B, and C. The positions of these three document identifiers in the corresponding portable document name are the first character position, the third character position, and the sixth character position, respectively.

[0016] Generally speaking, the distance between adjacent characters in each line of a document is the same. Therefore, the spacing between A and B is fixed, which means the spacing between A and B is one unit. Similarly, the spacing between B and C is two units. Thus, the total spacing of this portable document name is three units.

[0017] Similarly, the total interval distance corresponding to different identifier sets is calculated, and the scan information is sorted according to the length of the total interval distance to generate an audit sequence list. The scan information with the shortest total interval distance is displayed first, and the scan information with the longest total interval distance is displayed last. This can clarify the arrangement order of multiple scan information and arrange the scan information that users may want first, thereby improving the efficiency of user retrieval.

[0018] In a preferred embodiment, this application can be further configured such that obtaining the interval distance between adjacent document identifiers in the portable document name includes the following processing steps: If the number of lines in a portable document name is greater than one, then the portable document names in different lines are concatenated from top to bottom to form a linear document name. Get the distance between adjacent document identifiers in a linear document name.

[0019] In obtaining the distance between adjacent document identifiers in a portable document name, if the portable document name contains multiple lines of text, it may lead to inaccurate calculation of the distance between adjacent document identifiers. For example, a portable document name may contain two lines of text, and the document identifiers may be located in the same column of text in both lines of text. For example, the actual distance between the positions of D and E is 16 units, but the distance between them on the plane is 1 unit. This greatly reduces the accuracy of the distance measurement.

[0020] With the above technical solution, if the number of lines of the portable document name is greater than one, and the document identifier is distributed in two lines of name text, then the portable document names in different lines are concatenated from top to bottom into a linear document name, which is displayed in one line of text, effectively improving the accuracy of the interval distance measurement.

[0021] In a preferred embodiment, this application can be further configured such that, based on the total interval distance corresponding to multiple different identifier sets, the portable documents corresponding to different identifier sets are sorted to generate an audit sequence list, including the following processing steps: If there are portable documents with the same total interval distance, then obtain the interval distance between the first two document identifiers in the different portable document names; A preliminary review sequence list is generated based on the interval between the first two document identifiers in different portable document names.

[0022] Using the above technical solution, if there are portable documents with the same total interval distance, it is difficult to sort them according to the total interval distance. In this case, the interval distance between the first two document identifiers in the portable document name is obtained.

[0023] For example, if the three document identifiers A, B, and C are located at the first character position, the third character position, and the sixth character position in one of the corresponding portable document names, then the distance between A and B is obtained, which is one unit length.

[0024] If the three document identifiers A, B, and C are located in the first, fourth, and sixth character positions respectively in the name of another corresponding portable document, then the interval between A and B, which is two units, is obtained. In this case, the scan information corresponding to the portable document with the shorter interval between the first two document identifiers is displayed first in the primary review sequence list.

[0025] In a preferred embodiment, this application can be further configured such that, after generating the primary audit sequence list, the following processing steps are included: If there are portable documents with the same interval between the first two document identifiers, then obtain the interval between the nth document identifier and the (n+1)th document identifier in turn, where n is a positive integer greater than one; An n-level review sequence list is generated based on the interval between the nth document identifier and the (n+1)th document identifier.

[0026] Using the above technical solution, if there are portable documents with the same interval distance between the first two document identifiers, then continue to obtain the interval distance between adjacent document identifiers in turn, that is, obtain the interval distance between the nth document identifier and the (n+1)th document identifier in turn, where n is a positive integer greater than one.

[0027] For example, first obtain the interval between the second document identifier and the third document identifier, and compare the interval between the second document identifier and the third document identifier in different portable document names, and prioritize displaying the scan information corresponding to portable documents with smaller interval between the second document identifier and the third document identifier.

[0028] If the interval between the second document identifier and the third document identifier is also the same, then continue to obtain the distance between the third document identifier and the fourth document identifier, compare them, and prioritize them.

[0029] Similarly, if the distance between the third and fourth document identifiers is still equal, continue comparing the identifier spacing until the distance between the last two document identifiers is compared. After the distance between the last two document identifiers is compared, the resulting audit sequence list is the final level audit sequence list.

[0030] In a preferred example, this application can be further configured to include the following processing steps: If there are portable documents in the last-level review sequence list where the interval between the last two document identifiers is the same, then calculate the sum of the intervals between adjacent document identifiers among the first three document identifiers; A two-level audit sequence list is generated based on the sum of the interval distances.

[0031] Using the above technical solution, if there are portable documents with the same spacing between the last two document identifiers, then the sum of the spacing between adjacent document identifiers in the first three document identifiers is calculated. For example, if the spacing between the first and second document identifiers in the first three document identifiers of a portable document name is one unit, and the spacing between the second and third document identifiers is two units, then the sum of the spacing between adjacent document identifiers in the first three document identifiers of this portable document name is three units.

[0032] If the sum of the intervals between adjacent document identifiers in the first three document identifiers of another portable document name is four units, then the scan information of portable documents with a sum of intervals of three units will be displayed first.

[0033] In a preferred embodiment, this application can be further configured such that, after matching the identifier set containing the document identifiers corresponding to the search information, the following processing steps are included: If multiple different identifier sets are output, then based on each identifier set and the corresponding portable document name, as well as the multiple document identifiers corresponding to the retrieval information, the position of different document identifiers in the same portable document name is obtained; Based on the positions of different document identifiers within the same portable document name, obtain the number of spacing characters between adjacent document identifiers in the portable document name; Calculate the total number of characters based on the number of multiple spacer characters corresponding to the portable document; Based on the total number of characters corresponding to multiple different identifier sets, the portable documents corresponding to different identifier sets are sorted to generate an audit sequence list.

[0034] Through the above technical solution, if the system matches multiple corresponding identifier sets, it obtains the position of different document identifiers in the portable document name corresponding to each identifier set. For example, after the search information entered by a user is converted into document identifiers, it includes three document identifiers, namely A, B, and C. The positions of these three document identifiers in the corresponding portable document name are the first character position, the third character position, and the sixth character position, respectively.

[0035] The number of characters separating A and B is fixed. Compared to comparing the spacing between adjacent identifiers, there is no need to organize multi-line name text, effectively improving the accuracy of the order. Moreover, in order to maintain the consistency and neatness of the text, the character spacing between different lines is fine-tuned by the document control system, while the number of characters remains constant, further improving the accuracy of the order.

[0036] In summary, this application includes the following beneficial technical effects: 1. If a user questions the retrieved text document, they can click to view the scan information of the corresponding portable document to verify whether the previously retrieved text document is correct, thereby improving the accuracy of the user's file retrieval and the efficiency of the retrieval process. 2. If the system can match multiple identifier sets based on the user's input search information, that is, match the scan information of multiple portable documents, then calculate the sum of the interval distances between adjacent document identifiers in each portable document name, and sort the scan information of multiple portable documents in this way to improve the accuracy of the user's file retrieval and improve the search efficiency. Attached Figure Description

[0037] Figure 1 This is a flowchart illustrating the document verification method.

[0038] Figure 2 This is a flowchart illustrating the process of generating the audit sequence list.

[0039] Figure 3 This is a flowchart illustrating the process of obtaining the distance between adjacent document identifiers in a linear document name.

[0040] Figure 4 This is a flowchart illustrating the process of generating a preliminary audit sequence list.

[0041] Figure 5 This is a flowchart illustrating the process of comparing document identifier intervals at each level.

[0042] Figure 6 This is a flowchart illustrating the process of generating a secondary audit sequence list.

[0043] Figure 7 This is a flowchart illustrating another process for generating an audit sequence list.

[0044] Figure 8 This is a schematic diagram illustrating the distribution structure of document identifiers in portable document names.

[0045] Figure 9 This is a schematic diagram of the different line distribution structure of document identifiers in portable document names.

[0046] Figure 10 This is a schematic diagram of the structure in an embodiment of this application. Detailed Implementation

[0047] The following is in conjunction with the appendix Figure 1 - Appendix Figure 10 This application will be described in further detail.

[0048] This application discloses a multi-source electronic document verification method based on large model and OCR technology. This method can be applied to information retrieval systems, and the executing entity of this method can be a server terminal in the information retrieval system.

[0049] See attached document Figure 1 As shown, the multi-source electronic document verification method based on large model and OCR technology includes the following processing steps: S101. Based on the portable document, generate the scan information of the portable document.

[0050] In implementation, the content of a portable document may consist of multiple pages, each containing text, images, formulas, and other information. Generating scanned information of the portable document based on this information means scanning and saving each page of the portable document (e.g., in image format or other readable formats) to preserve the specific information of the portable document and prevent issues such as document loss or inaccessible paths.

[0051] S102. Based on the name of the portable document, generate multiple document identifiers corresponding to the portable document name.

[0052] In implementation, multiple document identifiers are generated based on the name of the portable document. For example, the name of a portable document may contain multiple keywords or key words, and a keyword or key word may correspond to a document identifier, document identifier A (A may be a word or phrase). Therefore, the name of a portable document may correspond to multiple document identifiers.

[0053] S103. Generate an identifier set based on multiple document identifiers corresponding to the portable document name.

[0054] In implementation, since the name of a portable document may include multiple document identifiers, for example, if the name of a portable document includes three document identifiers, namely A, B, and C, then these three document identifiers are included in the identifier set (that is, the set of multiple document identifiers corresponding to this portable document name) so that users can search by entering keywords.

[0055] S104. Establish the relationship chain between the portable document, the text document corresponding to the portable document, the scan information corresponding to the portable document, and the identifier set corresponding to the portable document.

[0056] In implementation, a relationship chain is established between portable documents, their corresponding text documents, their corresponding scanned information, and their corresponding identifier sets. The text document corresponding to the portable document refers to the text document converted from the portable document; the scanned information refers to the scanned information formed from the specific content of the portable document; and the identifier set refers to the set of multiple document identifiers contained in the portable document name. This relationship chain is generated so that when users verify and retrieve information, they can find the original information based on the one-to-one correspondence, complete information verification, and improve the accuracy of document retrieval.

[0057] S105. Obtain the search information input by the user.

[0058] S106. Based on the search information, match the identifier set containing the document identifier corresponding to the search information.

[0059] In practice, after a user enters search information (keywords) into the search system, the system will convert the search information into corresponding document identifiers and match the identifier set containing the document identifiers corresponding to the search information in the system using this method.

[0060] S107. Based on the identifier set and relationship chain, output the scanning information corresponding to the identifier set.

[0061] In practice, after the system matches the corresponding identifier set based on the search information, it can output the scan information corresponding to this identifier set based on the identifier set and the relationship chain corresponding to the identifier set, and display it on the screen. At this time, the user can click to view the scan information to verify whether the text document that popped up in the previous search is correct, thereby improving the accuracy of the user's file search and the efficiency of the search work.

[0062] See attached document Figure 2 As shown, step S106, after matching the identifier set containing the document identifier corresponding to the search information, may include the following processing steps: S201. If multiple different identifier sets are output, then based on each identifier set and the corresponding portable document name, as well as the multiple document identifiers corresponding to the retrieval information, obtain the position of different document identifiers in the same portable document name.

[0063] In implementation, based on the search information entered by the user, if the system matches multiple corresponding identifier sets, it means that multiple scan information will be displayed on the screen, and at this time, it is necessary to sort the multiple scan information. If the system matches multiple corresponding identifier sets, it obtains the position of different document identifiers in the portable document name corresponding to each identifier set.

[0064] For example, refer to the appendix Figure 8 As shown, for example, after a user's search information is converted into document identifiers, it includes three document identifiers, namely A, B, and C. The positions of these three document identifiers in the corresponding portable document name are the first character position, the third character position, and the sixth character position, respectively.

[0065] S202. Based on the positions of different document identifiers in the same portable document name, obtain the interval distance between adjacent document identifiers in the portable document name.

[0066] S203. Calculate the total interval distance based on the multiple interval distances corresponding to the portable document.

[0067] In implementation, refer to the appendix. Figure 8 As shown, the distance between adjacent characters in each line of a document is usually the same. Therefore, the spacing between A and B is fixed, which is one unit. Similarly, the spacing between B and C is two units. Thus, the total spacing of this portable document name is three units.

[0068] S204. Based on the total interval distance corresponding to multiple different identifier sets, sort the portable documents corresponding to different identifier sets and generate an audit sequence list.

[0069] In implementation, the total interval distance corresponding to different identifier sets is calculated, and the scan information is sorted according to the length of the total interval distance to generate an audit sequence list. The scan information with the shortest total interval distance is displayed first, and the scan information with the longest total interval distance is displayed last. This clarifies the arrangement order of multiple scan information and arranges the scan information that users may want first, thereby improving the efficiency of user retrieval.

[0070] See attached document Figure 3 As shown, in step S202, obtaining the interval distance between adjacent document identifiers in the portable document name includes the following processing steps: S301. If the number of lines in a portable document name is greater than one, then the portable document names in different lines are concatenated from top to bottom to form a linear document name.

[0071] In practice, when obtaining the distance between adjacent document identifiers in a portable document name, if the portable document name contains multiple lines of text, the calculation of the distance between adjacent document identifiers may be inaccurate. (See attached document.) Figure 9 As shown, for example, a portable document name consists of two lines of text, and the document identifier is distributed in the same column of text in the two lines of name text. For example, the actual distance between the positions of D and E is 16 units, but the distance between them on the plane is 1 unit. This greatly reduces the accuracy of the distance measurement.

[0072] If the number of lines in a portable document name is greater than one, and the document identifier is distributed across two lines of name text, then the portable document names on different lines are concatenated from top to bottom to form a linear document name. A linear document name is a portable document name arranged in a straight line.

[0073] S302. Obtain the distance between adjacent document identifiers in the linear document name.

[0074] In practice, displaying the document name of the portable document in a single line of text avoids line breaks and effectively improves the accuracy of interval distance calculation.

[0075] See attached document Figure 4 As shown, step S204, which sorts the portable documents corresponding to different identifier sets based on the total interval distance of multiple different identifier sets and generates an audit sequence list, may include the following processing steps: S401. If there are portable documents with the same total interval distance, obtain the interval distance between the first two document identifiers in the different portable document names.

[0076] In implementation, if portable documents with the same total interval distance exist, it is difficult to sort them based on the total interval distance. In this case, the interval distance between the first two document identifiers in the portable document name is obtained. See Appendix. Figure 10 As shown, for example, if the three document identifiers A, B, and C are located at the first character position, the third character position, and the sixth character position in one of the corresponding portable document names, then the distance between A and B, which is one unit length, can be obtained.

[0077] If the three document identifiers A, B, and C are located in the first, fourth, and sixth character positions respectively in the other corresponding portable document name, then the interval between A and B is obtained, which is two units.

[0078] S402. Generate a primary review sequence list based on the interval between the first two document identifiers in different portable document names.

[0079] In practice, in the primary review sequence list, scan information corresponding to portable documents with a short interval between the first two document identifiers is displayed first.

[0080] See attached document Figure 5 As shown, after generating the primary review sequence list in step S402, the following processing steps may be included: S501. If there are portable documents with the same interval between the first two document identifiers, then obtain the interval between the nth document identifier and the (n+1)th document identifier in sequence, where n is a positive integer greater than one.

[0081] S502. Generate an n-level audit sequence list based on the interval between the nth document identifier and the (n+1)th document identifier.

[0082] In implementation, for example, the interval between the second document identifier and the third document identifier is first obtained, and the interval between the second document identifier and the third document identifier in different portable document names is compared. The scan information corresponding to the portable document with the smaller interval between the second document identifier and the third document identifier is displayed first.

[0083] If the interval between the second document identifier and the third document identifier is also the same, then continue to obtain the distance between the third document identifier and the fourth document identifier, compare them, and prioritize them.

[0084] Similarly, if the distance between the third and fourth document identifiers is still equal, continue comparing the identifier spacing until the distance between the last two document identifiers is compared. After the distance between the last two document identifiers is compared, the resulting audit sequence list is the final level audit sequence list.

[0085] See attached document Figure 6 As shown, the processing steps include the following: S601. If there are portable documents in the last-level review sequence list where the interval distance between the last two document identifiers is the same, then calculate the sum of the interval distances between adjacent document identifiers among the first three document identifiers.

[0086] S602. Generate a two-level audit sequence list based on the sum of the interval distances.

[0087] In implementation, the final-level review sequence list is the review sequence list arranged based on the interval distance between the last two document identifiers. If there are portable documents with the same interval distance between the last two document identifiers, then the sum of the interval distances between adjacent document identifiers among the first three document identifiers is calculated.

[0088] For example, if the first and second document identifiers of a portable document name are spaced one unit apart, and the second and third document identifiers are spaced two units apart, then the sum of the intervals between adjacent document identifiers in the first three document identifiers of this portable document name is three units apart.

[0089] If the sum of the intervals between adjacent document identifiers in the first three document identifiers of another portable document name is four units, then the scan information of portable documents with a sum of intervals of three units will be displayed first.

[0090] See attached document Figure 7 As shown, step S106, after matching the identifier set containing the document identifier corresponding to the search information, may include the following processing steps: S701. If multiple different identifier sets are output, then based on each identifier set and the corresponding portable document name, as well as the multiple document identifiers corresponding to the retrieval information, obtain the position of the different document identifiers in the same portable document name.

[0091] S702. Based on the positions of different document identifiers in the same portable document name, obtain the number of spacing characters between adjacent document identifiers in the portable document name.

[0092] S703. Calculate the total number of characters based on the number of multiple spacer characters corresponding to the portable document.

[0093] S704. Based on the total number of characters corresponding to multiple different identifier sets, sort the portable documents corresponding to different identifier sets and generate an audit sequence list.

[0094] In implementation, the difference between step S7 and step S2 is that if the system matches multiple corresponding identifier sets, it obtains the position of different document identifiers within the portable document name corresponding to each identifier set. (See Appendix) Figure 8 As shown, for example, after a user's search information is converted into document identifiers, it includes three document identifiers, namely A, B, and C. The positions of these three document identifiers in the corresponding portable document name are the first character position, the third character position, and the sixth character position, respectively.

[0095] The number of characters separating A and B is fixed. Compared to comparing the spacing between adjacent identifiers, there is no need to organize multi-line name text, effectively improving the accuracy of the order. Moreover, in order to maintain the consistency and neatness of the text, the character spacing between different lines is fine-tuned by the document control system, while the number of characters remains constant, further improving the accuracy of the order.

[0096] The embodiments described in this specific implementation are all preferred embodiments of this application and are not intended to limit the scope of protection of this application. Therefore, all equivalent changes made in accordance with the structure, shape and principle of this application should be covered within the scope of protection of this application.

Claims

1. A multi-source electronic document verification method based on large-scale model and OCR technology, characterized in that, include: Based on the portable document, generate scan information of the portable document; Based on the name of the portable document, generate multiple document identifiers corresponding to the portable document name; Generate an identifier set based on multiple document identifiers corresponding to the portable document name; Establish a relationship chain between portable documents, their corresponding text documents, their corresponding scan information, and their corresponding identifier sets; Obtain the search information input by the user; Based on the search information, a set of identifiers containing the document identifiers corresponding to the search information is matched; Based on the identifier set and relationship chain, the scanning information corresponding to the identifier set is output.

2. The multi-source electronic document verification method based on large model and OCR technology according to claim 1, characterized in that, The matching process includes, after defining the identifier set of document identifiers corresponding to the retrieved information, the following: If multiple different identifier sets are output, then based on each identifier set and the corresponding portable document name, as well as the multiple document identifiers corresponding to the retrieval information, the position of different document identifiers in the same portable document name is obtained; Based on the positions of different document identifiers within the same portable document name, obtain the spacing between adjacent document identifiers within the portable document name; Calculate the total interval distance based on multiple interval distances corresponding to the portable document; Based on the total interval distance corresponding to multiple different identifier sets, the portable documents corresponding to different identifier sets are sorted to generate an audit sequence list.

3. The multi-source electronic document verification method based on large model and OCR technology according to claim 2, characterized in that, The step of obtaining the interval distance between adjacent document identifiers in the portable document name includes: If the number of lines in a portable document name is greater than one, then the portable document names in different lines are concatenated from top to bottom to form a linear document name. Get the distance between adjacent document identifiers in a linear document name.

4. The multi-source electronic document verification method based on large model and OCR technology according to claim 2, characterized in that, The step of sorting portable documents corresponding to different identifier sets based on the total interval distance of multiple different identifier sets to generate an audit sequence list includes: If there are portable documents with the same total interval distance, then obtain the interval distance between the first two document identifiers in the different portable document names; A preliminary review sequence list is generated based on the interval between the first two document identifiers in different portable document names.

5. The multi-source electronic document verification method based on large model and OCR technology according to claim 4, characterized in that, After generating the primary audit sequence list, the following steps are included: If there are portable documents with the same interval between the first two document identifiers, then obtain the interval between the nth document identifier and the (n+1)th document identifier in turn, where n is a positive integer greater than one; An n-level review sequence list is generated based on the interval between the nth document identifier and the (n+1)th document identifier.

6. The multi-source electronic document verification method based on large model and OCR technology according to claim 5, characterized in that, include: If there are portable documents in the last-level review sequence list where the interval between the last two document identifiers is the same, then calculate the sum of the intervals between adjacent document identifiers among the first three document identifiers; A two-level audit sequence list is generated based on the sum of the interval distances.

7. The multi-source electronic document verification method based on large model and OCR technology according to claim 1, characterized in that, The matching process includes, after defining the identifier set of document identifiers corresponding to the retrieved information, the following: If multiple different identifier sets are output, then based on each identifier set and the corresponding portable document name, as well as the multiple document identifiers corresponding to the retrieval information, the position of different document identifiers in the same portable document name is obtained; Based on the positions of different document identifiers within the same portable document name, obtain the number of spacing characters between adjacent document identifiers in the portable document name; Calculate the total number of characters based on the number of multiple spacer characters corresponding to the portable document; Based on the total number of characters corresponding to multiple different identifier sets, the portable documents corresponding to different identifier sets are sorted to generate an audit sequence list.