Character structured processing method based on character recognition large model and related device

By constructing a scene graph of logical relationships based on a large text recognition model and pre-setting a text recognition model, the problem of inaccurate keyword generation in document processing in the power industry was solved, and the accuracy of structured documents was improved.

CN122024262APending Publication Date: 2026-05-12WENSHAN POWER SUPPLY BUREAU YUNNAN GRID
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
WENSHAN POWER SUPPLY BUREAU YUNNAN GRID
Filing Date
2025-12-18
Publication Date
2026-05-12

AI Technical Summary

Technical Problem

In existing technologies for document processing in the power industry, keyword generation is inaccurate and redundant, leading to a decrease in the accuracy of structured processing.

Method used

A method based on a large text recognition model is adopted to perform structured processing of power documents by constructing a scene graph with logical relationships and a preset text recognition model.

Benefits of technology

The accuracy of structured documents has been improved. By extracting text regions, obtaining key information, and constructing scene graphs of logical relationships, the structured processing effect of power document data has been enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122024262A_ABST
    Figure CN122024262A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of data processing, and provides a character structured processing method based on a character recognition large model and a related device, and the method comprises the steps: obtaining to-be-processed power document data; performing document data extraction on the to-be-processed power document data according to a text region to obtain k region document data; key information of the document data of the k areas is obtained, and k document key information is obtained; according to the k pieces of document key information and the corresponding regional document data, performing logic association relationship scene graph construction to obtain a target logic association relationship scene graph; and according to the target logic association relationship scene graph, performing structured processing on the to-be-processed power document data by adopting a preset character recognition model to obtain a structured document, and performing structured processing based on the logic association relationship scene graph and the preset character recognition model to obtain the structured document. And the accuracy of determining the structured document is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of data processing technology, specifically to a text structuring processing method and related apparatus based on a large character recognition model. Background Technology

[0002] The power industry currently has a wide variety of document and invoice types. When processing this type of data, the common practice is to manually define keyword rules for feature extraction, and then perform structured processing based on these extracted features. However, this method faces limitations in keyword extraction, including inaccurate keyword generation and redundant keywords, which reduces the accuracy of the structured data processing. Summary of the Invention

[0003] This application provides a text structuring method and related apparatus based on a large character recognition model. It can construct a logical relationship scene graph of the power document data to be processed, and perform structuring processing based on the logical relationship scene graph and a preset character recognition model to obtain a structured document, thereby improving the accuracy of determining the structured document.

[0004] The first aspect of this application provides a text structuring processing method based on a large character recognition model, the method comprising: Retrieve power document data to be processed; The power document data to be processed is extracted according to text regions to obtain k regions of document data; Obtain key information from k region document data to obtain key information for k documents; Based on the key information of k documents and the corresponding regional document data, a logical relationship scenario diagram is constructed to obtain the target logical relationship scenario diagram. Based on the target logical relationship scenario diagram, a preset text recognition model is used to perform structured processing on the power document data to be processed, resulting in a structured document.

[0005] In one possible implementation, the step of extracting document data from the power document data to be processed according to text regions to obtain k regions of document data includes: Obtain the document attribute information of the power document data to be processed; The text area layout information of the power document data to be processed is determined based on the document attribute information. Based on the text region layout information, determine the k text region information corresponding to the power document data to be processed; Based on the information from the k text regions, document data is extracted from the power document data to be processed, resulting in k regions of document data.

[0006] In one possible implementation, the step of extracting document data from the power document data to be processed based on the information of the k text regions to obtain the document data of the k regions includes: Based on the information of k text regions, the power document data to be processed is labeled to obtain labeled power document data; The labeled power document data is extracted according to the labeled areas to obtain k area document data.

[0007] In one possible implementation, the step of constructing a logical relationship scene graph based on key information from k documents and corresponding regional document data to obtain the target logical relationship scene graph includes: Perform correlation analysis on the key information of k documents to obtain the target correlation information between the key information of every two documents; Based on the target association information, construct a logical association description between the key information of the two corresponding documents; Based on the logical relationship description information between key information of every two documents and key information of k documents, a logical relationship scenario diagram is constructed to obtain the target logical relationship scenario diagram.

[0008] In one possible implementation, the step of constructing a logical relationship scene graph based on key information from k documents and corresponding regional document data to obtain the target logical relationship scene graph includes: Feature extraction is performed on key information from k documents to obtain k key feature information; Randomly extract key feature information of the target from k key feature information; Extract the correlation scores between k key feature information and target key feature information to obtain k correlation score values; The scene relationship distance is determined based on k relevance scores, resulting in k scene relationship distance values; Extract k key feature information and describe the correlation between them and the target key feature information to obtain k correlation description information; A logical relationship scene graph is constructed based on k relational descriptions, k scene relationship distance values, k key document information, and corresponding regional document data to obtain the target logical relationship scene graph.

[0009] A second aspect of this application provides a text structuring processing apparatus based on a large character recognition model, the apparatus comprising: The first acquisition unit is used to acquire power document data to be processed. The extraction unit is used to extract document data from the power document data to be processed according to text regions, and obtain document data in k regions; The second acquisition unit is used to acquire key information from k region document data to obtain key information for k documents; The construction unit is used to construct a logical relationship scene graph based on key information of k documents and corresponding regional document data, thereby obtaining the target logical relationship scene graph. The processing unit is used to perform structured processing on the power document data to be processed using a preset text recognition model based on the target logical relationship scenario diagram, so as to obtain a structured document.

[0010] In one possible implementation, the extraction unit is specifically used for: Obtain the document attribute information of the power document data to be processed; The text area layout information of the power document data to be processed is determined based on the document attribute information. Based on the text region layout information, determine the k text region information corresponding to the power document data to be processed; Based on the information from the k text regions, document data is extracted from the power document data to be processed, resulting in k regions of document data.

[0011] In one possible implementation, in the step of extracting document data from the power document data to be processed based on the information of the k text regions to obtain document data of the k regions, the extraction unit is specifically used for: Based on the information of k text regions, the power document data to be processed is labeled to obtain labeled power document data; The labeled power document data is extracted according to the labeled areas to obtain k area document data.

[0012] In one possible implementation, the building unit is specifically used for: Perform correlation analysis on the key information of k documents to obtain the target correlation information between the key information of every two documents; Based on the target association information, construct a logical association description between the key information of the two corresponding documents; Based on the logical relationship description information between key information of every two documents and key information of k documents, a logical relationship scenario diagram is constructed to obtain the target logical relationship scenario diagram.

[0013] In one possible implementation, the building unit is specifically used for: Feature extraction is performed on key information from k documents to obtain k key feature information; Randomly extract key feature information of the target from k key feature information; Extract the correlation scores between k key feature information and target key feature information to obtain k correlation score values; The scene relationship distance is determined based on k relevance scores, resulting in k scene relationship distance values; Extract k key feature information and describe the correlation between them and the target key feature information to obtain k correlation description information; A logical relationship scene graph is constructed based on k relational descriptions, k scene relationship distance values, k key document information, and corresponding regional document data to obtain the target logical relationship scene graph.

[0014] A third aspect of this application provides a terminal including a processor, an input device, an output device, and a memory, wherein the processor, input device, output device, and memory are interconnected, wherein the memory is used to store a computer program, the computer program including program instructions, and the processor is configured to invoke the program instructions to execute the step instructions as described in the first aspect of this application.

[0015] A fourth aspect of this application provides a computer-readable storage medium storing a computer program for electronic data interchange, wherein the computer program causes a computer to perform some or all of the steps described in the first aspect of this application.

[0016] A fifth aspect of this application provides a computer program product, wherein the computer program product includes a non-transitory computer-readable storage medium storing a computer program operable to cause a computer to perform some or all of the steps described in the first aspect of this application. The computer program product may be a software installation package.

[0017] Implementing the embodiments of this application has at least the following beneficial effects: By acquiring power document data to be processed, document data is extracted from the data according to text regions to obtain k region document data. Key information of the k region document data is obtained to obtain k document key information. Based on the k document key information and the corresponding region document data, a logical relationship scene graph is constructed to obtain a target logical relationship scene graph. Based on the target logical relationship scene graph, a preset text recognition model is used to perform structured processing on the power document data to be processed to obtain a structured document. Therefore, it is possible to construct a logical relationship scene graph on the power document data to be processed, and perform structured processing based on the logical relationship scene graph and the preset text recognition model to obtain a structured document, thereby improving the accuracy of determining the structured document. Attached Figure Description

[0018] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0019] Figure 1 This application provides a flowchart illustrating a text structuring method based on a large character recognition model. Figure 2 This is a schematic diagram of the structure of a terminal provided in an embodiment of this application; Figure 3 This application provides a schematic diagram of a text structuring processing device based on a large character recognition model. Detailed Implementation

[0020] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0021] The terms "first," "second," etc., in the specification, claims, and accompanying drawings of this application are used to distinguish different objects, not to describe a specific order. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or apparatus that includes a series of steps or units is not limited to the listed steps or units, but may optionally include steps or units not listed, or may optionally include other steps or units inherent to these processes, methods, products, or apparatuses.

[0022] In this application, the reference to "embodiment" means that a specific feature, structure, or characteristic described in connection with an embodiment may be included in at least one embodiment of this application. The appearance of this phrase in various places throughout the specification does not necessarily refer to the same embodiment, nor is it a mutually exclusive, independent, or alternative embodiment. It will be explicitly and implicitly understood by those skilled in the art that the embodiments described in this application can be combined with other embodiments.

[0023] To better understand the text structuring method based on a large character recognition model provided in this application, a brief introduction to existing text structuring methods is given below. In existing methods, feature extraction is typically performed by manually setting keyword rules, followed by structuring based on the extracted features. However, this method faces limitations in keyword extraction, including inaccurate keyword generation and redundant keywords.

[0024] To address the aforementioned technical problems, this application provides a text structuring method based on a large character recognition model. This method can construct a logical relationship scene graph of the power document data to be processed, and perform structuring processing based on the logical relationship scene graph and a preset character recognition model to obtain a structured document, thereby improving the accuracy of determining the structured document.

[0025] Please see Figure 1 , Figure 1 This application provides a flowchart illustrating a text structuring method based on a large character recognition model, as illustrated in this embodiment. Figure 1 As shown, this method is applied to a business product generation system, and the method includes: 101. Obtain the power document data to be processed.

[0026] The power document data to be processed can include power industry invoices and related documents, such as power invoices and power information processing forms. This data can be obtained through user input or retrieved from a database. The power document data can be understood as relevant invoices or documents, such as power document diagrams generated according to a specific format.

[0027] 102. Extract document data from the power document data to be processed according to the text region to obtain k region document data.

[0028] This process involves extracting document attribute information from the power document data to be processed, determining the text region layout information based on this attribute information, and finally extracting k region document data based on the text region layout information. Specifically, document attribute information can be understood as the document type information of the power document data to be processed. Different types of information have their corresponding text region layout information, which can then be extracted. The text region layout information can correspond to region boxes or region box lines. Different document regions have their corresponding region box lines, thus the data corresponding to the region box lines can be identified as region document data.

[0029] 103. Obtain key information from k region document data to obtain key information for k documents.

[0030] Key information can be extracted using general keyword extraction methods to obtain k key information points for the document. For example, a general keyword extraction model can be used to extract the key information for the document.

[0031] 104. Construct a logical relationship scenario diagram based on the key information of k documents and the corresponding regional document data to obtain the target logical relationship scenario diagram.

[0032] Several different construction methods can be used, and one method can be selected for specific construction. Specifically, one method for constructing a logical relationship scenario graph based on key information from k documents and corresponding regional document data can be as follows: Correlation analysis can be performed on the key information from the k documents to obtain the target relationship information between each pair of key information from the documents. Then, a logical relationship scenario graph can be constructed based on this target relationship information to obtain the target logical relationship scenario graph.

[0033] Another method to construct a logical relationship scene graph based on key information from k documents and corresponding regional document data, and obtain the target logical relationship scene graph, is as follows: Feature extraction can be performed on key information from k documents to obtain k key feature information. Then, target key feature information is randomly selected from the k key feature information, and the corresponding relevance score is performed. Finally, a logical relationship scene graph is constructed based on the relevance score to obtain the target logical relationship scene graph.

[0034] This allows for the accurate acquisition of the target logical relationship scenario diagram, which can be used for subsequent document structuring processing to obtain a structured document.

[0035] 105. Based on the target logical relationship scenario diagram, a preset text recognition model is used to perform structured processing on the power document data to be processed, resulting in a structured document.

[0036] The preset text recognition model is a pre-trained model used to perform structured processing of corresponding power document data based on a logical relationship scene graph. For example, it could be a convolutional neural network model.

[0037] In this example, by acquiring the power document data to be processed, document data is extracted from the power document data according to text regions to obtain k region document data. Key information of the k region document data is obtained to obtain k document key information. Based on the k document key information and the corresponding region document data, a logical relationship scene graph is constructed to obtain a target logical relationship scene graph. Based on the target logical relationship scene graph, a preset text recognition model is used to perform structured processing on the power document data to be processed to obtain a structured document. Therefore, it is possible to construct a logical relationship scene graph on the power document data to be processed, and perform structured processing based on the logical relationship scene graph and the preset text recognition model to obtain a structured document, thereby improving the accuracy of determining the structured document.

[0038] In one possible implementation, the step of extracting document data from the power document data to be processed according to text regions to obtain k regions of document data includes: A1. Obtain the document attribute information of the power document data to be processed; A2. Determine the text area layout information of the power document data to be processed based on the document attribute information; A3. Determine the k text region information corresponding to the power document data to be processed based on the text region layout information; A4. Based on the information of the k text regions, extract the document data of the power document data to be processed to obtain the document data of the k regions.

[0039] The document attribute information of the power document data to be processed can be understood as the document type information. Different types of information have their corresponding text area layout information, which can then be extracted. Specifically, this correspondence can be stored and represented through a mapping table.

[0040] This allows the text area layout information of the power document data to be processed to be determined by querying the mapping table based on the document attribute information, thereby quickly determining the text area layout information and improving processing efficiency.

[0041] The text region layout information includes the positional relationship of each text region, or the coordinate information of each text region. Therefore, k text region information can be determined based on the text layout information. When extracting document data based on text region information, regions can be labeled based on the text region information to obtain labeled power document data. Finally, data extraction can be performed according to the labeled regions of the labeled power document data to obtain k region document data.

[0042] Therefore, when determining the regional document data using the above method, the text region layout information can be determined based on the document attribute information. Then, k text region information can be determined using the text region layout information. Finally, data extraction can be performed using the k text region information to obtain the k regional document data, thus improving the efficiency and accuracy of determining the k regional document data.

[0043] In one possible implementation, the step of extracting document data from the power document data to be processed based on the information of the k text regions to obtain the document data of the k regions includes: B1. Based on the information of k text regions, perform region labeling on the power document data to be processed to obtain the labeled power document data; B2. Extract data from the labeled power document data according to the labeled areas to obtain k area document data.

[0044] When annotating regions, the boundaries of the regions can be marked to obtain annotated power document data. Alternatively, all contours within a region can be marked. Specifically, different annotation methods can be used for region boundaries, and the annotation methods for contours within each region can differ from those for the region boundaries. This allows for better differentiation of regions and sub-regions within a single region. For example, different regions can be annotated with different boundary colors, enabling targeted annotation and improving recognizability.

[0045] After annotation, data can be extracted quickly based on the annotation boundaries to obtain the corresponding regional document data, thereby improving the efficiency of regional document data extraction.

[0046] In one possible implementation, the step of constructing a logical relationship scene graph based on key information from k documents and corresponding regional document data to obtain the target logical relationship scene graph includes: C1. Perform correlation analysis on the key information of k documents to obtain the target correlation information between the key information of each pair of documents; C2. Construct a logical relationship description between the key information of the two documents based on the target association information; C3. Construct a logical relationship scenario diagram based on the logical association description information between key information of every two documents and key information of k documents to obtain the target logical relationship scenario diagram.

[0047] Specifically, the cosine similarity between key information in each pair of documents can be calculated to obtain the target association information. The higher the cosine similarity, the higher the association; the lower the cosine similarity, the lower the association.

[0048] The method for constructing logical association description information based on target association information can be as follows: the logical association description information is determined based on the numerical range in which the target association information is located. Specifically, different numerical ranges have different logical association description information, such as high association degree, medium association degree, and low association degree.

[0049] Finally, the connection annotations are processed based on the logical association description information and the key information of the document. The higher the association, the shorter the connection distance; the lower the association, the longer the connection distance.

[0050] By constructing the logical relationship scenario diagram using the above method, the target logical relationship scenario diagram can be obtained, which can more intuitively display the relationship between document data in each region, improving the convenience and accuracy of subsequent data processing.

[0051] In one possible implementation, the step of constructing a logical relationship scene graph based on key information from k documents and corresponding regional document data to obtain the target logical relationship scene graph includes: D1. Extract key information from k documents to obtain k key feature information; D2. Randomly extract key feature information of the target from k key feature information; D3. Extract the correlation scores between k key feature information and target key feature information to obtain k correlation score values; D4. Determine the scene relationship distance based on the k relevance scores to obtain the k scene relationship distance values; D5. Extract the correlation description between k key feature information and the target key feature information to obtain k correlation description information; D6. Construct a logical relationship scene graph based on k relational description information, k scene relationship distance values, k document key information and corresponding regional document data to obtain the target logical relationship scene graph.

[0052] One approach is to use general feature extraction methods to extract key information from the document. For example, word frequency analysis can be used. Alternatively, a target key feature can be extracted from k key features, where the target key feature is any one of the k key features.

[0053] The relevance score can be determined by calculating the cosine similarity between k key features and the target key features. Alternatively, other similarity calculation methods can be used to calculate the similarity between the k key features and the target key features, and the similarity score can be determined accordingly.

[0054] The specific relationship between the scene relationship distance value and the relevance score value can be linear. For example, the larger the relevance score value, the smaller the scene relationship distance value, and vice versa. This allows us to determine the scene relationship distance value.

[0055] When determining the extraction of correlation description information, the correlation description information can be determined by combining the type information and correlation score value of the power document data to be processed. Specifically, a set of correlation description information can be determined based on the type information, and then the corresponding correlation description information can be determined from this set of correlation description information based on the mapping relationship between correlation score value and description information.

[0056] A scene graph construction template can be extracted based on type information. By inputting k relational descriptions, k scene relationship distance values, k key document information, and corresponding regional document data into this template, a target logical relational scene graph can be obtained. This allows for the rapid and accurate acquisition of the target logical relational scene graph.

[0057] For examples consistent with the above embodiments, please refer to... Figure 2 , Figure 2 This is a schematic diagram of the structure of a terminal provided in an embodiment of this application, such as... Figure 2As shown, it includes a processor, an input device, an output device, and a memory, which are interconnected. The memory is used to store a computer program, which includes program instructions. The processor is configured to call the program instructions. The program includes instructions for performing the following steps. Retrieve power document data to be processed; The power document data to be processed is extracted according to text regions to obtain k regions of document data; Obtain key information from k region document data to obtain key information for k documents; Based on the key information of k documents and the corresponding regional document data, a logical relationship scenario diagram is constructed to obtain the target logical relationship scenario diagram. Based on the target logical relationship scenario diagram, a preset text recognition model is used to perform structured processing on the power document data to be processed, resulting in a structured document.

[0058] In this example, by acquiring the power document data to be processed, document data is extracted from the power document data according to text regions to obtain k region document data. Key information of the k region document data is obtained to obtain k document key information. Based on the k document key information and the corresponding region document data, a logical relationship scene graph is constructed to obtain a target logical relationship scene graph. Based on the target logical relationship scene graph, a preset text recognition model is used to perform structured processing on the power document data to be processed to obtain a structured document. Therefore, it is possible to construct a logical relationship scene graph on the power document data to be processed, and perform structured processing based on the logical relationship scene graph and the preset text recognition model to obtain a structured document, thereby improving the accuracy of determining the structured document.

[0059] The above mainly describes the solutions of the embodiments of this application from the perspective of the method execution process. It is understood that, in order to achieve the above functions, the terminal includes the corresponding hardware structure and / or software modules for executing each function. Those skilled in the art should readily recognize that, in conjunction with the units and algorithm steps of the various examples described in the embodiments provided herein, this application can be implemented in hardware or a combination of hardware and computer software. Whether a function is executed in hardware or by computer software driving hardware depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0060] This application embodiment can divide the terminal into functional units according to the above method example. For example, each function can be divided into a separate functional unit, or two or more functions can be integrated into one processing unit. The integrated unit can be implemented in hardware or as a software functional unit. It should be noted that the unit division in this application embodiment is illustrative and only represents one logical functional division. In actual implementation, there may be other division methods.

[0061] For those consistent with the above, please refer to Figure 3 , Figure 3 This application provides a schematic diagram of a text structuring processing device based on a large character recognition model, as illustrated in this embodiment. Figure 3 As shown, the device includes: The first acquisition unit 301 is used to acquire power document data to be processed; Extraction unit 302 is used to extract document data from the power document data to be processed according to text regions to obtain k region document data; The second acquisition unit 303 is used to acquire key information of k region document data to obtain key information of k documents; Construction unit 304 is used to construct a logical relationship scene graph based on key information of k documents and corresponding regional document data, so as to obtain the target logical relationship scene graph. The processing unit 305 is used to perform structured processing on the power document data to be processed using a preset text recognition model based on the target logical relationship scene diagram, so as to obtain a structured document.

[0062] In this example, by acquiring the power document data to be processed, document data is extracted from the power document data according to text regions to obtain k region document data. Key information of the k region document data is obtained to obtain k document key information. Based on the k document key information and the corresponding region document data, a logical relationship scene graph is constructed to obtain a target logical relationship scene graph. Based on the target logical relationship scene graph, a preset text recognition model is used to perform structured processing on the power document data to be processed to obtain a structured document. Therefore, it is possible to construct a logical relationship scene graph on the power document data to be processed, and perform structured processing based on the logical relationship scene graph and the preset text recognition model to obtain a structured document, thereby improving the accuracy of determining the structured document.

[0063] In one possible implementation, the extraction unit 302 is specifically used for: Obtain the document attribute information of the power document data to be processed; The text area layout information of the power document data to be processed is determined based on the document attribute information. Based on the text region layout information, determine the k text region information corresponding to the power document data to be processed; Based on the information from the k text regions, document data is extracted from the power document data to be processed, resulting in k regions of document data.

[0064] In one possible implementation, in extracting document data from the power document data to be processed based on the information of the k text regions to obtain document data of the k regions, the extraction unit 302 is specifically used for: Based on the information of k text regions, the power document data to be processed is labeled to obtain labeled power document data; The labeled power document data is extracted according to the labeled areas to obtain k area document data.

[0065] In one possible implementation, the building unit 304 is specifically used for: Perform correlation analysis on the key information of k documents to obtain the target correlation information between the key information of every two documents; Based on the target association information, construct a logical association description between the key information of the two corresponding documents; Based on the logical relationship description information between key information of every two documents and key information of k documents, a logical relationship scenario diagram is constructed to obtain the target logical relationship scenario diagram.

[0066] In one possible implementation, the building unit 304 is specifically used for: Feature extraction is performed on key information from k documents to obtain k key feature information; Randomly extract key feature information of the target from k key feature information; Extract the correlation scores between k key feature information and target key feature information to obtain k correlation score values; The scene relationship distance is determined based on k relevance scores, resulting in k scene relationship distance values; Extract k key feature information and describe the correlation between them and the target key feature information to obtain k correlation description information; A logical relationship scene graph is constructed based on k relational descriptions, k scene relationship distance values, k key document information, and corresponding regional document data to obtain the target logical relationship scene graph.

[0067] This application also provides a computer storage medium storing a computer program for electronic data interchange, which causes a computer to perform some or all of the steps of any of the text structuring methods based on a large character recognition model as described in the above method embodiments.

[0068] This application also provides a computer program product, which includes a non-transitory computer-readable storage medium storing a computer program that causes a computer to perform some or all of the steps of any of the text structuring processing methods based on a large character recognition model as described in the above method embodiments.

[0069] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are preferred embodiments, and the actions and modules involved are not necessarily essential to this application.

[0070] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.

[0071] In the several embodiments provided in this application, it should be understood that the disclosed apparatus can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical or other forms.

[0072] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.

[0073] Furthermore, the functional units in the various embodiments of the application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software program module.

[0074] If the integrated unit is implemented as a software program module and sold or used as an independent product, it can be stored in a computer-readable storage device (CMD). Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a memory and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned memory includes various media capable of storing program code, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.

[0075] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be implemented by a program instructing related hardware. The program can be stored in a computer-readable storage device, which may include: a flash drive, a read-only memory, a random access memory, a magnetic disk, or an optical disk, etc.

[0076] The embodiments of this application have been described in detail above. Specific examples have been used to illustrate the principles and implementation methods of this application. The description of the above embodiments is only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A text structuring method based on a large-scale character recognition model, characterized in that, The method includes: Retrieve power document data to be processed; The power document data to be processed is extracted according to text regions to obtain k regions of document data; Obtain key information from k region document data to obtain key information for k documents; Based on the key information of k documents and the corresponding regional document data, a logical relationship scenario diagram is constructed to obtain the target logical relationship scenario diagram. Based on the target logical relationship scenario diagram, a preset text recognition model is used to perform structured processing on the power document data to be processed, resulting in a structured document.

2. The text structuring processing method based on a large character recognition model according to claim 1, characterized in that, The process of extracting document data from the power document data to be processed according to text regions yields k regions of document data, including: Obtain the document attribute information of the power document data to be processed; The text area layout information of the power document data to be processed is determined based on the document attribute information. Based on the text region layout information, determine the k text region information corresponding to the power document data to be processed; Based on the information from the k text regions, document data is extracted from the power document data to be processed, resulting in k regions of document data.

3. The text structuring processing method based on a large character recognition model according to claim 2, characterized in that, The step of extracting document data from the power document data to be processed based on the information of the k text regions, to obtain the document data of the k regions, includes: Based on the information of k text regions, the power document data to be processed is labeled to obtain labeled power document data; The labeled power document data is extracted according to the labeled areas to obtain k area document data.

4. The text structuring processing method based on a large character recognition model according to any one of claims 1-3, characterized in that, The step of constructing a logical relationship scene graph based on key information from k documents and corresponding regional document data to obtain the target logical relationship scene graph includes: Perform correlation analysis on the key information of k documents to obtain the target correlation information between the key information of every two documents; Based on the target association information, construct a logical association description between the key information of the two corresponding documents; Based on the logical relationship description information between key information of every two documents and key information of k documents, a logical relationship scenario diagram is constructed to obtain the target logical relationship scenario diagram.

5. The text structuring method based on a large character recognition model according to any one of claims 1-3, characterized in that, The step of constructing a logical relationship scene graph based on key information from k documents and corresponding regional document data to obtain the target logical relationship scene graph includes: Feature extraction is performed on key information from k documents to obtain k key feature information; Randomly extract key feature information of the target from k key feature information; Extract the correlation scores between k key feature information and target key feature information to obtain k correlation score values; The scene relationship distance is determined based on k relevance scores, resulting in k scene relationship distance values; Extract k key feature information and describe the correlation between them and the target key feature information to obtain k correlation description information; A logical relationship scene graph is constructed based on k relational descriptions, k scene relationship distance values, k key document information, and corresponding regional document data to obtain the target logical relationship scene graph.

6. A text structuring processing device based on a large-scale character recognition model, characterized in that, The device includes: The first acquisition unit is used to acquire power document data to be processed; The extraction unit is used to extract document data from the power document data to be processed according to text regions, and obtain document data in k regions; The second acquisition unit is used to acquire key information from k region document data to obtain key information for k documents; The construction unit is used to construct a logical relationship scene graph based on key information of k documents and corresponding regional document data, thereby obtaining the target logical relationship scene graph. The processing unit is used to perform structured processing on the power document data to be processed using a preset text recognition model based on the target logical relationship scenario diagram, so as to obtain a structured document.

7. The text structuring processing device based on a large character recognition model according to claim 6, characterized in that, The extraction unit is specifically used for: Obtain the document attribute information of the power document data to be processed; The text area layout information of the power document data to be processed is determined based on the document attribute information. Based on the text region layout information, determine the k text region information corresponding to the power document data to be processed; Based on the information from the k text regions, document data is extracted from the power document data to be processed, resulting in k regions of document data.

8. The text structuring processing device based on a large character recognition model according to claim 7, characterized in that, In the process of extracting document data from the power document data to be processed based on the information of the k text regions to obtain document data in the k regions, the extraction unit is specifically used for: Based on the information of k text regions, the power document data to be processed is labeled to obtain labeled power document data; The labeled power document data is extracted according to the labeled areas to obtain k area document data.

9. A terminal, characterized in that, The system includes a processor, an input device, an output device, and a memory, which are interconnected. The memory stores a computer program, which includes program instructions. The processor is configured to invoke the program instructions to execute the text structuring processing method based on a large character recognition model as described in any one of claims 1-5.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, the computer program including program instructions, which, when executed by a processor, cause the processor to perform the text structuring processing method based on a large character recognition model as described in any one of claims 1-5.