AI-based document data extraction method and device and storage medium

Through the AI-based document data extraction method, document analysis is performed using preprocessing and natural language algorithms, and feature extraction and fusion is combined with AI algorithms, the problem of difficulty in efficient extraction of complex document data in the existing technology is solved, and efficient and accurate document data extraction is achieved.

CN120197609APending Publication Date: 2025-06-24LIVEFAN INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510292393.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-12
Publication Date
2025-06-24

AI Technical Summary

Technical Problem

The prior art is difficult to efficiently and accurately automatically and intelligently extract data from complex documents from massive documents, especially when processing non-text content and diversified documents, which are inefficient and difficult to meet the requirements of depth and accuracy.

Method used

AI-based document data extraction method is used to analyze documents through pre-processing algorithms and natural language algorithms, generate entity relationship diagrams, and use AI algorithms to perform feature extraction and feature fusion, and finally generate extracted format data.

Benefits of technology

It realizes efficient and accurate extraction of complex document data, can handle heterogeneous documents and diversified documents, improves extraction efficiency and accuracy, and solves the problem of insufficient efficiency and accuracy of traditional methods when dealing with complex documents.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120197609A_ABST
    Figure CN120197609A_ABST
Patent Text Reader

Abstract

The invention relates to the field of document extraction, and discloses an AI-based document data extraction method and device and a storage medium. The method comprises the steps of obtaining to-be-processed document data; according to a preprocessing algorithm, preprocessing the to-be-processed document data to obtain preprocessed data; based on a natural language algorithm, performing document analysis processing on the preprocessed data to obtain an entity relation graph; according to an AI algorithm, feature extraction processing is carried out on the preprocessed data, entity features are obtained, and the entity features correspond to attribute features; according to the entity relation graph, performing feature fusion processing on the entity features and the attribute features corresponding to the entity features to obtain feature fusion data; and performing format conversion processing on the feature fusion data to generate extraction format data. In the embodiment of the invention, the data can be intelligently and automatically extracted from the document data, and the technical problem that the data of the complex document cannot be automatically and intelligently extracted at present is solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of document extraction, and particularly to an AI-based document data extraction method, device, and storage medium. Background Art

[0002] With the advent of the information age, the amount of various document data has exploded. How to efficiently and accurately extract valuable data from a large number of documents has become an urgent problem to be solved. Most traditional document data extraction methods rely on manual operations or simple keyword matching, which are not only inefficient but also prone to missing important information. Manual operations require a large amount of human and time resources and are overwhelmed when faced with a large number of documents. Simple keyword matching is too crude to handle complex semantics and diverse expressions and cannot meet the requirements for depth and accuracy.

[0003] In addition, with the increasing diversification of document types, including various forms such as pictures, tables, and charts, traditional methods are even more inadequate when dealing with these non-text contents. In practical applications, in fields such as enterprise document management, academic research data collation, and official document processing, problems such as information delay and decision-making errors caused by the limitations of traditional methods are common. Therefore, in order to solve the problem of being unable to automatically and intelligently extract data from complex documents currently, a new technology is needed to solve the current problem. Summary of the Invention

[0004] The main object of the present invention is to solve the technical problem of being unable to automatically and intelligently extract data from complex documents currently.

[0005] The first aspect of the present invention provides an AI-based document data extraction method, including the steps of:

[0006] Obtaining the document data to be processed;

[0007] Preprocessing the document data to be processed according to a preprocessing algorithm to obtain preprocessed data;

[0008] Performing document analysis processing on the preprocessed data based on a natural language algorithm to obtain an entity relationship graph;

[0009] Performing feature extraction processing on the preprocessed data according to an AI algorithm to obtain entity features and attribute features corresponding to the entity features;

[0010] Performing feature fusion processing on the entity features and the attribute features corresponding to the entity features according to the entity relationship graph to obtain feature fusion data;

[0011] Performing format conversion processing on the feature fusion data to generate extraction format data.

[0012] Optionally, in the first implementation manner of the first aspect of the present invention, the document analysis and processing of the preprocessed data based on the natural language algorithm to obtain the entity relationship graph includes:

[0013] Performing entity recognition processing on the preprocessed data based on the natural language algorithm to obtain an entity vector graph;

[0014] Performing relationship annotation processing on the entity vector data according to the text meaning of the preprocessed data to obtain an entity relationship graph.

[0015] Optionally, in the second implementation manner of the first aspect of the present invention, the preprocessing of the to-be-processed document data according to the preprocessing algorithm to obtain the preprocessed data includes:

[0016] Judging whether the to-be-processed document data is a heterogeneous document;

[0017] When it is a heterogeneous document, performing type splitting processing on the to-be-processed document to obtain N type documents and the type document sorting parameter, where N is a positive integer;

[0018] Inputting the N type documents into corresponding type parallel processing channels, performing parallel feature recognition on the N type documents to obtain the recognition characters corresponding to the N type documents;

[0019] Assembling the recognition characters corresponding to the N type documents according to the type document sorting parameter to obtain the preprocessed data.

[0020] Optionally, in the third implementation manner of the first aspect of the present invention, after judging whether the to-be-processed document data is a heterogeneous document, it further includes:

[0021] When it is not a heterogeneous document, performing recognition processing on the to-be-processed document data according to the document type of the to-be-processed document data to obtain the preprocessed data.

[0022] Optionally, in the fourth implementation manner of the first aspect of the present invention, the feature extraction processing of the preprocessed data according to the AI algorithm to obtain the entity features and the attribute features corresponding to the entity features includes:

[0023] Performing word vector conversion processing on the preprocessed data to obtain a word vector set;

[0024] Performing step-by-step hierarchical convolution processing on the word vector set based on a convolution kernel set to obtain a convolution vector set;

[0025] Performing full connection processing on the convolution vector set to obtain a one-dimensional flattened vector;

[0026] According to the activation function, perform activation processing on the one-dimensional flattened vector to obtain entity features and the attribute features corresponding to the entity features.

[0027] Optionally, in the fifth implementation manner of the first aspect of the present invention, the performing fully connected processing on the convolution vector set to obtain a one-dimensional flattened vector includes:

[0028] Based on the first projection weight, the second projection weight, and the third projection weight, perform convolution on the sequence form of the convolution vector set in sequence to obtain a Q matrix, a K matrix, and a V matrix;

[0029] According to the scaling factor, perform convolution normalization processing on the transpose of the Q matrix and the K matrix to obtain an attention normalization matrix;

[0030] Perform convolution processing on the attention normalization matrix and the V matrix to obtain a global aggregation matrix;

[0031] Perform one-dimensional projection processing on the global aggregation matrix to obtain a one-dimensional flattened vector.

[0032] Optionally, in the sixth implementation manner of the first aspect of the present invention, the performing feature fusion processing on the entity features and the attribute features corresponding to the entity features according to the entity relationship graph to obtain feature fusion data includes:

[0033] Based on the entity features and the attribute features corresponding to the entity features, perform attribute marking on the corresponding entity features in the entity relationship graph to generate feature fusion data.

[0034] Optionally, in the seventh implementation manner of the first aspect of the present invention, the performing format conversion processing on the feature fusion data to generate extraction format data includes:

[0035] Based on the parent-child relationship of the entity relationship graph, perform conversion processing on the feature fusion data into JOSN format or XML format to generate JOSN extraction format data or XML extraction format data.

[0036] The second aspect of the present invention provides an AI-based document data extraction device, including: a memory and at least one processor, wherein instructions are stored in the memory, and the memory and the at least one processor are interconnected by a line; the at least one processor calls the instructions in the memory so that the AI-based document data extraction device executes the above-mentioned AI-based document data extraction method.

[0037] The third aspect of the present invention provides a computer-readable storage medium storing instructions that, when run on a computer, cause the computer to execute the above-mentioned AI-based document data extraction method.

[0038] In the embodiments of the present invention, through a prediction processing method, heterogeneous data of a document to be processed is uniformly converted. Using natural language algorithms, the preprocessed data is analyzed and processed to obtain entity relationship data. Then, using AI algorithms, the attribute feature data and entity feature data of the preprocessed data are separated. Finally, by feature fusion and format conversion, the extracted data is generated. The extracted data can be used for web display, digital documents, and specific readers to achieve the extraction of document data, solving the technical problem that current complex document data cannot be automatically and intelligently extracted. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is a schematic diagram of an embodiment of the AI-based document data extraction method in the embodiments of the present invention;

[0040] Figure 2 It is a schematic diagram of a specific implementation manner of step 104 of the AI-based document data extraction method in the embodiments of the present invention;

[0041] Figure 3 It is a schematic diagram of a specific implementation manner of step 1044 of the AI-based document data extraction method in the embodiments of the present invention;

[0042] Figure 4 It is a schematic diagram of an embodiment of the AI-based document data extraction device in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0043] The embodiments of the present invention provide an AI-based document data extraction method, device, and storage medium.

[0044] The embodiments disclosed by the present invention will be described in more detail below with reference to the accompanying drawings. Although some embodiments of the present invention are shown in the drawings, it should be understood that the present invention can be implemented in various forms and should not be construed as limited to the embodiments set forth herein. Instead, these embodiments are provided to more thoroughly and completely understand the present disclosure. It should be understood that the drawings and embodiments of the present invention are only for exemplary purposes and are not used to limit the protection scope of the present invention.

[0045] In the description of the embodiments disclosed in the present invention, the term "including" and its similar terms should be understood as open inclusion, that is, "including but not limited to". The term "based on" should be understood as "at least partially based on". The term "one embodiment" or "the embodiment" should be understood as "at least one embodiment". The terms "first", "second", etc. may refer to different or the same objects. There may also be other explicit and implicit definitions hereinafter.

[0046] For ease of understanding, the specific process of the embodiments of the present invention will be described below. Please refer to Figure 1 One embodiment of the AI-based document data extraction method in the embodiments of the present invention includes:

[0047] 101. Obtain the document data to be processed;

[0048] In this embodiment, the document data to be processed is obtained. The content of this document data may include text, formulas, charts, etc. The format of this document data may be various formats such as text, image, PDF, Word, etc.

[0049] 102. Preprocess the document data to be processed according to the preprocessing algorithm to obtain preprocessed data;

[0050] In this embodiment, based on the preprocessing algorithm, document format conversion, text cleaning, denoising, etc. are performed to obtain preprocessed data to improve the accuracy and efficiency of subsequent processing.

[0051] Specifically, the following specific implementation manners are included in step 102:

[0052] 1021. Determine whether the document data to be processed is a heterogeneous document;

[0053] 1022. When it is a heterogeneous document, perform type splitting processing on the document to be processed to obtain N type documents and the sorting parameters of the type documents, where N is a positive integer;

[0054] 1024. Input the N type documents into the corresponding type parallel processing channels, and perform parallel feature recognition on the N type documents to obtain the recognition characters corresponding to the N type documents;

[0055] 1025. Assemble the recognition characters corresponding to the N type documents according to the sorting parameters of the type documents to obtain preprocessed data.

[0056] In steps 1021 - 1025, first analyze whether the document data to be processed is a heterogeneous document. A heterogeneous document is not a single - document form and is composed of different types of data structures or formats, including various media types such as text, images, audio, and video. These document data come from different sources and have different formats, structures, and semantic descriptions. For example, if the document data includes two types of document data, text and image, sort them according to the content of the original document and assign relevant sorting coordinate parameters to the text and image. Then input the text data with coordinate parameters into the document recognition processing module, and input the picture data with coordinate parameters into the picture ORC recognition processing module. The two processing modules process in parallel to obtain the recognition strings corresponding to the two processing modules. Combine the sorting parameters of the type documents corresponding to the recognition strings into a DataFrame - formatted table, and assemble and process the recognition characters based on the sorting parameters of the type documents to obtain pre - processed data.

[0057] Specifically, after step 1021, the following specific implementation manners are further included:

[0058] 10211. When it is not a heterogeneous document, identify and process the document data to be processed according to the document type of the document data to be processed to obtain pre - processed data.

[0059] In step 10211, if it is not a heterogeneous document, based on the type of the document data such as pictures, text, videos, etc., find the corresponding processing method for the document type, and identify and process the document to be processed to obtain pre - processed data.

[0060] 103. Based on natural language algorithms, perform document analysis processing on the pre - processed data to obtain an entity relationship diagram;

[0061] In this embodiment, through natural language algorithms such as the RNN algorithm or the CNN algorithm, perform document analysis processing on the pre - processed data to obtain an entity relationship diagram.

[0062] Suppose we have a document about the employees of a certain company, and the content is as follows:

[0063] "Zhang San is the superior of Li Si, and they both belong to the Marketing Department. Wang Wu works in the Marketing Department and is the subordinate of Zhang San. In addition, Zhao Liu is an employee of the Human Resources Department."

[0064] We can construct an entity relationship diagram according to the following steps:

[0065] 1. Data pre - processing: Perform word segmentation, part - of - speech tagging, etc. on the document to identify the entities and relevant information in the text. For example, "Zhang San", "Li Si", "Marketing Department", "Wang Wu", "Zhao Liu", etc. are identified as entities.

[0066] 2. Entity Recognition: Using NER technology, identify person name entities (Zhang San, Li Si, Wang Wu, Zhao Liu) and organization name entities (Marketing Department, Human Resources Department) from the text.

[0067] 3. Relationship Extraction: Extract the relationships between entities from the text through relationship extraction technology. For example, "Zhang San is the superior of Li Si" represents a superior-subordinate relationship, "They both belong to the Marketing Department" represents an affiliation relationship, "Wang Wu is the subordinate of Zhang San" also represents a superior-subordinate relationship, and "Zhao Liu is an employee of the Human Resources Department" represents an affiliation relationship.

[0068] 4. Entity Relationship Diagram Construction: Represent the identified entities and relationships in a graphical way. For example, an entity relationship diagram can be constructed with the following nodes and edges:

[0069] Nodes: Zhang San, Li Si, Wang Wu, Zhao Liu, Marketing Department, Human Resources Department.

[0070] Edges: Zhang San → Li Si (superior-subordinate relationship), Zhang San, Li Si → Marketing Department (affiliation relationship), Zhang San ← Wang Wu (superior-subordinate relationship), Zhao Liu → Human Resources Department (affiliation relationship).

[0071] Specifically, the following specific implementation manners are included in step 103:

[0072] 1031. Based on the natural language algorithm, perform entity recognition processing on the preprocessed data to obtain an entity vector diagram;

[0073] 1032. According to the text meaning of the preprocessed data, perform relationship annotation processing on the entity vector data to obtain an entity relationship diagram.

[0074] In steps 1031 - 1032, first, use the natural language algorithm to perform entity recognition processing on the preprocessed data to obtain an entity vector diagram, that is, identify the nodes: Zhang San, Li Si, Wang Wu, Zhao Liu, Marketing Department, Human Resources Department, and obtain the vectors: Zhang San → Li Si, Zhang San, Li Si → Marketing Department, Zhang San ← Wang Wu, Zhao Liu → Human Resources Department.

[0075] Then, based on the text meaning, annotate the meanings such as the superior-subordinate relationship and the affiliation relationship to obtain the entity relationship diagram:

[0076] Zhang San → Li Si (superior-subordinate relationship), Zhang San, Li Si → Marketing Department (affiliation relationship), Zhang San ← Wang Wu (superior-subordinate relationship), Zhao Liu → Human Resources Department (affiliation relationship).

[0077] 104. According to the AI algorithm, perform feature extraction processing on the preprocessed data to obtain entity features and the corresponding attribute features of the entity features;

[0078] In this embodiment, an AI algorithm is used to extract features from the preprocessed data, obtaining entity features of Zhang San, Li Si, Wang Wu, Zhao Liu, the Marketing Department, and the Human Resources Department, as well as attribute features corresponding to employees, employees, employees, departments, departments, etc. of Zhang San, Li Si, Wang Wu, Zhao Liu, the Marketing Department, and the Human Resources Department.

[0079] The AI algorithm can use a Convolutional Neural Network (CNN):

[0080] 1. Text feature extraction:

[0081] Word embedding layer: Convert each word in the text into a low-dimensional real-valued vector, i.e., a word vector. These word vectors can be obtained through a pre-trained word embedding model or learned together in the network.

[0082] Convolutional layer: Use multiple convolutional kernels of different sizes to perform sliding convolutional operations on the word vector sequence to extract local features. These features can be n-gram features, phrase features, etc. The output of the convolutional layer is usually a feature map, where each element represents the feature of the input text at a certain local position.

[0083] Pooling layer: Perform a pooling operation on the output of the convolutional layer, such as max pooling or average pooling, to reduce the dimension of the feature map and retain the most important information. The pooling layer can help the model capture the most significant features in the text while reducing the computational amount.

[0084] 2. Text classification:

[0085] Fully connected layer: Flatten the output of the pooling layer into a one-dimensional vector and pass it to the fully connected layer. The fully connected layer can learn the mapping relationship between the text features and the class labels.

[0086] Output layer: Usually use the softmax function as the activation function of the output layer to convert the output of the fully connected layer into a probability distribution. The probability of each class represents the likelihood that the text belongs to that class.

[0087] Loss function and optimization: Use the cross-entropy loss function to measure the difference between the probability distribution predicted by the model and the true class labels, and adjust the network weights through the backpropagation algorithm and an optimizer (such as Adam, SGD, etc.) to minimize the loss function.

[0088] Specifically, please refer to Figure 2 , Figure 2 which is a schematic diagram of a specific implementation manner of step 104 in the embodiment of the present invention. The following specific implementation manners are included in step 104:

[0089] 1041. Perform word vector conversion processing on the preprocessed data to obtain a word vector set;

[0090] 1042. Based on the convolution kernel set, perform step-by-step hierarchical convolution processing on the word vector set to obtain a convolution vector set;

[0091] 1043. Perform a fully connected process on the convolution vector set to obtain a one-dimensional flattened vector;

[0092] 1044. According to the activation function, perform activation processing on the one-dimensional flattened vector to obtain entity features and the corresponding attribute features of the entity features.

[0093] In steps 1041 - 1044, first convert the preprocessed data into word vectors to obtain a set of multiple word vectors, such as {A1, A2, A3, A4, A5, A6}, and then for the convolution kernel set {{B1, B2}, {B3, B4}, {B5, B6}, {B7, B8}, {B9, B 10}, {B 11 , B 12}}, and then convolve A 1* B 1* B2,....., A 6* B 11* B 12 to form a convolution vector set {C1, C2, C3, C4, C5, C6}, and perform a one-dimensional flattening process on {C1, C2, C3, C4, C5, C6} to obtain a one-dimensional flattened vector of X[1*n].

[0094] Based on the softmax activation function or the ReLU function, perform activation processing on the one-dimensional flattened vector X[1*n] to obtain entity features and the corresponding attribute features of the entity features.

[0095] Specifically, please refer to Figure 3 , Figure 3 which is a schematic diagram of a specific implementation manner of step 1044 in the embodiments of the present invention. The following specific implementation manners are included in step 1044:

[0096] 10441. Based on the first projection weight, the second projection weight, and the third projection weight, perform convolution on the sequence form of the convolution vector set in sequence to obtain a Q matrix, a K matrix, and a V matrix;

[0097] 10442. According to the scaling factor, perform convolution normalization processing on the transpose of the Q matrix and the K matrix to obtain an attention normalization matrix;

[0098] 10443. Perform convolution processing on the attention normalization matrix and the V matrix to obtain a global aggregation matrix;

[0099] 10444. Perform one-dimensional projection processing on the global aggregation matrix to obtain a one-dimensional flattened vector.

[0100] In steps 10441 - 10444, the first projection weight W Q , the second projection weight W K , the third projection weight W V , flatten the convolutional vector set X into a sequence form X', and obtain the Q matrix = X' * W Q , the K matrix = X' * W K , the V matrix == X' * W V .

[0101] Set √d k as the scaling factor, where d k is the key dimension, A = softmax(Q * K T ) / (√d k ) attention normalization matrix, and then obtain the global aggregation matrix Z = A * V. Perform a one-dimensional projection process on the global aggregation matrix Z to obtain a one-dimensional flattened vector, that is, the one-dimensional flattened vector output = Z * W o , where, W o is the projection matrix, and this projection matrix is a matrix for single-dimensional projection, that is, finally obtain a one-dimensional flattened vector output of one dimension.

[0102] 105. According to the entity relationship diagram, perform feature fusion processing on the entity features and the corresponding attribute features of the entity features to obtain feature fusion data;

[0103] In this embodiment, label each entity feature data in the entity relationship diagram of Zhang San → Li Si (superior-subordinate relationship), Zhang San, Li Si → Marketing Department (subordination relationship), Zhang San ← Wang Wu (superior-subordinate relationship), Zhao Liu → Human Resources Department (subordination relationship) with attribute labels to obtain the following feature fusion data: Zhang San [employee] → Li Si [employee] (superior-subordinate relationship), Zhang San [employee], Li Si [employee] → Marketing Department [department] (subordination relationship), Zhang San [employee] ← Wang Wu [employee] (superior-subordinate relationship), Zhao Liu [employee] → Human Resources Department [department] (subordination relationship).

[0104] Specifically, step 105 includes the following specific implementation manners:

[0105] 1051. Based on the entity features and the corresponding attribute features of the entity features, perform attribute labeling on the corresponding entity features in the entity relationship diagram to generate feature fusion data.

[0106] In this embodiment, the mapping relationships of Zhang San, Li Si, Wang Wu, Zhao Liu, the Marketing Department, and the Human Resources Department corresponding to employees, employees, employees, departments, and departments are used. Taking Zhang San, Li Si, Wang Wu, Zhao Liu, the Marketing Department, and the Human Resources Department as keys and employees, employees, employees, departments, and departments as values, they are matched to the entity relationship diagram of Zhang San → Li Si (superior-subordinate relationship), Zhang San, Li Si → Marketing Department (subordination relationship), Zhang San ← Wang Wu (superior-subordinate relationship), Zhao Liu → Human Resources Department (subordination relationship) to obtain feature fusion data.

[0107] 106. Perform format conversion processing on the feature fusion data to generate extraction format data.

[0108] In this embodiment, the feature fusion data is as follows: Zhang San [employee] → Li Si [employee] (superior-subordinate relationship), Zhang San [employee], Li Si [employee] → Marketing Department [department] (subordination relationship), Zhang San [employee] ← Wang Wu [employee] (superior-subordinate relationship), Zhao Liu [employee] → Human Resources Department [department] (subordination relationship). Converting the data with the above relationship logic can be converted into JOSN format or XML format.

[0109] Specifically, the following specific implementation manners are included in step 106:

[0110] 1061. Based on the parent-child relationship of the entity relationship diagram, perform JOSN format or XML format conversion processing on the feature fusion data to generate JOSN extraction format data or XML extraction format data.

[0111] In step 1061, the subordinate relationships of Zhang San [employee] → Li Si [employee] (superior-subordinate relationship), Zhang San [employee], Li Si [employee] → Marketing Department [department] (subordination relationship), Zhang San [employee] ← Wang Wu [employee] (superior-subordinate relationship), Zhao Liu [employee] → Human Resources Department [department] (subordination relationship) are converted into a structured JOSN format data of the following content:

[0112] {

[0113] "employees":

[0114] {

[0115]

[0116] In an embodiment of the present invention, through a prediction processing method, heterogeneous data of a document to be processed is uniformly converted. Using a natural language algorithm, the preprocessed data is subjected to document analysis processing to obtain entity relationship data. Then, an AI algorithm is used to separate the attribute feature data and entity feature data of the preprocessed data. Finally, feature fusion and format conversion are used to generate extracted data. The extracted data can be used for web display, digital documents, and specific readers, realizing the extraction of document data and solving the technical problem that current complex document data cannot be automatically and intelligently extracted.

[0117] Figure 4 FIG. 4 is a schematic structural diagram of an AI-based document data extraction device provided by an embodiment of the present invention. The AI-based document data extraction device 400 may vary greatly due to configuration or performance differences, and may include one or more processors (central processing units, CPUs) 410 (for example, one or more processors) and a memory 420, and one or more storage media 430 (for example, one or more mass storage devices) for storing application programs 433 or data 432. Among them, the memory 420 and the storage media 430 may be transient storage or persistent storage. The program stored in the storage media 430 may include one or more modules (not shown in the figure), and each module may include a series of instruction operations on the AI-based document data extraction device 400. Further, the processor 410 may be configured to communicate with the storage media 430 and execute a series of instruction operations in the storage media 430 on the AI-based document data extraction device 400.

[0118] The AI-based document data extraction device 400 may further include one or more power supplies 440, one or more wired or wireless network interfaces 450, one or more input / output interfaces 460, and / or one or more operating systems 431, such as Windows Server, Mac OS X, Unix, Linux, FreeBSD, and so on. Those skilled in the art can understand that Figure 4 the shown structural diagram of the AI-based document data extraction device does not limit the AI-based document data extraction device, and it may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0119] The present invention also provides a computer-readable storage medium. The computer-readable storage medium may be a non-volatile computer-readable storage medium or a volatile computer-readable storage medium. Instructions are stored in the computer-readable storage medium. When the instructions run on a computer, the computer is caused to execute the steps of the AI-based document data extraction method.

[0120] In the context of the present disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0121] Moreover, although the operations are depicted in a particular order, this should be understood as requiring that such operations be performed in the particular order shown or in sequential order, or that all illustrated operations be performed to achieve the desired result. In certain circumstances, multitasking and parallel processing may be advantageous. Likewise, although several specific implementation details are included in the above discussion, these should not be construed as limiting the scope of the present disclosure. Certain features that are described in the context of separate embodiments can also be implemented in combination in a single implementation. Conversely, the various features that are described in the context of a single implementation can also be implemented separately or in any suitable subcombination in multiple implementations.

[0122] Although the subject matter has been described in language specific to structural features and / or methodological acts, it should be understood that the subject matter defined in the appended claims is not necessarily limited to the specific features or acts described above. Rather, the specific features and acts described above are merely example forms of implementing the claims.

Claims

1. A document data extraction method based on AI, characterized in that: Includes steps: Get the document data to be processed; Preprocessing the document data to be processed according to a preprocessing algorithm to obtain preprocessed data; Based on a natural language algorithm, document analysis is performed on the preprocessed data to obtain an entity relationship diagram; According to the AI ​​algorithm, feature extraction processing is performed on the preprocessed data to obtain entity features and attribute features corresponding to the entity features; According to the entity relationship diagram, the entity features and the attribute features corresponding to the entity features are subjected to feature fusion processing to obtain feature fusion data; The feature fusion data is format converted to generate extraction format data.

2. The AI-based document data extraction method according to claim 1, characterized in that: The document analysis and processing of the pre-processed data based on the natural language algorithm to obtain the entity relationship diagram includes: Based on a natural language algorithm, entity recognition processing is performed on the preprocessed data to obtain an entity vector graph; According to the textual meaning of the preprocessed data, the entity vector data is subjected to relationship annotation processing to obtain an entity relationship graph.

3. The AI-based document data extraction method according to claim 1, characterized in that: The preprocessing of the document data to be processed according to the preprocessing algorithm to obtain the preprocessed data includes: Determining whether the document data to be processed is a heterogeneous document; When it is a heterogeneous document, the document to be processed is split into types to obtain N types of documents and sorting parameters of the types of documents, where N is a positive integer; Inputting N documents of the type into the corresponding type parallel processing channel, performing parallel feature recognition on the N documents of the type, and obtaining recognition characters corresponding to the N documents of the type; According to the sorting parameters of the type of documents, the recognition characters corresponding to N documents of the type are assembled to obtain preprocessed data.

4. The AI-based document data extraction method according to claim 3, characterized in that: After determining whether the document data to be processed is a heterogeneous document, the method further includes: When it is not a heterogeneous document, the document data to be processed is identified and processed according to the document type of the document data to be processed to obtain pre-processed data.

5. The AI-based document data extraction method according to claim 1, characterized in that: The preprocessed data is subjected to feature extraction processing according to the AI ​​algorithm to obtain entity features, and the attribute features corresponding to the entity features include: Performing word vector conversion processing on the preprocessed data to obtain a word vector set; Based on the convolution kernel set, the word vector set is subjected to step-by-step hierarchical convolution processing to obtain a convolution vector set; Performing full connection processing on the convolution vector set to obtain a one-dimensional flattened vector; According to the activation function, the one-dimensional flattened vector is activated to obtain entity features and attribute features corresponding to the entity features.

6. The AI-based document data extraction method according to claim 5, characterized in that: The step of performing full connection processing on the convolution vector set to obtain a one-dimensional flattened vector comprises: Based on the first projection weight, the second projection weight, and the third projection weight, the sequence form of the convolution vector set is sequentially convolved to obtain a Q matrix, a K matrix, and a V matrix; According to the scaling factor, performing convolution normalization processing on the transpose of the Q matrix and the K matrix to obtain an attention normalization matrix; Convolving the attention normalization matrix with the V matrix to obtain a global aggregation matrix; A one-dimensional projection process is performed on the global aggregation matrix to obtain a one-dimensional flattened vector.

7. The AI-based document data extraction method according to claim 1, characterized in that: The step of performing feature fusion processing on the entity features and the attribute features corresponding to the entity features according to the entity relationship diagram to obtain feature fusion data includes: Based on the entity features and the attribute features corresponding to the entity features, attribute marking is performed on the corresponding entity features in the entity relationship graph to generate feature fusion data.

8. The AI-based document data extraction method according to claim 1, characterized in that: The performing format conversion processing on the feature fusion data to generate the extraction format data comprises: Based on the parent-child relationship of the entity relationship diagram, the feature fusion data is converted into a JOSN format or an XML format to generate JOSN extraction format data or XML extraction format data.

9. An AI-based document data extraction device, characterized in that: The AI-based document data extraction device comprises: a memory and at least one processor, the memory stores instructions, and the memory and the at least one processor are interconnected via a line; The at least one processor calls the instructions in the memory to enable the AI-based document data extraction device to execute the AI-based document data extraction method as described in any one of claims 1-8.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the AI-based document data extraction method as described in any one of claims 1 to 8 is implemented.