Visual rich document information extraction method and device, equipment, medium and program product

By lightening the initial information extraction model and replacing the network structure, the problem of high computing resources and operation and maintenance costs when dealing with multiple card and certificate scenarios is solved, and the effect of reducing the system processing pressure and cost is achieved.

CN119940353APending Publication Date: 2025-05-06BANK OF COMMUNICATIONS
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202411695870.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2024-11-25
Publication Date
2025-05-06

AI Technical Summary

Technical Problem

When dealing with a wide variety of card and certificate scenarios, the existing technology requires the deployment of multiple different models, resulting in a significant increase in computing resource costs and system operation and maintenance costs.

Method used

By lightening the trained initial information extraction model, the network parameters of the model are reduced, the information extraction model is built, the system's processing pressure is reduced, and some network structures are replaced by other network structures with smaller network parameters, further reducing network parameters.

Benefits of technology

While maintaining the accuracy of information extraction, it significantly reduces the model size and reduces the system's computing resource costs and operation and maintenance costs.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119940353A_ABST
    Figure CN119940353A_ABST
Patent Text Reader

Abstract

The invention provides a visual rich document information extraction method and device, equipment, a medium and a program product. Relates to the technical field of text information extraction. The method comprises the following steps: acquiring a visual rich document to be processed, and determining at least one group of text information contained in the visual rich document; performing information extraction on each piece of text information by adopting a trained information extraction model to obtain an extraction result of the visual rich document; wherein the information extraction model is constructed based on a preset neural network structure and a network structure obtained by distilling the trained initial information extraction model. According to the method, the technical problem that the processing pressure is large in the text information extraction process of the system is solved, the trained initial model is subjected to lightweight processing, the accuracy is kept as much as possible, meanwhile, the model size is greatly reduced, the data processing pressure of the system is reduced, and the user experience is improved. The computing resource cost and the operation and maintenance cost of the system are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the technical field of text information extraction, and in particular to a method, device, equipment, medium and program product for extracting visual rich document information. Background Art

[0002] Visually-Rich Document (VRD) refers to text that contains rich visual and structured information, usually including not only text content, but also visual elements such as images, tables, graphics, layout, format, etc. Key Information Extraction is a technology that automatically identifies and extracts valuable information from unstructured data. This technology is usually used to extract key information such as entities, relationships, events, etc. from different types of visually rich documents such as text, tables, images, etc., for subsequent analysis and processing.

[0003] In recent years, with the development of text artificial intelligence technology, new solutions have been provided for visual rich document processing and data analysis. In particular, the introduction of neural network models can better understand the contextual semantics of text and automatically extract key data from it, thus achieving more effective processing of scanned visual rich documents and greatly improving the accuracy and efficiency of information extraction.

[0004] However, in actual applications, processing a wide variety of card scenarios requires deploying multiple different models, which will significantly increase computing resource costs and system operation and maintenance costs. Summary of the invention

[0005] The present application provides a method, device, equipment, medium and program product for extracting visual rich document information, which are used to solve the technical problem of high processing pressure in the system during text information extraction in scenarios with a large number of card texts. By lightweight processing the trained initial model, the model size can be greatly reduced while maintaining accuracy as much as possible, reducing the data processing pressure of the system, and reducing the computing resource cost and operation and maintenance cost of the system.

[0006] In a first aspect, the present application provides a method for extracting information from a visually rich document, comprising:

[0007] Acquire a visually rich document to be processed, and determine at least one set of text information contained in the visually rich document;

[0008] A trained information extraction model is used to extract information from each of the text information to obtain an extraction result of the visual rich document; wherein the information extraction model includes a first network structure in a model obtained by distilling the trained initial information extraction model, and a preset neural network structure obtained by replacing a second network structure in a model obtained by distilling the trained initial information extraction model.

[0009] In an optional implementation, the model parameters of the initial information extraction model are distilled to obtain a shallow encoding module and a fusion module;

[0010] Determine a deep encoding module and a decoding module based on the preset neural network structure;

[0011] The information extraction model is obtained according to the shallow encoding module, the fusion module, the deep encoding module and the decoding module.

[0012] In an optional implementation, the shallow encoding module in the initial information extraction model includes an embedding layer and a position encoding layer; the fusion module in the initial information extraction model is a self-attention encoding network;

[0013] The model parameters of the initial information extraction model are distilled to obtain the shallow encoding module and the fusion module, including:

[0014] Performing distillation processing on the embedding layer and the position encoding layer to obtain the shallow encoding module;

[0015] The self-attention encoding network is distilled to obtain the fusion module.

[0016] In an optional implementation, the deep encoding module is a two-layer bidirectional long short-term memory network; the number of output feature dimensions of the two-layer bidirectional long short-term memory network is half the number of output feature dimensions of the deep encoding module in the initial information extraction model.

[0017] In an optional implementation, the training of the information extraction model includes:

[0018] The constructed information extraction model is trained using the training data set to obtain a trained information extraction model; the loss function used in the training includes information extraction loss and model distillation loss; among them,

[0019] The information extraction loss represents the difference between a first predicted extraction result obtained by performing information extraction on any training text in the training data set based on the information extraction model and the label of the training text;

[0020] The model distillation loss represents the difference between a second predicted extraction result obtained by extracting information from the training text by the initial information extraction model and the first predicted extraction result.

[0021] In an optional implementation manner, any set of text information includes text content information and text position information corresponding to the text content information;

[0022] The trained information extraction model is used to extract information from each of the text information to obtain the extraction results of the visually rich document, including:

[0023] For any set of text information, the text content information and the text position information are respectively encoded to obtain the corresponding text content vector and text position vector;

[0024] Fusing the text content vector and the text position vector to obtain a multimodal fusion vector;

[0025] Performing deep feature encoding on the multimodal fusion vector to obtain a multimodal deep feature vector;

[0026] Information extraction is performed on the multimodal deep feature vector to obtain an extraction result corresponding to the visual rich document.

[0027] In a second aspect, the present application provides a visual rich document information extraction device, comprising:

[0028] A file information acquisition module, used to acquire a visually rich document to be processed and determine at least one set of text information contained in the visually rich document;

[0029] An extraction result acquisition module is used to use a trained information extraction model to extract information from each of the text information to obtain an extraction result of the visual rich document; wherein the information extraction model includes a first network structure in a model obtained by distilling the trained initial information extraction model, and a preset neural network structure obtained by replacing the second network structure in a model obtained by distilling the trained initial information extraction model.

[0030] In a third aspect, the present application provides an electronic device, comprising: a processor, and a memory communicatively connected to the processor;

[0031] The memory stores computer-executable instructions;

[0032] The processor executes the computer-executable instructions stored in the memory to implement the method according to the first aspect.

[0033] In a fourth aspect, the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores computer-executable instructions, and when the computer-executable instructions are executed by a processor, they are used to implement the method described in the first aspect.

[0034] In a fifth aspect, the present application provides a computer program product, including a computer program, which implements the method described in the first aspect when executed by a processor.

[0035] The visual rich document information extraction technology provided in the present application obtains a visual rich document to be processed, determines at least one group of text information contained in the visual rich document; uses a trained information extraction model to extract information from each of the text information, and obtains an extraction result of the visual rich document; wherein the information extraction model is constructed based on a preset neural network structure and a network structure obtained by distilling the trained initial information extraction model; in the above scheme, since the information extraction model is constructed based on a network structure obtained by distilling a large model with information extraction capability, the network parameters of the model can be reduced, so that the subsequent model can reduce the processing pressure of the system when extracting information; further, during construction, the network structure obtained after distillation is optimized by other network structures with smaller network parameters, so as to further reduce the network parameters and reduce the computing resource cost and system operation and maintenance cost of the model when extracting information. BRIEF DESCRIPTION OF THE DRAWINGS

[0036] The accompanying drawings, which are incorporated in and constitute a part of this specification, illustrate embodiments consistent with the present application and, together with the description, serve to explain the principles of the present application.

[0037] Figure 1 An application scenario diagram of the visual rich document information extraction method provided in this application;

[0038] Figure 2 A flowchart of a method for extracting visually rich document information provided in an embodiment of the present application;

[0039] Figure 3 A schematic diagram of the structure of an information extraction model provided in an embodiment of the present application;

[0040] Figure 4 A schematic diagram of the structure of a visual rich document information extraction device provided in an embodiment of the present application;

[0041] Figure 5 It is a block diagram of an electronic device shown in an embodiment of the present application.

[0042] The above drawings have shown clear embodiments of the present application, which will be described in more detail later. These drawings and text descriptions are not intended to limit the scope of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. DETAILED DESCRIPTION

[0043] Exemplary embodiments will be described in detail herein, examples of which are shown in the accompanying drawings. When the following description refers to the drawings, the same numbers in different drawings represent the same or similar elements unless otherwise indicated. The implementations described in the following exemplary embodiments do not represent all implementations consistent with the present application. Instead, they are merely examples of devices and methods consistent with some aspects of the present application as detailed in the appended claims.

[0044] Visually-Rich Document (VRD) refers to text that contains rich visual and structured information, usually including not only text content, but also visual elements such as images, tables, graphics, layout, format, etc. Key Information Extraction is a technology that automatically identifies and extracts valuable information from unstructured data. This technology is usually used to extract key information such as entities, relationships, events, etc. from different types of visually rich documents such as text, tables, images, etc., for subsequent analysis and processing.

[0045] In recent years, with the development of text artificial intelligence technology, new solutions have been provided for extracting key information from visually rich documents. In particular, the introduction of neural network models can better understand the contextual semantics of text and automatically extract key data from it, thus achieving more effective processing of scanned visually rich documents and greatly improving the accuracy and efficiency of information extraction.

[0046] In practical applications, for scenarios where a wide variety of cards are processed, it is required to deploy multiple different models, that is, each card type corresponds to a corresponding information extraction model. However, since information extraction models usually contain many parameters, the deployment of multiple models means that more computing resources will be consumed by the system, resulting in a significant increase in computing resource costs and system operation and maintenance costs.

[0047] The visual rich document information extraction method provided in this application is intended to solve the above technical problems of the prior art. Specifically, by distilling a pre-fine-tuned large model with information extraction capabilities, an information extraction model is constructed, and by reducing the network parameters of the model, the processing pressure of the system is reduced; on this basis, other network structures with smaller network parameters are used to replace part of the network structure in the distilled information extraction model, further reducing network parameters, and reducing computing resource costs and system operation and maintenance costs.

[0048] Figure 1 An application scenario diagram of the visual rich document information extraction method provided in this application. The technical solution provided in this application can be applied to application scenarios involving information extraction in the process of bill data processing in the financial industry. Banks and financial institutions need to process a large number of visual texts such as bills, invoices and contracts. These visual rich documents are usually digitally archived by scanning or photographing. The visual rich document information extraction method provided in this application extracts key information such as entities, relationships, events, etc. from visual rich documents. Specifically, it automatically extracts key information (such as amount, date, supplier information) in invoices and receipts for financial record and reimbursement processing.

[0049] For ease of understanding, the following Figure 1 The application scenarios to which the embodiments of the present application are applicable are described. Figure 1 The information extraction technology provided in this application involves a text information collection device 11 and a text information extraction device 12; specifically, in the text information collection device 11, optical character recognition (OCR) technology can be used to extract text information from visually rich documents containing text information such as financial bills, wherein each set of text information contains text content information and text position information corresponding to the text content information. The extracted text information is transmitted to the text information extraction device 12 for information extraction processing.

[0050] In the text information extraction device 12, a lightweight information extraction model is pre-set, and the information extraction model is used to process the text information of the visual rich document transmitted by the information acquisition device 11 to obtain the corresponding extraction result; during the extraction process, the information extraction model after lightweight processing is used to extract information from the visual rich document, so as to reduce the consumption of system computing resources during the extraction process and reduce the computing resource cost and operation and maintenance cost of the system.

[0051] It should be understood that the information extraction technology provided in this application can be applied to various scenarios such as file management, mobile office, educational examinations and legal contract management in addition to financial instrument processing scenarios. The embodiments of this application do not specifically limit the application scenarios of the information extraction technology provided.

[0052] The technical solution of the present application and how the technical solution of the present application solves the above-mentioned technical problems are described in detail below with specific embodiments. The following specific embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below in conjunction with the accompanying drawings.

[0053] Figure 2 A flowchart of a method for extracting visual rich document information provided in an embodiment of the present application. The method can be executed by a visual rich document information extraction device, which can be a server or an electronic device. The following description takes an electronic device as an example. The method in this embodiment can be implemented by software, hardware, or a combination of software and hardware, such as Figure 2 As shown, the method includes the following steps.

[0054] S201. Obtain a visually rich document to be processed, and determine at least one set of text information contained in the visually rich document.

[0055] In this application, a visually rich document can be understood as a text containing visual elements such as images, tables, graphics, layouts, formats, etc. The acquisition method of a visually rich document includes but is not limited to text images captured or scanned by a scanner, a digital camera, or a mobile phone camera.

[0056] When a visually rich document is obtained, optical character recognition technology can be used to identify and process the visually rich document to obtain at least one set of text information contained in the document; wherein, any set of text information can include text content information in the document and text position information corresponding to the text content information.

[0057] Exemplarily, text content information can be understood as a group of characters obtained by identifying any image area in a visually rich document, that is, information representing the text content in the document; text position information can be understood as the corresponding coordinate positions of the group of characters in the text image, that is, information representing the layout of the text content in the document.

[0058] S202: Use the trained information extraction model to extract information from each text information to obtain an extraction result of a visually rich document.

[0059] In this application, the information extraction model can be implemented using deep learning techniques such as convolutional neural networks, recurrent neural networks, long short-term memory networks, and models based on Transformer architectures. Specifically, the information extraction model can be pre-built by a neural network structure and trained by pre-set information tags in visually rich documents and texts. Of course, the information extraction model can also be other types of models, which are not specifically limited.

[0060] It can be understood that in order to reduce system resource consumption and reduce system processing pressure, the information extraction model used in this application includes a first network structure in a model obtained by distilling the trained initial information extraction model, and a preset neural network structure obtained by replacing the second network structure in the model obtained by distilling the trained initial information extraction model.

[0061] Specifically, the network structure obtained after distillation of a large model with information extraction capabilities can be used as a basis, and the network structure obtained after distillation can be further optimized through other network structures with smaller network parameters to obtain a completed information extraction model, thereby reducing network parameters and reducing computing resource costs and system operation and maintenance costs.

[0062] On this basis, the pre-acquired target data set is used to train the constructed information extraction model to obtain an information extraction model with better extraction performance.

[0063] Specifically, when extracting any visually rich document, the text information corresponding to the visually rich document is input into the information extraction model for feature extraction and prediction processing to obtain the extraction result corresponding to the visually rich document.

[0064] It is understandable that the extraction results generally include structured or semi-structured data extracted from visually rich documents, and these results may vary depending on the specific application and objectives. For example, entity recognition, that is, naming entities, such as names of people, places, organization names, dates, monetary amounts, etc.; another example is relationship extraction, that is, obtaining the relationship between entities, such as "Company A acquires Company B", "Person X was born in Place Y", etc. The extraction results in this application also include other types of extraction results, which are not specifically limited and are not exemplified one by one here.

[0065] It should be understood that these extraction results can be used for further data analysis, decision support, information retrieval and automated processing, etc.

[0066] In the above technical scheme, a visually rich document to be processed is obtained, and at least one set of text information contained in the visually rich document is determined; a trained information extraction model is used to extract information from each text information to obtain an extraction result of the visually rich document; wherein the information extraction model is constructed based on a preset neural network structure and a network structure obtained by distilling the trained initial information extraction model; in the above scheme, since the information extraction model is constructed based on a network structure obtained by distilling a large model with information extraction capability, the network parameters of the model can be reduced, so that the subsequent model can reduce the processing pressure of the system when extracting information; further, during construction, the network structure obtained after distillation is optimized by other network structures with smaller network parameters, so as to further reduce the network parameters and reduce the computing resource cost and system operation and maintenance cost of the model when extracting information.

[0067] Next, the construction process of the information extraction model is introduced in detail.

[0068] Optionally, the present application can reduce network parameters by compressing the initial information extraction model after pre-training and fine-tuning, or by replacing the network structure and adopting a network structure with smaller network parameters to obtain an information extraction model after lightweight processing.

[0069] In some optional embodiments, the present application constructs an information extraction model based on a preset neural network structure and a network structure after distillation of a trained initial information extraction model, and the process includes: distilling the model parameters of the shallow encoding module and the fusion module in the initial information extraction model to obtain a shallow encoding module and a fusion module; determining a deep encoding module and a decoding module based on the preset neural network structure; and obtaining an information extraction model based on the shallow encoding module, the fusion module, the deep encoding module and the decoding module.

[0070] It can be explained that model distillation aims to reduce the requirements of computing resources and improve inference speed by transferring the knowledge of a large, complex "teacher" model to a smaller "student" model.

[0071] That is, in the present application, it can be understood that a series of processes such as distillation are performed on the initial information extraction model to obtain a constructed information extraction model.

[0072] In some optional implementations, the open source LILT-Base (Language-independentLayout Transformer) model is usually selected as the neural network model for processing visually rich documents. This is because the model is based on the Transformer architecture and integrates visual information such as text and layout in the document to achieve a more comprehensive understanding of the document content and thus better extract information from the document.

[0073] Therefore, in this application, the pre-trained LILT-Base model can be selected as the teacher model, and the target data set can be used to fine-tune the teacher model to obtain a teacher model with information extraction capabilities, that is, the initial information extraction model in this application.

[0074] During fine-tuning training, all text regions identified in each visually rich document and the text position information corresponding to the text regions are used as input, and the teacher model is fine-tuned for the Named Entity Recognition (NER) task. The cross entropy is used as the loss function during the fine-tuning process.

[0075] When the fine-tuned teacher model is obtained, the student model is obtained by distilling the teacher model, and the student model is used as the basic network structure for constructing the information extraction model in this application.

[0076] Since the distilled model is not fully used as the information extraction model in this application, part of the network structure in the teacher model can be distilled to construct an information extraction model. This can reduce resource consumption in the model distillation process, improve the model distillation efficiency, and thus improve the model construction efficiency.

[0077] Optionally, the shallow coding module in the initial information extraction model in the present application includes an embedding layer and a position coding layer; the fusion module in the initial information extraction model is a self-attention coding network; accordingly, the model parameters of the initial information extraction model are distilled to obtain a shallow coding module and a fusion module, including: copying the embedding layer and the position coding layer to obtain a shallow coding module; and distilling the self-attention coding network to obtain a fusion module.

[0078] Since one of the advantages of the LILT-Base model is that it can integrate text content and layout information to achieve more accurate information extraction, the embedding layer of the shallow encoding module of the teacher model in this application, that is, the encoding part, can use a dual-link encoder. Specifically, on the one hand, it includes a text encoder that processes text content, and on the other hand, it also includes a layout encoder that processes layout information.

[0079] In order to preserve the temporal information carried by the text content in the document, it is also necessary to positionally encode the text content information and the text position information corresponding to the text content information to obtain a feature vector with temporal information. Therefore, the shallow encoding module in the teacher model also includes a position encoding layer.

[0080] In the process of model distillation, in order to ensure the extraction performance of the student model after distillation, the encoder of the student model in this application adopts the encoder chain form of the type teacher model. That is, the embedding layer of the student model also includes two encoders, and a corresponding position encoding layer is connected after each encoder; further, the network parameters of the teacher model are used to initialize the embedding layer and position encoding layer of the student model. In order to reduce the network parameters of the model, the number of network layers of the preceding layer and the position encoding layer can be reduced when initializing the student model to achieve model distillation.

[0081] Exemplarily, the embedding layer and position encoding layer in the teacher model are copied to realize multimodal features that can be extracted from the document; of course, in order to reduce the network parameters of the model, the copied embedding layer and position encoding layer can also be distilled to obtain embedding layers and position encoding layers with a preset number of layers; exemplarily, for the embedding layer that processes the text content, the text encoder contained in the embedding layer is copied and distilled to obtain an encoder with a Transformer encoding layer of 2 layers; for another example, for the embedding layer that processes the text position, the layout encoder contained in the embedding layer is copied and distilled to obtain an encoder with a Transformer encoding layer of 2 layers.

[0082] On this basis, the teacher model in the present application also includes a fusion module, that is, the fusion module is used to fuse the multimodal features output in the above-mentioned shallow coding module to obtain the fused fusion features. Exemplarily, in the present application, a self-attention coding network is used for the fusion module in the teacher model, and the process of distilling it can be a process of reducing the number of network layers. The specific implementation method of distilling the fusion module in the teacher model to obtain the fusion module in the student model can refer to the distillation process of the above-mentioned shallow coding module, which will not be described in detail here.

[0083] Furthermore, in order to obtain deeper features and increase the accuracy of subsequent extraction results, the teacher model in this application also includes a deep encoding module for deep feature extraction to obtain deep features. In the teacher model, a network layer based on the Transformer architecture is usually used for processing. However, the researchers of this application discovered during the research process that the bidirectional long short-term memory network (Bidirectional Long Short-Term Memory, BiLSTM) can also effectively capture long-term dependencies in time series data, and the network parameters of the double-layer bidirectional LSTM network are less than the network parameters based on the multi-layer Transformer encoder in the teacher model, which can reduce computing overhead, further reduce system resource consumption, and improve reasoning speed. Therefore, in this application, a double-layer bidirectional LSTM network is used to replace the original deep encoding module in the teacher model to achieve deep feature extraction.

[0084] It should be noted that since the long short-term memory network used in this application is a two-layer bidirectional long short-term memory network, the number of output feature dimensions of the network can be set to half of the number of output feature dimensions of the deep encoding module in the initial information extraction model, so that the output feature dimension of the original deep encoding module in the teacher model can be achieved.

[0085] Since the number of feature dimensions corresponds to the number of network parameter settings halved, the overall network parameters of the model are further reduced.

[0086] Since the deep encoding module of the student model in this application has changed, the decoding module connected to the deep encoding module is also encoded accordingly. Specifically, in the decoding module, a fully connected layer can be used to convert the output features of the double-layer bidirectional long short-term memory network into a feature vector with a dimension consistent with the number of preset extraction categories, and then the output probability is obtained through the Softmax activation function.

[0087] On this basis, the shallow encoding module, fusion module, deep encoding module and decoding module obtained in the above implementation are integrated to obtain the following: Figure 3 The information extraction model shown.

[0088] By distilling the teacher model and optimizing the network structure after distillation, the network parameters of the model can be reduced, and the system processing pressure and computing resource consumption can be reduced when the model performs extraction tasks.

[0089] On this basis, the constructed information extraction model needs to be trained using the target data set to obtain a lightweight model with information extraction capabilities.

[0090] Optionally, the training of the information extraction model in the present application includes: using a training data set to train the constructed information extraction model to obtain a trained information extraction model; the loss function used in the training includes information extraction loss and model distillation loss; wherein, the information extraction loss is based on the information extraction model, and the difference between the first predicted extraction result obtained by extracting information from any training text in the training data set and the label of the training text; the model distillation loss characterizes the difference between the second predicted extraction result obtained by extracting information from the training text by the initial information extraction model and the first predicted extraction result.

[0091] Specifically, the public EATEN ticket dataset is used as the target dataset to train the information extraction model. The target dataset contains 300,000 synthetic domestic train ticket data as the training set, and 400 real train tickets with private information hidden as the test set. The dataset is annotated with eight key information tags, namely: ticket number, departure station, arrival station, train number, fare, seat class, departure date and passenger name.

[0092] During the training process of the constructed information extraction model, the loss obtained by weighted summation of the information extraction loss and the model distillation loss is used as the total loss of the model, and multiple iterative training is performed based on the total loss until the iteration stops, and the trained information extraction model is obtained.

[0093] In this application, the information extraction loss represents the difference between the first predicted extraction result obtained by extracting information from any training text in the training data set based on the information extraction model and the label of the training text; illustratively, the difference can be obtained by the following expression:

[0094]

[0095] in, Represents information extraction loss; CrossEntropy (·,·) represents function expression; q s Represents the first prediction extraction result; y label Labels representing the training text.

[0096] Furthermore, the model distillation loss represents the difference between the second predicted extraction result obtained by the initial information extraction model for information extraction of the training text and the first predicted extraction result; illustratively, the difference can be obtained by the following expression:

[0097]

[0098] Among them, L KDrepresents the distillation loss of the model; T represents the distillation temperature (T>=1); CrossEntropy (·,·) represents the function expression; q s Characterizes the first prediction extraction result; q t Characterize the second prediction extraction result.

[0099] In some embodiments, q may be expressed in the form of a probability.

[0100] On this basis, the following expression is used to perform weighted summation on the above two loss functions to obtain the total loss of the model. Exemplarily, the expression for obtaining the total loss of the model may include:

[0101]

[0102] Among them, L represents the total loss of the model; Represents the ratio of two losses.

[0103] Furthermore, after the information extraction model is trained, it is deployed in a practical application system to perform information extraction tasks.

[0104] Optionally, the present application uses a trained information extraction model to extract information from each text information to obtain an extraction result of a visual rich document, including: the information type of any group of text information includes text content information and text position information corresponding to the text content information; for any group of text information, the text content information and the text position information are encoded and processed respectively to obtain corresponding text content vectors and text position vectors; the text content vector and the text position vector are fused to obtain a multimodal fusion vector; deep feature encoding is performed on the multimodal fusion vector to obtain a multimodal deep feature vector; information extraction is performed on the multimodal deep feature vector to obtain an extraction result corresponding to the visual rich document.

[0105] Specifically, the process of encoding text content information and text position information can be understood as the process of converting them into low-dimensional vectors. Specifically, the process of converting text content information into low-dimensional vectors may include: performing word segmentation on the text content information to obtain multiple phrases; for any phrase, performing low-dimensional vector conversion on the phrase to obtain a corresponding word vector, and performing position encoding on the word vector to obtain a corresponding word position vector; according to the word vector and word position vector corresponding to each phrase in each text information, respectively, obtaining the text word vector corresponding to each text information.

[0106] In the present application, in order to facilitate the low-dimensional vector conversion processing of the text content, a word segmenter can be used to first segment the text content information to obtain multiple phrases.

[0107] Since the text content information will be calculated in matrix form after low-dimensional vector conversion, but different groups of text content information have different word lengths after word segmentation, in order to facilitate subsequent calculations, for text content information with shorter length, a padding marker can be added after the word group after word segmentation.

[0108] When converting the phrases into low-dimensional vectors, a learnable text embedding matrix can be used to convert the phrases after word segmentation into low-dimensional vectors to obtain the word vectors corresponding to each phrase.

[0109] In order to comprehensively consider the spatial position and global information of the text content information, the word vector of each phrase can be positionally encoded to obtain the word position vector.

[0110] On this basis, the word vector and word position vector obtained above can be fused to obtain a text content vector corresponding to the text content information.

[0111] The process of converting text position information into a low-dimensional vector in the present application may specifically include: determining the word position information corresponding to each phrase in at least one group of text information based on each text position information; performing low-dimensional vector conversion processing on each word position information to obtain a text position vector corresponding to each text information.

[0112] In this application, in order to integrate more information, when encoding the text position information, in addition to encoding the position coordinates of the text content, the area size of the text area corresponding to the text content and the area shape are also encoded.

[0113] Specifically, before processing, the text position coordinates corresponding to the text content will be determined according to the results of optical character recognition, and then the area size, that is, the height h and width w of the area, will be determined according to the text position coordinates. On this basis, the shape of the area can be represented by the distance from the edge of the area to the center of mass of the area. For example, the horizontal coordinate distance from the upper left corner of the text area to the center of mass of the text block, the vertical coordinate distance from the upper left corner of the text area to the center of mass of the text block, the horizontal coordinate distance from the lower right corner of the text area to the center of mass of the text block, and the vertical coordinate distance from the lower corner of the text area to the center of mass of the text block are calculated.

[0114] The above text position information is converted into a low-dimensional vector to obtain a corresponding text position vector.

[0115] It can be explained that, since the text position information itself carries timing information, in some optional implementations, the vector obtained after encoding the text position information is no longer position-encoded, and the result is directly used as the text position vector.

[0116] On the basis of the above implementation, an encoder can be used to encode each vector separately to obtain the self-attention vector corresponding to each vector, and then the respective attention vectors are fused to obtain a multimodal fusion vector.

[0117] In an optional implementation, a multimodal Transformer self-attention mechanism may be used in the encoder to encode each vector to obtain a multimodal fusion vector.

[0118] Furthermore, in order to obtain deeper feature information, a deep module is used to perform deep feature encoding on the multimodal fusion vector to obtain a multimodal deep feature vector.

[0119] At this point, the encoding process of the visual rich document has been completed, and the corresponding multimodal deep feature vector has been obtained. The fusion vector of multiple modalities helps to extract the key features in the text, making it easier to obtain more accurate extraction results in the subsequent information extraction process.

[0120] In this application, the process of information extraction can also be understood as the process of decoding the obtained feature vector to obtain the final probability value. Specifically, the obtained multimodal deep feature vector is input into the LSTM network for feature processing, and the vector feature output by the network is input into the fully connected layer for dimensional conversion, and then the output probability matrix after SoftMax. It should be noted that the probability matrix is ​​the probability value corresponding to the preset extraction labels. Finally, the probability value that meets the conditions in the probability matrix, such as the maximum probability value, or the extraction label corresponding to the probability value of the top preset position is selected as the extraction result corresponding to the currently processed visual rich document.

[0121] Figure 4 A schematic diagram of the structure of a visual rich document information extraction device provided in an embodiment of the present application. Figure 4 The visual rich document information extraction device 40 comprises: a file information acquisition module 401 and an extraction result acquisition module 402; wherein,

[0122] The file information acquisition module 401 is used to acquire a visually rich document to be processed and determine at least one set of text information contained in the visually rich document;

[0123] The extraction result acquisition module 402 is used to use the trained information extraction model to extract information from each of the text information to obtain the extraction result of the visual rich document; wherein the information extraction model includes a first network structure in a model obtained by distilling the trained initial information extraction model, and a preset neural network structure obtained by replacing the second network structure in the model obtained by distilling the trained initial information extraction model.

[0124] In an optional embodiment, the device further includes: a model building module, specifically including:

[0125] A first distillation submodule, used for performing distillation processing on the model parameters of the initial information extraction model to obtain a shallow encoding module and a fusion module;

[0126] A second distillation submodule, used to determine a deep encoding module and a decoding module based on the preset neural network structure;

[0127] The model construction submodule is used to obtain the information extraction model according to the shallow encoding module, the fusion module, the deep encoding module and the decoding module.

[0128] In an optional implementation, the shallow encoding module in the initial information extraction model includes an embedding layer and a position encoding layer; the fusion module in the initial information extraction model is a self-attention encoding network;

[0129] The first module is a distillation submodule, including:

[0130] A shallow coding module distillation unit, used for performing distillation processing on the embedding layer and the position coding layer to obtain the shallow coding module;

[0131] The fusion module distillation unit is used to perform distillation processing on the self-attention encoding network to obtain the fusion module.

[0132] In an optional implementation, the deep encoding module is a two-layer bidirectional long short-term memory network; the number of output feature dimensions of the two-layer bidirectional long short-term memory network is half the number of output feature dimensions of the deep encoding module in the initial information extraction model.

[0133] In an optional implementation, the device further includes: a model training module, specifically including:

[0134] The model training unit is used to train the constructed information extraction model using the training data set to obtain a trained information extraction model; the loss function used in the training includes information extraction loss and model distillation loss; wherein,

[0135] The information extraction loss represents the difference between a first predicted extraction result obtained by performing information extraction on any training text in the training data set based on the information extraction model and the label of the training text;

[0136] The model distillation loss represents the difference between a second predicted extraction result obtained by extracting information from the training text by the initial information extraction model and the first predicted extraction result.

[0137] In an optional implementation manner, any set of text information includes text content information and text position information corresponding to the text content information;

[0138] The extraction result obtaining module 402 includes:

[0139] A text vector obtaining unit, for encoding the text content information and the text position information for any set of text information, to obtain the corresponding text content vector and text position vector respectively;

[0140] A multimodal fusion vector obtaining unit, used for fusing the text content vector and the text position vector to obtain a multimodal fusion vector;

[0141] A deep feature vector obtaining unit, used for performing deep feature encoding on the multimodal fusion vector to obtain a multimodal deep feature vector;

[0142] The extraction result obtaining unit is used to perform information extraction processing on the multimodal deep feature vector to obtain the extraction result corresponding to the visual rich document.

[0143] Figure 5 is a block diagram of an electronic device shown in an embodiment of the present application, and the device may be a computer, a digital broadcast terminal, etc. Figure 5 , the device 800 may include one or more of the following components: a processing component 802 , a memory 804 , a power component 806 , a multimedia component 808 , an audio component 810 , an input / output interface 812 , a sensor component 814 , and a communication component 816 .

[0144] The processing component 802 generally controls the overall operation of the device 800, such as operations associated with display, phone calls, data communications, camera operations, and recording operations. The processing component 802 may include one or more processors 820 to execute instructions to complete all or part of the steps of the above-mentioned method. In addition, the processing component 802 may include one or more modules to facilitate the interaction between the processing component 802 and other components. For example, the processing component 802 may include a multimedia module to facilitate the interaction between the multimedia component 808 and the processing component 802.

[0145] The memory 804 is configured to store various types of data to support operations on the device 800. Examples of such data include instructions for any application or method operating on the device 800, contact data, phone book data, messages, pictures, videos, etc. The memory 804 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0146] The power supply component 806 provides power to the various components of the device 800. The power supply component 806 can include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power for the device 800.

[0147] The multimedia component 808 includes a screen that provides an output interface between the device 800 and the user. In some embodiments, the screen may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor may not only sense the boundaries of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation. In some embodiments, the multimedia component 808 includes a front camera and / or a rear camera. When the device 800 is in an operating mode, such as a shooting mode or a video mode, the front camera and / or the rear camera may receive external multimedia data. Each front camera and rear camera may be a fixed optical lens system or have a focal length and optical zoom capability.

[0148] The audio component 810 is configured to output and / or input audio signals. For example, the audio component 810 includes a microphone (MIC), and when the device 800 is in an operating mode, such as a call mode, a recording mode, and a speech recognition mode, the microphone is configured to receive an external audio signal. The received audio signal can be further stored in the memory 804 or sent via the communication component 816. In some embodiments, the audio component 810 also includes a speaker for outputting audio signals.

[0149] The input / output interface 812 provides an interface between the processing component 802 and the peripheral interface modules, which may be keyboards, click wheels, buttons, etc. These buttons may include but are not limited to: a home button, a volume button, a start button, and a lock button.

[0150] The sensor assembly 814 includes one or more sensors for providing various aspects of status assessment for the device 800. For example, the sensor assembly 814 can detect the open / closed state of the device 800, the relative positioning of components, such as the display and keypad of the device 800, and the sensor assembly 814 can also detect the position change of the device 800 or a component of the device 800, the presence or absence of user contact with the device 800, the orientation or acceleration / deceleration of the device 800, and the temperature change of the device 800. The sensor assembly 814 may include a proximity sensor configured to detect the presence of nearby objects without any physical contact. The sensor assembly 814 may also include an optical sensor, such as a solid image (Complementary Metal Oxide Semiconductor, CMOS) sensor or a semiconductor image (Charge-coupled Device, CCD) sensor, for use in imaging applications. In some embodiments, the sensor assembly 814 may also include an acceleration sensor, a gyroscope sensor, a magnetic sensor, a pressure sensor, or a temperature sensor.

[0151] The communication component 816 is configured to facilitate wired or wireless communication between the device 800 and other devices. The device 800 can access a wireless network based on a communication standard, such as WiFi, 4G or 5G, or a combination thereof. In an exemplary embodiment, the communication component 816 receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component 816 also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wideband (UWB) technology, Bluetooth (BT) technology and other technologies.

[0152] In an exemplary embodiment, the device 800 may be implemented by one or more application specific integrated circuits (ASIC), digital signal processors (DSP), digital signal processing devices (DSPD), programmable logic devices (PLD), field programmable gate arrays (FPGA), controllers, microcontrollers, microprocessors or other electronic components to perform the above method.

[0153] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 804 including instructions, and the instructions can be executed by a processor 820 of the device 800 to perform the above method. For example, the non-transitory computer-readable storage medium can be a ROM, a random access memory (RAM), a CD-ROM, a magnetic tape, a floppy disk, an optical data storage device, etc.

[0154] A non-transitory computer-readable storage medium, when the instructions in the storage medium are executed by a processor of a server, enables the server to execute the above-mentioned physical backup method of the database.

[0155] An embodiment of the present application also provides a chip for running instructions, which is used to execute the technical solution of the physical backup method of the database in the above embodiment.

[0156] An embodiment of the present application further provides a computer-readable storage medium, in which computer execution instructions are stored. When the computer execution instructions are executed on a computer, the computer executes the technical solution of the physical backup method of the database in the above embodiment.

[0157] An embodiment of the present application also provides a computer program product, which includes a computer program stored in a computer-readable storage medium. At least one processor can read the computer program from the computer-readable storage medium. When at least one processor executes the computer program, the technical solution of the physical backup method of the database in the above embodiment can be implemented.

[0158] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.

[0159] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

[0160] Those skilled in the art will readily appreciate other embodiments of the present application after considering the specification and practicing the invention disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The specification and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the following claims.

[0161] It should be understood that the present application is not limited to the precise structures that have been described above and shown in the drawings, and that various modifications and changes may be made without departing from the scope thereof. The scope of the present application is limited only by the appended claims.

Claims

1. A method for extracting information from visually rich documents, characterized in that: The method comprises: Acquire a visually rich document to be processed, and determine at least one set of text information contained in the visually rich document; A trained information extraction model is used to extract information from each of the text information to obtain an extraction result of the visual rich document; wherein the information extraction model includes a first network structure in a model obtained by distilling the trained initial information extraction model, and a preset neural network structure obtained by replacing a second network structure in a model obtained by distilling the trained initial information extraction model.

2. The method according to claim 1, characterized in that The construction of the information extraction model includes: Performing distillation processing on the model parameters of the initial information extraction model to obtain a shallow encoding module and a fusion module; Determine a deep encoding module and a decoding module based on the preset neural network structure; The information extraction model is obtained according to the shallow encoding module, the fusion module, the deep encoding module and the decoding module.

3. The method according to claim 2, characterized in that The shallow encoding module in the initial information extraction model includes an embedding layer and a position encoding layer; the fusion module in the initial information extraction model is a self-attention encoding network; The model parameters of the initial information extraction model are distilled to obtain the shallow encoding module and the fusion module, including: Performing distillation processing on the embedding layer and the position encoding layer to obtain the shallow encoding module; The self-attention encoding network is distilled to obtain the fusion module.

4. The method according to claim 2, characterized in that: The deep encoding module is a double-layer bidirectional long short-term memory network; the number of output feature dimensions of the double-layer bidirectional long short-term memory network is half the number of output feature dimensions of the deep encoding module in the initial information extraction model.

5. The method according to any one of claims 1 to 4, characterized in that: The training of the information extraction model includes: The constructed information extraction model is trained using the training data set to obtain a trained information extraction model; the loss function used in the training includes information extraction loss and model distillation loss; among them, The information extraction loss represents the difference between a first predicted extraction result obtained by performing information extraction on any training text in the training data set based on the information extraction model and the label of the training text; The model distillation loss represents the difference between a second predicted extraction result obtained by extracting information from the training text by the initial information extraction model and the first predicted extraction result.

6. The method according to any one of claims 1 to 4, characterized in that: Any set of text information includes text content information and text position information corresponding to the text content information; The trained information extraction model is used to extract information from each of the text information to obtain the extraction results of the visually rich document, including: For any set of text information, the text content information and the text position information are respectively encoded to obtain the corresponding text content vector and text position vector; Fusing the text content vector and the text position vector to obtain a multimodal fusion vector; Performing deep feature encoding on the multimodal fusion vector to obtain a multimodal deep feature vector; Information extraction is performed on the multimodal deep feature vector to obtain an extraction result corresponding to the visual rich document.

7. A visual rich document information extraction device, characterized in that: The device comprises: A file information acquisition module, used to acquire a visually rich document to be processed and determine at least one set of text information contained in the visually rich document; An extraction result acquisition module is used to use a trained information extraction model to extract information from each of the text information to obtain an extraction result of the visual rich document; wherein the information extraction model includes a first network structure in a model obtained by distilling the trained initial information extraction model, and a preset neural network structure obtained by replacing the second network structure in a model obtained by distilling the trained initial information extraction model.

8. An electronic device, characterized in that: include: A processor and a memory communicatively connected to the processor; The memory stores computer-executable instructions; When executing the computer-executable instructions, the processor is used to implement the visual rich document information extraction method as described in any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the visual rich document information extraction method as described in any one of claims 1 to 6.

10. A computer program product, characterized in that The method comprises a computer program, which implements the method according to any one of claims 1 to 6 when the computer program is executed by a processor.