Data processing method and device and visual question and answer model training method and device
By replacing the text in the initial sample document image with reading comprehension sample document text in the visual question-and-answer model, the target sample document image is solved, and the problem of requiring a large amount of sample data in the prior art is achieved, and data processing with lower cost and higher accuracy is achieved.
Patent Information
- Application Number
- CN202311814863.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-26
- Publication Date
- 2025-06-27
AI Technical Summary
Existing visual question and answer models require a large amount of sample data to ensure high accuracy when processing material image data, but this increases cost and time consuming and is poor generalization.
By determining the sample prompt text, initial sample document image, and reading comprehension sample text, the text in the initial sample document image is replaced to generate the target sample document image, thereby amplifying the sample data and model training is performed based on only a small number of initial sample document images.
It reduces the cost and time-consuming of data processing, improves the generalization of the model and the accuracy of data processing.
Smart Images

Figure CN120216712A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technology, and particularly to a data processing method. One or more embodiments of this specification simultaneously relate to two visual question answering model training methods, a data processing device, a visual question answering model training device, a computing device, a computer-readable storage medium, and a computer program. Background Art
[0002] Currently, there is a need to understand material image data and answer questions related to the material image data according to user questions in many industries. For example, in the government affairs industry, a lot of material image data needs to be processed. Manual processing of material data is time-consuming, laborious, and inefficient. Therefore, people often train visual question answering models to achieve automated processing of material image data, such as the Ernie-layout model (cross-modal document understanding model) / Layoutlmv3 model (document intelligent multi-modal pre-training model).
[0003] However, in the process of automatically processing material image data by the above models, if high accuracy is required, a large amount of sample data needs to be obtained during model training. The large amount of sample data brings high costs, high time consumption, and poor generalization for material image data outside the sample data. If only a small amount of sample data is used during model training, the accuracy of the model data processing obtained by training will be greatly reduced. Therefore, a technical solution is urgently needed to solve the above problems. Summary of the Invention
[0004] In view of this, the embodiments of this specification provide a data processing method. One or more embodiments of this specification simultaneously relate to two visual question answering model training methods, a data processing device, a visual question answering model training device, a computing device, a computer-readable storage medium, and a computer program to solve the technical defects existing in the prior art.
[0005] According to the first aspect of the embodiments of this specification, a data processing method is provided, including:
[0006] Determine a target prompt text and a target document image corresponding to the target prompt text;
[0007] Input the target prompt text and the target document image into a visual question answering model to obtain a text answer corresponding to the target prompt text determined from the target document image,
[0008] wherein the visual question answering model is trained by sample prompt texts, sample document images, and sample text answers corresponding to the sample prompt texts determined from the sample document images.
[0009] According to the second aspect of the embodiments of the present specification, a method for training a visual question answering model is provided, including:
[0010] Determine a sample prompt text, an initial sample document image, and a reading comprehension sample text;
[0011] Replace the text in the initial sample document image with the reading comprehension sample text to obtain a target sample document image;
[0012] Determine the sample document image according to the initial sample document image and the target sample document image;
[0013] Train an initial visual question answering model according to the sample prompt text, the sample document image, and the sample text answer corresponding to the sample prompt text determined from the sample document image until a training stop condition is met, and obtain the visual question answering model.
[0014] According to the third aspect of the embodiments of the present specification, a data processing device is provided, including:
[0015] A determination module configured to determine a target prompt text and a target document image corresponding to the target prompt text;
[0016] A processing module configured to input the target prompt text and the target document image into a visual question answering model to obtain a text answer corresponding to the target prompt text determined from the target document image, where the visual question answering model is trained by a sample prompt text, a sample document image, and a sample text answer corresponding to the sample prompt text determined from the sample document image, and the sample document image includes an initial sample document image and a target sample document image amplified according to the initial sample document image and the reading comprehension sample text.
[0017] According to the fourth aspect of the embodiments of the present specification, a visual question answering model training device is provided, including:
[0018] A sample determination module configured to determine a sample prompt text, an initial sample document image, and a reading comprehension sample text;
[0019] A target sample document image obtaining module configured to replace the text in the initial sample document image with the reading comprehension sample text to obtain a target sample document image;
[0020] A sample document image determination module configured to determine the sample document image according to the initial sample document image and the target sample document image;
[0021] A visual question answering model obtaining module, configured to train an initial visual question answering model according to the sample prompt text, the sample document image, and the sample text answer corresponding to the sample prompt text determined from the sample document image until a training stop condition is met, and obtain the visual question answering model.
[0022] According to a fifth aspect of the embodiments of the present specification, there is provided a method for training a visual question answering model, which is applied to the cloud and includes:
[0023] Receiving a sample prompt text, an initial sample document image, and a reading comprehension sample text sent by the front end;
[0024] Replacing the text in the initial sample document image with the reading comprehension sample text to obtain a target sample document image;
[0025] Determining the sample document image according to the initial sample document image and the target sample document image;
[0026] Training an initial visual question answering model according to the sample prompt text, the sample document image, and the sample text answer corresponding to the sample prompt text determined from the sample document image until a training stop condition is met, obtaining the visual question answering model, and sending the visual question answering model to the front end.
[0027] According to a sixth aspect of the embodiments of the present specification, there is provided a computing device, including:
[0028] A memory and a processor;
[0029] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the above data processing method, visual question answering model training method, or visual question answering model training method applied to the cloud are implemented.
[0030] According to a seventh aspect of the embodiments of the present specification, there is provided a computer-readable storage medium, which stores computer-executable instructions. When the instructions are executed by a processor, the steps of the above data processing method, visual question answering model training method, or visual question answering model training method applied to the cloud are implemented.
[0031] According to an eighth aspect of the embodiments of the present specification, there is provided a computer program, wherein when the computer program is executed on a computer, the computer is made to execute the steps of the above data processing method, visual question answering model training method, or visual question answering model training method applied to the cloud.
[0032] An embodiment of this specification provides a data processing method, including: determining a target prompt text and a target document image corresponding to the target prompt text; inputting the target prompt text and the target document image into a visual question answering model to obtain a text answer corresponding to the target prompt text determined from the target document image, where the visual question answering model is trained by sample prompt texts, sample document images, and sample text answers corresponding to the sample prompt texts determined from the sample document images.
[0033] Specifically, this method inputs the target prompt text and the target document image into the visual question answering model to obtain a text answer corresponding to the target prompt text determined from the target document image, realizing the understanding of the target document image and answering the answer corresponding to the target prompt text from the target document; moreover, the training of this visual question answering model is achieved by training with sample prompt texts, sample document images, and sample text answers corresponding to the sample prompt texts determined from the sample document images. Considering the training of sample prompt texts, the accuracy of the visual question answering model in recognizing sample prompt texts is improved to enhance the accuracy of subsequent data processing. Additionally, when training this visual question answering model, it is only based on a limited number of initial sample document images, reducing the cost of sample data. In summary, this method achieves the technical effects of lower data processing time consumption, lower cost, and improved data processing accuracy. Brief Description of the Drawings
[0034] Figure 1 is a flowchart of a method for training a visual question answering model provided by an embodiment of this specification;
[0035] Figure 2 is a flowchart of a data processing method provided by an embodiment of this specification;
[0036] Figure 3 is a flowchart of another method for training a visual question answering model provided by an embodiment of this specification;
[0037] Figure 4 is a schematic structural diagram of a visual question answering model provided by an embodiment of this specification;
[0038] Figure 5 is a schematic structural diagram of a device for training a visual question answering model provided by an embodiment of this specification;
[0039] Figure 6 is a schematic structural diagram of a data processing device provided by an embodiment of this specification;
[0040] Figure 7 is a structural block diagram of a computing device provided by an embodiment of this specification. Detailed implementation manners
[0041] In the following description, numerous specific details are set forth in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.
[0042] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and encompasses any and all possible combinations of one or more of the associated listed items.
[0043] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first may also be referred to as the second, and similarly, the second may also be referred to as the first. Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining".
[0044] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data that have been authorized by the user or fully authorized by all parties, and the collection, use, and processing of the relevant data need to comply with the relevant laws, regulations, and standards of the relevant countries and regions, and corresponding operation entrances are provided for the user to choose to authorize or refuse.
[0045] First, the noun terms involved in one or more embodiments of this specification are explained.
[0046] seq2seq model: Sequence to Sequence model, a sequence-to-sequence model.
[0047] T5: Transformers-based Text-to-Text Transfer Transformer, a sequence-to-sequence model based on the Transformer architecture.
[0048] BART: Bidirectional and Auto-Regressive Transformers, a sequence-to-sequence model for generative pre-training.
[0049] OCR: Optical Character Recognition. It refers to the process in which electronic devices (such as scanners or digital cameras) check the characters printed on paper, determine their shapes by detecting dark and bright patterns, and then translate the shapes into computer text using character recognition methods. When using OCR to recognize text, the recognized text is often first annotated with text boxes, and the text shapes within each text box are translated into computer text.
[0050] Ernie-layout: A cross-modal document understanding model.
[0051] Layoutlmv3: A document intelligent multi-modal pre-training model.
[0052] Paddle OCR: A text recognition model suite that realizes the function of recognizing text in images by integrating a three-stage model: text box detection - angle classification - text recognition. Using Paddle OCR for text recognition is one of the OCR recognition algorithms.
[0053] nlp: Natural Language Processing. NLP, natural language processing, uses computers to analyze and generate natural languages (text, speech) with the aim of enabling humans to interact with computer systems in natural language forms, thereby facilitating and effectively managing information.
[0054] CNN: Convolutional Neural Networks, a class of feedforward neural networks that contain convolutional computations and have a deep structure.
[0055] Currently, in the government affairs industry, there is a need to process a large amount of material data. The material visual question-answering technical solution that automatically understands material image data and answers questions related to the material image according to the user's intention has great value for the government affairs industry.
[0056] There are some difficulties in material visual question answering. Firstly, the material content is not only related to the text above, but also related to the layout, images, and table information therein, which requires the model to understand multi-modal information. Secondly, when answering questions, the model also needs to understand the meaning and intention of the questions, which involves semantic parsing and understanding of the questions, as well as reasoning about the relationship between the images and the questions. Finally, the annotation and acquisition of the dataset are also a difficulty, and it is necessary to collect data with rich material features and diverse questions and perform accurate annotation for model training and evaluation.
[0057] To address the above difficulties and challenges, the current technical solution is to use an encoder-based visual question answering model, such as the Ernie-layout model (a cross-modal document understanding model) and the Layoutlmv3 model (a document intelligent multi-modal pre-training model). Through the encoder-based visual question answering model framework, they have achieved the understanding of multi-modal information in the material content. Secondly, when using the encoder-based visual question answering model to answer questions, they first understand an input question and, for this question, associate and infer the relationship between the information in the image and this question to answer the question. Finally, they obtain a large number of annotated sample data to enrich and accurately annotate the sample data, thereby ensuring the accuracy of model training and evaluation.
[0058] However, in actual application scenarios, since the above models can only answer one question at a time, when the number of questions to be answered is greater than 1 or increases, the total time taken for the model to answer questions will increase, and the greater the number of questions answered, the higher the total time taken. Moreover, the training of the above models requires obtaining a large number of annotated sample data, which will also lead to a high cost of model training. The greater the number of annotated sample data obtained, the higher the cost.
[0059] In view of this, in this specification, a data processing method is provided. One or more embodiments of this specification are simultaneously related to two visual question answering model training methods, a data processing device, a visual question answering model training device, a computing device, a computer-readable storage medium, and a computer program to solve the above technical problems, which will be described in detail one by one in the following embodiments.
[0060] See Figure 1 , Figure 1 is a flowchart of a visual question answering model training method provided by an embodiment of this specification, which specifically includes the following steps.
[0061] Step 102: Determine the sample prompt text, the initial sample document image, and the reading comprehension sample text.
[0062] Among them, the sample prompt text can be understood as the guiding text input into the model. For example, in the scenario of reviewing the materials of members applying to go abroad, the sample prompt text can be understood as "Why do the people applying to go abroad want to go abroad?", "Should the people applying to go abroad be allowed to go abroad? Why?", etc. Another example is that in the scenario of financial material analysis, the sample prompt text can be understood as "What is the current situation of the stock market?", "Can you tell us what the future stock market will look like based on the current situation?", etc.
[0063] The initial sample document image can be understood as a pre-collected sample document image, and this sample document image can come from multiple pre-collected documents with different styles and texts (such as picture files, pdf files, etc.).
[0064] The reading comprehension sample text can be understood as the text containing the text and the questions related to this text or the text for interpreting the text. For example, this text can be understood as an article, an ancient poem or a sentence, etc., and the questions or interpretations related to the text can be understood as answering questions related to the text, summarizing the main content of the text, the annotations of the text, etc.
[0065] In practical applications, in order to reduce the cost of obtaining samples, the initial sample document image can be a small number of images (for example, 5000 images), and only relying on this small number of initial sample document images for model training, the effect of the trained model is poor. For example, it cannot accurately understand the document image and cannot obtain the semantics of the document image, etc.
[0066] Therefore, the embodiments of this specification adopt the method of amplifying samples to improve the accuracy of model training without increasing the cost of obtaining samples.
[0067] Step 104: Replace the text in the initial sample document image with the reading comprehension sample text to obtain the target sample document image.
[0068] Among them, the text in the initial sample document image can be understood as the original and unprocessed text in the initial sample document image.
[0069] Then, replacing the text in the initial sample document image with the reading comprehension sample text to obtain the target sample document image can be understood as replacing the text in the initial sample document image with the reading comprehension sample text, thereby generating a new sample document image with the same text layout as the text in the initial sample document image but different text content, that is, the target sample document image.
[0070] In practical applications, in order to obtain richer target sample images, the initial sample document image may include multiple document images that contain text and have different text layouts, and use the reading comprehension sample text to replace the text in the initial sample document image to generate multiple target sample document images with different text layouts. The specific implementation method is as follows:
[0071] The initial sample document image includes multiple document images that contain text and have different text layouts;
[0072] Correspondingly, obtaining the target sample document image by replacing the text in the initial sample document image with the reading comprehension sample text includes:
[0073] By using a character recognition algorithm, recognize the sample document text in each initial sample document image, and perform text box annotation on the sample document text;
[0074] Replace the sample document text marked with the text box in each initial sample document image with the reading comprehension sample text to obtain the target sample document image.
[0075] Among them, the text layout can be understood as the way text is arranged in a document image. The text arrangement can be understood as the arrangement order of each paragraph of text, each character, each sentence, etc. relative to each other. For example, the text layout can be understood as a landscape layout, a portrait layout, a grid layout, a free layout, etc.
[0076] The character recognition algorithm can be understood as an OCR (Optical Character Recognition) recognition algorithm. In practical applications, there are various OCR recognition algorithms, which are not limited in the embodiments of this specification. For example, in the embodiments of this specification, the character recognition algorithm can be exemplified by the PaddleOCR (a text recognition model suite) algorithm in the OCR recognition algorithm as the OCR recognition algorithm.
[0077] The text box annotation can be understood as an operation of adding identifiers such as borders and border lines around the text.
[0078] On this basis, by using a character recognition algorithm to recognize the sample document text in each initial sample document image and perform text box annotation on the sample document text, it can be understood that the PaddleOCR algorithm is first used to recognize each initial sample document image, and the text contained in each initial sample document image, that is, the sample document text in each initial sample document image, is recognized. After recognizing the sample document text in each initial sample document image, text boxes are marked around each text in the sample document text, including but not limited to performing text box annotation on each sentence, each paragraph, each individual word, etc.
[0079] The text of the sample document marked by the text box can be understood as the text within each text box after the above-mentioned text box marking of the text in the initial sample document image.
[0080] Then, by understanding and reading the sample text and replacing the text of the sample document marked by the text box in each initial sample document image, obtaining the target sample document image can be understood as replacing the text of the sample document marked by the text box in each initial sample document image with the text content of the understood and read sample text, resulting in a new sample document image, which is the target sample document image.
[0081] For example, the text of the sample document included in an initial sample document image is "Name: Wang XX; Expected departure time: 2021-01-01", and the text content included in an understood and read sample text is: "Ancient poem: XX; Author: XXX". By using this understood and read sample text to replace the text of the sample document marked by the text box in this initial sample document image, a target sample image containing "Ancient poem: XX; Author XXX" is obtained.
[0082] It should be noted that the embodiments of this specification do not limit the number of understood and read sample texts either. That is to say, in order to obtain more target sample images, the above steps can also use multiple different understood and read sample texts to replace the texts in multiple initial sample document images to generate multiple target sample document images.
[0083] For example, for two initial sample document images with different layouts and three different understood and read sample texts, at least 2 * 3 = 6 target sample document images can be generated, thus making the target sample images richer.
[0084] The visual question answering model training method provided by the embodiments of this specification replaces the text of the sample document in each initial sample document image through the understood and read sample text, expanding the scale of the document image, enabling the training of this visual question answering model to be completed with only a small number of initial sample images and achieving better model training effects. Moreover, due to the replacement using the understood and read sample text, the generalization of this visual question answering model to new document images is better and the inference latency is lower. Furthermore, during the replacement process, a character recognition algorithm is used for text recognition, increasing the accuracy of text recognition, thereby ensuring the accuracy of the replacement.
[0085] Step 106: Determine the sample document image according to the initial sample document image and the target sample document image.
[0086] Specifically, after obtaining the target sample document image, since the target sample document image is generated according to an algorithm and its authenticity relative to the real initial sample document image is reduced, in order to ensure the accuracy of model training, the sample document image for actual training can be obtained by combining the initial sample document image and the target sample document image to reduce the deviation in model training.
[0087] For example, if the initial sample document images are multiple images containing the registration information text of people going abroad, and the target sample document images are multiple images containing reading comprehension text generated through the above steps, then the sample document images determined according to the initial sample document images and the target sample document images can be understood as multiple images containing the registration information text of people going abroad and multiple images containing reading comprehension text.
[0088] Step 108: Train the initial visual question answering model according to the sample prompt text, the sample document image, and the sample text answer corresponding to the sample prompt text determined from the sample document image until the training stop condition is met, and obtain the visual question answering model.
[0089] Among them, the initial visual question answering model can be understood as a pre-constructed visual question answering model. For example, the initial visual question answering model can be pre-constructed through the model parameters of other existing machine learning models so that the initial visual question answering model has the ability to understand the samples for model training. Other existing visual question answering models include, but are not limited to, the Ernie-layout model (a cross-modal document understanding model), the T5 model (Transformers-based Text-to-Text Transfer Transformer, a sequence-to-sequence model based on the Transformer structure), etc.; for example, the initial visual question answering model can be understood as a visual question answering model that includes an encoding layer constructed by loading the pre-trained parameters of the Ernie-layout model and a decoding layer constructed by loading the pre-trained parameters of the T5 model.
[0090] The training stop condition can be understood as the condition for the visual question answering model to stop training. For example, the training stop condition can be understood as including, but not limited to, reaching a preset number of training iterations, such as 200 times, and the difference between the model output result and the sample label (that is, the sample text answer corresponding to the sample prompt text determined from the sample document image) being lower than a preset threshold, etc. Specifically in implementation, the training stop condition can be set according to actual requirements, actual goals, etc.
[0091] In practical applications, training the initial visual question answering model can be achieved through the encoding network and decoding network of the initial visual question answering model. The specific implementation method is as follows:
[0092] The initial visual question answering model includes an encoding network and a decoding network;
[0093] Correspondingly, training the initial visual question answering model according to the sample prompt text, the sample document image, and the sample text answer corresponding to the sample prompt text determined from the sample document image includes:
[0094] In the encoding network, according to the sample prompt text, determine the prompt text encoding vector corresponding to the sample prompt text;
[0095] According to the sample document image, determine the sample document encoding vector and the sample image encoding vector corresponding to the sample document image;
[0096] According to the prompt text encoding vector, the sample document encoding vector, and the sample image encoding vector, determine the semantic feature vector of the sample document image;
[0097] In the decoding network, decode the semantic feature vector to obtain the predicted text answer corresponding to the sample prompt text determined from the sample document image;
[0098] Train the initial visual question answering model according to the sample text answer corresponding to the sample prompt text determined from the sample document image and the predicted text answer.
[0099] Among them, the encoding network can be understood as the network layer in the initial visual question answering model for encoding the data input to this initial visual question answering model (such as the above sample prompt text, sample document image, etc.); the decoding network can be understood as the network layer in the initial visual question answering model for decoding the encoding result output by the encoding network.
[0100] On this basis, determining the prompt text encoding vector corresponding to the sample prompt text according to the sample prompt text can be understood as encoding the sample prompt text in the encoding network and obtaining the prompt text encoding vector corresponding to the encoded sample prompt text.
[0101] In practical applications, the specific implementation method for encoding the sample prompt text to determine the prompt text encoding vector corresponding to the sample prompt text is as follows:
[0102] Perform word encoding on the sample prompt text to determine the word encoding vector of the sample prompt text;
[0103] Perform segment encoding on the sample prompt text to determine the segment encoding vector of the sample prompt text;
[0104] Perform one-dimensional positional encoding on the sample prompt text to determine the one-dimensional positional encoding vector of the sample prompt text;
[0105] According to the word encoding vector, the segment encoding vector, and the one-dimensional positional encoding vector, determine the prompt text encoding vector corresponding to the sample prompt text.
[0106] Among them, performing word encoding on the sample prompt text can be understood as an operation of encoding the characters or words included in the sample prompt text. Performing segment encoding on the sample prompt text can be understood as an operation of performing segment encoding on the paragraphs included in the sample prompt text. Performing one-dimensional positional encoding on the sample prompt text can be understood as an operation of encoding the sample prompt text according to the positions of each character or word in the sentence where it is located.
[0107] On this basis, perform word encoding, segment encoding, and one-dimensional positional encoding on the sample prompt text respectively, and obtain the encoded word encoding vector, segment encoding vector, and one-dimensional positional encoding vector respectively. Then, combine the obtained word encoding vector, segment encoding vector, and one-dimensional positional encoding vector to obtain the prompt text encoding vector corresponding to the combined sample prompt text. It should be noted that the specific method of vector combination in the embodiments of this specification can be determined according to the actual situation, including but not limited to sequential connection, vector splicing, etc.
[0108] The visual question answering model training method provided in the embodiments of this specification performs word encoding, segment encoding, and one-dimensional positional encoding on the sample prompt text, and combines the word encoding, the segment encoding, and the one-dimensional positional encoding to obtain the prompt text encoding vector of the sample prompt text. Through comprehensive encoding from multiple dimensions, it realizes the comprehensive training of the initial visual question answering model according to the information contained in the sample text, making the training of the model more comprehensive. Moreover, since segment encoding is performed on the sample prompt text during the training of the initial visual question answering model, the obtained visual question answering model can recognize multiple prompt texts at one time, improving the speed of subsequent data processing using this visual question answering model.
[0109] After completing the encoding of the sample prompt text, it is also necessary to encode the sample document image, that is, according to the sample document image, determine the document encoding vector and the sample image encoding vector corresponding to the sample document image. It can be understood as determining the document encoding vector and the sample image encoding vector corresponding to the sample document image according to the sample document information and the sample picture information contained in the sample document image.
[0110] In practical applications, to determine the document encoding vector and the sample image encoding vector corresponding to the sample document image according to the sample document image, it can be achieved by encoding the sample document and the sample image included in the sample document image respectively. The specific implementation method is as follows:
[0111] Determining the sample document encoding vector and the sample image encoding vector corresponding to the sample document image according to the sample document image includes:
[0112] Identifying the sample document text in the sample document image through a character recognition algorithm, and performing text box annotation on the sample document text;
[0113] Determining the sample document encoding vector corresponding to the sample document image according to the sample document text;
[0114] Using an image feature extraction network to extract features from the sample document image to obtain sample image features;
[0115] Determining the sample image encoding vector corresponding to the sample document image according to the sample image features.
[0116] Among them, the character recognition algorithm can be understood as an OCR (Optical Character Recognition) recognition algorithm. In practical applications, there are various OCR recognition algorithms, which are not limited in the embodiments of this specification; the image feature extraction network can be understood as a CNN network (Convolutional Neural Networks).
[0117] Specifically, for the specific implementation of identifying the sample document text in the sample document image through a character recognition algorithm and performing text box annotation on the sample document text, reference can be made to the specific implementation of "identifying the sample document text in each initial sample document image through a character recognition algorithm and performing text box annotation on the sample document text" in the above-mentioned embodiments of the specification. The difference is only that the recognized document images are different.
[0118] On this basis, after identifying the sample document text in the sample document image and performing text box annotation, the sample document text can be encoded in the encoding network layer to determine the sample document encoding vector corresponding to the sample document image.
[0119] Specifically, the specific implementation method of encoding the sample document text in the encoding network layer is as follows:
[0120] Determining the sample document encoding vector corresponding to the sample document image according to the sample document text includes:
[0121] Performing word encoding on the sample document text to determine the word encoding vector of the sample document text;
[0122] Performing segment encoding on the sample document text to determine the segment encoding vector of the sample document text;
[0123] Perform one-dimensional position encoding on the text of the sample document to determine the one-dimensional position encoding vector of the text of the sample document;
[0124] Perform two-dimensional position encoding on the text of the sample document to determine the two-dimensional position encoding vector of the text of the sample document;
[0125] According to the word encoding vector, the segment encoding vector, the one-dimensional position encoding vector, and the two-dimensional position encoding vector, determine the sample document encoding vector corresponding to the sample document image.
[0126] Among them, performing word encoding on the text of the sample document can be understood as an operation of encoding the characters or words included in the text of the sample document, performing segment encoding on the text of the sample document can be understood as an operation of performing segment encoding on the paragraphs included in the text of the sample document, and performing one-dimensional position encoding on the text of the sample document can be understood as an operation of encoding the text of the sample document according to the positions of each character or word in the sentence where it is located; performing two-dimensional position encoding on the text of the sample document can be understood as an operation of encoding the text of the sample document according to the positions of each character or word in the image where it is located.
[0127] Specifically, the specific implementation methods of performing word encoding, segment encoding, and one-dimensional position encoding on the text of the sample document can refer to the specific implementation methods of performing word encoding, segment encoding, and one-dimensional position encoding on the sample prompt text in the above-mentioned specification embodiments; and performing two-dimensional position encoding on the text of the sample document can be understood as an operation of performing two-dimensional position encoding on the sample prompt text according to the two-dimensional positions of each character or word in the sample document image.
[0128] Then, after completing the word encoding, segment encoding, one-dimensional position encoding, and two-dimensional position encoding on the text of the sample document, and respectively obtaining the encoded word encoding vector, segment encoding vector, one-dimensional position encoding vector, and two-dimensional position encoding vector, combine the encoding vectors, segment encoding vectors, one-dimensional position encoding vectors, and two-dimensional position encoding vectors of the sample prompt text to obtain the sample document encoding vector corresponding to the combined text of the sample document.
[0129] The visual question answering model training method provided by the embodiments of this specification, by performing word encoding, segment encoding, one-dimensional position encoding, and two-dimensional position encoding on the text of the sample document, and combining the obtained word encoding vector, segment encoding vector, one-dimensional position encoding vector, and two-dimensional position encoding vector to obtain the prompt text encoding vector of the text of the sample document, comprehensively analyzes the text of the sample document in the sample document image from multiple dimensions, and improves the multimodality and comprehensiveness of model training.
[0130] In practical applications, to improve the accuracy of two-dimensional position encoding, a two-dimensional position encoding vector algorithm can be used to implement two-dimensional position encoding. The specific implementation method is as follows:
[0131] Performing two-dimensional position encoding on the sample document text to determine the two-dimensional position encoding vector of the sample document text includes:
[0132] Using the two-dimensional position encoding vector algorithm and the position information of the text box of the sample document text, performing two-dimensional position encoding on the sample document text to determine the two-dimensional position encoding vector of the sample document text.
[0133] Among them, the two-dimensional position encoding vector algorithm can be understood as an algorithm that uses the formula for performing two-dimensional position encoding on the sample document text and the position information of the text box of the sample document text to perform two-dimensional position encoding on the sample document text.
[0134] Specifically, the formula for performing two-dimensional position encoding on the sample document text is Formula 1 shown below:
[0135] L = E 2x (x0, x1, w) + E 2y (y0, y1, h) Formula 1
[0136] Among them, E 2x is the embedding layer in the x direction (the embedding layer, that is, the above-mentioned encoding layer);
[0137] E 2y is the embedding layer in the y direction (the embedding layer, that is, the above-mentioned encoding layer);
[0138] x and y refer to the position coordinates in two directions of the sample document image. Specifically, the x direction refers to the horizontal direction coordinate of the sample document image, and the y direction refers to the vertical direction coordinate of the sample document image;
[0139] (x0, y0) represents the upper left corner point of the bounding box (the boundary box, that is, the above-mentioned text box) of the sample document text;
[0140] (x1, y1) represents the lower right corner point of the bounding box (the boundary box, that is, the above-mentioned text box) of the sample document text;
[0141] (w, h) represents the width and height of the bounding box (the boundary box, that is, the above-mentioned text box) of the sample document text.
[0142] In specific implementation, according to the positions of the text boxes of each text in the above sample document text and the above formula for two-dimensional position encoding of the sample document text, the two-dimensional position encoding of the sample document text can be achieved by using the two-dimensional position encoding vector algorithm, and the two-dimensional position encoding vector corresponding to the sample document text can be obtained.
[0143] The visual question answering model training method provided by the embodiments of this specification further improves the accuracy of the two-dimensional position encoding vector of the sample document text by performing two-dimensional position encoding on the sample document text by using the two-dimensional position encoding vector algorithm and the position information of the text boxes of the sample document text.
[0144] And using the image feature extraction network to extract features from the sample document image and obtain the sample image features can be understood as obtaining the sample image features of one or more small block regions after the sample document image is segmented, feature extracted, and resized by using a CNN network, where the small block regions can be understood as one or more image small block regions obtained after the sample document image is segmented.
[0145] Further, after obtaining the sample image features, the specific implementation manner of determining the sample image encoding vector corresponding to the sample document image is as follows:
[0146] Determining the sample image encoding vector corresponding to the sample document image according to the sample image features includes:
[0147] Performing segment encoding on the sample image features to determine the segment encoding vector of the sample image features;
[0148] Performing one-dimensional position encoding on the sample image features to determine the one-dimensional position encoding vector of the sample image features;
[0149] Performing two-dimensional position encoding on the sample image features to determine the two-dimensional position encoding vector of the sample image features;
[0150] Determining the sample image encoding vector corresponding to the sample document image according to the sample image features, the segment encoding vector, the one-dimensional position encoding vector, and the two-dimensional position encoding vector.
[0151] Specifically, the specific implementation manner of encoding the sample image features can refer to the specific implementation manner of encoding the sample document text above, which will not be elaborated here. The difference is only that the encoding object is different, which leads to different encoding results.
[0152] The visual question answering model training method provided by the embodiments of this specification performs word encoding, segment encoding, one-dimensional position encoding, and two-dimensional position encoding on the sample image features, and combines the obtained sample image features, corresponding word encoding vectors, segment encoding vectors, one-dimensional position encoding vectors, and two-dimensional position encoding vectors to obtain the sample image encoding vector of the sample document image. By comprehensively analyzing the sample document image in the sample document image from multiple dimensions, the analysis of the sample document image is made more comprehensive, improving the multimodality and comprehensiveness of model training.
[0153] In practical applications, in order to improve the accuracy of the two-dimensional position encoding vector of the sample image features, a two-dimensional position encoding vector algorithm can also be used for encoding. The specific implementation method is as follows:
[0154] The two-dimensional position encoding of the sample image features to determine the two-dimensional position encoding vector of the sample image features includes:
[0155] Using the two-dimensional position encoding vector algorithm and the position information of the sample image features in the sample document image, perform two-dimensional position encoding on the sample image features to determine the two-dimensional position encoding vector of the sample image features.
[0156] Among them, the two-dimensional position encoding vector algorithm can be understood as an algorithm for performing two-dimensional position encoding on the sample document text by using the formula for two-dimensional position encoding of the sample image features and the position information of the sample image features in the sample document image.
[0157] Specifically, the formula for two-dimensional position encoding of the sample image features is the following formula 2:
[0158] L = E 2x (x0′, x1′, w′) + E 2y (y0′, y1′, h′) Formula 2
[0159] Among them, E 2x is the embedding layer in the x direction (the embedding layer, that is, the above-mentioned encoding layer);
[0160] E 2y is the embedding layer in the y direction (the embedding layer, that is, the above-mentioned encoding layer);
[0161] x and y refer to the position coordinates in two directions of the sample document image. Specifically, the x direction refers to the horizontal direction coordinate of the sample document image, and the y direction refers to the vertical direction coordinate of the sample document image;
[0162] (x0′, y0′) represents the upper left corner point of the patch (that is, the small block area corresponding to the sample image feature) in the sample document image;
[0163] (x1′, y1′) represents the lower right corner point of the patch (i.e., the small block area corresponding to the sample image feature) in the sample document image;
[0164] (w′, h′) represents the width and height of the patch (i.e., the small block area corresponding to the sample image feature).
[0165] In specific implementation, according to the position information of the above sample image feature in the sample document image and the above formula for two-dimensional position encoding of the sample image feature, the two-dimensional position encoding of the sample image feature can be realized by using the two-dimensional position encoding vector algorithm, and the two-dimensional position encoding vector corresponding to the sample document image can be obtained.
[0166] The visual question answering model training method provided by the embodiments of this specification further improves the accuracy of the two-dimensional position encoding vector of the sample document text by performing two-dimensional position encoding on the sample document text by using the two-dimensional position encoding vector algorithm and the position information of the text box of the sample document text.
[0167] The visual question answering model training method provided by the embodiments of this specification performs character recognition on the sample document image through a character recognition algorithm to obtain the sample document text corresponding to the sample document image, and then determines the sample document encoding vector according to the sample document text, ensuring the accuracy of the sample document encoding vector. Moreover, the sample document image is feature-extracted by using an image feature extraction network to obtain the sample image feature, and the sample image encoding vector is determined according to the sample image feature, ensuring the accuracy of the sample image encoding vector.
[0168] Then, after obtaining the prompt text encoding vector, the sample document encoding vector, and the sample image encoding vector, according to the prompt text encoding vector, the sample document encoding vector, and the sample image encoding vector, the semantic feature vector of the sample document image is determined. It can be understood that the prompt text encoding vector, the sample document encoding vector, and the sample image encoding vector are combined to obtain a combined vector, and the combined vector is encoded to obtain the semantic feature vector corresponding to the sample document image. This semantic feature vector represents the information of the sample prompt text, the text information in the sample document image, and the image information in the sample document image.
[0169] After obtaining the semantic feature vector, it is also necessary to decode this semantic feature vector.
[0170] Specifically, in the decoding network, the semantic feature vector is decoded to obtain a predicted text answer corresponding to the sample prompt text determined from the sample document image. It can be understood that after the encoding network determines the semantic feature vector of the sample document image, the semantic feature vector is input into the decoding network, and the decoding network decodes the semantic feature vector to obtain the output result of the model, which is also the predicted text answer corresponding to the sample prompt text determined from the sample document image.
[0171] Then, the sample text answer corresponding to the sample prompt text determined from the sample document image is used as the sample label, and the predicted text answer is the output result of model training. According to the difference between the sample label and the output result of model training, the training of the above initial visual question answering model can be achieved.
[0172] The visual question answering model training method provided by the embodiments of this specification specifically obtains the semantic feature vector of the sample document image based on the prompt text encoding vector, the sample document encoding vector, and the sample image encoding vector, ensuring the richness of the semantic feature vector. And through the model architecture of the encoding network - decoding network, the prompt text information is incorporated into the model training stage, enabling the subsequent visual question answering model to process multiple user questions at one time, reducing the processing time consumption of the visual question answering model in actual applications.
[0173] The visual question answering model training method provided by the embodiments of this specification is realized by training with the sample prompt text, the sample document image, and the sample text answer corresponding to the sample prompt text determined from the sample document image. Considering the influence of the sample prompt text, the accuracy of the visual question answering model in recognizing the sample prompt text is improved, so as to improve the accuracy of subsequent data processing using the visual question answering model. And when training this visual question answering model, it is only based on a limited number of initial sample document images, reducing the cost of sample data. By using the initial sample document images and the reading comprehension sample text to augment the target sample document images, and using the augmented target sample document images and the initial sample document images as the sample document images for training, the accuracy of the subsequent trained visual question answering model in actual applications is improved. In addition, due to the use of the reading comprehension sample text for sample augmentation, the generalization ability of the subsequent trained visual question answering model for new sample document images is increased. In summary, the visual question answering model trained by this method can achieve the technical effects of lower data processing time consumption, lower cost, higher generalization ability, and higher data processing accuracy.
[0174] See Figure 2 , Figure 2 which is a flowchart of a data processing method provided by an embodiment of this specification.
[0175] Step 202: Determine the target prompt text and the target document image corresponding to the target prompt text.
[0176] Step 204: Input the target prompt text and the target document image into a visual question answering model to obtain a text answer corresponding to the target prompt text determined from the target document image.
[0177] Wherein, the visual question answering model is trained by sample prompt texts, sample document images, and sample text answers corresponding to the sample prompt texts determined from the sample document images.
[0178] Optionally, the sample document images include initial sample document images and target sample document images amplified according to the initial sample document images and reading comprehension sample texts;
[0179] Correspondingly, the step of inputting the target prompt text and the target document image into the visual question answering model to obtain a text answer corresponding to the target prompt text determined from the target document image includes:
[0180] Input the target prompt text and the target document image into the visual question answering model, wherein the visual question answering model includes an encoding network and a decoding network;
[0181] In the encoding network, according to the target prompt text, determine a prompt text encoding vector corresponding to the target prompt text;
[0182] According to the target document image, determine a document encoding vector and an image encoding vector corresponding to the target document image;
[0183] According to the prompt text encoding vector, the document encoding vector, and the image encoding vector, determine a semantic feature vector of the target document image;
[0184] In the decoding network, decode the semantic feature vector to obtain a text answer corresponding to the target prompt text determined from the target document image.
[0185] The data processing method provided by the embodiments of this specification uses the initial sample document image and the reading comprehension sample text to augment and obtain the target sample document image, and uses the augmented target sample document image and the initial sample document image as the sample document images for training, further improving the accuracy of data processing using the visual question answering model. In addition, due to the use of the reading comprehension sample text for sample augmentation, the generalization of the visual question answering model to new sample document images is increased. Specifically, the semantic feature vector of the target document image is obtained based on the prompt text encoding vector, the document encoding vector, and the image encoding vector, ensuring the richness of the semantic feature vector. And in the training stage, a model architecture of an encoding network - decoding network is adopted, incorporating the prompt text information into the model training stage, enabling the visual question answering model to process multiple user questions at one time, reducing the processing time of the visual question answering model in actual applications, and thus reducing the total time for data processing according to this data processing method.
[0186] Optionally, the determining the prompt text encoding vector corresponding to the target prompt text according to the target prompt text includes:
[0187] Performing word encoding on the target prompt text to determine the word encoding vector of the target prompt text;
[0188] Performing segment encoding on the target prompt text to determine the segment encoding vector of the target prompt text;
[0189] Performing one - dimensional position encoding on the target prompt text to determine the one - dimensional position encoding vector of the target prompt text;
[0190] Determining the prompt text encoding vector corresponding to the target prompt text according to the word encoding vector, the segment encoding vector, and the one - dimensional position encoding vector.
[0191] The visual question answering model training method provided by the embodiments of this specification performs word encoding, segment encoding, and one - dimensional position encoding on the target prompt text, and combines the word encoding, the segment encoding, and the one - dimensional position encoding to obtain the prompt text encoding vector of the target prompt text, achieving comprehensive analysis of the information contained in the target prompt text from multiple dimensions through comprehensive encoding, making the data processing results more comprehensive.
[0192] Optionally, the determining the document encoding vector and the image encoding vector corresponding to the target document image according to the target document image includes:
[0193] Identifying the document text in the target document image through an optical character recognition algorithm, and performing text box annotation on the document text;
[0194] Determine the document encoding vector corresponding to the target document image according to the document text;
[0195] Use an image feature extraction network to extract features from the target document image to obtain image features;
[0196] Determine the image encoding vector corresponding to the target document image according to the image features.
[0197] The data processing method provided by the embodiments of this specification performs character recognition on the target document image through a character recognition algorithm, and obtains the sample document text corresponding to the target document image. Then, according to the document text, the document encoding vector is determined, ensuring the accuracy of the document encoding vector. Moreover, an image feature extraction network is used to extract features from the target document image to obtain image features, and according to the image features, the image encoding vector is determined, ensuring the accuracy of the image encoding vector.
[0198] Optionally, the determining the document encoding vector corresponding to the target document image according to the document text includes:
[0199] Perform word encoding on the document text to determine the word encoding vector of the document text;
[0200] Perform segment encoding on the document text to determine the segment encoding vector of the document text;
[0201] Perform one-dimensional position encoding on the document text to determine the one-dimensional position encoding vector of the document text;
[0202] Perform two-dimensional position encoding on the document text to determine the two-dimensional position encoding vector of the document text;
[0203] Determine the document encoding vector corresponding to the target document image according to the word encoding vector, the segment encoding vector, the one-dimensional position encoding vector, and the two-dimensional position encoding vector.
[0204] The data processing method provided by the embodiments of this specification performs word encoding, segment encoding, one-dimensional position encoding, and two-dimensional position encoding on the document text, and combines the obtained word encoding vector, segment encoding vector, one-dimensional position encoding vector, and two-dimensional position encoding vector to obtain the prompt text encoding vector of the document text, comprehensively analyzing the document text in the target document image from multiple dimensions, improving the accuracy of the finally generated text answer and the comprehensiveness of data processing.
[0205] Optionally, the determining the image encoding vector corresponding to the target document image according to the image features includes:
[0206] Perform segment encoding on the image features to determine the segment encoding vector of the image features;
[0207] Perform one-dimensional position encoding on the image features to determine the one-dimensional position encoding vector of the image features;
[0208] Perform two-dimensional position encoding on the image features to determine the two-dimensional position encoding vector of the image features;
[0209] According to the image features, the segment encoding vector, the one-dimensional position encoding vector, and the two-dimensional position encoding vector, determine the image encoding vector corresponding to the target document image.
[0210] The data processing method provided by the embodiments of this specification performs word encoding, segment encoding, one-dimensional position encoding, and two-dimensional position encoding on image features, and combines the obtained image features, corresponding word encoding vectors, segment encoding vectors, one-dimensional position encoding vectors, and two-dimensional position encoding vectors to obtain the image encoding vector of the document image. By comprehensively analyzing the document image in the target document image from multiple dimensions, the analysis of the target document image is made more comprehensive, improving the accuracy of the finally generated text answer and the comprehensiveness of data processing.
[0211] The data processing method provided by the embodiments of this specification inputs the target prompt text and the target document image into the visual question answering model to obtain the text answer corresponding to the target prompt text determined from the target document image, realizing the understanding of the target document image and answering the answer corresponding to the target prompt text from the target document; moreover, the training of the visual question answering model is achieved by training with sample prompt texts, sample document images, and sample text answers corresponding to the sample prompt texts determined from the sample document images. Considering the training of the sample prompt texts, the accuracy of the visual question answering model in recognizing the sample prompt texts is improved to improve the accuracy of subsequent data processing. Moreover, when the visual question answering model is trained, only a limited number of initial sample document images are used, reducing the cost of sample data. In summary, this method achieves the technical effects of lower data processing time consumption, lower cost, and improved data processing accuracy.
[0212] The above is a schematic solution of a data processing method of this embodiment. It should be noted that the technical solution for training the visual question answering model in this data processing method belongs to the same concept as the technical solution for training the visual question answering model in the above-mentioned visual question answering model training method. For the details not described in the technical solution of the data processing method, reference can be made to the description of the technical solution of the above-mentioned visual question answering model training method.
[0213] See Figure 3 , Figure 3 which is a flowchart of another visual question answering model training method provided by an embodiment of this specification, applied to the cloud, and specifically includes the following steps:
[0214] Step 302: Receive the sample prompt text, the initial sample document image, and the reading comprehension sample text sent by the front end;
[0215] Step 304: Replace the text in the initial sample document image with the reading comprehension sample text to obtain the target sample document image;
[0216] Step 306: Determine the sample document image according to the initial sample document image and the target sample document image;
[0217] Step 308: Train the initial visual question answering model according to the sample prompt text, the sample document image, and the sample text answer corresponding to the sample prompt text determined from the sample document image until the training stop condition is met, obtain the visual question answering model, and send the visual question answering model to the front end.
[0218] The visual question answering model training method provided in the embodiments of this specification performs model training in the cloud after receiving the sample prompt text, the sample document image, and the sample text answer corresponding to the sample prompt text determined from the sample document image sent by the front end, reducing the storage pressure and operation pressure of the front end, and considering the influence on the sample prompt text. Therefore, the effect of the subsequent trained visual question answering model processing multiple user questions at one time is achieved. Moreover, when training this visual question answering model, only a limited number of initial sample document images are used, reducing the cost of sample data. And by using the initial sample document image and the reading comprehension sample text to amplify and obtain the target sample document image, and using the amplified target sample document image and the initial sample document image as the sample document images for training, the accuracy of the subsequent trained visual question answering model in practical applications is ensured. In addition, due to the use of the reading comprehension sample text for sample augmentation, the generalization of the subsequent trained visual question answering model to new sample document images is increased.
[0219] The above is a schematic solution of a visual question answering model training method applied to the cloud in this embodiment. It should be noted that the technical solution of visual question answering model training in the visual question answering model training method applied to the cloud belongs to the same concept as the technical solution of visual question answering model training in the above visual question answering model training method. For the details not described in the technical solution of the visual question answering model training method applied to the cloud, reference can be made to the description of the technical solution of the visual question answering model training method above.
[0220] See Figure 4 , Figure 4 is a schematic diagram of the structure of a visual question answering model provided by an embodiment of this specification.
[0221] As shown Figure 4 in the figure Figure 4 it includes prompt text (hint), document image, encoder (encoding network), decoder (decoding network), and output (text answer corresponding to the prompt text determined from the document image).
[0222] Specifically, when the visual question answering model processes data, it inputs the prompt text and the document image into the visual question answering model. After processing the prompt text and the document image, the visual question answering model outputs the text answer corresponding to the prompt text determined from the document image.
[0223] Specifically, the specific processing process of the visual question answering model for the prompt text and the document image is as follows:
[0224] 1. Encode the prompt text.
[0225] Among them, encoding the prompt text includes three parts, specifically as follows:
[0226] (1) Perform word encoding (content embedding) on the prompt text to generate a prompt text word encoding vector;
[0227] (2) Perform segment encoding (Token type embedding) on the prompt text to generate a prompt text segment encoding vector;
[0228] (3) Perform one-dimensional position encoding (1D position embedding) on the prompt text to generate a prompt text one-dimensional position encoding vector.
[0229] 2. Perform OCR recognition on the document image to obtain the document text.
[0230] 3. Encode the document text.
[0231] (1) Perform word encoding (content embedding) on the document text to generate a document text word encoding vector;
[0232] (2) Perform segment encoding (Token type embedding) on the document text to generate a document text segment encoding vector;
[0233] (3) Perform one-dimensional position encoding (1D position embedding) on the document text to generate a document text one-dimensional position encoding vector;
[0234] (4) Use the above-mentioned document text two-dimensional position encoding vector calculation formula to perform two-dimensional position encoding (2D position embedding) on the document text to generate a document text two-dimensional position encoding vector;
[0235] Specifically, the calculation formula for the document text two-dimensional position encoding vector can be seen in the above formula 1.
[0236] 4. Encode the document image.
[0237] (1) Use the CNN network to first divide the document image into multiple patches (i.e., small regions in the document image), and extract the image features of each document image patch;
[0238] (2) Perform segment encoding (Token type embedding) on each document image patch to generate segment encoding vectors;
[0239] (3) Perform one-dimensional position encoding (1D position embedding) on each document image patch to generate one-dimensional position encoding vectors;
[0240] (4) Use the calculation formula of the two-dimensional position encoding vector of the document image to perform two-dimensional position encoding vector (2D position embedding) on each document image patch to generate two-dimensional position encoding vectors;
[0241] Specifically, the calculation formula of the two-dimensional position encoding vector of the document image can refer to the above formula 2.
[0242] (5) Take the sum of the image features, segment encoding vectors, one-dimensional position encoding vectors, and two-dimensional position encoding vectors of the document image as the document image encoding vector.
[0243] 5. Input the concatenation result of the prompt text encoding vector, document text encoding vector, and document image encoding vector into the 12-layer multi-modal Transformer layer of the encoder (encoding network) of the visual question answering model for encoding to obtain rich semantic features, and input the rich semantic features into the decoder (decoding network).
[0244] 6. Input the rich semantic feature vector into the 12-layer Transformer decoding layer of the decoder (decoding network) of the visual question answering model, and the encoder (encoding network) outputs the text answer in the document image corresponding to the prompt text.
[0245] It should be noted that the above steps are the data processing process of the visual question answering model. The training process of the visual question answering model can refer to the above steps, with the only difference being that during training, it will compare the output text answer with the preset text answer used as the sample label to train the model. In addition, the source of the sample document images during model training can be obtained through the following methods:
[0246] First, collect multiple text-containing files with different styles and a text reading comprehension dataset. Perform OCR detection on the document images of these files to detect the positions of the text and obtain initial layout samples (the layout samples can be understood as document images with text and text layout). Then, replace the text in the above layout samples with the text in the text reading comprehension dataset to obtain new document image samples. Combine the new document image samples with the initial layout samples to obtain the model training samples for the visual question answering model.
[0247] The above is a schematic structural diagram of a visual question answering model according to this embodiment. It should be noted that this visual question answering model has the same concept as the visual question answering model in the above data processing method, or the above visual question answering model training method, or the above visual question answering model training method applied to the cloud. For the details not described in detail in the technical solutions in the embodiments of this specification, reference can be made to the descriptions of the technical solutions in the above data processing method, or the above visual question answering model training method, or the above visual question answering model training method applied to the cloud.
[0248] Corresponding to the above embodiment of the visual question answering model training method, this specification also provides an embodiment of the visual question answering model training device. Refer to Figure 5 , Figure 5 which is a schematic structural diagram of a visual question answering model training device provided by an embodiment of this specification. As Figure 5 shown, the device includes:
[0249] A sample determination module 502, configured to determine a sample prompt text, an initial sample document image, and a reading comprehension sample text;
[0250] A target sample document image obtaining module 504, configured to replace the text in the initial sample document image with the reading comprehension sample text to obtain a target sample document image;
[0251] A sample document image determination module 506, configured to determine the sample document image according to the initial sample document image and the target sample document image;
[0252] A visual question answering model obtaining module 508, configured to train an initial visual question answering model according to the sample prompt text, the sample document image, and the sample text answer corresponding to the sample prompt text determined from the sample document image until the training stop condition is met, and obtain the visual question answering model.
[0253] Optionally, the initial sample document image includes multiple document images containing text with different text layouts;
[0254] Accordingly, the target sample document image obtaining module 504 is further configured to:
[0255] Identify the sample document text in each initial sample document image through a character recognition algorithm, and perform text box annotation on the sample document text;
[0256] Replace the sample document text annotated by the text box in each initial sample document image through the reading comprehension sample text to obtain a target sample document image.
[0257] Optionally, the initial visual question answering model includes an encoding network and a decoding network;
[0258] Accordingly, the visual question answering model obtaining module 508 is further configured to:
[0259] In the encoding network, determine the prompt text encoding vector corresponding to the sample prompt text according to the sample prompt text;
[0260] Determine the sample document encoding vector and the sample image encoding vector corresponding to the sample document image according to the sample document image;
[0261] Determine the semantic feature vector of the sample document image according to the prompt text encoding vector, the sample document encoding vector, and the sample image encoding vector;
[0262] In the decoding network, decode the semantic feature vector to obtain the predicted text answer corresponding to the sample prompt text determined from the sample document image;
[0263] Train the initial visual question answering model according to the sample text answer corresponding to the sample prompt text determined from the sample document image and the predicted text answer.
[0264] Optionally, the visual question answering model obtaining module 508 is further configured to:
[0265] Perform word encoding on the sample prompt text to determine the word encoding vector of the sample prompt text;
[0266] Perform segment encoding on the sample prompt text to determine the segment encoding vector of the sample prompt text;
[0267] Perform one-dimensional position encoding on the sample prompt text to determine the one-dimensional position encoding vector of the sample prompt text;
[0268] Determine the prompt text encoding vector corresponding to the sample prompt text according to the word encoding vector, the segment encoding vector, and the one-dimensional position encoding vector.
[0269] Optionally, the visual question answering model obtaining module 508 is further configured to:
[0270] Identify the sample document text in the sample document image through a character recognition algorithm, and perform text box annotation on the sample document text;
[0271] Determine the sample document encoding vector corresponding to the sample document image according to the sample document text;
[0272] Use an image feature extraction network to extract features from the sample document image to obtain sample image features;
[0273] Determine the sample image encoding vector corresponding to the sample document image according to the sample image features.
[0274] Optionally, the visual question answering model obtaining module 508 is further configured to:
[0275] Perform word encoding on the sample document text to determine the word encoding vector of the sample document text;
[0276] Perform segment encoding on the sample document text to determine the segment encoding vector of the sample document text;
[0277] Perform one-dimensional position encoding on the sample document text to determine the one-dimensional position encoding vector of the sample document text;
[0278] Perform two-dimensional position encoding on the sample document text to determine the two-dimensional position encoding vector of the sample document text;
[0279] Determine the sample document encoding vector corresponding to the sample document image according to the word encoding vector, the segment encoding vector, the one-dimensional position encoding vector, and the two-dimensional position encoding vector.
[0280] Optionally, the visual question answering model obtaining module 508 is further configured to:
[0281] Perform segment encoding on the sample image features to determine the segment encoding vector of the sample image features;
[0282] Perform one-dimensional position encoding on the sample image features to determine the one-dimensional position encoding vector of the sample image features;
[0283] Perform two-dimensional position encoding on the sample image features to determine the two-dimensional position encoding vector of the sample image features;
[0284] Determine the sample image encoding vector corresponding to the sample document image according to the sample image features, the segment encoding vector, the one-dimensional position encoding vector, and the two-dimensional position encoding vector.
[0285] Optionally, the visual question answering model obtaining module 508 is further configured to:
[0286] Perform two-dimensional position encoding on the sample document text by using the two-dimensional position encoding vector algorithm and the position information of the text box of the sample document text, and determine the two-dimensional position encoding vector of the sample document text.
[0287] Optionally, the visual question answering model obtaining module 508 is further configured to:
[0288] Perform two-dimensional position encoding on the sample image features by using the two-dimensional position encoding vector algorithm and the position information of the sample image features in the sample document image, and determine the two-dimensional position encoding vector of the sample image features.
[0289] The above is a schematic solution of a visual question answering model training device according to this embodiment. It should be noted that the technical solution of this visual question answering model training device and the technical solution of the above visual question answering model training method belong to the same concept. For the details not described in the technical solution of the visual question answering model training device, reference can be made to the description of the technical solution of the above visual question answering model training method.
[0290] Corresponding to the above data processing method embodiment, this specification also provides a data processing device embodiment. Figure 6 It is a schematic structural diagram of a data processing device provided by an embodiment of this specification. As Figure 6 shown, the device includes:
[0291] A determination module 602, configured to determine a target prompt text and a target document image corresponding to the target prompt text;
[0292] A processing module 604, configured to input the target prompt text and the target document image into a visual question answering model, and obtain a text answer corresponding to the target prompt text determined from the target document image, where the visual question answering model is trained by using a sample prompt text, a sample document image, and a sample text answer corresponding to the sample prompt text determined from the sample document image.
[0293] Optionally, the sample document image includes an initial sample document image and a target sample document image amplified according to the initial sample document image and a reading comprehension sample text;
[0294] Correspondingly, the processing module 604 is further configured to:
[0295] Input the target prompt text and the target document image into a visual question answering model, where the visual question answering model includes an encoding network and a decoding network;
[0296] In the encoding network, according to the target prompt text, determine a prompt text encoding vector corresponding to the target prompt text;
[0297] According to the target document image, determine a document encoding vector and an image encoding vector corresponding to the target document image;
[0298] According to the prompt text encoding vector, the document encoding vector, and the image encoding vector, determine a semantic feature vector of the target document image;
[0299] In the decoding network, decode the semantic feature vector to obtain a text answer corresponding to the target prompt text determined from the target document image.
[0300] Optionally, the processing module 604 is further configured to:
[0301] Perform word encoding on the target prompt text to determine a word encoding vector of the target prompt text;
[0302] Perform segment encoding on the target prompt text to determine a segment encoding vector of the target prompt text;
[0303] Perform one-dimensional position encoding on the target prompt text to determine a one-dimensional position encoding vector of the target prompt text;
[0304] According to the word encoding vector, the segment encoding vector, and the one-dimensional position encoding vector, determine a prompt text encoding vector corresponding to the target prompt text.
[0305] Optionally, the processing module 604 is further configured to:
[0306] Identify the document text in the target document image through an optical character recognition algorithm, and perform text box annotation on the document text;
[0307] According to the document text, determine a document encoding vector corresponding to the target document image;
[0308] Use an image feature extraction network to extract features from the target document image to obtain image features;
[0309] According to the image features, determine an image encoding vector corresponding to the target document image.
[0310] Optionally, the processing module 604 is further configured to:
[0311] Perform word encoding on the document text to determine the word encoding vector of the document text;
[0312] Perform segment encoding on the document text to determine the segment encoding vector of the document text;
[0313] Perform one-dimensional position encoding on the document text to determine the one-dimensional position encoding vector of the document text;
[0314] Perform two-dimensional position encoding on the document text to determine the two-dimensional position encoding vector of the document text;
[0315] Determine the document encoding vector corresponding to the target document image according to the word encoding vector, the segment encoding vector, the one-dimensional position encoding vector, and the two-dimensional position encoding vector.
[0316] Optionally, the processing module 604 is further configured to:
[0317] Perform segment encoding on the image features to determine the segment encoding vector of the image features;
[0318] Perform one-dimensional position encoding on the image features to determine the one-dimensional position encoding vector of the image features;
[0319] Perform two-dimensional position encoding on the image features to determine the two-dimensional position encoding vector of the image features;
[0320] Determine the image encoding vector corresponding to the target document image according to the image features, the segment encoding vector, the one-dimensional position encoding vector, and the two-dimensional position encoding vector.
[0321] The above is a schematic solution of a data processing device according to this embodiment. It should be noted that the technical solution of this data processing device and the technical solution of the above data processing method belong to the same concept. For the details not described in detail in the technical solution of the data processing device, reference can be made to the description of the technical solution of the above data processing method.
[0322] See Figure 7 , Figure 7 shows a structural block diagram of a computing device 700 according to an embodiment of this specification. The components of the computing device 700 include, but are not limited to, a memory 710 and a processor 720. The processor 720 is connected to the memory 710 through a bus 730, and a database 750 is used to store data.
[0323] The computing device 700 also includes an access device 740, which enables the computing device 700 to communicate via one or more networks 760. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 740 may include one or more of any type of wired or wireless network interface (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, a Near Field Communication (NFC) interface, and so on.
[0324] In one embodiment of the present specification, the above components of the computing device 700, as well as Figure 7 other components not shown, may also be connected to each other, for example, via a bus. It should be understood that Figure 7 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present specification. Those skilled in the art can add or replace other components as needed.
[0325] The computing device 700 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.), or other types of mobile devices, or a stationary computing device such as a desktop computer or a Personal Computer (PC). The computing device 700 can also be a mobile or stationary server.
[0326] Wherein, the processor 720 is used to execute the following computer-executable instructions, and when the computer-executable instructions are executed by the processor, the steps of the above data processing method, or the above visual question answering model training method, or the above visual question answering model training method applied to the cloud are implemented.
[0327] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solutions of the above data processing method, or the above visual question answering model training method, or the above visual question answering model training method applied to the cloud belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the descriptions of the technical solutions of the above data processing method, or the above visual question answering model training method, or the above visual question answering model training method applied to the cloud.
[0328] An embodiment of this specification also provides a computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the steps of the above data processing method, or the above visual question answering model training method, or the above visual question answering model training method applied to the cloud are implemented.
[0329] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solutions of the above data processing method, or the above visual question answering model training method, or the above visual question answering model training method applied to the cloud belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the descriptions of the technical solutions of the above data processing method, or the above visual question answering model training method, or the above visual question answering model training method applied to the cloud.
[0330] An embodiment of this specification also provides a computer program. When the computer program is executed on a computer, the computer is made to execute the steps of the above data processing method, or the above visual question answering model training method, or the above visual question answering model training method applied to the cloud.
[0331] The above is a schematic solution of a computer program according to this embodiment. It should be noted that the technical solution of this computer program and the technical solutions of the above data processing method, or the above visual question answering model training method, or the above visual question answering model training method applied to the cloud belong to the same concept. For the details not described in detail in the technical solution of the computer program, reference can be made to the descriptions of the technical solutions of the above data processing method, or the above visual question answering model training method, or the above visual question answering model training method applied to the cloud.
[0332] The above describes specific embodiments of this specification. Other embodiments are within the scope of the appended claims. In some cases, the acts or steps recited in the claims may be performed in a different order than in the embodiments and still achieve the desired result. Additionally, the processes depicted in the drawings do not necessarily require the particular order or sequential order shown to achieve the desired result. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0333] The computer instructions include computer program code, which may be in source code form, object code form, executable file, or some intermediate form, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording medium, USB flash drive, removable hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0334] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described order of actions, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0335] In the above embodiments, the descriptions of the various embodiments have their own focuses. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0336] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not elaborate on all the details, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and variations can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can well understand and utilize this specification. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A data processing method, comprising: Determining a target prompt text and a target document image corresponding to the target prompt text; Inputting the target prompt text and the target document image into a visual question answering model to obtain a text answer determined from the target document image and corresponding to the target prompt text, wherein the visual question answering model is trained by sample prompt texts, sample document images, and sample text answers determined from the sample document images and corresponding to the sample prompt texts.
2. The data processing method according to claim 1, wherein the sample document image includes an initial sample document image and a target sample document image amplified according to the initial sample document image and a reading comprehension sample text; Accordingly, the step of inputting the target prompt text and the target document image into the visual question answering model to obtain a text answer determined from the target document image and corresponding to the target prompt text includes: Inputting the target prompt text and the target document image into the visual question answering model, wherein the visual question answering model includes an encoding network and a decoding network; In the encoding network, determining a prompt text encoding vector corresponding to the target prompt text according to the target prompt text; Determining a document encoding vector and an image encoding vector corresponding to the target document image according to the target document image; Determining a semantic feature vector of the target document image according to the prompt text encoding vector, the document encoding vector, and the image encoding vector; In the decoding network, decoding the semantic feature vector to obtain a text answer determined from the target document image and corresponding to the target prompt text.
3. The data processing method according to claim 2, wherein the step of determining a prompt text encoding vector corresponding to the target prompt text according to the target prompt text includes: Performing word encoding on the target prompt text to determine a word encoding vector of the target prompt text; Performing segment encoding on the target prompt text to determine a segment encoding vector of the target prompt text; Performing one-dimensional position encoding on the target prompt text to determine a one-dimensional position encoding vector of the target prompt text; Determining a prompt text encoding vector corresponding to the target prompt text according to the word encoding vector, the segment encoding vector, and the one-dimensional position encoding vector.
4. The data processing method according to claim 2, wherein the step of determining a document encoding vector and an image encoding vector corresponding to the target document image according to the target document image includes: Identifying document text in the target document image through a character recognition algorithm and performing text box annotation on the document text; Determining a document encoding vector corresponding to the target document image according to the document text; Using an image feature extraction network to perform feature extraction on the target document image to obtain image features; Determining an image encoding vector corresponding to the target document image according to the image features.
5. A method for training a visual question answering model, comprising: Determining sample prompt texts, an initial sample document image, and reading comprehension sample texts; Replacing the text in the initial sample document image with the reading comprehension sample text to obtain a target sample document image; Determining the sample document image according to the initial sample document image and the target sample document image; Training the initial visual question answering model according to the sample prompt text, the sample document image, and the sample text answer corresponding to the sample prompt text determined from the sample document image until a training stop condition is satisfied to obtain the visual question answering model.
6. The method for training a visual question answering model according to claim 5, wherein the initial sample document image includes a plurality of document images containing text with different text layouts; Correspondingly, the replacing the text in the initial sample document image with the reading comprehension sample text to obtain a target sample document image includes: Identifying the sample document text in each initial sample document image through a character recognition algorithm and performing text box annotation on the sample document text; Replacing the sample document text annotated with the text box in each initial sample document image with the reading comprehension sample text to obtain a target sample document image.
7. The method for training a visual question answering model according to claim 5, wherein the initial visual question answering model includes an encoding network and a decoding network; Correspondingly, the training the initial visual question answering model according to the sample prompt text, the sample document image, and the sample text answer corresponding to the sample prompt text determined from the sample document image includes: In the encoding network, determining a prompt text encoding vector corresponding to the sample prompt text according to the sample prompt text; Determining a sample document encoding vector and a sample image encoding vector corresponding to the sample document image according to the sample document image; Determining a semantic feature vector of the sample document image according to the prompt text encoding vector, the sample document encoding vector, and the sample image encoding vector; In the decoding network, decoding the semantic feature vector to obtain a predicted text answer corresponding to the sample prompt text determined from the sample document image; Training the initial visual question answering model according to the sample text answer corresponding to the sample prompt text determined from the sample document image and the predicted text answer.
8. The method for training a visual question answering model according to claim 7, wherein the determining the prompt text encoding vector corresponding to the sample prompt text according to the sample prompt text includes: Performing word encoding on the sample prompt text to determine a word encoding vector of the sample prompt text; Performing segment encoding on the sample prompt text to determine a segment encoding vector of the sample prompt text; Performing one-dimensional position encoding on the sample prompt text to determine a one-dimensional position encoding vector of the sample prompt text; Determining the prompt text encoding vector corresponding to the sample prompt text according to the word encoding vector, the segment encoding vector, and the one-dimensional position encoding vector.
9. The method for training a visual question answering model according to claim 7, wherein determining the sample document encoding vector and the sample image encoding vector corresponding to the sample document image according to the sample document image comprises: Identifying the sample document text in the sample document image through a character recognition algorithm, and performing text box annotation on the sample document text; Determining the sample document encoding vector corresponding to the sample document image according to the sample document text; Using an image feature extraction network to extract features from the sample document image to obtain sample image features; Determining the sample image encoding vector corresponding to the sample document image according to the sample image features.
10. The method for training a visual question answering model according to claim 9, wherein determining the sample document encoding vector corresponding to the sample document image according to the sample document text comprises: Performing word encoding on the sample document text to determine the word encoding vector of the sample document text; Performing segment encoding on the sample document text to determine the segment encoding vector of the sample document text; Performing one-dimensional position encoding on the sample document text to determine the one-dimensional position encoding vector of the sample document text; Performing two-dimensional position encoding on the sample document text to determine the two-dimensional position encoding vector of the sample document text; Determining the sample document encoding vector corresponding to the sample document image according to the word encoding vector, the segment encoding vector, the one-dimensional position encoding vector, and the two-dimensional position encoding vector.
11. The method for training a visual question answering model according to claim 9, wherein determining the sample image encoding vector corresponding to the sample document image according to the sample image features comprises: Performing segment encoding on the sample image features to determine the segment encoding vector of the sample image features; Performing one-dimensional position encoding on the sample image features to determine the one-dimensional position encoding vector of the sample image features; Performing two-dimensional position encoding on the sample image features to determine the two-dimensional position encoding vector of the sample image features; Determining the sample image encoding vector corresponding to the sample document image according to the sample image features, the segment encoding vector, the one-dimensional position encoding vector, and the two-dimensional position encoding vector.
12. The method for training a visual question answering model according to claim 10, wherein performing two-dimensional position encoding on the sample document text to determine the two-dimensional position encoding vector of the sample document text comprises: Performing two-dimensional position encoding on the sample document text by using a two-dimensional position encoding vector algorithm and the position information of the text box of the sample document text to determine the two-dimensional position encoding vector of the sample document text.
13. The method for training a visual question answering model according to claim 11, wherein performing two-dimensional position encoding on the sample image features to determine the two-dimensional position encoding vector of the sample image features comprises: Using the two-dimensional position encoding vector algorithm and the position information of the sample image feature in the sample document image, perform two-dimensional position encoding on the sample image feature to determine the two-dimensional position encoding vector of the sample image feature.
14. A data processing device, comprising: A determination module, configured to determine a target prompt text and a target document image corresponding to the target prompt text; A processing module, configured to input the target prompt text and the target document image into a visual question answering model to obtain a text answer corresponding to the target prompt text determined from the target document image, where the visual question answering model is trained by a sample prompt text, a sample document image, and a sample text answer corresponding to the sample prompt text determined from the sample document image.
15. A computing device, comprising: A memory and a processor; The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions. When the computer-executable instructions are executed by the processor, the steps of the data processing method according to any one of claims 1 to 4 or the visual question answering model training method according to any one of claims 5 to 13 are implemented.
16. A computer-readable storage medium, which stores computer-executable instructions. When the computer-executable instructions are executed by a processor, the steps of the data processing method according to any one of claims 1 to 4 or the visual question answering model training method according to any one of claims 5 to 13 are implemented.