Data processing method, question and answer processing method and document understanding method
By annotating and layout prediction of the original sample set, the layout sample set is generated and data enhancement is performed, the problem of insufficient data labeling in complex environments of multimodal pre-trained models is solved, and the accuracy of data processing is improved.
Patent Information
- Application Number
- CN202410095648.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-01-23
- Publication Date
- 2025-07-29
AI Technical Summary
The multimodal pretrained model lacks complex data annotations in complex environments, resulting in insufficient accuracy of data processing results.
By obtaining the original sample set, annotating and layout prediction, the layout adjustment is made to obtain the target sample set, and data augmentation is performed on the target sample set to generate an enhanced sample set to train the data processing model.
Improve the data richness of the data processing model and ensure the accuracy of processing results.
Smart Images

Figure CN120386833A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of this specification relate to the field of computer technologies, and particularly to a data processing method, and also to a question-answering processing method, a document understanding method, a data processing system, a computing device, and a computer-readable storage medium. Background Art
[0002] With the in-depth research on multi-modal pre-trained models, the applications of multi-modal pre-trained models are becoming more and more extensive. For example, multi-modal pre-trained models have emerged in scenario fields such as question-answering, search, and human-computer interaction.
[0003] Before a multi-modal pre-trained model processes data for each scenario field, it is necessary to train the multi-modal pre-trained model using multi-modal sample data to ensure the accuracy of the results obtained by processing data based on the multi-modal pre-trained model. However, in real scenarios, multi-modal pre-trained models are mostly applied to complex data processing in various complex environments. Since it is difficult to obtain the corresponding complex data annotations in each complex environment, the multi-modal sample data is seriously insufficient, resulting in insufficient accuracy of the results obtained by processing using the trained multi-modal pre-trained model. Therefore, there is an urgent need for a method to improve the accuracy of processing results when using a multi-modal pre-trained model for data processing. Summary of the Invention
[0004] In view of this, the embodiments of this specification provide a data processing method. One or more embodiments of this specification also relate to a question-answering processing method, a document understanding method, a data processing device, a question-answering processing device, a document understanding device, a computing device, a computer-readable storage medium, and a computer program to solve the technical defects existing in the prior art.
[0005] According to the first aspect of the embodiments of this specification, a data processing method is provided, including:
[0006] Obtain target data;
[0007] Input the target data into a data processing model to obtain a processing result of the target data, where the data processing model is trained using an enhanced sample set, the enhanced sample set is obtained by performing data augmentation on a target sample set, the target sample set is obtained by performing layout adjustment on layout sample images of a layout sample set based on original sample images of an original sample set, the original sample set is obtained by annotating sample data, and the layout sample set is obtained by performing layout prediction on the sample data.
[0008] According to the second aspect of the embodiments of this specification, a question-answering processing method is provided, including:
[0009] Obtain the first question information;
[0010] Input the first question information into a data processing model to obtain the answer information of the first question information, where the data processing model is trained using an enhanced sample set, the enhanced sample set is obtained by data augmentation of a target sample set, the target sample set is obtained by layout adjustment of layout sample images in a layout sample set based on original sample images in an original sample set, the original sample set is obtained by annotating sample data, and the layout sample set is obtained by layout prediction of the sample data.
[0011] According to the third aspect of the embodiments of this specification, a document understanding method is provided, including:
[0012] Obtain a target document and second question information for the target document;
[0013] Input the target document and the second question information into a data processing model to obtain the answer information of the second question information, where the data processing model is trained using an enhanced sample set, the enhanced sample set is obtained by data augmentation of a target sample set, the target sample set is obtained by layout adjustment of layout sample images in a layout sample set based on original sample images in an original sample set, the original sample set is obtained by annotating sample data, and the layout sample set is obtained by layout prediction of the sample data.
[0014] According to the fourth aspect of the embodiments of this specification, a data processing device is provided, including:
[0015] A first acquisition module configured to acquire target data;
[0016] A first processing module configured to input the target data into a data processing model to obtain the processing result of the target data, where the data processing model is trained using an enhanced sample set, the enhanced sample set is obtained by data augmentation of a target sample set, the target sample set is obtained by layout adjustment of layout sample images in a layout sample set based on original sample images in an original sample set, the original sample set is obtained by annotating sample data, and the layout sample set is obtained by layout prediction of the sample data.
[0017] According to the fifth aspect of the embodiments of this specification, a question and answer processing device is provided, including:
[0018] A second acquisition module configured to acquire the first question information;
[0019] A second processing module, configured to input the first problem information into a data processing model to obtain answer information for the first problem information, where the data processing model is trained using an augmented sample set, the augmented sample set is obtained by performing data augmentation on a target sample set, the target sample set is obtained by performing layout adjustment on layout sample images of a layout sample set based on original sample images of an original sample set, the original sample set is obtained by annotating sample data, and the layout sample set is obtained by performing layout prediction on the sample data.
[0020] According to a sixth aspect of the embodiments of the present specification, there is provided a document understanding apparatus, including:
[0021] A third acquisition module, configured to acquire a target document and second problem information for the target document;
[0022] A third processing module, configured to input the target document and the second problem information into a data processing model to obtain answer information for the second problem information, where the data processing model is trained using an augmented sample set, the augmented sample set is obtained by performing data augmentation on a target sample set, the target sample set is obtained by performing layout adjustment on layout sample images of a layout sample set based on original sample images of an original sample set, the original sample set is obtained by annotating sample data, and the layout sample set is obtained by performing layout prediction on the sample data.
[0023] According to a seventh aspect of the embodiments of the present specification, there is provided a data processing system, where the data processing system includes an edge device and a cloud device;
[0024] The edge device is used to initiate a data processing request, where the target data is carried in the data processing request;
[0025] The cloud device is used to respond to the data processing request, acquire the target data, input the target data into a data processing model, and obtain a processing result of the target data, where the data processing model is trained using an augmented sample set, the augmented sample set is obtained by performing data augmentation on a target sample set, the target sample set is obtained by performing layout adjustment on layout sample images of a layout sample set based on original sample images of an original sample set, the original sample set is obtained by annotating sample data, and the layout sample set is obtained by performing layout prediction on the sample data.
[0026] According to an eighth aspect of the embodiments of the present specification, there is provided a computing device, including:
[0027] A memory and a processor;
[0028] The memory is used to store computer programs / instructions, and the processor is used to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the above-mentioned data processing method, question-and-answer processing method, and document understanding method are implemented.
[0029] According to the ninth aspect of the embodiments of the present specification, a computer-readable storage medium is provided, which stores computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the above-mentioned data processing method, question-and-answer processing method, and document understanding method are implemented.
[0030] According to the tenth aspect of the embodiments of the present specification, a computer program product is provided, including computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the above-mentioned data processing method, question-and-answer processing method, and document understanding method are implemented.
[0031] In one embodiment of the present specification, target data is obtained; the target data is input into a data processing model to obtain a processing result of the target data. Among them, the data processing model is trained using an enhanced sample set, and the enhanced sample set is obtained by data augmentation of a target sample set. The target sample set is obtained by performing layout adjustment on the layout sample images of a layout sample set based on the original sample images of an original sample set. The original sample set is obtained by annotating sample data, and the layout sample set is obtained by performing layout prediction on sample data. The target data is input into the data processing model trained using the enhanced sample set to obtain the processing result output by the data processing model. The enhanced sample set is obtained by performing annotation on the basis of sample data to obtain the original sample set, performing layout prediction on the basis of sample data to obtain the layout sample set, performing layout adjustment on the basis of the original sample set and the layout sample set to obtain the target sample set, and performing data augmentation on the basis of the target sample set. This fully ensures the data richness of the enhanced sample set, and further ensures the accuracy of the processing result when the data processing model trained using the enhanced sample set performs data processing. Description of the Drawings
[0032] Figure 1 is a schematic diagram of an interaction process under a data processing system architecture provided by one embodiment of the present specification;
[0033] Figure 2 is a framework diagram of a data processing system provided by one embodiment of the present specification;
[0034] Figure 3 is a flowchart of a data processing method provided by one embodiment of the present specification;
[0035] Figure 4 is a flowchart of a question-and-answer processing method provided by one embodiment of the present specification;
[0036] Figure 5 is a flowchart of a document understanding method provided by an embodiment of this specification;
[0037] Figure 6 is a processing flowchart of a data processing method provided by an embodiment of this specification;
[0038] Figure 7 is a schematic structural diagram of a data processing device provided by an embodiment of this specification;
[0039] Figure 8 is a schematic structural diagram of a question and answer processing device provided by an embodiment of this specification;
[0040] Figure 9 is a schematic structural diagram of a document understanding device provided by an embodiment of this specification;
[0041] Figure 10 is a structural block diagram of a computing device provided by an embodiment of this specification. Detailed implementation manners
[0042] Many specific details are set forth in the following description in order to provide a thorough understanding of this specification. However, this specification can be implemented in many other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the connotation of this specification. Therefore, this specification is not limited by the specific implementations disclosed below.
[0043] The terms used in one or more embodiments of this specification are for the purpose of describing specific embodiments only and are not intended to limit one or more embodiments of this specification. The singular forms "a", "the", and "said" used in one or more embodiments of this specification and the appended claims are also intended to include the plural forms unless the context clearly dictates otherwise. It should also be understood that the term "and / or" used in one or more embodiments of this specification refers to and includes any or all possible combinations of one or more of the associated listed items.
[0044] It should be understood that although the terms first, second, etc. may be used in one or more embodiments of this specification to describe various information, such information should not be limited to these terms. These terms are only used to distinguish the same type of information from each other. For example, without departing from the scope of one or more embodiments of this specification, the first can also be referred to as the second, and similarly, the second can also be referred to as the first. Depending on the context, the word "if" as used herein can be interpreted as "when" or "while" or "in response to determining".
[0045] In addition, it should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in one or more embodiments of this specification are all information and data authorized by the user or fully authorized by all parties. Moreover, the collection, use, and processing of relevant data need to comply with the relevant laws, regulations, and standards of relevant countries and regions, and corresponding operation entrances are provided for users to choose to authorize or reject.
[0046] In one or more embodiments of this specification, a large model refers to a deep learning model with a large number of model parameters, usually including hundreds of millions, tens of billions, hundreds of billions, trillions, or even more than one quadrillion model parameters. A large model can also be referred to as a Foundation Model. Through pre-training of the large model with a large amount of unlabeled corpus, a pre-trained model with more than one hundred million parameters is produced. This model can adapt to a wide range of downstream tasks and has good generalization ability. For example, large language models (LLMs), multi-modal pre-training models, etc.
[0047] When a large model is actually applied, it only needs to be fine-tuned with a small number of samples on the pre-trained model to be applied to different tasks. Large models can be widely applied in the fields of natural language processing (NLP) and computer vision. Specifically, they can be applied to tasks in the field of computer vision such as visual question answering (VQA), image captioning (IC), and image generation, as well as tasks in the field of natural language processing such as text-based sentiment classification, text summary generation, and machine translation. The main application scenarios of large models include digital assistants, intelligent robots, search, online education, office software, e-commerce, intelligent design, etc.
[0048] First, the noun terms involved in one or more embodiments of this specification are explained.
[0049] Optical Character Recognition (OCR) is an artificial intelligence (AI) tool that converts portable document format (PDF), document format (Word), spreadsheet document format (Excel), or text images into machine-encoded text.
[0050] For a long time, the optical character recognition and document understanding capabilities of multimodal pre-training models, especially the lack of optical character recognition and document understanding capabilities of Chinese, Japanese and Korean unified ideographic characters in complex environments, have been a major problem plaguing the training of multimodal pre-training models. An important part of solving the optical character recognition and document understanding capabilities of multimodal pre-training models lies in the generation of multilingual optical character recognition and document understanding data for the multimodal pre-training model. In this specification, a large amount of multilingual optical character recognition and document understanding data is obtained in batches by obtaining data from multiple sources and performing automatic annotation and data enhancement, which is used to improve the multilingual optical character recognition and document understanding capabilities of the multimodal pre-training model.
[0051] Specifically, in this specification, a data processing method is provided. This specification also involves a question and answer processing method, a document understanding method, a data processing device, a question and answer processing device, a document understanding device, a computing device, and a computer-readable storage medium, which are described in detail one by one in the following embodiments.
[0052] See also Figure 1 , Figure 1 FIG. 1 shows a schematic diagram of an interaction process under a data processing system architecture provided by an embodiment of this specification. Figure 1 As shown, the system includes a cloud-side device 100 and a terminal-side device 200 .
[0053] The end-side device 200 is used to initiate a data processing request, wherein the data processing request carries target data;
[0054] The cloud-side device 100 is configured to respond to the data processing request, obtain the target data, input the target data into a data processing model, and obtain a processing result of the target data, wherein the data processing model is trained using an enhanced sample set, the enhanced sample set is obtained by performing data enhancement on the target sample set, the target sample set is obtained by performing layout adjustment on layout sample images of a layout sample set based on original sample images of an original sample set, the original sample set is obtained by annotating sample data, and the layout sample set is obtained by performing layout prediction on the sample data;
[0055] Optionally, the end-side device 200 is further configured to receive processing results.
[0056] Applying the solution of the embodiments of this specification, input the target data into the data processing model obtained by training with the enhanced sample set to obtain the processing result output by the data processing model. The enhanced sample set is obtained by performing annotation on the basis of the sample data to obtain the original sample set, performing layout prediction on the basis of the sample data to obtain the layout sample set, performing layout adjustment on the basis of the original sample set and the layout sample set to obtain the target sample set, and performing data enhancement on the basis of the target sample set, which fully ensures the data richness of the enhanced sample set. Furthermore, when the data processing model obtained by training with the enhanced sample set performs data processing, the accuracy of the processing result is fully ensured.
[0057] See Figure 2 , Figure 2 which shows a framework diagram of a data processing system provided by an embodiment of this specification. The system may include a cloud-side device 100 and multiple end-side devices 200. Communication connections can be established between the multiple end-side devices 200 through the cloud-side device 100. In a data processing scenario, the cloud-side device 100 is used to provide data processing services between the multiple end-side devices 200. The multiple end-side devices 200 can respectively act as a sending end or a receiving end and communicate through the cloud-side device 100.
[0058] The user can interact with the cloud-side device 100 through the end-side device 200 to receive data sent by other end-side devices 200, or send data to other end-side devices 200, etc. In a data processing scenario, it may be that the user issues a data processing request to the cloud-side device 100 through the end-side device 200, and the cloud-side device 100 generates a data processing result according to the data processing request and pushes the data processing result to other end-side devices 200 that have established communication.
[0059] Among them, a connection is established between the end-side device 200 and the cloud-side device 100 through a network. The network provides a medium for the communication link between the end-side device 200 and the cloud-side device 100. The network may include various connection types, such as wired, wireless communication links, or fiber optic cables, etc. The data transmitted by the end-side device 200 may need to be processed such as encoded, transcoded, compressed, etc. before being published to the cloud-side device 100.
[0060] The edge device 200 can be a browser, an application (APP), or a web application such as a HyperText Markup Language 5 (H5) application, or a light application (also known as a mini-program, a lightweight application), or a cloud application, etc. The edge device 200 can be developed based on the software development kit (SDK) of the corresponding service provided by the cloud device, such as developed based on the Real Time Communication (RTC) SDK. The edge device 200 can be deployed in an electronic device and needs to rely on the device or certain APPs in the device to run, etc. The electronic device can, for example, have a display screen and support information browsing, etc., such as a personal mobile terminal such as a mobile phone, a tablet computer, a personal computer, etc. Various other types of applications can usually be configured in the electronic device, such as human-computer dialogue applications, model training applications, text processing applications, web browser applications, shopping applications, search applications, instant messaging tools, email edge devices, social platform software, etc.
[0061] The cloud device 100 can include servers that provide various services. For example, a server that provides communication services for multiple edge devices, or a server for background training that provides support for models used on edge devices, or a server that processes data sent by edge devices, etc. It should be noted that the cloud device 100 can be implemented as a distributed server cluster composed of multiple servers, or can be implemented as a single server. The server can also be a server of a distributed system, or a server combined with a blockchain. The server can also be a cloud server of basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, Content Delivery Network (CDN), and big data and artificial intelligence platforms, or an intelligent cloud computing server or an intelligent cloud host with artificial intelligence technology.
[0062] It is worth noting that the data processing method provided in the embodiments of this specification is generally executed by the cloud device 100. However, in other embodiments of this specification, the edge device 200 can also have a similar function to the cloud device, so as to execute the data processing method provided in the embodiments of this specification. In other embodiments, the data processing method provided in the embodiments of this specification can also be jointly executed by the edge device 200 and the cloud device 100.
[0063] See Figure 3 , Figure 3The flowchart of a data processing method provided by an embodiment of this specification is shown, which specifically includes the following steps.
[0064] Step 302: Obtain target data.
[0065] The embodiment of this specification is applied to an end-side device and / or a cloud-side device with data processing capabilities. The following takes the cloud-side device as an example for illustration.
[0066] When there is a need for data processing, the cloud-side device will obtain target data to input the target data into the data processing model and obtain the processing result output by the data processing model.
[0067] Specifically, the target data refers to data of various types, languages, and formats. For example, the target data is data in formats such as images, PDFs, Excel, Word, web archives (MHTML MIME Encapsulation of Aggregate HTML Documents), plain text, etc. that include one or more of languages X, Y, and Z. The server can perform optical character recognition, question-and-answer processing, document information processing, dialogue, sign recognition in autonomous driving, etc. based on this data. Exemplarily, when having a dialogue, the target data is a question, so that the data processing model answers the question. In the embodiment of this specification, the question and the answer can be the first question information and the answer information of the first question information as follows Figure 4 When performing document information processing, it can specifically be performing reading comprehension processing, then the target data includes a document and a question about the document. In the embodiment of this specification, the document and the question about the document can be the target document and the second question information about the target document as follows Figure 5 When performing sign recognition in autonomous driving, the target data is an image of a roadside sign during driving. The image of the sign is input into the data processing model to obtain the sign in the image of the sign.
[0068] There are many ways to obtain the target data. It can be that the user uploads a data processing request through the front end, and the data processing request carries the target data; it can also be that the data processing request carries the index address of the target data, and the cloud-side device obtains the target data based on the index address carried in the received data processing request to perform data processing on the target data.
[0069] Step 304: Input the target data into the data processing model to obtain the processing result of the target data. The data processing model is trained using the augmented sample set, which is obtained by data augmentation of the target sample set. The target sample set is obtained by adjusting the layout of the layout sample images in the layout sample set based on the original sample images in the original sample set. The original sample set is obtained by annotating the sample data, and the layout sample set is obtained by predicting the layout of the sample data.
[0070] Specifically, the data processing model refers to the model obtained after fine-tuning the multi-modal pre-trained model. The multi-modal pre-trained model can learn the semantic correspondence between different modalities through pre-training on large-scale data. Among them, the multi-modal pre-trained model is a multi-modal large model. The augmented sample set refers to a data set including rich sample data. The data processing model obtained by fine-tuning the multi-modal pre-trained model based on the augmented sample set can fully process multi-modal data and obtain accurate processing results.
[0071] The target sample set, the original sample set, and the layout sample set refer to the basic sample sets obtained based on the sample data, so as to use these basic sample sets to obtain the augmented sample set. Among them, the original sample set is directly obtained by annotating the sample data, the layout sample set is the sample set indirectly obtained based on the sample data and the original sample set, and the target sample set is the sample set obtained after layout adjustment based on the original sample set and the layout sample set. The target sample set has a certain degree of data richness, but the richness is relatively low compared with the augmented sample set. The sample data refers to the samples providing the original data for fine-tuning the multi-modal pre-trained model to obtain the data processing model. The sample data can be various types, various languages, and various formats of data, such as data in PDF, Word, Excel, MHTML, plain text format, etc., including one or more of languages X, Y, and Z.
[0072] The augmented sample image, the original sample image, and the layout sample image are the sample data included in the augmented sample set, the original sample set, and the layout sample set respectively. The augmented sample image, the original sample image, and the layout sample image also include corresponding annotation information. The annotation information refers to the information annotating the layout of the image. The annotation information can include annotation boxes and text annotations. Among them, the annotation box is the bounding box of the text on the sample image, and the representation form of the annotation box is the coordinates of the annotation box on the image. The representation form of the text annotation is the text and the coordinates of the text in the sample image.
[0073] The augmented sample set includes multiple augmented sample images and the annotation information of each augmented sample image. Among them, the annotation information is a box enclosing the text in the augmented sample image. The original sample set includes multiple original sample images and the annotation information of each original sample image. Among them, the annotation information is a box enclosing the text in the original sample image. The layout sample set includes multiple layout sample images and the annotation information of each layout sample image. Among them, the annotation information is a box enclosing the text in the layout sample image.
[0074] Optionally, to ensure the accuracy of data processing by the data processing model, in addition to training the data processing model using the augmented sample set, it can also be trained simultaneously based on the original sample set, the layout sample set, and the target sample set.
[0075] In an optional embodiment of this specification, the data processing method further includes the following steps:
[0076] Determine the sample data, and obtain the original sample set and the layout sample set of the sample data. Among them, the original sample set includes the original sample images and the annotation information on the original sample images, and the layout sample set includes the layout sample images and the annotation information on the layout sample images;
[0077] Based on the original sample images, the annotation information on the original sample images, and the annotation information on the layout sample images, perform layout adjustment on the layout sample images to obtain the target sample set. Among them, the target sample set includes the target sample images and the annotation information on the target sample images.
[0078] Specifically, the target sample image is the sample data included in the target sample set, and the target sample image also includes annotation information. The target sample set includes multiple target sample images and the annotation information of each target sample image. Among them, the annotation information is a box enclosing the text in the target sample image.
[0079] The implementation manner of determining the sample data can be to receive a sample retrieval task, and based on the sample retrieval task, perform data retrieval to obtain the sample data. The sample data retrieval task can include one or more of information such as the type of the sample data, the retrieval strategy of the sample data, and the keywords of the sample data.
[0080] The implementation method of obtaining the original sample set of sample data can be to render the sample data to obtain the sample images and sample titles of the sample data, make annotations on the sample images, integrate the sample images to obtain multiple original sample images and the annotation information on each original sample image, so as to obtain the original sample set; it can also be to input the sample data into the sample set generation model to obtain the sample set output by the sample set generation model, and determine this sample set as the original sample set, where the sample set generation model is a model that is pre-trained based on multiple sample data and the label sample sets corresponding to each sample data.
[0081] Among them, the implementation method of integrating the sample images to obtain multiple original sample images and the annotation information on each original sample image to obtain the original sample set can be to sort the sample images and the annotations on the sample images according to the order of each data in the sample data, so as to obtain multiple sorted original sample images and the annotation information on each original sample image, and obtain the original sample set.
[0082] The implementation method of obtaining the layout sample set of sample data can be to perform layout prediction on the sample data based on the original sample set and the sample title to obtain the layout sample set of the sample data.
[0083] There are many implementation methods for adjusting the layout of the layout sample image based on the original sample image, the annotation information on the original sample image, and the annotation information on the layout sample image. It is specifically determined according to the actual situation, and this specification does not limit it here.
[0084] In a possible implementation method of this specification, adjusting the layout of the layout sample image based on the original sample image, the annotation information on the original sample image, and the annotation information on the layout sample image can be to determine the text content on the original sample image and the layout sample image based on the annotation information on the original sample image and the annotation information on the layout sample image; obtain the first correspondence between the original sample image and the layout sample image, extract the original text content in the annotation information from the original sample image, and replace the layout text content in the annotation information on the layout sample image with the original text content according to the first correspondence.
[0085] In another possible implementation method of this specification, adjusting the layout of the layout sample image based on the original sample image, the annotation information on the original sample image, and the annotation information on the layout sample image can also be to determine the text content on the layout sample image and the original sample image based on the annotation information on the layout sample image and the annotation information on the original sample image, obtain the second correspondence between the layout sample image and the original sample image, extract the layout text content in the annotation information from the layout sample image, and replace the original text content in the annotation information on the original sample image with the layout text content according to the second correspondence.
[0086] By determining sample data, determining an original sample set and a layout sample set based on the sample data, and determining a target sample set based on the original sample set and the layout sample set, so as to perform data augmentation using the target sample set to obtain an augmented sample set, and performing model training based on the augmented sample set to obtain a data processing model, such that when target data is obtained, the data processing model can be used to process the target data to obtain a processing result, that is, using the data processing model trained with the highly rich augmented sample set to process the target data, which ensures the accuracy of the processing result.
[0087] In an optional embodiment of this specification, the above step of determining sample data includes the following steps:
[0088] Obtain a sample retrieval task, where at least one data type of the sample data is carried in the sample retrieval task;
[0089] Based on the at least one data type, determine a retrieval strategy for retrieving the sample data;
[0090] According to the retrieval strategy, perform data retrieval to obtain the sample data of each data type.
[0091] Specifically, the sample retrieval task refers to a task generated based on the need to retrieve sample data. For example, it can be that the user uploads retrieval requirements and model training requirements through the front end, and the cloud-side device generates a sample retrieval task based on one or more of these requirements. At least one data type of the sample data can be carried in these requirements, or at least one data type of the sample data may not be carried. In the case where the data type is not carried, the cloud-side device, based on a predefined program for processing the requirements, obtains at least one data type of the sample data, so as to generate a sample retrieval task based on the at least one data type and the requirements, where the predefined program is used to instruct the cloud-side device to determine at least one data type corresponding to the requirements. The retrieval strategy refers to the method strategy for data retrieval. For example, the retrieval strategy can be to purchase access rights to various databases to obtain sample data, or to obtain a typesetting system with unrestricted license types from a data website to obtain sample data.
[0092] Optionally, the sample retrieval task may further include the language of the content included in the sample data, such as language A, language B, language C, and the sample retrieval task may further include keywords of the content included in the sample data. For example, image processing data, data in the field of robotics, etc. The acquisition methods of the language and keywords are the same as or similar to those of the data type.
[0093] The data type can be a type for storing various different data. For example, the data type can be PDF, Word, Excel, MHTML, plain text.
[0094] Based on at least one data type, an implementation manner of determining a retrieval strategy for retrieving sample data may be to, based on at least one data type, look up the retrieval strategies corresponding to the respective data types in a strategy database, and determine the determined retrieval strategy as the retrieval strategy for retrieving sample data.
[0095] An implementation manner of performing data retrieval according to the retrieval strategy to obtain sample data of each data type may be to obtain one or more of the language and keywords of the sample data, and retrieve the sample data of each data type according to the retrieval strategy.
[0096] Exemplarily, purchase access rights to various databases and obtain a large number of PDF files therefrom; obtain a large amount of web page data from websites developed by the company or other authorized affiliated companies, and package the text and images in the web pages and save them as MHTML files; purchase access rights to various databases and obtain a large amount of text information, and obtain the text content after cleaning; obtain typesetting system (LaTeX) source files with unrestricted license types from data websites.
[0097] Through the obtained sample retrieval task, determine the retrieval strategy, and determine the sample data of each data type according to the retrieval strategy, so that the retrieved sample data corresponds to the sample retrieval task, that is, the sample data corresponds to the sample retrieval task with high richness requirements, ensuring the richness of the retrieved sample data, and the data processing model is trained with a high-richness enhanced sample set generated based on the sample data, ensuring the accuracy of the results obtained by processing each modality of data using the data processing model.
[0098] In an optional embodiment of this specification, the above steps of obtaining the original sample set and the layout sample set of the sample data include the following steps:
[0099] Based on the sample data, determine the rendering strategy corresponding to the sample data;
[0100] Use the rendering strategy to render the sample data to obtain the sample image and sample title of the sample data;
[0101] Annotate the sample image to obtain the sample image carrying annotation information;
[0102] Based on the sample image carrying annotation information, determine the original sample set;
[0103] According to the sample title and the original sample set, perform layout prediction on the sample data to obtain the layout sample set of the sample data.
[0104] Specifically, a rendering strategy refers to strategies such as methods and ways of rendering data. The rendering strategy can be rendering software, a neural network model, etc. A sample title refers to the sample title corresponding to the sample data.
[0105] Based on the sample data, determine the implementation manner of the rendering strategy corresponding to the sample data. It can be based on the sample data to determine the data type of the sample data, and based on the data type to determine the rendering strategy of the sample data.
[0106] Use the rendering strategy to render the sample data to obtain the sample image of the sample data and the implementation manner of the sample title. It can be to use the rendering strategy to render the sample data to obtain the sample image of the sample data and the sample title.
[0107] Annotate the sample image to obtain the implementation manner of the sample image carrying annotation information. It can be to use the rendering strategy to annotate the sample image to obtain the sample image carrying annotation information.
[0108] Exemplarily, for the obtained PDF file, we use rendering software A for rendering and annotation. Rendering software A can generate the rendered effect image file A of each page of the PDF file, as well as the bounding boxes of the corresponding text therein, and determine the rendered effect image file A and the bounding boxes as the sample image and the annotation information on the sample image; for the obtained MHTML file, use rendering software B to load the MHTML file for rendering and annotation. Rendering software B can generate the rendered effect image file B of the MHTML file, as well as generate the bounding boxes of all the corresponding text, and determine the rendered effect image file B and the bounding boxes as the sample image and the annotation information on the sample image; for the obtained LaTeX file, use rendering software C to render the LaTeX file to obtain the corresponding PDF file and the annotation of the compilation tool, use rendering software A to generate the rendered effect image file C of each page of the PDF file, use the compilation tool annotation to obtain the bounding boxes of each page of text, and determine the rendered effect image file C and the bounding boxes as the sample image and the annotation information on the sample image.
[0109] Based on the sample image carrying annotation information, determine the implementation manner of the original sample set. It can be to integrate the sample images carrying annotation information to obtain the original sample set.
[0110] According to the sample title and the original sample set, perform layout prediction on the sample data to obtain the implementation manner of the layout sample set of the sample data. It can be to train a layout generation model based on the sample title and the original sample set, input the sample data and the sample title into the layout generation model for layout prediction to obtain the layout sample image output by the layout generation model, and integrate the layout sample images to obtain the layout sample set.
[0111] Among them, the implementation method of integrating the layout sample images to obtain the layout sample set can be to sort the layout sample images carrying annotation information according to the order of each data in the sample data, obtain multiple sorted layout sample images and the annotation information on each layout sample image, and obtain the layout sample set.
[0112] Based on the sample data, determine the rendering strategy of the sample data, and based on the rendering strategy, render and annotate the sample data to obtain the original sample set and sample title of the sample data. Then, based on the sample title and the original sample set, perform layout prediction on the sample data to obtain the layout sample set, so as to subsequently use the original sample set obtained by rendering and annotation and the layout sample set obtained by layout prediction for layout adjustment and data augmentation to train the data processing model, thereby improving the accuracy of the data processing model in processing data.
[0113] In an optional embodiment of this specification, the above step of performing layout prediction on the sample data according to the sample title and the original sample set to obtain the layout sample set of the sample data includes the following steps:
[0114] Train the layout generation model according to the sample title and the annotation information on the original sample images in the original sample set to obtain the trained layout generation model;
[0115] Input the sample data and the sample title of the sample data into the layout generation model to obtain the layout sample set of the sample data.
[0116] Specifically, the annotation information is the bounding box surrounding the text on the sample image. The layout generation model refers to a neural network for layout prediction trained based on the sample data title, the sample image, and the annotation information on the sample image.
[0117] The implementation method of training the layout generation model based on the sample title and the annotation information on the original sample images in the original sample set to obtain the trained layout generation model can be to determine the target original sample image and the annotation information on the target original sample image from the original sample images in the original sample set, where the target original sample image is any original sample image in the original sample set; input the sample title and the target original sample image into the initial layout generation model to obtain the predicted layout image output by the initial layout generation model; calculate the loss value based on the annotation information on the predicted layout image and the annotation information on the target original sample image, adjust the initial layout generation model based on the loss value, and return to execute the steps of determining the target sample data from the sample data, determining the target original sample image and the annotation information on the target original sample image from the original sample images in the target sample data until the training stop condition is reached to obtain the layout generation model.
[0118] The implementation method of inputting sample data and the sample title of the sample data into a layout generation model to obtain a layout sample set of the sample data. Input the sample data and the sample title of the sample data into the layout generation model. Through layout prediction by the layout generation model, obtain multiple layout sample images of the sample data and the annotation information on the layout sample images, and determine the layout sample set.
[0119] Based on the sample title and the original sample set, train a layout generation model so that the layout generation model performs layout prediction on the sample data to obtain the predicted layout sample set, realizing the improvement of the richness of the sample set by using the original sample set and the layout sample set, and realizing the enhancement of the high richness of the sample data included in the sample set.
[0120] In an optional embodiment of this specification, the above steps are based on the original sample image, the annotation information on the original sample image, and the annotation information on the layout sample image to perform layout adjustment on the layout sample image to obtain a target sample set, including the following steps:
[0121] Based on the original sample image and the annotation information on the original sample image, determine the original text content;
[0122] Based on the layout sample image and the annotation information on the layout sample image, determine the layout text content;
[0123] Replace the layout text content in the layout sample image with the original text content to obtain a target sample image;
[0124] Integrate the target sample images carrying annotation information to obtain a target sample set.
[0125] Specifically, the original text content refers to the text content surrounded by the annotation information in the original sample image. The layout text content is the text content surrounded by the annotation information in the layout sample image.
[0126] The implementation method of determining the original text content based on the original sample image and the annotation information on the original sample image can be to identify the original text content surrounded by the annotation information in the original sample image based on the annotation information on the original sample image.
[0127] The implementation method of determining the layout text content based on the layout sample image and the annotation information on the layout sample image can be to identify the layout text content surrounded by the annotation information in the layout sample image based on the annotation information on the layout sample image.
[0128] The implementation method of replacing the layout text content in the layout sample image with the original text content to obtain the target sample image may be to extract the original text content from the original sample image, extract the layout text content from the layout sample image, obtain the original text content and an empty layout sample image without the layout text content, and fill the extracted original text content into the empty layout sample image to obtain the target sample image.
[0129] The implementation method of integrating the target sample images carrying annotation information to obtain the target sample set may be to integrate the target sample images carrying annotation information according to the data order in the sample data to obtain the target sample set.
[0130] Based on the original sample image and the layout sample image, determine the original text content and the layout text content, replace the layout text content with the original text content to obtain the target sample image, and then determine the target sample set, that is, improve the richness of the sample data included in the target sample set by means of layout adjustment.
[0131] In an optional embodiment of this specification, the data processing method further includes the following steps:
[0132] Perform data augmentation on the target sample image and the annotation information on the target sample image to obtain an augmented sample set, where the augmented sample set includes augmented sample images and the annotation information on the augmented sample images;
[0133] Use the augmented sample set to train the initial data processing model to obtain the trained data processing model.
[0134] Specifically, the initial data processing model is a data processing model that has not been trained with the augmented sample set. The trained data processing model is a data processing model obtained by training with the augmented sample set.
[0135] The implementation method of performing data augmentation on the target sample image and the annotation information on the target sample image to obtain an augmented sample set may be to obtain a data augmentation method and perform data augmentation on the target sample image and the annotation information on the target sample image based on the data augmentation method; it may also be to obtain at least two data augmentation methods and the execution order of the at least two data augmentation methods; according to the execution order, perform data augmentation on the target sample image and the annotation information on the target sample image based on the at least two data augmentation methods to obtain the augmented sample set.
[0136] Optionally, the data augmentation method may be one or more of a basic data augmentation method, a perspective transformation data augmentation method, and a mask data augmentation method.
[0137] Using the enhanced sample set to train an initial data processing model to obtain the implementation manner of the trained data processing model may be to extract a target enhanced sample image and the annotation information on the target enhanced sample image from the enhanced sample set, input the target enhanced sample image into the initial data processing model to obtain predicted annotation information, calculate a loss value based on the annotation information and the predicted annotation information on the target enhanced sample image, adjust the model parameters of the initial data processing model based on the loss value, and return to execute the step of extracting the target enhanced sample image and the annotation information on the target enhanced sample image from the enhanced sample set until the training stop condition is reached to obtain the trained data processing model, where the target enhanced sample image is any enhanced sample image in the enhanced sample set.
[0138] Optionally, the training stop condition may be that the number of iterations reaches a preset number threshold, the loss value reaches a loss value threshold, etc.
[0139] By annotating the target sample images and the annotation information on the target sample images in the target sample set to obtain an enhanced sample set, so as to use the enhanced sample set for model training to obtain a trained data processing model, that is, to obtain a data processing model trained using an enhanced sample set with high richness. Using the data processing model obtained through training to process the target data ensures the accuracy of the processing result.
[0140] In an optional embodiment of this specification, the above step of performing data augmentation on the target sample image and the annotation information on the target sample image to obtain an enhanced sample set includes the following steps:
[0141] Obtain at least two data augmentation methods and the execution order of the at least two data augmentation methods;
[0142] According to the execution order, perform data augmentation on the target sample image and the annotation information on the target sample image based on the at least two data augmentation methods to obtain an enhanced sample set.
[0143] Specifically, the execution order refers to the stage of performing data augmentation on the sample data in the target sample set, and different data augmentation methods correspond to different execution stages. The data to be augmented corresponding to different execution stages is different and there is an association relationship. For example, the data augmentation result of the previous stage is the data to be augmented in the next stage. The data augmentation method refers to the method of augmenting data, and the data augmentation method can be a basic data augmentation method, a perspective transformation data augmentation method, a mask data augmentation method.
[0144] The implementation method for obtaining at least two data augmentation methods and the execution order of at least two data augmentation methods can be to determine at least two data augmentation methods and the execution order based on a selection instruction; it can also be to determine at least two data augmentation methods based on a selection instruction and determine the execution order of at least two data augmentation methods based on the strong logical relationship between at least two data augmentation methods.
[0145] The implementation method for performing data augmentation on the target sample image and the annotation information on the target sample image based on at least two data augmentation methods in the execution order to obtain an augmented sample set can be to determine the execution stages corresponding to at least two data augmentation methods in the execution order and perform data augmentation on the target sample image and the annotation information on the target sample image based on the data augmentation methods corresponding to at least two execution stages in the execution order to obtain the augmented results corresponding to each execution stage.
[0146] When performing data augmentation on the target sample image and the annotation information on the target sample image in the target dataset, obtain at least two data augmentation methods and the execution order of at least two data augmentation methods, and perform data augmentation on the target sample image and the annotation information on the target sample image using at least two data augmentation methods based on the execution order to obtain an augmented sample set, realizing the sequential augmentation of the sample data in the target sample set and ensuring the data richness of the augmented sample set after data augmentation.
[0147] In an optional embodiment of this specification, performing data augmentation on the target sample image and the annotation information on the target sample image based on at least two data augmentation methods in the execution order to obtain an augmented sample set includes the following steps:
[0148] Determine the augmentation stages of at least two data augmentation methods in the execution order, where the augmentation stages at least include a first augmentation stage and a second augmentation stage;
[0149] For the first augmentation stage, perform data augmentation on the target sample image and the annotation information on the target sample image based on the data augmentation method corresponding to the first augmentation stage to obtain the first augmented sample image in the first augmentation stage and the annotation information on the first augmented sample image, where the first augmentation stage is the augmentation stage ranked first in the execution order;
[0150] For the second augmentation stage, perform data augmentation on the previous augmented sample image and the annotation information on the previous augmented sample image in the previous augmentation stage of the second augmentation stage based on the data augmentation method corresponding to the second augmentation stage to obtain the second augmented sample image in the second augmentation stage and the annotation information on the second augmented sample image, where the second augmentation stage is the augmentation stage other than the first augmentation stage.
[0151] Specifically, the first enhancement stage is the enhancement stage ranked first in the execution order, and there is a corresponding data enhancement method for the first enhancement stage.
[0152] For the first enhancement stage, based on the data enhancement method corresponding to the first enhancement stage, perform data enhancement on the target sample image and the annotation information on the target sample image to obtain the first enhanced sample image and the annotation information on the first enhanced sample image in the first enhancement stage. Among them, the implementation method of the first enhancement stage being the enhancement stage ranked first in the execution order can be to determine the first data enhancement method corresponding to the first enhancement stage for the first enhancement stage, and based on the first data enhancement method, perform data enhancement on the target sample image and the annotation information on the target sample image to obtain the first enhanced sample image and the annotation information on the first enhanced sample image in the first enhancement stage.
[0153] For the second enhancement stage, the implementation method of performing data enhancement on the previous enhanced sample image and the annotation information on the previous enhanced sample image in the previous enhancement stage based on the data enhancement method corresponding to the second enhancement stage to obtain the second enhanced sample image and the annotation information on the second enhanced sample image in the second enhancement stage can be to determine the second data enhancement method corresponding to the second enhancement stage for the second enhancement stage, and based on the second data enhancement method, perform data enhancement on the previous enhanced sample image and the annotation information on the previous enhanced sample image in the previous enhancement stage to obtain the second enhanced sample image and the annotation information on the second enhanced sample image in the second enhancement stage.
[0154] According to the execution order, divide the data enhancement process of the sample data in the target sample set into at least two enhancement stages, so that based on the data enhancement methods of the at least two enhancement stages in the execution order, perform data enhancement on the sample data in the target sample set, improving the data richness of the sample set.
[0155] In an optional embodiment of this specification, the data enhancement methods include at least two of the following: basic data enhancement method, perspective transformation data enhancement method, mask data enhancement method; the above steps, according to the execution order, perform data enhancement on the target sample image and the annotation information on the target sample image based on at least two data enhancement methods to obtain an enhanced sample set, including the following steps:
[0156] Based on the basic data enhancement method, obtain the filter to be superimposed, and superimpose the target sample image and the annotation information on the target sample image with the filter to be superimposed to obtain the basic enhanced sample image and the annotation information on the basic enhanced sample image;
[0157] Based on the perspective transformation data augmentation method, select the target perspective transformation parameters, and based on the target perspective transformation parameters, perform perspective transformation on the basic augmented sample image and the annotation information on the basic augmented sample image to obtain the transformed augmented sample image and the annotation information on the transformed augmented sample image;
[0158] Based on the mask data augmentation method, obtain the foreground image and the background image, and superimpose the foreground image, the background image, the transformed augmented sample image, and the annotation information on the transformed augmented sample image to obtain the masked augmented sample image and the annotation information on the masked augmented sample image;
[0159] Determine the augmented sample set according to the masked augmented sample image and the annotation information on the masked augmented sample image.
[0160] Specifically, the basic data augmentation method refers to the augmentation method of performing conventional augmentation on the data to be augmented. For example, adding filters to the data to be augmented, and the filters can be randomly modifying brightness, randomly adjusting contrast, randomly modifying hue, Gaussian blur, motion blur, elastic distortion, etc. The perspective transformation data augmentation method refers to the augmentation method of distorting the image based on the perspective for the data to be augmented. For example, based on the perspective transformation parameters, generate a perspective transformation strategy, and based on the perspective transformation strategy, perform perspective transformation on the data to be augmented. The mask data augmentation method refers to the augmentation method of adding a mask to the data to be augmented. For example, obtain the foreground image and the background image, and use the foreground image and the background image as masks to superimpose with the data to be augmented.
[0161] The implementation method of obtaining the filter to be superimposed based on the basic data augmentation method and superimposing the target sample image and the annotation information on the target sample image with the filter to be superimposed to obtain the basic augmented sample image and the annotation information on the basic augmented sample image can be to randomly obtain at least one filter to be superimposed based on the basic data augmentation method, and superimpose at least one filter to be superimposed with the target sample image marked with the annotation information to obtain the basic augmented sample image marked with the annotation information.
[0162] The implementation method of selecting the target perspective transformation parameters based on the perspective transformation data augmentation method and performing perspective transformation on the basic augmented sample image and the annotation information on the basic augmented sample image based on the target perspective transformation parameters to obtain the transformed augmented sample image and the annotation information on the transformed augmented sample image can be to obtain the perspective transformation parameters, generate a perspective transformation strategy based on the perspective transformation parameters and the perspective transformation data augmentation method, and use the perspective transformation strategy to perform data augmentation on the basic augmented sample image marked with the annotation information to obtain the transformed augmented sample image marked with the annotation information.
[0163] Optionally, the perspective transformation parameters are such that after the perspective transformation, the image will not exceed the page range. For example, not exceeding the page range means flipping the sample image, and the range obtained by projecting the flipped page visually cannot be larger than the original one.
[0164] Based on the mask data augmentation method, a foreground image and a background image are obtained. The foreground image, the background image, the transformed and augmented sample image, and the annotation information on the transformed and augmented sample image are superimposed to obtain a masked and augmented sample image and the annotation information on the masked and augmented sample image. The implementation method is to randomly select at least one foreground image and at least one background image based on the mask data augmentation method, and superimpose the at least one foreground image, the at least one background image, and the transformed and augmented sample image annotated with annotation information to obtain a masked and augmented sample image annotated with annotation information.
[0165] Optionally, based on the mask data augmentation method, data augmentation is performed on the transformed and augmented sample image annotated with annotation information to obtain a masked and augmented sample image annotated with annotation information. Refer to formula (1): I' = IA+(1.0 - I)B, where A is the foreground image, B is the background image, I is the transformed and augmented sample image annotated with annotation information, and I' is the masked and augmented sample image annotated with annotation information.
[0166] The data augmentation methods include basic data augmentation method, perspective transformation data augmentation method, and mask data augmentation method. The sample data in the target sample set is augmented according to the execution order based on the data augmentation methods, so as to improve the richness of the sample data in the target sample set. Furthermore, the processing richness of the data processing model obtained by training the data processing model based on the augmented sample set is improved, and thus the accuracy of the data processing model in processing multi-modal data is achieved.
[0167] In one or more embodiments of this specification, a large number of collected sample data, including file, text, and picture data, are used for rendering annotation, automatic layout generation, and layout randomization through a large-scale multi-modal pre-training model. On this basis, through data augmentation technology, a large amount of training data is provided for the OCR and document understanding capabilities of the multi-modal pre-training model. Automatic layout generation based on a large-scale multi-modal pre-training model can greatly improve the diversity of the generated data, so that the trained model can simultaneously have OCR capabilities, as well as document understanding capabilities and multi-modal dialogue capabilities brought by the language model.
[0168] See Figure 4 , Figure 4 shows a flowchart of a question-and-answer processing method provided by an embodiment of this specification, which specifically includes the following steps.
[0169] Step 402: Obtain the first question information.
[0170] The embodiments of this specification are applied to end-side devices and / or cloud-side devices with question-and-answer processing functions. The following takes the cloud-side device as an example for illustration.
[0171] When there is a need for question-and-answer processing, the cloud-side device will obtain the first question information so as to input the first question information into the data processing model and obtain the answer information of the first question information output by the data processing model.
[0172] Specifically, the first question information refers to the text corresponding to the question that the front-end user needs to know the corresponding answer. For example, the first question information can be "What is the day after yesterday the day before tomorrow?", "There are 9 birds in the tree, 1 is shot down, how many are left?", and so on.
[0173] The method for obtaining the first question information can refer to the method for obtaining the target data in the above Figure 3 and will not be elaborated here.
[0174] Step 404: Input the first question information into the data processing model to obtain the answer information of the first question information, where the data processing model is trained using an enhanced sample set, the enhanced sample set is obtained by data augmentation of the target sample set, the target sample set is obtained by layout adjustment of the layout sample images of the layout sample set based on the original sample images of the original sample set, the original sample set is obtained by annotating sample data, and the layout sample set is obtained by layout prediction of sample data.
[0175] The implementation manner of inputting the first question information into the data processing model to obtain the answer information of the first question information can be to input the first question information into the data processing model, and after the data processing model performs data processing, obtain the answer information of the first question information output by the data processing model.
[0176] Exemplarily, input "What is the day after yesterday the day before tomorrow?" into the data processing model to obtain "The current day after yesterday is the day before the 'day after tomorrow' " output by the data processing model.
[0177] Among them, the process of training the data processing model using the enhanced sample set is the same as or similar to the process of Figure 3 training the data processing model in and will not be elaborated here.
[0178] Applying the solution of the embodiment of the present specification, input the first question information into the data processing model obtained by training with the enhanced sample set, and obtain the answer information of the first question information output by the data processing model. The enhanced sample set is obtained by annotating on the basis of the sample data to obtain the original sample set, predicting the layout on the basis of the sample data to obtain the layout sample set, adjusting the layout on the basis of the original sample set and the layout sample set to obtain the target sample set, and performing data augmentation on the basis of the target sample set. This fully ensures the data richness of the enhanced sample set, and further ensures the accuracy of the answer information of the first question information when the data processing model obtained by training with the enhanced sample set performs data processing.
[0179] See Figure 5 , Figure 5 which shows a flowchart of a document understanding method provided by an embodiment of the present specification, specifically including the following steps.
[0180] Step 502: Obtain a target document and second question information for the target document.
[0181] The embodiment of the present specification is applied to an end-side device and / or a cloud-side device with a document understanding function. The following takes the cloud-side device as an example for illustration.
[0182] When there is a need for document understanding, the cloud-side device will obtain a target document and second question information for the target document, so as to input the target document and the second question information for the target document into the data processing model, and obtain the answer information of the second question information output by the data processing model.
[0183] Specifically, the second question information refers to the text corresponding to the question that the front-end user needs to know the corresponding answer for the target document. For example, the second question information can be "How many paragraphs are included in the target document?", "What services will be provided on the plane in the target document?", and so on.
[0184] The manner of obtaining the target document and the second question information for the target document can refer to the manner of obtaining the target data in the above Figure 3 and will not be elaborated here.
[0185] Step 504: Input the target document and the second question information into the data processing model, and obtain the answer information of the second question information, where the data processing model is obtained by training with an enhanced sample set, the enhanced sample set is obtained by performing data augmentation on the target sample set, the target sample set is obtained by performing layout adjustment on the layout sample image of the layout sample set based on the original sample image of the original sample set, the original sample set is obtained by annotating the sample data, and the layout sample set is obtained by performing layout prediction on the sample data.
[0186] The implementation manner of inputting the target document and the second question information into the data processing model to obtain the answer information of the second question information may be to input the target document and the second question information into the data processing model, and after being processed by the data processing model, obtain the answer information of the second question information output by the data processing model. Among them, the answer information is obtained by the data processing model through document understanding of the target document.
[0187] Among them, the process of training the data processing model using the enhanced sample set is the same as or similar to the process of Figure 3 training the data processing model in [reference], and will not be elaborated here.
[0188] Applying the solution of this embodiment of the present specification, input the target document and the second question information into the data processing model trained using the enhanced sample set, and obtain the answer information of the second question information output by the data processing model. The enhanced sample set is obtained by performing annotation on the basis of the sample data to obtain the original sample set, performing layout prediction on the basis of the sample data to obtain the layout sample set, performing layout adjustment on the basis of the original sample set and the layout sample set to obtain the target sample set, and performing data augmentation on the basis of the target sample set, which fully guarantees the data richness of the enhanced sample set. Furthermore, when the data processing model trained using the enhanced sample set performs data processing, the accuracy of the answer information of the second question information is fully guaranteed.
[0189] The following Figure 6 , taking the application of the data processing method provided in this specification in document understanding as an example, further describes the data processing method. Among them, Figure 6 shows the processing procedure flowchart of a data processing method provided by an embodiment of this specification, which specifically includes the following steps.
[0190] Step 602: Obtain a sample retrieval task, where at least one data type of the sample data is carried in the sample retrieval task.
[0191] Step 604: Determine a retrieval strategy for retrieving the sample data based on at least one data type.
[0192] Step 606: Perform data retrieval according to the retrieval strategy to obtain the sample data of each data type.
[0193] Purchase access rights to various databases and obtain a large number of PDF files from them; obtain a large amount of web page data from websites developed by the company or other authorized affiliated companies, and package the text and images in the web pages and save them as MHTML files; purchase access rights to various databases and obtain a large amount of text information from them, clean it to obtain the text content, and generate plain text files; obtain LaTeX source files with unrestricted license types from academic websites.
[0194] Step 608: Based on the sample data, determine the rendering strategy corresponding to the sample data.
[0195] Step 610: Perform rendering annotation based on the rendering strategy to obtain a sample image and a sample title carrying annotation information.
[0196] For the obtained PDF files, we directly use rendering software 1 for rendering and annotation. The final data format includes the title of the obtained file and the bounding boxes of all corresponding texts therein as annotations; for MHTML files, we use rendering software 1 to load the MHTML file for rendering and obtain annotations. The final data format includes the title of the obtained file and the bounding boxes of all corresponding texts therein as annotations; for plain text files, we use a layout prediction model to obtain a sample image output by the layout prediction model; for LaTeX files, we use rendering software 2 to render the LaTeX file to obtain the corresponding PDF file and compile tool annotations. We use the rendering effect image file of each page of the generated PDF file and use the compile tool annotations to obtain the text bounding boxes of each page as annotations.
[0197] Step 612: Integrate the sample images carrying annotation information to obtain an original sample set.
[0198] Step 614: Based on the sample title and the annotation information on the original sample images in the original sample set, train the layout generation model to obtain a trained layout generation model.
[0199] Step 616: Input the sample data and the sample title of the sample data into the layout generation model to obtain a layout sample set of the sample data.
[0200] Step 618: Based on the original sample images, the annotation information on the original sample images, and the annotation information on the layout sample images, perform layout randomization on the layout sample images to obtain a target sample set.
[0201] Step 620: Perform data augmentation on the target sample images and the annotation information on the target sample images to obtain an augmented sample set.
[0202] Step 622: Use the augmented sample set to train the data processing model to obtain a trained data processing model.
[0203] Extract the target enhanced sample image and the annotation information on the target enhanced sample image from the enhanced sample set, input the target enhanced sample image into the data processing model to obtain the predicted annotation information, calculate the loss value based on the annotation information and the predicted annotation information on the target enhanced sample image, adjust the model parameters of the data processing model based on the loss value, and return to the step of extracting the target enhanced sample image and the annotation information on the target enhanced sample image from the enhanced sample set until the training stop condition is reached, and obtain the data processing model after training, where the target enhanced sample image is any enhanced sample image in the enhanced sample set.
[0204] Step 624: Obtain the target document and the second question information for the target document.
[0205] Step 626: Input the target document and the second question information into the data processing model to obtain the answer information for the second question information.
[0206] Applying the solution of the embodiments of this specification, input the target data into the data processing model trained by the enhanced sample set to obtain the processing result output by the data processing model. The enhanced sample set is obtained by performing annotation on the sample data to obtain the original sample set, performing layout prediction on the sample data to obtain the layout sample set, performing layout adjustment on the basis of the original sample set and the layout sample set to obtain the target sample set, and performing data augmentation on the target sample set, which fully guarantees the data richness of the enhanced sample set. Furthermore, when the data processing model trained by the enhanced sample set performs data processing, the accuracy of the processing result is fully guaranteed.
[0207] Corresponding to the above method embodiments, this specification also provides embodiments of a data processing device. Figure 7 The structure diagram of a data processing device provided by an embodiment of this specification is shown. As Figure 7 shown, the device includes:
[0208] The first acquisition module 702 is configured to acquire target data;
[0209] The first processing module 704 is configured to input the target data into the data processing model to obtain the processing result of the target data, where the data processing model is trained by using the enhanced sample set, the enhanced sample set includes enhanced sample images and annotation information on the enhanced sample images, the enhanced sample set is obtained by performing data augmentation on the target sample set, and the target sample set is obtained by performing layout adjustment on the layout sample images of the layout sample set based on the original sample images of the original sample set, and the original sample set is obtained by performing annotation on the sample data, and the layout sample set is obtained by performing layout prediction on the sample data.
[0210] Optionally, the data processing device further includes a layout adjustment module configured to determine sample data and obtain an original sample set and a layout sample set of the sample data, where the original sample set includes an original sample image and annotation information on the original sample image, and the layout sample set includes a layout sample image and annotation information on the layout sample image; based on the original sample image, the annotation information on the original sample image, and the annotation information on the layout sample image, perform layout adjustment on the layout sample image to obtain a target sample set, where the target sample set includes a target sample image and annotation information on the target sample image.
[0211] Optionally, the layout adjustment module is further configured to obtain a sample retrieval task, where at least one data type of the sample data is carried in the sample retrieval task; based on the at least one data type, determine a retrieval strategy for retrieving the sample data; and perform data retrieval according to the retrieval strategy to obtain the sample data of each data type.
[0212] Optionally, the layout adjustment module is further configured to determine a rendering strategy corresponding to the sample data based on the sample data; use the rendering strategy to render the sample data to obtain a sample image and a sample title of the sample data; annotate the sample image to obtain a sample image carrying annotation information; determine the original sample set based on the sample image carrying annotation information; and perform layout prediction on the sample data according to the sample title and the original sample set to obtain the layout sample set of the sample data.
[0213] Optionally, the layout adjustment module is further configured to train a layout generation model according to the sample title and the annotation information on the original sample image in the original sample set to obtain a trained layout generation model; and input the sample data and the sample title of the sample data into the layout generation model to obtain the layout sample set of the sample data.
[0214] Optionally, the layout adjustment module is further configured to determine original text content based on the original sample image and the annotation information on the original sample image; determine layout text content based on the layout sample image and the annotation information on the layout sample image; replace the layout text content in the layout sample image with the original text content to obtain a target sample image; and integrate the target sample image carrying annotation information to obtain a target sample set.
[0215] Optionally, the data processing device further includes a training module configured to perform data augmentation on the target sample image and the annotation information on the target sample image to obtain an augmented sample set, where the augmented sample set includes an augmented sample image and annotation information on the augmented sample image; and use the augmented sample set to train an initial data processing model to obtain a trained data processing model.
[0216] Optionally, the training module is further configured to obtain at least two data augmentation methods and the execution order of the at least two data augmentation methods; according to the execution order, perform data augmentation on the target sample image and the annotation information on the target sample image based on the at least two data augmentation methods to obtain an augmented sample set.
[0217] Optionally, the training module is further configured to determine the augmentation stages of the at least two data augmentation methods according to the execution order, where the augmentation stages at least include a first augmentation stage and a second augmentation stage; for the first augmentation stage, perform data augmentation on the target sample image and the annotation information on the target sample image based on the data augmentation method corresponding to the first augmentation stage to obtain the first augmented sample image in the first augmentation stage and the annotation information on the first augmented sample image, where the first augmentation stage is the augmentation stage ranked first in the execution order; for the second augmentation stage, perform data augmentation on the previous augmented sample image in the previous augmentation stage of the second augmentation stage and the annotation information on the previous augmented sample image based on the data augmentation method corresponding to the second augmentation stage to obtain the second augmented sample image in the second augmentation stage and the annotation information on the second augmented sample image, where the second augmentation stage is the augmentation stage other than the first augmentation stage.
[0218] Optionally, the data augmentation methods include at least two of the following: basic data augmentation method, perspective transformation data augmentation method, mask data augmentation method; the training module is further configured to obtain a filter to be superimposed based on the basic data augmentation method, and superimpose the target sample image and the annotation information on the target sample image with the filter to be superimposed to obtain a basic augmented sample image and the annotation information on the basic augmented sample image; based on the perspective transformation data augmentation method, select target perspective transformation parameters, and perform perspective transformation on the basic augmented sample image and the annotation information on the basic augmented sample image based on the target perspective transformation parameters to obtain a transformed augmented sample image and the annotation information on the transformed augmented sample image; based on the mask data augmentation method, obtain a foreground image and a background image, and superimpose the foreground image, the background image, the transformed augmented sample image and the annotation information on the transformed augmented sample image to obtain a mask augmented sample image and the annotation information on the mask augmented sample image; determine the augmented sample set according to the mask augmented sample image and the annotation information on the mask augmented sample image.
[0219] Applying the solution of the embodiment of this specification, input the target data into the data processing model obtained by training with the enhanced sample set, and obtain the processing result output by the data processing model. The enhanced sample set is obtained by performing annotation on the sample data to obtain the original sample set, performing layout prediction on the sample data to obtain the layout sample set, performing layout adjustment on the basis of the original sample set and the layout sample set to obtain the target sample set, and performing data augmentation on the target sample set. This fully ensures the data richness of the enhanced sample set, and further ensures the accuracy of the processing result when the data processing model obtained by training with the enhanced sample set performs data processing.
[0220] The above is a schematic solution of a data processing device according to this embodiment. It should be noted that the technical solution of this data processing device and the technical solution of the above data processing method belong to the same concept. For the details not described in detail in the technical solution of the data processing device, reference can be made to the description of the technical solution of the above data processing method.
[0221] Corresponding to the above method embodiment, this specification also provides an embodiment of a question-answering processing device. Figure 8 It shows a schematic structural diagram of a question-answering processing device provided by an embodiment of this specification. As Figure 8 shown, the device includes:
[0222] A second acquisition module 802, configured to acquire first question information;
[0223] A second processing module 804, configured to input the first question information into the data processing model to obtain answer information of the first question information, where the data processing model is obtained by training with an enhanced sample set, the enhanced sample set includes enhanced sample images and annotation information on the enhanced sample images, the enhanced sample set is obtained by performing data augmentation on the target sample set, the target sample set is obtained by performing layout adjustment on the layout sample images of the layout sample set based on the original sample images of the original sample set, the original sample set is obtained by performing annotation on the sample data, and the layout sample set is obtained by performing layout prediction on the sample data.
[0224] Applying the solution of the embodiment of this specification, input the first problem information into the data processing model obtained by training with the enhanced sample set, and obtain the answer information of the first problem information output by the data processing model. The enhanced sample set is obtained by performing annotation on the basis of sample data to obtain the original sample set, performing layout prediction on the basis of sample data to obtain the layout sample set, performing layout adjustment on the basis of the original sample set and the layout sample set to obtain the target sample set, and performing data enhancement on the basis of the target sample set. This fully ensures the data richness of the enhanced sample set, and further ensures the accuracy of the answer information of the first problem information when the data processing model obtained by training with the enhanced sample set performs data processing.
[0225] The above is a schematic solution of a question-answering processing device according to this embodiment. It should be noted that the technical solution of this question-answering processing device and the technical solution of the above question-answering processing method belong to the same concept. For the details not described in detail in the technical solution of the question-answering processing device, reference can be made to the description of the technical solution of the above question-answering processing method.
[0226] Corresponding to the above method embodiment, this specification also provides an embodiment of a document understanding device. Figure 9 It shows a schematic structural diagram of a document understanding device provided by an embodiment of this specification. As Figure 9 shown, the device includes:
[0227] A third acquisition module 902, configured to acquire a target document and second problem information for the target document;
[0228] A third processing module 904, configured to input the target document and the second problem information into the data processing model to obtain the answer information of the second problem information, where the data processing model is obtained by training with an enhanced sample set, the enhanced sample set includes enhanced sample images and annotation information on the enhanced sample images, the enhanced sample set is obtained by performing data enhancement on the target sample set, the target sample set is obtained by performing layout adjustment on the layout sample images of the layout sample set based on the original sample images of the original sample set, the original sample set is obtained by performing annotation on sample data, and the layout sample set is obtained by performing layout prediction on sample data.
[0229] Applying the solution of the embodiments of this specification, input the target document and the second question information into the data processing model obtained by training with the enhanced sample set, and obtain the answer information of the second question information output by the data processing model. The enhanced sample set is obtained by performing annotation on the basis of sample data to obtain the original sample set, performing layout prediction on the basis of sample data to obtain the layout sample set, performing layout adjustment on the basis of the original sample set and the layout sample set to obtain the target sample set, and performing data enhancement on the basis of the target sample set. This fully ensures the data richness of the enhanced sample set, and further ensures the accuracy of the answer information of the second question information when the data processing model obtained by training with the enhanced sample set performs data processing.
[0230] The above is a schematic solution of a document understanding device according to this embodiment. It should be noted that the technical solution of this document understanding device and the technical solution of the above document understanding method belong to the same concept. For the details not described in detail in the technical solution of this document understanding device, reference can be made to the description of the technical solution of the above document understanding method.
[0231] Figure 10 The structural block diagram of a computing device provided by an embodiment of this specification is shown. The components of the computing device 1000 include but are not limited to a memory 1010 and a processor 1020. The processor 1020 is connected to the memory 1010 through a bus 1030, and the database 1050 is used to store data.
[0232] The computing device 1000 further includes an access device 1040, which enables the computing device 1000 to communicate via one or more networks 1060. Examples of such networks include the Public Switched Telephone Network (PSTN), Local Area Network (LAN), Wide Area Network (WAN), Personal Area Network (PAN), or a combination of communication networks such as the Internet. The access device 540 may include one or more of any type of wired or wireless network interfaces (e.g., a network interface card (NIC)), such as an IEEE 802.11 Wireless Local Area Network (WLAN) wireless interface, a Worldwide Interoperability for Microwave Access (Wi-MAX) interface, an Ethernet interface, a Universal Serial Bus (USB) interface, a cellular network interface, a Bluetooth interface, or a Near Field Communication (NFC) interface.
[0233] In one embodiment of the present specification, the above components of the computing device 1000 and Figure 10 other components not shown may also be connected to each other, for example, via a bus. It should be understood that Figure 10 the block diagram of the computing device shown is for illustrative purposes only and is not a limitation on the scope of the present specification. Those skilled in the art can add or replace other components as needed.
[0234] The computing device 1000 can be any type of stationary or mobile computing device, including a mobile computer or mobile computing device (e.g., a tablet computer, a personal digital assistant, a laptop computer, a notebook computer, a netbook, etc.), a mobile phone (e.g., a smartphone), a wearable computing device (e.g., a smartwatch, smart glasses, etc.) or other types of mobile devices, or a stationary computing device such as a desktop computer or a Personal Computer (PC). The computing device 1000 can also be a mobile or stationary server.
[0235] Wherein, the processor 1020 is used to execute the following computer program / instructions, and when the computer program / instructions are executed by the processor, the steps of the above data processing method, question-answering processing method, and document understanding method are implemented.
[0236] The above is a schematic solution of a computing device according to this embodiment. It should be noted that the technical solution of this computing device and the technical solutions of the above data processing method, question and answer processing method, and document understanding method belong to the same concept. For the details not described in detail in the technical solution of the computing device, reference can be made to the descriptions of the technical solutions of the above data processing method, question and answer processing method, and document understanding method.
[0237] An embodiment of this specification also provides a computer-readable storage medium, which stores computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the above data processing method, question and answer processing method, and document understanding method are implemented.
[0238] The above is a schematic solution of a computer-readable storage medium according to this embodiment. It should be noted that the technical solution of this storage medium and the technical solutions of the above data processing method, question and answer processing method, and document understanding method belong to the same concept. For the details not described in detail in the technical solution of the storage medium, reference can be made to the descriptions of the technical solutions of the above data processing method, question and answer processing method, and document understanding method.
[0239] An embodiment of this specification also provides a computer program product, including computer programs / instructions. When the computer programs / instructions are executed by a processor, the steps of the above data processing method, question and answer processing method, and document understanding method are implemented.
[0240] The above is a schematic solution of a computer program product according to this embodiment. It should be noted that the technical solution of this computer program product and the technical solutions of the above data processing method, question and answer processing method, and document understanding method belong to the same concept. For the details not described in detail in the technical solution of the computer program product, reference can be made to the descriptions of the technical solutions of the above data processing method, question and answer processing method, and document understanding method.
[0241] The above describes a specific embodiment of this specification. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recited in the claims can be executed in a different order than in the embodiments and still achieve the desired results. Additionally, the processes depicted in the figures do not necessarily require the specific order or sequential order shown to achieve the desired results. In certain embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0242] The computer instructions include computer program code, which may be in the form of source code, object code, executable files, or some intermediate forms, etc. The computer-readable medium may include: any entity or device capable of carrying the computer program code, recording media, USB flash drives, mobile hard disks, magnetic disks, optical disks, computer memories, read-only memories (ROMs), random access memories (RAMs), electrical carrier signals, telecommunication signals, and software distribution media, etc. It should be noted that the content included in the computer-readable medium can be appropriately increased or decreased according to the requirements of patent practice. For example, in some regions, according to patent practice, the computer-readable medium does not include electrical carrier signals and telecommunication signals.
[0243] It should be noted that for the foregoing method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the embodiments of this specification are not limited by the described order of actions, because according to the embodiments of this specification, some steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the embodiments of this specification.
[0244] In the above embodiments, the descriptions of the various embodiments have their own emphases. For the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0245] The preferred embodiments of this specification disclosed above are only used to help explain this specification. The alternative embodiments do not describe all the details in detail, nor do they limit the invention to the specific embodiments described. Obviously, many modifications and changes can be made according to the content of the embodiments of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the embodiments of this specification, so that those skilled in the art can understand and utilize this specification well. This specification is only limited by the claims and their full scope and equivalents.
Claims
1. A data processing method, comprising: Obtaining target data; Inputting the target data into a data processing model to obtain a processing result of the target data, wherein the data processing model is trained using an enhanced sample set, the enhanced sample set is obtained by performing data enhancement on a target sample set, the target sample set is obtained by performing layout adjustment on layout sample images of a layout sample set based on original sample images of an original sample set, the original sample set is obtained by annotating sample data, and the layout sample set is obtained by performing layout prediction on the sample data.
2. The data processing method according to claim 1, further comprising: Determining the sample data, and obtaining the original sample set and the layout sample set of the sample data, wherein the original sample set includes the original sample images and annotation information on the original sample images, and the layout sample set includes the layout sample images and annotation information on the layout sample images; Performing layout adjustment on the layout sample images based on the original sample images, the annotation information on the original sample images, and the annotation information on the layout sample images to obtain the target sample set, wherein the target sample set includes target sample images and annotation information on the target sample images.
3. The data processing method according to claim 2, wherein the determining the sample data comprises: Obtaining a sample retrieval task, wherein at least one data type of the sample data is carried in the sample retrieval task; Determining a retrieval strategy for retrieving the sample data based on the at least one data type, wherein the retrieval strategy is a retrieval method for performing data retrieval; Performing data retrieval according to the retrieval strategy to obtain the sample data of each of the data types.
4. The data processing method according to claim 2, wherein the obtaining the original sample set and the layout sample set of the sample data comprises: Determining a rendering strategy corresponding to the sample data based on the sample data, wherein the rendering strategy is a strategy method for rendering data; Rendering the sample data using the rendering strategy to obtain sample images and sample titles of the sample data; Annotating the sample images to obtain sample images carrying annotation information; Determining the original sample set based on the sample images carrying annotation information; Performing layout prediction on the sample data according to the sample titles and the original sample set to obtain the layout sample set of the sample data.
5. The data processing method according to claim 4, wherein the performing layout prediction on the sample data according to the sample titles and the original sample set to obtain the layout sample set of the sample data comprises: Training a layout generation model according to the sample titles and the annotation information on the original sample images in the original sample set to obtain the trained layout generation model; Inputting the sample data and the sample titles of the sample data into the trained layout generation model to obtain the layout sample set of the sample data.
6. The data processing method according to claim 2, wherein the layout adjustment of the layout sample image based on the original sample image, the annotation information on the original sample image, and the annotation information on the layout sample image to obtain the target sample set includes: Determining the original text content based on the original sample image and the annotation information on the original sample image; Determining the layout text content based on the layout sample image and the annotation information on the layout sample image; Replacing the layout text content in the layout sample image with the original text content to obtain a target sample image; Integrating the target sample images with annotation information to obtain the target sample set.
7. The data processing method according to claim 2 further includes: Performing data augmentation on the target sample image and the annotation information on the target sample image to obtain the augmented sample set, where the augmented sample set includes augmented sample images and the annotation information on the augmented sample images; Using the augmented sample set to train an initial data processing model to obtain the trained data processing model.
8. The data processing method according to claim 7, wherein the performing data augmentation on the target sample image and the annotation information on the target sample image to obtain the augmented sample set includes: Obtaining at least two data augmentation methods and the execution order of the at least two data augmentation methods; According to the execution order, performing data augmentation on the target sample image and the annotation information on the target sample image based on the at least two data augmentation methods to obtain the augmented sample set.
9. The data processing method according to claim 8, wherein the performing data augmentation on the target sample image and the annotation information on the target sample image based on the at least two data augmentation methods according to the execution order to obtain the augmented sample set includes: According to the execution order, determining the augmentation stages of the at least two data augmentation methods, where the augmentation stages at least include a first augmentation stage and a second augmentation stage; For the first augmentation stage, performing data augmentation on the target sample image and the annotation information on the target sample image based on the data augmentation method corresponding to the first augmentation stage to obtain the first augmented sample image of the first augmentation stage and the annotation information on the first augmented sample image, where the first augmentation stage is the augmentation stage ranked first in the execution order; For the second augmentation stage, performing data augmentation on the previous augmented sample image of the previous augmentation stage of the second augmentation stage and the annotation information on the previous augmented sample image based on the data augmentation method corresponding to the second augmentation stage to obtain the second augmented sample image of the second augmentation stage and the annotation information on the second augmented sample image, where the second augmentation stage is the augmentation stage other than the first augmentation stage.
10. The data processing method according to claim 8, wherein the data augmentation methods include at least two of the following: basic data augmentation method, perspective transformation data augmentation method, and masking data augmentation method; Performing data augmentation on the target sample image and the annotation information on the target sample image based on the at least two data augmentation methods according to the execution order to obtain the augmented sample set, including: Based on the basic data augmentation method, obtaining a filter to be superimposed, and superimposing the target sample image and the annotation information on the target sample image with the filter to be superimposed to obtain a basic augmented sample image and the annotation information on the basic augmented sample image; Based on the perspective transformation data augmentation method, selecting target perspective transformation parameters, and performing perspective transformation on the basic augmented sample image and the annotation information on the basic augmented sample image based on the target perspective transformation parameters to obtain a transformed augmented sample image and the annotation information on the transformed augmented sample image; Based on the masking data augmentation method, obtaining a foreground image and a background image, and superimposing the foreground image, the background image, the transformed augmented sample image, and the annotation information on the transformed augmented sample image to obtain a masked augmented sample image and the annotation information on the masked augmented sample image; Determining the augmented sample set according to the masked augmented sample image and the annotation information on the masked augmented sample image.
11. A question-and-answer processing method, including: Obtaining first question information; Inputting the first question information into a data processing model to obtain answer information for the first question information, wherein the data processing model is trained using an augmented sample set, the augmented sample set is obtained by performing data augmentation on a target sample set, the target sample set is obtained by performing layout adjustment on layout sample images of a layout sample set based on original sample images of an original sample set, the original sample set is obtained by annotating sample data, and the layout sample set is obtained by performing layout prediction on the sample data.
12. A document understanding method, including: Obtaining a target document and second question information for the target document; Inputting the target document and the second question information into a data processing model to obtain answer information for the second question information, wherein the data processing model is trained using an augmented sample set, the augmented sample set is obtained by performing data augmentation on a target sample set, the target sample set is obtained by performing layout adjustment on layout sample images of a layout sample set based on original sample images of an original sample set, the original sample set is obtained by annotating sample data, and the layout sample set is obtained by performing layout prediction on the sample data.
13. A data processing system, the data processing system includes an edge device and a cloud device; The edge device is used to initiate a data processing request, where The target data is carried in the data processing request; The cloud-side device is configured to obtain the target data in response to the data processing request, input the target data into a data processing model, and obtain a processing result of the target data. The data processing model is trained using an augmented sample set, which is obtained by augmenting a target sample set. The target sample set is obtained by performing layout adjustment on layout sample images of a layout sample set based on original sample images of an original sample set. The original sample set is obtained by annotating sample data, and the layout sample set is obtained by performing layout prediction on the sample data.
14. A computing device, comprising: a memory and a processor; The memory is configured to store computer programs / instructions, and the processor is configured to execute the computer programs / instructions. When the computer programs / instructions are executed by the processor, the steps of the method according to any one of claims 1 to 12 are implemented.
15. A computer-readable storage medium storing computer programs / instructions, which when executed by a processor implement the steps of the method according to any one of claims 1 to 12.
16. A computer program product comprising computer programs / instructions, which when executed by a processor implement the steps of the method according to any one of claims 1 to 12.