Data annotation method and device, multi-modal large model training method and device and data processing method and device

Through the automatic labeling method of multimodal large model, OCR and target recognition technology are used to generate rich comprehensive description text, solving the problem of inefficient manual labeling in fine-tuning of multimodal large model, and achieving efficient and accurate labeling and model training.

CN120256948APending Publication Date: 2025-07-04BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510215727.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The prior art requires manual annotation during the fine-tuning of multimodal large model, resulting in high labor and time costs and low efficiency, and difficult to ensure accuracy.

Method used

The multimodal large model automatic labeling method is adopted to generate comprehensive description text by obtaining sample images and the first prompt text, and inputting a multimodal large model to generate labeling text. The input content is enriched using OCR and target recognition technology, and the model is trained in combination with text alignment loss.

Benefits of technology

It reduces labor and time costs, improves labeling efficiency and accuracy, and enhances the special domain knowledge learning ability of multimodal large models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256948A_ABST
    Figure CN120256948A_ABST
Patent Text Reader

Abstract

The invention provides a data annotation method and device, a multi-modal large model training method and device and a data processing method and device, relates to the artificial intelligence field of computer vision, deep learning, large models and the like, and can be applied to scenes of content generation and the like based on artificial intelligence. The data annotation method comprises the steps of obtaining a sample image and a corresponding first prompt text, wherein the first prompt text comprises first demand description information proposed for the sample image; generating a comprehensive description text according to the sample image and the first prompt text; and inputting the sample image and the comprehensive description text into a multi-modal large model to obtain an output annotation text, the annotation text comprising first response information corresponding to the first demand description information.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, particularly to technical fields such as computer vision, deep learning, and large models, and specifically relates to data annotation, multi-modal large model training, and data processing methods and devices. Background Art

[0002] Currently, the development of deep learning is gradually evolving towards large models. Especially for multi-modal large models, due to their ability to process multiple types of data simultaneously, they have been increasingly widely used. In practical applications, for a pre-trained multi-modal large model, it is usually necessary to fine-tune it according to different application scenarios to achieve better usage effects. Summary of the Invention

[0003] The present disclosure provides data annotation, multi-modal large model training, and data processing methods and devices.

[0004] A data annotation method includes:

[0005] Obtain a sample image and a corresponding first prompt text, where the first prompt text includes first requirement description information proposed for the sample image;

[0006] Generate a comprehensive description text based on the sample image and the first prompt text;

[0007] Input the sample image and the comprehensive description text into the multi-modal large model to obtain an output annotation text, where the annotation text includes a first response information corresponding to the first requirement description information.

[0008] A multi-modal large model training method includes:

[0009] Obtain training samples, where the training samples include: a sample image, a first prompt text, and an annotation text. The first prompt text includes first requirement description information proposed for the sample image, and the annotation text includes a first response information corresponding to the first requirement description information;

[0010] Train the multi-modal large model according to the training samples.

[0011] A data processing method includes:

[0012] Obtain an image to be processed and a corresponding second prompt text, where the second prompt text includes second requirement description information proposed for the image to be processed;

[0013] Input the to-be-processed image and the second prompt text into the multimodal large model to obtain the output second generated text, where the second generated text includes the third response information corresponding to the second requirement description information. The multimodal large model is trained based on the obtained training samples, and the training samples include: sample images, first prompt texts, and annotation texts. The first prompt text includes the first requirement description information proposed for the sample image, and the annotation text includes the first response information corresponding to the first requirement description information.

[0014] A data annotation device includes: an information acquisition module, a text generation module, and a data annotation module;

[0015] The information acquisition module is used to acquire a sample image and the corresponding first prompt text, where the first prompt text includes the first requirement description information proposed for the sample image;

[0016] The text generation module is used to generate a comprehensive description text based on the sample image and the first prompt text;

[0017] The data annotation module is used to input the sample image and the comprehensive description text into the multimodal large model to obtain the output annotation text, where the annotation text includes the first response information corresponding to the first requirement description information.

[0018] A multimodal large model training device includes: a sample acquisition module and a model training module;

[0019] The sample acquisition module is used to acquire training samples, and the training samples include: sample images, first prompt texts, and annotation texts. The first prompt text includes the first requirement description information proposed for the sample image, and the annotation text includes the first response information corresponding to the first requirement description information;

[0020] The model training module is used to train the multimodal large model according to the training samples.

[0021] A data processing device includes: a data acquisition module and a data processing module;

[0022] The data acquisition module is used to acquire a to-be-processed image and the corresponding second prompt text, where the second prompt text includes the second requirement description information proposed for the to-be-processed image;

[0023] The data processing module is configured to input the image to be processed and the second prompt text into a multimodal large model to obtain an output second generated text, where the second generated text includes a third response message corresponding to the second requirement description information. The multimodal large model is trained based on acquired training samples, and the training samples include: sample images, first prompt texts, and annotation texts. The first prompt texts include first requirement description information proposed for the sample images, and the annotation texts include first response messages corresponding to the first requirement description information.

[0024] An electronic device includes:

[0025] At least one processor; and

[0026] A memory communicatively connected to the at least one processor; wherein,

[0027] The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method as described above.

[0028] A non-transitory computer-readable storage medium storing computer instructions for causing a computer to execute the method as described above.

[0029] A computer program product includes computer programs / instructions that, when executed by a processor, implement the method as described above.

[0030] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understandable through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0032] Figure 1 is a flowchart of an embodiment of the data annotation method described in the present disclosure;

[0033] Figure 2 is a schematic diagram of the process of performing OCR recognition on a sample image described in the present disclosure;

[0034] Figure 3 is a schematic diagram of the process of performing object recognition on a sample image described in the present disclosure;

[0035] Figure 4 is a flowchart of an embodiment of the multimodal large model training method described in the present disclosure;

[0036] Figure 5 Flowchart of the embodiment of the data processing method described in this disclosure;

[0037] Figure 6 Schematic diagram of the composition structure of Embodiment 600 of the data annotation device described in this disclosure;

[0038] Figure 7 Schematic diagram of the composition structure of Embodiment 700 of the multi-modal large model training device described in this disclosure;

[0039] Figure 8 Schematic diagram of the composition structure of Embodiment 800 of the data processing device described in this disclosure;

[0040] Figure 9 Schematic block diagram of an electronic device 900 that can be used to implement the embodiments of this disclosure is shown. Detailed implementation manners

[0041] The following describes exemplary embodiments of this disclosure with reference to the accompanying drawings. Various details of the embodiments of this disclosure are included to assist understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of this disclosure. Similarly, for clarity and conciseness, descriptions of well-known functions and structures are omitted below.

[0042] In addition, it should be understood that the term "and / or" herein is merely a description of the association relationship of associated objects, indicating that there can be three relationships. For example, A and / or B can represent: A exists alone, A and B exist simultaneously, and B exists alone. In addition, the character " / " herein generally represents an "or" relationship between the associated objects before and after.

[0043] Figure 1 Flowchart of the embodiment of the data annotation method described in this disclosure. As Figure 1 shown, it includes the following specific implementation manners.

[0044] In step 101, a sample image and a corresponding first prompt text are obtained, and the first prompt text includes first demand description information proposed for the sample image.

[0045] In step 102, a comprehensive description text is generated according to the sample image and the first prompt text.

[0046] In step 103, the sample image and the comprehensive description text are input into a multi-modal large model to obtain an output annotation text, and the annotation text includes a first response information corresponding to the first demand description information.

[0047] Currently, when fine-tuning a multi-modal large model, it is necessary to generate training samples by means of manual annotation, which requires a large amount of human and time costs, is inefficient, and it is also difficult to guarantee accuracy.

[0048] By adopting the solution described in the above method embodiments, a multi-modal large model can be used to generate the required annotation text, that is, the manual annotation method is transformed into an automatic annotation method, thus saving human and time costs and improving the processing efficiency. Moreover, for the obtained sample image and the first prompt text, a comprehensive description text can be first generated by combining the sample image and the first prompt text, and then an annotation text can be further generated by combining the sample image and the comprehensive description text, thereby enriching the input content of the multi-modal large model and further improving the accuracy of the generated annotation text, etc.

[0049] There is no limitation on how to obtain the sample image and the corresponding first prompt text. For example, the images actually input by users during the use of online products and the corresponding prompt texts can be obtained, etc.

[0050] For the convenience of distinction, the prompt text corresponding to the sample image is called the first prompt text, and similar situations will not be elaborated hereinafter.

[0051] According to the different multi-modal tasks, the content of the first requirement description information included in the first prompt text can also be different. Multi-modal tasks can include image-text question answering, visual description, creative generation, etc. For example, for an image-text question answering multi-modal task, the first prompt text can be: How to lower the seat height of this bicycle in the image; for a visual description multi-modal task, the first prompt text can be: Please describe this image in detail; for a creative generation multi-modal task, the first prompt text can be: Please write a 200-word composition based on the image.

[0052] After that, a comprehensive description text can be generated according to the sample image and the first prompt text. In some embodiments of the present disclosure, the sample image can be subjected to optical character recognition (OCR, Optical Character Recognition) to obtain a text recognition result, and the sample image can be subjected to object recognition to obtain an object recognition result. Then, a comprehensive description text can be generated according to the text recognition result, the object recognition result, and the first prompt text.

[0053] It can be seen that the comprehensive description text includes both the text information in the sample image, the object information in the sample image, and the first prompt text, that is, compared with the first prompt text, the comprehensive description text has richer image information, thus improving the accuracy of subsequent processing based on the comprehensive description text, etc.

[0054] In some embodiments of the present disclosure, the method for performing OCR recognition on a sample image to obtain a text recognition result may include: using an OCR classification model to determine a target classification corresponding to the sample image, and using the OCR model corresponding to the target classification to perform OCR recognition on the sample image to obtain a text recognition result.

[0055] For example, the target classification corresponding to the sample image may be determined from M vertical category classifications, where M is a positive integer greater than 1, and different vertical category classifications respectively correspond to their own OCR models. The specific value of M may be determined according to actual needs, and which specific vertical category classifications are included in the M vertical category classifications may also be determined according to actual needs.

[0056] Figure 2 It is a schematic diagram of the process of performing OCR recognition on the sample image according to the present disclosure. As Figure 2 shown, for the sample image, it is possible to first classify it using an OCR classification model, that is, to determine the target classification corresponding to the sample image from M vertical category classifications. The number of target classifications is usually 1. Further, the sample image can be input into the OCR model corresponding to the target classification. Assuming that the value of M is 7, that is, it includes 7 parallel OCR models, namely: a general OCR model, a financial OCR model, a card and certificate OCR model, an education OCR model, a table OCR model, an office OCR model, and a medical OCR model. The corresponding vertical category classifications are general, financial, card and certificate, education, table, office, and medical respectively. Assuming that the target classification is education, then the sample image can be input into the education OCR model to obtain the output text recognition result, that is, the text information extracted from the sample image by the education OCR model. Each OCR model can be a pre-trained model.

[0057] Correspondingly, through the above processing method, it is possible to use an OCR model suitable for the sample image to perform OCR recognition on the sample image, thereby improving the accuracy of the recognition result and so on.

[0058] In some embodiments of the present disclosure, the method for performing target recognition on a sample image to obtain a target recognition result may include: respectively using N target recognition models to perform target recognition on the sample image, where N is a positive integer greater than 1, and generating a target recognition result by combining the recognition results of the N target recognition models.

[0059] Among them, different target recognition models may respectively correspond to different vertical category classifications. In addition, the specific value of N may be determined according to actual needs, and which specific vertical category classifications are included in the N vertical category classifications may also be determined according to actual needs.

[0060] Figure 3 It is a schematic diagram of the process of performing target recognition on the sample image according to the present disclosure. As Figure 3As shown, it may include multiple parallel target recognition models, such as general target recognition models, celebrity recognition models, animal recognition models, plant recognition models, vehicle recognition models, road sign recognition models, and logo recognition models. The corresponding vertical category classifications are: general target recognition, celebrity recognition, animal recognition, plant recognition, vehicle recognition, road sign recognition, and logo recognition. The sample image can be input into each target recognition model in parallel to obtain the recognition results of each target recognition model respectively. The recognition results may include the type of the recognized target and the coordinate frame, etc. For example: type: goat, coordinates: [31, 20, 90, 110]. Further, the N recognition results can be fused to obtain the target recognition result. For example, the target recognition result may include: there is a goat at the upper left corner of the image, and there is a poplar tree at the middle of the image, etc. Each target recognition model can be a pre-trained model.

[0061] In the above processing method, the target recognition models with different vertical category classifications are respectively used to perform target recognition on the sample image, and the recognition results of each target recognition model are integrated to generate the target recognition result, thereby improving the comprehensiveness and accuracy of the target recognition result, etc.

[0062] After respectively obtaining the text recognition result and the target recognition result of the sample image, the comprehensive description text can be generated according to the text recognition result, the target recognition result, and the first prompt text. In some embodiments of the present disclosure, the text recognition result, the target recognition result, and the first prompt text can be spliced according to a predetermined splicing method to obtain the comprehensive description text.

[0063] There is no limitation on how to splice. For example, the text recognition result, the target recognition result, and the first prompt text can be spliced according to a pre-set splicing template. For example, the comprehensive description text obtained after splicing may include: # The text in the image includes: aquarium; # The target recognition information in the image includes: puffer fish (specific location information may also be included); # Prompt text: Please describe this image.

[0064] The above splicing method is simple and convenient to implement, thereby improving the generation efficiency of the comprehensive description text, etc.

[0065] Further, the sample image and the comprehensive description text can be input into the multi-modal large model to obtain the annotation text output by the multi-modal large model. The annotation text may include the first response information corresponding to the first requirement description information.

[0066] The multimodal large model may include an image encoder and a text decoder. The image encoder may be used to extract image features from input image data, and the text decoder may be used to extract text features from input text data, and the output result of the multimodal large model may be generated based on the image feature extraction results and the text feature extraction results. The input image data may be a sample image, the input text data may be a comprehensive description text, and the output result may be annotated text.

[0067] For example, when the first prompt text is "Please describe this image", the annotation text generated by the multimodal large model can be: This image shows a unique marine creature with a colorful coral reef in the background. The main body is a transparent creature with yellow stripes. The shape is similar to an oval balloon with small white spots on the surface. The corals in the background are red, green and orange, adding to the beauty of the picture. For another example, the annotation text generated by the multimodal large model can be: This image shows a pufferfish. The pufferfish has a round body with stripes and spots on the surface. The colors are mainly yellow and white. Some corals and marine plants can be seen in the background. The colors are rich, including red, green and orange. The pufferfish is in the foreground and the background is slightly blurred, highlighting the details of the pufferfish.

[0068] In the above processing method, since a comprehensive description text containing richer image information is used instead of the first prompt text as the input of the multimodal large model, the multimodal large model can better identify certain targets that require special domain knowledge, thereby improving the accuracy of the annotation text generated by the multimodal large model.

[0069] After obtaining the annotated text, the sample image, the first prompt text and the annotated text can be used to form a training sample. When a sufficient number of training samples are obtained, the training samples can also be used to train the multimodal large model.

[0070] Accordingly, Figure 4 Flow chart of an embodiment of the multimodal large model training method described in the present disclosure. Figure 4 As shown, the following specific implementation methods are included.

[0071] In step 401, a training sample is obtained, the training sample includes: a sample image, a first prompt text and annotated text, the first prompt text includes first requirement description information proposed for the sample image, and the annotated text includes first response information corresponding to the first requirement description information.

[0072] In step 402, the multimodal large model is trained according to the training samples.

[0073] The annotation text may be Figure 1 The annotation text generated by the method shown.

[0074] By training the multi-modal large model, the performance of the multi-modal large model can be improved, and it can learn more knowledge in special fields, etc.

[0075] Among them, when training the multi-modal large model, the method of predicting the next token can be adopted, that is, using the previously predicted token to predict the next token, and the token is discretized text.

[0076] In addition, in some embodiments of the present disclosure, the sample image and the first prompt text can be input into the multi-modal large model to obtain the output first generated text, and the first generated text includes the second response information corresponding to the first demand description information. Then, the text alignment loss between the first generated text and the annotated text can be obtained, and further, the parameters of the multi-modal large model can be updated according to the text alignment loss.

[0077] Input the sample image and the first prompt text into the multi-modal large model, use the multi-modal large model to generate a new generated text (i.e., the first generated text), and align the new generated text with the annotated text, so that the multi-modal large model can be trained to generate response information that is the same as or similar to the response information obtained after adding additional special field knowledge without adding additional special field knowledge.

[0078] The second response information included in the first generated text may be the same as or different from the first response information included in the annotated text.

[0079] The text alignment loss between the first generated text and the annotated text can be obtained. In some embodiments of the present disclosure, according to a preset text alignment loss function, such as the Cross-Entropy Loss function, the text alignment loss between the first generated text and the annotated text can be calculated.

[0080] Among them, the calculated text alignment loss may be:

[0081] L1 = -∑ t logP θ (x t+1 |x t:1 ); (1)

[0082] Among them, x t:1 represents the token that has been predicted from the initial moment to the current moment, x t+1 represents the next token to be predicted, and P θ (x t+1 |x t:1 ) represents x t+1The corresponding text alignment loss is used to train the multi-modal large model by minimizing the total text alignment loss.

[0083] The cross-entropy loss function is a commonly used loss function in deep learning. Using this loss function to calculate the text alignment loss can improve the calculation efficiency and the accuracy of the calculation results, etc.

[0084] According to the text alignment loss, the parameters of the multi-modal large model can be updated. In some embodiments of the present disclosure, the multi-modal large model may include: an image encoder and a text decoder. The image encoder can be used to extract image features from the input image data, and the text decoder can be used to extract text features from the input text data, and can generate the output result of the multi-modal large model according to the image feature extraction result and the text feature extraction result. Among them, the input image data can be a sample image, the input text data can be a first prompt text, and the output result can be a first generated text. Correspondingly, updating the parameters of the multi-modal large model may include: updating the parameters of the image encoder and the text decoder in the multi-modal large model.

[0085] The parameters of the image encoder and the text decoder in the multi-modal large model can be continuously updated according to the obtained text alignment loss to obtain an increasingly optimized multi-modal large model.

[0086] In practical applications, the data annotation method and the multi-modal large model training method described in the present disclosure can be iteratively executed. For example, for the pre-trained multi-modal large model, the data annotation method described in the present disclosure can be first performed for a batch of data annotation to obtain multiple training samples, such as obtaining 100,000 training samples. Then, these 100,000 training samples can be used to train the multi-modal large model according to the multi-modal large model training method described in the present disclosure to obtain an updated multi-modal large model. Then, the updated multi-modal large model can be used to perform a new batch of data annotation according to the data annotation method described in the present disclosure. Assuming that another 100,000 training samples are obtained, then the newly obtained 100,000 training samples can be accumulated to the previously obtained 100,000 training samples, and the accumulated 200,000 training samples can be used to retrain the multi-modal large model according to the multi-modal large model training method described in the present disclosure. In this way, continuous cyclic iteration can continuously improve the data annotation effect and continuously improve the training effect of the multi-modal large model, such as improving the ability boundary of the multi-modal large model, etc.

[0087] Furthermore, the trained multi-modal large model can be applied to actual data processing scenarios.

[0088] Correspondingly, Figure 5 is a flowchart of an embodiment of the data processing method described in the present disclosure. As Figure 5As shown, it includes the following specific implementation manners.

[0089] In step 501, obtain the image to be processed and the corresponding second prompt text, where the second prompt text includes the second requirement description information proposed for the image to be processed.

[0090] In step 502, input the image to be processed and the second prompt text into the multimodal large model to obtain the output second generated text. The second generated text includes the third response information corresponding to the second requirement description information. The multimodal large model is trained based on the obtained training samples, and the training samples include: sample images, first prompt texts, and annotation texts. The first prompt text includes the first requirement description information proposed for the sample image, and the annotation text includes the first response information corresponding to the first requirement description information.

[0091] It can generate annotation texts according to the Figure 1 shown method, and training samples can be constructed based on the annotation texts. Then, the multimodal large model can be trained according to the training samples, and then the trained multimodal large model can be used for data processing. For example, the image to be processed and the second prompt text can be input into the multimodal large model to obtain the output second generated text, and the accuracy of the second generated text is correspondingly improved, etc.

[0092] It should be noted that for the foregoing method embodiments, for the sake of simple description, they are all expressed as a series of action combinations. However, those skilled in the art should know that the present disclosure is not limited by the described action sequence, because according to the present disclosure, certain steps can be performed in other sequences or simultaneously. Secondly, those skilled in the art should also know that the embodiments described in the specification are all preferred embodiments, and the actions and modules involved are not necessarily essential to the present disclosure. In addition, for the parts not detailed in a certain embodiment, reference can be made to the relevant descriptions in other embodiments.

[0093] The above is the introduction of the method embodiments. The following further illustrates the solution of the present disclosure through device embodiments.

[0094] Figure 6 It is a schematic structural diagram of the data annotation device embodiment 600 of the present disclosure. As Figure 6 shown, it includes: an information acquisition module 601, a text generation module 602, and a data annotation module 603.

[0095] The information acquisition module 601 is used to obtain sample images and the corresponding first prompt texts, where the first prompt text includes the first requirement description information proposed for the sample image.

[0096] The text generation module 602 is configured to generate a comprehensive description text based on the sample image and the first prompt text.

[0097] The data annotation module 603 is configured to input the sample image and the comprehensive description text into a multi-modal large model to obtain an output annotation text, where the annotation text includes a first response message corresponding to the first requirement description information.

[0098] In some embodiments of the present disclosure, the text generation module 602 may perform OCR recognition on the sample image to obtain a text recognition result, and may perform object recognition on the sample image to obtain an object recognition result. Then, a comprehensive description text may be generated based on the text recognition result, the object recognition result, and the first prompt text.

[0099] In some embodiments of the present disclosure, the manner in which the text generation module 602 performs OCR recognition on the sample image to obtain a text recognition result may include: using an OCR classification model to determine the target classification corresponding to the sample image, and using the OCR model corresponding to the target classification to perform OCR recognition on the sample image to obtain a text recognition result.

[0100] For example, the target classification corresponding to the sample image may be determined from M vertical category classifications, where M is a positive integer greater than 1, and different vertical category classifications respectively correspond to their own OCR models.

[0101] In some embodiments of the present disclosure, the manner in which the text generation module 602 performs object recognition on the sample image to obtain an object recognition result may include: respectively using N object recognition models to perform object recognition on the sample image, where N is a positive integer greater than 1, and combining the recognition results of the N object recognition models to generate an object recognition result. Among them, different object recognition models may respectively correspond to different vertical category classifications.

[0102] After respectively obtaining the text recognition result and the object recognition result of the sample image, the text generation module 602 may generate a comprehensive description text based on the text recognition result, the object recognition result, and the first prompt text. In some embodiments of the present disclosure, the text generation module 602 may splice the text recognition result, the object recognition result, and the first prompt text according to a predetermined splicing method to obtain a comprehensive description text.

[0103] Further, the data annotation module 603 may input the sample image and the comprehensive description text into a multi-modal large model to obtain an output annotation text.

[0104] Figure 7 It is a schematic structural diagram of the composition of the multi-modal large model training device embodiment 700 of the present disclosure. As Figure 7 shown, it includes: a sample acquisition module 701 and a model training module 702.

[0105] A sample acquisition module 701 for acquiring training samples, where the training samples include: sample images, first prompt texts, and annotation texts. The first prompt texts include first demand description information proposed for the sample images, and the annotation texts include first response information corresponding to the first demand description information.

[0106] A model training module 702 for training a multi-modal large model based on the training samples.

[0107] In some embodiments of the present disclosure, the model training module 702 may input the sample images and the first prompt texts into the multi-modal large model to obtain an output first generated text, where the first generated text includes second response information corresponding to the first demand description information. Then, the text alignment loss between the first generated text and the annotation text may be obtained, and further, the parameters of the multi-modal large model may be updated according to the text alignment loss.

[0108] In some embodiments of the present disclosure, the model training module 702 may calculate the text alignment loss between the first generated text and the annotation text according to a preset text alignment loss function, such as a cross-entropy loss function.

[0109] In addition, in some embodiments of the present disclosure, the multi-modal large model may include: an image encoder and a text decoder. The image encoder may be used to extract image features from the input image data, and the text decoder may be used to extract text features from the input text data and generate an output result of the multi-modal large model according to the image feature extraction result and the text feature extraction result. Among them, the input image data may be the sample images, the input text data may be the first prompt texts, and the output result may be the first generated text. Correspondingly, the model training module 702 updating the parameters of the multi-modal large model may include: updating the parameters of the image encoder and the text decoder.

[0110] Figure 8 It is a schematic structural diagram of the composition of the data processing device embodiment 800 described in the present disclosure. As Figure 8 shown, it includes: a data acquisition module 801 and a data processing module 802.

[0111] A data acquisition module 801 for acquiring an image to be processed and a corresponding second prompt text, where the second prompt text includes second demand description information proposed for the image to be processed.

[0112] The data processing module 802 is configured to input the image to be processed and the second prompt text into the multimodal large model to obtain the output second generated text, where the second generated text includes the third response information corresponding to the second requirement description information. The multimodal large model is trained based on the obtained training samples, and the training samples include: sample images, first prompt texts, and annotation texts. The first prompt text includes the first requirement description information proposed for the sample image, and the annotation text includes the first response information corresponding to the first requirement description information.

[0113] For the specific working processes of the above device embodiments, reference may be made to the relevant descriptions in the foregoing method embodiments, which will not be elaborated herein.

[0114] The solution described in this disclosure can be applied to the field of artificial intelligence, particularly in the fields of computer vision, deep learning, and large models. Artificial intelligence is a discipline that studies how to make computers simulate certain thinking processes and intelligent behaviors of humans (such as learning, reasoning, thinking, planning, etc.). It includes both hardware-level technologies and software-level technologies. Artificial intelligence hardware technologies generally include technologies such as sensors, dedicated artificial intelligence chips, cloud computing, distributed storage, and big data processing. Artificial intelligence software technologies mainly include several major directions such as computer vision technology, speech recognition technology, natural language processing technology, and machine learning / deep learning, big data processing technology, and knowledge graph technology.

[0115] In addition, the images, texts, etc. in the embodiments described in this disclosure are not targeted at a specific user and do not reflect the personal information of a specific user. In the technical solution of this disclosure, the processing of the collection, storage, use, processing, transmission, provision, and disclosure of user personal information complies with the provisions of relevant laws and regulations and does not violate public order and good customs.

[0116] According to the embodiments of this disclosure, this disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0117] Figure 9 FIG. shows a schematic block diagram of an electronic device 900 that can be used to implement the embodiments of this disclosure. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely illustrative and are not intended to limit the implementation of this disclosure described and / or claimed herein.

[0118] As Figure 9As shown, the electronic device 900 includes a computing unit 901, which can perform various appropriate actions and processes according to computer programs stored in a read-only memory (ROM) 902 or computer programs loaded from a storage unit 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the electronic device 900 can also be stored. The computing unit 901, the ROM 902, and the RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0119] Multiple components in the electronic device 900 are connected to the I / O interface 905, including: an input unit 906, such as a keyboard, a mouse, etc.; an output unit 907, such as various types of displays, speakers, etc.; a storage unit 908, such as a magnetic disk, an optical disc, etc.; and a communication unit 909, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 909 allows the electronic device 900 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0120] The computing unit 901 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 901 include but are not limited to a central processing unit (CPU), a graphic processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 901 executes the various methods and processes described above, such as the methods described in the present disclosure. For example, in some embodiments, the methods described in the present disclosure can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 908. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 900 via the ROM 902 and / or the communication unit 909. When the computer program is loaded into the RAM 903 and executed by the computing unit 901, one or more steps of the methods described in the present disclosure can be executed. Alternatively, in other embodiments, the computing unit 901 can be configured to execute the methods described in the present disclosure by any other appropriate means (e.g., by means of firmware).

[0121] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which may be a special-purpose or general-purpose programmable processor that can receive data and instructions from, and transmit data and instructions to, a storage system, at least one input device, and at least one output device.

[0122] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable data processing apparatus, such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0123] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer disk, a hard disk, a random access memory, a read-only memory, an erasable programmable read-only memory (EPROM), a flash memory, an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0124] To provide for interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide for interaction with the user; for example, feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, speech input, or tactile input).

[0125] The systems and techniques described herein can be implemented in a computing system that includes backend components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system that includes any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), and the Internet.

[0126] A computer system may include a client and a server. The client and the server are generally far from each other and usually interact via a communication network. The client-server relationship is generated by computer programs running on respective computers and having a client-server relationship with each other. The server may be a cloud server, a server of a distributed system, or a server incorporating a blockchain.

[0127] It should be understood that the various forms of processes shown above can be used, steps can be reordered, added, or deleted. For example, the steps described in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution disclosed in this disclosure can be achieved, and no limitations are imposed herein.

[0128] The above specific embodiments do not constitute a limitation on the protection scope of this disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of this disclosure shall be included within the protection scope of this disclosure.

Claims

1. A data annotation method, comprising: Obtaining a sample image and a corresponding first prompt text, wherein the first prompt text includes first demand description information proposed for the sample image; Generating a comprehensive description text according to the sample image and the first prompt text; Inputting the sample image and the comprehensive description text into the multi-modal large model to obtain an output annotation text, wherein the annotation text includes a first response information corresponding to the first demand description information.

2. The method according to claim 1, wherein, The generating a comprehensive description text according to the sample image and the first prompt text includes: Performing optical character recognition on the sample image to obtain a text recognition result; Performing object recognition on the sample image to obtain an object recognition result; Generating the comprehensive description text according to the text recognition result, the object recognition result, and the first prompt text.

3. The method according to claim 2, wherein The performing optical character recognition on the sample image to obtain a text recognition result includes: Using an optical character recognition classification model to determine a target classification corresponding to the sample image; Performing optical character recognition on the sample image by using an optical character recognition model corresponding to the target classification to obtain the text recognition result.

4. The method according to claim 2, wherein, The performing object recognition on the sample image to obtain an object recognition result includes: Respectively performing object recognition on the sample image by using N object recognition models, where N is a positive integer greater than 1; Generating the object recognition result by combining the recognition results of the N object recognition models.

5. The method according to claim 2, wherein The generating the comprehensive description text according to the text recognition result, the object recognition result, and the first prompt text includes: Concatenating the text recognition result, the object recognition result, and the first prompt text according to a predetermined concatenation method to obtain the comprehensive description text.

6. A multi-modal large model training method, comprising: Obtaining a training sample, where the training sample includes: a sample image, a first prompt text, and an annotation text, the first prompt text includes first demand description information proposed for the sample image, and the annotation text includes a first response information corresponding to the first demand description information; Training the multi-modal large model according to the training sample.

7. The method according to claim 6, wherein, The training the multi-modal large model according to the training sample includes: Inputting the sample image and the first prompt text into the multi-modal large model to obtain a first generated text, where the first generated text includes a second response information corresponding to the first demand description information; Obtaining a text alignment loss between the first generated text and the annotation text; Updating the parameters of the multi-modal large model according to the text alignment loss.

8. The method according to claim 7, wherein, The obtaining a text alignment loss between the first generated text and the annotation text includes: Calculating the text alignment loss between the first generated text and the annotation text according to a cross-entropy loss function.

9. According to the method of claim 7, wherein, The multimodal large model includes: an image encoder and a text decoder. The image encoder is used to extract image features from the input image data, and the text decoder is used to extract text features from the input text data, and generate the output result of the multimodal large model according to the image feature extraction result and the text feature extraction result; The parameter update of the multimodal large model includes: updating the parameters of the image encoder and the text decoder.

10. A data processing method, including: Obtaining an image to be processed and a corresponding second prompt text, where the second prompt text includes second requirement description information proposed for the image to be processed; Inputting the image to be processed and the second prompt text into the multimodal large model to obtain an output second generated text, where the second generated text includes a third response information corresponding to the second requirement description information. The multimodal large model is trained according to the obtained training samples. The training samples include: sample images, first prompt texts, and annotation texts. The first prompt text includes first requirement description information proposed for the sample image, and the annotation text includes first response information corresponding to the first requirement description information.

11. A data annotation device, comprising: An information acquisition module, a text generation module, and a data annotation module; The information acquisition module is used to obtain a sample image and a corresponding first prompt text, where the first prompt text includes first requirement description information proposed for the sample image; The text generation module is used to generate a comprehensive description text according to the sample image and the first prompt text; The data annotation module is used to input the sample image and the comprehensive description text into the multimodal large model to obtain an output annotation text, where the annotation text includes first response information corresponding to the first requirement description information.

12. A multi-modal large model training device, comprising: A sample acquisition module and a model training module; The sample acquisition module is used to obtain training samples, where the training samples include: sample images, first prompt texts, and annotation texts. The first prompt text includes first requirement description information proposed for the sample image, and the annotation text includes first response information corresponding to the first requirement description information; The model training module is used to train the multimodal large model according to the training samples.

13. A data processing device, comprising: A data acquisition module and a data processing module; The data acquisition module is used to obtain an image to be processed and a corresponding second prompt text, where the second prompt text includes second requirement description information proposed for the image to be processed; The data processing module is configured to input the image to be processed and the second prompt text into a multi-modal large model to obtain an output second generated text, where the second generated text includes a third response information corresponding to the second requirement description information. The multi-modal large model is trained according to the obtained training samples, and the training samples include: sample images, first prompt texts, and annotation texts. The first prompt texts include first requirement description information proposed for the sample images, and the annotation texts include first response information corresponding to the first requirement description information.

14. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and when the instructions are executed by the at least one processor, the at least one processor is enabled to execute the method according to any one of claims 1-10.

15. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause a computer to execute the method according to any one of claims 1-10.

16. A computer program product, comprising computer programs / instructions, where when the computer programs / instructions are executed by a processor, the method according to any one of claims 1-10 is implemented.

Citation Information

Cited By

  • Multi-modal sample data generation method and device, electronic equipment and storage medium

    CN121365246A

  • Image multi-attribute automatic pre-labeling method based on multi-modal large language model

    CN121661646A