Data generation method and device, multi-modal model training method and device, multi-modal model processing method and device, and equipment

Through the combination of large language model and multimodal model, graphic description data is extracted and corrected, and the illusion problem of multimodal model in graphic description data processing is solved, and high-quality graphic description data generation is achieved.

CN120256949APending Publication Date: 2025-07-04BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510215754.7
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-26
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

Existing multimodal models have hallucinations in the processing of graphic description data, making it difficult to generate high-quality graphic description data.

Method used

The target problem of the original text data is extracted using a large language model, and the original image data is identified through a multimodal model to generate target answers. The target picture and text description data are generated based on the target answers and the original picture and text description data, and the consistency is improved by correcting the image or text data.

Benefits of technology

The quality of the graphic and text description data is improved, the hallucination problem of multimodal models is reduced, and the accuracy and efficiency of data generation are enhanced.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120256949A_ABST
    Figure CN120256949A_ABST
Patent Text Reader

Abstract

The invention provides a data generation method and device, a multi-modal model training and processing method and device, equipment, a medium and a product, relates to the technical field of artificial intelligence, in particular to the technical fields of computer vision, deep learning, large models and the like, and can be applied to scenes of AIGC content generation based on artificial intelligence and the like. The data generation method comprises the following steps: acquiring original image-text description data, wherein the original image-text description data comprises original image data and original text data; extracting the original text data by adopting a large language model to obtain a target problem; identifying the original image data based on the target question by adopting a multi-modal model to obtain a target answer; and generating target image-text description data based on the target answer and the original image-text description data.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of artificial intelligence technology, in particular to technical fields such as computer vision, deep learning, and large models, and can be applied to scenarios such as AIGC (Artificial Intelligence Generated Content). Specifically, it relates to a data generation, multi-modal model training and processing method, device, equipment, medium, and product. Background Art

[0002] Artificial Intelligence Generated Content (AIGC) is a technology that uses artificial intelligence (AI) technology to automatically generate various forms of content such as text, images, audio, and video.

[0003] A multi-modal model is an AIGC model that can process and fuse multi-modal data. In the text-image scenario, the multi-modal model is trained based on text-image description data.

[0004] The text-image description data includes image data and its corresponding text data to comprehensively and accurately express the image content and semantic information. Summary of the Invention

[0005] The present disclosure provides a data generation, multi-modal model training and processing method, device, equipment, medium, and product.

[0006] According to one aspect of the present disclosure, a data generation method is provided, including: obtaining original text-image description data, where the original text-image description data includes: original image data and original text data; using a large language model to extract the original text data to obtain a target question; using a multi-modal model to identify the original image data based on the target question to obtain a target answer; and generating target text-image description data based on the target answer and the original text-image description data.

[0007] According to another aspect of the present disclosure, a multi-modal model training method is provided, including: obtaining target text-image description data; using the target text-image description data to train a multi-modal model; where the target text-image description data is generated by using the method described in any one of the above.

[0008] According to another aspect of the present disclosure, there is provided a method for processing a multimodal model, including: obtaining first-modal data; using the multimodal model to process the first-modal data to obtain second-modal data; the first-modal data and the second-modal data are text data and image data respectively; wherein, the multimodal model is trained using target text-image description data, and the target text-image description data is generated by the method described in any one of the above.

[0009] According to another aspect of the present disclosure, there is provided a data generation device, including: an acquisition module for acquiring original text-image description data, the original text-image description data including: original image data and original text data; an extraction module for using a large language model to extract the original text data to obtain a target question; a question-answering module for using the multimodal model to identify the original image data based on the target question to obtain a target answer; a generation module for generating target text-image description data based on the target answer and the original text-image description data.

[0010] According to another aspect of the present disclosure, there is provided a multimodal model training device, including: an acquisition module for acquiring target text-image description data; a training module for using the target text-image description data to train the multimodal model; wherein, the target text-image description data is generated by the method described in any one of the above.

[0011] According to another aspect of the present disclosure, there is provided a multimodal model processing device, including: an acquisition module for acquiring first-modal data; a processing module for using the multimodal model to process the first-modal data to obtain second-modal data; the first-modal data and the second-modal data are text data and image data respectively; wherein, the multimodal model is trained using target text-image description data, and the target text-image description data is generated by the method described in any one of the above.

[0012] According to another aspect of the present disclosure, there is provided an electronic device, including: at least one processor; and a memory communicatively connected to the at least one processor; wherein, the memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor to enable the at least one processor to execute the method described in any one of the above aspects.

[0013] According to another aspect of the present disclosure, there is provided a non-transitory computer-readable storage medium storing computer instructions, wherein the computer instructions are used to cause the computer to execute the method described in any one of the above aspects.

[0014] According to another aspect of the present disclosure, there is provided a computer program product including a computer program which, when executed by a processor, implements the method according to any one of the above aspects.

[0015] According to the embodiments of the present disclosure, the quality of graphic and text description data can be improved.

[0016] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present disclosure, nor is it used to limit the scope of the present disclosure. Other features of the present disclosure will become easily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0017] The drawings are used to better understand the solution and do not constitute a limitation to the present disclosure. Among them:

[0018] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure;

[0019] Figure 2 is a schematic diagram of the overall system for implementing the embodiments of the present disclosure;

[0020] Figure 3 is a schematic diagram of the overall data generation process according to the embodiments of the present disclosure;

[0021] Figure 4 is a schematic diagram according to the second embodiment of the present disclosure;

[0022] Figure 5 is a schematic diagram according to the third embodiment of the present disclosure;

[0023] Figure 6 is a schematic diagram according to the fourth embodiment of the present disclosure;

[0024] Figure 7 is a schematic diagram according to the fifth embodiment of the present disclosure;

[0025] Figure 8 is a schematic diagram according to the sixth embodiment of the present disclosure;

[0026] Figure 9 is a schematic diagram according to the seventh embodiment of the present disclosure;

[0027] Figure 10 is a schematic diagram of an electronic device for implementing the data generation method, multi-modal model training method or multi-modal model processing method according to the embodiments of the present disclosure. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0028] The following describes exemplary embodiments of the present disclosure with reference to the accompanying drawings. Various details of the embodiments of the present disclosure are included to facilitate understanding, and they should be considered merely exemplary. Therefore, those of ordinary skill in the art should recognize that various changes and modifications can be made to the embodiments described herein without departing from the scope and spirit of the present disclosure. Similarly, descriptions of well-known functions and structures are omitted in the following description for clarity and conciseness.

[0029] To mitigate the hallucination problem of multimodal models, high-quality image-text description data is required.

[0030] To generate high-quality image-text description data, the present disclosure provides the following embodiments.

[0031] Figure 1 is a schematic diagram according to the first embodiment of the present disclosure. This embodiment provides a data generation method, as Figure 1 shown, the method includes:

[0032] 101. Obtain original image-text description data, where the original image-text description data includes: original image data and original text data.

[0033] 102. Use a large language model to extract the original text data to obtain a target question.

[0034] 103. Use a multimodal model to identify the original image data based on the target question to obtain a target answer.

[0035] 104. Generate target image-text description data based on the target answer and the original image-text description data.

[0036] Among them, the original image-text description data is existing image-text description data, which includes: original image data, and original text data used to describe the original image data.

[0037] For the original text data, use a large language model (Large Language Model, LLM) to extract the target question.

[0038] The large language model (Large Language Model, LLM) is a hot technology in the field of AI in recent years. The LLM is a natural language processing model based on deep learning, which has a huge number of parameters and a complex structure, enabling it to process and understand a large amount of natural language data.

[0039] Specifically, the original text data and a preset first prompt statement can be input into the LLM, and the output is the target question.

[0040] After obtaining the target question, a multimodal model is used to identify the above-mentioned original image data to obtain the target answer corresponding to the target question.

[0041] Specifically, the original image data, the target question, and a preset second prompt statement can be input into the multimodal model, and the output is the target answer.

[0042] Among them, the number of target answers and target questions is one or more, and each target answer corresponds to one target question.

[0043] After obtaining the target answer, based on the target answer and the original text and image description data, the target text and image description data is obtained.

[0044] In this embodiment, a large language model is used to obtain the target question, a multimodal model is used to obtain the target answer, and the target text and image description data is obtained based on the target answer and the original text and image description data. The excellent performance of the large language model and the multimodal model can be utilized to improve the quality of the target text and image description data, thereby reducing the hallucination problem of the multimodal model.

[0045] In some embodiments, if the target answer is consistent with the original text data. For example, if the original text data indicates that the clothes are green and the target answer also indicates that the clothes are green, then the original text and image description data can be used as the target text and image description data, that is, the target text and image description data includes the original image data and the original text data. Or rather, the target text and image description data is generated based on the original image data and the original text data. Specifically, the original image data and the original text data are combined to form the target text and image description data. For example, the original image data and the original text data are combined to form an original text-image pair, and both the original text and image description data and the target text and image description data include this original text-image pair.

[0046] If the target answer is inconsistent with the original text data. For example, if the original text data indicates that the clothes are green and the target answer indicates that the clothes are not green, then the original text and image description data can be corrected to obtain the target text and image description data.

[0047] In this embodiment, since the target answer is obtained based on the original image data, when the target answer is consistent with the original text data, it indicates that the original image data and the original text data are matched. At this time, the original text and image description data is used as the target text and image description data. On the contrary, when the target answer is inconsistent with the original text data, it indicates that the original image data and the original text data are not matched. At this time, correcting the original text and image description data can accurately and efficiently obtain the target text and image description data.

[0048] Furthermore, the target answer is a judgment on the target question. Based on this, if the judgment result indicates positive information, it is determined that the target answer is consistent with the original text data, and / or, if the judgment result indicates negative information, it is determined that the target answer is inconsistent with the original text data.

[0049] Specifically, the target question at this time can be called a judgment question, and the target answer can include a positive answer, and / or, a negative answer. Among them, the positive answer indicates positive information, such as expressed by "yes", and the negative answer indicates negative information, such as expressed by "no".

[0050] In this embodiment, since the target answer is a judgment on the target question, based on positive information or negative information, it can be simply and efficiently known whether the target answer is consistent with the original text data, improving the processing efficiency.

[0051] During correction, the original text data can be corrected, and the corrected text data is called the target text data. After that, the target graphic and text description data is generated based on the original image data and the target text data. Specifically, the original image data and the target text data are combined to form the target graphic and text description data. For example, the original image data and the target text data form a first graphic and text pair, and the target graphic and text description data includes this first graphic and text pair. Or,

[0052] The original image data can also be corrected, and the corrected image data can be called the target image data. For example, if the clothes color obtained based on the original image data is red, and the clothes color in the original text data is green, then the clothes color in the original image data can be corrected from red to green. After that, the target graphic and text description data is generated based on the target image data and the original text data. Specifically, the target image data and the original text data are combined to form the target graphic and text description data. For example, the target image data and the original text data form a second graphic and text pair, and the target graphic and text description data includes this second graphic and text pair.

[0053] In this way, by correcting the original image data or the original text data, and then obtaining the target graphic and text description data, the flexibility can be improved.

[0054] The correction method can include deletion, modification, etc. For example, the description of the clothes color in the original text data can be deleted; or, assuming that the target answer obtained through the multimodal model indicates that the clothes color is red, the clothes color in the original text data can be modified from green to red.

[0055] In this way, by performing corrections through various correction methods, the flexibility can be improved.

[0056] To better understand the embodiments of the present disclosure, the relevant application scenarios are described.

[0057] Figure 2It is a schematic diagram of the overall system for implementing the embodiments of the present disclosure.

[0058] As Figure 2 shown, the overall system includes: a user terminal 201 and a server 202.

[0059] The user terminal 201 may include: mobile devices such as a personal computer (PC), a laptop, a mobile phone, etc. The server 202 may be a local server or a cloud server, and may be a single server or a cluster server.

[0060] A client is deployed on the user terminal 201, and a server-side is deployed on the server 202.

[0061] The user can send the original graphic and text description data to the server-side through the client; the server-side processes the original graphic and text description data to obtain the target graphic and text description data; after obtaining the target graphic and text description data, the server-side can send it to the client and display it to the user through the client.

[0062] Among them, the server-side can process the original graphic and text description data based on a large language model and a multi-modal model to obtain the target graphic and text description data.

[0063] Figure 3 It is a schematic diagram of the overall data generation process provided by the embodiments of the present disclosure.

[0064] In this embodiment, taking the generation of judgment questions and correcting the original text data as an example.

[0065] As Figure 3 shown, the server-side receives the original graphic and text description data 301 sent by the client. The original graphic and text description data includes original image data and original text data; extracts the target question based on the original text data, specifically the judgment question 302; then obtains the target answer 303 corresponding to the target question based on the original image data; corrects the original text data based on the target answer to obtain the target text data 304. After that, the target graphic and text description data can be generated based on the original image data and the target text data.

[0066] Among them, for the original text data, a large language model is called, and the large language model is used to extract the original text data to obtain the target question.

[0067] Specifically, the original text data and a preset first prompt statement can be input into the large language model, and the output is the target question.

[0068] Taking the target question as a true - false question as an example, the first prompt statement can specifically be: "Here is a description of an image generated by a model, and some parts may be incorrect. Please read this description and try to come up with some true - false questions about the key information in it to help me check and improve the accuracy of the description. Each true - false question should be specific and contain only one question. Each true - false question is an interrogative sentence that can be answered with 'yes' or 'no'. Do not give the answers."

[0069] Based on the original text data and the first prompt statement, the large - language model can generate target questions, specifically true - false questions. The number of target questions can be one or more. Figure 3 Taking 10 true - false questions as an example.

[0070] After obtaining the target questions, call the multi - modal model to obtain the target answers corresponding to the target questions.

[0071] Specifically, the original image data, the extracted target questions, and the preset second prompt statement can be input into the multi - modal model, and the output is the target answer.

[0072] The second prompt statement can specifically be: "Please carefully analyze the picture, find the answer to the corresponding question in the image, and give a judgment, answering 'yes' or 'no', without any other explanations."

[0073] Based on the original image data, the target questions, and the second prompt statement, the multi - modal model can generate the target answers.

[0074] When the target question is a true - false question, the target answer includes an affirmative answer or a negative answer. The affirmative answer is represented by 'yes', and the negative answer is represented by 'no'. The target questions and the target answers are in one - to - one correspondence. For example, Figure 3 As shown, 10 target answers can be obtained based on 10 target questions.

[0075] After obtaining the target answers, correct the text content corresponding to the negative answers, specifically by using the deletion method.

[0076] The content to be deleted is part or all of the text content corresponding to the negative answer. For example, if the answer to the second question shown is a negative answer (no), and the text content corresponding to this question is "white shorts", then the entire text content "white shorts" can be deleted. Or, as Figure 3 shown, if the answer to the second question is a negative answer (no), and the text content corresponding to this question is "white shorts", then the entire text content "white shorts" can be deleted. Or, as Figure 3The answer corresponding to the fourth question shown is a negative answer (No). The text content corresponding to this question is "with hair tied up, fluttering in the wind". At this time, some content can be deleted, such as deleting "fluttering in the wind" while retaining "with hair tied up". Among them, when the multi-modal model outputs a negative answer, it can also output a prompt message indicating the content with an error (such as the error of "fluttering in the wind"), so that when deleting, the part with the error can be deleted based on the prompt message.

[0077] Generally speaking, the original graphic and text description data to be processed has a certain quality basis, and it will not happen that all target answers are "No". Therefore, even if all the text content corresponding to each negative answer is deleted, since there are at least some positive answers and the text content corresponding to the positive answers is retained, the target text data is also non-empty.

[0078] After obtaining the target text data, generate target graphic and text description data based on the original image data and the target text data. For example, form a graphic-text pair with the original image data and the target text data, and the target graphic and text description data includes this graphic-text pair.

[0079] Combined with the above application scenarios, the present disclosure also provides the following embodiments.

[0080] Figure 4 It is a schematic diagram according to the second embodiment of the present disclosure. This embodiment provides a data generation method, as Figure 4 shown, the method includes:

[0081] 401. Obtain original graphic and text description data, where the original graphic and text description data includes: original image data and original text data.

[0082] 402. Use a large language model to extract the original text data to obtain target questions.

[0083] 403. Use a multi-modal model to identify the original image data based on the target questions to obtain target answers.

[0084] Among them, referring to Figure 3 , the target questions can specifically be judgment questions. Correspondingly, the target answers include positive answers (Yes) or negative answers (No).

[0085] In this embodiment, by setting the target questions as judgment questions, it can be simply known whether the target answers are consistent with the original text data, improving the data generation efficiency.

[0086] 404. Correct the text content corresponding to the negative answers in the original text data to obtain target text data.

[0087] 405. Generate target graphic-text description data based on the original image data and the target text data.

[0088] In this embodiment, since text data is easier to correct than image data, obtaining the target graphic-text description data by correcting the original text data can reduce resource overhead.

[0089] Specifically, in the original text data, part or all of the text content corresponding to the negative answer can be deleted to obtain the target text data.

[0090] In this embodiment, obtaining the target text data by deletion can simply and efficiently obtain the target text data, thereby improving the data generation efficiency and simplicity.

[0091] Furthermore, a large language model can be used to perform the correction operation. That is, the large language model can be used to correct the text content corresponding to the above negative answer to obtain the target text data.

[0092] Specifically, the original text data, the target question corresponding to the negative answer, and a preset third prompt statement can be input into the large language model, and the output is the target text data. The third prompt statement is, for example, an instruction to the large language model to delete the relevant text content of the target question corresponding to the negative answer.

[0093] In this embodiment, obtaining the target text data by correcting the original text data with a large language model can improve the accuracy of the target text data, and thus improve the accuracy of the target graphic-text description data.

[0094] Figure 5 It is a schematic diagram according to the third embodiment of the present disclosure. This embodiment provides a multi-modal model training method, as Figure 5 shown, the method includes:

[0095] 501. Obtain target graphic-text description data.

[0096] 502. Use the target graphic-text description data to train a multi-modal model; wherein, the target graphic-text description data is generated by using the data generation method of any of the above embodiments.

[0097] In this embodiment, since the above embodiments can obtain high-quality target graphic-text description data, therefore, training a multi-modal model based on this high-quality target graphic-text description data can improve the performance of the multi-modal model and alleviate the hallucination problem of the multi-modal model.

[0098] Figure 6 It is a schematic diagram according to the fourth embodiment of the present disclosure. This embodiment provides a multi-modal model processing method, as Figure 6 shown, the method includes:

[0099] 601. Obtain the first modal data.

[0100] 602. Use a multimodal model to process the first modal data to obtain second modal data; the first modal data is text data and the second modal data is image data; or, the first modal data is image data and the second modal data is text data; wherein, the multimodal model is trained with target text-image description data, and the target text-image description data is generated by the data generation method of any of the above embodiments.

[0101] Specifically, the multimodal model can be used in the text-to-image scenario. In this case, the first modal data is text data and the second modal data is image data; or, it can also be used in the image-to-text scenario. In this case, the first modal data is image data and the second modal data is text data.

[0102] In this embodiment, since high-quality target text-image description data can be obtained in the above embodiment, the multimodal model trained based on the high-quality target text-image description data has good performance, which can further reduce the hallucination problem of the multimodal model and improve the model processing performance.

[0103] Figure 7 It is a schematic diagram according to the fifth embodiment of the present disclosure. This embodiment provides a data generation device 700, which includes: an acquisition module 701, an extraction module 702, a question-answering module 703, and a generation module 704.

[0104] The acquisition module 701 is used to acquire original text-image description data, and the original text-image description data includes: original image data and original text data; the extraction module 702 is used to extract the original text data with a large language model to obtain a target question; the question-answering module 703 is used to identify the original image data based on the target question with a multimodal model to obtain a target answer; the generation module 704 is used to generate target text-image description data based on the target answer and the original text-image description data.

[0105] In this embodiment, a target question is obtained by using a large language model, a target answer is obtained by using a multimodal model, and target text-image description data is obtained based on the target answer and the original text-image description data. The excellent performance of the large language model and the multimodal model can be utilized to improve the quality of the target text-image description data, thereby reducing the hallucination problem of the multimodal model.

[0106] In some embodiments, the generation module 704 is further used for:

[0107] If the target answer is consistent with the original text data, use the original graphic and text description data as the target graphic and text description data; or,

[0108] If the target answer is inconsistent with the original text data, correct the original graphic and text description data to obtain the target graphic and text description data.

[0109] In this embodiment, since the target answer is obtained based on the original image data, when the target answer is consistent with the original text data, it indicates that the original image data and the original text data are matched. At this time, use the original graphic and text description data as the target graphic and text description data. On the contrary, when the target answer is inconsistent with the original text data, it indicates that the original image data and the original text data are not matched. At this time, correct the original graphic and text description data to accurately and efficiently obtain the target graphic and text description data.

[0110] In some embodiments, the target answer is a judgment on the target question; the method further includes: if the judgment result indicates positive information, determine that the target answer is consistent with the original text data; and / or, if the judgment result indicates negative information, determine that the target answer is inconsistent with the original text data.

[0111] In this embodiment, by setting the target question as a judgment question, it is possible to simply and efficiently know whether the target answer is consistent with the original text data based on the positive answer or the negative answer, improving the processing efficiency.

[0112] In some embodiments, the generating module 904 is further configured to:

[0113] If the target answer is inconsistent with the original text data, correct the text content corresponding to the target answer in the original text data to obtain the target text data;

[0114] Generate the target graphic and text description data based on the original image data and the target text data.

[0115] In this embodiment, since the text data is easier to correct than the image data, obtaining the target graphic and text description data by correcting the original text data can reduce the resource overhead.

[0116] In some embodiments, the generating module 704 is further configured to:

[0117] In the original text data, delete the part or all of the text content corresponding to the target answer to obtain the target text data.

[0118] In this embodiment, obtaining the target text data by deletion can simply and efficiently obtain the target text data, thereby improving the data generation efficiency and simplicity.

[0119] In some embodiments, the generation module 704 is further configured to:

[0120] Use the large language model to correct the text content corresponding to the target answer in the original text data to obtain target text data.

[0121] In this embodiment, by using the large language model to correct the original text data to obtain the target text data, the accuracy of the target text data can be improved, and further the accuracy of the target graphic description data can be improved.

[0122] Figure 8 FIG. is a schematic diagram according to the sixth embodiment of the present disclosure. This embodiment provides a multi-modal model training device, and the device 800 includes: an acquisition module 801 and a training module 802.

[0123] The acquisition module 801 is configured to acquire target graphic description data; the training module 802 is configured to use the target graphic description data to train a multi-modal model; wherein, the target graphic description data is generated by using the data generation method of any of the above embodiments.

[0124] In this embodiment, since the above embodiments can obtain high-quality target graphic description data, therefore, training the multi-modal model based on the high-quality target graphic description data can improve the performance of the multi-modal model and alleviate the hallucination problem of the multi-modal model.

[0125] Figure 9 FIG. is a schematic diagram according to the seventh embodiment of the present disclosure. This embodiment provides a multi-modal model processing device, and the device 900 includes: an acquisition module 901 and a processing module 902.

[0126] The acquisition module 901 is configured to acquire first-modal data; the processing module 902 is configured to use the multi-modal model to process the first-modal data to obtain second-modal data; the first-modal data is text data and the second-modal data is image data; or, the first-modal data is image data and the second-modal data is text data; wherein, the multi-modal model is trained by using the target graphic description data, and the target graphic description data is generated by using the data generation method of any of the above embodiments.

[0127] In this embodiment, since the above embodiments can obtain high-quality target graphic description data, the multi-modal model trained based on the high-quality target graphic description data has good performance, and further the hallucination problem of the multi-modal model can be alleviated and the model processing performance can be improved.

[0128] It can be understood that in the embodiments of the present disclosure, the same or similar content in different embodiments can be referred to each other.

[0129] It is understood that the "first", "second", etc. in the embodiments of the present disclosure are only used for distinction and do not represent the level of importance, the sequence of time, etc.

[0130] It is understood that if there is no special limitation or explanation for the sequence of steps involved in the process, it indicates that the temporal relationship between these steps is not limited.

[0131] In the technical solution of the present disclosure, the processing of the collection, storage, use, processing, transmission, provision, and disclosure of the user's personal information complies with the provisions of relevant laws and regulations and does not violate public order and good customs.

[0132] According to the embodiments of the present disclosure, the present disclosure also provides an electronic device, a readable storage medium, and a computer program product.

[0133] Figure 10 A schematic block diagram of an exemplary electronic device 1000 that can be used to implement the embodiments of the present disclosure is shown. The electronic device 1000 is intended to represent various forms of digital computers, such as, a laptop computer, a desktop computer, a workbench, a server, a blade server, a mainframe computer, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, a personal digital assistant, a cellular phone, a smart phone, a wearable device, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present disclosure described and / or claimed herein.

[0134] As Figure 10 shown, the electronic device 1000 includes a computing unit 1001, which can execute various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 1002 or the computer program loaded from the storage unit 1008 into the random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation of the electronic device 1000 can also be stored. The computing unit 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. The input / output (I / O) interface 1005 is also connected to the bus 1004.

[0135] Multiple components in the electronic device 1000 are connected to the I / O interface 1005, including: an input unit 1006, such as a keyboard, a mouse, etc.; an output unit 1007, such as various types of displays, speakers, etc.; a storage unit 1008, such as a disk, an optical disc, etc.; and a communication unit 1009, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 1009 allows the electronic device 1000 to exchange information / data with other devices through a computer network such as the Internet and / or various telecommunication networks.

[0136] The computing unit 1001 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the computing unit 1001 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various computing units running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The computing unit 1001 executes the various methods and processes described above, such as the data generation method, the multi-modal model training method, or the multi-modal model processing method. For example, in some embodiments, the data generation method, the multi-modal model training method, or the multi-modal model processing method can be implemented as a computer software program, which is tangibly contained in a machine-readable medium, such as the storage unit 1008. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 1000 via the ROM 1002 and / or the communication unit 1009. When the computer program is loaded into the RAM 1003 and executed by the computing unit 1001, one or more steps of the data generation method, the multi-modal model training method, or the multi-modal model processing method described above can be executed. Alternatively, in other embodiments, the computing unit 1001 can be configured to execute the data generation method, the multi-modal model training method, or the multi-modal model processing method by any other suitable means (e.g., by means of firmware).

[0137] The various embodiments of the systems and techniques described above in this specification can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from, and transmits data and instructions to, a storage system, at least one input device, and at least one output device.

[0138] The program code for implementing the methods of the present disclosure can be written in any combination of one or more programming languages. These program codes can be provided to a processor or controller of a general purpose computer, special purpose computer, or other programmable task processing device such that the program codes, when executed by the processor or controller, cause the functions / operations specified in the flowchart and / or block diagram to be implemented. The program code can be executed entirely on the machine, partly on the machine, as a stand-alone software package partly on the machine and partly on a remote machine or entirely on the remote machine or server.

[0139] In the context of this disclosure, a machine-readable medium can be a tangible medium that can contain or store a program for use by or in connection with an instruction execution system, apparatus, or device. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. More specific examples of a machine-readable storage medium would include an electrical connection based on one or more wires, a portable computer diskette, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or Flash memory), an optical fiber, a portable compact disc read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.

[0140] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) by which the user can provide input to the computer. Other kinds of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0141] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), and the Internet.

[0142] A computer system can include a client and a server. The client and the server are generally remote from each other and typically interact through a communication network. The relationship between the client and the server is generated by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system, and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services ("Virtual Private Server", or simply "VPS"). The server can also be a server of a distributed system, or a server combined with a blockchain.

[0143] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in this disclosure can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solutions disclosed in this disclosure can be achieved, and no limitation is made herein.

[0144] The above specific embodiments do not constitute a limitation on the protection scope of the present disclosure. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present disclosure shall be included within the protection scope of the present disclosure.

Claims

1. A data generation method, comprising: Obtaining original graphic and text description data, where the original graphic and text description data includes: original image data and original text data; Using a large language model to extract the original text data to obtain a target question; Using a multi-modal model to identify the original image data based on the target question to obtain a target answer; Generating target graphic and text description data based on the target answer and the original graphic and text description data.

2. The method according to claim 1, wherein The generating target graphic and text description data based on the answer and the original graphic and text description data includes: If the target answer is consistent with the original text data, using the original graphic and text description data as the target graphic and text description data; or, If the target answer is inconsistent with the original text data, correcting the original graphic and text description data to obtain the target graphic and text description data.

3. The method according to claim 2, wherein, The target answer is a judgment on the target question; The method further includes: If the judgment result indicates positive information, determining that the target answer is consistent with the original text data; and / or, If the judgment result indicates negative information, determining that the target answer is inconsistent with the original text data.

4. The method according to claim 2, wherein, The if the target answer is inconsistent with the original text data, correcting the original graphic and text description data to obtain the target graphic and text description data includes: If the target answer is inconsistent with the original text data, correcting the text content corresponding to the target answer in the original text data to obtain target text data; Generating the target graphic and text description data based on the original image data and the target text data.

5. The method according to claim 4, wherein The correcting the text content corresponding to the target answer in the original text data to obtain target text data includes: In the original text data, deleting part or all of the text content corresponding to the target answer to obtain the target text data.

6. The method according to claim 4, wherein The correcting the text content corresponding to the target answer in the original text data to obtain target text data includes: Using the large language model to correct the text content corresponding to the target answer in the original text data to obtain target text data.

7. A multi-modal model training method, comprising: Obtaining target graphic and text description data; Using the target graphic and text description data to train a multi-modal model; Wherein, the target graphic and text description data is generated by using the method according to any one of claims 1-6.

8. A multi-modal model processing method, comprising: Obtaining first modal data; Using a multi-modal model to process the first modal data to obtain second modal data; The first modal data is text data and the second modal data is image data; or, the first modal data is image data and the second modal data is text data; Wherein, the multi-modal model is trained by using target graphic and text description data, and the target graphic and text description data is generated by using the method according to any one of claims 1-6.

9. A data generation device, comprising: An acquisition module, configured to acquire original graphic and text description data, where the original graphic and text description data includes: original image data and original text data; An extraction module, configured to use a large language model to extract the original text data to obtain a target question; A question and answer module, configured to use a multimodal model to identify the original image data based on the target question to obtain a target answer; A generation module, configured to generate target graphic and text description data based on the target answer and the original graphic and text description data.

10. A multimodal model training device, comprising: An acquisition module, configured to acquire target graphic and text description data; A training module, configured to use the target graphic and text description data to train a multimodal model; wherein the target graphic and text description data is generated by using the method according to any one of claims 1-6.

11. A multimodal model processing device, comprising: An acquisition module, configured to acquire first modality data; A processing module, configured to use a multimodal model to process the first modality data to obtain second modality data; the first modality data and the second modality data are text data and image data respectively; wherein the multimodal model is trained by using target graphic and text description data, and the target graphic and text description data is generated by using the method according to any one of claims 1-6.

12. An electronic device, comprising: At least one processor; And A memory communicatively connected to the at least one processor; wherein, The memory stores instructions executable by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the method according to any one of claims 1-8.

13. A non-transitory computer-readable storage medium storing computer instructions, wherein, The computer instructions are used to cause the computer to execute the method according to any one of claims 1-8.

14. A computer program product, comprising a computer program, where the computer program, when executed by a processor, implements the method according to any one of claims 1-8.

Citation Information

Cited By

  • Construction method of multi-modal image-text data set and data generation method

    CN121256069A