Content generation method and device, equipment, storage medium and program product

By training a target generative model and combining visual reproduction with multimodal information, the shortcomings of large multimodal language models in visual information processing are addressed, achieving more efficient content generation.

CN120673039APending Publication Date: 2025-09-19BEIJING ZITIAO NETWORK TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510788245.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-12
Publication Date
2025-09-19

AI Technical Summary

Technical Problem

Existing multimodal large language models rely insufficiently on language clues when processing visual information, making it difficult to effectively integrate multimodal data, resulting in poor content generation.

Method used

By introducing selective visual reproduction mechanism and visual clue dataset to train target generative model, its visual perception ability is enhanced and effective integration of multimodal data is achieved.

Benefits of technology

It improves the effect of content generation, and enhances the accuracy and efficiency of visual information processing capabilities and multimodal reasoning.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120673039A_ABST
    Figure CN120673039A_ABST
Patent Text Reader

Abstract

The invention discloses a content generation method and device, equipment, a storage medium and a program product, and relates to the technical field of large language model visual perception, and the method comprises the steps: obtaining image data and problem information corresponding to the image data; identifying semantic features of the problem information by using the target generative model, and positioning a visual area corresponding to the problem information in the image data according to the semantic features; identifying visual information in the visual area by using a target generative model, and generating target text content matched with the question information according to the visual information; wherein the target generative model is generated based on visual reproduction and multi-modal information training. According to the technical scheme, the visual features in the image data can be fully recognized, the processing capacity of the visual information is improved, reasoning of content generation is carried out in combination with the semantic features and the visual information, effective integration of multi-modal data is achieved, and therefore the content generation effect is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the field of large language model visual perception technology, and in particular to a content generation method, apparatus, device, storage medium, and program product. Background Art

[0002] With the development of large multimodal language models, generative applications based on these models have emerged. Current generative applications primarily rely on language cues to generate content, but they lack sufficient processing of visual information and struggle to effectively integrate multimodal data, resulting in poor overall comprehension and hindering content generation. Summary of the Invention

[0003] In view of this, the present disclosure provides a content generation method, apparatus, device, storage medium, and program product to solve the problem of poor content generation effect.

[0004] In a first aspect, the present disclosure provides a content generation method, including: obtaining image data and question information corresponding to the image data; using a target generative model to identify semantic features of the question information, and locating a visual area corresponding to the question information in the image data according to the semantic features; using a target generative model to identify visual information in the visual area, and generating target text content that matches the question information according to the visual information; wherein the target generative model is generated based on visual reproduction and multimodal information training.

[0005] In a second aspect, the present disclosure provides a content generation device, including: an acquisition module for acquiring image data and question information corresponding to the image data; a visual area positioning module for using a target generative model to identify semantic features of the question information, and positioning the visual area corresponding to the question information in the image data according to the semantic features; a content generation module for using a target generative model to identify visual information in the visual area, and generating target text content matching the question information according to the visual information; wherein the target generative model is generated based on visual reproduction and multimodal information training.

[0006] In a third aspect, the present disclosure provides an electronic device comprising: a memory and a processor, the memory and the processor being communicatively connected to each other, the memory storing computer instructions, and the processor executing the content generation method of the first aspect or any corresponding embodiment thereof by executing the computer instructions.

[0007] In a fourth aspect, the present disclosure provides a computer-readable storage medium having computer instructions stored thereon, the computer instructions being used to enable a computer to execute the content generation method of the first aspect or any corresponding embodiment thereof.

[0008] In a fifth aspect, the present disclosure provides a computer program product, including computer instructions, which are used to enable a computer to execute the content generation method of the first aspect or any corresponding embodiment thereof.

[0009] The content generation method, apparatus, device, storage medium and program product provided by the embodiments of the present disclosure generate a target generative model based on visual reproduction and multimodal information training, so that the target generative model has the ability to reproduce image features and the ability to reason about multimodal information, thereby enhancing the visual perception ability of the target generative model. Therefore, after obtaining the image data and the question information corresponding to the image data, the target generative model is used to identify the semantic features of the question information, and the visual area corresponding to the question information is located in the image data according to the semantic features, so as to fully identify the visual features in the image data and improve the processing ability of visual information. Subsequently, the target generative model is used to identify the visual information in the visual area, and the target text content that matches the question information is generated according to the visual information. Thus, the semantic features and visual information can be combined to perform reasoning for content generation, thereby achieving effective integration of multimodal data and improving the content generation effect. BRIEF DESCRIPTION OF THE DRAWINGS

[0010] In order to more clearly illustrate the specific embodiments of the present disclosure or the technical solutions in the related technologies, the following briefly introduces the drawings required for use in the specific embodiments or related technical descriptions. Obviously, the drawings described below are some embodiments of the present disclosure. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0011] Figure 1 is a schematic diagram of an application scenario according to an embodiment of the present disclosure;

[0012] Figure 2 is a flowchart of a content generation method according to an embodiment of the present disclosure;

[0013] Figure 3 is a flow chart of a target generative model training method according to an embodiment of the present disclosure;

[0014] Figure 4 is a schematic diagram of a visual training sample according to an embodiment of the present disclosure;

[0015] Figure 5 is a schematic diagram of a training scenario of a target generative model according to an embodiment of the present disclosure;

[0016] Figure 6 is a schematic diagram of constructing a visual training data set according to an embodiment of the present disclosure;

[0017] Figure 7is a performance comparison diagram according to an embodiment of the present disclosure;

[0018] Figure 8 is a schematic diagram of effective verification according to an embodiment of the present disclosure;

[0019] Figure 9 is a structural block diagram of a content generation device according to an embodiment of the present disclosure;

[0020] Figure 10 Schematic diagram of the hardware structure of an electronic device according to an embodiment of the present disclosure. DETAILED DESCRIPTION

[0021] To make the purpose, technical solutions, and advantages of the embodiments of the present disclosure more clear, the technical solutions in the embodiments of the present disclosure will be clearly and completely described below in conjunction with the drawings in the embodiments of the present disclosure. Obviously, the described embodiments are part of the embodiments of the present disclosure, not all of the embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the present disclosure.

[0022] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0023] For example, in response to a user's active request, a prompt message is sent to the user to clearly inform the user that the operation requested will require the acquisition and use of the user's personal information. This allows the user to independently choose whether to provide personal information to the electronic device, application, server, storage medium, or other software or hardware that performs the operations of the disclosed technical solution based on the prompt message.

[0024] As an optional but non-limiting implementation, in response to receiving a user's active request, the prompt information may be sent to the user in the form of a pop-up window, in which the prompt information may be presented in text form. Furthermore, the pop-up window may also contain a selection control for the user to select "agree" or "disagree" to provide personal information to the electronic device.

[0025] It is understandable that the above notification and user authorization process are merely illustrative and do not limit the implementation of the present disclosure. Other methods that comply with relevant laws and regulations may also be applied to the implementation of the present disclosure.

[0026] It is understandable that the data involved in this technical solution (including but not limited to the data itself, the acquisition or use of the data) must comply with the requirements of relevant laws, regulations and relevant provisions.

[0027] Currently, when large multimodal language models perform generation tasks, the following main problems exist: (1) Current large multimodal language models rely heavily on language clues during reasoning, which makes them prone to misjudgment when faced with language bias, which is particularly evident when dealing with tasks that require visual understanding. (2) They have limited ability to analyze image details, especially in complex scenes, and are unable to effectively identify and utilize important visual information in images. (3) There are deficiencies in the integration of language and visual information, and the synergistic effect of multimodal data cannot be fully utilized. (4) When processing multimodal data, large computing resources are often required, which affects its feasibility in practical applications. (5) Existing datasets lack annotations that can support fine-grained visual reasoning, which limits the training and reasoning capabilities of the model.

[0028] To address the above pain points, the solutions commonly used in the current field include the following: (1) Reasoning based on language models. Current multimodal reasoning usually relies on powerful language models (such as the GPT series) to perform reasoning in the language space. These models generate answers by encoding language information. However, this method is not capable of processing visual information, is easily affected by language bias, and cannot fully utilize the detailed information in the image. (2) Visual Question Answering (VQA) model. This method needs to combine images and question texts and output answers through deep learning models (such as a combination of convolutional neural networks and recurrent neural networks). However, VQA models can usually only handle simple visual reasoning and have difficulty dealing with complex scenes or tasks that require in-depth understanding. In addition, most VQA models are static and lack dynamic adjustment capabilities. (3) Multimodal attention mechanism, that is, using attention mechanisms to focus on and combine relevant information in multimodal data to improve reasoning performance. Although the attention mechanism can improve information integration to a certain extent, it still has the problem of insufficient grasp of important visual details when facing complex scenes. (4) Combining image feature extraction with language description: By extracting image features and combining them with language description, the model attempts to understand and answer questions. Since the quality and accuracy of feature extraction directly affect the reasoning effect, related methods have difficulty extracting sufficiently detailed and accurate features in complex scenes. (5) Large-scale pre-trained models: By pre-training on large-scale multimodal datasets, the model can show better versatility in reasoning tasks. However, they usually require a lot of computing resources and still have limitations when processing fine-grained visual tasks.

[0029] Based on this, when training a target generative model based on a multimodal large language model architecture, the present invention introduces a selective visual reproduction mechanism and a visual clue data set for multimodal thinking and reasoning, thereby enhancing the visual perception ability of the target generative model, achieving effective integration of multimodal data, and improving the accuracy and efficiency of multimodal reasoning.

[0030] As an optional application scenario of the embodiment of the present disclosure, Figure 1 As shown, the optional application scenario includes a target generative model 101 and an electronic device 102, wherein the target generative model 101 is deployed in a generative application 103 to provide visual perception based on questions and input images, and generate corresponding answer content, and the generative application 103 is deployed in the electronic device 102.

[0031] The electronic device 102 may be a device with computing capabilities, for example, the electronic device 102 may be provided with a processor and memory, or may be equipped with a dedicated accelerator (such as a graphics processing unit (GPU)). In addition, the electronic device 102 may store and maintain data.

[0032] Examples of electronic devices 102 may include supercomputers, personal computers, laptop computers, in-vehicle computing devices, mobile devices (such as smartphones, tablet computers, etc.), or a combination of any one or more of the above devices. It should be understood that the electronic devices described herein are merely exemplary and non-limiting, and for example, other different types of electronic devices may also be used.

[0033] According to an embodiment of the present disclosure, an embodiment of a content generation method is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions, and although a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in an order different from that shown here.

[0034] In this embodiment, a content generation method is provided, which can be used in electronic devices such as computers, tablet computers, etc. Figure 2 is a flow chart of a content generation method according to an embodiment of the present disclosure, such as Figure 2 As shown, the process includes the following steps:

[0035] Step S201: Acquire image data and question information corresponding to the image data.

[0036] The image data is the image that currently needs to be visually recognized; the question information is the instruction information for visual perception of the image data, such as "What is the economic difference between area A and area B?"

[0037] Specifically, a generative application is deployed in the electronic device, and the generative application provides an interactive page. Image data and corresponding question information can be input through the interactive page. Correspondingly, the generative application can obtain the image data and question information.

[0038] Among them, the image data can be selected and uploaded from the local electronic device to the interactive page, or it can be photographed and uploaded to the interactive page by calling the camera function, or it can be uploaded to the interactive page by taking a screenshot from a video or other file. There is no specific limitation on the way of uploading the image data here.

[0039] Step S202 : using the target generation model to identify semantic features of the question information, and locating the visual area corresponding to the question information in the image data according to the semantic features.

[0040] The target generative model is generated based on visual reproduction and multimodal information training. This target generative model can be trained based on a large language model architecture or a machine learning model architecture. As long as it can generate visual content from image data, the target generative model architecture is not limited here.

[0041] Visual reproduction refers to the target generative model dynamically selecting and reproducing visual area information related to the current question information during training and inference. Multimodal information refers to the target generative model generating content that matches the current question information based on multimodal features during training and inference. By training the target generative model based on visual reproduction and multimodal information, the target generative model performs visual perception on image data and generates corresponding response content based on the results of visual perception. In other words, the target generative model is configured to enhance fine-grained visual perception and multimodal reasoning capabilities, enabling the target generative model to fully perceive and understand the visual information in image data.

[0042] Specifically, the target generative model is deployed within a generative application to provide visual content generation for the application. When visual content generation is required for image data and question information, the image data and question information are input into the target generative model. The target generative model then identifies the semantic features carried by the question information and, based on these semantic features, detects and identifies the key regions corresponding to these features within the image data. These key regions are the visual regions corresponding to the question information.

[0043] Step S203: using the target generation model to identify visual information in the visual area, and generating target text content that matches the question information according to the visual information.

[0044] The target text content is the answer generated in response to the question information. This target text content is generated by reasoning based on the relationship between the question information and the visual information. Specifically, after determining the visual area, the target generation model analyzes the local image within the visual area, identifies the visual information carried in the local image (such as objects, scenes, colors, textures, data, text, etc.), parses the question information corresponding to the image data, and infers and generates the corresponding target text content based on the relationship between the question information and the visual information.

[0045] The content generation method provided in this embodiment generates a target generative model based on visual reproduction and multimodal information training, so that the target generative model has the ability to reproduce image features and the ability to reason about multimodal information, thereby enhancing the visual perception ability of the target generative model. Therefore, after obtaining the image data and the question information corresponding to the image data, the target generative model is used to identify the semantic features of the question information, and the visual area corresponding to the question information is located in the image data according to the semantic features, so as to fully identify the visual features in the image data and improve the processing ability of visual information. Subsequently, the target generative model is used to identify the visual information in the visual area, and the target text content that matches the question information is generated according to the visual information. Thus, the semantic features and visual information can be combined to perform reasoning for content generation, thereby achieving the effective integration of multimodal data and improving the content generation effect.

[0046] In this embodiment, a training method for a target generative model is provided, wherein the target generative model is applied to visual content generation. Figure 3 is a flow chart of a method for training a target generative model based on visual reproduction and multimodal information according to an embodiment of the present disclosure, such as Figure 3 As shown, the process includes the following steps:

[0047] Step S301: Acquire a visual training data set.

[0048] The visual training data set includes multiple visual training samples, each visual training sample includes a question information sample and its corresponding original image sample, an answer content sample and its corresponding image area sample, and the answer content sample includes a position signal of the image area sample.

[0049] The original image sample is the image sample that currently needs to be visually perceived; the answer content sample is the answer content for the question information sample, and there can be one or more answer content samples; each answer content sample has a corresponding image area sample, which is a feature map of any resolution in the original image sample.

[0050] Specifically, the answer content sample includes the text answer content and the position signal. The position signal indicates the position of the image area sample in the original image sample. The position signal includes the coordinate position corresponding to the image area sample. The position signal can be placed after the answer content, such as Figure 4 shown.

[0051] The visual training dataset is a pre-constructed dataset that includes visual clue annotations, including annotations of answer content, image area, position signal, etc. The target generative model is trained through the visual training dataset so that the target generative model can support visual content generation reasoning in complex scenes.

[0052] Step S302 : Slicing and visually encoding the original image samples to obtain first visual features corresponding to each image slice.

[0053] The first visual feature is an image feature generated by visually encoding the image slice. Specifically, the original image sample is sliced ​​by the target generative model to obtain multiple image slices, and the visual encoder in the target generative model is used to visually encode each image slice to extract the visual features corresponding to each image slice, such as Figure 5 shown.

[0054] Step S303: perform text encoding on the answer content sample to obtain a first text feature corresponding to the answer content sample.

[0055] The first text feature is a text semantic feature generated by encoding the answer content sample. Specifically, the text encoder in the target generative model encodes the answer content sample in the visual training sample to extract the text features corresponding to the answer content sample, such as Figure 5 shown.

[0056] Step S304: locating the visually relevant area corresponding to the answer content sample in the original image sample according to the position signal.

[0057] The visually relevant area is the key area where the answer content sample is located in the original image sample. Figure 5 As shown, by analyzing the position signal carried in the answer content sample, the coordinate position corresponding to the position information is determined, and the corresponding area in the original image sample is located according to the coordinate position. The corresponding area is the visually related area corresponding to the answer content sample in the original image sample.

[0058] Step S305 , extracting a second visual feature corresponding to the visually relevant area from the first visual feature, and performing visual reproduction according to the visually relevant area and the second visual feature, so as to match the visually relevant area with the answer content sample.

[0059] The second visual feature is the image feature of the image in the visually relevant area. The visually relevant area is composed of different areas in the original image sample. Here, the feature space of different image slices can be reintegrated to obtain image features of any resolution, so that the corresponding features can be directly extracted from the visual features obtained by the original visual encoder. Based on this, the visual features of the corresponding positions can be extracted from each first visual feature in combination with the position signal as the second visual feature, such as Figure 5 Then, during the model inference process, when the position signal of the visually relevant area is supervised, the second visual feature is inserted into the position of the position signal for visual reproduction, so that the visually relevant area matches the answer content sample.

[0060] Step S306: Combine the second visual feature and the first text feature into a multimodal feature, and perform training according to the multimodal feature so that the target generative model predicts the question-related area according to the question information sample, and generates visual content according to the visual information carried by the question-related area.

[0061] The second visual feature captured from the visually relevant region is combined with the first text feature corresponding to the visually relevant region to generate a multimodal feature. This multimodal feature is then used to train the target generative model's content generation process, enabling it to predict the question-related region based on the question information and its corresponding image data, thereby achieving visual reproduction of the question-related region. Subsequently, the visual information in the question-related region is identified, and the association between this visual information and the question information is combined to generate a reasoning chain for the visual content. This reasoning chain is then used to generate visual content that matches the question information.

[0062] The training method of the target generative model provided in this embodiment improves the target generative model's ability to focus on important visual information and reduces the interference of irrelevant information by locating the visually relevant area corresponding to the answer content sample in the original image sample according to the position signal, so as to make full use of the detailed information in the image. At the same time, the second visual feature is directly extracted from the first visual feature without the need for separate forward calculation or reasoning, thereby reducing the loss of computing resources and reasoning resources and facilitating more efficient training and reasoning. Combining text features and visual features for content generation training enables the target generative model to enhance visual perception by means of a multimodal thinking chain, facilitates the extraction of sufficiently detailed and accurate multimodal features in complex scenes, improves the accuracy and efficiency of multimodal reasoning, and enhances the accuracy of content generation.

[0063] In some optional implementations, locating the visually relevant area corresponding to the answer content sample in the original image sample according to the position signal includes:

[0064] Step a1: Identify the global coordinate position of the position signal in the original image sample.

[0065] Step a2: locate the enclosing area of ​​the global coordinate position in the original image sample, and determine the enclosing area as the visually relevant area of ​​the answer content sample in the original image sample.

[0066] The global coordinate position indicates the coordinate position determined by constructing a coordinate system based on the original image sample. Specifically, the position signal consists of a start mark, a global coordinate position, and an end mark, such as Figure 4 Shown as " <sot><h2 style=";text-align:left;direction:ltr">[x1,y1,x2,y2]<h2 style=";text-align:left;direction:ltr"> <eot>". Where [x1, y1] represents the coordinates of the upper left corner of the region, and [x2, y2] represents the coordinates of the lower right corner of the region; <sot>Indicates the start mark; <eot>Indicates the end marker.

[0067] The bounding area is the area covered by the global coordinate position. Specifically, the coverage of the bounding box (such as the length and width of the bounding box) is determined according to the global coordinate position. A box is selected in the original image sample according to the coverage of the bounding box to obtain the corresponding bounding area. This bounding area is the visually relevant area associated with the content of the answer content sample.

[0068] In the above embodiment, the global coordinate position is carried in the position signal so that the visual related area can be directly located using the global coordinate position, thereby eliminating the need for coordinate conversion and improving the positioning accuracy of the visual related area.

[0069] In some optional embodiments, the above method further includes:

[0070] Step b1: Obtain the autoregressive loss corresponding to the target generative model.

[0071] In step b2, the detection loss is integrated into the autoregressive loss to obtain the target detection loss.

[0072] In step b3, the object detection loss is used to supervise the visually relevant areas of the answer content sample in the original image sample.

[0073] The autoregressive loss is used to evaluate the autoregressive performance of the target generative model, which uses the information of a time step to predict the value of the next time step. This autoregressive loss can be represented by the cross entropy loss as follows:

[0074]

[0075] Among them, Loss represents the cross entropy loss value, which indicates the gap between the predicted result and the true label; C is the total number of samples; i represents the i-th predicted token; y i represents the i-th real token; p i represents the probability of correct prediction, y i log(p i ) represents the log-likelihood loss of the true label and the predicted probability.

[0076] Since the visual related area is actually represented by coordinate position, the detection loss is very important. det As a simple regression task to accurately align feature spaces, the cross entropy of visually relevant regions may be affected by quantization errors and discontinuous predictions. Therefore, it is necessary to integrate the detection loss into the autoregressive loss so that the target generative model can accurately locate visually relevant regions using continuous regression.

[0077] Specifically, the detection loss L det It is a combination of l1 and GIoU loss, expressed as follows:

[0078] L det =l1+βl GIoU ,

[0079] Among them, l1 is used to measure the absolute difference between the predicted bounding box coordinates and the true bounding box coordinates of the visual related area; the parameter β is the loss parameter, which can be set to β = 2 according to the empirical value.

[0080] For the center coordinates (x c ,y c ), a bounding box parameterized by width w and height h, l1 can be expressed as: in, is the predicted value.

[0081] The expression of GIOU loss is:

[0082]

[0083] Among them, UnionArea represents the combined area of ​​the predicted bounding box and the true bounding box; InterArea represents the intersection area of ​​the predicted bounding box and the true bounding box; C is the smallest box containing the predicted bounding box and the true bounding box, which can be expressed as:

[0084]

[0085] in, [x1, y1] represents the coordinates of the upper left corner of the true bounding box, and [x2, y2] represents the coordinates of the lower right corner of the true bounding box; represents the coordinate of the upper left corner of the predicted bounding box, Represents the coordinates of the lower right corner of the predicted bounding box.

[0086] In the above implementation, autoregressive loss training is combined with detection loss in training to supervise the coordinate area of ​​the answer content sample in the original image sample, so that the target generative model can more accurately detect and identify visually related areas from the image, facilitate visual reproduction in combination with visually related areas, so as to cope with complex scenes and deeply understand visual information.

[0087] In some optional implementations, obtaining a visual training data set includes:

[0088] Step c1: obtaining an original training data set, where the original training data set includes a plurality of original training samples, each of which includes a question information sample, an answer content sample, and an image sample.

[0089] The question information sample is the question content for visual perception of the image sample; the answer content sample is the real answer content corresponding to the question information sample; the image sample matches the question information sample, and is the image that currently needs to be visually perceived.

[0090] like Figure 6 As shown, each original training sample in the original training sample set includes a question information sample, an answer content sample, and an image sample. This original training sample set can be obtained from multiple public datasets, generated based on prompt descriptions using existing visual language models, manually constructed using a combination of data augmentation, or obtained using other hybrid methods. The method for obtaining the original training sample set is not specifically limited here.

[0091] In step c2, the annotation model is used to generate reasoning chains and predicted answers for question information samples and image samples, locate the key areas of the predicted answers in the image samples, and construct the initial annotation data with the reasoning chains, key areas, and predicted answers.

[0092] The labeling model is a cold-start model that is pre-trained to perform labeling tasks. The labeling model must have instruction-following capabilities and output diversity, and be able to perform well in target detection and visual reasoning tasks. Specifically, the labeling model can be trained based on deep learning models such as neural networks, convolutional neural networks (CNNs), recurrent neural networks (RNNs), or long short-term memory networks (LSTMs); it can also be trained based on a machine learning model architecture; it can also be trained based on a large language model architecture. There is no specific limitation on the labeling model here, as long as it can capture the correlation between the problem information and the image samples.

[0093] like Figure 6 As shown in the figure, for a given image sample and a corresponding question information sample, the annotation model generates a corresponding reasoning chain, which is then combined to give a predicted answer. At the same time, the reasoning chain is combined to locate the key areas in the image sample that are relevant to the predicted answer. The reasoning chain, key areas, and predicted answer constitute the initial annotation data.

[0094] Step c3: verify the initial annotation data and select valid annotation data from the initial annotation data according to the data verification result.

[0095] Encode the initial annotation data into a pre-set standard format (e.g., JSON), including the bounding boxes and semantic labels for each key region, to achieve standardization of the annotation data format. Verify the standardized initial annotation data to determine if the initial annotation data is valid.

[0096] like Figure 6 As shown in the figure, if it is determined that the initial annotation data fails to pass the verification, it means that the initial annotation data is invalid, and the invalid initial annotation data is deleted to improve the annotation quality of the initial annotation data. If it is determined that the initial annotation data passes the verification, it means that the initial annotation data is valid, and the valid initial annotation data is used as the valid annotation data for training the annotation model.

[0097] In step c4, the labeling model is trained using the valid labeling data and the preset inference data, and the valid training data generated by the trained labeling model is used to form a visual training data set.

[0098] The preset reasoning data is the refined text reasoning content selected from the open source dataset. Since the initial annotation model has not performed the reasoning task of the key area, the reasoning chain it generates, the predicted answers generated based on the reasoning chain, and the key areas contain a lot of erroneous data, resulting in a high rejection rate of the initial annotation data, which affects the speed of generating training data. Therefore, the annotation model is trained according to the valid annotation data, and is trained in combination with the preset reasoning data to enhance the reasoning mode of the annotation model and improve the annotation effectiveness of the annotation model for the training samples. Figure 6 As shown in the figure. The trained annotation model is then used to annotate the original training dataset, generating annotated data consisting of inference chains, predicted answers, and their corresponding key regions. Based on the data verification results, valid annotated data is extracted from the annotated data. Valid training data covering multiple public datasets is then extracted from the valid annotated data to form the training dataset.

[0099] In the above implementation, by using the annotation model to construct the initial annotation data and screening the effective annotation data from the initial annotation data to perform iterative optimization of the annotation model, the annotation pass rate of the annotation data is improved and the generation of effective training data is accelerated.

[0100] In some optional implementations, in step c2 above, constructing initial annotation data using the reasoning chain, key regions, and predicted answers includes:

[0101] Step c21: Obtain the inference description content of the key area in the inference chain.

[0102] In step c22, the key region is referenced in the inference description content, and the location information and semantic labels of the key region are integrated.

[0103] Step c23: construct initial annotation data using location information, semantic labels, key areas, and reasoning description content.

[0104] The inference description content is the textual description of the image content involved in the inference chain. Combining the semantic relevance between visual information and textual semantics, the inference description content corresponding to the visual information of the key areas is identified from the inference content of the inference chain. Subsequently, the identified key areas are explicitly referenced before the inference description content. These key areas are designated as visual reproduction areas during training. At the same time, the position information and semantic labels of the key areas are integrated into the inference description content so that the key areas can be located during training using the position information and visual reproduction can be determined using the semantic labels. This allows for annotation of position information, semantic labels, key areas, and inference description content, forming initial annotation data.

[0105] In some optional implementations, verifying the initial annotation data includes:

[0106] Step d1: verify the annotation format of the initial annotation data and generate a format verification result of the initial annotation data.

[0107] Step d2: Compare the predicted answer in the initial labeled data with the answer content sample to generate the answer correctness verification result.

[0108] Step d3: Visually verify the location information and semantic labels in the initial annotated data to generate a visual verification result.

[0109] Validation of the annotation format is used to ensure the parsability of the predicted answers, such as Figure 6 As shown, the answer can be extracted by locating the "FinalAnswer" field, and the format validity of the bounding box and semantic label can be verified to generate the corresponding format verification result.

[0110] For closed tasks (such as OCR and multiple-choice questions), the average normalized edit distance can be used to compare the similarity between the predicted answer and the answer content sample to determine whether the predicted answer is consistent with the true answer in the answer content sample. If the predicted answer is consistent with the true answer in the answer content sample, the answer generation is considered accurate. For example, if the similarity exceeds 98%, the answer generation can be considered accurate.

[0111] For open tasks, the semantic consistency between the inference chain and the answer content sample can be verified. If the inference chain is semantically inconsistent with the true answer in the answer content sample, the answer generation is judged to be incorrect and the annotated data is directly discarded. Semantic consistency can be set to a semantic similarity. If the semantic similarity exceeds a set value (such as 90%), it can be determined to be semantically consistent. For generated answers that are semantically consistent but have differences from the true answer, iterative corrections are made to make them closer to the true answer.

[0112] The position information in the initial annotation data and the image area corresponding to the semantic label should match. Visual verification is performed on the bounding box generated by the position information and the image area corresponding to the semantic label to determine whether the two match. If they do not match, the visual verification fails, otherwise the visual verification succeeds.

[0113] In the above implementation, the annotation effect of the initial annotation data is verified by combining format verification, correctness verification and visual verification, thereby screening effective annotation data through the verification results of multi-dimensional data verification to improve the quality of training data.

[0114] In some optional implementations, the above step d3 includes:

[0115] Step d31 : expanding the bounding box corresponding to the position information according to a preset rule to obtain a target bounding box.

[0116] In step d32, the target bounding box is matched with the semantic label to generate a visual verification result.

[0117] The preset rule is a pre-set boundary expansion rule, for example, expanding the bounding box by 20%. The bounding box corresponding to the position information is expanded according to the preset rule to obtain a target bounding box that meets the requirements, so as to retain more context information through the target bounding box.

[0118] Identify the visual information in the target bounding box, match the visual information with the image area corresponding to the semantic label, and generate a corresponding matching result. If the matching result indicates that the visual information matches the image area corresponding to the semantic label, the visual verification is judged to have passed. Otherwise, the visual verification is judged to have failed.

[0119] In the above implementation, the ability of the target generative model to handle complex visual semantic dependencies is enhanced by enlarging the bounding box to retain contextual information.

[0120] In some optional embodiments, the above method further includes:

[0121] In step e1, the reasoning chain is revised using the trained annotation model to obtain the target reasoning chain.

[0122] In step e2, key regions are labeled and predicted answers are generated according to the target reasoning chain.

[0123] In the process of data labeling using the trained labeling model, the original reasoning chain is revised through the trained labeling model, the derivation task is performed according to the revised target reasoning chain, the corresponding key areas are labeled, and the predicted answers are generated according to the visual information of the key areas.

[0124] In the above implementation, the data structure consistency and content generation accuracy in the reasoning process are ensured by revising the reasoning chain.

[0125] As a specific application example of the embodiment of the present disclosure, the performance of the target generative model trained by the present disclosure is compared with other models with the same number of parameters, such as Figure 7 As shown in the figure, experimental verification shows that compared to the baseline, the disclosed method can achieve VQA downstream task performance far higher than other model baselines using fewer tokens. At the same time, as the number of visual tokens increases, the visual perception performance of the target generative model of the disclosed method can be further improved, and the content generation effect based on visual perception will also be further improved.

[0126] The proposed scheme of jointly training the detection loss and the autoregressive loss of the large language model itself ( Figure 8 8-a), visual feature reproduction ( Figure 8 8-b), multimodal reasoning chain data ( Figure 8 8-c) in Figure 1 have been effectively verified, and even when one of the above three components is missing, the results are still better than those of the original baseline, proving the effectiveness of the three parts of detection loss, visual feature reproduction, and multimodal reasoning.

[0127] This embodiment also provides a content generation device for implementing the above-mentioned embodiments and preferred implementations. Details already described will not be repeated. As used below, the term "module" may refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, implementation using hardware, or a combination of software and hardware, is also possible and contemplated.

[0128] This embodiment provides a content generation device, such as Figure 9 Shown, including:

[0129] The acquisition module 501 is used to acquire image data and problem information corresponding to the image data.

[0130] The visual area positioning module 502 is used to identify the semantic features of the question information using the target generation model, and locate the visual area corresponding to the question information in the image data according to the semantic features.

[0131] The content generation module 503 is used to identify visual information in the visual area using a target generation model and generate target text content that matches the question information based on the visual information. The target generation model is generated based on visual reproduction and multimodal information training.

[0132] In some optional embodiments, the above device further includes:

[0133] A training module is used to train and generate a target generative model based on visual representation and multimodal information. Specifically, the training module includes:

[0134] The training data set acquisition unit is used to acquire a visual training data set, which includes multiple visual training samples. Each visual training sample includes a question information sample and its corresponding original image sample, an answer content sample and its corresponding image area sample, and the answer content sample includes a position signal of the image area sample.

[0135] The visual feature extraction unit is used to slice and visually encode the original image samples to obtain the first visual features corresponding to each image slice.

[0136] The text feature extraction unit is used to perform text encoding on the answer content sample to obtain a first text feature corresponding to the answer content sample.

[0137] The region positioning unit is used to locate the visually relevant region of the answer content sample in the original image sample according to the position signal.

[0138] The visual reproduction unit is used to extract the second visual feature corresponding to the visually relevant area from the first visual feature, and perform visual reproduction according to the visually relevant area and the second visual feature to match the visually relevant area with the answer content sample.

[0139] The multimodal training unit is used to combine the second visual feature and the first text feature into a multimodal feature, and perform training according to the multimodal feature so that the target generative model predicts the problem-related area according to the problem information sample, and generates visual content according to the visual information carried by the problem-related area.

[0140] In some optional implementations, the area positioning unit includes:

[0141] The coordinate recognition subunit is used to identify the global coordinate position of the position signal in the original image sample.

[0142] The positioning subunit is used to locate the enclosing area of ​​the global coordinate position in the original image sample, and determine the enclosing area as the visually relevant area of ​​the answer content sample in the original image sample.

[0143] In some optional implementations, the training module further includes:

[0144] The autoregressive loss acquisition unit is used to obtain the autoregressive loss corresponding to the target generative model.

[0145] The detection loss determination unit is used to fuse the detection loss into the autoregressive loss to obtain the target detection loss.

[0146] The supervision unit is used to supervise the visually relevant areas of the answer content sample in the original image sample using the target detection loss.

[0147] In some optional implementations, the training data set acquisition unit includes:

[0148] The original data set acquisition subunit is used to acquire an original training data set, which includes multiple original training samples. Each original training sample includes a question information sample, an answer content sample, and an image sample.

[0149] The data annotation subunit is used to use the annotation model to generate reasoning chains and predicted answers for question information samples and image samples, locate the key areas of the predicted answers in the image samples, and construct initial annotation data with the reasoning chains, key areas and predicted answers.

[0150] The data verification subunit is used to verify the initial annotation data and filter out valid annotation data from the initial annotation data according to the data verification results.

[0151] The training data extraction subunit is used to train the annotation model using valid annotation data and preset inference data, and to form a visual training data set using the valid training data generated by the trained annotation model.

[0152] In some optional embodiments, the data annotation subunit is specifically used to obtain the reasoning description content of the key area in the reasoning chain; reference the key area in the reasoning description content, and integrate the location information and semantic labels of the key area; and construct initial annotation data with the location information, semantic labels, key areas, and reasoning description content.

[0153] In some optional embodiments, the data verification subunit is specifically used to verify the annotation format of the initial annotation data to generate a format verification result of the initial annotation data; compare the predicted answer in the initial annotation data with the answer content sample to generate an answer correctness verification result; and visually verify the position information and semantic labels in the initial annotation data to generate a visual verification result.

[0154] In some optional implementations, the data verification subunit is further configured to expand the bounding box corresponding to the location information according to preset rules to obtain a target bounding box; and match the target bounding box with a semantic label to generate a visual verification result.

[0155] In some optional implementations, the training data set acquisition unit further includes:

[0156] The revision subunit is used to revise the reasoning chain using the trained annotation model to obtain the target reasoning chain.

[0157] The annotation subunit is used to annotate key areas and generate predicted answers according to the target reasoning chain.

[0158] The content generation device provided by the embodiments of the present disclosure can execute the content generation method provided by any embodiment of the present disclosure, and has the functional modules and beneficial effects corresponding to the execution method. A target generative model is generated based on visual reproduction and multimodal information training, so that the target generative model has the ability to reproduce image features and the ability to reason about multimodal information, thereby enhancing the visual perception ability of the target generative model. Therefore, after obtaining the image data and the question information corresponding to the image data, the target generative model is used to identify the semantic features of the question information, and the visual area corresponding to the question information is located in the image data according to the semantic features, so as to fully identify the visual features in the image data and improve the processing ability of visual information. Subsequently, the target generative model is used to identify the visual information in the visual area, and the target text content matching the question information is generated according to the visual information. Thus, the semantic features and visual information can be combined to perform reasoning for content generation, thereby achieving effective integration of multimodal data and improving the content generation effect.

[0159] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.

[0160] Figure 10 A schematic structural diagram of an electronic device provided in an embodiment of the present disclosure.

[0161] The following specific reference Figure 10 , which shows a schematic diagram of the structure of an electronic device suitable for implementing the embodiments of the present disclosure. The electronic device may include a processor (e.g., a central processing unit, a graphics processing unit, etc.) 601, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 602 or a program loaded from a memory 608 into a random access memory (RAM) 603. Various programs and data required for the operation of the electronic device are also stored in the RAM 603. The processor 601, ROM 602, and RAM 603 are connected to each other via a bus 604. An input / output (I / O) interface 605 is also connected to the bus 604.

[0162] Typically, the following devices may be connected to the I / O interface 605: an input device 606 including, for example, a touch screen, a touchpad, a keyboard, a mouse, a camera, a microphone, an accelerometer, a gyroscope, etc.; an output device 607 including, for example, a liquid crystal display (LCD), a speaker, a vibrator, etc.; a memory 608 including, for example, a magnetic tape, a hard disk, etc.; and a communication device 609. The communication device 609 may allow the electronic device to communicate with other devices wirelessly or by wire to exchange data. Although Figure 10 An electronic device having various devices is shown, but it should be understood that it is not required to implement or possess all of the devices shown, and more or fewer devices may be implemented or possessed instead.

[0163] In particular, according to an embodiment of the present disclosure, the process described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a non-transitory computer-readable medium, and the computer program includes program code for executing the method shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication device 609, or installed from the memory 608, or installed from the ROM 602. When the computer program is executed by the processor 601, the above-mentioned functions defined in the content generation method of the embodiment of the present disclosure are performed.

[0164] Figure 10 The electronic device shown is only an example and should not limit the functions and scope of use of the embodiments of the present disclosure.

[0165] The embodiments of the present disclosure also provide a computer-readable storage medium. The above-mentioned method according to the embodiments of the present disclosure can be implemented in hardware, firmware, or implemented as a computer code that can be recorded in a storage medium, or implemented as a computer code that is originally stored in a remote storage medium or a non-temporary machine-readable storage medium and downloaded through a network and will be stored in a local storage medium, so that the method described herein can be stored in such software processing on a storage medium using a general-purpose computer, a dedicated processor, or programmable or dedicated hardware. Among them, the storage medium can be a magnetic disk, an optical disk, a read-only storage memory, a random access memory, a flash memory, a hard disk or a solid-state drive, etc.; further, the storage medium can also include a combination of the above-mentioned types of memory. It can be understood that a computer, a processor, a microprocessor controller or programmable hardware includes a storage component that can store or receive software or computer code. When the software or computer code is accessed and executed by a computer, a processor or hardware, the content generation method shown in the above embodiment is implemented.

[0166] A portion of the present disclosure may be applied as a computer program product, such as a computer program instruction, which, when executed by a computer, can call or provide the method and / or technical solution according to the present disclosure through the operation of the computer. Those skilled in the art should understand that the form in which the computer program instruction exists in a computer-readable medium includes but is not limited to a source file, an executable file, an installation package file, etc. Accordingly, the way in which the computer program instruction is executed by the computer includes but is not limited to: the computer directly executes the instruction, or the computer compiles the instruction and then executes the corresponding compiled program, or the computer reads and executes the instruction, or the computer reads and installs the instruction and then executes the corresponding installed program. Here, the computer-readable medium can be any available computer-readable storage medium or communication medium that can be accessed by the computer.

[0167] Although the embodiments of the present disclosure have been described with reference to the accompanying drawings, those skilled in the art may make various modifications and variations without departing from the spirit and scope of the present disclosure, and such modifications and variations are all within the scope defined by the appended claims.< / eot> < / sot> < / eot> < / sot>

Claims

1. A content generation method, characterized in that: The method comprises: Acquiring image data and problem information corresponding to the image data; identifying semantic features of the question information using a target generative model, and locating a visual area corresponding to the question information in the image data according to the semantic features; Using the target generation model to identify visual information in the visual area, and generating target text content that matches the question information according to the visual information; The target generative model is generated based on visual reproduction and multimodal information training.

2. The method according to claim 1, characterized in that Training the target generative model based on visual representation and multimodal information includes: Acquire a visual training data set, wherein the visual training data set includes a plurality of visual training samples, each of the visual training samples includes a question information sample and its corresponding original image sample, an answer content sample and its corresponding image region sample, and the answer content sample includes a position signal of the image region sample; Slicing and visually encoding the original image samples to obtain first visual features corresponding to each image slice; Performing text encoding on the answer content sample to obtain a first text feature corresponding to the answer content sample; Locating a visually relevant area of ​​the answer content sample in the original image sample according to the position signal; extracting a second visual feature corresponding to the visually relevant area from the first visual feature, and performing visual reproduction according to the visually relevant area and the second visual feature, so as to match the visually relevant area with the answer content sample; The second visual feature and the first text feature are combined into a multimodal feature, and training is performed according to the multimodal feature so that the target generative model predicts the question-related area according to the question information sample and generates visual content according to the visual information carried by the question-related area.

3. The method according to claim 2, characterized in that The locating the visually relevant area corresponding to the answer content sample in the original image sample according to the position signal includes: Identifying the global coordinate position of the position signal in the original image sample; An enclosing area of ​​the global coordinate position is located in the original image sample, and the enclosing area is determined as the visually relevant area of ​​the answer content sample in the original image sample.

4. The method according to claim 3, characterized in that Also includes: Obtaining the autoregressive loss corresponding to the target generative model; The detection loss is integrated into the autoregressive loss to obtain the target detection loss; The object detection loss is used to supervise the visually relevant region of the answer content sample in the original image sample.

5. The method according to claim 2, characterized in that The obtaining of a visual training data set includes: Acquire an original training data set, wherein the original training data set includes a plurality of original training samples, each of the original training samples includes a question information sample, an answer content sample, and an image sample; generating a reasoning chain and a predicted answer for the question information sample and the image sample using a labeling model, locating a key area of ​​the predicted answer in the image sample, and constructing initial labeling data using the reasoning chain, the key area, and the predicted answer; Verifying the initial annotation data, and selecting valid annotation data from the initial annotation data according to the data verification result; The labeling model is trained using the valid labeling data and the preset inference data, and the valid training data generated by the trained labeling model is used to form a visual training data set.

6. The method according to claim 5, characterized in that Constructing initial annotation data based on the reasoning chain, the key area, and the predicted answer includes: Obtaining the reasoning description content of the key area in the reasoning chain; Reference the key region in the reasoning description content, and integrate the position information and semantic label of the key region; The initial annotation data is constructed using the position information, the semantic label, the key area, and the reasoning description content.

7. The method according to claim 6, characterized in that Verifying the initial annotation data includes: Verifying the annotation format of the initial annotation data and generating a format verification result of the initial annotation data; Comparing the predicted answer in the initial annotated data with the answer content sample to generate an answer correctness verification result; Visual verification is performed on the position information and the semantic label in the initial annotation data to generate a visual verification result.

8. The method according to claim 7, characterized in that The visual verification of the position information and the semantic label in the initial annotation data to generate a visual verification result includes: Expanding the bounding box corresponding to the position information according to a preset rule to obtain a target bounding box; The target bounding box is matched with the semantic label to generate the visual verification result.

9. The method according to claim 5, characterized in that Also includes: Revising the reasoning chain using the trained annotation model to obtain a target reasoning chain; The key areas are labeled and the predicted answers are generated according to the target reasoning chain.

10. A content generating device, characterized in that: The device comprises: An acquisition module, configured to acquire image data and problem information corresponding to the image data; a visual area positioning module, configured to identify semantic features of the question information using a target generation model, and locate a visual area corresponding to the question information in the image data according to the semantic features; a content generation module, configured to use the target generation model to identify visual information in the visual area and generate target text content matching the question information according to the visual information; The target generative model is generated based on visual reproduction and multimodal information training.

11. An electronic device, characterized in that: include: A memory and a processor, wherein the memory and the processor are communicatively connected to each other, the memory stores computer instructions, and the processor executes the content generation method according to any one of claims 1 to 9 by executing the computer instructions.

12. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a computer to execute the content generation method according to any one of claims 1 to 9.

13. A computer program product, characterized in that The method comprises computer instructions for causing a computer to execute the content generation method according to any one of claims 1 to 9.