Model training and image description generation method and device, equipment and medium
By iteratively fine-tuning the image description generation model, combining incremental optimization loss function and positive and negative image description, the problem of image description deviation after the introduction of context information is solved, and the accuracy of description and generalization ability of the model are improved.
Patent Information
- Application Number
- CN202510179350.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-05-16
AI Technical Summary
When the existing image description generation technology introduces context information, it is easy for the generated description to deviate from the image content, reducing the accuracy and practicality of the description.
The pre-trained basic image description generation model is iteratively fine-tuned by obtaining the image sample set, including image samples, context information and processing instructions, as well as positive and negative image descriptions. Using incremental optimization loss function, combined with the probability distribution of the current model and the underlying model, optimize the model to generate a more accurate image description.
Improve the accuracy and generalization ability of the image description generation model, ensure that the generated description is more in line with the image content, and reduce the probability of incorrect or irrelevant descriptions.
Smart Images

Figure CN120014384A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of multimodal image processing and deep learning technology, and in particular to a training method, device, equipment and medium for image description generation model and image description generation. Background Art
[0002] Image description generation technology is a key research direction in the current field of artificial intelligence. It integrates the technologies of natural language processing (NLP) and computer vision (CV). Its goal is to automatically output the corresponding text description based on a given image.
[0003] At present, the image description generation method mainly relies on the deep learning framework, which organically integrates the two major technical directions of image processing and natural language generation. When using the deep learning framework to generate image descriptions, the encoder first extracts and encodes the image features, and then the decoder generates the corresponding text description based on the encoded information. For example, visual models such as Vision Transformer (ViT) and Residual Network (ResNet) are often used for image encoding, while language models such as qwen and llama are used for decoding. In this method, a picture is usually paired with a simple instruction (such as "Please briefly describe the picture") as the model input, and the model generates the image description text. However, in actual application scenarios, the same picture often has multiple different interpretations, which mainly depends on the observer's perspective and background knowledge.
[0004] Therefore, in order to more accurately generate image descriptions that meet user needs, contextual information can be introduced as the background of the image application, so as to achieve a fusion understanding of the context and image content, allowing the model to capture key information more specifically and generate descriptions that are more in line with user expectations. However, if the contextual information is too strong or misleading, the model may generate text that is similar to or biased towards the context, and thus ignore the description of the image itself. This deviation will cause the generated description to be inconsistent with the image content, reducing the accuracy and practicality of the description.
[0005] Based on the above situation, how to ensure accurate description of the image itself while introducing contextual information has become one of the important challenges facing current image description generation technology. Summary of the invention
[0006] The present application provides a method, apparatus, device and medium for training an image description generation model and generating image description, which are used to solve the problem that the answers generated by the model in the existing image description generation task have angular deviation or ambiguity.
[0007] In a first aspect, the present application provides a method for training an image description generation model, the method comprising:
[0008] Acquire an image sample set; wherein the image sample set includes model input data and reference image descriptions corresponding to the model input data, any model input data includes an image sample, context information of the image sample and a processing instruction, and the reference image description includes a positive image description and a negative image description;
[0009] Based on the image sample set, iteratively fine-tuning the pre-trained basic image description generation model; wherein the basic image description generation model is a model that already has the ability to generate image descriptions;
[0010] Among them, in any iterative fine-tuning process:
[0011] For any model input data, a training data pair is determined based on the model input data and any reference image description corresponding to the image sample; a first probability distribution of the image description is obtained based on the model input data by using the currently fine-tuned image description generation model; and a second probability distribution of the image description is obtained based on the model input data by using the basic image description generation model; a loss value is determined based on the first probability distribution, the second probability distribution and the reference image description by using an incremental optimization loss function;
[0012] Fine-tune the currently fine-tuned image description generation model according to each loss value to obtain a trained image description generation model;
[0013] The incremental optimization loss function is expressed by the following formula:
[0014] L(π θ ,π ref )=E x,y∈D [w(y)(1-v(x,y;β))] (1)
[0015]
[0016] z ref =E x∈D [βKL(π θ (y|x)||π ref (y|x))] (4)
[0017] Wherein, L() represents the incremental optimization loss function, π θ represents the currently fine-tuned image description generation model, π ref represents the basic image description generation model, E() represents the expected function, x represents the model input data, y represents the reference image description, D represents the image sample set, w(y) represents the weight function of the reference image description, λ + represents the preset weight of the positive image description, λ - represents the preset weight of the negative image description, v(x, y; β) represents the confidence of the reference image description output by the currently fine-tuned image description generation model, σ represents the Sigmoid function, β represents the temperature hyperparameter, which is used to adjust the smoothness of the output of the currently fine-tuned image description generation model, and π θ (y|x) represents the first probability distribution, π ref (y|x) represents the second probability distribution, z ref represents the expected log-likelihood ratio of the basic image description generation model, KL(π θ (y|x)||π ref (y|x)) represents the Kullback-Leibler dispersion between the first probability distribution and the second probability distribution.
[0018] In a second aspect, the present application also provides a method for generating image description based on the above-mentioned model, the method comprising:
[0019] Acquire input data of the model to be processed; wherein the input data of the model to be processed includes an image to be processed, a context of the image to be processed, and a processing instruction of the image to be processed;
[0020] The image description information of the image to be processed is obtained by generating a model through a pre-trained image description based on the model input data to be processed.
[0021] In a third aspect, the present application further provides a training device for an image description generation model, the device comprising:
[0022] An acquisition module, used to acquire an image sample set; wherein the image sample set includes model input data and reference image descriptions corresponding to the model input data, any model input data includes an image sample, context information of the image sample and a processing instruction, and the reference image description includes a positive image description and a negative image description;
[0023] A training module, used for iteratively fine-tuning a pre-trained basic image description generation model based on the image sample set; wherein the basic image description generation model is a model that already has the ability to generate image descriptions;
[0024] Among them, in any iterative fine-tuning process:
[0025] For any model input data, a training data pair is determined based on the model input data and any reference image description corresponding to the image sample; a first probability distribution of the image description is obtained based on the model input data by using the currently fine-tuned image description generation model; and a second probability distribution of the image description is obtained based on the model input data by using the basic image description generation model; a loss value is determined based on the first probability distribution, the second probability distribution and the reference image description by using an incremental optimization loss function;
[0026] Fine-tune the currently fine-tuned image description generation model according to each loss value to obtain a trained image description generation model;
[0027] The incremental optimization loss function is expressed by the following formula:
[0028] L(π θ ,π ref )=E x,y∈D [w(y)(1-v(x,y;β))] (1)
[0029]
[0030] z ref =E x∈D [βKL(π θ (y|x)||π ref (y|x))] (4)
[0031] Wherein, L() represents the incremental optimization loss function, π θ represents the currently fine-tuned image description generation model, π ref represents the basic image description generation model, E() represents the expected function, x represents the model input data, y represents the reference image description, D represents the image sample set, w(y) represents the weight function of the reference image description, λ + represents the preset weight of the positive image description, λ - represents the preset weight of the negative image description, v(x, y; β) represents the confidence of the reference image description output by the currently fine-tuned image description generation model, σ represents the Sigmoid function, β represents the temperature hyperparameter, which is used to adjust the smoothness of the output of the currently fine-tuned image description generation model, and π θ (y|x) represents the first probability distribution, π ref (y|x) represents the second probability distribution, z refrepresents the expected log-likelihood ratio of the basic image description generation model, KL(π θ (y|x)||π ref (y|x)) represents the Kullback-Leibler dispersion between the first probability distribution and the second probability distribution.
[0032] In a fourth aspect, the present application further provides an image description generating device based on the above-mentioned model, the device comprising:
[0033] An acquisition unit, configured to acquire model input data to be processed; wherein the model input data to be processed includes an image to be processed, a context of the image to be processed, and a processing instruction of the image to be processed;
[0034] The processing unit is used to generate a model through pre-trained image description and obtain image description information of the image to be processed based on the input data of the model to be processed.
[0035] In a fifth aspect, the present application provides a computer device, comprising a processor, wherein the processor is used to implement the steps of the training method of the image description generation model as described above when executing a computer program stored in a memory, or to implement the steps of the image description generation method as described above.
[0036] In a sixth aspect, the present application provides a computer-readable storage medium storing a computer program, which, when executed by a processor, implements the steps of the training method of the image description generation model as described above, or implements the steps of the image description generation method as described above.
[0037] The beneficial effects of this application are as follows:
[0038] 1. By introducing positive and negative image descriptions as reference image descriptions, the image description generation model not only learns how to generate positive descriptions that match the image content during training, but also learns to distinguish negative descriptions that do not match the image. This makes the image descriptions generated by the model more accurate, detailed, and in line with the actual image content, reducing the probability of generating incorrect or irrelevant descriptions.
[0039] 2. The iterative fine-tuning process combined with the incremental optimization loss function enables the model to continuously optimize on top of the capabilities of the base model. By learning and adjusting different training data pairs in multiple iterations, the model can adapt to various types of images and diverse contextual information, thereby improving its generalization capabilities in different data sets and scenarios. This means that the model can generate reasonable and accurate descriptions when faced with unseen images.
[0040] 3. The trained image description generation model can introduce the contextual information of the image when generating image descriptions, thereby avoiding description deviation caused by a single image input. BRIEF DESCRIPTION OF THE DRAWINGS
[0041] In order to more clearly illustrate the technical solutions in the embodiments of the present application, the drawings required for use in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying creative labor.
[0042] Figure 1 A schematic diagram of a training process of an image description generation model provided in an embodiment of the present application;
[0043] Figure 2 A schematic diagram of a process for generating an image description provided in an embodiment of the present application;
[0044] Figure 3 A schematic diagram of the structure of a training device for an image description generation model provided in an embodiment of the present application;
[0045] Figure 4 A schematic diagram of the structure of an apparatus for generating image description provided in an embodiment of the present application;
[0046] Figure 5 It is a structural schematic diagram of a computer device provided in an optional embodiment of the present application. DETAILED DESCRIPTION
[0047] In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present application.
[0048] In order to introduce contextual details as auxiliary information of images in the image description generation task, and at the same time avoid the interference of context on image information and improve the accuracy of image description generation, the present application provides a training method, device, equipment and medium for image description generation model and image description generation.
[0049] Embodiment 1:
[0050] This application provides a training method for an image description generation model. Figure 1 A schematic diagram of a training process of an image description generation model provided in an embodiment of the present application, the process comprising:
[0051] S101: Obtain an image sample set; wherein the image sample set includes model input data and a reference image description corresponding to the model input data, any model input data includes an image sample, context information of the image sample and a processing instruction, and the reference image description includes a positive image description and a negative image description.
[0052] In the present application, the training method of the image description generation model is applied to a computer device, which may be an intelligent terminal, such as a computer, a robot, etc., or a server, such as an application server, a business server, etc.
[0053] The present application provides a training method for an image description generation model, which obtains a specific image sample set, iteratively fine-tunes a pre-trained basic image description generation model, and finally obtains a trained image description generation model, so that the image description generation model can effectively combine the context information of the image to generate an image description. The image sample set is a basic data set for the training model, which includes model input data and reference image descriptions corresponding to the model input data, and any model input data includes at least the following three types:
[0054] A. Image samples: They can be images of various types to ensure the diversity and representativeness of image samples. For example, natural scenery images, human images, animal images, etc. These images can be obtained from public image datasets (such as COCO, ImageNet, etc.) or document datasets (such as product manuals, papers, reports, financial statements, etc.), or from images or documents in specific fields collected by users themselves. The format of the image sample can be a common image format, such as JPEG, PNG, etc.
[0055] In a possible implementation, the image samples are obtained by preprocessing the images in the original data; wherein the preprocessing includes but is not limited to: enhancement, noise reduction, and size normalization.
[0056] In the training process of the image description generation model, the quality of the image samples has an important impact on the performance of the model. The images in the original data often have problems such as uneven quality and diverse formats, and direct use may affect the training effect of the model. Therefore, it is necessary to preprocess the images in the original data to improve the quality and consistency of the images. The preprocessing operations in this embodiment include but are not limited to enhancement, noise reduction and size normalization. Through these processes, the image can be made clearer, the features are more prominent, and the size of the image is unified, which helps the model to better learn the features of the image and improve the accuracy and quality of image description generation.
[0057] B. Contextual information of the image sample, which is used to provide additional information about the image sample to help the image description generation model better understand the image content. The contextual information can include metadata such as the shooting time, shooting location, shooting equipment, etc. of the image sample, or it can be a text description, label, etc. related to the image sample. For example, for a picture containing an indicator icon of the high beam, the contextual information can be "The method to turn on the high beam on the car is as follows: 1. Start the vehicle: Make sure the vehicle is in the starting state. 2. Find the light control lever: Located on the left side of the steering wheel, or on the left side of the dashboard. 3. Turn on the lights: Turn the light control lever to the low beam position (a light bulb icon). 4. Switch to high beam: Push the light control lever forward (away from you). 4. Confirm the operation: The high beam indicator light (as shown below) usually appears on the dashboard, indicating that the high beam is turned on.", or "The following is a picture from the car user manual.".
[0058] In a possible implementation, the context information corresponding to any image sample is obtained in the following manner:
[0059] If the original data to which the image sample belongs is of the image-text interlaced type, determining the context content of the position where the image sample is located in the original data as the context information of the image sample;
[0060] If the original data to which the image sample belongs is of a pure picture type, the application background information of the original data is determined as the context information of the image sample.
[0061] When obtaining the context information of any image sample, the context information of the image sample can be determined based on the type of original data to which the image sample belongs. The following two types are described:
[0062] Type 1: Interlaced text and images.
[0063] The original data of interlaced text and images usually refers to documents, web pages, e-books, etc. that contain mixed arrangement of text and images. In this type of data, text and images complement each other and convey information together. For example, a user manual about a product will be interspersed with pictures of the product, and the text part may describe the parts, functions, usage methods and other information of the product, and these texts are closely related to the corresponding pictures. Therefore, when obtaining the context information of the image sample in the original data of interlaced text and images, the context content of the location of the image sample in the original data can be determined as the context information of the image sample. Exemplarily, first, determine the specific location of the image sample in the original data. For electronic documents, the image can be located by the structural information of the document (such as page number, paragraph number, etc.); for web pages, the location of the image can be determined by HTML tags and CSS styles. After determining the location of any image sample, extract the text content around the location as context information. The extraction range can be adjusted according to specific needs, and generally can include paragraphs, titles, comments, etc. before and after the image. For example, in a news report, the image may be located in the middle of a paragraph. At this time, the paragraph and the adjacent paragraphs before and after can be extracted as context information.
[0064] In a possible implementation manner, determining the context content of the location of the image sample in the original data as the context information of the image sample includes:
[0065] The context content at the location of the image sample in the original data and within a pre-configured text length range is determined as the context information of the image sample.
[0066] In order to obtain context information more reasonably and effectively, the context content within the pre-configured text length range at the location of the image sample in the original data can be determined as the required context information to avoid the problem of extracting too much irrelevant information or insufficient information extraction, making the context information more targeted and effective, thereby improving the effect of model training. Among them, the configuration of the text length range needs to take into account many factors. On the one hand, it is necessary to ensure that the extracted context information is rich enough to contain key information that is important for image understanding, such as the characteristics of objects in the image, the background of the scene, etc. On the other hand, it is necessary to avoid extracting too long text, which leads to too much irrelevant information and increases the burden of model processing. Therefore, the appropriate length range can be determined according to different types of original data and specific task requirements. For example, for short news pictures, the text length range can be set relatively small; while for images in academic documents with rich content, the length range can be appropriately increased. Specifically, when configuring the text length range, it can be initially set through manual experience, or it can be used by machine learning methods to find the optimal text length range by experimenting and analyzing a large amount of data. In addition, the text length range can also be configured in the form of number of characters (for example, 1000 characters above and below the image sample, etc.), number of words (for example, 1000 words above and below the image sample, etc.), or number of paragraphs (for example, the content of the next paragraph above and below the image sample, etc.).
[0067] Type 2: Pure picture type.
[0068] Pure image type raw data refers to a data set that only contains images without directly related text descriptions, such as image libraries, photo albums, statistical charts in financial statements, etc. The images in this type of data may have specific purposes, backgrounds or themes, but relevant information needs to be obtained from external channels to understand the meaning of the images. Therefore, when obtaining image samples in pure image type raw data, the application background information of the raw data can be determined as the context information of the image sample. Exemplarily, the source and purpose of the raw data are collected in advance for pure image type raw data. For example, the application background information can be obtained by communicating with the data provider, consulting relevant documents or studying the data collection method. For example, if the image is collected from a commercial advertising campaign, then the theme, target audience, delivery channel and other information of the advertisement can be understood. The parts related to the image sample are screened out from the obtained application background information and sorted. For example, if the image is an advertising picture used to promote a certain mobile phone, the model, features, time to market and other information of the mobile phone can be extracted as context information.
[0069] C. Processing instructions are used to instruct the model to perform specific processing on the image or generate a specific type of description. The processing instructions can be simple text instructions, such as "Please refer to the surrounding text to briefly describe the image content", "Generate a detailed image description, including color and shape information", etc.
[0070] For each model input data obtained, a reference image description can be compiled for the pattern sample in the model input data. The reference image description includes a positive image description and a negative image description. The positive image description refers to an accurate and appropriate description of the content in the image sample from the correct perspective. For example, for an image containing a cloud and two mobile devices, the positive image description can be "the two mobile devices are communicating via the Internet." The negative image description refers to a description that is inconsistent with or inaccurate with the content in the image sample, and is used to enhance the model's recognition ability. For the above image containing a cloud and two mobile devices, the negative image description can be "a cloud is floating between the two mobile devices."
[0071] In a possible implementation, the negative image description includes one or more of the following: an incorrect perspective negative image description and a context-interference negative image description; wherein the incorrect perspective negative image description is a correct description of the image sample under an incorrect perspective, and the context-interference negative image description is an image description that is interfered by the context information of the image sample and deviates from the core information of the image sample or the description is not specific enough.
[0072] In the training process of the image description generation model, negative image description plays an important role. It can help the model better distinguish between correct and incorrect image descriptions, thereby improving the model's ability to generate accurate image descriptions. Therefore, in this application, negative image descriptions can be subdivided into two types: wrong perspective negative image descriptions and context interference negative image descriptions. By clarifying the definitions and characteristics of these two types of negative image descriptions, negative image descriptions can be generated and utilized more specifically, further optimizing the training effect of the image description generation model. The following is an introduction to the two types of negative image descriptions:
[0073] Type 1: Wrong perspective negative image description.
[0074] Wrong perspective negative image description refers to the correct description of the image sample under the wrong perspective. The "wrong perspective" here means that the observation angle, emphasis or understanding method based on the description is inconsistent with the perspective required to correctly understand the image, but the description content itself is accurate for the wrong perspective. For example, for an image of an artificial intelligence server, its core display focus may be the high-performance chip and cooling system of the server, but describing it from the color or appearance texture of the server shell is a wrong perspective negative image description. By introducing this type of negative description, the model can learn the differences in descriptions under different perspectives and enhance its sensitivity to correct perspective descriptions. Therefore, when generating wrong perspective negative image descriptions, wrong perspective negative image descriptions can be generated by simulating different observation positions. For example, for a smart watch product, its front dial is the core display element, and the correct description may revolve around the design, function and display effect of the dial. If it is described from the perspective of the watch strap, such as "the strap of this watch is soft, comfortable to wear, and the color is coordinated with the dial", although the description content is accurate, it does not highlight the core element of the dial, which is a wrong perspective negative image description. In addition, the description can be deliberately changed to deviate from the core content of the image. For example, for a virtual reality (VR) helmet, its key features are high-resolution display, accurate tracking system and comfortable wearing experience. However, if the description is "the outer packaging of this VR helmet is beautifully designed, and cool patterns and product models are printed on the box", the focus is placed on the outer packaging, while the core function of the helmet itself is ignored, which is a wrong perspective negative image description.
[0075] Type 2: Context interferes with negative image description.
[0076] Context-interfered negative image description refers to an image description that deviates from the core information of the image sample or is not specific enough when it is interfered by the context information of the image sample. Although context information helps to understand the image to a certain extent, if it is over-reliant on or misinterpreted, it may cause the description to deviate from the core content of the image. For example, for an image showing a 5G base station, the context information mentions the broad application prospects of 5G technology, and the description may over-emphasize the application prospects, while insufficiently describing the core information such as the equipment structure and signal coverage of the base station. Therefore, when generating a context-interfered negative image description of any image sample, the context information can be deliberately over-associated with the image description to deviate from the core content. For example, for an image of an intelligent robot, the context information mentions the development trend of artificial intelligence in the medical field. The context-interfered negative image description may be "With the rapid development of artificial intelligence in the medical field, this intelligent robot is expected to play an important role in medical diagnosis in the future", without describing the core information such as the robot's appearance and functional modules in detail. In addition, a broad and general description can be made based on the context information, lacking the characterization of the specific characteristics of the technology product or scene. For example, for an image of a blockchain data center, the context information mentions the decentralized characteristics of blockchain technology. A negative image description with contextual interference may be "This blockchain data center embodies the decentralized advantages of blockchain technology and provides reliable guarantees for data storage and processing", without specifying key features such as the data center's server configuration and data processing capabilities.
[0077] For example, for a model input data, the image sample in the model input data is a picture containing an indicator light icon of a high beam, and its context information is "The method to turn on the high beam on the car is as follows: 1. Start the vehicle: Make sure the vehicle is in the started state. 2. Find the light control lever: Located on the left side of the steering wheel, or on the left side of the dashboard. 3. Turn on the lights: Turn the light control lever to the low beam position (a light bulb icon). 4. Switch to high beam: Push the light control lever forward (away from you). 4. Confirm the operation: The high beam indicator light usually appears on the dashboard (as shown below), indicating that the high beam is on." The processing instruction is "Please refer to the surrounding text and briefly describe the content of the picture.". A positive image of the model's input data is described as "This is a schematic diagram of a car's high beam. The entire picture has a blue background. There are 5 short horizontal lines on the left, representing the high-beam light, and an elliptical circle on the right, representing the lamp body." A wrong perspective negative image is described as "This is a road schematic diagram. There are 5 short horizontal lines on the left, representing the zebra crossing, and a circle on the right, representing the circular green belt at the intersection." A context interference negative image is described as "This is a schematic diagram of a car's high beam. When this image appears on the car's dial, it means that the high beam is on."
[0078] S102: Iteratively fine-tune a pre-trained basic image description generation model based on the image sample set; wherein the basic image description generation model is a model that already has image description generation capability.
[0079] After obtaining the image sample set based on the above embodiment, the pre-trained basic image description generation model can be iteratively fine-tuned based on the image sample set. The basic image description generation model is pre-trained on a large-scale image dataset, has learned rich image features and language knowledge, and has the ability to generate image descriptions. For example, MLLM (Multimodal Large Language Model), ViT, ResNet, etc.
[0080] Exemplarily, in any iterative fine-tuning process, for any model input data, a training data pair is determined based on the model input data and any reference image description corresponding to the image sample. For example, for a model input data (including an image sample, context information and processing instructions) and a positive image description, they are combined into a training data pair; similarly, for the same model input data and a negative image description, a training data pair can also be formed. Through the currently fine-tuned image description generation model, based on the model input data, a first probability distribution of the image description is obtained. Specifically, the model input data is input into the currently fine-tuned image description generation model, and the model outputs a probability distribution of different image descriptions. For example, for a model input data, the model may output that the probability of "a cat on the sofa" is 0.8, the probability of "a dog on the sofa" is 0.2, and so on. At the same time, through the basic image description generation model, based on the model input data, a second probability distribution of the image description is obtained. Similarly, the model input data is input into the basic image description generation model to obtain the model's probability distribution of different image descriptions. Then, the loss value is determined based on the first probability distribution, the second probability distribution, and the reference image description by incrementally optimizing the loss function. The incremental optimization loss function is expressed by the following formula:
[0081] L(π θ ,π ref )=E x,y∈D [w(y)(1-v(x,y;β))] (1)
[0082]
[0083] z ref =E x∈D [βKL(π θ (y|x)||πref (y|x))] (4)
[0084] Wherein, L() represents the incremental optimization loss function, which is used to measure the difference between the current fine-tuned image description generation model and the basic image description generation model. θ represents the currently fine-tuned image description generation model, π ref represents the basic image description generation model. E() represents an expectation function, which is used to average the data input to the expectation function. x represents the model input data, y represents the reference image description, and D represents the image sample set. w(y) represents the weight function of the reference image description, λ + Represents the preset weight of the positive image description, represents the preset weight of the negative image description. v(x, y; β) represents the confidence of the reference image description output by the currently fine-tuned image description generation model. σ represents the Sigmoid function, which is used to map the input value to the [0,1] interval. β represents the temperature hyperparameter, which is used to adjust the smoothness of the output of the currently fine-tuned image description generation model. When β is larger, the probability distribution of the model output will be sharper; when β is smaller, the probability distribution of the model output will be smoother. π θ (y|x) represents the first probability distribution, π ref (y|x) represents the second probability distribution, z ref represents the expected log-likelihood ratio of the basic image description generation model. KL(π θ (y|x)||π ref (y|x)) represents the Kullback-Leibler discreteness between the first probability distribution and the second probability distribution, which is used to measure the difference between the two probability distributions.
[0085] It should be noted that when setting λ + and λ - When , you can set these two weights according to the specific task requirements and data characteristics. For example, if you want the model to pay more attention to the positive image description, you can set λ + If you want the model to have a better ability to distinguish negative image descriptions, you can increase λ appropriately. - .
[0086] Since these two negative image descriptions are used together with the positive image description as reference image descriptions during the training process of the image description generation model to determine the training data pairs and calculate the loss value, by introducing the wrong perspective negative image description, the image description generation model can learn the impact of different perspectives on image descriptions and avoid generating descriptions of the wrong perspective; by introducing the context interference negative image description, the model can learn to correctly process the context information and avoid being disturbed by the context and deviating from the core information of the image.
[0087] Through the above embodiment, several pieces of model input data can be obtained, and the above operation is performed for each piece of model input data. When the preset convergence condition is met, the image description generation model training is completed.
[0088] Among them, the preset convergence condition can be that the sum of the loss values determined based on the current iteration is less than the pre-configured loss threshold, or the sum of the loss values determined by the current iteration has been in a downward trend and tends to be flat, or the number of iterations for training the basic image description generation model reaches the set maximum number of iterations, etc. In the specific implementation, it can be flexibly set and is not specifically limited here.
[0089] As a possible implementation, when training the basic image description generation model, the model input data can be divided into training samples and test samples, the basic image description generation model is first trained based on the training samples, and then the reliability of the trained image description generation model is verified based on the test samples. After obtaining the loss values corresponding to the input data of each model in the current iteration based on the above embodiment, the currently fine-tuned image description generation model can be fine-tuned according to each loss value until the preset convergence condition is met, thereby obtaining a trained image description generation model.
[0090] The beneficial effects of this application are as follows:
[0091] 1. By introducing positive and negative image descriptions as reference image descriptions, the image description generation model not only learns how to generate positive descriptions that match the image content during training, but also learns to distinguish negative descriptions that do not match the image. This makes the image descriptions generated by the model more accurate, detailed, and in line with the actual image content, reducing the probability of generating incorrect or irrelevant descriptions.
[0092] 2. The iterative fine-tuning process combined with the incremental optimization loss function enables the model to continuously optimize on top of the capabilities of the base model. By learning and adjusting different training data pairs in multiple iterations, the model can adapt to various types of images and diverse contextual information, thereby improving its generalization capabilities in different data sets and scenarios. This means that the model can generate reasonable and accurate descriptions when faced with unseen images.
[0093] 3. The trained image description generation model can introduce the contextual information of the image when generating image descriptions, thereby avoiding description deviation caused by a single image input.
[0094] Embodiment 2:
[0095] The present application provides an image description generation method based on the image description generation model trained by the above embodiment, Figure 2 A schematic diagram of a process of generating an image description provided in an embodiment of the present application, the process comprising:
[0096] S201: Acquire model input data to be processed; wherein the model input data to be processed includes an image to be processed, a context of the image to be processed, and a processing instruction of the image to be processed.
[0097] S202: Generate an image description model through a pre-trained model, and obtain image description information of the image to be processed based on the model input data to be processed.
[0098] The image description generation method provided in the present application is applied to a computer device, which may be an intelligent device or a server. The computer device for image description generation in the present application may be the same as or different from the computer device for image description generation model training.
[0099] In a possible implementation, the image description generation model is generally trained in an offline manner. After the trained image description generation model is acquired, the image description generation model can be deployed to the computer device for image description generation.
[0100] In this application, the image to be processed can come from a variety of different sources. For example, it can be obtained from a public image dataset (such as COCO, ImageNet, etc.) or a document dataset (such as a product manual, paper, report, financial statement, etc.), or it can be obtained from images or documents in a specific field collected by the user. The format of the image to be processed can be a common image format, such as JPEG, PNG, etc.
[0101] In a possible implementation, after obtaining the image to be processed, the image to be processed may be preprocessed. The preprocessing includes, but is not limited to, enhancement, noise reduction, and size normalization. For example, the image size may be adjusted, the image may be cropped to remove unnecessary edge portions, and the image may be converted to a specific color mode (such as RGB), etc., to ensure that the image meets the input requirements of the pre-trained model.
[0102] The context information of the image to be processed can provide the model with additional background knowledge about the image and help the model better understand the image content. The context information can include metadata such as the shooting time, shooting location, shooting equipment, etc. of the image to be processed, or text descriptions and tags related to the image to be processed.
[0103] The processing instructions of the image to be processed are used to clarify the specific requirements for the model to generate the image description, wherein the processing instructions can be input by the user according to the requirements.
[0104] Before generating image descriptions, you need to load the image description generation model obtained by the previous training method. The model has learned the mapping relationship between image features and language expressions.
[0105] It should be noted that the specific training process of the image description generation model has been described in the above embodiment, and the repeated parts will not be repeated here.
[0106] Input the model input data (the image to be processed, context information, and processing instructions) to the pre-trained image description generation model for inference. The image description generation model will generate image description information for the image to be processed based on the input model input data and the knowledge it has learned. After obtaining the image description information to be processed, the generated image description information can be output in a suitable format. For example, it can be displayed on the user interface for the user to view; it can also be saved to a file for subsequent analysis and processing.
[0107] Embodiment 3:
[0108] The present application also provides a training device for an image description generation model. Figure 3 A schematic diagram of the structure of a training device for an image description generation model provided in an embodiment of the present application, the device comprising:
[0109] The acquisition module 31 is used to acquire an image sample set; wherein the image sample set includes model input data and reference image descriptions corresponding to the model input data, any model input data includes an image sample, context information of the image sample and a processing instruction, and the reference image description includes a positive image description and a negative image description;
[0110] A training module 32, configured to iteratively fine-tune a pre-trained basic image description generation model based on the image sample set; wherein the basic image description generation model is a model that already has the ability to generate image descriptions;
[0111] Among them, in any iterative fine-tuning process:
[0112] For any model input data, a training data pair is determined based on the model input data and any reference image description corresponding to the image sample; a first probability distribution of the image description is obtained based on the model input data by using the currently fine-tuned image description generation model; and a second probability distribution of the image description is obtained based on the model input data by using the basic image description generation model; a loss value is determined based on the first probability distribution, the second probability distribution and the reference image description by using an incremental optimization loss function;
[0113] Fine-tune the currently fine-tuned image description generation model according to each loss value to obtain a trained image description generation model;
[0114] The incremental optimization loss function is expressed by the following formula:
[0115] L(π θ ,π ref )=E x,y∈D [w(y)(1-v(x,y;β))] (1)
[0116]
[0117] z ref =E x∈D [βKL(π θ (y|x)||π ref (y|x))] (4)
[0118] Wherein, L() represents the incremental optimization loss function, π θ represents the currently fine-tuned image description generation model, π ref represents the basic image description generation model, E() represents the expected function, x represents the model input data, y represents the reference image description, D represents the image sample set, w(y) represents the weight function of the reference image description, λ + represents the preset weight of the positive image description, λ - represents the preset weight of the negative image description, v(x, y; β) represents the confidence of the reference image description output by the currently fine-tuned image description generation model, σ represents the Sigmoid function, β represents the temperature hyperparameter, which is used to adjust the smoothness of the output of the currently fine-tuned image description generation model, and π θ (y|x) represents the first probability distribution, π ref (y|x) represents the second probability distribution, z ref represents the expected log-likelihood ratio of the basic image description generation model, KL(π θ (y|x)||π ref(y|x)) represents the Kullback-Leibler dispersion between the first probability distribution and the second probability distribution.
[0119] The training device for the image description generation model in this embodiment is presented in the form of functional modules, where the modules refer to application specific integrated circuits (ASICs), processors and memories that execute one or more software or fixed programs, and / or other devices that can provide the above functions.
[0120] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0121] Embodiment 4:
[0122] The present application also provides an image description generation device based on the image description generation model trained by the above embodiment 1, Figure 4 A schematic diagram of a device structure for generating an image description provided in an embodiment of the present application, the device comprising:
[0123] An acquisition unit 41 is used to acquire model input data to be processed; wherein the model input data to be processed includes an image to be processed, a context of the image to be processed, and a processing instruction of the image to be processed;
[0124] The processing unit 42 is used to generate a model through a pre-trained image description and obtain image description information of the image to be processed based on the model input data to be processed.
[0125] The image description generating device in this embodiment is presented in the form of a functional module, where the module refers to an application specific integrated circuit (ASIC), a processor and memory that executes one or more software or fixed programs, and / or other devices that can provide the above functions.
[0126] The further functional description of each of the above modules and units is the same as that of the above corresponding embodiments and will not be repeated here.
[0127] Embodiment 5:
[0128] See also Figure 5 , Figure 5 is a schematic diagram of the structure of a computer device provided by an optional embodiment of the present application, such as Figure 5As shown, the computer device includes: one or more processors 10, a memory 20, and interfaces for connecting various components, including high-speed interfaces and low-speed interfaces. Various components are connected to each other using different buses for communication, and can be installed on a common mainboard or installed in other ways as needed. The processor can process the instructions executed in the computer device, including instructions stored in or on the memory to display the graphical information of the GUI on an external input / output device (such as, a display device coupled to the interface). In some optional embodiments, if necessary, multiple processors and / or multiple buses can be used together with multiple memories and multiple memories. Similarly, multiple computer devices can be connected, and each device provides some necessary operations (for example, as a server array, a group of blade servers, or a multi-processor system). Figure 5 A processor 10 is taken as an example.
[0129] The processor 10 may be a central processing unit, a network processor or a combination thereof. The processor 10 may further include a hardware chip. The hardware chip may be a dedicated integrated circuit, a programmable logic device or a combination thereof. The programmable logic device may be a complex programmable logic device, a field programmable gate array, a general purpose array logic or any combination thereof.
[0130] The memory 20 stores instructions executable by at least one processor 10, so that at least one processor 10 executes the method shown in the above embodiment.
[0131] The memory 20 may include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function; the data storage area may store data created by the use of a computer device based on the presentation of a small program landing page, etc. In addition, the memory 20 may include a high-speed random access memory, and may also include a non-transient memory, such as at least one disk storage device, a flash memory device, or other non-transient solid-state storage device. In some optional embodiments, the memory 20 may optionally include a memory remotely arranged relative to the processor 10, and these remote memories may be connected to the computer device via a network. Examples of the above-mentioned network include, but are not limited to, the Internet, an intranet, a local area network, a mobile communication network, and combinations thereof.
[0132] The memory 20 may include a volatile memory, such as a random access memory; the memory may also include a non-volatile memory, such as a flash memory, a hard disk or a solid state drive; the memory 20 may also include a combination of the above types of memory.
[0133] The computer device also includes an input device 30 and an output device 40. The processor 10, the memory 20, the input device 30 and the output device 40 may be connected via a bus or other means. Figure 5 The example of connecting through bus is taken in the following.
[0134] The input device 30 can receive input digital or character information, and generate key signal input related to the user settings and function control of the computer device, such as a touch screen, a keypad, a mouse, a track pad, a touch pad, an indicator bar, one or more mouse buttons, a trackball, a joystick, etc. The output device 40 may include a display device, an auxiliary lighting device (e.g., an LED) and a tactile feedback device (e.g., a vibration motor), etc. The above-mentioned display device includes but is not limited to a liquid crystal display, a light emitting diode, a display and a plasma display. In some optional embodiments, the display device can be a touch screen.
[0135] Embodiment 6:
[0136] On the basis of the above embodiments, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program executable by a processor is stored. When the program runs on the processor, the processor implements the following steps when executing:
[0137] Acquire an image sample set; wherein the image sample set includes model input data and reference image descriptions corresponding to the model input data, any model input data includes an image sample, context information of the image sample and a processing instruction, and the reference image description includes a positive image description and a negative image description;
[0138] Based on the image sample set, iteratively fine-tuning the pre-trained basic image description generation model; wherein the basic image description generation model is a model that already has the ability to generate image descriptions;
[0139] Among them, in any iterative fine-tuning process:
[0140] For any model input data, a training data pair is determined based on the model input data and any reference image description corresponding to the image sample; a first probability distribution of the image description is obtained based on the model input data by using the currently fine-tuned image description generation model; and a second probability distribution of the image description is obtained based on the model input data by using the basic image description generation model; a loss value is determined based on the first probability distribution, the second probability distribution and the reference image description by using an incremental optimization loss function;
[0141] Fine-tune the currently fine-tuned image description generation model according to each loss value to obtain a trained image description generation model;
[0142] The incremental optimization loss function is expressed by the following formula:
[0143] L(π θ ,π ref )=E x,y∈D [w(y)(1-v(x,y;β))] (1)
[0144]
[0145] z ref =E x∈D [βKL(π θ (y|x)||π ref (y|x))] (4)
[0146] Wherein, L() represents the incremental optimization loss function, π θ represents the currently fine-tuned image description generation model, π ref represents the basic image description generation model, E() represents the expected function, x represents the model input data, y represents the reference image description, D represents the image sample set, w(y) represents the weight function of the reference image description, λ + represents the preset weight of the positive image description, λ - represents the preset weight of the negative image description, v(x, y; β) represents the confidence of the reference image description output by the currently fine-tuned image description generation model, σ represents the Sigmoid function, β represents the temperature hyperparameter, which is used to adjust the smoothness of the output of the currently fine-tuned image description generation model, and π θ (y|x) represents the first probability distribution, π ref (y|x) represents the second probability distribution, z ref represents the expected log-likelihood ratio of the basic image description generation model, KL(π θ (y|x)||π ref (y|x)) represents the Kullback-Leibler dispersion between the first probability distribution and the second probability distribution.
[0147] Since the principle of solving the problem by the above-mentioned computer-readable storage medium is similar to the training method of the image description generation model, the implementation of the above-mentioned computer-readable storage medium can refer to Example 1 of the method, and the repeated parts will not be repeated.
[0148] Embodiment 7:
[0149] On the basis of the above embodiments, an embodiment of the present application further provides a computer-readable storage medium, in which a computer program executable by a processor is stored. When the program runs on the processor, the processor implements the following steps when executing:
[0150] Acquire input data of the model to be processed; wherein the input data of the model to be processed includes an image to be processed, a context of the image to be processed, and a processing instruction of the image to be processed;
[0151] The image description information of the image to be processed is obtained by generating a model through a pre-trained image description based on the model input data to be processed.
[0152] Since the principle of solving the problem by the above-mentioned computer-readable storage medium is similar to that of the image description generating method, the implementation of the above-mentioned computer-readable storage medium can refer to Example 2 of the method, and the repeated parts will not be repeated.
[0153] Obviously, those skilled in the art can make various changes and modifications to the present application without departing from the spirit and scope of the present application. Thus, if these modifications and variations of the present application fall within the scope of the claims of the present application and their equivalents, the present application is also intended to include these modifications and variations.
Claims
1. A training method for an image description generation model, characterized in that: The method comprises: Acquire an image sample set; wherein the image sample set includes model input data and reference image descriptions corresponding to the model input data, any model input data includes an image sample, context information of the image sample and a processing instruction, and the reference image description includes a positive image description and a negative image description; Based on the image sample set, iteratively fine-tuning the pre-trained basic image description generation model; wherein the basic image description generation model is a model that already has the ability to generate image descriptions; Among them, in any iterative fine-tuning process: For any model input data, a training data pair is determined based on the model input data and any reference image description corresponding to the image sample; a first probability distribution of the image description is obtained based on the model input data by using the currently fine-tuned image description generation model; and a second probability distribution of the image description is obtained based on the model input data by using the basic image description generation model; a loss value is determined based on the first probability distribution, the second probability distribution and the reference image description by using an incremental optimization loss function; Fine-tune the currently fine-tuned image description generation model according to each loss value to obtain a trained image description generation model; The incremental optimization loss function is expressed by the following formula: L(π θ ,p ref )=E x,y∈D [w(y)(1-v(x,y;β))] (1) With ref =E x∈D [βKL(π θ (y|x)||π ref (y|x))] (4) Wherein, L() represents the incremental optimization loss function, π θ represents the currently fine-tuned image description generation model, π ref represents the basic image description generation model, E() represents the expected function, x represents the model input data, y represents the reference image description, D represents the image sample set, w(y) represents the weight function of the reference image description, λ + represents the preset weight of the positive image description, λ - represents the preset weight of the negative image description, v(x, y; β) represents the confidence of the reference image description output by the currently fine-tuned image description generation model, σ represents the Sigmoid function, β represents the temperature hyperparameter, which is used to adjust the smoothness of the output of the currently fine-tuned image description generation model, and π θ (y|x) represents the first probability distribution, π ref (y|x) represents the second probability distribution, z ref represents the expected log-likelihood ratio of the basic image description generation model, KL(π θ (y|x)||π ref (y|x)) represents the Kullback-Leibler dispersion between the first probability distribution and the second probability distribution.
2. The method according to claim 1, characterized in that The context information corresponding to any image sample is obtained as follows: If the original data to which the image sample belongs is of the image-text interlaced type, determining the context content of the position where the image sample is located in the original data as the context information of the image sample; If the original data to which the image sample belongs is of a pure picture type, the application background information of the original data is determined as the context information of the image sample.
3. The method according to claim 2, characterized in that The determining the context content of the location of the image sample in the original data as the context information of the image sample includes: The context content at the location of the image sample in the original data and within a pre-configured text length range is determined as the context information of the image sample.
4. The method according to claim 2, characterized in that The image samples are obtained by preprocessing the images in the original data; wherein the preprocessing includes but is not limited to: enhancement, noise reduction, and size normalization.
5. The method according to claim 1, characterized in that The negative image description includes one or more of the following: wrong perspective negative image description and context interference negative image description; wherein the wrong perspective negative image description is a correct description of the image sample under a wrong perspective, and the context interference negative image description is an image description that deviates from the core information of the image sample or is not specific enough due to interference from the context information of the image sample.
6. A method for generating image description based on an image description generation model trained by the method according to any one of claims 1 to 5, characterized in that: The method comprises: Acquire input data of the model to be processed; wherein the input data of the model to be processed includes an image to be processed, a context of the image to be processed, and a processing instruction of the image to be processed; The image description information of the image to be processed is obtained by generating a model with pre-trained image description based on the model input data to be processed.
7. A training device for an image description generation model, characterized in that: The device comprises: An acquisition module, used to acquire an image sample set; wherein the image sample set includes model input data and reference image descriptions corresponding to the model input data, any model input data includes an image sample, context information of the image sample and a processing instruction, and the reference image description includes a positive image description and a negative image description; A training module, used for iteratively fine-tuning a pre-trained basic image description generation model based on the image sample set; wherein the basic image description generation model is a model that already has the ability to generate image descriptions; Among them, in any iterative fine-tuning process: For any model input data, a training data pair is determined based on the model input data and any reference image description corresponding to the image sample; a first probability distribution of the image description is obtained based on the model input data by using the currently fine-tuned image description generation model; and a second probability distribution of the image description is obtained based on the model input data by using the basic image description generation model; a loss value is determined based on the first probability distribution, the second probability distribution and the reference image description by using an incremental optimization loss function; Fine-tune the currently fine-tuned image description generation model according to each loss value to obtain a trained image description generation model; The incremental optimization loss function is expressed by the following formula: L(π θ ,p ref )=E x,y∈D [w(y)(1-v(x,y;β))] (1) With ref =E x∈D [βKL(π θ (y|x)||π ref (y|x))] (4) Wherein, L() represents the incremental optimization loss function, π θ represents the currently fine-tuned image description generation model, π ref represents the basic image description generation model, E() represents the expected function, x represents the model input data, y represents the reference image description, D represents the image sample set, w(y) represents the weight function of the reference image description, λ + represents the preset weight of the positive image description, λ - represents the preset weight of the negative image description, v(x, y; β) represents the confidence of the reference image description output by the currently fine-tuned image description generation model, σ represents the Sigmoid function, β represents the temperature hyperparameter, which is used to adjust the smoothness of the output of the currently fine-tuned image description generation model, and π θ (y|x) represents the first probability distribution, π ref (y|x) represents the second probability distribution, z ref represents the expected log-likelihood ratio of the basic image description generation model, KL(π θ (y|x)||π ref (y|x)) represents the Kullback-Leibler dispersion between the first probability distribution and the second probability distribution.
8. An image description generating device based on an image generation model trained by any one of the methods of claims 1-5, characterized in that: The device comprises: An acquisition unit, configured to acquire model input data to be processed; wherein the model input data to be processed includes an image to be processed, a context of the image to be processed, and a processing instruction of the image to be processed; The processing unit is used to generate a model through pre-trained image description and obtain image description information of the image to be processed based on the input data of the model to be processed.
9. A computer device, characterized in that: The computer device includes a processor, and the processor is used to implement the steps of the training method of the image description generation model as described in any one of claims 1 to 5 above when executing the computer program stored in the memory, or to implement the steps of the image description generation method as described in claim 6 above.
10. A computer-readable storage medium, characterized in that: It stores a computer program executable by a computer device. When the program is run on the computer device, the computer device executes the steps of the training method of the image description generation model as described in any one of claims 1 to 5 above, or implements the steps of the image description generation method as described in claim 6 above.