Image description processing method, computer device and storage medium

By extracting targets from images and matching descriptive information to generate fine-grained descriptions, this technology solves the problem of low-quality image description datasets in existing technologies, achieving efficient generation of high-quality image description datasets suitable for live background image processing.

CN116486188BActive Publication Date: 2026-02-03GUANGZHOU FANGGUI INFORMATION TECHNOLOGY CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202310420296.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2023-04-18
Publication Date
2026-02-03
Estimated Expiration
2043-04-18

AI Technical Summary

Technical Problem

Existing technologies generate image description datasets of low quality and weak correlation with images, relying mainly on manual annotation or online label cleaning, resulting in high cost and low efficiency.

Method used

By performing target extraction processing on the image to be described, a target region image is generated and described. By combining the matching verification of the image and the description information, fine-grained description information is generated and stored in the live background image library, and a matching image description is provided in response to user requests.

Benefits of technology

It generates high-quality image fine-grained description datasets, improving the accuracy and efficiency of description information, reducing manual annotation costs, and is suitable for live background image processing.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116486188B_ABST
    Figure CN116486188B_ABST
Patent Text Reader

Abstract

The application discloses an image description processing method, a computer device and a storage medium. The method comprises the following steps: obtaining image description information by describing a to-be-described image; obtaining at least one target region image by performing target extraction processing on the to-be-described image, and obtaining target description information by describing the at least one target region image; judging whether the to-be-described image and the image description information match, and whether each target region image and the corresponding target description information match; and if the to-be-described image and the image description information match, and the at least one target region image and the corresponding target description information match, generating fine-grained description information including the image description information and the corresponding target description information for the to-be-described image. In the foregoing manner, the application can generate fine-grained description information of high-quality images.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of image processing, in particular to an image description processing method, a computer device and a storage medium. BACKGROUND

[0002] As the visual basis for human perception of the world, images are an important means for humans to obtain, express and transmit information. In image description, image retrieval, image generation and other tasks, a description dataset is usually used, which is a dataset composed of picture-picture corresponding text descriptions. The magnitude and quality of the dataset generation directly and profoundly affect the performance of subsequent image processing algorithms, so it is crucial to generate an efficient data generation scheme. However, the establishment of the current description dataset is mainly through manual annotation or cleaning generated according to network labels, which relies on human-set image-related description information, resulting in a description dataset that may not be highly relevant to the image and may have low quality. SUMMARY

[0003] The technical problem solved by the present application is to provide an image description processing method, a computer device and a storage medium, which can generate high-quality image fine-grained description information.

[0004] To solve the above technical problems, the first technical solution adopted by the present application is to provide an image description processing method, which comprises: describing a to-be-described image to obtain image description information; performing target extraction processing on the to-be-described image to obtain at least one target region image, and describing the at least one target region image to obtain target description information; determining whether the to-be-described image and the image description information match, and whether each target region image and the corresponding target description information match; if the to-be-described image and the image description information match, and the at least one target region image and the corresponding target description information match, generating fine-grained description information including the image description information and the corresponding target description information for the to-be-described image.

[0005] To address the aforementioned technical problems, the second technical solution adopted in this application is: providing a method for describing a live background image. This method includes: performing target extraction processing on the image to be described to obtain at least one target region image, and describing the at least one target region image to obtain target description information; determining whether the image to be described and the image description information match, and whether each target region image and its corresponding target description information match; if the image to be described and the image description information match, and at least one target region image and its corresponding target description information match, then generating fine-grained description information for the image to be described, including the image description information and the corresponding target description information; associating and storing the fine-grained information and the image to be described in a live background image library; responding to a background image request instruction from a live user, determining the fine-grained description information matching the background image request instruction in the live background image library, and obtaining the image to be described corresponding to the fine-grained description information; and sending the obtained image to be described to the client terminal corresponding to the live user, enabling the client terminal to set the image to be described as the live background on the live room interface.

[0006] To solve the above-mentioned technical problems, the third technical solution adopted in this application is: to provide a computer device, which includes: a processor, a memory, and a communication circuit; the communication circuit and the memory are respectively coupled to the processor, the memory is used to store computer programs, and the processor is used to read and execute computer programs to implement the methods provided by the first and second technical solutions of this application as described above.

[0007] To solve the above-mentioned technical problems, the fourth technical solution adopted in this application is to provide a computer-readable storage medium that stores a computer program that can be read and executed by a processor to implement the methods provided by the first and second technical solutions of this application.

[0008] The beneficial effects of this application are as follows: Unlike existing technologies, this method obtains image description information by describing the image to be described, extracts targets from the image to be described to obtain at least one target region image, and describes the at least one target region image to obtain target description information. After obtaining the image description information and target description information, it determines whether the image to be described and the image description information match, and whether each target region image matches the corresponding target description information. This allows for the verification of the obtained description information, ensuring that the final generated image description information matches the image and the target description information matches the target region image. If the image to be described and the image description information match, and at least one target region image matches the corresponding target description information, then fine-grained description information including image description information and corresponding target description information is generated for the image to be described. This enables the generation of efficient and detailed fine-grained description information for the image to be described, thereby generating a high-quality image fine-grained description dataset. Attached Figure Description

[0009] Figure 1 This is a first flowchart illustrating an embodiment of the image description processing method of this application;

[0010] Figure 2 This is a second flowchart illustrating an embodiment of the image description processing method of this application;

[0011] Figure 3 This is a schematic diagram of the image to be processed and the target bounding box in an embodiment of the image description processing method of this application;

[0012] Figure 4 This is a schematic diagram of the interaction between the server and the client terminal in an embodiment of the image description processing method of this application.

[0013] Figure 5 This is a flowchart illustrating an embodiment of the live streaming background image description processing method of this application;

[0014] Figure 6 This is a schematic diagram of the circuit structure of an embodiment of the computer device of this application;

[0015] Figure 7 This is a schematic diagram of the circuit structure of an embodiment of the computer-readable storage medium of this application. Detailed Implementation

[0016] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.

[0017] With the rapid development of technology, people have increasingly higher demands for computer image processing. For example, after inputting the name of a specific object, the computer needs to select images containing that object from a set of images. The computer can obtain images containing the specific object by matching the user-input name with the corresponding descriptive terms of these images. The quality of the image description can determine the quality of some downstream tasks; therefore, processing images to generate high-quality descriptions is crucial.

[0018] The inventors of this application, through long-term research, have discovered that current methods for obtaining image descriptions primarily rely on manual annotation or cleaning based on web tags. However, manual annotation is extremely time-consuming and costly, making it difficult to generate large-scale datasets. As for cleaning based on web tags, since the vast majority of existing images on the internet do not contain descriptions, and the web tags may partially include promotional tags, the generated description sets may suffer from textual inconsistencies and numerous grammatical errors. To improve or solve these technical problems, this application proposes at least the following embodiments.

[0019] like Figure 1 and Figure 2 As shown, the first embodiment of the image description processing method of this application can be executed by a server, and the image description processing method described therein can include: S100: describing the image to be described to obtain image description information. S200: performing target extraction processing on the image to be described to obtain at least one target region image, and describing the at least one target region image to obtain target description information. S300: determining whether the image to be described and the image description information match, and whether each target region image and the corresponding target description information match. S400: if the image to be described and the image description information match, and at least one target region image and the corresponding target description information match, then generating fine-grained description information including image description information and corresponding target description information for the image to be described.

[0020] Image description information is obtained by describing the image to be described. Target extraction is then performed on the image to be described to obtain at least one target region image, and this target region image is described to obtain target description information. After obtaining the image description information and target description information, it is determined whether the image to be described and the image description information match, and whether each target region image matches its corresponding target description information. The obtained description information can be verified to ensure that the final generated image description information matches the image, and the target description information matches the target region image. If the image to be described and the image description information match, and at least one target region image matches its corresponding target description information, then fine-grained description information, including both image description information and corresponding target description information, is generated for the image to be described. This allows for the generation of an efficient and detailed image description dataset for the image to be described.

[0021] The following is a detailed description of the embodiments of the image description processing method of this application.

[0022] S100: Describe the image to be described to obtain image description information.

[0023] For example, if the image to be described is a living room with a sofa, dining table, and stools, then the corresponding image description information could be "a living room with a sofa, dining table, and stools".

[0024] In this step, the BLIP2 model, based on the VIT framework, can be used to generate image description information for the image to be described. During the generation of image description information, the image can first be decoded to obtain the corresponding image data. Image decoding can employ nucleus sampling or top-p sampling strategies to increase syntactic or vocabulary diversity while ensuring semantic similarity in the generated image description information. In other words, multiple image description terms can be obtained when describing the image, and the server can randomly select one of these terms as the image description information for subsequent processing.

[0025] Optionally, the steps preceding S100 can be referred to for instructions on how to obtain the image to be described:

[0026] S110: Obtain the original images from the network and perform classification processing to obtain at least one image set.

[0027] The server can acquire raw images from the internet and classify them to obtain at least one image set. For example, raw images collected from the internet can be divided into indoor scenes, outdoor scenes, and landscape photography, thus obtaining image sets for indoor scenes, outdoor scenes, and landscape photography, etc.

[0028] S120: Perform watermark removal processing on each original image in each image set to obtain the image to be described.

[0029] Since images on the internet may contain watermarks, the server can remove the watermarks from each original image in each image set, obtaining each image to be described corresponding to each original image. Optionally, the server can first roughly estimate the watermark location, identify the watermark content and segment the watermark content, and then use a repair network to repair the original image to obtain the watermark-free image to be described. This can improve the quality of the image to be described and reduce the adverse effects of watermarks on the quality of the generated image description information.

[0030] The server can therefore obtain a large number of images to be described, thereby generating a large number of fine-grained image description datasets, which in turn improves the quality of the final generated fine-grained image description datasets.

[0031] By acquiring raw images from the internet and classifying them, we can obtain a large number of image samples for training to generate fine-grained descriptions. On the other hand, in the subsequent step of generating image description information for the image to be described, we can generate category-related description information based on the category.

[0032] To make the description of the image to be described more detailed and specific, the target objects included in the image to be described can be extracted and the corresponding descriptive information of the target objects can be obtained. For details, please refer to the following steps after S100:

[0033] S200: Perform target extraction processing on the image to be described to obtain at least one target region image, and describe the at least one target region image to obtain target description information.

[0034] For example, if the image to be described is a living room with a sofa, dining table, and stools, then the server, after performing target extraction processing on the image to be described, can obtain at least one target region image, which can be an image of the sofa, a dining table, and stools. The server can also describe these target region images separately to obtain descriptive information for these target region images, i.e., target description information. For example, "a white sofa with cushions" or "a round dining table with small yellow flowers."

[0035] In some embodiments, this step can be executed by the server simultaneously with step S100.

[0036] After performing target extraction processing on the image to be described to obtain at least one target region image, the obtained target region image can be described to obtain target description information. See the following steps included in S200:

[0037] S201: Describe at least one target region image to obtain target description information.

[0038] In one implementation, the server can use a Generative Image-to-text Transformer for Vision and Language algorithm to obtain target description information for each target region image. The target description information can describe the items and their attributes in the target region image, such as a white plush chair or a circle of heart-shaped colored balloons.

[0039] Optionally, how to perform target extraction processing on the image to be described to obtain at least one target region image can be found in the following steps included in S200:

[0040] S210: Perform target detection on the image to be described to identify at least one target object.

[0041] In this step, the image to be described can be input into an object detection network to detect target objects in the image. For example, target objects could be sofas, dining tables, alarm clocks, etc. The object detection network could be a Detic network, which can detect 20,000 types of targets, essentially covering most of the main targets in everyday photographs, thus detecting virtually all target objects included in the image to be described.

[0042] Optionally, when determining at least one target object, the server can determine the bounding box and confidence level of that target object, as detailed in the following steps included in S210:

[0043] S211: Determine at least one target bounding box for selecting a target object in the image to be described, and calculate the confidence level of the target object.

[0044] like Figure 3 As shown, during object detection in the image to be described, the outline of the target object can be drawn to determine at least one target bounding box, and the confidence score of the bounding box, i.e., the confidence score of the target object, can be calculated. Simultaneously, the server can also obtain the size of the bounding box or its position information in the image to be described. However, during object detection, there may be cases where the selected bounding box is incorrect, i.e., the bounding box does not contain a target object. Figure 3 The second target box 12 shown may not completely enclose the outline of the target object, such as... Figure 3The third target box 13 is shown. Therefore, when determining the target object, the target box and confidence score can be obtained simultaneously so that the determined target object can be filtered later.

[0045] In addition, the server can also determine the category of the target object when detecting it. For example, it can determine that the category of the first target box 11 is furniture.

[0046] By performing object detection on the image to be described to identify at least one target object, the server can describe the objects included in the image, thereby generating a more detailed fine-grained description of the image and expanding the scope of application of the fine-grained description.

[0047] S220: Determine whether each target object meets the preset clipping conditions.

[0048] Before describing the target object in the image to be described to generate target description information, the server can crop out the image corresponding to the target object in the image to be described to obtain a target region image that is close to the outline of the target object, and then describe the target region image.

[0049] However, for some target objects for which it is not expected to obtain corresponding target description information, such as the background part of the image or a target object that occupies a very small area in the image, the target area image corresponding to the target object may not be cropped out. Therefore, the obtained target objects can be judged, that is, whether each target object meets the preset cropping conditions, so as to effectively avoid the situation of cropping out unnecessary target area images.

[0050] Optionally, the target bounding box and confidence score can be used to filter the target objects corresponding to the image to be described, as detailed in the following steps included in S220:

[0051] S221: Determine whether the target object meets the preset clipping conditions using at least one of the target bounding box and the confidence score.

[0052] For background parts or target objects that occupy a very small area in an image, the target area image corresponding to the target object may not be cropped. Therefore, the target bounding box can be used to determine whether the target object needs to be cropped, and the confidence score can be used to determine whether the target object is a valid object, such as not being a background part similar to the third target bounding box 13.

[0053] Optionally, the server determines whether the target object meets the preset clipping criteria using at least one of the target bounding box and the confidence score, as described in the following steps included in S221:

[0054] S2211: Determine whether the size of the target bounding box is greater than or equal to the preset size, and / or whether the confidence level is greater than or equal to the preset confidence level.

[0055] When determining the bounding box from the image to be described, the bounding box may select some smaller target objects. For these smaller target objects, the corresponding target area image can be left uncropped from the image to be described, thus reducing the resource consumption caused by cropping target objects smaller than a preset size. Therefore, before cropping the target area image from the image to be described, the server can first determine whether the size of the bounding box is greater than or equal to the preset size. The preset size can be a preset area, such as 0.1 times, 0.15 times, or 0.2 times the area of ​​the image to be described.

[0056] Furthermore, since the bounding boxes determined from the image to be described may not completely encompass the outline of the target object—that is, most of the area within the bounding box may be or entirely composed of background imagery, such as a white wall in a living room—cropping the corresponding target region image based on such bounding boxes is not very meaningful for generating a valid descriptive dataset. Therefore, bounding boxes that do not meet expectations can be filtered by judging whether their confidence level is greater than or equal to a preset confidence level. The preset confidence level can be 0.6, 0.7, 0.8, etc.

[0057] S2212: If the size of the target box is smaller than the preset size, and / or the confidence level is less than the preset confidence level, then the target object is determined to not meet the preset clipping conditions.

[0058] S2213: If the size of the target box is greater than or equal to the preset size, and / or the confidence level is greater than or equal to the preset confidence level, then the target object is determined to meet the preset clipping conditions.

[0059] By cropping the target region image from the image to be described that meets the preset cropping conditions, the quality of the target region image can be improved. This can reduce the occurrence of low quality target description information due to low quality of the target region image, and thus help generate efficient and high-quality description datasets.

[0060] If the target object meets the preset cropping conditions, the server can crop the target region image corresponding to the target object from the image to be described. See the steps after S220:

[0061] S230: If the preset cropping conditions are met, at least one target region image is cropped from the image to be described.

[0062] For a target object that meets the preset cropping conditions—that is, the target bounding box size is greater than or equal to the preset size, and / or the confidence level is greater than or equal to the preset confidence level—the server can crop the corresponding target region image from the image to be described. Since there can be multiple target objects in the image to be described, the server can crop at least one target region image from the image to be described.

[0063] In some embodiments, the target bounding box may include only one target object, so that the cropped target region image includes only one target object. In other embodiments, the target may select more than one target object, that is, the target bounding box includes at least one target object, so that the cropped target region image may include at least one target object.

[0064] Optionally, the server may use the target bounding box in step S211 to obtain at least one target region image, as can be seen in the following steps included in S230:

[0065] S231: Cropping from the image to be described according to the target bounding box to crop at least one target region image.

[0066] Since there can be multiple target objects in the image to be described, i.e., multiple target bounding boxes, the server can crop the image according to the target bounding boxes to obtain at least one target region image.

[0067] For example, if the image to be described is a living room with a sofa, dining table, and stools, then after object detection, the bounding boxes in the image to be described can frame the outlines of the sofa, dining table, and stools. Thus, the cropped target area image can be a sofa image, a dining table image, and a stool image.

[0068] By cropping at least one target region image from the image to be described, it is easier to generate descriptive information for the target region image in the future, thereby making the descriptive dataset generated for the image to be described more detailed.

[0069] Optionally, when extracting the target region image from the image to be described, the server can also obtain the location information of the target region image within the image to be described. See the following steps included in S200 for details:

[0070] S240: Perform target extraction on the image to be described to obtain at least one target region image and the position information of each target region image in the image to be described.

[0071] Location information can be the coordinates of the target region image within the image to be described. For example, it could be the coordinates of the lower left and upper right corners of the target region image in a coordinate system with the lower left corner of the image to be described as the origin. In some implementations, the generated fine-grained description dataset can be specific to the image to be described. That is, in some application scenarios, such as inputting a description and obtaining the corresponding image, when the image corresponding to a description input by the user is the target region image in the image to be described, there is no readily available cropped target region image to use directly. Therefore, the target region image can be obtained by obtaining its position within the image to be described.

[0072] By extracting targets from the image to be described, at least one target region image and the position information of each target region image in the image to be described can be obtained. This allows the target region image to be extracted from the image to be described using the position information, thereby eliminating the need to store target region images in the image and corresponding description set, and thus reducing the size of the set of images and corresponding descriptions.

[0073] Since the algorithm for generating descriptive information may make mistakes, the generated descriptive information can be matched with the image to verify whether the generated descriptive information can be stored in the fine-grained descriptive dataset. See the following steps after S200:

[0074] S300: Determine whether the image to be described matches the image description information, and whether the image of each target region matches the corresponding target description information.

[0075] The image to be described is subjected to target extraction processing to obtain at least one target region image, and the at least one target region image is described to obtain target description information.

[0076] After obtaining the image description information of the image to be described and the target region image and the target description information of the target region image, since there may be cases where the generated description information does not match the image, it is possible to determine whether the image to be described and the image description information match, and whether each target region image and the corresponding target description information match, so as to ensure that the image to be described and the image description information, and the target region image and the target description information are corresponding.

[0077] In some embodiments, the actions of determining whether the image to be described and the image description match, and determining whether each target region image and its corresponding target description information match, can be performed simultaneously. In other embodiments, after obtaining the image description information, it can be directly determined whether the image to be described and the image description information match, and after obtaining the target description information of the target region image, it can also be directly determined whether the target region image matches the target description information.

[0078] Optionally, how the server determines whether the image and the corresponding description information match can be found in the following steps included in S300:

[0079] S310: Calculate the first cross-modal similarity between the image to be described and the image description information, and the second cross-modal similarity between each target region image and the corresponding target description information.

[0080] Optionally, the image and description information can be converted into corresponding encoded values, and the similarity between the encoded values ​​can be calculated to obtain the similarity between the image and description information. See steps S310 for details:

[0081] S311: Visually encode the image to be described and each target region image to obtain a first image encoding value and a second image encoding value, and perform language encoding on the image description information and each target description information to obtain a first language encoding value and a second language encoding value.

[0082] Optionally, an encoder, such as the CLIP language encoder, can be used to obtain the encoded values ​​corresponding to the image and description. The CLIP algorithm is trained on a large-scale image-text dataset using an alignment learning method, and its generated embeddings can effectively measure the cross-modal similarity between images and text. For details, please refer to the following steps included in S311:

[0083] S3111: Input the image to be described and the image of each target region into the CLIP visual encoder for visual encoding to obtain the first image encoding value and the second image encoding value.

[0084] For example, the image to be described and the image of each target region are input into the CLIP visual encoder for visual encoding to obtain the embeddings of the image to be described and the embeddings of each target region image.

[0085] S3112: Input the image description information and each target description information into the CLIP language encoder respectively to obtain the first language encoding value and the second language encoding value.

[0086] For example, the image description information and the description information of each target are input into the CLIP language encoder for language encoding to obtain the embeddings of the image description information and the embeddings of each target description information.

[0087] Steps S3111 and S3112 can be performed synchronously or asynchronously.

[0088] By converting images and descriptive information into encoded values ​​respectively, the similarity between images and descriptive information can be calculated using image encoded values ​​and language encoded values, thereby realizing the calculation of the similarity between images and descriptive information.

[0089] S312: Calculate the cosine similarity between the first image encoding value and the first language encoding value to obtain the first cross-modal similarity.

[0090] For example, the cosine similarity between the embeddings of the image to be described and the embeddings of the image description information can be calculated to obtain the first cross-modal similarity.

[0091] S313: Calculate the cosine similarity between the second image encoding value and the second language encoding value to obtain the second cross-modal similarity.

[0092] For example, the cosine similarity between the embeddings of the target region image and the corresponding target description information embeddings can be calculated to obtain the second cross-modal similarity.

[0093] Since servers cannot intuitively determine whether an image matches its description like humans can, the similarity between the image and description can be obtained by converting both the image and description into encoded values ​​and then using the similarity between these encoded values.

[0094] S320: Determine whether the first cross-modal similarity meets the first preset condition, and determine whether the second cross-modal similarity meets the second preset condition.

[0095] The first and second preset conditions can be greater than a certain threshold, or between two thresholds. The first preset condition can be the same as the second preset condition.

[0096] Optionally, the preset condition can be greater than or equal to a certain similarity threshold. How the server determines whether the cross-modal similarity obtained in steps S312 and S313 meets the preset condition can be found in the following steps included in S320:

[0097] S321: Determine whether the first cross-modal similarity is greater than or equal to the first similarity threshold.

[0098] S322: Determine whether the second cross-modal similarity is greater than or equal to the second similarity threshold.

[0099] S323: If the first cross-modal similarity is greater than or equal to the first similarity threshold, then the first cross-modal similarity is determined to satisfy the first preset condition.

[0100] S324: If the first cross-modal similarity is less than the first similarity threshold, then it is determined that the first cross-modal similarity does not meet the first preset condition.

[0101] S325: If the second cross-modal similarity is greater than or equal to the second similarity threshold, then the second cross-modal similarity is determined to meet the second preset condition.

[0102] S326: If the second cross-modal similarity is less than the second similarity threshold, then the second cross-modal similarity is determined not to meet the second preset condition.

[0103] For example, the first similarity threshold could be 0.2. After obtaining the first cross-modal similarity between the image to be described and its description, the server can determine whether this first cross-modal similarity is greater than 0.2. If the calculated first cross-modal similarity is 0.8, then the server determines that the first cross-modal similarity meets the first preset condition.

[0104] For example, the second similarity threshold could be 0.3. After obtaining the second cross-modal similarity between the image to be described and its description, the server can determine whether this second cross-modal similarity is greater than 0.3. If the calculated second cross-modal similarity is 0.2, then the server determines that the second cross-modal similarity meets the second preset condition.

[0105] The results of determining whether the first cross-modal similarity meets the first preset condition and whether the second cross-modal similarity meets the second preset condition can be found in the following steps after S320:

[0106] S330: If the first cross-modal similarity satisfies the first preset condition, then it is determined that the image to be described and the image description information match.

[0107] S340: If the first cross-modal similarity does not meet the first preset condition, it is determined that the image to be described and the image description information do not match.

[0108] S350: If the second cross-modal similarity satisfies the second preset condition, then the target region image and the corresponding target description information are determined to match.

[0109] S360: If the second cross-modal similarity does not meet the second preset condition, then it is determined that the target region image and the corresponding target description information do not match.

[0110] By calculating the first cross-modal similarity between the image to be described and the image description information, and the second cross-modal similarity between the target region image and the target description information, and determining whether the obtained similarity meets the preset conditions, it is possible to determine whether the image to be described and the image description information match, and whether the target region image and the target description information match. In this way, description information with low matching degree between description information and image can be filtered out, thereby improving the matching degree between the final generated description dataset and the image, and improving the quality of the description dataset.

[0111] Optionally, if it is determined that the image and its corresponding description do not match, a new description can be generated for the image, and it can be determined whether the newly generated description matches the image. For details, please refer to the following steps after S300:

[0112] S371: If the image to be described and the image description information do not match, then the image to be described is approximated to regenerate the image description information, and the process returns to determine whether the image to be described and the image description information match.

[0113] To improve the utilization rate of the images to be described, ensuring that the image description information can ultimately be stored in the fine-grained image description dataset, image description information different from the previously generated image description information can be regenerated for images to be described that do not match the original image description information. Then, the newly generated image description information is re-evaluated to determine whether it matches the original image. The specific matching operation can be referred to the steps described above, and will not be repeated here.

[0114] In some embodiments, due to the algorithm used to generate image description information, the newly generated image description information may be the same as the previous image description information. In this case, the server can directly determine that the newly generated image description information does not match the information to be described, and regenerate the image description information of the image to be described.

[0115] S372: If the target region image and the corresponding target description information do not match, then the target region image is approximated to regenerate the target description information, and the process returns to determine whether each target region image and the corresponding target description information match.

[0116] For the processing of target region images that do not match the target description information, please refer to the server's processing of the image to be described in S371, which will not be repeated here.

[0117] When a mismatch is found between the image to be described / target region image and the corresponding target region image / target region description, the image to be described / target region image can be re-approximated to regenerate the target description information. This can improve the utilization rate of the image to be described / target region image and reduce the probability of discarding usable images to be described / target region images due to algorithm errors or other factors.

[0118] Optionally, if the image to be described and its corresponding image description information do not match, and the server keeps regenerating the corresponding image description information for the image to be described, it will cause unnecessary resource consumption. Therefore, the number of times the server regenerates the description information can be set so that the server will not keep regenerating and matching the image description information for an image to be described. For details, please refer to the steps before S350:

[0119] S381: Determine whether the number of times the image to be described has been approximated has reached the first preset threshold.

[0120] S382: If the first preset threshold is not reached, continue to perform an approximate description of the image to be described and regenerate image description information.

[0121] S383: If the first preset threshold has been reached, discard the image to be described and the image description information.

[0122] For example, the first preset threshold could be 5 times. If the image to be described and the image description information do not match in the first to fourth calculations, the server can resubmit the image to be described into the image description algorithm for approximate description. However, if the image to be described is approximated and regenerated with image description information in the fifth calculation, and the calculated image description information does not match the image to be described, the server can discard the image to be described to reduce resource waste.

[0123] Optionally, if the target region image and its corresponding target region description do not match, and the server keeps regenerating the corresponding target region description for the target region image, it will cause unnecessary resource consumption. Therefore, the number of times the server regenerates the description information can be set so that the server will not keep regenerating the target description information and matching it for a target region image. For details, please refer to the steps before S350:

[0124] S391: Determine whether the number of times each target region image has been approximated has reached the second preset threshold.

[0125] S392: If the second preset threshold is not reached, continue to perform an approximate description of the target region image to regenerate the target description information.

[0126] S393: If the second preset threshold has been reached, discard the target area image and the corresponding target description information.

[0127] For a description of how the server regenerates the target region description from the target region image until the number of regenerations reaches a preset threshold, please refer to the description of steps S381 to S382, which will not be repeated here.

[0128] S400: If the image to be described matches the image description information, and at least one target region image matches the corresponding target description information, then generate fine-grained description information for the image to be described, including the image description information and the corresponding target description information.

[0129] If the number of times image description information is generated for the image to be described reaches the first preset threshold, and the image to be described still does not match suitable image description information, then the image to be described can be discarded. At this time, at least one target region image corresponding to the image to be described does not have an image to be described, so the at least one target region image and the corresponding target description information can also be discarded.

[0130] If the image to be described matches the image description information, and at least one target region image associated with the image to be described matches the corresponding target description information, then the image description information and the target description information can be processed to generate fine-grained description information for the image to be described, including the image description information and the corresponding target description information. In this way, the fine-grained description dataset can be supplemented with the image to be described and the fine-grained description information of the image to be described.

[0131] In one implementation, if the image to be described matches the image description information, but none of the target region images associated with the image to be described match the corresponding target description information, then the image to be described can be discarded.

[0132] Optionally, since the server can obtain the location information of the target region image in step S240, the fine-grained description information corresponding to the image to be described can also include location information to supplement the fine-grained description information. For details, please refer to the following steps included in S400:

[0133] S410: Generate fine-grained description information for the image to be described, including image description information, corresponding target description information, and corresponding location information.

[0134] For example, such as Figure 3 As shown, the image description information for the image to be described is: A photo of a living room with a white couch and a table with chairs and a vase of flowers on it.

[0135] The corresponding target description information and the corresponding location information can be:

[0136] <0.0000><0.5785><0.4157><0.7604>a white couch with pillows and astuffed animal, where <0.0000><0.5785> are the coordinates of the bottom left corner of the couch, and <0.4157><0.7604> are the coordinates of the bottom right corner of the couch.

[0137] <0.7574><0.6264><0.9741><0.8694>A wooden chair sits next to a table.

[0138] <0.2426><0.6215><0.4352><0.8785>a modern chair in a living room with a white couch.

[0139] <0.5259><0.6646><0.7472><0.9590>a chair with a metal frame and orangeseat.

[0140] <0.5556><0.0576><0.7444><0.2951>a modern white ceiling light fixture with a circular shape.

[0141] Therefore, by performing the above processing on a large number of images to be described downloaded from the network, fine-grained description information can be obtained for each image to be described, each corresponding target description information, and each corresponding location information, so as to form a high-quality fine-grained description dataset.

[0142] Optionally, such as Figure 4 As shown, after generating fine-grained descriptive information for the image to be described, this fine-grained descriptive information can be stored in the live background image library for application in live streaming scenarios. For details, please refer to the following steps after S400:

[0143] S420: Associate and store fine-grained description information and the image to be described in the live background image library.

[0144] The live background image library includes multiple images to be described and corresponding fine-grained descriptive information, so that images or parts of the images to be described can be selected from the live background image library based on certain descriptions provided by the user.

[0145] S430: In response to the background image request command from the live streaming user, determine the fine-grained description information that matches the background image request command in the live streaming background image library, and obtain the image to be described corresponding to the fine-grained description information.

[0146] S440: Send the acquired image to be described to the client terminal corresponding to the live stream user, so that the client terminal can set the image to be described as the live stream background on the live stream interface.

[0147] For example, if a live stream user wants to set the background to a living room with a white sofa, the client can send a background image request command to the live stream server. This request carries information about the living room with the white sofa. In response, the live stream server can determine the fine-grained description information matching "living room with a white sofa" from the live stream background image library and obtain the corresponding image to be described. Alternatively, the fine-grained image descriptions generated by the above method can be used for training when creating virtual backgrounds to generate more realistic and vivid virtual background images.

[0148] In some embodiments, if a live stream user expects the background to be a portion of the image to be described, the server can use the location information included in the image to crop that portion of the image from the image to be described to generate the live stream background.

[0149] By using fine-grained descriptive datasets to generate live streaming backgrounds, it is possible to customize live streaming backgrounds according to the needs of live streaming users, reduce the manpower consumption caused by creating live streaming backgrounds, and also enable live streaming users to attract viewers with interesting live streaming backgrounds, thereby increasing the interaction rate of the live streaming room.

[0150] like Figure 5 As shown, the embodiment of the description processing method for live background images in this application can be executed by a live server. The description processing method can include: M100: performing target extraction processing on the image to be described to obtain at least one target region image, and describing the at least one target region image to obtain target description information. M200: determining whether the image to be described and the image description information match, and whether each target region image and the corresponding target description information match. M300: if the image to be described and the image description information match, and at least one target region image and the corresponding target description information match, then generating fine-grained description information including image description information and corresponding target description information for the image to be described. M400: associating and storing the fine-grained description information and the image to be described in a live background image library. M500: responding to the background image request instruction from the live user, determining the fine-grained description information that matches the background image request instruction in the live background image library, and obtaining the image to be described corresponding to the fine-grained description information. M600: sending the obtained image to be described to the client terminal corresponding to the live user, so that the client terminal can set the image to be described as the live background on the live room interface.

[0151] The description of the live background image processing method embodiment of this application can be found in the above description of the server in the image description processing method embodiment of this application, and will not be repeated here.

[0152] like Figure 6As shown in the embodiments of the computer device described in this application, the computer device 100 can be the server described above. The computer device 100 may include a processor 110, a memory 120, and a communication circuit 130.

[0153] The memory 120 is used to store computer programs and may be ROM (Read-Only Memory), RAM (Random Access Memory), or other types of storage devices. Specifically, the memory may include one or more computer-readable storage media, which may be non-transitory. The memory may also include high-speed random access memory and non-volatile memory, such as one or more disk storage devices or flash memory devices. In some embodiments, the non-transitory computer-readable storage media in the memory is used to store at least one line of program code.

[0154] Processor 110 is used to control the operation of computer device 100. Processor 110 may also be referred to as CPU (Central Processing Unit). Processor 110 may be an integrated circuit chip with signal processing capabilities. Processor 110 may also be a general-purpose processor, digital signal processor (DSP), application-specific integrated circuit (ASIC), off-the-shelf programmable gate array (FPGA), or other programmable logic device, discrete gate or transistor logic device, or discrete hardware component. The general-purpose processor may be a microprocessor, or processor 110 may be any conventional processor.

[0155] The processor 110 is used to execute the computer program stored in the memory 120 to implement the image description processing method described in the embodiments of the image description processing method of this application and the live background image description processing method.

[0156] The computer device 100 may also include a communication circuit 130, which is a communication connection device or circuit used by the computer device 100 to communicate with external devices, so that the processor 110 can interact with external devices via the communication circuit 130.

[0157] For a detailed description of the functions and execution processes of each functional module or component in the computer device embodiments of this application, please refer to the descriptions in the above-described embodiments of the image description processing method and the description processing method for live background images, which will not be repeated here.

[0158] In the several embodiments provided in this application, it should be understood that the disclosed computer device 100 and image description processing method can be implemented in other ways. For example, the embodiments of the computer device 100 described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the mutual coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.

[0159] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment, depending on actual needs.

[0160] Furthermore, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.

[0161] See Figure 7 If the integrated units described above are implemented as software functional units and sold or used as independent products, they can be stored in computer-readable storage medium 200. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions / computer programs to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute all or part of the steps of the methods of various embodiments of this invention. The aforementioned storage medium includes various media such as USB flash drives, portable hard drives, read-only memory, random access memory, magnetic disks, or optical disks, as well as electronic terminals such as computers, mobile phones, laptops, tablets, and cameras that have the aforementioned storage media.

[0162] The description of the execution process of program data in a computer-readable storage medium can be found in the above embodiments of the image description processing method and the description processing method for live background images, and will not be repeated here.

[0163] The above description is merely an embodiment of this application and does not limit the patent scope of this application. Any equivalent structural or procedural transformations made using the content of this application's specification and drawings, or direct or indirect applications in other related technical fields, are similarly included within the patent protection scope of this application.

Claims

1. An image description processing method, characterized in that, include: The image to be described is described to obtain image description information; The image to be described is subjected to target extraction processing to obtain at least one target region image, and the at least one target region image is described to obtain target description information; Determine whether the image to be described and the image description information match, and whether each target region image and the corresponding target description information match; If the image to be described matches the image description information, and at least one of the target region images matches the corresponding target description information, then fine-grained description information including the image description information and the corresponding target description information is generated for the image to be described. After determining whether the image to be described and the image description information match, and whether each target region image and the corresponding target description information match, the process includes: If the image to be described and the image description information do not match, then an approximate description is performed on the image to be described to regenerate the image description information, and the process of determining whether the image to be described and the image description information match is returned. If the target region image and the corresponding target description information do not match, then the target region image is approximated to regenerate the target description information, and the process of determining whether each target region image and the corresponding target description information match is returned.

2. The method according to claim 1, characterized in that, Determining whether the image to be described matches the image description information, and whether each target region image matches the corresponding target description information, includes: Calculate the first cross-modal similarity between the image to be described and the image description information, and the second cross-modal similarity between each target region image and the corresponding target description information; Determine whether the first cross-modal similarity satisfies a first preset condition, and determine whether the second cross-modal similarity satisfies a second preset condition; If the first cross-modal similarity satisfies the first preset condition, then the image to be described and the image description information are determined to match; otherwise, they do not match. If the second cross-modal similarity satisfies the second preset condition, then the target region image and the corresponding target description information are determined to match; otherwise, they do not match.

3. The method according to claim 2, characterized in that, The calculation of the first cross-modal similarity between the image to be described and the image description information, and the second cross-modal similarity between each target region image and the corresponding target description information, includes: Visual encoding is performed on the image to be described and each of the target regions to obtain a first image encoding value and a second image encoding value, respectively; and language encoding is performed on the image description information and each of the target description information to obtain a first language encoding value and a second language encoding value, respectively. The first cross-modal similarity is obtained by calculating the cosine similarity between the first image encoding value and the first language encoding value. The second cross-modal similarity is obtained by calculating the cosine similarity between the second image encoding value and the second language encoding value.

4. The method according to claim 3, characterized in that, The step of visually encoding the image to be described and each target region image to obtain a first image encoding value and a second image encoding value, and performing language encoding on the image description information and each target description information to obtain a first language encoding value and a second language encoding value, includes: The image to be described and each of the target regions are respectively input into the CLIP visual encoder to perform visual encoding to obtain the first image encoding value and the second image encoding value; The image description information and each of the target description information are respectively input into the CLIP language encoder for language encoding to obtain the first language encoding value and the second language encoding value.

5. The method according to claim 2, characterized in that, The determination of whether the first cross-modal similarity satisfies the first preset condition, and the determination of whether the second cross-modal similarity satisfies the second preset condition, include: Determine whether the first cross-modal similarity is greater than or equal to a first similarity threshold; and determine whether the second cross-modal similarity is greater than or equal to a second similarity threshold; If the first cross-modal similarity is greater than or equal to the first similarity threshold, then the first cross-modal similarity is determined to satisfy the first preset condition; otherwise, the first preset condition is not satisfied. If the second cross-modal similarity is greater than or equal to the second similarity threshold, then the second cross-modal similarity is determined to satisfy the second preset condition; otherwise, the second preset condition is not satisfied.

6. The method according to claim 1, characterized in that, Before performing an approximate description of the image to be described to regenerate the image description information, the process includes: Determine whether the number of times the image to be described has been approximated has reached a first preset threshold. If the first preset threshold is not reached, the process of approximating the image to be described and regenerating the image description information continues; if the first preset threshold has been reached, the image to be described and the image description information are discarded. Before performing an approximate description of the target region image to regenerate the target description information, the process includes: Determine whether the number of times each target region image has been approximated has reached a second preset threshold; If the second preset threshold is not reached, the process of approximating the target region image and regenerating the target description information continues; if the second preset threshold has been reached, the target region image and the corresponding target description information are discarded.

7. The method according to claim 1, characterized in that: The step of performing target extraction processing on the image to be described to obtain at least one target region image includes: Perform target detection on the image to be described to identify at least one target object; Determine whether each target object meets the preset clipping conditions; If the preset cropping conditions are met, at least one target region image is cropped from the image to be described; each target region image includes at least one target object.

8. The method according to claim 7, characterized in that, Determining at least one target object includes: In the image to be described, at least one target bounding box is determined to select the target object, and the confidence level of the target object is calculated; The step of determining whether each target object meets the preset clipping conditions includes: Whether the target object meets the preset clipping conditions is determined using at least one of the target bounding box and the confidence level.

9. The method according to claim 8, characterized in that, Determining whether the target object meets the preset clipping conditions using at least one of the target bounding box and the confidence score includes: Determine whether the size of the target box is greater than or equal to a preset size, and / or whether the confidence level is greater than or equal to a preset confidence level; If the size of the target box is smaller than the preset size, and / or the confidence level is less than the preset confidence level, then the target object is determined to not meet the preset clipping conditions; If the size of the target box is greater than or equal to the preset size, and / or the confidence level is greater than or equal to the preset confidence level, then the target object is determined to meet the preset clipping conditions.

10. The method according to claim 8, characterized in that, The step of cropping the at least one target region image from the image to be described includes: The image to be described is cropped according to the target bounding box to crop out the at least one target region image.

11. The method according to claim 1, characterized in that, The step of performing target extraction processing on the image to be described to obtain at least one target region image includes: Target extraction is performed on the image to be described to obtain at least one target region image and the position information of each target region image in the image to be described; The step of generating fine-grained description information for the image to be described, including the image description information and the corresponding target description information, includes: Fine-grained description information is generated for the image to be described, including the image description information, the corresponding target description information, and the corresponding location information.

12. The method according to claim 1, characterized in that, Before describing the image to be described to obtain image description information, the process includes: Obtain raw images from the internet and perform classification processing to obtain at least one image set; Watermark removal is performed on each original image in each of the image sets to obtain the image to be described.

13. The method according to claim 1, characterized in that, After generating fine-grained description information for the image to be described, including the image description information and the corresponding target description information, the process includes: The fine-grained description information and the image to be described are associated and stored in the live background image library; In response to a background image request instruction from a live streaming user, the fine-grained description information matching the background image request instruction is determined from the live streaming background image library, and the image to be described corresponding to the fine-grained description information is obtained. The acquired image to be described is sent to the client terminal corresponding to the live stream user, so that the client terminal can set the image to be described as the live stream background on the live stream interface.

14. A method for describing and processing a live streaming background image, characterized in that: The image to be described is subjected to target extraction processing to obtain at least one target region image, and the at least one target region image is described to obtain target description information; Determine whether the image to be described and the image description information match, and whether each target region image and the corresponding target description information match; If the image to be described matches the image description information, and at least one of the target region images matches the corresponding target description information, then fine-grained description information including the image description information and the corresponding target description information is generated for the image to be described. The fine-grained description information and the image to be described are associated and stored in the live background image library; In response to a background image request instruction from a live streaming user, the fine-grained description information matching the background image request instruction is determined from the live streaming background image library, and the image to be described corresponding to the fine-grained description information is obtained. The acquired image to be described is sent to the client terminal corresponding to the live stream user, so that the client terminal can set the image to be described as the live stream background on the live stream interface. After determining whether the image to be described and the image description information match, and whether each target region image and the corresponding target description information match, the process includes: If the image to be described and the image description information do not match, then an approximate description is performed on the image to be described to regenerate the image description information, and the process of determining whether the image to be described and the image description information match is returned. If the target region image and the corresponding target description information do not match, then the target region image is approximated to regenerate the target description information, and the process of determining whether each target region image and the corresponding target description information match is returned.

15. A computer device, characterized in that, include: Processor, memory, and communication circuitry; The communication circuit and the memory are respectively coupled to the processor. The memory is used to store a computer program, and the processor is used to read and execute the computer program to implement the method as described in any one of claims 1-14.

16. A computer-readable storage medium, characterized in that, The device contains a computer program that can be read and executed by a processor to implement the method as described in any one of claims 1-14.

Citation Information

Patent Citations

  • Image description generation method and device, electronic equipment and storage medium

    CN114648631A

  • Training method and device of image description information generation model, equipment and medium

    CN114842299A