A method and an electronic device for detecting a target described by natural language based on an image

Through the multimodal large language model and the reference matching model, the detailed positioning description data is generated, and the problem of low accuracy of complex natural language description object detection in the prior art is solved, and more efficient image information detection and generation is achieved.

CN120032149BActive Publication Date: 2025-07-11HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510469031.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-07-11
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The existing open-set object detection and reference expression understanding technology has low detection accuracy when facing complex natural language descriptions, and cannot effectively identify multi-category names and multi-reference expression goals.

Method used

The multimodal large language model and the reference matching model are used to generate detailed positioning description data by training the expert model, and candidate targets in the image that match the natural language description target are obtained, and detailed positioning description data are used for detection and generation.

Benefits of technology

It improves the detection accuracy of natural language description targets, can process complex natural language descriptions, and generates more accurate image information descriptions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032149B_ABST
    Figure CN120032149B_ABST
Patent Text Reader

Abstract

The present application discloses a method for detecting a natural language description target based on an image, including: inputting an image to be detected into a trained expert model for converting the input image into detailed localization description data with image detailed description data and performing localization description on text instances in the image detailed description data, obtaining the detailed localization description data through the inference of the expert model, where the detailed localization description data includes: image detailed description data, and image instance description data corresponding to the text instances in the image detailed description data, and using the detailed localization description data of the image to be detected to obtain candidate targets in the image to be detected that match the natural language description target characterized by the text instances. The present application is beneficial to improving the accuracy of detecting the target described by the natural language.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular, to a method for detecting a target described by natural language based on an image. Background Art

[0002] Detecting the target described by natural language in an image has always been an important research field. Currently, the technologies for detecting the target described by natural language include: open-set object detection technology (Open Vocabulary object Detection, OVD) and referring expression comprehension (Referring Expression Comprehension, REC) technology. Among them,

[0003] Open-set object detection can detect the target described by any category name. The category name is usually a short category name without attributes or relationships. For example, instances such as cats, dogs, trees, cars, etc. in an image, and the described target is described by the category name, that is, using cats, dogs, trees, cars as the category names.

[0004] Referring expression comprehension can understand the referring expression and detect and forcefully detect the target of a bounding box. The referring expression is a word or a phrase used to identify a specific person, place, or thing, usually a noun, a noun phrase, or a pronoun, also known as a phrasal natural language description. For example, an isolated phrasal description such as a man carrying a schoolbag.

[0005] Although open-set object detection and referring expression comprehension can detect the target referred to by simple natural language descriptions, users need to customize the natural language descriptions of the targets to be detected. Due to the different natural language description abilities of different users, the detection accuracy is relatively low. Once at least one of multi-category names and multi-referring expressions is used to describe the target, it will lead to the inability to detect the target referred to by at least one of multi-category names and multi-referring expressions, resulting in their inability to cope with more open application scenarios.

[0006] For example, when detecting the man wearing a white mask in "There is a lady wearing red clothes and blue pants on the left side of the image, and there is a man wearing a white mask behind her.", current detection means are powerless for such natural language descriptions that require long-context understanding. Summary of the Invention

[0007] The present invention provides a method for detecting a target described by natural language based on an image to improve the accuracy of detecting the target described by natural language.

[0008] The present invention provides a method for detecting a target described by natural language based on an image, and the method includes:

[0009] Input the image to be detected into a trained expert model for converting the input image into detailed localization description data with image detailed description data and locating and describing text instances in the image detailed description data. Through the inference of the expert model, obtain the detailed localization description data of the image to be detected.

[0010] Among them,

[0011] The detailed localization description data of the image to be detected includes: the image detailed description data of the image to be detected, and the image instance description data corresponding to the text instances in the image detailed description data.

[0012] The image instance description data is used to characterize the position information of the image instance corresponding to the text instance in the image.

[0013] The text instance is used to characterize the natural language description target referred to in the image detailed description data.

[0014] Utilize the detailed localization description data of the image to be detected to obtain candidate targets in the image to be detected that match the natural language description target characterized by the text instance.

[0015] Preferably, the step of utilizing the detailed localization description data of the image to be detected to obtain candidate targets in the image to be detected that match the natural language description target characterized by the text instance includes:

[0016] Based on the detailed localization description data of the image to be detected, obtain the position information of the image instance located by the image instance description data in the image, and the text instances in the image detailed description data.

[0017] According to the position information of the image instance in the image, determine the image content of the image instance, detect whether the determined image content matches the natural language description target, and use the image instance that matches the natural language description target as the candidate target.

[0018] The detailed localization description data of the image to be detected further includes: text instance description data corresponding to the text instances in the image detailed description data.

[0019] Preferably, the text instance description data includes: description data for referring to the instance noun of the natural language description target.

[0020] The image instance description data includes: description data of the target box for characterizing the image position where the image instance is located.

[0021] The step of, based on the detailed localization description data of the image to be detected, obtaining the position information of the image instance located by the image instance description data in the image, and the text instances in the image detailed description data includes:

[0022] Input the detailed localization description data of the image to be detected into the first large language model. Through the inference of the first large language model, obtain the target box description data in the detailed localization description data and the description data of the instance nouns in the detailed localization description data;

[0023] Determining the image content of the image instance according to the position information of the image instance in the image, and detecting whether the determined image content matches the natural language description target includes:

[0024] Input the image to be detected, the target box description data in the detailed localization description data, and the description data of the instance nouns in the detailed localization description data into the trained first referential expression matching model for matching the content information within the target box area with the referential expressions representing the description data of the instance nouns. Through the inference of the first referential expression matching model, filter out the incorrect target box description data in the detailed localization description data to obtain the denoised detailed localization description data of the image to be detected.

[0025] Preferably, the description data of the instance nouns includes: the first description data of the instance nouns and / or the second description data of the instance nouns, where the first description data is used to represent the simple description of the instance nouns, and the second description data is used to represent the complex description of the instance nouns. The simple description of the instance nouns includes the instance category name and / or the description of the content of interest. The complex description of the instance nouns includes: multiple category names and / or multiple phrasal arbitrary natural language descriptions other than the single category name description data;

[0026] The expert model is trained in the following manner:

[0027] Obtain the first-round detailed localization description data,

[0028] Use the first-round detailed localization description data as sample data to train the first multimodal large model until the first multimodal large model reaches the expectation, and obtain the trained first multimodal large model,

[0029] Use the trained first multimodal large model as the first-round trained expert model.

[0030] Preferably, the first referential expression matching model is trained in the following manner:

[0031] Input the first-round detailed localization description data into the second large language model. Through the inference of the second large language model, obtain the target box description data in the first-round detailed localization description data and the description data of the instance nouns,

[0032] Based on the target box description data and the description data of instance nouns in the first-round detailed localization description data, obtain positive and negative sample data.

[0033] Use the target box description data and the description data of instance nouns in the first-round detailed localization description data, as well as the image data from which the first-round detailed localization description data is sourced, and the positive and negative sample data as sample data to train the second multi-modal large model until the second multi-modal large model meets the expectations, obtaining the trained second multi-modal large model.

[0034] Use the trained second multi-modal large model as the first referential expression matching model that has been trained in the first round.

[0035] Preferably, the training of the expert model further includes:

[0036] Use the denoised detailed localization description data as the sample data for this round to train the previous-round first multi-modal large model until the previous-round first multi-modal large model meets the expectations, obtaining the first multi-modal large model trained in this round.

[0037] Use the first multi-modal large model trained in this round as the currently trained expert model for detecting the next image to be detected.

[0038] The training of the first referential expression matching model further includes:

[0039] Input the denoised detailed localization description data into the second large language model, and through the reasoning of the second large language model, obtain the target box description data and the description data of instance nouns in the denoised detailed localization description data.

[0040] Based on the target box description data and the description data of instance nouns in the denoised detailed localization description data, obtain positive and negative sample data.

[0041] Use the target box description data and the description data of instance nouns in the denoised detailed localization description data, as well as the image data from which the denoised detailed localization description data is sourced, and the positive and negative sample data as the sample data for this round to train the previous-round second multi-modal large model until the previous-round second multi-modal large model meets the expectations, obtaining the second multi-modal large model trained in this round.

[0042] Use the second multi-modal large model trained in this round as the currently trained first referential expression matching model for detecting the next image to be detected.

[0043] Preferably, the first-round detailed localization description data is obtained in the following manner:

[0044] Obtain the image detailed description data of the sample image.

[0045] Based on the image detailed description data of the sample image, perform text instance extraction to obtain the first description data and the second description data of the instance nouns included in the text instances in the sample image.

[0046] Perform object detection based on the sample image data to obtain a first target box that matches the first description data.

[0047] Match the sample image content of the first target box with the second description data to obtain a second target box that matches the second description data.

[0048] Fuse the second target box, the first description data of the instance nouns, the second description data, and the image detailed description data of the sample image to obtain detailed localization description pseudo-label data, which includes: image detailed description data, the first description data of the instance nouns of the second target box, the second description data, and the pseudo-label target box description data of the second target box.

[0049] Preferably, the obtaining of the image detailed description data of the sample image includes:

[0050] Input the sample image data into a third multi-modal large model, and through the inference of the third multi-modal large model, obtain the image detailed description data of the sample image.

[0051] The performing of text instance extraction based on the image detailed description data of the sample image includes:

[0052] Input the image detailed description data of the sample image into a third language large model, and through the inference of the third language large model, obtain the first description data of the instance nouns of the text instances in the sample image.

[0053] Input the text where the first description data of the instance nouns of the text instances in the sample image is located into a fourth language large model, and through the inference of the fourth language large model, obtain the second description data of the instance nouns of the text instances in the sample image.

[0054] The performing of object detection based on the sample image data to obtain a first target box that matches the first description data includes:

[0055] Input the sample image and the first description data of its instance nouns into an open-set object detection model, and through the inference of the open-set object detection model, obtain the position information of the image where the first target box is located and the instance nouns corresponding to the first target box.

[0056] The matching of the sample image content of the first target box with the second description data includes:

[0057] Input the sample image, the position information of the image where the first target box is located, the instance noun data corresponding to the first target box, and the second description data into the second deictic expression matching model. Through the inference of the second deictic expression matching model, obtain the position information of the image where the second target box that matches the second description data is located;

[0058] After fusing the second target box, the first description data of the instance noun, the second description data, and the detailed image description data of the sample image, it further includes:

[0059] Remove the incorrect data in the detailed localization description pseudo-label data to obtain the first-round detailed localization description data.

[0060] Preferably, inputting the detailed image description data of the sample image into the third large language model further includes:

[0061] Input the detailed image description data of the sample image and the first text prompt for prompting the instance nouns in the detailed image description data of the sample image into the third large language model to determine the positions of the instance nouns in the text in the detailed image description data of the sample image;

[0062] Inputting the text where the first description data of the instance noun of the text instance in the sample image is located into the fourth large language model further includes:

[0063] Input the paragraph text where the instance noun is located and the second text prompt for prompting the complex descriptions in the detailed image description data of the sample image into the fourth large language model to determine the complex descriptions of the instance nouns in the detailed image description data of the sample image.

[0064] The second aspect of the present invention provides a method for generating natural language description data of an image. The generation method includes:

[0065] Input the target image for which natural language description data is to be generated and / or the detailed target image description data for which natural language description data is to be generated into the trained expert model. Through the inference of the expert model, obtain the detailed localization description data,

[0066] Among them,

[0067] The expert model is used to convert the input image into detailed localization description data with detailed image description data and perform image localization description on the text instances in the detailed image description data, and / or perform text localization description on the text instances in the input detailed image description data,

[0068] The detailed location description data includes: image detailed description data, and image instance description data corresponding to the text instances in the image detailed description data, and / or, the detailed location description data includes: image detailed description data, and text instance description data corresponding to the text instances in the image detailed description data.

[0069] The image instance description data is used to characterize the position information of the image instance corresponding to the text instance in the image.

[0070] The text instance is used to characterize the natural language description target referred to in the image detailed description data.

[0071] A third aspect of the present application provides an electronic device, including a memory and a processor. The memory stores a computer program, and the processor is configured to execute the computer program to implement the steps of the method for detecting a natural language description target based on an image, and / or the steps of the method for generating natural language description data of an image.

[0072] The method for detecting a natural language description target based on an image provided by the present application obtains a candidate target in the image to be detected that matches the natural language description target characterized by the text instance by generating detailed location description data of the image to be detected. The present application neither needs to preset the natural language description target nor can detect the target referred to by the text instance with any complex description in the detailed location description data, which is beneficial to improving the accuracy of natural language description target detection; and, the present application is beneficial to improving the intelligence of generating natural language description data of an image through the generated detailed location description data, so that the information contained in the image is more accurately described. Description of the Drawings

[0073] Figure 1 It is a schematic flowchart of the method for detecting a natural language description target based on an image according to an embodiment of the present application.

[0074] Figure 2 It is a schematic flowchart of obtaining the first-round detailed location description data for training the first multi-modal large model in this embodiment.

[0075] Figure 3 It is a schematic diagram of extracting the first description data of instance nouns in this embodiment.

[0076] Figure 4 It is a schematic diagram of extracting the second description data of instance nouns in this embodiment.

[0077] Figure 5 It is a schematic diagram of obtaining the first-round detailed location description data based on an artificial intelligence model in this embodiment.

[0078] Figure 6A schematic flowchart of a method for generating natural language description data of an image in this embodiment and detecting a natural language description target.

[0079] Figure 7 A schematic diagram of generating natural language description data of an image in this embodiment.

[0080] Figure 8 A schematic diagram of a frame of image in this embodiment.

[0081] Figure 9 For this embodiment based on Figure 8 A schematic diagram of a target box corresponding to detailed localization data.

[0082] Figure 10 A schematic flowchart of a device for detecting a natural language description target based on an image in an embodiment of this application.

[0083] Figure 11 A schematic flowchart of a device for generating natural language description of an image in an embodiment of this application.

[0084] Figure 12 Another schematic diagram of a device for detecting a natural language description target based on an image and / or a device for generating natural language description data of an image in an embodiment of this application. Detailed implementation manners

[0085] In order to make the purpose, technical means and advantages of this application clearer, the following further describes this application in detail with reference to the accompanying drawings.

[0086] To facilitate the understanding of the embodiments of this application, the following explains the technical terms in the embodiments of this application.

[0087] Dense caption: Natural language description of each local detail in an image.

[0088] Complex description: Natural language description of multiple phrases and / or any combination of multiple category names other than a single category name, including but not limited to, natural language descriptions that require long context understanding, and natural language descriptions with at least one of certain logic, deduction, and induction.

[0089] Detailed localization description data: Includes image detailed description data and description data for locating and describing text instances in the image detailed description data. Among them, the localization description includes: image localization description for describing the image position in the image, and text localization description for describing the text position in the image detailed description data, either one or a combination of both.

[0090] Text instance: The instance described by the text instance description data in the image detailed description data.

[0091] Text instance description data: A set of description data in the image detailed description data for describing the instance nouns used to describe text instances. Among them, the description data of the instance nouns includes one or a combination of the following two: the first description data for simply describing the instance nouns and the second description data for complexly describing the instance nouns. In this application, the simple description of the instance nouns can be understood as the instance noun description referred to by the category name with attribute description. For example, a single phrasal instance noun. The complex description of the instance nouns can be understood as the instance nouns referred to in a complex description manner. For example, the context formed by the combination of multiple different phrasal instance nouns. The text instance description data can be used for text location description.

[0092] Image instance: An entity target in an image.

[0093] Image instance description data: The description data in the image detailed description data that describes the image area where the image instance is located in the form of image position information (such as pixel coordinates); the image instance description data can be used for image location description; the image area where the image instance is located is usually represented by a target box, so the image instance description data includes target box description data.

[0094] Referring Expression Matching (REM): Detect whether the image content in the target box area matches the content described by the referring expression.

[0095] The method for detecting a target described in natural language based on an image provided in an embodiment of this application generates detailed location description data with image detailed description data and locates and describes each text instance in the image detailed description data through the inference of a trained multi-modal large model for the to-be-detected image input to the trained multi-modal large model, and uses the detailed location description data to implement the detection of the target described in natural language.

[0096] See Figure 1 as described Figure 1 which is a schematic flowchart of a method for detecting a target described in natural language based on an image according to an embodiment of this application. The method includes:

[0097] Step 101, input the to-be-detected image into a trained expert model for converting the input image into detailed location description data with image detailed description data and locating and describing the text instances in the image detailed description data, and through the inference of the expert model, obtain the detailed location description data of the to-be-detected image.

[0098] Among them,

[0099] The detailed localization description data of the image to be detected includes: the detailed image description data of the image to be detected, and the image instance description data corresponding to the text instance in the detailed image description data.

[0100] The image instance description data is used to characterize the position information of the image instance corresponding to the text instance in the image.

[0101] The text instance is used to characterize the natural language description target referred to in the detailed image description data.

[0102] As an example,

[0103] The expert model is trained in the following manner:

[0104] Obtain the first-round detailed localization description data.

[0105] Use the first-round detailed localization description data as sample data to train the first multi-modal large model until the first multi-modal large model meets the expectations, and obtain the trained first multi-modal large model.

[0106] Use the trained first multi-modal large model as the first-round trained expert model.

[0107] Furthermore, use the denoised detailed localization description data as the sample data for this round to train the previous-round first multi-modal large model until the previous-round first multi-modal large model meets the expectations, obtain the trained first multi-modal large model for this round, and use the trained first multi-modal large model for this round as the currently trained expert model to detect the next image to be detected.

[0108] As an example, the detailed localization description data of the image to be detected further includes: the text instance description data corresponding to the text instance in the detailed image description data.

[0109] As an example, the text instance description data includes: the description data of the instance noun used to refer to the natural language description target, and the image instance description data includes: the description data of the target box used to characterize the position of the image instance in the image.

[0110] Step 102, using the detailed localization description data of the image to be detected, obtain the candidate target in the image to be detected that matches the natural language description target characterized by the text instance.

[0111] As an example,

[0112] Based on the detailed localization description data of the image to be detected, obtain the position information of the image instance located by the image instance description data in the image, and the text instance in the detailed image description data.

[0113] Based on the position information of the image instance in the image, determine the image content of the image instance, detect whether the determined image content matches the natural language description target, and use the image instance that matches the natural language description target as a candidate target.

[0114] For example, input the detailed localization description data of the image to be detected into the first large language model. Through the inference of the first large language model, obtain the target box description data in the detailed localization description data and the description data of the instance nouns in the detailed localization description data. Input the image to be detected, the target box description data in the detailed localization description data, and the description data of the instance nouns in the detailed localization description data into the trained first referential expression matching model for matching the content information within the target box area with the referential expression representing the description data of the instance nouns. Through the inference of the first referential expression matching model, filter out the incorrect target box description data in the detailed localization description data to obtain the denoised detailed localization description data of the image to be detected, and the denoised detailed localization description data includes candidate target box description data.

[0115] The method for detecting natural language description targets based on images provided by the embodiments of the present application can, by generating the detailed localization description data of the image, obtain candidate targets in the image that match the natural language description targets represented by text instances based on the detailed localization description data. The embodiments of the present application do not require a preset natural language description for the target to be detected, and moreover, can detect targets with any complex descriptions in the detailed localization description data.

[0116] For ease of understanding the embodiments of the present application, the following takes the implementation manner of an artificial intelligence model as an example for illustration. It should be understood that the embodiments of the present application are not limited to specific models.

[0117] See Figure 2 as shown Figure 2 is a schematic flowchart of a process for obtaining the first-round detailed localization description data for training the first multimodal large model in this embodiment. Specifically, it includes:

[0118] Step 201, based on the sample image data, obtain the image detailed description data of the sample image.

[0119] As an example, input a frame of sample image into the third multimodal large model. Through the inference of the third multimodal large model, obtain the image detailed description data of the input sample image.

[0120] Among them, the third multimodal large model is a trained model that has the ability to understand the input image and generate a natural language description, that is, the third multimodal large model already has the ability to generate text from images.

[0121] Step 202: Based on the generated detailed image description data, perform text instance extraction to obtain the description data of the instance nouns of the text instances in the detailed image description data of the sample image.

[0122] Among them, the instance noun is used to represent the category name of the text instance in the detailed image description data. The description data of the instance noun includes: the first description data used to represent the simple description of the instance noun, and the second description data used to represent the complex description of the instance noun. The simple description can describe the instance by its category name and / or the content of interest. For example, optical character recognition (OCR) content, relationships, attributes, etc., which is equivalent to the description of the instance name itself. The complex description uses natural language descriptions of multiple phrases and / or any combination of multiple category names other than a single category name. For example, the description of context understanding.

[0123] In this way, the first description data of the instance noun represents the description data of the instance noun and its related information. For example, a red balloon, etc., to achieve the local detail description of the text instance, and the second description data of the instance noun realizes the summary and refinement of the text instance.

[0124] See Figure 3 as shown Figure 3 This is a schematic diagram for extracting the first description data of the instance noun in this embodiment. The detailed image description data of the sample image and its corresponding first text prompt are input into the third language large model. Through the inference of the third language large model, the text position of the instance noun in the detailed image description data is determined to obtain the first description data of the instance noun of the input detailed image description data. Among them, the first text prompt is used to prompt the instance noun in the detailed image description data, and the third language large model is a trained model that has the ability to understand the input text and generate the first description of the instance noun.

[0125] For example, the detailed image description data of the sample image is: This photo was taken at time yyyy - mm - dd hh:ff:ss, and the location is pppp. The picture shows a part of the city street. There are shops and facilities on the left side of the picture, and parked cars on the right side, including a silver sedan with license plate number AAA. On the left side of the picture, there is a middle - aged woman wearing a gray coat and black tight - fitting pants and black shoes. She is leading a dog with tan hair...

[0126] Input the detailed image description data of this sample image into the third language large model, and the output text marked with the first description data (underlined) of the instance noun can be obtained as follows:

[0127] This photo was taken at Time yyyy - mm - dd hh:ff:ss , Location is pppp。The picture shows a part of a city street. There are shops and facilities on the left side of the picture, and parked Automobile , including a License Plate Number for AAA of Silver Sedan 。On the left side of the picture, there is a Middle - aged Woman , wearing Grey Coat and Black Tight Trousers , with Black Shoes on her feet. She is leading a Brown Brown Hair of Dog …

[0128] See Figure 4 as shown. Figure 4 This is a schematic diagram for extracting the second description data of instance nouns in this embodiment. The text of the paragraph where the instance noun is located and its second text prompt are input into the fourth language model. Through the inference of the fourth language model, the context text position of the instance noun in the image detailed description data of the sample image is determined, that is, the text positions of multiple phrasal descriptions of the instance noun in the image detailed description data of the sample image are determined, and the second description data of the instance noun is obtained. Among them, the text of the paragraph where the instance noun is located can be marked, and the second text prompt is used to prompt the phrasal description of the instance noun in the image detailed description data. The fourth language model is a trained model that has the ability to understand the input text and generate the second description of the instance noun.

[0129] The third language model and the fourth language model can be the same language model or different language models, and this application does not limit this.

[0130] For example, input the description paragraph "On the left side of the picture, there is a middle-aged woman, wearing a gray coat and black tight pants, with black shoes on her feet. She is leading a dog with tan hair" in the image detailed description data of the sample image and the text prompt "woman" into the fourth language model, and the second description data of the instance noun "woman" can be obtained:

[0131] A middle-aged woman wearing a gray coat

[0132] A woman wearing black tight pants and black shoes, leading a dog

[0133] Step 203, based on the first description data and the second description data of the instance noun, perform object detection and anaphora matching to obtain the target box information that matches the second description data.

[0134] As an example, the sample image and the first description data of the instance noun are input into the open-set object detection model. Through the inference of the open-set object detection model, the first target box that matches the first description data of the instance noun is obtained from the sample image, and the first target box information is obtained. The first target box information characterizes the position information of the first target box in the image and the instance noun corresponding to the first target box. Among them, the open-set object detection model is a trained model, which has the ability to detect the object described by the instance noun based on the image.

[0135] Since the open-set object detection model uses the instance noun for object detection, resulting in low accuracy and many false recalls, therefore, the instance noun, its first target box information, and the second description data of the instance noun are input into the second referential expression matching model. Through the inference of the second referential expression matching model, the second target box that matches the second description data of the instance noun in the image in the first target box is obtained, and the second target box information is obtained, thereby removing the incorrect target boxes in the target boxes obtained by the open-set object detection model. Therefore, the second target box is a subset of the first target box. Among them, the second target box information characterizes the position information of the second target box in the image and the second description data of the instance noun corresponding to the second target box; the second referential expression matching model is a trained model, which has the ability to detect whether the content of the image in the target box area matches the content described by the referential expression. During the matching process, the second description data is the referential expression.

[0136] Step 204: Integrate the first description data of the instance noun of the text instance in the sample image, the second description data of the instance noun, its second target box information, and the image detailed description data to obtain the detailed localization description pseudo-label data.

[0137] As an example, the position information of the second target box in the image in the second target box information obtained in step 203 is added to the image detailed description data as the target box description data, and the corresponding second target box information is added to the instance noun in the image detailed description data.

[0138] Due to the reasons of model and tool capabilities, during the first-round generation of the detailed localization description data, the embodiments of the present application rely on manual review to correct some incorrect content in the detailed localization description pseudo-label data. For example, the description data, target box, and instance noun in the detailed localization description pseudo-label data are checked for correctness to obtain the detailed localization description data. If the tool performance is improved, manual review may not be required.

[0139] Steps 201 to 204 are repeatedly executed to obtain the detailed localization description data of each sample image in the sample image set as the first-round detailed localization description data of each sample image in the sample image set.

[0140] See Figure 5 as shown Figure 5 This is a schematic diagram of obtaining the first-round detailed location description data based on an artificial intelligence model in this embodiment. The sample image is input into the third multi-modal large model, and through the inference of the third multi-modal large model, image detailed description data is obtained to realize the generation of image detailed description data; the image detailed description data is input into the third language large model, and through the inference of the third language large model, the first description data of the instance nouns in the image detailed description data is obtained, and the paragraph where the first description data of the instance nouns is located is input into the fourth language large model, and through the inference of the fourth language large model, the second description data of the instance nouns is obtained to realize text instance extraction; the first description data of the instance nouns and the sample image are input into the open-set object detection model, and through the inference of the open-set object detection model, the instance nouns and the first target box corresponding to the instance nouns are obtained to realize the rough detection of the target. The instance nouns, the image where the first target box is located, and the second description data of the instance nouns are input into the second referential expression matching model, and through the inference of the second referential expression matching model, the second target box matching the second description data of the instance nouns is obtained to realize the fine detection of the target; the second target box, the description data of its instance nouns, and the image detailed data are fused to obtain the detailed location description pseudo-label data, and after filtering, the first-round detailed location description data is obtained to realize the generation of the first-round detailed location description data.

[0141] See Figure 6 As shown, Figure 6 is a schematic flowchart of the method for generating the natural language description data of the image in this embodiment and detecting the natural language description target. It includes:

[0142] Step 601, using the first-round detailed location description data, perform the first-round training on the first multi-modal large model to obtain an expert model for generating detailed location description data.

[0143] As an example, the first-round detailed location description data and the sample image from which the first-round detailed location description data is derived are used as sample data and input into the first multi-modal large model to train the first multi-modal large model until the first multi-modal large model reaches the expectation, and the first multi-modal large model after the first-round training is obtained. The first multi-modal large model after the first-round training is used as the currently trained expert model, and this expert model has the ability to convert the input image into detailed location description data with image detailed description data and perform location description on the text instances in the image detailed description data;

[0144] Step 602, using the first-round detailed location description data, perform the first-round training on the second multi-modal large model to obtain the first referential expression matching model.

[0145] As an example, the first-round detailed localization description data is input into the first large language model, and through the inference of the first large language model, the target box description data is obtained. Among them, the first large language model is a trained model, which has the ability to generate target box description data based on the detailed localization description data. The first large language model, the third large language model, and the fourth large language model can be the same model or different models, and this application does not limit this.

[0146] Based on the obtained target box description data, positive and negative samples are obtained. Among them, the positive sample is that the target box matches the second description data of the instance noun of the target box, and the negative sample is that the target box does not match the second description data of the instance noun of the target box.

[0147] The negative samples can be constructed in the following ways:

[0148] One way is to use the first large language model to modify the second description data of the instance noun. For example, modify "the girl wearing a black mask and leading a dog" to "the girl wearing a white mask and leading a dog", and the target box that matches the original second description data and the modified second description data can be used as a negative sample pair.

[0149] Another way is that in the same image, different second description data of the same category and the target boxes that do not correspond to them form negative sample pairs. For example, the target boxes of "the lady wearing a black hat on the left side of the image" and "the lady wearing a white shirt and green trousers and making a phone call" are negative sample pairs, and the target boxes of "the lady wearing a white shirt and green trousers and making a phone call" and "the lady wearing a black hat on the left side of the image" are negative sample pairs.

[0150] The obtained target box description data, positive and negative samples, and the sample image data from which the first-round detailed localization description data is derived are used as sample data to train the second multi-modal large model until prediction is achieved, and the trained second multi-modal large model is obtained. This trained second multi-modal large model is used as the currently trained first referential expression matching model.

[0151] The above steps 601 and 602 have no strict order and can be executed in parallel.

[0152] The first referential expression matching model and the second referential expression matching model can be the same model or different models, and this embodiment does not limit this.

[0153] Step 603, input the target image for which detailed localization description data is to be generated and / or the detailed description data of the target image for which detailed localization description data is to be generated into the currently trained expert model, and through the inference of the currently trained expert model, obtain the detailed localization description data as the natural language description data of the image.

[0154] In the case of inputting the target image, the obtained detailed localization description data includes: image detailed description data, text instance description data corresponding to the text instances in the image detailed description data, and target box description data corresponding to the text instances. This detailed localization description data is noisy detailed localization description data, that is, the detailed localization description data contains incorrect target box description data;

[0155] In the case of inputting the detailed description data of the target image, the obtained detailed localization description data includes: image detailed description data, and text instance description data corresponding to the text instances in the image detailed description data. In this way, more accurate image detailed description data and text instance description data can be obtained through the expert model.

[0156] Since the detailed localization description data includes text instance description data corresponding to the text instances, the text instance description data can be used as the target described in natural language. In this way, there is no need to preset the target described in natural language.

[0157] Step 604: Input the noisy detailed localization description data into the first large language model. Through the inference of the first large language model, obtain the target box description data in the noisy detailed localization description data and the second description data of the instance nouns of the target boxes.

[0158] Among them, the first large language model is a trained model, which has the ability to generate target box description data and the second description data of the instance nouns of the target boxes based on the detailed localization description data.

[0159] Step 605: Input the target box description data in the noisy detailed localization description data, the second description data of the instance nouns of the target boxes, and the image from which the noisy detailed localization description data is derived into the currently trained first referential expression matching model. Through the inference of the first referential expression matching model, filter out the incorrect target box description data in the noisy detailed localization description data to obtain the detailed localization description data of the target image. This detailed localization description data is denoised detailed localization description data. The target box description data in the denoised detailed localization description data is the description data of the target box that matches the second description data, and the target box in the target box description data is the candidate target.

[0160] During the inference process of the first referential expression matching model, perform matching detection on the image content within the target box in the image corresponding to the target box description data. Among them, the referential expression is the second description data of the instance noun corresponding to the target box, and the second description data of the instance noun can be used as the target described in natural language.

[0161] Through the above steps 603-605, accurate detailed location description data can be generated as the natural language description data of the image, and the detection of natural language description targets can also be performed.

[0162] Step 606: Use the detailed location description data of the target image and / or the detailed location description data of the detailed description data of the target image obtained in step 605 to perform this round of training on the currently trained expert model and the currently trained first reference expression matching model, so as to update the model.

[0163] As an example, use the detailed location description data of the target image and / or the detailed location description data of the detailed description data of the target image obtained in step 605 as the sample data for this round, and perform training on the currently trained expert model and the currently trained first reference expression matching model to update the currently trained expert model and the currently trained first reference expression matching model, and obtain the next round of trained expert model and the next round of trained first reference expression matching model.

[0164] Among them, the training method is the same as that in steps 601 and 602, that is:

[0165] Use the detailed location description data obtained in step 605 and the image from which it is derived as the sample data for this round, and input it into the first multi-modal large model to train the first multi-modal large model until the first multi-modal large model meets the expectations, obtain the first multi-modal large model after this round of training, and use the multi-modal large model after this round of training as the currently trained expert model.

[0166] In the iterative training process of this embodiment, the image is not directly input into the expert model to obtain the detailed location description data, but into the first multi-modal large model to ensure the diversity of the detailed description, avoid the style of the detailed location description data being too unified, and reduce the diversity of expressions.

[0167] Use the detailed location description data obtained in step 605 as the sample data for this round, input it into the first large language model, through the inference of the first large language model, obtain the target box description data, based on the obtained target box description data, obtain positive and negative sample data, and use the obtained target box description data, positive and negative sample data, and the sample image data from which the detailed location description data obtained in step 605 is derived as the sample data for this round to train the second multi-modal large model until the prediction is achieved, obtain the second multi-modal large model after this round of training, and use the second multi-modal large model after this round of training as the currently trained first reference expression matching model.

[0168] Return to execute step 603 to generate the target image and / or the detailed description data of the target image in the next round, and update the sample data in the way of a data flywheel, and continuously iteratively update the model.

[0169] See Figure 7 as shown Figure 7 This is a schematic diagram for generating natural language description data of the image in this embodiment. The first multi-modal large model is trained as an expert model, and the second multi-modal large model is trained as the first referential expression matching model. Among them, in the first round of training, the first-round detailed localization description data is used as sample data. In subsequent iterative training, the final result of the inference, that is, the denoised detailed localization description data, is used as sample data. In the inference process, the target image is input into the currently trained expert model, and through the inference of the currently trained expert model, detailed localization description data is generated to obtain noisy detailed localization description data; the noisy detailed localization description data is input into the first language large model, and through the inference of the first language large model, target box description data matching the second description data of the instance noun and the second description data of the instance noun are obtained; the target box description data matching the second description data of the instance noun, the second description data of the instance noun, and the input target image are input into the currently trained first referential expression matching model, and through the inference of the currently trained first referential expression matching model, denoised detailed localization description data is obtained.

[0170] To facilitate understanding of the iterative process of this application, the following uses an example to illustrate.

[0171] In the first round of the inference process, image description data is input into the currently trained expert model:

[0172] This photo was taken at yyyy-mm-dd hh:ff:ss, and the location is pppp. The picture shows a part of the city street. There are shops and facilities on the left side of the picture, and parked cars on the right side, including a silver sedan with license plate number AAA. On the left side of the picture, there is a middle-aged woman wearing a gray coat and black tight pants and black shoes. She is leading a dog with tan hair...

[0173] Through the inference of the currently trained expert model, instance nouns can be extracted to obtain text description data containing description data of the instance nouns:

[0174] This photo was taken at Time yyyy - mm - dd hh:ff:ss , Location is pppp . The picture shows a part of the city street. There are shops and facilities on the left side of the picture, and parked Automobile , including a License Plate Number for AAA of Silver Sedan . On the left side of the picture, there is a Middle - aged Woman , wearing Grey Coat and Black Tight Trousers , and wearing Black Shoes on her feet. She is leading aBrown Brown Hair of Dog …

[0175] In the second round of the reasoning process, the image (as shown in Figure 8 ) and the text description data obtained in the previous round are input into the currently trained expert model. Through the reasoning of the currently trained expert model, the targets matching the instance nouns are detected, and detailed localization description data including the position information of each target box of each instance noun is obtained:

[0176] This photo was taken in Time yyyy - mm - dd hh:ff:ss [Target Bounding Box Position Information] , Location is pppp [Target Bounding Box Position Information] . The picture shows a part of the city street. There are shops and facilities on the left side of the picture, and parked cars on the right side, including a License Plate Number for AAA Target Bounding Box Position Information] of Silver Sedan Target Bounding Box Position Information] …

[0177] As can be seen from the example, the detailed localization description data includes: image detailed description data and target box description data of text instances. Among them, the target box description data is used to describe the position information of the target box corresponding to the text instance in the image in the image detailed description data, and multiple target boxes correspond to the text instance.

[0178] In the third round of the reasoning process, the image (as shown in Figure 8 ) and its image description data are input into the currently trained expert model. Through the reasoning of the currently trained expert model, the image description data is expanded into detailed localization description data, and detailed localization description data including the position information of each target box of each instance noun is obtained. This detailed localization description data is the same as the detailed localization description data obtained in the second round.

[0179] In the fourth round, the image (as shown in Figure 8 ) is input into the currently trained expert model. Through the reasoning of the currently trained expert model, detailed localization description data is generated, and detailed localization description data including the position information of each target box of each instance noun is obtained. This detailed localization description data is the same as the detailed localization description data obtained in the second round.

[0180] See Figure 9 shown, Figure 9 This embodiment is based on Figure 8 ​​A schematic diagram of the target box corresponding to the detailed positioning data. Among them, the green target is the image content including time information formed by OCR, the lake blue target is the image content including address information formed by OCR, and the instances composed of instance nouns correspond to multiple target boxes. For example, the instance composed of the instance nouns "license plate number", "AAA", and "silver car" corresponds to 2 target boxes, and the instance composed of the instance nouns "middle-aged woman", "gray coat", "black tight pants", "black shoes", and "dog" corresponds to 5 target boxes. This embodiment can enable the multimodal large model to output a detailed description of the image and match N targets for any complex description in the detailed description of the image, where N is a natural number greater than or equal to 0.

[0181] See Figure 10 as shown Figure 10 This is a schematic diagram of an apparatus for describing a target in natural language based on image detection according to an embodiment of the present application. The apparatus includes:

[0182] A detailed positioning description data generation module, configured to input the image to be detected into a trained expert model for converting the input image into detailed positioning description data with image detailed description data and performing positioning description on the text instances in the image detailed description data, and obtain the detailed positioning description data of the image to be detected through the inference of the expert model.

[0183] A detection module, configured to use the detailed positioning description data of the image to be detected to obtain candidate targets in the image to be detected that match the natural language description target represented by the text instance.

[0184] As an example, the detection module includes:

[0185] An instance acquisition sub-module, configured to obtain the position information of the image instance located by the image instance description data in the image and the text instances in the image detailed description data based on the detailed positioning description data of the image to be detected.

[0186] A matching sub-module, determines the image content of the image instance according to the position information of the image instance in the image, detects whether the determined image content matches the natural language description target, and uses the image instance that matches the natural language description target as a candidate target.

[0187] See Figure 11 as shown Figure 11 This is a schematic diagram of an apparatus for generating natural language description data of an image according to an embodiment of the present application. The apparatus includes:

[0188] The detailed location description data generation module inputs the target image for which natural language description data is to be generated and / or the detailed description data of the target image for which natural language description data is to be generated into the trained expert model. Through the inference of the expert model, noisy detailed location description data is obtained as the natural language description data of the image.

[0189] As an example, the device further includes:

[0190] The instance acquisition module is used to obtain the position information of the image instance located by the image instance description data and the text instance in the image detailed description data based on the generated detailed location description data.

[0191] The matching module is used to determine the image content of the image instance according to the position information of the image instance in the image, detect whether the determined image content matches the natural language description target, so as to remove the image instance description data of the image instance that does not match the natural language description target, and obtain the denoised detailed location description data as the natural language description data of the image.

[0192] See Figure 12 as shown Figure 12 This is another schematic diagram of the device for detecting the natural language description target based on an image and / or the device for generating the natural language description data of an image according to an embodiment of the present application. The device includes a memory and a processor. The memory stores a computer program, and the processor is configured to execute the computer program to implement the steps of any one of the methods for detecting the natural language description target based on an image and / or the steps of the method for generating the natural language description data of an image.

[0193] The memory may include a random access memory (RAM), or may also include a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located far from the aforementioned processor.

[0194] The above-mentioned processor may be a general-purpose processor, including a Central Processing Unit (CPU), a Network Processor (NP), etc.; it may also be a Digital Signal Processor (DSP), an Application Specific Integrated Circuit (ASIC), a Field-Programmable Gate Array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0195] An embodiment of the present invention further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of any one of the above-mentioned methods for naturally describing a target based on image detection and / or the steps of the method for generating natural language description data of an image are implemented.

[0196] For the embodiments of the device / network-side device / storage medium, since they are basically similar to the method embodiments, the description is relatively simple. For the relevant parts, please refer to the partial description of the method embodiments.

[0197] In this article, relational terms such as first and second are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any actual relationship or order between these entities or operations. Moreover, the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements not only includes those elements, but also includes other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "including one..." does not exclude the existence of additional identical elements in the process, method, article or device including the element.

[0198] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the scope of protection of the present invention.

Claims

1. A method for describing an object in natural language based on image detection, characterized in that, The method includes: Input the image to be detected into a pre-trained expert model for converting the input image into detailed localization description data with image detailed description data and localizing and describing text instances in the image detailed description data. Through the inference of the expert model, obtain the detailed localization description data of the image to be detected. The detailed localization description data of the image to be detected includes: the image detailed description data of the image to be detected, the image instance description data corresponding to the text instances in the image detailed description data, and the text instance description data corresponding to the text instances in the image detailed description data. The image instance description data is used to characterize the position information of the image instance corresponding to the text instance in the image. The image instance description data includes: the description data of the target box for characterizing the position of the image instance in the image. The text instance is used to characterize the natural language description target referred to in the image detailed description data. The text instance description data includes: the description data of the instance noun for referring to the natural language description target. Utilize the detailed localization description data of the image to be detected to obtain candidate targets in the image to be detected that match the natural language description target characterized by the text instance. Wherein, The step of utilizing the detailed localization description data of the image to be detected to obtain candidate targets in the image to be detected that match the natural language description target characterized by the text instance includes: Input the detailed localization description data of the image to be detected into the first large language model. Through the inference of the first large language model, obtain the target box description data in the detailed localization description data and the description data of the instance noun in the detailed localization description data. Input the image to be detected, the target box description data in the detailed localization description data, and the description data of the instance noun in the detailed localization description data into a pre-trained first referential expression matching model for matching the content information within the target box area with the referential expression for characterizing the description data of the instance noun. Through the inference of the first referential expression matching model, filter out the incorrect target box description data in the detailed localization description data and obtain the denoised detailed localization description data of the image to be detected.

2. The method according to claim 1, characterized in that, The description data of the instance noun includes: the first description data of the instance noun and / or the second description data of the instance noun. Wherein, The first description data is used to characterize the simple description of the instance noun. The second description data is used to characterize the complex description of the instance noun. The simple description of the instance noun includes the instance category name and / or the description of the content of interest. The complex description of the instance noun includes: multiple category names and / or multiple phrasal arbitrary natural language descriptions other than the single category name description data.

3. The method according to claim 1, characterized in that, The expert model is trained in the following manner: Obtain the first-round detailed localization description data. Use the first-round detailed localization description data as sample data to train the first multi-modal large model until the first multi-modal large model reaches the expectation, and obtain the trained first multi-modal large model. Use the trained first multi-modal large model as the first-round pre-trained expert model.

4. The method according to claim 3, characterized in that, The first referential expression matching model is trained in the following manner: Input the first-round detailed localization description data into the second large language model. Through the inference of the second large language model, obtain the target box description data in the first-round detailed localization description data and the description data of instance nouns. Based on the target box description data and the description data of instance nouns in the first-round detailed localization description data, obtain positive and negative sample data. Use the target box description data and the description data of instance nouns in the first-round detailed localization description data, the image data from which the first-round detailed localization description data is sourced, and the positive and negative sample data as sample data to train the second multimodal large model until the second multimodal large model meets the expectations, and obtain the trained second multimodal large model. Use the trained second multimodal large model as the first reference expression matching model that has been trained in the first round.

5. The method according to claim 4, wherein The training of the expert model further includes: Use the denoised detailed localization description data as the sample data for this round to train the previous-round first multimodal large model until the previous-round first multimodal large model meets the expectations, and obtain the first multimodal large model trained in this round. Use the first multimodal large model trained in this round as the currently trained expert model to detect the next image to be detected.

6. The method according to claim 4, characterized in that The training of the first reference expression matching model further includes: Input the denoised detailed localization description data into the second large language model. Through the inference of the second large language model, obtain the target box description data in the denoised detailed localization description data and the description data of instance nouns. Based on the target box description data and the description data of instance nouns in the denoised detailed localization description data, obtain positive and negative sample data. Use the target box description data and the description data of instance nouns in the denoised detailed localization description data, the image data from which the denoised detailed localization description data is sourced, and the positive and negative sample data as the sample data for this round to train the previous-round second multimodal large model until the previous-round second multimodal large model meets the expectations, and obtain the second multimodal large model trained in this round. Use the second multimodal large model trained in this round as the currently trained first reference expression matching model to detect the next image to be detected.

7. The method according to claim 6, wherein The first-round detailed localization description data is obtained in the following manner: Obtain the image detailed description data of the sample image. Based on the image detailed description data of the sample image, perform text instance extraction to obtain the first description data and the second description data of the instance nouns included in the text instances in the sample image. Perform target detection based on the sample image data to obtain the first target box that matches the first description data. Match the sample image content of the first target box with the second description data to obtain the second target box that matches the second description data. Fuse the second target box, the first description data of the instance nouns, the second description data, and the image detailed description data of the sample image to obtain the detailed localization description pseudo-label data, which includes: image detailed description data, the first description data of the instance nouns of the second target box, the second description data, and the pseudo-label target box description data of the second target box.

8. The method according to claim 7, characterized in that The image detailed description data for obtaining the sample image includes: Inputting the sample image data into the third multi-modal large model, and through the inference of the third multi-modal large model, obtaining the image detailed description data of the sample image; Performing text instance extraction based on the image detailed description data of the sample image, including: Inputting the image detailed description data of the sample image into the third language large model, and through the inference of the third language large model, obtaining the first description data of the instance nouns of the text instances in the sample image; Inputting the text where the first description data of the instance nouns of the text instances in the sample image is located into the fourth language large model, and through the inference of the fourth language large model, obtaining the second description data of the instance nouns of the text instances in the sample image; Performing object detection based on the sample image data to obtain a first target box that matches the first description data, including: Inputting the sample image and the first description data of its instance nouns into the open-set object detection model, and through the inference of the open-set object detection model, obtaining the position information of the image where the first target box is located and the instance nouns corresponding to the first target box; Matching the sample image content of the first target box with the second description data, including: Inputting the sample image, the position information of the image where the first target box is located, the instance noun data corresponding to the first target box, and the second description data into the second referential expression matching model, and through the inference of the second referential expression matching model, obtaining the position information of the image where the second target box that matches the second description data is located; After fusing the second target box, the first description data of the instance nouns, the second description data, and the image detailed description data of the sample image, it further includes: Removing the incorrect data in the detailed localization description pseudo-label data to obtain the first-round detailed localization description data.

9. The method according to claim 8, characterized in that, Inputting the image detailed description data of the sample image into the third language large model, further including: Inputting the image detailed description data of the sample image and the first text prompt for prompting the instance nouns in the image detailed description data of the sample image into the third language large model to determine the positions of the instance nouns in the text in the image detailed description data of the sample image; Inputting the text where the first description data of the instance nouns of the text instances in the sample image is located into the fourth language large model, further including: Inputting the paragraph text where the instance nouns are located and the second text prompt for prompting the complex descriptions in the image detailed description data of the sample image into the fourth language large model to determine the complex descriptions of the instance nouns in the image detailed description data of the sample image.

10. An electronic device, characterized in that, Including a memory and a processor, the memory stores a computer program, and the processor is configured to execute the computer program to implement the steps of a method for detecting a natural language description target based on an image according to any one of claims 1 to 9.

Citation Information

Patent Citations

  • Image description processing method, computer equipment and storage medium

    CN116486188A