Method for describing target by natural language based on image detection and electronic equipment

By generating detailed positioning description data, the problem of low accuracy of natural language description object detection in the prior art is solved, and effective detection of complex descriptions and more accurate generation of image natural language description data is achieved.

CN120032149AActive Publication Date: 2025-05-23HANGZHOU HIKVISION DIGITAL TECHNOLOGY CO LTD
View PDF 8 Cites 0 Cited by

Patent Information

Application Number
CN202510469031.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-23
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The prior art has low accuracy when detecting targets referred to by natural language descriptions, especially when dealing with complex descriptions represented by multi-category names and multi-references, and cannot effectively detect targets.

Method used

By inputting the image to be detected into the trained expert model, detailed positioning description data is generated, which includes the image detailed description data and location information of the text instance, and the candidate target matching the natural language description target is obtained.

Benefits of technology

Improves the accuracy of natural language description object detection, can process complex natural language descriptions, and generate more accurate image natural language description data.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032149A_ABST
    Figure CN120032149A_ABST
Patent Text Reader

Abstract

The invention discloses a natural language description target method based on image detection, and the method comprises the steps: inputting a to-be-detected image into a trained expert model of detailed positioning description data which is used for converting the input image into detailed image description data and carrying out the positioning description of a text instance in the detailed image description data, detailed positioning description data are obtained through reasoning of the expert model, the detailed positioning description data comprise image detailed description data and image instance description data corresponding to text instances in the image detailed description data, and the detailed positioning description data of the to-be-detected image are utilized to detect the to-be-detected image according to the detailed positioning description data of the to-be-detected image. And obtaining a candidate target matched with the natural language description target represented by the text instance in the to-be-detected image. The detection accuracy of the target described by the natural language can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence, and in particular, to a method for detecting a natural language description target based on an image. Background Art

[0002] Detecting objects described by natural language in images has always been an important research area. Currently, the technologies for detecting objects described by natural language include: Open Vocabulary Object Detection (OVD) and Referring Expression Comprehension (REC). Open set object detection can detect objects described by any category name. Category names are usually short category names without attributes or relationship descriptions. For example, cats, dogs, trees, cars and other instances in an image are described by category names, that is, cats, dogs, trees, and cars are category names. Referential expression understanding can understand referential expressions and detect and enforce the detection of a bounding box target. Referential expressions are words or phrases used to identify a specific person, place, or thing. They are usually a noun, noun phrase, or pronoun, also known as phrased natural language descriptions. For example, a man with a backpack is an isolated phrased description.

[0003] Although open-set target detection and referential expression understanding can detect targets referred to by simple natural language descriptions, users are required to customize the natural language descriptions of the targets to be detected. Due to the different natural language description capabilities of different users, the detection accuracy is low. Once at least one of the multi-category names and multi-referential expressions is used to describe the target, it will be impossible to detect the target referred to by at least one of the multi-category names and multi-referential expressions, which makes them unable to cope with more open application scenarios.

[0004] For example, when detecting the man wearing a white mask in "There is a lady in red clothes and blue pants on the left side of the image, and behind her there is a man wearing a white mask." Current detection methods are powerless for this kind of natural language description that requires long context understanding. Summary of the invention

[0005] The present invention provides a method for detecting a target described in natural language based on an image, so as to improve the accuracy of detecting the target described in natural language.

[0006] The present invention provides a method for detecting a natural language description target based on an image, the method comprising: The image to be detected is input into an expert model that has been trained to convert the input image into detailed positioning description data having detailed image description data and positioning description of text instances in the detailed image description data, and the detailed positioning description data of the image to be detected is obtained through reasoning of the expert model. in, The detailed positioning description data of the image to be detected includes: image detailed description data of the image to be detected, and image instance description data corresponding to the text instance in the image detailed description data. The image instance description data is used to represent the position information of the image instance corresponding to the text instance in the image. The text instance is used to represent the natural language description target referred to in the image detailed description data, By using the detailed positioning description data of the image to be detected, candidate targets in the image to be detected that match the natural language description targets represented by the text instance are obtained.

[0007] Preferably, the method of using the detailed positioning description data of the image to be detected to obtain a candidate target in the image to be detected that matches the natural language description target represented by the text instance includes: Based on the detailed positioning description data of the image to be detected, the position information of the image instance located by the image instance description data in the image and the text instance in the detailed image description data are obtained, Determine the image content of the image instance according to the position information of the image instance in the image, detect whether the determined image content matches the natural language description target, and use the image instance matching the natural language description target as a candidate target; The detailed positioning description data of the image to be detected also includes: text instance description data corresponding to the text instance in the image detailed description data.

[0008] Preferably, the text instance description data includes: description data for referring to an instance noun of a natural language description target, The image instance description data includes: description data of a target frame for representing an image position where the image instance is located, The step of obtaining the position information of the image instance located by the image instance description data in the image and the text instance in the image detailed description data based on the detailed positioning description data of the image to be detected includes: Inputting the detailed positioning description data of the image to be detected into the first language model, and obtaining the target frame description data in the detailed positioning description data and the description data of the instance noun in the detailed positioning description data through the reasoning of the first language model; The step of determining the image content of the image instance according to the position information of the image instance in the image, and detecting whether the determined image content matches the natural language description target, includes: The image to be detected, the target frame description data in the detailed positioning description data, and the description data of the instance nouns in the detailed positioning description data are input into a first referential expression matching model that has been trained to match the content information in the target frame area with the referential expression of the description data used to characterize the instance nouns. Through the reasoning of the first referential expression matching model, the erroneous target frame description data in the detailed positioning description data is filtered out to obtain the denoised detailed positioning description data of the image to be detected.

[0009] Preferably, the description data of the instance noun includes: first description data of the instance noun and / or second description data of the instance noun, wherein the first description data is used to represent a simple description of the instance noun, and the second description data is used to represent a complex description of the instance noun, the simple description of the instance noun includes the instance class name and / or the description of the content of interest, and the complex description of the instance noun includes: multiple class names and / or multiple phrased arbitrary natural language descriptions other than the single class name description data; The expert model is trained in the following way: Get the first round of detailed positioning description data, The first round of detailed positioning description data is used as sample data to train the first multimodal large model until the first multimodal large model meets expectations, thereby obtaining the trained first multimodal large model. The trained first multimodal large model is used as the first round of trained expert model.

[0010] Preferably, the first index expression matching model is trained in the following manner: The first round of detailed positioning description data is input into the second largest language model. Through the reasoning of the second largest language model, the target box description data and the description data of the instance noun in the first round of detailed positioning description data are obtained. Based on the target box description data and instance noun description data in the first round of detailed positioning description data, positive and negative sample data are obtained. The target frame description data and the description data of the instance noun in the first round of detailed positioning description data, as well as the image data and the positive and negative sample data from which the first round of detailed positioning description data originate are used as sample data to train the second multimodal large model until the second multimodal large model meets expectations, thereby obtaining a trained second multimodal large model. The trained second multimodal large model is used as the first index expression matching model trained in the first round.

[0011] Preferably, the training of the expert model further includes: The denoised detailed positioning description data is used as the sample data of this round, and the first multimodal large model of the previous round is trained until the first multimodal large model of the previous round meets the expectations, and the first multimodal large model after this round of training is obtained. The first multimodal large model after this round of training is used as the currently trained expert model to detect the next image to be detected; The training of the first index expression matching model further includes: The denoised detailed positioning description data is input into the second largest language model, and the target box description data and the description data of the instance noun in the denoised detailed positioning description data are obtained through the reasoning of the second largest language model. Based on the target box description data in the denoised detailed positioning description data and the description data of the instance noun, positive and negative sample data are obtained. The target frame description data and the description data of the instance nouns in the denoised detailed positioning description data, as well as the image data and the positive and negative sample data from which the denoised detailed positioning description data comes, are used as sample data for this round, and the second multimodal large model of the previous round is trained until the second multimodal large model of the previous round meets expectations, thereby obtaining the second multimodal large model after this round of training. The second multimodal large model after this round of training is used as the currently trained first index expression matching model to detect the next image to be detected.

[0012] Preferably, the first round of detailed positioning description data is obtained in the following manner: Get the image detailed description data of the sample image, Based on the detailed image description data of the sample image, text instance extraction is performed to obtain first description data and second description data of the instance noun included in the text instance in the sample image. Performing target detection based on the sample image data to obtain a first target frame matching the first description data, Matching the sample image content of the first target frame with the second description data to obtain a second target frame matching the second description data, The second target frame, the first description data of the instance noun, the second description data, and the detailed image description data of the sample image are fused to obtain detailed positioning description pseudo-label data, which includes: detailed image description data, the first description data of the instance noun of the second target frame, the second description data, and the pseudo-label target frame description data of the second target frame.

[0013] Preferably, the acquiring of detailed image description data of the sample image comprises: Inputting the sample image data into the third multimodal large model, and obtaining detailed image description data of the sample image through reasoning of the third multimodal large model; The text instance extraction based on the detailed image description data of the sample image includes: The detailed image description data of the sample image is input into the third language model, and the first description data of the instance noun of the text instance in the sample image is obtained through the reasoning of the third language model. Inputting the text containing the first description data of the instance noun of the text instance in the sample image into the fourth language macro model, and obtaining the second description data of the instance noun of the text instance in the sample image through the reasoning of the fourth language macro model; The performing target detection based on the sample image data to obtain a first target frame matching the first description data includes: Inputting the first description data of the sample image and its instance noun into the open set object detection model, and obtaining the position information of the image where the first object frame is located and the instance noun corresponding to the first object frame through the reasoning of the open set object detection model; The matching of the sample image content of the first target frame with the second description data includes: Inputting the sample image, the position information of the image where the first target frame is located, the instance noun data corresponding to the first target frame, and the second description data into the second referential expression matching model, and obtaining the position information of the image where the second target frame is located that matches the second description data through the reasoning of the second referential expression matching model; After fusing the second target frame, the first description data of the instance noun, the second description data, and the detailed image description data of the sample image, the method further includes: Remove the erroneous data in the detailed positioning description pseudo-label data to obtain the first round of detailed positioning description data.

[0014] Preferably, the step of inputting the detailed image description data of the sample image into the third language macro model further comprises: Inputting the detailed image description data of the sample image and a first text prompt for prompting an instance noun in the detailed image description data of the sample image into the third language macro model to determine the position of the instance noun in the detailed image description data of the sample image in the text; The step of inputting the text containing the first description data of the instance noun of the text instance in the sample image into the fourth language model further comprises: The paragraph text where the instance noun is located and the second text prompt for prompting the complex description in the detailed image description data of the sample image are input into the fourth language macro model to determine the complex description of the instance noun in the detailed image description data of the sample image.

[0015] A second aspect of the present invention provides a method for generating natural language description data of an image, the generating method comprising: Input the target image to be generated with natural language description data and / or the target image detailed description data to be generated with natural language description data into the trained expert model, and obtain the detailed positioning description data through the reasoning of the expert model. in, The expert model is used to convert the input image into detailed positioning description data having image detailed description data and performing image positioning description on text instances in the image detailed description data, and / or, detailed positioning description data performing text positioning description on text instances in the input image detailed description data, The detailed positioning description data includes: image detailed description data, and image instance description data corresponding to the text instance in the image detailed description data, and / or the detailed positioning description data includes: image detailed description data, and text instance description data corresponding to the text instance in the image detailed description data, The image instance description data is used to represent the position information of the image instance corresponding to the text instance in the image. The text instance is used to represent the natural language description target referred to in the image detailed description data.

[0016] A third aspect of the present application provides an electronic device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor is configured to execute the computer program to implement the steps of a method for describing a target in natural language based on image detection, and / or the steps of a method for generating natural language description data of an image.

[0017] The method for detecting natural language description targets based on images provided in the present application generates detailed positioning description data of the image to be detected to obtain candidate targets in the image to be detected that match the natural language description targets represented by the text instances. The present application does not require presetting the natural language description targets, and can detect the targets referred to by the text instances of arbitrarily complex descriptions in the detailed positioning description data, which is beneficial to improving the accuracy of natural language description target detection. Moreover, the present application is beneficial to improving the intelligence of the generation of natural language description data of the image through the generated detailed positioning description data, so that the information contained in the image can be described more accurately. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] Figure 1 A flowchart of a method for detecting a natural language description target based on an image in an embodiment of the present application.

[0019] Figure 2A schematic diagram of a process for obtaining a first round of detailed positioning description data for training a first multimodal large model in this embodiment.

[0020] Figure 3 A schematic diagram of extracting first description data of an instance noun in this embodiment.

[0021] Figure 4 A schematic diagram of extracting the second description data of the instance noun in this embodiment.

[0022] Figure 5 This is a schematic diagram of obtaining the first round of detailed positioning description data based on the artificial intelligence model in this embodiment.

[0023] Figure 6 A flowchart of the method for generating natural language description data of an image and detecting a natural language description target is shown in the figure.

[0024] Figure 7 A schematic diagram of generating natural language description data of an image in this embodiment.

[0025] Figure 8 FIG. 4 is a schematic diagram of a frame of image in this embodiment.

[0026] Fig. 9 This embodiment is based on Figure 8 A schematic diagram of the target box corresponding to the detailed positioning data.

[0027] Fig.10 A schematic diagram of a flow chart of an apparatus for detecting a target described in natural language based on an image according to an embodiment of the present application.

[0028] Fig.11 A schematic diagram of a flow chart of a device for generating a natural language description of an image according to an embodiment of the present application.

[0029] Fig.12 Another schematic diagram of an apparatus for detecting a natural language description target based on an image and / or an apparatus for generating natural language description data of an image according to an embodiment of the present application. DETAILED DESCRIPTION

[0030] In order to make the objectives, technical means and advantages of the present application more clearly understood, the present application is further described in detail below in conjunction with the accompanying drawings.

[0031] To facilitate understanding of the embodiments of the present application, the technical terms in the embodiments of the present application are explained below.

[0032] Dense caption: A natural language description of the local details in an image.

[0033] Complex description: A natural language description of any combination of multiple phrases and / or multiple category names other than a single category name, including but not limited to natural language descriptions that require a long context to understand, and natural language descriptions that have at least one of certain logic, deductive nature, and inductive nature.

[0034] Detailed positioning description data: includes image detailed description data, and description data for positioning description of text instances in the image detailed description data, wherein the positioning description includes: an image positioning description for describing the image position in the image, a text positioning description for describing the text position in the image detailed description data, or a combination thereof.

[0035] Text instance: The instance described by the text instance description data in the image detailed description data. Text instance description data: a collection of description data of instance nouns used to describe text instances in image detailed description data, wherein the description data of instance nouns include: first description data used to simply describe instance nouns, second description data used to complexly describe instance nouns, one of them or a combination thereof. In this application, a simple description of an instance noun can be understood as a description of an instance noun referred to by a category name with an attribute description, for example, a single phrased instance noun, and a complex description of an instance noun can be understood as an instance noun referred to in a complex description manner, for example, a context formed by a combination of multiple different phrased instance nouns. Text instance description data can be used for text positioning description.

[0036] Image instance: A physical object in an image.

[0037] Image instance description data: description data that describes the image region where the image instance is located in the detailed image description data in the form of image position information (such as pixel coordinates); the image instance description data can be used for image positioning description; the image region where the image instance is located is usually represented by a target box, so the image instance description data includes the target box description data.

[0038] Referring Expression Matching (REM): Detects whether the image content in the target box area matches the content described by the referring expression.

[0039] The method for detecting a target described in natural language based on an image provided in an embodiment of the present application generates, through the inference of a trained multimodal large model, detailed location description data having image detailed description data for an image to be detected that is input into the trained multimodal large model, and location description data for each text instance in the image detailed description data, and uses the detailed location description data to realize the detection of the target described in natural language.

[0040] See also Figure 1 As stated, Figure 1 A flowchart of a method for detecting a natural language description target based on an image in an embodiment of the present application is provided. The method includes: Step 101: input the image to be detected into a trained expert model for converting the input image into detailed location description data having detailed image description data and performing location description on text instances in the detailed image description data, and obtain detailed location description data of the image to be detected through reasoning of the expert model. in, The detailed positioning description data of the image to be detected includes: image detailed description data of the image to be detected, and image instance description data corresponding to the text instance in the image detailed description data. The image instance description data is used to represent the location information of the image instance corresponding to the text instance in the image. The text instance is used to represent the natural language description target referred to in the image detailed description data. As an example, The expert model is trained as follows: Get the first round of detailed positioning description data, The first round of detailed positioning description data is used as sample data to train the first multimodal large model until the first multimodal large model meets expectations, thereby obtaining the trained first multimodal large model. The trained first multimodal large model is used as the first round of trained expert model.

[0041] Furthermore, the denoised detailed positioning description data is used as the sample data of this round, and the first multimodal large model of the previous round is trained until the first multimodal large model of the previous round meets expectations, so as to obtain the first multimodal large model after this round of training, and the first multimodal large model after this round of training is used as the currently trained expert model for detecting the next image to be detected.

[0042] As an example, the detailed positioning description data of the image to be detected further includes: text instance description data corresponding to the text instance in the image detailed description data.

[0043] As an example, the text instance description data includes: description data of an instance noun used to refer to a natural language description target, and the image instance description data includes: description data of a target box used to characterize an image position where the image instance is located.

[0044] Step 102, using the detailed positioning description data of the image to be detected, obtain candidate targets in the image to be detected that match the natural language description targets represented by the text instance.

[0045] As an example, Based on the detailed positioning description data of the image to be detected, the position information of the image instance located by the image instance description data in the image and the text instance in the detailed image description data are obtained, The image content of the image instance is determined according to the position information of the image instance in the image, and it is detected whether the determined image content matches the natural language description target, and the image instance matching the natural language description target is taken as a candidate target.

[0046] For example, the detailed location description data of the image to be detected is input into the first largest language model, and through the reasoning of the first largest language model, the target frame description data in the detailed location description data and the description data of the instance nouns in the detailed location description data are obtained; the image to be detected, the target frame description data in the detailed location description data, and the description data of the instance nouns in the detailed location description data are input into the first referential expression matching model that has been trained to match the content information in the target frame area with the referential expression of the description data used to characterize the instance nouns, and through the reasoning of the first referential expression matching model, the erroneous target frame description data in the detailed location description data is filtered out to obtain the denoised detailed location description data of the image to be detected, and the denoised detailed location description data includes candidate target frame description data.

[0047] The method for detecting natural language description targets based on images provided in the embodiments of the present application can obtain candidate targets in the image that match the natural language description targets represented by text instances based on the detailed location description data by generating detailed location description data of the image. The embodiments of the present application do not require a preset natural language description of the target to be detected, and can detect targets with arbitrarily complex descriptions in the detailed location description data.

[0048] To facilitate understanding of the embodiments of the present application, the following is an explanation of the implementation of an artificial intelligence model as an example. It should be understood that the embodiments of the present application are not limited to a specific model.

[0049] See also Figure 2 As shown, Figure 2 A schematic diagram of a process for obtaining the first round of detailed positioning description data for training the first multimodal large model in this embodiment. Specifically, it includes: Step 201, based on the sample image data, obtaining detailed image description data of the sample image, As an example, a frame of sample image is input into the third multimodal large model, and through reasoning by the third multimodal large model, detailed image description data of the input sample image is obtained.

[0050] Among them, the third multimodal large model is a trained model, which has the ability to understand the input image and generate natural language description, that is, the third multimodal large model has the ability to generate text from images.

[0051] Step 202 : extracting text instances based on the generated detailed image description data to obtain description data of instance nouns of text instances in the detailed image description data of the sample image.

[0052] Among them, the instance noun is used to represent the category name of the text instance in the detailed description data of the image. The description data of the instance noun includes: first description data for representing a simple description of the instance noun, and second description data for representing a complex description of the instance noun. The simple description can be described in terms of the category name and / or content of interest of the instance, such as optical character recognition (OCR) content, relationships, attributes, etc., which is equivalent to the description of the instance name itself. The complex description is a natural language description of multiple phrases other than a single category name, and / or any combination of multiple category names, such as a context-understood description.

[0053] In this way, the first description data of the instance noun represents the description data of the instance noun and its related information, such as a red balloon, etc., to achieve a local detailed description of the text instance, and the second description data of the instance noun achieves the summary and refinement of the text instance.

[0054] See also Figure 3 As shown, Figure 3 A schematic diagram of extracting the first description data of the instance noun in this embodiment. The image detailed description data of the sample image and its corresponding first text prompt are input into the third language model, and the text position of the instance noun in the image detailed description data is determined through the reasoning of the third language model to obtain the first description data of the instance noun of the input image detailed description data. The first text prompt is used to prompt the instance noun in the image detailed description data, and the third language model is a trained model that has the ability to understand the input text and generate the first description of the instance noun.

[0055] For example, the detailed image description data of the sample image is: This photo was taken at time yyyy-mm-dd hh:ff:ss, location pppp. The picture shows a part of a city street, with shops and facilities on the left side of the picture and parked cars on the right side, including a silver sedan with a license plate number of AAA. On the left side of the picture, there is a middle-aged woman wearing a gray coat, black leggings, and black shoes. She is walking a tan-haired dog... The detailed image description data of the sample image is input into the third language model, and the output text of the first description data (underlined) marked with the instance noun is obtained as follows: This photo was taken on Time yyyy-mm-dd hh:ff:ss , Place yes ppppThe image shows a section of a city street, with shops and facilities on the left and parked cars on the right. car , including a License plate number for AAA of Silver sedan On the left side of the screen, there is a Middle-aged women , wearing Gray coat and Black leggings , foot wear Black shoes She was holding a Brown Brown hair of dog … See also Figure 4 As shown, Figure 4 A schematic diagram of extracting the second description data of the instance noun in this embodiment. The paragraph text where the instance noun is located and its second text prompt are input into the fourth language big model, and the fourth language big model is reasoned to determine the context text position of the instance noun in the detailed image description data of the sample image, that is, to determine the text positions of multiple phrased descriptions of the instance noun in the detailed image description data of the sample image, and obtain the second description data of the instance noun. The paragraph text where the instance noun is located can be marked, and the second text prompt is used to prompt the phrased description of the instance noun in the detailed image description data. The fourth language big model is a trained model that has the ability to understand the input text and generate the second description of the instance noun.

[0056] The third language macro model and the fourth language macro model may be the same language macro model or different language macro models, and this application does not impose any limitation on this.

[0057] For example, the description paragraph "On the left side of the picture, there is a middle-aged woman wearing a gray coat, black tights, and black shoes. She is holding a brown-haired dog" and the text prompt "female" in the detailed description data of the sample image are input into the fourth language large model, and the second description data of the instance noun "female" can be obtained: Middle-aged woman in gray coat Woman wearing black leggings and black shoes, walking with dog Step 203: perform target detection and referential expression matching based on the first description data and the second description data of the instance noun to obtain target frame information matching the second description data.

[0058] As an example, a sample image and the first description data of an instance noun are input into an open set target detection model, and through the reasoning of the open set target detection model, a first target frame matching the first description data of the instance noun is obtained from the sample image, and first target frame information is obtained, wherein the first target frame information represents the position information of the first target frame in the image, and the instance noun corresponding to the first target frame. The open set target detection model is a trained model that has the ability to detect targets described by instance nouns based on images.

[0059] In view of the fact that the open set target detection model uses instance nouns for target detection, resulting in low accuracy and many false recalls, the instance noun and its first target frame information, as well as the second description data of the instance noun are input into the second referential expression matching model, and the second target frame information is obtained from the image in the first target frame through the reasoning of the second referential expression matching model, thereby removing the erroneous target frame in the target frame obtained by the open set target detection model, and therefore, the second target frame is a subset of the first target frame. Among them, the second target frame information represents the position information of the second target frame in the image, and the second description data of the instance noun corresponding to the second target frame; the second referential expression matching model is a trained model, which has the ability to detect whether the image content of the target frame area matches the content described by the referential expression. In the matching process, the second description data is the referential expression.

[0060] Step 204, the first description data of the instance noun of the text instance in the sample image, the second description data of the instance noun and its second target frame information, and the image detailed description data are fused to obtain detailed positioning description pseudo-label data. As an example, the position information of the second target frame in the image obtained in step 203 is added to the image detailed description data as target frame description data, and the corresponding second target frame information is added to the instance noun in the image detailed description data.

[0061] Due to the model and tool capabilities, in the first round of detailed location description data generation process, the embodiment of the present application uses manual review to correct some erroneous content in the detailed location description pseudo-label data, for example, the description data, target box, and instance noun in the detailed location description pseudo-label data are checked for correctness to obtain the detailed location description data. If the tool performance is improved, manual review may not be required.

[0062] Steps 201 to 204 are repeatedly performed to obtain detailed positioning description data of each sample image in the sample image set as the first round of detailed positioning description data of each sample image in the sample image set.

[0063] See also Figure 5 As shown, Figure 5 This is a schematic diagram of obtaining the first round of detailed positioning description data based on the artificial intelligence model in this embodiment. The sample image is input into the third multimodal large model, and the image detailed description data is obtained through the reasoning of the third multimodal large model to realize the generation of the image detailed description data; the image detailed description data is input into the third language large model, and the first description data of the instance noun in the image detailed description data is obtained through the reasoning of the third language large model, and the paragraph where the first description data of the instance noun is located is input into the fourth language large model, and the second description data of the instance noun is obtained through the reasoning of the fourth language large model to realize text instance extraction; the first description data of the instance noun and the sample image are input into the open set target detection model, and the instance noun and the first target frame corresponding to the instance noun are obtained through the reasoning of the open set target detection model to realize the rough detection of the target, and the instance noun and the image where the first target frame is located, as well as the second description data of the instance noun are input into the second referential expression matching model, and the second target frame matching the second description data of the instance noun is obtained through the reasoning of the second referential expression matching model to realize the precise detection of the target; the description data of the second target frame and its instance noun and the image detailed data are fused to obtain the detailed positioning description pseudo-label data, and after filtering, the first round of detailed positioning description data is obtained to realize the generation of the first round of detailed positioning description data.

[0064] See also Figure 6 As shown, FIG6 is a flow chart of a method for generating natural language description data of an image and detecting a natural language description target in this embodiment. It includes: Step 601, using the first round of detailed positioning description data, a first round of training is performed on the first multimodal large model to obtain an expert model for generating detailed positioning description data.

[0065] As an example, the first round of detailed positioning description data and the sample image from which the first round of detailed positioning description data is derived are used as sample data and input into the first multimodal large model to train the first multimodal large model until the first multimodal large model meets expectations, thereby obtaining the first multimodal large model after the first round of training, and using the multimodal large model after the first round of training as the currently trained expert model, which has the ability to convert the input image into detailed positioning description data having image detailed description data and positioning description of text instances in the image detailed description data; Step 602: Perform a first round of training on the second multimodal large model using the first round of detailed positioning description data to obtain a first index expression matching model. As an example, the first round of detailed positioning description data is input into the first largest language model, and the target box description data is obtained through reasoning with the first largest language model, wherein the first largest language model is a trained model, which has the ability to generate target box description data based on the detailed positioning description data. The first largest language model, the third largest language model, and the fourth largest language model can be the same model or different models, and this application does not impose any restrictions on this.

[0066] Based on the obtained target frame description data, positive and negative samples are obtained, wherein the positive sample is that the target frame matches the second description data of the instance noun of the target frame, and the negative sample is that the target frame does not match the second description data of the instance noun of the target frame. Negative samples can be constructed as follows: One method is to use the first language model to modify the second description data of the instance noun, for example, to modify "a girl wearing a black mask and holding a dog" to "a girl wearing a white mask and holding a dog", and the target frame matching the original second description data and the modified second description data can be used as a negative sample pair; The second method is that in the same image, different second description data of the same category and the target boxes that do not correspond to them constitute negative sample pairs. For example, the target boxes of "the lady wearing a black hat on the left side of the image" and "the lady wearing a white shirt, green trousers and talking on the phone" are negative sample pairs, and the target boxes of "the lady wearing a white shirt, green trousers and talking on the phone" and "the lady wearing a black hat on the left side of the image" are negative sample pairs.

[0067] The obtained target box description data, positive and negative samples, and sample image data from which the first round of detailed positioning description data comes are used as sample data to train the second multimodal large model until prediction is achieved, and the trained second multimodal large model is obtained. The trained second multimodal large model is used as the currently trained first reference expression matching model.

[0068] There is no strict sequence for the above steps 601 and 602 and they can be executed in parallel.

[0069] The first finger expression matching model and the second finger expression matching model may be the same model or different models, which is not limited in this embodiment.

[0070] Step 603: input the target image for which detailed positioning description data is to be generated and / or the detailed description data of the target image for which detailed positioning description data is to be generated into the currently trained expert model, and obtain the detailed positioning description data through the reasoning of the currently trained expert model as the natural language description data of the image. In the case of inputting a target image, the obtained detailed positioning description data includes: image detailed description data, text instance description data corresponding to the text instance in the image detailed description data, and target frame description data corresponding to the text instance, and the detailed positioning description data is noisy detailed positioning description data, that is, the detailed positioning description data contains erroneous target frame description data; When the target image detailed description data is input, the obtained detailed positioning description data includes: image detailed description data, and text instance description data corresponding to the text instance in the image detailed description data. In this way, more accurate image detailed description data and text instance description data can be obtained through the expert model.

[0071] Since the detailed positioning description data includes the text instance description data corresponding to the text instance, the text instance description data can be used as the target described by the natural language, so there is no need to preset the target described by the natural language.

[0072] Step 604: input the noisy detailed positioning description data into the first large language model, and obtain the target box description data in the noisy detailed positioning description data and the second description data of the instance noun of the target box through reasoning by the first large language model. The first language model is a trained model, which has the ability to generate target frame description data and second description data of instance nouns of the target frame based on the detailed positioning description data.

[0073] Step 605: input the target frame description data in the noisy detailed positioning description data, the second description data of the instance noun of the target frame, and the image from which the noisy detailed positioning description data comes into the currently trained first reference expression matching model, and filter out the erroneous target frame description data in the noisy detailed positioning description data through the reasoning of the first reference expression matching model to obtain the target image detailed positioning description data, which is the denoised detailed positioning description data. The target frame description data in the denoised detailed positioning description data is the description data of the target frame that matches the second description data, and the target frame in the description data of the target frame is a candidate target.

[0074] During the inference process of the first referential expression matching model, the image content within the target box in the image corresponding to the target box description data is matched and detected with the referential expression, wherein the referential expression is the second description data of the instance noun corresponding to the target box, and the second description data of the instance noun can be used as the target described by natural language.

[0075] Through the above steps 603 to 605, accurate and detailed positioning description data can be generated as natural language description data of the image, and natural language description targets can be detected.

[0076] Step 606, using the detailed positioning description data of the target image and / or the target image detailed description data obtained in step 605, the currently trained expert model and the currently trained first reference expression matching model are trained in this round to update the model.

[0077] As an example, the detailed positioning description data of the target image and / or the detailed description data of the target image obtained in step 605 is used as the sample data of this round, and the currently trained expert model and the currently trained first-index expression matching model are trained to update the currently trained expert model and the currently trained first-index expression matching model to obtain the next round of trained expert model and the next round of trained first-index expression matching model.

[0078] The training method is the same as steps 601 and 602, namely: The detailed positioning description data obtained in step 605 and the image from which it is derived are used as sample data for this round and input into the first multimodal large model to train the first multimodal large model until the first multimodal large model meets expectations, thereby obtaining the first multimodal large model after this round of training, and using the multimodal large model after this round of training as the currently trained expert model.

[0079] In the iterative training process, this embodiment does not directly input the image into the expert model to obtain detailed positioning description data, but inputs the image into the first multimodal large model. This is to ensure the diversity of the detailed description and avoid the detailed positioning description data having an overly unified style, thereby reducing the diversity of expression.

[0080] The detailed positioning description data obtained in step 605 is used as the sample data of this round, and is input into the first large language model. The target box description data is obtained through reasoning with the first large language model. Based on the obtained target box description data, positive and negative sample data are obtained. The obtained target box description data, positive and negative sample data, and the sample image data from which the detailed positioning description data obtained in step 605 are derived are used as the sample data of this round. The second multimodal large model is trained until prediction is achieved, and the second multimodal large model after this round of training is obtained. The second multimodal large model after this round of training is used as the currently trained first reference expression matching model.

[0081] Return to step 603 to generate the next round of target image and / or target image detailed description data, and update the sample data through the data flywheel method to continuously iterate and update the model.

[0082] See also Figure 7 As shown, Figure 7 A schematic diagram of the generation of natural language description data of the image of this embodiment. The first multimodal large model is trained as an expert model, and the second multimodal large model is trained as a first referential expression matching model, wherein the first round of detailed positioning description data is used as sample data in the first round of training, and the final result of reasoning, i.e., denoised detailed positioning description data, is used as sample data in subsequent iterative training. In the reasoning process, the target image is input into the currently trained expert model, and detailed positioning description data is generated through reasoning of the currently trained expert model to obtain noisy detailed positioning description data; the noisy detailed positioning description data is input into the first language large model, and the target frame description data and the second description data of the instance noun matching the second description data of the instance noun are obtained through reasoning of the first language large model; the target frame description data and the second description data of the instance noun matching the second description data of the instance noun and the input target image are input into the currently trained first referential expression matching model, and the denoised detailed positioning description data is obtained through reasoning of the currently trained first referential expression matching model.

[0083] To facilitate understanding of the iterative process of this application, an example is used below to illustrate.

[0084] During the first round of inference, the image description data is input to the currently trained expert model: This photo was taken at yyyy-mm-dd hh:ff:ss, in pppp. It shows a portion of a city street, with shops and facilities on the left and parked cars on the right, including a silver sedan with a license plate of AAA. On the left, there is a middle-aged woman wearing a gray coat, black leggings, and black shoes. She is walking a tan-haired dog… Through the reasoning of the currently trained expert model, the instance noun can be extracted to obtain text description data containing the description data of the instance noun: This photo was taken on Time yyyy-mm-dd hh:ff:ss , Place yes pppp The image shows a section of a city street, with shops and facilities on the left and parked cars on the right. car , including a License plate number for AAA of Silver sedan On the left side of the screen, there is a Middle-aged women , wearing Gray coat and Black leggings , foot wear Black shoes She was holding a Brown Brown hair of dog … In the second round of reasoning, the image (e.g. Figure 8 As shown in the figure, the text description data obtained in the previous round is used to detect the target that matches the instance noun through the reasoning of the currently trained expert model, and obtain the detailed positioning description data including the position information of each target box of each instance noun: This photo was taken on Time yyyy-mm-dd hh:ff:ss [target frame location information] , Place yes pppp[target Frame location information] The image shows a section of a city street with shops and facilities on the left and parked cars on the right, including a License plate number for AAA [ Target frame position information] of Silver sedan [ Target frame position information] … As can be seen from the example, the detailed positioning description data includes: image detailed description data, and target box description data of the text instance, wherein the target box description data is used to describe the position information of the image where the target box corresponding to the text instance in the image detailed description data is located, and the text instance corresponds to multiple target boxes.

[0085] In the third round of reasoning, the image (e.g. Figure 8 As shown) and its image description data, the image description data is expanded into detailed positioning description data through the reasoning of the currently trained expert model, so as to obtain detailed positioning description data including the position information of each target box of each instance noun, and the detailed positioning description data is the same as the detailed positioning description data obtained in the second round.

[0086] In the fourth round, the image (e.g. Figure 8 As shown), through the reasoning of the currently trained expert model, detailed positioning description data is generated to obtain detailed positioning description data including the position information of each target box of each instance noun, and the detailed positioning description data is the same as the detailed positioning description data obtained in the second round.

[0087] See also Fig. 9 As shown, Fig. 9 This embodiment is based on Figure 8A schematic diagram of the target frame corresponding to the detailed positioning data. Among them, the green target is the image content including time information formed by OCR, the lake blue target is the image content including address information formed by OCR, and the instance constituted by the instance noun corresponds to multiple target frames. For example, the instance constituted by the instance nouns "license plate number", "AAA", and "silver car" corresponds to 2 target frames, and the instance constituted by the instance nouns "middle-aged woman", "gray coat", "black tights", "black shoes", and "dog" corresponds to 5 target frames. This embodiment enables the multimodal large model to output a detailed description of the image, and match N targets to any complex description in the detailed description of the image, where N is a natural number greater than or equal to 0.

[0088] See also Fig.10 As shown, Fig.10 This is a schematic diagram of a device for detecting a natural language description target based on an image according to an embodiment of the present application. The device includes: The detailed positioning description data generation module is used to input the image to be detected into the expert model that has been trained to convert the input image into detailed positioning description data with image detailed description data and to perform positioning description on the text instances in the image detailed description data, and obtain the detailed positioning description data of the image to be detected through the reasoning of the expert model. The detection module is used to use the detailed positioning description data of the image to be detected to obtain candidate targets in the image to be detected that match the natural language description targets represented by the text instance.

[0089] As an example, the detection module includes: The instance acquisition submodule is used to acquire the position information of the image instance located by the image instance description data in the image and the text instance in the image detailed description data based on the detailed location description data of the image to be detected. The matching submodule determines the image content of the image instance according to the position information of the image instance in the image, detects whether the determined image content matches the natural language description target, and takes the image instance that matches the natural language description target as a candidate target.

[0090] See also Fig.11 As shown, Fig.11 A schematic diagram of a device for generating natural language description data of an image in an embodiment of the present application. The device includes: The detailed positioning description data generation module inputs the target image for which natural language description data is to be generated and / or the detailed description data of the target image for which natural language description data is to be generated into the trained expert model, and obtains the noisy detailed positioning description data through the reasoning of the expert model as the natural language description data of the image.

[0091] As an example, the device further includes: An instance acquisition module is used to acquire the position information of the image instance located by the image instance description data in the image and the text instance in the image detailed description data based on the generated detailed location description data, The matching module is used to determine the image content of the image instance according to the position information of the image instance in the image, detect whether the determined image content matches the natural language description target, so as to remove the image instance description data of the image instance that does not match the natural language description target, and obtain denoised detailed positioning description data as the natural language description data of the image.

[0092] See also Fig.12 As shown, Fig.12 Another schematic diagram of an apparatus for detecting a natural language description target based on an image and / or an apparatus for generating natural language description data of an image according to an embodiment of the present application. The apparatus includes a memory and a processor, the memory stores a computer program, and the processor is configured to execute the computer program to implement any of the steps of the method for detecting a natural language description target based on an image and / or the steps of the method for generating natural language description data of an image.

[0093] The memory may include a random access memory (RAM) or a non-volatile memory (NVM), such as at least one disk memory. Optionally, the memory may also be at least one storage device located away from the aforementioned processor.

[0094] The above-mentioned processor can be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), etc.; it can also be a digital signal processor (DSP), an application specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components.

[0095] An embodiment of the present invention also provides a computer-readable storage medium, which stores a computer program. When the computer program is executed by a processor, it implements any step of the method for describing a target in natural language based on image detection, and / or any step of the method for generating natural language description data of an image.

[0096] As for the apparatus / network-side device / storage medium embodiment, since it is basically similar to the method embodiment, the description is relatively simple, and the relevant parts can be referred to the partial description of the method embodiment.

[0097] In this article, relational terms such as first and second, etc. are used only to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "comprise", "include" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "comprise a ..." do not exclude the presence of other identical elements in the process, method, article or device including the elements.

[0098] The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A method for detecting a natural language description target based on an image, characterized in that: The method includes: The image to be detected is input into an expert model that has been trained to convert the input image into detailed positioning description data having detailed image description data and positioning description of text instances in the detailed image description data, and the detailed positioning description data of the image to be detected is obtained through reasoning of the expert model. in, The detailed positioning description data of the image to be detected includes: image detailed description data of the image to be detected, and image instance description data corresponding to the text instance in the image detailed description data. The image instance description data is used to represent the position information of the image instance corresponding to the text instance in the image. The text instance is used to represent the natural language description target referred to in the image detailed description data, By using the detailed positioning description data of the image to be detected, candidate targets in the image to be detected that match the natural language description targets represented by the text instance are obtained.

2. The method according to claim 1, characterized in that The method of using the detailed positioning description data of the image to be detected to obtain a candidate target in the image to be detected that matches the natural language description target represented by the text instance includes: Based on the detailed positioning description data of the image to be detected, the position information of the image instance located by the image instance description data in the image and the text instance in the detailed image description data are obtained, Determine the image content of the image instance according to the position information of the image instance in the image, detect whether the determined image content matches the natural language description target, and use the image instance matching the natural language description target as a candidate target; The detailed positioning description data of the image to be detected also includes: text instance description data corresponding to the text instance in the image detailed description data.

3. The method according to claim 2, characterized in that The text instance description data includes: description data for referring to instance nouns of natural language description targets, The image instance description data includes: description data of a target frame for representing an image position where the image instance is located, The step of obtaining the position information of the image instance located by the image instance description data in the image and the text instance in the image detailed description data based on the detailed positioning description data of the image to be detected includes: Inputting the detailed positioning description data of the image to be detected into the first language model, and obtaining the target frame description data in the detailed positioning description data and the description data of the instance noun in the detailed positioning description data through the reasoning of the first language model; The step of determining the image content of the image instance according to the position information of the image instance in the image, and detecting whether the determined image content matches the natural language description target, includes: The image to be detected, the target frame description data in the detailed positioning description data, and the description data of the instance nouns in the detailed positioning description data are input into a first referential expression matching model that has been trained to match the content information in the target frame area with the referential expression of the description data used to characterize the instance nouns. Through the reasoning of the first referential expression matching model, the erroneous target frame description data in the detailed positioning description data is filtered out to obtain the denoised detailed positioning description data of the image to be detected.

4. The method according to claim 3, characterized in that The description data of the instance noun includes: first description data of the instance noun and / or second description data of the instance noun, wherein the first description data is used to represent a simple description of the instance noun, and the second description data is used to represent a complex description of the instance noun, the simple description of the instance noun includes the instance class name and / or the description of the content of interest, and the complex description of the instance noun includes: multiple class names and / or multiple phrased arbitrary natural language descriptions other than the single class name description data; The expert model is trained in the following way: Get the first round of detailed positioning description data, The first round of detailed positioning description data is used as sample data to train the first multimodal large model until the first multimodal large model meets expectations, thereby obtaining the trained first multimodal large model. The trained first multimodal large model is used as the first round of trained expert model.

5. The method according to claim 4, characterized in that The first index expression matching model is trained in the following manner: The first round of detailed positioning description data is input into the second largest language model. Through the reasoning of the second largest language model, the target box description data and the description data of the instance noun in the first round of detailed positioning description data are obtained. Based on the target box description data and instance noun description data in the first round of detailed positioning description data, positive and negative sample data are obtained. The target frame description data and the description data of the instance noun in the first round of detailed positioning description data, as well as the image data and the positive and negative sample data from which the first round of detailed positioning description data originate are used as sample data to train the second multimodal large model until the second multimodal large model meets expectations, thereby obtaining a trained second multimodal large model. The trained second multimodal large model is used as the first index expression matching model trained in the first round.

6. The method according to claim 5, characterized in that The training of the expert model further includes: The denoised detailed positioning description data is used as the sample data of this round, and the first multimodal large model of the previous round is trained until the first multimodal large model of the previous round meets the expectations, and the first multimodal large model after this round of training is obtained. The first multimodal large model after this round of training is used as the currently trained expert model to detect the next image to be detected; The training of the first index expression matching model further includes: The denoised detailed positioning description data is input into the second largest language model, and the target box description data and the description data of the instance noun in the denoised detailed positioning description data are obtained through the reasoning of the second largest language model. Based on the target box description data in the denoised detailed positioning description data and the description data of the instance noun, positive and negative sample data are obtained. The target frame description data and the description data of the instance nouns in the denoised detailed positioning description data, as well as the image data and the positive and negative sample data from which the denoised detailed positioning description data comes, are used as sample data for this round, and the second multimodal large model of the previous round is trained until the second multimodal large model of the previous round meets expectations, thereby obtaining the second multimodal large model after this round of training. The second multimodal large model after this round of training is used as the currently trained first index expression matching model to detect the next image to be detected.

7. The method according to claim 6, characterized in that The first round of detailed positioning description data is obtained in the following manner: Get the image detailed description data of the sample image, Based on the detailed image description data of the sample image, text instance extraction is performed to obtain first description data and second description data of the instance noun included in the text instance in the sample image. Performing target detection based on the sample image data to obtain a first target frame matching the first description data, Matching the sample image content of the first target frame with the second description data to obtain a second target frame matching the second description data, The second target frame, the first description data of the instance noun, the second description data, and the detailed image description data of the sample image are fused to obtain detailed positioning description pseudo-label data, which includes: detailed image description data, the first description data of the instance noun of the second target frame, the second description data, and the pseudo-label target frame description data of the second target frame.

8. The method according to claim 7, characterized in that The acquiring of detailed image description data of the sample image comprises: Inputting the sample image data into the third multimodal large model, and obtaining detailed image description data of the sample image through reasoning of the third multimodal large model; The text instance extraction based on the detailed image description data of the sample image includes: The detailed image description data of the sample image is input into the third language model, and the first description data of the instance noun of the text instance in the sample image is obtained through the reasoning of the third language model. Inputting the text containing the first description data of the instance noun of the text instance in the sample image into the fourth language macro model, and obtaining the second description data of the instance noun of the text instance in the sample image through the reasoning of the fourth language macro model; The performing target detection based on the sample image data to obtain a first target frame matching the first description data includes: Inputting the first description data of the sample image and its instance noun into the open set object detection model, and obtaining the position information of the image where the first object frame is located and the instance noun corresponding to the first object frame through the reasoning of the open set object detection model; The matching of the sample image content of the first target frame with the second description data includes: Inputting the sample image, the position information of the image where the first target frame is located, the instance noun data corresponding to the first target frame, and the second description data into the second referential expression matching model, and obtaining the position information of the image where the second target frame is located that matches the second description data through the reasoning of the second referential expression matching model; After fusing the second target frame, the first description data of the instance noun, the second description data, and the detailed image description data of the sample image, the method further includes: Remove the erroneous data in the detailed positioning description pseudo-label data to obtain the first round of detailed positioning description data.

9. The method according to claim 8, characterized in that The step of inputting the detailed description data of the sample image into the third language model further comprises: Inputting the detailed image description data of the sample image and the first text prompt for prompting the instance noun in the detailed image description data of the sample image into the third language macro model to determine the position of the instance noun in the detailed image description data of the sample image in the text; The step of inputting the text containing the first description data of the instance noun of the text instance in the sample image into the fourth language model further comprises: The paragraph text where the instance noun is located and the second text prompt for prompting the complex description in the detailed image description data of the sample image are input into the fourth language macro model to determine the complex description of the instance noun in the detailed image description data of the sample image.

10. A method for generating natural language description data of an image, characterized in that: The generation method includes: Input the target image to be generated with natural language description data and / or the target image detailed description data to be generated with natural language description data into the trained expert model, and obtain the detailed positioning description data through the reasoning of the expert model. in, The expert model is used to convert the input image into detailed positioning description data having image detailed description data and performing image positioning description on text instances in the image detailed description data, and / or, detailed positioning description data performing text positioning description on text instances in the input image detailed description data, The detailed positioning description data includes: image detailed description data, and image instance description data corresponding to the text instance in the image detailed description data, and / or the detailed positioning description data includes: image detailed description data, and text instance description data corresponding to the text instance in the image detailed description data, The image instance description data is used to represent the position information of the image instance corresponding to the text instance in the image. The text instance is used to represent the natural language description target referred to in the image detailed description data.

11. An electronic device, characterized in that: It includes a memory and a processor, the memory stores a computer program, and the processor is configured to execute the computer program to implement the steps of a method for natural language description of a target based on image detection as described in any one of claims 1 to 9, and / or the steps of a method for generating natural language description data of an image as described in claim 10.

Citation Information

Patent Citations

  • Image restoration method and device, electronic device and storage medium

    CN109886891A

  • Target identification method and model thereof, electronic equipment and storage medium

    CN115496895A

  • Image description processing method, computer equipment and storage medium

    CN116486188A

  • Model training method and device, target detection method and device, electronic equipment and medium

    CN117829243A

  • Description generation model training method and device, description generation method and device and electronic equipment

    CN119810593A