Image annotation data generation method and apparatus, model training method and apparatus, device and medium

WO2025185014A8PCT designated stage Publication Date: 2025-10-02BEIJING BAIDU NETCOM SCI & TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
PCT/CN2024/100320
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2024-03-08
Filing Date
2024-06-20
Publication Date
2025-10-02

AI Technical Summary

Technical Problem

Existing technologies have problems of low efficiency and high cost in the image annotation process. Especially in large-scale image data processing, manual annotation consumes a lot of human resources, while machine learning methods have uncertainty and inaccuracy.

Method used

By automatically extracting visual description information of target recognition objects in images, performing word verification and adjustment, accurate target description information is generated as annotation data, which is then used to train visual semantic understanding models to improve the efficiency and accuracy of annotation data generation.

Benefits of technology

It achieves efficient generation of image annotation data, reduces annotation costs, and improves the accuracy of annotation data and the accuracy of the output description text of the visual semantic understanding model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024100320_02102025_PF_FP_ABST
    Figure CN2024100320_02102025_PF_FP_ABST
Patent Text Reader

Abstract

An image annotation data generation method and apparatus, a model training method and apparatus, a device and a medium. The image annotation data generation method comprises: acquiring a sample image, a target recognition object in the sample image, and visual description information of the target recognition object (S101); verifying words in the visual description information on the basis of the sample image and the target recognition object to obtain a verification result of the visual description information (S102); adjusting the words in the visual description information on the basis of the verification result of the visual description information to obtain target description information of the target recognition object (S103); and taking the target description information of the target recognition object as annotation data of the target recognition object (S104).
Need to check novelty before this filing date? Find Prior Art

Description

Image annotation data generation, model training method, device, equipment and medium

[0001] This application claims priority to the Chinese patent application filed with the China Patent Office on March 8, 2024, with application number 202410268967.1, the entire contents of which are incorporated by reference into this application. Technical Field

[0002] This application relates to the fields of image processing, artificial intelligence, computer vision, large language models and smart cities, for example, to a method, device, equipment and medium for generating image annotation data, training a model. Background Art

[0003] With the rapid increase in the number of images, people urgently need to achieve efficient annotation of image content to achieve effective retrieval and management of large-scale images.

[0004] With the continuous advancement of artificial intelligence technology, the emergence of large-scale models has attracted widespread attention in the industry. In the ecological architecture of the new wave of artificial intelligence, the opportunities identified in the computing infrastructure layer and the application innovation of large-scale models have brought new development opportunities to the industry.

[0005] Summary of the Invention

[0006] The present application provides a method, apparatus, device and medium for generating image annotation data and model training.

[0007] According to one aspect of the present application, a method for generating image annotation data is provided, comprising:

[0008] Acquire a sample image, a target object in the sample image, and visual description information of the target object;

[0009] Verifying the words in the visual description information according to the sample image and the target recognition object to obtain a verification result of the visual description information;

[0010] Adjusting the words in the visual description information according to the verification result of the visual description information to obtain target description information of the target recognition object;

[0011] The target description information of the target recognition object is used as the annotation data of the target recognition object; wherein the annotation data is used to train a visual semantic understanding model, and the visual semantic understanding model is used to output a description text of the visual content based on the input visual content.

[0012] According to one aspect of the present application, a method for training a visual semantic understanding model is provided, comprising:

[0013] Acquire sample data, wherein the sample data includes a sample image and annotation data, and the annotation data is generated by the image annotation data generation method described in any embodiment of the present application;

[0014] The sample data is used to train a visual semantic understanding model.

[0015] According to one aspect of the present application, a device for generating image annotation data is provided, comprising:

[0016] A sample data acquisition module is configured to acquire a sample image, a target object in the sample image, and visual description information of the target object;

[0017] a visual description verification module configured to verify the words in the visual description information based on the sample image and the target recognition object, and obtain a verification result of the visual description information;

[0018] a target description information adjustment module, configured to adjust the words in the visual description information according to the verification result of the visual description information to obtain the target description information of the target recognition object;

[0019] The annotation data generation module is configured to use the target description information of the target recognition object as the annotation data of the target recognition object; wherein the annotation data is used to train the visual semantic understanding model, and the visual semantic understanding model is used to output the description text of the visual content based on the input visual content.

[0020] According to one aspect of the present application, a training device for a visual semantic understanding model is provided, comprising:

[0021] a sample data acquisition module configured to acquire sample data, wherein the sample data includes a sample image and annotation data, and the annotation data is generated by the image annotation data generation method as described in any embodiment of the present application;

[0022] The model training module is configured to use the sample data to train a visual semantic understanding model.

[0023] According to another aspect of the present application, an electronic device is provided, including:

[0024] at least one processor; and

[0025] a memory communicatively connected to the at least one processor; wherein,

[0026] The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the image annotation data generation method described in any embodiment of the present application, or the visual semantic understanding model training method described in any embodiment of the present application.

[0027] According to another aspect of the present application, a non-transitory computer-readable storage medium storing computer instructions is provided, wherein the computer instructions are used to enable the computer to execute the image annotation data generation method described in any embodiment of the present application, or the visual semantic understanding model training method described in any embodiment of the present application.

[0028] According to another aspect of the present application, a computer program product is provided, comprising a computer program, which, when executed by a processor, implements the image annotation data generating method described in any embodiment of the present application, or the visual semantic understanding model training method described in any embodiment of the present application. BRIEF DESCRIPTION OF THE DRAWINGS

[0029] FIG1 is a flowchart of a method for generating image annotation data according to an embodiment of the present application;

[0030] FIG2 is a flowchart of another method for generating image annotation data according to an embodiment of the present application;

[0031] FIG3 is a flowchart of another method for generating image annotation data according to an embodiment of the present application;

[0032] FIG4 is a flowchart of a method for training a visual semantic understanding model according to an embodiment of the present application;

[0033] FIG5 is a scene diagram of a method for generating image annotation data according to an embodiment of the present application;

[0034] FIG6 is a schematic diagram of the structure of an apparatus for generating image annotation data according to an embodiment of the present application;

[0035] FIG7 is a schematic diagram of the structure of a training device for a visual semantic understanding model according to an embodiment of the present application;

[0036] FIG8 is a block diagram of an electronic device provided according to an embodiment of the present application. DETAILED DESCRIPTION

[0037] The following description of exemplary embodiments of the present application is made in conjunction with the accompanying drawings, which include various details of the embodiments of the present application to facilitate understanding. These details should be considered as merely exemplary. Therefore, various changes and modifications may be made to the embodiments described herein. Similarly, for the sake of clarity and conciseness, descriptions of well-known functions and structures are omitted in the following description.

[0038] Figure 1 is a flowchart of a method for generating image annotation data according to an embodiment of the present application. This embodiment can be applied to the case of generating descriptive content of an image. The method of this embodiment can be executed by an image annotation data generating device, which can be implemented in software and / or hardware, and can be configured in an electronic device with a certain data computing capability. The electronic device can be a client device or a server device. The client device can include: a personal computer, a laptop computer, a smart phone, a tablet computer, an Internet of Things device or a portable wearable device, etc. The Internet of Things device can be a smart speaker, a smart TV, a smart air conditioner or a smart car device, etc. The portable wearable device can be a smart watch, a smart bracelet or a head-mounted device, etc.

[0039] S101: Acquire a sample image, a target object in the sample image, and visual description information of the target object.

[0040] The sample image includes at least one target recognition object. If there are multiple target recognition objects, different target recognition objects may correspond to different objects, and / or different target recognition objects may correspond to different types of objects. For example, the sample image includes a target recognition object of a person, a target recognition object of a cat, and a target recognition object of a table, etc. For another example, the sample image includes at least one target recognition object of a cat. Target detection can be performed on the sample image to obtain a target detection frame of the object in the sample image, and a target detection frame is determined as a target recognition object. Alternatively, the sample image can be segmented to obtain a segmented area in the sample image, and an area of ​​an object is determined as a target recognition object, while the background area does not include the object. In addition, while obtaining the target recognition object, the target recognition object can be classified to obtain the category of the target recognition object.

[0041] Visual description information is used to describe the image content of the target object and represent the semantic information of the image corresponding to the target object. Visual description information can refer to a statement describing the visually visible content. The visual description information can include at least one of the following: target object content, subject content, and relationship content. The subject content can refer to key information about the sample image. Exemplarily, the subject content can include scenes and / or events. The target object content can refer to the content of the target object itself. For example, the target object can include at least one of the following: size, state, position, and color. Relationship content can refer to the relationship between the target object and other objects in the image. For example, other objects can include objects within the target object and / or objects other than the target object in the sample image. Relationships can include at least one of positional relationships, dynamic and static relationships, and state relationships. Furthermore, the visual description information can include other content, which is not limited to this. Exemplarily, the sample image is an image of the surrounding environment captured by a vehicle while driving; the target object can include pedestrians, vehicles, road signs, indicator lights, and guardrails. For example, if the target object is a pedestrian, the corresponding visual description information can include: "There is a pedestrian directly ahead."

[0042] Sample images and target recognition objects can be input into a pre-trained generative model to obtain visual description information. Exemplarily, the generative model's processing process is as follows: the sample image and target recognition object are encoded separately to obtain vector representations corresponding to the sample image and target recognition object, respectively; prompt word text is obtained and encoded to obtain a vector representation of the prompt word text; the two vector representations are fused to obtain a fusion result; the fusion result is decoded to obtain visual description information. Furthermore, sample images, target recognition objects, and the category of the target recognition object can also be input into the generative model to obtain visual description information. The generative model can perform semantic understanding of visual content and generate description information.

[0043] S102: Verify the words in the visual description information according to the sample image and the target recognition object to obtain a verification result of the visual description information.

[0044] The visual description information is segmented to obtain at least one word. Each word is verified to determine if it is correct. Image information can be extracted from a sample image and a target object for verification of the words in the visual description information. Verification can include detecting whether a word is correct or exists based on the target object and the sample image. The verification result of the visual description information can determine if each word in the visual description information is correct, retaining correct words and removing incorrect words.

[0045] S103: Adjust the words in the visual description information according to the verification result of the visual description information to obtain target description information of the target recognition object.

[0046] The visual description information may refer to a sentence. Accordingly, the words in the visual description information may be adjusted by retaining correct words and modifying incorrect words to obtain a new sentence, which is then determined as the target description information for the target recognition object. Furthermore, synonyms of at least one word in the new sentence may be obtained, and the at least one word and the synonyms may be reorganized to form more new sentences. These multiple new sentences may then be screened to obtain more accurate and standardized sentences, which are then determined as the target description information.

[0047] S104. Use the target description information of the target recognition object as the annotation data of the target recognition object; the annotation data is used to train a visual semantic understanding model, and the visual semantic understanding model is used to output a description text of the visual content based on the input visual content.

[0048] The target description information is added to the sample data as annotation data. This can be understood as the true value of the description information of the target recognition object. The sample images and their target description information are used to train the visual semantic understanding model. The trained visual semantic understanding model can semantically understand the input image and output a description text of the image as the semantic understanding content of the image.

[0049] In fact, the sample image may include at least one target recognition object, and each target recognition object generates target description information. The sample image and the target description information of at least one target recognition object in the sample image can be used as training data to train the visual semantic understanding model.

[0050] In one embodiment, while manual image annotation can provide optimal labeling, it requires significant manpower and resources. Large models require enormous amounts of data, and some large enterprises have spent thousands of people working together for months to complete all required annotations, resulting in expensive labeling costs. When faced with new business needs, many enterprises face the dilemma of data and large amounts of data annotation, such as in the current era of large models. Alternatively, machine learning and deep learning methods can be used for automated annotation in related fields (such as detection and segmentation), but this approach carries significant uncertainty and inaccuracy.

[0051] According to the technical solution of the present application, the visual description information of the target recognition object is automatically extracted, and the words in the visual description information are verified to obtain the verification result. At the same time, based on the verification result, the words in the visual description information are adjusted to obtain the target description information as the annotation data of the target recognition object. Based on this, the image semantic understanding model is trained to obtain a model for semantic understanding of the input visual content, thereby improving the generation efficiency of image annotation data, reducing the annotation cost, and taking into account the annotation accuracy, and then quickly generating correct annotation data to train the model and improve the accuracy of the description text output by the visual semantic understanding model.

[0052] FIG2 is a flowchart of another method for generating image annotation data according to an embodiment of the present application, which is described based on the above technical solution and can be combined with the above multiple optional implementation methods. The method of verifying the words in the visual description information according to the sample image and the target recognition object to obtain the verification result of the visual description information includes: dividing the visual description information into at least one visual noun and at least one visual adjective; verifying the at least one visual noun according to the sample image and the target recognition object to obtain a noun verification result corresponding to the at least one visual noun; verifying the at least one visual adjective according to the sample image and the target recognition object to obtain an adjective verification result corresponding to the at least one visual adjective; and determining the noun verification result and the adjective verification result as the verification result of the visual description information.

[0053] S201: Acquire a sample image, a target object in the sample image, and visual description information of the target object.

[0054] S202: Divide the visual description information into at least one visual noun and at least one visual adjective.

[0055] Visual nouns refer to visually visible nouns related to the target object. Visual adjectives refer to visually visible adjectives related to visual nouns. Visual nouns are used to describe the category of the target object. Visual adjectives are used to limit the target object. Visual adjectives can be divided into material, color, or location, etc. For example, the visual description information is a white ragdoll cat, the visual noun can include cat, and the visual adjectives can include white and ragdoll. For another example, the visual description information is a black dog lying on the carpet. Visual nouns include: dog and carpet. Visual adjectives can include black and dog on the carpet.

[0056] For example, the visual description information can be divided into at least one noun-adjective pair. For example, the visual description information is a red apple, and the noun-adjective pair extracted is { <red> , <apple>}. For another example, a black dog is lying on the carpet, and the noun-adjective pairs extracted are {<black>, <dog>}, {<lying>, <dog>}, {<up>, <dog>}, and {<down>, <carpet>}. This visual description information corresponds to the target recognition object. In addition, a sample image can contain multiple target recognition objects. For example, the visual description information of object 1 is: a white ragdoll cat, and the visual description information of object 2 is a black dog lying on the carpet. Correspondingly, the noun-adjective pairs that can be formed in a sample image include: {<white>, <cat>, <object 1>}, {<ragdoll>, <cat>, <object 1>}, {<black>, <dog>, <object 2>}, {<lying>, <dog>, <object 2>}, {<up>, <dog>, <object 2>}, and {<down>, <carpet>, <object 2>}.

[0057] S203 : Verify the at least one visual noun according to the sample image and the target recognition object to obtain a noun verification result corresponding to the at least one visual noun.

[0058] Based on the sample image and the target object, the category of the target object can be obtained and the visual noun can be verified. For example, this can be done to determine whether the category of the target object is similar to the visual noun. The noun verification result can indicate whether the visual noun is correct. Each visual noun corresponds to one noun verification result.

[0059] Optionally, the at least one visual noun is verified based on the sample image and the target recognition object to obtain a noun verification result corresponding to the at least one visual noun, including: obtaining category description information of the target recognition object; for the at least one visual noun, detecting the consistency between the category description information of the target recognition object and the at least one visual noun to obtain a noun verification result corresponding to the at least one visual noun.

[0060] The category description information may be the target object and the category of the target object detected when detecting the object in the sample image, and the category of the target object determined as the category description information. For example, a detection frame of the target object is detected in the sample image, along with the category of the detection frame, and the image region corresponding to the detection frame is determined as the target object; the category of the detection frame is determined as the category description information of the target object. The category description information and the target object can be stored in association to facilitate subsequent retrieval of the category description information of the target object.

[0061] The category description information is compared with the visual noun. If they are consistent, the noun verification result of the visual noun is determined to be correct; if they are inconsistent, the noun verification result of the visual noun is determined to be incorrect. In addition, the similarity between the category description information and the visual noun can be calculated and the similarity can be determined as the noun verification result. Alternatively, when the similarity is greater than or equal to a corresponding similarity threshold, the noun verification result of the visual noun is determined to be correct; when the similarity is less than the corresponding similarity threshold, the noun verification result of the visual noun is determined to be incorrect.

[0062] By comparing the category description information of the target recognition object with the visual noun for consistency, a comparison result is obtained, and based on the comparison result, the noun verification result of the visual noun is determined. The visual noun can be verified according to the category obtained by the sample image when detecting the object, avoiding the verification of the visual noun through additional complex and redundant steps, simplifying the verification operation of the visual noun, and improving the verification efficiency of the visual noun. At the same time, the accuracy of the category of the target recognition object is high, and the category of the target recognition object is used for verification to improve the verification accuracy of the visual noun.

[0063] S204: Verify the at least one visual adjective according to the sample image and the target recognition object to obtain an adjective verification result corresponding to the at least one visual adjective.

[0064] Visual adjectives typically qualify or describe a visual noun. Verification of visual adjectives typically involves checking whether the characteristics of the visual noun they describe are consistent with the visual adjective. If they are consistent, the visual adjective is considered correct; if not, it is considered incorrect. Each visual adjective corresponds to one adjective verification result.

[0065] Optionally, the verifying of the at least one visual adjective according to the sample image and the target recognition object to obtain an adjective verification result corresponding to the at least one visual adjective includes: generating interactive dialogue content for the at least one visual adjective according to the at least one visual adjective and a visual noun corresponding to the at least one visual adjective; using the sample image and the target recognition object as context, and inputting the interactive dialogue content into a large language model to obtain an existence judgment result of the at least one visual adjective; and generating an adjective verification result of the at least one visual adjective according to the existence judgment result of the at least one visual adjective.

[0066] The interactive dialogue content is used to detect whether the description content of the visual noun contains content corresponding to the corresponding visual adjective. The interactive dialogue content is usually a correctness judgment statement. The interactive dialogue content includes a visual adjective, the visual noun corresponding to the visual adjective, and the judgment content of whether the visual noun corresponding to the visual adjective is associated with the visual adjective. Exemplarily, the judgment content includes: a judgment word or a predicate; the association between a visual noun and a visual adjective can mean that the visual noun has the property corresponding to the visual adjective, for example, the visual noun is cat, the visual adjective is white, and the color of the cat in the sample image is white, that is, the visual noun has the color property represented by the visual adjective, and it is determined that the visual noun is associated with the visual adjective. A visual noun and a visual adjective are not associated. This can mean that the visual noun does not possess the property corresponding to the visual adjective. On the one hand, the property represented by the visual adjective is different from the property of the visual noun. For example, in the sample image, the color of the cat is black, meaning that the visual noun does not possess the color property represented by the visual adjective, thus determining that the visual noun and the visual adjective are not associated. On the other hand, in the sample image, the property represented by the visual noun is unrelated to the property represented by the visual adjective. For example, in the sample image, the visual noun is "cat" and the visual adjective is "cat's eyes are blue." In the sample image, the cat is shown from behind, so it is impossible to determine the eye color of the visual noun, thus determining that the visual noun and the visual adjective are not associated. The interactive dialogue content is used as input to the large language model. The large language model relies on its semantic understanding of the target recognition object in the sample image to determine whether the interactive dialogue content is correct. The large language model outputs the result of whether the interactive dialogue content is correct. The visual adjective presence determination result can refer to whether the visual noun corresponding to the visual adjective possesses the property of the visual adjective.

[0067] Inputting the sample image and target recognition object as context into the large language model can mean that the large language model responds to the interactive dialogue content based on the sample image and target recognition object, and outputs a visual adjective presence judgment result. Alternatively, the sample image can be input as context into the large language model. The sample image and target recognition object can be input as context into the large language model, and the interactive dialogue content can be input into the large language model. The large language model then judges the correctness of the interactive dialogue content based on the semantic understanding of the sample image and target recognition object, and obtains a visual adjective presence judgment result. The sample image and target recognition object can be input into the large language model simultaneously with the interactive dialogue content, or a multi-round dialogue can be conducted, where the sample image and target recognition object are first input into the large language model for dialogue, and then the interactive dialogue content is input into the large language model. In other words, in a multi-round dialogue, the first dialogue includes the sample image and target recognition object, and the second dialogue includes the interactive dialogue content. Among them, the large language model can be the aforementioned generative model, which is used to input sample images and target recognition objects and output visual description information. After inputting the sample images and target recognition objects, the next round of dialogue is carried out, the interactive dialogue content is input, and the judgment result of the interactive dialogue content is output, that is, the judgment result of the existence of visual adjectives.

[0068] For example, the visual description information is: white ragdoll cat. The visual adjective is white, and the corresponding visual noun is cat. The interactive dialogue content is: Is the cat white? For another example, the interactive dialogue content is: Do white cats exist? Taking the sample image and the target recognition object as the context, if the large language model outputs the result of "yes" for the aforementioned interactive dialogue content, then the existence judgment result of the visual adjective is "yes", and the corresponding adjective verification result of the visual adjective is correct; if the large language model outputs the result of "does not exist" for the aforementioned interactive dialogue content, then the existence judgment result of the visual adjective is "does not exist", and the corresponding adjective verification result of the visual adjective is incorrect.

[0069] A visual noun can correspond to multiple visual adjectives, and one visual adjective corresponds to one visual noun. For multiple visual adjectives, the existence judgment result is obtained for each visual adjective, that is, each visual adjective generates corresponding interactive dialogue content, and the existence judgment result is obtained through the large language model. If the verification result of a visual adjective is wrong, only the visual adjective will be eliminated or corrected. Since there may be other visual adjectives corresponding to the same visual noun as the visual adjective, the visual noun corresponding to the visual adjective cannot be directly eliminated or corrected, and the corresponding visual noun needs to be verified using other verification methods.

[0070] By combining visual adjectives with corresponding visual nouns, interactive dialogue content is generated for judging whether visual adjectives exist, and the interactive dialogue content is input into a large language model with sample images and target recognition objects as context to detect whether the visual noun has the properties of the visual adjective, thereby judging whether the visual adjective exists, and then determining the verification result of the visual adjective. The judgment can be made based on the visual content of the sample image itself, avoiding the verification of visual adjectives through additional complex and redundant steps, simplifying the verification operation of visual adjectives, and improving the verification efficiency of visual adjectives. At the same time, the visual nouns are associated with the visual adjectives, and the association judgment of the visual nouns and visual adjectives is made, making full use of the characteristic that visual adjectives cannot exist alone, verifying both the existence of the visual adjective and the association of the visual noun with the visual adjective, thereby improving the verification accuracy of the visual noun.

[0071] S205: Determine the noun verification result and the adjective verification result as the verification result of the visual description information.

[0072] The verification results of all nouns and all adjectives are determined as the verification results of the visual description information.

[0073] S206: Adjust the words in the visual description information according to the verification result of the visual description information to obtain target description information of the target recognition object.

[0074] Correct or delete the words corresponding to the incorrect verification results. Obtain similar words for the words corresponding to the correct verification results. Reorganize the correct words, similar words, and corrected words to obtain new visual description information. Filter the multiple description information to obtain the target description information of the target recognition object.

[0075] Optionally, the wording in the visual description information is adjusted according to the verification result of the visual description information to obtain the target description information of the target recognition object, including: screening the at least one visual noun according to the verification result of each noun to obtain at least one target noun; screening the at least one visual adjective according to the verification result of each adjective to obtain at least one target adjective; generating at least one descriptive phrase according to the at least one target noun and the at least one target adjective; and generating the target description information of the target recognition object according to the at least one descriptive phrase.

[0076] Filter the words with correct noun verification results from at least one visual noun and determine them as target nouns. Filter the words with correct adjective verification results from at least one visual adjective and determine them as target adjectives. A description phrase includes a noun and an adjective. For each target adjective, the target adjective can be combined with the defined target noun to obtain a description phrase. Alternatively, in the aforementioned example, when verifying the visual adjectives, a visual adjective and the visual noun corresponding to the visual adjective are formed into a phrase, and the phrase is filtered according to the target noun and the target adjective to determine the description phrase. The target noun and the target adjective in the description phrase correspond to each other, that is, the target adjective defines and describes the target noun in the same description phrase. Sentences are constructed based on the nouns and adjectives in the description phrase to form sentences, which are determined as target description information of the target recognition object.

[0077] A target recognition object can generate at least one target description information. All target description information can be used as the annotation data of the target recognition object, or the target description information can be filtered, for example, by scoring and filtering based on grammar, standardization and expression richness, to obtain a target description information as the annotation data of the target recognition object.

[0078] In addition, corresponding synonyms can be added to at least one target noun and at least one target adjective as target nouns or target adjectives, and based on the added target nouns and target adjectives, they can be arranged and combined to form descriptive phrases, which can enrich the content of the descriptive phrases, thereby enriching the content of the target description information and increasing the representativeness of the target description content.

[0079] By screening visual nouns and visual adjectives respectively according to the verification results of the words, the target nouns and target adjectives are obtained, and the descriptive phrases are generated in combination. According to the words in the descriptive phrases, the target description information of the target recognition object is generated. The words of different parts of speech can be screened separately to achieve accurate verification, and the screened and retained words are combined accordingly, that is, the correlation between nouns and adjectives is screened to achieve correlation verification, and the target description information is generated according to the verified descriptive phrases to improve the accuracy of the target description information.

[0080] S207. Use the target description information of the target recognition object as the annotation data of the target recognition object; the annotation data is used to train a visual semantic understanding model, and the visual semantic understanding model is used to output a description text of the visual content based on the input visual content.

[0081] According to the technical solution of the present application, by segmenting the visual description information, visual nouns and visual adjectives are obtained, and the visual nouns and visual adjectives are verified respectively to obtain the verification results of the visual description information. The visual description information can be disassembled and verified based on words as units, and the description information can be verified at a finer granularity. At the same time, words with different parts of speech are verified separately to achieve targeted description word verification and improve the verification accuracy of the description information.

[0082] FIG3 is a flowchart of another method for generating image annotation data according to an embodiment of the present application, which is described based on the above technical solution and can be combined with the above multiple optional implementation methods. The sample image and target recognition object are obtained by: obtaining a real-time updated object category and adding the obtained real-time updated object category to a dynamic vocabulary; obtaining a sample image; identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image, and obtaining the target recognition object in the sample image.

[0083] S301: Acquire a sample image, a target object in the sample image, and visual description information of the target object.

[0084] S302: Verify the words in the visual description information according to the sample image and the target recognition object to obtain a verification result of the visual description information.

[0085] S303: Adjust the words in the visual description information according to the verification result of the visual description information to obtain target description information of the target recognition object.

[0086] S304. Use the target description information of the target recognition object as the annotation data of the target recognition object; the annotation data is used to train a visual semantic understanding model, and the visual semantic understanding model is used to output a description text of the visual content based on the input visual content.

[0087] S305 , the sample image and the target recognition object are obtained by: acquiring the object category updated in real time, and adding the object category updated in real time to a dynamic vocabulary.

[0088] The object category can be the category of an object that can be identified in an image. The object category that is updated in real time refers to the category of an object that can be identified in an image and collected from a network hotspot or a real-time event. The dynamic vocabulary is used to store the object category. For example, real-time images or hotspot images can be periodically acquired, and objects can be identified therefrom, and the object category of the object can be added to the dynamic vocabulary. The steps of obtaining the sample image and the target identification object when generating the annotation data can be independent of each other. The sample image and the target identification object can be updated periodically, and a large number of sample images and target identification objects can be stored. The sample image and at least one target identification object identified in the sample image can be used as a sample data; or the sample image and a target identification object can be used as a sample data. If multiple target identification objects are identified in the sample image, multiple sample data are generated accordingly. When it is necessary to generate annotation data, at least one sample data is obtained from a large amount of pre-stored sample data. For example, the latest multiple sample data can be obtained.

[0089] It is also possible to obtain hot words or real-time words as object categories, but these words may not be visualized, making them difficult to use as object categories. Furthermore, the latest identifiable object categories have little impact. Therefore, real-time images or hot-spot images can be selected and detected on these latest images to obtain new object categories. These new object categories can then be added to the dynamic vocabulary to accurately update object categories.

[0090] In addition, a static vocabulary can be set up. For example, by extracting categories from the OBJECT365 dataset (a general object detection dataset), the Common Objects in Context (COCO) dataset, and the OPENIMAGES dataset (an open large-scale image annotation dataset), and combining them with business needs, the collected categories are added to the static vocabulary. However, static vocabulary is not open to images, that is, in the future, there is no way to extract more diverse frame information in new business needs. Therefore, tagging is used to extract unique words from each image and add them to the dynamic vocabulary. The combination of static and dynamic vocabulary greatly enriches the content of object categories.

[0091] Optionally, obtaining the object category updated in real time includes: obtaining a visual set updated in real time; performing target detection on the visual data in the visual set to obtain at least one first object category; and / or performing semantic understanding on the visual data in the visual set to obtain at least one second object category; and determining the first object category and / or the second object category as the object category updated in real time.

[0092] A visual set may refer to a collection of images or videos. The visual data may include a captured image or a frame of video. The first object category may be an object category obtained by object detection in the latest visual data. Object detection may be performed on the visual data to obtain at least one first object and a first object category for the at least one first object.

[0093] In one example, a Recognize Anything Plus model (RAM++) can be proposed to perform target detection on visual data through the Inject Semantic Concepts Into Image Tagging Open-Set Recognition method to obtain a first object category. The extracted first object category can ensure the word diversity of the dynamic vocabulary and the different extracted visual data, avoiding the homogenization of the first object category caused by extracting the first object category from the same or similar visual data, thereby increasing the richness and representativeness of the first object category.

[0094] The second object category may be an object category derived from semantically understanding the latest visual data. Semantic understanding of the visual data yields at least one second object and a second object category for the at least one second object. The first object category and the second object category are obtained in different ways, but both are object categories extracted from the visual data.

[0095] In one example, a large language model can be used to describe visual data. For example, if the large language model is fed the question "What objects might be included in the following scene?", it will respond based on the visual data, extracting the category terms from the response to obtain a second object category. Describing the second object category based on the visual data ensures the integrity and diversity of the dynamic vocabulary.

[0096] Only the first object category or only the second object category may be acquired and added to the dynamic vocabulary.

[0097] By acquiring real-time updated visual data and performing target detection and / or semantic understanding on the visual data, a first object category and / or a second object category is obtained, which is added to a dynamic vocabulary to enrich the words in the dynamic vocabulary. Based on more complete, rich and representative words, objects in sample images are identified, thereby increasing the diversity and representativeness of the identified objects, thereby making the labeled data real-time, reliable and diverse.

[0098] S306: Acquire a sample image.

[0099] The sample image may be the same as or different from the image used to identify the object category updated in real time. The sample image may be collected, obtained from a public channel, or obtained from an undisclosed channel with authorization.

[0100] S307 : Identify the target recognition object corresponding to each word in the dynamic vocabulary in the sample image to obtain the target recognition object in the sample image.

[0101] Each word in the dynamic vocabulary that is updated in real time is used to identify whether there is an object corresponding to each word in the sample image. If there is an object, the object is determined as the target recognition object.

[0102] The corresponding word can be determined as the target category of the target recognition object, or the target category of at least one target recognition object can be identified while identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image. The sample image and all target recognition objects identified in the sample image are combined to generate a piece of sample data. The target recognition object can be a target detection frame. In addition, the sample image, at least one target recognition object identified in the sample image, and the target category of the at least one identified target recognition object can be combined to generate sample data.

[0103] The target identification object can also be pre-processed to reduce erroneous redundant data.

[0104] Optionally, after identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image, it also includes: detecting the intersection-and-union ratio between multiple target recognition objects in the sample image; detecting the multiple target recognition objects based on the intersection-and-union ratio between the multiple target recognition objects to obtain similar multiple target recognition objects; fusing the similar multiple target recognition objects to obtain a fusion result; and updating the multiple target recognition objects in the sample image based on the fusion result.

[0105] In practice, there are multiple target recognition objects in the sample image, and these multiple target recognition objects may contain redundant data. If there is only one target recognition object, this step can be omitted. Multiple target recognition objects can be deduplicated. Deduplication can be performed using the intersection-of-union (IoU) ratio between multiple target recognition objects. For each of the multiple target recognition objects, the IoU ratio is calculated between the target detection frames or image regions of each pair of target recognition objects. When the IoU ratio is greater than or equal to a preset IoU threshold, the two target recognition objects are determined to be similar, resulting in two similar target recognition objects. A similarity comparison is performed between each pair of target recognition objects to obtain multiple similar target recognition objects. Furthermore, semantic understanding can be performed on the detection frames of the target recognition objects to obtain target description information for the detection frames and divide it into visual noun-visual adjective pairs. When the IoU ratio of two target recognition objects is greater than or equal to a preset IoU threshold, and the corresponding visual noun-visual adjective pairs in the target description information are the same, the two target recognition objects are determined to be similar. If only the IoU ratio is greater than or equal to the preset IoU threshold, or if only the word pairs are the same, the two target recognition objects cannot be determined to be similar.

[0106] Exemplarily, a fusion method for a group of similar multiple target recognition objects may be: selecting one of the target recognition objects as the fusion result, eliminating the remaining target recognition objects, or calculating the intersection or union of the multiple target recognition objects to determine the fusion result. At the same time, in this group of similar multiple target recognition objects, each target recognition object has a target category, and the number of identical target categories is counted as the number of occurrences of the identical target category. The target category with the largest number of occurrences in the group of similar multiple target recognition objects is counted and determined as the target category of the fusion result. Updating the target recognition object in the sample image according to the fusion result may be to use the fusion result as the target recognition object finally identified in the sample image to update the target recognition object in the sample image. Alternatively, based on the fusion result, redundant target recognition objects in the sample image are eliminated to update the target recognition object in the sample image.

[0107] By using the intersection-over-union method to eliminate redundant data from multiple target recognition objects identified in the sample image, the redundant data of the target recognition objects is reduced, the multiple annotations of the same object are reduced, and the amount of annotation data is reduced, thereby improving the annotation efficiency and accuracy.

[0108] Optionally, after identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image, it also includes: detecting the number of pixels of at least one target recognition object in the sample image; and updating the target recognition object in the sample image according to the number of pixels of the at least one target recognition object.

[0109] The number of pixels is used to represent the amount of information about a target object. In practice, if a target object has too few pixels—for example, a single pixel—it's difficult to distinguish it from other targets and to obtain valid information about it. Consequently, the identification of that target object is of limited significance. Even if it can be annotated, the annotated data won't provide sufficient information for model training.

[0110] Compare the number of pixels of each target object to the corresponding pixel count threshold. Target objects with a pixel count less than the threshold are eliminated, while those with a pixel count greater than or equal to the threshold are retained. Different pixel count thresholds can be set for different categories.

[0111] Updating the target recognition objects in the sample image according to the number of pixels of at least one target recognition object may be eliminating redundant target recognition objects in the sample image according to the number of pixels of at least one target recognition object. For example, target recognition objects with too few pixels are eliminated to achieve the updating of the target recognition objects in the sample image.

[0112] By screening the target recognition objects according to the number of pixels of the target recognition objects, the target recognition objects with more effective information and more key information can be retained, so as to increase the representativeness of the target recognition objects, increase the content richness and representativeness of the annotation data, and make the annotation data contain more effective information.

[0113] The aforementioned target identification objects are screened by the dimensions of redundancy and effectiveness. In addition, other dimensions can also be used to screen the target identification objects, which is not limited.

[0114] In fact, in order to extract more effective content from the target recognition object and increase the accuracy of the labeled data, the relative area ratio of the target recognition object and the sample image can be adjusted before generating sample data.

[0115] Optionally, after identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image, it also includes: detecting the size ratio of at least one target recognition object in the sample image to the sample image; when the size ratio of the target recognition object is less than a preset threshold, expanding the boundary of the target recognition object to increase the size ratio of the expanded target recognition object to the sample image.

[0116] The size ratio of the target object is used to indicate its importance to the sample image. To extract richer and more complete data from the target object, the size ratio of the target object is typically kept within a certain range. The size ratio refers to the ratio of the area of ​​the target object's detection box or image region to the sample image area. If the size ratio of the target object is less than the preset threshold, it indicates that the size ratio is too small and should generally be increased.

[0117] You can expand the boundaries of a target object. Typically, the target object's detection box is rectangular. You can move the edges of the rectangle outward by a distance x. The expanded target object includes the original target object. Typically, the center of the target object remains unchanged. The target object can be enlarged proportionally.

[0118] Optionally, after identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image, it also includes: detecting the size ratio of at least one target recognition object in the sample image to the sample image; when the size ratio of the target recognition object is less than a preset threshold, cropping the sample image and updating the sample image to increase the size ratio of the target recognition object to the cropped sample image.

[0119] Crop the sample image to reduce its size, ensuring that the cropped area does not intersect with the target object. The image can be cropped to avoid the location of at least one target object, while preserving the shape, angle, and position of the sample image.

[0120] You can only enlarge the size of the target recognition object, or only crop the sample image and reduce the size of the sample image, or you can simultaneously enlarge the size of the target recognition object and crop the sample image to make the size ratio larger, which can increase the image area used in the annotation process to be the most informative, thereby improving the accuracy of the annotation.

[0121] According to the technical solution of the present application, by obtaining the latest object category and adding it to a dynamic vocabulary, the target recognition object is recognized according to each word in the dynamic vocabulary to obtain the target recognition object, which can increase the diversity of identifiable objects, thereby increasing the diversity of annotation data and enriching the content of the annotation task. In addition, the object category can be updated in real time, and the recognized object can be updated accordingly, thereby updating the annotation data in real time, which can make the annotated data open and dynamically changeable, thereby improving the accuracy and real-time performance of the annotation data.

[0122] Figure 4 is a flow chart of a method for training a visual semantic understanding model according to an embodiment of the present application. This embodiment can be applicable to the case where a visual semantic understanding model is trained based on the aforementioned generated annotation data in combination with sample images, and the visual semantic understanding model is used to input an image and output a description of the image content. The method of this embodiment can be executed by a training device for a visual semantic understanding model, which can be implemented in software and / or hardware and configured in an electronic device with a certain data computing capability, which can be a client device or a server device, and the client device can include: a personal computer, a laptop computer, a smart phone, a tablet computer, an Internet of Things device or a portable wearable device, etc.

[0123] S401. Acquire sample data, where the sample data includes a sample image and annotation data, and the annotation data is generated by the image annotation data generation method described in any embodiment of the present application.

[0124] The labeled data is used as the true description content of the sample image, and the sample image is input into the visual semantic understanding model to train the visual semantic understanding model.

[0125] S402: Using the sample data, train a visual semantic understanding model.

[0126] The visual semantic understanding model can be a large language model.

[0127] According to the technical solution of the present application, by using verified annotated data as training data to train the visual semantic understanding model, the accuracy of the descriptive text output by the visual semantic understanding model can be improved. At the same time, the automatic generation of annotated data can reduce the model training cost and improve the model training efficiency.

[0128] Figure 5 is a scene diagram of a method for generating image annotation data according to an embodiment of the present application. This embodiment of the present application is mainly divided into two phases: Phase 1 and Phase 2. Phase 1 includes five small steps, such as generating sample data, which includes an image, a detection box in the image, and a category label. Phase 2 includes four small steps, such as generating target description information for the detection box, that is, implementing the annotation of the detection box.

[0129] Overview of the first and second stages: RAM++ is used to perform target detection on sample images, and category labels are added to the detected objects, and Open-World Localization version 2 (OWLv2) is used to implement open vocabulary object detection (OVD). The generated detection box is subjected to box suppression by the filter module; the input of the visual semantic understanding model is generated, and the visual semantic understanding model is used to generate description information (Caption). The input needs to include the detection box (BOX) and label (LABEL) of the object (which can refer to the category label). The image, the object in the image, that is, the BOX and LABEL of the target recognition object are input into the large language model to obtain the initial version of the target description information Caption. The large language model can be a multimodal large language model (MLLM). Exemplarily, the MLLM model is a shikra model. Typically, the MLLM model includes a visual encoding layer, a fusion layer, a text encoding layer, and a decoding layer. The initial version of the target caption is decomposed into semantic terms to obtain corresponding nouns and adjectives, namely visual nouns and visual adjectives. The word meanings are parsed, and adjectives are categorized into material, color, behavior, and location. Each adjective is double-checked to ensure the accuracy of the filtered adjectives. The nouns are matched with tagging information (i.e., category labels) to ensure the accuracy of the noun pairs. The generated parsed terms are semantically concatenated and diversified to obtain the target caption. This is verified through MLLM to improve the ability to generate captions.

[0130] For example, 1. Manually set at least one static vocabulary: By extracting categories from the OBJECT365 dataset, the COCO dataset, and the OPENIMAGES dataset, and combining them with business needs, we collected more than 500 categories as a static vocabulary. However, static vocabulary is not open to images. That is, in the future, when new business needs arise, there is no way to extract more diverse frame information. Therefore, tagging is used to extract unique vocabulary information for each image.

[0131] 2. Obtain a dynamic vocabulary for the image through tagging: Execute step 1-1 in Figure 5 to add tags (tagging). RAM++ crawls millions of visual data (images or images in videos) from the Internet, automatically annotates the visual data, and implements coarse-grained tagging operations. It adds image-level tags to the visual data, obtains a large number of data image-text pairs, and extracts corresponding word information based on the automatically annotated tags. It can recognize more than 6,000 categories, and can identify the categories in the image and the background, and add the extracted word information to the dynamic vocabulary. Tagging can ensure the diversity of the image vocabulary and ensure that the multiple extracted images are different. For example, the dynamic vocabulary can include: people, advertisements, cars, animals, scenes, food, furniture, and appliances. In addition, there are many more and richer words, which are not limited to this.

[0132] 3. Imaginative vocabulary obtained through a large language model: We also extracted background categories from RAM++, such as park and kitchen, and used the large language model to generate the question "What objects might be included in the following scenes?" This ensures the integrity of the vocabulary and the diversity of the generated image vocabulary.

[0133] Execute steps 1-2 to add the words in the dynamic vocabulary to the preset vocabulary. In addition, the words in the static vocabulary and imaginary vocabulary obtained above can also be added to the preset vocabulary.

[0134] 4. Generate detection frames using open domain detection capabilities: Using the open domain vocabulary model based on the aforementioned vocabulary, perform reasoning in steps 1-3 based on the preset vocabulary to detect and obtain detection frames in the sample image. In the embodiment of the present application, OWLv2 is used to generate detection frames corresponding to the preset vocabulary. For example, by inputting the original image, i.e., the sample image, and combining it with the preset vocabulary, detection frames corresponding to the words in the preset vocabulary in the sample image can be generated. For example, in the sample image, object 1 and object 2 are identified, where the categories of object 1 and object 2 can be different or the same.

[0135] 5. The generated detection frames are input into the Filter module after executing steps 1-4. The Filter module performs frame suppression and outputs multiple sets of images, detection frames, and category labels based on steps 1-5. Through manual labeling, static categories are divided into at least one major category, such as at least one of food, tools, people, traffic scenes, and tools. The main suppression content is food and people. The operation is as follows:

[0136] a) Extract the corresponding nouns and adjectives from the caption. For example, if the description is a red apple, { <red> , <apple>}.

[0137] b) An image will generate at least one adjective-noun pair and the corresponding box position, and the adjective, noun and box will be combined to generate a list [{<adj.1> ,<n.1> ,box1},{adj.2>,<n.2> ,box2}…].

[0138] c) Loop through the list and calculate the Intersection over Union (IOU) between the boxes in the list. If the IOU of two boxes is greater than or equal to a preset IOU threshold, and the adjective (adj) and noun (n) are the same, then the two boxes are merged into one box. Alternatively, if the IOU of two boxes is greater than or equal to a preset IOU threshold, then the two boxes are merged into one box. Otherwise, no fusion is performed.

[0139] For example,

[0140] i.Box_new=

[0141] [min(box1.x,box2.x),min(box1.y,box2.y),max(box1.w,box2.w),max(box1.h,box2.h)

[0142] Where Box_new is the fused detection box. The center coordinates of the fused detection box are the minimum of the center coordinates of the two fused detection boxes. The width and height of the fused detection box are the maximum of the width and height of the two fused detection boxes.

[0143] ii. Merge the corresponding label content. For example, you can extract the enhanced label as the label corresponding to the fused detection box. The enhanced label refers to the label content with the richest semantics. You can perform semantic segmentation on the label to obtain noun-adjective pairs and retain the largest number of labels.

[0144] At this point, the first phase of the task, from steps 1-1 to 1-5, has been completed, generating sample data. The sample data includes a sample image, at least one detection box, and at least one class label for the detection box. Each detection box represents a target object. Next, proceed to the second phase of the task.

[0145] 6. Execute step 2-1 and input the filtered box, label, and image into the MLLM to obtain the initial target description information caption. At this time, the caption has many problems, including at least one of the following: incorrect expression, incorrect color description, and incorrect object description. The above problems can be solved by combining semantic segmentation with semantic reorganization.

[0146] 7. Before inputting the filtered box, label, and image into the automatic annotation MLLM, you can detect the size ratio of the detection box in the sample image to the sample image, and determine the subsequent processing flow based on this ratio.

[0147] First, if the size ratio is below a preset threshold, the sample image is cropped into a new sample image of appropriate size based on the position and ratio of the detection frame. This ensures that the sample image used in the annotation process is the most informative, improving the accuracy of the annotation.

[0148] On the contrary, if the size ratio is higher than or equal to the preset threshold, the original sample image is used directly without cropping.

[0149] 8. Execute: Input the cropped sample image, the corresponding detection bounding box, and the category label into the MLLM. The model generates object description information related to the corresponding detection bounding box in the sample image, including the object body, descriptive words, and the relationship between the object body and surrounding objects. To improve the quality of the annotation results, the annotation filtering module and the annotation correction module are introduced.

[0150] Execute step 2-2 and output the visual description information to the annotation filtering module and the annotation correction module. The generated visual description information is preliminarily judged by the annotation filtering module to detect whether there is any inconsistency with the actual sample image content. If there is a difference, it will enter the processing flow of the annotation correction module. The annotation filtering module performs semantic division on the visual description to obtain nouns and adjectives. In addition, relations can also be regarded as adjectives or separated from adjectives. Among them, nouns can refer to the object body of the detection frame, adjectives can be adjective descriptive words, and relations can refer to the relationship between the object subject and the surrounding objects. Semantic division is combined with semantic reorganization to achieve the cleaning of nouns and adjectives, as well as the reorganization of correct nouns and correct adjectives, to obtain accurate visual description information and realize the cleaning of visual description information.

[0151] The annotation correction module is responsible for correcting nouns and adjectives accordingly. This module ensures that the final annotation results more accurately reflect the actual image content. In step 2-3, the annotation correction module inputs the corrected terms into the MLLM, performs semantic reorganization of the description information, and obtains multiple corrected visual descriptions. In step 2-2, the corrected visual descriptions are output to the annotation filtering module. The annotation filtering module filters the multiple corrected descriptions to obtain the final target description. In step 2-4, the filtered target description is output as the final annotation result, completing the entire image annotation process.

[0152] If the visual description information in the previous stage is consistent with the actual image content, execute steps 2-4, directly use the visual description information in the previous stage as the target description information, and output it as the final annotation result to complete the entire image annotation process.

[0153] At this point, the second phase, from steps 2-1 to 2-4, is complete, generating target descriptions for the sample data. This systematic process improves the automation and accuracy of annotation. The resulting annotated data can be used for both MLLM and OWLv2 training.

[0154] The technical solution of this application, through automatic tagging technology, part-of-speech checking and box filtering, and making full use of the calibration process of the Generative Pre-Trained Transformer (GPT), can effectively improve the accuracy of the tagging process and effectively improve the credibility of the automated tagging process.

[0155] According to an embodiment of the present application, Figure 6 is a structural diagram of an apparatus for generating image annotation data in an embodiment of the present application. This embodiment of the present application is applicable to generating image descriptions. The apparatus is implemented using software and / or hardware and is configured in an electronic device with certain data processing capabilities.

[0156] As shown in FIG6 , an image annotation data generating device 600 includes: a sample data acquisition module 601, a visual description verification module 602, an object description information adjustment module 603 and an annotation data generating module 604.

[0157] The sample data acquisition module 601 is configured to acquire a sample image, a target object in the sample image, and visual description information of the target object;

[0158] a visual description verification module 602 configured to verify the words in the visual description information based on the sample image and the target recognition object, and obtain a verification result of the visual description information;

[0159] The target description information adjustment module 603 is configured to adjust the words in the visual description information according to the verification result of the visual description information to obtain the target description information of the target recognition object;

[0160] The annotation data generation module 604 is configured to use the target description information of the target recognition object as the annotation data of the target recognition object; the annotation data is used to train the visual semantic understanding model, and the visual semantic understanding model is used to output the description text of the visual content based on the input visual content.

[0161] According to the technical solution of the present application, the visual description information of the target recognition object is automatically extracted, and the words in the visual description information are verified to obtain the verification result. At the same time, according to the verification result, the words in the visual description information are adjusted to obtain the target description information, which is used as the annotation data of the target recognition object and added to the sample data to train the image semantic understanding model to obtain a model for semantic understanding of the input visual content, thereby improving the generation efficiency of image annotation data, reducing the annotation cost, and taking into account the annotation accuracy, thereby quickly generating correct annotation data to train the model and improve the accuracy of the description text of the visual semantic understanding model.

[0162] In one or more embodiments, the visual description verification module 602 includes: a description segmentation unit, configured to divide the visual description information into at least one visual noun and at least one visual adjective; a noun verification unit, configured to verify the at least one visual noun based on the sample image and the target recognition object, and obtain a noun verification result corresponding to the at least one visual noun; an adjective verification unit, configured to verify the at least one visual adjective based on the sample image and the target recognition object, and obtain an adjective verification result corresponding to the at least one visual adjective; and a verification result determination unit, configured to determine the noun verification result and the adjective verification result as the verification result of the visual description information.

[0163] In one or more embodiments, the noun verification unit includes: a category acquisition subunit, configured to obtain category description information of the target recognition object based on the sample image and the target recognition object; a category verification subunit, configured to detect the consistency between the category description information of the target recognition object and the visual noun for at least one visual noun, and obtain a noun verification result corresponding to the visual noun.

[0164] In one or more embodiments, the adjective verification unit includes: an interaction generation subunit, configured to generate interactive dialogue content for the at least one visual adjective based on the visual adjective and the visual noun corresponding to the visual adjective; an adjective semantic understanding subunit, configured to use the sample image and the target recognition object as context, and input the interactive dialogue content into a large language model to obtain an existence judgment result of the visual adjective; and an adjective verification subunit, configured to generate an adjective verification result of the visual adjective based on the existence judgment result of the visual adjective.

[0165] In one or more embodiments, the target description information adjustment module 603 includes: a noun screening unit, configured to screen the at least one visual noun based on the verification result of each noun to obtain at least one target noun; an adjective screening unit, configured to screen the at least one visual adjective based on the verification result of each adjective to obtain at least one target adjective; a description phrase generation unit, configured to generate at least one description phrase based on the at least one target noun and the at least one target adjective; and a description information generation unit, configured to generate target description information of the target recognition object based on the at least one description phrase.

[0166] In one or more embodiments, the image annotation data generating device further includes: a real-time category acquisition module, configured to acquire real-time updated object categories and add the acquired real-time updated object categories to a dynamic vocabulary; an image acquisition module, configured to acquire a sample image; and an object update acquisition module, configured to identify, in the sample image, a target recognition object corresponding to each word in the dynamic vocabulary, and obtain the target recognition object in the sample image.

[0167] In one or more embodiments, the real-time category acquisition module includes: a visual set acquisition unit, configured to acquire a visual set updated in real time; a target detection unit, configured to perform target detection on the visual data in the visual set to obtain at least one first object category; and / or a semantic understanding unit, configured to perform semantic understanding on the visual data in the visual set to obtain at least one second object category; and a category update unit, configured to determine the first object category and / or the second object category as the real-time updated object category.

[0168] In one or more embodiments, the image annotation data generating device further includes: an intersection-over-union (IoU) detection module, configured to detect the IoU between multiple target recognition objects in the sample image after identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image; a similar object detection module, configured to detect the multiple target recognition objects based on the IoU between the multiple target recognition objects to obtain similar multiple target recognition objects; an object fusion module, configured to fuse the similar multiple target recognition objects to obtain a fusion result; and a redundancy screening module, configured to update the multiple target recognition objects in the sample image based on the fusion result.

[0169] In one or more embodiments, the image annotation data generating device further includes: an intersection-over-union detection module, configured to detect the number of pixels of at least one target recognition object in the sample image after identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image; and a pixel number screening module, configured to update the target recognition object in the sample image based on the number of pixels of the at least one target recognition object.

[0170] In one or more embodiments, the image annotation data generating device further includes: a size ratio detection module, configured to detect the size ratio of at least one target recognition object in the sample image to the sample image after identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image; an object boundary expansion module, configured to expand the boundary of the target recognition object when the size ratio of the target recognition object is less than a preset threshold, so as to increase the size ratio of the expanded target recognition object to the sample image.

[0171] In one or more embodiments, the image annotation data generating device further includes: a size ratio detection module, configured to detect the size ratio of at least one target recognition object in the sample image and the sample image after identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image; an image cropping module, configured to crop the sample image when the size ratio of the target recognition object is less than a preset threshold, and update the sample image to increase the size ratio of the target recognition object and the cropped sample image.

[0172] The above-mentioned image annotation data generation device can execute the image annotation data generation method provided by any embodiment of the present application, and has the corresponding functional modules and effects of executing the image annotation data generation method.

[0173] According to an embodiment of the present application, FIG7 is a structural diagram of a training device for a visual semantic understanding model in an embodiment of the present application. The embodiment of the present application is applicable to training a visual semantic understanding model. The device is implemented using software and / or hardware and is configured in an electronic device with certain data processing capabilities.

[0174] As shown in FIG7 , a visual semantic understanding model training device 700 includes: a sample data acquisition module 701 and a model training module 702 .

[0175] A sample data acquisition module 701 is configured to acquire sample data, where the sample data includes a sample image and annotation data, and the annotation data is generated by the image annotation data generation method described in any embodiment of the present application;

[0176] The model training module 702 is configured to use the sample data to train a visual semantic understanding model.

[0177] According to the technical solution of the present application, by using verified annotated data as training data to train the visual semantic understanding model, the accuracy of the descriptive text output by the visual semantic understanding model can be improved. At the same time, the automatic generation of annotated data can reduce the model training cost and improve the model training efficiency.

[0178] The above-mentioned training device for the visual semantic understanding model can execute the training method for the visual semantic understanding model provided in any embodiment of the present application, and has the corresponding functional modules and effects for executing the training method for the visual semantic understanding model.

[0179] In the technical solution of this application, the collection, storage, use, processing, transmission, provision and disclosure of user personal information involved comply with the provisions of relevant laws and regulations and do not violate public order and good morals.

[0180] According to an embodiment of the present application, the present application also provides an electronic device, a readable storage medium and a computer program product.

[0181] FIG8 shows a schematic area diagram of an example electronic device 800 that can be used to implement an embodiment of the present application. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital processing, cellular phones, smart phones, wearable devices, and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present application described and / or required herein.

[0182] As shown in Figure 8, the electronic device 800 includes a computing unit 801, which can perform various appropriate actions and processes according to a computer program stored in a read-only memory (ROM) 802 or a computer program loaded from a storage unit 808 into a random access memory (RAM) 803. In the RAM 803, various programs and data required for the operation of the electronic device 800 can also be stored. The computing unit 801, the ROM 802, and the RAM 803 are connected to each other via a bus 804. An input / output (I / O) interface 805 is also connected to the bus 804.

[0183] Multiple components in the electronic device 800 are connected to the I / O interface 805, including an input unit 806, such as a keyboard, a mouse, etc.; an output unit 807, such as various types of displays, speakers, etc.; a storage unit 808, such as a magnetic disk, an optical disk, etc.; and a communication unit 809, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 809 allows the electronic device 800 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0184] The computing unit 801 can be a variety of general and / or special processing components with processing and computing capabilities. Some examples of the computing unit 801 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), a variety of dedicated artificial intelligence (AI) computing chips, a variety of computing units that run machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The computing unit 801 performs the multiple methods and processes described above, such as an image annotation data generation method or a visual semantic understanding model training method. For example, in some embodiments, the image annotation data generation method or the visual semantic understanding model training method can be implemented as a computer software program, which is tangibly included in a machine-readable medium, such as a storage unit 808. In some embodiments, part or all of the computer program can be loaded and / or installed on the electronic device 800 via the ROM 802 and / or the communication unit 809. When the computer program is loaded into RAM 803 and executed by the computing unit 801, one or more steps of the method for generating image annotation data or the method for training a visual semantic understanding model described above may be performed. Alternatively, in other embodiments, the computing unit 801 may be configured to perform the method for generating image annotation data or the method for training a visual semantic understanding model by any other appropriate means (e.g., by means of firmware).

[0185] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard parts (ASSPs), system on chips (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0186] The program code for implementing the method of the present application can be written in any combination of one or more programming languages. Such program code can be provided to a processor or controller of a general-purpose computer, a special-purpose computer, or other programmable data processing device so that when the program code is executed by the processor or controller, the functions / instructions specified in the flow chart and / or area diagram are implemented. The program code can be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0187] In the context of the present application, a machine-readable medium can be a tangible medium that can contain or store a program for use by an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. A machine-readable medium can be a machine-readable signal medium or a machine-readable storage medium. A machine-readable medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. A machine-readable storage medium includes an electrical connection based on one or more lines, a portable computer disk, a hard disk, RAM, ROM, erasable programmable read-only memory (EPROM), flash memory, optical fiber, portable compact disc read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination of the foregoing. A storage medium can be a non-transitory storage medium.

[0188] To provide interaction with a user, the systems and techniques described herein can be implemented on a computer having: a display device (e.g., a cathode ray tube (CRT) or a liquid crystal display (LCD) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the computer. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0189] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with embodiments of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: Local Area Networks (LANs), Wide Area Networks (WANs), blockchain networks, and the Internet.

[0190] A computer system may include a client and a server. The client and server are generally remote from each other and typically interact via a communication network. This client-server relationship is established by computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host, a host product within a cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosts and virtual private servers (VPS) services. The server may also be a server in a distributed system or a server integrated with blockchain.

[0191] Artificial intelligence (AI) is the study of how computers can simulate certain human thought processes and intelligent behaviors (such as learning, reasoning, thinking, and planning). It encompasses both hardware and software technologies. AI hardware technologies generally include sensors, specialized AI chips, cloud computing, distributed storage, and big data processing. AI software technologies primarily encompass computer vision, speech recognition, natural language processing, machine learning / deep learning, big data processing, and knowledge graphs.

[0192] Cloud computing refers to a technology system that provides network access to elastically scalable shared pools of physical or virtual resources. These resources can include servers, instruction sets, networks, software, applications, and storage devices, and can be deployed and managed on-demand in a self-service manner. Cloud computing technology provides efficient and powerful data processing capabilities for the application of technologies such as artificial intelligence and blockchain, as well as for model training.

[0193] The various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the multiple steps described in this application can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution provided by this application can be achieved. This document is not limited here.< / apple> < / red> < / apple> < / red>

Claims

1. A method for generating image annotation data, comprising: Acquire a sample image, a target object in the sample image, and visual description information of the target object; Verifying the words in the visual description information according to the sample image and the target recognition object to obtain a verification result of the visual description information; Adjusting the words in the visual description information according to the verification result of the visual description information to obtain target description information of the target recognition object; The target description information of the target recognition object is used as the annotation data of the target recognition object; wherein the annotation data is used to train a visual semantic understanding model, and the visual semantic understanding model is used to output a description text of the visual content based on the input visual content.

2. The method according to claim 1, wherein The verifying the words in the visual description information according to the sample image and the target recognition object to obtain the verification result of the visual description information includes: dividing the visual description information into at least one visual noun and at least one visual adjective; Verifying the at least one visual noun according to the sample image and the target recognition object to obtain a noun verification result corresponding to the at least one visual noun; Verifying the at least one visual adjective according to the sample image and the target recognition object to obtain an adjective verification result corresponding to the at least one visual adjective; The noun verification result and the adjective verification result are determined as the verification result of the visual description information.

3. The method according to claim 2, wherein: Verifying the at least one visual noun based on the sample image and the target recognition object to obtain a noun verification result corresponding to the at least one visual noun includes: Acquiring category description information of the target recognition object according to the sample image and the target recognition object; For the at least one visual noun, the consistency between the category description information of the target recognition object and the at least one visual noun is detected to obtain a noun verification result corresponding to the at least one visual noun.

4. The method according to claim 2, wherein: Verifying the at least one visual adjective according to the sample image and the target recognition object to obtain an adjective verification result corresponding to the at least one visual adjective includes: generating interactive dialogue content for the at least one visual adjective according to the at least one visual adjective and a visual noun corresponding to the at least one visual adjective; Using the sample image and the target recognition object as context, and inputting the interactive dialogue content into a large language model, to obtain a result of determining the presence of the at least one visual adjective; An adjective verification result of the at least one visual adjective is generated according to the existence determination result of the at least one visual adjective.

5. The method according to claim 2, wherein: The step of adjusting the words in the visual description information according to the verification result of the visual description information to obtain the target description information of the target recognition object includes: According to the verification result of each noun, screening the at least one visual noun to obtain at least one target noun; According to the verification result of each adjective, screening the at least one visual adjective to obtain at least one target adjective; generating at least one descriptive phrase based on the at least one target noun and the at least one target adjective; Generate target description information of the target recognition object according to the at least one description phrase.

6. The method according to claim 1, wherein The sample image and the target recognition object are obtained in the following manner: Acquire the real-time updated object categories, and add the acquired real-time updated object categories to the dynamic vocabulary; acquiring the sample image; The target recognition object corresponding to each word in the dynamic vocabulary is identified in the sample image to obtain the target recognition object in the sample image.

7. The method according to claim 6, wherein: The object categories for obtaining real-time updates include: Get a visual set that updates in real time; performing object detection on the visual data in the visual set to obtain at least one first object category; or performing semantic understanding on the visual data in the visual set to obtain at least one second object category; or performing object detection on the visual data in the visual set to obtain at least one first object category and performing semantic understanding on the visual data in the visual set to obtain at least one second object category; Determine at least one of the first object category or the second object category as a real-time updated Object category.

8. The method according to claim 6, after identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image, further comprising: Detecting the intersection-over-union ratio between a plurality of target recognition objects in the sample image; Detecting the multiple target recognition objects according to the intersection-over-union ratio between the multiple target recognition objects to obtain multiple similar target recognition objects; fusing the multiple similar target recognition objects to obtain a fusion result; A plurality of target recognition objects in the sample image are updated according to the fusion result.

9. The method according to claim 6, after identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image, further comprising: Detecting the number of pixels of at least one target recognition object in the sample image; The target recognition object in the sample image is updated according to the number of pixels of the at least one target recognition object.

10. The method according to claim 6, wherein: After identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image, the method further includes: Detecting a size ratio between at least one target recognition object in the sample image and the sample image; When the size ratio of the target recognition object is smaller than a preset threshold, the boundary of the target recognition object is expanded to increase the size ratio between the expanded target recognition object and the sample image.

11. The method according to claim 6, after identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image, further comprising: Detecting a size ratio between at least one target recognition object in the sample image and the sample image; When the size ratio of the target recognition object is smaller than a preset threshold, the sample image is cropped and the sample image is updated to increase the size ratio between the target recognition object and the cropped sample image.

12. A method for training a visual semantic understanding model, comprising: Acquire sample data, wherein the sample data includes a sample image and annotation data, and the annotation data is generated by the image annotation data generation method according to any one of claims 1 to 11; The sample data is used to train a visual semantic understanding model.

13. A device for generating image annotation data, comprising: A sample data acquisition module is configured to acquire a sample image, a target object in the sample image, and visual description information of the target object; a visual description verification module configured to verify the words in the visual description information based on the sample image and the target recognition object, and obtain a verification result of the visual description information; a target description information adjustment module, configured to adjust the words in the visual description information according to the verification result of the visual description information to obtain the target description information of the target recognition object; The annotation data generation module is configured to use the target description information of the target recognition object as the annotation data of the target recognition object; wherein the annotation data is used to train the visual semantic understanding model, and the visual semantic understanding model is used to output the description text of the visual content based on the input visual content.

14. The device according to claim 13, wherein The visual description verification module includes: a description segmentation unit configured to divide the visual description information into at least one visual noun and at least one visual adjective; a noun verification unit, configured to verify the at least one visual noun based on the sample image and the target recognition object, and obtain a noun verification result corresponding to the at least one visual noun; an adjective verification unit, configured to verify the at least one visual adjective based on the sample image and the target recognition object, and obtain an adjective verification result corresponding to the at least one visual adjective; The verification result determining unit is configured to determine the noun verification result and the adjective verification result as the verification result of the visual description information.

15. The device according to claim 14, wherein The noun verification unit includes: a category acquisition subunit, configured to acquire category description information of the target recognition object based on the sample image and the target recognition object; The category verification subunit is configured to detect the consistency between the category description information of the target recognition object and the at least one visual noun, and obtain a noun verification result corresponding to the at least one visual noun.

16. The device according to claim 14, wherein The adjective verification unit includes: an interaction generating subunit configured to generate interactive dialogue content for the at least one visual adjective according to the at least one visual adjective and a visual noun corresponding to the at least one visual adjective; an adjective semantic understanding subunit, configured to use the sample image and the target recognition object as context, and input the interactive dialogue content into a large language model to obtain a result of determining the existence of the at least one visual adjective; The adjective verification subunit is configured to generate an adjective verification result of the at least one visual adjective according to the existence judgment result of the at least one visual adjective.

17. The device according to claim 14, wherein The target description information adjustment module includes: a noun screening unit configured to screen the at least one visual noun based on the verification result of each noun to obtain at least one target noun; an adjective screening unit configured to screen the at least one visual adjective according to the verification result of each adjective to obtain at least one target adjective; a description phrase generating unit configured to generate at least one description phrase based on the at least one target noun and the at least one target adjective; The description information generating unit is configured to generate target description information of the target recognition object according to the at least one description phrase.

18. The apparatus according to claim 13, further comprising: A real-time category acquisition module is configured to acquire the real-time updated object categories and add the acquired real-time updated object categories to the dynamic vocabulary; An image acquisition module, configured to acquire the sample image; The object update acquisition module is configured to identify the target recognition object corresponding to each word in the dynamic vocabulary in the sample image, and obtain the target recognition object in the sample image.

19. The device according to claim 18, wherein The real-time category acquisition module includes: A visual set acquisition unit, configured to acquire a visual set updated in real time; At least one of an object detection unit or a semantic understanding unit: wherein the object detection unit is configured to perform object detection on the visual data in the visual set to obtain at least one first object category; and the semantic understanding unit is configured to perform semantic understanding on the visual data in the visual set to obtain at least one second object category; The category updating unit is configured to determine at least one of the first object category or the second object category as an object category to be updated in real time.

20. The apparatus according to claim 18, further comprising: An IoU detection module configured to identify each word in the dynamic vocabulary in the sample image After the corresponding target recognition objects are obtained, the intersection-over-union ratios between the multiple target recognition objects in the sample image are detected; a similar object detection module configured to detect the plurality of target recognition objects based on an intersection-over-union ratio between the plurality of target recognition objects to obtain a plurality of similar target recognition objects; An object fusion module is configured to fuse the multiple similar target recognition objects to obtain a fusion result; The redundancy screening module is configured to update multiple target recognition objects in the sample image according to the fusion result.

21. The apparatus according to claim 18, further comprising: an IoU detection module configured to detect the number of pixels of at least one target recognition object in the sample image after identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image; The pixel number screening module is configured to update the target recognition object in the sample image according to the pixel number of the at least one target recognition object.

22. The apparatus of claim 18, further comprising: a size ratio detection module configured to detect a size ratio between at least one target recognition object in the sample image and the sample image after identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image; The object boundary expansion module is configured to expand the boundary of the target recognition object when the size ratio of the target recognition object is less than a preset threshold, so as to increase the size ratio between the expanded target recognition object and the sample image.

23. The apparatus of claim 18, further comprising: a size ratio detection module configured to detect a size ratio between at least one target recognition object in the sample image and the sample image after identifying the target recognition object corresponding to each word in the dynamic vocabulary in the sample image; The image cropping module is configured to crop the sample image when the size ratio of the target recognition object is less than a preset threshold, and update the sample image to increase the size ratio between the target recognition object and the cropped sample image.

24. A training device for a visual semantic understanding model, comprising: a sample data acquisition module configured to acquire sample data, wherein the sample data includes a sample image and annotation data, and the annotation data is generated by the image annotation data generation method according to any one of claims 1 to 11; The model training module is configured to use the sample data to train a visual semantic understanding model.

25. An electronic device comprising: at least one processor; as well as a memory communicatively connected to the at least one processor; wherein, The memory stores instructions that can be executed by the at least one processor, and the instructions are executed by the at least one processor so that the at least one processor can execute the image annotation data generation method described in any one of claims 1 to 11, or the visual semantic understanding model training method described in claim 12.

26. A non-transitory computer-readable storage medium storing computer instructions, wherein: The computer instructions are used to enable a computer to execute the method for generating image annotation data according to any one of claims 1 to 11, or the method for training a visual semantic understanding model according to claim 12.

27. A computer program product, comprising a computer program, which, when executed by a processor, implements the method for generating image annotation data according to any one of claims 1 to 11, or the method for training a visual semantic understanding model according to claim 12.