Image labeling method and device, computer readable storage medium and electronic device
By combining multiple label description generation models and large language models, image labels for game resource libraries are automatically generated, solving the problem of low efficiency in traditional manual annotation and achieving efficient and accurate image annotation.
Patent Information
- Application Number
- CN202311708797.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-12
- Publication Date
- 2026-08-25
- Estimated Expiration
- 2043-12-12
AI Technical Summary
Traditional game resource library tags require a lot of manual processing, resulting in low tag generation efficiency and subjectivity and inconsistency.
Multiple label description generation models are used to generate descriptions for target images. A large language model is used to extract label description information, automatically generate target labels, and annotate the images.
It enables automated generation of image labels, improving label generation efficiency, reducing manual intervention, and enhancing the accuracy and consistency of labels.
Smart Images

Figure CN117883784B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of games, and more specifically, to an image annotation method and apparatus, a computer-readable storage medium, and an electronic device. Background Technology
[0002] Currently, game resource library tags in traditional games are manually categorized, requiring significant manual processing and specialized knowledge. The traditional manual tagging process includes: data preparation, establishing tagging standards, training taggers, tagging data, tagging quality control, data cleaning and integration, data verification and evaluation, and data publication and updates. Therefore, the traditional manual tagging process is highly complex, requires substantial time and manpower, and the results may be subjective and inconsistent, leading to low tag generation efficiency.
[0003] There is currently no effective solution to the above problems. Summary of the Invention
[0004] This disclosure provides at least some embodiments of an image annotation method and apparatus, a computer-readable storage medium and an electronic device, to at least solve the technical problem that the prior art relies on manual image annotation, resulting in low label generation efficiency.
[0005] According to one embodiment of this disclosure, an image annotation method is provided, comprising: acquiring a target image; inputting the target image into multiple label description generation models, performing label description generation on the target image respectively, and obtaining label description information of the target image, wherein the label description information is used to characterize information describing the labels of the target image; extracting the label description information using a large language model to obtain target labels of the target image; and annotating the target image based on the target labels.
[0006] According to one embodiment of this disclosure, a label generation apparatus is also provided, comprising: an image acquisition module for acquiring a target image; a label description generation module for inputting the target image into multiple label description generation models, performing label description generation on the target image respectively, and obtaining label description information of the target image, wherein the label description information is used to characterize information describing the labels of the target image; a label extraction module for extracting the label description information using a large language model to obtain target labels of the target image; and an image annotation module for annotating the target image based on the target labels.
[0007] According to one embodiment of the present disclosure, a computer-readable storage medium is also provided, wherein a computer program is stored in the computer program, wherein the computer program is configured to execute the image annotation method of any of the above claims when it is run.
[0008] According to one embodiment of this disclosure, an electronic device is also provided, including a memory and a processor, wherein the memory stores a computer program and the processor is configured to run the computer program to perform the image annotation method of any of the above claims.
[0009] In at least some embodiments of this disclosure, a target image is acquired; the target image is input into multiple label description generation models, and label description generation is performed on the target image respectively to obtain label description information of the target image, wherein the label description information is used to characterize the information describing the label of the target image; the label description information is extracted using a large language model to obtain the target label of the target image; and the target image is labeled based on the target label. It should be noted that by using multiple label description generation models to generate label description information of the target image respectively, and using a large language model to extract the label description information to obtain the target label, the target label matching the target image can be accurately and efficiently obtained, thereby improving the label generation efficiency and realizing the technical effect of automatically generating labels from the target image and labeling the target image, thus solving the technical problem of low label generation efficiency caused by manual image labeling in the prior art. Attached Figure Description
[0010] The accompanying drawings, which are included to provide a further understanding of this disclosure and form part of this application, illustrate exemplary embodiments of this disclosure and are used to explain this disclosure, but do not constitute an undue limitation of this disclosure. In the drawings:
[0011] Figure 1 This is a hardware structure block diagram of a mobile terminal for an image annotation method according to an embodiment of this disclosure;
[0012] Figure 2 This is a flowchart of an image annotation method according to one embodiment of the present disclosure;
[0013] Figure 3 This is a schematic diagram of an optional tagging page according to one embodiment of the present disclosure;
[0014] Figure 4 This is a schematic diagram of an optional homepage according to one embodiment of this disclosure;
[0015] Figure 5 This is a schematic diagram of an optional preset tag library according to one embodiment of the present disclosure;
[0016] Figure 6 This is a schematic diagram of an optional search page according to one embodiment of the present disclosure;
[0017] Figure 7 This is a schematic diagram of an optional label confirmation interface according to one embodiment of the present disclosure;
[0018] Figure 8 This is a flowchart of an optional storage target tag according to one embodiment of the present disclosure;
[0019] Figure 9 This is a schematic diagram of an optional multilingual library page according to one embodiment of this disclosure;
[0020] Figure 10 This is a flowchart of an optional translation target label according to one embodiment of the present disclosure;
[0021] Figure 11 This is a flowchart of an optional image annotation method according to an embodiment of the present disclosure;
[0022] Figure 12 This is a structural block diagram of an image annotation apparatus according to one embodiment of the present disclosure;
[0023] Figure 13 This is a schematic diagram of an electronic device according to one embodiment of the present disclosure. Detailed Implementation
[0024] To enable those skilled in the art to better understand the present disclosure, the technical solutions of the present disclosure will be clearly and completely described below with reference to the accompanying drawings of the embodiments. Obviously, the described embodiments are only some embodiments of the present disclosure, and not all embodiments. Based on the embodiments of the present disclosure, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present disclosure.
[0025] It should be noted that the terms "first," "second," etc., in the specification, claims, and accompanying drawings of this disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this disclosure described herein can be implemented in orders other than those illustrated or described herein. Furthermore, the terms "comprising" and "having," and any variations thereof, are intended to cover non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0026] In one possible implementation, addressing the persistent problem of low tag generation efficiency in commonly used computer science methods, which the inventors, after practical experience and careful research, propose an image annotation method. This method addresses the issue of low tag generation efficiency caused by manual image annotation in existing technologies. Specifically, it applies to games involving virtual data tagging, typically those with databases or resource libraries. The method involves acquiring a target image; inputting the target image into multiple tag description generation models, each generating a tag description to obtain tag description information for the target image. This tag description information characterizes the information describing the target image; extracting the tag description information using a large language model to obtain target tags for the target image; and then annotating the target image based on these target tags. It is noteworthy that by using multiple tag description generation models to generate tag description information for the target image and then extracting the target tags using a large language model, the method accurately and efficiently obtains target tags matching the target image, thereby improving tag generation efficiency. This achieves the technical effect of automatically generating tags from and annotating target images, thus solving the problem of low tag generation efficiency caused by manual image annotation in existing technologies.
[0027] The methods and embodiments described above in this disclosure can be executed on mobile terminals, computer terminals, or similar computing devices. Taking a mobile terminal as an example, the mobile terminal can be a smartphone, tablet computer, PDA, mobile internet device, PAD, game console, or other terminal device. Figure 1 This is a hardware structure block diagram of a mobile terminal for an image annotation method according to an embodiment of this disclosure. For example... Figure 1 As shown, a mobile terminal may include one or more ( Figure 1 Only one is shown in the diagram. Processor 102 (processor 102 may include, but is not limited to, a central processing unit (CPU), graphics processing unit (GPU), digital signal processing (DSP) chip, microprocessor (MCU), programmable logic device (FPGA), neural network processor (NPU), tensor processor (TPU), artificial intelligence (AI) type processor, etc.) and memory 104 for storing data. In one embodiment of this disclosure, it may also include: input / output device 108 and display device 110.
[0028] In some optional embodiments primarily focused on gaming scenarios, the aforementioned device may also provide a human-computer interaction interface with a touch-sensitive surface. This interface can sense finger contact and / or gestures to interact with a graphical user interface (GUI). The human-computer interaction functions may include the following: creating web pages, drawing, word processing, creating electronic documents, playing games, video conferencing, instant messaging, sending and receiving emails, call interfaces, playing digital videos, playing digital music, and / or web browsing, etc. Executable instructions for performing the aforementioned human-computer interaction functions are configured / stored in one or more processor-executable computer program products or readable storage media.
[0029] Those skilled in the art will understand that Figure 1 The structure shown is for illustrative purposes only and does not limit the structure of the mobile terminal described above. For example, the mobile terminal may also include components that are more... Figure 1 The more or fewer components shown, or having the same Figure 1 The different configurations shown.
[0030] According to one embodiment of this disclosure, an embodiment of an image annotation method is provided. It should be noted that the steps shown in the flowchart in the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. Furthermore, although a logical order is shown in the flowchart, in some cases, the steps shown or described may be executed in a different order than that shown here.
[0031] In one possible implementation, this disclosure provides an image annotation method that provides a graphical user interface through a terminal device, wherein the terminal device may be the aforementioned local terminal device or a client device in the aforementioned cloud interactive system. Figure 2 This is a flowchart of an image annotation method according to one embodiment of the present disclosure. A graphical user interface is provided through a terminal device. The content displayed by the graphical user interface includes at least images of multiple virtual objects and a first operation control, such as... Figure 2 As shown, the method includes the following steps:
[0032] Step S202: Obtain the target image.
[0033] The target image mentioned above can be an image that needs to be labeled.
[0034] In one optional embodiment, the user selects images of multiple virtual objects as needed. These images can be multiple pieces of virtual data stored in a database or resource library, including but not limited to: labeled images and unlabeled images. The virtual objects can be virtual assets appearing in the game, and can be, but are not limited to: boxes, ginkgo trees, cats, thatched huts, and chimneys. Thus, the image to be labeled is determined from the multiple virtual objects as the target image. In this embodiment, Figure 3 This is a schematic diagram of an optional tagging page according to one embodiment of the present disclosure, such as... Figure 3 As shown, the left side contains images of multiple virtual objects, such as Image 1, Image 2, Image 3, Image 5, Image 6, and Image 7. Users can directly select and click on the images they want to label on the left side to confirm them as target images. Users can click the "select all" box to confirm that all images are selected for labeling, and can also adjust the scaling ratio of multiple virtual object images as needed.
[0035] Step S204: Input the target image into multiple label description generation models, perform label description generation on the target image respectively, and obtain the label description information of the target image. The label description information is used to characterize the information describing the labels of the target image.
[0036] The aforementioned multiple label description generation models can be multiple models capable of describing target images and generating corresponding labels, and can be, but are not limited to: PaddlePaddle version model (Danbooru), distributed version control system (Git), multimodal translation model (Blip), and multimodal image model (Coca).
[0037] The aforementioned label description information can be information that describes the target image.
[0038] In one optional embodiment, multiple different label description generation models are used to describe the target image and generate description labels, thus obtaining label description information generated by multiple label description generation models. For example, a photo of a person wearing armor is input into multiple label description generation models. The multiple label description generation models describe the target image and generate corresponding label description information. For example, the label description information generated by the git-large model after fine-tuning coco is "Emperor's Golden Armor"; the label description information generated by the git-large model after fine-tuning the text is "A man wearing clothing with the words 'on it'"; the label description information generated by the Blip-large model is "A man wearing armor"; the label description information generated by the coca model is "A person wearing clothing holding a sword"; and the label description information generated by Blip-2-OPT6.7b is "A person wearing clothing standing on the ground".
[0039] Step S206: Use a large language model to extract the label description information to obtain the target label of the target image.
[0040] The large-scale language model mentioned above can be a language model that can convert label description information into labels, and can be, but is not limited to, a generative pretrained transform model (Chat Generative Pretrained Transformer, or Chat GPT for short).
[0041] The target labels mentioned above can be labels used to annotate target images.
[0042] In one optional embodiment, tag description information generated by multiple tag description generation models is collected, and a large language model is used to extract the tag description information to obtain tags related to the target image. In another optional embodiment, tag description information generated by multiple tag description generation models is collected, and a large language model determines the key text in the tag description information based on the overall semantics of the tag description information, and then converts the key text into corresponding tags.
[0043] Because there is a lot of tag description information, the number of tags that can be converted is also large. To improve the accuracy of the target tags, we can select tags with higher accuracy from the converted tags and use them as target tags. Optionally, we can sort the tags according to their frequency of occurrence in the tag description information, select the tags with the highest frequency as target tags, and store the remaining tags for selection when manually modifying the tags. For example, for a photo of a person wearing armor, we can identify "Chinese emperor's clothing," "golden armor," and "royal attire" as target tags, and store "emperor," "standing," and "white background" as other tags for backup.
[0044] Step S208: Label the target image based on the target label.
[0045] In an alternative embodiment, the target image can be labeled according to a target tag, for example, the target tag can be displayed below the target image.
[0046] In another optional embodiment, since the target image may have been previously manually labeled or tagged using intelligent language tools, these historical tags need to be considered during the labeling process. Optionally, after obtaining the target tag, historical tags of the target image can be obtained. These historical tags can be tags that were previously used to annotate the target image, and can be, but are not limited to, manually annotated tags or tags used to annotate the target image in the past. The target tag and historical tags are sorted simultaneously, and tags with a high correlation to the target image are determined as first tags, while tags with a low correlation to the target image are determined as second tags. The target image is then labeled using the first tags, and the second tags are stored to provide a reference for users to modify the tags of the target image.
[0047] In this embodiment of the invention, a target image is acquired; the target image is input into multiple label description generation models, and label descriptions are generated for the target image respectively to obtain label description information of the target image. The label description information is used to characterize the information describing the labels of the target image; a large language model is used to extract the label description information to obtain the target label of the target image; and the target image is labeled based on the target label. It is worth noting that by using multiple label description generation models to generate label description information for the target image respectively, and then using a large language model to extract the label description information to obtain the target label, the target label matching the target image can be accurately and efficiently obtained. This achieves the goal of improving label generation efficiency and realizes the technical effect of automatically generating labels from the target image and labeling the target image, thereby solving the technical problem of low label generation efficiency caused by manual image labeling in the prior art.
[0048] It should be noted that, Figure 4 This is a schematic diagram of an optional homepage according to one embodiment of this disclosure, such as... Figure 4 As shown, users can select intelligent tagging and multilingual libraries as needed, and enter the corresponding interface by inputting the path to the multilingual table, the path to the model master table, and the path to the image folder.
[0049] Optionally, a large language model is used to extract label description information to obtain target labels for the target image, including: extracting multiple initial labels from the label description information using a large language model; and selecting target labels from the multiple initial labels based on their frequency of occurrence in the label description information.
[0050] The aforementioned initial labels can be obtained by transforming the label description information using a large language model.
[0051] In one alternative embodiment, label description information can be input into a large language model. The large language model can treat labels with similar meanings as the same label to obtain multiple initial labels. Then, the labels can be sorted from high to low according to their frequency of occurrence, thereby filtering out target labels from multiple initial labels. This enables the rapid extraction of effective target labels through a large language model, improves the accuracy of target labels, and makes the target labels more correlated with the target image.
[0052] For example, you could input the following into a large language model: "Please extract the following tag descriptions to obtain multiple tags. Merge tags with similar meanings and sort them by frequency of occurrence, obtaining the top three tags. Place the remaining tags in the 'Other Tags' section. Please answer in the following format."
[0053] 1. The first tag
[0054] 2. The second tag
[0055] 3. The third tag
[0056] Other tags:
[0057] 1. The fourth tag
[0058] 2. The fifth tag
[0059] 3. The sixth tag
[0060] Finally, the label was translated into Chinese.
[0061] The tag description information is as follows:
[0062] "The Emperor's Golden Armor"
[0063] "A man wearing clothing with the words 'on it' on it"
[0064] "A man dressed in armor"
[0065] "A person dressed in clothes holding a sword."
[0066] A person dressed in clothes is standing on the ground.
[0067] The target labels of the responses from the large language model are then obtained, as follows:
[0068] 1. Clothing of Chinese Emperors
[0069] 2. Golden Armor
[0070] 3. Royal Attire
[0071] Other tags:
[0072] 1. Giraffe
[0073] 2. Toys
[0074] 3. Blue tail
[0075] Optionally, a target tag is obtained by filtering from multiple initial tags based on the frequency of occurrence of multiple initial tags in the tag description information, including: filtering out initial tags whose occurrence frequency is greater than a preset occurrence frequency based on the frequency of occurrence of multiple initial tags in the tag description information, and obtaining the target tag.
[0076] The preset frequency of occurrence mentioned above can be set in advance according to specific needs to filter the frequency of occurrence of multiple initial tags.
[0077] In one optional embodiment, a large language model sorts multiple initial labels according to their frequency of occurrence, and uses the initial labels with a frequency greater than a preset frequency as target labels, and uses the remaining initial labels as other labels for users to refer to when modifying labels. This makes the target labels more relevant to the target image, the target labels more accurate, and also facilitates subsequent label modifications.
[0078] Optionally, the method further includes: responding to an input command on a graphical user interface, obtaining a selection ratio corresponding to the input command, wherein the selection ratio is used to characterize the proportion of the target label among multiple initial labels; and determining a preset occurrence frequency based on the selection ratio.
[0079] The input instructions mentioned above can be operation instructions corresponding to the selection ratio entered by the user. For example, there is a text input box or a drop-down menu on the graphical user interface, in which the user can select the selection ratio.
[0080] The selection ratio mentioned above can be set according to specific circumstances to select the proportion of labels that are highly relevant to the target image, and can be, but is not limited to, 70%, 80%, or 90%.
[0081] In an alternative embodiment, when the user needs to set the number of target tags, such as Figure 3 As shown, the target image is displayed on the right to confirm the image currently being labeled. There are operation controls for intelligent label generation at the top right and the selection ratio can be set below according to user needs, thereby determining the preset occurrence frequency based on the selection ratio.
[0082] For example, if a user needs to obtain initial tags with a selection rate of more than 60% as target tags, the user only needs to enter 60% in the selection rate input box. If the tag description information contains 30 description texts, then the initial tags with an occurrence frequency of more than 60% can be determined as target tags. That is, the initial tags that appear in 18 description texts are target tags, and the remaining initial tags are stored as other tags.
[0083] Optionally, labeling the target image based on the target label includes: obtaining historical labels of the target image; classifying the historical labels and target labels based on the allocation configuration for the historical labels and target labels to obtain a first label and a second label, wherein the first label is used to label the target image and the second label is used to update the first label; labeling the target image based on the first label and storing the second label.
[0084] The first label mentioned above can be a default label that is highly relevant to the target image and provides a more accurate description, and is used to annotate the target image.
[0085] The second tag mentioned above can be a backup tag with a lower relevance compared to the first tag, used as a reference for subsequent manual modification of the tag.
[0086] In an optional embodiment, since the corresponding labels may have been pre-input by a human based on the target image, but the manually input labels may not be correct, the manually annotated labels can be confirmed as either the first label or the second label, that is, the labels generated by the large language model can be used as either the second label or the first label. Figure 3As shown, the upper right side features a smart label generation control panel with four options: 1. Retain historical labels to default, with the target label as a backup; 2. Retain historical labels to default, and add the target label to default; 3. Retain historical labels to backup, with the target label as default; 4. Discard historical labels, and use the target label as default. Users can configure these settings to categorize historical and target labels, obtain corresponding first and second labels, and then use the first label to annotate the target image. The second label is stored for manual label modification, thus improving the labeling accuracy of the target image.
[0087] For example, if the historical tags contain user-annotated labels for the target image, and the user confirms that the user-annotated labels are correct, selecting "Keep historical tags to default, target tag as backup" will determine the historical tags as the first tag and the target tag as the second tag. If the historical tags contain user-annotated labels for the target image, and the user believes the user-annotated labels are correct and also wants to add the target tag to the first tag, selecting "Keep historical tags to default, add target tag to default" will determine both the historical tags and the target tag as the first tag. If the historical tags contain user-annotated labels for the target image, and the user believes the user-annotated labels are not entirely correct, selecting "Keep historical tags to backup, target tag as default" will determine the historical tags as the second tag and the target tag as the first tag. If the historical tags contain user-annotated labels for the target image, and the user believes the user-annotated labels are incorrect, selecting "Discard historical tags, use target tag as default" will determine the target tag as the first tag.
[0088] Optionally, the first label is displayed in the first display area of the graphical user interface, and the second label is displayed in the second display area of the graphical user interface.
[0089] The first display area mentioned above can be an area in the image user interface that displays labels that are highly relevant to the target image.
[0090] The second display area mentioned above can be an area in the image user interface that displays labels that are less relevant to the target image.
[0091] In one alternative embodiment, such as Figure 3As shown, the right side center includes a first display area for a first label and a second display area for a second label. The first display area consists of multiple label boxes within the first label display frame; the second display area consists of multiple label boxes within the second label display frame. The first and second labels can be displayed correspondingly in the first and second display areas. The first and second labels can be displayed from top to bottom or from left to right according to their frequency of appearance, facilitating the display of the first and second labels and allowing the user to make a selection.
[0092] Optionally, the method further includes: determining whether a target label exists in a preset label library; in response to the existence of a target label in the preset label library, displaying the target label at a first position in the target image according to a first display method; and in response to the absence of a target label in the preset label library, displaying the target label at the first position according to a second display method.
[0093] The aforementioned preset tag library can be a database that is pre-configured to store all tags according to requirements.
[0094] The first display method mentioned above can be a method of pre-setting the display of the target label, which may include, but is not limited to: setting the size, font color, and font weight of the target label.
[0095] The first position mentioned above can be a location on the target image used to display the target label, for example, it can be below the target image, but it is not limited to this.
[0096] The second display method mentioned above can be a display method different from the first display method.
[0097] In one optional embodiment, after obtaining the target tag, a scanning algorithm is used to scan a preset tag library to determine whether the target tag exists in the preset tag library. If the target tag exists in the preset tag library, the target tag is displayed in a first display mode in a first position, such as making the target tag bold, setting the font size to 18, and displaying the target tag in red. If the target tag does not exist in the preset tag library, the target tag can be displayed in a second display mode in the first position, such as not making the target tag bold, setting the font color to gray, and setting the font size to 10.
[0098] In another optional embodiment, after obtaining the target tag, a large language model is used to determine the major and minor categories of the target image based on the target tag. For example, if the target tag is "treasure chest," then the major category is "container"; and the minor category is "functional item." A search is then performed in a preset tag library according to the major and minor categories to determine if the target tag exists in the library. If the target tag exists, it is directly used to label the target image; if the target tag does not exist, it is stored after the corresponding category according to the major and minor categories, and the target image is then labeled using the target tag.
[0099] Specifically, such as Figure 3 As shown, the bottom right corner displays a preset tag library, which allows you to quickly determine if the target tag exists. Furthermore, Figure 5 This is a schematic diagram of an optional preset tag library according to one embodiment of the present disclosure, such as... Figure 5 As shown, a preset tag library is generated based on the tag category. Broad categories can include, but are not limited to: common, style, environment, facial expressions / features, theme, hairstyle, clothing, etc. Subcategories can include, but are not limited to: realistic, cartoon, fantasy, modern, traditional Chinese style, science fiction, architecture, etc. Subcategories can be further refined; for example, architecture can be further subdivided into building as a whole, building components, land plots, etc. Corresponding tags are then placed after the category. For example, the cartoon category includes tags such as Western cartoons, pixel art, and anime; the fantasy category includes tags such as fantasy, medieval, desert civilization, and Nordic civilization; the modern category includes tags such as western, European, urban, and diary; and the traditional Chinese style category includes... The system includes tags such as Three Kingdoms, Journey to the West, Wuxia, Xuanhuan, and Xianxia; the science fiction category includes tags such as post-apocalyptic, cyberpunk, steampunk, space opera, and Cthulhu; the overall architecture category includes tags such as towers, training grounds, defensive towers, barracks, bases, stone buildings, attics, archways, mud houses, sacrificial altars, tin sheds, workshops, warehouses, teachers, stone houses, mines, stone seats, docks, furnaces, kitchens, watchtowers, and bridges; the building component category includes tags such as walls, doors, windows, ground, pillars, roofs, platforms, chimneys, bases, brackets, curtains, bricks, floors, reliefs, construction scaffolds, ruins, handrails, and passageways; and the land category includes tags such as brick plots and pedestrian crossings. This allows for rapid determination of whether the target tag exists in the preset tag library based on the tag category, improving the tag generation rate. Users can directly input the target tag into a large language model to obtain the target tag's major and minor categories and corresponding tags. For example, if the target label is entered, the classification result will be as follows:
[0100] 1. Chinese Emperor's Costume
[0101] Major category: Chinese style;
[0102] Subcategory: Ancient;
[0103] Tag: Clothing.
[0104] 2. Golden Armor
[0105] Main category: Fantasy;
[0106] Subcategory: Fantasy;
[0107] Tag: Armor.
[0108] 3. Royal attire
[0109] Major category: Europe;
[0110] Subcategory: Ancient;
[0111] Tag: Clothing.
[0112] 4. Giraffe
[0113] Major category: Animals;
[0114] Subcategory: Land;
[0115] Tag: Wild animals.
[0116] 5. Toy
[0117] Main category: Modern;
[0118] Subcategory: Toys;
[0119] Tags: children's toys.
[0120] Optionally, the method further includes: displaying a search page in the graphical user interface in response to a preset operation performed on the graphical user interface, wherein the search page is used to search for images of multiple virtual objects; and determining images of multiple virtual objects corresponding to the search operation from a set of images of virtual objects in response to a search operation performed on the search page.
[0121] The aforementioned preset operations can be operations that are pre-set according to specific circumstances to display the search page, such as clicking the right mouse button, but are not limited to this.
[0122] The search page mentioned above can be a page that allows users to search for the data or information they need by entering or selecting.
[0123] The search operation described above can be performed by the user entering or selecting information on the search interface.
[0124] In one optional embodiment, in response to a user needing to retrieve a target image from images of multiple virtual objects, a search page can be accessed by right-clicking the mouse or using a keyboard shortcut. Figure 6 This is a schematic diagram of an optional search page according to one embodiment of the present disclosure, such as... Figure 6 As shown, the search page includes a search bar for search tags / object names, pending tags, confirmed tags, animals, and buildings. Pending tags further include tagged, untagged, and tagged with changes. Users can select from the search page as needed, and the search engine will then select multiple images of virtual objects that match the user's chosen criteria from a corresponding set of images. Using the search page to search for images of multiple desired virtual objects allows for rapid acquisition and processing of these images, improving the efficiency of tag processing.
[0125] Optionally, the method further includes: obtaining flag information corresponding to the images of multiple virtual objects, wherein the flag information is used to characterize whether the images of multiple virtual objects are labeled, or whether the labels of the images of multiple virtual objects are modified or confirmed; generating target subscript elements corresponding to the images of multiple virtual objects based on the flag information; and displaying the target subscript elements at a second position in the images of multiple virtual objects.
[0126] The aforementioned flag information can be used to indicate whether the label of the target image is correct.
[0127] The aforementioned target subscript element can be a subscript element used to indicate the labeling status of the target image. The target subscript element can be used to determine whether the label of the target image has been modified, whether there is a label, or whether it has been modified.
[0128] In one optional embodiment, flag information corresponding to the images of multiple virtual objects is obtained. Table 1 is an optional model summary report according to one embodiment of the present disclosure, as shown in Table 1 below. The model summary report includes image code (Identity Document, abbreviated as ID), target label, target image and flag information. When the flag information is 1, it indicates that the label corresponding to the image has been confirmed as correct or the labels of the images of multiple virtual objects have been confirmed as modified. When the flag is 0, it indicates that the label corresponding to the image has not been confirmed as correct or the labels of the images of multiple virtual objects have not been modified.
[0129] Table 1 Model Summary Report
[0130]
[0131] Therefore, target index elements corresponding to the images of multiple virtual objects are generated based on the flag information, and the target index elements are displayed in the second position of the images of multiple virtual objects, such as... Figure 3 As shown, the upper left corner of multiple images on the left displays target label elements to indicate whether the images of multiple virtual objects have been labeled and whether the labeling is correct. A circle in the upper right corner indicates a labeled image; an image without a circle in the upper right corner is unlabeled; a solid black circle indicates that the labeling of that image has been confirmed as correct; and an empty circle in the upper right corner indicates that the labeling of that image has not been confirmed as correct. Furthermore, a solid red target label element indicates that the target label has been modified or is missing, while a solid yellow target label element indicates that the target label has not been modified. This allows users to immediately see whether a target image has a target label and whether the target label has been modified, improving label processing efficiency.
[0132] Optionally, in response to the number of target images being greater than a preset number, the method further includes: in response to the processing operation performed on the target images, determining the operation mode and operation label corresponding to the processing operation; and processing the target labels of the target images according to the operation label based on the operation mode, wherein the operation mode includes one of the following: replacement and addition.
[0133] The aforementioned preset quantity can be the number of target images set in advance according to specific circumstances, and can be, but is not limited to, 1.
[0134] The above processing operation can be the user inputting the tag to be replaced or added. If the user needs to replace a tag, they need to input the old tag to be replaced and the new tag to be replaced; if the user needs to add a tag, they can directly input a new tag to be added.
[0135] The operation labels mentioned above can be labels that the user wants to replace or add. If the user needs to replace a label, the operation labels are the old label and the new label; if the user needs to add a label, the operation label is the new label entered by the user.
[0136] In one alternative embodiment, such as Figure 3 As shown, the operation controls include: a batch add label option and a corresponding label content box; a batch replace label option and corresponding old label content boxes and new label content boxes. If a user needs to batch process multiple target images, that is, if the number of target images selected by the user is greater than the preset number, the user can determine the operation method as replace or add according to their needs, and select the labels to be batch processed. The add operation adds the corresponding label to the first position of multiple target images, while the replace operation requires the user to select the old label and the new label, thereby replacing the old label of multiple target images with the new label, making it convenient for the user to operate on the target labels.
[0137] For example, if a user selects five images of chairs from multiple images of virtual objects, they can choose the option to add labels in batches and confirm the label as "chair." This will add the label "chair" to the first position of the corresponding five target images. The user can also replace the old label "stool" in the five target images with the new label "chair."
[0138] Optionally, the method further includes: determining whether a target tag exists in a preset tag library; in response to the absence of a target tag in the preset tag library, displaying a tag confirmation interface in a graphical user interface, wherein the tag confirmation interface displays a target tag and a storage control; and in response to a storage operation performed on the storage control, storing the target tag in the preset tag library.
[0139] The aforementioned label confirmation interface can be an interface used to confirm the target label, including but not limited to: the target image corresponding to the target label, the storage control, the export intermediate file control, and the display area of the intermediate file export path.
[0140] The aforementioned storage control can be a save control used to determine the target tag for storage.
[0141] The aforementioned storage operations can be clicks, double-clicks, long presses, drags, swipes, etc., performed by the user on the storage control, but are not limited to these, and this application does not make any specific limitations on them.
[0142] In one alternative embodiment, such as Figure 3 As shown, the upper left corner includes controls for a full scan and a one-click view of letter vocabulary. In response to the user clicking the full scan control, a full scan of the preset tag library is performed to determine if the target tag exists. If the target tag exists, it is displayed; otherwise, clicking the one-click view of letter vocabulary control redirects to the tag confirmation screen. Figure 7 This is a schematic diagram of an optional label confirmation interface according to one embodiment of the present disclosure, such as... Figure 7 As shown, the label confirmation interface includes target labels for the target image, such as the main category label for wood: "Building," and the subcategory label: "Building Components"; and the main category label for brick ground: "Pedestrian Crossing." It also includes a save control. In response to the user's confirmation to save the corresponding target label, clicking the save control stores the target label in a preset label library, facilitating subsequent label processing.
[0143] Figure 8 This is a flowchart of an optional storage target tag according to one embodiment of the present disclosure, such as... Figure 8 As shown, the method for this step is as follows:
[0144] Step S801: In response to the user clicking the full scan operation control, a full scan of the preset tag library is performed.
[0145] Step S802: Determine whether the target tag exists in the preset tag library. If yes, proceed to step S803; otherwise, proceed to step S804.
[0146] Step S803: Display the target label.
[0147] In step S804, in response to the user clicking the operation control to view the vocabulary of the letter with one click, the user is redirected to the tag confirmation interface.
[0148] Step S805: Determine the corresponding target label;
[0149] Step S806: Store the target tag in the preset tag library.
[0150] Optionally, the method further includes: in response to a mapping operation on a target tag, determining a mapping tag corresponding to the mapping operation, wherein the mapping tag is used to represent a tag in a preset tag library; and mapping the target tag to a mapping tag.
[0151] The mapping operation described above can be performed by the user entering mapping labels in the mapping control or selecting labels from the preset label library.
[0152] The mapping labels mentioned above can be labels entered or selected by the user.
[0153] In one alternative embodiment, such as Figure 7 As shown, the label confirmation interface also includes controls for mapping existing labels. Users can determine whether new label categories need to be modified or whether existing labels need to be mapped. For example, the newly added "Brickland" label can be mapped to the existing "Brickland" label. Users only need to enter "Brickland" in the blank area of the corresponding column and the row corresponding to "Brickland". This allows the target label to be mapped to the mapped label.
[0154] Optionally, storing the target tag in a preset tag library includes: generating a target identifier corresponding to the target tag; translating the target tag to obtain a multilingual tag corresponding to the target tag; and storing the target identifier and the multilingual tag in the preset tag library.
[0155] The above confirmation operation can be used to confirm the storage of the target tag in the preset tag library.
[0156] The aforementioned target identifier can be an identifier for keyword feature annotation of target tags, and the target identifier corresponding to each target tag is unique.
[0157] The above-mentioned multilingual tags can be tags corresponding to a target tag in multiple languages, including but not limited to: Chinese tags, English tags, Japanese tags, Korean tags.
[0158] In an optional embodiment, while storing the target tag into the preset tag library, the target identifier corresponding to the target tag is obtained. For example, the target identifier of the treasure chest is code_tag_model_Treasure_chest. And the large language model is used to translate the target tag to obtain multiple multilingual tags. For example, Chinese: treasure chest, English: Treasure chest, Japanese: 宝箱, Korean: etc. And the target identifier and the multilingual tags are stored into the preset tag library. Among them, Figure 9 is a schematic diagram of an optional multilingual library page according to one embodiment of the present disclosure. As Figure 9 shown, the upper left corner includes importing an intermediate file and the corresponding file path to obtain the target tag, and displaying the target identifier corresponding to the target tag, and the corresponding Chinese tag, English tag, Japanese tag, and Korean tag in the area corresponding to the conflicting multilingual keywords. Among them, the user can also modify the target identifier and the corresponding multilingual tags in the area of the to-be-confirmed newly added multilingual keywords and map them in the originally corresponding area. After confirming the configuration table, the control for exporting the configuration table can be clicked to store the target identifier and the multilingual tags correspondingly into the preset tag library, which provides convenience for subsequent tag processing.
[0159] It should be noted that the generation format of the target identifier is a prefix plus a tag, such as code_effect_tag_XXX. Among them, the special effect prefix is code_effect_tag_, the model prefix is code_tag_model_, the GUI prefix is code_GUI_Tag_, and the audio prefix is code_audio_tag_. As Figure 9 shown, the multilingual library page also includes options for adding a suffix sequence number and customizing the suffix, as well as corresponding edit boxes, and the user can determine the format of the target identifier according to needs.
[0160] Figure 10 is a flowchart for translating a target tag according to one embodiment of the present disclosure. As Figure 10 shown, the steps of the method are as follows:
[0161] Step S1001, obtain the newly stored target tag.
[0162] Step S1002, determine whether there are multilingual tags corresponding to the target tag stored. If so, execute step S1003; if not, execute step S1004.
[0163] Step S1003 calls the multilingual tag and target identifier corresponding to the target tag.
[0164] Step S1004: Generate multilingual tags and target identifiers corresponding to the target tags.
[0165] Optionally, the target tag is translated to obtain a multilingual tag corresponding to the target tag, including one of the following: translating the target tag according to a preset language to obtain a multilingual tag corresponding to the preset language; or translating the target tag according to multiple languages to obtain a multilingual tag corresponding to multiple languages.
[0166] The preset languages mentioned above can be pre-set according to specific circumstances, and users can modify them at any time as needed. They can be, but are not limited to, Chinese, Korean, Japanese, and English.
[0167] In one optional embodiment, a preset language can be set according to user needs. The target label is then translated into multiple languages to obtain multilingual labels. This avoids displaying too many unnecessary multilingual labels on the screen, which would affect screen cleanliness and prevent users from quickly and accurately finding labels they can understand. Furthermore, to facilitate understanding of the target image by users of different languages, promote cross-cultural communication and understanding, and expand the audience, the target label can also be translated into multiple languages to obtain multilingual labels corresponding to those languages. This allows users who do not know how to set a preset language to directly find labels they can understand.
[0168] Optionally, the method further includes: performing an export operation based on the target image and target label to generate a first export file; and generating a second export file based on the target label and a preset label library.
[0169] The aforementioned export operation can be any operation performed by the user on the export control, such as clicking, double-clicking, long-pressing, dragging, or swiping, but it is not limited to these, and this application does not make any specific limitations on this.
[0170] The first exported file mentioned above can be an exported file consisting of the target image and the target label.
[0171] The second exported file mentioned above can be an exported file consisting of target tags and a preset tag library.
[0172] In one alternative embodiment, such as Figure 7As shown, the label confirmation interface also displays controls for exporting intermediate files. When the user clicks the export control and confirms the export of the target image, target labels, and preset label library, a first export file is generated for the target image and target labels, and a second export file is generated for the target labels and preset label library. The first export file can be a model summary table, displaying the target image, its corresponding image code, target label, and flag information, allowing the user to easily confirm the correctness of the target labels. The second export file can be the preset label library, facilitating subsequent user operations.
[0173] It should be noted that after exporting the first and second export files, the flag information of the corresponding target image is converted to 1.
[0174] Figure 11 This is a flowchart of an optional image annotation method according to an embodiment of the present disclosure, such as... Figure 11 As shown, the specific implementation method includes the following steps:
[0175] Step S1101: Determine the target image;
[0176] Step S1102: Generate label description information using multiple label description generation models;
[0177] Step S1103: Use a large language model to convert the label description information into target labels;
[0178] Step S1104: Obtain the frequency of occurrence of the target tag;
[0179] Step S1105: Obtain the first tag and the second tag based on their frequency of occurrence;
[0180] Step S1106: Display the first label in the first display area corresponding to the target image, and store the second label.
[0181] Through the above description of the embodiments, those skilled in the art can clearly understand that the methods according to the above embodiments can be implemented by means of software plus necessary general-purpose hardware platforms. Of course, they can also be implemented by hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product is stored in a storage medium (such as ROM / RAM, magnetic disk, optical disk), and includes several instructions to cause a terminal device (which may be a mobile phone, computer, server, or network device, etc.) to execute the methods described in the various embodiments of this disclosure.
[0182] This embodiment also provides a label generation apparatus for implementing the above embodiments and preferred embodiments; details already described will not be repeated. As used below, the terms "unit" and "module" can refer to a combination of software and / or hardware that performs a predetermined function. Although the apparatus described in the following embodiments is preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated.
[0183] Figure 12 This is a structural block diagram of an image annotation apparatus according to one embodiment of the present disclosure. A graphical user interface is provided through a terminal device. The content displayed by the graphical user interface includes at least images of multiple virtual objects and a first operation control, such as... Figure 12 As shown, the device includes: an image acquisition module 1202 for acquiring a target image; a label description generation module 1204 for inputting the target image into multiple label description generation models, performing label generation on the target image respectively, and obtaining label description information of the target image, wherein the label description information is used to characterize the information describing the label of the target image; a label extraction module 1206 for extracting the label description information using a large language model to obtain the target label of the target image; and an image annotation module 1208 for annotating the target image based on the target label.
[0184] Optionally, the tag conversion module includes: a first conversion unit for extracting multiple initial tags from tag description information using a large language model; and a filtering unit for filtering target tags from the multiple initial tags based on the frequency of occurrence of the multiple initial tags in the tag description information.
[0185] Optionally, the filtering unit includes: a filtering subunit, used to filter out initial tags whose occurrence frequency is greater than a preset occurrence frequency from multiple initial tags based on the occurrence frequency of multiple initial tags in the tag description information, so as to obtain the target tag.
[0186] Optionally, the filtering unit further includes: an acquisition subunit, used to acquire the selection ratio corresponding to the input command in response to an input command on the image user interface, wherein the selection ratio is used to characterize the proportion of the target label among multiple initial labels; and a determination subunit, used to determine a preset occurrence frequency based on the selection ratio.
[0187] Optionally, the image annotation module includes: a first acquisition unit for acquiring historical labels of the target image; a label processing unit for classifying the historical labels and the target label based on a classification configuration for the historical labels and the target label to obtain a first label and a second label, wherein the first label is used to annotate the target image, and the second label is used to update the first label; and an annotation unit for annotating the target image based on the first label and storing the second label. The first label is displayed in a first display area of the graphical user interface, and the second label is displayed in a second display area of the graphical user interface.
[0188] Optionally, the device further includes: a first determining module, configured to determine whether a target label exists in a preset label library; a first display module, configured to display the target label at a first position on the target image in a first display manner in response to the existence of the target label in the preset label library; and a second display module, configured to display the target label at a first position in a second display manner in response to the absence of the target label in the preset label library.
[0189] Optionally, the device further includes: a third display module, configured to display a search page in the graphical user interface in response to a preset operation performed on the graphical user interface, wherein the search page is used to search for images of multiple virtual objects; and an image determination module, configured to determine images of multiple virtual objects corresponding to the search operation from a set of images of virtual objects in response to a search operation performed on the search page.
[0190] Optionally, the device further includes: a first acquisition module, configured to acquire flag information corresponding to the images of multiple virtual objects, wherein the flag information is used to indicate whether the images of multiple virtual objects are labeled, or whether the labels of the images of multiple virtual objects are modified or confirmed; a first generation module, configured to generate target subscript elements corresponding to the images of multiple virtual objects based on the flag information; and a fourth display module, configured to display the target subscript elements at a second position in the images of multiple virtual objects.
[0191] Optionally, the device further includes: an operation determination module, configured to determine the operation mode and operation label corresponding to the processing operation in response to the processing operation on the target image; and a label processing module, configured to process the target label of the target image according to the operation label based on the operation mode, wherein the operation mode includes one of the following: replacement and addition.
[0192] Optionally, the device further includes: a second determining module for determining whether a target label exists in a preset label library; a fifth display module for displaying a label confirmation interface in a graphical user interface in response to the absence of a target label in the preset label library, wherein the label confirmation interface displays a target label and a storage control; and a label storage module for storing the target label in the preset label library in response to a storage operation performed on the storage control.
[0193] Optionally, the device further includes: a third determining module, configured to determine a mapping label corresponding to the mapping operation in response to the mapping operation for the target label, wherein the mapping label is used to represent a label in a preset label library; and a mapping module, configured to map the target label to a mapping label.
[0194] Optionally, the tag storage module includes: an identifier generation unit for generating a target identifier corresponding to the target tag; a tag translation unit for translating the target tag to obtain a multilingual tag corresponding to the target tag; and a tag storage unit for storing the target identifier and the multilingual tag in a preset tag library.
[0195] Optionally, the device further includes: a second generation module for performing an export operation based on the target image and target labels to generate a first export file; and a third generation module for generating a second export file based on the target labels and a preset label library.
[0196] It should be noted that the above-mentioned units and modules can be implemented by software or hardware. For the latter, they can be implemented in the following ways, but not limited to these: all the above-mentioned units and modules are located in the same processor; or, the above-mentioned units and modules are located in different processors in any combination.
[0197] Embodiments of this disclosure also provide a computer-readable storage medium storing a computer program configured to perform the steps in any of the above method embodiments when executed.
[0198] Optionally, in this embodiment, the computer-readable storage medium may include, but is not limited to, various media capable of storing computer programs, such as USB flash drives, read-only memory (ROM), random access memory (RAM), portable hard drives, magnetic disks, or optical disks.
[0199] Optionally, in this embodiment, the computer-readable storage medium may be located in any computer terminal in a group of computer terminals in a computer network, or in any mobile terminal in a group of mobile terminals.
[0200] Optionally, in this embodiment, the computer-readable storage medium may be configured to store a computer program for performing the following steps:
[0201] Step S1: Obtain the target image;
[0202] Step S2: Input the target image into multiple label description generation models, perform label generation on the target image respectively, and obtain the label description information of the target image. The label description information is used to characterize the information that describes the labels of the target image.
[0203] Step S3: Use a large language model to extract the label description information to obtain the target label of the target image;
[0204] Step S4: Label the target image based on the target label.
[0205] Optionally, the aforementioned computer-readable storage medium is further configured to store program code for performing the following steps: extracting multiple initial tags from tag description information using a large language model; and selecting target tags from the multiple initial tags based on the frequency of occurrence of the multiple initial tags in the tag description information.
[0206] Optionally, the aforementioned computer-readable storage medium is further configured to store program code for performing the following steps: based on the frequency of occurrence of multiple initial tags in tag description information, selecting initial tags from multiple initial tags whose frequency of occurrence is greater than a preset frequency of occurrence, and obtaining target tags.
[0207] Optionally, the aforementioned computer-readable storage medium is further configured to store program code for performing the following steps: in response to an input instruction on a graphical user interface, obtaining a selection ratio corresponding to the input instruction, wherein the selection ratio is used to characterize the proportion of the target label among multiple initial labels; and determining a preset occurrence frequency based on the selection ratio.
[0208] Optionally, the aforementioned computer-readable storage medium is further configured to store program code for performing the following steps: obtaining historical labels of a target image; classifying the historical labels and target labels based on a classification configuration for the historical labels and target labels to obtain a first label and a second label, wherein the first label is used to annotate the target image and the second label is used to update the first label; annotating the target image based on the first label and storing the second label.
[0209] Optionally, the aforementioned computer-readable storage medium is further configured to store program code for performing the following steps: a first label is displayed in a first display area of the graphical user interface, and a second label is displayed in a second display area of the graphical user interface.
[0210] Optionally, the aforementioned computer-readable storage medium is further configured to store program code for performing the following steps: determining whether a target tag exists in a preset tag library; in response to the existence of a target tag in the preset tag library, displaying the target tag at a first position in a first display mode; and in response to the absence of a target tag in the preset tag library, displaying the target tag at a first position in a second display mode.
[0211] Optionally, the aforementioned computer-readable storage medium is further configured to store program code for performing the following steps: displaying a search page in the graphical user interface in response to a preset operation performed on the graphical user interface, wherein the search page is used to search for images of multiple virtual objects; and determining images of multiple virtual objects corresponding to the search operation from a set of images of virtual objects in response to a search operation performed on the search page.
[0212] Optionally, the aforementioned computer-readable storage medium is further configured to store program code for performing the following steps: obtaining flag information corresponding to the images of multiple virtual objects, wherein the flag information is used to characterize whether the images of multiple virtual objects are labeled, or whether the labels of the images of multiple virtual objects are modified or confirmed; generating target subscript elements corresponding to the images of multiple virtual objects based on the flag information; and displaying the target subscript elements at a second position in the images of multiple virtual objects.
[0213] Optionally, the aforementioned computer-readable storage medium is further configured to store program code for performing the following steps: in response to a processing operation performed on a target image, determining an operation mode and operation label corresponding to the processing operation; processing the target label of the target image according to the operation label based on the operation mode, wherein the operation mode includes one of the following: replacement and addition.
[0214] Optionally, the aforementioned computer-readable storage medium is further configured to store program code for performing the following steps: determining whether a target tag exists in a preset tag library; in response to the absence of a target tag in the preset tag library, displaying a tag confirmation interface in a graphical user interface, wherein the tag confirmation interface displays a target tag and a storage control; and in response to a storage operation performed on the storage control, storing the target tag in the preset tag library.
[0215] Optionally, the aforementioned computer-readable storage medium is further configured to store program code for performing the following steps: in response to a mapping operation for a target tag, determining a mapping tag corresponding to the mapping operation, wherein the mapping tag is used to characterize a tag in a preset tag library; mapping the target tag to a mapping tag.
[0216] Optionally, the aforementioned computer-readable storage medium is further configured to store program code for performing the following steps: generating a target identifier corresponding to a target tag; translating the target tag to obtain a multilingual tag corresponding to the target tag; and storing the target identifier and the multilingual tag in a preset tag library.
[0217] Optionally, the aforementioned computer-readable storage medium is further configured to store program code for performing the following steps: performing an export operation based on the target image and target label to generate a first export file; generating a second export file based on the target label and a preset label library.
[0218] This embodiment provides a technical solution in a computer-readable storage medium. The process involves: acquiring a target image; inputting the target image into multiple label description generation models, performing label description generation on the target image respectively to obtain label description information for the target image, wherein the label description information is used to characterize the information describing the label of the target image; extracting the label description information using a large language model to obtain target labels for the target image; and labeling the target image based on the target labels. It is noteworthy that by using multiple label description generation models to generate label description information for the target image respectively, and then using a large language model to extract the label description information to obtain target labels, the target labels matching the target image can be accurately and efficiently obtained, achieving the goal of improving label generation efficiency. This realizes the technical effect of automatically generating labels from the target image and labeling the target image, thereby solving the technical problem of low label generation efficiency caused by manual image labeling in the prior art.
[0219] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions according to the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a computer-readable storage medium (such as a CD-ROM, USB flash drive, external hard drive, etc.) or on a network, including several instructions to cause a computing device (such as a personal computer, server, terminal device, or network device, etc.) to execute the methods according to the embodiments of this disclosure.
[0220] In exemplary embodiments of this application, a computer-readable storage medium stores a program product capable of implementing the methods described above in this embodiment. In some possible implementations, various aspects of the embodiments of this disclosure may also be implemented as a program product including program code, which, when the program product is run on a terminal device, causes the terminal device to perform the steps according to various exemplary embodiments of this disclosure described in the "Exemplary Methods" section above.
[0221] The program product for implementing the above-described method according to embodiments of the present disclosure may employ a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the program product of the embodiments of the present disclosure is not limited thereto. In the embodiments of the present disclosure, the computer-readable storage medium may be any tangible medium that contains or stores a program that may be used by or in conjunction with an instruction execution system, apparatus, or device.
[0222] The aforementioned program product may take the form of any combination of one or more computer-readable media. Such computer-readable storage media may be, for example, but not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatuses, or devices, or any combination thereof. More specific examples (not exhaustive) of computer-readable storage media include: electrical connections having one or more wires, portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof.
[0223] It should be noted that the program code contained on the computer-readable storage medium can be transmitted using any suitable medium, including but not limited to wireless, wired, optical fiber, RF, etc., or any suitable combination thereof.
[0224] Embodiments of this disclosure also provide an electronic device including a memory and a processor, the memory storing a computer program and the processor being configured to run the computer program to perform the steps in any of the above method embodiments.
[0225] Optionally, the electronic device may further include a transmission device and an input / output device, wherein the transmission device is connected to the processor and the input / output device is connected to the processor.
[0226] Optionally, in this embodiment, the processor can be configured to perform the following steps via a computer program:
[0227] Step S1: Obtain the target image;
[0228] Step S2: Input the target image into multiple label description generation models, perform label generation on the target image respectively, and obtain the label description information of the target image. The label description information is used to characterize the information that describes the labels of the target image.
[0229] Step S3: Use a large language model to extract the label description information to obtain the target label of the target image;
[0230] Step S4: Label the target image based on the target label.
[0231] Optionally, the processor described above may also be configured to perform the following steps via a computer program: extracting multiple initial tags from the tag description information using a large language model; and selecting target tags from the multiple initial tags based on the frequency of occurrence of the multiple initial tags in the tag description information.
[0232] Optionally, the processor may also be configured to perform the following steps via a computer program: based on the frequency of occurrence of multiple initial tags in the tag description information, select initial tags from the multiple initial tags whose frequency of occurrence is greater than a preset frequency of occurrence, and obtain target tags.
[0233] Optionally, the processor may also be configured to perform the following steps via a computer program: in response to an input command on a graphical user interface, obtain the selection ratio corresponding to the input command, wherein the selection ratio is used to characterize the proportion of the target label among multiple initial labels; and determine a preset occurrence frequency based on the selection ratio.
[0234] Optionally, the processor may also be configured to perform the following steps via a computer program: acquiring historical labels of the target image; classifying the historical labels and target labels based on a classification configuration for the historical labels and target labels to obtain a first label and a second label, wherein the first label is used to annotate the target image and the second label is used to update the first label; annotating the target image based on the first label and storing the second label.
[0235] Optionally, the processor may also be configured to perform the following steps via a computer program: a first label is displayed in a first display area of the graphical user interface, and a second label is displayed in a second display area of the graphical user interface.
[0236] Optionally, the processor may also be configured to perform the following steps via a computer program: determining whether a target tag exists in a preset tag library; in response to the existence of a target tag in the preset tag library, displaying the target tag at a first position in a first display mode; and in response to the absence of a target tag in the preset tag library, displaying the target tag at a first position in a second display mode.
[0237] Optionally, the processor may also be configured to perform the following steps via a computer program: in response to a preset operation performed on the graphical user interface, displaying a search page in the graphical user interface, wherein the search page is used to search for images of multiple virtual objects; in response to a search operation performed on the search page, determining images of multiple virtual objects corresponding to the search operation from a set of images of virtual objects.
[0238] Optionally, the processor may also be configured to perform the following steps via a computer program: obtaining flag information corresponding to the images of multiple virtual objects, wherein the flag information is used to characterize whether the images of multiple virtual objects are labeled, or whether the labels of the images of multiple virtual objects are modified or confirmed; generating target subscript elements corresponding to the images of multiple virtual objects based on the flag information; and displaying the target subscript elements at a second position in the images of multiple virtual objects.
[0239] Optionally, the processor may also be configured to perform the following steps via a computer program: in response to a processing operation performed on a target image, determining the operation mode and operation label corresponding to the processing operation; and processing the target label of the target image according to the operation label based on the operation mode, wherein the operation mode includes one of the following: replacement and addition.
[0240] Optionally, the processor may also be configured to perform the following steps via a computer program: determining whether a target tag exists in a preset tag library; in response to the absence of a target tag in the preset tag library, displaying a tag confirmation interface in a graphical user interface, wherein the tag confirmation interface displays a target tag and a storage control; and in response to a storage operation performed on the storage control, storing the target tag in the preset tag library.
[0241] Optionally, the processor may also be configured to perform the following steps via a computer program: in response to a mapping operation for a target tag, determine the mapping tag corresponding to the mapping operation, wherein the mapping tag is used to represent a tag in a preset tag library; and map the target tag to the mapping tag.
[0242] Optionally, the processor may also be configured to perform the following steps via a computer program: generating a target identifier corresponding to the target tag; translating the target tag to obtain a multilingual tag corresponding to the target tag; and storing the target identifier and the multilingual tag in a preset tag library.
[0243] Optionally, the processor may also be configured to perform the following steps via a computer program: perform an export operation based on the target image and target label to generate a first export file; and generate a second export file based on the target label and a preset label library.
[0244] In the electronic device of this embodiment, a technical solution is provided. The process involves: acquiring a target image; inputting the target image into multiple label description generation models, performing label description generation on the target image respectively to obtain label description information for the target image, wherein the label description information is used to characterize the information describing the label of the target image; extracting the label description information using a large language model to obtain target labels for the target image; and labeling the target image based on the target labels. It is noteworthy that by using multiple label description generation models to generate label description information for the target image respectively, and then using a large language model to extract the label description information to obtain target labels, the target labels matching the target image can be accurately and efficiently obtained, achieving the goal of improving label generation efficiency. This realizes the technical effect of automatically generating labels from the target image and labeling the target image, thereby solving the technical problem of low label generation efficiency caused by manual image labeling in the prior art.
[0245] Figure 13 This is a schematic diagram of an electronic device according to an embodiment of the present disclosure. Figure 13 As shown, the electronic device 1300 is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments disclosed herein.
[0246] like Figure 13 As shown, the electronic device 1300 is presented in the form of a general-purpose computing device. The components of the electronic device 1300 may include, but are not limited to: at least one processor 1310, at least one memory 1320, a bus 1330 connecting different system components (including memory 1320 and processor 1310), and a display 1340.
[0247] The memory 1320 stores program code that can be executed by the processor 1310, causing the processor 1310 to perform the steps described in the method section of the embodiments of this application according to various exemplary implementations of this disclosure.
[0248] The memory 1320 may include a readable medium in the form of volatile memory cells, such as random access memory (RAM) 13201 and / or cache memory 13202, and may further include read-only memory (ROM) 13203, and may also include non-volatile memory, such as one or more magnetic storage devices, flash memory, or other non-volatile solid-state memory.
[0249] In some instances, memory 1320 may also include programs / utilities 13204 having a set (at least one) of program modules 13205, including but not limited to: an operating system, one or more application programs, other program modules, and program data. Each or some combination of these examples may include an implementation of a network environment. Memory 1320 may further include memory remotely located relative to processor 1310, which can be connected to electronic device 1300 via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.
[0250] Bus 1330 can represent one or more of several types of bus structures, including a memory cell bus or memory cell controller, peripheral bus, graphics acceleration port, processor 1310, or a local bus using any of the various bus structures.
[0251] The display 1340 may be, for example, a touch screen liquid crystal display (LCD) that allows a user to interact with the user interface of the electronic device 1300.
[0252] Optionally, the electronic device 1300 can also communicate with one or more external devices 1400 (e.g., keyboard, pointing device, Bluetooth device, etc.), one or more devices that enable a user to interact with the electronic device 1300, and / or any device that enables the electronic device 1300 to communicate with one or more other computing devices (e.g., router, modem, etc.). This communication can be performed via the input / output (I / O) interface 1350. Furthermore, the electronic device 1300 can also communicate with one or more networks (e.g., local area network (LAN), wide area network (WAN), and / or public networks, such as the Internet) via the network adapter 1360. Figure 13 As shown, network adapter 1360 communicates with other modules of electronic device 1300 via bus 1330. It should be understood that, although... Figure 13 As not shown, other hardware and / or software modules may be used in conjunction with electronic device 1300, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems.
[0253] The aforementioned electronic device 1300 may further include: a keyboard, a cursor control device (such as a mouse), an input / output interface (I / O interface), a network interface, a power supply, and / or a camera.
[0254] Those skilled in the art will understand that Figure 13The structure shown is for illustrative purposes only and does not limit the structure of the electronic device described above. For example, the electronic device 1300 may also include components that are more... Figure 13 The more or fewer components shown, or having the same Figure 1 Different configurations are shown. The memory 1320 can be used to store computer programs and corresponding data, such as the computer program and corresponding data corresponding to the image annotation method in this embodiment. The processor 1310 executes various functional applications and data processing by running the computer program stored in the memory 1320, thereby implementing the image annotation method described above.
[0255] The sequence numbers of the embodiments disclosed above are for descriptive purposes only and do not represent the superiority or inferiority of the embodiments.
[0256] In the above embodiments of this disclosure, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0257] In the several embodiments provided in this application, it should be understood that the disclosed technical content can be implemented in other ways. The device embodiments described above are merely illustrative; for example, the division of units can be a logical functional division, and in actual implementation, there may be other division methods. For instance, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the displayed or discussed mutual coupling, direct coupling, or communication connection may be through some interfaces; the indirect coupling or communication connection between units or modules may be electrical or other forms.
[0258] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0259] Furthermore, the functional units in the various embodiments of this disclosure can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0260] If the integrated unit is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this disclosure, in essence, or the part that contributes to the prior art, or all or part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this disclosure. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, read-only memory (ROM), random access memory (RAM), portable hard drive, magnetic disk, or optical disk.
[0261] The above description is only a preferred embodiment of this disclosure. It should be noted that for those skilled in the art, several improvements and modifications can be made without departing from the principles of this disclosure, and these improvements and modifications should also be considered within the scope of protection of this disclosure.
Claims
1. An image annotation method, characterized in that, include: Acquire the target image; The target image is input into multiple label description generation models, and label description generation is performed on the target image respectively to obtain label description information of the target image, wherein the label description information is used to characterize the information describing the labels of the target image; The target label of the target image is obtained by extracting the label description information using a large language model. Obtain the historical tags of the target image; Based on the classification configuration for the historical tags and the target tags, the historical tags and the target tags are classified to obtain a first tag and a second tag, wherein the first tag is used to annotate the target image, and the second tag is used to update the first tag; The target image is labeled based on the first label, and the second label is stored. The first label is displayed in a first display area of the graphical user interface, and the second label is displayed in a second display area of the graphical user interface.
2. The method according to claim 1, characterized in that, The target image is obtained by extracting the label description information using the large language model, including: The large language model is used to extract multiple initial tags from the tag description information; The target tag is obtained by filtering from the multiple initial tags based on their frequency of occurrence in the tag description information.
3. The method according to claim 2, characterized in that, Based on the frequency of occurrence of the multiple initial tags in the tag description information, the target tag is obtained by filtering from the multiple initial tags, including: Based on the frequency of occurrence of the multiple initial tags in the tag description information, the initial tags with a frequency greater than a preset frequency are selected from the multiple initial tags to obtain the target tag.
4. The method according to claim 3, characterized in that, The method further includes: In response to an input command on a graphical user interface, the selection ratio corresponding to the input command is obtained, wherein the selection ratio is used to characterize the proportion of the target label among the plurality of initial labels; Based on the selected ratio, the preset occurrence frequency is determined.
5. The method according to claim 1, characterized in that, The method further includes: Determine whether the target tag exists in the preset tag library; In response to the presence of the target tag in the preset tag library, the target tag is displayed at a first position on the target image according to a first display method; In response to the absence of the target tag in the preset tag library, the target tag is displayed at the first position according to the second display method.
6. The method according to claim 1, characterized in that, The method further includes: In response to a preset operation performed on the graphical user interface, a search page is displayed in the graphical user interface, wherein the search page is used to search for images of multiple virtual objects; In response to a search operation performed on the search page, images of the plurality of virtual objects corresponding to the search operation are determined from a set of images of virtual objects.
7. The method according to claim 1, characterized in that, The method further includes: Obtain flag information corresponding to the images of multiple virtual objects, wherein the flag information is used to indicate whether the images of the multiple virtual objects are labeled, or whether the labels of the images of the multiple virtual objects are modified or confirmed; Based on the flag information, target label elements corresponding to the images of the plurality of virtual objects are generated; The target superscript element is displayed at a second position on the image of the plurality of virtual objects.
8. The method according to claim 1, characterized in that, In response to the number of target images being greater than a preset number, the method further includes: In response to a processing operation performed on the target image, determine the operation mode and operation label corresponding to the processing operation; Based on the operation method, the target label of the target image is processed according to the operation label, wherein the operation method includes one of the following: replacement and addition.
9. The method according to claim 1, characterized in that, The method further includes: Determine whether the target tag exists in the preset tag library; In response to the fact that the target tag does not exist in the preset tag library, a tag confirmation interface is displayed in the graphical user interface, wherein the target tag and the storage control are displayed in the tag confirmation interface; In response to a storage operation performed on the storage control, the target tag is stored in the preset tag library.
10. The method according to claim 9, characterized in that, The method further includes: In response to a mapping operation for the target tag, a mapping tag corresponding to the mapping operation is determined, wherein the mapping tag is used to represent a tag in the preset tag library; Map the target label to the mapped label.
11. The method according to claim 9, characterized in that, Storing the target tag into the preset tag library includes: Generate the target identifier corresponding to the target label; The target tag is translated to obtain the multilingual tag corresponding to the target tag; The target identifier and the corresponding multilingual tag are stored in the preset tag library.
12. The method according to claim 11, characterized in that, The target tag is translated to obtain a multilingual tag corresponding to the target tag, including one of the following: The target tag is translated according to a preset language to obtain the multilingual tag corresponding to the preset language; The target tag is translated into multiple languages to obtain the multilingual tag corresponding to each language.
13. The method according to claim 9, characterized in that, The method further includes: Based on the target image and the target label, an export operation is performed to generate a first exported file; A second exported file is generated based on the target tag and the preset tag library.
14. An image annotation device, characterized in that, include: The image acquisition module is used to acquire the target image; The label description generation module is used to input the target image into multiple label description generation models to generate label descriptions for the target image, thereby obtaining label description information for the target image. The label description information is used to characterize the information describing the labels of the target image. The tag extraction module is used to extract the tag description information using a large language model to obtain the target tag of the target image; An image annotation module is used to obtain historical tags of the target image; classify the historical tags and the target tags based on the classification configuration for the historical tags and the target tags to obtain a first tag and a second tag, wherein the first tag is used to annotate the target image, and the second tag is used to update the first tag; annotate the target image based on the first tag and store the second tag, wherein the first tag is displayed in a first display area of the graphical user interface, and the second tag is displayed in a second display area of the graphical user interface.
15. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, wherein the computer program is configured to execute the method described in any one of claims 1 to 13 when run by a processor.
16. An electronic device comprising a memory and a processor, characterized in that, The memory stores a computer program, and the processor is configured to run the computer program to perform the method as described in any one of claims 1 to 13.
Citation Information
Patent Citations
Image description generation method and device, equipment, medium and product
CN114627353A
Data processing method and device, electronic equipment and storage medium
CN117079299A