Image marking method and device, equipment and storage medium
By inputting sample images into multiple visual language models and topic classification models, generating and adjusting labels, the problem of low marking efficiency of sample images in the prior art is solved, and the training efficiency of literary and artistic graph models is improved.
Patent Information
- Application Number
- CN202311660061.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2023-12-05
- Publication Date
- 2025-06-06
AI Technical Summary
In the prior art, the sample image marking processing efficiency is low, resulting in corresponding reduction in the training efficiency of literary and artistic graphics models.
Descriptive information is obtained by inputting the sample images into at least two visual language models, generating description data, and performing deduplication word processing. Then, the description information is input into the topic classification model to generate the subject words as labels for the image.
The efficiency and accuracy of image marking are improved, so that sample images with labels can be suitable for the training of various vernacular graphics models, and the training efficiency of vernacular graphics models is improved.
Smart Images

Figure CN120107967A_ABST
Abstract
Description
Technical Field
[0001] The present application belongs to the field of model training technology, and in particular, relates to an image labeling method, device, equipment and storage medium. Background Art
[0002] The rapid development of modern science and technology has promoted the progress of computer vision and natural language processing. Generating images based on text descriptions is a comprehensive task that spans the two fields of computer vision and natural language processing. It has great application potential and is expected to play an important role in criminal investigation, data enhancement, design, etc. in the future.
[0003] In the related art, sample images with labels can be input into the text graph model, and the text graph model can be trained to obtain a text graph model that meets the requirements. When in use, text is input into the text graph model, and the text graph model outputs an image corresponding to the text. When training the text graph model, a large number of sample images are required, and the sample images need to be manually labeled to obtain sample images with labels, which makes the training efficiency low. Summary of the invention
[0004] The embodiments of the present application provide an image marking method, device, equipment and storage medium, which can solve the problem of low efficiency in the existing sample image marking process.
[0005] In a first aspect, an embodiment of the present application provides an image marking method, the method comprising:
[0006] Inputting the sample images into at least two visual language models respectively, and each visual language model outputs description data corresponding to the sample images respectively;
[0007] Deleting duplicate words from the description data to obtain description information corresponding to the sample image;
[0008] The description information is input into the topic classification model, a topic word corresponding to the description information is generated, and the topic word is set as a label corresponding to the sample image.
[0009] In some embodiments, the description data includes a description paragraph; and deduplication processing is performed on the description data to obtain description information corresponding to the sample image, including:
[0010] Perform word segmentation on the description paragraph to obtain a word segmentation phrase corresponding to the description paragraph;
[0011] Remove duplicate words from the word segmentation phrases corresponding to the sample image to obtain the description information corresponding to the sample image
[0012] In some embodiments, the description data includes a description paragraph and a description phrase, and the at least two visual language models include a first visual language model and a second visual language model; inputting the sample image into the at least two visual language models, and generating description data corresponding to the sample image includes:
[0013] Inputting the sample image into the first visual language model to generate a description segment corresponding to the sample image;
[0014] Inputting the sample image into the second visual language model to generate a description phrase corresponding to the sample image;
[0015] The description data is processed to remove duplicate words, and the description information corresponding to the sample image is obtained, including:
[0016] Perform word segmentation on the description paragraph to obtain a word segmentation phrase corresponding to the description paragraph;
[0017] The segmentation phrases and description phrases corresponding to the sample image are processed to remove duplicate words, so as to obtain description information corresponding to the sample image.
[0018] In some embodiments, after inputting the description information into a topic classification model, generating a topic word corresponding to the description information, and setting the topic word as a label corresponding to the sample image, the method further includes:
[0019] According to the feature information of the Wensheng graph model to be trained, one or more operations of adding, replacing and deleting are performed on the labels corresponding to the sample images to obtain training labels matching the feature information, wherein the feature information includes the type of the Wensheng graph model to be trained and the image style output by the Wensheng graph model to be trained;
[0020] The training sample set is input into the to-be-trained Wensheng graph model to train the to-be-trained Wensheng graph model to obtain a trained Wensheng graph model, wherein the training sample set includes a plurality of sample images and training labels corresponding to the sample images.
[0021] In some embodiments, the image style includes a target training element; when the to-be-trained text image model is a Lora model, one or more operations of adding, replacing and deleting a label corresponding to the sample image according to feature information of the to-be-trained text image model include:
[0022] The target training element in the feature information is obtained, and the label that is the same as the target training element is deleted from the label corresponding to the sample image to obtain the training label that matches the feature information.
[0023] In some embodiments, the image style includes a target training element; when the to-be-trained text image model is a Dreambooth model, one or more operations of adding, replacing and deleting a label corresponding to the sample image according to feature information of the to-be-trained text image model include:
[0024] The target training element in the feature information and the activation label corresponding to the target training element are obtained, and the label corresponding to the sample image is replaced with the activation label to obtain the training label matching the feature information.
[0025] In some embodiments, inputting the training sample set into the to-be-trained document graph model to train the to-be-trained document graph model to obtain the trained document graph model includes:
[0026] Inputting a sample image of a training sample set into an image decoder to obtain a first vector corresponding to the sample image;
[0027] Input the training labels of the training sample set into the text decoder to obtain a second vector corresponding to the training labels;
[0028] The first vector and the second vector corresponding to the same sample image form a vector pair;
[0029] The vector pair is input into the to-be-trained text graph model, and the to-be-trained text graph model is trained to obtain a trained text graph model.
[0030] In a second aspect, an embodiment of the present application further provides an image marking device, the device comprising:
[0031] A description information generation module, used to input the sample image into at least two visual language models respectively, and each visual language model outputs description data corresponding to the sample image respectively;
[0032] A de-duplication word processing module is used to perform de-duplication word processing on the description data to obtain description information corresponding to the sample image;
[0033] The label generation module is used to input the description information into the topic classification model, generate the topic words corresponding to the description information, and set the topic words as labels corresponding to the sample images.
[0034] In a third aspect, an embodiment of the present application provides an image marking device, the device comprising: a processor and a memory storing computer program instructions;
[0035] The above image marking method is implemented when the processor executes the computer program instructions.
[0036] In a fourth aspect, an embodiment of the present application provides a computer storage medium, on which computer program instructions are stored, and when the computer program instructions are executed by a processor, the above-mentioned image marking method is implemented.
[0037] In the present application, by inputting sample images into at least two visual language models respectively, each visual language model outputs description data corresponding to the sample image respectively, and deduplication processing is performed on the description data to obtain description information corresponding to the sample image, so as to enrich the content and description method of the description information, which is conducive to improving the richness of the labels generated according to the description information, so that the sample images with labels can be suitable for the training of various cultural graph models; by inputting the description information into the topic classification model, generating the subject words corresponding to the description information, and setting the subject words as the labels corresponding to the sample images, so that the labels can accurately represent the content of the sample images, thereby improving the accuracy of the labels; through the image labeling method provided in the present application, manual labeling of sample images can be eliminated, thereby improving the labeling efficiency and the training efficiency of cultural graph models. BRIEF DESCRIPTION OF THE DRAWINGS
[0038] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the drawings required for use in the embodiments of the present application will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.
[0039] Figure 1 It is a flowchart of an image marking method provided by an embodiment of the present application;
[0040] Figure 2 This is a partial flow chart of an image marking method provided by an embodiment of the present application;
[0041] Figure 3 This is a partial flow chart of an image marking method provided by an embodiment of the present application;
[0042] Figure 4 It is a schematic diagram of the hardware structure of an image marking device provided in one embodiment of the present application;
[0043] Figure 5 It is a structural schematic diagram of an image marking device provided in one embodiment of the present application. DETAILED DESCRIPTION
[0044] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is only to provide a better understanding of the present application by illustrating the examples of the present application.
[0045] It should be noted that, in this article, relational terms such as first and second, etc. are only used to distinguish one entity or operation from another entity or operation, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Moreover, the terms "include", "comprise" or any other variants thereof are intended to cover non-exclusive inclusion, so that a process, method, article or device including a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article or device. In the absence of further restrictions, the elements defined by the statement "include..." do not exclude the existence of other identical elements in the process, method, article or device including the elements.
[0046] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The embodiments will be described in detail below in conjunction with the accompanying drawings.
[0047] Specifically, in order to solve the problems of the prior art, the embodiments of the present application provide an image marking method, device, equipment and storage medium. The image marking method provided by the embodiments of the present application is first introduced below.
[0048] Figure 1 A schematic diagram of a process flow of an image marking method provided by an embodiment of the present application is shown. The method comprises the following steps:
[0049] S110, inputting the sample image into at least two visual language models respectively, and each visual language model outputs description data corresponding to the sample image respectively;
[0050] The sample image is a pre-selected sample for training the model, and the sample can be in image format. The sample image can be randomly acquired, or it can be an image with target elements and style selected according to the text-based image model to be trained. For example, if the text-based image model to be trained is used to output an anime-style image based on the input text, then for the text-based image model to be trained, an anime-style image is selected as the sample image.
[0051] The visual language model is a pre-trained visual language model. A sample image is input into the visual language model, and description information corresponding to the sample image can be output. The description information can characterize the elements, content, style, etc. displayed in the sample image. The visual language model can be a visual language model in the related art, or a visual language model trained by itself as needed. In the present application, the visual language model can be Clip, Blip-2, Alt Clip, etc.
[0052] S120, performing deduplication processing on the description data to obtain description information corresponding to the sample image;
[0053] For the same sample image, the description data output by different visual language models are different, and the difference can be reflected in the different formats of the output description data. The description data can be a description paragraph, which describes the sample image through a paragraph including one or more sentences. The description data can also be a description phrase, which describes the sample image through one or more words.
[0054] For the same sample image, the description data output by different visual language models may contain the same or different information. Deduplication of description data can be understood as union processing of multiple description data, mixing multiple description data and deleting duplicate items in the description data. Different types of description data are output by different visual language models to enrich the description information; deduplication of description data is performed to simplify the description information and reduce the computational complexity of the subsequent topic classification model.
[0055] For example; the sample image shows "a car parked in a garage, and in the yard outside the garage is an older man holding a water pipe and a rag". Visual language model A outputs description data based on the sample image as "car parked in the garage", visual language model B outputs description data based on the sample image as "garage, car, man, water pipe, rag", and visual language model C outputs description data based on the sample image as "older man holding a water pipe and a rag". Multiple description data are mixed to obtain "car parked in garage, garage, car, man, water pipe, rag, older man holding a water pipe and a rag". The description data is de-duplicated to obtain "car parked in garage, older man holding a water pipe and a rag".
[0056] S130, inputting the description information into a topic classification model, generating topic words corresponding to the description information, and setting the topic words as labels corresponding to the sample images.
[0057] The topic classification model is a pre-trained model that can classify description information and identify the main content of the description information, which is the label of the sample image.
[0058] For the same sample image, the description information output by different visual language models is different. The description information output by multiple visual language models reflects different natural language text and element arrangement schemes, enriching the labels obtained through the description information. The topic classification model can be a text convolutional neural network model ("Convolutional Neural Network, Text CNN).
[0059] For example; the sample image shows "a car parked in a garage, and outside the garage in the yard is an older man holding a water pipe and a rag". Visual language model A outputs description information based on the sample image as "the car is parked in the garage, and outside the garage in the yard is a man holding a water pipe and a rag". Visual language model B outputs description information based on the sample image as "garage, car, man, water pipe, rag". Visual language model C outputs description information based on the sample image as "older man holding a water pipe and a rag". The above description information is input into the topic classification model, and the label corresponding to the sample image is "car parked in the garage, man standing in the yard, man holding a water pipe and a rag, older man".
[0060] A plurality of sample images with labels can be composed into a training sample set and a test sample set. The training sample set is used to train the Wensheng graph model to be trained, and the test sample set is used to test the Wensheng graph model to be trained after training, so as to detect whether the Wensheng graph model to be trained after training meets the accuracy requirements of the model. If it meets the requirements, it can be applied in the actual Wensheng graph scene. If not, the Wensheng graph model to be trained is iteratively trained by adjusting the model parameters until a model that meets the accuracy requirements is obtained.
[0061] In the embodiments provided in the present application, by inputting sample images into at least two visual language models respectively, each visual language model outputs description data corresponding to the sample image respectively, and deduplication processing is performed on the description data to obtain description information corresponding to the sample image, so as to enrich the content and description method of the description information, which is conducive to improving the richness of the labels generated according to the description information, so that the sample images with labels can be suitable for the training of various text graph models; by inputting the description information into the topic classification model, generating the subject words corresponding to the description information, and setting the subject words as the labels corresponding to the sample images, so that the labels can accurately represent the content of the sample images, thereby improving the accuracy of the labels; through the image labeling method provided in the present application, manual labeling of sample images can be eliminated, thereby improving the labeling efficiency and the training efficiency of the text graph model.
[0062] Please refer to Figure 2 In some embodiments, the description data includes a description paragraph; S120 includes:
[0063] S210, performing word segmentation processing on the description segment to obtain a word segmentation phrase corresponding to the description segment;
[0064] S220, performing a deduplication process on the segmented word group corresponding to the sample image to obtain description information corresponding to the sample image.
[0065] The description segment includes one or more sentences, each of which can be divided into multiple phrases. The description segment can be segmented by a word segmentation tool or a word segmentation algorithm to obtain one or more segmented phrases.
[0066] For example, JIEBA word segmentation can be used to segment the description paragraph to obtain multiple segmentation phrases.
[0067] The segmentation phrases corresponding to the sample images are compared, and repeated segmentation phrases are deleted to reduce the volume of the obtained description information. The segmentation phrases for comparison can be obtained by segmenting description paragraphs output by different visual language models, or by segmenting description paragraphs output by the same visual language model.
[0068] By performing word segmentation on the description paragraph, it is possible to remove duplicate words at the phrase level, thereby improving the efficiency of removing duplicate words.
[0069] In some embodiments, the description data includes a description paragraph and a description phrase, and the at least two visual language models include a first visual language model and a second visual language model; S120 includes:
[0070] S310, inputting the sample image into a first visual language model to generate a description segment corresponding to the sample image;
[0071] The first visual language model can be Clip, Blip-2, Alt Clip, etc. The first visual language model outputs a description segment, which is a natural language text describing the sample image. The description segment can include elements, subjects, scenes, artistic styles, etc. in the sample image, for example: "A car parked in a garage, and in the yard outside the garage is an older man holding a water pipe and a rag."
[0072] The first visual language model can also be a model based on visual question answering. Questions such as "What is in the picture?" and "Please describe the picture" are pre-set, and the pre-set questions and sample images are input to the first visual language model. The answer output by the first visual language model is the description paragraph corresponding to the sample image.
[0073] S320, inputting the sample image into a second visual language model to generate a description phrase corresponding to the sample image;
[0074] The second visual language model may be ImageNet, etc. The second visual language model outputs a description phrase, where the description phrase is one or more phrases describing the sample image. The description phrase may include elements in the sample image, such as: "car, man, water pipe, rag, garage, yard".
[0075] S220 includes:
[0076] S330, performing word segmentation processing on the description segment to obtain a word segmentation phrase corresponding to the description segment;
[0077] S340 , performing a deduplication process on the segmentation phrases and description phrases corresponding to the sample image to obtain description information corresponding to the sample image.
[0078] Both the segmentation phrase and the description phrase are phrases, so that the two can be compared in pairs and repeated phrases can be deleted with high efficiency.
[0079] For example; the sample image is a documentary photo, showing "a car parked in a garage, and an older man holding a water pipe and a rag in the yard outside the garage". Visual language model A outputs description data based on the sample image as "a photo, the photo content is a car parked in a garage, and an older man holding a water pipe and a rag". The segmentation phrase obtained after word segmentation is "a, photo, photo, content, car, parked, in, garage, older, man, holding, water pipe, and, rag". Visual language model B outputs description data based on the sample image as "garage, car, man, water pipe, rag". Multiple description data are mixed to obtain "a, photo, photo, content, car, parked, in, garage, older, man, holding, water pipe, and, rag, garage, car, man, water pipe, rag". Deduplication processing is performed on the description elements in the description data to obtain "a, photo, content, car, parked, in, garage, older, man, holding, water pipe, and, rag".
[0080] In some embodiments, after S120, the method further includes:
[0081] S410, performing one or more operations of adding, replacing and deleting labels corresponding to sample images according to feature information of the Wensheng graph model to be trained, to obtain training labels matching the feature information, wherein the feature information includes the type of the Wensheng graph model to be trained and the image style output by the Wensheng graph model to be trained;
[0082] S420, inputting the training sample set into the to-be-trained Wensheng graph model to train the to-be-trained Wensheng graph model to obtain a trained Wensheng graph model, wherein the training sample set includes a plurality of sample images and training labels corresponding to the sample images.
[0083] Different Vincent graph models to be trained have different feature information, which can be inherent in the Vincent graph model to be trained, or can be set by technicians in this field according to the operation rules of the Vincent graph model to be trained. Different types of Vincent graph models to be trained can have different architectures, so that the effects of training different types of Vincent graph models to be trained using sample images with the same label are different. Therefore, labels can be added, replaced, and deleted according to the type of Vincent graph model to be trained, so that the obtained training labels are more suitable for the training of the Vincent graph model to be trained. The image style output by the Vincent graph model can be the drawing art style of the output image, or the output image can display the target element.
[0084] In this embodiment, one or more operations of adding, replacing and deleting are performed on the labels corresponding to the sample images to obtain training labels matching the feature information, so that the obtained training labels all satisfy the feature information of the Vincent graph model to be trained. The Vincent graph model to be trained is trained by inputting the sample images carrying the training labels into the Vincent graph model to be trained. The Vincent graph model to be trained can be any one of the models such as decision tree, artificial neural network, support vector machine, random forest and logistic regression.
[0085] By performing one or more operations of adding, replacing and deleting the labels corresponding to the sample images according to the feature information of the text graph model to be trained, the training labels matching the feature information are obtained, so that the sample image-label image-text pairs can be applied to the training of different text graph models to be trained.
[0086] In some embodiments, the image style includes a target training element; when the image model to be trained is a Lora model, S410 includes:
[0087] S510, obtaining a target training element in the feature information, deleting labels identical to the target training element from labels corresponding to the sample images, and obtaining training labels matching the feature information.
[0088] When the language graph model to be trained is a Lora (Low-Rank Adaptation of Large Language Models) model, it is necessary to delete the labels corresponding to the sample images so that there are no target training elements in the training labels carried by the sample images, so that the images generated by the trained Lora model by default should carry the target training elements.
[0089] For example: multiple sample images show "portraits with big eyes", the label of sample image A is "boy, big eyes, wearing blue clothes", the label of sample image B is "girl, big eyes, wearing pink clothes", and the label of sample image C is "middle-aged man, big eyes, wearing black clothes". When the target training element is "big eyes", delete the "big eyes" in the label. The label of sample image A is "boy, wearing blue clothes", the label of sample image B is "girl, wearing pink clothes", and the label of sample image C is "middle-aged man, wearing black clothes". Sample images A, B, and C all show the element of big eyes, which is not present in the training labels. Then the training sample set with sample images A, B, and C is input into the Lora model. The trained Lora model assumes that the "big eyes" element must exist for portraits. The trained Lora model outputs images with the "big eyes" element for any input text.
[0090] In some embodiments, the image style includes a target training element; when the image model to be trained is a Dreambooth model, S410 includes:
[0091] S610, obtaining a target training element in the feature information and an activation label corresponding to the target training element, and replacing the label identical to the target training element with the activation label in the label corresponding to the sample image to obtain a training label matching the feature information.
[0092] When the image model to be trained is a Dreambooth model, the labels corresponding to the sample images need to be replaced so that the training labels carried by the sample images contain activation labels, so that the trained Dreambooth can recognize and generate images with elements corresponding to the activation labels.
[0093] For example, multiple sample images show "portrait of a girl with pink hair", the label of sample image A is "girl, big eyes, wearing blue clothes", the label of sample image B is "girl, big eyes, wearing pink clothes", and the label of sample image C is "girl, small eyes, wearing black clothes". When the target training element is "girl" and the activation label is "pink-haired girl", the "girl" in the label is replaced. The label of sample image A is "pink-haired girl, big eyes, wearing blue clothes", the label of sample image B is "pink-haired girl, big eyes, wearing pink clothes", and the label of sample image C is "pink-haired girl, small eyes, wearing black clothes". Sample images A, B, and C all show the element of girl, and the training label contains the element of "pink-haired girl". Then the training sample set with sample images A, B, and C is input into the Dreambooth model, and the trained Dreambooth model learns the element of "pink-haired girl". The images output by the trained Dreambooth model for the input text of "pink-haired girl" all have the element of "pink-haired girl".
[0094] In some embodiments, for activation words that are not displayed in the label, you can also add labels as needed to obtain labels with activation words, for example:
[0095] Multiple sample images show "Girl Oil Painting Portrait", sample image A is labeled "Girl, big eyes, wearing blue clothes", sample image B is labeled "Girl, big eyes, wearing pink clothes", and sample image C is labeled "Girl, small eyes, wearing black clothes". It is necessary to train the oil painting style. When the activated label is "Oil Painting", "Oil Painting" is added to the label. The label of sample image A is "Girl, big eyes, wearing blue clothes, oil painting", the label of sample image B is "Girl, big eyes, wearing pink clothes, oil painting", and the label of sample image C is "Girl, small eyes, wearing black clothes, oil painting". Sample images A, B, and C all show the element of oil painting, and the training label contains the element of "Oil Painting". Then the training sample set with sample images A, B, and C is input into the Vincent graph model, and the trained Vincent graph model learns the element of "Oil Painting". The trained Vincent graph model outputs an oil painting image for the input text "Oil Painting".
[0096] The sample image-label image-text pairs provided in the present application can be used to train the Wensheng graph model based on SD and SDXL. In some embodiments, the labels corresponding to the sample images can also be adjusted based on SDXL, so that the adjusted training labels are consistent with or close to the expression of the SDXL original training set image-text pairs. Since the original training set image-text pairs of SDXL and SD1.5 are not exactly the same in terms of expression and image size, some focus more on elements, and some focus more on the correlation between objects. Those skilled in the art can adjust the sample image-label image-text pairs as needed, so that the adjusted sample image-training label image-text pairs meet the feature information of the Wensheng graph model to be trained.
[0097] Through the image labeling method provided in this application, long-term training can be performed on massive sample images. By using public weight files to fine-tune labels on the established small batch data set, complex and changeable images can be understood, and automatic data labeling of the SD and SDXL-based Wensheng graph model training sample sets can be efficiently completed.
[0098] Please refer to Figure 3 In some embodiments, S420 includes:
[0099] S710, inputting a sample image of a training sample set into an image decoder to obtain a first vector corresponding to the sample image;
[0100] S720, inputting the training label of the training sample set into the text decoder to obtain a second vector corresponding to the training label;
[0101] S730, forming a vector pair from a first vector and a second vector corresponding to the same sample image;
[0102] S740, input the vector pair into the to-be-trained text graph model, train the to-be-trained text graph model, and obtain a trained text graph model.
[0103] In this embodiment, decoders are set for images and labels respectively, and the extracted feature vectors are paired so that the part available for model training is a text-image corresponding feature module. S710-S730 can be a network structure designed based on transformer, which includes an image decoder (Image Encoder) for image visual feature extraction, a text decoder (Text Encoder) for text visual feature decoding, and then the first vector and the second vector are matched one by one. In the training process of the text-graph model to be trained, the parameters of the text encoder and the visual encoder are frozen and remain unchanged. By forming a vector pair with the first vector and the second vector corresponding to the same sample image, it is not necessary to classify the label of each sample image, only the description text segment of the image is required. These text-image pairs can obtain a large amount of information from the Internet; the vector pairs can be applied to multi-task scenarios, and can maintain the same or better performance without additional training, have extremely strong generalization ability, solve the problem of retraining downstream tasks; solve the multimodal differences between image vision and natural language text.
[0104] Based on the image marking method provided in the above embodiment, the present application also provides a specific implementation of the image marking device. Please refer to the following embodiment.
[0105] See first Figure 4 , the image marking device 400 provided in the embodiment of the present application includes the following modules:
[0106] A description information generating module 401 is used to input the sample image into at least two visual language models respectively, and each visual language model outputs description data corresponding to the sample image respectively;
[0107] A de-duplication processing module 402 is used to perform de-duplication processing on the description data to obtain description information corresponding to the sample image;
[0108] The label generation module 405 is used to input the description information into the topic classification model, generate topic words corresponding to the description information, and set the topic words as labels corresponding to the sample images.
[0109] In this embodiment, by inputting the sample image into at least two visual language models respectively, each visual language model outputs the description data corresponding to the sample image respectively, and the description data is deduplicated to obtain the description information corresponding to the sample image, so as to enrich the content and description method of the description information, which is conducive to improving the richness of the labels generated according to the description information, so that the sample image with the label can be suitable for the training of various text graph models; by inputting the description information into the topic classification model, generating the subject words corresponding to the description information, and setting the subject words as the label corresponding to the sample image, so that the label can accurately characterize the content of the sample image, thereby improving the accuracy of the label; through the image labeling method provided by the present application, manual labeling of sample images can be eliminated, thereby improving the labeling efficiency and the training efficiency of the text graph model.
[0110] As an implementation of the present application, the description data includes a description paragraph; the above-mentioned de-duplication word processing module 402 may include:
[0111] The word segmentation unit is used to perform word segmentation processing on the description paragraph to obtain the word segmentation phrase corresponding to the description paragraph;
[0112] The repeated word processing unit is used to remove repeated words from the segmented word groups corresponding to the sample image to obtain description information corresponding to the sample image.
[0113] As an implementation of the present application, the description data includes a description paragraph and a description phrase, and the at least two visual language models include a first visual language model and a second visual language model; the description information generation module 401 is further used to:
[0114] A first output unit, used to input the sample image into a first visual language model to generate a description segment corresponding to the sample image;
[0115] A second output unit, used to input the sample image into a second visual language model to generate a description phrase corresponding to the sample image;
[0116] The de-duplication processing module 402 is also used to perform word segmentation processing on the description paragraph to obtain the word segmentation phrase corresponding to the description paragraph; and perform word segmentation processing on the word segmentation phrase and the description phrase corresponding to the sample image to obtain the description information corresponding to the sample image.
[0117] As an implementation of the present application, the image marking device 400 further includes:
[0118] The label adjustment module 403 is used to perform one or more operations of adding, replacing and deleting labels corresponding to sample images according to feature information of the Wensheng graph model to be trained, so as to obtain training labels matching the feature information, wherein the feature information includes the type of the Wensheng graph model to be trained and the image style output by the Wensheng graph model to be trained;
[0119] The training module 404 is used to input the training sample set into the to-be-trained Wensheng graph model to train the to-be-trained Wensheng graph model to obtain a trained Wensheng graph model. The training sample set includes a plurality of sample images and training labels corresponding to the sample images.
[0120] As an implementation of the present application, when the text graph model to be trained is a Lora model, the label adjustment module 403 includes:
[0121] The deletion unit is used to obtain the target training element in the feature information, delete the label that is the same as the target training element in the label corresponding to the sample image, and obtain the training label that matches the feature information.
[0122] As an implementation of the present application, when the text graph model to be trained is a Dreambooth model, the label adjustment module 403 includes:
[0123] The replacement unit is used to obtain the target training element in the feature information and the activation label corresponding to the target training element, and replace the label identical to the target training element with the activation label in the label corresponding to the sample image to obtain the training label matching the feature information.
[0124] As an implementation of the present application, the training module 404 includes:
[0125] A first decoding unit, used for inputting a sample image of a training sample set into an image decoder to obtain a first vector corresponding to the sample image;
[0126] A second decoding unit, used for inputting the training label of the training sample set into a text decoder to obtain a second vector corresponding to the training label;
[0127] A matching unit, used for forming a vector pair from a first vector and a second vector corresponding to the same sample image;
[0128] The training unit is used to input the vector pair into the to-be-trained text graph model, train the to-be-trained text graph model, and obtain a trained text graph model.
[0129] The image marking device provided in the embodiment of the present invention can implement each step in the above method embodiment, and will not be described again here to avoid repetition.
[0130] Figure 5A schematic diagram of the hardware structure of the image marking device provided in an embodiment of the present application is shown.
[0131] The image marking device may include a processor 1001 and a memory 1002 storing computer program instructions.
[0132] Specifically, the processor 1001 may include a central processing unit (CPU), or an application specific integrated circuit (ASIC), or may be configured to implement one or more integrated circuits of the embodiments of the present application.
[0133] The memory 1002 may include a large capacity memory for data or instructions. By way of example and not limitation, the memory 1002 may include a hard disk drive (HDD), a floppy disk drive, a flash memory, an optical disk, a magneto-optical disk, a magnetic tape, or a universal serial bus (USB) drive or a combination of two or more of these. In appropriate cases, the memory 1002 may include a removable or non-removable (or fixed) medium. In appropriate cases, the memory 1002 may be inside or outside the integrated gateway disaster recovery device. In a specific embodiment, the memory 1002 is a non-volatile solid-state memory.
[0134] The memory may include read-only memory (ROM), random access memory (RAM), magnetic disk storage media devices, optical storage media devices, flash memory devices, electrical, optical or other physical / tangible memory storage devices. Thus, typically, the memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to an aspect of the present disclosure.
[0135] The processor 1001 implements any one of the image marking methods in the above embodiments by reading and executing computer program instructions stored in the memory 1002 .
[0136] In one example, the image marking device may further include a communication interface 1003 and a bus 1010. Figure 5 As shown, the processor 1001, the memory 1002, and the communication interface 1003 are connected via a bus 1010 and communicate with each other.
[0137] The communication interface 1003 is mainly used to implement communication between various modules, devices, units and / or equipment in the embodiments of the present application.
[0138] Bus 1010 includes hardware, software or both, and the parts of image marking equipment are coupled to each other. For example, but not limitation, bus may include accelerated graphics port (AGP) or other graphics bus, enhanced industrial standard architecture (EISA) bus, front-end bus (FSB), hypertransport (HT) interconnection, industrial standard architecture (ISA) bus, infinite bandwidth interconnection, low pin count (LPC) bus, memory bus, micro channel architecture (MCA) bus, peripheral component interconnection (PCI) bus, PCI-Express (PCI-X) bus, serial advanced technology attachment (SATA) bus, video electronics standard association local (VLB) bus or other suitable bus or two or more of these combinations. In appropriate cases, bus 1010 may include one or more buses. Although the present application embodiment describes and shows a specific bus, the present application considers any suitable bus or interconnection.
[0139] The image marking device may be based on the above embodiment, thereby realizing the image marking method and device combined with the above embodiment.
[0140] In addition, in combination with the image marking method in the above embodiment, the embodiment of the present application may provide a computer storage medium to implement it. The computer storage medium stores computer program instructions; when the computer program instructions are executed by the processor, any one of the image marking methods in the above embodiment is implemented, and the same technical effect can be achieved. In order to avoid repetition, it will not be repeated here. Among them, the above-mentioned computer-readable storage medium may include non-transitory computer-readable storage media, such as read-only memory (Read-Only Memory, referred to as ROM), random access memory (Random Access Memory, referred to as RAM), a disk or an optical disk, etc., which are not limited here.
[0141] In addition, an embodiment of the present application also provides a vehicle, including computer program instructions, which, when executed by a processor, can implement the steps and corresponding contents of the aforementioned method embodiment.
[0142] It should be clear that the present application is not limited to the specific configuration and processing described above and shown in the figures. For the sake of simplicity, a detailed description of the known method is omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of the present application is not limited to the specific steps described and shown, and those skilled in the art can make various changes, modifications and additions, or change the order between the steps after understanding the spirit of the present application.
[0143] The functional blocks shown in the above block diagram can be implemented as hardware, software, firmware or a combination thereof. When implemented in hardware, it can be, for example, an electronic circuit, an application specific integrated circuit (ASIC), appropriate firmware, a plug-in, a function card, etc. When implemented in software, the elements of the present application are programs or code segments that are used to perform the required tasks. The program or code segment can be stored in a machine-readable medium, or transmitted on a transmission medium or a communication link by a data signal carried in a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, optical fiber media, radio frequency (RF) links, etc. The code segment can be downloaded via a computer network such as the Internet, an intranet, etc.
[0144] It should also be noted that the exemplary embodiments mentioned in this application describe some methods or systems based on a series of steps or devices. However, this application is not limited to the order of the above steps, that is, the steps can be performed in the order mentioned in the embodiment, or in a different order from the embodiment, or several steps can be performed simultaneously.
[0145] Aspects of the present disclosure are described above with reference to the flowcharts and / or block diagrams of the methods, devices and vehicles according to the embodiments of the present disclosure. It should be understood that each box in the flowchart and / or block diagram and the combination of each box in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device to produce a machine so that these instructions executed by the processor of the computer or other programmable data processing device enable the implementation of the function / action specified in one or more boxes of the flowchart and / or block diagram. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field programmable logic circuit. It can also be understood that each box in the block diagram and / or flowchart and the combination of boxes in the block diagram and / or flowchart can also be implemented by dedicated hardware that performs a specified function or action, or can be implemented by a combination of dedicated hardware and computer instructions.
[0146] The above are only specific implementation methods of the present application. Those skilled in the art can clearly understand that for the convenience and simplicity of description, the specific working processes of the systems, modules and units described above can refer to the corresponding processes in the aforementioned method embodiments, and will not be repeated here. It should be understood that the protection scope of the present application is not limited to this. Any technician familiar with the technical field can easily think of various equivalent modifications or replacements within the technical scope disclosed in this application, and these modifications or replacements should be included in the protection scope of this application.
Claims
1. An image labeling method, It is characterized in that The method comprises: Inputting the sample images into at least two visual language models respectively, and each of the visual language models outputs description data corresponding to the sample images respectively; Deleting duplicate words from the description data to obtain description information corresponding to the sample image; The description information is input into a topic classification model, a topic word corresponding to the description information is generated, and the topic word is set as a label corresponding to the sample image.
2. The image marking method according to claim 1, It is characterized in that The description data includes a description paragraph; the deduplication processing of the description data to obtain description information corresponding to the sample image includes: Performing word segmentation processing on the description segment to obtain a word segmentation phrase corresponding to the description segment; The segmented word group corresponding to the sample image is processed to remove duplicate words to obtain description information corresponding to the sample image.
3. The image marking method according to claim 1, It is characterized in that The description data includes a description paragraph and a description phrase, and the at least two visual language models include a first visual language model and a second visual language model; and the step of inputting the sample image into the at least two visual language models to generate description data corresponding to the sample image includes: Inputting the sample image into the first visual language model to generate a description segment corresponding to the sample image; Inputting the sample image into the second visual language model to generate a description phrase corresponding to the sample image; The performing of deduplication processing on the description data to obtain description information corresponding to the sample image includes: Performing word segmentation processing on the description segment to obtain a word segmentation phrase corresponding to the description segment; The segmentation phrases and the description phrases corresponding to the sample image are processed to remove duplicate words, so as to obtain description information corresponding to the sample image.
4. The image marking method according to any one of claims 1 to 3, It is characterized in that After inputting the description information into a topic classification model, generating a topic word corresponding to the description information, and setting the topic word as a label corresponding to the sample image, the method further includes: According to the feature information of the Wensheng graph model to be trained, one or more operations of adding, replacing and deleting are performed on the label corresponding to the sample image to obtain a training label matching the feature information, wherein the feature information includes the type of the Wensheng graph model to be trained and the image style output by the Wensheng graph model to be trained; The training sample set is input into the to-be-trained Vincent graph model to train the to-be-trained Vincent graph model to obtain a trained Vincent graph model, wherein the training sample set includes a plurality of the sample images and the training labels corresponding to the sample images.
5. The image marking method according to claim 4, It is characterized in that The image style includes a target training element; when the to-be-trained Wensheng graph model is a Lora model, the adding, replacing and deleting one or more operations of the label corresponding to the sample image according to the feature information of the to-be-trained Wensheng graph model includes: A target training element in the feature information is obtained, and a label identical to the target training element is deleted from labels corresponding to the sample images to obtain a training label matching the feature information.
6. The image marking method according to claim 5, It is characterized in that The image style includes a target training element; when the to-be-trained Wensheng graph model is a Dreambooth model, the adding, replacing and deleting one or more operations of the label corresponding to the sample image according to the feature information of the to-be-trained Wensheng graph model includes: A target training element in the feature information and an activation label corresponding to the target training element are obtained, and among the labels corresponding to the sample images, the label identical to the target training element is replaced with the activation label to obtain a training label matching the feature information.
7. The image marking method according to claim 5, It is characterized in that The step of inputting the training sample set into the to-be-trained Wensheng graph model to train the to-be-trained Wensheng graph model to obtain the trained Wensheng graph model comprises: Inputting a sample image of the training sample set into an image decoder to obtain a first vector corresponding to the sample image; Inputting the training labels of the training sample set into a text decoder to obtain a second vector corresponding to the training labels; The first vector and the second vector corresponding to the same sample image form a vector pair; The vector pair is input into the to-be-trained Wensheng graph model, and the to-be-trained Wensheng graph model is trained to obtain the trained Wensheng graph model.
8. An image marking device, It is characterized in that The device comprises: A description information generation module, used to input the sample image into at least two visual language models respectively, and each of the visual language models outputs description data corresponding to the sample image respectively; A de-duplication word processing module, used for performing de-duplication word processing on the description data to obtain description information corresponding to the sample image; The label generation module is used to input the description information into a topic classification model, generate a topic word corresponding to the description information, and set the topic word as a label corresponding to the sample image.
9. An image marking device, It is characterized in that The image marking device includes: a processor and a memory storing computer program instructions; When the processor executes the computer program instructions, the image marking method according to any one of claims 1 to 7 is implemented.
10. A computer storage medium, It is characterized in that The computer storage medium stores computer program instructions, and when the computer program instructions are executed by the processor, the image marking method according to any one of claims 1 to 7 is implemented.