Large model picture training set generation method and device, electronic equipment and storage medium
By manually labeling some images and training the visual-language model, most images can be automatically labeled, solving the problem of high cost and low efficiency in generating large model image training sets, and achieving efficient and low-cost generation of image training sets.
Patent Information
- Application Number
- CN202510703151.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-29
- Publication Date
- 2025-10-28
AI Technical Summary
In the existing technology, the generation cost of large model image training sets is high and the efficiency is low, mainly due to the inefficiency of manual labeling.
Manually annotate some images and train the labeling model, use the visual-language model to automatically annotate data, generate synthetic images, and establish an image training set. Only some images need to be manually annotated, and the visual-language model can automatically complete most of the annotations.
It reduces the cost of generating large model image training sets, improves generation efficiency, and enhances the quality of image training sets and the generalization ability of large models.
Smart Images

Figure CN120853172A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of large model technology, and in particular to a method, apparatus, electronic device, and storage medium for generating image training sets for large models. Background Art
[0002] Images used to train large models usually need to be labeled. Currently, the labeling method is to manually label each image. Manual labeling is costly and inefficient, which leads to high cost and low efficiency in generating image training sets. Summary of the Invention
[0003] This invention provides a method, apparatus, electronic device, and storage medium for generating image training sets of large models, in order to solve the technical problems of high generation cost and low generation efficiency of image training sets of large models in the prior art.
[0004] This invention provides a method for generating image training sets for large models, comprising:
[0005] Obtain the image dataset built for the large model;
[0006] After manually annotating a portion of the images in the image dataset, the manually annotated images are used to train a labeling model.
[0007] Each image in the image dataset is input into the trained labeling model to obtain the labeled data output by the labeling model;
[0008] Establish an image training set consisting of each image and its corresponding labeled data.
[0009] According to the present invention, a method for generating image training sets for a large model is provided, wherein the labeling model is a visual-language model.
[0010] According to the present invention, a method for generating image training sets for a large model is provided, wherein the labeled data includes first-level labels;
[0011] After establishing the image training set consisting of each image and its corresponding labeled data, the method further includes:
[0012] Count the number of samples for each of the first-level labels in the image training set;
[0013] The primary labels whose sample size is less than a preset threshold are identified as candidate labels;
[0014] Determine the target label among the candidate labels;
[0015] Generate a composite image corresponding to the target label;
[0016] Add the synthesized image to the image dataset.
[0017] According to the improved method for generating image training sets for large models based on the present invention, the labeled data further includes secondary labels under the primary labels;
[0018] Determining the target label among the candidate labels includes:
[0019] Obtain the secondary labels of each sample of the candidate labels and the preset secondary label library corresponding to the candidate labels;
[0020] If there exists a label in the preset secondary label library that is different from the secondary labels of each sample of the candidate label, then the candidate label is determined to be the target label.
[0021] According to the improved image training set generation method for a large model of the present invention, determining the target label among the candidate labels includes:
[0022] The large model is trained using the image training set;
[0023] Determine the classification accuracy of the trained large model for each of the candidate labels;
[0024] The candidate labels whose classification accuracy is lower than a preset percentage threshold are determined as the target labels.
[0025] According to the improved method for generating image training sets for a large model based on the present invention, the step of generating a synthetic image corresponding to the target label includes:
[0026] Determine the keywords corresponding to the target tags;
[0027] The topic terms are input into the inference model to obtain the subtopic terms output by the inference model;
[0028] Each of the sub-topic terms is input into the text-to-image model to obtain the synthesized image output by the text-to-image model.
[0029] According to the present invention, a method for generating image training sets for a large model is provided, wherein the inference model is DeepseekR1 or the text-to-image model is Stable Diffusion XL.
[0030] The present invention also provides an apparatus for generating image training sets for large models, comprising:
[0031] The acquisition module is used to acquire the image dataset built for the large model;
[0032] The training module is used to train a labeling model using manually labeled images after manually labeling a portion of the images in the image dataset.
[0033] The annotation module is used to input each image in the image dataset into the trained labeling model to obtain the annotation data output by the labeling model;
[0034] A module is established to create an image training set consisting of each image and its corresponding labeled data.
[0035] The present invention also provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the program to implement the image training set generation method for the large model as described above.
[0036] The present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the image training set generation method for a large model as described above.
[0037] The present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the image training set generation method for a large model as described above.
[0038] The present invention provides a method, apparatus, electronic device, and storage medium for generating image training sets for large models. It trains a labeling model using manually labeled images, enabling the model to annotate images. By inputting each image from the image dataset into the trained labeling model, the model outputs labeled data, which is then used to build an image training set. Only a portion of the images need manual annotation during this process; the annotation of most images is automatically completed by the labeling model. This reduces the cost and improves the efficiency of generating large model image training sets. Attached Figure Description
[0039] To more clearly illustrate the technical solutions in this invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are some embodiments of this invention. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.
[0040] Figure 1 This is one of the flowcharts illustrating the method for generating image training sets for large models provided by this invention.
[0041] Figure 2This is the second flowchart illustrating the method for generating image training sets for large models provided by this invention.
[0042] Figure 3 This is a schematic diagram of the structure of the image training set generation device for large models provided by the present invention.
[0043] Figure 4 It is a structural schematic diagram of the electronic device provided by the present invention. DETAILED DESCRIPTION
[0044] To make the objectives, technical solutions, and advantages of this invention clearer, the technical solutions of this invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without creative effort are within the scope of protection of this invention.
[0045] It should be noted that in the description of this invention, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element. The terms "upper," "lower," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Unless otherwise expressly specified and limited, the terms "installed," "connected," and "linked" should be interpreted broadly, for example, as a fixed connection, a detachable connection, or an integral connection; a mechanical connection or an electrical connection; a direct connection or an indirect connection through an intermediate medium; or a connection within two elements. Those skilled in the art can understand the specific meaning of the above terms in this invention according to the specific circumstances.
[0046] The terms "first," "second," etc., used in this invention are used to distinguish similar objects, not to describe a specific order or sequence. It should be understood that such data can be interchanged where appropriate so that embodiments of the invention can be implemented in orders other than those illustrated or described herein, and the objects distinguished by "first," "second," etc., are generally of the same class, without limiting the number of objects; for example, a first object can be one or more. Furthermore, "and / or" indicates at least one of the connected objects, and the character " / " generally indicates that the preceding and following objects are in an "or" relationship.
[0047] The following is combined Figures 1-4 This invention describes the method, apparatus, electronic device, and storage medium for generating image training sets for large models.
[0048] like Figure 1 As shown, the method for generating image training sets for large models provided by the present invention includes steps S1-S4.
[0049] Step S1: Obtain the image dataset for the large model.
[0050] Keywords can be generated using LLM (Large Language Model), and images can be batch-fetched from the internet using a web crawler. After cleaning the crawled images, a dataset of remaining images is obtained. The web crawler of this invention can crawl 10,000 images from the internet per day, with a throughput of up to 200 MB / s.
[0051] Step S2: After manually annotating some images in the image dataset, use the manually annotated images to train the labeling model.
[0052] Images with prominent features in the image dataset can be manually selected and labeled, including primary labels and secondary labels (primary label - secondary label), such as: "Child protection - inappropriate interactions with children", "Bullying and harassment - cyberbullying", and "Bloody - hanging".
[0053] After training the labeling model with manually labeled images, the labeling model will have the ability to label images.
[0054] Step S3: Input each image in the image dataset into the trained labeling model to obtain the labeled data output by the labeling model.
[0055] The labeled data can include primary tags and secondary tags under the primary tags, such as "bloody - hanging", "bloody - execution", "child protection - inappropriate interaction with children", "child protection - sexual exploitation / abuse", "child protection - inappropriate exposure and nudity (children)", "bullying and harassment - cyberbullying", "bullying and harassment - personal attacks", etc.
[0056] The tagging model of this invention can tag up to 100 images per minute.
[0057] Step S4: Establish an image training set consisting of each image and its corresponding labeled data.
[0058] A sample in the image training set consists of an image and its primary and secondary labels.
[0059] As described above, the image training set generation method of the present invention uses manually annotated images to train a labeling model, enabling the labeling model to annotate images. By inputting each image in the image dataset into the trained labeling model, the labeled data output by the labeling model can be obtained. This allows for the further establishment of an image training set composed of each image and its corresponding labeled data. In the process of establishing the image training set, only a portion of the images need to be manually annotated, while the annotation of most images is automatically completed by the labeling model. This reduces the generation cost of large model image training sets and improves the generation efficiency of large model image training sets.
[0060] In some embodiments, the labeling model of the present invention can be a vision-language model (VLM).
[0061] Visual Imagery Model (VLM) is an artificial intelligence model that integrates image (visual) and text (linguistic) data, aiming to achieve cross-modal semantic understanding and generation. The core objective of VLM is to enable the model to simultaneously process visual signals (such as images and videos) and linguistic signals (such as text and speech), and establish the correlation between the two. The VLM architecture mainly consists of three parts: a visual encoder, a linguistic decoder, and a cross-modal interaction module. The visual encoder converts visual inputs such as images / videos into numerical feature vectors (visual embeddings). The linguistic decoder generates text based on visual-linguistic associations. The cross-modal interaction module establishes semantic alignment between visual and linguistic features, achieving cross-modal information fusion.
[0062] The loss function used in this invention to train the labeling model can be the autoregressive cross-entropy loss function.
[0063] After step S4, the present invention will use the image training set to train a large model. Considering that if the number of samples of a certain first-level label in the image training set is insufficient, it may lead to poor generalization ability of the large model after training.
[0064] To avoid poor generalization ability of large models due to insufficient sample size, in some implementations, such as Figure 2 As shown, after step S4, the image training set generation method of the present invention may further include:
[0065] Step S5: Count the number of samples for each primary label in the image training set;
[0066] Step S6: Determine the primary labels whose sample size is less than a preset threshold as candidate labels;
[0067] Step S7: Determine the target label among the candidate labels;
[0068] Step S8: Generate the composite image corresponding to the target label;
[0069] Step S9: Add the synthesized image to the image dataset.
[0070] In this context, a sample size of less than a preset threshold for a primary label indicates an insufficient number of samples for that primary label. The preset threshold can be 1000.
[0071] For example, suppose there are primary labels "bloody", "child protection" and "bullying and harassment". The number of samples for the primary labels "bloody", "child protection" and "bullying and harassment" in the image training set are 500, 1000 and 2000 respectively. If the number of samples for the primary label "bloody" is less than 1000, then the primary label "bloody" is determined as a candidate label.
[0072] Step S7 can directly use each candidate label as the target label, for example, directly use the first-level label "bloody" as the target label.
[0073] Step S8 can generate corresponding synthetic images based on the target label. The primary label of the subsequently labeled synthetic images should be the corresponding target label. The number of synthetic images to be generated is determined based on the principle that the number of target label samples in the image training set can reach a preset threshold.
[0074] After adding the synthesized image to the image dataset in step S9, the number of target label samples in the image training set obtained after steps S1-S4 will increase.
[0075] For example, 500 synthetic images corresponding to the primary label "bloody" are generated. After steps S1-S4, these 500 synthetic images are labeled to obtain 500 primary label "bloody" samples. Adding these 500 primary label "bloody" samples to the original 500 primary label "bloody" samples in the image training set, the number of primary label "bloody" samples reaches 1000.
[0076] Of course, if the number of samples for a primary label exceeds the preset threshold, the primary label will not be considered as a candidate label, and the number of samples for that primary label will not be increased.
[0077] This can increase the number of primary label samples that are insufficient, improve the quality of the image training set, and help improve the generalization ability of large models.
[0078] As mentioned above, step S7 can directly use each candidate label as the target label. However, if the sample of the candidate label covers all the secondary labels in the preset secondary label library corresponding to the candidate label, even if the number of samples of the candidate label is insufficient, the generalization ability of the trained large model can be better. In this case, it is unnecessary to use the candidate label as the target label to increase the number of samples of the candidate label.
[0079] Therefore, in order to avoid unnecessarily increasing the number of candidate label samples, in some implementations, step S7 may further include:
[0080] Obtain the secondary labels of each sample of the candidate labels and the preset secondary label library corresponding to the candidate labels;
[0081] If a pre-defined secondary label library contains a label that is different from the secondary labels of all samples of the candidate label, then the candidate label is determined as the target label.
[0082] Each primary tag has a pre-defined secondary tag library containing all secondary tags under that primary tag. For example, the pre-defined secondary tag library for the primary tag "bloody" contains all secondary tags including "human bloody and corpse", "animal bloody and corpse", "blood", "car accident scene", "human burning", "hanging", "execution", "violent crushing", and "explosion".
[0083] If the pre-defined secondary label library contains labels that are different from the secondary labels of all samples with the candidate label, it means that the samples with the candidate label do not completely cover all the secondary labels in the pre-defined secondary label library corresponding to the candidate label, which will result in poor generalization ability of the trained large model. For example, if the secondary labels of the samples with the candidate label "bloody" include "hanging" and "execution", but the secondary labels of the samples with the candidate label "bloody" do not include "human bloody and corpse", "animal bloody and corpse", "blood", "car accident scene", "human burning", "violent crushing", and "explosion", then the candidate label "bloody" will be used as the target label.
[0084] Of course, if the sample of the candidate label completely covers all the secondary labels in the preset secondary label library corresponding to the candidate label, then the candidate label can be excluded as the target label.
[0085] This allows candidate labels to be used as target labels only when the sample of candidate labels does not completely cover all secondary labels, thus avoiding an unnecessary increase in the number of candidate label samples.
[0086] The above discussion determines whether to use a candidate label as the target label based on the coverage of secondary labels by the candidate label's samples. In some cases, even if the candidate label's samples do not completely cover all secondary labels, the generalization ability of the trained large model may still be good. In such cases, it is unnecessary to use the candidate label as the target label to increase the number of samples for that candidate label.
[0087] To more accurately determine the target label and thus improve the generalization ability of large models, in some implementations, step S7, determining the target label among the candidate labels, may further include:
[0088] Train a large model using an image training set;
[0089] Determine the classification accuracy of the trained large model for each candidate label;
[0090] Candidate labels with a classification accuracy rate lower than a preset percentage threshold are identified as target labels.
[0091] Specifically, if the classification accuracy of the large model for the candidate label is lower than a preset percentage threshold, it indicates that the large model has a low classification accuracy for that candidate label. The preset percentage threshold can be 70%.
[0092] If a large model has a low classification accuracy for a certain candidate label, increasing the number of samples for that candidate label will definitely improve the classification accuracy of the large model for that candidate label. Therefore, it is necessary to use that candidate label as the target label to increase the number of samples for that candidate label.
[0093] Of course, if the classification accuracy of the trained large model for the candidate label is higher than the preset percentage threshold, the candidate label may not be used as the target label.
[0094] This way, candidate labels are only used as target labels when the classification accuracy of the large model is below a preset percentage threshold, which avoids unnecessarily increasing the number of candidate label samples.
[0095] The method described above, which determines whether to use a candidate label as the target label based on the coverage of secondary labels by the candidate label samples, does not depend on the output of the large model. However, the method that determines whether to use a candidate label as the target label based on the classification accuracy of the large model for the candidate label depends on the output of the large model.
[0096] In some embodiments, step S8 of the present invention may further include:
[0097] Identify the keywords corresponding to the target tags;
[0098] Input the keywords into the inference model to obtain the sub-keywords output by the inference model;
[0099] Input each sub-topic word into the text-to-image model to obtain the synthesized image output by the text-to-image model.
[0100] For example, if the target label is "bloody," the corresponding topic could be "Bloody and Violence Behaviors." Inputting "Bloody and Violence Behaviors" into the inference model, the output sub-topics could include "Battlefield Casualty Scenes," "Armed Violence at Crime Scenes," "Domestic Violence Physical Altercations," "Historical Execution Devices Display," "Terrorist Attack Explosion Sites," "Animal Predation Bloody Process," "Medical Accident Trauma Close-ups," "Religious Sacrifice Ceremony Scenes," "Sports Competition Violence Injuries," and "Gang Wars Gunfight." Scenes include gang warfare and gunfights, school bullying extreme cases, natural disaster casualties, psychopathic torture and murder behavior, film and game blood special effects, and cold weapon era combat injuries. Because the text-based image model has better support for English, English descriptions are used.
[0101] It is understandable that the more sub-topics a keyword corresponds to and the stronger their relevance, the more relevant the generated composite image will be to the target label.
[0102] In some implementations, the inference model of this invention can be Deepseek R1, and the text-to-image model can be Stable Diffusion XL. Stable Diffusion XL is the latest generation text-to-image generation model developed by Stability AI, utilizing deep learning and neural networks to learn and analyze based on large amounts of image data. The Stable Diffusion XL model generates new images by understanding the features and semantic information of images, based on prompts or conditions provided by the user. This invention achieves a throughput of 8.5 images per minute when generating synthetic images using Deepseek R1 and Stable Diffusion XL.
[0103] Furthermore, after generating the synthesized image, this invention can also perform image topic similarity comparison based on the image text embedding model (CLIP-base-p16), and simultaneously calculate the image quality index (FID) to ensure image quality and topic relevance.
[0104] like Figure 3 As shown, the large model image training set generation device provided by the present invention includes:
[0105] The acquisition module is used to acquire the image dataset built for the large model;
[0106] The training module is used to train a labeling model using manually labeled images after manually annotating a portion of the images in the image dataset.
[0107] The annotation module is used to input each image in the image dataset into the trained labeling model and obtain the labeled data output by the labeling model.
[0108] The module is used to create an image training set consisting of each image and its corresponding annotation data.
[0109] It should be noted that the large model image training set generation device provided by the present invention can execute the large model image training set generation method of any of the above embodiments during specific operation, which will not be elaborated in this embodiment.
[0110] In some embodiments, the labeling model of the present invention can be a visual-language model.
[0111] In some implementations, the annotation data may include first-level labels;
[0112] The image training set generation device may also include:
[0113] The statistics module is used to count the number of samples for each primary label in the image training set;
[0114] The determination module is used to identify primary labels whose sample size is less than a preset threshold as candidate labels.
[0115] And the target label used to determine the candidate label;
[0116] The generation module is used to generate composite images corresponding to the target labels;
[0117] Add a module to add composite images to the image dataset.
[0118] In some implementations, the labeled data may also include secondary labels under the primary labels;
[0119] The determination module can also be used for:
[0120] Obtain the secondary labels of each sample of the candidate labels and the preset secondary label library corresponding to the candidate labels;
[0121] If a pre-defined secondary label library contains a label that is different from the secondary labels of all samples of the candidate label, then the candidate label is determined as the target label.
[0122] In some implementations, the determining module can also be used for:
[0123] Train a large model using an image training set;
[0124] Determine the classification accuracy of the trained large model for each candidate label;
[0125] Candidate labels with a classification accuracy rate lower than a preset percentage threshold are identified as target labels.
[0126] In some implementations, the generation module can also be used for:
[0127] Identify the keywords corresponding to the target tags;
[0128] Input the keywords into the inference model to obtain the sub-keywords output by the inference model;
[0129] Input each sub-topic word into the text-to-image model to obtain the synthesized image output by the text-to-image model.
[0130] In some implementations, the inference model can be Deepseek R1 or the text-based graph model can be StableDiffusion XL.
[0131] Figure 4 This is a schematic diagram of the structure of the electronic device provided by the present invention, such as... Figure 4As shown, the electronic device may include a processor, a communications interface, memory, and a communication bus, wherein the processor, communications interface, and memory communicate with each other via the communication bus. The processor can call logical instructions in the memory to execute a method for generating image training sets for a large model.
[0132] Furthermore, the logical instructions in the aforementioned memory can be implemented as software functional units and sold or used as independent products, and can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0133] On the other hand, the present invention also provides a computer program product, the computer program product including a computer program stored on a non-transitory computer-readable storage medium, the computer program including program instructions, and when the program instructions are executed by a computer, the computer is able to execute the image training set generation method for large models provided in the above embodiments.
[0134] In another aspect, the present invention also provides a non-transitory computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, is implemented to perform the image training set generation method for the large model provided in the above embodiments.
[0135] The device embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without any creative effort.
[0136] Through the description of the above embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus a necessary general hardware platform, or of course, by hardware. Based on this understanding, the essence of the above technical solution or the part that contributes to the existing technology can be embodied in the form of a software product. The computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a magnetic disk, an optical disk, etc., and includes a number of instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or certain parts of the embodiments.
[0137] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.
Claims
1. A method for generating image training sets for a large model, characterized in that, include: Obtain the image dataset built for the large model; After manually annotating a portion of the images in the image dataset, the manually annotated images are used to train a labeling model. Each image in the image dataset is input into the trained labeling model to obtain the labeled data output by the labeling model; Establish an image training set consisting of each image and its corresponding labeled data.
2. The method for generating image training sets for large models according to claim 1, characterized in that, The labeling model is a visual-language model.
3. The method for generating image training sets for large models according to claim 1, characterized in that, The labeled data includes primary labels; After establishing the image training set consisting of each image and its corresponding labeled data, the method further includes: Count the number of samples for each of the first-level labels in the image training set; The primary labels whose sample size is less than a preset threshold are identified as candidate labels; Determine the target label among the candidate labels; Generate a composite image corresponding to the target label; Add the synthesized image to the image dataset.
4. The method for generating image training sets for large models according to claim 3, characterized in that, The labeled data also includes secondary labels under the primary labels; Determining the target label among the candidate labels includes: Obtain the secondary labels of each sample of the candidate labels and the preset secondary label library corresponding to the candidate labels; If there exists a label in the preset secondary label library that is different from the secondary labels of each sample of the candidate label, then the candidate label is determined to be the target label.
5. The method for generating image training sets for large models according to claim 3, characterized in that, Determining the target label among the candidate labels includes: The large model is trained using the image training set; Determine the classification accuracy of the trained large model for each of the candidate labels; The candidate labels whose classification accuracy is lower than a preset percentage threshold are determined as the target labels.
6. The method for generating image training sets for large models according to claim 3, characterized in that, The process of generating the composite image corresponding to the target label includes: Determine the keywords corresponding to the target tags; The topic terms are input into the inference model to obtain the subtopic terms output by the inference model; Each of the sub-topic terms is input into the text-to-image model to obtain the synthesized image output by the text-to-image model.
7. The method for generating image training sets for large models according to claim 6, characterized in that, The inference model is Deepseek R1 or the text-based graph model is Stable Diffusion XL.
8. A device for generating image training sets for a large model, characterized in that, include: The acquisition module is used to acquire the image dataset built for the large model; The training module is used to train a labeling model using manually labeled images after manually labeling a portion of the images in the image dataset. The annotation module is used to input each image in the image dataset into the trained labeling model to obtain the annotation data output by the labeling model; A module is established to create an image training set consisting of each image and its corresponding labeled data.
9. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the image training set generation method for the large model as described in any one of claims 1 to 7.
10. A non-transitory computer-readable storage medium having a computer program stored thereon, characterized in that, When the computer program is executed by the processor, it implements the image training set generation method for the large model as described in any one of claims 1 to 7.