Training method and device of low-rank adaptive model, electronic equipment and storage medium
By fusing the initial image and the background image to generate an extended image set, and combining iterative training with the pedestal model, the problem of insufficient data in the training of low-rank adaptive models is solved, thereby improving the accuracy of the model and the effect of target subject generation.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- BEIJING CO WHEELS TECH CO LTD
- Filing Date
- 2024-11-26
- Publication Date
- 2026-06-02
AI Technical Summary
The limited availability of image data containing specific subjects in existing technologies leads to insufficient accuracy during the training of low-rank adaptive models.
By fusing the target subject of each initial image in the initial image set with a preset background image, a first extended image set is generated. The initial low-rank adaptation model is then trained using the first pedestal model. Extended images are generated by combining text prompts, and iterative training is performed until the conditions are met, resulting in a trained low-rank adaptation model.
This method enables the training of a low-rank adaptive model using a small amount of image data, improving the accuracy of model training and the image generation effect of the target subject.
Smart Images

Figure CN122134849A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of data processing technology, and in particular to a training method, apparatus, electronic device, and storage medium for a low-rank adaptive model. Background Technology
[0002] With the development of artificial intelligence technology, text-to-image generation technology has developed rapidly. Text-to-image generation technology is a technique that generates corresponding images based on user-input text descriptions.
[0003] While text-based image models can generate aesthetically pleasing images containing specific objects, with technological advancements, users increasingly desire images with specific subjects. This necessitates training low-rank adaptive models using image data containing such subjects. However, the availability of images with specific subjects in current technologies is limited, thus reducing the accuracy of model training. Summary of the Invention
[0004] This application provides a training method, apparatus, electronic device, and storage medium for a low-rank adaptive model, which can improve the accuracy of model training.
[0005] To address the aforementioned problems, firstly, embodiments of this application provide a training method for a low-rank adaptation model, comprising:
[0006] The target subject of each initial image in the initial image set is fused with a preset background image to obtain the first extended image set;
[0007] Based on the initial image set and the first extended image set, the initial low-rank adaptation model is trained on the first pedestal model to obtain the first low-rank adaptation model.
[0008] Based on multiple text prompts, multiple extended images are generated using a text-based graph model composed of the first low-rank adaptation model and the first base model, and a second set of extended images is determined based on the generated multiple extended images.
[0009] Based on the initial image set, the first extended image set, and the second extended image set, the first low-rank adaptive model is iteratively trained and new extended images are generated based on the first pedestal model. The new extended images are then added to the second extended image set until the training termination condition is met, resulting in a trained low-rank adaptive model.
[0010] Secondly, embodiments of this application provide a training apparatus for a low-rank adaptation model, comprising:
[0011] The data expansion module is used to fuse the target subject of each initial image in the initial image set with a preset background image to obtain the first expanded image set;
[0012] An initial training module is used to train an initial low-rank adaptation model based on a first pedestal model, according to the initial image set and the first extended image set, to obtain a first low-rank adaptation model.
[0013] The image generation module is used to generate multiple extended images based on multiple text prompts, using a text-based image model composed of the first low-rank adaptation model and the first base model, and to determine a second set of extended images based on the generated multiple extended images.
[0014] The iterative training module is used to iteratively train the first low-rank adaptive model based on the first pedestal model and generate new extended images according to the initial image set, the first extended image set and the second extended image set, and add the new extended images to the second extended image set until the training termination condition is met, so as to obtain the trained low-rank adaptive model.
[0015] Thirdly, embodiments of this application also provide an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the training method of the low-rank adaptive model described in embodiments of this application.
[0016] Fourthly, embodiments of this application provide a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, represents the steps of the training method for the low-rank adaptive model disclosed in embodiments of this application.
[0017] The low-rank adaptation model training method, apparatus, electronic device, and storage medium provided in this application embodiment fuse the target subject of each initial image in the initial image set with a preset background image to obtain a first extended image set. Based on the initial image combination and the first extended image set, the initial low-rank adaptation model is trained on a first pedestal model to obtain a first low-rank adaptation model. Multiple extended images are generated based on multiple text prompts through a text-based image model composed of the first low-rank adaptation model and the first pedestal model. A second extended image set is determined based on the generated multiple extended images. Based on the initial image set, the first extended image set, and the second extended image set, the first low-rank adaptation model is iteratively trained on the first pedestal model and new extended images are generated. The new extended images are added to the second extended image set until the training termination condition is met, resulting in a trained low-rank adaptation model. By fusing the target subject of each initial image in the initial image set with a preset background image, the initial image set is expanded. The initial low-rank adaptation model is then trained based on the expanded image set. After training, the first low-rank adaptation model and the base model are used to generate a second expanded image set, further expanding the image set. Then, based on the initial image set, the first expanded image set, and the second expanded image set, the first low-rank adaptation model is iteratively trained and new expanded images are generated. During the iteration process, new expanded images containing the target subject can be continuously generated and used to train the low-rank adaptation model. This allows the low-rank adaptation model to be trained using a small amount of image data (the initial image set), which can improve the accuracy of model training. Moreover, the low-rank adaptation model trained in this way can improve the image generation effect for the target subject. Attached Figure Description
[0018] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0019] Figure 1 This is a flowchart of a training method for a low-rank adaptation model provided in an embodiment of this application;
[0020] Figure 2 This is an example image of an extended image from the first extended image set in the embodiments of this application;
[0021] Figure 3 This is an example image of an extended image from the second extended image set in the embodiments of this application;
[0022] Figure 4 This is a schematic diagram of the structure of a training device for a low-rank adaptation model provided in an embodiment of this application;
[0023] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation
[0024] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.
[0025] Figure 1 This is a flowchart of a training method for a low-rank adaptation model provided in an embodiment of this application, such as... Figure 1 As shown, the method includes steps 110 to 140.
[0026] Step 110: The target subject of each initial image in the initial image set is fused with a preset background image to obtain the first extended image set.
[0027] In one exemplary embodiment, the initial image set is a collection of original images containing the target subject. The state of the target subject in different initial images within the initial image set can be the same or different; however, all initial images in the initial image set must include multiple different subject states. The background in the initial image set can be simple and clean to facilitate the extraction of the target subject; for example, the background can be transparent. The target subject can be a target person, a target animated character, a target object, etc.
[0028] In one exemplary embodiment, the preset background image can be a relatively clean background image. That is, the preset background image needs to avoid clutter in the middle, avoid too many lines and details, and avoid highlighting the subject, so that it can stand out after merging with the target subject, avoiding any negative impact on the subject and preventing training collapse. For example, the preset background image can be a solid color background, the color of which is distinct from the color of the target subject. For instance, if the target subject is white, the solid color background could be black, red, orange, green, etc. For example, the preset background image can also be a clean background image such as the sky, aurora borealis, desert, forest, or grassland.
[0029] In one exemplary embodiment, the initial image set is expanded by extracting the target subject from each initial image in the initial image set, and then stitching the target subject in each initial image with a preset background image to obtain multiple expanded images with higher aesthetic appeal. These expanded images constitute a first expanded image set. For example, when stitching the target subject in each initial image with the preset background image, the target subject can be placed before the background image of the preset background image. For instance, based on the placement position of the target subject, the pixel value of the target subject is averaged with the pixel value of the overlapping position in the preset background image to obtain the pixel value of the corresponding pixel, thus achieving the stitching of the target subject with the preset background image.
[0030] In an optional embodiment, the step of fusing the target subject of each initial image in the initial image set with a preset background image to obtain a first expanded image set includes: scaling the target subject of the initial image and / or moving the position of the target subject to obtain an adjusted target subject, and fusing the adjusted target subject with the preset background image to obtain the first expanded image set.
[0031] For example, when stitching the target subject with a preset background image, the target subject can be randomly scaled and / or its position randomly moved. Then, the adjusted target subject is stitched with the preset background image. This avoids overfitting of the dataset. The initial image set typically contains a small number of images, such as 5 to 100. After expansion, a larger amount of image data can be obtained, for example, by a factor of 20. Figure 2 This is an example image of an extended image from the first extended image set in the embodiments of this application, such as... Figure 2 As shown, the extended images in the first extended image set are obtained by adding the target subject 1 to the preset background image 2.
[0032] Step 120: Based on the initial image set and the first extended image set, train the initial low-rank adaptation model on the first pedestal model to obtain the first low-rank adaptation model.
[0033] In one exemplary embodiment, the pedestal model is a text-generated graph model, which may be based on SDXL (StableDiffusion XL). SDXL is an open-source text-generated graph framework designed to address computational efficiency and memory consumption issues in large-scale text generation tasks. The first pedestal model is the pedestal model that participates in the training of the low-rank adaptation model.
[0034] In an exemplary embodiment, after obtaining the first expanded image set, text descriptions can be added to each expanded image in the first expanded image set. The expanded images and text descriptions in the first image set form image-text pairs. Based on the image-text pairs in the initial image set and the first image set, a low-rank adaptation (LoRA) technique is used to train an initial low-rank adaptation model based on a first pedestal model. During training, the network parameters of the first pedestal model are fixed, and the network parameters of the initial low-rank adaptation model are adjusted to obtain the first low-rank adaptation model. Low-rank adaptation is a method for fine-tuning large pre-trained models using low-rank matrices. Specifically, LoRA inserts low-rank matrices into the pre-trained model, thereby achieving model customization and improvement without significantly increasing computational complexity and the number of parameters. The low-rank adaptation method is particularly suitable for large-scale pre-trained models, enabling refined optimization for specific tasks or datasets while preserving the powerful performance of the base model. LoRA provides flexible adaptability while ensuring low computational overhead, making it particularly useful in resource-constrained environments.
[0035] In one embodiment of this application, the step of training an initial low-rank adaptation model based on a first pedestal model according to the initial image set and the first extended image set to obtain a first low-rank adaptation model includes:
[0036] Obtain the text description information corresponding to each extended image in the first extended image set;
[0037] The text description information corresponding to each initial image in the initial image set and the text description information corresponding to each extended image in the first extended image set are respectively used as input text, and each initial image in the initial image set has corresponding text description information;
[0038] An image corresponding to the input text is generated by the text-generated image model composed of the initial low-rank adaptation model and the first base model. Based on the image and the initial or extended image corresponding to the input text, the network parameters of the initial low-rank adaptation model are adjusted to obtain the first low-rank adaptation model.
[0039] In an exemplary embodiment, each extended image in the first extended image set is labeled with text to obtain corresponding text description information. The text description information may include the target subject, subject state, the target subject's location within the image, and the background color and composition. This labeling scheme allows the model to learn what object in the image the target subject is during training, better decoupling it from other elements in the image. In practice, more detailed labeling (text description) results in better overall model performance. An example of text description information is: "A smiley face icon of XX is in the center of the image, with a lake and snow-capped mountains in the background, and blue-green aurora in the sky; the image is very beautiful." For example, each initial image in the initial image set also corresponds to text description information. The text description information of the initial image provides the target subject, subject state, etc., while the text description information of the preset background image provides information such as the background color and composition. The text description information of the initial image and the preset background image can be concatenated to obtain the text description information of the extended image. Of course, in order to make the text description information of the extended images more accurate, the generated text description information can be manually proofread, or the text description information of each extended image can be manually marked. This application embodiment does not limit this.
[0040] The text descriptions corresponding to each initial image in the initial image set and each extended image in the first expanded image set are used as input text. These input texts are then fed into the text-generated graph model, which is composed of the initial low-rank adaptation model and the first pedestal model. Each text description is used as input text, and one text description is input to the text-generated graph model at a time. The model generates an image corresponding to the input text. Based on the generated image and the corresponding initial or expanded image, the network parameters of the initial low-rank adaptation model are adjusted. The network parameters of the first pedestal model are fixed. This training process is iteratively executed until training is complete, resulting in the first low-rank adaptation model. The training of the initial low-rank adaptation model can end when all images in both the initial and first expanded image sets have participated in the training.
[0041] By acquiring the text description information corresponding to each extended image, and using the text description information corresponding to the initial image and each extended image as input text, the text-generated image model composed of the initial low-rank adaptation model and the first base model generates an image corresponding to the input text. Then, based on the generated image and the initial or extended image corresponding to the input text, the network parameters of the initial low-rank adaptation model are adjusted. The first low-rank adaptation model obtained through such training can have a certain degree of subject reconstruction, subject generalization, and background generalization.
[0042] Step 130: Based on multiple text prompts, generate multiple extended images using the text-generated image model composed of the first low-rank adaptation model and the first base model, and determine a second set of extended images based on the generated multiple extended images.
[0043] In one exemplary embodiment, multiple text prompts can be pre-generated, with the subject of each text prompt being the target subject. Each text prompt is input into a text-based graph model composed of a first low-rank adaptation model and a first base model. The text-based graph model converts the text prompts into visual information, obtaining an extended image corresponding to each text prompt. Multiple text prompts can generate multiple extended images, which can form a second set of extended images.
[0044] Step 140: Based on the initial image set, the first extended image set, and the second extended image set, iteratively train the first low-rank adaptive model and generate new extended images based on the first pedestal model, and add the new extended images to the second extended image set until the training termination condition is met, thereby obtaining the trained low-rank adaptive model.
[0045] In an exemplary embodiment, an initial image set, a first extended image set, and a second extended image set are used as training data. Using this training data, a first low-rank adaptation model is trained based on a first pedestal model. After one training iteration, the newly trained low-rank adaptation model generates new extended images based on text prompts, and these new extended images are added to the second extended image set. The second extended image set is then used for the next training iteration. This process of determining training data, training the low-rank adaptation model, generating new extended images, and adding new extended images to the second extended image set is iteratively executed until the training termination condition is met, resulting in a trained low-rank adaptation model. During the training of the low-rank adaptation model, the generated extended images can be used as training data for other models (e.g., other pedestal models). The training termination condition can be the convergence of the network parameters of the low-rank adaptation model, or the sufficiently high target subject reconstruction and good generalization in the new extended images generated by the low-rank adaptation model. For example, the target subject reconstruction and generalization can be determined manually.
[0046] The training method for the low-rank adaptation model provided in this application embodiment involves fusing the target subject of each initial image in the initial image set with a preset background image to obtain a first extended image set. Based on the initial image combination and the first extended image set, the initial low-rank adaptation model is trained using a first pedestal model to obtain a first low-rank adaptation model. Multiple extended images are generated using a text-based image model composed of the first low-rank adaptation model and the first pedestal model based on multiple text prompts. A second extended image set is determined based on the generated multiple extended images. Based on the initial image set, the first extended image set, and the second extended image set, the first low-rank adaptation model is iteratively trained and new extended images are generated using the first pedestal model. The new extended images are added to the second extended image set until the training termination condition is met, resulting in a trained low-rank adaptation model. By fusing the target subject of each initial image in the initial image set with a preset background image, the initial image set is expanded. The initial low-rank adaptation model is then trained based on the expanded image set. After training, the first low-rank adaptation model and the first base model are used to generate a second expanded image set, further expanding the image set. Then, based on the initial image set, the first expanded image set, and the second expanded image set, the first low-rank adaptation model is iteratively trained and new expanded images are generated. During the iteration process, new expanded images containing the target subject can be continuously generated and used to train the low-rank adaptation model. This allows the low-rank adaptation model to be trained using a small amount of image data (the initial image set), which can improve the accuracy of model training. Moreover, the low-rank adaptation model trained in this way can improve the image generation effect for the target subject.
[0047] Based on the above technical solution, before generating multiple extended images using the text-based image model composed of the first low-rank adaptation model and the first base model according to multiple text prompts, the method further includes: generating multiple text prompts based on multiple background words, multiple states of the target subject, and multiple alternative colors.
[0048] In one exemplary embodiment, background words are words used to describe the background. The various states of the target subject can include, for example, happiness, sadness, looking left, looking right, expert mode, calmness, etc. Alternate colors are alternative colors for the target subject, meaning the target subject can be different colors in different images.
[0049] In one exemplary embodiment, a background word library can be prepared in advance. The background word library includes multiple background words, such as 120 commonly used background words. Background words can be, for example, sky, desert, lawn, flowers and grass, etc.
[0050] In one exemplary embodiment, different background words, the state of the target subject, and alternative colors are combined to generate multiple text prompts. Subsequently, a text-based graph model composed of a first low-rank adaptation model and a first pedestal model can be used to generate extended images corresponding to the text prompts, resulting in a large number of extended images.
[0051] Multiple text prompts are generated based on various background words, multiple states of the target subject, and multiple alternative colors, so that corresponding extended images can be generated based on the text prompts, thereby expanding the image data.
[0052] Based on the above technical solution, the step of iteratively training the first low-rank adaptation model and generating new extended images based on the first pedestal model according to the initial image set, the first extended image set, and the second extended image set, and adding the new extended images to the second extended image set, until the training termination condition is met to obtain the trained low-rank adaptation model, may include:
[0053] The initial image set, a portion of the expanded images in the first expanded image set, and the second expanded image set are used as the training data set; wherein, for each iteration of training of the first low-rank adaptive model, the determined training data set is reduced by a preset number or a preset proportion of the expanded images in the first expanded image set compared to the training data set of the previous iteration.
[0054] Based on the training data set, the first low-rank adaptive model is trained on the first pedestal model, and a new extended image is generated by the text-based image model composed of the trained first low-rank adaptive model and the first pedestal model, and the new extended image is added to the second extended image set.
[0055] The process iteratively executes the operations of determining the training data set, training, generating new extended images, and adding the new extended images to the second extended image set until the training termination condition is met, resulting in a trained low-rank adaptive model.
[0056] In an exemplary embodiment, during each training iteration, a training dataset is first determined. While ensuring the main image collapse rate does not increase, the initial image set can be completely added to the training dataset. This allows the model to learn the specific form of the target subject, achieving the insertion of the original concept and ensuring the model's reproducibility (reproducibility of the target subject). A portion of the expanded images from the first expanded image set are then added to the training dataset. The training dataset determined in this iteration has a predetermined number or proportion fewer expanded images from the first expanded image set compared to the training dataset from the previous iteration. This gradually reduces the number of expanded images from the first expanded image set in the training dataset during iteration. All expanded images from the second expanded image set can also be added to the training dataset. Since the amount of data in the second expanded image set increases with each training iteration, the number of expanded images in the second expanded image set can be gradually increased during iteration. This gradually improves the generalization of the training model and enhances its aesthetics, as the data combined with the first expanded image set has lower aesthetics, while the data combined with the second expanded image set has higher aesthetics and a higher degree of generalization. The image collapse rate can be manually judged based on the images generated by the model after one iteration of training.
[0057] In one exemplary embodiment, since each extended image in the generated second extended image set may not completely correspond to the text prompt information, text marking can be performed on each extended image in the second extended image set. The marking method can be manual fine marking to obtain the text description information corresponding to each extended image.
[0058] In an exemplary embodiment, after determining the training dataset, the text description information corresponding to each image in the training dataset can be used as input text. This input text is then input into a text-based image model composed of a first low-rank adaptation model and a base model. The text-based image model generates an image corresponding to the input text. Based on the generated image and the image corresponding to the input text, the network parameters of the first low-rank adaptation model are adjusted, while the network parameters of the first base model are fixed. After training, a trained first low-rank adaptation model is obtained. The text-based image model composed of the trained first low-rank adaptation model and the first base model generates a new extended image based on the text prompt information. The new extended image is added to a second extended image set. Then, the above operations of determining the training dataset, training, generating new extended images, and adding new extended images to the second extended image set are iteratively executed until the training termination condition is met, resulting in a trained low-rank adaptation model.
[0059] By gradually reducing the number of expanded images in the first expanded image set and gradually increasing the number of expanded images in the second expanded image set during the iterative training of the first low-rank adaptive model, the generalization of the trained model can be gradually improved and the aesthetics can be enhanced. Moreover, by using all the initial image sets at the same time, the model's ability to reproduce the target subject can be improved.
[0060] Based on the above technical solution, the method may further include: if, during the iterative execution of determining the training data set, training, generating a new extended image, and adding the new extended image to the second extended image set, the target subject in the generated new extended image does not meet the target condition, then the new extended image is deleted, and the extended image used in this training is deleted from the second extended image set.
[0061] After each iteration of training the first low-rank adaptation model, a new extended image is generated based on the text prompt information using the trained low-rank adaptation model. The quality of the generated new extended images is verified to determine whether the target subject in the new extended image meets the target conditions. If the target subject in the new extended image does not meet the target conditions, the new extended image is deleted, meaning it is not added to the second extended image set. Simultaneously, the extended images used in this training are deleted from the second extended image set. This process removes data that degrades training performance, thereby improving the model's training effect. For example, whether the target subject in the new extended image meets the target conditions can be determined by whether the collapse rate of the target subject in the generated batch of new extended images exceeds a collapse rate threshold. If the collapse rate exceeds the threshold, it is determined that this batch of generated new extended images does not meet the target conditions. Alternatively, whether the target subject in the new extended image meets the target conditions can also be determined based on the target subject's fidelity and generalization ability; this judgment can be made manually.
[0062] Based on the above technical solution, the step of determining the second set of extended images according to the generated multiple extended images may include: selecting extended images that meet the target conditions from the generated multiple extended images to form the second set of extended images.
[0063] For example, the requirements for the second extended image set may include: the aesthetics must be high enough, the main subject in the image must be accurately reproduced, and there must be no artifacts or distortions in the image. Minor artifacts that do not affect the overall image quality can be manually repaired. Figure 3 This is an example image of an extended image from the second extended image set in the embodiments of this application, such as... Figure 3 As shown, relative to Figure 2The extended images in the first extended image set shown have shown some generalization in the image of the target subject 1, and the background 2 also has a certain degree of generalization. They differ somewhat from the original textures in the first extended image set, and the image blending is more refined (thanks to the text labeling of each extended image in the first extended image set). After training, when the data volume of the second extended image set is 1 / 4 of that of the first extended image set, the performance of the generated model has been greatly improved.
[0064] In an exemplary embodiment, after generating multiple extended images using a text-based image model composed of a first low-rank adaptation model and a first pedestal model, some of these extended images may exhibit degradation issues, and some may have low subject reproduction accuracy, subject generalization, and background generalization. Therefore, it is necessary to filter the generated extended images to select those that meet the target conditions. These target-condition extended images are then grouped into a second extended image set, which is used as training data for the low-rank adaptation model. This improves the model's training performance and enhances its subject reproduction accuracy, subject generalization, and background generalization. For example, when filtering extended images that meet the target conditions from the generated multiple extended images, manual filtering can be performed to select extended images with high aesthetic appeal, sufficient reproduction accuracy, and sufficient generalization.
[0065] Based on the above technical solution, the method may further include: training the second pedestal model according to the initial image set, the first extended image set, and the second extended image set to obtain the trained second pedestal model.
[0066] In an exemplary embodiment, the second base model can be the first base model described above, or it can be another base model other than the first base model.
[0067] During the training of the low-rank adaptive model, training data with high aesthetic appeal and fidelity was generated, including the first extended image set and the second extended image set. These data can be used as training data for the second pedestal model, thus eliminating the need for manual preparation of data for pedestal model training and saving data preparation time and costs.
[0068] The embodiments of this application can produce a low-rank adaptive model that generates stable subjects and training data with high aesthetics and fidelity. The generated training data can also be used to fine-tune the model base and train, thus realizing the generation of stable subjects with high aesthetics and fidelity even with only a small batch of data.
[0069] Figure 4This is a schematic diagram of the structure of a training device for a low-rank adaptation model provided in an embodiment of this application, as shown below. Figure 4 As shown, the device includes:
[0070] Data expansion module 410 is used to fuse the target subject of each initial image in the initial image set with a preset background image to obtain a first expanded image set;
[0071] The initial training module 420 is used to train the initial low-rank adaptation model based on the first pedestal model according to the initial image set and the first extended image set, so as to obtain the first low-rank adaptation model.
[0072] The image generation module 430 is used to generate multiple extended images based on multiple text prompts, using a text-based image model composed of the first low-rank adaptation model and the first base model, and to determine a second set of extended images based on the generated multiple extended images.
[0073] The iterative training module 440 is used to iteratively train the first low-rank adaptive model based on the first pedestal model and generate new extended images according to the initial image set, the first extended image set and the second extended image set, and add the new extended images to the second extended image set until the training termination condition is met, so as to obtain the trained low-rank adaptive model.
[0074] Optionally, the initial training module includes:
[0075] The text description acquisition unit is used to acquire text description information corresponding to each extended image in the first extended image set;
[0076] The input text determination unit is used to take the text description information corresponding to each initial image in the initial image set and the text description information corresponding to each extended image in the first extended image set as input text, respectively. Each initial image in the initial image set has corresponding text description information.
[0077] An initial training unit is used to generate an image corresponding to the input text using a text-based image model composed of the initial low-rank adaptation model and the first base model, and to adjust the network parameters of the initial low-rank adaptation model according to the image and the initial or extended image corresponding to the input text to obtain the first low-rank adaptation model.
[0078] Optionally, the device further includes:
[0079] The text prompt generation module is used to generate multiple text prompt messages based on multiple background words, multiple states of the target subject, and multiple alternative colors.
[0080] Optionally, the iterative training module is specifically used for:
[0081] The initial image set, a portion of the expanded images in the first expanded image set, and the second expanded image set are used as the training data set; wherein, for each iteration of training of the first low-rank adaptive model, the determined training data set is reduced by a preset number or a preset proportion of the expanded images in the first expanded image set compared to the training data set of the previous iteration.
[0082] Based on the training data set, the first low-rank adaptive model is trained on the first pedestal model, and a new extended image is generated by the text-based image model composed of the trained first low-rank adaptive model and the first pedestal model, and the new extended image is added to the second extended image set.
[0083] The process iteratively executes the operations of determining the training data set, training, generating new extended images, and adding the new extended images to the second extended image set until the training termination condition is met, resulting in a trained low-rank adaptive model.
[0084] Optionally, the device further includes:
[0085] The data deletion module is used to delete the new extended image if, during the iterative execution of determining the training data set, training, generating a new extended image, and adding the new extended image to the second extended image set, the target subject in the generated new extended image does not meet the target condition, and to delete the extended image used in this training from the second extended image set.
[0086] Optionally, the image generation module includes:
[0087] An image filtering unit is used to filter extended images that meet the target conditions from the generated multiple extended images to form the second extended image set.
[0088] Optionally, the data expansion module is specifically used for:
[0089] The target subject of the initial image is scaled and / or its position is moved to obtain an adjusted target subject. The adjusted target subject is then merged with the preset background image to obtain the first extended image set.
[0090] Optionally, the device further includes:
[0091] The second pedestal model is trained based on the initial image set, the first extended image set, and the second extended image set to obtain the trained second pedestal model.
[0092] The training apparatus for the low-rank adaptive model provided in this application embodiment is used to implement the steps of the training method for the low-rank adaptive model described in this application embodiment. The specific implementation of each module of the apparatus is described in the corresponding steps, and will not be repeated here.
[0093] The training apparatus for the low-rank adaptation model provided in this application embodiment fuses the target subject of each initial image in the initial image set with a preset background image to obtain a first extended image set. Based on the initial image combination and the first extended image set, the initial low-rank adaptation model is trained on a first pedestal model to obtain a first low-rank adaptation model. Based on multiple text prompts, multiple extended images are generated through a text-based image model composed of the first low-rank adaptation model and the first pedestal model. A second extended image set is determined based on the generated multiple extended images. Based on the initial image set, the first extended image set, and the second extended image set, the first low-rank adaptation model is iteratively trained on the first pedestal model and new extended images are generated. The new extended images are added to the second extended image set until the training termination condition is met, and a trained low-rank adaptation model is obtained. By fusing the target subject of each initial image in the initial image set with a preset background image, the initial image set is expanded. The initial low-rank adaptation model is then trained based on the expanded image set. After training, the first low-rank adaptation model and the first base model are used to generate a second expanded image set, further expanding the image set. Then, based on the initial image set, the first expanded image set, and the second expanded image set, the first low-rank adaptation model is iteratively trained and new expanded images are generated. During the iteration process, new expanded images containing the target subject can be continuously generated and used to train the low-rank adaptation model. This allows the low-rank adaptation model to be trained using a small amount of image data (the initial image set), which can improve the accuracy of model training. Moreover, the low-rank adaptation model trained in this way can improve the image generation effect for the target subject.
[0094] Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application, such as... Figure 5 As shown, the electronic device 500 may include one or more processors 510 and one or more memories 520 connected to the processors 510. The electronic device 500 may also include an input interface 530 and an output interface 540 for communicating with another device or system. Program code executed by the processor 510 may be stored in the memory 520.
[0095] The processor 510 in the electronic device 500 calls the program code stored in the memory 520 to execute the training method of the low-rank adaptation model in the above embodiment.
[0096] This application also provides a computer-readable storage medium storing a computer program thereon, which, when executed by a processor, implements the steps of the training method for a low-rank adaptive model as described in this application.
[0097] This application also provides a computer program product that, when executed by a processor, implements the steps of the training method for the low-rank adaptive model as described in this application.
[0098] The various embodiments in this specification are described in a progressive manner, with each embodiment focusing on its differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the apparatus embodiments, since they are fundamentally similar to the method embodiments, the description is relatively simple; relevant parts can be referred to the descriptions in the method embodiments.
[0099] The above provides a detailed description of a training method, apparatus, electronic device, and storage medium for a low-rank adaptive model provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of this application. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of this application. Therefore, the content of this specification should not be construed as a limitation of this application.
[0100] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented by means of software plus necessary general-purpose hardware platforms, and of course, it can also be implemented by hardware. Based on this understanding, the above technical solutions, in essence or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods described in the various embodiments or some parts of the embodiments.
Claims
1. A training method for a low-rank adaptive model, characterized in that, include: The target subject of each initial image in the initial image set is fused with a preset background image to obtain the first extended image set; Based on the initial image set and the first extended image set, the initial low-rank adaptation model is trained on the first pedestal model to obtain the first low-rank adaptation model. Based on multiple text prompts, multiple extended images are generated using a text-based graph model composed of the first low-rank adaptation model and the first base model, and a second set of extended images is determined based on the generated multiple extended images. Based on the initial image set, the first extended image set, and the second extended image set, the first low-rank adaptive model is iteratively trained and new extended images are generated based on the first pedestal model. The new extended images are then added to the second extended image set until the training termination condition is met, resulting in a trained low-rank adaptive model.
2. The method according to claim 1, characterized in that, The step of training an initial low-rank adaptation model based on a first pedestal model using the initial image set and the first expanded image set to obtain a first low-rank adaptation model includes: Obtain the text description information corresponding to each extended image in the first extended image set; The text description information corresponding to each initial image in the initial image set and the text description information corresponding to each extended image in the first extended image set are respectively used as input text, and each initial image in the initial image set has corresponding text description information; An image corresponding to the input text is generated by the text-generated image model composed of the initial low-rank adaptation model and the first base model. Based on the image and the initial or extended image corresponding to the input text, the network parameters of the initial low-rank adaptation model are adjusted to obtain the first low-rank adaptation model.
3. The method according to claim 1, characterized in that, Before generating multiple extended images based on multiple text prompts using the text-based image model composed of the first low-rank adaptation model and the first pedestal model, the method further includes: Multiple text prompts are generated based on multiple background words, multiple states of the target subject, and multiple alternative colors.
4. The method according to claim 1, characterized in that, The step of iteratively training the first low-rank adaptation model based on the first pedestal model and generating new extended images according to the initial image set, the first extended image set, and the second extended image set, and adding the new extended images to the second extended image set, until the training termination condition is met, to obtain the trained low-rank adaptation model, includes: The initial image set, a portion of the expanded images in the first expanded image set, and the second expanded image set are used as the training data set; wherein, for each iteration of training of the first low-rank adaptive model, the determined training data set is reduced by a preset number or a preset proportion of the expanded images in the first expanded image set compared to the training data set of the previous iteration. Based on the training data set, the first low-rank adaptive model is trained on the first pedestal model, and a new extended image is generated by the text-based image model composed of the trained first low-rank adaptive model and the first pedestal model, and the new extended image is added to the second extended image set. The process iteratively executes the operations of determining the training data set, training, generating new extended images, and adding the new extended images to the second extended image set until the training termination condition is met, resulting in a trained low-rank adaptive model.
5. The method according to claim 4, characterized in that, Also includes: If, during the iterative execution of determining the training data set, training, generating new extended images, and adding the new extended images to the second extended image set, the target subject in the generated new extended image does not meet the target conditions, then the new extended image is deleted, and the extended image used in this training is also deleted from the second extended image set.
6. The method according to any one of claims 1-5, characterized in that, The step of determining the second set of extended images based on the generated multiple extended images includes: From the generated multiple extended images, extended images that meet the target conditions are selected to form the second extended image set.
7. The method according to any one of claims 1-5, characterized in that, The step of fusing the target subject of each initial image in the initial image set with a preset background image to obtain a first expanded image set includes: The target subject of the initial image is scaled and / or its position is moved to obtain an adjusted target subject. The adjusted target subject is then merged with the preset background image to obtain the first extended image set.
8. The method according to any one of claims 1-5, characterized in that, Also includes: The second pedestal model is trained based on the initial image set, the first extended image set, and the second extended image set to obtain the trained second pedestal model.
9. A training device for a low-rank adaptive model, characterized in that, include: The data expansion module is used to fuse the target subject of each initial image in the initial image set with a preset background image to obtain the first expanded image set; An initial training module is used to train an initial low-rank adaptation model based on a first pedestal model, according to the initial image set and the first extended image set, to obtain a first low-rank adaptation model. The image generation module is used to generate multiple extended images based on multiple text prompts, using a text-based image model composed of the first low-rank adaptation model and the first base model, and to determine a second set of extended images based on the generated multiple extended images. The iterative training module is used to iteratively train the first low-rank adaptive model based on the first pedestal model and generate new extended images according to the initial image set, the first extended image set and the second extended image set, and add the new extended images to the second extended image set until the training termination condition is met, so as to obtain the trained low-rank adaptive model.
10. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the training method for the low-rank adaptive model according to any one of claims 1 to 8.
11. A computer-readable storage medium having a computer program stored thereon, characterized in that, When executed by a processor, the program implements the steps of the training method for the low-rank adaptive model as described in any one of claims 1 to 8.