Method for generating lora with separated main body and style
By dividing the text annotation of the data set of the LoRA model into content subject and style, and introducing column weight vectors p and q into the calculation formula, the problem of interfering with each other in multi-task learning is solved, and the effect of flexible generation of images and videos is achieved.
Patent Information
- Application Number
- CN202510410711.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-02
- Publication Date
- 2025-08-15
AI Technical Summary
In the prior art, LoRA models are prone to interfering with each other in multi-task learning, resulting in poor generation results and the inability to flexibly generate images or video styles and contents of multiple tasks.
The LoRA generation method with the separation of subject and style is adopted, the data set text annotation is divided into content subject and style, and two independent LoRA tasks are trained separately. The calculation formula h=W0x+pΔWsxs+qΔWcxc is adjusted to avoid feature interference, and the column weight vectors p and q are introduced to ensure compatibility.
It realizes that the LoRA model can flexibly select subjects and styles when generating images and videos, improves generation efficiency and economy, and can generate customized image styles or content, avoiding interfering with each other with features.
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of large models, and in particular to a LoRa generation method with separated subject and style. Background Art
[0002] In the field of large models, LoRA is a technology for efficiently fine-tuning models; its core idea is to reduce the number of parameters and computational cost during model fine-tuning by introducing low-rank decomposition while maintaining the performance of the model.
[0003] The number of parameters in large models is usually very large (for example, GPT-3 has 175 billion parameters). Directly fine-tuning the entire model requires a lot of computing resources and storage space, which is extremely costly. LoRA introduces low-rank decomposition in the model's weight matrix, converting full parameter fine-tuning into fine-tuning of the low-rank matrix. This method significantly reduces the number of parameters that need to be updated, thereby reducing computing and storage costs.
[0004] Directly fine-tuning a large model may cause the model to lose the versatility of the pre-training stage and overfit to a specific task; LoRA only fine-tunes the low-rank matrix, while the weights of the original model remain unchanged; this allows the model to retain the general knowledge of the pre-training stage after fine-tuning while being able to adapt to the needs of specific tasks.
[0005] In multi-task learning, LoRA introduces an independent low-rank matrix for each task, allowing the model to share pre-training knowledge between different tasks while retaining specific fine-tuning parameters for each task; this approach enables the model to efficiently support multi-task learning.
[0006] However, in the field of raw images and raw videos, Lora is used for a single task, that is, a large model is fine-tuned once to train Lora for a single task; for example, to train character Lora, corresponding character photos are collected and Lora is trained. After the training of this Lora is completed, it can only be used for the generation task of this character; when the Lora model of another person is needed, the character Lora needs to be retrained; similarly, style Lora, light and shadow Lora, etc., after a Lora is trained, it can only be used for a single task; when a large model calls multiple Lora at the same time, the effects of multiple Lora are easily interfered with each other after superposition, resulting in poor results. In view of this, the present invention is proposed. Summary of the Invention
[0007] The purpose of the present invention is to solve the problems existing in the prior art and to propose a lora generation method with separated subject and style.
[0008] In order to achieve the above object, the present invention adopts the following technical solutions:
[0009] A lora generation method with separated subject and style is performed by the following steps:
[0010] Step 1: Separate the dataset text annotations and divide the dataset text annotations of the Lora model into content and style, and then train two Lora models for different tasks together;
[0011] Step 2: Calculate the final output of the Lora model using the following formula:
[0012] h=W0x+ΔW s x s +ΔW c x c
[0013] Where h is the final output of the Lora model, ΔW s is the style matrix, x s is the style vector; ΔW c is the content matrix, x c is the content vector.
[0014] Preferably, in step 1, the content body of the Lora model includes the specific content of the data; the style of the Lora model includes the style features of the data.
[0015] Furthermore, the style vector and content vector are obtained by decomposing the original text annotations. It is necessary to train two independent tasks in one model to eliminate the contamination of texts and avoid cross-contamination during training.
[0016] Furthermore, to eliminate the interference caused by feature addition and ensure compatibility, the calculation formula in step 2 is adjusted to:
[0017] h=W0x+pΔW s x s +qΔW c x c
[0018] Where h is the final output of the Lora model, ΔW s is the style matrix, x s is the style vector; ΔW c is the content matrix, x c is the content vector, p is the column weight vector of style features, and q is the column weight vector of content features.
[0019] Furthermore, the formula h=W0x+ΔW s x s +ΔW c x c The formula h=W0x+pΔW s x s +qΔWc x c middle
[0020] W0 is the original large model matrix with a dimension of m*n, and x is the original network feature; the gradient of W0 is locked during training.
[0021] Preferably, in step 1, when training Lora, the first part of the content body is converted into a content vector x by the pre-trained Clip model. c and content matrix ΔW c The operation is performed, and then the feature matrix is multiplied by the corresponding weight coefficient q of each column; the second part of the style information is converted into a style vector x through the pre-trained Clip model s and style matrix ΔW s After the operation is performed, the feature matrix is multiplied by the corresponding weight coefficient p of each column; then the two parts of the feature are summed with the original large model.
[0022] Preferably, the formula:
[0023] h=W0x+pΔW s x s +qΔW c x c
[0024] In the equation, p and q correspond to ΔW respectively. s , ΔW c The weight coefficient vector of the column vector of , which means the importance of each column of the matrix. If some columns are more important to the final result, the corresponding numbers will be large, and vice versa. The sum of the internal numbers of p and q is 1.
[0025] Furthermore, the closer the cosine similarity value of vectors p and q is to 0, the more orthogonal the two vectors are.
[0026] Preferably, in step one, two Lora machines of different tasks respectively learn the content body and corresponding style of a data.
[0027] Furthermore, when the input text only has content but no style, only one content Lora is run, which can directly extract the main content of the image; when the input text only has style but no content, the style of the generated image remains highly similar to the style of the original image; when the input text has both style and content, the model runs two independent Lora simultaneously, and the generated image remains highly similar to the original image, with both content and style being similar.
[0028] Compared with the existing technology, the present invention provides a lora generation method with separated subject and style, which has the following beneficial effects:
[0029] 1. The subject and style separation lora generation method of the present invention can not only decouple the style and subject of pictures, but also can be used for videos. The lora of the video is trained, and the subject and style in the video are learned and generated separately. By separating the content and style of the data, two independent loras are trained together, making the model more flexible and the generated subject and style can be freely selected. It is more efficient and economical to generate customized image styles or content for specific users.
[0030] 2. The lora generation method with separated subject and style has updated the calculation formula to eliminate the interference of feature addition and ensure compatibility. Different from the original lora, the improved lora of the present invention needs to separate the text annotation of the data during training. The first part of the content body is converted into the content vector x through the pre-trained Clip model. c , and the content matrix ΔW c The operation is performed, and then the feature matrix is multiplied by the corresponding weight coefficient q of each column; the second part of the style information is converted into a style vector x through the pre-trained Clip model s , and the style matrix ΔW s Perform the operation, then multiply the feature matrix by the corresponding weight coefficient p of each column; then the two parts of the feature are summed with the original large model, and p and q correspond to ΔW respectively s , ΔW c The weight coefficient vector of the column vector of , which means the importance of each column of the matrix. If some columns are more important to the final result, the corresponding numbers will be slightly larger, otherwise the corresponding numbers will be smaller; and the sum of the internal numbers of p and q is 1. DETAILED DESCRIPTION
[0031] In the description of the present invention, it should be understood that the terms "up", "down", "front", "back", "left", "right", "top", "bottom", "inside", "outside", etc., indicating directions or positional relationships, are only for the convenience of describing the present invention and simplifying the description, and therefore cannot be understood as limiting the present invention.
[0032] Example 1:
[0033] A lora generation method with separated subject and style is performed by the following steps:
[0034] Step 1: Separate the dataset text annotations and divide the dataset text annotations of the Lora model into content and style, and then train two Lora models for different tasks together;
[0035] Step 2: Calculate the final output of the Lora model using the following formula:
[0036] h=W0x+ΔW s x s+ΔW c x c
[0037] Where h is the final output of the Lora model, ΔW s is the style matrix, x s is the style vector; ΔW c is the content matrix, x c is the content vector.
[0038] In step 1, the content body of the Lora model contains the specific content of the data; the style of the Lora model contains the style characteristics of the data.
[0039] The style vector and content vector are obtained by decomposing the original text annotations. It is necessary to train two independent tasks in one model to eliminate the contamination of texts and avoid cross-contamination during training.
[0040] To eliminate the interference caused by feature addition and ensure compatibility, the calculation formula in step 2 is adjusted to:
[0041] h=W0x+pΔW s x s +qΔW c x c
[0042] Where h is the final output of the Lora model, ΔW s is the style matrix, x s is the style vector; ΔW c is the content matrix, x c is the content vector, p is the column weight vector of style features, and q is the column weight vector of content features.
[0043] In the field of raw images and videos, the present invention decomposes the content body and style of the original input, and trains two independent LoRAs at the same time. After training once, Lora can be used for multiple tasks. It should be noted that the original Lora calculation method is h=W0x+BAx, where h is the final output of the Lora model, W0 is the original large model matrix with a dimension of m*n, B and A are additional matrices with dimensions of m*r and r*n respectively, and x is the original network feature. The gradient of W0 is locked during training, and the two matrices AB need to be trained and saved as the Lora model.
[0044] Specifically, the present invention divides the text annotation of the Lora dataset into two parts: content body and style; the content body mainly includes the specific content of the data, such as "kitten", "ship", etc.; the style mainly includes the style characteristics of the data, such as "cartoon style", "two-dimensional style", "national style" and other style information.
[0045] This invention improves the original Lora and initially updates it to:
[0046] h=W0x+ΔW s x s +ΔW c x c , specifically, ΔW s is the style matrix, x s is the style vector; ΔW c is the content matrix, x c is the content vector; the style vector and content vector are obtained by decomposing the original text annotation. Because two independent tasks need to be trained in one model, the texts are excluded from contamination and cross contamination during training is avoided. In order to eliminate the interference of feature addition and ensure compatibility, the final calculation formula of the present invention is: h = W0x + pΔW s x s +qΔW c x c , where p is the column weight vector of style features and q is the column weight vector of content features.
[0047] There are four parts that need to be trained and optimized in this invention, namely: style matrix ΔW s , content matrix ΔW c , the column weight vector p of the style feature and the column weight vector q of the content feature; for other variables, they can be obtained through pre-processing or pre-training models; and p and q can only select a part of the parameters for training according to their importance. The amount of parameters and the amount of calculation introduced by the present invention are very small compared to the large model itself; when training a large model at a time, by decomposing the original data, two independent LoRAs are trained at the same time, and the main body and corresponding style of a data are learned respectively; because the two Loras are independent, targeted input can be made during model inference; when the input text has only the main content but no style, only one content Lora is running at this time, which can directly extract the main content of the picture, such as "kitten", "ship" and other content, and the generated picture is consistent with the training The "kitten" and "ship" in the image have similar appearances and contours. When the input text only has a style but no main content, the generated image style remains highly similar to the original image's style, such as a "cartoon style." When the input text has both style and main content, the model simultaneously runs two independent LoRa models, and the generated image remains highly similar to the original image, with similar content and style. This invention not only decouples style and subject for images, but also operates similarly for videos, training LoRa for videos to learn and generate the subject and style separately. By separating the content and style of the data and training two independent LoRa models together, the model becomes more flexible, allowing the generated subject and style to be freely selected. Customized image styles or content can be generated for specific users, making it more efficient and economical.
[0048] Example 2:
[0049] A lora generation method with separated subject and style is performed by the following steps:
[0050] Step 1: Separate the dataset text annotations and divide the dataset text annotations of the Lora model into content and style, and then train two Lora models for different tasks together;
[0051] Step 2: Calculate the final output of the Lora model using the following formula:
[0052] h=W0x+ΔW s x s +ΔW c x c
[0053] Where h is the final output of the Lora model, ΔW s is the style matrix, x s is the style vector; ΔW c is the content matrix, x c is the content vector.
[0054] In step 1, the content body of the Lora model contains the specific content of the data; the style of the Lora model contains the style characteristics of the data.
[0055] The style vector and content vector are obtained by decomposing the original text annotations. It is necessary to train two independent tasks in one model to eliminate the contamination of texts and avoid cross-contamination during training.
[0056] To eliminate the interference caused by feature addition and ensure compatibility, the calculation formula in step 2 is adjusted to:
[0057] h=W0x+pΔW s x s +qΔW c x c
[0058] Where h is the final output of the Lora model, ΔW s is the style matrix, x s is the style vector; ΔW c is the content matrix, x c is the content vector, p is the column weight vector of style features, and q is the column weight vector of content features.
[0059] Formula h=W0x+ΔW s x s +ΔW c x cThe formula h=W0x+pΔW s x s +qΔW c x c middle
[0060] W0 is the original large model matrix with a dimension of m*n, and x is the original network feature; the gradient of W0 is locked during training.
[0061] In step 1, when training Lora, the first part of the content body is converted into a content vector x through the pre-trained Clip model c and content matrix ΔW c The operation is performed, and then the feature matrix is multiplied by the corresponding weight coefficient q of each column; the second part of the style information is converted into a style vector x through the pre-trained Clip model s and style matrix ΔW s After the operation is performed, the feature matrix is multiplied by the corresponding weight coefficient p of each column; then the two parts of the feature are summed with the original large model.
[0062] Mode:
[0063] h=W0x+pΔW s x s +qΔW c x c
[0064] In the equation, p and q correspond to ΔW respectively. s , ΔW c The weight coefficient vector of the column vector of , which means the importance of each column of the matrix. If some columns are more important to the final result, the corresponding numbers will be large, and vice versa. The sum of the internal numbers of p and q is 1.
[0065] The closer the cosine similarity value of vectors p and q is to 0, the more orthogonal the two vectors are.
[0066] In step 1, two Lora with different tasks learn the content and corresponding style of a data respectively.
[0067] When the input text only has content but no style, only one content Lora is run, which can directly extract the main content of the image; when the input text only has style but no content, the style of the generated image remains highly similar to the style of the original image; when the input text has both style and content, the model runs two independent Lora simultaneously, and the generated image remains highly similar to the original image, with both content and style being similar.
[0068] In this embodiment, in order to eliminate the interference of feature addition and ensure compatibility, the final calculation formula of the present invention is obtained as follows: h = W0x + pΔW s xs +qΔW c x c , where p is the column weight vector of style features and q is the column weight vector of content features; Different from the original Lora, the improved Lora of the present invention needs to separate the text annotations of the data during training. The first part of the content body is converted into the content vector x through the pre-trained Clip model. c , and the content matrix ΔW c The operation is performed, and then the feature matrix is multiplied by the corresponding weight coefficient q of each column; the second part of the style information is converted into a style vector x through the pre-trained Clip model s , and the style matrix ΔW s After the operation is performed, the feature matrix is multiplied by the corresponding weight coefficient p of each column; then the two parts of the feature are summed with the original large model.
[0069] Specifically, p and q correspond to ΔW respectively. s , ΔW c The weight coefficient vector of the column vector of the matrix means the importance of each column of the matrix. Some columns are more important to the final result, so the corresponding numbers are large, and vice versa. The sum of the internal numbers of p and q is 1; in order to avoid overfitting, the dimensions of p and q in the present invention can be flexibly changed, such as only selecting ΔW s , ΔW c The most important 50% of the column vectors are optimized and trained, and the less important columns can be discarded directly, thereby further reducing the number of parameters and computation.
[0070] Since the present invention trains two LoRas for different tasks in one model, and the features are directly summed, in order to avoid mutual interference between features, an orthogonality loss of the weight coefficient vector p and the weight coefficient vector q is added to the loss function; specifically, the closer the cosine similarity value of vectors p and q is to 0, the more orthogonal the two vectors are. This will ensure that the matrix summation additionally introduced by the present invention does not interfere with each other, which is beneficial to the robustness of the model.
[0071] Specifically, there are four things that need to be trained and optimized in this invention: the style matrix ΔW s , content matrix ΔW c , the column weight vector p of the style features, and the column weight vector q of the content features; other variables can be obtained through pre-processing or pre-training models. Furthermore, only a subset of p and q parameters can be selected for training based on their importance. Compared to large models, this method introduces very few parameters and computational complexity.
[0072] The present invention separates the content and style of the data, thereby training two independent LoRa models together, making the model more flexible and allowing the generation subject and style to be freely selected; generating customized image styles or content for specific users is more efficient and economical.
[0073] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.
Claims
1. A lora generation method with separated subject and style, characterized in that: Use the following steps: Step 1: Separate the dataset text annotations and divide the dataset text annotations of the Lora model into content and style, and then train two Lora models for different tasks together; Step 2: Calculate the final output of the Lora model using the following formula: h=W0x+ΔW s x s +ΔW c x c Where h is the final output of the Lora model, ΔW s is the style matrix, x s is the style vector; ΔW c is the content matrix, x c is the content vector.
2. A lora generation method with separated subject and style according to claim 1, characterized in that In step 1, the content body of the Lora model contains the specific content of the data; the style of the Lora model contains the style characteristics of the data.
3. A lora generation method with separated subject and style according to claim 2, characterized in that The style vector and content vector are obtained by decomposing the original text annotations. It is necessary to train two independent tasks in one model to eliminate the contamination of texts and avoid cross-contamination during training.
4. A lora generation method with separated subject and style according to claim 3, characterized in that To eliminate the interference caused by feature addition and ensure compatibility, the calculation formula in step 2 is adjusted to: h=W0x+pΔW s x s +qΔW c x c Where h is the final output of the Lora model, ΔW s is the style matrix, x s is the style vector; ΔW c is the content matrix, x c is the content vector, p is the column weight vector of style features, and q is the column weight vector of content features.
5. A lora generation method with separated subject and style according to claim 4, characterized in that: Formula h=W0x+ΔW s x s +ΔW c x c The formula h=W0x+pΔW s x s +qΔW c x c Where W0 is the original large model matrix with a dimension of m*n, and x is the original network feature; the gradient of W0 is locked during training.
6. The method for generating lora with separation of subject and style according to claim 4, characterized in that: In step 1, when training Lora, the first part of the content body is converted into a content vector x through the pre-trained Clip model c and content matrix ΔW c The operation is performed, and then the feature matrix is multiplied by the corresponding weight coefficient q of each column; the second part of the style information is converted into a style vector x through the pre-trained Clip model s and style matrix ΔW s After the operation is performed, the feature matrix is multiplied by the corresponding weight coefficient p of each column; then the two parts of the feature are summed with the original large model.
7. The method for generating lora with separation of subject and style according to claim 4, characterized in that: Mode: h=W0x+pΔW s x s +qΔW c x c In the equation, p and q correspond to ΔW respectively. s , ΔW c The weight coefficient vector of the column vector of , which means the importance of each column of the matrix. If some columns are more important to the final result, the corresponding numbers will be large, and vice versa. The sum of the internal numbers of p and q is 1.
8. The method for generating lora with separated subject and style according to claim 7, characterized in that: The closer the cosine similarity value of vectors p and q is to 0, the more orthogonal the two vectors are.
9. The method for generating lora with separation of subject and style according to claim 4, characterized in that: In step 1, two Lora with different tasks learn the content and corresponding style of a data respectively.
10. A lora generation method with separated subject and style according to claim 9, characterized in that: When the input text only has content but no style, only one content Lora is run, which can directly extract the main content of the image; when the input text only has style but no content, the style of the generated image remains highly similar to the style of the original image; when the input text has both style and content, the model runs two independent Lora simultaneously, and the generated image remains highly similar to the original image, with both content and style being similar.