Collaborative feature perception-oriented large model fine tuning method and system
Through a large-modal fine-tuning method for collaborative feature perception, multimodal data and low-rank adaptation LoRA technology are used to solve the trade-off between computing resource consumption and model performance of existing methods, achieving efficient fine-tuning and good task adaptability.
Patent Information
- Application Number
- CN202510541166.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-28
- Publication Date
- 2025-05-30
- Estimated Expiration
- 2045-04-28
AI Technical Summary
Existing large model fine-tuning methods are difficult to make a good trade-off between computing resource consumption and the full use of the existing knowledge of the model, resulting in poor performance.
A large-model fine-tuning method for collaborative feature perception is adopted, through iterative fine-tuning of multimodal data sets and low-rank adaptation LoRA technology, the potential correlations between different data sources are mined, synergistic features are perceived, and the fine-tuning process is optimized using multimodal knowledge migration.
It improves the performance of large models on specific tasks, reduces training resource consumption, maintains the original capabilities of pre-trained models, and improves the migration performance on different tasks.
Smart Images

Figure CN120067618A_ABST
Abstract
Description
Technical Field
[0001] The present disclosure relates to the technical field of large language model fine-tuning, and particularly to a large model fine-tuning method and system for collaborative feature perception. Background Art
[0002] In recent years, with the rapid development of artificial intelligence technology, large language models (LLMs) have been widely applied in multiple fields such as natural language processing (NLP) and computer vision (CV). By training on large-scale datasets, large language models can learn rich semantic information and feature representations, and demonstrate excellent performance in various tasks. However, due to the usually large number of parameters in pre-trained models, directly fine-tuning on specific tasks often faces problems such as high computational resource consumption, high fine-tuning cost, and easy overfitting. Therefore, how to efficiently fine-tune large models has become one of the current research hotspots.
[0003] Existing large model fine-tuning methods mainly include full fine-tuning, adapter-based methods, low-rank adaptation (LoRA), and prompt tuning, etc. Among them, although the full fine-tuning method can obtain good performance on specific tasks, due to the need to update the parameters of the entire model, the computational cost is extremely high; the adapter-based method reduces the computational overhead by inserting small-scale adapter layers inside the model and only fine-tuning the parameters of the adapter layers. However, this method still requires additional parameters and may affect the original capabilities of the model; LoRA reduces the computational cost while maintaining the expressive power of the model by performing low-rank decomposition on the weight matrix and only adjusting the parameters of the low-rank matrix; the prompt tuning method guides the performance of the model on specific tasks by optimizing the input prompt (Prompt), avoiding direct modification of the model parameters, but its optimization process is relatively complex and there are still limitations in performance on some tasks.
[0004] In summary, among the existing large model fine-tuning methods, directly adjusting all the parameters of the model (such as full fine-tuning) has extremely high computational costs, while only adjusting some parameters (such as LoRA, Adapter, Prompt Tuning, etc.) may cause the model to not fully utilize the existing features, thus affecting the performance of downstream tasks; therefore, the existing methods do not make a good balance between computational overhead and making full use of the existing knowledge of the model, resulting in poor performance. Summary of the Invention
[0005] To solve the above problems, the present disclosure proposes a large model fine-tuning method and system for collaborative feature perception, which can improve the performance of the large model in specific tasks and reduce the consumption of training resources, and is applicable to various large-scale artificial intelligence application scenarios.
[0006] According to some embodiments, the present disclosure adopts the following technical solutions: A large model fine-tuning method for collaborative feature perception iteratively fine-tunes a large model using a multi-modal dataset composed of different data sources until the fine-tuning verification result meets a preset condition. The specific steps of each round of iterative fine-tuning are as follows: Under the guidance of the previous round of fine-tuning verification result, screen text attributes from the multi-modal data and construct a fine-tuning prompt including task objectives, execution steps, and constraints; Using the fine-tuning prompt and the Low-Rank Adaptation (LoRA) technique, explore the potential associations between different data sources, perceive collaborative features, and fine-tune the large language model; Utilize multi-modal knowledge transfer to extract multi-modal features from the multi-modal data, and generate a fine-tuning verification prompt based on the multi-modal features; Input the fine-tuning verification prompt into the fine-tuned large model for inference to generate a fine-tuning verification result. If the fine-tuning verification result does not meet the preset condition, perform the next round of iterative fine-tuning.
[0007] According to some embodiments, the present disclosure adopts the following technical solutions: A large model fine-tuning system for collaborative feature perception iteratively fine-tunes a large model using a multi-modal dataset composed of different data sources until the fine-tuning verification result meets a preset condition, including: A fine-tuning prompt construction module configured to screen text attributes from the multi-modal data under the guidance of the previous round of fine-tuning verification result and construct a fine-tuning prompt including task objectives, execution steps, and constraints; A large model fine-tuning module configured to use the fine-tuning prompt and the Low-Rank Adaptation (LoRA) technique to explore the potential associations between different data sources, perceive collaborative features, and fine-tune the large language model; A verification prompt generation module configured to utilize multi-modal knowledge transfer to extract multi-modal features from the multi-modal data and generate a fine-tuning verification prompt based on the multi-modal features; A fine-tuning verification module configured to input the fine-tuning verification prompt into the fine-tuned large model for inference to generate a fine-tuning verification result. If the fine-tuning verification result does not meet the preset condition, perform the next round of iterative fine-tuning.
[0008] According to some embodiments, the present disclosure adopts the following technical solutions: A computer program product includes a computer program, and when the computer program is executed by a processor, the method for fine-tuning a large model for collaborative feature perception is implemented.
[0009] According to some embodiments, the present disclosure adopts the following technical solutions: A non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the large model fine-tuning method for collaborative feature perception is implemented.
[0010] According to some embodiments, the present disclosure adopts the following technical solutions: An electronic device comprises: a processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory so that the electronic device executes the large model fine-tuning method for collaborative feature perception.
[0011] Compared with the prior art, the present invention has the following beneficial effects: This paper proposes a large model fine-tuning method for collaborative feature perception, which aims to mine collaborative features in tasks, improve the adaptability of the model on specific tasks, and reduce computational overhead. This method uses the feature collaboration mechanism to enable the model to adaptively select text attributes from multimodal data during the fine-tuning process, and verify it in combination with feature information of different modalities, optimize the fine-tuning process, and thus improve the generalization ability of the large language model; in addition, this method can maintain the original ability of the pre-trained model while reducing the amount of parameter updates, and improve the migration performance on different tasks.
[0012] The present invention fully considers the multimodal information involved in the task, uses CLIP and SBERT to extract features from images and text information, and introduces a multi-layer perceptron as an encoder to map features of different modalities so that they are in a unified feature representation space, thereby ensuring the comparability and fusion of features of different modalities.
[0013] The present invention constructs a fine-tuning prompt containing task objectives, execution steps, and constraints using the text information of the task to guide the large language model to accurately understand the task content. This method uses text attributes as the core elements and combines a dynamic adjustment mechanism. Under the guidance of fine-tuning verification feedback, it selects and optimizes the optimal features to adapt to different task environments. By gradually introducing text attributes such as task objectives, example data, user preferences, and environmental information, it enhances the collaborative perception ability of the model, improves the task execution effect, reduces computational resource consumption, and ensures the efficient adaptability and expressive ability of the large model. The LoRA (Low-Rank Adaptation) technology is used to efficiently fine-tune the large language model, only adjusting the parameters of some layers, reducing computational resource consumption while ensuring the expressive ability of the large model.
[0014] Based on the representation after multi-modal feature mapping, the present invention calculates the mapping loss using MSE to achieve multi-modal knowledge transfer. At the same time, a decoder with the same structure is used to map the features of different modalities back to the original feature space, and the MSE loss (i.e., the reconstruction loss) before and after the transfer is calculated to ensure the balance between multi-modal knowledge transfer and modal information, avoid the phenomenon of over-smoothing in the embedding, thereby enhancing the collaborative perception ability of different modal features, providing richer context information for the fine-tuning verification process, and improving the execution efficiency of the fine-tuning verification process.
[0015] The present invention generates the final fine-tuning verification prompt by collaborating with the large language model after feature perception fine-tuning and combining multi-modal feature embedding, inputs it into the large language model to execute the task, ensures the accurate execution of the fine-tuning inference verification task, calculates the verification metrics after the task execution, provides feedback optimization for the fine-tuning process, continuously improves the fine-tuning effect, and enhances the adaptability and generalization ability of the large language model's inference task. BRIEF DESCRIPTION OF THE DRAWINGS
[0016] The accompanying drawings forming a part of this disclosure are used to provide a further understanding of this disclosure. The illustrative embodiments and descriptions thereof of this disclosure are used to explain this disclosure and do not constitute an improper limitation of this disclosure.
[0017] Figure 1 It is the overall flowchart of a large model fine-tuning method for collaborative feature perception in Embodiment 1. Figure 2 It is the data flow processing process diagram of a large model fine-tuning method for collaborative feature perception in Embodiment 1. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0018] The present disclosure will be further described below in conjunction with the accompanying drawings and embodiments.
[0019] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure pertains.
[0020] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprise" and / or "comprising" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0021] Example 1 An embodiment of the present disclosure, relying on the Key Laboratory of Artificial Intelligence Application for People's Livelihood Services in Shandong Province and the Future Industry Laboratory of General Artificial Intelligence in Shandong Province's institutions of higher learning, provides a large model fine-tuning method for collaborative feature perception, which can improve the performance of the large model on specific tasks and reduce the consumption of training resources, and is applicable to various large-scale artificial intelligence application scenarios. Taking the product recommendation task as an example below, Figure 1 and Figure 2 are respectively the overall flowchart and the data flow processing process diagram of the large model fine-tuning method. Based on Figure 1 and 2 , the specific implementation process will be described.
[0022] For the product recommendation task of an online shopping mall, intelligent recommendations need to be made based on the user's interaction history records in the online mall. Therefore, information related to the task is collected to construct a multimodal dataset, including all product information, user information, and the user's interaction history records such as browsing, clicking, and purchasing obtained from the mall, as well as task auxiliary information such as task descriptions, environmental information, and domain knowledge. Among them, the product information includes two modalities: pictures and text; the product information here is used to extract product feature vectors, and the user information, interaction history records, and task auxiliary information are used to analyze user preferences.
[0023] After obtaining the multimodal dataset related to the task, it needs to be preprocessed. Specifically: Based on fact analysis, it is necessary to preprocess the multimodal dataset of the task to improve the data quality, reduce the impact of noise on model training, and provide more accurate and high-quality input for the fine-tuning of the large language model.
[0024] The data preprocessing here includes missing value filling, text cleaning, and data standardization, etc. For text data, multiple cleaning operations need to be performed, including removing special characters, stop word filtering, etc., and at the same time removing text attributes with more noise and more missing values.
[0025] After the preprocessing is completed, a multi-modal dataset composed of different data sources is used to iteratively fine-tune the large model until the fine-tuning verification result meets the preset conditions. The specific steps for each round of iterative fine-tuning are as follows: A multi-modal dataset composed of different data sources is used to iteratively fine-tune the large model until the fine-tuning verification result meets the preset conditions. The specific steps for each iterative fine-tuning are as follows: Step 1: Under the guidance of the previous round of fine-tuning verification results, screen text attributes from the multi-modal data to construct a fine-tuning prompt that includes the task objective, execution steps, and constraints.
[0026] Based on the text information in the multi-modal data, construct a fine-tuning prompt to guide the large language model to more accurately understand the task content, and select the best text attributes under the guidance of the fine-tuning verification results. Based on the best text attributes, construct a fine-tuning prompt. The fine-tuning prompt includes key information such as task objectives, execution steps, and constraints. At the same time, use a dynamic adjustment mechanism to dynamically adjust the prompt content according to different task requirements to make it adapt to different task environments and improve the collaborative perception ability of the large model. The specific process is as follows: (1) Initialize the task prompt: Use the most basic task description as the initial fine-tuning prompt, such as "Recommend a product for the user".
[0027] (2) Select a text attribute: Select an unused text attribute from the feature attribute pool, such as "title of the product".
[0028] (3) Enhance the fine-tuning prompt: Add this text attribute to the current fine-tuning prompt, such as "The user has previously played 'Title of Historical Product 1', 'Title of Historical Product 2', and 'Title of Historical Product 3'. Please recommend a game for the user to play next from the following options: 'Title of Candidate Product 1', 'Title of Candidate Product 1', and 'Title of Candidate Product 1'. What is the recommended game?".
[0029] The guidance of the fine-tuning verification results here is to dynamically adjust the selected text attributes through the following fine-tuning verification results. Specifically: If the newly selected text attribute improves the task execution effect, retain this feature and select a new text attribute for fine-tuning in the next round of iteration; if the newly selected text attribute does not bring significant improvement, fallback to the previous best version and sequentially select other text attributes; continue to execute the above steps until all text attributes have been tried.
[0030] This step uses text attributes as the core elements for constructing the fine-tuning prompt. Each text attribute represents a different dimension that can be used to enhance the understanding and execution ability of the LLM in the task description. The text attributes include: Task goal information: clearly describe the ultimate goal of the task, such as “classify text sentiment”, “predict the next action”, etc. Sample data: Provide specific input-output examples to allow LLM to learn how to perform tasks through cases; Domain knowledge: Introducing background information in a specific field, such as medicine, finance, law, etc., to enhance the professionalism of the task.
[0031] User preferences: In personalized recommendations or interactive tasks, optimization is performed based on user habits and historical behaviors. The "product titles that users have interacted with" mentioned above belong to the information in user preferences.
[0032] Environmental information: Combined with the context of task execution, such as time, equipment, and region, to generate more scenario-appropriate output for LLM.
[0033] Taking the task of product recommendation as an example, the goal is to generate personalized recommendation fine-tuning prompts based on user interaction history, product title, description, brand, category and other information and other auxiliary information. The method of this embodiment can combine user interaction history, product title and other information to generate fine-tuning prompts as follows: “This user has previously played Singularity, The Secret World, and Beyond: Two Souls. Please recommend a game for this user to play next from the following options: Cyberpunk 2077, Baa Baa Robots, and Gears of War: Ultimate Edition. Which game would you recommend?” Step 2: Using fine-tuning hints, use low-rank adaptation LoRA technology to explore potential correlations between different data sources, perceive collaborative features, and fine-tune the large language model.
[0034] Based on the generated fine-tuning hints, the LoRA (Low-Rank Adaptation) technology is used to adapt the parameters of some layers of the large language model, perform low-rank decomposition on only some of the model's attention layer weights, and only train low-rank matrices to reduce video memory usage and improve the model's fine-tuning efficiency, thereby reducing computing resource consumption while maintaining the model's expressiveness. The optimization goals are as follows: (1) in, Represents the parameters of LoRA training, that is, the low-rank decomposition matrix newly added in the LLM Transformer structure, which is used to freeze the original LLM parameters Lightweight fine-tuning in the case of It represents the tth token of the text sequence y output by the large language model, that is, the token to be predicted currently. , y represents all tokens before the t-th token, i.e., the context information; P is the conditional probability, indicating the probability that the model predicts the t-th token given the fine-tuning prompt x of the input and the previous t - 1 tokens.
[0035] In this process, the large model can learn collaborative features, that is, by using information such as the user's interaction history, the title, description, brand, classification, etc. of the product, and other auxiliary information, to explore the potential associations between different data sources; this ability enables the model to more accurately understand user preferences and generate results that better meet personalized needs, improving the inference effect; for example, the model can predict the trend of the user's interest changes based on the user's browsing, clicking, and purchase records, and combine the text descriptions and title information of the products to judge the potential connections between different products; the model can also comprehensively refer to the similar behavior patterns of other users and use collaborative features to mine possible points of interest.
[0036] Step 3: Use multi-modal knowledge transfer to extract multi-modal features from multi-modal data, and generate a fine-tuning verification prompt based on the multi-modal features.
[0037] After the fine-tuning of this round, it is necessary to verify the fine-tuning effect. The verification method is to generate a fine-tuning verification prompt, input it into the fine-tuned large model for inference, and generate a fine-tuning verification result. Therefore, this step is to generate a fine-tuning verification prompt, which is specifically divided into two sub-steps: 1. Extract multi-modal features from multi-modal data In different task scenarios, the information of a single modality may have limitations. Therefore, it is necessary to fuse multi-modal information to ensure that the model can make full use of various data features, improve the large model's understanding and adaptation ability to tasks during the fine-tuning process, and improve the accuracy and efficiency of the fine-tuning result verification process.
[0038] Therefore, in this embodiment, feature extraction is performed on the preprocessed multi-modal data, namely product text features, product image features, and task features: SBERT (Sentence-BERT) is used to extract features from text data to obtain the feature vector of the product text modality , CLIP (Contrastive Language-Image Pretraining) is used to extract features from image data to obtain the feature vector of the image modality , one-hot encoding (One-Hot Encoding) is performed on user information, the user's interaction history with products, and task auxiliary information respectively, and the encoded feature vectors are concatenated to form the final task feature .
[0039] Since the data features of different modalities have different representation methods and scales, for example, there may be significant differences in the feature dimensions of text feature vectors, image feature vectors, and task feature vectors, which may affect the alignment and fusion of multimodal information and the execution efficiency of the fine-tuning verification process; therefore, in order to eliminate the scale differences between different modal features and ensure that they are in the same feature space, a multi-layer perceptron (MLP) is used as a mapping network (encoder) to transform the feature vectors of different modalities into a unified feature vector space, making them comparable and fusible, which is expressed by the formula as follows:
[0040] (2)
[0041] Among them, and respectively represent the encoders for tasks, text, and images, 、 and are respectively the task feature vector, text feature vector, and image feature vector after the i-th commodity coding (i.e., mapping), 、 and are respectively the task feature vector, text feature vector, and image feature vector before the i-th commodity coding (i.e., mapping).
[0042] After the multimodal feature mapping, in order to further optimize the fusion effect of multimodal information, this embodiment uses the mean squared error (MSE) to calculate the mapping loss between different modal features to measure the information retention during the modal mapping process. Specifically, after using the encoder to map the text feature vector, image feature vector, and task feature vector into the same feature vector space, the MSE loss after mapping is calculated as follows: (3) Among them, represents the set of all users, represents the set of tasks executed by user u, represents taking the expectation over the tasks executed by all users, that is, averaging over all users, represents averaging over all the commodities interacted by user u, Calculate the mean squared error to measure their similarity in the latent space. Subsequently, a decoder with the same structure is used to convert the mapped features back to the original feature vector space, that is, to reconstruct the task feature vector, text feature vector, and image feature vector, and calculate the MSE loss again to ensure that key information is not lost during the knowledge transfer process and to avoid the phenomenon of over-smoothing in the embedding, thereby enhancing the model's perception and understanding abilities of different modality features. The reconstruction loss function is as follows:
[0043] (4)
[0044] where, , and represent the reconstruction losses of the task, text, and image respectively, and represent the decoders of the task, text, and image respectively.
[0045] The decoded task feature vector is mapped through a two-layer multi-layer perceptron to obtain the multi-modal feature vector of the i-th commodity, that is, the multi-modal feature vector of the commodity integrating task information, which is expressed by the formula: (5) In the formula, is the mapping process of the multi-layer perceptron, is the vector concatenation.
[0046] 2. Generate fine-tuning verification prompts Based on the multi-modal feature vector of each commodity i generated in formula (5) and the format of the fine-tuning prompts obtained in step one, construct fine-tuning verification prompts; on the basis of the fine-tuned large language model, use the fine-tuning verification prompts to guide the fine-tuned large language model to complete the corresponding fine-tuning verification.
[0047] Taking the fine-tuning prompts in step one as an example, it is as follows: “This user has played "Singularity" , "The Secret World" , and "Beyond: Two Souls" . Please recommend a game from the following options for this user to play next: "Cyberpunk 2077" , "Bugsnax" , and "Gears of War: Ultimate Edition" . Which game is the recommended one?” where, Multimodal feature vectors of the products "Singularity", "The Secret World", "Beyond: Two Souls", "Cyberpunk 2077", "Bleating Sheep Robot", and "Gears of War: Ultimate Edition" under the recommended scenario examples respectively.
[0048] Step 4: Input the fine-tuning verification prompt into the fine-tuned large model for inference to generate the fine-tuning verification result. If the fine-tuning verification result does not meet the preset conditions, perform the next round of iterative fine-tuning.
[0049] After the fine-tuning prompt is constructed, it is input into the fine-tuned large model for inference, the inference result is parsed to generate the fine-tuning verification result, the effectiveness of the fine-tuning prompt is evaluated, and feedback information is extracted to iteratively optimize the task execution process to ensure that the large language model can continuously improve its generalization ability. The optimization objectives are as follows:
[0050] Among them, represents the trainable parameters of represents the frozen parameters of the large language model, all the task sets in the training set, t is the current task instance, the task execution object all the tokens before the k-th token, is the task execution object the k-th token of is the input prompt for task t, including the task historical execution information and candidate information. The role of this objective function is to adjust the mapping of the multi-layer perceptron to maximize the probability of the large language model generating the correct result and improve the efficiency of the fine-tuning verification process of the large language model.
[0051] The specific parsing method is as follows: Based on the inference result, using the hit rate (Hit@1) as the evaluation index, when the hit rate meets the threshold, the fine-tuning process ends; otherwise, return to Step 1 and continue to iteratively optimize the construction of the fine-tuning prompt.
[0052] In this embodiment, the recommendation task test is carried out on three types of websites, and the test results are compared with those of other baseline models. Using the hit rate (Hit@1) as the evaluation index, the comparison results are shown in Table 1: Table 1 Test comparison
[0053] It can be seen from the table that the large model fine-tuning method for collaborative feature perception proposed in this embodiment has better effects than other fine-tuning methods, and is superior to traditional recommendation systems and large model recommendation systems in the specific recommendation task execution effects.
[0054] Embodiment 2 In one embodiment of the present disclosure, a large model fine-tuning system for collaborative feature perception is provided, which uses a multimodal data set composed of different data sources to iteratively fine-tune the large model until the fine-tuning verification result meets the preset conditions, including: The fine-tuning prompt construction module is configured to: filter text attributes from multimodal data under the guidance of the previous round of fine-tuning verification results, and construct fine-tuning prompts containing task goals, execution steps, and constraints; The large model fine-tuning module is configured to: utilize fine-tuning hints, use low-rank adaptation LoRA technology, mine potential associations between different data sources, perceive collaborative features, and fine-tune the large language model; The verification prompt generation module is configured to: utilize multimodal knowledge transfer to extract multimodal features in multimodal data, and generate fine-tuning verification prompts based on the multimodal features; The fine-tuning verification module is configured to: input the fine-tuning verification prompts into the fine-tuned large model for reasoning, generate fine-tuning verification results, and perform the next round of iterative fine-tuning if the fine-tuning verification results do not meet the preset conditions.
[0055] Example 3 In one embodiment of the present disclosure, a computer program product is provided, including a computer program, which, when executed by a processor, implements the large model fine-tuning method for collaborative feature perception.
[0056] Example 4 In one embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided, wherein the non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the large model fine-tuning method for collaborative feature perception is implemented.
[0057] Example 5 In one embodiment of the present disclosure, an electronic device is provided, including: a processor, a memory, and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory, so that the electronic device executes the large model fine-tuning method for collaborative feature perception.
[0058] This disclosure is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to embodiments of the disclosure. It should be understood that each flow and / or block in the flowcharts and / or block diagrams, and the combination of flows and / or blocks in the flowcharts and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to the processors of general-purpose computers, special-purpose computers, embedded processors, or other programmable data processing devices to generate a machine, such that the instructions executed by the processors of the computer or other programmable data processing devices produce means for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0059] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one flow Figure 1 one flow or multiple flows and / or blocks Figure 1 or multiple blocks.
[0060] Although the specific embodiments of the disclosure have been described above in conjunction with the accompanying drawings, they are not limitations on the protection scope of the disclosure. Those skilled in the art should understand that, based on the technical solutions of the disclosure, various modifications or variations that can be made by those skilled in the art without creative efforts are still within the protection scope of the disclosure.
Claims
1. A large model fine-tuning method for collaborative feature perception, characterized in that: Using a multimodal dataset composed of different data sources, the large model is iteratively fine-tuned until the fine-tuning verification results meet the preset conditions. The specific steps of each round of iterative fine-tuning are: Guided by the verification results of the previous round of fine-tuning, we filter text attributes from multimodal data and construct fine-tuning prompts containing task objectives, execution steps, and constraints. Using fine-tuning hints, we use low-rank adaptation LoRA technology to mine potential correlations between different data sources, perceive collaborative features, and fine-tune large language models; Utilize multimodal knowledge transfer to extract multimodal features from multimodal data, and generate fine-tuning verification prompts based on multimodal features; The fine-tuning verification prompts are input into the fine-tuned large model for reasoning to generate fine-tuning verification results. If the fine-tuning verification results do not meet the preset conditions, the next round of iterative fine-tuning is performed.
2. A large model fine-tuning method for collaborative feature perception as claimed in claim 1, characterized in that: The large model is used for product recommendation tasks, and the data sources include product information, user information, historical records of user-product interactions, and task auxiliary information, wherein the product information includes two modes: picture and text.
3. The large model fine-tuning method for collaborative feature perception according to claim 1, characterized in that: It also includes preprocessing of multimodal data, including data cleaning, missing value filling and format standardization.
4. The large model fine-tuning method for collaborative feature perception according to claim 1, characterized in that: The step of extracting multimodal features from multimodal data is specifically as follows: Extract product text features and product image features from product information, and extract task features from user information, user-product interaction history, and task auxiliary information; The encoder is used to map product text features, product image features, and task features to unify the feature representation space. During the mapping process, MSE is used to calculate the mapping loss and perform multimodal knowledge transfer. A decoder with the same structure is used to decode the product text features, product image features and task features, and reconstruct the original feature space. During the reconstruction process, MSE is used to calculate the reconstruction loss to achieve a balance between multimodal knowledge transfer and modal information, and finally obtain the multimodal features of the product.
5. A large model fine-tuning method for collaborative feature perception as claimed in claim 4, characterized in that: The mapping loss, which calculates the MSE loss after mapping, is expressed as follows: in, represents the set of all users, represents the set of tasks performed by user u, It means taking the expectation on all the tasks executed by all users. It means the average of all the products interacted by the user u. Compute the mean squared error, a measure of how similar they are in the latent space.
6. A large model fine-tuning method for collaborative feature perception as claimed in claim 4, characterized in that: The reconstruction loss calculates the MSE loss before and after reconstruction and is expressed as follows: in, , and Represent the reconstruction losses of tasks, texts, and images respectively.
7. A large model fine-tuning system for collaborative feature perception, characterized in that: Using multimodal datasets composed of different data sources, the large model is iteratively fine-tuned until the fine-tuning verification results meet the preset conditions, including: The fine-tuning prompt construction module is configured to: filter text attributes from multimodal data under the guidance of the previous round of fine-tuning verification results, and construct fine-tuning prompts containing task goals, execution steps, and constraints; The large model fine-tuning module is configured to: utilize fine-tuning hints, use low-rank adaptation LoRA technology, mine potential associations between different data sources, perceive collaborative features, and fine-tune the large language model; The verification prompt generation module is configured to: utilize multimodal knowledge transfer to extract multimodal features in multimodal data, and generate fine-tuning verification prompts based on the multimodal features; The fine-tuning verification module is configured to: input the fine-tuning verification prompts into the fine-tuned large model for reasoning, generate fine-tuning verification results, and perform the next round of iterative fine-tuning if the fine-tuning verification results do not meet the preset conditions.
8. A computer program product, comprising a computer program, characterized in that When the computer program is executed by a processor, it implements a large model fine-tuning method for collaborative feature perception as described in any one of claims 1 to 6.
9. A non-transitory computer-readable storage medium, characterized in that: The non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by the processor, a large model fine-tuning method for collaborative feature perception as described in any one of claims 1-6 is implemented.
10. An electronic device, characterized in that: include: A processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory to enable the electronic device to implement a large model fine-tuning method for collaborative feature perception as described in any one of claims 1-6.
Citation Information
Patent Citations
Task execution method and device for large model, electronic equipment, storage medium and program product
CN118519779A
Large language model fine tuning method and device based on causal relationship perception
CN119443182A
Large model technology-based ultrahigh-speed optical module digital manufacturing scene knowledge reasoning method, device, equipment and medium
CN119692476A
Case question answering method based on large language model, medium and equipment
CN119692484A
Multi-modal large model optimization method and system based on defect detection and analysis
CN119721165A