A Fine-tuning Method and System for Large Models Oriented to Collaborative Feature Perception

The collaborative feature-aware fine-tuning method enhances LLM performance on specific tasks by using multi-modal data and LoRA for efficient adaptation, addressing the computational and feature utilization imbalance in existing methods.

CN120067618BActive Publication Date: 2025-07-15SHANDONG UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510541166.2
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-28
Publication Date
2025-07-15
Estimated Expiration
2045-04-28

AI Technical Summary

Technical Problem

Existing large model fine-tuning methods fail to trade off well between computing resource consumption and the full use of the existing knowledge of the model, resulting in poor performance.

Method used

The large-model fine-tuning method of collaborative feature perception is adopted, and iterative fine-tuning is used to use multimodal data sets. Combined with LoRA technology and multimodal feature mapping, different modal features are processed through low-rank adaptation and multi-layer perceptrons to optimize the fine-tuning process.

Benefits of technology

It improves the adaptability and generalization ability of the large model on specific tasks, reduces the consumption of computing resources, and maintains the expression ability and migration performance of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120067618B_ABST
    Figure CN120067618B_ABST
Patent Text Reader

Abstract

The present disclosure provides a large model fine-tuning method and system for collaborative feature perception, which relates to the technical field of large language model fine-tuning. A multi-modal data set composed of different data sources is used to iteratively fine-tune the large model until the fine-tuning verification result meets the preset conditions. Each iteration is as follows: under the guidance of the previous fine-tuning verification result, text attributes are screened from the multi-modal data to construct a fine-tuning prompt; using the fine-tuning prompt and the low-rank adaptation technology, collaborative features are perceived to fine-tune the large language model; using multi-modal knowledge transfer, multi-modal features in the multi-modal data are extracted to generate a fine-tuning verification prompt; the fine-tuning verification prompt is input into the fine-tuned large model for inference to generate a fine-tuning verification result. If the fine-tuning verification result does not meet the preset conditions, the next round of iterative fine-tuning is performed; the present invention can improve the performance of the large model on specific tasks and reduce the consumption of training resources, and is applicable to various large-scale artificial intelligence application scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present disclosure relates to the technical field of large language model fine-tuning, and particularly to a large model fine-tuning method and system for collaborative feature perception. Background Art

[0002] In recent years, with the rapid development of artificial intelligence technology, large language models (LLMs) have been widely applied in multiple fields such as natural language processing (NLP) and computer vision (CV). By training on large-scale datasets, large language models can learn rich semantic information and feature representations, and demonstrate excellent performance in various tasks. However, due to the usually large number of parameters in pre-trained models, directly fine-tuning on specific tasks often faces problems such as high computational resource consumption, high fine-tuning cost, and easy overfitting. Therefore, how to efficiently fine-tune large models has become one of the current research hotspots.

[0003] Existing large model fine-tuning methods mainly include full fine-tuning, adapter-based methods, low-rank adaptation (LoRA), and prompt tuning, etc. Among them, although the full fine-tuning method can achieve good performance on specific tasks, due to the need to update the parameters of the entire model, the computational cost is extremely high; the adapter-based method inserts small-scale adapter layers inside the model and only fine-tunes the parameters of the adapter layers to reduce the computational overhead. However, this method still requires additional parameters and may affect the original capabilities of the model; LoRA reduces the computational cost by performing low-rank decomposition on the weight matrix and only adjusting the parameters of the low-rank matrix, while maintaining the expressive power of the model; the prompt tuning method guides the performance of the model on specific tasks by optimizing the input prompt, avoiding direct modification of the model parameters, but its optimization process is relatively complex, and there are still limitations in performance on some tasks.

[0004] In summary, among the existing large model fine-tuning methods, directly adjusting all the parameters of the model (such as full fine-tuning) has extremely high computational costs, while only adjusting some parameters (such as LoRA, Adapter, Prompt Tuning, etc.) may cause the model to not fully utilize the existing features, thus affecting the performance of downstream tasks; therefore, the existing methods do not make a good trade-off between computational overhead and fully utilizing the existing knowledge of the model, resulting in poor performance. Summary of the invention

[0005] In order to solve the above problems, the present invention proposes a large model fine-tuning method and system for collaborative feature perception, which can improve the performance of large models on specific tasks and reduce the consumption of training resources, and is suitable for various large-scale artificial intelligence application scenarios.

[0006] According to some embodiments, the present disclosure adopts the following technical solutions:

[0007] A large model fine-tuning method for collaborative feature perception uses a multimodal dataset composed of different data sources to iteratively fine-tune the large model until the fine-tuning verification result meets the preset conditions. The specific steps of each round of iterative fine-tuning are:

[0008] Guided by the verification results of the previous round of fine-tuning, we filter text attributes from multimodal data and construct fine-tuning prompts containing task objectives, execution steps, and constraints.

[0009] Using fine-tuning hints, we use low-rank adaptation LoRA technology to mine potential correlations between different data sources, perceive collaborative features, and fine-tune large language models;

[0010] Utilize multimodal knowledge transfer to extract multimodal features from multimodal data, and generate fine-tuning verification prompts based on multimodal features;

[0011] The fine-tuning verification prompts are input into the fine-tuned large model for reasoning to generate fine-tuning verification results. If the fine-tuning verification results do not meet the preset conditions, the next round of iterative fine-tuning is performed.

[0012] According to some embodiments, the present disclosure adopts the following technical solutions:

[0013] A large model fine-tuning system for collaborative feature perception uses a multimodal dataset composed of different data sources to iteratively fine-tune the large model until the fine-tuning verification result meets the preset conditions, including:

[0014] The fine-tuning prompt construction module is configured to: filter text attributes from multimodal data under the guidance of the previous round of fine-tuning verification results, and construct fine-tuning prompts containing task goals, execution steps, and constraints;

[0015] The large model fine-tuning module is configured to: utilize fine-tuning hints, use low-rank adaptation LoRA technology, mine potential associations between different data sources, perceive collaborative features, and fine-tune the large language model;

[0016] The verification prompt generation module is configured to: utilize multimodal knowledge transfer to extract multimodal features in multimodal data, and generate fine-tuning verification prompts based on the multimodal features;

[0017] The fine-tuning verification module is configured to: input the fine-tuning verification prompts into the fine-tuned large model for reasoning, generate fine-tuning verification results, and perform the next round of iterative fine-tuning if the fine-tuning verification results do not meet the preset conditions.

[0018] According to some embodiments, the present disclosure adopts the following technical solutions:

[0019] A computer program product includes a computer program, and when the computer program is executed by a processor, the method for fine-tuning a large model for collaborative feature perception is implemented.

[0020] According to some embodiments, the present disclosure adopts the following technical solutions:

[0021] A non-transitory computer-readable storage medium is used to store computer instructions. When the computer instructions are executed by a processor, the large model fine-tuning method for collaborative feature perception is implemented.

[0022] According to some embodiments, the present disclosure adopts the following technical solutions:

[0023] An electronic device comprises: a processor, a memory and a computer program; wherein the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device is running, the processor executes the computer program stored in the memory so that the electronic device executes the large model fine-tuning method for collaborative feature perception.

[0024] Compared with the prior art, the present invention has the following beneficial effects:

[0025] This paper proposes a large model fine-tuning method for collaborative feature perception, which aims to mine collaborative features in tasks, improve the adaptability of the model on specific tasks, and reduce computational overhead. This method uses the feature collaboration mechanism to enable the model to adaptively select text attributes from multimodal data during the fine-tuning process, and verify it in combination with feature information of different modalities, optimize the fine-tuning process, and thus improve the generalization ability of the large language model; in addition, this method can maintain the original ability of the pre-trained model while reducing the amount of parameter updates, and improve the migration performance on different tasks.

[0026] The present invention fully considers the multimodal information involved in the task, uses CLIP and SBERT to extract features from images and text information, and introduces a multi-layer perceptron as an encoder to map features of different modalities so that they are in a unified feature representation space, thereby ensuring the comparability and fusion of features of different modalities.

[0027] The present invention constructs a fine-tuning prompt containing task objectives, execution steps, and constraints using the text information of the task, guiding the large language model to accurately understand the task content. This method uses text attributes as the core elements and combines a dynamic adjustment mechanism. Under the guidance of fine-tuning verification feedback, it selects and optimizes the optimal features to adapt to different task environments. By gradually introducing text attributes such as task objectives, example data, user preferences, and environmental information, it enhances the collaborative perception ability of the model, improves the task execution effect, reduces the consumption of computing resources, and ensures the efficient adaptability and expression ability of the large model. The LoRA (Low-Rank Adaptation) technology is used to efficiently fine-tune the large language model, adjusting only the parameters of some layers, reducing the consumption of computing resources while ensuring the expression ability of the large model.

[0028] Based on the representation after multi-modal feature mapping, the present invention calculates the mapping loss using MSE to achieve multi-modal knowledge transfer. At the same time, a decoder with the same structure is used to map the features of different modalities back to the original feature space, and the MSE loss (i.e., the reconstruction loss) before and after the transfer is calculated to ensure the balance between multi-modal knowledge transfer and modal information, avoid the phenomenon of over-smoothing in the embedding, thereby enhancing the collaborative perception ability of different modal features, providing richer context information for the fine-tuning verification process, and improving the execution efficiency of the fine-tuning verification process.

[0029] The present invention generates the final fine-tuning verification prompt by collaborating with the large language model after feature perception and combining multi-modal feature embedding, inputs it into the large language model to execute the task, ensures the accurate execution of the fine-tuning inference verification task, calculates the verification metrics after the task execution, provides feedback optimization for the fine-tuning process, continuously improves the fine-tuning effect, and enhances the adaptability and generalization ability of the large language model's inference task. Brief Description of the Drawings

[0030] The schematic diagrams in the specification that form a part of this disclosure are used to provide a further understanding of this disclosure. The illustrative embodiments and their descriptions of this disclosure are used to explain this disclosure and do not constitute an improper limitation of this disclosure.

[0031] Figure 1 It is the overall flowchart of a large model fine-tuning method for collaborative feature perception in Embodiment 1.

[0032] Figure 2 It is the process diagram of data flow processing for a large model fine-tuning method for collaborative feature perception in Embodiment 1. Detailed Embodiments

[0033] The following further describes this disclosure in conjunction with the drawings and embodiments.

[0034] It should be noted that the following detailed description is exemplary and is intended to provide further illustration of the present disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which the present disclosure pertains.

[0035] It should be noted that the terms used herein are only for describing specific embodiments and are not intended to limit the exemplary embodiments according to the present disclosure. As used herein, unless the context clearly indicates otherwise, the singular form is also intended to include the plural form. In addition, it should be understood that when the terms "comprising" and / or "comprise" are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.

[0036] Example 1

[0037] An embodiment of the present disclosure, relying on the Key Laboratory of Artificial Intelligence Application for People's Livelihood Services in Shandong Province and the Future Industry Laboratory of General Artificial Intelligence in Shandong Province's institutions of higher learning, provides a large model fine-tuning method for collaborative feature perception, which can improve the performance of the large model on specific tasks and reduce the consumption of training resources, and is applicable to various large-scale artificial intelligence application scenarios. Taking the product recommendation task as an example below, Figure 1 and Figure 2 are respectively the overall flowchart and the data flow processing process diagram of the large model fine-tuning method. Based on Figure 1 and 2 , the specific implementation process will be described.

[0038] For the product recommendation task of an online shopping mall, intelligent recommendations need to be made based on the user's interaction history records in the online mall. Therefore, information related to the task is collected to construct a multimodal dataset, including all product information, user information, and the user's interaction history records such as browsing, clicking, and purchasing obtained from the mall, as well as task auxiliary information such as task descriptions, environmental information, and domain knowledge. Among them, the product information includes two modalities: pictures and texts; the product information here is used to extract product feature vectors, and the user information, interaction history records, and task auxiliary information are used to analyze user preferences.

[0039] After obtaining the multimodal dataset related to the task, it needs to be preprocessed. Specifically:

[0040] Based on fact analysis, it is necessary to perform data preprocessing on the multimodal dataset of the task to improve the data quality, reduce the impact of noise on model training, and provide more accurate and high-quality input for the fine-tuning of the large language model.

[0041] Data preprocessing here includes missing value filling, text cleaning, and data normalization, etc. For text data, multiple cleaning operations need to be performed, including removing special characters, stop word filtering, etc., and at the same time, text attributes with a lot of noise and many missing values are removed.

[0042] After the preprocessing is completed, use the multi-modal dataset composed of different data sources to iteratively fine-tune the large model until the fine-tuning verification result meets the preset conditions. The specific steps for each round of iterative fine-tuning are as follows:

[0043] Use the multi-modal dataset composed of different data sources to iteratively fine-tune the large model until the fine-tuning verification result meets the preset conditions. The specific steps for each iterative fine-tuning are as follows:

[0044] Step 1: Guided by the fine-tuning verification result of the previous round, screen text attributes from the multi-modal data to construct a fine-tuning prompt that includes the task objective, execution steps, and constraints.

[0045] Based on the text information in the multi-modal data, construct a fine-tuning prompt to guide the large language model to more accurately understand the task content, and select the best text attributes under the guidance of the fine-tuning verification result. Based on the best text attributes, construct a fine-tuning prompt. The fine-tuning prompt includes key information such as task objectives, execution steps, and constraints. At the same time, use a dynamic adjustment mechanism to dynamically adjust the prompt content according to different task requirements to make it adapt to different task environments and improve the collaborative perception ability of the large model. The specific process is as follows:

[0046] (1) Initialize the task prompt: Use the most basic task description as the initial fine-tuning prompt, such as "Recommend a product for the user".

[0047] (2) Select a text attribute: Select an unused text attribute from the feature attribute pool, such as "the title of the product".

[0048] (3) Enhance the fine-tuning prompt: Add this text attribute to the current fine-tuning prompt, such as "The user has played 'Title of Historical Product 1', 'Title of Historical Product 2', and 'Title of Historical Product 3' before. Please recommend a game from the following options for the user to play next: 'Title of Candidate Product 1', 'Title of Candidate Product 1', and 'Title of Candidate Product 1'. What is the recommended game?".

[0049] The guidance for fine-tuning verification results here is to dynamically adjust the selected text attributes through the following fine-tuning verification results. Specifically, if the newly selected text attribute improves the task execution effect, then retain the feature and select a new text attribute for fine-tuning in the next iteration; if the newly selected text attribute does not bring significant improvement, then fall back to the previous best version and select other text attributes sequentially; continue to perform the above steps until all text attributes have been tried.

[0050] This step uses text attributes as the core elements for constructing fine-tuning prompts. Each text attribute represents a different dimension in the task description that can be used to enhance LLM understanding and execution. The text attributes include:

[0051] Task goal information: clearly describe the ultimate goal of the task, such as “classify text sentiment”, “predict the next action”, etc.

[0052] Sample data: Provide specific input-output examples to allow LLM to learn how to perform tasks through cases;

[0053] Domain knowledge: Introducing background information in a specific field, such as medicine, finance, law, etc., to enhance the professionalism of the task.

[0054] User preferences: In personalized recommendations or interactive tasks, optimization is performed based on user habits and historical behaviors. The "product titles that users have interacted with" mentioned above belong to the information in user preferences.

[0055] Environmental information: Combined with the context of task execution, such as time, equipment, and region, to generate more scenario-appropriate output for LLM.

[0056] Taking the task of product recommendation as an example, the goal is to generate personalized recommendation fine-tuning prompts based on user interaction history, product title, description, brand, category and other information and other auxiliary information. The method of this embodiment can combine user interaction history, product title and other information to generate fine-tuning prompts as follows:

[0057] “This user has previously played Singularity, The Secret World, and Beyond: Two Souls. Please recommend a game for this user to play next from the following options: Cyberpunk 2077, Baa Baa Robots, and Gears of War: Ultimate Edition. Which game would you recommend?”

[0058] Step 2: Using fine-tuning hints, use low-rank adaptation LoRA technology to explore potential correlations between different data sources, perceive collaborative features, and fine-tune the large language model.

[0059] Based on the generated fine-tuning hints, the LoRA (Low-Rank Adaptation) technology is used to adapt the parameters of some layers of the large language model, perform low-rank decomposition on only some of the model's attention layer weights, and only train low-rank matrices to reduce video memory usage and improve the model's fine-tuning efficiency, thereby reducing computing resource consumption while maintaining the model's expressiveness. The optimization goals are as follows:

[0060] (1)

[0061] in, Represents the parameters of LoRA training, that is, the low-rank decomposition matrix newly added in the LLM Transformer structure, which is used to freeze the original LLM parameters Lightweight fine-tuning in the case of It represents the tth token of the text sequence y output by the large language model, that is, the token to be predicted currently. , y is all tokens before the t-th token, i.e., context information; P is the conditional probability, which means the probability of the model predicting the t-th token given the input fine-tuning hint x and the first t−1 tokens.

[0062] In this process, the large model is able to learn collaborative features, that is, to explore potential connections between different data sources through the user's interaction history, product title, description, brand, classification and other information and other auxiliary information; this ability enables the model to understand user preferences more accurately, and generate results that are more in line with personalized needs, improving the reasoning effect; for example, the model can predict the user's interest change trend based on the user's browsing, clicking, and purchasing records, and combine the product's text description and title information to determine the potential connection between different products; the model can also comprehensively refer to similar behavior patterns of other users and use collaborative features to explore possible points of interest.

[0063] Step 3: Utilize multimodal knowledge transfer to extract multimodal features from multimodal data, and generate fine-tuning verification prompts based on the multimodal features.

[0064] After this round of fine-tuning is completed, the fine-tuning effect needs to be verified. The verification method is to generate fine-tuning verification prompts, input them into the fine-tuned large model for inference, and generate fine-tuning verification results. Therefore, this step is to generate fine-tuning verification prompts, which is specifically divided into two sub-steps:

[0065] 1. Extracting multimodal features from multimodal data

[0066] In different task scenarios, the information of a single modality may have limitations. Therefore, it is necessary to fuse multi-modal information to ensure that the model can make full use of various data features, improve the understanding and adaptation ability of the large model to tasks during the fine-tuning process, and improve the accuracy and efficiency of the fine-tuning result verification process.

[0067] Therefore, in this embodiment, feature extraction is performed on the preprocessed multi-modal data, namely commodity text features, commodity picture features, and task features: SBERT (Sentence-BERT) is used to extract features from the text data to obtain the feature vectors of the commodity text modality. CLIP (Contrastive Language-Image Pretraining) is used to extract features from the picture data to obtain the feature vectors of the picture modality. One-Hot Encoding is performed on the user information, the interaction history between the user and the commodity, and the task auxiliary information respectively, and the encoded feature vectors are concatenated to form the final task features. .

[0068] Since the data features of different modalities have different representation methods and scales, for example, there may be large differences in the feature dimensions of text feature vectors, picture feature vectors, and task feature vectors, which may affect the alignment and fusion of multi-modal information and the execution efficiency of the fine-tuning verification process; therefore, in order to eliminate the scale differences between different modal features and ensure that they are in the same feature space, a multi-layer perceptron (MLP) is used as a mapping network (encoder) to transform the feature vectors of different modalities into a unified feature vector space, making them comparable and fusible. It is expressed by the formula:

[0069]

[0070] (2)

[0071]

[0072] Among them, and respectively represent the encoders of tasks, text, and pictures. , and are respectively the task feature vector, text feature vector, and picture feature vector of the i-th commodity after encoding (i.e., mapping). , and are respectively the task feature vector, text feature vector, and picture feature vector of the i-th commodity before encoding (i.e., mapping).

[0073] After multi-modal feature mapping, in order to further optimize the fusion effect of multi-modal information, this embodiment uses the mean square error (MSE) to calculate the mapping loss between different modal features to measure the preservation of information during the modal mapping process. Specifically, after using the encoder to map the text feature vector, the image feature vector, and the task feature vector to the same feature vector space, the MSE loss after mapping is calculated as follows:

[0074] (3)

[0075] Among them, represents the set of all users, represents the set of tasks executed by user u, represents taking the expectation over all tasks executed by all users, that is, averaging over all users, represents averaging over all the products interacted with by user u, Calculate the mean square error to measure their similarity in the latent space.

[0076] Subsequently, use a decoder with the same structure to convert the mapped features back to the original feature vector space, that is, reconstruct the task feature vector, the text feature vector, and the image feature vector, and calculate the MSE loss again to ensure that key information will not be lost during the knowledge transfer process and avoid the phenomenon of over-smoothing of the embedding, thereby improving the model's perception and understanding ability of different modal features. The reconstruction loss function is:

[0077]

[0078] (4)

[0079]

[0080] Among them, , and respectively represent the reconstruction losses of tasks, text, and images, and respectively represent the decoders of tasks, text, and images.

[0081] The decoded task feature vector , after being mapped through two layers of multi-layer perceptrons, obtains the multi-modal feature vector of the i-th product, that is, the multi-modal feature vector of the product integrating task information, which is expressed by the formula as:

[0082] (5)

[0083] In the formula, For the mapping process of the multi-layer perceptron, For vector concatenation.

[0084] 2. Generate fine-tuning verification prompts

[0085] Based on the multi-modal feature vectors of each commodity i generated in formula (5) And the format of the fine-tuning prompts obtained in the first step, construct fine-tuning verification prompts; based on the fine-tuned large language model, use the fine-tuning verification prompts to guide the fine-tuned large language model to complete the corresponding fine-tuning verification.

[0086] Taking the fine-tuning prompts in the first step as an example, it is as follows:

[0087] "This user has played "Singularity" , "The Secret World" and "Beyond: Two Souls" . Please recommend a game from the following options for this user to play next: "Cyberpunk 2077" , "Bloop Bloop" , and "Gears of War: Ultimate Edition" . What is the recommended game?"

[0088] Among them, Are the multi-modal feature vectors of the commodities "Singularity", "The Secret World", "Beyond: Two Souls", "Cyberpunk 2077", "Bloop Bloop", and "Gears of War: Ultimate Edition" respectively under the recommended scenario example.

[0089] Step 4: Input the fine-tuning verification prompts into the fine-tuned large model for reasoning to generate fine-tuning verification results. If the fine-tuning verification results do not meet the preset conditions, perform the next round of iterative fine-tuning.

[0090] After the fine-tuning prompts are constructed, input them into the fine-tuned large model for reasoning, analyze the reasoning results, generate fine-tuning verification results, evaluate the effectiveness of the fine-tuning prompts, and extract feedback information to iteratively optimize the task execution process to ensure that the large language model can continuously improve its generalization ability. The optimization goals are as follows:

[0091]

[0092] Among them, Represents The trainable parameters of, Represents the frozen parameters of the large language model, All task sets in the training set, t is the current task instance, Task execution object All tokens before the kth token, Task execution object The kth token of is the input prompt for task t, which contains task history execution information and candidate information. The objective function is to adjust the mapping of the multi-layer perceptron to maximize the probability of the large language model generating correct results and improve the efficiency of the fine-tuning verification process of the large language model.

[0093] The specific analysis method is: based on the inference results, the hit rate (Hit@1) is used as the evaluation indicator. When the hit rate meets the threshold, the fine-tuning process ends. Otherwise, it returns to step 1 and continues to iteratively optimize the construction of fine-tuning prompts.

[0094] This example performs recommendation task tests on three types of websites and compares the test results with those of other baseline models. The hit rate (Hit@1) is used as the evaluation indicator. The comparison results are shown in Table 1:

[0095] Table 1 Test comparison

[0096]

[0097] It can be seen from the table that the large model fine-tuning method for collaborative feature perception proposed in this embodiment is better than other fine-tuning methods, and is better than the traditional recommendation system and the large model recommendation system in terms of the execution effect of specific recommendation tasks.

[0098] Example 2

[0099] In one embodiment of the present disclosure, a large model fine-tuning system for collaborative feature perception is provided, which uses a multimodal data set composed of different data sources to iteratively fine-tune the large model until the fine-tuning verification result meets the preset conditions, including:

[0100] The fine-tuning prompt construction module is configured to: filter text attributes from multimodal data under the guidance of the previous round of fine-tuning verification results, and construct fine-tuning prompts containing task goals, execution steps and constraints;

[0101] The large model fine-tuning module is configured to: utilize fine-tuning hints, use low-rank adaptation LoRA technology, mine potential associations between different data sources, perceive collaborative features, and fine-tune the large language model;

[0102] The verification prompt generation module is configured to: utilize multimodal knowledge transfer to extract multimodal features in multimodal data, and generate fine-tuning verification prompts based on the multimodal features;

[0103] The fine-tuning verification module is configured to: input the fine-tuning verification prompts into the fine-tuned large model for reasoning, generate fine-tuning verification results, and perform the next round of iterative fine-tuning if the fine-tuning verification results do not meet the preset conditions.

[0104] Example 3

[0105] In one embodiment of the present disclosure, a computer program product is provided, including a computer program, and when the computer program is executed by a processor, it implements the method for fine-tuning a large model for collaborative feature perception as described above.

[0106] Example 4

[0107] In one embodiment of the present disclosure, a non-transitory computer-readable storage medium is provided, and the non-transitory computer-readable storage medium is used to store computer instructions, and when the computer instructions are executed by a processor, it implements the method for fine-tuning a large model for collaborative feature perception as described above.

[0108] Example 5

[0109] In one embodiment of the present disclosure, an electronic device is provided, including: a processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes and implements the method for fine-tuning a large model for collaborative feature perception as described above.

[0110] The present disclosure is described with reference to the flowcharts and / or block diagrams of methods, devices (systems), and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or block in the flowchart and / or block diagram, and the combination of processes and / or blocks in the flowchart and / or block diagram, can be implemented by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, so that the instructions executed by the processor of the computer or other programmable data processing devices generate a device for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0111] These computer program instructions can also be loaded onto a computer or other programmable data processing device, so that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one process Figure 1 one process or multiple processes and / or blocks Figure 1 one block or multiple blocks.

[0112] Although the specific embodiments of the present disclosure have been described above in conjunction with the accompanying drawings, they are not intended to limit the scope of protection of the present disclosure. Those skilled in the art should understand that various modifications or variations that can be made without creative efforts on the basis of the technical solutions of the present disclosure are still within the scope of protection of the present disclosure.

Claims

1. A large model fine-tuning method for collaborative feature perception, characterized in that Iteratively fine-tune the large model using a multi-modal dataset composed of different data sources until the fine-tuning verification result meets the preset conditions. The specific steps for each round of iterative fine-tuning are as follows: Under the guidance of the previous round's fine-tuning verification result, screen text attributes from the multi-modal data and construct a fine-tuning prompt that includes the task objective, execution steps, and constraints; Use the low-rank adaptation (LoRA) technique with the fine-tuning prompt to explore the potential associations between different data sources, perceive collaborative features, and fine-tune the large language model; Utilize multi-modal knowledge transfer to extract multi-modal features from the multi-modal data and generate a fine-tuning verification prompt based on the multi-modal features; Input the fine-tuning verification prompt into the fine-tuned large model for inference to generate a fine-tuning verification result. If the fine-tuning verification result does not meet the preset conditions, perform the next round of iterative fine-tuning.

2. The fine-tuning method for a large model for collaborative feature perception according to claim 1, wherein The large model is used for the commodity recommendation task. The data sources include commodity information, user information, the interaction history between users and commodities, and task assistance information. Among them, the commodity information includes two modalities: pictures and text.

3. The fine-tuning method for a large model for collaborative feature perception according to claim 1, wherein, It also includes preprocessing the multi-modal data, including data cleaning, missing value completion, and format standardization.

4. The fine-tuning method for a large model for collaborative feature perception according to claim 1, wherein The extraction of multi-modal features from the multi-modal data is specifically as follows: Extract commodity text features and commodity picture features from the commodity information, and extract task features from the user information, the interaction history between users and commodities, and the task assistance information; Map the commodity text features, commodity picture features, and task features through an encoder to unify the feature representation space. During the mapping process, use the mean squared error (MSE) to calculate the mapping loss and perform multi-modal knowledge transfer; Use a decoder with the same structure to decode the commodity text features, commodity picture features, and task features to reconstruct the original feature space. During the reconstruction process, use the MSE to calculate the reconstruction loss to achieve the balance between multi-modal knowledge transfer and modal information, and finally obtain the multi-modal features of the commodity.

5. The fine-tuning method for a large model for collaborative feature perception according to claim 4, wherein The mapping loss calculates the MSE loss after mapping and is expressed by the formula: Among them, represents the set of all users, represents the set of tasks executed by user u, represents taking the expectation over the tasks executed by all users, represents averaging all the commodities interacted by the user u, Calculate the mean squared error to measure their similarity in the latent space.

6. The fine-tuning method for large models for collaborative feature perception according to claim 4, wherein The reconstruction loss calculates the MSE loss before and after reconstruction and is expressed by the formula: Among them, , and respectively represent the reconstruction losses of tasks, texts, and pictures.

7. A large model fine-tuning system for collaborative feature perception, characterized in that, Iteratively fine-tune the large model using a multi-modal dataset composed of different data sources until the fine-tuning verification result meets the preset conditions, including: A fine-tuning prompt construction module, configured to: under the guidance of the previous round's fine-tuning verification result, screen text attributes from the multi-modal data and construct a fine-tuning prompt that includes the task objective, execution steps, and constraints; A large model fine-tuning module, configured to: use the low-rank adaptation (LoRA) technique with the fine-tuning prompt to explore the potential associations between different data sources, perceive collaborative features, and fine-tune the large language model; A verification prompt generation module, configured to: utilize multi-modal knowledge transfer to extract multi-modal features from the multi-modal data and generate a fine-tuning verification prompt based on the multi-modal features; A fine-tuning verification module, configured to: input the fine-tuning verification prompt into the fine-tuned large model for inference to generate a fine-tuning verification result. If the fine-tuning verification result does not meet the preset conditions, perform the next round of iterative fine-tuning.

8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements a large model fine-tuning method for collaborative feature perception according to any one of claims 1-6.

9. A non-transitory computer-readable storage medium, characterized in that, The non-transitory computer-readable storage medium is used to store computer instructions, which, when executed by a processor, implement a large model fine-tuning method for collaborative feature perception according to any one of claims 1-6.

10. An electronic device, characterized in that, Comprising: A processor, a memory, and a computer program; wherein, the processor is connected to the memory, the computer program is stored in the memory, and when the electronic device runs, the processor executes the computer program stored in the memory so that the electronic device executes and implements a large model fine-tuning method for collaborative feature perception according to any one of claims 1-6.

Citation Information

Patent Citations

  • Task execution method and device for large model, electronic equipment, storage medium and program product

    CN118519779A

  • Model fine tuning method and device based on fine-grained knowledge perception and medium

    CN119886276A