Bidirectional iterative method for pre-trained model and downstream sequence tasks, device, and medium

Through the two-way iteration method of pre-training model and downstream tasks, the fine-tuning training is used with soft prompt words, which improves the adaptability and performance of the pre-trained model in downstream tasks, and solves the problems of insufficient downstream tasks and insufficient field migration, especially in the scenario of small samples.

WO2025139017A1PCT designated stage expired Publication Date: 2025-07-03TAOBAO CHINA SOFTWARE

Patent Information

Application Number
PCT/CN2024/116725
Authority / Receiving Office
WO · WO
Patent Type
Applications
Current Assignee / Owner
Priority Date
2023-12-25
Filing Date
2024-09-04
Publication Date
2025-07-03

AI Technical Summary

Technical Problem

When existing pre-trained models are implemented in downstream tasks, there are problems of insufficient understanding and insufficient field mobility, especially in small sample scenarios.

Method used

The two-way iteration method of pre-training model and downstream tasks is adopted to improve the adaptability and performance of the pre-trained model through fine-tuning training of historical downstream task feedback links and current downstream task adaptation links.

Benefits of technology

The performance of pre-trained models in downstream tasks is improved, especially in the case of small sample sizes to achieve better model performance, solving the problems of insufficient downstream tasks and insufficient field migration.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN2024116725_03072025_PF_FP_ABST
    Figure CN2024116725_03072025_PF_FP_ABST
Patent Text Reader

Abstract

Embodiments of the present application provide a bidirectional iterative method for a pre-trained model and downstream sequence tasks, a device, and a medium. In the embodiments of the present application, a novel soft-prompt-based fine-tuning training approach is provided, and each round of fine-tuning training comprises both a feedback link from historical downstream tasks to a pre-trained model and an adaption link from the pre-trained model to a current downstream task. For each current round, on the basis of the feedback link, the pre-trained model undergoes a first fine-tuning by using historical downstream tasks that have appeared prior to the current round, so as to enhance the capability of the pre-trained model; and on the basis of the adaption link, the already fine-tuned pre-trained model undergoes a second fine-tuning by using the current downstream task of the current round, so as to train a task model more adapted to the downstream task. Thus, the pre-trained model can be practically applied to downstream tasks, and improved model performance can be achieved especially in few-shot scenarios.
Need to check novelty before this filing date? Find Prior Art

Description

Bidirectional iteration method, device and medium for pre-training model and downstream sequence task

[0001] Cross-references

[0002] This application refers to Chinese patent application No. 2023117970127, filed on December 25, 2023, entitled “Bidirectional Iterative Method, Device and Medium for Pre-trained Model and Downstream Sequence Tasks”, which is incorporated into this application in its entirety by reference. Technical Field

[0003] The present application relates to the field of machine learning technology, and in particular to a method, device, and medium for bidirectional iteration of a pre-training model and downstream sequence tasks. Background Art

[0004] Pre-trained models are trained on large amounts of training data, either supervised or unsupervised, to learn general common knowledge. These models can then be applied to downstream tasks through transfer learning or direct application, reducing the learning burden on the model. For example, popular large language models (LLMs) and general language models (GLMs) are examples of pre-trained models.

[0005] The training process of a pretrained model consists of two phases: pretraining and fine-tuning, also known as the pretrain-then-fine-tuning paradigm. The pretraining phase utilizes a large amount of training data to learn general, context-free features to assist in learning downstream tasks. The fine-tuning phase focuses on learning context-aware feature representations to develop a task model suitable for downstream tasks.

[0006] With the widespread adoption of pre-trained models, their vertical application and downstream migration have become important research areas, focusing on improving the capabilities of pre-trained models when applied in downstream tasks. However, current research, both domestically and internationally, is largely limited to the pre-training and fine-tuning phases, model complexity, and dataset expansion. There is an urgent need for innovative solutions that can improve the performance of pre-trained models and enable their application in downstream tasks.

[0007] Summary of the Invention

[0008] Multiple aspects of the present application provide a method, device, and medium for bidirectional iteration of a pre-trained model and downstream sequence tasks, so as to enable the pre-trained model to be better applied in downstream tasks.

[0009] An embodiment of the present application provides a bidirectional iteration method for a pre-trained model and a downstream sequence task, including: determining an initial pre-trained model for the current round, where the initial pre-trained model for the current round is a target pre-trained model obtained by fine-tuning in the previous round; using training data of historical downstream tasks that appeared before the current round, fine-tuning the initial pre-trained model for the current round based on soft prompt words to obtain a target pre-trained model for the current round; and using training data of the current downstream task that appeared in the current round, fine-tuning the target pre-trained model for the current round based on soft prompt words to obtain a task model corresponding to the current downstream task.

[0010] An embodiment of the present application also provides a downstream task processing method, including: obtaining task data of the downstream task to be processed, the task model corresponding to the downstream task to be processed, and the soft prompt words used for task model reasoning; generating model input data based on the soft prompt words and task data; inputting the model input data into the task model to obtain model output data; wherein the task model is trained according to the pre-training model provided in the embodiment of the present application and the bidirectional iterative method of the downstream sequence task.

[0011] An embodiment of the present application also provides an electronic device comprising: a memory and a processor; the memory is used to store a computer program; the processor is coupled to the memory, and is used to execute the computer program to execute the steps in the bidirectional iteration method between the pre-training model and the downstream sequence task or the downstream task processing method of the embodiment of the present application.

[0012] An embodiment of the present application also provides a computer storage medium storing a computer program. When the computer program is executed by a processor, the processor is enabled to implement the steps in the bidirectional iteration method of the pre-training model and downstream sequence tasks or the downstream task processing method provided in the embodiment of the present application.

[0013] In an embodiment of the present application, a new fine-tuning training method based on soft prompt words is provided. In each round of fine-tuning training, it includes both a feedback link from historical downstream tasks to pre-trained models and an adaptation link from pre-trained models to current downstream tasks. For each current round, the pre-trained model is fine-tuned once based on the feedback link using the historical downstream tasks that occurred before the current round to improve the capabilities of the pre-trained model; based on the adaptation link, the pre-trained model that has been fine-tuned is fine-tuned for a second time using the current downstream tasks of the current round to train a task model that is more adapted to downstream tasks. In an embodiment of the present application, the capabilities of the pre-trained model are improved by feeding back the downstream tasks, so that the pre-trained model has better performance on vertical downstream tasks, presenting a two-way iterative relationship between the pre-trained model and the downstream tasks, solving the problems of insufficient understanding of downstream tasks and insufficient domain transferability of the pre-trained model in vertical downstream tasks, so that the pre-trained model can be better applied in downstream tasks, especially in scenarios with few samples, to achieve better model performance. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] The drawings described herein are used to provide a further understanding of the present application and constitute a part of the present application. The illustrative embodiments of the present application and their descriptions are used to explain the present application and do not constitute an improper limitation on the present application. In the drawings:

[0015] FIG1 is a flowchart of a bidirectional iteration method for a pre-training model and downstream sequence tasks provided by an embodiment of the present application;

[0016] FIG2a is a schematic diagram showing the relationship between soft prompt words and model accuracy under a 16-shot sample using a 5T model as an example provided in an embodiment of the present application;

[0017] FIG2 b is a schematic diagram showing the relationship between soft prompt words and model accuracy under 32-shot samples using a 5T model as an example provided in an embodiment of the present application;

[0018] FIG2c is a schematic diagram showing the relationship between soft prompt words and model accuracy under a 100-shot sample using a 5T model as an example provided in an embodiment of the present application;

[0019] FIG3a is a schematic diagram of an exemplary process of bidirectional iteration of a pre-training model and downstream tasks provided in an embodiment of the present application;

[0020] FIG3 b is a schematic diagram of an exemplary fine-tuning of a feedback link and an adaptation link in each round provided by an embodiment of the present application;

[0021] FIG4 is a flowchart of a downstream task processing method provided in an embodiment of the present application;

[0022] FIG5 is a diagram of an exemplary application scenario provided by an embodiment of the present application;

[0023] FIG6 is a schematic structural diagram of a bidirectional iteration device provided in an embodiment of the present application;

[0024] FIG7 is a schematic structural diagram of an electronic device provided in an embodiment of the present application. DETAILED DESCRIPTION

[0025] To make the purpose, technical solutions, and advantages of this application more clear, the technical solutions of this application will be clearly and completely described below in conjunction with the specific embodiments of this application and the corresponding drawings. Obviously, the embodiments described are only part of the embodiments of this application, not all of them. Based on the embodiments in this application, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of this application.

[0026] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, stored data, displayed data, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions, and provide corresponding operation entrances for users to choose to authorize or refuse.

[0027] With the rapid development of pre-trained models, the number of model layers, the number of model parameters, and the amount of training data are constantly increasing, enabling new capabilities such as progressive understanding and human-like reasoning. This gradual improvement in model capabilities allows pre-trained models to incorporate more general knowledge. However, this also leads to a sharp increase in the model's consumption of computing resources and a continuous increase in the demand for data. Consequently, zero-shot learning has become a key metric for evaluating the capabilities of pre-trained models. This is partly because the computational cost of fine-tuning very large pre-trained models is often prohibitive, and partly because very large pre-trained models can achieve similar results to fine-tuning in zero-shot or few-shot learning scenarios.

[0028] In order to solve the problem of improving the capabilities of pre-trained models when they are applied in downstream tasks, relevant research on pre-training has been conducted both at home and abroad. Some studies focus on the two-stage process of fine-tuning after pre-training, and aim to propose better and more efficient pre-training models; some studies improve the capabilities of pre-training models by increasing model complexity, expanding data sets, etc., in order to obtain more general pre-training models, and directly apply them to downstream tasks in zero-shot learning scenarios.

[0029] Regardless of which studies, none of them jointly modeled the pre-training task and the downstream task, ignored the impact of the downstream task on the pre-training model, and also ignored the performance improvement that the training data of the downstream task may bring to the pre-training model. In addition, since the differences between the pre-training task and the downstream task are not taken into account, it is difficult to achieve good few-sample or zero-sample performance when pre-training with continuous high-quality labeled samples. In an embodiment of the present application, a downstream task is introduced in the fine-tuning process of the pre-training model to achieve two-way iteration of the pre-training task and the downstream task, and the pre-training model is fine-tuned using historical downstream tasks. The ability of the pre-training model is improved by feeding back the downstream tasks, and then the pre-training model after feeding back is used to fine-tune the current downstream task to obtain the task model of the current downstream task, so that the pre-training model has better performance in vertical downstream tasks, presenting the relationship between the two-way iteration of the pre-training model and the sequential downstream task, solving the problems of insufficient understanding of downstream tasks and insufficient domain transferability of pre-training models in vertical downstream tasks, so that the pre-training model can be better applied in downstream tasks, especially in few-sample scenarios. Better model performance can be achieved.

[0030] The following describes in detail the technical solutions provided by various embodiments of the present application in conjunction with the accompanying drawings.

[0031] FIG1 is a flowchart of a bidirectional iteration method for a pre-training model and downstream sequence tasks provided by an embodiment of the present application. Referring to FIG1 , the method may include the following steps:

[0032] 101. Determine the initial pre-training model for the current round. The initial pre-training model for the current round is the target pre-training model obtained by fine-tuning in the previous round.

[0033] 102. Using the training data of historical downstream tasks that occurred before the current round, fine-tune the initial pre-trained model of the current round based on soft prompt words to obtain the target pre-trained model of the current round;

[0034] 103. Using the training data of the current downstream task that appears in the current round, fine-tune the target pre-trained model of the current round based on the soft prompt words to obtain the task model corresponding to the current downstream task.

[0035] Among them, historical downstream tasks and current downstream tasks appear in different rounds according to different task triggering times.

[0036] The bidirectional iterative method for pre-training models and downstream sequence tasks provided in this embodiment involves a post-pre-training fine-tuning paradigm. This paradigm first trains a pre-trained model with strong generalization capabilities on a large dataset, then fine-tunes it on a specific downstream task to obtain a task model that adapts to different scenarios and requirements.

[0037] In an embodiment of the present application, the original pre-training model can be obtained in advance, and there is no limitation on the generalization training process of the original pre-training model. For example, the training data generated by some basic tasks can be used to adopt an unsupervised learning method to construct the original pre-training model by automatically learning the statistical laws and semantic information of the training data. Among them, the process of pre-training based on the training data of the basic task will be different depending on the function of the pre-training model. Taking the pre-training model as a large language model as an example, the process of pre-training based on the training data of the basic task includes predicting the next word and mask filling based on the previous word. In an embodiment of the present application, the original pre-training model can learn richer general knowledge and related representations through the training process of the basic task, and has higher generalization ability and robustness.

[0038] In this embodiment, the model obtained by fine-tuning the pre-trained model using the training data of a specific downstream task can be called a task model. The downstream task refers to a task that is specifically applied to actual problems based on the pre-trained model, which is a more specific application form. For example, the downstream task is a downstream task with commodity category prediction as a task goal, a task with commodity feature extraction as a task goal, or a downstream task with commodity content understanding as a task goal. The task goals of these different downstream tasks are different, presenting highly diversified goals. The data distribution of different downstream tasks has its own characteristics, that is, it presents a heterogeneous data distribution. In other words, the downstream tasks have highly diversified goals and heterogeneous data distributions. The task model obtained by fine-tuning the pre-trained model for downstream tasks performs better on downstream tasks, and the model performance and adaptability are improved.

[0039] In this embodiment, the downstream tasks have a timing characteristic, and each downstream task appears in a sequential pattern instead of appearing all at once. The multiple downstream tasks that appear in chronological order are recorded as downstream sequence tasks (also referred to as sequential downstream tasks). In other words, the application requirements of the pre-trained model are discovered from different scenarios over time. In this embodiment, the problem of bidirectional knowledge transfer between the pre-trained model and the downstream tasks that appear in sequence is modeled, which is referred to as the bidirectional iterative relationship between the pre-trained model and the downstream tasks, and the bidirectional iterative relationship between the pre-trained model and the downstream tasks is introduced into the fine-tuning stage of the pre-trained model to obtain a new fine-tuning method. In this new fine-tuning method, a feedback link (Feedback) from the downstream task to the pre-trained model and an adaptation link (Adaption) from the pre-trained model to the downstream task are included. Among them, the feedback link refers to the link that uses downstream tasks to fine-tune the pre-trained model to obtain a new pre-trained model. The fine-tuning on this link is mainly to update the pre-trained model so that the pre-trained model can accumulate common knowledge shared between different downstream tasks; the adaptation link refers to the link that fine-tunes the pre-trained model for downstream tasks to obtain a task model suitable for downstream tasks. The main purpose of fine-tuning on this link is to allow the pre-trained model to learn the specific knowledge of downstream tasks so that the pre-trained model can perform better on downstream tasks.

[0040] Among them, the bidirectional iteration of the pre-training model and the downstream task can be understood as a mutually reinforcing relationship between the pre-training model and the downstream task. In this embodiment, the downstream tasks appear in a sequential pattern instead of appearing all at once, and the application requirements of the pre-training model vary over time. For a specific downstream task, unlike the traditional solution that is independent of other downstream tasks, during the fine-tuning of the pre-training model, the specific downstream task will participate in the adaptation link and feedback link in the fine-tuning process respectively. Specifically, on the one hand, it can participate in the adaptation link from the pre-training model to the downstream task, that is, it can be used to fine-tune the pre-training model to obtain the task model of the downstream task; on the other hand, it can participate in the feedback link from the downstream task to the pre-training model, that is, it can be used to update the pre-training model, assisting the pre-training model to learn the common knowledge shared between downstream tasks, realizing the feedback of downstream tasks to the pre-training tasks, and then improving the ability of the pre-training model. The pre-training model after feedback will have better performance on subsequent downstream tasks, presenting a bidirectional iterative relationship, solving the problems of insufficient understanding of downstream tasks and insufficient domain transferability of pre-training models in vertical downstream tasks.

[0041] In this embodiment, to model the temporal characteristics of downstream tasks, each downstream task is divided into different rounds in the fine-tuning training process according to the order in which the downstream tasks arrive. In this way, as downstream tasks continue to arrive, the pre-trained model can be continuously updated and iterated, and the common knowledge between downstream tasks learned by the pre-trained model is continuously enriched, so that the pre-trained tasks have better performance on downstream tasks. In each round, one or more downstream tasks are included.

[0042] In this embodiment, one or more target application scenarios using the method provided in the embodiment of the present application can be predetermined. For example, from the e-commerce dimension, the application scenarios involved in various e-commerce applications (such as second-hand trading applications, comprehensive applications, overseas shopping applications, etc.) provided by a certain e-commerce company can be used as the target application scenarios in this embodiment, or, from the application type as an example, the application scenarios involved in a certain type of application (such as instant messaging applications, e-commerce applications or game applications) can be used as the target application scenarios in the basic embodiment; of course, the target application scenarios in this embodiment can also be determined from other dimensions. For the target application scenario, as time goes by, downstream tasks that depend on the pre-training model can be continuously collected. For example, for e-commerce application scenarios, downstream tasks such as product feature extraction tasks, user portrait generation tasks, product category prediction tasks, product image generation tasks, and 3D digital human generation tasks may appear.

[0043] In this embodiment, the method used to divide downstream tasks into different rounds is not limited. For example, the number of incoming downstream tasks can be continuously accumulated. When the number of incoming downstream tasks reaches a specified number or is within a specified number range, these downstream tasks are divided into a round. Then, as downstream tasks continue to arrive, these downstream tasks are accumulated again until the specified number is reached again or is within a specified number range, and then a new round is divided, and so on. For another example, downstream tasks can be divided into rounds according to set time intervals. Assuming the set time interval is 1 hour, as time passes, downstream tasks arriving within each hour are divided into a round. For another example, downstream tasks can be divided into rounds according to event triggers. For example, whenever a set trigger event occurs, downstream tasks arriving between two adjacent set trigger events are divided into a round. The trigger event can be determined based on application requirements. For another example, downstream tasks can be divided into rounds according to the division instructions of the model trainer. For example, whenever a division instruction is received, downstream tasks arriving between two adjacent split instructions are divided into the same round.

[0044] Regardless of the method used to divide downstream tasks into rounds, downstream tasks appearing in the same round have the same or similar task trigger times, while downstream tasks appearing in different rounds have different task trigger times, showing significant variability. The task trigger time refers to the time when the downstream task appears or arrives. The earlier the task trigger time, the higher the order of the round it belongs to. For example, the downstream tasks include the product category prediction task at 9:00 AM on November 13, 2023, the product feature extraction task at 9:00 AM on November 13, 2023, the product content understanding task at 10:00 AM on November 13, 2023, the e-commerce query understanding task at 10:00 AM on November 13, 2023, the similar product recommendation task at 11:00 AM on November 13, 2023, and the product scene graph generation task at 11:00 AM on November 13, 2023. The rounds of fine-tuning training are chronologically as follows: Round 1, Round 2, and Round 3. The downstream tasks in Round 1 include the product category prediction task at 9:00 AM on November 13, 2023, and the product feature extraction task at 9:00 AM on November 13, 2023. The downstream tasks in Round 2 include the product content understanding task at 10:00 AM on November 13, 2023, and the e-commerce query understanding task at 10:00 AM on November 13, 2023. The downstream tasks in Round 3 include the similar product recommendation task at 11:00 AM on November 13, 2023, and the product scene graph generation task at 11:00 AM on November 13, 2023.

[0045] In this embodiment, the pre-trained model is subjected to multiple rounds of fine-tuning training. In each round, the pre-trained model is first fine-tuned on the feedback link, and then the pre-trained model fine-tuned on the feedback link is fine-tuned on the adaptation link. For each round of fine-tuning training, on the one hand, the historical downstream tasks that appeared before the current round are determined. These historical downstream tasks are downstream tasks that appeared in historical rounds. On the other hand, the current downstream tasks that appear in the current round are determined. Among them, the historical downstream tasks and the current downstream tasks appear in different rounds according to different task trigger times. For the convenience of description and distinction, the downstream tasks that appear in the current round are called current downstream tasks. The task trigger time of the current downstream tasks is within the time range corresponding to the current round. The downstream tasks that appear before the current round are called historical downstream tasks. The historical downstream tasks are downstream tasks that appeared in the corresponding historical rounds before the current round. Alternatively, the historical downstream tasks that appeared before the current round can be understood as source tasks (Source Task), and the current downstream tasks that appear in the current round can be understood as target tasks (Target Task). For example, relative to Round 1, the product category prediction task at 9:00 AM on November 13, 2023 and the product feature extraction task at 9:00 AM on November 13, 2023 are the current downstream tasks of Round 1. However, as time goes by, when Rounds 2 and 3 appear, relative to Rounds 2 and 3, the downstream tasks in Round 1 will become historical downstream tasks that appeared before Rounds 2 and 3.

[0046] In this embodiment, the two-way knowledge transfer between the pre-training model and the downstream task is modeled, including a feedback link from the downstream task to the pre-training model and an adaptation link from the pre-training model to the downstream task. Wherein, when the downstream task is divided into different rounds, a feedback link and an adaptation link are included in each round, as shown in Figure 2a. Specifically, the feedback link refers to fine-tuning the pre-training model using the historical downstream tasks that appeared before the current round, so that the pre-training model learns the common knowledge shared between these historical downstream tasks; the adaptation link refers to fine-tuning the pre-training model (also referred to as the feedback model) that has been fine-tuned on the feedback link using the current downstream task that appears in the current round, so that the pre-training model learns the specific knowledge of the current downstream task, and then obtains the task model of the current downstream task. In this way, for each current round, the general capabilities of the pre-training model are improved through the feedback link, and the task model that is more adapted to the downstream task is trained through the adaptation link.

[0047] In this embodiment, when using the feedback link to improve the general capabilities of the pre-trained model, for the convenience of description and distinction, the pre-trained model in each round is divided into an initial pre-trained model and a target pre-trained model; wherein, the initial pre-trained model in each round is the pre-trained model that needs to be fine-tuned through the feedback link in this round, and is also the target pre-trained model in the previous round; the target pre-trained model in each round is the pre-trained model obtained by fine-tuning the feedback link in this round, and is also the initial pre-trained model in the next round. For rounds other than the first, the initial pre-trained model in the round is the target pre-trained model in the previous round; for rounds other than the last, the target pre-trained model in the round is also the initial pre-trained model in the next round, as shown in Figure 2a.

[0048] It is explained here that for the first round, a universal pre-trained model trained with a large data set is used as the target pre-trained model obtained by fine-tuning in the 0th round, that is, as the initial pre-trained model for the first round. Accordingly, for the first round, several tasks can be selected from the tasks for training the universal pre-trained model as historical downstream tasks that occurred before the first round. In this embodiment, the initial pre-trained model of each round is different, and the pre-trained model is continuously fine-tuned through a feedback link.

[0049] Taking the current round as an example, the process of fine-tuning the pre-trained model using the feedback link is explained. Specifically, for the current round, the initial pre-trained model of the current round is determined; using the training data of the historical downstream tasks that appeared before the current round, the initial pre-trained model of the current round is fine-tuned based on soft prompts to obtain the target pre-trained model of the current round. Fine-tuning based on soft prompts belongs to the prompt tuning method. Soft prompts can be continuously optimized during the fine-tuning process. Soft prompts usually contain an embedded vector or a string of digital data. By adding the embedded vector or digital data to the beginning of the model input as prompt information to guide the model's output results, as the model is continuously fine-tuned, the soft prompts can be continuously optimized, and more accurate knowledge can be learned from the model. In an embodiment of the present application, soft prompt words are used to reflect the proprietary knowledge of downstream tasks, such as keywords, context and other information related to the downstream tasks, to guide the pre-trained model to better understand and process the model input and give output that is more in line with the needs of downstream tasks; and as the model is continuously fine-tuned, the soft prompt words can be continuously optimized, which can more accurately reflect the proprietary knowledge of downstream tasks, continuously guide the output of the pre-trained model, make the pre-trained model more suitable for downstream tasks, and improve the performance of the pre-trained model on downstream tasks.

[0050] In an embodiment of the present application, the pre-training model is fine-tuned using historical downstream tasks, with the goal of accumulating general knowledge shared between different downstream tasks. However, not every downstream task's knowledge is helpful to other downstream tasks. In order to separate task-specific knowledge from general knowledge, in an embodiment of the present application, soft prompt words are introduced for each downstream task, and a multi-task feedback algorithm based on learnable prompt words is proposed to utilize the training data of historical downstream tasks that appeared before the current round to fine-tune the initial pre-training model of the current round based on the soft prompt words to obtain the target pre-training model of the current round. In an embodiment of the present application, in the process of fine-tuning the initial pre-training model of the current round based on the soft prompt words, the order between the historical downstream tasks is not limited. The order between these historical downstream tasks can be randomly mixed or sorted in the order of appearance.

[0051] In this embodiment, each downstream task, whether historical or current, has training data. This training data is labeled and includes model input data and labeling results. For example, the training data for a historical downstream task includes model input data and labeling results, and accordingly, the training data for the current downstream task also includes model input data and labeling results. It should be noted that the types and objectives of downstream tasks that appear over time vary. Depending on the type of downstream task, the model input data included in the training data for each downstream task will also vary. For example, for a product category prediction task, the corresponding model input data can be multimodal data such as product text descriptions, product details, or product images, and the labeling results represent product categories such as clothing, alcohol, or beverages. For a product content understanding task, the corresponding model input data can be multimodal data such as product text descriptions, product details, or product images, and the labeling results include product knowledge such as the product category system, attribute system, and product description. Based on this, the training data of historical downstream tasks that occurred before the current round can be used to fine-tune the initial pre-trained model of the current round based on soft prompt words to obtain the target pre-trained model for the current round. This embodiment does not limit the manner in which the initial pre-training model of the current round is fine-tuned based on the soft prompt words.

[0052] Further optionally, the training data of historical downstream tasks that occurred before the current round are used to fine-tune the initial pre-trained model of the current round based on soft prompt words to obtain the target pre-trained model of the current round. The implementation method is as follows: generate first soft prompt words corresponding to each historical downstream task that occurred before the current round; embed the first soft prompt words into the training data of the corresponding historical downstream task to obtain new training data of the historical downstream task; and fine-tune the initial pre-trained model of the current round according to the new training data of the historical downstream task to obtain the target pre-trained model of the current round.

[0053] Specifically, for any historical downstream task that occurred before the current round, a soft prompt word is generated for the historical downstream task. Here, the soft prompt word generated for any historical downstream task that occurred before the current round is referred to as the first soft prompt word. This embodiment does not limit the method for generating the first soft prompt word. The following are several optional generation methods:

[0054] Method 1: For any historical downstream task, randomly initialize the first soft prompt word for it.

[0055] Method 2: For any historical downstream task, a first soft prompt word is generated for it based on the fourth soft prompt word used by the task model inference corresponding to any historical downstream task.

[0056] Specifically, the fourth soft prompt is used by the task model for inference for any historical downstream task. Compared to the first soft prompt obtained in Method 1, the first soft prompt obtained in Method 2 better reflects the specific knowledge of the historical downstream task, facilitating better and faster model convergence, thereby improving the performance of the target pre-trained model obtained through fine-tuning.

[0057] Method 3: For any historical downstream task, a first soft prompt word is generated for it according to the second soft prompt word corresponding to the historical downstream task before the historical downstream task.

[0058] Specifically, for any historical downstream task, its second soft prompt word is a soft prompt word obtained by optimizing the first soft prompt word corresponding to the previous historical downstream task. The historical downstream task before any historical downstream task refers to the historical downstream task that appeared in the round before the round in which the historical downstream task appeared, and the task trigger time of the historical downstream task before any historical downstream task is earlier than the task trigger time of any historical downstream task. Various statistical analyses such as weighted summation, averaging, and accumulation are performed on the second soft prompt word corresponding to the historical downstream task before any historical downstream task to obtain the first soft prompt word of the historical downstream task. Compared with the first soft prompt word obtained by method 2, the first soft prompt word obtained by method 3 can reflect the specific knowledge of the previous historical downstream task, which is conducive to better and faster convergence of the model, thereby improving the performance of the target pre-training model obtained based on this fine-tuning.

[0059] In this embodiment, after obtaining the first soft prompt word corresponding to each of the historical downstream tasks that appeared before the current round, the first soft prompt word is embedded in the training data of the corresponding historical downstream task to obtain new training data for the historical downstream task. Specifically, the training data of the historical downstream task that appeared before the current round includes model input data and annotation results. The first soft prompt word is embedded in the corresponding model input data to obtain new model input data to update the training data of the historical downstream task that appeared before the current round. The new training data of the historical downstream task includes the annotation results and the new model input data that has been embedded with the first soft prompt word. For example, the first soft prompt word (specifically a vector) of the historical downstream task can be embedded in the front part of the model input data (specifically a vector) corresponding to the historical downstream task to obtain new model input data (specifically a vector). The new model input data used in the fine-tuning process actually includes the original model input data acting on the main network part of the model and the task-specific soft prompt word part.

[0060] Optionally, in order to improve model performance, the initial pre-trained model of the current round is fine-tuned according to the new training data of the historical downstream task to obtain the target pre-trained model of the current round. The implementation method is as follows: the model input data in the new training data of the historical downstream task is input into the initial pre-trained model of the current round to obtain the prediction result of the model input data of the historical downstream task; the first loss value is calculated according to the prediction result of the model input data of the historical downstream task and the corresponding labeling result in the new training data of the historical downstream task; the model parameters of the initial pre-trained model of the current round are adjusted according to the first loss value to obtain the target pre-trained model of the current round; and the first soft prompt word corresponding to the historical downstream task is optimized according to the first loss value to obtain the second soft prompt word.

[0061] Specifically, the first loss value reflects the error information between the prediction result of the model input data of the historical downstream task and the corresponding annotation result in the new training data of the historical downstream task. The loss function used to calculate the first loss value includes, but is not limited to: log loss function, cross-entropy loss function and Focal loss function for solving data imbalance problems. With the goal of minimizing the first loss value, the model parameters of the initial pre-trained model of the current round are updated through back propagation and gradient descent to obtain the target pre-trained model of the current round. In practical applications, a target loss value can be set as needed. If the first loss value is greater than the target loss value, it means that the first loss value has not yet been minimized, and it is necessary to continue to adjust the model parameters and continue model training. If the first loss value is less than or equal to the target loss value, it means that the first loss value has met the end condition of model training, and the fine-tuning training of the pre-trained model on the feedback link in the current round is completed.

[0062] In this embodiment, during the fine-tuning of the initial pre-trained model on the feedback link, in addition to adjusting the model parameters of the main network in the initial pre-trained model, the first soft prompt words corresponding to each historical downstream task participating in the model training can also be optimized according to the first loss value to obtain the second soft prompt words corresponding to each historical downstream task. In this process, the first soft prompt words can be regarded as an extension of the model parameters to be optimized, and the gradient is calculated in the same way as the main network model parameters, and updates such as backpropagation and gradient descent are performed. With the optimization and learning of the soft prompt words, the misleading effect of downstream tasks on the model parameters can be reduced, which is conducive to further improving the performance of the target pre-trained model obtained.

[0063] In actual applications, the first soft prompt words corresponding to each historical downstream task before the current round can be optimized, or the first soft prompt words corresponding to some historical downstream tasks before the current round can be optimized, without limitation. Further optionally, when optimizing the first soft prompt words corresponding to the historical downstream tasks according to the first loss value to obtain the second soft prompt words, the first soft prompt words corresponding to the historical downstream tasks that appeared in at least one of the most recent historical rounds can be optimized according to the first loss value to obtain the second soft prompt words, and the first soft prompt words corresponding to the historical downstream tasks that appeared in other historical rounds before the current round can be frozen, that is, the first soft prompt words corresponding to the historical downstream tasks that appeared in other historical rounds before the current round can be kept fixed. Among them, at least one historical round can be the previous historical round or the two most recent historical rounds, and the specific selection can be based on application requirements and is not limited to this. For example, if the current round is the 5th round, the soft prompt words corresponding to the historical downstream tasks that appeared in the 4th round are optimized, while the soft prompt words corresponding to the historical downstream tasks that appeared in the 1st, 2nd, and 3rd rounds are kept fixed. In this embodiment, the soft prompt words of recent historical downstream tasks are updated, and the soft prompt words of historical downstream tasks in the distant past are frozen. This can reduce the overfitting of the model parameters of the main network for downstream tasks to a certain extent; and the continuous learning of the model parameters of the initial pre-trained model through the training data of historical downstream tasks can enable the model to achieve stronger generalization ability and better performance in few-sample scenarios.

[0064] In this embodiment, for each round, after fine-tuning the initial pre-training model based on soft prompt words using historical downstream tasks on the feedback link, the target pre-training task of the round can be obtained, and then the target pre-training task can be fine-tuned on the adaptation link for the current downstream task that appears in the round to obtain a task model that is adapted to the current downstream task. Wherein, when fine-tuning the target pre-training model on the adaptation link to obtain a task model that is adapted to the current downstream task, the training data of the current downstream task that appears in the current round is used to fine-tune the target pre-training model of the current round based on soft prompt words to obtain the task model corresponding to the current downstream task. It is worth noting that if multiple current downstream tasks appear in the current round, the target pre-training task can be fine-tuned on the adaptation link for each current downstream task to obtain a task model that is adapted to each current downstream task, specifically: using the training data of each current downstream task that appears in the current round, the target pre-training model of the current round is fine-tuned based on soft prompt words to obtain a task model corresponding to each current downstream task. For example, the product category prediction task at 9:00 AM on November 13, 2023, and the product feature extraction task at 9:00 AM on November 13, 2023, are the current downstream tasks of round 1. The task models output by round 1 include the product category prediction model and the product feature extraction model.

[0065] This embodiment does not limit the manner in which the target pre-trained model of the current round is fine-tuned based on soft prompt words. Further, optionally, to improve model performance, the target pre-trained model of the current round is fine-tuned based on soft prompt words using the training data of the current downstream task that appears in the current round to obtain a task model corresponding to the current downstream task. This is achieved by: generating a third soft prompt word corresponding to the current downstream task; embedding the third soft prompt word into the training data of the current downstream task to obtain new training data for the current downstream task; and fine-tuning the target pre-trained model of the current round based on the new training data of the current downstream task to obtain a task model corresponding to the current downstream task and a fourth soft prompt word used for task model inference.

[0066] Specifically, for any current downstream task that appears in the current round, a soft prompt word for the current downstream task is generated. Here, the soft prompt word generated for any current downstream task that appears in the current round is referred to as the third soft prompt word. This embodiment does not limit the method for generating the third soft prompt word. The following are several optional generation methods:

[0067] Method 1: Randomly initialize the third soft prompt word for the current downstream task.

[0068] Method 2: Generate a third soft prompt for the current downstream task based on the first soft prompt for the downstream task that appeared before the current round. Compared to the third soft prompt obtained in Method 1, the third soft prompt obtained in Method 2 can better improve model performance.

[0069] Optionally, the first soft prompt words corresponding to historical downstream tasks can be weighted summed or averaged to obtain the third soft prompt word corresponding to the current downstream task; or, one of the first soft prompt words corresponding to historical downstream tasks can be randomly selected, or one with better quality can be selected, as the third soft prompt word corresponding to the current downstream task.

[0070] Method 3: Generate a third soft prompt for the current downstream task based on the second soft prompt for the downstream task before the current round. Compared to the third soft prompt obtained in Method 2, the third soft prompt obtained in Method 3 can better improve model performance.

[0071] Optionally, the second soft prompt words corresponding to the historical downstream tasks can be weighted summed or averaged to obtain the third soft prompt word corresponding to the current downstream task; or, one of the second soft prompt words corresponding to the historical downstream tasks can be randomly selected, or one with better quality can be selected, as the third soft prompt word corresponding to the current downstream task.

[0072] In this embodiment, after obtaining the third soft prompt word corresponding to each of the current downstream tasks appearing in the current round, the third soft prompt word is embedded in the training data corresponding to the current downstream task to obtain new training data for the current downstream task. Specifically, the training data for the current downstream task appearing in the current round includes model input data and annotation results. The third soft prompt word is embedded in the corresponding model input data to obtain new model input data to update the training data for the current downstream task appearing in the current round. The new training data for the current downstream task includes the annotation results and the new model input data in which the third soft prompt word has been embedded. When training the task model of the current downstream task, the third soft prompt word can be optimized. The soft prompt word used for task model inference obtained by optimizing the third soft prompt word is referred to as the fourth soft prompt word.

[0073] Optionally, in order to improve model performance, when training the task model of the current downstream task, the model parameters of the target pre-trained model of the current round can be frozen, and the third soft prompt word of the current downstream task can be optimized. Thus, as an example, the target pre-trained model of the current round is fine-tuned according to the new training data of the current downstream task to obtain the task model corresponding to the current downstream task and the fourth soft prompt word used for task model reasoning. The implementation method is as follows: the model input data in the new training data of the current downstream task is input into the target pre-trained model of the current round to obtain the prediction result of the model input data of the current downstream task; the second loss value is calculated based on the prediction result of the model input data of the current downstream task and the corresponding annotation result in the new training data of the current downstream task; the model parameters of the target pre-trained model of the current round are frozen to make the target pre-trained model the task model of the current downstream task, and according to the second loss value, the third soft prompt word corresponding to the current downstream task is optimized to obtain the fourth soft prompt word used for task model reasoning.

[0074] Specifically, the second loss value reflects the error information between the prediction result of the model input data of the current downstream task and the corresponding annotation result in the new training data of the current downstream task. The loss function used to calculate the second loss value includes, but is not limited to: log loss function, cross entropy loss function and Focalloss loss function for solving data imbalance problems. With the goal of minimizing the second loss value, the third soft prompt word corresponding to the current downstream task is optimized. In practical applications, a target loss value can be set as needed. If the second loss value is greater than the target loss value, it means that the second loss value has not yet been minimized, and the model training continues. If the second loss value is less than or equal to the target loss value, it means that the second loss value has been minimized, and the fine-tuning training of the feedback link of the current round is completed. The introduction of learnable soft prompt words reduces the misleading effect of downstream tasks on model parameters.

[0075] Optionally, to improve model performance, when training the task model of the current downstream task, the model parameters of the target pre-trained model of the current round can be adjusted, and the third soft prompt word of the current downstream task can be optimized. Thus, as another example, the target pre-trained model of the current round is fine-tuned based on the new training data of the current downstream task to obtain the task model corresponding to the current downstream task and the fourth soft prompt word used for task model reasoning. This is achieved by: inputting the model input data in the new training data of the current downstream task into the target pre-trained model of the current round to obtain the prediction result of the model input data of the current downstream task; calculating a second loss value based on the prediction result of the model input data of the current downstream task and the corresponding annotation result in the new training data of the current downstream task; and optimizing the model parameters of the target pre-trained model of the current round and the third soft prompt word corresponding to the current downstream task based on the second loss value to obtain the task model corresponding to the current downstream task and the fourth soft prompt word used for task model reasoning.

[0076] Specifically, the second loss value reflects the error information between the prediction result of the model input data of the current downstream task and the corresponding annotation result in the new training data of the current downstream task. The loss function used to calculate the second loss value includes, but is not limited to: log loss function, cross entropy loss function and Focal loss function for solving data imbalance problems. With the goal of minimizing the second loss value, the model parameters of the target pre-trained model of the current round are updated through back propagation and gradient descent until the second loss value is minimized. After the second loss value is minimized, the current target pre-trained model is used as the task model corresponding to the current downstream task. In actual applications, a target loss value can be set as needed. If the second loss value is greater than the target loss value, it means that the second loss value has not yet been minimized, and it is necessary to continue to adjust the model parameters and continue model training. If the second loss value is less than or equal to the target loss value, it means that the second loss value has been minimized, and the fine-tuning training of the current round of adaptation link is completed.

[0077] In this embodiment, the third soft prompt word corresponding to the current downstream task is optimized with the goal of minimizing the second loss value. The introduction of learnable soft prompt words mitigates the misleading effect of the downstream task on the model parameters. It should be noted that minimizing the second loss value is only one example of an optimization termination condition and is not limited to this. For example, the goal can also be to reduce the second loss value to within a set range.

[0078] In this embodiment, the corresponding task models vary depending on the downstream task. For example, if the downstream task is a product category prediction task, the task model is a product category prediction model; another example, if the downstream task is a product feature extraction task, the task model is a product feature extraction model; another example, if the downstream task is a product content understanding task, the task model is a product content understanding model; but the present invention is not limited to the above examples.

[0079] The technical solution provided by the embodiment of the present application provides a new fine-tuning training method based on soft prompt words. In each round of fine-tuning training, it includes both a feedback link from historical downstream tasks to pre-trained models and an adaptation link from pre-trained models to current downstream tasks. For each current round, the pre-trained model is fine-tuned using the historical downstream tasks that occurred before the current round based on the feedback link to improve the capabilities of the pre-trained model; based on the adaptation link, the pre-trained model that has been fine-tuned is fine-tuned using the current downstream tasks of the current round to train a task model that is more adapted to downstream tasks. The downstream tasks are fed back to improve the capabilities of the pre-trained model, and the pre-trained model after the feedback has better performance on vertical downstream tasks, presenting a two-way iterative relationship between the pre-trained model and the downstream sequence tasks, solving the problems of insufficient understanding of downstream tasks and insufficient domain transferability of pre-trained models in vertical downstream tasks. The pre-trained model can be better applied in downstream tasks, especially in scenarios with few samples, to achieve better model performance.

[0080] To demonstrate the performance of the model fine-tuning method based on bidirectional iteration of pre-trained models and downstream tasks provided by the embodiments of the present application, extensive experiments were conducted on multiple discriminant models and generative models using the model fine-tuning method provided by the embodiments of the present application and some traditional pre-training methods. The experimental results compared the model fine-tuning method provided by the embodiments of the present application with the traditional pre-training methods in terms of model fine-tuning effect on the target task and model fine-tuning efficiency, fully demonstrating the effectiveness and necessity of the model fine-tuning method based on bidirectional iteration of pre-trained models and downstream tasks provided by the embodiments of the present application. The relevant experimental data and comparative analysis are as follows:

[0081] (1) Based on the BERT (Bidirectional Encoder Representation from Transformers) model architecture and the Roberta (A Robustly Optimized BERT) model architecture, model parameter fine-tuning or prompt word fine-tuning was selected for various types of downstream tasks, and the performance of the pre-trained model obtained by the model fine-tuning method provided in the embodiment of the present application and the traditional pre-trained model was compared, as shown in Table 1.

[0082] The BERT model is a pre-trained language representation model based on a bidirectional encoder representation of the Transformer model. The Transformer model uses an attention mechanism to speed up model training. The RoBERTa model is an improved version of the BERT model.

[0083] In Table 1, FT indicates fine-tuning only the model parameters of the pre-trained model during model training; PT indicates fine-tuning only the soft prompt words during model training; BiKTFT indicates fine-tuning the model parameters during model training using a bidirectional iterative method between the pre-trained model and the downstream task; and FTBiKTPT indicates fine-tuning the soft prompt words during model training using a bidirectional iterative method between the pre-trained model and the downstream task. BERT-base(FT) represents a traditional BERT model trained by fine-tuning model parameters, BERT-base(FTPT) represents a traditional BERT model trained by fine-tuning both model parameters and soft prompt words, and BERT-base(BiKTPT) represents a BERT model obtained by fine-tuning model parameters in the method provided in the embodiments of this application. BERT-base(PT) represents the traditional BERT model trained by fine-tuning the soft prompt words. BERT-base(PTMT) represents the traditional BERT model trained by fine-tuning the soft prompt words. The difference from BERT-base(PT) is that the initialization of the soft prompt words is different. The soft prompt words of BERT-base(PTMT) are obtained by multi-task learning on the source task, and the soft prompt words of BERT-base(PT) are obtained by random initialization. BERT-base(BiKTPT) represents the BERT model obtained by fine-tuning the soft prompt words in the method provided in the embodiment of the present application. RoBERTa-base(PT) represents the traditional RoBERTa model trained by fine-tuning the soft prompt words. RoBERTa-base(PTMT) represents the traditional RoBERTa model trained by fine-tuning the soft prompt words. The difference between it and RoBERTa-base(PT) is that the initialization of the soft prompt words is different. The soft prompt words of RoBERTa-base(PTMT) are obtained by multi-task learning on the source task, and the soft prompt words of RoBERTa-base(PT) are obtained by random initialization. RoBERTa-base(BiKTPT) represents the RoBERTa model obtained by fine-tuning the soft prompt words in the method provided in the embodiment of the present application.

[0084] BoolQ, CB, COPA, MRC, RTE, WiC, SNLI, PAWS, and IMDB in Table 1 represent different types of downstream tasks. BoolQ is a question-answering task whose input consists of a question and a text, with labels in the form of {no, yes} and {0, 1} for the discriminant model. CB is a textual entailment task whose input consists of a premise and a hypothesis, with labels in the form of {neutral, contradiction, entailment} and {0, 1, 2} for the discriminant model. COPA is a causal reasoning task whose input consists of a question topic, a premise, and two options, with labels in the form of {choice1, choice2}. For the discriminant model, during the experiment, a sample is converted into two data, each containing only one option. If the option is correct, the label is 1; otherwise, it is 0.

[0085] MRC is a question-answering task whose input consists of a paragraph, a question, and an answer. Labels are linguistically {false, true} and discriminatively {0, 1}. RTE is a textual entailment analysis task whose input consists of two sentences. Labels are linguistically {not entailment, entailment} and discriminatively {0, 1}. WiC is a word sense disambiguation task whose input consists of two sentences. Labels are linguistically {false, true} and discriminatively {0, 1}. SNLI is a natural language inference task whose input consists of two sentences. Labels are linguistically {neutral, contradiction, entailment} and discriminatively {0, 1, 2}. PAWS is a synonym detection task whose input consists of two sentences from a Wikipedia page. Labels are linguistically {not entailment, entailment} and discriminatively {0, 1}. IMDB is a sentiment classification task whose input consists of a sentence from a movie review. Labels are linguistically {negative, positive} and discriminatively {0, 1}.

[0086] As shown in Table 1 below, by fine-tuning model parameters or prompt words based on the BERT model, the pre-trained model trained in the present embodiment achieved an average improvement of 3.0% and 6.7% on each experimental task. Furthermore, by fine-tuning prompt words based on the Roberta model, the pre-trained model trained in the present embodiment achieved an average improvement of 9.2% on each experimental task.

[0087] Table 1

[0088] (2) In the few-shot learning scenario, the prompt word fine-tuning is performed based on the generative T5 model. In the few-shot scenarios such as 16-shot (16-sample learning), 32-shot (32-sample learning), and 100-shot (100-sample learning), the performance of the pre-trained model obtained by the model fine-tuning method provided in the embodiment of the present application is compared with the traditional pre-trained model. The pre-trained model trained in the embodiment of the present application has an improvement of 8.8%, 10.2%, and 8.5% in terms of different sample sizes, respectively, as shown in Table 2. Among them, the T5 (Text-to-Text Transfer Transformer) model regards each text processing problem as a "Text-to-Text" problem, that is, taking text as input and generating new text as output.

[0089] Table 2

[0090] (3) In order to verify the scalability of the bidirectional iterative method for the pre-trained model and downstream tasks provided by the embodiment of the present application, experiments were conducted on models with different model parameter amounts. When prompt word fine-tuning was performed on the BERT-base model or the BERT-large model, the pre-trained model obtained by the model fine-tuning method provided by the embodiment of the present application achieved an improvement of 6.7% and 6.5% respectively compared with the traditional model; when prompt word fine-tuning was performed on the T5-small model and the T5-base model, the pre-trained model obtained by the model fine-tuning method provided by the embodiment of the present application achieved an improvement of 6.0% and 8.5% compared with the traditional model. For details, please refer to Table 3. Different model parameter amounts refer to different parameter amounts of the BERT-base model and the BERT-large model, and different parameter amounts of the T5-small model and the T5-base model. Relatively speaking, the BERT-base model is a BERT model with fewer model parameters; the BERT-large model is a BERT model with more model parameters; the T5-small model is a T5 model with fewer model parameters, and the T5-base model is a T5 model with more model parameters. Table 3 is a comparison table of the effects between the pre-trained model obtained by the model fine-tuning method provided in the embodiment of the present application and the traditional model under different model parameter values.

[0091] Table 3

[0092] (4) An ablation experiment was conducted on whether to introduce learnable soft prompt words in the feedback link of the model fine-tuning method provided in the embodiment of the present application. The performance of the models with and without introducing soft prompt words was compared and analyzed on the BERT-base model, RoBERTa-base model and T5 model. The method of introducing soft prompt words can achieve an improvement of 1.0%, 0.3% and more than 3% respectively. See Table 4 for details.

[0093] Table 4

[0094] In the embodiment of the present application, the length of the soft prompt words is not limited, and the length of the soft prompt words can be flexibly set according to the application requirements and the requirements for model performance. The length of the soft prompt words belongs to a hyperparameter, and the length of the soft prompt words can be determined or modified by setting the hyperparameter. The length of the soft prompt words is different, which has a certain impact on the model performance (mainly the model accuracy). In the present embodiment, taking the T5 model as an example, the relationship diagram between the length of the soft prompt words and the model accuracy is analyzed for different few sample scenarios.

[0095] Figure 2a compares the performance of soft-prompt fine-tuning based on the T5 model for the 16-shot scenario at different soft-prompt lengths. The horizontal axis in Figure 2a represents the soft-prompt length, and the vertical axis represents the average accuracy of the model.

[0096] Figure 2b compares the performance of soft-prompt fine-tuning based on the T5 model for a 32-shot scenario at different soft-prompt lengths. The horizontal axis in Figure 2b represents the soft-prompt length, and the vertical axis represents the average accuracy of the model.

[0097] Figure 2c compares the performance of soft-prompt fine-tuning based on the T5 model for the 100-shot scenario at different soft-prompt lengths. The horizontal axis in Figure 2c represents the soft-prompt length, and the vertical axis represents the average accuracy of the model.

[0098] To better understand the bidirectional iterative method between the pre-trained model and the downstream sequence task, the following is an introduction with reference to Figures 3a and 3b.

[0099] First, prepare downstream sequential tasks. These tasks include multiple tasks with different trigger times. For example, the downstream tasks include product category prediction at 9:00 AM on November 13, 2023, product feature extraction at 9:00 AM on November 13, 2023, product content understanding at 10:00 AM on November 13, 2023, e-commerce query understanding at 10:00 AM on November 13, 2023, similar product recommendation at 11:00 AM on November 13, 2023, and model image generation at 11:00 AM on November 13, 2023. The rounds of fine-tuning training are, in chronological order: Round 1, Round 2, and Round 3. Among them, the downstream tasks appearing in Round 1 include the product category prediction task at 9:00 AM on November 13, 2023, and the product feature extraction task at 9:00 AM on November 13, 2023. The downstream tasks appearing in Round 2 include the product content understanding task at 10:00 AM on November 13, 2023, and the e-commerce query understanding task at 10:00 AM on November 13, 2023. The downstream tasks appearing in Round 3 include the similar product recommendation task at 11:00 AM on November 13, 2023, and the model image generation task at 11:00 AM on November 13, 2023.

[0100] Next, referring to FIG3a , in the pre-training stage, pre-training is performed based on training data of multiple different basic tasks to obtain a pre-training model. The pre-training model is a universal model trained with a large data set. The pre-training model can be understood as the original pre-training model.

[0101] Next, referring to Figure 3a, multiple rounds of fine-tuning training are performed after pre-training. Each round of fine-tuning training includes both a feedback link and an adaptation link.

[0102] As shown in Figure 3b, the feedback loop uses training data from several historical downstream tasks that occurred before the current round to fine-tune the initial pre-trained model of the current round based on soft prompt words, resulting in the target pre-trained model of the current round. Fine-tuning training based on the feedback loop involves adjusting the model parameters of the initial pre-trained model and optimizing the soft prompt words corresponding to the historical downstream tasks. The four-pointed star in Figure 3b represents "fine-tuning," which refers to adjusting model parameters or optimizing soft prompt words.

[0103] It is worth noting that the initial pre-training model of each round is different. Except for the initial pre-training model of the first round, which is the original pre-training model output in the pre-training phase, the initial pre-training model of other rounds is the target pre-training model of the previous round.

[0104] As shown in Figure 3b, the new model input data for each historical downstream task is model input data embedded with soft prompt words. This new model input data for each historical downstream task is fed into the initial pre-trained model for the current round, where it is fine-tuned based on the soft prompt words. This yields the prediction results output by the initial pre-trained model for the current round. The model parameters of the initial pre-trained model for the current round are then adjusted based on the loss between the labeled results and the predicted results for each historical downstream task, yielding the target pre-trained model for the current round. Furthermore, the soft prompt words for each historical downstream task are optimized based on the loss between the labeled results and the predicted results for each historical downstream task.

[0105] Referring to Figure 3b, when fine-tuning training is performed based on the adaptation link, one method is to fine-tune the model parameters of the target pre-training model of the current round and optimize the soft prompt words of the current downstream task. The other method is to freeze the model parameters of the target pre-training model of the current round and optimize the soft prompt words of the current downstream task. Regardless of which fine-tuning training is used, it is first necessary to determine the soft prompt words of the current downstream task that appear in the current round. This can be obtained by various methods such as weighted summation and averaging of the soft prompt words of each historical downstream task that appeared before the current round. In Figure 3b, the gray-filled squares represent the soft prompt words of the historical downstream tasks, the black-filled squares represent the soft prompt words of the current downstream task, and the black hexagonal star represents freezing.

[0106] Referring to Figure 3b, the new model input data for the current downstream task is model input data embedded with soft prompt words. The new model input data for each current downstream task is input into the target pre-trained model for the current round, where it is fine-tuned based on the soft prompt words. This yields the prediction results output by the target pre-trained model for the current round. Based on the loss between the annotation results and the prediction results of the current downstream task, the model parameters of the target pre-trained model for the current round are adjusted or frozen to obtain the task model corresponding to the current downstream task. Furthermore, based on the loss between the annotation results and the prediction results of the current downstream task, the soft prompt words for the current downstream task are optimized to obtain the soft prompt words used by the current downstream task during task model inference.

[0107] After obtaining the task model of the downstream task based on the above-mentioned pre-trained model and the bidirectional iterative method of the downstream sequence task, the task model can be used to process the downstream task. To this end, the embodiment of the present application also provides a downstream task processing method. Figure 4 is a flow chart of a downstream task processing method provided by the embodiment of the present application. Referring to Figure 4, the method may include the following steps:

[0108] 401. Obtain task data of a downstream task to be processed, a task model corresponding to the downstream task to be processed, and soft prompt words used for task model reasoning.

[0109] 402. Generate model input data based on the soft prompt words and task data.

[0110] 403. Input the model input data into the task model to obtain model output data.

[0111] In some optional embodiments, the task data includes multimodal data, and the task model is a large multimodal model that can perform reasoning using the multimodal data as input data. The large multimodal model has good reasoning performance.

[0112] In this embodiment, the corresponding task models vary depending on the downstream task. For example, if the downstream task is a product category prediction task, the task model is a product category prediction model. When the downstream task to be processed is a product category prediction task, the multimodal data related to the product category prediction and the soft prompt words are spliced ​​to obtain model input data. The model input data is input into the product category prediction model to perform category prediction and obtain the product category prediction result.

[0113] For another example, the downstream task is a product feature extraction task, and the task model is a product feature extraction model; when the downstream task to be processed is a product feature extraction task, the multimodal data and soft prompt words related to the product feature extraction task are spliced ​​to obtain model input data, and the model input data is input into the product feature extraction model for feature extraction to obtain product features.

[0114] For another example, the downstream task is a product content understanding task, and the task model is a product content understanding model. When the downstream task to be processed is a product content understanding task, the multimodal data and soft prompt words related to the product content understanding task are spliced ​​to obtain model input data, and the model input data is input into the product content understanding model for content understanding to obtain the product content understanding result.

[0115] The downstream task processing method provided in the embodiment of the present application is able to be better applied in downstream tasks and has better model performance because the task model is trained using a bidirectional iterative method of a pre-trained model and downstream sequence tasks.

[0116] It is worth noting that in the aforementioned embodiments, the data input into various models may be in vector form. When obtaining another soft prompt word based on one soft prompt word, it may be based on a soft prompt word in a vector form to obtain a soft prompt word in another vector form. There is no limitation on this.

[0117] In order to better understand the technical solutions provided by the embodiments of the present application, specific scenario embodiments are introduced below.

[0118] Scenario Example 1:

[0119] The number of users using e-commerce platforms that offer second-hand goods transactions (referred to as second-hand e-commerce platforms) is increasing. However, these platforms face some challenges in structuring product information, such as difficulty extracting product attributes and product categories. At the same time, these platforms have accumulated a large amount of multimodal product data and search queries, which lays the foundation for training multimodal pre-trained models in the second-hand e-commerce field. Although many open-source multimodal pre-trained models are trained on general domain data, they do not fully understand product information or provide a complete product feature space, and they do not fully utilize the feedback provided by downstream tasks to pre-trained models.

[0120] To this end, as shown in Figure 5 (1), the cloud server first collects various data, including multimodal product data and search questions, from e-commerce platforms that provide second-hand goods trading services. Multimodal product data contains multimodal information, such as images and text. This data enables the model to better understand product information and provide a more comprehensive representation of product features. Next, as shown in Figure 5 (2), the cloud server uses the multimodal product data and search questions to pre-train a general multimodal pre-trained model for the second-hand e-commerce domain. This general multimodal pre-trained model can be a general multimodal large model. Finally, as shown in Figure 5 (3), the cloud server performs multiple rounds of fine-tuning training to obtain task models for downstream tasks in different verticals. The multimodal pre-trained model and the downstream task feedback mechanism are then combined to generate task models for different downstream tasks. For example, these task models include, but are not limited to, product category prediction models, product feature extraction models, and product content understanding models. These downstream task models can be deployed on the second-hand e-commerce platform to perform product category prediction, product feature extraction, and product content understanding tasks, thereby improving the user experience and service quality of the second-hand e-commerce platform.

[0121] Scenario Example 2:

[0122] The cloud server uses multimodal product data and search problems to pre-train a general multimodal pre-training model in the second-hand e-commerce field, which can also be applied to various vertical downstream tasks in e-commerce scenarios. For example, it can be applied to e-commerce platforms that provide first-hand product trading services (which can be called first-hand e-commerce platforms). For example, the first-hand e-commerce platform has video content understanding tasks, model main picture generation tasks, and similar product recommendation tasks. The cloud server then performs multiple rounds of fine-tuning training to obtain task models applied to different vertical downstream tasks. Then, combining the multimodal pre-training model and the downstream task feedback mechanism requires task models for different downstream tasks. For example, task models include but are not limited to: video content understanding model, model main picture generation model, and similar product recommendation model. The task models of these downstream tasks can be deployed in the first-hand e-commerce platform to enable the first-hand e-commerce platform to perform video content understanding tasks, model main picture generation tasks, and similar product recommendation tasks, thereby improving the user experience and service quality of the first-hand e-commerce platform.

[0123] It should be noted that the execution entity of each step of the method provided in the above embodiment can be the same device, or the method can be executed by different devices. For example, the execution entity of steps 101 to 103 can be device A; for another example, the execution entity of steps 101 and 102 can be device A, and the execution entity of step 103 can be device B; and so on.

[0124] In addition, some of the processes described in the above embodiments and the accompanying drawings include multiple operations that appear in a specific order, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or may be executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any order of execution. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., and do not represent a sequential order, nor do they limit "first" and "second" to different types.

[0125] FIG6 is a schematic diagram of the structure of a bidirectional iteration device provided in an embodiment of the present application. As shown in FIG6 , the device may include:

[0126] Determination module 61, used to determine the initial pre-training model of the current round, the initial pre-training model of the current round is the target pre-training model obtained by fine-tuning in the previous round;

[0127] Feedback training module 62, configured to utilize training data of historical downstream tasks that occurred before the current round to fine-tune the initial pre-trained model of the current round based on soft prompt words to obtain a target pre-trained model of the current round;

[0128] The adaptive training module 63 is used to use the training data of the current downstream task appearing in the current round to fine-tune the target pre-training model of the current round based on the soft prompt words to obtain the task model corresponding to the current downstream task.

[0129] Further optionally, the feedback training module 62 is specifically used to: generate a first soft prompt word corresponding to each historical downstream task that appeared before the current round; embed the first soft prompt word into the training data of the corresponding historical downstream task to obtain new training data for the historical downstream task; and fine-tune the initial pre-training model of the current round according to the new training data of the historical downstream task to obtain the target pre-training model of the current round.

[0130] Further optionally, when the feedback training module 62 generates the first soft prompt word corresponding to each historical downstream task that appeared before the current round, it is specifically used to: for any historical downstream task, generate a first soft prompt word for it based on the second soft prompt word corresponding to the historical downstream task before any historical downstream task; or for any historical downstream task, generate a first soft prompt word for it based on the fourth soft prompt word used for task model reasoning corresponding to any historical downstream task; or for any historical downstream task, randomly initialize the first soft prompt word for it.

[0131] Further optionally, the feedback training module 62 fine-tunes the initial pre-training model of the current round based on the new training data of the historical downstream tasks to obtain the target pre-training model of the current round, specifically for:

[0132] Input the model input data in the new training data of the historical downstream task into the initial pre-trained model of the current round to obtain the prediction result of the model input data of the historical downstream task; calculate the first loss value according to the prediction result of the model input data of the historical downstream task and the corresponding annotation result in the new training data of the historical downstream task; adjust the model parameters of the initial pre-trained model of the current round according to the first loss value to obtain the target pre-trained model of the current round; and optimize the first soft prompt word corresponding to the historical downstream task according to the first loss value to obtain the second soft prompt word.

[0133] Further optionally, when the feedback training module 62 optimizes the first soft prompt word corresponding to the historical downstream task according to the first loss value to obtain the second soft prompt word, it is specifically used to: optimize the first soft prompt word corresponding to the historical downstream task that appeared in at least one recent historical round according to the first loss value to obtain the second soft prompt word.

[0134] Further optionally, the adaptive training module 63 is specifically used to: generate a third soft prompt word corresponding to the current downstream task; embed the third soft prompt word into the training data of the current downstream task to obtain new training data for the current downstream task; and fine-tune the target pre-training model of the current round according to the new training data of the current downstream task to obtain the task model corresponding to the current downstream task and the fourth soft prompt word used for task model inference.

[0135] Further optionally, when the adaptive training module 63 generates the third soft prompt word corresponding to the current downstream task, it is specifically used to: generate the third soft prompt word corresponding to the current downstream task based on the second soft prompt word corresponding to the historical downstream task that appeared before the current round; or generate the third soft prompt word corresponding to the current downstream task based on the first soft prompt word corresponding to the historical downstream task that appeared before the current round; or randomly initialize the third soft prompt word for the current downstream task.

[0136] Further optionally, the adaptive training module 63 fine-tunes the target pre-training model of the current round according to the new training data of the current downstream task to obtain the task model corresponding to the current downstream task and the fourth soft prompt word used for task model reasoning. It is specifically used to: input the model input data in the new training data of the current downstream task into the target pre-training model of the current round to obtain the prediction result of the model input data of the current downstream task; calculate the second loss value according to the prediction result of the model input data of the current downstream task and the corresponding annotation result in the new training data of the current downstream task; freeze the model parameters of the target pre-training model of the current round to make the target pre-training model the task model of the current downstream task, and optimize the third soft prompt word corresponding to the current downstream task according to the second loss value to obtain the fourth soft prompt word used for task model reasoning.

[0137] Further optionally, the adaptive training module 63 fine-tunes the target pre-training model of the current round according to the new training data of the current downstream task to obtain the task model corresponding to the current downstream task and the fourth soft prompt word used for task model reasoning. It is specifically used to: input the model input data in the new training data of the current downstream task into the target pre-training model of the current round to obtain the prediction result of the model input data of the current downstream task; calculate the second loss value according to the prediction result of the model input data of the current downstream task and the corresponding annotation result in the new training data of the current downstream task; and optimize the model parameters of the target pre-training model of the current round and the third soft prompt word corresponding to the current downstream task according to the second loss value to obtain the task model corresponding to the current downstream task and the fourth soft prompt word used for task model reasoning.

[0138] Further optionally, the current downstream task includes at least one of the following: a product category prediction task, a product feature extraction task, and a product content understanding task; and / or the task model is a large multimodal model.

[0139] The device shown in FIG6 can execute the method shown in the embodiment shown in FIG1, and its implementation principle and technical effects are not described in detail here. The specific manner in which each module and unit performs the operation of the device shown in FIG6 in the above embodiment has been described in detail in the embodiment of the method, and will not be elaborated here.

[0140] FIG7 is a schematic diagram of the structure of an electronic device provided in an embodiment of the present application. As shown in FIG7 , the electronic device includes: a memory 71 and a processor 72;

[0141] The memory 71 is used to store computer programs and can be configured to store various other data to support operations on the computing platform. Examples of such data include instructions for any application or method operating on the computing platform, contact data, phone book data, messages, pictures, videos, etc.

[0142] The memory 71 can be implemented by any type of volatile or non-volatile memory device or a combination thereof, such as static random-access memory (SRAM), electrically erasable programmable read only memory (EEPROM), erasable programmable read only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic memory, flash memory, magnetic disk or optical disk.

[0143] The processor 72 is coupled to the memory 71 and is used to execute the computer program in the memory 71 to: execute the steps in the bidirectional iterative method of the pre-training model and the downstream sequence task or the downstream task processing method.

[0144] Further optionally, as shown in Figure 7, the electronic device also includes: a communication component 73, a display 74, a power supply component 75, an audio component 76 and other components. Figure 7 only schematically shows some components, which does not mean that the electronic device only includes the components shown in Figure 7. In addition, the components in the dotted box in Figure 7 are optional components, not mandatory components, and the specific product form of the electronic device may depend on the product form. The electronic device of this embodiment can be implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone or an IOT (Internet of things) device, or it can be a server-side device such as a conventional server, a cloud server or a server array. If the electronic device of this embodiment is implemented as a terminal device such as a desktop computer, a laptop computer, a smart phone, etc., it may include the components in the dotted box in Figure 7; if the electronic device of this embodiment is implemented as a server-side device such as a conventional server, a cloud server or a server array, it may not include the components in the dotted box in Figure 7.

[0145] The detailed implementation process of the processor executing each action can be found in the relevant description in the aforementioned method embodiment or device embodiment, and will not be repeated here.

[0146] Accordingly, an embodiment of the present application further provides a computer-readable storage medium storing a computer program, which, when executed, can implement the steps that can be performed by the electronic device in the above method embodiment.

[0147] Accordingly, an embodiment of the present application also provides a computer program product, including a computer program / instruction. When the computer program / instruction is executed by a processor, the processor is enabled to implement the steps in the above method embodiment that can be performed by an electronic device.

[0148] The above-mentioned communication component is configured to facilitate wired or wireless communication between the device where the communication component is located and other devices. The device where the communication component is located can access a wireless network based on a communication standard, such as WiFi (Wireless Fidelity), 2G (2Generation, 2nd generation), 3G (3Generation, 3rd generation), 4G (4Generation, 4th generation) / LTE (long Term Evolution, long term evolution), 5G (5Generation, 5th generation) and other mobile communication networks, or a combination thereof. In an exemplary embodiment, the communication component receives a broadcast signal or broadcast-related information from an external broadcast management system via a broadcast channel. In an exemplary embodiment, the communication component also includes a near field communication (NFC) module to facilitate short-range communication. For example, the NFC module can be implemented based on radio frequency identification (RFID) technology, infrared data association (IrDA) technology, ultra-wide band (UWB) technology, Bluetooth (BT) technology and other technologies.

[0149] The above-mentioned display includes a screen, which may include a liquid crystal display (LCD) and a touch panel (TP). If the screen includes a touch panel, the screen may be implemented as a touch screen to receive input signals from the user. The touch panel includes one or more touch sensors to sense touch, slide, and gestures on the touch panel. The touch sensor can not only sense the boundary of the touch or slide action, but also detect the duration and pressure associated with the touch or slide operation.

[0150] The power supply assembly provides power to various components of the device in which the power supply assembly is located. The power supply assembly may include a power management system, one or more power supplies, and other components associated with generating, managing, and distributing power to the device in which the power supply assembly is located.

[0151] The above-mentioned audio component can be configured to output and / or input audio signals. For example, the audio component includes a microphone (MIC), and when the device where the audio component is located is in an operating mode, such as call mode, recording mode, and voice recognition mode, the microphone is configured to receive external audio signals. The received audio signal can be further stored in a memory or sent via a communication component. In some embodiments, the audio component also includes a speaker for outputting audio signals.

[0152] Those skilled in the art will appreciate that the embodiments of the present application can be provided as methods, systems, or computer program products. Therefore, the present application can adopt the form of a complete hardware embodiment, a complete software embodiment, or an embodiment in combination with software and hardware. Moreover, the present application can adopt the form of a computer program product implemented on one or more computer-readable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) that contain computer-usable program code.

[0153] The present application is described with reference to the flow chart and / or block diagram of the method, device (system), and computer program product according to the embodiment of the present application. It should be understood that each flow process and / or box in the flow chart and / or block diagram and the combination of the flow process and / or box in the flow chart and / or block diagram can be realized by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processing machine or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device produce a device for realizing the function specified in one flow chart flow or multiple flows and / or one box or multiple boxes of the block diagram.

[0154] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0155] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.

[0156] In a typical configuration, a computing device includes one or more processors (Central Processing Unit, CPU), input / output interfaces, network interfaces, and memory.

[0157] Memory may include non-permanent storage in a computer-readable medium, random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash RAM. Memory is an example of a computer-readable medium.

[0158] Computer-readable media includes permanent and non-permanent, removable and non-removable media that can be used to store information using any method or technology. Information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change RAM (PRAM), static random-access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, magnetic disk storage or other magnetic storage devices, or any other non-transmission media that can be used to store information that can be accessed by a computing device. As defined herein, computer-readable media does not include transitory computer-readable media, such as modulated data signals and carrier waves.

[0159] It should also be noted that the terms "comprises," "includes," or any other variations thereof are intended to encompass non-exclusive inclusion, such that a process, method, commodity, or apparatus that includes a series of elements includes not only those elements but also other elements not explicitly listed, or includes elements inherent to such process, method, commodity, or apparatus. In the absence of further limitations, an element defined by the phrase "comprises a ..." does not exclude the presence of other identical elements in the process, method, commodity, or apparatus that includes the element.

[0160] The above are merely embodiments of the present application and are not intended to limit the present application. For those skilled in the art, the present application may have various changes and variations. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present application should all be included within the scope of the claims of the present application.

Claims

1. A bidirectional iterative method for a pre-trained model and a downstream sequence task, characterized in that, Including: Determine the initial pre-trained model for the current round. The initial pre-trained model for the current round is the target pre-trained model obtained by fine-tuning in the previous round; Use the training data of historical downstream tasks that appeared before the current round to fine-tune the initial pre-trained model of the current round based on soft prompts to obtain the target pre-trained model of the current round; Use the training data of the current downstream task that appears in the current round to fine-tune the target pre-trained model of the current round based on soft prompts to obtain the task model corresponding to the current downstream task.

2. The method according to claim 1, characterized in that, Using the training data of historical downstream tasks that appeared before the current round to fine-tune the initial pre-trained model of the current round based on soft prompts to obtain the target pre-trained model of the current round, including: Generate the first soft prompt corresponding to each historical downstream task that appeared before the current round; Embed the first soft prompt into the training data of the corresponding historical downstream task to obtain the new training data of the historical downstream task; Fine-tune the initial pre-trained model of the current round according to the new training data of the historical downstream task to obtain the target pre-trained model of the current round.

3. The method according to claim 2, wherein Generating the first soft prompt corresponding to each historical downstream task that appeared before the current round, including: For any historical downstream task, generate the first soft prompt for it according to the second soft prompt corresponding to the historical downstream task before the any historical downstream task; Or For any historical downstream task, generate the first soft prompt for it according to the fourth soft prompt used for inference of the task model corresponding to the any historical downstream task; Or For any historical downstream task, randomly initialize the first soft prompt for it.

4. The method according to claim 2, wherein Fine-tuning the initial pre-trained model of the current round according to the new training data of the historical downstream task to obtain the target pre-trained model of the current round, including: Input the model input data in the new training data of the historical downstream task into the initial pre-trained model of the current round to obtain the prediction result of the model input data of the historical downstream task; Calculate the first loss value according to the prediction result of the model input data of the historical downstream task and the corresponding annotation result in the new training data of the historical downstream task; Adjust the model parameters of the initial pre-trained model of the current round according to the first loss value to obtain the target pre-trained model of the current round; and Optimize the first soft prompt corresponding to the historical downstream task according to the first loss value to obtain the second soft prompt.

5. The method according to claim 4, characterized in that, Optimizing the first soft prompt corresponding to the historical downstream task according to the first loss value to obtain the second soft prompt, including: Optimize the first soft prompt corresponding to the historical downstream task that appeared in at least the most recent historical rounds according to the first loss value to obtain the second soft prompt.

6. The method according to any one of claims 1-5, characterized in that, Using the training data of the current downstream task that appears in the current round to fine-tune the target pre-trained model of the current round based on soft prompts to obtain the task model corresponding to the current downstream task, including: Generate the third soft prompt corresponding to the current downstream task; Embed the third soft prompt into the training data of the current downstream task to obtain new training data for the current downstream task; Fine-tune the target pre-trained model of the current round according to the new training data of the current downstream task to obtain a task model corresponding to the current downstream task and a fourth soft prompt used for inference of the task model.

7. The method according to claim 6, characterized in that, Generate a third soft prompt corresponding to the current downstream task, including: Generate a third soft prompt corresponding to the current downstream task according to the second soft prompt corresponding to the historical downstream tasks that appeared before the current round; Alternatively, generate a third soft prompt corresponding to the current downstream task according to the first soft prompt corresponding to the historical downstream tasks that appeared before the current round; Alternatively, randomly initialize a third soft prompt for the current downstream task.

8. The method according to claim 7, characterized in that Fine-tune the target pre-trained model of the current round according to the new training data of the current downstream task to obtain a task model corresponding to the current downstream task and a fourth soft prompt used for inference of the task model, including: Input the model input data in the new training data of the current downstream task into the target pre-trained model of the current round to obtain a prediction result of the model input data of the current downstream task; Calculate a second loss value according to the prediction result of the model input data of the current downstream task and the corresponding annotation result in the new training data of the current downstream task; Freeze the model parameters of the target pre-trained model of the current round to use the target pre-trained model as the task model for the current downstream task, and optimize the third soft prompt corresponding to the current downstream task according to the second loss value to obtain the fourth soft prompt used for inference of the task model.

9. The method according to claim 7, characterized in that, Fine-tune the target pre-trained model of the current round according to the new training data of the current downstream task to obtain a task model corresponding to the current downstream task and a fourth soft prompt used for inference of the task model, including: Input the model input data in the new training data of the current downstream task into the target pre-trained model of the current round to obtain a prediction result of the model input data of the current downstream task; Calculate a second loss value according to the prediction result of the model input data of the current downstream task and the corresponding annotation result in the new training data of the current downstream task; According to the second loss value, optimize the model parameters of the target pre-trained model of the current round and the third soft prompt corresponding to the current downstream task to obtain a task model corresponding to the current downstream task and a fourth soft prompt used for inference of the task model.

10. The method according to any one of claims 1 to 5, characterized in that, The current downstream task includes at least one of the following: product category prediction task, product feature extraction task, and product content understanding task; and / or, the task model is a multimodal large model.

11. A downstream task processing method, characterized in that, Including: Obtain the task data of the downstream task to be processed, the task model corresponding to the downstream task to be processed, and the soft prompt used for inference of the task model; Generate model input data according to the soft prompt and the task data; Input the model input data into the task model to obtain model output data; Wherein, the task model is trained according to the method described in any one of claims 1-10.

12. The method according to claim 11, wherein The task data includes multimodal data, and the task model is a multimodal large model.

13. The method according to claim 11 or 12, characterized in that Inputting the model input data into the task model to obtain model output data, including: When the downstream task to be processed is a commodity category prediction task, inputting the model input data into a commodity category prediction model for category prediction to obtain a commodity category prediction result, where the model input data is obtained by splicing multimodal data related to commodity category prediction and the soft prompt words; When the downstream task to be processed is a commodity feature extraction task, inputting the model input data into a commodity feature extraction model for feature extraction to obtain commodity features, where the model input data is obtained by splicing multimodal data related to the commodity feature extraction task and the soft prompt words; When the downstream task to be processed is a commodity content understanding task, inputting the model input data into a commodity content understanding model for content understanding to obtain a commodity content understanding result, where the model input data is obtained by splicing multimodal data related to the commodity content understanding task and the soft prompt words.

14. An electronic device, characterized in that, Including: A memory and a processor; The memory is used for storing a computer program; The processor is coupled to the memory and is used for executing the computer program to execute the steps in the method according to any one of claims 1-13.

15. A computer storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it causes the processor to be able to implement the steps in the method according to any one of claims 1-13.

Citation Information

Patent Citations

  • Pre-training model fine tuning method and device, equipment and storage medium

    CN115049003A

  • Pre-training language model fine tuning method, system and device

    CN115423118A

  • Health risk assessment method based on pre-training and multi-task bidirectional regulation mechanism

    CN116110582A

  • Model training method and device, image processing method and device, equipment and storage medium

    CN116580261A

  • Language perception multi-language pre-training and fine tuning method based on language discrimination prompt

    CN116861242A

Cited By

  • Multi-model cooperative processing method and system

    CN120560865A

  • Multi-agent cooperative abnormal information processing method and device, storage medium and program product

    CN120639482A

  • Large model application optimization method and device based on user feedback, equipment and medium

    CN120806172A

  • Prompt word optimization method and device, electronic equipment and storage medium

    CN121920385A

  • Prompt word optimization method and device, electronic equipment and storage medium

    CN121920385B