Transfer learning model training
By initializing the source domain and target domain models at the target migration moment, and using the parameter migration network to obtain and migrate hidden parameters from the source domain model, the problem of large amount of model training and low training efficiency in the existing technology is solved, and efficient transfer learning is achieved.
Patent Information
- Application Number
- PCT/CN2024/128060
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2023-11-14
- Filing Date
- 2024-10-29
- Publication Date
- 2025-05-22
AI Technical Summary
In the prior art, model training is large in computational volume and low in training efficiency, making it difficult to effectively migrate the source domain model to the target domain model.
The source domain model and the target domain model are initialized at the target migration moment to meet the preset similarity conditions. Through the parameter migration network, obtain the target source domain hidden parameters from the source domain model and migrate it to the target domain hidden layer to form the target domain migration model. The target domain migration model is then trained based on the target domain sample data until convergence.
Directly obtaining deep feature parameters through parameter migration network, avoiding the time and resource consumption of retraining, improving training efficiency, and not interfering with each other between the source domain and the target domain model.
Smart Images

Figure CN2024128060_22052025_PF_FP_ABST
Abstract
Description
Transfer learning model training Technical Field
[0001] The embodiments of this specification relate to the field of computer artificial intelligence technology, and in particular, to a transfer learning model training method, device, storage medium, and terminal. Background Art
[0002] In today's diverse internet environments, there are numerous scenarios with high similarity. Within these scenarios, there are numerous similar computing tasks with shared computational processing logic. To effectively improve model performance in the target scenario, model knowledge from these similar scenarios can be transferred to optimize the target model. This model training method is called transfer learning. Therefore, there is an urgent need for a transfer learning model training method that allows the target domain model to learn key knowledge without interfering with the source domain model.
[0003] Summary of the Invention
[0004] The embodiments of this specification provide a transfer learning model training method, device, storage medium and terminal, which can solve the technical problems of large model training computational complexity and low training efficiency in related technologies.
[0005] In a first aspect, an embodiment of the present specification provides a method for training a transfer learning model, the method comprising: at a target migration moment, simultaneously initializing a source domain model and a target domain model, wherein the source domain and the target domain meet a preset similarity condition; based on a parameter migration network corresponding to each target domain hidden layer in the target domain model, respectively obtaining target source domain hidden parameters of each target domain hidden layer corresponding to each source domain hidden layer in the source domain model; migrating each target source domain hidden parameter to each target domain hidden layer to obtain a target domain migration model, and training the target domain migration model based on target sample data corresponding to the target domain until the target domain migration model converges.
[0006] In a second aspect, an embodiment of the present specification provides a transfer learning model training device, which includes: an initialization module for simultaneously initializing a source domain model and a target domain model at a target migration moment, wherein the source domain and the target domain meet a preset similarity condition; a parameter migration module for respectively obtaining target source domain hidden parameters of each target domain hidden layer corresponding to each source domain hidden layer in the source domain model based on a parameter migration network corresponding to each target domain hidden layer in the target domain model; a target domain training module for migrating each target source domain hidden parameter to each target domain hidden layer to obtain a target domain migration model, and training the target domain migration model based on target sample data corresponding to the target domain until the target domain migration model converges.
[0007] In a third aspect, an embodiment of this specification provides a computer program product comprising instructions, which, when executed on a computer or a processor, enables the computer or the processor to execute the steps of the above method.
[0008] In a fourth aspect, an embodiment of this specification provides a computer storage medium, wherein the computer storage medium stores a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the steps of the above method.
[0009] In a fifth aspect, an embodiment of this specification provides a terminal, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the computer program is suitable for being loaded by the processor and executing the steps of the above method.
[0010] The beneficial effects brought about by the technical solutions provided by some embodiments of this specification include at least: the embodiments of this specification provide a transfer learning model training method, at the target migration moment, the source domain model and the target domain model are initialized simultaneously, and the source domain and the target domain meet the preset similarity condition; based on the parameter migration network corresponding to each target domain hidden layer in the target domain model, the target source domain hidden parameters of each target domain hidden layer corresponding to each source domain hidden layer in the source domain model are obtained respectively; each target source domain hidden parameter is migrated to each target domain hidden layer to obtain a target domain migration model, and the target domain migration model is trained based on the target sample data corresponding to the target domain until the target domain migration model converges. When a certain similarity relationship is satisfied between the target domain and the source domain, the target source domain hidden parameters that reflect the deep network characteristics are selected from each source domain hidden layer of the source domain model through the parameter transfer network, and are transferred to the corresponding target domain hidden layer of the target domain model, so that the target domain hidden layer can directly obtain effective deep feature parameters without the need to retrain from the initial state. Moreover, the target domain transfer model obtained after parameter transfer can still be trained using the target sample data corresponding to the target domain, so that the target domain transfer model is suitable for completing the prediction task in the target domain. This not only saves a lot of training time and training computing resources of the target domain transfer model, but also prevents the source domain model and the target domain model from interfering with each other. BRIEF DESCRIPTION OF THE DRAWINGS
[0011] In order to more clearly illustrate the embodiments of this specification or the technical solutions in the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are only some embodiments of this specification. For those skilled in the art, other drawings can be obtained based on these drawings without paying any creative work.
[0012] FIG1 is an exemplary system architecture diagram of a transfer learning model training method provided in an embodiment of this specification.
[0013] FIG2 is a flow chart of a transfer learning model training method provided in an embodiment of this specification.
[0014] FIG3 is a schematic diagram of a model deployment for cross-domain transfer learning provided in an embodiment of this specification.
[0015] FIG4 is a flow chart of a transfer learning model training method provided in an embodiment of this specification.
[0016] FIG5 is a schematic diagram of a continuous transfer learning model training provided in an embodiment of this specification.
[0017] FIG6 is a structural block diagram of a transfer learning model training device provided in an embodiment of this specification.
[0018] FIG7 is a schematic diagram of the structure of a terminal provided in an embodiment of this specification. DETAILED DESCRIPTION
[0019] To make the features and advantages of the embodiments of this specification more obvious and easy to understand, the technical solutions in the embodiments of this specification will be clearly and completely described below in conjunction with the drawings in the embodiments of this specification. Obviously, the embodiments described are only part of the embodiments of this specification, not all of the embodiments. Based on the embodiments in this specification, all other embodiments obtained by those skilled in the art without making creative efforts shall fall within the scope of protection of the embodiments of this specification.
[0020] When the following description refers to the accompanying drawings, unless otherwise indicated, the same numbers in different drawings represent the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the embodiments of this specification. Instead, they are merely examples of devices and methods consistent with certain aspects of the embodiments of this specification, as detailed in the appended claims.
[0021] With the rapid development of artificial intelligence (AI) and machine learning technologies, neural network models, owing to their fast and accurate computing capabilities, have been widely applied in many fields. The typical network model training process involves inputting a large amount of sample data from the target scenario after the initial modeling is completed. This allows the model to learn knowledge during training and continuously adjust the network feature parameters in the hidden layer. When the model's output performance reaches the expected state, the model is considered converged and can be deployed in real-world scenarios. When preparing sample data, it is often necessary to manually assign the corresponding classification labels to the sample data so that the model can correctly learn the relationship between data features and classification prediction results.
[0022] However, if the target scenario model is continuously optimized solely based on data from a single target scenario, the model's optimization performance will eventually reach a bottleneck as the amount and detail of information contained in the data samples within that scenario stabilizes. At this point, other approaches can be considered to obtain richer network features to further iterate on model performance. In today's diverse internet environment, many scenarios share similarities, and similar computing tasks with shared processing logic exist within similar scenarios. Therefore, it can be observed that after convergence, models in these similar scenarios also possess similar feature information. Therefore, a model training method has emerged that transfers model knowledge from similar scenarios to help optimize the target scenario model, known as transfer learning model training. Transfer learning allows existing information from one model to be directly applied to another. This not only saves the training time and computing resources required to acquire knowledge from similar models, but also enriches the network information in the current scenario with information from other scenarios.
[0023] Basic transfer learning involves various approaches, including sample-based transfer, feature-based transfer, model-based transfer, and relationship-based transfer. Sample-based and model-based transfers primarily focus on directly reusing sample data and models, relationship-based transfers primarily explore and leverage existing relationships within scenarios for analogical transfer, and feature-based transfers prioritize the use of feature knowledge within the model network. Therefore, to quickly and efficiently optimize the target domain model, cross-domain transfer learning model training can be employed. This involves transferring knowledge from a source domain model that has a certain degree of similarity to the target domain, thereby conserving training resources for the target domain model. Existing cross-domain transfer learning model training methods in academia and industry can be categorized into two main types: joint training and pre-training & fine-tuning. Joint training methods simultaneously optimize both the source and target domain models. However, this type of method requires the introduction of source domain data during training, and the source domain data is usually large in scale, which requires huge computing and storage resources. Therefore, the pre-training-fine-tuning method is more widely used. However, the pre-training-fine-tuning method is an offline knowledge transfer method that only completes transfer learning once, and the model cannot perform transfer learning on newly added data.
[0024] Therefore, the embodiments of this specification provide a transfer learning model training method, which initializes the source domain model and the target domain model at the same time at the target migration moment, and the source domain and the target domain meet the preset similarity condition; based on the parameter migration network corresponding to each target domain hidden layer in the target domain model, obtain the target source domain hidden parameters of each target domain hidden layer corresponding to each source domain hidden layer in the source domain model; migrate each target source domain hidden parameter to each target domain hidden layer to obtain a target domain migration model, and train the target domain migration model based on the target sample data corresponding to the target domain until the target domain migration model converges, so as to solve the technical problems of large model training computational complexity and low training efficiency.
[0025] Please refer to Figure 1, which is an exemplary system architecture diagram of a transfer learning model training method provided in an embodiment of this specification.
[0026] As shown in Figure 1, the system architecture may include a terminal 101, a network 102, and a server 103. The network 102 is used to provide a medium for a communication link between the terminal 101 and the server 103. The network 102 may include various types of wired communication links or wireless communication links, for example, a wired communication link including an optical fiber, a twisted pair, or a coaxial cable, and a wireless communication link including a Bluetooth communication link, a Wireless-Fidelity (Wi-Fi) communication link, or a microwave communication link.
[0027] The terminal 101 can interact with the server 103 through the network 102 to receive a message from the server 103 or send a message to the server 103, or the terminal 101 can interact with the server 103 through the network 102 to receive a message or data sent by other users to the server 103. The terminal 101 can be hardware or software. When the terminal 101 is hardware, it can be various electronic devices, including but not limited to smart watches, smart phones, tablet computers, laptop portable computers and desktop computers. When the terminal 101 is software, it can be installed in the electronic devices listed above, which can be implemented as multiple software or software modules (for example: for providing distributed services), or it can be implemented as a single software or software module, which is not specifically limited here.
[0028] In an embodiment of the present specification, a transfer learning environment between a source domain and a target domain is deployed in the terminal 101, so that the terminal 101 can initialize the source domain model and the target domain model at the same time at the target migration moment, and the source domain and the target domain meet a preset similarity condition; after the initialization is completed, the terminal 101 migrates the network based on the parameters corresponding to each target domain hidden layer in the target domain model, and obtains the target source domain hidden parameters of each target domain hidden layer corresponding to each source domain hidden layer in the source domain model; finally, after obtaining the target source domain hidden parameters, the terminal 101 migrates each target source domain hidden parameter to each target domain hidden layer to obtain a target domain migration model, and trains the target domain migration model based on the target sample data corresponding to the target domain until the target domain migration model converges.
[0029] The server 103 may be a business server that provides various services. It should be noted that the server 103 may be hardware or software. When the server 103 is hardware, it may be implemented as a distributed server cluster consisting of multiple servers, or it may be implemented as a single server. When the server 103 is software, it may be implemented as multiple software or software modules (for example, for providing distributed services), or it may be implemented as a single software or software module, which is not specifically limited herein.
[0030] Alternatively, the system architecture may also not include the server 103. In other words, the server 103 may be an optional device in the embodiments of this specification, that is, the method provided in the embodiments of this specification may be applied to a system structure that only includes the terminal 101, and the embodiments of this specification do not limit this.
[0031] It should be understood that the number of terminals, networks, and servers in FIG1 is merely illustrative, and any number of terminals, networks, and servers may be used according to implementation requirements.
[0032] Please refer to Figure 2, which is a flowchart illustrating a transfer learning model training method provided in an embodiment of this specification. The execution subject of an embodiment of this specification can be a terminal that performs transfer learning model training, a processor in a terminal that performs the transfer learning model training method, or a transfer learning model training service in a terminal that performs the transfer learning model training method. For ease of description, the specific execution process of the transfer learning model training method is described below using the example of a processor in a terminal as the execution subject.
[0033] As shown in FIG2 , the transfer learning model training method may include at least the following steps.
[0034] S202: At the target migration moment, the source domain model and the target domain model are initialized simultaneously, and the source domain and the target domain meet a preset similarity condition.
[0035] Optionally, during model algorithm development, scenarios requiring online deployment and dealing with large amounts of computational data are called industrial scenarios. With the recent application of deep models, the effectiveness of recommendation systems in industrial scenarios has significantly improved. However, due to the rapid iteration of online environmental information in industrial scenarios, models must use the latest samples and learn the latest data distribution through methods such as offline incremental learning or online learning. Therefore, compared to ordinary models that undergo a single offline training phase and are then deployed online, a key feature of recommendation models for industrial scenarios is that they are trained using a continuous learning paradigm.
[0036] Alternatively, since both the source domain model and the target domain model in the industrial scenario follow the continuous learning training method in their respective domains, which means that the knowledge in the model is also constantly iteratively updated, then the knowledge obtained by only one transfer learning will obviously not remain valid for a long time. Therefore, in order to obtain a stable knowledge transfer effect, continuous learning can be introduced on the basis of transfer learning, so that the migration process is continuous and multiple times. At this time, the two cross-domain models respectively maintain the knowledge of receiving new data, and the target domain model can also receive new effective knowledge in the source domain model through continuous transfer learning, thereby achieving efficient target domain model optimization. That is, continuous transfer learning is performed between cross-domain models to achieve knowledge transfer from one domain that changes over time to another domain that also changes over time, and then a stable migration effect is ensured through continuous migration to adapt to the rapidly changing data distribution in the scenario.
[0037] Optionally, when implementing continuous transfer learning, a continuous cross-domain transfer training method is used in the target domain environment and source domain environment where the single domain model was originally deployed. Several target migration moments are set, and the target domain model and the source domain model are instructed to perform transfer learning at each target migration moment. In this way, continuous cross-domain transfer training is achieved based on multiple target migration moments.
[0038] Furthermore, because the source and target domain models continuously make predictions in their respective scenarios and continuously receive new data, this process may cause bias in the weights of some nodes in the model. Therefore, when transfer learning is required, the source and target domain models must be initialized. Before starting transfer learning training, the weights and biases of each node in the model must be initialized to ensure the stability of transfer learning training. In other words, at the target transfer time, the source and target domain models are initialized simultaneously.
[0039] It should be noted that the premise of transfer learning is that the source domain and the target domain meet the preset similarity conditions. If there are no similar and common features between the two domains for cross-domain transfer learning, negative transfer problems will occur, resulting in negative optimization of the model. Considering that the cross-domain scenarios that can be used for transfer learning can be of multiple similar types, in the embodiment of this specification, the preset similarity conditions include at least one of the user similarity conditions, the prediction task similarity conditions, and the scene attribute similarity conditions. Among them, the scene attribute (Content-level) similarity refers to the similar attributes of the prediction tasks (items) and / or users (users) in different domains. For example, the services of Amazon Music and Netflix are relatively similar. Although the two do not have many common prediction tasks and users, the attributes of the prediction tasks and users are similar; user (User) similarity refers to the fact that the two domains have a large number of the same users. For example, if two video platforms have many common users, then the two video platforms are user similar domains; the prediction task (Item) similarity condition refers to the fact that the two domains have a large number of the same products. Therefore, when the source domain and the target domain meet the preset similarity conditions, it means that the source domain and the target domain have reached a certain similarity, and effective transfer learning can be achieved.
[0040] S204 , based on the parameter transfer network corresponding to each target domain hidden layer in the target domain model, respectively obtain target source domain hidden parameters of each source domain hidden layer corresponding to each target domain hidden layer in the source domain model.
[0041] Optionally, it can be understood from the description of the above embodiments that in existing model training methods, if continuous transfer learning is to be achieved, it is necessary to re-fine-tune the model using the latest source domain model at regular intervals. However, due to the scene differences between the source domain and the target domain, a large number of samples are usually required to obtain a better result by fine-tuning the source domain model, which results in a very high training cost. In fact, such a training method is also difficult to deploy online. In addition, fine-tuning using a large number of samples may also cause the model to forget useful knowledge that has been retained for a long time, that is, catastrophic forgetting problem occurs, which means that transfer learning causes the model to be disturbed and lose some of its original performance.
[0042] Similarly, for the target domain model, although the model used in the recommendation scenario has a certain degree of common structural logic, the model details in different specific scenarios will vary greatly. Therefore, if the source domain model parameters are directly used to replace the parameters that have been learned in the original target domain, the useful knowledge acquired historically by the original target domain model will be discarded. This will put the cart before the horse and not only fail to achieve the expected transfer learning effect, but also lose the original effective performance.
[0043] Optionally, the original parameters accumulated by the source and target domain models during the single-domain continuous learning process preserve the effective knowledge acquired by the model through very long historical data learning. The loss or forgetting of this knowledge will directly affect the model performance. To avoid forgetting and discarding the knowledge acquired historically by the target and source domain models, all the parameters of the original source and target domain models are retained. An additional parameter transfer network can be used to connect the source and target domain models, and the parameter representation results in the intermediate hidden layer of the source domain model are mapped and used as additional knowledge for the target domain model. This process does not require the source domain model to be re-fine-tuned each time, and this additional knowledge will not affect the existing historical parameters of the target domain model. Unlike other methods that require backtracking data to readjust parameters to achieve continuous transfer learning, after using the parameter transfer network to achieve knowledge transfer, the source and target domain models only need to be updated with the incremental data of their respective domains, thereby achieving efficient and accurate continuous transfer learning.
[0044] Optionally, the parameter transfer network acts as an adapter between the source and target domain models. To enable the target domain model to directly accept some of the source model's parameters as additional knowledge, the source model's network features can be embedded in the original target domain's refined model, ensuring that the target and source domain models have the same network structure. This ensures that the target domain model can directly use the source domain parameters transferred by the parameter transfer network without affecting its original network parameter structure.
[0045] Furthermore, since the network structures of the target domain model and the source domain model correspond, the target domain model and the source domain model can form a double-tower structure, wherein each target domain hidden layer located in the middle of the target domain model structure has a corresponding source domain hidden layer in the source domain model. The network structure characteristics of the corresponding source domain hidden layer and target domain hidden layer are consistent, ensuring the accurate transfer of parameters between them. Considering that the dimensions, information content, and network characteristics of different hidden layers in the target domain model and the source domain model may be different, in order to more accurately complete the parameter transfer between hidden layers, an independent parameter transfer network can be used between each pair of target domain hidden layers and source domain hidden layers, so that each pair of source domain-target domain hidden layers has a dedicated parameter transfer network to achieve transfer learning. That is, when the target domain model performs transfer learning, based on the parameter transfer network corresponding to each target domain hidden layer in the target domain model, the target source domain hidden parameters of each target domain hidden layer corresponding to each source domain hidden layer in the source domain model are obtained respectively.
[0046] It can be understood that the target source domain hidden parameters obtained in this way are the parameter representation results in each source domain hidden layer of the source domain model that are useful to the target domain model, which contains the knowledge information required by the target domain model. Therefore, the target domain model no longer needs to obtain this part of knowledge through large-scale data training, which greatly saves training time and computing resources. In addition, in the process of acquiring additional knowledge, the original knowledge is protected, and effective knowledge transfer is achieved.
[0047] S206 , migrating each target source domain hidden parameter to each target domain hidden layer to obtain a target domain migration model, and training the target domain migration model based on target sample data corresponding to the target domain until the target domain migration model converges.
[0048] Optionally, after obtaining the target source domain hidden parameters in each source domain hidden layer in the source domain model through the parameter transfer network, each parameter transfer network can map the target source domain hidden parameters to the corresponding target domain hidden layers to obtain a target domain transfer model. At this time, the target domain transfer model is the combination of the original historical knowledge and the transfer knowledge from the source domain model. Such a network structure adopts an additive design, so that the dimension and structure of the hidden layer of the model will not change during the migration process. The target domain transfer model is completely initialized by the original target domain model, avoiding the random re-initialization of the hidden layer of the target domain model, which can ensure that the original performance of the model is not damaged to the greatest extent.
[0049] Furthermore, the target domain transfer model ensures knowledge validity. By incrementally training the target domain transfer model based solely on target sample data in the target domain until convergence, a target domain transfer model suitable for the target domain scenario can be directly obtained. Based on cross-domain transfer learning performed at multiple target transfer moments, the target domain transfer model maintains the original knowledge while absorbing new source domain knowledge, achieving excellent training results with only a small amount of incremental data from the current domain. This enables a hot start of the model, saving significant training time and computing resources for the target domain transfer model and preventing cross-interference in the streaming training between the source and target domain models.
[0050] Please refer to Figure 3, which is a schematic diagram of a model deployment for cross-domain transfer learning provided in an embodiment of this specification. As shown in Figure 3, before time t0+1, the source domain and target domain, which have been deployed as single-domain model training environments, each have a source domain model and a target domain model that are separately and continuously incrementally trained using only the supervision data of their respective scenarios; starting from time t0, a cross-domain transfer learning model training environment is deployed on the target domain, and a parameter migration network is added to migrate the parameters in the source domain model to the target domain model. Without forgetting the knowledge acquired historically, the target domain data will continue to be incrementally trained, while continuously migrating knowledge from the source domain model.
[0051] Specifically, the source domain model is represented as Φ S , source domain model Φ S The parameters at time t are Denote the target domain model as Φ T , target domain model Φ T The parameters at time t are The parameters Θ of the parameter transfer network can be used to calculate the loss value using a common loss function during model training. The parameters Θ of the parameter transfer network are also continuously fitted based on the loss function value during the transfer learning training process until convergence.
[0052] In an embodiment of the present specification, a method for training a transfer learning model is provided. At the target migration moment, a source domain model and a target domain model are initialized simultaneously, and the source domain and the target domain meet a preset similarity condition; based on a parameter migration network corresponding to each target domain hidden layer in the target domain model, target source domain hidden parameters of each target domain hidden layer corresponding to each source domain hidden layer in the source domain model are obtained respectively; each target source domain hidden parameter is migrated to each target domain hidden layer to obtain a target domain migration model, and the target domain migration model is trained based on target sample data corresponding to the target domain until the target domain migration model converges. When a certain similarity relationship is satisfied between the target domain and the source domain, the target source domain hidden parameters that reflect the deep network characteristics are selected from each source domain hidden layer of the source domain model through the parameter transfer network, and are transferred to the corresponding target domain hidden layer of the target domain model, so that the target domain hidden layer can directly obtain effective deep feature parameters without the need to retrain from the initial state. Moreover, the target domain transfer model obtained after parameter transfer can still be trained using the target sample data corresponding to the target domain, so that the target domain transfer model is suitable for completing the prediction task in the target domain. This not only saves a lot of training time and training computing resources of the target domain transfer model, but also prevents the source domain model and the target domain model from interfering with each other.
[0053] Please refer to Figure 4, which is a flow chart of a transfer learning model training method provided in an embodiment of this specification.
[0054] As shown in FIG4 , the transfer learning model training method may include at least the following steps.
[0055] S402. The source domain model is updated incrementally based on the source domain incremental data corresponding to the source domain. At the target migration time, the source domain model is initialized to a first updated state before the target migration time, and the target domain model is initialized to a second updated state before the target migration time.
[0056] Optionally, in an embodiment of the present specification, the target domain model obtains useful knowledge in the source domain model by continuously performing transfer learning from the source domain model. In order to adapt to the scenario environment of high-speed iteration of data in the target domain and the source domain, the source domain model will maintain a separate incremental update based on the source domain incremental data corresponding to the source domain. Then, in order to enable the target domain model to continuously learn the latest knowledge in the source domain model to adapt to the latest data distribution changes, at each target migration moment, when initializing the source domain model, the source domain model can be initialized to a first update state before the target migration moment. Similarly, the target domain model can also be initialized to a second update state before the target migration moment, where the first update state and the second update state can be the latest update state as of the current target migration moment, or can be other versions of the update state. The specific versions of the first update state and the second update state can be adaptively set according to the needs of the actual scenario.
[0057] Optionally, when setting multiple target migration moments, the target migration moments can be set to be triggered continuously at a preset period. The preset period can be set based on the scenario needs or the incremental training time of the source and target domain models. Preferably, transfer learning can be performed after the source domain model completes one incremental training, so that the target domain model can obtain the latest valid knowledge from the source domain model in real time or near real time.
[0058] S404 , based on the parameter transfer network corresponding to each target domain hidden layer in the target domain model, and based on the prediction task of the target domain model, determining the parameter characteristic conditions corresponding to each target domain hidden layer.
[0059] Optionally, for the target domain model, the ultimate goal of all its training is to complete the prediction task. Then, when performing knowledge transfer, the most needed network parameters are those that can help the target domain model complete the prediction task. Therefore, when determining the target source domain hidden parameters corresponding to each target domain hidden layer based on the parameter transfer network, it is hoped that the parameter transfer network can adaptively select more useful target source domain hidden parameters from the source domain hidden layer according to the prediction task.
[0060] Optionally, first, the parameter migration network corresponding to each target domain hidden layer needs to determine the parameter characteristic conditions corresponding to each target domain hidden layer based on the prediction task of the target domain model. After determining the parameter characteristics of the effective parameters, the target source domain hidden parameters can be further accurately selected from the source domain hidden layer. Specifically, a gating network can be used as a parameter migration network. The gating network can map all the intermediate hidden layers of the multi-layer perceptron (MLP) in the source domain model, especially the representation results of the high-order feature interaction information of the user and the prediction task contained in the deep layer of the source domain model MLP, to the target domain model, and add the results to the corresponding hidden layer of the target domain model. Using a gating network for parameter migration is different from the common method of only using the final score of the source domain model or only using some shallow representations, such as embedding feature representation (Embedding). Instead, it can realize the migration of deep representation information in the MLP through the linear layer of the gating network, and finally realize the migration and transformation of deep hidden information, which greatly improves the performance of the target domain model.
[0061] Furthermore, when the target domain model is required to focus on a certain type of feature when performing a prediction task, this can be used as the preset key feature aspect of the target domain model's prediction task, also known as the focus. For example, if the target domain model focuses on time, then it is hoped that the parameter transfer network can gain insight into the periodicity of historical data. In this case, an attention gating network can be used. The attention gating network can perceive the specified focus. For example, if the model is for the current day off, the attention mechanism can give more weight to the source domain model of historical days off. In this case, the attention gating network can focus on the time period of the historical source domain model, paying more attention to the historical source domain model of days off. The attention gating network adaptively adjusts the attention weight of the source domain model.
[0062] Specifically, when the parameter transfer network is an attention-gated network, each attention-gated network can determine the parameter characteristic conditions corresponding to each target domain hidden layer based on the prediction task of the target domain model and the preset key feature aspects of the prediction task. The preset key feature aspects are the feature aspects that meet the target importance conditions when the target domain model makes predictions, and the target importance conditions are a measure of the importance of the feature aspects to the prediction task. In addition, compared to ordinary gating networks, attention-gated networks can model the entire amount of historical data, which can further solve the forgetting problem in continuous learning and achieve better knowledge transfer effects.
[0063] S406 , obtaining target source domain hidden parameters that meet parameter feature conditions from each source domain hidden layer corresponding to each target domain hidden layer in the source domain model.
[0064] Optionally, after determining the parameter characteristic conditions corresponding to each target domain hidden layer, the target source domain hidden parameters that meet the parameter characteristic conditions can be obtained from the corresponding source domain hidden layers in the source domain model based on an attention gating network. These parameters are then transferred to the target domain model to obtain a target domain transfer model. Based on the attention gating network, adaptive feature selection of network characteristic parameters in the source domain model can be effectively implemented, transferring useful knowledge from the model while filtering out information that does not conform to the characteristics of the target domain scenario. Because the source domain model is continuously pre-trained using the latest source domain supervision data, the target domain model will also continuously load the newly updated source domain model parameters and keep them fixed during the backpropagation process. This ensures the efficient implementation of continuous transfer learning, allowing the target domain model to continuously learn the latest knowledge provided by the source domain model to adapt to the latest changes in the data distribution. At the same time, because the target domain model is trained only on target domain data, it is guaranteed to be unaffected by the source domain training objectives and does not require source domain data training at all, avoiding significant storage and computational overhead. Moreover, the attention network can actually focus on the preset key features of the prediction task, that is, pay attention to the focus of the prediction task, and focus on perceiving historical knowledge related to the focus. With the help of attention perception ability, it can more focusedly integrate historical knowledge during the training process of transfer learning.
[0065] Please refer to Figure 5, which is a schematic diagram of a continuous transfer learning model training provided by an embodiment of this specification. As shown in Figure 5, in the original single-domain continuous incremental training environment, the input of the source domain model is the user features, item features, and other features corresponding to the source domain, and there is a source domain hidden layer in the source domain model. Source domain hidden layer Source domain hidden layer The input of the target domain model is the user features, item features, and other features corresponding to the target domain, and there are target domain hidden layers w1, w2, and w3 in the target domain model; a continuous transfer learning environment is deployed between the target domain model and the source domain model, and an attention gate network g1 is used to transfer the hidden layers from the source domain to the target domain. The target source domain hidden parameters corresponding to the target domain hidden layer w1 are selected, and the obtained target source domain hidden parameters are mapped to the target domain hidden layer w1 as additional knowledge; similarly, the attention gate network g2 is used to map the target source domain hidden parameters from the source domain hidden layer The target source domain hidden parameters corresponding to the target domain hidden layer w2 are selected, and the obtained target source domain hidden parameters are mapped to the target domain hidden layer w2 as additional knowledge; the attention gate network g3 is used to map the target source domain hidden parameters from the source domain hidden layer The target source domain hidden parameters corresponding to the target domain hidden layer w3 are selected, and the obtained target source domain hidden parameters are mapped to the target domain hidden layer w3 as additional knowledge; finally, the target domain migration model obtained by receiving the target source domain hidden parameters is used for training and convergence of prediction tasks for the target domain, and deployed in the actual environment after convergence.
[0066] S408 , migrating each target source domain hidden parameter to each target domain hidden layer to obtain a target domain migration model, and training the target domain migration model based on target sample data corresponding to the target domain until the target domain migration model converges.
[0067] Regarding step S408 , please refer to the detailed description in step S206 , which will not be repeated here.
[0068] In an embodiment of the present specification, a transfer learning model training method is provided. Through the linear layer of the gated network, the transfer of deep representation information is realized, and ultimately the transfer transformation of deep hidden information is realized, which greatly improves the performance of the target domain model. The source domain model continuously uses the latest source domain supervision data for continuous pre-training, and the target domain model will also continuously load the latest updated source domain model parameters and keep them fixed during the backpropagation process, ensuring the efficient implementation of continuous transfer learning, so that the target domain model continuously learns the latest knowledge provided by the source domain model to adapt to the latest data distribution changes. In this case, the streaming training of the source domain model and the target domain model will not interfere with each other. The target domain model is only trained on the target domain data, ensuring that the model is not affected by the source domain training objectives and does not require source domain data training at all, avoiding a large amount of storage and computing overhead. The attention network can actually perceive historical knowledge related to the preset key features of the prediction task based on the preset key features. With the help of attention perception capabilities, the historical knowledge can be more focused on integrating during the transfer learning training process.
[0069] Please refer to Figure 6, which is a block diagram of a transfer learning model training device provided in an embodiment of this specification. As shown in Figure 6, the transfer learning model training device 600 includes: an initialization module 610, which is used to initialize the source domain model and the target domain model at the target migration moment, so that the source domain and the target domain meet the preset similarity condition; a parameter migration module 620, which is used to obtain the target source domain hidden parameters of each target domain hidden layer corresponding to each source domain hidden layer in the source domain model based on the parameter migration network corresponding to each target domain hidden layer in the target domain model; a target domain training module 630, which is used to migrate each target source domain hidden parameter to each target domain hidden layer to obtain a target domain migration model, and train the target domain migration model based on the target sample data corresponding to the target domain until the target domain migration model converges.
[0070] Optionally, the parameter migration module 620 is further used to determine the parameter characteristic conditions corresponding to each target domain hidden layer based on the prediction task of the target domain model; and obtain the target source domain hidden parameters that meet the parameter characteristic conditions from each source domain hidden layer corresponding to each target domain hidden layer in the source domain model.
[0071] Optionally, the parameter migration module 620 is also used to determine the parameter feature conditions corresponding to each target domain hidden layer based on the prediction task of the target domain model and the preset key feature aspects of the prediction task when the parameter migration network is an attention gated network. The preset key feature aspects are feature aspects that meet the target importance conditions when the target domain model makes predictions.
[0072] Optionally, the target migration moment is triggered continuously according to a preset period.
[0073] Optionally, the initialization module 610 is further used to initialize the source domain model to a first update state before the target migration time when the source domain model maintains a separate incremental update based on the source domain incremental data corresponding to the source domain, and to initialize the target domain model to a second update state before the target migration time.
[0074] Optionally, the preset similarity condition includes at least one of a user similarity condition, a prediction task similarity condition, and a scene attribute similarity condition.
[0075] In an embodiment of the present specification, a transfer learning model training device is provided, wherein an initialization module is used to simultaneously initialize a source domain model and a target domain model at a target migration moment, so that the source domain and the target domain meet a preset similarity condition; a parameter migration module is used to obtain target source domain hidden parameters of each target domain hidden layer corresponding to each source domain hidden layer in the source domain model based on a parameter migration network corresponding to each target domain hidden layer in the target domain model; a target domain training module is used to migrate each target source domain hidden parameter to each target domain hidden layer to obtain a target domain migration model, and train the target domain migration model based on target sample data corresponding to the target domain until the target domain migration model converges. When a certain similarity relationship is satisfied between the target domain and the source domain, the target source domain hidden parameters that reflect the deep network characteristics are selected from each source domain hidden layer of the source domain model through the parameter transfer network, and are transferred to the corresponding target domain hidden layer of the target domain model, so that the target domain hidden layer can directly obtain effective deep feature parameters without the need to retrain from the initial state. Moreover, the target domain transfer model obtained after parameter transfer can still be trained using the target sample data corresponding to the target domain, so that the target domain transfer model is suitable for completing the prediction task in the target domain. This not only saves a lot of training time and training computing resources of the target domain transfer model, but also prevents the source domain model and the target domain model from interfering with each other.
[0076] The embodiments of this specification provide a computer program product including instructions. When the computer program product is run on a computer or a processor, the computer or the processor is caused to perform the steps of any one of the methods in the above embodiments.
[0077] The embodiments of this specification also provide a computer storage medium, which can store multiple instructions, and the instructions are suitable for being loaded by a processor and executing the steps of any method in the above embodiments.
[0078] Please refer to Figure 7, which is a schematic diagram of the structure of a terminal provided in an embodiment of this specification. As shown in Figure 7, the terminal 700 may include: at least one terminal processor 701, at least one network interface 704, a user interface 703, a memory 705, and at least one communication bus 702.
[0079] The communication bus 702 is used to implement the connection and communication between these components.
[0080] The user interface 703 may include a display screen (Display) and a camera (Camera). Optionally, the user interface 703 may also include a standard wired interface and a wireless interface.
[0081] The network interface 704 may optionally include a standard wired interface or a wireless interface (such as a WI-FI interface).
[0082] The terminal processor 701 may include one or more processing cores. The terminal processor 701 utilizes various interfaces and circuits to connect various components within the terminal 700. It executes instructions, programs, code sets, or instruction sets stored in the memory 705, and accesses data stored in the memory 705 to perform various functions and process data for the terminal 700. Optionally, the terminal processor 701 may be implemented using at least one hardware form factor selected from the group consisting of a digital signal processing (DSP), a field-programmable gate array (FPGA), and a programmable logic array (PLA). The terminal processor 701 may integrate one or a combination of a central processing unit (CPU), a graphics processing unit (GPU), and a modem. The CPU primarily processes the operating system, user interface, and application programs; the GPU is responsible for rendering and drawing the content displayed on the display screen; and the modem handles wireless communications. It is understood that the modem may also be implemented independently of the terminal processor 701 via a separate chip.
[0083] Among them, the memory 705 may include a random access memory (RAM) or a read-only memory (ROM). Optionally, the memory 705 includes a non-transitory computer-readable storage medium. The memory 705 can be used to store instructions, programs, codes, code sets or instruction sets. The memory 705 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function (such as a touch function, a sound playback function, an image playback function, etc.), instructions for implementing the above-mentioned various method embodiments, etc.; the data storage area may store data involved in the above-mentioned various method embodiments, etc. The memory 705 may also be at least one storage device located away from the aforementioned terminal processor 701. As shown in Figure 7, the memory 705 as a computer storage medium may include an operating system, a network communication module, a user interface module and a transfer learning model training program.
[0084] In the terminal 700 shown in Figure 7, the user interface 703 is mainly used to provide an input interface for the user and obtain the data input by the user; and the terminal processor 701 can be used to call the transfer learning model training program stored in the memory 705, and specifically perform the following operations: at the target migration moment, initialize the source domain model and the target domain model at the same time, and the source domain and the target domain meet the preset similarity condition; based on the parameter migration network corresponding to each target domain hidden layer in the target domain model, obtain the target source domain hidden parameters of each source domain hidden layer corresponding to the source domain hidden layer in the source domain model; migrate each target source domain hidden parameter to each target domain hidden layer to obtain a target domain migration model, and train the target domain migration model based on the target sample data corresponding to the target domain until the target domain migration model converges.
[0085] In some embodiments, when the terminal processor 701 executes the steps of respectively obtaining the target source domain hidden parameters of each source domain hidden layer corresponding to each target domain hidden layer in the source domain model, it specifically performs the following steps: determining the parameter characteristic conditions corresponding to each target domain hidden layer based on the prediction tasks of the target domain model; and obtaining the target source domain hidden parameters that meet the parameter characteristic conditions from each source domain hidden layer corresponding to each target domain hidden layer in the source domain model.
[0086] In some embodiments, when the parameter transfer network is an attention gated network, the terminal processor 701 performs the following steps when executing prediction tasks based on the target domain model and determining the parameter feature conditions corresponding to each target domain hidden layer: determining the parameter feature conditions corresponding to each target domain hidden layer based on the prediction task of the target domain model and the preset key feature aspects of the prediction task, respectively. The preset key feature aspects are feature aspects that meet the target importance conditions when the target domain model makes predictions.
[0087] In some embodiments, the target migration moment is triggered continuously according to a preset period.
[0088] In some embodiments, when the source domain model maintains a separate incremental update based on the source domain incremental data corresponding to the source domain, the terminal processor 701, when performing the simultaneous initialization of the source domain model and the target domain model, specifically performs the following steps: initializing the source domain model to a first update state before the target migration moment, and initializing the target domain model to a second update state before the target migration moment.
[0089] In some embodiments, the preset similarity condition includes at least one of a user similarity condition, a prediction task similarity condition, and a scene attribute similarity condition.
[0090] In the several embodiments provided in this specification, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of modules is merely a logical function division. In actual implementation, there may be other division methods, such as multiple modules or components can be combined or integrated into another system, or some features can be ignored or not executed. In addition, the mutual coupling or direct coupling or communication connection shown or discussed can be through some interface, indirect coupling or communication connection of devices or modules, and can be electrical, mechanical or other forms.
[0091] Modules described as separate components may or may not be physically separate, and components shown as modules may or may not be physical modules, that is, they may be located in one place or distributed across multiple network modules. Some or all of these modules may be selected to achieve the purpose of this embodiment based on actual needs.
[0092] In the above embodiments, it can be implemented in whole or in part by software, hardware, firmware or any combination thereof. When software is used for implementation, it can be implemented in whole or in part in the form of a computer program product. The above-mentioned computer program product includes one or more computer instructions. When the above-mentioned computer program instructions are loaded and executed on a computer, the above-mentioned process or function according to the embodiment of this specification is generated in whole or in part. The above-mentioned computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The above-mentioned computer instructions can be stored in a computer-readable storage medium or transmitted by the above-mentioned computer-readable storage medium. The above-mentioned computer instructions can be transmitted from a website, computer, server or data center to another website, computer, server or data center by wired (such as coaxial cable, optical fiber, digital subscriber line (DSL)) or wireless (such as infrared, wireless, microwave, etc.). The above-mentioned computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that includes one or more available media integrations. The available media may be magnetic media (eg, floppy disks, hard disks, tapes), optical media (eg, digital versatile discs (DVDs)), or semiconductor media (eg, solid state disks (SSDs)).
[0093] It should be noted that for the aforementioned method embodiments, for the sake of simplicity of description, they are all expressed as a series of action combinations, but those skilled in the art should be aware that the embodiments of this specification are not limited by the order of the actions described, because according to the embodiments of this specification, certain steps can be performed in other orders or simultaneously. Secondly, those skilled in the art should also be aware that the embodiments described in this specification are all preferred embodiments, and the actions and modules involved are not necessarily required by the embodiments of this specification.
[0094] In addition, it should be noted that the information (including but not limited to user device information, user personal information, etc.), data (including but not limited to data used for analysis, stored data, displayed data, etc.) and signals involved in the embodiments of this specification are all authorized by the user or fully authorized by all parties, and the collection, use and processing of relevant data must comply with the relevant laws, regulations and standards of relevant countries and regions.
[0095] In the above embodiments, the description of each embodiment has its own focus. For parts that are not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.
[0096] The above is a description of a transfer learning model training method, device, storage medium, and terminal provided in the embodiments of this specification. For technical personnel in this field, based on the ideas of the embodiments of this specification, there may be changes in the specific implementation methods and application scopes. In summary, the content of this specification should not be understood as a limitation on the embodiments of this specification.
Claims
1. A transfer learning model training method, the method comprising: At the target migration moment, the source domain model and the target domain model are initialized simultaneously, and the source domain and the target domain meet a preset similarity condition; Based on the parameter transfer network corresponding to each target domain hidden layer in the target domain model, respectively obtaining target source domain hidden parameters of each source domain hidden layer corresponding to each target domain hidden layer in the source domain model; The target domain migration model is obtained by migrating the hidden parameters of each target source domain to the hidden layers of each target domain, and the target domain migration model is trained based on the target sample data corresponding to the target domain until the target domain migration model converges.
2. According to the method of claim 1, the step of respectively obtaining target source domain hidden parameters of each source domain hidden layer corresponding to each target domain hidden layer in the source domain model comprises: Determining parameter characteristic conditions corresponding to each target domain hidden layer based on the prediction task of the target domain model respectively; Obtain target source domain hidden parameters that meet the parameter characteristic condition from the source domain hidden layers corresponding to the target domain hidden layers in the source domain model.
3. According to the method of claim 2, the parameter transfer network is an attention gated network, and the parameter characteristic conditions corresponding to each target domain hidden layer are determined based on the prediction tasks of the target domain model, including: The parameter characteristic conditions corresponding to each target domain hidden layer are determined based on the prediction task of the target domain model and the preset key feature aspects of the prediction task, respectively. The preset key feature aspects are feature aspects that meet the target importance conditions when the target domain model makes predictions.
4. According to the method of claim 1, the target migration time is continuously triggered according to a preset period.
5. The method according to claim 1, wherein the source domain model is updated incrementally based on the source domain incremental data corresponding to the source domain, and the simultaneous initialization of the source domain model and the target domain model comprises: The source domain model is initialized to a first updated state before the target migration time, and the target domain model is initialized to a second updated state before the target migration time.
6. According to the method of claim 1, the preset similarity condition includes at least one of a user similarity condition, a prediction task similarity condition, and a scene attribute similarity condition.
7. A transfer learning model training device, the device comprising: An initialization module, used to initialize a source domain model and a target domain model at the target migration moment, wherein the source domain and the target domain meet a preset similarity condition; A parameter migration module, used to obtain target source domain hidden parameters of each source domain hidden layer corresponding to each target domain hidden layer in the source domain model based on a parameter migration network corresponding to each target domain hidden layer in the target domain model; The target domain training module is used to transfer the hidden parameters of each target source domain to each target domain hidden layer to obtain the target domain A target domain migration model is provided, and the target domain migration model is trained based on the target sample data corresponding to the target domain until the target domain migration model converges.
8. A computer program product comprising instructions, which, when executed on a computer or a processor, causes the computer or the processor to execute the steps of the method according to any one of claims 1 to 6.
9. A computer storage medium storing a plurality of instructions, wherein the instructions are suitable for being loaded by a processor and executing the steps of the method according to any one of claims 1 to 6.
10. A terminal comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the steps of the method according to any one of claims 1 to 6 when executing the program.
Citation Information
Patent Citations
Data center energy efficiency optimization method based on transfer learning
CN110703899A
Distributed deep learning reasoning deployment method for protecting data privacy
CN111832729A
Pruning method and device for neural network model
CN112749797A
Method and device for training transfer learning model and recommendation model
CN113222073A
Hyper-parameter adaptive optimization system and method for automatic machine learning
CN113392983A
Cited By
Domain adaptation width learning industrial penicillin concentration prediction method based on parameter migration
CN120600134A
Mobile multi-agent knowledge migration method based on replacement strategy network
CN120893518A
Transfer learning model and recommendation model training method and device
CN121051468A
Tire vulcanization quality prediction method and system based on time sequence characteristic migration
CN121117624A
Nuclear security supervision data migration method and device, computer equipment, storage medium and computer program product
CN121542247A