Model training method and device and related equipment
By introducing adaptive temperature regulation function and domain difference function in knowledge distillation method, the limitations of existing methods in temperature regulation and cross-domain adaptability are solved, and the performance of student models in cross-domain tasks is significantly improved.
Patent Information
- Application Number
- CN202411997200.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-12-31
- Publication Date
- 2025-05-06
AI Technical Summary
The existing knowledge distillation method has limitations in temperature parameter regulation and adaptability of cross-domain tasks, and it is difficult to meet the needs of complex application scenarios.
By constructing adaptive temperature regulation functions and domain differences functions, dynamically adjust temperature parameters and reduce data feature distribution differences between source and target areas to optimize training of student models.
The student model quickly captures the overall characteristics of the teacher model in the early stage and refines the learning in the later stage, improving the performance in cross-domain tasks and seamlessly transfers the knowledge of the teacher model to the target field.
Smart Images

Figure CN119940423A_ABST
Abstract
Description
Background Art
[0002] In the current field of artificial intelligence and deep learning, with the increase in model size, especially pre-trained models such as Bidirectional Encoder Representations from Transformers (BERT) and Generative Pre-trained Transformer (GPT), although these models perform well, their high computing and storage costs make it difficult to deploy in many application scenarios. Therefore, knowledge distillation is gradually being used as an effective model compression method. By transferring the knowledge of large-scale teacher models to smaller student models, the consumption of computing resources can be significantly reduced while ensuring accuracy. However, existing knowledge distillation methods have certain limitations in temperature parameter adjustment and adaptability to cross-domain tasks, and it is difficult to meet the needs of complex application scenarios. Therefore, optimizing the temperature adjustment mechanism and domain adaptability of knowledge distillation has become the key to improving the training efficiency and transfer learning ability of student models.
[0003] The existing technology mainly has the following problems:
[0004] 1) Limitations of fixed temperature adjustment: Existing knowledge distillation techniques usually use fixed temperature parameters to smooth the output of the teacher model. However, different training stages have different requirements for temperature, and fixed temperature cannot effectively adapt to the changes in model requirements during training, making it difficult for the student model to obtain sufficient information in the early stages, and it is also difficult to finely imitate the output of the teacher model in the later stages.
[0005] 2) Insufficient adaptability to cross-domain tasks: In many application scenarios, there are significant differences between the source domain data and the target domain data. Existing knowledge distillation methods cannot effectively deal with such differences, resulting in poor performance of student models in cross-domain tasks and difficulty in achieving efficient knowledge transfer.
[0006] It should be noted that the information disclosed in the above background technology section is only used to enhance the understanding of the background of the present disclosure, and therefore may include information that does not constitute the prior art known to ordinary technicians in the field. Summary of the invention
[0007] The present disclosure provides a model training method, apparatus and related equipment, which at least to a certain extent overcome the problem that the knowledge distillation method in the related art has limitations in temperature parameter adjustment and adaptability to cross-domain tasks.
[0008] Other features and advantages of the present disclosure will become apparent from the following detailed description, or may be learned in part by the practice of the present disclosure.
[0009] According to one aspect of the present disclosure, a model training method is provided, including: obtaining a source domain dataset and a target domain dataset; training a first deep learning model based on the source domain dataset to obtain a teacher model; constructing a second deep learning model corresponding to the teacher model through the target domain dataset; in the process of training the second deep learning model, optimizing the second deep learning model based on a pre-constructed target loss function until a student model that meets the conditions is obtained; wherein the target loss function is constructed based on an adaptive temperature adjustment function and a domain difference function, the adaptive temperature adjustment function is used to dynamically adjust the temperature parameter according to the model loss or gradient change of the second deep learning model during the training process, and the domain difference function is used to minimize the difference in data feature distribution between the source domain and the target domain.
[0010] In some exemplary embodiments of the present disclosure, based on the aforementioned scheme, in the process of training the second deep learning model, the second deep learning model is optimized based on a pre-constructed target loss function until a student model that meets the conditions is obtained, and the method also includes: constructing the adaptive temperature adjustment function based on the initial temperature of the second deep learning model, a preset temperature decay rate, and a preset training time step of the second deep learning model; performing adversarial learning on the output features of the source domain and the target domain through a domain discriminator, or aligning the output distributions of the source domain and the target domain through a maximum mean difference method to obtain the domain difference function.
[0011] In some exemplary embodiments of the present disclosure, based on the above scheme, based on the initial temperature of the second deep learning model, the preset temperature decay rate and the preset training time step of the second deep learning model, the adaptive temperature adjustment function is constructed, including: the adaptive temperature adjustment function is obtained by the following formula Wherein, T0 represents the initial temperature of the second deep learning model; α represents the hyperparameter of the temperature decay rate; and t represents the training time step of the second deep learning model.
[0012] In some exemplary embodiments of the present disclosure, based on the aforementioned scheme, the objective loss function also includes: distillation loss. During the training of the second deep learning model, the second deep learning model is optimized based on a pre-constructed objective loss function until a student model that meets the conditions is obtained. The method also includes: calculating the difference between the output of the second deep learning model and the teacher model through the soft label of the teacher model and the output result of the second deep learning model, and determining the distillation loss.
[0013] In some exemplary embodiments of the present disclosure, based on the aforementioned solution, in the process of training the second deep learning model, based on the pre-constructed target loss function, the second deep learning model is optimized until a student model that meets the conditions is obtained, the method further includes: obtaining the target loss function L by the following formula total :L total =L distill +λ1×L adv , where L distill represents the knowledge distillation loss; L adv represents the adversarial loss; λ1 represents the hyperparameter that balances the distillation loss and the adversarial loss.
[0014] In some exemplary embodiments of the present disclosure, based on the aforementioned solution, in the process of training the second deep learning model, based on the pre-constructed target loss function, the second deep learning model is optimized until a student model that meets the conditions is obtained, the method further includes: obtaining the target loss function L by the following formula total :L total =L distill +λ2×L MMD , where λ2 represents the hyperparameter that balances the distillation loss and the maximum mean difference loss; L MMD Denotes the maximum mean difference loss.
[0015] In some exemplary embodiments of the present disclosure, based on the aforementioned scheme, during the process of training the second deep learning model, the second deep learning model is optimized based on a pre-constructed target loss function until a student model that meets the conditions is obtained. The method also includes: inputting preset verification data into the student model, and optimizing and adjusting the adaptive temperature adjustment function and the domain difference function according to the output verification result.
[0016] According to another aspect of the present disclosure, a model training device is also provided, including: a data set acquisition module, used to acquire a source domain data set and a target domain data set; a teacher model determination module, used to train a first deep learning model based on the source domain data set to obtain a teacher model; a second deep learning model construction module, used to construct a second deep learning model corresponding to the teacher model through the target domain data set; a student model determination module, used to optimize the second deep learning model based on a pre-constructed target loss function during the training of the second deep learning model until a student model that meets the conditions is obtained; wherein the target loss function is constructed based on an adaptive temperature adjustment function and a domain difference function, the adaptive temperature adjustment function is used to dynamically adjust the temperature parameter according to the model loss or gradient change of the second deep learning model during the training process, and the domain difference function is used to minimize the difference in data feature distribution between the source domain and the target domain.
[0017] According to another aspect of the present disclosure, an electronic device is provided, comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute any one of the above-mentioned model training methods by executing the executable instructions.
[0018] According to another aspect of the present disclosure, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, any one of the above-mentioned model training methods is implemented.
[0019] According to another aspect of the present disclosure, a computer program product is also provided, including: a computer program or instructions, which implement any of the above-mentioned model training methods when executed by a processor.
[0020] A model training method, apparatus and related equipment provided in the embodiments of the present disclosure can adjust the temperature parameters in real time according to the loss changes or gradient trends of the model during the training process through pre-built adaptive temperature adjustment functions and domain difference functions, so that the student model can more accurately imitate the teacher model; and the embodiments of the present disclosure reduce the data feature distribution differences between the source domain and the target domain through the domain difference function, ensuring that the student model can effectively learn domain-independent features. This mechanism can seamlessly transfer the knowledge of the teacher model to the target domain, significantly improving the performance of the student model in cross-domain tasks.
[0021] It is to be understood that the foregoing general description and the following detailed description are exemplary and explanatory only and are not restrictive of the present disclosure. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The accompanying drawings herein are incorporated into the specification and constitute a part of the specification, illustrate embodiments consistent with the present disclosure, and together with the specification are used to explain the principles of the present disclosure. Obviously, the accompanying drawings described below are only some embodiments of the present disclosure, and for ordinary technicians in this field, other accompanying drawings can be obtained based on these accompanying drawings without creative work.
[0023] Figure 1 A schematic diagram showing an exemplary application system architecture of a model training method in an embodiment of the present disclosure;
[0024] Figure 2 A schematic diagram of a model training method in an embodiment of the present disclosure is shown;
[0025] Figure 3 A schematic diagram of an adaptive temperature adjustment mechanism in an embodiment of the present disclosure is shown;
[0026] Figure 4 A schematic diagram of the overall training process of a model in an embodiment of the present disclosure is shown;
[0027] Figure 5 A schematic diagram of a model training device in an embodiment of the present disclosure is shown;
[0028] Figure 6 A schematic diagram of an electronic device to which a model training method is applied in an embodiment of the present disclosure is shown. DETAILED DESCRIPTION
[0029] Example embodiments will now be described more fully with reference to the accompanying drawings. However, example embodiments can be implemented in a variety of forms and should not be construed as limited to the examples set forth herein; rather, these embodiments are provided so that the disclosure will be more comprehensive and complete and to fully convey the concepts of the example embodiments to those skilled in the art. The described features, structures, or characteristics may be combined in any suitable manner in one or more embodiments.
[0030] In addition, the described features, structures or characteristics may be combined in one or more embodiments in any suitable manner. In the following description, many specific details are provided to provide a full understanding of the embodiments of the present disclosure. However, those skilled in the art will appreciate that the technical solutions of the present disclosure may be practiced without one or more of the specific details, or other methods, components, devices, steps, etc. may be adopted. In other cases, known methods, devices, implementations or operations are not shown or described in detail to avoid blurring the various aspects of the present disclosure.
[0031] The flowcharts shown in the accompanying drawings are only exemplary and do not necessarily include all the contents and operations / steps, nor must they be executed in the order described. For example, some operations / steps can be decomposed, and some operations / steps can be combined or partially combined, so the actual execution order may change according to actual conditions.
[0032] For ease of understanding, the terms involved in this disclosure are first explained as follows:
[0033] Knowledge distillation: It is a model compression and optimization technique that reduces the complexity and computational overhead of the model by transferring knowledge from a large teacher model to a small student model while retaining the performance of the large model as much as possible.
[0034] Figure 1 FIG. 1 shows an exemplary application system architecture diagram to which the model training method in the embodiment of the present disclosure can be applied. Figure 1 As shown, the system architecture may include a terminal device 101 , a network 102 and a server 103 .
[0035] The network 102 is a medium for providing a communication link between the terminal device 101 and the server 103, and may be a wired network or a wireless network.
[0036] Optionally, the wireless network or wired network described above uses standard communication technology and / or protocol. The network is usually the Internet, but it can also be any network, including but not limited to a local area network (LAN), a metropolitan area network (MAN), a wide area network (WAN), a mobile, wired or wireless network, a dedicated network or any combination of a virtual private network). In some embodiments, technologies and / or formats including Hyper Text Mark-up Language (HTML), Extensible Markup Language (XML), etc. are used to represent data exchanged through the network. In addition, conventional encryption technologies such as Secure Socket Layer (SSL), Transport Layer Security (TLS), Virtual Private Network (VPN), Internet Protocol Security (IPsec) can also be used to encrypt all or some links. In other embodiments, customized and / or dedicated data communication technologies can also be used to replace or supplement the above data communication technologies.
[0037] The terminal device 101 can be various electronic devices, including but not limited to smart phones, tablet computers, laptop computers, desktop computers, wearable devices, augmented reality devices, virtual reality devices, etc.
[0038] Optionally, the client of the application installed in different terminal devices 101 is the same, or the client of the same type of application based on different operating systems. Based on the different terminal platforms, the specific form of the client of the application can also be different, for example, the application client can be a mobile client, a PC client, etc.
[0039] The server 103 may be a server that provides various services, such as a background management server that provides support for the device operated by the user using the terminal device 301. The background management server may analyze and process the received request and other data, and feed back the processing results to the terminal device.
[0040] Optionally, the server can be an independent physical server, or a server cluster or distributed system composed of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. The terminal can be a smart phone, tablet computer, laptop computer, desktop computer, smart speaker, smart watch, etc., but is not limited to this. The terminal and the server can be directly or indirectly connected via wired or wireless communication, which is not limited in this application.
[0041] Those skilled in the art will know that Figure 1 The number of terminal devices, networks and servers in the embodiment is only for illustration, and any number of terminal devices, networks and servers may be provided according to actual needs, and the embodiments of the present disclosure do not limit this.
[0042] Under the above system architecture, a model training method is provided in an embodiment of the present disclosure, which can be executed by any electronic device with computing and processing capabilities.
[0043] In some embodiments, the model training method provided in the embodiments of the present disclosure can be executed by a terminal device of the above-mentioned system architecture; in other embodiments, the model training method provided in the embodiments of the present disclosure can be executed by a server in the above-mentioned system architecture; in other embodiments, the model training method provided in the embodiments of the present disclosure can be implemented by the terminal device and the server in the above-mentioned system architecture through interaction.
[0044] It should be noted that the embodiments of the present disclosure can be used in the following scenarios:
[0045] Edge computing and the Internet of Things: On edge devices, due to limited resources, the deployment of models needs to be more lightweight. The model training method proposed in the embodiment of the present disclosure can effectively compress large models while ensuring their reasoning performance on edge devices. It is suitable for smart homes, autonomous driving, smart monitoring and other fields.
[0046] Cross-language natural language processing: Due to the significant differences between data in different languages, traditional models are difficult to effectively perform cross-language migration. In the embodiments of the present disclosure, domain difference functions can be used for cross-language natural language processing tasks, realizing transfer learning and deployment of multi-language models, and improving the performance of applications such as translation and text generation.
[0047] Figure 2 A schematic diagram of a model training method in an embodiment of the present disclosure is shown, and the method includes the following steps:
[0048] S202: Acquire a source domain dataset and a target domain dataset.
[0049] It should be noted that the source domain datasets in the embodiments of the present disclosure are two key datasets used in transfer learning and domain adaptation tasks. The source domain refers to the environment, conditions or context in which the original data is collected and annotated. The source domain dataset is a dataset used to train the initial model, which usually has rich label information and can represent the typical characteristics and patterns of the source domain; the target domain refers to the new environment or scene in which the model is expected to be applied, which may be different from the source domain. The target domain dataset is a dataset used to evaluate the performance of the model in the new domain, and is sometimes also used to fine-tune the model to adapt to the new environment. For example, the embodiments of the present disclosure take a large number of photos of objects in an indoor environment and mark each object with an accurate bounding box and category label. The indoor environment is the source domain, and these photos constitute the source domain dataset in the embodiments of the present disclosure. If the user wants to apply this model to outdoor scenes, such as identifying vehicles and pedestrians on the street, then the target domain is the outdoor scene. In this case, it is necessary to collect labeled outdoor scene pictures as the target domain dataset.
[0050] S204: Train the first deep learning model based on the source domain dataset to obtain a teacher model.
[0051] It should be noted that the teacher model in the embodiment of the present disclosure refers to a model that has been trained and has strong performance. Its main function is to guide and help another smaller or simpler model (called the student model) to learn. In this way, the knowledge of the teacher model can be passed on to the student model, enabling the latter to achieve better performance while maintaining high efficiency.
[0052] S206, constructing a second deep learning model corresponding to the teacher model through the target domain dataset.
[0053] It should be noted that the second deep learning model in the embodiment of the present disclosure is a prototype of the student model, and the subsequent student model is obtained by processing the second deep learning model.
[0054] S208. During the training of the second deep learning model, the second deep learning model is optimized based on a pre-constructed target loss function until a student model that meets the conditions is obtained, wherein the target loss function is constructed based on an adaptive temperature adjustment function and a domain difference function, and the adaptive temperature adjustment function is used to dynamically adjust the temperature parameters according to the model loss or gradient change of the second deep learning model during the training process, and the domain difference function is used to minimize the difference in data feature distribution between the source domain and the target domain.
[0055] It should be noted that the embodiment of the present disclosure does not use a fixed temperature parameter to smooth the output of the teacher model, but uses an adaptive temperature adjustment function to adjust the temperature parameter in real time according to the loss change or gradient trend of the model during training, so that the student model can capture the overall characteristics of the teacher model through a higher temperature in the early stage, and gradually lower the temperature in the later stage to achieve refined learning. This method effectively improves the learning efficiency of the student model and imitates the teacher model more accurately.
[0056] Furthermore, the embodiments of the present disclosure also reduce the feature distribution differences between the source domain and the target domain through domain difference functions in cross-domain tasks, ensuring that the student model can effectively learn domain-independent features. This mechanism enables the knowledge of the teacher model to be seamlessly transferred to the target domain, significantly improving the performance of the student model in cross-domain tasks.
[0057] The model training method provided in the embodiments of the present disclosure first obtains a source domain dataset and a target domain dataset; secondly, trains a first deep learning model based on the source domain dataset to obtain a teacher model; then, constructs a second deep learning model corresponding to the teacher model through the target domain dataset; finally, in the process of training the second deep learning model, optimizes the second deep learning model based on a pre-constructed target loss function until a student model that meets the conditions is obtained, wherein the target loss function is constructed based on an adaptive temperature adjustment function and a domain difference function, the adaptive temperature adjustment function is used to dynamically adjust the temperature parameters according to the model loss or gradient change of the second deep learning model during the training process, and the domain difference function is used to minimize the difference in data feature distribution between the source domain and the target domain. Compared with the knowledge distillation method in the related art, which has certain limitations in temperature parameter adjustment and adaptability to cross-domain tasks and is difficult to meet the needs of complex application scenarios, the embodiment of the present disclosure uses a pre-built adaptive temperature adjustment function and domain difference function to adjust the temperature parameters in real time according to the loss change or gradient trend of the model during training, so that the student model can more accurately imitate the teacher model; and the embodiment of the present disclosure uses the domain difference function to reduce the data feature distribution difference between the source domain and the target domain, ensuring that the student model can effectively learn domain-independent features. This mechanism enables the knowledge of the teacher model to be seamlessly transferred to the target domain, significantly improving the performance of the student model in cross-domain tasks.
[0058] In some embodiments, in the process of training the second deep learning model, the embodiment of the present disclosure optimizes the second deep learning model based on a pre-constructed target loss function until a student model that meets the conditions is obtained. The model training method in the embodiment of the present disclosure also includes: constructing an adaptive temperature adjustment function based on the initial temperature of the second deep learning model, a preset temperature decay rate, and a preset training time step of the second deep learning model; performing adversarial learning on the output features of the source domain and the target domain through a domain discriminator, or aligning the output distributions of the source domain and the target domain through a maximum mean difference method to obtain a domain difference function. Specifically, the embodiment of the present disclosure not only solves the limitations of the fixed temperature parameters in the prior art by constructing an adaptive temperature adjustment function and a domain difference function, but also enhances the adaptability of knowledge distillation in cross-domain tasks, so that large-scale models can be quickly deployed in more practical applications.
[0059] In some embodiments, the embodiments of the present disclosure construct an adaptive temperature adjustment function based on the initial temperature of the second deep learning model, the preset temperature decay rate, and the preset training time step of the second deep learning model, including: obtaining the adaptive temperature adjustment function T by formula (1): (t) :
[0060]
[0061] Among them, T0 represents the initial temperature of the second deep learning model; α represents the hyperparameter of the temperature decay rate; t represents the training time step of the second deep learning model.
[0062] In some embodiments, in the traditional knowledge distillation process, the teacher model usually generates soft labels (SoftTargets) and smoothes the output probability distribution through a fixed temperature parameter. However, the fixed temperature may not be suitable at different training stages. In the early stage of training, a higher temperature can help the student model better capture the overall pattern of the teacher model and increase the recognition ability of the main features; in the later stage of training, a lower temperature helps the student model to finely imitate the output of the teacher model. Based on this phenomenon, the embodiment of the present disclosure proposes an adaptive temperature adjustment mechanism so that the temperature parameter can be dynamically adjusted as the training progresses, such as Figure 3 As shown, it specifically includes the following steps:
[0063] S302, design of adaptive temperature adjustment function: an adaptive temperature adjustment function is adopted, which adjusts the temperature parameter in real time according to the change rate of the model loss function or the amplitude change of the gradient during the training process. Specifically, the adaptive temperature adjustment function is shown in formula (1). The adaptive temperature adjustment function in the embodiment of the present disclosure can gradually reduce the temperature parameter as the training time increases, thereby ensuring that the model can converge faster in the early stage and can be effectively refined in the later stage.
[0064] S304, adaptive adjustment driven by loss function: The temperature parameter can also be adjusted dynamically in combination with the changing trend of the loss function. When the training loss decreases faster, the temperature reduction rate can be accelerated to encourage the student model to focus more on refined learning; when the loss decreases slower, the temperature reduction rate will slow down, encouraging the model to continue exploring and learning a wide range of features.
[0065] This adaptive temperature regulation mechanism can dynamically optimize the distillation process according to the specific progress of training, which not only improves the convergence speed of training, but also ensures that the student model can learn the knowledge of the teacher model more accurately, especially in large model compression scenarios, which can significantly improve training efficiency.
[0066] In some embodiments, the target loss function in the disclosed embodiment also includes: distillation loss. In the process of training the second deep learning model, the second deep learning model is optimized based on the pre-constructed target loss function until a student model that meets the conditions is obtained. The model training method in the disclosed embodiment also includes: calculating the difference between the second deep learning model and the output of the teacher model through the soft label of the teacher model and the output of the second deep learning model, and determining the distillation loss. Specifically, the disclosed embodiment can calculate the difference between the output of the student model and the teacher model through the soft label of the teacher model and the output of the student model, using the cross entropy loss or other loss functions, thereby outputting the distillation loss.
[0067] In some embodiments, in the process of training the second deep learning model, the embodiment of the present disclosure optimizes the second deep learning model based on the pre-constructed target loss function until a student model that meets the conditions is obtained. The model training method in the embodiment of the present disclosure also includes: obtaining the target loss function L by formula (2) total :
[0068] L total =L distill +λ1×L adv (2)
[0069] Among them, L distill represents the knowledge distillation loss; L adv represents the adversarial loss; λ1 represents the hyperparameter that balances the distillation loss and the adversarial loss.
[0070] In some embodiments, in the process of training the second deep learning model, the embodiment of the present disclosure optimizes the second deep learning model based on the pre-constructed target loss function until a student model that meets the conditions is obtained. The model training method in the embodiment of the present disclosure also includes: obtaining the target loss function L by formula (3) total :
[0071] L total =L distill +λ2×L MMD (3)
[0072] Among them, λ2 represents the hyperparameter that balances the distillation loss and the maximum mean difference loss; L MMD Denotes the maximum mean difference loss.
[0073] In cross-domain tasks, there may be differences in data distribution between the source domain (teacher model) and the target domain (student model), which makes it difficult for traditional knowledge distillation methods to obtain ideal model performance in the target domain. To solve this problem, the embodiments of the present disclosure reduce the difference in feature distribution between the source domain and the target domain through a domain difference function, which is equivalent to domain alignment, so that the student model can effectively learn domain-independent features.
[0074] In some embodiments, the disclosed embodiments may use adversarial training technology to perform adversarial learning on the output features of the source domain and the target domain through a domain discriminator, and minimize the feature distribution difference between the two domains through adversarial loss. The domain discriminator attempts to distinguish the features of the source domain and the target domain, and the student model finally generates a domain-independent feature representation through adversarial learning with the discriminator. The specific target loss function is shown in formula (2). In another embodiment, the disclosed embodiments may also directly align the output distribution of the source domain and the target domain through the maximum mean difference (MMD) method. MMD measures the mean difference between the two distributions. Minimizing this difference can effectively reduce the feature distribution gap between the domains, thereby enabling the student model to maintain high performance in the target domain. The specific target loss function is shown in formula (3).
[0075] The disclosed embodiment solves the adaptability problem in cross-domain knowledge distillation through this domain alignment mechanism, so that the teacher model's knowledge in the source domain can be seamlessly transferred to the target domain. It is particularly suitable for scenarios with significant domain differences such as multi-task learning, cross-language translation, and image classification.
[0076] In some embodiments, during the training of the second deep learning model, the embodiment of the present disclosure optimizes the second deep learning model based on a pre-constructed target loss function until a student model that meets the conditions is obtained. The model training method in the embodiment of the present disclosure also includes: inputting the pre-set verification data into the student model, and optimizing and adjusting the adaptive temperature adjustment function and the domain difference function according to the output verification result. Specifically, the embodiment of the present disclosure evaluates the performance of the student model by inputting a verification set, adjusting the hyperparameters of the adaptive temperature adjustment module and the domain alignment module when necessary, repeating the training until the ideal performance is achieved, and finally outputting the evaluation results and the optimized hyperparameters (such as the initial temperature value T0, the decay rate α, the adversarial loss weight λ1, etc.).
[0077] In some embodiments, Figure 4 As shown, the overall training process of the model training method of the embodiment of the present disclosure includes:
[0078] S402: Train the first deep learning model to obtain a teacher model.
[0079] S404, constructing a second deep learning model corresponding to the teacher model through the target domain dataset.
[0080] S406, performing adaptive temperature adjustment on the second deep learning model.
[0081] S408, based on the domain alignment mechanism, reducing the difference in feature distribution between the source domain and the target domain.
[0082] S410, calculating distillation loss.
[0083] S412, determine the target loss function.
[0084] S414, evaluating and adjusting the student model based on the pre-set verification data.
[0085] In some embodiments, the disclosed embodiments can accelerate the model training and compression process through an adaptive temperature adjustment mechanism, while maintaining the high accuracy of the model, making the deployment of large-scale models more efficient and reducing the computing cost of enterprises. In cross-domain tasks, the disclosed embodiments give cross-domain adaptability to knowledge distillation technology through a domain alignment mechanism, so that the model can better perform knowledge transfer in cross-domain tasks, thereby improving performance in multi-task and cross-domain applications. This has broad commercial value in the fields of multi-language processing, image recognition, etc.
[0086] Based on the same inventive concept, the present disclosure also provides a model training device in the following embodiments. Since the principle of solving the problem in the device embodiment is similar to that in the above method embodiment, the implementation of the device embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be repeated.
[0087] Figure 5 A schematic diagram of a model training device in an embodiment of the present disclosure is shown. Figure 6 As shown, the device comprises:
[0088] The data set acquisition module 501 is used to acquire the source domain data set and the target domain data set;
[0089] A teacher model determination module 502 is used to train the first deep learning model based on the source domain data set to obtain a teacher model;
[0090] A second deep learning model construction module 503, used to construct a second deep learning model corresponding to the teacher model through a target domain dataset;
[0091] A student model determination module 504 is used to optimize the second deep learning model based on a pre-constructed target loss function during the training of the second deep learning model until a student model that meets the conditions is obtained;
[0092] Among them, the target loss function is constructed based on the adaptive temperature adjustment function and the domain difference function. The adaptive temperature adjustment function is used to dynamically adjust the temperature parameters according to the model loss or gradient change of the second deep learning model during the training process. The domain difference function is used to minimize the difference in data feature distribution between the source domain and the target domain.
[0093] The model training device provided in the embodiments of the present disclosure obtains a source domain dataset and a target domain dataset through a dataset acquisition module; trains a first deep learning model based on the source domain dataset to obtain a teacher model through a teacher model determination module; constructs a second deep learning model corresponding to the teacher model through the target domain dataset through a second deep learning model construction module; and optimizes the second deep learning model based on a pre-constructed target loss function during the training of the second deep learning model through a student model determination module until a student model that meets the conditions is obtained, wherein the target loss function is constructed based on an adaptive temperature adjustment function and a domain difference function, the adaptive temperature adjustment function is used to dynamically adjust the temperature parameters according to the model loss or gradient change of the second deep learning model during the training process, and the domain difference function is used to minimize the data feature distribution difference between the source domain and the target domain. Compared with the knowledge distillation method in the related art, which has certain limitations in temperature parameter adjustment and adaptability to cross-domain tasks and is difficult to meet the needs of complex application scenarios, the embodiment of the present disclosure uses a pre-built adaptive temperature adjustment function and domain difference function to adjust the temperature parameters in real time according to the loss change or gradient trend of the model during training, so that the student model can more accurately imitate the teacher model; and the embodiment of the present disclosure uses the domain difference function to reduce the data feature distribution difference between the source domain and the target domain, ensuring that the student model can effectively learn domain-independent features. This mechanism enables the knowledge of the teacher model to be seamlessly transferred to the target domain, significantly improving the performance of the student model in cross-domain tasks.
[0094] In some embodiments, the model training model in the embodiments of the present disclosure also includes: an adaptive temperature adjustment function construction module, which is used to optimize the second deep learning model based on a pre-constructed target loss function during the training of the second deep learning model, until a student model that meets the conditions is obtained, and an adaptive temperature adjustment function is constructed based on the initial temperature of the second deep learning model, a preset temperature decay rate, and a preset training time step of the second deep learning model; a domain difference function construction module, which is used to perform adversarial learning on the output features of the source domain and the target domain through a domain discriminator, or to align the output distributions of the source domain and the target domain through the maximum mean difference method to obtain a domain difference function.
[0095] In some embodiments, the adaptive temperature adjustment function building module in the embodiment of the present disclosure is also used to obtain the adaptive temperature adjustment function T through formula (1): (t) .
[0096] In some embodiments, the target loss function in the embodiment of the present disclosure also includes: distillation loss. The model training device in the embodiment of the present disclosure also includes: a distillation loss determination module, which is used to optimize the second deep learning model based on a pre-constructed target loss function during the training of the second deep learning model until a student model that meets the conditions is obtained. The difference between the output of the second deep learning model and the teacher model is calculated through the soft label of the teacher model and the output result of the second deep learning model to determine the distillation loss.
[0097] In some embodiments, the model training device in the embodiment of the present disclosure further includes: a first target loss function determination module, which is used to optimize the second deep learning model based on a pre-constructed target loss function during the training of the second deep learning model, until a student model that meets the conditions is obtained, and the target loss function L is obtained by formula (2) total .
[0098] In some embodiments, the model training device in the embodiment of the present disclosure further includes: a second target loss function determination module, which is used to optimize the second deep learning model based on the pre-constructed target loss function during the training of the second deep learning model, until a student model that meets the conditions is obtained, and the target loss function L is obtained by formula (3) total .
[0099] In some embodiments, the model training device in the embodiments of the present disclosure also includes: an optimization and adjustment module, which is used to optimize the second deep learning model based on a pre-constructed target loss function during the training of the second deep learning model, until a student model that meets the conditions is obtained, and then the pre-set verification data is input into the student model, and the adaptive temperature adjustment function and the domain difference function are optimized and adjusted according to the output verification results.
[0100] Those skilled in the art will appreciate that various aspects of the present disclosure may be implemented as systems, methods or program products. Therefore, various aspects of the present disclosure may be specifically implemented in the following forms, namely: complete hardware implementation, complete software implementation (including firmware, microcode, etc.), or a combination of hardware and software, which may be collectively referred to herein as "circuits", "modules" or "systems".
[0101] Based on the same inventive concept, an electronic device is also provided in an embodiment of the present disclosure, the electronic device comprising: a processor; and a memory for storing executable instructions of the processor; wherein the processor is configured to execute any one of the above model training methods by executing the executable instructions. Since the principle of solving the problem in the electronic device embodiment is similar to that in the above method embodiment, the implementation of the electronic device embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be repeated.
[0102] Refer to the following Figure 6 The electronic device 600 according to this embodiment of the present disclosure is described. Figure 6 The electronic device 600 shown is merely an example and should not bring any limitation to the functions and scope of use of the embodiments of the present disclosure.
[0103] like Figure 6 As shown, the electronic device 600 is in the form of a general computing device. The components of the electronic device 600 may include but are not limited to: at least one processing unit 601, at least one storage unit 602, and a bus 603 connecting different system components (including the storage unit 602 and the processing unit 601).
[0104] The storage unit stores program codes, which can be executed by the processing unit 601, so that the processing unit 601 executes the steps described in the above “exemplary method” section of this specification according to various exemplary embodiments of the present disclosure.
[0105] In some embodiments, when the electronic device is used to control, for example, the model training method disclosed above, the processing unit 601 may perform the following steps of the above method embodiment:
[0106] Acquire a source domain dataset and a target domain dataset; train a first deep learning model based on the source domain dataset to obtain a teacher model; construct a second deep learning model corresponding to the teacher model through the target domain dataset; in the process of training the second deep learning model, optimize the second deep learning model based on a pre-constructed target loss function until a student model that meets the conditions is obtained; wherein the target loss function is constructed based on an adaptive temperature adjustment function and a domain difference function, the adaptive temperature adjustment function is used to dynamically adjust the temperature parameter according to the model loss or gradient change of the second deep learning model during the training process, and the domain difference function is used to minimize the difference in data feature distribution between the source domain and the target domain.
[0107] The storage unit 602 may include a readable medium in the form of a volatile storage unit, such as a random access storage unit (RAM) 6021 and / or a cache storage unit 6022 , and may further include a read-only storage unit (ROM) 6023 .
[0108] The storage unit 602 may also include a program / utility 6024 having a set (at least one) of program modules 6025, such program modules 6025 including but not limited to: an operating system, one or more application programs, other program modules, and program data, each of which or some combination may include an implementation of a network environment.
[0109] Bus 603 may represent one or more of several types of bus structures, including a memory unit bus or memory unit controller, a peripheral bus, an accelerated graphics port, a processing unit, or a local bus using any of a variety of bus architectures.
[0110] The electronic device 600 may also communicate with one or more external devices 604 (e.g., keyboards, pointing devices, Bluetooth devices, etc.), may also communicate with one or more devices that enable a user to interact with the electronic device 600, and / or communicate with any device that enables the electronic device 600 to communicate with one or more other computing devices (e.g., routers, modems, etc.). Such communication may be performed via an input / output (I / O) interface 605. Furthermore, the electronic device 600 may also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 606. As shown, the network adapter 606 communicates with other modules of the electronic device 600 via a bus 603. It should be understood that, although not shown in the figure, other hardware and / or software modules may be used in conjunction with the electronic device 600, including but not limited to: microcode, device drivers, redundant processing units, external disk drive arrays, RAID systems, tape drives, and data backup storage systems, etc.
[0111] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a terminal device, or a network device, etc.) to execute the method according to the implementation of the present disclosure.
[0112] Based on the same inventive concept, a computer-readable storage medium is also provided in the embodiment of the present disclosure, on which a computer program is stored, and when the computer program is executed by a processor, any of the above-mentioned model training methods is implemented. Since the principle of solving the problem in the embodiment of the computer-readable storage medium is similar to that in the above-mentioned method embodiment, the implementation of the embodiment of the computer-readable storage medium can refer to the implementation of the above-mentioned method embodiment, and the repeated parts will not be repeated.
[0113] More specific examples of computer-readable storage media in the present disclosure may include, but are not limited to, an electrical connection having one or more conductors, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the foregoing.
[0114] In the present disclosure, a computer readable storage medium may include a data signal propagated in baseband or as part of a carrier wave, wherein a readable program code is carried. Such propagated data signals may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A readable signal medium may also be any readable medium other than a readable storage medium, which may send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device.
[0115] Alternatively, the program code contained on the computer-readable storage medium may be transmitted using any appropriate medium, including but not limited to wireless, wired, optical cable, RF, etc., or any suitable combination of the foregoing.
[0116] In a specific implementation, the program code for performing the operations of the present disclosure may be written in any combination of one or more programming languages, including object-oriented programming languages such as Java, C++, etc., and conventional procedural programming languages such as "C" or similar programming languages. The program code may be executed entirely on the user computing device, partially on the user device, as a separate software package, partially on the user computing device and partially on a remote computing device, or entirely on a remote computing device or server. In the case of a remote computing device, the remote computing device may be connected to the user computing device through any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computing device (e.g., using an Internet service provider to connect through the Internet).
[0117] Based on the same inventive concept, a computer program product is also provided in the embodiments of the present disclosure, including a computer program product, including: a computer program or an instruction, which implements the model training method of any one of the above method embodiments when the computer program or the instruction is executed by the processor. Since the principle of solving the problem in the computer program product embodiment is similar to that in the above method embodiment, the implementation of the computer program product embodiment can refer to the implementation of the above method embodiment, and the repeated parts will not be repeated.
[0118] It should be noted that, although several modules or units of the device for action execution are mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.
[0119] In addition, although the steps of the method in the present disclosure are described in a specific order in the drawings, this does not require or imply that the steps must be performed in this specific order, or that all the steps shown must be performed to achieve the desired results. Additionally or alternatively, some steps may be omitted, multiple steps may be combined into one step, and / or one step may be decomposed into multiple steps, etc.
[0120] Through the description of the above implementation, it is easy for those skilled in the art to understand that the example implementation described here can be implemented by software, or by software combined with necessary hardware. Therefore, the technical solution according to the implementation of the present disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (which can be a CD-ROM, a USB flash drive, a mobile hard disk, etc.) or on a network, including several instructions to enable a computing device (which can be a personal computer, a server, a mobile terminal, or a network device, etc.) to execute the method according to the implementation of the present disclosure.
[0121] Those skilled in the art will readily appreciate other embodiments of the present disclosure after considering the specification and practicing the invention disclosed herein. The present disclosure is intended to cover any variations, uses or adaptations of the present disclosure, which follow the general principles of the present disclosure and include common knowledge or customary techniques in the art that are not disclosed in the present disclosure. The description and examples are intended to be exemplary only, and the true scope and spirit of the present disclosure are indicated by the appended claims.
Claims
1. A model training method, characterized in that: include: Obtain source domain datasets and target domain datasets; Training a first deep learning model based on the source domain dataset to obtain a teacher model; Constructing a second deep learning model corresponding to the teacher model through the target domain dataset; In the process of training the second deep learning model, based on a pre-constructed target loss function, optimizing the second deep learning model until a student model that meets the conditions is obtained; Among them, the target loss function is constructed based on an adaptive temperature adjustment function and a domain difference function. The adaptive temperature adjustment function is used to dynamically adjust the temperature parameters according to the model loss or gradient change of the second deep learning model during the training process. The domain difference function is used to minimize the difference in data feature distribution between the source domain and the target domain.
2. The model training method according to claim 1, characterized in that: In the process of training the second deep learning model, based on the pre-constructed target loss function, optimizing the second deep learning model until a student model that meets the conditions is obtained, the method further includes: Constructing the adaptive temperature adjustment function based on the initial temperature of the second deep learning model, the preset temperature decay rate, and the preset training time step of the second deep learning model; The domain difference function is obtained by performing adversarial learning on the output features of the source domain and the target domain through a domain discriminator, or aligning the output distributions of the source domain and the target domain through a maximum mean difference method.
3. The model training method according to claim 2, characterized in that: Based on the initial temperature of the second deep learning model, the preset temperature decay rate and the preset training time step of the second deep learning model, constructing the adaptive temperature adjustment function, including: obtaining the adaptive temperature adjustment function T by the following formula (t) : Wherein, T0 represents the initial temperature of the second deep learning model; α represents the hyperparameter of the temperature decay rate; and t represents the training time step of the second deep learning model.
4. The model training method according to claim 2, characterized in that: The target loss function further includes: distillation loss. In the process of training the second deep learning model, the second deep learning model is optimized based on the pre-constructed target loss function until a student model that meets the conditions is obtained. The method further includes: The difference between the output of the second deep learning model and the output of the teacher model is calculated by using the soft label of the teacher model and the output result of the second deep learning model to determine the distillation loss.
5. The model training method according to claim 4, characterized in that: In the process of training the second deep learning model, based on the pre-constructed target loss function, the second deep learning model is optimized until a student model that meets the conditions is obtained. The method further includes: obtaining the target loss function L by the following formula total : L total =L distill +λ1×L adv Among them, L distill represents the knowledge distillation loss; L adv represents the adversarial loss; λ1 represents the hyperparameter that balances the distillation loss and the adversarial loss.
6. The model training method according to claim 4, characterized in that: In the process of training the second deep learning model, based on the pre-constructed target loss function, the second deep learning model is optimized until a student model that meets the conditions is obtained. The method further includes: obtaining the target loss function L by the following formula total : L total =L distill +λ2×L MMD Among them, λ2 represents the hyperparameter that balances the distillation loss and the maximum mean difference loss; L MMD Denotes the maximum mean difference loss.
7. The model training method according to claim 1, characterized in that: In the process of training the second deep learning model, based on the pre-constructed target loss function, optimizing the second deep learning model until a student model that meets the conditions is obtained, the method further includes: The preset verification data is input into the student model, and the adaptive temperature adjustment function and the domain difference function are optimized and adjusted according to the output verification result.
8. A model training device, characterized in that: include: A data set acquisition module is used to acquire source domain data sets and target domain data sets; A teacher model determination module, used to train the first deep learning model based on the source domain data set to obtain a teacher model; A second deep learning model construction module, used to construct a second deep learning model corresponding to the teacher model through the target domain dataset; A student model determination module, used to optimize the second deep learning model based on a pre-constructed target loss function during the training of the second deep learning model until a student model that meets the conditions is obtained; Among them, the target loss function is constructed based on an adaptive temperature adjustment function and a domain difference function. The adaptive temperature adjustment function is used to dynamically adjust the temperature parameters according to the model loss or gradient change of the second deep learning model during the training process. The domain difference function is used to minimize the difference in data feature distribution between the source domain and the target domain.
9. An electronic device, characterized in that: include: processor; as well as A memory, configured to store executable instructions of the processor; Wherein, the processor is configured to execute the model training method described in any one of claims 1 to 7 by executing the executable instructions.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the model training method described in any one of claims 1 to 7 is implemented.
11. A computer program product comprising: A computer program or instruction, characterized in that when the computer program or instruction is executed by a processor, it implements the model training method described in any one of claims 1 to 7.