Knowledge distillation method and device for parameter search matching

By employing a knowledge distillation method based on parameter search and matching, dynamically adjusting and controlling parameters to filter out invalid parameters, and performing multi-scale feature alignment pooling, the problem of uneven knowledge acquisition during the learning process of student network models is solved, thereby achieving efficient transmission of teacher network model information and optimization of the learning process.

CN121303255APending Publication Date: 2026-01-09E SURFING VISION TECHNOLOGY CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511591483.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-11-03
Publication Date
2026-01-09

AI Technical Summary

Technical Problem

In existing technologies, student network models struggle to progressively acquire knowledge from teacher network models during the learning process, particularly in the matching of shallow and deep semantic information, which increases the learning difficulty.

Method used

By employing a knowledge distillation method based on parameter search and matching, the threshold control parameters are dynamically adjusted to filter out invalid parameters, the magnitude control parameters are maximized to match intermediate layers, multi-scale feature alignment pooling operations are performed, and the inter-layer loss function is updated in real time to ensure that the student network model gradually learns the knowledge of the teacher network model.

Benefits of technology

This approach enables a gradual knowledge acquisition process for student network models, improves knowledge transfer efficiency, reduces the impact of hyperparameters on training results, and ensures the full transmission of information from teacher network models.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121303255A_ABST
    Figure CN121303255A_ABST
Patent Text Reader

Abstract

The invention relates to a knowledge distillation method and device for parameter search matching, and belongs to the technical field of knowledge distillation. The method comprises the following steps: randomly initializing student network model parameters, and loading a teacher network model; initializing a threshold control parameter and a magnitude control parameter, and calculating parameter quantities contained in all network layers in each feature output layer of the student network model and the teacher network model; in the distillation training process, threshold control parameters are dynamically adjusted to filter invalid parameters in the student network model and the teacher network model; controlling magnitude control parameters to be matched with the number of a group of middle layers of the teacher network model and the student network model participating in distillation training to the maximum extent; performing multi-scale feature alignment pooling operation on a group of matched network intermediate layer output features, fusing different scale features, and updating an interlayer loss function in real time; and determining an overall loss function of distillation training, training iteration for multiple rounds until the overall loss function of the distillation training is minimum, and ending the distillation training.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application belongs to the technical field of knowledge distillation, and particularly relates to a knowledge distillation method and device for parameter search matching. BACKGROUND

[0002] The closest prior art to the technical solution of the present application is the technical solution described in the paper FitNets: Hints for Thin Deep Nets, which trains the network by guiding the student network model to learn the features of the intermediate layers of the teacher network model. However, this simple introduction of intermediate layer distillation is obviously unreasonable in the training process. The shallow semantic information of the teacher network model belongs to simple knowledge, and the deep semantic information belongs to abstract knowledge. When guiding the training of the student network model, it is necessary to consider the matching of the knowledge learned by the guided layer of the student network model with the depth of the network layer. If the guided layer of the student network is shallow and the hint layer of the teacher network model is deep, it will lead to that the student network model accepts more abstract knowledge in the early stage of learning, which is obviously very difficult and difficult to learn. SUMMARY

[0003] In view of the deficiencies of the prior art, the purpose of the present application is to provide a knowledge distillation method and device for parameter search matching, so that in the distillation training process, the student network learns simpler knowledge from the shallow layer of the teacher network in the early stage of training, and with the continuous distillation training, the student network model learns more abstract knowledge from the deep layer of the teacher network model. The whole learning process is gradual, and the student network model can learn as much knowledge as possible from the teacher network model.

[0004] In a first aspect of the present application, a knowledge distillation method for parameter search matching is provided, comprising:

[0005] Randomly initializing the parameters of the student network model and loading the teacher network model;

[0006] Initializing the threshold control parameter and the order control parameter, and calculating the parameter amount contained in each network layer of all feature output layers of the student network model and the teacher network model;

[0007] In the distillation training process, the threshold control parameter is dynamically adjusted to filter the invalid parameters in the student network model and the teacher network model;

[0008] The order control parameter is controlled to maximize the number of intermediate layers in the matched group of the teacher network model and the student network model participating in the distillation training;

[0009] Performing multi-scale feature alignment pooling operation on the output features of the matched group of network intermediate layers, and fusing different scale features to update the interlayer loss function in real time;

[0010] The overall loss function of the distillation training is determined according to the last layer features of the student network model, the last layer features of the teacher network model and the interlayer loss function, and the distillation training is iterated for multiple rounds until the overall loss function of the distillation training reaches the minimum, and the distillation training is ended.

[0011] Further, in the above-mentioned parameter search matching knowledge distillation method, the threshold control parameter and the order control parameter are initialized, and the parameter quantity contained in each feature output layer of the student network model and the teacher network model is calculated to initialize and calculate the parameter search matching network.

[0012] Further, in the above-mentioned parameter search matching knowledge distillation method, in the distillation training process, the threshold control parameter is dynamically adjusted to filter the invalid parameters in the student network model and the teacher network model, including:

[0013] The distillation training is iterated for t rounds, and the threshold control parameter is automatically updated as

[0014] In each round, the parameter quantity of the network layer parameters greater than a t in the teacher network model and the parameter quantity of the network layer parameters greater than a t in the student network model are selected.

[0015] Wherein, a0 represents the initialized threshold control parameter, a t represents the threshold control parameter of the tth round, and t represents an integer.

[0016] Further, in the above-mentioned parameter search matching knowledge distillation method, the maximum matching of the number of intermediate layers of the teacher network model and the student network model participating in the distillation training is controlled by the parameter adaptive matching module in the parameter search matching network, and the control method is represented by the following formula:

[0017]

[0018] Wherein, η represents the order control parameter, i represents the i th layer of the student network model, represents the parameter quantity of the parameters greater than a in the j th layer of the teacher network model, represents the parameter quantity of the parameters greater than a in the i th layer of the student network model, j represents the j th layer of the teacher network model, S(i,j) represents the set of matching layer pairs of the i th layer of the student network model and the j th layer of the teacher network model with the same parameter quantity, M represents the last layer of the student network model, and L represents the last layer of the teacher network model.

[0019] Further, in the parameter search matching knowledge distillation method, the multi-scale feature alignment pooling operation is performed on the matched group of network intermediate layer output features, and the different scale features are fused through a multi-scale pooling fusion network.

[0020] Further, in the parameter search matching knowledge distillation method, the real-time inter-layer loss function is represented by the following formula:

[0021]

[0022] Wherein, and The features output by the i-th layer of the student network model and the j-th layer of the teacher network model are of the same order of magnitude, and PoF S and PoF T respectively represent that the multi-scale pooling fusion network PoF performs a multi-scale feature alignment pooling operation on the output features of the student network model and the teacher network model, and L layer represents the inter-layer loss function, and alpha t represents the threshold control parameter of the t-th round, and t represents an integer.

[0023] Further, in the parameter search matching knowledge distillation method, the overall loss function of the distillation training is determined according to the last layer features of the student network model, the last layer features of the teacher network model and the inter-layer loss function, and is represented by the following formula:

[0024]

[0025] Wherein, and The features output by the last layer of the student network model and the last layer of the teacher network model are of the same order of magnitude, F represents a function of converting the features, D represents a feature distance function, and L layer represents the inter-layer loss function, i represents the i-th layer of the student network model, j represents the j-th layer of the teacher network model, M represents the last layer of the student network model, L represents the last layer of the teacher network model, and L layer represents the inter-layer loss function, and L KD represents the overall loss function of the distillation training.

[0026] The second aspect of the present application also provides a parameter search matching knowledge distillation device, comprising:

[0027] The loading module is used for randomly initializing the parameters of the student network model and loading the teacher network model.

[0028] An initialization module is configured to initialize threshold control parameters and magnitude control parameters, and calculate the parameter quantity contained in all network layers in each feature output layer of the student network model and the teacher network model;

[0029] A filtering module is configured to dynamically adjust the threshold control parameters to filter the invalid parameters in the student network model and the teacher network model during distillation training;

[0030] A control module is configured to control the magnitude control parameters to maximize the number of intermediate layers in the matched set of the teacher network model and the student network model participating in distillation training;

[0031] An updating module is configured to perform multi-scale feature alignment pooling operation on the output features of the matched set of network intermediate layers, fuse different scale features, and update the inter-layer loss function in real time;

[0032] A determination module is configured to determine the overall loss function of distillation training according to the last layer features of the student network model with consistent magnitude, the last layer features of the teacher network model, and the inter-layer loss function, and train multiple rounds until the overall loss function of distillation training reaches the minimum, and end the distillation training.

[0033] The third aspect of the present application also provides an electronic device, comprising a processor and a memory.

[0034] The processor is configured to execute the parameter search matching knowledge distillation method according to any one of the above by calling the program or instructions stored in the memory.

[0035] The fourth aspect of the present application also provides a computer readable storage medium, which stores programs or instructions, and the programs or instructions enable the computer to execute the parameter search matching knowledge distillation method according to any one of the above.

[0036] The present application has the following beneficial effects: a dynamic parameter filtering mechanism is constructed by the parameter adaptive matching module to filter the invalid parameters in the teacher-student network, and the magnitude control parameters can maximize the number of intermediate layers in the matched set of the teacher-student model participating in distillation training, so as to ensure that the information in the teacher network is fully transmitted to the student network, the feature pooling fusion alignment module designs a multi-scale pooling fusion network and an inter-layer loss function of the teacher-student network model, the multi-scale pooling fusion network aligns the output features of the teacher-student network through multi-scale feature fusion, the inter-layer loss function introduces the threshold control parameters, and the matched set of intermediate layers participating in the loss function calculation in the training process is updated in real time, so as to improve the efficiency of knowledge transmission from the teacher network to the student network in the training process, the automatic optimization training strategy can ensure that all artificially introduced parameters are automatically optimized and updated in the training process, and the influence of hyperparameters on the training effect is reduced. BRIEF DESCRIPTION OF DRAWINGS

[0037] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Throughout the drawings, the same reference numerals denote the same parts. It is obvious that the drawings described below are merely some embodiments of the present invention, and those skilled in the art can obtain other drawings based on these drawings.

[0038] Figure 1 A diagram illustrating a knowledge distillation method for parameter search and matching provided in an embodiment of the present invention;

[0039] Figure 2 A diagram of a knowledge distillation apparatus for parameter search and matching provided in an embodiment of the present invention;

[0040] Figure 3 This is a schematic block diagram of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0041] To enable those skilled in the art to better understand the technical solutions in the embodiments of the present invention, the technical solutions of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, not all embodiments. It should be understood that these descriptions are merely exemplary and are not intended to limit the scope of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of the present invention.

[0042] Furthermore, descriptions of well-known structures and techniques are omitted in the following description to avoid unnecessarily obscuring the concepts disclosed in this invention.

[0043] In the description of this invention, the terms "first," "second," and "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. The terms "installed," "connected," and "linked" should be interpreted broadly; for example, they can refer to a fixed connection, a detachable connection, or an integral connection; they can refer to a mechanical connection or an electrical connection; they can refer to a direct connection or an indirect connection through an intermediate medium; and they can refer to the internal connection of two components. Those skilled in the art will understand the specific meaning of the above terms in this invention based on the specific circumstances.

[0044] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numerals in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with the present invention. Rather, they are merely examples of methods and systems consistent with some aspects of the invention as detailed in the appended claims.

[0045] This invention proposes a knowledge distillation method, apparatus, electronic device, and storage medium for parameter search and matching. In the early stages of distillation training, the student network learns simpler knowledge from the shallow layer of the teacher network. As distillation training continues, the student network model learns more abstract knowledge from the deeper layer of the teacher network model. The entire learning process is gradual, and the student network model can learn as much knowledge as possible from the teacher network model.

[0046] Before introducing the embodiments of the present invention, the technical terms involved in the present invention will be introduced first.

[0047] KD (Knowledge Distillation) is a deep learning model compression technique. Its core idea is to improve the performance of a student network model by guiding a lightweight student network model to "imitate" a teacher network model with better performance and a more complex structure, without changing the structure of the student network model.

[0048] PMSN (Parameter Match Search Net): In model distillation, it is a network that ensures the number of parameters in the intermediate layer eigenvalues ​​of the student network model is on the same order of magnitude as the number of parameters in the intermediate layer eigenvalues ​​of the teacher network model in calculating the loss function.

[0049] Method Implementation Examples

[0050] Figure 1 A diagram illustrating a knowledge distillation method for parameter search and matching provided in an embodiment of the present invention.

[0051] In a first aspect, the present invention proposes a knowledge distillation method for parameter search and matching, combined with Figure 1 It includes six steps, S1 to S6:

[0052] S1: Randomly initialize the parameters of the student network model and load the teacher network model.

[0053] Specifically, in this embodiment of the invention, the teacher network model is known, and the parameters of the student network model can be initialized in various ways, such as normal distribution.

[0054] S2: Initialize threshold control parameters and magnitude control parameters, and calculate the number of parameters contained in all network layers in each feature output layer of the student network model and the teacher network model.

[0055] Specifically, in this embodiment of the invention, if it is calculated that the number of parameters contained in all network layers of one feature output layer of the student network model is 50 numbers between 0 and 0.9, and the number of parameters contained in all network layers of one feature output layer of the teacher network model is 100 numbers between 0 and 0.9, the initial threshold control parameter is 0.5, and the initial magnitude control parameter is 0.5 so that 100 * 0.5 = 50, thus the number of parameters contained in one feature output layer of the teacher network model and one feature output layer of the student network model are on the same order of magnitude.

[0056] S3: During the distillation training process, dynamically adjust the threshold control parameters to filter out invalid parameters in the student network model and the teacher network model.

[0057] Specifically, in this embodiment of the invention, since there are parameters with extremely small values ​​in the structure of the teacher network model and the student network model, these parameters have minimal impact on the feature representation ability of the network and can be regarded as invalid parameters. During the distillation training process, the threshold control parameter is dynamically adjusted to filter the invalid parameters in the teacher network model and the student network model. The specific filtering method is described in detail below.

[0058] S4: Control magnitude control parameters maximize the number of intermediate layers in a set of teacher and student network models participating in distillation training.

[0059] Specifically, in this embodiment of the invention, since the number of parameters in the teacher network model and the student network model may differ by several orders of magnitude, for example, the number of parameters in a feature output layer in the teacher network model is 200, while the number of parameters in a feature output layer in the student network model is 500, the parameter adaptive matching module of this invention has a parameter magnitude adaptive mechanism. The magnitude control parameter η can maximize the number of intermediate layers in a set of teacher and student models participating in distillation training. The specific matching method is described in detail below.

[0060] S5: Perform multi-scale feature alignment pooling on a set of matched intermediate network output features, fuse features at different scales, and update the inter-layer loss function in real time.

[0061] Specifically, in this embodiment of the invention, multi-scale feature alignment pooling is performed on the output features of a set of matched intermediate network layers, and features of different scales are fused. The inter-layer loss function is updated in real time through the feature pooling fusion alignment module. The feature pooling fusion alignment module involves the inter-layer loss function of the multi-scale pooling fusion network and the teacher-student model network. Multi-scale feature alignment pooling is performed on the output features of a set of matched intermediate network layers, and features of different scales are fused. The inter-layer loss function introduces a threshold control parameter, and the matched set of intermediate layers participating in the loss function calculation during training is updated in real time. This improves the efficiency of knowledge transfer from the teacher network to the student network during training. The specific method is described in detail below.

[0062] S6: Determine the overall loss function for distillation training based on the last layer features of the student network model and the last layer features of the teacher network model, which are of the same order of magnitude, and the inter-layer loss function. Iterate the training for multiple rounds until the overall loss function of distillation training reaches its minimum, and then end the distillation training.

[0063] Specifically, in this embodiment of the invention, the order of magnitude is consistent. For example, the number of parameters in the final feature output layer of the teacher network model is 50, and the number of parameters in the final feature output layer of the student network model is 50. The method for determining the overall loss function of distillation training is described in detail below. The overall loss function of multi-round distillation training is minimized, thereby improving the distillation effect on the student network model's ability to accept feature representations from the teacher network model.

[0064] It should be understood that the parameter adaptive matching module connects all feature output layers of the teacher network model and the student network model, and is used to obtain the effective weight parameter quantity in the weight of the teacher and student network models, and match a set of intermediate network layers participating in the calculation of the loss function of the teacher and student network models. The feature pooling fusion alignment module is used to process and calculate the feature loss function of the paired teacher and student models.

[0065] Furthermore, in the aforementioned knowledge distillation method for parameter search matching, the initialization of threshold control parameters and magnitude control parameters, and the calculation of the number of parameters contained in all network layers in each feature output layer of the student network model and the teacher network model are initialized and calculated by the parameter search matching network.

[0066] Specifically, in this embodiment of the invention, for example: the parameter search matching network calculates that the number of parameters contained in all network layers of one feature output layer of the student network model is 50 numbers between 0 and 0.9, and the number of parameters contained in all network layers of one feature output layer of the teacher network model is 100 numbers between 0 and 0.9. The parameter search matching network initialization threshold control parameter is 0.5, and the parameter search matching network initialization magnitude control parameter is 0.5 so that 100 * 0.5 = 50. Thus, the number of parameters contained in one feature output layer of the teacher network model and one feature output layer of the student network model are on the same order of magnitude.

[0067] Furthermore, in the aforementioned knowledge distillation method for parameter search and matching, during the distillation training process, the threshold control parameter is dynamically adjusted to filter invalid parameters in the student network model and the teacher network model, including:

[0068] Distillation training iterations t times, the threshold control parameter is automatically updated to

[0069] In each round, network layer parameters in the teacher network model that are greater than α are selected. t The number of parameters and the number of network layer parameters in the student network model are greater than α. t The number of parameters;

[0070] Where α0 represents the initialized threshold control parameter, α t This represents the threshold control parameter for round t, where t is an integer.

[0071] Specifically, in this embodiment of the invention, the threshold control parameter is dynamically adjusted and gradually decreased, such as in the first round, α t =0.5, α in the second round t =0.5 / 2, α in the third round t =0.5 / 3, gradually filtering out invalid parameters to improve the expressive power of features during the knowledge distillation training process.

[0072] Furthermore, in the aforementioned knowledge distillation method using parameter search and matching, the control magnitude—specifically, the number of intermediate layers in the teacher and student network models participating in the distillation training—is controlled by the parameter adaptive matching module within the parameter search and matching network. This control method is expressed by the following formula:

[0073]

[0074] Where η represents the magnitude control parameter, and i represents the i-th layer of the student network model. This represents the number of parameters in the j-th layer of the teacher network model that are greater than α. Let α represent the number of parameters in the i-th layer of the student network model that are greater than α, j represent the j-th layer of the teacher network model, S(i,j) represent the set of matching layer pairs of the i-th layer of the student network model and the j-th layer of the teacher network model with the same number of parameters, M represent the last layer of the student network model, and L represent the last layer of the teacher network model.

[0075] Specifically, in this embodiment of the invention, the parameter adaptive matching module matches parameters that satisfy the magnitude control parameter η. and During the t-th round of iterative training, the threshold control parameter is automatically updated to α. t The magnitude control parameter η is automatically rematched according to the parameter magnitude adaptive mechanism. and

[0076] Furthermore, in the knowledge distillation method for parameter search matching described above, multi-scale feature alignment pooling is performed on the output features of a set of intermediate network layers that are matched, and the fusion of features at different scales is achieved through a multi-scale pooling fusion network.

[0077] Furthermore, in the knowledge distillation method for parameter search and matching described above, the inter-layer loss function is updated in real time using the following formula:

[0078]

[0079] in, and Features output by the i-th layer of the student network model and the j-th layer of the teacher network model of the same order of magnitude, PoF S and PoF T L represents the multi-scale feature alignment pooling operation performed on the output features of the student network model and the teacher network model by the multi-scale pooling fusion network PoF. layer Let α represent the interlayer loss function. t This represents the threshold control parameter for round t, where t is an integer.

[0080] Specifically, in this embodiment of the invention, the order of magnitude is consistent. For example, the number of parameters in the second feature output layer of the teacher network model is 50, and the number of parameters in the second feature output layer of the student network model is 50. The multi-scale pooling fusion network PoF aligns the output features of the teacher and student networks through multi-scale feature fusion, and the inter-layer loss function L... layer By introducing a threshold control parameter, a set of intermediate layers that participate in the loss function calculation during training are updated in real time, which improves the efficiency of knowledge transfer from the teacher network model to the student network model during training.

[0081] Furthermore, in the aforementioned knowledge distillation method for parameter search and matching, the overall loss function for distillation training, determined based on the last-layer features of the student network model, the last-layer features of the teacher network model, and the inter-layer loss function of consistent order of magnitude, is expressed by the following formula:

[0082]

[0083] in, and The features output from the last layer of the student network model and the teacher network model are of the same order of magnitude. F represents the function that transforms the dimensions of the features, D represents the feature distance function, and L... layer Let L represent the inter-layer loss function, i represent the i-th layer of the student network model, j represent the j-th layer of the teacher network model, M represent the last layer of the student network model, and L represent the last layer of the teacher network model. layer L represents the interlayer loss function. KD This represents the overall loss function for distillation training.

[0084] Specifically, in the embodiments of the present invention, This involves transforming the dimensions of the features output from the last layer of the student network model, such as converting them from 10x5 dimensions to 5x5 dimensions. This involves transforming the dimensions of the features output from the last layer of the teacher network model, such as from 10*5 dimensions to 10*10 dimensions. The purpose of this transformation is to convert the features output from the last layer of the student network model and the features output from the last layer of the teacher network model to the same dimension. The distance between the features can be used to intuitively determine whether the student network model has learned the knowledge from the teacher network model.

[0085] Device Examples

[0086] Figure 2 A diagram illustrating a knowledge distillation method for parameter search and matching provided in an embodiment of the present invention.

[0087] A second aspect of the present invention also proposes a knowledge distillation apparatus for parameter search and matching, comprising:

[0088] Loading module 21: Used to randomly initialize the parameters of the student network model and load the teacher network model.

[0089] Specifically, in this embodiment of the invention, the teacher network model is known, and the parameters of the student network model can be initialized in various ways, such as normal distribution.

[0090] Initialization module 22: Used to initialize threshold control parameters and magnitude control parameters, and to calculate the number of parameters contained in all network layers in each feature output layer of the student network model and the teacher network model.

[0091] Specifically, in this embodiment of the invention, if the initialization module 22 calculates that the number of parameters contained in all network layers of one feature output layer of the student network model is 50 numbers between 0 and 0.9, and the number of parameters contained in all network layers of one feature output layer of the teacher network model is 100 numbers between 0 and 0.9, the initialization module 22 initializes the threshold control parameter to 0.5 and the initialization module 22 initializes the magnitude control parameter to 0.5 so that 100 * 0.5 = 50, thereby the number of parameters contained in one feature output layer of the teacher network model and one feature output layer of the student network model are on the same order of magnitude.

[0092] Filtering module 23: Used to dynamically adjust threshold control parameters during distillation training to filter invalid parameters in the student network model and the teacher network model.

[0093] Specifically, in this embodiment of the invention, since there are parameters with extremely small values ​​in the structure of the teacher network model and the student network model, these parameters have minimal impact on the feature representation ability of the network and can be regarded as invalid parameters. During the distillation training process, the filtering module 23 dynamically adjusts the threshold control parameters to filter invalid parameters in the teacher network model and the student network model. The specific filtering method is described in detail below.

[0094] Control module 24: Used to control the magnitude of the control parameters to maximize the number of intermediate layers in a set of teacher and student network models participating in the distillation training.

[0095] Specifically, in this embodiment of the invention, since the number of parameters in the teacher network model and the student network model may differ by several orders of magnitude, for example, the number of parameters in a feature output layer in the teacher network model is 200, while the number of parameters in a feature output layer in the student network model is 500, the parameter adaptive matching module of the present invention has a parameter magnitude adaptive mechanism, and the control module 24 controls the magnitude control parameter η to maximize the number of intermediate layers in a set of teacher and student models participating in distillation training.

[0096] Update module 25: Used to perform multi-scale feature alignment pooling on a set of matched network intermediate layer output features, fuse features at different scales, and update the inter-layer loss function in real time.

[0097] Specifically, in this embodiment of the invention, the update module 25 performs multi-scale feature alignment pooling on the output features of a set of matched intermediate network layers and fuses features of different scales. The real-time update of the inter-layer loss function is achieved through the feature pooling fusion alignment module, which involves multi-scale pooling fusion of the network and the inter-layer loss function of the teacher-student model network. It performs multi-scale feature alignment pooling on the output features of a set of matched intermediate network layers and fuses features of different scales. The inter-layer loss function introduces a threshold control parameter and updates the matched set of intermediate layers participating in the loss function calculation in real time, thereby improving the efficiency of knowledge transfer from the teacher network to the student network during the training process.

[0098] Module 26: This module determines the overall loss function for distillation training based on the last layer features of the student network model, the last layer features of the teacher network model, and the interlayer loss function, all of which are of similar magnitude. The training is iterated for multiple rounds until the overall loss function for distillation training reaches its minimum, at which point the distillation training ends.

[0099] Specifically, in the embodiments of the present invention, the order of magnitude is consistent. For example, the number of parameters in the final feature output layer of the teacher network model is 50, and the number of parameters in the final feature output layer of the student network model is 50. The method for determining the overall loss function of distillation training is introduced in the method class embodiments. The overall loss function of training iterations and multiple rounds of distillation training is minimized, thereby improving the distillation effect on the student network model's ability to accept feature representations from the teacher network model.

[0100] A third aspect of the present invention also provides an electronic device comprising: a processor and a memory;

[0101] The processor executes a knowledge distillation method, such as one of the above parameter search matching methods, by calling programs or instructions stored in memory.

[0102] In a fourth aspect, the present invention also provides a computer-readable storage medium storing a program or instructions that cause a computer to perform a knowledge distillation method for parameter search and matching as described in any of the preceding inventions.

[0103] Figure 3 This is a schematic block diagram of an electronic device provided in an embodiment of the present invention.

[0104] like Figure 3As shown, the electronic device includes at least one processor 301, at least one memory 302, and at least one communication interface 303. The various components of the electronic device are coupled together via a bus system 304. The communication interface 303 is used for information transmission with external devices. It is understood that the bus system 304 is used to implement communication between these components. In addition to a data bus, the bus system 304 also includes a power bus, a control bus, and a status signal bus. However, for clarity, ... Figure 3 The general designated all buses as Bus System 304.

[0105] It is understood that the memory 302 in this embodiment may be volatile memory or non-volatile memory, or may include both volatile and non-volatile memory.

[0106] In some implementations, memory 302 stores elements such as executable units or data structures, or subsets thereof, or extended sets thereof: operating systems and applications.

[0107] The operating system includes various system programs, such as the framework layer, core library layer, and driver layer, used to implement various basic business functions and handle hardware-based tasks. The application programs include various applications, such as media players and browsers, used to implement various application functions. A program implementing any method in the knowledge distillation method for parameter search and matching provided in this embodiment of the invention can be included in the application programs.

[0108] In this embodiment of the invention, the processor 301 executes the steps of various embodiments of the knowledge distillation method for parameter search and matching provided by the present invention by calling the program or instructions stored in the memory 302, specifically, the program or instructions stored in the application program.

[0109] Randomly initialize the parameters of the student network model and load the teacher network model;

[0110] Initialize threshold control parameters and magnitude control parameters, and calculate the number of parameters contained in all network layers in each feature output layer of the student network model and the teacher network model;

[0111] During the distillation training process, the threshold control parameters are dynamically adjusted to filter out invalid parameters in the student network model and the teacher network model.

[0112] The control magnitude and control parameters maximize the number of intermediate layers in a set of teacher and student network models participating in distillation training;

[0113] Multi-scale feature alignment pooling is performed on a set of matched intermediate network output features, and features at different scales are fused to update the inter-layer loss function in real time.

[0114] The overall loss function for distillation training is determined based on the last layer features of the student network model and the last layer features of the teacher network model, which are of the same order of magnitude, and the inter-layer loss function. The training is iterated for multiple rounds until the overall loss function of distillation training reaches its minimum, at which point the distillation training ends.

[0115] Any method in the knowledge distillation method for parameter search and matching provided in this embodiment of the invention can be applied to, or implemented by, the processor 301. The processor 301 can be an integrated circuit chip with signal processing capabilities. During implementation, each step of the above method can be completed by the integrated logic circuitry in the hardware of the processor 301 or by instructions in software form. The processor 301 can be a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field-programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. The general-purpose processor can be a microprocessor or any conventional processor, etc.

[0116] The steps of any method in the knowledge distillation method for parameter search and matching provided in this embodiment of the invention can be directly implemented by a hardware decoding processor, or implemented by a combination of hardware and software units in the decoding processor. The software units can reside in random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, registers, or other mature storage media in the art. This storage medium is located in memory 302, and processor 301 reads the information in memory 302 and combines it with its hardware to complete the steps of the method.

[0117] Those skilled in the art will understand that although some embodiments described herein include certain features included in other embodiments but not others, combinations of features from different embodiments are meant to be within the scope of the invention and form different embodiments.

[0118] Those skilled in the art will understand that the descriptions of the various embodiments have different focuses, and for parts not described in detail in a certain embodiment, reference can be made to the relevant descriptions of other embodiments.

[0119] Although embodiments of the present invention have been described in conjunction with the accompanying drawings, those skilled in the art can make various modifications and variations without departing from the spirit and scope of the invention. All such modifications and variations fall within the scope defined by the appended claims. The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

[0120] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in the present invention, and these modifications or substitutions should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.

Claims

1. A knowledge distillation method for parameter search and matching, characterized in that, include: Randomly initialize the parameters of the student network model and load the teacher network model; Initialize threshold control parameters and magnitude control parameters, and calculate the number of parameters contained in all network layers in each feature output layer of the student network model and the teacher network model; During the distillation training process, the threshold control parameters are dynamically adjusted to filter out invalid parameters in the student network model and the teacher network model. The control magnitude and control parameters maximize the number of intermediate layers in a set of teacher and student network models participating in distillation training; Multi-scale feature alignment pooling is performed on a set of matched intermediate network output features, and features at different scales are fused to update the inter-layer loss function in real time. The overall loss function for distillation training is determined based on the last layer features of the student network model and the last layer features of the teacher network model, which are of the same order of magnitude, and the inter-layer loss function. The training is iterated for multiple rounds until the overall loss function of distillation training reaches its minimum, at which point the distillation training ends.

2. The knowledge distillation method for parameter search and matching according to claim 1, characterized in that, The initialization threshold control parameters and magnitude control parameters, and the calculation of the parameter search matching network for the number of parameters contained in all network layers in each feature output layer of the student network model and the teacher network model, are initialized and calculated by the parameter search matching network.

3. The knowledge distillation method for parameter search and matching according to claim 1, characterized in that, During the distillation training process, threshold control parameters are dynamically adjusted to filter invalid parameters in both the student and teacher network models, including: Distillation training iterations t times, the threshold control parameter is automatically updated to In each round, network layer parameters in the teacher network model that are greater than α are selected. t The number of parameters and the number of network layer parameters in the student network model are greater than α. t The number of parameters; Where α0 represents the initialized threshold control parameter, α t This represents the threshold control parameter for round t, where t is an integer.

4. The knowledge distillation method for parameter search and matching according to claim 1, characterized in that, The number of intermediate layers in a set of teacher and student network models participating in distillation training is controlled by the parameter adaptive matching module in the parameter search matching network, and the control method is expressed by the following formula: η = arg max(|S(i,j)|) Where η represents the magnitude control parameter, and i represents the i-th layer of the student network model. This represents the number of parameters in the j-th layer of the teacher network model that are greater than α. Let α represent the number of parameters in the i-th layer of the student network model that are greater than α, j represent the j-th layer of the teacher network model, S(i,j) represent the set of matching layer pairs of the i-th layer of the student network model and the j-th layer of the teacher network model with the same number of parameters, M represent the last layer of the student network model, and L represent the last layer of the teacher network model.

5. The knowledge distillation method for parameter search and matching according to claim 1, characterized in that, Multi-scale feature alignment pooling is performed on the output features of a set of matched intermediate network layers, and features at different scales are fused through a multi-scale pooling fusion network.

6. The knowledge distillation method for parameter search and matching according to claim 1, characterized in that, The real-time update of the inter-layer loss function is expressed by the following formula: in, and Features output by the i-th layer of the student network model and the j-th layer of the teacher network model of the same order of magnitude, PoF S and PoF T L represents the multi-scale feature alignment pooling operation performed on the output features of the student network model and the teacher network model by the multi-scale pooling fusion network PoF. layer Let α represent the interlayer loss function. t This represents the threshold control parameter for round t, where t is an integer.

7. The knowledge distillation method for parameter search and matching according to claim 1, characterized in that, The overall loss function for distillation training, determined based on the last-layer features of the student network model and the last-layer features of the teacher network model (both of similar magnitude) and the inter-layer loss function, is expressed by the following formula: in, and The features output from the last layer of the student network model and the teacher network model are of the same order of magnitude. F represents the function that transforms the dimensions of the features, D represents the feature distance function, and L... layer Let L represent the inter-layer loss function, i represent the i-th layer of the student network model, j represent the j-th layer of the teacher network model, M represent the last layer of the student network model, and L represent the last layer of the teacher network model. layer L represents the interlayer loss function. KD This represents the overall loss function for distillation training.

8. A knowledge distillation apparatus for parameter search and matching, characterized in that, include: Loading module: Used to randomly initialize the parameters of the student network model and load the teacher network model; Initialization module: Used to initialize threshold control parameters and magnitude control parameters, and calculate the number of parameters contained in all network layers in each feature output layer of the student network model and the teacher network model; Filtering module: Used to dynamically adjust threshold control parameters during distillation training to filter invalid parameters in the student network model and the teacher network model; Control module: Used to control the magnitude of the control parameters to maximize the number of intermediate layers in a set of teacher and student network models participating in distillation training; Update module: Used to perform multi-scale feature alignment pooling on a set of matched intermediate network output features, fuse features at different scales, and update the inter-layer loss function in real time; Determine the overall loss function for distillation training based on the last layer features of the student network model, the last layer features of the teacher network model, and the inter-layer loss function, which are of similar order of magnitude. The training is iterated for multiple rounds until the overall loss function of distillation training reaches its minimum, at which point the distillation training ends.

9. An electronic device, characterized in that, include: Processor and memory; The processor executes a knowledge distillation method for parameter search and matching as described in any one of claims 1 to 7 by calling programs or instructions stored in the memory.

10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores a program or instructions that cause a computer to perform a knowledge distillation method for parameter search matching as described in any one of claims 1 to 7.