Model training method and device, electronic equipment and storage medium
Through dynamic token selection and multi-scale mask training at high spatiotemporal resolution, the teacher-student model distillation method is used to optimize the inference performance of the student model under different computing resource conditions, solving the problem of inflexible calculation amount in downstream applications of the video understanding model, and achieving efficient inference on the end-side device.
Patent Information
- Application Number
- CN202510488771.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-18
- Publication Date
- 2025-07-18
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
The video understanding model in downstream applications has huge computing overhead due to the additional time dimension, which is difficult to effectively utilize on end-side devices. The existing model does not fully consider the flexibility of downstream scenarios during development, resulting in insufficient computing volume.
Through dynamic token selection and multi-scale mask training based on high spatiotemporal resolution, the teacher-student model distillation method is used to generate student training masks of different token counts to optimize the performance of student models under different computational volume limitations.
It improves the inference performance of the student model under various downstream computing volume limitations, and realizes efficient and flexible inference under different computing resource conditions.
Smart Images

Figure CN120339802A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer technology, and in particular, to a model training method, apparatus, electronic device, and storage medium. Background Art
[0002] Video understanding shows great application potential in the field of artificial intelligence, such as its potential in embodied systems and world models, as well as its applications in video retrieval, recommendation, etc. However, the additional time dimension leads to a huge computational overhead, and it is difficult for many downstream applications, such as end-side devices, to utilize the latest large video understanding models. When developing large video understanding models, the best performance under the maximum computational capacity is often pursued, and downstream application scenarios are rarely considered, so it is not flexible enough. Summary of the Invention
[0003] The present invention provides a model training method, apparatus, electronic device, and storage medium to solve the deficiency of video understanding models in downstream flexible inference. It can utilize dynamic token selection at high spatio-temporal resolution to maximize the information content in the input token set corresponding to each input token count, thereby maximizing the utilization of limited computational resources, enabling the model to achieve better inference performance under various downstream computational limitations.
[0004] According to an aspect of the present invention, a model training method is provided. The method includes:
[0005] Sampling the original video frame sequence based on a target training spatio-temporal resolution to obtain target training data; the target training spatio-temporal resolution is higher than the spatio-temporal resolution during downstream inference of the student model;
[0006] Inputting the target training data into a teacher model to generate a target token set through dynamic token selection and generating teacher training features through forward propagation;
[0007] Performing multi-scale cropping on the target token set according to the target self-attention weights of the teacher model to generate at least three different student training masks with different token counts; the target self-attention weights are the weight matrices corresponding to the last multi-head self-attention layer of the teacher model;
[0008] Inputting the target training data and different student training masks into the student model for forward propagation to generate student training features;
[0009] Aligning and distilling the student training features with the teacher training features to obtain the target student model.
[0010] According to another aspect of the present invention, a model training apparatus is provided. The apparatus includes:
[0011] A sampling module, configured to sample an original video frame sequence based on a target training spatio-temporal resolution to obtain target training data; the target training spatio-temporal resolution is higher than the spatio-temporal resolution during downstream inference of the student model;
[0012] A teacher model training module, configured to input the target training data into the teacher model, generate a target token set through dynamic token selection, and generate teacher training features through forward propagation;
[0013] A multi-scale student mask generation module, configured to perform multi-scale cropping on the target token set according to the target self-attention weights of the teacher model to generate at least three student training masks with different token numbers; the target self-attention weights are the weight matrices corresponding to the last multi-head self-attention layer of the teacher model;
[0014] A student model training module, configured to input the target training data and different student training masks into the student model for forward propagation to generate student training features;
[0015] An alignment distillation module, configured to perform alignment distillation on the student training features and the teacher training features to obtain a target student model.
[0016] According to another aspect of the present invention, there is provided an electronic device, which includes:
[0017] At least one processor; and
[0018] A memory communicatively connected to the at least one processor; wherein,
[0019] The memory stores a computer program executable by the at least one processor, and when the computer program is executed by the at least one processor, the at least one processor is enabled to execute the model training method according to any embodiment of the present invention.
[0020] According to another aspect of the present invention, there is provided a computer-readable storage medium storing computer instructions for causing a processor to implement the model training method according to any embodiment of the present invention when executed.
[0021] According to another aspect of the present invention, there is provided a computer program product including a computer program, which implements the model training method according to any embodiment of the present invention when executed by a processor.
[0022] The technical solution of the embodiment of the present invention samples the original video frame sequence based on the target training spatio-temporal resolution to obtain target training data; the target training spatio-temporal resolution is higher than the spatio-temporal resolution during downstream inference of the student model; the target training data is input into the teacher model, and the target token set is generated through dynamic token selection and the teacher training features are generated through forward propagation; the target token set is multi-scale cropped according to the target self-attention weights of the teacher model to generate at least three student training masks with different token numbers; the target self-attention weights are the weight matrices corresponding to the last multi-head self-attention layer of the teacher model; the target training data and different student training masks are input into the student model for forward propagation to generate student training features; the student training features are aligned and distilled with the teacher training features to obtain the target student model. This technical solution takes into account the dynamic nature of downstream inference costs during training, uses dynamic token selection at high spatio-temporal resolution to maximize the information content in the input token set corresponding to each input token number, and effectively improves the generalization ability of the student model under different token input numbers, that is, computational constraints, through the teacher-student model distillation method, solves the deficiency of the video understanding model in flexible downstream inference, and the trained student model can achieve better inference performance under various downstream computational constraints.
[0023] It should be understood that the content described in this part is not intended to identify the key or important features of the embodiments of the present invention, nor is it used to limit the scope of the present invention. Other features of the present invention will become easily understood through the following description. Brief Description of the Drawings
[0024] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention, and those of ordinary skill in the art can obtain other drawings without creative efforts based on these drawings.
[0025] Figure 1 is a flowchart of a model training method provided in Embodiment 1 of the present invention;
[0026] Figure 2 is a flowchart of a model training method provided in Embodiment 2 of the present invention;
[0027] Figure 3 is a schematic diagram of dynamic token selection provided in Embodiment 2 of the present invention;
[0028] Figure 4 is a schematic structural diagram of the teacher model provided in Embodiment 2 of the present invention;
[0029] Figure 5It is a schematic structural diagram of a student model provided in Embodiment 2 of the present invention;
[0030] Figure 6 It is a schematic diagram of feature alignment provided in Embodiment 2 of the present invention;
[0031] Figure 7 It is a schematic structural diagram of a downstream student model provided in Embodiment 2 of the present invention;
[0032] Figure 8 It is a schematic structural diagram of a model training device provided in Embodiment 3 of the present invention;
[0033] Figure 9 It is a schematic structural diagram of an electronic device for implementing the model training method of the embodiments of the present invention. Detailed implementation manners
[0034] In order to enable those skilled in the art to better understand the solution of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.
[0035] It should be noted that the terms "first", "second", etc. in the specification and claims of the present invention and the above drawings are used to distinguish similar objects, and do not necessarily need to describe a specific order or sequence. It should be understood that such used data can be interchanged under appropriate circumstances so that the embodiments of the present invention described here can be implemented in an order other than those illustrated or described here. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device including a series of steps or units does not necessarily need to be limited to those clearly listed steps or units, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.
[0036] It can be understood that the acquisition, storage, use, processing, etc. of data in the technical solution of the present invention all comply with the relevant regulations of national laws and regulations.
[0037] It should be noted that in the embodiments of the present invention, some industry-existing solutions such as certain software, components, models, etc. may be mentioned. They should be regarded as exemplary, and their purpose is only to illustrate the feasibility in the implementation of the technical solution of the present application, but it does not mean that the applicant has already or necessarily used this solution.
[0038] Embodiment 1
[0039] Figure 1 The following is a flowchart of a model training method provided in Embodiment 1 of the present invention. This embodiment is applicable to the situation of dynamic token selection under high spatio-temporal resolution input and model distillation training based on multi-scale masks. This method can be executed by a model training device, which can be implemented in the form of hardware and / or software. The model training device can be configured in an electronic device, which can be a mobile terminal, a PC, a server, etc. As Figure 1 shown, a model training method provided in Embodiment 1 specifically includes the following steps:
[0040] S110. Sample the original video frame sequence based on the target training spatio-temporal resolution to obtain target training data.
[0041] Among them, the teacher model is usually a larger and more complex model with good performance and generalization ability; the network scale of the student model is usually smaller and the expressive ability is limited. Guiding the training of the student model based on the knowledge learned by the teacher model, that is, model distillation, can enable the student model to have performance comparable to that of the teacher model, but the number of parameters is greatly reduced, so as to adapt to downstream applications such as edge devices and embedded systems. Exemplarily, the teacher model and the student model can be video understanding models, which can implement tasks such as action recognition and video text retrieval, and both the teacher model and the student model can include Transformer encoders.
[0042] The original video frame sequence can refer to the training data used for model training of the teacher model and the student model without high spatio-temporal resolution sampling. The target training data can be understood as the video frame sequence obtained by dynamically sampling the original video frame sequence based on a higher target training spatio-temporal resolution.
[0043] The target training spatio-temporal resolution can be understood as the spatio-temporal resolution adopted during model training, and this target training spatio-temporal resolution is higher than the spatio-temporal resolution adopted during downstream inference of the student model. The spatio-temporal resolution can include temporal resolution and spatial resolution. Exemplarily, the spatio-temporal resolution adopted during model training can be 16 (number of frames) × 224 (height) × 224 (width), and the spatio-temporal resolution adopted during downstream inference of the (student) model is 4 × 112 × 112. In this embodiment, the (teacher model) obtains richer spatio-temporal information through high spatio-temporal resolution training, and then compresses it into a low-resolution inference model (student model) through knowledge distillation, so as to maintain model performance while reducing the amount of calculation.
[0044] In an embodiment of the present invention, the pre-set target training spatio-temporal resolution T_train×H_train×W_train and the original video frame sequence used for model training can be obtained, and then image frames corresponding to a corresponding number of frames are extracted from the original video frame sequence according to the target training time resolution T_train, wherein the image frame extraction method can be not limited to equal-interval extraction, random extraction, etc.; and each of the extracted image frames is resized according to the target spatial resolution H_train×W_train, so as to obtain the target training data after dynamic sampling.
[0045] S120. Input the target training data into the teacher model, and generate a target token set through dynamic token selection and generate teacher training features through forward propagation.
[0046] Among them, dynamic token selection can be understood as a process for dynamically selecting the token(s) with the largest amount of information from a high spatio-temporal resolution input. A token is the basic processing unit of the model, and is used to convert high-dimensional video data into a low-dimensional serialized representation suitable for neural network processing, so as to balance information density and computing efficiency.
[0047] The target token set can be understood as the token set with the largest amount of information dynamically selected by the teacher model from a high spatio-temporal resolution input. The target token set will be set to a specified length according to requirements during model training. For example, 2048 tokens are dynamically selected from a high spatio-temporal resolution input: 16×224×224, that is, 3136 tokens (16×16 blocks), etc. The target token set generated through dynamic token selection is the approximate optimal input set, which can achieve information maximization under a given input length and avoid the loss of key information caused by random token reduction.
[0048] The teacher training features can be understood as the feature representations generated by the teacher model when processing high spatio-temporal resolution inputs, and are used to guide the feature learning and parameter optimization of the student model, so as to transfer the understanding ability of the teacher model for complex data to the student model. The teacher training features can include teacher intermediate layer features and / or teacher final layer features; among them, the teacher intermediate layer features can include the output features of each encoding block in the Transformer encoder of the teacher model; the teacher final layer features can include the output features of the last classification layer of the teacher model, or the output features of the attention pooling layer of the teacher model, etc.
[0049] In an embodiment of the present invention, the sampled target training data can be input into a teacher model for processing. The teacher model first performs block encoding (patch embedding) on the target training data to convert the target training data from pixel format to token format, that is, to generate an initial teacher token set. Then, the initial teacher token set and the globally-position encoding pre-initialized and generated are input into the dynamic token selection module of the teacher model, and the target token set of the target length is generated by screening the tokens with the most information. Finally, the generated target token set and the globally-position encoding are input into the Transformer encoding block of the teacher model for forward propagation to generate corresponding teacher training features, which may include, for example, teacher intermediate layer features and teacher final layer features, etc.
[0050] S130. Multiscale cropping is performed on the target token set according to the target self-attention weights of the teacher model to generate at least three student training masks with different token numbers.
[0051] Among them, the target self-attention weights can be understood as the weight matrix corresponding to the last multi-head self-attention layer of the teacher model. The target self-attention weights can reflect the semantic association strength (i.e., importance) between tokens. Tokens with higher importance usually correspond to some key spatio-temporal features such as moving objects, etc.
[0052] Multiscale cropping can be understood as a token screening process for the target token set with different cropping ratios. Exemplarily, tokens corresponding to the corresponding numbers can be cropped from the target token set according to the cropping ratios of 100%, 75%, and 50% respectively.
[0053] The student training mask can be used to represent the retention situation of each token in the target token set. The student training mask can be a binary mask, that is, the position where the token is retained is 1, and the position where the token is not retained is 0. By using the student training mask, the number of input tokens of the student model can be controlled, so that it can adapt to different computational load limitations (such as the number of input tokens ranging from 2048 to 1024) during the training phase.
[0054] In an embodiment of the present invention, after the teacher model performs forward propagation, the target self-attention weights corresponding to its last multi-head self-attention layer can be extracted, and the token importance scores corresponding to each token in the target token set are determined based on the target self-attention weights. Then, they are sorted in descending order according to the token importance scores, so as to generate student training masks with different token numbers. For example, the first 100%, 75%, and 50% of the tokens can be retained respectively to generate the corresponding student training masks.
[0055] S140. The target training data and different student training masks are input into the student model for forward propagation to generate student training features.
[0056] Among them, the student training features can be understood as the feature representations generated by the student model based on different input tokens. Similarly to the teacher training features, the student training features can include student intermediate layer features and / or student final layer features. Among them, the student intermediate layer features can include the output features of each encoding block in the Transformer encoder of the student model; the student final layer features can include the output features of the last classifier of the student model, or the output features of the attention pooling layer of the student model, etc.
[0057] In the embodiment of the present invention, the sampled target training data and the multi-scale student training masks generated according to the teacher model can be respectively input into the student model for forward propagation, that is, corresponding forward propagations are performed according to the number of student training masks in a batch to obtain corresponding student training features, such as student intermediate layer features and student final layer features, etc.
[0058] S150. Align and distill the student training features with the teacher training features to obtain the target student model.
[0059] Among them, alignment distillation can be understood as the core training strategy in the teacher-student framework, aiming to make the feature representations generated by the student model under any number of input tokens semantically consistent with the high-resolution features of the teacher model through feature-level knowledge transfer; the alignment methods can include intermediate layer alignment, final layer alignment, etc. The target student model can refer to the student model after alignment distillation and model parameter optimization.
[0060] In the embodiment of the present invention, the student training features output by the student model can be feature-aligned with the teacher training features output by the teacher model, and then the inter-layer loss after alignment of each layer is used as the total alignment distillation loss, and the student model is backpropagated and parameter-updated using this total alignment distillation loss, so as to obtain the trained target student model. It should be noted that during the backpropagation process of the student model, the model parameters of the teacher model need to be frozen.
[0061] The technical solution of the embodiment of the present invention samples the original video frame sequence based on the target training spatio-temporal resolution to obtain target training data; the target training spatio-temporal resolution is higher than the spatio-temporal resolution during the downstream inference of the student model; the target training data is input into the teacher model, and the target token set is generated through dynamic token selection and the teacher training features are generated through forward propagation; according to the target self-attention weight of the teacher model, multi-scale cropping is performed on the target token set to generate student training masks with at least three different token numbers; the target self-attention weight is the weight matrix corresponding to the last multi-head self-attention layer of the teacher model; the target training data and different student training masks are input into the student model for forward propagation to generate student training features; the student training features are aligned and distilled with the teacher training features to obtain the target student model. This technical solution takes into account the dynamic nature of the downstream inference cost during training, uses dynamic token selection at high spatio-temporal resolution to maximize the information content in the input token set corresponding to each input token number, and effectively improves the generalization ability of the student model under different input token numbers, that is, computational constraints, through the teacher-student model distillation method, solves the deficiency of the video understanding model in flexible downstream inference, and the trained student model can achieve better inference performance under various downstream computational constraints.
[0062] Further, on the basis of the above embodiment of the invention, in S110, sampling the original video frame sequence based on the target training spatio-temporal resolution to obtain target training data specifically includes:
[0063] Randomly select a target training spatio-temporal resolution from the preset spatio-temporal resolution pool, and sample the original video frame sequence of the current training batch using the target training spatio-temporal resolution to obtain the sampled target training data.
[0064] Among them, the preset spatio-temporal resolution pool can be pre-set with spatio-temporal resolution combinations of various time resolutions (number of frames) and space resolutions (sizes) for dynamically adjusting the computational amount of the input data. Exemplarily, 8×224×224 corresponds to a low computational amount, 16×112×112 corresponds to a medium computational amount, and 32×64×64 corresponds to a high computational amount.
[0065] In the embodiment of the present invention, at the beginning of each training batch, a resolution combination can be randomly selected from the preset spatio-temporal resolution pool as the target training spatio-temporal resolution T_train×H_train×W_train, and then image frames corresponding to the corresponding number of frames are equally spacedly extracted from the original video frame sequence according to the target training time resolution T_train, and each extracted image frame is scaled to the target space resolution H_train×W_train, so as to obtain the dynamically sampled target training data.
[0066] Further, based on the above-described invention embodiments, in S120, inputting the target training data into the teacher model to generate the target token set through dynamic token selection and generate the teacher training features through forward propagation specifically includes the following steps:
[0067] S1201: Input the target training data into the patch embedding layer of the teacher model to generate the teacher initial token set;
[0068] S1202: Input the teacher initial token set and the preset global position encoding into the dynamic token selection module of the teacher model to generate the target token set with the target length;
[0069] S1203: Input the target token set and the preset global position encoding into the teacher Transformer encoder and the attention pooling layer of the teacher model to obtain the teacher intermediate layer features output by each first encoding block in the teacher Transformer encoder and the teacher final layer features output by the attention pooling layer, and use the teacher intermediate layer features and the teacher final layer features as the teacher training features.
[0070] Among them, the teacher initial token set can be understood as the token set obtained after the teacher model divides the input target training data into patches and maps them into token vectors.
[0071] The preset global position encoding can be understood as the global position encoding generated after the model is initialized, which is used to provide position information for the tokens. The preset global position encoding can be a learnable position embedding initialized by the sine-cosine method.
[0072] The teacher Transformer encoder can refer to the Transformer encoder configured in the teacher model. The teacher Transformer encoder includes a first number of first encoding blocks. The first encoding block can include a multi-head self-attention layer and a feed-forward neural network. The multi-head self-attention layer can be used to capture the spatio-temporal dependence relationship between tokens. The attention pooling layer can be used to aggregate the features output by the multi-head self-attention layer and retain the high semantic information to adapt to downstream tasks.
[0073] The teacher intermediate layer features can refer to the feature vectors output by the feed-forward neural network in each first encoding block of the teacher model. Each first encoding block corresponds to a teacher intermediate layer feature. The teacher final layer features can refer to the final feature vectors output by the attention pooling layer of the teacher model.
[0074] In an embodiment of the present invention, the target training data can be input into the patch embedding layer, i.e., the PatchEmbedding layer, of the teacher model for patch embedding, so as to convert the pixel format into the token format, i.e., generate a teacher initial token set; then the teacher initial token set and the pre-initialized preset global position encoding (e.g., sine-cosine encoding) are input into the dynamic token selection module of the teacher model, and the target token set of the target length is generated by screening the tokens with the largest amount of information; finally, the generated target token set and the global position encoding are input into the teacher Transformer encoder and the attention pooling layer of the teacher model for forward propagation, and the teacher intermediate layer features output by each first encoding block and the teacher final layer features output by the attention pooling layer are extracted, and the obtained teacher intermediate layer features and teacher final layer features are used as teacher training features.
[0075] Further, based on the above-mentioned embodiment of the invention, the dynamic token selection process in the dynamic token selection module in S1202 specifically includes the following steps:
[0076] Step 1: Call the dynamic token selection module to divide the teacher's initial token set into several sparse blocks according to the preset block size, and determine the token features corresponding to each token in each sparse block based on the preset global position encoding; wherein the token features include visual semantic information and spatial position information;
[0077] Step 2: In each sparse block, the target norm distance between token features at the same spatial position in adjacent frames is used to determine the dynamic value corresponding to each token;
[0078] Step 3. Use dynamic values to arrange the tokens in each sparse block in descending order, and select the first preset number of tokens from each sparse block to generate the corresponding teacher training mask, and determine the teacher's initial token set after the teacher training mask processing as the target token set; wherein the target length is equal to the product of the first preset number and the number of blocks in the sparse block.
[0079] Among them, the token feature may include the visual semantic information and spatial position information corresponding to the token. The visual semantic information may refer to the semantic feature vector information generated after the target training data is patch embedded; the spatial position information may refer to the position information assigned based on the global position encoding.
[0080] The target norm distance can be understood as an indicator for measuring the difference in token features at the same spatial position in adjacent frames. The target norm distance can at least include the L2 norm distance, etc.
[0081] In an embodiment of the present invention, the dynamic token selection process specifically includes:
[0082] ①Divide the initial teacher token set into several sparse blocks according to the preset block size. Each sparse block contains several frames, and each frame contains several tokens. For example, for the target training data {F1, F2,..., F T}(T represents the number of frames), uniformly sample it into N sparse blocks B i ={F t |t ∈ [iB, (i + 1)B - 1]}, where i ∈ [0, N).
[0083] ②Add position information to each token based on the preset global position information generated by pre-initialization, so as to obtain the token feature corresponding to each token.
[0084] ③For the token at the j-th spatial position in each sparse block B i , calculate the difference from the same position in the previous frame, that is, the dynamic value: D(F t+1,j ) = ||F t+1,j - F t,j || p , where F t,j represents the token feature of the token corresponding to the j-th spatial position in the t-th frame, and p represents the p-norm value of the distance. It can be understood that the larger the dynamic value, the more significant the temporal change and the higher the information content at this spatial position.
[0085] ④Sort all the tokens in each sparse block in descending order of the dynamic value, and then select the first preset number of tokens from each block respectively, that is, select TOP(M / N) tokens respectively, where M is the target length, that is, the total number of target tokens to be selected, and N is the number of sparse blocks.
[0086] ⑤Generate a corresponding teacher training mask based on the token selection situation of each sparse block to identify the selected token positions, and then apply the generated teacher training mask to the teacher initial token set to obtain the final target token set.
[0087] In the embodiment of the present invention, by adopting the sparse dynamic token selection strategy, tokens with significant temporal changes and rich semantics can be screened out, ensuring the maximization of the information density of the input tokens under limited downstream computing power.
[0088] Further, on the basis of the above-mentioned embodiment of the invention, in S130, multi-scale cropping is performed on the target token set according to the target self-attention weights of the teacher model to generate at least three student training masks with different token numbers, which specifically includes the following steps:
[0089] S1301. Extract the target self-attention weights, and determine the token importance scores corresponding to each token in the target token set based on the target self-attention weights.
[0090] S1302. Sort each token in the target token set in descending order according to the token importance score to generate a token importance ranking list.
[0091] S1303. Select corresponding numbers of tokens in the token importance ranking list according to at least three different preset token cropping ratios to generate at least three corresponding student training masks.
[0092] Among them, the token importance score can be used to characterize the critical contribution of a token at a specific spatio-temporal position in a video frame to model inference. Tokens with high scores (such as the edges of moving objects, significant semantic regions) are preferentially retained, and tokens with low scores (such as static backgrounds, redundant information) will be filtered out by the mask.
[0093] The token importance ranking list can be understood as a token list obtained by sorting each token in the target token set in descending order according to the token importance score.
[0094] The preset token cropping ratio can be understood as a pre-configured multi-scale token cropping ratio, and the preset token cropping ratio can at least include 100% (fully retained), 75% (retain the first 75% of the tokens), and 50% (retain the first 50% of the tokens), etc.
[0095] In the embodiment of the present invention, the generation process of the multi-scale student training mask includes:
[0096] ① Extract the target self-attention weight after the teacher model processes the target token set. The target self-attention weight can be the weight matrix corresponding to the last multi-head self-attention layer of the teacher model.
[0097] ② Use the target self-attention weight to determine the token importance score corresponding to each token in the target token set. For example, the following token importance score calculation formula can be adopted:
[0098]
[0099] In the formula, Score(T u ) represents the token importance score corresponding to the u-th token; A u,v represents the attention value of the u-th token to the v-th token in the target self-attention weight.
[0100] ③ Sort each token in the target token set in descending order according to the token importance score to generate a token importance ranking list.
[0101] ④According to at least three different preset token cropping ratios configured in advance, such as 100%, 75%, and 50%, extract the top corresponding ratio of tokens from the token importance ranking list respectively. For example, if the target token set contains 2048 tokens, the target sub-token sets obtained after multi-scale cropping can contain 2048, 1536, and 1024 tokens respectively.
[0102] ⑤Generate corresponding student training masks according to the token selection conditions of different cropping ratios to identify the positions of the selected tokens.
[0103] Through the token importance evaluation driven by self-attention weights and multi-scale mask generation in the embodiments of the present invention, dynamic training optimization of the subsequent student model under different computational constraints can be achieved.
[0104] Further, based on the above embodiments of the invention, in S140, inputting the target training data and different student training masks into the student model for forward propagation to generate student training features specifically includes the following steps:
[0105] S1401: Input the target training data into the dual normalization module of the student model to generate a student initial token set;
[0106] S1402: Input the preset global position encoding into the depthwise separable convolution module of the student model to generate a target enhanced position encoding;
[0107] S1403: Respectively determine the student initial token sets processed by different student training masks as target sub-token sets;
[0108] S1404: For each target sub-token set, input the target sub-token set and the target enhanced position encoding into the student Transformer encoder and the attention pooling layer of the student model to obtain the student intermediate layer features output by each second encoding block in the student Transformer encoder and the student final layer features output by the attention pooling layer, and use the student intermediate layer features and student final layer features obtained under all student training masks as student training features.
[0109] Among them, the dual normalization module can be used to alleviate the input distribution shift problem caused by the dynamic token selection of the model, stabilize model training, and thus enhance the robustness of the model to variable spatio-temporal resolution inputs. The dual normalization module can include: a first normalization layer, a patch embedding layer, and a second normalization layer. Among them, the first normalization layer is used to perform layer normalization (Layer Normalization) on the target training data (pixel data) and output the normalized pixel data, the purpose of which is to stabilize the original input distribution and alleviate the influence of noises such as illumination and contrast; the patch embedding layer is used to perform patch embedding on the normalized pixel data and output the corresponding token set, the purpose of which is to map the spatio-temporal local region into a high-dimensional feature vector to form an initial token representation; the second normalization layer is used to perform layer normalization (Layer Normalization) on the tokens after patch embedding and output the normalized token set, the purpose of which is to alleviate the distribution shift caused by dynamic token selection and enhance the stability of token embedding.
[0110] The depth-wise separable convolution (DWConv) module can refer to a component used to achieve enhanced position awareness of dynamic sparse inputs. Using the depth-wise separable convolution module can enhance the model's ability to perceive the spatial position of input tokens, that is, it can capture broader context relationships in the input data, thereby providing a more robust global position embedding.
[0111] The target sub-token set can be understood as the token set obtained by applying the student training mask to the student initial token set output by the dual normalization module. It is the token set with maximized information content filtered from high spatio-temporal resolution inputs according to the computational load limit of the downstream task (i.e., the number of tokens allowed to be processed).
[0112] The student Transformer encoder can refer to the Transformer encoder configured in the student model. The student Transformer encoder includes a second number of second encoding blocks, and the second encoding block can include an improved multi-head self-attention layer and a feed-forward neural network. Among them, the improved multi-head self-attention layer adds a local encoding component relative to the standard multi-head self-attention layer in the teacher model, that is, a linear projection function (Linear Projection Function, LPE) is introduced into the attention mechanism and applied to the value matrix V to adapt to token sets with different input lengths and sparse patterns.
[0113] Similarly to the teacher model, the student intermediate layer features can refer to the feature vectors output by the feed-forward neural network in each second encoding block of the student model, and each second encoding block corresponds to a student intermediate layer feature. The student final layer features can refer to the final feature vector output by the attention pooling layer of the student model.
[0114] In the embodiments of the present invention, the sampled target training data can be input into the dual normalization module of the student model, and after being processed by the first normalization layer, the patch embedding layer, and the second normalization layer in sequence, a set of student initial tokens is generated; at the same time, a preset global position encoding (such as sine-cosine encoding) generated by pre-initialization is input into the depthwise separable convolution module of the student model for position enhancement processing to generate a corresponding target enhanced position encoding; then the generated student training masks of different scales are applied to the set of student initial tokens to obtain a corresponding number of sets of target sub-tokens; finally, the generated multiple sets of target sub-tokens and the target enhanced position encoding are respectively input into the student Transformer encoder and the attention pooling layer of the student model for forward propagation, extracting the student intermediate layer features output by each second encoding block and the student final layer features output by the attention pooling layer, and taking the student intermediate layer features and the student final layer features obtained under all student training masks as student training features.
[0115] Furthermore, on the basis of the above embodiments of the invention, in the improved multi-head self-attention layer, the calculation formula of a single attention head can be expressed as:
[0116]
[0117] where Z represents the output of a single attention head; Q represents the query matrix; K represents the key matrix; V represents the value matrix; T represents the transpose symbol; D represents the feature dimension; Softmax represents the normalized exponential function; and LPE represents the linear projection function.
[0118] It should be understood that in the distillation training process of the teacher model - student model, there are the following differences in the structural compositions of the two models:
[0119] ① The teacher model uses a standard patch normalization layer to generate a set of initial tokens, while the student model uses a dual normalization module (the first normalization layer + the patch embedding layer + the second normalization layer) to generate a set of initial tokens;
[0120] ② The student model uses a depthwise separable convolution module to perform position enhancement processing on the preset global position encoding, while the teacher model does not use the above depthwise separable convolution module;
[0121] ③ The teacher model will use a dynamic token selection module to generate a set of target tokens of the target length, while the student model does not contain a dynamic token selection module. Instead, it generates a student training mask based on the self-attention weights output by the teacher model, thereby generating a set of target sub-tokens of different input lengths;
[0122] ④ The first encoding block in the teacher model uses a standard multi-head self-attention layer, while the second encoding block in the student model uses an improved multi-head self-attention layer. That is, a local encoding component LPE is added compared to the standard multi-head self-attention layer, which can enhance the dynamic perception of local positional relationships through linear projection.
[0123] Further, based on the above invention embodiments, in S150, the student training features and the teacher training features are aligned and distilled to obtain a target student model, which specifically includes the following steps:
[0124] S1501. Determine the alignment intermediate layer index corresponding to the teacher model according to the preset inter-layer alignment rule;
[0125] S1502. Extract the corresponding target teacher intermediate layer features from the teacher training features according to the alignment intermediate layer index;
[0126] S1503. For each student intermediate layer feature in the student training features, determine the intermediate layer alignment loss between the student intermediate layer feature and the corresponding target teacher intermediate layer feature;
[0127] S1504. Determine the final layer alignment loss between the student final layer feature in the student training features and the teacher final layer feature in the teacher training features;
[0128] S1505. Determine the weighted sum result of each intermediate layer alignment loss and the final layer alignment loss as the alignment distillation total loss;
[0129] S1506. Based on the alignment distillation total loss, call the preset gradient descent method to update the parameters of the student model to obtain the target student model.
[0130] Among them, the preset inter-layer alignment rule can be understood as the mapping relationship between the intermediate layers of the pre-configured teacher model and the learning model. The number of layers of the first encoding block included in the teacher model and the number of layers of the second encoding block included in the student model can be the same or different. Usually, the number of layers of the first encoding block in the teacher model is higher than the number of layers of the second encoding block in the student model. When the number of layers of both is the same, that is, each first encoding block included in the teacher model is aligned with the second encoding block included in the student model one by one; if the number of layers L T of the first encoding block is higher than the number of layers L S of the second encoding block, then the ratio L T / L SIn the teacher model, select the target teacher intermediate layer to be aligned, that is, the target first coding block. Exemplarily, if the teacher model has 12 layers and the student model has 6 layers, the indexes of the target teacher intermediate layers to be aligned (i.e., the alignment intermediate layer indexes) can be 1, 3, 5, 7, 9, and 11.
[0131] The target teacher intermediate layer features can be understood as the teacher intermediate layer features output by the target first coding block to be aligned, that is, the teacher intermediate layer features corresponding to the alignment intermediate layer indexes.
[0132] The intermediate layer alignment loss can be understood as the loss between the student intermediate layer features output by each second coding block of the student model and the teacher intermediate layer features output by the target first coding block to be correspondingly aligned, that is, the loss between the intermediate layers of the two models. Continuing with the above example where the teacher model has 12 layers and the student model has 6 layers, the intermediate layer alignment loss can include the alignment losses corresponding to 6 intermediate layers. Among them, the alignment loss corresponding to each intermediate layer can be determined by methods such as cosine similarity, KL divergence (Kullback-Leibler Divergence), and mean square error (MSE).
[0133] The final layer alignment loss can be understood as the alignment loss between the student final layer features and the teacher final layer features, and the final layer alignment loss can also be determined by methods such as cosine similarity, KL divergence, and MSE.
[0134] The alignment distillation total loss can be understood as the overall loss situation of the intermediate layer alignment losses and the final layer alignment loss under all student training masks, and it can be represented by the weighted sum result of the intermediate layer alignment losses and the final layer alignment loss under all student training masks.
[0135] In the embodiment of the present invention, the process of alignment distillation specifically includes:
[0136] ① Determine the indexes of the target teacher intermediate layers to be aligned in the teacher model based on the pre-configured preset inter-layer alignment rules, that is, determine the alignment intermediate layer indexes;
[0137] ② Extract the corresponding target teacher intermediate layer features from the previously generated teacher training features according to the determined alignment intermediate layer indexes;
[0138] ③ Under each student training mask, determine the intermediate layer alignment loss between each student intermediate layer feature and the corresponding target teacher intermediate layer feature, and determine the final layer alignment loss between the student final layer feature and the teacher final layer feature;
[0139] ④Determine the weighted sum of the intermediate layer alignment losses and the final layer alignment losses under all student training masks as the final total alignment distillation loss;
[0140] In one embodiment, the total alignment distillation loss can be expressed as follows:
[0141]
[0142] In the formula, L total represents the total alignment distillation loss; represents the intermediate layer alignment loss of the l-th intermediate layer under the e-th student training mask; represents the final layer alignment loss under the e-th student training mask; α l and β respectively represent the weight coefficients of each intermediate layer and the final layer; E represents the number of student training masks; L S represents the number of intermediate layers.
[0143] ⑤Based on the determined total alignment distillation loss, use a pre-configured preset gradient descent method (such as Adam gradient descent method, stochastic gradient descent method, etc.) to update the model parameters of the student model, so as to train the target student model. It should be understood that during the backpropagation and parameter update of the student model, the parameters of the teacher model do not need to be updated and will be frozen.
[0144] By considering the dynamics of the downstream inference cost in the above model training, using dynamic token selection at high spatio-temporal resolution, maximizing the information content in the input token set corresponding to each input token number, and effectively improving the generalization ability of the student model under different token input numbers (i.e., computational constraints) through the teacher-student model distillation method, the deficiency of the video understanding model in flexible downstream inference is solved, and the trained student model can achieve better inference performance under various downstream computational constraints.
[0145] Embodiment 2
[0146] Figure 2 The flowchart of a model training method provided in Embodiment 2 of the present invention. Based on the above embodiment, this embodiment provides an implementation manner of a model training method, which can achieve dynamic token selection under high spatio-temporal resolution input and model distillation training according to multi-scale masks. As Figure 2 shown, a model training method provided in Embodiment 2 of the present invention specifically includes the following steps:
[0147] S210. Randomly select a target training spatio-temporal resolution to sample the original video frame sequence of the current batch to obtain the sampled target training data.
[0148] S220. Input the target training data into the patch embedding layer of the teacher model and the dual normalization module of the student model respectively to generate the corresponding teacher initial token set and student initial token set.
[0149] In the embodiment of the present invention, since the token selection mechanism is based on block-based visual tokens, the patch embedding layer plays a crucial role in dynamic estimation and must be robust to various spatio-temporal resolution distributions. It can be observed through experiments that adding a Layer Normalization layer, i.e., the second normalization layer, after the standard patch embedding layer can improve the accuracy of dynamic estimation, thereby achieving more effective token selection. In addition, flexible sampling strategies lead to diverse distributions of input tokens. Therefore, compared with other modules, the gradient norm of the patch embedding layer may become too large. To stabilize the training and enhance the patch embedding layer, a second Layer Normalization layer, i.e., the first normalization layer, is introduced before the patch embedding operation. This dual normalization scheme effectively stabilizes the training dynamics and can enhance the robustness of the student model under different input conditions.
[0150] S230. Input the teacher initial token set and the preset global position encoding into the dynamic token selection module of the teacher model to generate a target token set with a target length.
[0151] As Figure 3 shown, the existing token pruning strategies can reduce data redundancy, but they will cause model performance loss due to inevitable information loss; while the dynamic token selection strategy proposed in the embodiment of the present invention can select the token set with the largest amount of information from higher spatio-temporal resolution data, thereby maximizing the input information density.
[0152] S240. Input the target token set and the preset global position encoding into the teacher Transformer encoder and the attention pooling layer of the teacher model for forward propagation to obtain the teacher intermediate layer features output by each first encoding block in the teacher Transformer encoder and the teacher final layer features output by the attention pooling layer.
[0153] In the embodiment of the present invention, as Figure 4 shown, the generated target token set and the preset global position encoding generated by initialization can be input into the teacher Transformer encoder and the attention pooling layer (AttenPool) of the teacher model for forward propagation to extract the teacher intermediate layer features output by each first encoding block and the teacher final layer features output by the attention pooling layer; wherein, the teacher Transformer encoder includes a first number of first encoding blocks, and the first encoding block includes a multi-head self-attention layer (MHSA) and a feed-forward neural network (FFN).
[0154] S250. Extract the target self-attention weights corresponding to the last multi-head self-attention layer of the teacher model, and generate a token importance ranking list corresponding to the target token set using the target self-attention weights.
[0155] S260. Select corresponding numbers of tokens in the token importance ranking list according to three different preset token pruning ratios to generate three corresponding student training masks.
[0156] In the embodiment of the present invention, the teacher model uses a dynamic token selection module to generate a target token set of 2048 tokens. After pruning by three preset token pruning ratios of 100%, 75%, and 50%, three different scales of student training masks of 2048, 1536, and 1024 will be obtained.
[0157] S270. Input the preset global position encoding into the depthwise separable convolution module of the student model to generate the target enhanced position encoding.
[0158] In the embodiment of the present invention, standard Positional Embeddings are difficult to handle the variable token numbers and sparse patterns inherent in the model training method proposed in the embodiment of the present invention. As Figure 5 shown, in order to enable the model to effectively utilize the position information in different resolutions and token distributions, this solution enhances the standard learnable position embedding (initialized using the sine-cosine method to adapt to the largest possible input size) by adopting a depthwise separable convolution (DWConv) module. This method enables the model to capture broader context relationships in the input sequence, thereby providing more robust global position embeddings.
[0159] S280. Respectively determine the student initial token sets processed by different student training masks as the target sub-token sets.
[0160] S290. For each target sub-token set, input the target sub-token set and the target enhanced position encoding into the student Transformer encoder and the attention pooling layer of the student model for forward propagation to obtain the student intermediate layer features output by each second encoding block in the student Transformer encoder and the student final layer features output by the attention pooling layer.
[0161] In the embodiment of the present invention, as Figure 5As shown, the generated enhanced global embedding, i.e., the target enhanced position encoding, is added to the output of the dual normalization module. Next, in order to further optimize the position information at a finer granularity, in this embodiment, an improved multi-head self-attention layer is adopted in the second encoding block, which adds a local encoding component compared with the standard multi-head self-attention layer, that is, a linear projection function (LPE) is introduced into the attention mechanism and applied to the value matrix V. Specifically, in the improved multi-head self-attention layer, the calculation formula of a single attention head can be expressed as:
[0162]
[0163] By adopting the above global-local position embedding mechanism, that is, the depthwise separable convolution (DWConv) module and the local encoding component (LPE), it is possible to adapt to token sets with different input lengths and sparse patterns, thereby improving the model performance.
[0164] S2100. Align and distill the student intermediate layer features and the student final layer features with the corresponding teacher intermediate layer features and teacher final layer features respectively to obtain the target student model with optimized model parameters.
[0165] In the embodiment of the present invention, during the distillation training process of the teacher-student model, three different input quantities are introduced for the student model in a single batch, so as to make full use of the computing power of the teacher model and improve the learning ability of the model under different input sizes. As Figure 6 shown, by performing intermediate layer alignment and final layer alignment (global alignment) on the student model and the teacher model, and optimizing the parameters of the student model based on the alignment distillation total loss, the trained student model can efficiently inherit the knowledge of the teacher model and significantly improve the performance and robustness under different computational constraints.
[0166] Furthermore, after obtaining the pre-trained student model, it can be fine-tuned with downstream task data to make the student model adapt to specific scenarios (such as action recognition, video retrieval, etc.), while maintaining its dynamic inference ability. In addition, as Figure 7 shown, compared with the student model in the distillation stage, the downstream student model at this time adds a dynamic token selection module, which can set the maximum number of input tokens according to the computing power of the deployment device. Further, during the fine-tuning process of the student model, different spatio-temporal resolutions can be selected for testing of downstream tasks, and the corresponding spatio-temporal resolution with the best test results of the downstream tasks is selected as the target spatio-temporal resolution of the student model input. In summary, through fine-tuning and deployment optimization, the student model can flexibly adapt to the computational requirements of different downstream tasks, while maintaining high accuracy and meeting the real-time and resource limitations of the actual scenario.
[0167] The model training method provided by the embodiments of the present invention is a new method that can improve the model performance through the token optimization stage with higher spatio-temporal resolution under any computational load. It has at least the following advantages: ① The method of sparse dynamic sampling can be used to select an approximately optimal input token set of a given input size from the input with higher spatio-temporal resolution, maximizing the information volume in the input token set corresponding to each input token number, so as to make the best use of the limited computational load, enabling the model to achieve better performance under various downstream computational load limitations; ② Using the model distillation method can support high-performance performance under token sets of any input length, and can be widely applied to single-modal and multi-modal training methods, and can be efficiently coupled with existing pre-training methods; ③ This method is a plug-in supplementary solution for existing pre-training methods, can be used in large-scale pre-training methods, and has achieved the current optimal performance on various downstream tasks through this pre-training.
[0168] Embodiment III
[0169] Figure 8 It is a schematic structural diagram of a model training device provided by Embodiment III of the present invention. As Figure 8 shown, the device includes:
[0170] A sampling module 31, configured to sample the original video frame sequence based on the target training spatio-temporal resolution to obtain target training data; the target training spatio-temporal resolution is higher than the spatio-temporal resolution during the downstream inference of the student model;
[0171] A teacher model training module 32, configured to input the target training data into the teacher model, generate a target token set through dynamic token selection, and generate teacher training features through forward propagation;
[0172] A multi-scale student mask generation module 33, configured to perform multi-scale cropping on the target token set according to the target self-attention weights of the teacher model to generate at least three student training masks with different token numbers; the target self-attention weights are the weight matrices corresponding to the last multi-head self-attention layer of the teacher model;
[0173] A student model training module 34, configured to input the target training data and different student training masks into the student model for forward propagation to generate student training features;
[0174] An alignment distillation module 35, configured to perform alignment distillation on the student training features and the teacher training features to obtain a target student model.
[0175] Further, on the basis of the above-mentioned embodiments of the invention, the sampling module 31 includes:
[0176] A sampling unit, configured to randomly select a target training spatio-temporal resolution from a preset spatio-temporal resolution pool, and sample the original video frame sequence of the current training batch using the target training spatio-temporal resolution to obtain the sampled target training data.
[0177] Further, on the basis of the above-mentioned invention embodiments, the teacher model training module 32 includes:
[0178] A teacher initial token set generation unit, configured to input the target training data into the patch embedding layer of the teacher model to generate a teacher initial token set;
[0179] A dynamic token selection unit, configured to input the teacher initial token set and a preset global position encoding into the dynamic token selection module of the teacher model to generate a target token set with a target length;
[0180] A teacher training feature generation unit, configured to input the target token set and a preset global position encoding into the teacher Transformer encoder and the attention pooling layer of the teacher model, obtain the teacher intermediate layer features output by each first encoding block in the teacher Transformer encoder and the teacher final layer features output by the attention pooling layer, and use the teacher intermediate layer features and the teacher final layer features as teacher training features; wherein, the teacher Transformer encoder includes a first number of first encoding blocks, and the first encoding block includes a multi-head self-attention layer and a feed-forward neural network.
[0181] Further, on the basis of the above-mentioned invention embodiments, the dynamic token selection unit is specifically configured to:
[0182] Call the dynamic token selection module to divide the teacher initial token set into several sparse blocks according to a preset block size, and determine the token features corresponding to each token in each sparse block based on the preset global position encoding; wherein, the token features include visual semantic information and spatial position information;
[0183] In each sparse block, use the target norm distance between the token features at the same spatial position of adjacent frames to determine the dynamic value corresponding to each token;
[0184] Use the dynamic values to sort the tokens in each sparse block in descending order respectively, and select the first preset number of tokens from each sparse block respectively to generate the corresponding teacher training mask, and determine the teacher initial token set processed by the teacher training mask as the target token set; wherein, the target length is equal to the product of the first preset number and the number of sparse blocks.
[0185] Further, on the basis of the above-mentioned invention embodiments, the multi-scale student mask generation module 33 includes:
[0186] A token importance score determination unit, configured to extract target self-attention weights and determine the token importance scores corresponding to each token in the target token set based on the target self-attention weights;
[0187] A token importance ranking list generation unit, configured to rank each token in the target token set in descending order according to the token importance scores, and generate a token importance ranking list;
[0188] A student training mask generation unit, configured to select corresponding numbers of tokens in the token importance ranking list according to at least three different preset token cropping ratios to generate at least three corresponding student training masks; wherein the preset token cropping ratios include at least 100%, 75%, and 50%.
[0189] Further, based on the above-mentioned invention embodiments, the student model training module 34 includes:
[0190] A student initial token set generation unit, configured to input target training data into the dual normalization module of the student model to generate a student initial token set; wherein the dual normalization module includes: a first normalization layer, a patch embedding layer, and a second normalization layer;
[0191] A target enhanced position encoding generation unit, configured to input a preset global position encoding into the depthwise separable convolution module of the student model to generate a target enhanced position encoding;
[0192] A mask processing unit, configured to respectively determine the student initial token sets processed by different student training masks as target sub-token sets;
[0193] A student training feature generation unit, configured to, for each target sub-token set, input the target sub-token set and the target enhanced position encoding into the student Transformer encoder and the attention pooling layer of the student model, obtain the student intermediate layer features output by each second encoding block in the student Transformer encoder and the student final layer features output by the attention pooling layer, and use the student intermediate layer features and the student final layer features obtained under all student training masks as student training features; wherein the student Transformer encoder includes a second number of second encoding blocks, and the second encoding block includes an improved multi-head self-attention layer and a feed-forward neural network.
[0194] Further, based on the above-mentioned invention embodiments, the alignment distillation module 35 includes:
[0195] An index determination unit, configured to determine the alignment intermediate layer index corresponding to the teacher model according to a preset inter-layer alignment rule;
[0196] The target teacher intermediate layer feature extraction unit is used to extract the corresponding target teacher intermediate layer features from the teacher training features according to the aligned intermediate layer index;
[0197] The intermediate layer alignment loss determination unit is used to determine the intermediate layer alignment loss between each student intermediate layer feature in the student training features and the corresponding target teacher intermediate layer feature;
[0198] The final layer alignment loss determination unit is used to determine the final layer alignment loss between the student final layer feature in the student training features and the teacher final layer feature in the teacher training features;
[0199] The alignment distillation total loss determination unit is used to determine the alignment distillation total loss by taking the weighted sum result of each intermediate layer alignment loss and the final layer alignment loss;
[0200] The student model parameter update unit is used to update the parameters of the student model by calling a preset gradient descent method based on the alignment distillation total loss to obtain the target student model.
[0201] Furthermore, on the basis of the above-mentioned invention embodiments, in the improved multi-head self-attention layer, the calculation formula of a single attention head is:
[0202]
[0203] where Z represents the output of a single attention head; Q represents the query matrix; K represents the key matrix; V represents the value matrix; T represents the transpose symbol; D represents the feature dimension; Softmax represents the normalized exponential function; LPE represents the linear projection function.
[0204] The model training device provided by the embodiments of the present invention can execute the model training method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects for executing the method.
[0205] Embodiment 4
[0206] Figure 9 FIG. shows a schematic structural diagram of an electronic device 40 that can be used to implement the embodiments of the present invention. The electronic device is intended to represent various forms of digital computers, such as, laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as, personal digital processing, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are only examples and are not intended to limit the implementation of the present invention described and / or claimed herein.
[0207] As shown Figure 9 in FIG. 1, the electronic device 40 includes at least one processor 41 and a memory communicatively connected to the at least one processor 41, such as a read-only memory (ROM) 42, a random access memory (RAM) 43, etc. The memory stores a computer program executable by the at least one processor. The processor 41 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 42 or the computer program loaded from the storage unit 48 into the random access memory (RAM) 43. In the RAM 43, various programs and data required for the operation of the electronic device 40 can also be stored. The processor 41, the ROM 42, and the RAM 43 are connected to each other via a bus 44. An input / output (I / O) interface 45 is also connected to the bus 44.
[0208] A plurality of components in the electronic device 40 are connected to the I / O interface 45, including: an input unit 46, such as a keyboard, a mouse, etc.; an output unit 47, such as various types of displays, speakers, etc.; a storage unit 48, such as a magnetic disk, an optical disk, etc.; and a communication unit 49, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 49 allows the electronic device 40 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.
[0209] The processor 41 can be various general-purpose and / or special-purpose processing components with processing and computing capabilities. Some examples of the processor 41 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various dedicated artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any suitable processor, controller, microcontroller, etc. The processor 41 executes the various methods and processes described above, such as the model training method.
[0210] In some embodiments, the model training method can be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 48. In some embodiments, part or all of the computer program can be loaded and / or installed onto the electronic device 40 via the ROM 42 and / or the communication unit 49. When the computer program is loaded into the RAM 43 and executed by the processor 41, one or more steps of the model training method described above can be executed. Alternatively, in other embodiments, the processor 41 can be configured to execute the model training method in any other suitable manner (e.g., by means of firmware).
[0211] The various embodiments of the systems and techniques described above in this document can be implemented in digital electronic circuitry, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), systems on a chip (SOCs), complex programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include: being implemented in one or more computer programs that are executable and / or interpretable on a programmable system including at least one programmable processor, which can be a special-purpose or general-purpose programmable processor that receives data and instructions from a storage system, at least one input device, and at least one output device, and transmits the data and instructions to the storage system, the at least one input device, and the at least one output device.
[0212] In some embodiments, the model training method can be implemented as a computer program that is tangibly embodied in a computer program product, the computer program implementing the model training method of the present invention when executed by a processor. The computer program product can be understood as a software product that mainly realizes its solution through a computer program. The computer program for implementing the method of the present invention can be written in any combination of one or more programming languages. These computer programs can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing devices, such that when the computer programs are executed by the processor, the functions / operations specified in the flowchart and / or block diagram are implemented. The computer programs can be executed entirely on the machine, partially on the machine, as a stand-alone software package partially on the machine and partially on a remote machine, or entirely on a remote machine or server.
[0213] In the context of the present invention, a computer-readable storage medium can be a tangible medium that can contain or store a computer program for use by or in connection with an instruction execution system, apparatus, or device. The computer-readable storage medium can include, but is not limited to, electronic, magnetic, optical, electromagnetic, infrared, or semiconductor systems, apparatus, or devices, or any suitable combination of the foregoing. Alternatively, the computer-readable storage medium can be a machine-readable signal medium. More specific examples of the machine-readable storage medium would include electrical connections based on one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.
[0214] To provide interaction with users, the systems and techniques described herein can be implemented on an electronic device having: a display device for displaying information to the user (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor); and a keyboard and a pointing device (e.g., a mouse or a trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with users; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).
[0215] The systems and techniques described herein can be implemented in a computing system including backend components (e.g., as a data server), or a computing system including middleware components (e.g., an application server), or a computing system including frontend components (e.g., a user computer having a graphical user interface or a web browser through which the user can interact with an implementation of the systems and techniques described herein), or a computing system including any combination of such backend components, middleware components, or frontend components. The components of the system can be interconnected with each other by digital data communication in any form or medium (e.g., a communication network). Examples of communication networks include: local area network (LAN), wide area network (WAN), blockchain network, and the Internet.
[0216] A computing system can include a client and a server. The client and the server are generally far from each other and usually interact through a communication network. The client-server relationship is created by computer programs running on the respective computers and having a client-server relationship with each other. The server can be a cloud server, also known as a cloud computing server or a cloud host, which is a host product in the cloud computing service system and solves the defects of difficult management and weak business scalability existing in traditional physical hosts and VPS services.
[0217] It should be understood that various forms of the processes shown above can be used, with steps reordered, added, or deleted. For example, the steps recited in the present invention can be executed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved, and no limitation is made herein.
[0218] The above specific embodiments do not constitute a limitation to the protection scope of the present invention. Those skilled in the art should understand that various modifications, combinations, sub-combinations, and substitutions can be made according to design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A model training method, characterized in that, The method includes: Sampling the original video frame sequence based on the target training spatio-temporal resolution to obtain target training data; the target training spatio-temporal resolution is higher than the spatio-temporal resolution during downstream inference of the student model. Inputting the target training data into the teacher model to generate a target token set through dynamic token selection and generate teacher training features through forward propagation. Performing multi-scale cropping on the target token set according to the target self-attention weights of the teacher model to generate at least three student training masks with different token numbers; the target self-attention weights are the weight matrices corresponding to the last multi-head self-attention layer of the teacher model. Inputting the target training data and different student training masks into the student model for forward propagation to generate student training features. Aligning and distilling the student training features with the teacher training features to obtain the target student model.
2. The method according to claim 1, characterized in that, The sampling the original video frame sequence based on the target training spatio-temporal resolution to obtain target training data includes: Randomly selecting one of the target training spatio-temporal resolutions from a preset spatio-temporal resolution pool, and sampling the original video frame sequence of the current training batch using the target training spatio-temporal resolution to obtain the sampled target training data.
3. The method according to claim 1, characterized in that, The inputting the target training data into the teacher model to generate a target token set through dynamic token selection and generate teacher training features through forward propagation includes: Inputting the target training data into the patch embedding layer of the teacher model to generate an initial teacher token set. Inputting the initial teacher token set and a preset global position encoding into the dynamic token selection module of the teacher model to generate the target token set with a target length. Inputting the target token set and the preset global position encoding into the teacher Transformer encoder and attention pooling layer of the teacher model to obtain the teacher intermediate layer features output by each first encoding block in the teacher Transformer encoder and the teacher final layer features output by the attention pooling layer, and using each of the teacher intermediate layer features and the teacher final layer features as the teacher training features; wherein, the teacher Transformer encoder includes a first number of the first encoding blocks, and the first encoding block includes a multi-head self-attention layer and a feed-forward neural network.
4. The method according to claim 3, characterized in that, The inputting the initial teacher token set and a preset global position encoding into the dynamic token selection module of the teacher model to generate the target token set with a target length includes: Invoking the dynamic token selection module to divide the initial teacher token set into a number of sparse blocks according to a preset block size, and determining the token features corresponding to each token in each sparse block based on the preset global position encoding; wherein, the token features include visual semantic information and spatial position information. In each sparse block, determining the dynamic value corresponding to each token using the target norm distance between the token features at the same spatial position in adjacent frames. Descendingly sort the tokens in each of the sparse blocks using the dynamic value, and respectively select the first preset number of tokens from each of the sparse blocks to generate corresponding teacher training masks, and determine the teacher initial token set processed by the teacher training masks as the target token set; wherein, the target length is equal to the product of the first preset number and the number of sparse blocks.
5. The method according to claim 1, wherein Performing multi-scale cropping on the target token set according to the target self-attention weights of the teacher model to generate at least three student training masks with different token numbers, including: Extracting the target self-attention weights and determining the token importance scores corresponding to the tokens in the target token set based on the target self-attention weights; Descendingly sort the tokens in the target token set according to the token importance scores to generate a token importance ranking list; Select corresponding numbers of tokens in the token importance ranking list according to at least three different preset token cropping ratios to generate at least three corresponding student training masks; wherein, the preset token cropping ratios include at least 100%, 75% and 50%.
6. The method according to claim 1, characterized in that, Inputting the target training data and different student training masks into the student model for forward propagation to generate student training features, including: Inputting the target training data into the dual normalization module of the student model to generate a student initial token set; wherein, the dual normalization module includes: a first normalization layer, a patch embedding layer and a second normalization layer; Inputting a preset global position encoding into the depthwise separable convolution module of the student model to generate a target enhanced position encoding; Respectively determining the student initial token sets processed by different student training masks as target sub-token sets; For each of the target sub-token sets, inputting the target sub-token set and the target enhanced position encoding into the student Transformer encoder and the attention pooling layer of the student model to obtain the student intermediate layer features output by each second encoding block in the student Transformer encoder and the student final layer features output by the attention pooling layer, and taking all the student intermediate layer features and the student final layer features obtained under all the student training masks as the student training features; wherein, the student Transformer encoder includes a second number of the second encoding blocks, and the second encoding block includes an improved multi-head self-attention layer and a feed-forward neural network.
7. The method according to claim 1, characterized in that Aligning and distilling the student training features with the teacher training features to obtain a target student model, including: Determining the corresponding alignment intermediate layer indices of the teacher model according to the preset inter-layer alignment rules; Extracting the corresponding target teacher intermediate layer features from the teacher training features according to the alignment intermediate layer indices; For each student intermediate layer feature in the student training features, determining the intermediate layer alignment loss between the student intermediate layer feature and the corresponding target teacher intermediate layer feature; Determine the final layer alignment loss between the student final layer features in the student training features and the teacher final layer features in the teacher training features; Determine the weighted sum result of each of the intermediate layer alignment losses and the final layer alignment loss as the alignment distillation total loss; Based on the alignment distillation total loss, call a preset gradient descent method to update the parameters of the student model to obtain the target student model.
8. The method according to claim 6, wherein In the improved multi-head self-attention layer, the calculation formula for a single attention head is: where Z represents the output of a single attention head; Q represents the query matrix; K represents the key matrix; V represents the value matrix; T represents the transpose symbol; D represents the feature dimension; Softmax represents the normalization exponential function; LPE represents the linear projection function.
9. A model training device, characterized in that The device includes: A sampling module, configured to sample the original video frame sequence based on a target training spatio-temporal resolution to obtain target training data; the target training spatio-temporal resolution is higher than the spatio-temporal resolution during downstream inference of the student model; A teacher model training module, configured to input the target training data into the teacher model, generate a target token set through dynamic token selection, and generate teacher training features through forward propagation; A multi-scale student mask generation module, configured to perform multi-scale cropping on the target token set according to the target self-attention weights of the teacher model to generate at least three student training masks with different token numbers; the target self-attention weights are the weight matrices corresponding to the last multi-head self-attention layer of the teacher model; A student model training module, configured to input the target training data and different student training masks into the student model for forward propagation to generate student training features; An alignment distillation module, configured to perform alignment distillation on the student training features and the teacher training features to obtain a target student model.
10. An electronic device, characterized in that, The electronic device includes: At least one processor; and A memory communicatively connected to the at least one processor; wherein, The memory stores a computer program executable by the at least one processor, and the computer program is executed by the at least one processor so that the at least one processor can execute the model training method according to any one of claims 1-8.
11. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer instructions, and the computer instructions are used to cause the processor to implement the model training method according to any one of claims 1-8 when executed.
Citation Information
Patent Citations
Asymmetric mask distillation method for small mask autoencoder pre-training
CN116704053A
Generating Pretrained Sparse Student Model for Transfer Learning
US20230010142A1
Cited By
Heterogeneous model alignment method and system based on cross-modal knowledge distillation
CN120724393A
High-fidelity lightweight world model construction method for end-to-end automatic driving test
CN120909949A
High-fidelity lightweight world model construction method for end-to-end autonomous driving test
CN120909949B
Model distillation method and device, model reasoning method and reasoning model
CN122334403A