Dynamic parameter adjustment small sample behavior recognition method based on text feature guidance
By reconstructing the linear layer into an extensible basis matrix library and introducing text feature guidance and regularization strategies, the problem of weak model generalization ability in small-sample learning is solved, efficient small-sample behavior recognition is achieved, and the model's generalization ability and action recognition accuracy are improved.
Patent Information
- Application Number
- CN202510952399.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-10
- Publication Date
- 2025-10-17
AI Technical Summary
现有小样本学习方法中模型参数固定导致泛化能力弱,易过拟合且对新领域适应能力低下,难以在训练样本数量有限的情况下实现稳定而准确的行为识别。
将传统线性层重构为可扩展的基矩阵库,每个线性层解耦为多组基参数矩阵,通过文本特征引导生成适用于特定任务的线性层参数,并引入质心互斥损失和对比聚类损失以增强基参数矩阵间的高内聚和低耦合。
The model has achieved efficient generalization capabilities under small sample conditions, improved classification performance, and can more accurately identify complex behavioral actions, showing significant generalization and action understanding capabilities.
Smart Images

Figure CN120808234A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to computer vision technology, in particular to a dynamic parameter adjustment small sample behavior recognition method based on text feature guidance. BACKGROUND
[0002] In recent years, with the continuous expansion of application scenarios such as intelligent monitoring, human-computer interaction and automatic driving, behavior recognition as a key technology in video understanding has attracted widespread attention. However, this task still faces many challenges, especially under the condition of limited training sample quantity, the model performance is often difficult to guarantee. In order to reduce the dependence on large-scale labeled video data, researchers gradually turn their attention to small sample learning methods, and strive to achieve stable and accurate classification effect under the condition of sample scarcity.
[0003] At present, most small sample learning methods aim to train the model to learn parameters to realize the generalization of new categories, and the model parameters are usually fixed after training. However, due to the limitation of data, the model is often difficult to learn parameters with generalization ability, and is easy to fall into overfitting of specific induction bias of source field. This will lead to catastrophic forgetting or low adaptability to new fields. SUMMARY
[0004] The purpose of the present application is to solve the problems of weak generalization ability, easy overfitting and low adaptability to new fields caused by fixed model parameters in the prior art, and provide a dynamic parameter adjustment small sample behavior recognition method based on text feature guidance which can enhance the generalization ability.
[0005] In order to achieve the above application purpose, the present application provides the following technical scheme.
[0006] A dynamic parameter adjustment small sample behavior recognition method based on text feature guidance, comprising the following steps:
[0007] A. Given a video dataset, randomly extract each video Frame to form a new video frame sequence;
[0008] B. The sampled video frame sequence obtained in step A is input into a visual feature extractor; in the visual extractor, the traditional linear layer is reconstructed into an expandable base matrix library: each linear layer is decoupled into multiple groups of base parameter matrices, wherein each base parameter matrix is similar to the base vector of the linear layer, and together constitutes the base of the parameter space;
[0009] C. The label of the sampled video frame sequence obtained in step A is input into a text feature extractor after a text template and text features are obtained;
[0010] D. Input the text features obtained in step C into the coordinate vector calculation module, use the text information as semantic guidance, generate a combination coefficient of multiple base parameter matrices, and construct a linear layer parameter suitable for a specific task;
[0011] E. Calculate the centroid repulsion loss and contrast clustering loss of the base parameter matrices generated in step D to achieve the characteristics of high cohesion and low coupling between groups of base parameter matrices;
[0012] F. Input the video features obtained by the feature extractor into the predictor to obtain a prediction result, calculate a loss function, and update the gradient.
[0013] In step A, the random extraction of each video frame to form a new video frame sequence, the specific steps are as follows: divide each video into segments, randomly select a frame from each segment, and construct a new frame sequence ; the video frame sequence formed by these randomly selected frames, wherein represents a video image frame randomly extracted from the th video segment.
[0014] In step B, the sampled video frame sequence obtained in step A is input into the visual feature extractor, and a TFPA-Adapter module based on text feature guided dynamic parameter adjustment is designed. The adapter module reconstructs the traditional linear layer into an expandable base matrix library; each fully connected layer is decoupled into multiple groups of base parameter matrices, and each parameter matrix can be regarded as a base vector of the linear layer, which together constitutes the basis of the parameter space; the multiple groups of parameter matrix coordinate vectors obtained through text feature guidance can be combined to generate new linear layer parameters suitable for specific tasks; therefore, the TFPA-Adapter based on text feature guided dynamic parameter adjustment can share knowledge between different tasks and quickly adapt to the characteristics of different tasks;
[0015] For an original weight matrix of a linear layer , the TFPA-Adapter based on text feature guided dynamic parameter adjustment expands it into M base parameter matrices , wherein each ; these base parameter matrices together constitute a base matrix library; for an input task feature , the TFPA-Adapter based on text feature guided dynamic parameter adjustment generates a coordinate vector under this task through a coordinate vector calculation module, wherein represents the weight of the i-th base parameter matrix; based on the coordinate vector , the TFPA-Adapter based on text feature guided dynamic parameter adjustment can generate a task-specific weight matrix :
[0016]
[0017] The model can dynamically adjust the parameters according to the task characteristics without completely replacing the parameters of the old task; the entire computing process of the adapter based on text feature guided dynamic parameter adjustment can be represented as:
[0018]
[0019] wherein, and are linear weights for downsampling and upsampling respectively, represent the video feature input (N represents the number of tokens, L represents the number of frames contained by the video, D is the feature dimension, is the feature dimension after sampling); TFPA-Adapter is an adapter based on text feature guided dynamic parameter adjustment, which is a variant of the standard "downsampling-activation function-upsampling" adapter structure, and the core improvement is to replace the activation function with a task-specific weight matrix mapping.
[0020] In step D, the text feature input coordinate vector calculation module obtained in step C uses text information as a semantic guide to construct linear layer parameters suitable for a specific task by generating a combination coefficient of multiple base parameter matrices, and the specific steps can be:
[0021] The multi-modal base model ensures effective alignment of images and text in the same embedding space through a visual-text alignment loss function; at the same time, the coordinate vector calculation module dynamically generates a linear combination coefficient (i.e., coordinate vector ) of the base parameter matrix based on the text feature of each video as a semantic guide signal, thereby ensuring the geometric consistency of the mapping results of the base matrix library and the text feature space;
[0022] The specific implementation of the coordinate vector calculation module is as follows:
[0023] Given the input video feature and its corresponding text feature , the CVC module first maps the feature dimensions of the two to the same dimensional space to achieve cross-modal feature alignment and interaction:
[0024]
[0025] wherein is the downsampling linear weight in the adapter based on text feature guided dynamic parameter adjustment, and It is a linear weight used to compress text features to the same feature dimension; for the support set video text features, the coordinate vector calculation module uses the text of the corresponding category of the video to generate features through the text branch of the multimodal basic model; and the query set video text features use the average value of all category codes in the task; then, the coordinate vector calculation module extracts the global features of each frame of the video (i.e., the first Token feature of ViT) and the predefined M basis parameter matrices The linear transformation is as follows:
[0026]
[0027] In the coordinate vector calculation module, in order to maintain the text features Mapping space with basis matrix The geometric consistency of the least squares method is used to quantify the alignment error between the two, so that the combination result of the basis matrix is closer to the text semantic space; by minimizing the alignment error, the text features implicitly guide the linear combination of the basis matrix, so that the generated dynamic parameters Capable of capturing task-specific semantic patterns; the specific process is as follows:
[0028] Using singular value decomposition to solve text features The projection coefficients in the basis matrix mapping space (i.e. the coordinate vector ):
[0029]
[0030] The closed-form solution to the least squares problem obtained by singular value decomposition ensures that the gradient can be back-propagated to the basis matrix library , avoiding the computational overhead caused by iterative optimization and is suitable for large-scale training.
[0031] In step E, the basis parameter matrix generated in step D is used to calculate the centroid mutual exclusion loss and contrastive clustering loss. The adapter for dynamic parameter adjustment guided by text features handles complex and diverse explicit task scenarios by combining multiple sets of parameter matrices. The key to improving the model effect is to ensure the discrimination between different parameter matrices. To avoid excessive coupling between different parameter matrices and strengthen the cohesion within each parameter matrix, two regularization strategies are introduced: centroid mutual exclusion loss and contrastive clustering loss.
[0032] Centroid mutual exclusion loss: The main goal of centroid mutual exclusion loss is to make the centers of each group of parameter matrices as evenly distributed as possible, thereby minimizing the dependencies between them. By introducing this regularization term, the mutual interference between different parameter matrix groups can be effectively avoided, so that each parameter matrix can complete its task independently. In terms of implementation, each group of basis parameter matrices is first The centroid vector is obtained by averaging the L2 normalized column vectors of each parameter matrix group where represents the th basis parameter matrix, is the feature dimension after sampling:
[0033]
[0034] At this time, the centroid repulsion loss expects the centroid vectors to be uniformly distributed in the vector space, which is equivalent to the inner product between each centroid vector being as small as possible, i.e. At this time, only the optimal solution between the vectors is required:
[0035]
[0036] From the formula, we can get , that is, the optimal angle between each group of centroid vectors is ; and the centroid repulsion loss achieves repulsion by punishing the difference from the theoretical optimal angle:
[0037]
[0038] where represents the angle between the centroid vectors and
[0039] Contrastive clustering loss: The purpose of the contrastive clustering loss is to promote close cooperation within each parameter matrix group, thereby achieving high cohesion of the parameter matrices; specifically, this loss performs contrastive clustering on the vectors of each parameter matrix, so that the vectors within the same parameter matrix are as close as possible in the vector space; this strategy can enhance the internal functional consistency of each parameter matrix and ensure that the parameter matrix group can focus on performing a single task; in implementation, the positive sample is defined as all column vectors within the same parameter matrix, and the negative sample is defined as the centroid vectors of different parameter matrices; for any column vector , where represents the th column vector, the contrastive clustering loss calculates the angle difference between it and the positive and negative samples as follows:
[0040]
[0041] where is the adaptive margin threshold and the scaling factor); this formula encourages vectors within the same matrix to be closely clustered around their respective centers, while maintaining separation from the centers of other matrices.
[0042] Compared with the prior art, the present application has the following outstanding advantages:
[0043] The present application realizes high-efficiency and strong generalization ability small sample behavior recognition. Compared with the current mainstream small sample behavior recognition method, the classification performance of the present application is improved. The present application is inspired by vector space basis decomposition, and the traditional linear layer is reconstructed into an extensible basis matrix library: each linear layer is decoupled into multiple groups of basis parameter matrices, each of which is similar to the basis vector of the linear layer, and collectively constitutes the basis of the parameter space. The coordinate vector calculation module uses text information as semantic guidance, and constructs the linear layer parameters suitable for specific tasks by generating the combination coefficients of multiple parameter matrices. In addition, in order to enhance the discrimination between the basis parameter matrices, two regularization terms are designed: centroid repulsion loss and contrastive clustering loss to realize the characteristics of high cohesion and low coupling between the groups of basis parameter matrices. The experimental results on five small sample action recognition benchmark datasets show that the method has significant advantages and excellent generalization ability. It provides a new direction and inspiration for the research and application of this field. BRIEF DESCRIPTION OF DRAWINGS
[0044] Figure 1 The overall flowchart of the embodiments of the present application.
[0045] Figure 2 The experimental result heat map. DETAILED DESCRIPTION
[0046] The method of the present application will be described in detail below in combination with the drawings and embodiments. The embodiments are implemented on the premise of the technical solutions of the present application, and the implementation modes and specific operation processes are given, but the protection scope of the present application is not limited to the following embodiments.
[0047] The embodiment of the present application is a small sample behavior recognition method based on text feature guided dynamic parameter adjustment, which comprises the following steps:
[0048] A. For the small sample behavior recognition task, the data set processed by the present application is composed of a plurality of video samples, and each video is composed of a plurality of frames. In order to extract representative time sequence features, a sparse time sampling method is used. The specific steps are as follows: first, divide each video into segments, and then randomly select one frame from each segment to construct a new frame sequence . The final video frame sequence is composed of these randomly selected frames, wherein, represents the video image frame randomly extracted from the th video segment. In this embodiment, the number of video sequence frames is set to .
[0049] B. As Figure 1As shown, the sampled video frame sequence obtained in step A is input to the visual feature extractor. In the visual extractor, the traditional linear layer is reconstructed into an expandable base matrix library. A text feature guide-based dynamic parameter adjustment adapter module is designed, which reconstructs the traditional linear layer into an expandable base matrix library. Each fully connected layer is decoupled into multiple sets of base parameter matrices, and each parameter matrix can be regarded as a basis vector of the linear layer, which collectively constitutes the basis of the parameter space. Through the text feature guide to obtain the multiple sets of parameter matrix coordinate vectors, the linear layer new parameters that adapt to specific tasks can be combined and generated. Therefore, the text feature guide-based dynamic parameter adjustment adapter can share knowledge between different tasks and quickly adapt to the characteristics of different tasks. For the original weight matrix of a linear layer , the text feature guide-based dynamic parameter adjustment adapter expands it into M base parameter matrices , where each . These base parameter matrices collectively constitute a base matrix library. For the input task feature , the text feature guide-based dynamic parameter adjustment adapter generates the coordinate vector under the task , where represents the weight of the i-th base parameter matrix. Based on the coordinate vector , the text feature guide-based dynamic parameter adjustment adapter can generate the task-specific weight matrix by linearly combining the base matrices:
[0050]
[0051] The model can dynamically adjust the parameters according to the task characteristics without completely replacing the parameters of the old task. The entire calculation process of the text feature guide-based dynamic parameter adjustment adapter can be represented as:
[0052]
[0053] where and are the linear weights for downsampling and upsampling, respectively, represents the video feature input (N represents the number of tokens, L represents the number of frames contained in the video, D is the feature dimension, is the sampled feature dimension). TFPA-Adapter is a text feature guide-based dynamic parameter adjustment adapter, which is a variant of the standard "downsampling-activation function-upsampling" adapter structure, and the core improvement is to replace the activation function with a task-specific weight matrix mapping.
[0054] C. As Figure 1As shown, the label of the sampled video frame sequence obtained in step A is input to the text feature extractor through a text template and the text feature is obtained.
[0055] D. As Figure 1 shown, the text feature obtained in step C is input to the coordinate vector calculation module, which uses the text information as a semantic guide to construct linear layer parameters suitable for specific tasks by generating combination coefficients of multiple parameter matrices. With the help of the coordinate vector calculation module, these task features are used to calculate the coordinate vectors of multiple groups of parameter matrices in each adapter based on text feature guided dynamic parameter adjustment . Specifically, the multi-modal base model ensures that the image and the text are effectively aligned in the same embedding space through the visual-text alignment loss function. At the same time, the coordinate vector calculation module dynamically generates the linear combination coefficients (i.e., coordinate vectors ) of the base parameter matrix using the text feature of each video as a semantic guide signal, thereby ensuring the geometric consistency of the mapping results of the base matrix library and the text feature space.
[0056] The specific implementation of the coordinate vector calculation module is as follows. Given the input video feature and its corresponding text feature , the CVC module first maps the feature dimensions of the two to the same dimensional space to realize the alignment and interaction of cross-modal features:
[0057]
[0058] wherein, is the down-sampling linear weight in the adapter based on text feature guided dynamic parameter adjustment, and is the linear weight used to compress the text feature to the same feature dimension. For the support set video text feature, the coordinate vector calculation module uses the text corresponding to the category of the video to generate the feature through the text branch of the multi-modal base model; while the query set video text feature adopts the average value of the encoding of all categories in the task. Subsequently, the coordinate vector calculation module extracts the global feature of each frame of the video (i.e., the first Token feature of ViT), and performs linear transformation with the predefined M base parameter matrices as follows:
[0059]
[0060] In linear algebra, the least squares method is used to minimize the sum of the squares of the differences between the observed data and the predicted values of the model to find the best fitting function of the data. In the coordinate vector calculation module, in order to maintain the text feature and the base matrix mapping space geometric consistency, the alignment error between them can be quantified by the optimization objective of least square method, so that the combination result of base matrix is closer to the text semantic space. By minimizing the alignment error, the text feature implicitly guides the linear combination way of base matrix, so that the generated dynamic parameters can capture task-specific semantic patterns. The specific process is as follows, the singular value decomposition is used to solve the text feature In the projection coefficient (i.e. coordinate vector ) of the base matrix mapping space:
[0061]
[0062] The closed-form solution obtained by singular value decomposition to solve the least square problem not only ensures that the gradient can be back-propagated to the base matrix library , but also avoids the computational overhead brought by iterative optimization, which is suitable for large-scale training.
[0063] E. Calculate the centroid repulsion loss and contrast clustering loss of the matrix parameters generated in step D to realize the characteristics of high cohesion and low coupling between groups of base parameter matrices. The adapter of dynamic parameter adjustment guided by text features processes complex and diverse explicit task scenarios through the combination of multiple groups of parameter matrices, and the key to improving the model effect lies in ensuring the discrimination between different parameter matrices. The key to achieving this goal is to avoid excessive coupling between different parameter matrices while strengthening the cohesion within each parameter matrix. For this purpose, two regularization strategies are introduced: centroid repulsion loss and contrast clustering loss.
[0064] Centroid repulsion loss: The main goal of centroid repulsion loss is to make the centers of each group of parameter matrices as evenly distributed as possible, thereby minimizing their dependence on each other. By introducing this regularization term, the mutual interference between different parameter matrix groups can be effectively avoided, so that each parameter matrix can independently complete its task. In implementation, first, the column vectors of each group of base parameter matrices are L2 normalized and averaged to obtain the centroid vector , where represents the th base parameter matrix, and is the feature dimension after sampling:
[0065]
[0066] At this time, the centroid repulsion loss expects the centroid vector to be evenly distributed in the vector space, which is equivalent to the inner product between each centroid vector being as small as possible, i.e. , at this time only the optimal solution between the vector angle is required:
[0067]
[0068] From the formula, we can get , that is, the optimal angle between each group of centroid vectors is And the centroid repulsion loss achieves repulsion by penalizing the difference from the theoretical optimal angle:
[0069]
[0070] where, denotes the angle between the centroid vectors and .
[0071] Contrastive clustering loss: The purpose of the contrastive clustering loss is to promote close cooperation within each group of parameter matrices, thereby achieving high cohesion of parameter matrices. Specifically, this loss performs contrastive clustering on the vectors of each parameter matrix, so that the vectors within the same parameter matrix are as close as possible in the vector space. This strategy can enhance the internal functional consistency of each parameter matrix and ensure that the group of parameter matrices can focus on performing a single task. In implementation, the positive samples are all column vectors within the same parameter matrix, and the negative samples are the centroid vectors of different parameter matrices. For any column vector , where denotes the column vector, the contrastive clustering loss calculates the angle difference between it and the positive and negative samples as follows:
[0072]
[0073] where, is the adaptive margin threshold ( is the scaling factor). This formula encourages vectors within the same matrix to be closely clustered around their respective centers, while keeping separate from the centers of other matrices.
[0074] F. Input the video features obtained by the feature extractor into the predictor to obtain the prediction result, calculate the loss function and update the gradient. Compared with the current mainstream small sample behavior recognition method, the classification performance of the present application is improved.
[0075] Figure 2 Baseline model AIM and the attention visualization results obtained by the text feature guided dynamic parameter adjustment (TFPA) proposed in this paper are shown to compare the differences in action area focusing ability between the two. Specifically, Figure 2 (a) in (a) shows the visualization effect of the action "riding a horse", Figure 2(b) in FIG. 6 corresponds to the visualization of the action "Ice Dance". In each group of images, the first is the sampled frame of the original video, providing the visual context; the second is the visualization of the activation region generated by the AIM method; and the third is the visualization output of the dynamic parameter adjustment method guided by text features. From Figure 2 As can be seen from (a) in FIG. 6, although the AIM can preliminarily identify the character target in the video, its attention region is scattered and difficult to accurately focus on the key action region, so there is a deviation in understanding the complex interactive action "Riding Horse". In contrast, the dynamic parameter adjustment method guided by text features not only successfully identifies the key entity "horse", but also effectively captures the interaction between the character and the horse, showing stronger action understanding ability. In Figure 2 In (b) in FIG. 6, for the detailed dynamic action "Ice Dance", the attention of the AIM is focused on the outline of the character, ignoring the time sequence features of the action itself; while the dynamic parameter adjustment method guided by text features can more accurately focus on the footstep movement trajectory of the character, showing its high sensitivity to key action details. These visualization results show that the dynamic parameter adjustment method guided by text features can more accurately focus on the key action region, which benefits from the dynamic parameter generation mechanism and regularization constraints. At the same time, the dynamic parameter adjustment method guided by text features maintains the consistency of cross-frame attention. Compared with the AIM, the dynamic parameter adjustment method guided by text features performs significantly better in capturing time sequence actions and subtle details, which provides an intuitive explanation for its performance improvement in complex action recognition tasks.
[0076] The present application proposes a dynamic parameter adjustment method guided by text features for small sample behavior recognition, to realize efficient and strong generalization capability small sample behavior recognition. Inspired by vector space basis decomposition, the method reconstructs the traditional linear layer into an expandable basis matrix library: each linear layer is decoupled into multiple groups of basis parameter matrices, where each basis parameter matrix is similar to the basis vector of the linear layer, which together constitutes the basis of the parameter space. The coordinate vector calculation module uses text information as semantic guidance to construct linear layer parameters suitable for specific tasks by generating combination coefficients of multiple parameter matrices. In addition, to enhance the discrimination between basis parameter matrices, two regularization terms are designed: centroid repulsion loss and contrastive clustering loss to realize the characteristics of high cohesion and low coupling between groups of basis parameter matrices. Table 1 gives the recognition accuracy results of the present application and existing methods on the small sample behavior recognition datasets UCF101, HMDB51, Kinetics, SSv2-Small and SSv2-Full under 1-shot and 5-shot, which has a relatively obvious improvement, proving that the present application can fully understand semantic information and realize more accurate small sample behavior recognition.
[0077] Table 1
[0078]
[0079] The 1-shot and 5-shot recognition accuracy results of the present application (TFPA) and the prior art on five data sets of small sample behavior recognition data sets UCF101, HMDB51, Kinetics, SSv2-Small and SSv2-Full have obvious improvement, specifically, compared with the method MoLo, on the HMDB51 data set, the present application is improved by 17.7% on the 1shot index and 11.5% on the 5-shot index. On the Kinetics data set, it is improved by 15.9% on the 1shot index and 10.1% on the 5-shot index. The above results verify that the present application can fully understand semantic information and realize more accurate small sample behavior recognition.
[0080] The above embodiments are only preferred embodiments of the present application and cannot be considered as limiting the scope of the present application. Any equivalent changes and improvements made within the scope of the present application should still belong to the patent scope of the present application.
Claims
1. A small sample behavior recognition method based on dynamic parameter adjustment guided by text features, characterized by The following steps are involved: A. Given a video dataset, randomly extract each video The frames constitute a new video frame sequence; B. Input the sampled video frame sequence obtained in step A into the visual feature extractor. In the visual feature extractor, the traditional linear layer is reconstructed into a scalable basis matrix library: each linear layer is decoupled into multiple sets of basis parameter matrices, where each basis parameter matrix is similar to the basis vector of the linear layer and together constitutes the basis of the parameter space. C. Input the labels of the sampled video frame sequence obtained in step A into the text feature extractor through the text template to obtain text features; D. Input the text features obtained in step C into the coordinate vector calculation module. Using the text information as a semantic guide, the linear layer parameters suitable for the specific task are constructed by generating the combined coefficients of multiple basis parameter matrices. E. Calculate the centroid mutual exclusion loss and contrast clustering loss for the basis parameter matrix generated in step D to achieve high cohesion and low coupling between basis parameter matrix groups; F. Input the video features obtained by the feature extractor into the predictor to obtain the prediction results, calculate the loss function and update the gradient.
2. A small sample behavior recognition method based on dynamic parameter adjustment guided by text features as described in claim 1, characterized in that In step A, each video is randomly selected Frames constitute a new video frame sequence. The specific steps are as follows: Divide each video into segments, randomly select a frame from each segment to construct a new video frame sequence ; The video frame sequence formed is composed of these randomly selected frames, among which, Indicates that from A video image frame is randomly selected from a video clip.
3. A small sample behavior recognition method based on dynamic parameter adjustment guided by text features as described in claim 1, characterized in that In step A, .
4. A small sample behavior recognition method based on dynamic parameter adjustment guided by text features as described in claim 1, characterized in that In step B, the sampled video frame sequence obtained in step A is input into a visual feature extractor, and an adapter module for dynamic parameter adjustment guided by text features is designed. The adapter module reconstructs the traditional linear layer into an extensible basis matrix library; each fully connected layer is decoupled into multiple groups of basis parameter matrices, and each parameter matrix can be regarded as the basis vector of the linear layer, which together constitute the basis of the parameter space; the multiple groups of parameter matrix coordinate vectors obtained through text feature guidance can be combined to generate new linear layer parameters adapted to a specific task; therefore, the adapter for dynamic parameter adjustment guided by text features can share knowledge between different tasks and quickly adapt to the characteristics of different tasks; The original weight matrix for a linear layer is , the adapter based on dynamic parameter adjustment guided by text features expands it into M basis parameter matrices , where each ; These basis parameter matrices together constitute a basis matrix library; for the input task features The adapter with dynamic parameter adjustment guided by text features generates the coordinate vector for this task through the coordinate vector calculation module. ,in represents the weight of the i-th basis parameter matrix; Based on any coordinate vector The adapter with dynamic parameter adjustment guided by text features generates task-specific weight matrices by linearly combining the basis matrices : The model dynamically adjusts parameters according to the task characteristics without completely replacing the parameters of the old task; the calculation process of the entire adapter with dynamic parameter adjustment guided by text features is expressed as: in, and are the linear weights for downsampling and upsampling, respectively, Represents the video feature input, N represents the number of tokens, L represents the number of frames contained in the video, and D is the feature dimension. is the feature dimension after sampling.
5. A small sample behavior recognition method based on dynamic parameter adjustment guided by text features as described in claim 1, characterized in that In step D, the text features obtained in step C are input into the coordinate vector calculation module, and the text information is used as a semantic guide to construct linear layer parameters suitable for a specific task by generating the combination coefficients of multiple basis parameter matrices. The specific steps are as follows: The multimodal base model ensures that images and texts are effectively aligned in the same embedding space through the visual-text alignment loss function. At the same time, the coordinate vector calculation module uses the text features of each video as semantic guidance signals to dynamically generate the linear combination coefficients of the basis parameter matrix, i.e., the coordinate vector , thereby ensuring the geometric consistency between the mapping results of the base matrix library and the text feature space; The specific implementation of the coordinate vector calculation module is as follows: Given the input video features and its corresponding text features , the CVC module first maps the feature dimensions of the two to the same dimensional space to achieve alignment and interaction of cross-modal features: in The downsampled linear weights in the adapter are adjusted based on dynamic parameters guided by text features, while It is a linear weight used to compress text features to the same feature dimension; for the support set video text features, the coordinate vector calculation module uses the text of the corresponding category of the video to generate features through the text branch of the multimodal basic model; and the query set video text features use the average value of all category codes in the task; then, the coordinate vector calculation module extracts the global features of each frame of the video , which is the first Token feature of ViT, and the predefined M basis parameter matrix The linear transformation is as follows: In the coordinate vector calculation module, in order to maintain the text features Mapping space with basis matrix The geometric consistency of the least squares method is used to quantify the alignment error between the two, so that the combination result of the basis matrix is closer to the text semantic space; by minimizing the alignment error, the text features implicitly guide the linear combination of the basis matrix, so that the generated dynamic parameters Capable of capturing task-specific semantic patterns; the specific process is as follows: Using singular value decomposition to solve text features The projection coefficient in the basis matrix mapping space, that is, the coordinate vector : The closed-form solution to the least squares problem obtained by singular value decomposition ensures that the gradient can be back-propagated to the basis matrix library , avoiding the computational overhead caused by iterative optimization and is suitable for large-scale training.
6. A small sample behavior recognition method based on dynamic parameter adjustment guided by text features as claimed in claim 1, characterized in that In step E, the centroid mutual exclusion loss is: for each set of basis parameter matrices The column vector is L2 normalized and the mean is obtained to obtain the centroid vector ,in Indicates the basis parameter matrices, is the feature dimension after sampling: At this time, the centroid mutual exclusion loss expects the centroid vector Uniform distribution in the vector space is equivalent to the inner product between the centroid vectors being as small as possible, that is, , at this time only the optimal solution angle between the vectors is required: From this formula we get , that is, the optimal angle between the centroid vectors of each group is ; The centroid mutual exclusion loss achieves mutual exclusion by penalizing the difference from the theoretical optimal angle: in, Represents the centroid vector and Angle.
7. A small sample behavior recognition method based on dynamic parameter adjustment guided by text features as described in claim 1, characterized in that In step E, the contrast clustering loss specifies that the positive samples are all column vectors in the same parameter matrix, and the negative samples are the centroid vectors of different parameter matrices; for any column vector ,in express No. Column vector, the difference in angles between the positive and negative samples is calculated by comparing the clustering loss: in is the adaptive interval threshold, is the scaling factor; This formula forces vectors within the same matrix to cluster tightly around their respective centers while remaining separated from the centers of other matrices.