Video-oriented self-evolution target recognition method

CN117911789BActive Publication Date: 2026-09-25NORTHWESTERN POLYTECHNICAL UNIV +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202410218220.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-02-28
Publication Date
2026-09-25
Estimated Expiration
2044-02-28

AI Technical Summary

Technical Problem

目前,传统的基于内存管理策略的自进化学习技术,无法满足视频监控系统对于实时性,异或是由于过强的正则化减轻分类模型遗忘的同时但又导致分类模型很难学习新任务,这些方法所存在的问题导致其在工程应用上不切实际

Benefits of technology

[0043]本发明通过模型正本在线响应分类需求和模型副本离线自进化学习来满足监控视频系统的实时性和准确性,仅需额外一个临时的内存缓冲区和一定量标注成本。并且,在自进化学习上,本发明进一步提出了更加高效的学习策略,通过简单的logit屏蔽(logitmask)和批次注意力机制,极大地缓解了分类模型在自进化学习的时候,存在的类别不平衡问题导致模型旧知识的遗忘,并且交叉融合的批次注意力机制提高了分类模型快速整合新知识的能力。此外,通过更低存储成本的教师保留来进行特征蒸馏的正则化,进一步通过正则化手段缓解了分类模型对旧类特征表示的急剧变化,从而使得分类模型的知识遗忘的更慢。总之,通过本发明的自进化学习技术,在满足系统实时响应的同时,对比传统自进化学习技术,本发明自进化学习的模型具备更强大的分类性能,同时该自进化学习的方法具备更高效的学习效率。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN117911789B_ABST
    Figure CN117911789B_ABST
Patent Text Reader

Abstract

The application discloses a video-oriented self-evolution target recognition method, which firstly utilizes a classification model to respond to services in real time, then utilizes a labeled temporary buffer to perform self-evolution learning on a copy of the classification model, and finally updates the copy of the classification model after self-evolution learning back to the original classification model. The self-evolution learning model has stronger classification performance, and the self-evolution learning method has higher learning efficiency.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of pattern recognition technology, specifically relating to a self-evolving target recognition method for video. Background Technology

[0002] With the rapid development of video surveillance technology, the demand for video target recognition is growing. Video surveillance systems are widely used in urban security, intelligent transportation, industrial production, and other fields to capture and analyze (moving) targets in video in real time, providing critical information for decision-making and incident response.

[0003] Traditional video target recognition methods primarily rely on offline training using static datasets, lacking effective mechanisms to handle newly emerging targets during system operation. However, in practical applications, changes in the monitoring scene, the emergence of new target categories, and fluctuations in environmental conditions can lead to a decline in the performance of existing classification models, as these changes often exceed the scope of the model's prior training.

[0004] To address this issue, self-evolving learning techniques have gradually attracted attention. This technique dynamically adjusts model parameters by continuously learning from new data, enabling the model to adapt to constantly changing scenarios. Self-evolving target recognition technology for video has emerged in this context, aiming to improve the adaptability of video surveillance systems to new (moving) targets, mitigate the forgetting problem of classification models, and maintain accurate classification of known targets.

[0005] Against this backdrop, developing a self-evolving target recognition technology for video has become crucial. Currently, traditional self-evolving learning techniques based on memory management strategies cannot meet the real-time requirements of video surveillance systems. Alternatively, excessive regularization may reduce the forgetting of classification models but also make it difficult for them to learn new tasks. These problems make these methods impractical for engineering applications. Summary of the Invention

[0006] To overcome the shortcomings of existing technologies, this invention provides a self-evolving target recognition method for video. First, a classification model is used in real-time response service. Then, a labeled temporary buffer is used to perform self-evolving learning on a copy of the classification model. Finally, the self-evolved learning copy of the classification model is used to update the original classification model. The self-evolving learning model of this invention possesses stronger classification performance, and the self-evolving learning method also has higher learning efficiency.

[0007] The technical solution adopted by this invention to solve its technical problem is as follows:

[0008] Step 1: Convert an input video into video frames, and use the ViBe detection algorithm to extract detection boxes of moving targets; if there is a moving target, crop the detection box from the video frame and scale it to a pre-classified image to be processed with a size of 224*224;

[0009] Step 2: Send the preprocessed image to be classified into a classification model, then predict the final category through forward propagation of the classification model, add the predicted category label to the detection box of the moving target and return the result; since there are unseen categories or same-label samples from different domains in the self-evolution technology, prediction errors may occur in the model. Therefore, the pre-classified image to be processed obtained in Step 1 is further stored in a temporary buffer without annotation for subsequent self-evolution training based on manual annotation.

[0010] Step 3: When the cumulative number of images in the unannotated temporary buffer exceeds N, perform manual annotation on the images in the buffer and divide them into M batches, where M<N and M are new class batches; then first copy a copy of the online prediction classification model to the local, and then update and optimize the classification model copy batch by batch using the manually annotated samples;

[0011] Step 4: Let each new class batch in Step 3 be B i , retrieve memory batch samples from the memory buffer then perform data augmentation on batch B i and to obtain B I and B cat respectively; the memory buffer refers to a buffer that stores annotated images and caches part of historical images; the data augmentation includes random cropping and random flipping;

[0012] Step 5: Use the feature extraction network f in the classification model copy θ to extract the original feature E of batch B I I and the original feature E of batch B cat cat respectively, then send the feature and the feature into a Transformer Encoder Layer respectively for nonlinear cross-fusion based on batch attention mechanism to obtain virtual features and

[0013] Step 6: and are concatenated to obtain E I , and are concatenated to obtain E​​​cat A shared classifier is applied to both the original and virtual features, meaning they pass through the same classification layer; then feature E is calculated separately. I and E cat Cross-entropy loss based on logit mask and

[0014] Step 7: Filter out memory batch samples The old class samples—that is, samples that do not belong to the current task label set—are used as the new batch B. old and use feature extraction network f θ Extract the corresponding feature E old ;

[0015] Step 8: Using E old Perform feature distillation on the class prototype with the cumulative update of the corresponding label, and obtain the corresponding loss. If there is no corresponding class prototype for cumulative updates, then set And wait for step 11 to update before performing characteristic distillation;

[0016] Step 9: Weight the losses from the above steps to obtain the final loss. And then perform stochastic gradient descent to update the model;

[0017] Step 10: Finally, use the reservoir update algorithm to update the new batch of samples B. i Update to memory buffer middle;

[0018] Step 11: If the current new class batch is the last batch, update the class prototype of each corresponding category using all samples in the memory buffer; otherwise, skip this step.

[0019] Step 12: Keep the classification model from Step 2 predicting the category of the detected moving target in real time; then use the labeled batch data of moving targets to repeat the training and update of Step 2 to Step 10 on the copy of the classification model until all labeled batches are processed, and then update the copy of the model to the classification model in Step 2; through this self-evolutionary learning technique that alleviates class imbalance, the previously learned knowledge of the classification model is greatly preserved, and it is also ensured that it can learn to distinguish new categories.

[0020] Furthermore, step 5 specifically includes:

[0021] Step 5-1: Let feature extraction network f be used. θ The extracted batch features are in d represents the feature dimension;

[0022] Step 5-2: Perform a dimensionality expansion operation on the currently extracted feature E, so that the attention mechanism of the Transformer EncoderLayer can be applied to the batch dimension of the current feature; the operation is formally expressed as:

[0023]

[0024] Step 5-3: Expand the dimensions of the features from Step 5-2 The input to the single-layer Transformer EncoderLayer performs a batch-level attention mechanism, enabling cross-fusion of features within the batch; its operation can be formally expressed as:

[0025]

[0026] Step 5-4: For The dimensionality compression operation can be formally represented as follows:

[0027]

[0028] Furthermore, in step 6, the formulaic expression for the cross-entropy loss based on the logit mask is as follows:

[0029]

[0030] Wherein, the label corresponding to feature E is C cur W is the unique set of labels Y. c The weights of class c in the fully connected layer. For category y in the fully connected layer i Weights, bias c For the bias term of category c in the fully connected layer, For category y in the fully connected layer i The bias term; b is the size of the current batch; e i Let i represent the features of the i-th sample.

[0031] Furthermore, the characteristic distillation in step 8 is formally described as follows:

[0032]

[0033] Among them, proto c It is the prototype of the cumulative update class of category c. It is category y i The cumulative update class prototype performs cosine normalization on each feature before calculating cosine similarity; Y old Indicates feature E oldThe label corresponding to each feature of the old class sample.

[0034] Furthermore, the specific steps for updating the corresponding class prototype in step 11 are as follows:

[0035] Step 11-1: First calculate the memory buffer. The current class mean for each category is expressed as follows:

[0036]

[0037] Here, 1{·} is an indicator function that has a value of 1 if and only if the condition within the parentheses is met; in the memory buffer The number of samples labeled as category c is m. c 1; μ c This means that m c The category m calculated from each sample c Mean characteristics; (x i ,y i This represents the image and label of each sample in the memory buffer;

[0038] Step 11-2: Accumulate the average memory usage of each class into the class prototype, which can be expressed as:

[0039]

[0040] Among them, proto c It is the cumulative update class prototype of category c, n c As of now, proto c How many sample features were used to calculate it;

[0041] Step 11-3: Update the counting array n so that the count value n for category c in array n is updated. c Add m c .

[0042] The beneficial effects of this invention are as follows:

[0043] This invention satisfies the real-time performance and accuracy requirements of surveillance video systems by using an online response model to classify needs and an offline self-evolutionary learning model copy, requiring only an additional temporary memory buffer and a certain amount of annotation cost. Furthermore, regarding self-evolutionary learning, this invention proposes a more efficient learning strategy. Through simple logit masking and batch attention mechanisms, it significantly alleviates the problem of class imbalance leading to the forgetting of old knowledge in classification models during self-evolutionary learning. The cross-fusion batch attention mechanism also improves the classification model's ability to quickly integrate new knowledge. In addition, regularization for feature distillation using teacher retention, which has lower storage costs, further mitigates the drastic changes in the classification model's representation of old class features, thus slowing down the forgetting of knowledge. In summary, this invention's self-evolutionary learning technology, while meeting the system's real-time response requirements, demonstrates stronger classification performance compared to traditional self-evolutionary learning techniques, and the method itself offers higher learning efficiency. Attached Figure Description

[0044] Figure 1 This is a schematic diagram of the method flow of the present invention.

[0045] Figure 2 This is a diagram showing the nodes involved in gradient update in the Logit mask of this invention. Detailed Implementation

[0046] The present invention will be further described below with reference to the accompanying drawings and embodiments.

[0047] This invention proposes a self-evolving target recognition technology for video. To meet the real-time and accuracy requirements of monitoring systems, this invention first utilizes a classification model for real-time response services, then uses a labeled temporary buffer to perform self-evolving learning on a copy of the classification model, and finally updates the original classification model with the self-evolved learning copy. The specific steps of this invention are as follows:

[0048] (1) Convert the input video into video frames and use the pixel-level ViBe detection algorithm to extract the detection boxes of moving targets. If a moving target exists, crop the image block according to the detection box and scale it to a 224*224 image size;

[0049] (2) inputting the preprocessed image to be classified in step (1) into a classification model, then predicting the final category through forward propagation of the model, and attaching a predicted category label to the detection box of the moving target to return the result; since there are unseen categories or same-label samples from different domains in the self-evolution technology, the model may have prediction errors, therefore, further storing the preprocessed image to be classified in step (1) into a temporary buffer area with no labels for subsequent self-evolution training of manual labeling.

[0050] (3) after the number of images in the unlabeled buffer area accumulates to more than N images, performing manual labeling on the images in the buffer area, and dividing the images into M batches (new class batches, M<N); then first copying a copy of the online prediction classification model to the local, then updating and optimizing the copied classification model batch by batch using the manually labeled samples;

[0051] (4) setting each new class batch in step (3) as B i , and retrieving memory batch samples from a memory buffer (which stores labeled images and caches partial historical images) then performing data enhancement (including random cropping and random flipping) on batch B i and to obtain B I and B cat respectively;

[0052] (5) extracting features E of batch B θ and features E of batch B I by using the feature extraction network f I in the copied classification model cat respectively cat , then feeding the features and features into a Transformer Encoder Layer respectively for non-linear cross fusion on batch attention mechanism to obtain virtual features and

[0053] (6) and are spliced to obtain E I , and are spliced to obtain E cat , executing a shared classifier (i.e., passing through the same classification layer) on the original features and the virtual features, so as to learn more robust features, then respectively calculating the cross entropy loss of features E I and E cat in step (5) under the logit mask and

[0054] (7) Filter out memory batch samples The old class samples (i.e., samples that do not belong to the current task label set) are used as the new batch B. old and use feature extraction network f θ Extract the corresponding features E respectively old .

[0055] (8) Utilizing the current E old Perform feature distillation on the class prototype with the cumulative update of the corresponding label, and obtain the corresponding loss. If there is no class prototype that corresponds to the cumulative update of all tags, then set Then wait for step 11 to update before performing characteristic distillation.

[0056] (9) Weight the losses from the above steps to obtain the final loss. (Right now ), and then update the model using stochastic gradient descent;

[0057] (10) Finally, the reservoir update algorithm is used to update the new batch of samples B. i Update to memory buffer middle;

[0058] (11) If the current new class batch is the last batch, then update the corresponding class prototype using all samples in the memory buffer. Otherwise, skip this step.

[0059] The classification model in step (2) is maintained to predict the category of the detected moving target in real time. Then, using the labeled batch data of moving targets, the training and update of steps (2) to (10) are repeated on the copy of the classification model until all labeled batches are processed. Then, the copy is updated to the classification model in step (2). Through this self-evolutionary learning technique that alleviates class imbalance, the previously learned knowledge of the classification model is largely preserved, and it is also ensured that it can learn to distinguish new categories.

[0060] Furthermore, considering that the operations performed in different batches are consistent, let's assume a feature extraction network f is used. θ The extracted batch features are in d represents the feature dimension. Therefore, the specific steps of step (5) include:

[0061] 5.1 Perform a dimensionality expansion operation (DimUnsqueeze) on the currently extracted feature E, allowing the attention mechanism of the TransformerEncoder Layer to be applied to the batch dimension of the current feature. The formula for this operation is as follows:

[0062]

[0063] 5.2 Expand the dimensions of the features from step 5.1 The input to the single-layer Transformer Encoder Layer performs a batch-level attention mechanism, enabling feature cross-fusion within the batch. This operation can be formally represented as:

[0064]

[0065] 5.3 Regarding the features in step 5.2 Perform a dimensionality compression operation (DimSqueeze). Its formulaic expression is as follows:

[0066]

[0067] Furthermore, according to the symbol definition in step (5), the formulaic expression of the logit mask-based cross-entropy loss in step (6) is as follows:

[0068]

[0069] Wherein, the label corresponding to feature E is C cur W is the unique set of labels Y. c The weights of class c in the fully connected layer. For category y in the fully connected layer i Weights, bias c For the bias term of category c in the fully connected layer, For category y in the fully connected layer i The bias term; b is the size of the current batch; e i Let i represent the features of the i-th sample.

[0070] Furthermore, the characteristic distillation of step (8) is formally expressed as follows:

[0071]

[0072] Among them, proto c It is the prototype of the cumulative update class of category c. It is category y i The cumulative update class prototype performs cosine normalization on each feature before calculating cosine similarity; Y old Indicates feature E old The label corresponding to each feature of the old class sample.

[0073] Furthermore, when the current new class batch is the last batch, the specific steps for cumulatively updating the class prototype in step (11) are as follows:

[0074] 11.1 First, calculate the memory buffer. The current class mean for each category is expressed as follows:

[0075]

[0076] Here, 1{·} is an indicator function, which has a value of 1 if and only if the condition within the parentheses is satisfied; the number of samples labeled as category c in the memory buffer M is m. c 1; μ c This means that m c The category m calculated from each sample c Mean characteristics; (x i ,y i This represents the image and label of each sample in the memory buffer;

[0077] 11.2 The average memory usage of each class is accumulated into the class prototype, which can be expressed as follows:

[0078]

[0079] Among them, proto c It is the cumulative update class prototype of category c, n c As of now, proto c It is calculated from the features of a certain number of samples.

[0080] 11.3 Update the counting array n, which can be formally expressed as:

[0081] n c =n c +m c

[0082] Example:

[0083] Step 1: Convert the input video into video frames and use the pixel-level ViBe detection algorithm to extract detection boxes for moving targets. If a moving target is found, crop the image block based on the detection box and scale it to a 224*224 image size.

[0084] Step 2: Feed the preprocessed images to be classified from Step 1 into the classification model, then predict the final category through the forward propagation of the model, and label the detection boxes of moving targets with the predicted category and return the results; Since there are unseen categories or samples with the same label in different domains in the self-evolution technique, the model may make prediction errors. Therefore, the preprocessed images to be classified from Step 1 are then transferred to a temporary unlabeled buffer for subsequent manual labeling self-evolution training.

[0085] Step 3: When the cumulative number of images in the buffer area without unlabeled tags exceeds N images, perform corresponding manual annotation and divide them into M batches (new class batches, M < N). Then, first copy a copy of the online prediction classification model to the local, and then update and optimize the classification model copy batch by batch using the manually labeled samples;

[0086] Step 4: Let each new class batch in Step 3 be B i , and retrieve memory batch samples from the memory buffer (labeled images with some historical images cached); Then perform data augmentation operations (including random cropping and random flipping) on batch B i and to obtain B I and B cat ;

[0087] Step 5: Use the feature extraction network f in the classification model copy θ to respectively extract the features E of batch B I and the features E of batch B I and batch B cat of features E cat , then send the features and the features into the Transformer Encoder Layer respectively for nonlinear cross-fusion on batch attention mechanism to obtain virtual features and

[0088] Since the features of batch B I and the features of batch B cat undergo the same operation, for convenience, let the batch features extracted by the feature extraction network f θ be wherein d is the feature dimension.

[0089] 5.1 Perform a dimension expansion operation (DimUnsqueeze) on the currently extracted feature E, so that the attention mechanism of the Transformer Encoder Layer can be applied to the batch dimension of the current feature. The operation is formulated as:

[0090]

[0091] 5.2 Apply the features after dimension expansion in step 5.1 The input to the single-layer Transformer Encoder Layer performs a batch-level attention mechanism, enabling feature cross-fusion within the batch. This operation can be formally represented as:

[0092]

[0093] 5.3 Regarding the features in step 5.2 Perform a dimensionality compression operation (DimSqueeze). Its formulaic expression is as follows:

[0094]

[0095] Step 6: and E is obtained by splicing. I , and E is obtained by splicing. cat A shared classifier (i.e., passing through the same classification layer) is applied to both the original and virtual features to learn more robust features. Then, the features E from step 5 are calculated separately. I and E cat Cross-entropy loss based on logit mask and

[0096] According to the symbol definition in step 5, the formulaic expression of the cross-entropy loss based on the logit mask is as follows:

[0097]

[0098]

[0099] Among them, feature E I and E cat The corresponding tags are respectively and and Label Y I and Y cat The unique set, W c The bias is the weight of class c in the fully connected layer. c This is the bias term for category c in the fully connected layer.

[0100] Step 7: Filter out memory batch samples The old class samples (i.e., samples that do not belong to the current task label set) are used as the new batch B. old and use feature extraction network f θ Extract the corresponding features E respectively old .

[0101] Step 8: Utilize the current E old Perform feature distillation on the cumulative update class prototype of the corresponding label and obtain the corresponding loss. If there is no class prototype that corresponds to the cumulative update of all tags, then set Then wait for step 11 to update before performing characteristic distillation.

[0102]

[0103] Among them, proto c It is the prototype of the cumulative update class of category c. It is category y i The cumulative update class prototype performs cosine normalization on each feature before calculating cosine similarity; Y old Indicates feature E old The label corresponding to each feature of the old class sample.

[0104] Step 9: Weight the losses from the above steps to obtain the final loss. (Right now ), and then update the model using stochastic gradient descent;

[0105] Step 10: Finally, use the reservoir update algorithm to update the new batch of samples B. i Update to memory buffer middle;

[0106] Step 11: If the current batch of new classes is the last batch, update the corresponding class prototype using all samples in the memory buffer. Otherwise, skip this step.

[0107] Furthermore, when the current new batch is the last batch:

[0108] 11.1 First, calculate the memory buffer. Current class mean for each category:

[0109]

[0110] Here, 1{·} is an indicator function that has a value of 1 if and only if the condition within the parentheses is met; in the memory buffer The number of samples labeled as category c is m. c 1; μ c This means that m c The category m calculated from each sample c Mean characteristics; (x i ,y i This represents the image and label of each sample in the memory buffer;

[0111] 11.2 Accumulate the average memory usage of each class into the class prototype:

[0112]

[0113] Among them, proto c It is the cumulative update class prototype of category c, n c As of now, proto c It is calculated from the features of a certain number of samples.

[0114] 11.3 Finally, update the counting array n, which can be expressed as follows:

[0115] n c =n c +m c

[0116] Step 12: Keep the classification model from Step 2 predicting the category of detected moving targets in real time. Then, using the labeled batch data of moving targets, repeat the training and updating process from Step 2 to Step 10 on the copy of the classification model until all labeled batches have been processed. After that, update the copy to the classification model from Step 2.

Claims

1. A self-evolving target recognition method for video, characterized in that, Comprising the following steps: Step 1: converting an input video into video frames, and extracting detection bounding boxes of moving targets by using the ViBe detection algorithm; if there is a moving target, cropping the detection bounding box from the video frame and scaling it to a pre-classified image to be processed with a size of 224*224; Step 2: feeding the pre-classified image to be processed into a classification model, predicting a final category through forward propagation of the classification model, marking the detection bounding box of the moving target with the predicted category label and returning a result; since there are unseen categories or same-label samples from different domains in the self-evolution technology, the model may have prediction errors, therefore, further storing the pre-classified image to be processed obtained in step 1 into a temporary buffer area without labels for subsequent self-evolution training with manual labeling; Step 3: after the cumulative number of images in the unlabeled temporary buffer area exceeds N images, performing manual labeling on the images in the buffer area and dividing the images into M batches, wherein M<N, and M is a new-category batch; then firstly copying a copy of an online prediction classification model to local, and then updating and optimizing the classification model copy batch by batch with manually labeled samples; Step 4: Let each new batch in Step 3 be... and from memory buffer Retrieved memory batch samples Then for batches and Perform data augmentation operations to obtain the following results: and The memory buffer This refers to labeled images and cached historical images; the data augmentation includes random cropping and random flipping. Step 5: Use the feature extraction network from the classification model replica Extract batches separately original features and batch original features Then the features and characteristics The virtual features are fed into the Transformer Encoder Layer for non-linear cross-fusion using a batch attention mechanism. and ; Step 6: and By splicing , and By splicing A shared classifier is applied to both the original and virtual features, meaning they pass through the same classification layer; then, the features are calculated separately. and Cross-entropy loss based on logit mask and The formulaic expression for the cross-entropy loss based on the logit mask is as follows: Among them, features The corresponding tags are , For tags The unique set, Category in fully connected layers The weight, Category in fully connected layers The weight, Category in fully connected layers The bias term, Category in fully connected layers The bias term; The size of the current batch; Indicates the first Features of each sample; Step 7: Filter out memory batch samples The old class samples—that is, samples that do not belong to the current task label set—are used as the new batch. and use feature extraction network Extract the corresponding features ; Step 8: Utilize Perform feature distillation on the class prototype with the cumulative update of the corresponding label, and obtain the corresponding loss. If there is no corresponding cumulative update class prototype for the tag, then set... And wait for step 11 to update before performing characteristic distillation; Step 9: Weight the losses from the above steps to obtain the final loss. And update the model using stochastic gradient descent; Step 10: Finally, use the reservoir update algorithm to update the new batch of samples. Update to memory buffer middle; Step 11: if the current new-category batch is the last batch, updating class prototypes of each corresponding category by using all samples in a memory buffer area; otherwise, skipping this step; Step 12: keeping the classification model in step 2 predicting categories of detected moving targets in real time; then repeating the training and updating from step 2 to step 10 on the classification model copy by using the labeled moving target batch data until all labeled batches are processed, and then updating the classification model copy to the classification model in step 2; by virtue of the self-evolution learning technology for alleviating class imbalance, the knowledge previously learned by the classification model is greatly retained, and the capability of the classification model in learning to distinguish new categories is also ensured.

2. The self-evolving target recognition method for video according to claim 1, characterized in that, The step 5 is specifically: Step 5-1: Set up a feature extraction network The extracted batch features are ,in d is the feature dimension; Step 5-2: Analyze the currently extracted features Performing a dimensionality expansion operation allows the attention mechanism of the Transformer Encoder Layer to be applied to the batch dimension of the current features; its formulaic expression is as follows: Step 5-3: Expand the dimensions of the features from Step 5-2 The single-layer Transformer Encoder Layer performs a batch-level attention mechanism, enabling feature cross-fusion within the batch; its operation can be formally expressed as: Step 5-4: For The dimensionality compression operation can be formally represented as follows: 。 3. The self-evolving target recognition method for video according to claim 2, characterized in that, The formulation of feature distillation in said step 8 is: in, It is the prototype of the cumulative update class of category c. It is a category The cumulative update class prototype performs cosine normalization on each feature before calculating the cosine similarity. Representation of features The label corresponding to each feature of the old class sample.

4. The self-evolving target recognition method for video according to claim 3, characterized in that, The specific step of updating the corresponding class prototype in said step 11 is: Step 11-1: First calculate the memory buffer. The current class mean for each category is expressed as follows: in, It is an indicator function that has a value of 1 if and only if the condition within the parentheses is met; it resides in the memory buffer. The middle label is the category. The number of samples is indivual; This means The categories calculated from each sample Mean characteristics; This represents the image and label of each sample in the memory buffer; Step 11-2: accumulating the class mean in a memory into the class prototype, and the formulation thereof is represented as: in, It is a category The cumulative update class prototype, As of now How many sample features were used to calculate it; Step 11-3: Update the counting array Make the count value of category c in array n. .