Video analysis model generation method and device, video analysis method and device, equipment and medium

By using the teacher coding model to perform knowledge distillation and fine-tuning on the student coding model in the video analysis model, a target video analysis model is generated, which solves the problem of insufficient accuracy in video analysis in existing technologies and achieves accurate video analysis.

CN120877174APending Publication Date: 2025-10-31SHENZHEN INTELLIFUSION TECHNOLOGIES CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510860519.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-10-31

AI Technical Summary

Technical Problem

Existing video surveillance and analysis technologies cannot guarantee the accuracy of video analysis.

Method used

By acquiring distillation training videos and fine-tuning training videos, feature extraction and decoding of student coding models are performed using N teacher coding models. The student coding models are updated by combining the target loss function value to generate an initial video feature extraction model. The initial classification model is then fine-tuned to generate a target video analysis model.

Benefits of technology

It enables precise video analysis and improves the accuracy of video analysis.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120877174A_ABST
    Figure CN120877174A_ABST
Patent Text Reader

Abstract

The invention discloses a video analysis model generation method and device, a video analysis method and device, equipment and a medium. The method comprises the following steps: respectively carrying out feature extraction processing on a distillation training video by adopting N teacher coding models, and determining teacher feature extraction results corresponding to the teacher coding models; performing feature extraction processing on the distillation training video by adopting a student coding model, outputting a student feature extraction result, performing decoding processing on the student feature extraction result by adopting a decoding network corresponding to the teacher coding model, and determining a decoding result corresponding to the teacher coding model; determining a target loss function value based on N teacher feature extraction results and decoding results; updating the student coding model based on the target loss function value, and determining an initial video feature extraction model; and based on the fine-tuning training video, performing fine-tuning processing on the initial video feature extraction model and the initial classification model to generate a target video analysis model. According to the method, the purpose of accurately analyzing the video can be achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of artificial intelligence technology, and in particular to a video analysis model generation method, video analysis method, apparatus, device, and medium. Background Technology

[0002] With the acceleration of urbanization, video surveillance in public places plays a crucial role in public safety. For example, it involves monitoring and analyzing video footage collected in public places to determine whether any events endangering public safety exist. Current technologies typically use traditional motion detection algorithms (such as convolutional neural network algorithms) to monitor and analyze the collected video. However, existing video surveillance analysis technologies cannot guarantee the accuracy of video analysis. Therefore, how to accurately analyze video footage is a technical problem that needs to be solved. Summary of the Invention

[0003] This invention provides a video analysis model generation method, video analysis method, apparatus, device, and medium to solve the technical problem of how to accurately analyze videos.

[0004] A video analysis model generation method, comprising: Obtain distillation training videos and fine-tuning training videos; The distillation training video is processed by using N teacher coding models to extract features, and the teacher feature extraction result corresponding to each teacher coding model is determined, where N>1; The student coding model is used to extract features from the distillation training video, and the student feature extraction results are output. The decoding network corresponding to each teacher coding model is used to decode the student feature extraction results, and the decoding result corresponding to each teacher coding model is determined. Based on the teacher feature extraction and decoding results corresponding to N teacher coding models, the target loss function value is determined; Based on the target loss function value, the student coding model is updated to determine the initial video feature extraction model; Based on the fine-tuned training video, the initial video feature extraction model and the initial classification model are fine-tuned to generate a target video analysis model. The target video analysis model includes a target video feature extraction model and a target classification model connected in series with the target video feature extraction model.

[0005] A video analysis method, comprising: Obtain the video to be analyzed; The target video features are extracted from the video to be analyzed using a target video feature extraction model to determine the target features of the video; The target features of the video are classified using a target classification model to determine the target classification label corresponding to the video to be analyzed; The target video feature extraction model and the target classification model are models generated using the video analysis model generation method described above.

[0006] A video analysis model generation device, comprising: The video acquisition module is used to acquire distillation training videos and fine-tuning training videos; The teacher coding module is used to perform feature extraction processing on the distillation training video using N teacher coding models respectively, and to determine the teacher feature extraction result corresponding to each teacher coding model, where N>1; The decoding result determination module is used to perform feature extraction processing on the distillation training video using a student coding model, output student feature extraction results, and decode the student feature extraction results using a decoding network corresponding to each teacher coding model to determine the decoding result corresponding to each teacher coding model. The loss function value determination module determines the target loss function value based on the teacher feature extraction and decoding results corresponding to the N teacher coding models. The model update module updates the student coding model based on the target loss function value to determine the initial video feature extraction model; The fine-tuning module, based on the fine-tuning training video, fine-tunes the initial video feature extraction model and the initial classification model to generate a target video analysis model. The target video analysis model includes a target video feature extraction model and a target classification model connected in series with the target video feature extraction model.

[0007] A computer device includes a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the video analysis model generation method described above, or the processor implements the video analysis method described above when it executes the computer program.

[0008] A computer-readable storage medium storing a computer program that, when executed by a processor, implements the above-described video analysis model generation method, or, when executed by a processor, implements the above-described video analysis method.

[0009] The aforementioned video analysis model generation method, video analysis method, apparatus, equipment, and medium, in the distillation stage, distill the feature extraction knowledge of N teacher coding models with accurate extraction effects on different features into the student coding model, obtaining an initial video feature extraction model capable of accurately extracting multiple features of the video. In the video analysis task, the initial video feature extraction model and the initial classification model are fine-tuned to generate a target video analysis model including a target video feature extraction model and a target classification model set in series, used for accurate video analysis and to determine the video's category. This method uses N teacher coding models to perform knowledge distillation on the student coding model to obtain an initial video feature extraction model capable of accurately extracting multiple features from the video, and then fine-tunes the initial video feature extraction model and the initial classification model under the video analysis task to obtain a target video analysis model with high accuracy in video analysis, achieving the goal of accurate video analysis. Attached Figure Description

[0010] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0011] Figure 1 This is a flowchart of a video analysis model generation method in one embodiment of the present invention; Figure 2 This is another flowchart of the video analysis model generation method in one embodiment of the present invention; Figure 3 This is another flowchart of the video analysis model generation method in one embodiment of the present invention; Figure 4 This is another flowchart of the video analysis model generation method in one embodiment of the present invention; Figure 5 This is another flowchart of the video analysis model generation method in one embodiment of the present invention; Figure 6 This is another flowchart of the video analysis model generation method in one embodiment of the present invention; Figure 7 This is another flowchart of the video analysis model generation method in one embodiment of the present invention; Figure 8 This is another flowchart of the video analysis model generation method in one embodiment of the present invention; Figure 9 This is a schematic diagram of a video analysis model generation device in one embodiment of the present invention; Figure 10This is a schematic diagram of a computer device according to an embodiment of the present invention; Figure 11 This is a schematic diagram of a target video analysis model in one embodiment of the present invention; Figure 12 This is a schematic diagram of distillation training of a student coding model in one embodiment of the present invention; Figure 13 This is a structural diagram of the decoding network in one embodiment of the present invention; Figure 14 This is a schematic diagram of feature extraction of the DBMSE module in one embodiment of the present invention. Detailed Implementation

[0012] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0013] The video analysis model generation method provided in this embodiment of the invention can be applied to, for example... Figure 10 The computer device shown generates a video analysis model, which is used to analyze the acquired video to achieve accurate video analysis.

[0014] In one embodiment, such as Figure 1 As shown, a video analysis model generation method is provided, which is then applied to... Figure 10 Taking a computer device as an example, the explanation includes the following steps: S101: Obtain distillation training videos and fine-tuning training videos; S102: Use N teacher coding models to perform feature extraction processing on the distillation training video respectively, and determine the teacher feature extraction result corresponding to each teacher coding model, where N>1; S103: Use the student coding model to extract features from the distillation training video, output the student feature extraction results, use the decoding network corresponding to each teacher coding model to decode the student feature extraction results, and determine the decoding results corresponding to each teacher coding model. S104: Determine the target loss function value based on the teacher feature extraction and decoding results corresponding to N teacher coding models; S105: Update the student coding model based on the target loss function value to determine the initial video feature extraction model; S106: Based on the fine-tuned training video, the initial video feature extraction model and the initial classification model are fine-tuned to generate the target video analysis model. The target video analysis model includes the target video feature extraction model and the target classification model connected in series with the target video feature extraction model.

[0015] In this context, "distillation training video" refers to the video used to train the student coding model during the distillation phase. "Fine-tuning training video" refers to the video used to fine-tune the initial video feature extraction model and the initial classification model during the fine-tuning phase. The distillation phase refers to the stage of distilling the feature extraction knowledge from the teacher coding model into the student coding model. The teacher coding model is a pre-trained model used for feature extraction from videos, such as a teacher coding model for motion feature extraction, spatial feature extraction, and multimodal feature extraction. The student coding model is the model that needs to learn from the teacher coding model for feature extraction. The initial video feature extraction model is the feature extraction model generated by distilling the student coding model. The initial classification model is the model used to classify videos based on the features extracted from them.

[0016] In this embodiment, the initial video feature extraction model is fine-tuned to obtain the target video feature extraction model, and the initial classification model is fine-tuned to obtain the target classification model. The target video analysis model includes the target video feature extraction model and the target classification model set in series. The target video analysis model refers to the model used for classifying videos. The target video feature extraction model refers to the model, after distillation and fine-tuning, capable of extracting features from videos with relatively high accuracy. The target classification model refers to the model capable of classifying videos with relatively high accuracy based on their features.

[0017] In this embodiment, structurally, the student coding model, the initial video feature extraction model, and the target video feature extraction model all include a first vit-block module, a DBMSE module, and a second vit-block module arranged in series. The DBMSE module includes a vit-block sub-module and a mamba-block sub-module arranged in parallel. Figure 11 The image shown is a schematic diagram of a target video analysis model. Figure 11It is known that the target video feature extraction model includes a first vit-block module, a DBMSE module, and a second vit-block module arranged in series. The DBMSE module includes vit-block sub-modules and mamba-block sub-modules arranged in parallel. Specifically, the first vit-block module comprises M layers of vit-block sub-modules arranged in series, the target video feature extraction model comprises K layers of DBMSE modules arranged in series, each DBMSE module comprising vit-block sub-modules and mamba-block sub-modules arranged in parallel, and the second vit-block module comprises L layers of vit-block sub-modules arranged in series, where M > 1, K > 1, and L > 1. Understandably, the Transformer architecture has high accuracy in handling long-range dependencies and global information. Since video analysis has significant long-range dependencies and requires processing global information, the vit architecture (VisionTransformer) is applied to the student coding model, initial video feature extraction model, and target video feature extraction model for video analysis to improve the accuracy of video analysis. Meanwhile, the Mamba network, as a temporal modeling method, effectively integrates short-term and long-term memory information of videos to process dynamic video sequences. It preserves temporal information and enhances contextual understanding, thereby better capturing dynamic changes in video sequences. Furthermore, the Mamba network architecture, by introducing a state-space model, can more efficiently model complex temporal dependencies and sequence relationships, especially when processing long-sequence data, significantly reducing computational and memory overhead. Therefore, applying the VIT architecture and the Mamba-block architecture within the Mamba network to video analysis can reduce the computational complexity of video analysis while achieving relatively accurate results.

[0018] As an example, in step S101, the computer device acquires pre-collected distillation training videos and fine-tuning training videos, which are used to train the student coding model through the distillation training videos in the distillation stage to obtain an initial video feature extraction model, and in the fine-tuning stage, the initial video feature extraction model and the initial classification model set in series are fine-tuned through the fine-tuning training videos to obtain a target video analysis model including a target video feature extraction model and a target classification model.

[0019] Among them, the teacher feature extraction result refers to the feature extraction result output by the teacher coding model.

[0020] As an example, in step S102, the computer device inputs the distillation training video into N teacher coding models respectively, and outputs the teacher feature extraction results of each teacher coding model after performing feature extraction on the distillation training video. Understandably, since video analysis requires extracting multiple features from the video, including motion features, spatial features, and multimodal features, it is necessary to distill the feature extraction knowledge contained in multiple pre-trained teacher coding models into student coding models so that the distilled models can extract multiple features from the video more accurately. Specifically, the teacher coding model can extract one feature from the video more accurately. In this example, the distillation training video is processed by N teacher coding models respectively, and the teacher feature extraction result corresponding to each teacher coding model is determined. This allows for the subsequent determination of the target loss function value based on the teacher feature extraction result, and then the student coding model is updated based on the target loss function value, resulting in an initial video feature extraction model capable of accurately extracting multiple features from the video.

[0021] Here, student feature extraction results refer to the feature extraction results of the student coding model. Decoding results refer to the results of decoding the student feature extraction results.

[0022] As an example, in step S103, the computer device uses a student coding model to perform feature extraction processing on the distillation training video, outputs the student feature extraction results, and uses the decoding network corresponding to each teacher coding model to decode the student feature extraction results, determining the decoding result corresponding to each teacher coding model, resulting in N decoding results. Understandably, since each teacher coding model is used to accurately extract one feature from the distillation training video, in the distillation stage, the student feature extraction results are decoded according to each teacher coding model. This allows for the evaluation of the feature extraction effect of the student coding model based on the teacher feature extraction results corresponding to each teacher coding model, thereby updating the student coding model and achieving knowledge distillation of the student coding model. In this example, different teacher coding models decode different features in the student feature extraction results. For example, the decoding network corresponding to the teacher coding model that extracts motion features from the video is used to decode the motion features in the student feature extraction network. This decoding method, which decodes different features in the student feature extraction results using different teacher coding models, focuses on the features extracted by the corresponding teacher coding model. It can more accurately reflect the different features in the student feature extraction results. This method facilitates the subsequent determination of the target loss function value corresponding to the student coding model based on the decoding results of different teacher coding models. In turn, the knowledge distillation effect of the student coding model can be judged more accurately through the target loss function value.

[0023] The target loss function value refers to the loss function value during the distillation stage, which is used to characterize the degree of distillation training of the student coding model.

[0024] As an example, in step S104, the computer device processes the teacher feature extraction and decoding results corresponding to the same teacher coding model to determine the loss function value corresponding to each teacher coding model. The loss function values ​​corresponding to N teacher coding models are then processed to determine the target loss function value. For example, the computer device determines the difference between the teacher feature extraction result corresponding to each teacher coding model and the decoding result corresponding to that teacher coding model, and the sum of the differences corresponding to each teacher coding model is determined as the target loss function value. Understandably, since a teacher coding model can accurately extract one feature from the distillation training video, the target loss function value is used to characterize the difference between the student feature extraction result and the teacher feature extraction results corresponding to the N teacher coding models, thereby characterizing whether the student coding model accurately extracts multiple features from the video.

[0025] As an example, in step S105, the computer device determines whether the target loss function value meets the preset convergence condition. If the target loss function value meets the preset convergence condition, the student coding model is determined as the initial video feature extraction model. If the target loss function value does not meet the preset convergence condition, the student coding model is updated according to the preset gradient to obtain the updated student coding model. The updated student coding model is then used as the student coding model, and steps S102 to S105 are repeated until the target loss function value meets the preset convergence condition, at which point the student coding model is determined as the initial video feature extraction model. The preset convergence condition includes, but is not limited to, the target loss function value being less than a preset threshold or the difference between the target loss function values ​​of two consecutive updates of the student coding model being within a preset difference range.

[0026] In this example, when the target loss function value does not meet the preset convergence condition, the parameters in the first vit-block module, the DBMSE module, and the second vit-block module set in series in the student coding model are updated. Specifically, the parameters in the vit-block sub-module and the mamba-block sub-module set in parallel in the DBMSE module are updated until the target loss function value meets the preset convergence condition, resulting in the updated first vit-block module, the DBMSE module, and the second vit-block module set in series. Furthermore, the updated DBMSE module includes the vit-block sub-module and the mamba-block sub-module set in parallel. This not only has a high accuracy in extracting multiple features from the video, but also effectively reduces the computational complexity in the process of extracting multiple features from the video. It can achieve the goal of accurately extracting multiple features from the video without complex processing of the video.

[0027] In this example, the student coding model is distilled and trained based on the target loss function value. The student coding model is then updated so that the ability of N teacher coding models to accurately extract multiple features from the video is distilled into the student coding model, resulting in an initial video feature extraction model. This initial video feature extraction model can extract multiple features from the video more accurately.

[0028] As an example, in step S106, the computer device inputs the acquired fine-tuned training video into the initial video feature extraction model trained in the distillation stage for feature extraction, obtaining the feature extraction result. The feature extraction result is then input into the initial classification model for classification, and the loss function values ​​of the initial video feature extraction model and the initial classification model during the video classification process are obtained. Based on these loss function values, the initial video feature extraction model and the initial classification model are fine-tuned to obtain the target video feature extraction model and the target classification model. The target video feature extraction model and the target classification model are then concatenated to generate the target video analysis model. Here, the target video feature extraction model is the model fine-tuned from the initial video feature extraction model, and the target classification model is the model fine-tuned from the initial classification model.

[0029] In this example, the initial video feature extraction model obtained from distillation training is applied to a video analysis task for fine-tuning training. This yields a target video feature extraction model that can accurately and conveniently extract various video features from a video under the video analysis task. Furthermore, by fine-tuning the initial classification model under the video analysis task, a target classification model is obtained that can accurately classify videos based on the features extracted by the target video feature extraction model.

[0030] In this embodiment, during the distillation stage, the feature extraction knowledge of N teacher coding models that have accurate extraction effects on different features is distilled into the student coding model to obtain an initial video feature extraction model capable of accurately extracting multiple features of the video. In the video analysis task, the initial video feature extraction model and the initial classification model are fine-tuned to generate a target video analysis model including a target video feature extraction model and a target classification model set in series. This model is used for accurate video analysis to determine the video's category. This method uses N teacher coding models to perform knowledge distillation on the student coding model to obtain an initial video feature extraction model capable of accurately extracting multiple features from the video. The initial video feature extraction model and the initial classification model are then fine-tuned under the video analysis task to obtain a target video analysis model with high accuracy in video analysis, achieving the goal of accurate video analysis.

[0031] In one embodiment, the teacher coding model includes a video teacher model and an image teacher model.

[0032] The video teacher model refers to a teacher model used for motion feature extraction from videos. The image teacher model refers to a teacher model used for spatial feature and / or multimodal feature extraction from each frame of a video. Multimodal features include, but are not limited to, text modality, image modality, and audio modality.

[0033] Understandably, when extracting features from a video, in order to obtain the video's features more comprehensively, it is necessary to extract motion features from the video using a video teacher model and extract spatial and multimodal features from each image in the video using an image teacher model. This allows the student coding model to be updated based on the extracted motion, spatial, and multimodal features, enabling the student coding model to learn as many motion, spatial, and multimodal features as possible from the video and to extract these features more accurately.

[0034] In one embodiment, such as Figure 2 As shown, step S103 involves using a student coding model to extract features from the distillation training video, outputting the student feature extraction results, and using the decoding network corresponding to each teacher coding model to decode the student feature extraction results, determining the decoding result corresponding to each teacher coding model, including: S201: Perform vectorization and masking on the distillation training video to determine the mask vector matrix; S202: Use the student coding model to perform feature extraction on the mask vector matrix and output the student feature extraction results; S203: Using the decoding network corresponding to each video teacher model and the decoding network corresponding to each image teacher model, the student feature extraction results are decoded to determine the first decoding result corresponding to each video teacher model and the second decoding result corresponding to each image teacher model.

[0035] Vectorization refers to the process of converting video into vectors. Masking refers to the process of masking the vectors corresponding to the video to a certain extent.

[0036] As an example, in step S201, the computer device performs vectorization processing on the distillation training video to determine the vector matrix corresponding to the distillation training video. Then, according to a preset masking rate, it randomly masks the vector matrix corresponding to the distillation training video to obtain a masked vector matrix. In this example, as... Figure 12 The image shown illustrates the distillation training of the student coding model. Figure 12 As can be seen, for a collected distillation training video clip (video-clip), the computer device vectorizes the video clip using video embedding to obtain the corresponding vector matrix (makedinput). The computer device then performs random masking on the makedinput vector matrix using a preset masking rate of 90%, resulting in the masked vector matrix (visible token). Video embedding refers to converting the video into a matrix with a preset number of rows and columns. For example, video embedding converts the distillation training video into a vector matrix with 512 rows and 768 columns.

[0037] Feature extraction processing refers to the processing methods used to extract features.

[0038] As an example, in step S202, the computer device inputs the mask vector matrix corresponding to the distillation training video into the student coding model, and uses the student coding model to extract features from the mask vector matrix to determine the student feature extraction result corresponding to the distillation training video. Figure 12 As shown, the computer equipment uses a student encoder model to extract features from the visible token mask vector matrix, and determines the student feature extraction results corresponding to the visible token mask vector matrix.

[0039] The first decoding result refers to the decoding result obtained by the decoding network corresponding to the video teacher model based on the student feature extraction result. The second decoding result refers to the decoding result obtained by the decoding network corresponding to the video teacher model based on the student feature extraction result.

[0040] As an example, in step S203, the computer device uses the decoding network corresponding to each video teacher model to decode the student feature extraction results, determining a first decoding result corresponding to each video teacher model. Then, it uses the decoding network corresponding to each image teacher model to decode the student feature extraction results, determining a second decoding result corresponding to each image teacher model. Understandably, by decoding the student feature extraction results using the decoder corresponding to each type of teacher coding model, a decoding result corresponding to each teacher coding model is obtained. This facilitates subsequent updates to the student coding model based on the teacher feature extraction results and decoding results corresponding to the same teacher coding model, enabling the student coding model to more comprehensively learn the feature extraction targets corresponding to each teacher coding model.

[0041] In this embodiment, the training video is vectorized and masked to obtain a mask vector matrix, so that the student coding model is in the mask distillation stage during the training process, and can accurately learn the knowledge of the video teacher model to extract motion features in the video, as well as the knowledge of the image teacher model to extract spatial features and multimodal features in the video.

[0042] In one embodiment, such as Figure 3 As shown, step S203, which involves using the decoding network corresponding to each video teacher model and the decoding network corresponding to each image teacher model to decode the student feature extraction results and determine the first decoding result corresponding to each video teacher model and the second decoding result corresponding to each image teacher model, includes: S301: The student feature extraction results are spliced ​​using the learnable mask matrix corresponding to each video teacher model to determine the first splicing matrix corresponding to each video teacher model. The student feature extraction results are spliced ​​using the learnable mask matrix corresponding to each image teacher model to determine the second splicing matrix corresponding to each image teacher model. S302: Use the decoding network corresponding to each video teacher model to decode the first splicing matrix corresponding to each video teacher model, and determine the first decoding result corresponding to each video teacher model. S303: The second concatenation matrix corresponding to each image teacher model is decoded using the decoding network corresponding to each image teacher model to determine the second decoding result corresponding to each image teacher model.

[0043] Here, the learnable mask matrix refers to the matrix used to ensure that the dimension of the student feature extraction results is consistent with the dimension of the teacher feature extraction results of the corresponding teacher coding model. The first concatenation matrix is ​​the matrix obtained by concatenating the student feature extraction results with the learnable mask matrix corresponding to the video teacher model. The second concatenation matrix is ​​the matrix obtained by concatenating the student feature extraction results with the learnable mask matrix corresponding to the image teacher model.

[0044] As an example, in step S301, the computer device uses a learnable mask matrix corresponding to a video teacher model to concatenate the student feature extraction results, obtaining a first concatenated matrix with the same dimension as the teacher feature extraction results of the video teacher model. The same operation is performed using the learnable mask matrix corresponding to each video teacher model to obtain a first concatenated matrix corresponding to each video teacher model. Similarly, the computer device uses a learnable mask matrix corresponding to an image teacher model to concatenate the student feature extraction results, obtaining a second concatenated matrix with the same dimension as the teacher feature extraction results of the image teacher model. The same operation is performed using the learnable mask matrix corresponding to each image teacher model to obtain a second concatenated matrix corresponding to each image teacher model. In this example, by... Figure 12It is known that the video teacher model is a pre-trained VideoMAEv2 model, specifically the videomaev2-g model, used to extract motion features from the distillation training video with relatively high accuracy. The image teacher models are the mae_h model and the InternViT-6B-224px model. The mae_h model is used to extract spatial features from each frame of the distillation training video, while the InternViT-6B-224px model is used to extract multimodal features from each frame. The computer equipment uses the learnable mask matrix corresponding to the videomaev2-g model to concatenate the student feature extraction results, obtaining the first concatenated matrix mask token-g. The learnable mask matrix corresponding to the mae_h model is then used to concatenate the student feature extraction results, obtaining the second concatenated matrix mask token-h. Finally, the learnable mask matrix corresponding to the InternViT-6B-224px model is used to concatenate the student feature extraction results, obtaining the second concatenated matrix mask token-Vit. At this point, N=3. For example, if the dimension of the teacher feature extraction result corresponding to the InternViT-6B-224px model is 1024*768, and the dimension of the student feature extraction result is 512*768, then the dimension of the learnable mask matrix corresponding to the InternViT-6B-224px model is 512*768. The computer device concatenates the learnable mask matrix corresponding to the InternViT-6B-224px model with the student feature extraction result to obtain a second concatenated matrix with the same dimension as the teacher feature extraction result of the InternViT-6B-224px model. In essence, the student feature extraction result is concatenated using the learnable mask matrix corresponding to each teacher encoding model to obtain a first and a second concatenated matrix with the same dimension as the teacher feature extraction result corresponding to each teacher encoding model. This facilitates subsequent decoding of the first and second concatenated matrices to obtain a decoding result with the same dimension as the teacher feature extraction result corresponding to each teacher encoding model, thus making it feasible to calculate the loss function using the teacher feature extraction result and decoding result with the same dimension corresponding to each teacher encoding model.

[0045] As an example, in step S302, the computer device inputs the first concatenation matrix into the decoding network corresponding to the video teacher model and outputs the first decoding result. In this example, the computer device inputs the first concatenation matrix mask token-g corresponding to the videomaev2-g model into the decoding network corresponding to the pre-trained videomaev2-g model and outputs the first decoding result.

[0046] As an example, in step S303, the computer device inputs the second concatenation matrix into the decoding network corresponding to the image teacher model and outputs the second decoding result. In this example, the image teacher model includes the mae_h model and the InternViT-6B-224px model. The computer device inputs the second concatenation matrix mask token-h corresponding to the mae_h model into the pre-trained decoding network corresponding to the mae_h model and outputs the second decoding result corresponding to the mae_h model. Similarly, the computer device inputs the second concatenation matrix mask token-Vit corresponding to the InternViT-6B-224px model into the pre-trained decoding network corresponding to the InternViT-6B-224px model and outputs the second decoding result corresponding to the InternViT-6B-224px model.

[0047] In this embodiment, learnable mask matrices corresponding to different teacher coding models are used to concatenate the student feature extraction results, so as to obtain a concatenated matrix and decoding results with the same dimension as the teacher feature extraction results corresponding to the teacher coding model. This allows the target loss function value to be determined based on the decoding results and teacher feature extraction results with the same dimension, thereby achieving accurate and effective updates to the student coding model.

[0048] In one embodiment, the decoding network includes multiple vit-block modules and a multilayer perceptron module arranged in series.

[0049] As an example, such as Figure 13 The diagram shows the structure of the decoding network, which includes a series-connected vit-block module and a multilayer perceptron module (MLP). The decoding network comprises two series-connected vit-block modules for decoding processing. In steps S302 and S303, the computer device sequentially inputs a first stitching matrix into the two series-connected vit-block modules and the MLP of the decoding network, outputting a first decoding result corresponding to the first stitching matrix. The computer device then sequentially inputs a second stitching matrix into the two series-connected vit-block modules and the MLP of the decoding network, outputting a second decoding result corresponding to the second stitching matrix.

[0050] In this embodiment, when the decoding network, which includes multiple vit-block modules and multilayer perceptron modules arranged in series, is used as the decoding network corresponding to the video teacher model, it can focus on motion features. When used as the decoding network corresponding to the image teacher model, it can focus on spatial features and multimodal features, thereby achieving accurate decoding.

[0051] In one embodiment, such as Figure 4As shown, step S104, which is to determine the target loss function value based on the teacher feature extraction and decoding results corresponding to N teacher coding models, includes: S401: Based on the teacher feature extraction and decoding results corresponding to N teacher coding models, determine the initial loss function value for each teacher coding model; S402: Use the target weights corresponding to each teacher coding model to weight the initial loss function values ​​corresponding to the teacher coding model, and determine the target loss function value.

[0052] The initial loss function value refers to the loss function value corresponding to each teacher's coding model.

[0053] As an example, in step S401, the computer device calculates the loss function value for the teacher feature extraction result corresponding to each teacher coding model and the decoding result corresponding to the same teacher coding model, determining the initial loss function value of the student coding model relative to each teacher coding model. In this example, the computer device determines the reconstruction loss function value (e.g., Smooth L1 Loss) between the teacher feature extraction result corresponding to each teacher coding model and the decoding result corresponding to the same teacher coding model, and uses the reconstruction loss function value as the initial loss function value of the student coding model relative to each teacher coding model. In this example, the teacher coding model includes a video teacher model and an image teacher model. Figure 12 As shown, the video teacher model is the videomaev2-g model, and the image teacher models are the mae_h model and the InternViT-6B-224px model. The computer equipment calculates the reconstruction loss function value based on the teacher feature extraction results and the first decoding results corresponding to the videomaev2-g model to determine the initial loss function value corresponding to the videomaev2-g model. The computer equipment calculates the reconstruction loss function value based on the teacher feature extraction results corresponding to the mae_h model and the second decoding results corresponding to the mae_h model, thus determining the initial loss function value corresponding to the mae_h model. The computer equipment calculates the reconstruction loss function value based on the teacher feature extraction results corresponding to the InternViT-6B-224px model and the second decoding results corresponding to the InternViT-6B-224px model, thus determining the initial loss function value corresponding to the InternViT-6B-224px model. In this example, when determining the initial loss function values ​​for the image teacher model mae_h and the InternViT-6B-224px model, the initial loss function value can be calculated for each frame of the distillation training video to accurately determine the initial loss function value. Alternatively, the initial loss function value can be calculated only for the last frame of the distillation training video to save computational resources.

[0054] The target weight is used to characterize the proportion of the teacher's coding model in the loss function value.

[0055] As an example, in step S402, the computer device uses the target weights corresponding to each teacher coding model to weight the initial loss function values ​​corresponding to the teacher coding models, thereby obtaining the target loss function value. In this example, the target loss function value... ,in, The target loss function value, Let the target weights be the encoding model corresponding to the i-th teacher. Let be the initial loss function value corresponding to the i-th teacher's coding model. The number of teacher coding models. For example... Figure 12 As shown, when the computer device determines that the teacher coding model includes a video teacher model and an image teacher model, and that the video teacher model is the videomaev2-g model, and the image teacher models are the mae_h model and the InternViT-6B-224px model, it determines... =3, and the target loss function value for: in, These are the target weights corresponding to the videomaev2-g model. These are the target weights corresponding to the mae_h model. The target weights for the InternViT-6B-224px model. The initial loss function value corresponding to the videomaev2-g model. These are the initial loss function values ​​corresponding to the mae_h model. These are the initial loss function values ​​for the InternViT-6B-224px model. In this example, =2, =1.5, =1.5.

[0056] In this embodiment, the initial loss function value corresponding to each teacher coding model is weighted according to the target weight corresponding to each teacher coding model. This takes into account the importance of each teacher coding model in the calculation of the loss function value, and can reasonably determine the target loss function value. This makes it feasible to update the student coding model based on the reasonably determined target loss function value to obtain a more accurate initial video feature extraction model.

[0057] In one embodiment, the fine-tuning training video includes standard classification labels; The standard classification label refers to the pre-assigned standard label for the fine-tuning training video, used to characterize the category corresponding to the fine-tuning training video. For example, for fine-tuning training videos collected in public places, according to the behavior of the monitored object (human or other animal), the standard classification labels include fighting, normal, and arguing labels, etc. After acquiring the fine-tuning training video, the fine-tuning training video is identified to determine the standard classification label corresponding to the fine-tuning training video, which is used to fine-tune the initial video feature extraction model and the initial classification model.

[0058] In one embodiment, such as Figure 5 As shown, step S106, which involves fine-tuning the initial video feature extraction model and the initial classification model based on the fine-tuned training video to generate the target video analysis model, includes: S501: Use the initial video feature extraction model to perform feature extraction processing on the fine-tuned training video to determine the feature matrix to be classified; S502: Use the initial classification model to classify the feature matrix to be classified, and determine the initial classification labels corresponding to the fine-tuned training videos; S503: Determine the classification loss function value based on the standard classification label and the initial classification label; S504: Based on the classification loss function value, the initial video feature extraction model and the initial classification model are fine-tuned to obtain the target video feature extraction model and the target classification model.

[0059] The feature matrix to be classified refers to the matrix formed by feature extraction from the initial video feature extraction model.

[0060] As an example, in step S501, the computer device inputs the fine-tuning training video used for model fine-tuning into the initial video feature extraction model obtained in the mask distillation stage, and extracts features such as motion features, spatial features and multimodal features in the fine-tuning training video to obtain the feature matrix to be classified corresponding to the fine-tuning training video.

[0061] The initial classification label refers to the label of the fine-tuned training video determined by the feature matrix to be classified.

[0062] As an example, in step S502, the computer device inputs the feature matrix to be classified into the initial classification model and outputs the initial classification label corresponding to the fine-tuned training video. In this example, the initial classification model is an untuned classification head used to classify the features corresponding to the video. If the task type of video analysis is to determine whether a video is a fight, then the initial classification model is used to output whether the initial classification label corresponding to the video is a fight label.

[0063] The classification loss function value refers to the function value used to characterize the difference between the initial classification label and the standard classification label.

[0064] As an example, in step S503, the computer device calculates the standard classification labels corresponding to the fine-tuning training video. Compared with the initial classification model The difference between the values ​​is used as the classification loss function value. In this example, the computer device will use the standard classification labels... Compared with the initial classification model The cross-entropy loss function value between them is used to determine the standard classification label. Compared with the initial classification model The difference between the values, i.e., the classification loss function value, is the standard classification label. Compared with the initial classification model The cross-entropy loss function value. In this example, the cross-entropy loss function value is LOSS=Cross Entropy( ).

[0065] As an example, in step S504, after determining the classification loss function values ​​of the initial video feature extraction model and the initial classification model for the fine-tuned training video, the computer device determines whether the classification loss function value has converged. If the classification loss function value has converged, the initial video feature extraction model is determined as the target video feature extraction model, and the initial classification model is determined as the target classification model. If the computer device determines that the classification loss function value has not converged, it fine-tunes the initial video feature extraction model and the initial classification model to obtain the fine-tuned initial video feature extraction model and the fine-tuned initial classification model. Based on the fine-tuned initial video feature extraction model and the fine-tuned initial classification model, steps S501 to S504 are re-executed until the classification loss function value converges. The initial video feature extraction model when the classification loss function value converges is determined as the target video feature extraction model, and the initial classification model when the classification loss function value converges is determined as the target classification model. Understandably, when the classification loss function converges, the difference between the standard classification label and the initial classification label is small. Therefore, determining the initial video feature extraction model and the initial classification model allows for relatively accurate video classification. Thus, the initial video feature extraction model at which the classification loss function converges is determined as the target video feature extraction model, and the initial classification model at which the classification loss function converges is determined as the target classification model. In this example, the computer device uses the Adamw algorithm to update the parameters in the initial video feature extraction model and the initial classification model, achieving fine-tuning of these models.

[0066] In this embodiment, the initial video feature extraction model and the initial classification model are fine-tuned by fine-tuning the training video to obtain a target video feature extraction model and a target classification model that can classify videos more accurately.

[0067] In another embodiment, such as Figure 6 As shown, a video analysis method is provided for accurate video analysis using a target video feature extraction model and a target classification model trained and fine-tuned by distillation masking. The method includes: S601: Obtain the video to be analyzed; S602: Use a target video feature extraction model to extract features from the video to be analyzed and determine the target features of the video; S603: Use a target classification model to classify the video target features and determine the target classification label corresponding to the video to be analyzed; In this embodiment, the target video feature extraction model and the target classification model are models generated using the video analysis model generation method described in the above embodiment.

[0068] Among them, the video to be analyzed refers to the video that needs to be labeled and classified using the target video feature extraction model and the target classification model.

[0069] As an example, in step S601, the computer device acquires the video to be analyzed that needs to be tagged and classified in the application scenario. For example, a video collected in a public place that needs to be analyzed to determine whether it is a fight.

[0070] Among them, video target features refer to the features corresponding to the video to be analyzed.

[0071] As an example, in step S602, the computer device uses the target video feature extraction model generated by the video analysis model generation method of the above embodiment to extract features from the video to be analyzed and outputs the video target features corresponding to the video to be analyzed. It can be understood that the target video feature extraction model obtained by distillation and fine-tuning can accurately analyze multiple features in the video. By using the target video feature extraction model, multiple features in the video to be analyzed can be extracted relatively accurately.

[0072] Target classification labels refer to the results of video analysis of the video being analyzed. Understandably, the types of target classification labels differ in different video surveillance analysis scenarios. For example, in a scenario analyzing whether there are situations endangering public safety in public places, target classification labels include labels for fighting, normal situations, and arguments.

[0073] As an example, in step S603, the computer device inputs the video target features corresponding to the video to be analyzed into the target classification model corresponding to the video surveillance analysis scenario, and outputs the target classification label corresponding to the video to be analyzed. For example, in a monitoring and analysis scenario of whether there are situations endangering public safety in public places, the target classification model classifies the video to be analyzed according to the video target features and outputs the target classification label corresponding to the video to be analyzed. For example, the target classification label is a fighting label, so as to facilitate real-time monitoring of public safety in public places and handling of sudden safety events (such as fighting events). Figure 11 As shown, the target classification model is the fine-tuned classification head, ClassificationHead. If the video surveillance analysis task is to identify fighting, the classification head is FightHead.

[0074] In this embodiment, a target video feature extraction model generated through distillation and training is used to accurately and quickly extract spatial and motion features from the video to be analyzed, thereby obtaining video target features. The target classification model generated through fine-tuning is then used to accurately classify the video target features, thereby obtaining the target classification label corresponding to the video to be analyzed. This method can not only accurately monitor and analyze videos, but also save manpower and resources. It can be widely applied in various video monitoring and analysis scenarios and has high application value.

[0075] In one embodiment, the target video feature extraction model includes a first vit-block module, a DBMSE module, and a second vit-block module set in series.

[0076] like Figure 11 As shown, the target video feature extraction model includes a first vit-block module, a DBMSE module, and a second vit-block module arranged in series. The DBMSE module includes vit-block sub-modules and mamba-block sub-modules arranged in parallel. Specifically, the first vit-block module comprises M layers of vit-block sub-modules arranged in series, the target video feature extraction model comprises K layers of DBMSE modules arranged in series, each DBMSE module comprising vit-block and mamba-block sub-modules arranged in parallel, and the second vit-block module comprises L layers of vit-block sub-modules arranged in series. Understandably, the Transformer architecture has high accuracy when handling long-range dependencies and global information. Since video analysis has significant long-range dependencies and requires processing global information, the Vision Transformer architecture is applied to the target video feature extraction model to improve the accuracy of video analysis. Meanwhile, the Mamba network, as a temporal modeling method, is used to process dynamic video sequences. It can effectively integrate short-term and long-term memory information of videos, preserving information in the temporal dimension and enhancing contextual understanding, thereby better capturing dynamic changes in video sequences. Furthermore, the Mamba network architecture, by introducing a state-space model, can more efficiently model complex temporal dependencies and sequence relationships, especially when processing long sequence data, significantly reducing computational and memory overhead. In this embodiment, the VIT architecture and the Mamba-block architecture in the Mamba network are applied to video analysis, which can reduce the computational complexity of video analysis while performing video analysis with relatively high accuracy.

[0077] In one embodiment, such as Figure 7 As shown, step S602, which involves using a target video feature extraction model to extract features from the video to be analyzed and determine the target features of the video, includes: S701: The first vit-blcok module is used to extract features from the video to be analyzed and determine the coarse features of the video. S702: The DBMSE module is used to extract coarse features from the video to determine the video spatial features and video motion features; S703: The second vit-block module is used to integrate video spatial features and video motion features to determine video target features.

[0078] Among them, video coarse features refer to the coarse-grained features obtained after feature extraction of the video to be analyzed.

[0079] As an example, in step S701, when the computer device uses the target video analysis model obtained after training and fine-tuning through distillation masks to classify the video to be analyzed, the video to be analyzed is input into the target video analysis model. The first vit-block module of the target video feature extraction model extracts coarse-grained features of the video to be analyzed, thus obtaining the coarse features of the video to be analyzed. In this example, such as... Figure 11 As shown, the first vit-block module includes M tandem vit-block sub-modules. The computer device uses the distilled and fine-tuned M tandem vit-block sub-modules to extract coarse-grained features from the video to be analyzed, which can extract the coarse features of the video to be analyzed relatively accurately.

[0080] Among them, video spatial features refer to the spatial features of the video to be analyzed. Video motion features refer to the motion features of the video to be analyzed.

[0081] As an example, in step S702, the computer device inputs the coarse video features output by the first vit-block module into the DBMSE module of the target video feature extraction model. The DBMSE module further extracts the spatial and motion features from the coarse video features to obtain the video spatial features and video motion features. For example, the coarse video features corresponding to the video to be analyzed, collected in a public scene, are input into the DBMSE module of the target video feature extraction model, which outputs the video spatial features and video motion features corresponding to the video to be analyzed. Understandably, by using the distilled and fine-tuned DBMSE module to extract features from the coarse video features, the video spatial features and video motion features corresponding to the video to be analyzed can be accurately extracted.

[0082] Among them, video target features refer to the results of extracting multiple features from the video to be analyzed through a target video feature extraction model that has been distilled and fine-tuned.

[0083] As an example, in step S703, the computer device inputs the video spatial features and video motion features corresponding to the video to be analyzed into the second vit-block module in the target video feature extraction model. The second vit-block module integrates and processes the video spatial features and video motion features, and outputs the video target features corresponding to the video to be analyzed. Figure 11 As shown, the second vit-block module includes L-layer cascaded vit-block sub-modules. By using L-layer cascaded vit-block sub-modules, video spatial features and video motion features can be comprehensively integrated and processed to obtain video target features that contain relatively comprehensive video spatial features and video motion features.

[0084] In this embodiment, the target video feature extraction model generated through distillation and training sequentially uses a first vit-block module, a DBMSE module, and a second vit-block module to accurately extract spatial and motion features from the video to be analyzed, thereby obtaining the video target features. This method can achieve the purpose of accurate video monitoring and analysis and has high application value.

[0085] In one embodiment, the DBMSE module includes a vit-block submodule and a mamba-block submodule configured in parallel.

[0086] like Figure 11 As shown, the target video feature extraction model includes a K-layer cascaded DBMSE module. Each DBMSE module includes a vit-block sub-module and a mamba-block sub-module set in parallel. The K-layer cascaded DBMSE module can accurately extract fine-grained features such as spatial features and motion features from the coarse features of the video. Furthermore, since the vit-block sub-module in the DBMSE module has the advantage of accurate video analysis and the mamba-block sub-module has the advantage of reducing the computational complexity during video analysis, the distilled and fine-tuned DBMSE module can accurately and quickly obtain the video spatial features and video motion features corresponding to the video to be analyzed by extracting features from the coarse features of the video.

[0087] In one embodiment, such as Figure 8 As shown, step S702, which involves using the DBMSE module to extract coarse features from the video and determine the video spatial features and video motion features, includes: S801: Classify the coarse features of the video according to a preset classification ratio, and determine the first input feature and the second input feature. The proportion of the first input feature is less than the proportion of the second input feature. S802: The vit-block submodule is used to extract spatial features from the first input features to determine the video spatial features; S803: The mamba-block submodule is used to extract motion features from the second input features to determine the video motion features.

[0088] The preset classification ratio refers to the proportion of coarse video features used for classification. The first input feature refers to the features used as input to the vit-block submodule of the DBMSE module. The second input feature refers to the features used as input to the mamba-block submodule of the DBMSE module.

[0089] As an example, in step S801, the computer device classifies the video coarse features according to a preset classification ratio, obtaining a first input feature for inputting into the vit-block submodule of the DBMSE module, and a second input feature for inputting into the mamba-block submodule of the DBMSE module. In this example, the preset classification ratio is 1:3, meaning the first input feature is a coarse video feature. The second input feature is the coarse feature of the video. .like Figure 14 The image shown is a schematic diagram of feature extraction from a DBMSE module. Figure 14 It can be seen that the computer device will use the video coarse feature token_list The input is fed into the vit-block submodule of the DBMSE module, which will convert the video coarse feature token_list into a single block. The input is fed into the `mamba-block` submodule of the DBMSE module. Understandably, the `vit-block` submodule is used to extract spatial features from the video to be analyzed, while the `mamba-block` submodule is used to extract motion features. Since motion features are more important than spatial features in video surveillance analysis tasks, a larger proportion of the coarse video features are used as the second input features to the `mamba-block` submodule for motion feature extraction. This allows for more accurate extraction of the more important motion features. Furthermore, the Mamba network corresponding to the `mamba-block` submodule has the advantage of reducing computational load and saving computing power. Using a larger proportion of the coarse video features as the second input features allows the `mamba-block` submodule to extract features from this larger proportion of the second input features, effectively reducing computational load and accelerating the video analysis process.

[0090] As an example, in step S802, the computer device inputs a relatively small proportion of the first input features into the vit-block submodule of the DBMSE module in the target video feature extraction model. The vit-block submodule of the DBMSE module extracts spatial fine-grained features from the first input features to obtain the video spatial features corresponding to the video to be analyzed. In this example, such as... Figure 14 As shown, the computer device will extract coarse features from the video. As the first input feature, it is input into the vit-block submodule of the DBMSE module in the target classification model, and outputs the video spatial features corresponding to the video to be analyzed.

[0091] As an example, in step S803, the computer device inputs the second input feature, which has a larger proportion, into the mamba-block submodule of the DBMSE module in the target video feature extraction model. The mamba-block submodule of the DBMSE module extracts fine-grained motion features from the second input feature to obtain the video motion features corresponding to the video to be analyzed. In this example, such as... Figure 14 As shown, the computer device will extract coarse features from the video. As the second input feature, it is fed into the mamba-block submodule of the DBMSE module in the target classification model, and outputs the video motion features corresponding to the video to be analyzed.

[0092] In this embodiment, as Figure 14 As shown, the computer device concatenates the video spatial features and video motion features output by the previous DBMSE module in the target video feature extraction model, updates them into coarse video features, and inputs them into the next DBMSE module to repeat steps S701 to S703 until the last DBMSE module, obtaining video spatial features and video motion features, which are then input into the second vit-block module of the target video feature extraction model for integration.

[0093] In this embodiment, a smaller proportion of the first input features and a larger proportion of the second input features are determined according to a preset classification ratio. The vit-block submodule of the DBMSE module extracts video spatial features from the smaller proportion of the first input features, and the mamba-block submodule of the DBMSE module extracts video motion features from the larger proportion of the second input features. This method accurately extracts the video spatial features and video motion features that are important for video analysis. Furthermore, while accurately extracting video spatial features and video motion features, this method makes good use of the advantages of the Mamba network corresponding to the mamba-block submodule in reducing computational load and saving computing power. By using the mamba-block submodule to extract motion features from the larger proportion of the second input features, the computational load can be effectively reduced, the video analysis process can be accelerated, and thus the efficiency of video analysis can be effectively improved.

[0094] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0095] In one embodiment, a video analysis model generation apparatus is provided, which corresponds one-to-one with the video analysis model generation method in the above embodiments. For example... Figure 9 As shown, the video analysis model generation device includes a video acquisition module 901, a teacher encoding module 902, a decoding result determination module 903, a loss function value determination module 904, a model update module 905, and a fine-tuning module 906. Detailed descriptions of each functional module are as follows: The video acquisition module 901 is used to acquire distillation training videos and fine-tuning training videos. Teacher coding module 902 is used to perform feature extraction processing on the distillation training video using N teacher coding models respectively, and to determine the teacher feature extraction result corresponding to each teacher coding model, where N>1; The decoding result determination module 903 is used to perform feature extraction processing on the distillation training video using the student coding model, output the student feature extraction result, and use the decoding network corresponding to each teacher coding model to decode the student feature extraction result and determine the decoding result corresponding to each teacher coding model. The loss function value determination module 904 determines the target loss function value based on the teacher feature extraction and decoding results corresponding to N teacher coding models. The model update module 905 updates the student coding model based on the target loss function value to determine the initial video feature extraction model; The fine-tuning module 906, based on the fine-tuning training video, performs fine-tuning on the initial video feature extraction model and the initial classification model to generate the target video analysis model. The target video analysis model includes the target video feature extraction model and the target classification model connected in series with the target video feature extraction model.

[0096] In one embodiment, the decoding result determination module 903 includes: The mask vector matrix determination submodule is used to perform vectorization and masking processing on the distillation training video to determine the mask vector matrix; The student feature extraction result determination submodule is used to perform feature extraction processing on the mask vector matrix using the student coding model and output the student feature extraction result. The decoding result determination submodule is used to decode the student feature extraction results using the decoding network corresponding to each video teacher model and the decoding network corresponding to each image teacher model, and to determine the first decoding result corresponding to each video teacher model and the second decoding result corresponding to each image teacher model.

[0097] In one embodiment, the decoding result determination submodule includes: The splicing matrix determination unit is used to splice the student feature extraction results using the learnable mask matrix corresponding to each video teacher model to determine the first splicing matrix corresponding to each video teacher model, and to splice the student feature extraction results using the learnable mask matrix corresponding to each image teacher model to determine the second splicing matrix corresponding to each image teacher model. The first decoding result determination unit is used to perform decoding processing on the first splicing matrix corresponding to each video teacher model using the decoding network corresponding to each video teacher model, and determine the first decoding result corresponding to each video teacher model. The second decoding result determination unit is used to perform decoding processing on the second splicing matrix corresponding to each image teacher model using the decoding network corresponding to each image teacher model, and to determine the second decoding result corresponding to each image teacher model.

[0098] In one embodiment, the loss function value determination module 904 includes: The initial loss function value determination submodule determines the initial loss function value for each teacher coding model based on the teacher feature extraction and decoding results corresponding to N teacher coding models. The target loss function value determination submodule is used to weight the initial loss function value corresponding to each teacher coding model using the target weights corresponding to each teacher coding model, and then determine the target loss function value.

[0099] In one embodiment, the fine-tuning module 906 includes: The submodule for determining the feature matrix to be classified is used to perform feature extraction processing on the fine-tuned training video using the initial video feature extraction model to determine the feature matrix to be classified. The initial classification label determination submodule is used to classify the feature matrix to be classified using the initial classification model and determine the initial classification label corresponding to the fine-tuning training video. The classification loss function value determination submodule determines the classification loss function value based on the standard classification label and the initial classification label; The fine-tuning submodule, based on the classification loss function value, fine-tunes the initial video feature extraction model and the initial classification model to obtain the target video feature extraction model and the target classification model.

[0100] In another embodiment, a video analysis device is provided, which corresponds one-to-one with the video analysis methods in the above embodiments. The video analysis device includes a video acquisition module, a video target feature determination module, and a target classification label determination module. Detailed descriptions of each functional module are as follows: The video acquisition module is used to acquire the video to be analyzed. The video target feature determination module is used to extract features from the video to be analyzed using a target video feature extraction model, and to determine the video target features. The target classification label determination module is used to classify the video target features using a target classification model and determine the target classification label corresponding to the video to be analyzed.

[0101] In one embodiment, the video target feature determination module includes: The first feature extraction submodule is used to extract features from the video to be analyzed using the first vit-blcok module, and determine the coarse features of the video. The second feature extraction submodule is used to extract coarse features from the video using the DBMSE module, and to determine the video spatial features and video motion features. The feature integration submodule is used to integrate video spatial features and video motion features using the second vit-block module to determine video target features.

[0102] In one embodiment, the second feature extraction submodule includes: The input feature determination unit is used to classify the coarse features of the video according to a preset classification ratio, and determine the first input feature and the second input feature, wherein the ratio of the first input feature is smaller than the ratio of the second input feature; The first feature extraction unit is used to extract spatial features from the first input features using the vit-block submodule to determine the video spatial features; The second feature extraction unit is used to extract motion features from the second input features using the mamba-block submodule to determine the video motion features.

[0103] Specific limitations regarding the video analysis model generation device can be found in the limitations regarding the video analysis model generation method described above, and specific limitations regarding the video analysis device can be found in the limitations regarding the video analysis method described above; they will not be repeated here. The various modules within the aforementioned video analysis model generation device and video analysis device can be implemented entirely or partially through software, hardware, or a combination thereof. These modules can be embedded in hardware or independently of the processor in the computer device, or stored in software in the memory of the computer device, so that the processor can call and execute the corresponding operations of each module.

[0104] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 10 As shown. The computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The database stores data used or generated during the execution of the method. The network interface communicates with external terminals via a network connection. When executed by the processor, the computer program implements a video analysis model generation method. Alternatively, when executed by the processor, the computer program implements a video analysis method.

[0105] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the video analysis model generation method described in the above embodiments, for example... Figure 1 As shown in S101-S106, or Figures 2 to 5 As shown, to avoid repetition, it will not be described again here. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in this embodiment of the video analysis model generation device, for example... Figure 9 The functions of the video acquisition module 901, teacher encoding module 902, decoding result determination module 903, loss function value determination module 904, model update module 905, and fine-tuning module 906 shown are not described in detail here to avoid repetition.

[0106] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, it implements the video analysis method described in the above embodiments, for example... Figure 6 As shown in S601-S603, or Figures 7 to 8 As shown, to avoid repetition, it will not be described again here. Alternatively, when the processor executes the computer program, it implements the functions of each module / unit in this embodiment of the video analysis device; to avoid repetition, this will not be described again here.

[0107] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When executed by a processor, the computer program implements the video analysis model generation method described in the above embodiments, for example... Figure 1 As shown in S101-S106, or Figures 2 to 5 As shown, to avoid repetition, it will not be described again here. Alternatively, when the computer program is executed by the processor, it implements the functions of each module / unit in this embodiment of the video analysis model generation apparatus, for example... Figure 9 The functions of the video acquisition module 901, teacher encoding module 902, decoding result determination module 903, loss function value determination module 904, model update module 905, and fine-tuning module 906 shown are not described again here to avoid repetition. The computer-readable storage medium may be non-volatile or volatile.

[0108] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When executed by a processor, the computer program implements the video analysis method described in the above embodiments, for example... Figure 6 As shown in S601-S603, or Figures 7 to 8 As shown, to avoid repetition, it will not be described again here. Alternatively, when the computer program is executed by the processor, it implements the functions of each module / unit in this embodiment of the video analysis device; to avoid repetition, this will not be described again here. The computer-readable storage medium can be non-volatile or volatile.

[0109] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. This computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory may include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory may include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in a variety of forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), RAMbus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0110] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0111] The above-described embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

Claims

1. A method for generating a video analysis model, characterized in that, include: Obtain distillation training videos and fine-tuning training videos; The distillation training video is processed by using N teacher coding models to extract features, and the teacher feature extraction result corresponding to each teacher coding model is determined, where N>1; The student coding model is used to extract features from the distillation training video, and the student feature extraction results are output. The decoding network corresponding to each teacher coding model is used to decode the student feature extraction results, and the decoding result corresponding to each teacher coding model is determined. Based on the teacher feature extraction and decoding results corresponding to N teacher coding models, the target loss function value is determined; Based on the target loss function value, the student coding model is updated to determine the initial video feature extraction model; Based on the fine-tuned training video, the initial video feature extraction model and the initial classification model are fine-tuned to generate a target video analysis model. The target video analysis model includes a target video feature extraction model and a target classification model connected in series with the target video feature extraction model.

2. The video analysis model generation method as described in claim 1, characterized in that, The teacher coding model includes a video teacher model and an image teacher model; The process involves using a student coding model to extract features from the distillation training video, outputting student feature extraction results, and then using a decoding network corresponding to each teacher coding model to decode the student feature extraction results, determining the decoding result corresponding to each teacher coding model, including: The distillation training video is vectorized and masked to determine the mask vector matrix; The student coding model is used to perform feature extraction on the mask vector matrix, and the student feature extraction results are output. The student feature extraction results are decoded using the decoding network corresponding to each video teacher model and the decoding network corresponding to each image teacher model to determine the first decoding result corresponding to each video teacher model and the second decoding result corresponding to each image teacher model.

3. The video analysis model generation method as described in claim 2, characterized in that, The step of using the decoding network corresponding to each video teacher model and the decoding network corresponding to each image teacher model to decode the student feature extraction results and determine the first decoding result corresponding to each video teacher model and the second decoding result corresponding to each image teacher model includes: The student feature extraction results are spliced ​​using the learnable mask matrix corresponding to each video teacher model to determine the first splicing matrix corresponding to each video teacher model. The student feature extraction results are spliced ​​using the learnable mask matrix corresponding to each image teacher model to determine the second splicing matrix corresponding to each image teacher model. The first concatenation matrix corresponding to each video teacher model is decoded using the decoding network corresponding to each video teacher model to determine the first decoding result corresponding to each video teacher model. The second concatenation matrix corresponding to each image teacher model is decoded using the decoding network corresponding to each image teacher model to determine the second decoding result corresponding to each image teacher model.

4. The video analysis model generation method according to any one of claims 1 to 3, characterized in that, The decoding network includes multiple vit-block modules and a multilayer perceptron module arranged in series.

5. The video analysis model generation method as described in claim 1, characterized in that, The determination of the target loss function value based on the teacher feature extraction and decoding results corresponding to N teacher coding models includes: Based on the teacher feature extraction and decoding results corresponding to N teacher coding models, determine the initial loss function value for each teacher coding model; The target loss function value is determined by weighting the initial loss function value corresponding to each teacher coding model with the target weights.

6. The video analysis model generation method as described in claim 1, characterized in that, The fine-tuning training videos include standard classification labels; The step of fine-tuning the initial video feature extraction model and the initial classification model based on the fine-tuned training video to generate the target video analysis model includes: The initial video feature extraction model is used to perform feature extraction processing on the fine-tuned training video to determine the feature matrix to be classified; The initial classification model is used to classify the feature matrix to be classified, and the initial classification label corresponding to the fine-tuned training video is determined. Based on the standard classification label and the initial classification label, determine the classification loss function value; Based on the classification loss function value, the initial video feature extraction model and the initial classification model are fine-tuned to obtain the target video feature extraction model and the target classification model.

7. A video analysis method, characterized in that, include: Obtain the video to be analyzed; The target video features are extracted from the video to be analyzed using a target video feature extraction model to determine the target features of the video; The target features of the video are classified using a target classification model to determine the target classification label corresponding to the video to be analyzed; The target video feature extraction model and the target classification model are models generated using the video analysis model generation method according to any one of claims 1-6.

8. The video analysis method as described in claim 7, characterized in that, The target video feature extraction model includes a first vit-block module, a DBMSE module, and a second vit-block module set in series. The step of using a target video feature extraction model to extract features from the video to be analyzed and determining the target features of the video includes: The first vit-blcok module is used to extract features from the video to be analyzed, and coarse features of the video are determined. The DBMSE module is used to extract coarse features from the video to determine the video spatial features and video motion features; The second vit-block module is used to integrate the video spatial features and the video motion features to determine the video target features.

9. The video analysis method as described in claim 8, characterized in that, The DBMSE module includes a vit-block submodule and a mamba-block submodule configured in parallel. The step of using the DBMSE module to extract coarse features from the video to determine video spatial features and video motion features includes: The video coarse features are classified according to a preset classification ratio to determine a first input feature and a second input feature, wherein the ratio of the first input feature is smaller than the ratio of the second input feature. The vit-block submodule is used to extract spatial features from the first input features to determine the video spatial features; The mamba-block submodule is used to extract motion features from the second input features to determine the video motion features.

10. A video analysis model generation device, characterized in that, include: The video acquisition module is used to acquire distillation training videos and fine-tuning training videos; The teacher coding module is used to perform feature extraction processing on the distillation training video using N teacher coding models respectively, and to determine the teacher feature extraction result corresponding to each teacher coding model, where N>1; The decoding result determination module is used to perform feature extraction processing on the distillation training video using a student coding model, output student feature extraction results, and decode the student feature extraction results using a decoding network corresponding to each teacher coding model to determine the decoding result corresponding to each teacher coding model. The loss function value determination module determines the target loss function value based on the teacher feature extraction and decoding results corresponding to the N teacher coding models. The model update module updates the student coding model based on the target loss function value to determine the initial video feature extraction model; The fine-tuning module, based on the fine-tuning training video, fine-tunes the initial video feature extraction model and the initial classification model to generate a target video analysis model. The target video analysis model includes a target video feature extraction model and a target classification model connected in series with the target video feature extraction model.

11. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the video analysis model generation method as described in any one of claims 1 to 6; or, when the processor executes the computer program, it implements the video analysis method as described in any one of claims 7 to 9.

12. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the video analysis model generation method as described in any one of claims 1 to 6, or when the computer program is executed by the processor, it implements the video analysis method as described in any one of claims 7 to 9.