Cross-scene action recognition method

By introducing ControlNet and dark to bright conditional diffusion model into the action recognition model, the bright information of low-light video is restored, and combined with the dual-channel action recognition backbone network of self-distillation branches, the problem of action recognition under low-light conditions is solved, significantly improving the recognition accuracy and network robustness.

CN120014706AActive Publication Date: 2025-05-16FUDAN UNIVERSITY
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510109511.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-01-23
Publication Date
2025-05-16
Estimated Expiration
2045-01-23

AI Technical Summary

Technical Problem

The existing action recognition models perform poorly under low-light conditions, are difficult to effectively extract action information, and lack action recognition backbone networks for low-light scenes.

Method used

Using a cross-scene action recognition method, dark images are trained using the ControlNet model to generate a preliminary restored light image, and the dark video frame is restored to continuous bright video frames through the dark to bright condition diffusion model. At the same time, a dual-channel action recognition backbone network based on self-distillation branches is designed to improve the network's recognition ability of dark videos.

Benefits of technology

It effectively solves the problem of low-visuality in low-light videos, enables the network to extract effective action information, improves the accuracy of action recognition, and enhances the robustness and generalization capabilities of the network.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120014706A_ABST
    Figure CN120014706A_ABST
Patent Text Reader

Abstract

The invention provides a cross-scene action recognition method, and belongs to the technical field of computer vision video understanding tasks. The cross-scene action recognition method comprises the following steps: firstly, training a dark-to-bright diffusion model, and converting an input dark video frame into a bright video frame in combination with priori knowledge obtained in a large-scale diffusion model pre-trained based on normal illumination data; in a sampling recovery process, a specific space-time attention mechanism is integrated into a trained conditional diffusion model, so that the discontinuity between video frames caused by a low-illumination enhancement method based on image training is relieved. And then designing a specific self-distillation branch and configuring the specific self-distillation branch into the action recognition backbone network, and extracting weighted spatial-temporal characteristics among layers of the backbone network so as to improve the generalization ability of the action recognition network. Compared with a mainstream method in the industry, the method has the advantages that the most advanced result is obtained on the basis of an existing dark video recognition data set, and meanwhile, the effect is greatly improved compared with a baseline result.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of computer vision video understanding tasks, and in particular relates to a cross-scene action recognition method. Background Art

[0002] Action recognition is a highly influential task in computer vision video understanding. Current research has made full use of advanced recognition network frameworks and data, such as convolutional neural network-based architectures. [1][2][3] , Transformer-based model [4][5] And large-scale action recognition dataset [6][7] Although these methods have achieved outstanding results, they are usually carried out under optimal lighting conditions, which limits the applicability of these models in real-world environments. The challenges of this task are: (1) the low visibility of dark videos makes it difficult for the network to extract effective action information; (2) there is a lack of action recognition backbone networks for this scenario. At the same time, due to the lack of large-scale low-light video training datasets, the effectiveness of existing models will be greatly reduced when identifying video content in low light conditions.

[0003] Previous research has focused on the recognition task of specific low-light scenes. [8][9]

[10]

[11] There are limitations. First, previous work over-relies on basic low-light image enhancement strategies, resulting in insufficient improvement in low-light video quality. Second, previous work often does not improve the action recognition backbone network for this task.

[0004] [1]Feichtenhofer,C.,Fan,H.,Malik,J.,&He,K..Slowfast networks for video recognition.In:CVPR(2019).

[0005] [2]Tran, D., Wang, H., Torresani, L., Ray, J., LeCun, Y., & Paluri, MA closer look at spatiotemporal convolutions for action recognition. In: CVPR (2018).

[0006] [3]Feichtenhofer, C.X3d:Expanding architectures for efficient videorecognition.In:CVPR(2020).

[0007] [4]Xing,Z.,Dai,Q.,Hu,H.,Chen,J.,Wu,Z.,&Jiang,Y.G.Svformer:Semi-supervised video transformer for action recognition.In:CVPR(2023).

[0008] [5]Liu,Z.,Ning,J.,Cao,Y.,Wei,Y.,Zhang,Z.,Lin,S.,&Hu,H.Video swintransformer.In:CVPR(2022).

[0009] [6]Goyal,R.,Ebrahimi Kahou,S.,Michalski,V.,Materzynska,J.,Westphal,S.,Kim,H.,...&Memisevic,R.The"something something"video database for learningand evaluating visual common sense.In:ICCV(2017).

[0010] [7]Kuehne,H.,Jhuang,H.,Garrote,E.,Poggio,T.,&Serre,T..HMDB:a largevideo database for human motion recognition.In:ICCV(2011).

[0011] [8]Hira,S.,Das,R.,Modi,A.,&Pakhomov,D.Delta Sampling R-BERT forlimited data and low-light action recognition.In:CVPR(2021).

[0012] [9]Chen,R.,Chen,J.,Liang,Z.,Gao,H.,&Lin,S..Darklight networks foraction recognition in the dark.In:CVPR(2021).

[0013]

[10] Tu, Z., Liu, Y., Zhang, Y., Mu, Q., & Yuan, J. DTCM: Joint Optimization of Dark Enhancement and Action Recognition in Videos. In: IEEE Transactions onImage Processing. (2023).

[0014]

[11] Xu, Y., Yang, J., Cao, H., Mao, K., Yin, J., & See, S. Arid: A new dataset for recognizing action in the dark. In: Deep Learning for Human Activity Recognition: Second International Workshop (2021). Summary of the invention

[0015] The present invention is made to solve the above-mentioned problem, and its purpose is to provide a cross-scenario action recognition method.

[0016] The present invention provides a cross-scene action recognition method with the following characteristics, which is used to improve the network's recognition ability for dark videos, and includes the following steps: S10, using a dark image as a condition and a corresponding real light image as a target, and training the learnable part of a ControlNet model pre-trained on a large-scale normal light data set; S20, sampling the ControlNet model trained in step S10 to obtain a dark image and its corresponding preliminary restored light image; S30, splicing the dark image, the corresponding real light image, and the preliminary restored light image in the channel dimension as input, and using a multi-head attention mechanism to train a dark-to-light conditional diffusion model; S40, using the multi-head attention mechanism Replace it with the spatiotemporal attention mechanism, and use the dark-to-light conditional diffusion model to sample the video frames of the dark video to generate continuous light video frames; S50, based on the dual-channel action recognition backbone network structure composed of a dark input channel and a light input channel, construct N self-distillation branches between its residual modules, and each self-distillation branch includes a spatiotemporal fusion module and an alignment module; S60, input the video frames of the dark video and the corresponding light video frames into the dark input channel and the light input channel respectively, after the spatiotemporal fusion module and the alignment module are used to extract features and calculate losses, they are merged along the channel dimension, and after merging, the output features are obtained through the self-attention module, and the output features are output from the fully connected layer to the corresponding logit of the shallow residual module.

[0017] The cross-scenario action recognition method provided by the present invention may also have the following feature: wherein, in step S30, the training loss function of the dark-to-light conditional diffusion model is: represents the weighted average of the loss under the entire data distribution, t represents the time step of gradual denoising in the diffusion model, θ represents the parameter optimized during model training, ∈ t represents the Gaussian noise added to the image by the diffusion model at time step t, ∈ θ represents the noise predicted by the diffusion model at time step t, represents the accumulated noise attenuation factor, x0 represents the original image data, which is the real illumination image without adding noise, d Denotes a dark image, X b represents the real lighting image, P(X d ,X b ) represents the preliminary restored illumination image.

[0018] The cross-scene action recognition method provided by the present invention may also have the following features: wherein, in step S40, the spatiotemporal attention mechanism is based on a multi-head attention mechanism, and is used to restore the video frame of the dark video input into the dark-to-light conditional diffusion model. For the currently restored frame j, the query is the feature of the currently restored frame j, and the value and key are reconstructed with the features of the first frame and the previous frame j-1 of the video frame of the input dark video. The calculation method of the spatiotemporal attention mechanism is: z represents the feature representation after the spatiotemporal attention mechanism, Q = Q j represents the query matrix of the j-th frame feature input, Τ represents the matrix transpose, K Τ represents the transpose of the reconstructed bond matrix, The key matrix under the spatiotemporal attention mechanism is concatenated by the key matrix of the first frame and the key matrix of the previous frame in the channel dimension. K0 represents the key matrix of the first frame, and K j-1 represents the key matrix of the previous frame, represents the scaling factor of the attention mechanism, The value matrix under the spatiotemporal attention mechanism is concatenated by the value matrix of the first frame and the value matrix of the previous frame in the channel dimension. V0 represents the value matrix of the first frame, V j-1 represents the value matrix of the previous frame, Represents a merge operation.

[0019] The cross-scenario action recognition method provided by the present invention may also have the following feature: in step S40, the preliminary restored illumination image is used as an additional control condition of the dark-to-light conditional diffusion model, thereby enhancing the authenticity and effectiveness of illumination restoration.

[0020] The cross-scenario action recognition method provided by the present invention may also have the following features: wherein, in steps S50 to S60, the spatiotemporal fusion module includes a separable 3D convolutional layer and a series of upsampling layers, the separable 3D convolutional layer is used to capture spatiotemporal information from the shallower residual module in the dual-channel action recognition backbone network structure, and after the upsampling operation, the dot product with the input feature is performed, and the extracted features are weighted according to their correlation.

[0021] The cross-scenario action recognition method provided by the present invention may also have the following feature: in steps S50 to S60, the alignment module is composed of a series of separable 3D convolutional layers to ensure that the feature size passed by the spatiotemporal fusion module matches the reference feature size, thereby calculating the loss between them.

[0022] The cross-scenario action recognition method provided by the present invention may also have the following features: wherein, in step S60, the dual-channel action recognition backbone network structure is used as a neural network model, and its loss function is: represents the classification loss, represents the loss for the label, It represents the loss for the feature, and α and λ are weight coefficients.

[0023] The cross-scenario action recognition method provided by the present invention may also have the following characteristics: wherein the classification loss Loss for labels Feature-specific loss and Together they make up the self-distillation loss, represents the cross entropy loss, h i represents the output logit of the i-th self-distillation branch, y represents the corresponding label, and h N+1 Indicates the final output logit of the neural network model, represents the Kullback-Leibler divergence loss, represents the L2 loss, F i represents the predicted features output by the i-th self-distillation branch, F N+1 Represents the predicted features of the final output of the neural network model.

[0024] Functions and Effects of the Invention

[0025] The present invention utilizes a diffusion model to restore dark video frames into continuous bright video frames. The restoration results are close to videos under real natural lighting, which solves the low visibility problem of dark videos and enables the subsequent network to extract effective action information.

[0026] The present invention adopts a specific spatiotemporal attention mechanism in the sampling stage of the dark-to-bright diffusion model, which can effectively utilize the temporal information of a continuous video frame being sampled, thereby ensuring the spatiotemporal consistency of the restored bright video frame. In addition, the spatiotemporal attention mechanism can be seamlessly inserted into the diffusion model trained based on image data, and continuous bright video frames with excellent restoration effects can be obtained without additional model training.

[0027] The dual-channel action recognition backbone network based on self-distillation branches designed in the present invention can effectively extract and learn the action features in the video sequence, and make the network more robust and generalizable through the distillation process, thereby enhancing the interaction between dark and light video information and improving the recognition accuracy of the network. BRIEF DESCRIPTION OF THE DRAWINGS

[0028] Figure 1 is a flow chart of a cross-scenario action recognition method according to an embodiment of the present invention;

[0029] Figure 2 is a schematic diagram of a spatiotemporal attention mechanism of an embodiment of the present invention;

[0030] Figure 3 Schematic diagram of the sampling process of the dark-to-light conditional diffusion model of an embodiment of the present invention;

[0031] Figure 4 It is a comparison diagram of the preliminary restored image generated by the fine-tuned ControlNet in the test example of the present invention and the final output result of the dark-to-bright conditional diffusion model;

[0032] Figure 5 It is a result diagram of the dark-to-light conditional diffusion model in the test example of the present invention on the dark action recognition dataset ARID and the dark scene recognition dataset Dark48 dataset. DETAILED DESCRIPTION

[0033] In order to make the technical means, creative features, objectives and effects achieved by the present invention easy to understand, the following embodiments and accompanying drawings specifically illustrate a cross-scenario action recognition method of the present invention.

[0034] <Example>

[0035] Figure 1 It is a flowchart of a cross-scenario action recognition method according to an embodiment of the present invention.

[0036] like Figure 1 As shown, this embodiment provides a cross-scenario action recognition method for improving the network's action recognition capability in dark videos, including the following steps:

[0037] S10, using the low-light image to be restored (dark image Xd ) as a condition, the corresponding real lighting image X b As a goal, we train the learnable parts of the ControlNet model pre-trained on a large-scale normal lighting dataset.

[0038] S20, sampling the ControlNet model trained in step S10 to obtain a dark image and its corresponding preliminary restored illumination image P(X d ,X b ).

[0039] S30, dark image X d , the corresponding real lighting image X b And the initial restored illumination image P(X d ,X b ) are concatenated in the channel dimension as input, and the dark-to-light conditional diffusion model is trained using a multi-head attention mechanism.

[0040] Since the existing annotated dark action recognition video dataset lacks its corresponding bright video, it is impossible to directly build a dark-to-bright conditional diffusion model based on the dataset. To overcome this problem, this embodiment uses the image dataset ExLPose for the task of human posture estimation under low light conditions.

[13] As training data for the dark-to-light conditional diffusion model.

[0041] The training loss function of the dark-to-light conditional diffusion model is:

[0042]

[0043] in, represents the weighted average of the loss under the entire data distribution, t represents the time step of gradual denoising in the diffusion model, θ represents the parameter optimized during model training, ∈ t represents the Gaussian noise added to the image by the diffusion model at time step t, ∈ θ represents the noise predicted by the diffusion model at time step t, represents the accumulated noise attenuation factor, x0 represents the original image data, which is the real illumination image without adding noise, d Denotes a dark image, X b represents the real lighting image, P(X d ,X b ) represents the preliminary restored illumination image.

[0044] S40, replacing the multi-head attention mechanism with the spatiotemporal attention mechanism, and using the dark-to-light conditional diffusion model to sample the video frames of the dark video, including the following sub-steps S41 to S42:

[0045] S41: Replace the multi-head attention mechanism with a spatiotemporal attention mechanism, which is seamlessly inserted into the trained dark-to-light conditional diffusion model. This mechanism does not require additional training to address the discontinuity between adjacent video frames caused by direct sampling of the image-based diffusion model.

[0046] Figure 2 Schematic diagram of the spatiotemporal attention mechanism of an embodiment of the present invention. Figure 2 As shown in the figure, the spatiotemporal attention mechanism is based on the multi-head attention mechanism and is used to restore the video frames of the dark video input into the dark-to-light conditional diffusion model. The specific operations are as follows:

[0047] A continuous video frame of length k is input and sampled. In the spatiotemporal attention mechanism, for the currently restored frame j, the query is the feature of the current frame j. The features of the first frame of the input dark video and the previous frame j-1 are concatenated to reconstruct the value and key, and then the spatiotemporal features of the current frame j are calculated through the multi-head attention calculation mechanism.

[0048] The calculation method of the spatiotemporal attention mechanism is:

[0049]

[0050] Among them, z represents the feature representation after the spatiotemporal attention mechanism, Q = Q j represents the query matrix of the j-th frame feature input, Τ represents the matrix transpose, K Τ represents the transpose of the reconstructed bond matrix, The key matrix under the spatiotemporal attention mechanism is concatenated by the key matrix of the first frame and the key matrix of the previous frame in the channel dimension. K0 represents the key matrix of the first frame, and K j-1 represents the key matrix of the previous frame, represents the scaling factor of the attention mechanism, The value matrix under the spatiotemporal attention mechanism is concatenated by the value matrix of the first frame and the value matrix of the previous frame in the channel dimension. V0 represents the value matrix of the first frame, V j-1 represents the value matrix of the previous frame, Represents a merge operation.

[0051] S42, using the preliminary restored illumination image P(X d ,X b ) and noise are spliced ​​in the dimension as additional control conditions for the input dark-to-light conditional diffusion model, and large-scale light training data is added as prior knowledge to the dark-to-light conditional diffusion model to enhance the authenticity and effectiveness of light restoration. Following the spatiotemporal attention mechanism in the sampling process, the video frames of the dark video are sampled from the dark-to-light conditional diffusion model to obtain the restored continuous light video frames Xe ,like Figure 3 shown.

[0052] In this step, the spatiotemporal attention mechanism samples a continuous segment of video frames and provides the information of the initial frame and the previous frame for the current frame to maintain the continuity of the bright video frames generated by the sampling and avoid the randomness and discontinuity of each frame in the video caused by operating a single frame.

[0053] S50, based on the dual-channel action recognition backbone network structure consisting of a dark input channel and a light input channel, constructs N self-distillation branches between its residual modules, each of which includes a spatiotemporal fusion module and an alignment module.

[0054] The spatiotemporal fusion module consists of a separable 3D convolutional layer and a series of upsampling layers. The separable 3D convolutional layer captures spatiotemporal information from the shallower residual modules in the backbone network, and after the upsampling operation, the dot product with the input features is performed to weight the extracted features according to their relevance.

[0055] The alignment module consists of a series of separable 3D convolutional layers to ensure that the feature size passing through the spatiotemporal fusion module matches the reference feature size so as to calculate the loss between them.

[0056] In this step, the dual-channel action recognition backbone network structure is used as a neural network model, and its loss function Including classification loss and self-distillation losses.

[0057] Classification Loss It consists of a cross-entropy loss function, including the loss between the final output of the action recognition network and the corresponding label, and the loss between the output of each distillation branch and the corresponding label:

[0058]

[0059] in, represents the cross entropy loss, h i represents the output logit of the i-th self-distillation branch, y represents the corresponding label, and h N+1 Represents the final output logit of the neural network model.

[0060] The self-distillation loss includes the loss for the label and the loss for the feature

[0061] Loss for labels The Kullback-Leibler divergence loss is calculated between the logit output of each distillation branch and the final output logit of the action recognition network:

[0062]

[0063] in, represents the Kullback-Leibler divergence loss.

[0064] Feature-specific loss The L2 loss is calculated between the predicted features output by each distillation branch and the predicted features output by the action recognition network:

[0065]

[0066] in, represents the L2 loss, F i represents the predicted features output by the i-th self-distillation branch, F N+1 Represents the predicted features of the final output of the neural network model.

[0067] Finally, the loss function is obtained

[0068]

[0069] Among them, α is the weight coefficient of Kullback-Leibler divergence loss, and λ is The weight coefficient of .

[0070] S60, the video frames of the dark video and the corresponding light video frames are input into the dark input channel and the light input channel respectively, and after passing through the spatiotemporal fusion module and the alignment module to extract features and calculate losses, they are merged along the channel dimension. After merging, the output features are obtained through the self-attention module, and the output features are output from the fully connected layer to the corresponding logit of the shallow residual module.

[0071] Then, the Top-1 accuracy of the final output action recognition is used as a quantitative evaluation indicator of the cross-scenario action recognition method of this embodiment.

[0072] <Test example>

[0073] This test example uses the cross-scenario action recognition method in the embodiment to perform corresponding tests.

[0074] In this test case, the dark-to-light conditional diffusion model accepts a 256×256 fixed-size image as input, which is composed of the original dark image X d 、ControlNet initially restored the illumination image P(X d,X b ) and the corresponding real lighting image X b Spliced ​​in the channel dimension.

[0075] When fine-tuning and training ControlNet, pairs of dark and light images are used as input to ControlNet, where the dark image X d As a condition, the bright image X b As a generation target, it is used to train the adaptable part of ControlNet to obtain a preliminary restored illumination image P(X d ,X b ). During fine-tuning, in order to avoid adding irrelevant information to the diffusion process, “A detailed high-quality professional image” was used as the default prompt. The number of diffusion steps in the dark-to-light conditional diffusion model training process was set to 1000 steps, and 47,000 iterations of training were performed.

[0076] In the sampling process of the dark-to-light conditional diffusion model, the attention mechanism in UNet is replaced with the spatiotemporal attention mechanism specified in the embodiment, DDIM sampling is adopted and the diffusion step number is set to 100. A segment of 16 dark video frames is used to generate a segment of continuous bright video frames as the input of the subsequent dual-channel action recognition backbone network structure.

[0077] Figure 4 It is a comparison diagram of the preliminary restored image generated by the fine-tuned ControlNet in the test example of the present invention and the final output result of the dark-to-bright conditional diffusion model.

[0078] like Figure 4 As shown, the dark-to-light conditional diffusion model is significantly better than the initially restored video frames in terms of detail and color restoration of the video frames.

[0079] Figure 5 It is a result diagram of the dark-to-light conditional diffusion model in the test example of the present invention on the dark action recognition dataset ARID and the dark scene recognition dataset Dark48 dataset.

[0080] like Figure 5 As shown in the figure, for the input dark video frame, the dark-to-light conditional diffusion model can effectively restore the color of the video frame with the assistance of a large amount of prior knowledge of illumination. At the same time, the spatiotemporal attention mechanism used in the sampling process can effectively maintain the continuity between the restored video frames.

[0081] Based on the self-distillation branch and dual-channel action recognition backbone network structure, video clips of fixed size 112×112×3×64 are accepted as input. Both dark and light videos are normalized to the above size by scaling, and multi-scale cropping, random horizontal flipping, and Cutout operations are used for video data enhancement. The dual-channel action recognition backbone network structure adopts the R(2+1)D-34 structure, and the dark video and the corresponding restored light video are input into the weight-sharing network. Three self-distillation branches are constructed between the residual modules, each of which consists of a spatiotemporal fusion module and an alignment module. The spatiotemporal fusion module consists of a 3D separable convolutional layer, a 3D batch normalization layer, and a ReLU activation function layer. At the same time, the extracted features are upsampled and dot-producted with the input features to calculate their correlation and then weighted. The alignment module uses a series of 3D separable convolutional layers to ensure that the feature size of the spatiotemporal fusion module matches the reference feature size, so as to calculate the loss between them. After passing through the spatiotemporal fusion module and the alignment module, the dark information and light information of the two channels will be merged along the channel dimension. The merged features will pass through the self-attention module to obtain the output features. The output features will be output from the fully connected layer to the corresponding logit of the shallow residual module.

[0082] For action recognition in dark scenes, recognition accuracy is used as the evaluation indicator.

[0083] The specific experimental results of the cross-scenario action recognition method in this test case on the ARID dataset and the Dark48 dataset are shown in Tables 1 and 2 respectively.

[0084] Table 1 (Specific experimental results on the ARID dataset)

[0085]

[0086] Table 2 (Specific experimental results on the Dark48 dataset)

[0087] method <![CDATA[DTCM

[10] ]]> <![CDATA[DarkLight-ResNext101 [9] ]]> <![CDATA[DarkLight-R(2+1)D [9] ]]> The method of this embodiment Top-1 Accuracy 46.68% 42.27% 39.08% 47.14% Top-5 Accuracy 75.92% 70.47% 71.24% 75.65%

[0088] As shown in Tables 1 and 2, the cross-scene action recognition method proposed in the embodiment has achieved state-of-the-art results on the existing dark video recognition dataset, and has a significant improvement over the baseline results.

[0089] Functions and Effects of the Embodiments

[0090] This embodiment uses the conditional diffusion model to enhance the video frames with insufficient light, and combines it with the ControlNet model

[12] The prior knowledge obtained in the image is integrated with normal lighting information, and the diffusion model is used to restore dark video frames into continuous bright video frames. The restored results are close to the videos under real natural lighting, which solves the low visibility problem of dark videos and enables the subsequent network to extract effective action information.

[0091] The specific spatiotemporal attention mechanism used in the dark-to-bright diffusion model sampling stage of this embodiment can effectively utilize the temporal information of a continuous video frame being sampled, alleviate the discontinuity between adjacent video frames caused by direct sampling of the image-based diffusion model, and thus ensure the spatiotemporal consistency of the restored bright video frame. In addition, the spatiotemporal attention mechanism can be seamlessly inserted into the diffusion model trained based on image data, and continuous bright video frames with excellent restoration effects can be obtained without additional model training.

[0092] The dual-channel action recognition backbone network based on self-distillation branches designed in this embodiment can effectively extract and learn the action features in the video sequence, and make the network more robust and generalized through the distillation process, thereby enhancing the interaction between dark and light video information and improving the recognition accuracy of the network.

[0093] This embodiment integrates the self-distillation branch with the dual-channel action recognition backbone network to extract weighted spatiotemporal features from different blocks of the backbone network, so as to encourage the self-distillation branch to focus on more important information in dark videos, thereby promoting information interaction between different blocks in the action recognition backbone architecture and effectively improving the network's recognition ability for dark videos.

[0094] The test case uses the expertise gained from the ExLPose dataset to enhance the performance of the dark-to-light diffusion model for dark scene action recognition tasks.

[0095] The test example uses the Top-1 accuracy of the final output action recognition of the model as a quantitative evaluation indicator of the cross-scene action recognition method, and compares the cross-scene action recognition method of this embodiment with other methods, verifying the advantages of this embodiment. In addition, the test example also visualizes the restored bright video frames and conducts visual qualitative evaluation, verifying that the restoration results of the model in the test example are close to the video under real natural lighting, while maintaining the continuity between adjacent video frames.

[0096]

[12] Zhang, L., Rao, A., & Agrawala, M. Adding conditional control to text-to-image diffusion models. In: ICCV (2023).

[0097]

[13] Lee, S., Rim, J., Jeong, B., Kim, G., Woo, B., Lee, H.,... & Kwak, S. Human poseestimation in extremely low-light conditions. In: CVPR (2023).

[0098] Those skilled in the art should understand that the present invention is not limited to the above embodiments, and the above embodiments and descriptions are only for explaining the principles of the present invention. Without departing from the spirit and scope of the present invention, the present invention may have various changes and improvements, and these changes and improvements fall within the scope of the present invention to be protected. The scope of protection of the present invention is defined by the attached claims and their equivalents.

Claims

1. A cross-scenario action recognition method, characterized in that: It is used to improve the network's ability to recognize dark videos, including the following steps: S10, using the dark image as the condition and the corresponding real-light image as the target, trains the learnable part of the ControlNet model pre-trained on a large-scale normal-light dataset; S20, sampling the ControlNet model trained in step S10 to obtain the dark image and its corresponding preliminary restored light image; S30, concatenating the dark image, the corresponding real light image, and the preliminary restored light image in the channel dimension as input, and using a multi-head attention mechanism to train to obtain a dark-to-light conditional diffusion model; S40, replacing the multi-head attention mechanism with a spatiotemporal attention mechanism, and using the dark-to-light conditional diffusion model to sample video frames of the dark video to generate continuous light video frames; S50, based on the dual-channel action recognition backbone network structure consisting of a dark input channel and a bright input channel, N self-distillation branches are constructed between its residual modules, each of which includes a spatiotemporal fusion module and an alignment module; S60, input the video frame of the dark video and the corresponding light video frame into the dark input channel and the light input channel respectively, and after passing through the spatiotemporal fusion module and the alignment module to extract features and calculate losses, merge them along the channel dimension, and after merging, obtain output features through the self-attention module, and the output features are output from the fully connected layer to the corresponding logit of the residual module in the shallow layer.

2. The cross-scenario action recognition method according to claim 1, characterized in that: in, In step S30, the training loss function of the dark-to-light conditional diffusion model is: represents the weighted average of the loss under the entire data distribution, t represents the time step of gradual denoising in the diffusion model, θ represents the optimized parameters during model training, ∈ t represents the Gaussian noise added to the image by the diffusion model at time step t, ∈ θ represents the noise predicted by the diffusion model at time step t, represents the accumulated noise attenuation factor, x0 represents the original image data, and the original image data is the real illumination image without adding noise, X d represents the dark image, X b represents the real illumination image, P(X d ,X b ) represents the preliminary restored illumination image.

3. The cross-scenario action recognition method according to claim 1, characterized in that: in, In step S40, the spatiotemporal attention mechanism is based on the multi-head attention mechanism to restore the video frame of the dark video input into the dark-to-light conditional diffusion model. For the currently restored frame j, query is the feature of the currently restored frame j, and the value and key are reconstructed by inputting the features of the first frame and the previous frame j-1 of the video frame of the dark video. The calculation method of the spatiotemporal attention mechanism is: z represents the feature representation after the spatiotemporal attention mechanism, Q = Q j represents the query matrix of the j-th frame feature input, Τ represents the matrix transpose, K Τ represents the transpose of the reconstructed bond matrix, The key matrix under the spatiotemporal attention mechanism is concatenated by the key matrix of the first frame and the key matrix of the previous frame in the channel dimension. K0 represents the key matrix of the first frame, and K j-1 represents the key matrix of the previous frame, represents the scaling factor of the attention mechanism, The value matrix under the spatiotemporal attention mechanism is concatenated by the value matrix of the first frame and the value matrix of the previous frame in the channel dimension. V0 represents the value matrix of the first frame, V j-1 represents the value matrix of the previous frame, Represents a merge operation.

4. The cross-scenario action recognition method according to claim 1, characterized in that: in, In step S40, the preliminary restored illumination image is also used as an additional control condition of the dark-to-light conditional diffusion model, thereby enhancing the authenticity and effectiveness of illumination restoration.

5. The cross-scenario action recognition method according to claim 1, characterized in that: in, In step S50 to step S60, the spatiotemporal fusion module includes a separable 3D convolutional layer and a series of upsampling layers. The separable 3D convolutional layer is used to capture spatiotemporal information from the shallower residual module in the dual-channel action recognition backbone network structure. After the upsampling operation, the dot product with the input feature is performed, and the extracted features are weighted according to their correlation.

6. The cross-scenario action recognition method according to claim 5, characterized in that: in, In step S50 to step S60, the alignment module is composed of a series of separable 3D convolutional layers to ensure that the feature size passed through the spatiotemporal fusion module matches the reference feature size, thereby calculating the loss between them.

7. The cross-scenario action recognition method according to claim 1, characterized in that: in, In step S60, the dual-channel action recognition backbone network structure is used as a neural network model, and its loss function is: represents the classification loss, represents the loss for the label, It represents the loss for the feature, and α and λ are weight coefficients.

8. The cross-scenario action recognition method according to claim 7, characterized in that: in, Classification Loss Loss for labels Feature-specific loss and Together they make up the self-distillation loss, represents the cross entropy loss, h i represents the output logit of the i-th self-distillation branch, y represents the corresponding label, and h N+1 Indicates the final output logit of the neural network model, represents the Kullback-Leibler divergence loss, represents L2 loss, F i represents the predicted features output by the i-th self-distillation branch, F N+1 Represents the predicted features of the final output of the neural network model.

Citation Information

Patent Citations

  • Video question and answer method in low-light scene

    CN117095336A

  • Dark scene-oriented end-to-end multi-task action recognition method and system

    CN117315774A

  • Multi-level cross-scene hyperspectral image classification method

    CN118334420A

  • Motion deblurring method and device based on hybrid visual sensor

    CN119090770A

  • KR20240146429A