Action recognition methods, systems, devices, and storage media based on multi-model aggregation

By optimizing the path generation network and source domain model through a multi-model aggregation method, the problem of low efficiency in multi-source model aggregation is solved, thereby improving the accuracy and efficiency of video action recognition.

CN119360434BActive Publication Date: 2026-01-30BEIJING UNIV OF POSTS & TELECOMM +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202411255283.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-09
Publication Date
2026-01-30
Estimated Expiration
2044-09-09

AI Technical Summary

Technical Problem

Existing unsupervised video domain adaptation methods are inefficient when aggregating multiple source models and fail to effectively utilize the knowledge provided by source domain models with different structures, ignoring instance-level differences and resulting in insufficient accuracy in video action recognition.

Method used

By employing a multi-model aggregation method, the path weights and aggregated action recognition results are calculated using an initial path generation network and an optimized source domain model. The model parameters are then optimized by combining instance-level transferability estimation metrics and loss functions, thereby improving the adaptability of the source domain model in the target domain.

Benefits of technology

It improves the accuracy of action recognition in the target domain scene and enhances the efficiency and accuracy of multi-model aggregation.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119360434B_ABST
    Figure CN119360434B_ABST
Patent Text Reader

Abstract

This invention relates to the field of computer vision technology, specifically disclosing a method, system, device, and storage medium for action recognition based on multi-model aggregation. The method includes: acquiring multiple target path weights and aggregated action recognition results corresponding to each sample video in the target domain; calculating a loss value and iteratively optimizing it based on the action recognition label, aggregated action recognition results, quantized value of the instance-level transferability estimation index, and multiple target path weights corresponding to each sample video; inputting the video to be tested into a trained path generation network to obtain multiple target path weights corresponding to the video to be tested, and obtaining the aggregated action recognition result of the video to be tested based on the multiple target path weights corresponding to the video to be tested and the corresponding trained source domain model. The method of this invention improves the adaptability of the source domain model to the target domain scene, thereby enhancing the accuracy of action recognition in the target domain scene.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of computer vision technology, and in particular to an action recognition method, system, device and storage medium based on multi-model aggregation. Background Technology

[0002] Video action recognition, a crucial task in video understanding, aims to identify the types of actions contained within an input video. However, constructing large-scale datasets of real-world scenarios incurs extremely high human and economic costs. Unsupervised video domain adaptation has been proposed as a solution. This method feeds labeled source domain data and unlabeled target domain data into the model. By aligning the source and target domain data, knowledge is transferred from the labeled source domain to the unlabeled target domain to overcome domain shift. While traditional unsupervised video domain adaptation methods effectively mitigate domain shift between video domains, the need to access the source data during the adaptation process introduces significant privacy risks and substantial transmission costs.

[0003] Existing benchmarks and tests for passive-domain video domain transfer are based on a single source model. These traditional benchmark configurations aim for fair ablation comparisons, neglecting the fact that in real-world scenarios, there are many source models with different architectures available, providing rich source domain knowledge. In reality, multiple source models are often available, and different source models can provide richer source domain knowledge, significantly improving adaptation results. It remains unclear how to extend single-model passive-domain video domain transfer to multiple pre-trained models. Simply aggregating all source domain pre-trained models is undoubtedly inefficient, as not all source domain models exhibit good transferability in the target domain.

[0004] Existing multi-model aggregation algorithms for passive domain transfer in the image domain assume that the source domain models have a uniform structure, such as all being ResNet models. This limits the knowledge that the model can learn in the source domain. It does not consider that there are various different model architectures for the source domain task, such as ResNet50 and ResNet101 based on CNN architecture, and Vision-Transformer and Swin-Transformer based on Transformer architecture.

[0005] Furthermore, videos possess more temporal information than images, adding an extra temporal dimension. This makes methods previously used for images inapplicable to the video domain. Specific designs are needed to address the temporal information inherent in videos. Moreover, images and videos differ more significantly at the instance level. Multi-model aggregation algorithms for passive domain transfer in the image domain only consider the domain level when assigning different weights to different models. This ignores the significant differences between different videos at the instance level, leading to varying model preferences for each video instance. Summary of the Invention

[0006] To address the aforementioned technical problems, this invention provides an action recognition method, system, device, and storage medium based on multi-model aggregation.

[0007] Firstly, the present invention provides an action recognition method based on multi-model aggregation, the technical solution of which is as follows:

[0008] S1. Input any sample video from the target domain into the initial path generation network to obtain N path weights corresponding to any sample video, and determine all path weights except the minimum path weight as target path weights; where N≥3 and N is a positive integer, and each path weight corresponds to a filtered source domain model used for action recognition.

[0009] S2. Input each sample video into the filtered source domain model corresponding to each target path weight to obtain multiple first action recognition results corresponding to each sample video, and based on each first action recognition result corresponding to each sample video and the corresponding target path weight, obtain the aggregated action recognition result corresponding to each sample video.

[0010] S3. Repeat S1-S2 to obtain multiple target path weights and aggregated action recognition results corresponding to each sample video in the target domain;

[0011] S4. Based on the action recognition label, aggregated action recognition result, quantized value of instance-level transferability estimation index, and multiple target path weights corresponding to each sample video, calculate and optimize the parameters of the filtered source domain model according to the current loss value, to obtain the optimized source domain model. Based on the quantized value of instance-level transferability estimation index corresponding to each sample video, optimize the parameters of the initial path generation network to obtain the optimized path generation network.

[0012] S5. Use the optimized path generation network as the initial path generation network and the optimized source domain model as the filtered source domain model, and return to execute S3 until the iterative optimization conditions are met. Then, determine the initial path generation network as the trained path generation network and the filtered source domain model as the trained source domain model.

[0013] S6. Input the video to be tested into the trained path generation network to obtain the target path weights corresponding to the video to be tested. Then, input the video to be tested into the trained source domain model corresponding to each target path weight to obtain multiple first action recognition results corresponding to the video to be tested. Based on each first action recognition result corresponding to the video to be tested and the corresponding target path weights, obtain the aggregate action recognition result corresponding to the video to be tested.

[0014] The beneficial effects of the action recognition method based on multi-model aggregation of the present invention are as follows:

[0015] The method of this invention improves the accuracy of action recognition in the target domain scene by enhancing the adaptability of the source domain model to the target domain scene.

[0016] Based on the above scheme, the action recognition method based on multi-model aggregation of the present invention can be further improved as follows.

[0017] In one alternative approach, it also includes:

[0018] Determine the quantization value of the transferability between each source domain model and the target domain, and sort them in descending order according to the size of the quantization value to obtain the target queue;

[0019] The source domain models corresponding to the first N quantization values ​​of the target queue are determined as the filtered source domain models.

[0020] In one alternative approach, the step of determining a quantized value of the transferability between any source domain model and the target domain includes:

[0021] Based on the first objective formula, a quantized value of the transferability between any source domain model and the target domain is determined; wherein, the first objective formula is: IC am =E(IC) i )×SUTE×TC÷E(IC i ); IC i =H(h(V) i )), IC am V represents the quantized value of the transferability between any source domain model and the target domain. iLet h(V) represent the i-th sample video of the target domain. i ) represents V i The output result obtained by inputting the data into any of the source domain models, H(h(V) i )) represents the calculation of h(V) i The entropy of ), SUTE represents the quantized value of the source-free domain transferability estimation index, TC represents the quantized value of time consistency, and IC i E(IC) represents the quantized value of the determinism of an instance of any source domain model. i () indicates the calculation of IC with different values i The mathematical expectation, Indicates V i The video after time perturbation processing, D T This represents the sample set of the target domain. This represents the calculation of H(h(V) with different values. i The average value of )) This indicates the calculation of h(V) with different values. i The average value of ) Represents the predicted semantics of the sample video. This represents the pseudo-label of the sample video. Indicates calculation KL divergence, h c (V i ) represents V i Predicted probability for class c.

[0022] In one alternative approach, the step of obtaining a quantified value of the instance-level transferability estimation index corresponding to any sample video includes:

[0023] Based on the second objective formula, the quantized value of the instance-level transferability estimation index corresponding to any sample video is calculated; wherein, the second objective formula is: Γ i This represents the quantized value of the instance-level transferability estimation index corresponding to the i-th sample video. Indicates to Perform softmax normalization.

[0024] In one alternative approach, the step of calculating the current loss value based on the action recognition label corresponding to each sample video, the aggregated action recognition result, the quantized value of the instance-level transferability estimation index, and the weights of multiple target paths includes:

[0025] The current loss value is calculated by substituting the action recognition label, aggregated action recognition result, quantized value of instance-level transferability estimation index, and multiple target path weights corresponding to each sample video into the target loss function; wherein, the target loss function is: L=θ1Lcor +L SHTC , L represents the current loss value, θ1 is a hyperparameter, and L SHTC L represents the original loss obtained based on the action recognition label corresponding to each sample video and the aggregated action recognition result. cor G1(V) represents the weighted loss, where n represents the number of sample videos. i ) represents the weights of multiple target paths in the i-th sample video.

[0026] In one alternative approach, the iterative optimization condition is: the maximum number of iterations or the convergence of the loss function.

[0027] Secondly, this invention provides an action recognition system based on multi-model aggregation, the technical solution of which is as follows:

[0028] The system comprises a first processing module, a second processing module, a third processing module, a fourth processing module, an iterative training module, and a detection module;

[0029] The first processing module is used to: input any sample video of the target domain into the initial path generation network, obtain N path weights corresponding to any sample video, and determine all path weights except the minimum path weight as target path weights; wherein, N≥3 and N is a positive integer, and each path weight corresponds to a filtered source domain model for action recognition;

[0030] The second processing module is used to: input the any sample video into the filtered source domain model corresponding to each target path weight, obtain multiple first action recognition results corresponding to the any sample video, and obtain the aggregated action recognition result corresponding to the any sample video based on each first action recognition result corresponding to the any sample video and the corresponding target path weight;

[0031] The third processing module is used to: repeatedly call the first processing module to the second processing module to obtain multiple target path weights and aggregated action recognition results corresponding to each sample video in the target domain;

[0032] The fourth processing module is used to: calculate and optimize the parameters of the filtered source domain model based on the action recognition label corresponding to each sample video, the aggregated action recognition result, the quantized value of the instance-level transferability estimation index, and the weights of multiple target paths, to obtain an optimized source domain model; and optimize the parameters of the initial path generation network based on the quantized value of the instance-level transferability estimation index corresponding to each sample video to obtain an optimized path generation network.

[0033] The iterative training module is used to: use the optimized path generation network as the initial path generation network and the optimized source domain model as the filtered source domain model, and return to call the third processing module until the iterative optimization conditions are met, and then determine the initial path generation network as the trained path generation network and the filtered source domain model as the trained source domain model.

[0034] The detection module is used to: input the video to be tested into the trained path generation network to obtain the target path weights corresponding to the video to be tested, and input the video to be tested into the trained source domain model corresponding to each target path weight to obtain multiple first action recognition results corresponding to the video to be tested, and based on each first action recognition result corresponding to the video to be tested and the corresponding target path weights, obtain the aggregate action recognition result corresponding to the video to be tested.

[0035] The beneficial effects of the action recognition system based on multi-model aggregation of the present invention are as follows:

[0036] The system of the present invention improves the accuracy of action recognition in the target domain scene by enhancing the adaptability of the source domain model to the target domain scene.

[0037] Based on the above scheme, the action recognition system based on multi-model aggregation of the present invention can be further improved as follows.

[0038] In an alternative embodiment, the method further includes a preprocessing module; the preprocessing module is used for:

[0039] Determine the quantization value of the transferability between each source domain model and the target domain, and sort them in descending order according to the size of the quantization value to obtain the target queue;

[0040] The source domain models corresponding to the first N quantization values ​​of the target queue are determined as the filtered source domain models.

[0041] Thirdly, the technical solution of an electronic device according to the present invention is as follows:

[0042] It includes a memory, a processor, and a program stored in the memory and running on the processor, wherein the processor executes the program to implement the steps of the action recognition method based on multi-model aggregation as described in this invention.

[0043] Fourthly, the technical solution of a computer-readable storage medium provided by the present invention is as follows:

[0044] The computer-readable storage medium stores instructions that, when read, cause the computer-readable storage medium to perform the steps of the action recognition method based on multi-model aggregation of the present invention.

[0045] The above description is merely an overview of the technical solution of the present invention. In order to better understand the technical means of the present invention and to implement it in accordance with the contents of the specification, and in order to make the above and other objects, features and advantages of the present invention more apparent and understandable, specific embodiments of the present invention are described below. Attached Figure Description

[0046] The accompanying drawings are for illustrative purposes only and are not intended to limit the invention. Furthermore, the same reference numerals denote the same parts throughout the drawings. In the drawings:

[0047] Figure 1 This is a flowchart illustrating an embodiment of an action recognition method based on multi-model aggregation according to the present invention.

[0048] Figure 2 This is a schematic diagram of the training process for the source domain model library;

[0049] Figure 3 This is a schematic diagram illustrating the overall principle.

[0050] Figure 4 This is a schematic diagram of an embodiment of the action recognition system based on multi-model aggregation according to the present invention;

[0051] Figure 5 This is a schematic diagram of an embodiment of an electronic device according to the present invention. Detailed Implementation

[0052] Exemplary embodiments of the invention will now be described in more detail with reference to the accompanying drawings. While exemplary embodiments of the invention are shown in the drawings, it should be understood that the invention can be implemented in various forms and should not be limited to the embodiments set forth herein.

[0053] Figure 1 This diagram illustrates a flowchart of an embodiment of an action recognition method based on multi-model aggregation provided by the present invention. This action recognition method based on multi-model aggregation can be executed by electronic devices such as terminal devices or servers. The terminal device can be any fixed or mobile terminal, such as user equipment (UE), mobile device, user terminal, terminal, cellular phone, cordless phone, personal digital assistant (PDA), handheld device, computing device, in-vehicle device, or wearable device. The server can be a single server or a server cluster consisting of multiple servers. Any electronic device can implement the action recognition method based on multi-model aggregation by having its processor call computer-readable instructions stored in its memory. Figure 1 As shown, it includes the following steps:

[0054] S1. Input any sample video from the target domain into the initial path generation network to obtain N path weights corresponding to any sample video, and determine all path weights except the minimum path weight as the target path weights.

[0055] Where N ≥ 3 and N is a positive integer, each path weight corresponds to a filtered source domain model used for action recognition. In this embodiment, the default value of N is 3. The initial path generation network is an untrained path generation network, which is a standard C3d network containing 8 convolutional layers, 5 max pooling layers, 2 fully connected layers, and 1 softmax output layer. The softmax output layer outputs 3 path weights, the smallest path weight is recorded as 0, and the remaining path weights are retained and determined as the target path weights. The path generation network calculates the path weights using the following formula: G1(V i )=f top2 (G0(V i ),k); k is the number of path weights, G0(V i G1(V) represents the N path weights corresponding to the i-th sample video. i ) represents the N-1 target path weights corresponding to the i-th sample video.

[0056] S2. Input each sample video into the filtered source domain model corresponding to each target path weight to obtain multiple first action recognition results corresponding to each sample video, and based on each first action recognition result corresponding to each sample video and the corresponding target path weight, obtain the aggregate action recognition result corresponding to each sample video.

[0057] The first action recognition result is the action recognition result output by the source domain model. Each target path weight corresponds to a filtered source domain model, and each filtered source domain model outputs a first action recognition result. The aggregated action recognition result for any sample video is calculated using the following formula: output i H represents the aggregated action recognition result corresponding to the i-th sample video. j (V i ) indicates that V i The first action recognition result obtained by inputting the j-th filtered source domain model, where m represents the number of filtered source domain models, and A represents the first action recognition result obtained by inputting the j-th filtered source domain model. Perform cumulative summation.

[0058] S3. Repeat S1-S2 to obtain multiple target path weights and aggregated action recognition results corresponding to each sample video in the target domain.

[0059] For each sample video in the target domain, S1-S2 are executed repeatedly until the multiple target path weights and aggregated action recognition results corresponding to each sample video in the target domain are obtained.

[0060] S4. Based on the action recognition label, aggregated action recognition result, quantized value of instance-level transferability estimation index, and multiple target path weights corresponding to each sample video, calculate and optimize the parameters of the filtered source domain model according to the current loss value to obtain the optimized source domain model. Then, based on the quantized value of the instance-level transferability estimation index corresponding to each sample video, optimize the parameters of the initial path generation network to obtain the optimized path generation network.

[0061] In S4, the step of calculating the current loss value based on the action recognition label corresponding to each sample video, the aggregated action recognition result, the quantized value of the instance-level transferability estimation index, and the weights of multiple target paths includes:

[0062] The current loss value is calculated by substituting the action recognition label, aggregated action recognition result, quantized value of instance-level transferability estimation index, and multiple target path weights corresponding to each sample video into the target loss function. The target loss function is: L = θ1L cor +L SHTC , L represents the current loss value, θ1 is a hyperparameter, and L SHTC L represents the original loss obtained based on the action recognition label corresponding to each sample video and the aggregated action recognition result. cor G1(V) represents the weighted loss, where n represents the number of sample videos. i θ1 represents the weights of multiple target paths in the i-th sample video. It should be noted that θ1 takes the value 0.01.

[0063] Based on the second objective formula, the quantized value of the instance-level transferability estimation index corresponding to any sample video is calculated. The second objective formula is: Γ i This represents the quantized value of the instance-level transferability estimation index corresponding to the i-th sample video. Indicates to Perform softmax normalization.

[0064] It should be noted that the process of optimizing the parameters of the selected source domain model based on the current loss value to obtain the optimized source domain model, and optimizing the parameters of the initial path generation network based on the quantized value of the instance-level transferability estimation index corresponding to each sample video to obtain the optimized path generation network, is existing technology and will not be elaborated here.

[0065] S5. Use the optimized path generation network as the initial path generation network and the optimized source domain model as the filtered source domain model, and return to execute S3 until the iterative optimization conditions are met. Then, determine the initial path generation network as the trained path generation network and the filtered source domain model as the trained source domain model.

[0066] The iterative optimization condition is: the maximum number of iterations or the convergence of the loss function.

[0067] S6. Input the video to be tested into the trained path generation network to obtain the target path weights corresponding to the video to be tested. Then, input the video to be tested into the trained source domain model corresponding to each target path weight to obtain multiple first action recognition results corresponding to the video to be tested. Based on each first action recognition result corresponding to the video to be tested and the corresponding target path weights, obtain the aggregate action recognition result corresponding to the video to be tested.

[0068] The video to be tested is the video data for which action recognition needs to be performed in this embodiment. Except for calculating the loss value, the detection method for the video to be tested is the same as that for the sample video, and will not be described in detail here.

[0069] In one alternative approach, it also includes:

[0070] Determine the quantization value of the transferability between each source domain model and the target domain, and sort them in descending order according to the size of the quantization value to obtain the target queue.

[0071] Among them, such as Figure 2 As shown, based on the mmaction2 framework, pre-trained models for video action recognition tasks are selected from a source domain model library with various structures, and trained on the source domain dataset to obtain multiple source domain models. Each source domain model corresponds to a quantization value of transferability.

[0072] Specifically, the step of determining the quantized value of the transferability between any source domain model and the target domain includes: determining the quantized value of the transferability between any source domain model and the target domain based on a first target formula.

[0073] The first objective formula is: IC am =E(IC) i )×SUTE×TC÷E(IC i );

[0074]

[0075] IC i =H(h(V) i )), IC am V represents the quantized value of the transferability between any source domain model and the target domain. i Let h(V) represent the i-th sample video of the target domain. i ) represents V i The output result obtained by inputting the data into any of the source domain models, H(h(V) i )) represents the calculation of h(V) i The entropy of ), SUTE represents the quantized value of the source-free domain transferability estimation index, TC represents the quantized value of time consistency, and IC i E(IC) represents the quantized value of the determinism of an instance of any source domain model. i () indicates the calculation of IC with different values i The mathematical expectation, Indicates V i Perform time perturbation (for V) i The video after random time enhancement of the sampled frames in the video, D T This represents the sample set of the target domain. This represents the calculation of H(h(V) with different values. i The average value of )) This indicates the calculation of h(V) with different values. i The average value of ) Represents the predicted semantics of the sample video. This represents the pseudo-label of the sample video. Indicates calculation KL divergence, h c (V i ) represents V i Predicted probability for class c.

[0076] The source domain models corresponding to the first N quantization values ​​of the target queue are determined as the filtered source domain models.

[0077] The default value for N is 3.

[0078] In one alternative approach, the step of obtaining a quantified value of the instance-level transferability estimation index corresponding to any sample video includes:

[0079] Based on the second objective formula, calculate the quantized value of the instance-level transferability estimation index corresponding to any sample video.

[0080] The formula for the second objective is: Γ i This represents the quantized value of the instance-level transferability estimation index corresponding to the i-th sample video. Indicates to Perform softmax normalization.

[0081] It should be noted that, Figure 3 The schematic diagram of this embodiment is shown. This embodiment verifies the effectiveness of the proposed scheme through experiments on the large public dataset Daily-DA for video action recognition. DailyDA is another large-scale cross-domain action recognition benchmark. It includes four datasets: ARID (A), HMDB51 (H), Moments-in-Time (M), and Kinetics (K). Videos from eight shared classes are used for cross-domain evaluation. Because the models in the source domain model library are pre-trained in Kinetics, the Kinetics dataset is removed to ensure performance accuracy. After removal, three datasets remain, generating a total of six tasks. For the accuracy metric, Mean-1 accuracy was chosen, representing the average accuracy of the Top-1 in each category. The method is compared with current state-of-the-art domain adaptation algorithms SHOT, STHC, and multi-model domain adaptation algorithms CAiDA, KD3A, and DECISION. The comparison results are shown in Table 1 below. As shown in Table 1, the method of this embodiment significantly improves the action recognition accuracy of the source domain model in the target domain.

[0082] Table 1:

[0083] Method M->A H->A A->M H->M A->H M->H AVG SHOT 58.95 39.74 44.50 48.00 60.00 69.17 53.39 STHC 57.72 49.67 45.00 48.25 59.17 69.58 54.90 Decision 52.47 49.73 45.75 47.50 57.92 68.75 53.69 kd3A) 57.12 50.09 45.25 48.25 59.17 69.17 54.84 Caida 57.57 48.13 45.25 47.75 59.17 68.75 54.44 This plan 63.37 48.25 51.00 51.25 68.75 79.58 60.37

[0084] The technical solution in this embodiment improves the adaptability of the source domain model to the target domain scene, thereby enhancing the accuracy of action recognition in the target domain scene.

[0085] Figure 4 A schematic diagram of an embodiment of an action recognition system 200 based on multi-model aggregation provided by the present invention is shown. Figure 4 As shown, the system 200 includes: a first processing module 210, a second processing module 220, a third processing module 230, a fourth processing module 240, an iterative training module 250, and a detection module 260;

[0086] The first processing module 210 is used to: input any sample video of the target domain into the initial path generation network, obtain N path weights corresponding to any sample video, and determine all path weights other than the minimum path weight as target path weights; wherein, N≥3 and N is a positive integer, and each path weight corresponds to a filtered source domain model for action recognition.

[0087] The second processing module 220 is used to: input the any sample video into the filtered source domain model corresponding to each target path weight, obtain multiple first action recognition results corresponding to the any sample video, and obtain the aggregated action recognition result corresponding to the any sample video based on each first action recognition result corresponding to the any sample video and the corresponding target path weight;

[0088] The third processing module 230 is used to: repeatedly call the first processing module 120 to the second processing module 220 to obtain multiple target path weights and aggregated action recognition results corresponding to each sample video in the target domain;

[0089] The fourth processing module 240 is used to: calculate and optimize the parameters of the filtered source domain model based on the action recognition label corresponding to each sample video, the aggregated action recognition result, the quantized value of the instance-level transferability estimation index, and multiple target path weights, to obtain an optimized source domain model; and optimize the parameters of the initial path generation network based on the quantized value of the instance-level transferability estimation index corresponding to each sample video to obtain an optimized path generation network.

[0090] The iterative training module 250 is used to: use the optimized path generation network as the initial path generation network and the optimized source domain model as the filtered source domain model, and return to call the third processing module 230 until the iterative optimization conditions are met, and then determine the initial path generation network as the trained path generation network and the filtered source domain model as the trained source domain model.

[0091] The detection module 260 is used to: input the video to be tested into the trained path generation network to obtain the target path weights corresponding to the video to be tested, and input the video to be tested into the trained source domain model corresponding to each target path weight to obtain multiple first action recognition results corresponding to the video to be tested, and obtain the aggregate action recognition result corresponding to the video to be tested based on each first action recognition result corresponding to the video to be tested and the corresponding target path weights.

[0092] In an alternative embodiment, the method further includes a preprocessing module; the preprocessing module is used for:

[0093] Determine the quantization value of the transferability between each source domain model and the target domain, and sort them in descending order according to the size of the quantization value to obtain the target queue;

[0094] The source domain models corresponding to the first N quantization values ​​of the target queue are determined as the filtered source domain models.

[0095] In one alternative approach, the preprocessing module is specifically used for:

[0096] Based on the first objective formula, a quantized value of the transferability between any source domain model and the target domain is determined; wherein, the first objective formula is: IC am =E(IC) i )×SUTE×TC÷E(IC i ); IC i =H(h(V) i )), IC am V represents the quantized value of the transferability between any source domain model and the target domain. i Let h(V) represent the i-th sample video of the target domain. i ) represents V i The output result obtained by inputting the data into any of the source domain models, H(h(V) i )) represents the calculation of h(V) i The entropy of ), SUTE represents the quantized value of the source-free domain transferability estimation index, TC represents the quantized value of time consistency, and IC i E(IC) represents the quantized value of the determinism of an instance of any source domain model. i () indicates the calculation of IC with different values i The mathematical expectation, Indicates V i The video after time perturbation processing, D T This represents the sample set of the target domain. This represents the calculation of H(h(V) with different values. i The average value of )) This indicates the calculation of h(V) with different values. i The average value of ) Represents the predicted semantics of the sample video. This represents the pseudo-label of the sample video. Indicates calculation KL divergence, h c (V i ) represents V i Predicted probability for class c.

[0097] In an alternative embodiment, the method further includes: an acquisition module; the acquisition module is used for:

[0098] The steps for obtaining the quantified value of the instance-level transferability estimation index corresponding to any sample video include:

[0099] Based on the second objective formula, the quantized value of the instance-level transferability estimation index corresponding to any sample video is calculated; wherein, the second objective formula is: Γ i This represents the quantized value of the instance-level transferability estimation index corresponding to the i-th sample video. Indicates to Perform softmax normalization.

[0100] In an alternative embodiment, the fourth processing module 240 is specifically used for:

[0101] The current loss value is calculated by substituting the action recognition label, aggregated action recognition result, quantized value of instance-level transferability estimation index, and multiple target path weights corresponding to each sample video into the target loss function; wherein, the target loss function is: L=θ1L cor +L SHTC , L represents the current loss value, θ1 is a hyperparameter, and L SHTC L represents the original loss obtained based on the action recognition label corresponding to each sample video and the aggregated action recognition result. cor G1(V) represents the weighted loss, where n represents the number of sample videos. i ) represents the weights of multiple target paths in the i-th sample video.

[0102] In one alternative approach, the iterative optimization condition is: the maximum number of iterations or the convergence of the loss function.

[0103] The technical solution in this embodiment improves the adaptability of the source domain model to the target domain scene, thereby enhancing the accuracy of action recognition in the target domain scene.

[0104] The parameters and steps for implementing the corresponding functions of each module in the multi-model aggregation-based action recognition system 200 of this embodiment can be referred to the parameters and steps in the embodiments of the multi-model aggregation-based action recognition method above, and will not be repeated here.

[0105] like Figure 5 As shown, an electronic device 300 according to an embodiment of the present invention includes a processor 320 coupled to a memory 310. The memory 310 stores at least one computer program 330, which is loaded and executed by the processor 320 to enable the electronic device 300 to implement any of the above-mentioned multi-model aggregation-based action recognition methods. Specifically:

[0106] The electronic device 300 can vary considerably due to differences in configuration or performance. It may include one or more processors 320 (Central Processing Units, CPUs) and one or more memories 310. The one or more memories 310 store at least one computer program 330, which is loaded and executed by the one or more processors 320 to enable the electronic device 300 to implement any of the multi-model aggregation-based action recognition methods provided in the above embodiments. Of course, the electronic device 300 may also have wired or wireless network interfaces, a keyboard, and input / output interfaces for input and output. The electronic device 300 may also include other components for implementing device functions, which will not be elaborated upon here.

[0107] An embodiment of the present invention provides a computer-readable storage medium storing at least one computer program, which is loaded and executed by a processor to enable a computer to implement any of the above-described multi-model aggregation-based action recognition methods.

[0108] Alternatively, the computer-readable storage medium may be a read-only memory (ROM), a random access memory (RAM), a compact disc read-only memory (CD-ROM), magnetic tape, a floppy disk, and an optical data storage device, etc.

[0109] In an exemplary embodiment, a computer program product or computer program is also provided, which includes computer instructions stored in a computer-readable storage medium. A processor of an electronic device reads the computer instructions from the computer-readable storage medium and executes the computer instructions, causing the electronic device to perform any of the aforementioned multi-model aggregation-based action recognition methods.

[0110] It should be noted that the terms "first," "second," etc., used in the specification and claims of this application are used to distinguish similar objects and represent a limitation on a specific order or sequence. Where appropriate, the order of use for similar objects can be interchanged so that the embodiments of this application described herein can be implemented in an order other than that shown or described.

[0111] Those skilled in the art will recognize that this invention can be implemented as a system, method, or computer program product. Therefore, this disclosure can be specifically implemented in the following forms: it can be entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software, generally referred to herein as a "circuit," "module," or "system." Furthermore, in some embodiments, this invention can also be implemented as a computer program product in one or more computer-readable media containing computer-readable program code.

[0112] Any combination of one or more computer-readable media can be used. A computer-readable medium can be a computer-readable signal medium or a computer-readable storage medium. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples (a non-exhaustive list) of computer-readable storage media include: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this application, a computer-readable storage medium can be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device.

[0113] Although embodiments of the present invention have been shown and described above, it is understood that the above embodiments are exemplary and should not be construed as limiting the present invention. Those skilled in the art can make changes, modifications, substitutions and variations to the above embodiments within the scope of the present invention.

Claims

1. A method for action recognition based on multi-model aggregation, characterized in that, Comprising: S1, input any sample video of the target domain to the initial path generation network to obtain N path weights corresponding to the any sample video, and determine all path weights except the minimum path weight as target path weights; wherein, Each path weight corresponds to a screened source domain model for action recognition. The initial path generation network is an untrained path generation network. The path generation network is a standard C3d network, which includes 8 convolution layers, 5 max pooling, 2 fully connected layers and 1 softmax output layer. The softmax output layer outputs 3 path weights. The minimum path weight is recorded as 0, and the remaining path weights are reserved and determined as target path weights. The path generation network uses the following formula to calculate the path weight: ; N is the number of path weights, N is the N path weights corresponding to the i-th sample video, N-1 is the N-1 target path weights corresponding to the i-th sample video. S2, inputting the any sample video into each target path weight corresponding screened source domain model respectively to obtain a plurality of first action recognition results corresponding to the any sample video, and based on each first action recognition result corresponding to the any sample video and the corresponding target path weight, obtaining an aggregated action recognition result corresponding to the any sample video; S3, repeatedly performing S1-S2 to obtain a plurality of target path weights and aggregated action recognition results corresponding to each sample video of the target domain; S4, based on the action recognition label corresponding to each sample video, the aggregated action recognition result, the quantized value of the instance-level transferability estimation indicator, and the plurality of target path weights, calculating and optimizing the parameters of the screened source domain model according to the current loss value, obtaining an optimized source domain model, and optimizing the parameters of the initial path generation network according to the quantized value of the instance-level transferability estimation indicator corresponding to each sample video, obtaining an optimized path generation network; S5, taking the optimized path generation network as the initial path generation network, and taking the optimized source domain model as the screened source domain model, and returning to perform S3 until the iteration optimization condition is met, determining the initial path generation network as a trained path generation network, and determining the screened source domain model as a trained source domain model; S6, inputting the to-be-tested video into the trained path generation network to obtain a target path weight corresponding to the to-be-tested video, and inputting the to-be-tested video into each trained source domain model corresponding to the target path weight to obtain a plurality of first action recognition results corresponding to the to-be-tested video, and based on each first action recognition result corresponding to the to-be-tested video and the corresponding target path weight, obtaining an aggregated action recognition result corresponding to the to-be-tested video; Further comprising: determining the quantized value of the transferability between each source domain model and the target domain, and sorting in descending order according to the size of the quantized value to obtain a target queue; and determining the first N source domain models corresponding to the quantized values in the target queue as the screened source domain models; The step of determining the quantized value of the transferability between any source domain model and the target domain comprises: determining a quantified value of transferability between the any source domain model and the target domain based on a first target formula, wherein the first target formula is: ; , , , ; denotes a quantified value of transferability between the any source domain model and the target domain, denotes an i-th sample video of the target domain, denotes an output result obtained by inputting the any source domain model, denotes an entropy of , denotes a quantified value of source-free transfer estimation indicator, denotes a quantified value of temporal consistency, denotes a quantified value of instance individual determinacy of the any source domain model, denotes a mathematical expectation of of different values, denotes a video after time disturbance processing on , denotes a sample set of the target domain, denotes an average value of of different values, denotes an average value of of different values, denotes a predicted semantic of a sample video, denotes a pseudo label of a sample video, denotes a KL divergence of , denotes a predicted probability on a class c; The step of obtaining the quantized value of the instance-level transferability estimation indicator corresponding to the any sample video comprises: Based on the second target formula, a quantitative value of an instance-level transferability estimation index corresponding to each sample video is calculated; wherein the second target formula is: ; represents a quantitative value of an instance-level transferability estimation index corresponding to the i th sample video, represents a softmax normalization processing on . The step of calculating the current loss value based on the action recognition label corresponding to each sample video, the aggregated action recognition result, the quantized value of the instance-level transferability estimation indicator, and the plurality of target path weights comprises: The action recognition label corresponding to each sample video, the aggregated action recognition result, the quantized value of the instance-level transferability estimation indicator, and the plurality of target path weights are substituted into a target loss function to calculate the current loss value; wherein the target loss function is: , ; represents the current loss value, is a hyperparameter, represents an original loss obtained according to the action recognition label corresponding to each sample video and the aggregated action recognition result, represents a weight loss, and n represents the number of sample videos, represents the plurality of target path weights of the i th sample video. 2.The multi-model aggregation based action recognition method of claim 1, wherein, The iteration optimization condition is: the maximum number of iterations or the convergence of the loss function.

3. A multi-model aggregation based action recognition system, characterized in that, Comprising: The first processing module, the second processing module, the third processing module, the fourth processing module, the iterative training module, and the detection module; The first processing module is configured to input any sample video of the target domain into an initial path generation network to obtain N path weights corresponding to the any sample video, and determine all path weights except the minimum path weight as target path weights; wherein Each path weight corresponds to a screened source domain model for action recognition; the initial path generation network is an untrained path generation network, the path generation network is a standard C3d network, and the path generation network includes 8 convolution layers, 5 maximum pooling layers, 2 full connection layers and 1 softmax output layer; the softmax output layer outputs 3 path weights, the minimum path weight is recorded as 0, and the remaining path weights are reserved and determined as target path weights; the path generation network uses the following formula to calculate the path weights: ; is the number of path weights, is the N path weights corresponding to the i th sample video, is the N-1 target path weights corresponding to the i th sample video; The second processing module is configured to input the any sample video into each target path weight corresponding screened source domain model respectively to obtain a plurality of first action recognition results corresponding to the any sample video, and obtain an aggregated action recognition result corresponding to the any sample video based on each first action recognition result corresponding to the any sample video and the corresponding target path weight; The third processing module is configured to repeatedly call the first processing module to the second processing module to obtain a plurality of target path weights and aggregated action recognition results corresponding to each sample video of the target domain; The fourth processing module is configured to calculate and optimize parameters of the screened source domain model based on the current loss value, and obtain an optimized source domain model based on the current loss value, and optimize parameters of the initial path generation network based on the quantified value of the instance-level transferability estimation index corresponding to each sample video to obtain an optimized path generation network; The iterative training module is configured to use the optimized path generation network as the initial path generation network, use the optimized source domain model as the screened source domain model, and call the third processing module until the initial path generation network is determined as a trained path generation network and the screened source domain model is determined as a trained source domain model when an iterative optimization condition is met; The detection module is configured to input a to-be-detected video into the trained path generation network to obtain a target path weight corresponding to the to-be-detected video, input the to-be-detected video into each trained source domain model corresponding to the target path weight respectively to obtain a plurality of first action recognition results corresponding to the to-be-detected video, and obtain an aggregated action recognition result corresponding to the to-be-detected video based on each first action recognition result corresponding to the to-be-detected video and the corresponding target path weight; Further comprising a preprocessing module; the preprocessing module is configured to: determine the quantified value of the transferability between each source domain model and the target domain respectively, and sort the quantified values in descending order of size to obtain a target queue; and determine the first N quantified values of the target queue as the screened source domain models respectively; The preprocessing module is specifically configured to: determining a quantified value of transferability between the any source domain model and the target domain based on a first target formula, wherein the first target formula is: ; , , , ; denotes a quantified value of transferability between the any source domain model and the target domain, denotes an i-th sample video of the target domain, denotes an output result obtained by inputting the any source domain model, denotes an entropy of , denotes a quantified value of source-free transfer estimation indicator, denotes a quantified value of temporal consistency, denotes a quantified value of instance individual determinacy of the any source domain model, denotes a mathematical expectation of of different values, denotes a video after time disturbance processing on , denotes a sample set of the target domain, denotes an average value of of different values, denotes an average value of of different values, denotes a predicted semantic of a sample video, denotes a pseudo label of a sample video, denotes a KL divergence of , denotes a predicted probability on a c-th class; Further comprising an acquisition module; the acquisition module is configured to: The step of acquiring the quantified value of the instance-level transferability estimation index corresponding to the any sample video comprises: Based on the second target formula, a quantitative value of an instance-level transferability estimation index corresponding to each sample video is calculated; wherein the second target formula is: ; represents a quantitative value of an instance-level transferability estimation index corresponding to the i th sample video, represents a softmax normalization processing on ; The fourth processing module is specifically configured to: The action recognition label corresponding to each sample video, the aggregated action recognition result, the quantized value of the instance-level transferability estimation indicator, and the plurality of target path weights are substituted into a target loss function to calculate the current loss value; wherein the target loss function is: , ; is represented as the current loss value, is a hyperparameter, represents an original loss obtained according to the action recognition label corresponding to each sample video and the aggregated action recognition result, represents a weight loss, and n represents the number of sample videos, represents the plurality of target path weights of the i-th sample video.

4. An electronic device, comprising: The electronic device comprises a processor and a memory, the memory stores at least one computer program, the at least one computer program is loaded and executed by the processor, so that the electronic device implements the action recognition method based on multi-model aggregation as claimed in claim 1 or 2.

5. A computer readable storage medium, characterized in that, The computer readable storage medium stores at least one computer program, and the at least one computer program is loaded and executed by the processor, so that the computer readable storage medium implements the multi-model aggregation-based action recognition method according to claim 1 or 2.

Citation Information

Patent Citations

  • Multi-source cross-domain expression recognition method and device and storage medium

    CN114612961A

  • Cross-domain video action recognition method, device and equipment and computer readable storage medium

    CN115439791A