Multi-object Scenario Spatiotemporal Action Localization Method and Device Based on Weakly Supervised Learning
Through the multi-object scene spatio-temporal action positioning method based on weak supervision learning, the convolutional neural network and pseudo-label generation technology are used to solve the annotation needs of action positioning in multi-person scenarios, efficient action recognition is achieved, and labeling costs are reduced.
Patent Information
- Application Number
- CN202210765837.X
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-07-01
- Publication Date
- 2025-07-11
- Estimated Expiration
- 2042-07-01
AI Technical Summary
The existing fully supervised space-time action positioning method requires a lot of tedious manual annotations in multi-person scenarios, and the existing weak supervision methods cannot effectively solve the action positioning problem in dense multi-person scenarios.
The multi-object scene spatio-temporal action positioning method based on weak supervision learning is adopted, and feature extraction is performed using convolutional neural networks. By generating pseudo-labels and cross-entropy loss functions, the annotation amount is reduced, multiple body package labels and frame interval pseudo-labels are generated, and parameters are updated to achieve action recognition.
In multi-player scenarios, the number of labels is effectively reduced, and the recognition effect similar to the full supervision method is achieved. It can identify the type, position and duration of the actions in the video without a lot of cumbersome manual labeling.
Smart Images

Figure CN115147925B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of computer vision, and particularly to a multi-object scene spatio-temporal action localization method and device based on weakly supervised learning. Background Art
[0002] Using computer vision methods to recognize the actions of target objects in videos provides support for many important downstream tasks (such as abnormal behavior monitoring). In this case, an important issue is the spatio-temporal localization of the actions of target objects in videos. This problem aims to determine the action category of the target object at a specific time point and specific spatial position. However, the commonly used fully supervised spatio-temporal action localization methods rely on a large amount of cumbersome manual annotations.
[0003] (1) Spatio-temporal action localization (see reference [1]): Spatio-temporal action localization requires localizing the duration and spatial position of an action and identifying the action category. Early spatio-temporal action localization methods mainly focused on action localization in a short video containing only one person or a small number of people, and were not applicable to crowded surveillance scenarios.
[0004] (2) Weakly supervised action localization (see reference [2]): Most previous studies on the action localization problem were fully supervised methods, which required complete and detailed annotations for each frame of the video. The annotation process was cumbersome and costly. Weakly supervised action localization methods can greatly reduce the dependence on complex annotations. However, existing weakly supervised action localization methods usually only focus on the situation where there are one or two people in the scene and cannot solve the action localization problem in crowded multi-person scenarios.
[0005] (3) Multi-instance learning technology: Multi-instance learning technology has been widely applied to various weakly supervised methods. However, existing multi-instance learning technology has the problem of low utilization rate of annotation information.
[0006] References
[0007] [1]Chunhui Gu,Chen Sun,David A Ross,Carl Vondrick,Caroline Pantofaru,Yeqing Li,Sudheendra Vijayanarasimhan,George Toderici,Susanna Ricco,RahulSukthankar:Ava:A video dataset of spatio-temporally localized atomic visualactions.CVPR 2018:6047–6056.
[0008] [2]Chéron, Guilhem, Alayrac J B, Laptev I: A flexible model for training action localization with varying levels of supervision. NIPS 2018:942–953. Summary of the Invention
[0009] The present invention provides a method and device for spatio-temporal action localization in a multi-object scene based on weakly supervised learning. The present invention uses less data annotation to complete spatio-temporal action localization in a dense crowd scene, effectively solving the problem of lack of large-scale training data in spatio-temporal action localization in a multi-person scene and achieving good results, as described in detail below:
[0010] A method for spatio-temporal action localization in a multi-object scene based on weakly supervised learning, the method comprising:
[0011] Using a convolutional neural network as a feature extraction network, performing frame-by-frame feature extraction on a video, and extracting target object features for each video frame and outputting feature vectors of all individuals through a fully connected layer;
[0012] Calculating the individual set X in the original dataset M of multi-individual package labels and calculating an action prediction result through the feature vectors of the individuals Calculating the multi-individual package labels and the action prediction result to obtain a cross-entropy loss therebetween;
[0013] Finding adjacent frames within a certain range for frame t, screening with a threshold ξ t to generate pseudo-labels for qualified individuals, examining unlabeled individual x, and calculating its maximum predicted value and the most likely label value to obtain a cross-entropy loss c therebetween, screening with a threshold ξ a , ξ c as a standard to add the qualified individuals as anchored individuals;
[0014] Comprehensively considering the cross-entropy loss function, using as a total loss function, updating parameters using the original training set and the new dataset after pseudo-labeling, and stopping after completing the specified number of training rounds;
[0015] Decomposing the video to be recognized into several frames, inputting them frame by frame into the trained network for action recognition, and outputting the action categories of all individuals in the video frames.
[0016] Among them, the individual set X in the original calculation dataset M with multiple individual packet labels and calculates the action prediction result through the feature vector of the individual Specifically:
[0017] The overall composed of M individuals X with labels M ={x1, x2, …, x M} is marked as a multiple individual packet Multiple individual packets The label of is obtained by the following formula:
[0018]
[0019] Among them, y m represents the action label of the individual x m and means performing an OR operation on the corresponding elements one by one;
[0020] For the packet with M individuals Calculate the action prediction result at the multiple individual packet level K represents the number of action categories, and the prediction result can be obtained by the following formula:
[0021]
[0022] Among them, φ represents the aggregation function; p m ∈R 1×K represents the action prediction result of the individual x m and is obtained through the Inception-v3 network with RoIAlign.
[0023] Among them, the method further includes: training the action prediction result Specifically:
[0024] Calculate the minimum loss
[0025]
[0026] Among them, when calculating the aggregation function φ adopts the min-pooling function;
[0027] Calculate the intermediate loss
[0028]
[0029] Among them, represents the set of individuals whose predicted values are not in the maximum loss and minimum loss, and the U function is defined as:
[0030]
[0031] Among them, \(c1, c2\in[0, 1]\), and \(d > 0\) is a preset parameter;
[0032] Finally, the improved multi-agent loss function is expressed as:
[0033]
[0034] Among them, the parameters \(\alpha\) and \(\beta\) are used to balance different situations, and the total loss function is obtained by summing over each packet ;
[0035] Introduce an uncertainty quantity for the matrix \(P\) Calculate the multi-agent packet loss
[0036]
[0037] Among them, the screening with the threshold \(\xi\) t is as follows:
[0038] For an anchored individual \(q\) t , assuming its label is action \(k\), for the adjacent frame \(t - 1\), if the predicted action category of the same individual \(q\) t-1 is also \(k\), and the corresponding predicted score is greater than the threshold \(\xi\) t , then for the unlabeled individual \(q\) t-1 generate a pseudo-label with the action category \(k\) and add it to the training set, and use the same method to expand the single-frame annotation into a multi-frame annotation frame by frame within the interval \([t - R, t + R]\).
[0039] Among them, the screening with the threshold \(\xi\) a , \(\xi\) c as the standard, the screening threshold is:
[0040] For a certain unlabeled individual \(x\) n , among the predicted scores of all its action categories, the predicted score of category \(k\) is the highest. Calculate the prediction confidence \(c\), and \(c\) is the cross-entropy loss between the predicted value and the most likely label value ; when and only when the predicted value and the confidence \(c\) satisfy the following relationship, add the individual \(x\) n as an anchored individual:
[0041]
[0042] Among them, \(\xi\) a , \(\xi\) c is a threshold parameter.
[0043] A multi-object scenario spatio-temporal action localization device based on weakly supervised learning, the device comprising: a processor and a memory, wherein program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to enable the device to execute the method steps described above.
[0044] The beneficial effects of the technical solution provided by the present invention are as follows:
[0045] 1. The present invention studies for the first time the problem of weakly supervised spatio-temporal action localization in a multi-person scenario, and can be effectively applied to video surveillance and sports scenario analysis. Compared with the fully supervised spatio-temporal action localization method, it does not require a large amount of training data and has high applicability;
[0046] 2. The weakly supervised action localization method adopted by the present invention achieves results similar to those of the fully supervised method on the basis of greatly reducing the required annotation amount, and has good recognition effect;
[0047] 3. The present invention can be used for video surveillance and sports scenario analysis, and can identify the types, positions, and durations of actions in the video without a large amount of cumbersome manual annotation;
[0048] 4. The present invention aims to locate the actions of each person in a multi-person scenario, which is more in line with the application of spatio-temporal action localization in the real world. Description of the Drawings
[0049] Figure 1 is a flowchart of a multi-object scenario spatio-temporal action localization method based on weakly supervised learning;
[0050] Figure 2 is a schematic structural diagram of a multi-object scenario spatio-temporal action localization device based on weakly supervised learning. Detailed Embodiments
[0051] To make the objectives, technical solutions, and advantages of the present invention clearer, the embodiments of the present invention will be further described in detail below.
[0052] Embodiment 1
[0053] The embodiment of the present invention provides a multi-object scenario spatio-temporal action localization method based on weakly supervised learning. Refer to Figure 1 , the method includes the following steps:
[0054] Step 101: Use a convolutional neural network as the backbone to build a feature extraction network, and perform frame-by-frame feature extraction on the video. For each video frame, use RoIAlign (a target object feature extraction technique) and a fully connected layer to obtain the feature vectors of all individuals;
[0055] Step 102: Calculate the individual set X in the original dataset M multiple individual packet labels of and calculate the action prediction result through the feature vectors of the individuals in Step 101 Calculate the multiple individual packet labels and the action prediction result The cross-entropy loss between them;
[0056] Step 103: First, for frame t, find the neighboring frames within a certain range, examine the corresponding action prediction values of the anchored individuals in the neighboring frames, and screen them with the threshold ξ t as the standard, generate pseudo-labels for the qualified individuals, then examine the unlabeled individual x, and calculate its maximum prediction value and the most likely label value The cross-entropy loss c between them, and screen them with the thresholds ξ a and ξ c as the standard, and add the qualified individuals as the anchored individuals;
[0057] Step 104: Comprehensively consider the cross-entropy loss functions in Steps 102 and 103, and use as the total loss function, update the parameters using the original training set and the new dataset after adding the pseudo-labels generated in Step 103, and stop after completing the specified number of training rounds;
[0058] Step 105: Decompose the video to be recognized into several frames, input them frame by frame into the network trained in Step 104 for action recognition, and the network will output the action categories of all individuals in the video frames.
[0059] In summary, the embodiments of the present invention are used in video surveillance and sports scene analysis through the above Steps 101-105 to recognize the types, positions, and durations of actions in the video in the case of lack of large-scale training data and cumbersome manual annotation processes.
[0060] Embodiment 2
[0061] The following combines Figure 1 specific calculation formulas and examples to further introduce the solution in Embodiment 1, as described in detail below:
[0062] The technical solution adopted in the embodiments of the present invention is: a weakly supervised spatio-temporal action localization method in a multi-person scenario, which is mainly described from two aspects: one is the problem definition, and the other is the weakly supervised training method based on the above definition.
[0063] I. Problem definition:
[0064] Given an unclipped video, the video has a total of K action categories and n individuals. For the individuals in the video, use Xn = {x1, x2, …, x n}, the true action label of an individual is represented as a vector y n ∈R 1×K . The objective of the embodiments of the present invention is to establish and train a weakly supervised action recognition network, and predict a matrix P ∈ {0, 1} for each individual N×K . The element in the n-th row and k-th column of this matrix represents the predicted probability of the K action of the corresponding individual n
[0065] II. Weakly supervised action localization method:
[0066] For the spatio-temporal action localization problem in a multi-person scenario with less data annotation, the weakly supervised action localization method of the embodiments of the present invention solves the problem of less annotation information in the dataset by generating multi-individual packet labels and frame interval pseudo-labels, and adopting corresponding weakly supervised learning methods. The label generation process is as Figure 1 shown
[0067] Step 201: Use a convolutional neural network as the backbone to build a feature extraction network, and perform frame-by-frame feature extraction on the video. For each video frame, use RoIAlign and a fully connected layer to obtain the feature vectors of all individuals
[0068] The embodiments of the present invention complete this operation by using the method of deep learning which is widely applied at the present stage. Specifically, use a convolutional neural network as the backbone to perform frame-by-frame feature extraction on the video, and use RoIAlign and a fully connected layer to obtain the features of each individual
[0069] Step 202: Calculate the multi-individual packet labels M of the individual set X in the original dataset and the action prediction result
[0070] Mark the whole composed of M individuals with labels X M = {x1, x2, …, x M} as a multi-individual packet The multi-individual packet label can be obtained by the following formula
[0071]
[0072] where y m represents the action label of the individual x m , represents performing the OR operation on the corresponding elements one by one
[0073] For the packet with M individuals Calculate the action prediction results at the multi - individual package level Let \(K\) denote the number of action categories, and the prediction result can be obtained by the following formula:
[0074]
[0075] where \(\varphi\) represents the aggregation function; \(p m \in\mathbb{R} 1×K represents the action prediction result of individual \(x m , which is obtained through the Inception - v3 network with RoIAlign.
[0076] Calculate the maximum loss It can be obtained by the following formula:
[0077]
[0078] where \(K\) is the number of action categories. When calculating , the aggregation function \(\varphi\) adopts the max - pooling function.
[0079] Calculate the minimum loss It can be obtained by the following formula:
[0080]
[0081] where, when calculating , the aggregation function \(\varphi\) adopts the min - pooling function.
[0082] Calculate the intermediate loss It can be obtained by the following formula:
[0083]
[0084] where, represents the set of individuals whose predicted values are not among the maximum loss and the minimum loss. The \(U\) function is defined as:
[0085]
[0086]
[0087] Finally, the improved multi - individual loss function (as shown in Figure 2 ) is expressed as:
[0088]
[0089] where the parameters \(\alpha\) and \(\beta\) are used to balance different situations, and the total loss function is obtained by summing over each package .
[0090] Based on the multi-agent loss, considering the problem of action label loss when decomposing a large packet into smaller packets, and the inevitable misselection problem when using max-pooling and min-pooling aggregation functions to select agents, there will be significant noise interference in actual situations. Therefore, an uncertainty quantity is introduced for matrix P Calculate the multi-agent packet loss
[0091]
[0092] Step 203: First, find the neighboring frames within a certain range for frame t, examine the corresponding anchor agent action prediction values in the neighboring frames, and use the threshold ξ t as the standard for screening. Generate pseudo-labels for the qualified agents, and then examine the unlabeled agent x and calculate its maximum prediction value and the most likely label value The cross-entropy loss c between them is used as the standard for screening with the thresholds ξ a , ξ c as the standard for screening, and the qualified agents are newly added as anchor agents.
[0093] a) First stage of Frame Periodic Label (FPL) generation:
[0094] Given a labeled frame t, randomly select Q agents with action labels on frame t as anchor agents. Among them, for an anchor agent q t , assume its label is action k. For the adjacent frame t - 1, if the predicted action category of the same agent q t-1 is also k and the corresponding prediction score is greater than the threshold ξ t , then generate a pseudo-label with action category k for the unlabeled agent q t-1 and add it to the training set. Similarly, use the same method to expand the single-frame annotation to multi-frame annotations frame by frame within the interval [t - R, t + R].
[0095] b) Second stage of Frame Periodic Label generation:
[0096] For a certain unlabeled agent x n , assume that among the prediction scores of all its action categories, the prediction score of category k is the highest. Calculate the prediction confidence c, where c is the cross-entropy loss between the prediction value and the most likely label value . When and only when the prediction value and the confidence c satisfy the following relationship, add the agent x n as an anchor agent:
[0097]
[0098] Among them, ξ a , ξ c is a threshold parameter.
[0099] After stage one and stage two, a number of pseudo-labels are generated and added to the training data set. Calculate the pseudo-label loss function
[0100]
[0101] Among them, respectively represent the true label - individual pair, the pseudo-label - individual pair, is the label value after completing the pseudo-label generation process.
[0102] Step 204: Considering the loss functions in steps 202 and 203 comprehensively, take as the total loss function, use the original training set and the new data set after adding the pseudo-labels generated in step 203 to update the parameters, and stop after completing the specified number of rounds of training;
[0103] The total loss function consists of the multi-individual packet loss and the pseudo-label loss function is given by:
[0104]
[0105] Use the original training set to train the action recognition network for 30 rounds, enter the first stage of frame interval pseudo-label generation, and generate labels on the unlabeled video frames. Use the new training sample pool with pseudo-labels to train the action recognition network, enter the second stage of frame interval pseudo-label generation, and add anchor individuals. Use the new training sample pool to train the action recognition network.
[0106] Step 205: Decompose the video to be recognized into several frames, and input them frame by frame into the network trained in step 204 for action recognition. The network will output the action categories of all individuals in the video frames.
[0107] Separate the video to be predicted frame by frame, use the action recognition network to recognize each frame, and predict the action categories of all individuals in it. By performing the above operations frame by frame on the video, the spatio-temporal action localization result of the video can be obtained.
[0108] In summary, the embodiments of the present invention achieve weakly supervised spatio-temporal action localization through the above steps 201-205 by means of two weakly supervised annotation methods and weakly supervised learning. On the one hand, multiple individuals in a single or multiple video frames are packaged for overall annotation to obtain weakly supervised labels for multiple individual packages. On the other hand, in two stages, anchor individuals and frame interval pseudo-labels are screened and generated based on action prediction values. Due to the characteristics of weakly supervised labels, this method simplifies the data annotation process, reduces the workload of manual annotation, and effectively solves the problem of requiring a large amount of cumbersome manual annotation.
[0109] Embodiment 3
[0110] A spatio-temporal action localization device for multi-object scenarios based on weakly supervised learning, see Figure 2 , the device includes: a processor and a memory, and program instructions are stored in the memory. The processor calls the program instructions stored in the memory to enable the device to execute the following method steps:
[0111] Use a convolutional neural network as a feature extraction network to perform frame-by-frame feature extraction on the video, and perform target object feature extraction on each video frame to output the feature vectors of all individuals through a fully connected layer;
[0112] Calculate the multiple individual package labels of the individual set X M in the original dataset and calculate the action prediction result through the feature vectors of the individuals Calculate the multiple individual package labels and the action prediction result to calculate the cross-entropy loss between them;
[0113] For frame t, find its neighboring frames within a certain range, and screen them with a threshold ξ t , generate pseudo-labels for qualified individuals, examine unlabeled individual x, and calculate its maximum prediction value and the most likely label value to calculate the cross-entropy loss c between them. Screen with a threshold ξ a , ξ c as the standard, and add the qualified individuals as anchor individuals;
[0114] Comprehensively consider the cross-entropy loss function, and use as the total loss function, and use the original training set and the new dataset after pseudo-labeling to update the parameters, and stop after completing the specified number of training rounds;
[0115] Decompose the video to be recognized into several frames, input them frame by frame into the trained network for action recognition, and output the action categories of all individuals in the video frame.
[0116] Among them, calculate the individual set X in the original datasetM Multi - individual package labels And calculate the action prediction result through the feature vector of each individual Specifically:
[0117] Label the whole composed of M individuals X with labels M ={x1, x2, …, x M} as a multi - individual package Multi - individual package Label Is obtained by the following formula:
[0118]
[0119] Where y m Represents the action label of individual x m , and Indicates to perform the OR operation element - by - element on the corresponding elements;
[0120] For the package with M individuals Calculate the action prediction result at the multi - individual - package level K represents the number of action categories, and the prediction result can be obtained by the following formula:
[0121]
[0122] Where φ represents the aggregation function; p m ∈R 1×K Represents the action prediction result of individual x m , which is obtained through the Inception - v3 network with RoIAlign.
[0123] Where the method also includes: training the action prediction result , specifically:
[0124] Calculate the minimum loss
[0125]
[0126] Where, when calculating , the aggregation function φ uses the min - pooling function;
[0127] Calculate the intermediate loss
[0128]
[0129] Where Represents the set of individuals whose predicted values are not among the maximum loss and the minimum loss, and the U function is defined as:
[0130]
[0131] Among them, \(c1, c2\in[0, 1]\), and \(d > 0\) is a preset parameter;
[0132] Finally, the improved multi-agent loss function is expressed as:
[0133]
[0134] Among them, the parameters \(\alpha\) and \(\beta\) are used to balance different situations, and the total loss function is obtained by summing over each package ;
[0135] For the matrix \(P\), an uncertainty is introduced to calculate the multi-agent package loss
[0136]
[0137] Among them, with a threshold \(\xi\) t for screening as:
[0138] For an anchored individual \(q\) t , assuming its label is action \(k\), for the adjacent frame \(t - 1\), if the predicted action category of the same individual \(q\) t-1 is also \(k\), and the corresponding predicted score is greater than the threshold \(\xi\) t , then for the unlabeled individual \(q\) t-1 a pseudo-label with action category \(k\) is generated and added to the training set, and in the same way, the single-frame annotation is extended to multi-frame annotation frame by frame within the interval \([t - R, t + R]\).
[0139] Among them, with a threshold \(\xi\) a , \(\xi\) c as the standard for screening, the threshold is:
[0140] For a certain unlabeled individual \(x\) n , among the predicted scores of all its action categories, the predicted score of category \(k\) is the highest. Calculate the prediction confidence \(c\), where \(c\) is the cross-entropy loss between the predicted value and the most likely label value ; when and only when the predicted value and the confidence \(c\) satisfy the following relationship, add the individual \(x\) n as an anchored individual:
[0141]
[0142] Among them, \(\xi\) a , \(\xi\) c are threshold parameters.
[0143] It should be noted here that the device descriptions in the above embodiments correspond to the method descriptions in the embodiments, and the embodiments of the present invention will not be elaborated herein.
[0144] The execution subjects of the above-mentioned processor and memory can be devices with computing functions such as a computer, a single-chip microcomputer, a microcontroller, etc. In specific implementation, the embodiments of the present invention do not limit the execution subject, and it can be selected according to the needs in actual applications.
[0145] Data signals are transmitted between the memory and the processor through a bus, and the embodiments of the present invention will not elaborate on this.
[0146] Except for those with special descriptions, the embodiments of the present invention do not limit the models of each device, and any device that can perform the above functions can be used.
[0147] Those skilled in the art can understand that the drawings are only schematic diagrams of a preferred embodiment, and the serial numbers of the above embodiments of the present invention are only for description and do not represent the advantages or disadvantages of the embodiments.
[0148] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.
Claims
1. A spatio-temporal action localization method for multi-object scenarios based on weakly supervised learning, characterized in that, The method includes: Using a convolutional neural network as a feature extraction network to perform frame-by-frame feature extraction on the video, and performing target object feature extraction on each video frame to output the feature vectors of all individuals through a fully connected layer; Calculate the individual set X in the original dataset M Multi - individual packet label And calculate the action prediction result through the feature vector of the individual Calculate the multi - individual packet label And the action prediction result Calculate the cross - entropy loss between them; Find neighboring frames within a certain range for frame t, with threshold ξ t Perform screening, generate pseudo-labels for qualified individuals, and examine unlabeled individual x n , calculate its maximum predicted value and the most likely label value to compute the cross-entropy loss c between them, and use thresholds ξ a , ξ c as criteria for screening, and newly add qualified individuals as anchor individuals; Taking into account the cross-entropy loss function, is used as the total loss function, and the parameters are updated using the original training set and the new data set after pseudo-labeling, and the training stops after completing the specified number of rounds; Decomposing the video to be recognized into several frames, inputting them frame by frame into the trained network for action recognition, and outputting the action categories of all individuals in the video frame; The threshold ξ t for screening is as follows: For an anchored individual q t , assuming its label is action k. For the adjacent frame t-1, if the predicted action category of the same individual q t-1 is also k, and the corresponding predicted score is greater than the threshold ξ t , then for the unlabeled individual q t-1 generate a pseudo-label with the action category k and add it to the training set. Use the same method to expand the single-frame annotation into a multi-frame annotation frame by frame within the interval of [t-A, t+A].
2. The multi-object scene spatio-temporal action localization method based on weak supervision learning according to claim 1, characterized in that The individual set X in the original calculation dataset M with multiple individual packet labels and calculates the action prediction result through the feature vectors of the individuals Specifically: Label the whole composed of M individuals X with labels M ={x1, x2, …, x M} as a multi - individual package Multi - individual package The label of is obtained by the following formula: Among them, y m represents the action label of individual x m , indicating to perform the OR operation on each corresponding element one by one; R represents the real number field; For a packet B with M individuals, calculate the action prediction results at the multi - individual packet level K represents the number of action categories, and the prediction result is obtained by the following formula: where φ represents an aggregation function; p m ∈R 1×K represents the action prediction result of individual x m which is obtained through the Inception-v3 network with RoIAlign.
3. A multi-object scenario spatio-temporal action localization method based on weakly supervised learning according to claim 1, characterized in that, The method further includes: training the action prediction result specifically as follows: Calculate the minimum loss Among them, when calculating , the aggregation function φ uses the min-pooling function; K represents the number of action categories; Calculate intermediate losses where p m ∈R 1×K represents the action prediction result of individual x m ; Among them, represents the set of individuals whose predicted values are not among the maximum loss and the minimum loss, and the U function is defined as: wherein, c1, c2 ∈ [0, 1], and d > 0 are preset parameters; Finally, the improved multi-agent loss function is expressed as: Among them, the parameters α and β are used to balance different situations, and the total loss function is obtained by summing over each packet ; is the maximum loss; R represents the real number field; For the matrix P, introduce an uncertainty σ ∈ R to calculate the multi-agent packet loss 4. A multi-object scene spatio-temporal action localization method based on weakly supervised learning according to claim 1, characterized in that The threshold ξ a , ξ c is used as the standard for screening, and the threshold is as follows: For an unlabeled individual x n , among the predicted scores of all its action categories, the predicted score of category k is the highest. Calculate the prediction confidence c, where c is the cross-entropy loss between the predicted value and the most likely label value . If and only if the predicted value and the confidence c satisfy the following relationship, add the individual x n as an anchor individual: Among them, ξ a , ξ c is a threshold parameter.
5. A multi-object scene spatio-temporal action localization device based on weakly supervised learning, characterized in that, The device includes: a processor and a memory. Program instructions are stored in the memory, and the processor calls the program instructions stored in the memory to enable the device to execute the method steps described in any one of claims 1-4.
Citation Information
Patent Citations
Basketball video event and target online detection method based on multiple tasks
CN111639563A
Training method and device of label proportion learning model based on self-supervised learning
CN113139651A