Method and apparatus for detecting timing actions

By introducing a Gaussian time-aware network into timing action detection, dynamically adjusting the time scale, the problem of insufficient robustness caused by fixed time scale in the prior art is solved, and more efficient timing action detection is achieved.

CN112052704BActive Publication Date: 2025-06-20BEIJING JINGDONG SHANGKE INFORMATION TECH CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN201910485368.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2019-06-05
Publication Date
2025-06-20
Estimated Expiration
2039-06-05

AI Technical Summary

Technical Problem

The existing timing action detection methods are not robust enough due to fixed time scales, making it difficult to effectively detect variable action videos.

Method used

The Gaussian time perception network is adopted to dynamically adjust the time scale of action video by introducing a Gaussian kernel, combining the basic feature network and multiple cascaded one-dimensional time convolution layers with Gaussian kernels to perform timing action detection.

Benefits of technology

It improves the accuracy of timing action detection, solves the problem of insufficient robustness caused by fixed time scales, and can more effectively detect variable action videos.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN112052704B_ABST
    Figure CN112052704B_ABST
Patent Text Reader

Abstract

Embodiments of the present application disclose a method and apparatus for detecting temporal actions. A specific implementation of the method includes: obtaining an action video; inputting the action video into a pre-trained Gaussian time perception network to obtain the location information of the temporal actions in the action video, where the Gaussian time perception network dynamically optimizes the time scale of the temporal actions by introducing a Gaussian kernel and is used for performing one-step temporal action detection on the action video, and the location information includes the start time, end time, and action category of the temporal actions. This implementation uses the Gaussian time perception network to detect the temporal actions in the action video, enabling the Gaussian time perception network to dynamically optimize the time scale of the temporal actions by introducing a Gaussian kernel, improving the detection performance of the temporal actions, and thus enhancing the detection accuracy of the temporal actions.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of computer technology, and particularly to methods and apparatuses for detecting temporal actions. Background Art

[0002] With the substantial increase in online and personal media archives, people are generating, storing, and consuming a large amount of videos. In this trend, the development of efficient algorithms for intelligently analyzing video data is encouraged. One of the fundamental challenges for the success of these improvements is the detection of actions in videos in terms of time and space, that is, temporal action detection.

[0003] Temporal action detection is a very challenging problem in the current field of computer video understanding. Most existing temporal action detection methods are inspired by image object detection methods (such as SSD and Faster R-CNN) and are extended to detect the action positions on a one-dimensional temporal sequence. However, these methods have the problem of low robustness because they fix the time scale in advance and ignore the inherent structure of action videos. For variable action videos, detection becomes difficult. Summary of the Invention

[0004] The embodiments of the present application propose methods and apparatuses for detecting temporal actions.

[0005] In a first aspect, the embodiments of the present application provide a method for detecting temporal actions, including: obtaining an action video; inputting the action video into a pre-trained Gaussian time perception network to obtain the positioning information of the temporal actions in the action video, where the Gaussian time perception network dynamically adjusts the time scale of the action video by introducing a Gaussian kernel and is used for one-step temporal action detection of the action video, and the positioning information includes the start time, end time, and action category of the temporal actions.

[0006] In some embodiments, the Gaussian time perception network includes a basic feature network and multiple cascaded one-dimensional temporal convolutional layers with Gaussian kernels.

[0007] In some embodiments, inputting the action video into a pre-trained Gaussian time perception network to obtain the positioning information of the temporal actions in the action video includes: inputting the action video into the basic feature network to obtain the feature map of the action video; inputting the feature map of the action video into multiple cascaded one-dimensional temporal convolutional layers with Gaussian kernels to obtain the positioning information of the temporal actions in the action video.

[0008] In some embodiments, the basic feature network includes a three-dimensional convolutional neural network, a one-dimensional convolutional layer, and a max pooling layer.

[0009] In some embodiments, inputting an action video into a basic feature network to obtain a feature map of the action video, including: inputting the action video into a three-dimensional convolutional neural network to obtain features of video segments in the action video; sequentially connecting the features of the video segments in the action video to generate a feature map of the action video; inputting the feature map of the action video into a one-dimensional convolutional layer and a max pooling layer to increase the temporal dimension of the receptive field of the feature map of the action video.

[0010] In some embodiments, inputting the feature map of the action video into multiple cascaded one-dimensional temporal convolutional layers with Gaussian kernels to obtain localization information of temporal actions in the action video, including: inputting the feature map of the action video into multiple cascaded one-dimensional temporal convolutional layers to obtain feature maps with multiple different temporal resolutions; for cells of the feature maps with multiple different temporal resolutions, learning Gaussian kernels to predict the temporal scales of action nominations corresponding to the cells; aggregating the context features of the cells to generate aggregated features of action nominations corresponding to the cells; inputting the aggregated features of action nominations corresponding to the cells into multiple parallel one-dimensional convolutional layers respectively to obtain scores of action categories, localization parameters, and overlap parameters of action nominations corresponding to the cells; minimizing the action classification loss and the regression loss to generate the localization information of temporal actions in the action video.

[0011] In some embodiments, after learning Gaussian kernels to predict the temporal scales of action nominations corresponding to the cells, it further includes: determining whether there are Gaussian kernels that highly overlap with each other in the Gaussian kernels; if there are Gaussian kernels that highly overlap with each other, using a Gaussian kernel aggregation algorithm to merge the Gaussian kernels that highly overlap with each other into a mixture Gaussian kernel to predict the temporal scale of a long action nomination corresponding to the cell.

[0012] In some embodiments, determining whether there are Gaussian kernels that highly overlap with each other in the Gaussian kernels includes: calculating the temporal intersection and temporal union of the overlapping Gaussian kernels in the Gaussian kernels; calculating the intersection-over-union ratio of the overlapping Gaussian kernels based on the length of the temporal intersection of the overlapping Gaussian kernels and the length of the temporal union of the overlapping Gaussian kernels; determining whether the intersection-over-union ratio of the overlapping Gaussian kernels is greater than a preset threshold; if it is greater than the preset threshold, determining that the overlapping Gaussian kernels are Gaussian kernels that highly overlap with each other; if it is not greater than the preset threshold, determining that the overlapping Gaussian kernels are not Gaussian kernels that highly overlap with each other.

[0013] In some embodiments, the regression loss includes a localization loss and an overlap loss.

[0014] Second aspect, an embodiment of the present application provides a device for detecting temporal actions, including: an acquisition unit configured to acquire an action video; a detection unit configured to input the action video into a pre-trained Gaussian time perception network to obtain the location information of the temporal actions in the action video, where the Gaussian time perception network dynamically adjusts the time scale of the action video by introducing a Gaussian kernel and is used to perform one-step temporal action detection on the action video, and the location information includes the start time, end time, and action category of the temporal actions.

[0015] In some embodiments, the Gaussian time perception network includes a basic feature network and multiple cascaded one-dimensional temporal convolutional layers with Gaussian kernels.

[0016] In some embodiments, the detection unit includes: an extraction subunit configured to input the action video into the basic feature network to obtain a feature map of the action video; a localization subunit configured to input the feature map of the action video into multiple cascaded one-dimensional temporal convolutional layers with Gaussian kernels to obtain the location information of the temporal actions in the action video.

[0017] In some embodiments, the basic feature network includes a three-dimensional convolutional neural network, a one-dimensional convolutional layer, and a max pooling layer.

[0018] In some embodiments, the extraction subunit includes: an extraction module configured to input the action video into the three-dimensional convolutional neural network to obtain the features of the video segments in the action video; a connection module configured to sequentially connect the features of the video segments in the action video to generate a feature map of the action video; an enhancement module configured to input the feature map of the action video into the one-dimensional convolutional layer and the max pooling layer to enhance the temporal dimension of the receptive field of the feature map of the action video.

[0019] In some embodiments, the localization subunit includes: a cascaded convolutional module configured to input the feature map of the action video into multiple cascaded one-dimensional temporal convolutional layers to obtain feature maps with multiple different temporal resolutions; a learning module configured to learn a Gaussian kernel for the cells of the feature maps with multiple different temporal resolutions to predict the time scale of the action nominations corresponding to the cells; an aggregation module configured to aggregate the context features of the cells to generate the aggregated features of the action nominations corresponding to the cells; a parallel convolutional module configured to input the aggregated features of the action nominations corresponding to the cells into multiple parallel one-dimensional convolutional layers to obtain the scores of the action categories, localization parameters, and overlap parameters of the action nominations corresponding to the cells; a localization module configured to minimize the action classification loss and the regression loss to generate the location information of the temporal actions in the action video.

[0020] In some embodiments, the positioning subunit further includes: a determination module configured to determine whether there are Gaussian kernels that highly overlap with each other in the Gaussian kernels; a merging module configured to, if there are Gaussian kernels that highly overlap with each other, use a Gaussian kernel aggregation algorithm to merge the Gaussian kernels that highly overlap with each other into a mixture Gaussian kernel to predict the time scale of the long action nomination corresponding to the cell.

[0021] In some embodiments, the determination module is further configured to: calculate the time intersection and time union of the overlapping Gaussian kernels in the Gaussian kernels; calculate the intersection-over-union ratio of the overlapping Gaussian kernels based on the length of the time intersection of the overlapping Gaussian kernels and the length of the time union of the overlapping Gaussian kernels; determine whether the intersection-over-union ratio of the overlapping Gaussian kernels is greater than a preset threshold; if it is greater than the preset threshold, determine that the overlapping Gaussian kernels are Gaussian kernels that highly overlap with each other; if it is not greater than the preset threshold, determine that the overlapping Gaussian kernels are not Gaussian kernels that highly overlap with each other.

[0022] In some embodiments, the regression loss includes a positioning loss and an overlap loss.

[0023] In a third aspect, an embodiment of the present application provides a server, which includes: one or more processors; a storage device storing one or more programs thereon; when the one or more programs are executed by the one or more processors, the one or more processors are caused to implement the method described in any implementation manner of the first aspect.

[0024] In a fourth aspect, an embodiment of the present application provides a computer-readable medium storing a computer program thereon, and when the computer program is executed by a processor, the method described in any implementation manner of the first aspect is implemented.

[0025] The method and device for detecting temporal actions provided by the embodiments of the present application first obtain an action video; then input the action video into a pre-trained Gaussian time perception network to obtain the positioning information of the temporal actions in the action video. Detecting the temporal actions in the action video by using the Gaussian time perception network enables the Gaussian time perception network to dynamically optimize the time scale of the temporal actions by introducing Gaussian kernels, improve the detection performance of the temporal actions, and thus improve the detection accuracy of the temporal actions. Moreover, the problem of insufficient robustness caused by a fixed time scale in the prior art is solved. BRIEF DESCRIPTION OF THE DRAWINGS

[0026] By reading the detailed description of the non-limiting embodiments with reference to the following drawings, other features, objects, and advantages of the present application will become more apparent:

[0027] Figure 1 is an exemplary system architecture to which the present application can be applied;

[0028] Figure 2 is a flowchart of an embodiment of a method for detecting temporal actions according to the present application;

[0029] Figure 3 is a flowchart of another embodiment of a method for detecting temporal actions according to the present application;

[0030] Figure 4 is a flowchart of another embodiment of a method for detecting temporal actions according to the present application;

[0031] Figure 5 is a schematic structural diagram of a Gaussian time perception network;

[0032] Figure 6 is a schematic diagram of a Gaussian aggregation process;

[0033] Figure 7 is a comparison diagram of the action localization processes of an extended image object detection method and a Gaussian time perception network;

[0034] Figure 8 is a schematic structural diagram of an embodiment of an apparatus for detecting temporal actions according to the present application;

[0035] Figure 9 is a schematic structural diagram of a computer system of a server suitable for implementing the embodiments of the present application. Detailed implementation manners

[0036] The present application will be further described in detail below with reference to the accompanying drawings and embodiments. It can be understood that the specific embodiments described herein are only used to explain the related invention and are not intended to limit the invention. Additionally, it should be noted that for the sake of description, only parts related to the relevant invention are shown in the drawings.

[0037] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments can be combined with each other. The present application will be described in detail below with reference to the drawings and embodiments.

[0038] Figure 1 Illustrates an exemplary system architecture 100 to which the embodiments of the method for detecting temporal actions or the apparatus for detecting temporal actions according to the present application can be applied.

[0039] As Figure 1 shown, the system architecture 100 may include a video capture device 101, a network 102, and a server 103. The network 102 is used to provide a medium for a communication link between the video capture device 101 and the server 103. The network 102 may include various connection categories, such as wired, wireless communication links, or fiber optic cables, etc.

[0040] The video acquisition device 101 can send the action video it captures to the server 103 via the network 102. The video acquisition device 101 can be hardware or software. When the video acquisition device 101 is hardware, it can be various electronic devices supporting video acquisition functions, including but not limited to cameras, video cameras, cameras, and smartphones, etc. When the video acquisition device 101 is software, it can be installed in the above-mentioned electronic devices. It can be implemented as multiple software or software modules, or as a single software or software module. No specific limitation is made here.

[0041] The server 103 can be a server providing various services, such as a temporal action detection server. The temporal action detection server can analyze and process data such as the obtained action video, and generate a processing result (such as the location information of the temporal action in the action video).

[0042] It should be noted that the server 103 can be hardware or software. When the server 103 is hardware, it can be implemented as a distributed server cluster composed of multiple servers, or as a single server. When the server 103 is software, it can be implemented as multiple software or software modules (such as for providing distributed services), or as a single software or software module. No specific limitation is made here.

[0043] It should be noted that the method for detecting temporal actions provided by the embodiments of the present application is generally executed by the server 103. Correspondingly, the device for detecting temporal actions is generally arranged in the server 103.

[0044] It should be understood that Figure 1 the numbers of the video acquisition device, the network, and the server in

[0045] are merely illustrative. According to the implementation requirements, there can be any number of video acquisition devices, networks, and servers.

[0045] Continuing to refer to Figure 2 , it shows a flow 200 of an embodiment of the method for detecting temporal actions according to the present application. The method for detecting temporal actions includes the following steps:

[0046] Step 201, obtain an action video.

[0047] In this embodiment, the execution subject of the method for detecting temporal actions (such as Figure 1 the server 103 shown) can obtain the action video captured by the video acquisition device (such as Figure 1 the video acquisition device 101 shown). Among them, the action video can be a video obtained by shooting the movement process of the target. The target can include but not limited to people, animals, vehicles, etc.

[0048] Step 202: Input the action video into a pre-trained Gaussian Temporal Awareness Network to obtain the localization information of the temporal actions in the action video.

[0049] In this embodiment, the above-mentioned execution entity can input the action video into a pre-trained Gaussian Temporal Awareness Network to obtain the localization information of the temporal actions in the action video. Among them, an action video usually includes a set of temporal actions. The localization information may include information such as the start time, end time, and action category of each temporal action.

[0050] In this embodiment, the Gaussian Temporal Awareness Network (GTAN) can be used for one-step temporal action detection of the action video. Generally, the Gaussian Temporal Awareness Network can dynamically optimize the time scale of each temporal action by introducing a Gaussian kernel. Specifically, the Gaussian Temporal Awareness Network can integrate the exploration of the action structure into a one-step temporal action detection framework and model the inherent structure of the action video by learning a series of Gaussian kernels. Among them, each Gaussian kernel can correspond to a specific action segment. In addition, by fusing multiple Gaussian kernels, it can also be used to detect action videos with time-varying characteristics.

[0051] The method for detecting temporal actions provided by the embodiments of the present application first obtains an action video; then inputs the action video into a pre-trained Gaussian Temporal Awareness Network to obtain the localization information of the temporal actions in the action video. Using the Gaussian Temporal Awareness Network to detect the temporal actions in the action video enables the Gaussian Temporal Awareness Network to dynamically optimize the time scale of the temporal actions by introducing a Gaussian kernel, improving the detection performance of the temporal actions, thereby improving the detection accuracy of the temporal actions. And it solves the problem of insufficient robustness caused by the fixed time scale in the prior art.

[0052] For further reference Figure 3 , which shows the flow 300 of another embodiment of the method for detecting temporal actions according to the present application.

[0053] In this embodiment, the Gaussian Temporal Awareness Network may include a basic feature network and multiple cascaded one-dimensional temporal convolutional layers with Gaussian kernels.

[0054] In this embodiment, the method for detecting temporal actions may include the following steps:

[0055] Step 301: Obtain an action video.

[0056] In this embodiment, the specific operation of step 301 has been introduced in detail in step 201 of the embodiment shown in Figure 2 and will not be elaborated here.

[0057] Step 302: Input the action video into the basic feature network to obtain the feature map of the action video.

[0058] In this embodiment, the execution subject of the method for detecting temporal actions (such as Figure 1 the server 103 shown) can input the action video into the basic feature network to obtain the feature map of the action video. Among them, the basic feature network can be used to extract the feature map of the action video from the video segments of the action video.

[0059] Step 303: Input the feature map of the action video into multiple cascaded one-dimensional temporal convolutional layers with Gaussian kernels to obtain the localization information of the temporal actions in the action video.

[0060] In this embodiment, the above execution subject can input the feature map of the action video into multiple cascaded one-dimensional temporal convolutional layers with Gaussian kernels to obtain the localization information of the temporal actions in the action video. Among them, the multiple cascaded one-dimensional temporal convolutional layers with Gaussian kernels can dynamically optimize the time scale of each temporal action to obtain the localization information of each temporal action in the action video.

[0061] As can be seen from Figure 3 , compared with the corresponding embodiment, the process 300 of the method for detecting temporal actions in this embodiment highlights the steps of detecting the action video. Thus, the Gaussian time perception network in the solution described in this embodiment can include a basic feature network and multiple cascaded one-dimensional temporal convolutional layers with Gaussian kernels. By introducing Gaussian kernels using multiple cascaded one-dimensional convolutional layers, a way of introducing Gaussian kernels is provided. Figure 2

[0062] Further referring to Figure 4 , which shows the process 400 of another embodiment of the method for detecting temporal actions according to the present application.

[0063] In this embodiment, the Gaussian time perception network can include a basic feature network and multiple cascaded one-dimensional temporal convolutional layers with Gaussian kernels. Among them, the basic feature network can include a three-dimensional convolutional neural network, a one-dimensional convolutional layer, and a max pooling layer. For ease of understanding, Figure 5 shows a schematic structural diagram of the Gaussian time perception network. Among them, Figure 5 the left half of shows the basic feature network, and the basic feature network can include 4 three-dimensional convolutional neural networks (3D CNNs), two one-dimensional convolutional layers (conv1 and conv2), and one max pooling layer (pool1). Figure 5 the right half of shows 8 cascaded one-dimensional temporal convolutional layers with Gaussian kernels (conv_a1, conv_a2,..., conv_a8).​

[0064] In this embodiment, the method for detecting temporal actions may include the following steps:

[0065] Step 401, obtain an action video.

[0066] In this embodiment, the specific operation of step 401 has been introduced in detail in step 201 of the embodiment shown in Figure 2 and will not be elaborated here.

[0067] Step 402, input the action video into a three-dimensional convolutional neural network to obtain the features of the video segments in the action video.

[0068] In this embodiment, the execution subject of the method for detecting temporal actions (such as Figure 1 the server 103 shown) may input the action video into a three-dimensional convolutional neural network to obtain the features of the video segments in the action video. Among them, the three-dimensional convolutional neural network can extract the segment-level features of the action video from the video segments of the action video.

[0069] In practice, the ultimate goal of temporal action detection is to detect actions in the time dimension. Specifically, for the input action video, the above execution subject may first divide the action video into continuous video segments; then input the continuous video segments into a three-dimensional convolutional neural network to obtain the features of each video segment, that is, segment-level features.

[0070] In some alternative implementation manners of this embodiment, the extraction of segment-level features from the action video by the three-dimensional convolutional neural network can be represented by the following feature sequence:

[0071]

[0072] where T is the time length of the action video. Here, the action video is divided into T continuous video segments and numbered from 0 to T - 1 in sequence. 0 ≤ i ≤ T - 1, and i is an integer, representing the video segments numbered from 0 to T - 1. f i is the feature of the video segment numbered i.

[0073] Step 403, sequentially connect the features of the video segments in the action video to generate a feature map of the action video.

[0074] In this embodiment, the above execution subject may sequentially connect the features of the video segments in the action video to generate a feature map of the action video.

[0075] Step 404, input the feature map of the action video into a one-dimensional convolutional layer and a max pooling layer to increase the temporal size of the receptive field of the feature map of the action video.

[0076] In this embodiment, the above-mentioned execution entity may input the feature map of the action video into a one-dimensional convolutional layer and a max pooling layer to increase the temporal dimension of the receptive field of the feature map of the action video.

[0077] In some optional implementation manners of this embodiment, the above-mentioned execution entity may use two one-dimensional convolutional layers plus a max pooling layer to increase the temporal dimension of the receptive field. Among them, the temporal kernel sizes of the two one-dimensional convolutional layers are both 3, and the strides are both 1. The temporal kernel size of a max pooling layer is 3, and the stride is 2.

[0078] Step 405: Input the feature map of the action video into multiple cascaded one-dimensional temporal convolutional layers to obtain feature maps with multiple different temporal resolutions.

[0079] In this embodiment, the above-mentioned execution entity may input the feature map of the action video into multiple cascaded one-dimensional temporal convolutional layers to obtain feature maps with multiple different temporal resolutions.

[0080] In some optional implementation manners of this embodiment, if 8 one-dimensional temporal convolutional layers are cascaded, then the feature map of the j-th convolutional layer can be represented by the following feature sequence:

[0081]

[0082] Among them, 1 ≤ j ≤ 8, and j is an integer, representing from the 1st to the 8th convolutional layer. Divide the feature map of the j-th convolutional layer into T j cells, and label them as 0 to T j -1 in sequence. 0 ≤ i ≤ T j -1, and i is an integer, representing the cells from label 0 to T j -1. f i is the feature of the cell with label i. T j is the temporal scale of the feature map of the j-th convolutional layer, and D j is the feature scale of the feature map of the j-th convolutional layer. is a vector group composed of all T j ×D j dimensional vectors.

[0083] Step 406: For the cells of the feature maps with multiple different temporal resolutions, learn Gaussian kernels to predict the temporal scale of the action nomination corresponding to the cells.

[0084] In this embodiment, for the cells of the feature maps with multiple different temporal resolutions, the above-mentioned execution entity may learn Gaussian kernels to predict the temporal scale of the action nomination corresponding to the cells. Among them, the action nomination focuses on locating the video segment containing the action. Each Gaussian kernel corresponds to a specific action segment.

[0085] In some alternative implementation manners of this embodiment, for an action nomination P with a central position of t t j , its time scale is controlled by a Gaussian kernel . The standard deviation of the Gaussian kernel is learned by a one-dimensional temporal convolutional layer on a 3*D j feature cell, and then its value is constrained within the range (0, 1) through a sigmoid operation. Among them, the weight of the Gaussian kernel is learned through the following formula:

[0086]

[0087] where i ∈ {0, 1,..., T j -1}, t ∈ {0, 1,..., T j -1}. Z is the normalization constant.

[0088] Inspired by this theory, the standard deviation can be regarded as the width measurement (root mean square width, RMS) of the Gaussian kernel , so is used for the interval measurement of the action nomination P t j . For example, can be multiplied by a certain ratio to represent the default time boundary:

[0089] a c = (t + 0.5) / T j ;

[0090]

[0091] where a c is the central position of the default time boundary, a w is the width of the default time boundary, and r d is the time scale ratio. For example, r d can be set to 2 0 , 2 1 / 3 and 2 2 / 3 .

[0092] Step 407, determine whether there are Gaussian kernels that highly overlap with each other in the Gaussian kernel.

[0093] In this embodiment, the above-mentioned execution entity can determine whether there are Gaussian kernels that highly overlap with each other in the Gaussian kernels. If there are Gaussian kernels that highly overlap with each other, it indicates that these Gaussian kernels that highly overlap with each other correspond to a long action. At this time, step 408 is continued; otherwise, step 409 is directly executed.

[0094] In practice, by learning Gaussian kernels, the time scales of most action nominations can be characterized by the predicted standard deviations. However, if the learned Gaussian kernels cross and overlap with each other, these Gaussian kernels correspond to a long action. At this time, a new set of Gaussian kernels needs to be generated to predict the central position and time scale of this long action.

[0095] In some optional implementation manners of this embodiment, the above-mentioned execution entity can first calculate the time intersection and time union of the overlapping Gaussian kernels in the Gaussian kernels; then, based on the length of the time intersection of the overlapping Gaussian kernels and the length of the time union of the overlapping Gaussian kernels, calculate the intersection-over-union ratio of the overlapping Gaussian kernels; finally, determine whether the intersection-over-union ratio of the overlapping Gaussian kernels is greater than a preset threshold; if it is greater than the preset threshold, it is determined that the overlapping Gaussian kernels are Gaussian kernels that highly overlap with each other, and these overlapping Gaussian kernels correspond to a long action; if it is not greater than the preset threshold, it is determined that the overlapping Gaussian kernels are not Gaussian kernels that highly overlap with each other, and these overlapping Gaussian kernels respectively correspond to an action.

[0096] In some optional implementation manners of this embodiment, given two adjacent Gaussian kernels G(t1,σ1) and G(t2,σ2), first calculate the widths of the default time boundaries corresponding to the Gaussian kernels G(t1,σ1) and G(t2,σ2) respectively; then, according to the widths of the default time boundaries corresponding to the Gaussian kernels G(t1,σ1) and G(t2,σ2), calculate the time intersection and time union between the Gaussian kernels G(t1,σ1) and G(t2,σ2); then, based on the length H of the time intersection and the length L of the time union between the Gaussian kernels G(t1,σ1) and G(t2,σ2), calculate the intersection-over-union ratio H / L between the Gaussian kernels G(t1,σ1) and G(t2,σ2); if the intersection-over-union ratio H / L exceeds the preset threshold ε (for example, 0.7), then the Gaussian kernels G(t1,σ1) and G(t2,σ2) highly overlap with each other. At this time, step 408 is continued to generate new Gaussian kernels. Wherein, t1 is the central position of the Gaussian kernel G(t1,σ1), and σ1 is the standard deviation of the Gaussian kernel G(t1,σ1). t2 is the central position of the Gaussian kernel G(t2,σ2), and σ2 is the standard deviation of the Gaussian kernel G(t2,σ2).

[0097] Step 408: Use the Gaussian kernel aggregation algorithm to merge the Gaussian kernels that highly overlap with each other into a mixture Gaussian kernel to predict the time scale of the long action nomination corresponding to the cell.

[0098] In this embodiment, if there are Gaussian kernels that highly overlap with each other in the Gaussian kernel, the Gaussian kernel aggregation algorithm is used to merge the Gaussian kernels that highly overlap with each other into a mixed Gaussian kernel to predict the time scale of the long action nomination corresponding to the cell.

[0099] In some alternative implementation manners of this embodiment, the mixed Gaussian kernel can be generated by merging through the following formula:

[0100]

[0101] Among them,

[0102] For ease of understanding, Figure 6 a schematic diagram of the Gaussian aggregation process is shown. Among them, Figure 6 the upper half of shows the Gaussian curves corresponding to two adjacent Gaussian kernels G(t1,σ1) and G(t2,σ2). The length of the time intersection between the Gaussian kernels G(t1,σ1) and G(t2,σ2) is H, and the length of the time union between the Gaussian kernels G(t1,σ1) and G(t2,σ2) is L. If the intersection-to-union ratio H / L between the Gaussian kernels G(t1,σ1) and G(t2,σ2) exceeds a preset threshold ε, the Gaussian kernels G(t1,σ1) and G(t2,σ2) are merged into a mixed Gaussian kernel. Figure 6 the lower half of shows the Gaussian curve corresponding to the mixed Gaussian kernel.

[0103] Step 409: Aggregate the context features of the cell to generate the aggregated features of the action nomination corresponding to the cell.

[0104] In this embodiment, the above-mentioned execution subject can aggregate the context features of the cell to generate the aggregated features of the action nomination corresponding to the cell.

[0105] In practice, the above-mentioned execution subject can use learning and the mixed Gaussian kernel to calculate the sum of features weighted by the values on the Gaussian curve to obtain the aggregated features.

[0106] In some alternative implementation manners of this embodiment, given the Gaussian kernel with the weighting coefficient W t j at the central position t of the feature map of the j-th convolutional layer, t j the aggregated features of the action nomination P

[0107]

[0108] Step 410: Input the aggregated features of the action nominations corresponding to the cells into multiple parallel one-dimensional convolutional layers respectively to obtain the scores of the action categories, the localization parameters, and the overlap parameters of the action nominations corresponding to the cells.

[0109] In this embodiment, the above-mentioned execution entity may input the aggregated features of the action nominations corresponding to the cells into multiple parallel one-dimensional convolutional layers respectively to obtain the scores of the action categories, the localization parameters, and the overlap parameters of the action nominations corresponding to the cells. Generally, the above-mentioned execution entity may use three one-dimensional convolutional layers in parallel to predict the scores of the action categories, the localization parameters, and the overlap parameters respectively.

[0110] In some alternative implementation manners of this embodiment, the scores of the action categories may represent the probabilities of belonging to each of the C action categories and the probability of belonging to the background. The localization parameters (Δc, Δw) may represent the central position a c relative to the default time boundary and the width a w of the default time boundary, and the time offset thereof adjusts the time coordinates through the following formula:

[0111]

[0112]

[0113] where is the central position of the action nomination, is the width of the action nomination. α1 and α2 can be used to control the influence of the time offset. For example, α1 and α2 can be set to 1.0.

[0114] Step 411: Minimize the action classification loss and the regression loss to generate the localization information of the temporal actions in the action video.

[0115] In this embodiment, the above-mentioned execution entity may generate the localization information of the temporal actions in the action video by minimizing the action classification loss and the regression loss. Generally, the regression loss may include the localization loss and the overlap loss.

[0116] In some alternative implementation manners of this embodiment, the above-mentioned execution entity may integrate the action classification loss L cls , the localization loss L loc , and the overlap loss L ov through the following formula to obtain the multi-task loss:

[0117] L = L cls + βL loc + γL ov ;

[0118] Among them, β and γ are trade-off parameters. For example, β can be set to 2.0 and γ can be set to 75.

[0119] The above-mentioned execution entity can measure the classification loss through the softmax loss:

[0120]

[0121] Among them, if n is equal to the action label C of the labeled region, the index function I n=c = 1; otherwise, I n=c = 0.

[0122] The above-mentioned execution entity can calculate the localization loss through the following formula:

[0123]

[0124] Among them, g iou is the intersection over union between the default time boundary of the action nomination and the closest labeled region corresponding to it. If g iou is greater than 0.8, the action nomination is set as a foreground sample; if g iou is less than 0.3, the action nomination is set as a background sample. During training, the ratio between the foreground sample and the background sample is set to 1.0. S L1 is the smooth loss between the predicted foreground action nomination and the closest labeled region corresponding to it. g c is the center position of the closest labeled region corresponding to the action nomination, and g w is the width of the closest labeled region corresponding to the action nomination.

[0125] The above-mentioned execution entity can use the mean squared error (MSE) loss to optimize the overlap loss:

[0126] L ov = (y ov - g iou ) 2 .

[0127] Finally, by minimizing the three loss functions, the entire network is trained in an end-to-end manner.

[0128] During the prediction of action localization, the final ranking score y f of each action nomination depends on the score y a of the action category and the overlap parameter y ov :

[0129] y f = max(y a ) · y ov .

[0130] Given a predicted action nomination with a fine boundary and a refined boundary (predicted by action label C a and ranking score y f ), post-processing is performed using soft-NMS. In each iteration of soft-NMS, the action nomination with the maximum ranking score y fm is represented as φ m . Whether the ranking scores y k of other action nominations φ fk are reduced depends on the intersection over union calculated using φ m :

[0131]

[0132] where ξ is the decay parameter and ρ is the NMS threshold. For example, ξ can be set to 0.8 and ρ can be set to 0.75.

[0133] Continue to refer to Figure 7 , which shows a comparison diagram of the action localization processes of the extended image object detection method and the Gaussian time perception network. Among them, Figure 7 the upper half shows the action localization process of the extended image object detection method. Specifically, first, the frame-level or segment-level features in the action video are aggregated into a feature map; then, multiple one-dimensional temporal convolutional layers (1D conv) are designed to increase the size of the temporal receptive field and predict action nominations. However, since the temporal scale of each cell in the feature map is fixed, this method cannot capture the inherent temporal structure of the action. Therefore, in this case, one action in the thick box is detected as three. Figure 7 the lower half shows the action localization process of the Gaussian time perception network. Specifically, by learning Gaussian kernels for each cell to explore the temporal structure of the action, the specific interval of the action is dynamically determined. Different Gaussian kernels can even be aggregated to represent a long action. Thus, a way to localize actions of different lengths is provided. More importantly, the context information of the action is also included along with the feature pooling of the Gaussian curve.

[0134] As can be seen from Figure 4 , compared with the corresponding embodiment in Figure 2 , the process 200 of the method for detecting temporal actions in this embodiment highlights the steps of detecting the action video. Thus, the solution described in this embodiment solves the problem of insufficient robustness caused by the fixed temporal scale in the prior art by introducing Gaussian kernels to dynamically optimize and accurately predict the temporal scale of action nominations. At the same time, a Gaussian kernel aggregation algorithm is introduced to accurately localize long actions by merging Gaussian kernels.

[0135] Further refer toFigure 8 , as an implementation of the methods shown in the above figures, the present application provides an embodiment of a device for detecting temporal actions. This device embodiment corresponds to Figure 2 the method embodiment shown, and this device can be specifically applied to various electronic devices.

[0136] As shown in Figure 8 , the device 800 for detecting temporal actions in this embodiment may include: an acquisition unit 801 and a detection unit 802. Among them, the acquisition unit 801 is configured to acquire an action video; the detection unit 802 is configured to input the action video into a pre-trained Gaussian time perception network to obtain the location information of the temporal actions in the action video. The Gaussian time perception network dynamically optimizes the time scale of the temporal actions by introducing a Gaussian kernel and is used for one-step temporal action detection of the action video. The location information includes the start time, end time, and action category of the temporal actions.

[0137] In this embodiment, in the device 800 for detecting temporal actions: the specific processing of the acquisition unit 801 and the detection unit 802 and the technical effects brought by them can respectively refer to Figure 2 the relevant descriptions of step 201 and step 202 in the corresponding embodiment, which will not be elaborated here.

[0138] In some optional implementation manners of this embodiment, the Gaussian time perception network includes a basic feature network and multiple cascaded one-dimensional time convolutional layers with Gaussian kernels.

[0139] In some optional implementation manners of this embodiment, the detection unit 802 includes: an extraction subunit (not shown in the figure), configured to input the action video into the basic feature network to obtain the feature map of the action video; a location subunit (not shown in the figure), configured to input the feature map of the action video into multiple cascaded one-dimensional time convolutional layers with Gaussian kernels to obtain the location information of the temporal actions in the action video.

[0140] In some optional implementation manners of this embodiment, the basic feature network includes a three-dimensional convolutional neural network, a one-dimensional convolutional layer, and a max pooling layer.

[0141] In some optional implementation manners of this embodiment, the extraction subunit includes: an extraction module (not shown in the figure), configured to input the action video into the three-dimensional convolutional neural network to obtain the features of the video segments in the action video; a connection module (not shown in the figure), configured to sequentially connect the features of the video segments in the action video to generate the feature map of the action video; an increasing module (not shown in the figure), configured to input the feature map of the action video into the one-dimensional convolutional layer and the max pooling layer to increase the time dimension of the receptive field of the feature map of the action video.

[0142] In some alternative implementation manners of this embodiment, the positioning subunit includes: a cascaded convolution module (not shown in the figure), configured to input a feature map of an action video into a plurality of cascaded one-dimensional temporal convolution layers to obtain feature maps with multiple different temporal resolutions; a learning module (not shown in the figure), configured to learn Gaussian kernels for cells of the feature maps with multiple different temporal resolutions to predict the temporal scales of action nominations corresponding to the cells; an aggregation module (not shown in the figure), configured to aggregate the context features of the cells to generate aggregated features of action nominations corresponding to the cells; a parallel convolution module (not shown in the figure), configured to input the aggregated features of action nominations corresponding to the cells into a plurality of parallel one-dimensional convolution layers respectively to obtain scores of action categories, positioning parameters, and overlap parameters of action nominations corresponding to the cells; a positioning module (not shown in the figure), configured to minimize an action classification loss and a regression loss to generate positioning information of temporal actions in the action video.

[0143] In some alternative implementation manners of this embodiment, the positioning subunit further includes: a determination module (not shown in the figure), configured to determine whether there are Gaussian kernels that highly overlap with each other in the Gaussian kernels; a merging module (not shown in the figure), configured to, if there are Gaussian kernels that highly overlap with each other, use a Gaussian kernel aggregation algorithm to merge the Gaussian kernels that highly overlap with each other into a mixture Gaussian kernel to predict the temporal scale of a long action nomination corresponding to the cell.

[0144] In some alternative implementation manners of this embodiment, the determination module is further configured to: calculate the temporal intersection and temporal union of the overlapping Gaussian kernels in the Gaussian kernels; calculate the intersection-over-union ratio of the overlapping Gaussian kernels based on the length of the temporal intersection of the overlapping Gaussian kernels and the length of the temporal union of the overlapping Gaussian kernels; determine whether the intersection-over-union ratio of the overlapping Gaussian kernels is greater than a preset threshold; if it is greater than the preset threshold, determine that the overlapping Gaussian kernels are Gaussian kernels that highly overlap with each other; if it is not greater than the preset threshold, determine that the overlapping Gaussian kernels are not Gaussian kernels that highly overlap with each other.

[0145] In some alternative implementation manners of this embodiment, the regression loss includes a positioning loss and an overlap loss.

[0146] Reference is made below to Figure 9 , which shows a schematic structural diagram of a computer system 900 of a server (such as Figure 1 the server 103 shown) suitable for implementing the embodiments of the present application. Figure 9 The server shown is only an example and should not bring any limitation to the functions and usage scopes of the embodiments of the present application.

[0147] As shown in Figure 9As shown, computer system 900 includes a central processing unit (CPU) 901, which can perform various appropriate actions and processes according to programs stored in a read-only memory (ROM) 902 or programs loaded from a storage section 908 into a random access memory (RAM) 903. In the RAM 903, various programs and data required for the operation of the system 900 are also stored. The CPU 901, ROM 902, and RAM 903 are connected to each other via a bus 904. An input / output (I / O) interface 905 is also connected to the bus 904.

[0148] The following components are connected to the I / O interface 905: an input section 906 including a keyboard, a mouse, etc.; an output section 907 including, for example, a cathode ray tube (CRT), a liquid crystal display (LCD), etc. and a speaker, etc.; a storage section 908 including a hard disk, etc.; and a communication section 909 including a network interface card such as a LAN card, a modem, etc. The communication section 909 performs communication processing via a network such as the Internet. A drive 910 is also connected to the I / O interface 905 as needed. A removable medium 911, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 910 as needed so that a computer program read from it can be installed into the storage section 908 as needed.

[0149] In particular, according to an embodiment of the present disclosure, the processes described above with reference to the flowchart can be implemented as a computer software program. For example, an embodiment of the present disclosure includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program contains program codes for performing the methods shown in the flowchart. In such an embodiment, the computer program can be downloaded and installed from a network via the communication section 909, and / or installed from the removable medium 911. When the computer program is executed by a central processing unit (CPU) 701, the above-described functions defined in the method of the present application are executed.

[0150] It should be noted that the computer-readable medium described in this application can be a computer-readable signal medium, a computer-readable storage medium, or any combination of the two. A computer-readable storage medium can be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination of the above. More specific examples of a computer-readable storage medium can include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In this application, a computer-readable storage medium can be any tangible medium that contains or stores a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. And in this application, a computer-readable signal medium can include a data signal propagated in a baseband or as part of a carrier wave, in which computer-readable program code is carried. Such a propagated data signal can take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. A computer-readable signal medium can also be any computer-readable medium other than a computer-readable storage medium that can send, propagate, or transmit a program for use by or in conjunction with an instruction execution system, apparatus, or device. The program code contained on a computer-readable medium can be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.

[0151] Computer program code for performing the operations of this application can be written in one or more programming languages or combinations thereof. The programming languages include object-oriented programming languages such as Java, Smalltalk, C++, and also include conventional procedural programming languages such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, executed as an independent software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In the case of a remote computer, the remote computer can be connected to the user's computer through any type of network, including a local area network (LAN) or a wide area network (WAN), or can be connected to an external computer (for example, by using an Internet service provider to connect through the Internet).

[0152] The flowcharts and block diagrams in the accompanying drawings illustrate the possible architectures, functions, and operations of systems, methods, and computer program products according to various embodiments of the present application. In this regard, each block in the flowchart or block diagram may represent a module, a segment of a program, or a portion of code that contains one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions noted in the blocks may occur in a different order than that noted in the accompanying drawings. For example, two consecutive blocks shown may actually be executed substantially in parallel, and they may sometimes be executed in the reverse order, depending on the functions involved. It should also be noted that each block in the block diagram and / or flowchart, and combinations of blocks in the block diagram and / or flowchart, can be implemented by a dedicated hardware-based system that performs the specified functions or operations, or can be implemented by a combination of dedicated hardware and computer instructions.

[0153] The units involved in the embodiments described in the present application can be implemented in software or in hardware. The described units can also be provided in a processor. For example, it can be described as: a processor includes an acquisition unit and a detection unit. Among them, the names of these units do not constitute a limitation on the unit itself in some cases. For example, the acquisition unit can also be described as "the unit for acquiring action videos".

[0154] On the other hand, the present application also provides a computer-readable medium, which can be included in the server described in the above embodiments; or it can exist alone and not be assembled into the server. The above computer-readable medium carries one or more programs. When the above one or more programs are executed by the server, the server is caused to: acquire an action video; input the action video into a pre-trained Gaussian time-aware network to obtain the localization information of the temporal actions in the action video, where the Gaussian time-aware network dynamically optimizes the time scale of the temporal actions by introducing a Gaussian kernel and is used for one-step temporal action detection of the action video, and the localization information includes the start time, end time, and action category of the temporal action.

[0155] The above description is only a preferred embodiment of the present application and an explanation of the applied technical principles. Those skilled in the art should understand that the scope of the invention involved in the present application is not limited to the technical solutions formed by the specific combination of the above technical features, and should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the above inventive concept. For example, the technical solutions formed by mutually replacing the above features with the (but not limited to) technical features with similar functions disclosed in the present application.

Claims

1. A method for detecting sequential actions, comprising: Obtain an action video; Input the action video into a pre-trained Gaussian time perception network to obtain the localization information of the temporal action in the action video. Among them, the Gaussian time perception network dynamically optimizes the time scale of the temporal action by introducing a Gaussian kernel and is used for one-step temporal action detection of the action video. The localization information includes the start time, end time, and action category of the temporal action; Among them, the step of inputting the action video into a pre-trained Gaussian time perception network to obtain the localization information of the temporal action in the action video includes: Input the action video into a basic feature network to obtain the feature map of the action video; Input the feature map of the action video into a plurality of cascaded one-dimensional time convolutional layers to obtain feature maps with multiple different time resolutions; For the cells of the feature maps with multiple different time resolutions, learn Gaussian kernels to predict the time scale of the action nomination corresponding to the cells; Aggregate the context features of the cells to generate the aggregated features of the action nomination corresponding to the cells; Input the aggregated features of the action nomination corresponding to the cells into a plurality of parallel one-dimensional convolutional layers respectively to obtain the scores of the action categories, localization parameters, and overlap parameters of the action nomination corresponding to the cells; Minimize the action classification loss and the regression loss to generate the localization information of the temporal action in the action video.

2. The method according to claim 1, wherein, The basic feature network includes a three-dimensional convolutional neural network, a one-dimensional convolutional layer, and a max pooling layer.

3. The method according to claim 2, wherein, The step of inputting the action video into the basic feature network to obtain the feature map of the action video includes: Input the action video into the three-dimensional convolutional neural network to obtain the features of the video segments in the action video; Sequentially connect the features of the video segments in the action video to generate the feature map of the action video; Input the feature map of the action video into the one-dimensional convolutional layer and the max pooling layer to increase the temporal dimension of the receptive field of the feature map of the action video.

4. The method according to claim 1, wherein, After learning Gaussian kernels to predict the time scale of the action nomination corresponding to the cells, it further includes: Determine whether there are Gaussian kernels that highly overlap with each other in the Gaussian kernels; If there are Gaussian kernels that highly overlap with each other, use the Gaussian kernel aggregation algorithm to merge the Gaussian kernels that highly overlap with each other into a mixture Gaussian kernel to predict the time scale of the long action nomination corresponding to the cells.

5. The method according to claim 4, wherein, The step of determining whether there are Gaussian kernels that highly overlap with each other in the Gaussian kernels includes: Calculate the time intersection and time union of the overlapping Gaussian kernels in the Gaussian kernels; Based on the length of the time intersection of the overlapping Gaussian kernels and the length of the time union of the overlapping Gaussian kernels, calculate the intersection over union of the overlapping Gaussian kernels; Determine whether the intersection over union of the overlapping Gaussian kernels is greater than a preset threshold; If it is greater than the preset threshold, determine that the overlapping Gaussian kernels are the Gaussian kernels that highly overlap with each other; If it is not greater than the preset threshold, determine that the overlapping Gaussian kernels are not the Gaussian kernels that highly overlap with each other.

6. The method according to claim 1, wherein, The regression loss includes a localization loss and an overlap loss.

7. A device for detecting sequential actions, comprising: An acquisition unit configured to acquire an action video; A detection unit configured to input the action video into a pre-trained Gaussian temporal perception network to obtain localization information of temporal actions in the action video, wherein the Gaussian temporal perception network dynamically optimizes the time scale of the temporal actions by introducing a Gaussian kernel and is used for performing one-step temporal action detection on the action video, and the localization information includes the start time, end time, and action category of the temporal actions; Wherein, the detection unit is further configured to: Input the action video into a basic feature network to obtain a feature map of the action video; Input the feature map of the action video into a plurality of cascaded one-dimensional temporal convolutional layers to obtain feature maps with multiple different temporal resolutions; For cells of the feature maps with multiple different temporal resolutions, learn a Gaussian kernel to predict the time scale of the action nomination corresponding to the cell; Aggregate the context features of the cells to generate aggregated features of the action nominations corresponding to the cells; Input the aggregated features of the action nominations corresponding to the cells into a plurality of parallel one-dimensional convolutional layers respectively to obtain scores of the action categories, localization parameters, and overlap parameters of the action nominations corresponding to the cells; Minimize the action classification loss and the regression loss to generate the localization information of the temporal actions in the action video.

8. A server, comprising: One or more processors; A storage device having stored thereon one or more programs, When the one or more programs are executed by the one or more processors, the one or more processors implement the method according to any one of claims 1-6.

9. A computer-readable medium having a computer program stored thereon, wherein, When the computer program is executed by a processor, the method according to any one of claims 1-6 is implemented.

Citation Information

Patent Citations

  • Power system fault diagnosis method comprehensively using electricity amount and timing sequence information

    CN104297637A

  • Three-dimensional convolution and Faster RCNN-based video action detection method

    CN108399380A