Time sequence action detection method based on multi-scale image convolutional network

Through the multi-scale graph convolution network method, combined with I3D and CNN networks, multi-scale feature modeling is used to capture the time-space relationship of the action unit, solving the problem of insufficient time-space relationship coordination in the existing methods, and improving the accuracy of timing action detection.

CN120260128APending Publication Date: 2025-07-04YANSHAN UNIV
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510336375.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-21
Publication Date
2025-07-04

AI Technical Summary

Technical Problem

The existing timing action detection method only considers the temporal semantic information of the video, and fails to effectively model the spatial hierarchical relationship between action units, resulting in low detection accuracy.

Method used

A multi-scale graph convolution network is used to extract video features through the I3D network, and after refining it in combination with the CNN network, multi-scale feature modeling is used to use the graph pyramid module to perform graph convolution, including time edges, similar edges and spatial edges, to capture the time and space relationship of the action unit.

Benefits of technology

It improves the accuracy of timing action detection, effectively solves the problem of inability to coordinate time and space relationships, realizes more full modeling of action units, and improves detection effect.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120260128A_ABST
    Figure CN120260128A_ABST
Patent Text Reader

Abstract

The invention discloses a time sequence action detection method based on a multi-scale image convolutional network, and belongs to the field of computer vision, and the method comprises the steps: carrying out the feature extraction of a video through employing an I3D network; further refining the extracted video features by using a CNN (Convolutional Neural Network); the refined video features are sent to a graph pyramid module for down-sampling to obtain video features of different scales, modeling of a time edge, a similar edge and a space edge is carried out on the video features of all the scales, and graph convolution is carried out on the video features after modeling to aggregate context information of the video features of the current scale; the video features after image convolution are sent to a classification head and a regression head for prediction; and integrating action prediction results of different scales to obtain a final time sequence action detection result. According to the invention, the motion unit in the video can be modeled more fully, the time-space relationship of the motion unit is effectively utilized, and the accuracy of time sequence motion detection is greatly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical fields of deep neural networks and computer vision, and in particular to a temporal action detection method based on a multi-scale graph convolutional network. Background Art

[0002] With the development of intelligent systems, video surveillance systems have gradually evolved into intelligent recognition and are widely used in industrial production environments and traffic accidents, such as construction worker safety monitoring and traffic accident detection. Therefore, understanding human behavior in videos has gradually become the focus of the industry.

[0003] Temporal Action Detection (TAD for short, with the full English name being Temporal Action Detection) aims to classify and locate actions in untrimmed videos. Current TAD methods mainly adopt the sliding window method. In order to cover all action predictions as much as possible, the sliding window method will generate a certain number of prediction proposals, and at the same time, it will also generate a large number of redundant predictions, increasing the time complexity. To solve the problem of redundant predictions, some researchers have proposed a boundary matching method to predict the probability and confidence of the boundaries. Although these methods use a bottom-up framework and utilize the duration to obtain possible action proposals by locally sliding around the boundaries, they still result in some redundant action proposals. To further address this problem, some researchers have proposed an adaptive matching method to generate a fixed number of action proposals. These methods use classification and regression heads to directly predict the action positions, saving time to a certain extent.

[0004] Dilated convolutions can further expand the receptive field to aggregate context, so methods based on dilated convolutions have achieved better detection performance in the latest TAD tasks. However, videos have temporal and semantic contexts. Only using dilated convolutions to expand the receptive field to aggregate temporal context ignores the intrinsic connections between action units. Relying solely on temporal context is not sufficient for effective action detection. To aggregate semantic context, it is crucial to represent the semantic relationships between action units. This includes capturing similar semantic relationships and spatial relationships. Similar semantic relationships are the similarities between action units, which are beneficial for understanding action content. For example, for the actions "cliff diving" and "diving", and "cricket bowling" and "cricket shooting", their action units show a high degree of similarity. The process of finding action units with similar semantics usually captures global information. Hierarchical semantic relationships refer to the hierarchical structure between action units in a complete action, which is beneficial for action localization. In other words, the semantics of an action unit are based on previous action units. If the semantics of an action unit change, the semantics of the remaining action units will also change. For example, in "high jump", the body landing action unit must be based on the jump. The process of finding action units with spatial semantics usually captures local modeling. In summary, it is particularly important to accurately capture the different relationships between action units for local and global modeling of video features.

[0005] Based on the above analysis, it can be seen that existing temporal action detection methods only consider the temporal semantic information of videos and do not consider modeling the spatial hierarchy of action units.

[0006] Therefore, it is necessary to provide a temporal action detection method based on a multi-scale graph convolutional network to solve the above problems. Summary of the Invention

[0007] The technical problem to be solved by the present invention is to provide a temporal action detection method based on a multi-scale graph convolutional network, which deeply analyzes and utilizes the spatio-temporal relationships of different action units, more fully models the action units in the video, effectively utilizes the spatio-temporal relationships of the action units, and thus greatly improves the accuracy of temporal action detection.

[0008] To solve the above technical problems, the technical solution adopted by the present invention is:

[0009] A temporal action detection method based on a multi-scale graph convolutional network, comprising the following steps:

[0010] Step 1, use the I3D network to extract features from the video;

[0011] Step 2, use the CNN network to further refine the extracted video features;

[0012] Step 3: Feed the refined video features into the graph pyramid module for downsampling to obtain video features at different scales. Model the temporal edges, similarity edges, and spatial edges for the video features at each scale respectively. After modeling, perform graph convolution on the video features to aggregate the context information of the video features at the current scale.

[0013] Step 4: Feed the video features after graph convolution into the classification head and regression head for prediction.

[0014] Step 5: Synthesize the action prediction results at different scales to obtain the final temporal action detection result.

[0015] A further improvement of the technical solution of the present invention lies in: In step 1, specifically including: Input the video sequence into the I3D network to extract video network features. Use the I3D network to obtain the scene spatial flow and temporal flow information respectively, and expand the output to a fixed size.

[0016] A further improvement of the technical solution of the present invention lies in: In step 3, the graph pyramid module includes temporal edges, similarity edges, and spatial edges; The temporal edge represents the temporal position information of the action units according to the temporal order of the action units, and can aggregate the temporal context; The similarity edge aims to model the similar actions in the video. The similarity between action units is captured by the Euclidean distance, and the top k action units are connected by K-nearest neighbor; The spatial edge uses hyperbolic space to capture the spatial semantic relationship between action units.

[0017] A further improvement of the technical solution of the present invention lies in: In step 3, it specifically includes the following steps:

[0018] Step 3.1: Given the input video feature map F, the graph pyramid module performs downsampling to obtain video features at different scales. The downsampling rate is 2, and max-pooling sampling is used to enhance the action units in complex backgrounds.

[0019] Step 3.2: For the feature map F at a given scale m , the graph pyramid module uses the temporal edge to capture the global sequential relationship between action units; Except for the last node, the out-degree of each node is 1, and there is a unique edge connecting to the next node; The temporal edge connects each action unit in temporal order.

[0020] Step 3.3: For the feature map F at a given scale m , the graph pyramid module captures the similarity between action units by calculating the Euclidean distance of the features at each scale, and uses the K-nearest neighbor algorithm to select the top k action units to establish a similarity relationship.

[0021] Step 3.4, the spatial edge preserves the local temporal continuity of the action units within the action; the graph pyramid module uses a hyperbolic space to calculate the semantic hierarchy of the action units and establish hierarchical edges;

[0022] Step 3.5, use a two-layer graph convolutional network to capture these relationships between action units;

[0023] For each layer of convolution, the graph convolution operation is implemented as follows:

[0024] X (k) ={[X T , AX (k) T - X T W (k) , k = 1, 2}

[0025] In the formula, A represents the adjacency matrix, [,] represents column concatenation, X (k) represents the hidden feature of the k-th layer, and W (k) represents the trainable weight of the k-th layer;

[0026] The graph convolution uses three layers of conv and residual connections to aggregate the temporal edge, spatial edge, and similarity edge information to obtain the final output:

[0027] Out = ReLU(conv(X, A t , W t ) + conv(X, A s , W s ) + conv(X, A h , W h ) + X)

[0028] In the formula, A t , A s , A h represent the adjacency matrix, W t , W s , W h represent the trainable weights, and conv represents a stacked convolutional layer with a mask matrix.

[0029] A further improvement of the technical solution of the present invention lies in: in Step 3.4, it specifically includes the following steps:

[0030] Step 3.4.1, calculate the Euclidean norm of the action units to measure the size of the action units in the hyperbolic space. The norm calculation formula is as follows:

[0031]

[0032] In the formula, D v represents the hidden dimension, set to 512; f ijRepresents the characteristics of the action unit;

[0033] Step 3.4.2. Due to the curvature property of the hyperbolic space, the magnitude of a vector in a two-dimensional space is different from that in the hyperbolic space. To further model the action unit in the hyperbolic space, it is necessary to calculate the vector magnitude of the action unit in the hyperbolic space, as shown in the following formula:

[0034]

[0035] In the formula, ρ represents the regularization parameter;

[0036] Step 3.4.3. Through the Euclidean norm and vector magnitude calculated in Steps 3.3.1 and 3.3.2, the action unit in the hyperbolic space is obtained, as shown in the following formula:

[0037]

[0038] Step 3.4.4. The Poincaré ball model is used to calculate the hierarchical distance between action units, as shown in the following formula:

[0039]

[0040] In the formula, f i h Represents the representation of the i-th action unit in the hyperbolic feature space; Represents the representation of the j-th action unit in the hyperbolic feature space;

[0041] The graph pyramid module will only embed the action units into the hyperbolic space to obtain the edges when calculating the hyperbolic space distance, without changing the original action features.

[0042] A further improvement of the technical solution of the present invention lies in: In Step 4, it specifically includes: Feature maps F of different given scales m Are first sent to the GN layer to reduce the bias caused by sample differences, and then sent to the classification head and regression head to obtain the action prediction results at different time series. Both the classification head and the regression head use 4 3×3 convolutions, and the final layer obtains the final classification and localization results.

[0043] Due to the adoption of the above technical solution, the technical progress achieved by the present invention is:

[0044] 1. A temporal action detection method based on a multi-scale graph convolutional network proposed by the present invention simultaneously considers the local and global relationships between action units, models the action units to obtain the local and global relationships of the action units, helps the network better understand and coordinate the spatio-temporal features of different action units, can improve the accuracy and effectiveness of action detection, effectively solve the problem that spatio-temporal simultaneous modeling cannot be performed in temporal action detection, and can accurately detect video actions.

[0045] 2. A novel Graph Pyramid Module (GPM) is proposed in the present invention to model each video, process video features to obtain multi-scale features, and capture the temporal semantics, similarity semantics, and spatial semantics relationships between action units through temporal edges, similarity edges, and spatial edges for different scales of video features, thereby refining the different relationships between action units, effectively solving the problem of the inability to coordinate spatio-temporal in the temporal action detection method, and effectively improving the accuracy of temporal action detection.

[0046] 3. The Graph Pyramid Module (GPM) proposed in the present invention can be easily integrated into existing methods, overcoming the problem that existing temporal action detection methods cannot model the relationships between action units during retraining. Specifically, the multi-scale features obtained by GPM are upsampled to achieve feature alignment. To prevent GPM from changing the dimension of the original input data, the features of each scale can be linearly added to obtain the final output, which enables GPM to be easily inserted into any position of existing temporal action detection methods. BRIEF DESCRIPTION OF THE DRAWINGS

[0047] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the following will briefly introduce the drawings required for the description of the embodiments or the prior art. Obviously, the drawings in the following description are some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can also be obtained based on these drawings.

[0048] Figure 1 is the overall flowchart of a temporal action detection method based on a multi-scale graph convolutional network provided in the embodiments of the present invention;

[0049] Figure 2 is the algorithm description diagram of the Graph Pyramid Module proposed in the embodiments of the present invention;

[0050] Figure 3 is the processing flow schematic diagram of the Graph Pyramid Module proposed in the embodiments of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0051] It should be noted that the terms "including" and "having" and any variations thereof in the specification and claims of the present invention and the above-mentioned drawings are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0052] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments:

[0053] As Figure 1 shown, a temporal action detection method based on a multi-scale graph convolutional network includes the following steps:

[0054] Step 1, use the I3D network to extract features from the video;

[0055] Specifically, input the video sequence into the I3D network (the full English name is Inflated 3D Convolutional Network) to extract video network features, use the I3D network to obtain scene spatial flow and temporal flow information respectively, and expand the output to a fixed size.

[0056] Step 2, use the CNN network (the full English name is Convolutional Neural Network, that is, convolutional neural network) to further refine the extracted video features;

[0057] Specifically, input the fixed-size I3D video features F output in Step 1 into the CNN network for feature refinement, extract richer semantic information, improve the discrimination ability for different categories, and help the network identify the subtle differences between different categories to improve the classification accuracy.

[0058] Step 3, send the refined video features into the graph pyramid module for downsampling to obtain video features of different scales, respectively model the temporal edges, similarity edges, and spatial edges of the video features at each scale, and perform graph convolution on the video features after modeling to aggregate the context information of the video features at the current scale;

[0059] Specifically, it includes: further send the video features refined by the CNN network into the graph pyramid module (the full English name is Graph Pyramid Module, abbreviated as GPM) for downsampling to obtain feature maps of different scales, perform spatio-temporal modeling on the action units at each scale, and use graph convolution to obtain the spatio-temporal relationships of different action units.

[0060] GPM has three types of edges: temporal edges, similarity edges, and spatial edges. Temporal edges represent the temporal position information of action units according to the temporal order of action units, and can aggregate temporal context. Similarity edges aim to model similar actions in the video. The similarity between action units is captured by the Euclidean distance, and the top-k action units are connected by K-Nearest Neighbors (KNN). Similar actions are beneficial for detection and classification. Spatial edges use hyperbolic space to capture the spatial semantic relationships between action units. Actions in the video often consist of multiple action units, which not only have an order relationship in temporal semantics but also have an order relationship in spatial semantics. Action units are embedded into hyperbolic space, and the spatial semantic relationships between action units are captured by calculating the distances between action units in hyperbolic space. KNN is used to select the top-k action units to establish graph edges. Similarity edges and spatial edges are used to aggregate semantic context.

[0061] The following combines Figure 2 and Figure 3 to illustrate the novel graph pyramid module used in the present invention:

[0062] As Figure 3 shown in the schematic diagram of the processing flow of the pyramid module (GPM), the video features are represented as:

[0063]

[0064] where L represents the number of action units, and f i represents the i-th action unit in the video.

[0065] Step 3.1, GPM first uses a downsampling method for the video feature F to obtain video features at different scales, where the downsampling rate is 2, and each scale level is a downsampling of the previous scale level. Different downsampling methods play important roles in the graph construction of subsequent scales. The max-pooling method captures local action units to model actions, that is, g i = max(f1, f2, …, f k ). The max-pooling method can enhance action units in complex backgrounds, which is more conducive to accurately finding the relationships between action units. In addition, GPM can perform local and global modeling by capturing different relationships between action units, which makes up for the limitations of the max-pooling operation.

[0066] Step 3.2, GPM uses temporal edges to capture the global order relationship between action units. Specifically, except for the last node, the out-degree of each node is 1, and there is a unique edge connecting to the next node. Temporal edges connect each action unit in temporal order, which preserves the global temporal continuity of action units. Temporal edges can be defined by the following formula:

[0067]

[0068] Wherein, m represents the number of levels of the multi-scale.

[0069] Step 3.3, a video often contains multiple similar actions, and since these actions are often located at different positions in the video, such as Figure 3 shown. GPM captures the similarity between action units by calculating the Euclidean distance of features at each scale, and uses KNN to select the top k action units; the similarity edge is defined as follows:

[0070]

[0071] Wherein, d e represents the Euclidean distance.

[0072] Step 3.4, the spatial edge preserves the local temporal continuity of action units within an action. Since the perimeter and area of the hyperbolic space are proportional to the radius, it can be regarded as a simulation of a tree and has the ability of hierarchical representation. GPM uses the hyperbolic space to calculate the semantic hierarchy of action units and establish hierarchical edges. Specifically, it includes the following steps:

[0073] Step 3.4.1, first calculate the Euclidean norm of the action unit, as shown in the following formula:

[0074]

[0075] Wherein, D v represents the hidden dimension, which is set to 512; f ij represents the feature of the action unit.

[0076] Step 3.4.2, due to the bending property of the hyperbolic space, the magnitude of a vector in the two-dimensional space is different from that in the hyperbolic space. To further model the action unit in the hyperbolic space, it is necessary to calculate the vector magnitude of the action unit in the hyperbolic space, as shown in the following formula:

[0077]

[0078] Wherein, ρ represents the regularization parameter.

[0079] After calculating the obtained Euclidean norm and vector magnitude, the action unit representation in the hyperbolic space can be obtained as follows:

[0080]

[0081] Step 3.4.4, The Poincaré ball model is a special hyperbolic space and can be applied to the field of action prediction as a special hyperbolic geometry. The Poincaré ball model is used to calculate the hierarchical distance between action units as shown in the following formula:

[0082]

[0083] In the formula, f i h represents the representation of the i-th action unit in the hyperbolic feature space; represents the representation of the j-th action unit in the hyperbolic feature space.

[0084] When calculating the hyperbolic space distance, GPM only embeds the action units into the hyperbolic space to obtain the edges without changing the original action features.

[0085] Step 3.5, Use two-layer GCN (Graph Convolutional Networks, abbreviated as GCN) to capture these relationships between action units. As Figure 3 shown, the different shades of gray in the convolutional layer represent different kernel sizes. Specifically, the lighter-colored blocks set the kernel size to 1, and the darker-colored blocks set the kernel size to 3. In addition, the numbers in each box refer to the input and output channels. For each layer of convolution, the graph convolution operation is implemented as follows:

[0086] X (k) ={[X T ,AX (k) T -X T W (k) ,k=1,2} (8)

[0087] In the formula, A represents the adjacency matrix, [,] represents column connection, X (k) represents the hidden feature of the k-th layer, and W (k) represents the trainable weight of the k-th layer. The graph convolution uses three layers of conv and residual connections to aggregate the time edge, space edge, and similarity edge information, and the final output is as follows:

[0088] Out=ReLU(conv(X,A t ,W t )+conv(X,A s ,W s )+conv(X,A h ,W h )+X) (9)

[0089] In the formula, A t 、A s 、A hdenotes the adjacency matrix, W t 、W s 、W h denote trainable weights, and conv denotes a stacked convolutional layer with a mask matrix. After aggregating the context of action units, GPM makes full use of the relationships between action units.

[0090] Step 4: Feed the video features after graph convolution into the classification head and the regression head for prediction;

[0091] Specifically, the features F of different scales m are first fed into the GN layer to reduce the bias caused by sample differences, and then sent to the classification head and the regression head to obtain action detection results at different time series. Both the classification head and the regression head use 4 3×3 convolutions, and the final layer obtains the final classification and localization results.

[0092] Step 5: Synthesize the action prediction results of different scales to obtain the final temporal action detection result.

[0093] In summary, in view of the problems in the existing temporal action detection methods, such as relying only on a single temporal context relationship and ignoring the hierarchical relationship of action units in the video, resulting in flaws in modeling and thus less than ideal detection effects, the present invention proposes a temporal action detection method based on a multi-scale graph convolutional network. Under the premise of retaining the temporal semantic information of action units in the video, it fully excavates and models the hierarchical semantic information of different action units to achieve more accurate temporal action detection. The video is fed into the I3D network to generate video features, and then the video features are further refined in the CNN network. After obtaining the refined video features, they are processed using the graph pyramid module. After the video features of different scales are processed by the graph pyramid module, they are fed into the classification head and the regression head for prediction to obtain the final prediction result. The present invention models the action units in the video more fully, effectively utilizes the spatio-temporal relationship of action units, and thus greatly improves the accuracy of temporal action detection.

[0094] Finally, it should be noted that: the above embodiments are only used to illustrate the technical solutions of the present invention, rather than to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions described in the foregoing embodiments, or perform equivalent replacements on some or all of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the scope of the technical solutions of the embodiments of the present invention.

Claims

1. A temporal action detection method based on a multi-scale graph convolutional network, characterized in that It includes the following steps: Step 1: Use the I3D network to extract features from the video; Step 2: Use the CNN network to further refine the extracted video features; Step 3: Feed the refined video features into the graph pyramid module for downsampling to obtain video features of different scales, respectively model the temporal edge, similarity edge, and spatial edge for each scale of video features, and after modeling, perform graph convolution on the video features to aggregate the context information of the current scale video features; Step 4: Feed the video features after graph convolution into the classification head and regression head for prediction; Step 5: Synthesize the action prediction results of different scales to obtain the final temporal action detection result.

2. The temporal action detection method based on a multi-scale graph convolutional network according to claim 1, wherein Specifically, Step 1 includes: Input the video sequence into the I3D network to extract video network features, use the I3D network to obtain the scene spatial flow and temporal flow information respectively, and expand the output to a fixed size.

3. A temporal action detection method based on a multi-scale graph convolutional network according to claim 1, characterized in that, In Step 3, the graph pyramid module includes a temporal edge, a similarity edge, and a spatial edge; the temporal edge represents the temporal position information of action units according to the temporal order of action units, and can aggregate temporal context; the similarity edge aims to model similar actions in the video, the similarity between action units is captured by the Euclidean distance, and the top k action units are connected by K-nearest neighbors; the spatial edge uses hyperbolic space to capture the spatial semantic relationship between action units.

4. A temporal action detection method based on a multi-scale graph convolutional network according to claim 1, characterized in that Step 3 specifically includes the following steps: Step 3.1: Given the input video feature map F, the graph pyramid module performs downsampling to obtain video features of different scales, adopts a downsampling rate of 2, and uses max-pooling sampling to enhance action units under complex backgrounds; Step 3.2, for the feature map F of a given scale m , the graph pyramid module uses temporal edges to capture the global sequential relationship between action units; except for the last node, the out-degree of each node is 1, and there is a unique edge connecting to the next node; the temporal edges connect each action unit in chronological order; Step 3.3, for the feature map F at a given scale m , the figure pyramid module captures the similarity between action units by calculating the Euclidean distance of features at each scale, and uses the K-nearest neighbor algorithm to select the top k action units to establish a similarity relationship; Step 3.4: The spatial edge preserves the local temporal continuity of action units within an action; the graph pyramid module uses hyperbolic space to calculate the semantic hierarchy of action units and establish hierarchical edges; Step 3.5: Use a two-layer graph convolutional network to capture these relationships between action units; For each layer of convolution, the graph convolution operation is implemented as follows: X (k) = {[X T , AX (k) T -X T W (k) , k = 1, 2} where A represents the adjacency matrix, [,] represents column connection, and X (k) represents the hidden feature of the k-th layer, and W (k) represents the trainable weight of the k-th layer; The graph convolution uses three layers of conv and residual connections to aggregate the temporal edge, spatial edge, and similarity edge information to obtain the final output: Out = ReLU(conv(X, A t , W t ) + conv(X, A s , W s ) + conv(X, A h , W h ) + X) where, A t , A s , A h represent the adjacency matrix, W t , W s , W h represent the trainable weights, and conv represents the stacked convolutional layer with a mask matrix.

5. The temporal action detection method based on a multi-scale graph convolutional network according to claim 4, characterized in that, Step 3.4 specifically includes the following steps: Step 3.4.1: Calculate the Euclidean norm of the action unit to measure the size of the action unit in hyperbolic space, and the norm calculation formula is expressed as follows: where D v represents the hidden dimension, which is set to 512; f ij represents the feature of the action unit; Step 3.4.2: Due to the curvature property of hyperbolic space, the size of a vector in two-dimensional space is different from the size of a vector in hyperbolic space. To further model the action unit in hyperbolic space, it is necessary to calculate the vector amplitude of the action unit in hyperbolic space, as expressed in the following formula: In the formula, ρ represents the regularization parameter; Step 3.4.3: Through the Euclidean norm and vector amplitude calculated in Steps 3.3.1 and 3.3.2, obtain the action unit in hyperbolic space, as expressed in the following formula: Step 3.4.4: Use the Poincaré ball model to calculate the hierarchical distance between action units, as expressed in the following formula: where f i h represents the representation of the i-th action unit in the hyperbolic feature space; represents the representation of the j-th action unit in the hyperbolic feature space; When calculating the hyperbolic space distance, the graph pyramid module only embeds the action units into hyperbolic space to obtain edges, without changing the original action features.

6. The temporal action detection method based on a multi-scale graph convolutional network according to claim 1, wherein Step 4 specifically includes: feature maps F of different given scales m First, it is sent to the GN layer to reduce the bias caused by sample differences, and then to the classification head and regression head to obtain the action prediction results of different time series. Both the classification head and the regression head use four 3×3 convolutions, and the final classification and localization results are obtained in the last layer.