Temporal action segmentation method and apparatus, model training method and apparatus, and storage medium
Through multimodal action recognition technology, combined with high-dimensional feature extraction, dimensionality reduction processing and attention mechanism, the problem of insufficient accuracy of action recognition and timing reasoning in the existing technology is solved, and a more efficient action segmentation effect is achieved.
Patent Information
- Application Number
- PCT/CN2025/071813
- Authority / Receiving Office
- WO · WO
- Patent Type
- Applications
- Current Assignee / Owner
- Priority Date
- 2024-01-15
- Filing Date
- 2025-01-10
- Publication Date
- 2025-07-24
AI Technical Summary
The existing timing action segmentation technology has the problem of insufficient accuracy in action recognition and timing reasoning, especially because the action representation and understanding angles caused by using only visual color information are limited, and the multimodal fusion method cannot fully utilize the complementary information of each mode, resulting in fusion of irrelevant information and introducing interference.
Multimodal action recognition technology is adopted to extract high-dimensional feature of modal data such as inertia, key points and bounding boxes, combine dimensionality reduction processing and attention mechanism, and acquire motion and spatial flow characteristics, use multi-head cross attention mechanism to perform feature fusion, and process action categories and timing boundary prediction through category encoder and boundary encoder networks, and optimize model parameters using multi-stage interaction method.
It improves the accuracy of action segmentation, fully integrates multimodal features, reduces the amount of interference information, enhances the comprehensiveness and robustness of action representation, improves timing modeling capabilities, and reduces the problem of oversegmentation.
Smart Images

Figure CN2025071813_24072025_PF_FP_ABST
Abstract
Description
Method and device for temporal action segmentation, model training method, device and storage medium
[0001] CROSS-REFERENCE TO RELATED APPLICATIONS
[0002] This application is based on the application with CN application number 202410057649.0 and application date January 15, 2024, and claims its priority. The disclosed content of the CN application is hereby introduced as a whole into this application. Technical Field
[0003] The present disclosure relates to the field of artificial intelligence technology, and in particular to a method and device for temporal action segmentation, a model training method and device, and a storage medium. Background Art
[0004] Temporal action segmentation technology refers to predicting the human action category in each frame of a given video. In related technologies, temporal action segmentation technology is usually divided into two parts: action recognition and temporal reasoning. In the action recognition part, the visual features of the video frame are first extracted, and then the video frame features are converted into feature vectors in the metric space through machine learning technology (such as deep neural networks). In the metric space, the distance between the features of video frames of different action categories is large, which facilitates action classification. In the temporal reasoning part, the focus will be on the temporal continuity of human actions, and all frames in a continuous time segment will be predicted to have the same action category.
[0005] In related technologies, temporal action segmentation technology usually uses deep convolutional neural networks as video feature extractors and metric space learners, uses recurrent neural networks, temporal convolutional networks and Transformer architectures to perform temporal modeling of frame-level action features, uses large-scale video annotation data to train deep neural networks, and uses cross-entropy loss functions to optimize deep convolutional neural networks. Summary of the Invention
[0006] One purpose of the present disclosure is to improve the accuracy of action segmentation.
[0007] According to one aspect of some embodiments of the present disclosure, a training method for a temporal action segmentation model is proposed, comprising: performing high-dimensional feature extraction on data of each modality in sample data to obtain a first feature of each modality; obtaining an intermediate loss function through dimensionality reduction processing based on the first feature, and obtaining a second feature of each modality through dimensionality recovery; obtaining a motion representation through the fusion of features of different modalities based on the second feature and the supplementary feature; processing the motion representation through a category encoder network and a boundary encoder network respectively to obtain an action category prediction vector and a temporal boundary prediction vector, determining an action category prediction result, and determining a prediction loss function, wherein the temporal action segmentation model adjusts parameters according to the intermediate loss function and the prediction loss function until the training is completed.
[0008] In some embodiments, based on the second feature and the supplementary feature, the motion representation is obtained by fusing features of different modalities, including: obtaining attention-enhanced motion flow features and attention-enhanced spatial flow features based on the second feature and the supplementary feature; obtaining cross-stream attention features through a multi-head cross-attention mechanism; obtaining motion feature vectors and spatial feature vectors based on the cross-stream attention features, the attention-enhanced motion flow features, and the attention-enhanced spatial flow features; and determining the motion representation based on the motion feature vectors and the spatial feature vectors.
[0009] In some embodiments, based on the first feature, an intermediate loss function is obtained through dimensionality reduction processing, and the second feature of each modality is obtained through dimensionality recovery, including: temporal encoding of the first feature, determining the intermediate category feature through a fully connected layer, and determining the intermediate category loss function; restoring the dimension of the intermediate category feature and adding it to the first feature to obtain the second feature; obtaining the intermediate boundary feature of the second feature through dimensionality reduction processing and temporal information extraction, and obtaining the intermediate boundary loss function based on the intermediate boundary feature, wherein the intermediate loss function includes the intermediate category loss function and the intermediate boundary loss function.
[0010] In some embodiments, motion representation is processed by a category encoder network and a boundary encoder network respectively to obtain an action category prediction vector and a temporal boundary prediction vector, and determining the action category prediction result includes: at the beginning of the encoder network, using the query feature vector in the category encoder network to retrieve the key-value pair feature vector in the boundary encoder network to obtain the boundary feature; at the end of the encoder network, using the query feature vector in the boundary encoder network to retrieve the key-value pair feature vector in the boundary encoder to obtain the category feature; according to the boundary feature and the category feature, obtaining the action category prediction vector and the temporal boundary prediction vector through dimensionality adjustment; filtering out the boundary value in the temporal boundary prediction vector according to the confidence level, and adjusting the action category prediction vector according to the boundary value to obtain the action category prediction result.
[0011] In some embodiments, determining the prediction loss function includes: obtaining an action classification loss function corresponding to the action category prediction result; obtaining a boundary probability vector based on the temporal boundary prediction vector, and determining the temporal boundary loss function based on the boundary probability vector, wherein the prediction loss function includes an action classification loss function and a temporal boundary loss function.
[0012] In some embodiments, obtaining attention-enhanced motion flow features and attention-enhanced spatial flow features based on the second features and the supplementary features includes: obtaining initial motion flow features based on the second features of the inertial modality, and obtaining initial spatial flow features based on the second features of the key point modality and the second features of the bounding box modality, wherein the modalities include inertia, key points and bounding boxes; adding the initial motion flow features and the initial spatial flow features to the supplementary features respectively to obtain auxiliary enhanced motion flow features and auxiliary enhanced spatial flow features; and obtaining attention-enhanced motion flow features and attention-enhanced spatial flow features based on the auxiliary enhanced motion flow features and the supplementary features using a multi-head attention mechanism.
[0013] In some embodiments, obtaining a motion feature vector and a spatial feature vector based on the cross-stream attention feature, the attention-enhanced motion stream feature and the attention-enhanced spatial stream feature includes: processing the first cross-stream attention feature and the attention-enhanced motion stream feature through a frame-by-frame convolution layer to obtain a third feature; processing the second cross-stream attention feature and the attention-enhanced spatial stream feature through a frame-by-frame convolution layer to obtain a fourth feature; adding the third feature, the first cross-stream attention feature and the attention-enhanced motion stream feature to obtain a motion feature vector; adding the fourth feature, the second cross-stream attention feature and the attention-enhanced spatial stream feature to obtain a spatial feature vector, wherein the first cross-stream attention feature is a cross-stream attention feature obtained by matching the query feature vector of the spatial stream with the key-value pair feature vector of the motion stream, and the second cross-stream attention feature is a cross-stream attention feature obtained by matching the query feature vector of the motion stream with the key-value pair feature vector of the spatial stream.
[0014] In some embodiments, determining the motion representation based on the motion feature vector and the spatial feature vector includes: adding the motion feature vector and the spatial feature vector as the motion representation.
[0015] In some embodiments, the second feature is processed through dimensionality reduction processing and time series information extraction to obtain an intermediate boundary feature, and the intermediate boundary loss function is obtained based on the intermediate boundary feature, including: processing the second feature through a layer of convolutional neural network to obtain a fifth feature; obtaining the time series information in the fifth feature through a multi-layer convolutional neural network with residual connections, and adding it to the fifth feature to determine the intermediate boundary feature; obtaining a boundary classification probability vector through an activation function for the intermediate boundary feature; and determining the intermediate boundary loss function based on the boundary classification probability vector.
[0016] In some embodiments, the training method also includes: obtaining the weighted sum of the intermediate category loss function, the intermediate boundary loss function, the action classification loss function and the temporal boundary loss function as the loss function value, wherein the temporal action segmentation model adjusts parameters according to the loss function value, the intermediate loss function includes the intermediate category loss function and the intermediate boundary loss function, and the prediction loss function includes the action classification loss function and the temporal boundary loss function.
[0017] In some embodiments, the supplementary features are determined based on a work log of the IoT device.
[0018] According to one aspect of some embodiments of the present disclosure, a temporal action segmentation method is proposed, including: obtaining data and supplementary data of each modality of a target; inputting the data of each modality and the supplementary dimension data into a temporal action segmentation model to obtain an action category prediction result of the target, wherein the temporal action segmentation model is generated according to any one of the training methods of the temporal action segmentation model mentioned above.
[0019] In some embodiments, the method complies with at least one of the following: the modality includes inertia, key points, and bounding boxes; or the supplementary data includes a work log of the IoT device.
[0020] According to one aspect of some embodiments of the present disclosure, a training device for a temporal action segmentation model is proposed, comprising: a feature extraction module, configured to perform high-dimensional feature extraction on data of each modality in sample data to obtain a first feature of each modality; a feature constraint module, configured to obtain an intermediate loss function through dimensionality reduction processing based on the first feature, and obtain a second feature of each modality through dimensionality recovery; a fusion module, configured to obtain a motion representation through the fusion of features of different modalities based on the second feature and a supplementary feature; and an interactive dual-branch module, configured to process the motion representation through a category encoder network and a boundary encoder network, respectively, to obtain an action category prediction vector and a temporal boundary prediction vector, determine an action category prediction result, and determine a prediction loss function, wherein the temporal action segmentation model adjusts parameters according to the intermediate loss function and the prediction loss function until the training is completed.
[0021] In some embodiments, the training device also includes: a loss value determination unit, configured to obtain the weighted sum of the intermediate category loss function, the intermediate boundary loss function, the action classification loss function and the temporal boundary loss function as the loss function value, wherein the temporal action segmentation model adjusts parameters according to the loss function value, the intermediate loss function includes the intermediate category loss function and the intermediate boundary loss function, and the prediction loss function includes the action classification loss function and the temporal boundary loss function.
[0022] According to one aspect of some embodiments of the present disclosure, a temporal action segmentation device is proposed, comprising: a data acquisition module, configured to acquire data and supplementary data of each modality of a target; and a prediction module, configured to input the data of each modality and the supplementary dimension data into a temporal action segmentation model to obtain a prediction result of the action category of the target, wherein the temporal action segmentation model is generated according to any one of the training methods of the temporal action segmentation model mentioned above.
[0023] According to one aspect of some embodiments of the present disclosure, a data processing device is provided, including: a memory; and a processor coupled to the memory, wherein the processor is configured to execute any one of the above-mentioned methods based on instructions stored in the memory.
[0024] According to one aspect of some embodiments of the present disclosure, a computer-readable storage medium is provided, on which computer program instructions are stored. When the instructions are executed by a processor, any one of the methods mentioned above is implemented.
[0025] According to one aspect of some embodiments of the present disclosure, a computer program product is provided, comprising a computer program or instructions, wherein the computer program or instructions implement any one of the methods mentioned above when executed by a processor.
[0026] According to one aspect of some embodiments of the present disclosure, a computer program is proposed, configured to cause a processor to execute any one of the methods mentioned above. BRIEF DESCRIPTION OF THE DRAWINGS
[0027] The drawings described herein are used to provide a further understanding of the present disclosure and constitute a part of the present disclosure. The exemplary embodiments of the present disclosure and their descriptions are used to explain the present disclosure and do not constitute an improper limitation of the present disclosure. In the drawings:
[0028] FIG1 is a flowchart of some embodiments of the training method of the temporal action segmentation model disclosed herein.
[0029] FIG2 is a flowchart of some embodiments of modal information constraints in the training method of the temporal action segmentation model disclosed herein.
[0030] FIG3 is a flowchart of some embodiments of the feature fusion process in the training method of the temporal action segmentation model disclosed in the present invention.
[0031] FIG4 is a flowchart of some embodiments of determining action category prediction results in the training method of the temporal action segmentation model disclosed herein.
[0032] FIG5 is a schematic diagram of some embodiments of the temporal action segmentation model disclosed herein.
[0033] FIG6 is a flowchart of some embodiments of the method for temporal action segmentation disclosed herein.
[0034] FIG7 is a schematic diagram of some embodiments of a training device for a temporal action segmentation model disclosed herein.
[0035] FIG8 is a schematic diagram of some embodiments of the time-sequence action segmentation device disclosed herein.
[0036] FIG9 is a schematic diagram of some embodiments of a data processing device according to the present disclosure.
[0037] FIG10 is a schematic diagram of some other embodiments of the data processing device disclosed herein. DETAILED DESCRIPTION
[0038] The technical solution of the present disclosure is further described in detail below through the accompanying drawings and examples.
[0039] Related temporal action segmentation techniques typically use video visual color information as samples for feature extraction and model training. These techniques extract action features for frame-level action category prediction and employ multi-stage adjustments or independent boundary branches to strengthen temporal relationships between actions. However, these approaches, relying solely on visual color information, limit their ability to represent and understand human motion.
[0040] Multimodal action recognition technology refers to the process of fusing the complementary information provided by multiple types of data to obtain a comprehensive representation of human actions, and then predict human actions. Related multimodal action recognition technologies usually use deep convolutional neural networks and Transformer structures. By training the network model with large-scale labeled data, the information of different modalities is fused to obtain action recognition results. The main multimodal fusion methods in related technologies are: 1) data-level fusion, which refers to collecting multi-source data and then inputting it into the feature learning network; 2) feature-level fusion, which refers to converting different modal data into high-dimensional feature expressions before fusing them; 3) decision-level fusion, which refers to training a classifier for each modal data separately and fusing the prediction scores of each classifier.
[0041] Related feature-level fusion multimodal fusion methods mainly use high-dimensional feature splicing, addition or attention mechanism, which cannot guarantee the quality of multimodal features before fusion. This may lead to the fusion of irrelevant information and introduce interference. In addition, the feature fusion method used is relatively simple and cannot fully utilize the complementary information of each modality.
[0042] In response to the above problems, the present disclosure proposes a temporal action segmentation method, device, model training method, device and storage medium to improve the accuracy of action segmentation.
[0043] A flowchart of some embodiments of the training method of the temporal action segmentation model disclosed in the present invention is shown in FIG1 .
[0044] In step S11, high-dimensional feature extraction is performed on the data of each modality in the sample data to obtain the first feature of each modality. In some embodiments, each piece of sample data includes multi-modal data for the processing object (such as a person) obtained through multiple channels and devices. In some embodiments, the sample data also includes the action category label corresponding to each frame.
[0045] In some embodiments, the above modalities may include inertia, key points, and bounding boxes, thereby avoiding the use and reliance on visual colors, which may easily expose sensitive information of people and the surrounding environment and the risk of leaking privacy, thereby improving security.
[0046] In some embodiments, the feature extraction operation includes performing high-dimensional feature extraction on the inertial data, keypoint data, and bounding box data, respectively, to obtain a first feature of the inertial data, a first feature of the keypoint data, and a first feature of the bounding box data. In some embodiments, an encoder can be set separately for each modality of data to perform high-dimensional feature extraction. In some embodiments, high-dimensional features refer to features with a dimension greater than or equal to a predetermined number, which in some embodiments is 1000.
[0047] In some embodiments, the first feature is a feature of L×D dimensions, where L is the time length and D is the modal feature dimension.
[0048] In step S13, an intermediate loss function is obtained based on the first feature through dimensionality reduction, and the second feature of each modality is obtained through dimensionality recovery. By constraining the intermediate loss function, the information of each modality is constrained separately, improving the quality of the features based on the subsequent fusion step and reducing the amount of interference information.
[0049] In some embodiments, the intermediate loss function includes an intermediate category loss function and an intermediate boundary loss function, thereby constraining the information of each modality from the perspectives of category and boundary, thereby improving the quality of feature optimization.
[0050] In some embodiments, the process of modal information constraint is shown in FIG2 .
[0051] In step 231, the first feature is temporally encoded through a long short-term memory network, and then processed through a fully connected layer to obtain an intermediate category feature, and an intermediate category loss function is determined.
[0052] In some embodiments, the dimension of the intermediate category feature is smaller than the dimension of the first feature. In some embodiments, the intermediate category feature is a feature of L×C dimensions, where C is the number of predicted action categories.
[0053] In some embodiments, a classification probability vector of L×C dimensions is obtained through a Softmax operation, and an intermediate category loss function is determined based on the classification probability vector. In some embodiments, a stochastic gradient descent algorithm and a backpropagation algorithm are used to optimize network parameters based on the intermediate category loss function.
[0054] In step 232, the intermediate category features are restored to their original dimensions and added to the first features to obtain the second features. In some embodiments, a fully connected neural network can be used to restore the L×C dimensional intermediate category features to L×D dimensions, and then add them to the first features to obtain the second features. In some embodiments, the feature addition is performed bit by bit. This method can avoid the loss of high-dimensional information caused by dimensionality reduction.
[0055] In step 233, the second feature is processed through dimensionality reduction and time series information extraction to obtain an intermediate boundary feature, and an intermediate boundary loss function is obtained according to the intermediate boundary feature.
[0056] In some embodiments, the second feature is used as input information, and the second feature is processed by a layer of convolutional neural network to obtain a fifth feature; further, the timing information in the fifth feature is obtained by a multi-layer convolutional neural network with residual connections, and is added to the fifth feature to determine the intermediate boundary feature.
[0057] In some embodiments, the intermediate boundary features are passed through an activation function to obtain a boundary classification probability vector, where the values in the vector range from 0 to 1. The closer the value is to 0, the lower the probability that the corresponding frame is at the action timing boundary, and the closer the value is to 1, the greater the probability that the frame is at the action timing boundary. The intermediate boundary loss function is determined based on the boundary classification probability vector. In some embodiments, the network parameters are optimized based on the intermediate category loss function using a stochastic gradient descent algorithm and a backpropagation algorithm.
[0058] Through this method, the information of each modality can be constrained separately through the constraints of the intermediate loss function, thereby improving the quality of the features based on which the subsequent fusion steps are based and reducing the amount of interference information; after the dimensionality reduction process, the dimension is increased and added to the original information to avoid the loss of high-dimensional information caused by dimensionality reduction and improve the comprehensiveness of the information; the information of each modality is constrained from the two perspectives of category and boundary to improve the quality of feature optimization.
[0059] In step S15, based on the second feature and the supplementary feature, a motion representation is obtained by fusing features from different modalities. In some embodiments, the supplementary feature may be a feature determined based on temporally sparse information. In some embodiments, the supplementary feature may be a feature related only to the characteristic action, such as a feature determined based on the work log of an IoT device. In some embodiments, supplementary data may be obtained as an additional modality. In step S11 above, high-dimensional feature extraction is not performed. Instead, the supplementary data is processed through a fully connected layer to obtain the supplementary features.
[0060] In some embodiments, the feature fusion process is shown in FIG3 .
[0061] In step 351, an attention-enhanced motion flow feature and an attention-enhanced spatial flow feature are obtained according to the second feature and the supplementary feature.
[0062] In some embodiments, the second features of different modalities can be integrated based on the characteristics of each modality. For example, initial motion flow features can be obtained based on the second features of the inertial modality, and initial spatial flow features can be obtained based on the second features of the keypoint modality and the second features of the bounding box modality. Further operations can be performed based on the integrated features. This approach can reduce the amount of subsequent computation and improve data processing efficiency.
[0063] In some embodiments, the initial motion flow features and the initial spatial flow features are added to the supplementary features to obtain auxiliary enhanced motion flow features and auxiliary enhanced spatial flow features. Through this method, supplementary feature information can be added to the features of the motion flow and spatial flow, respectively, thereby increasing the effective amount of information carried by the auxiliary enhanced motion flow features and the auxiliary enhanced spatial flow features, which is conducive to improving the accuracy of action segmentation.
[0064] In some embodiments, based on the auxiliary enhanced motion flow features and the supplementary features, a multi-head attention mechanism is utilized to obtain attention-enhanced motion flow features and attention-enhanced spatial flow features.
[0065] In step 352, cross-stream attention features are obtained through a multi-head cross-attention mechanism.
[0066] In some embodiments, based on the attention-enhanced motion stream features and attention-enhanced spatial stream features obtained above, a multi-head attention mechanism can be used to perform cross-stream processing on the two types of features. For example, a first cross-stream attention feature is obtained by matching the query feature vector of the spatial stream with the key-value feature vector of the motion stream, and a second cross-stream attention feature is obtained by matching the query feature vector of the motion stream with the key-value feature vector of the spatial stream. In some embodiments, the cross-stream attention features include the first cross-stream attention feature and the second cross-stream attention feature.
[0067] In step 353, a motion feature vector and a spatial feature vector are obtained based on the cross-stream attention features, the attention-enhanced motion stream features, and the attention-enhanced spatial stream features.
[0068] In some embodiments, fine-grained action feature vectors can be obtained as motion feature vectors and spatial feature vectors based on cross-stream attention features, attention-enhanced motion stream features, and attention-enhanced spatial stream features.
[0069] In some embodiments, a first cross-stream attention feature and an attention-enhanced motion stream feature are processed by a frame-by-frame convolution layer to obtain a third feature; a second cross-stream attention feature and an attention-enhanced spatial stream feature are processed by a frame-by-frame convolution layer to obtain a fourth feature; the third feature, the first cross-stream attention feature, and the attention-enhanced motion stream feature are added together to obtain a motion feature vector; and the fourth feature, the second cross-stream attention feature, and the attention-enhanced spatial stream feature are added together to obtain a spatial feature vector. Through this method, the motion feature vector includes the first cross-stream attention feature and the attention-enhanced motion stream feature processed by the frame-by-frame convolution layer, as well as the first cross-stream attention feature and the attention-enhanced motion stream feature themselves. The spatial feature vector includes the second cross-stream attention feature and the attention-enhanced spatial stream feature processed by the frame-by-frame convolution layer, as well as the second cross-stream attention feature and the attention-enhanced spatial stream feature themselves, thereby improving the comprehensiveness of the information in the vector.
[0070] In step 354, a motion representation is determined based on the motion feature vector and the spatial feature vector. In some embodiments, the motion feature vector and the spatial feature vector are added together to form the motion representation. This combination of fine-grained features improves the comprehensiveness and robustness of the motion representation.
[0071] Through the method in the embodiment shown above, it is possible to improve the fusion of information from different modalities based on the attention mechanism and supplementary information, encode the interactive context of the motion and spatial information of the action, thereby obtaining a complete action representation, improving the fusion between the motion flow and spatial flow information, and improving the ability to represent the action.
[0072] In step S17, the motion representation is processed by the class encoder network and the boundary encoder network, respectively, to obtain an action category prediction vector and a temporal boundary prediction vector, determine the action category prediction result, and determine a prediction loss function. In some embodiments, the prediction loss function may include an action classification loss function and a temporal boundary loss function, thereby constraining the model from both the classification and boundary perspectives to improve the model's accuracy.
[0073] In some embodiments, a classification probability vector can be obtained based on the action category prediction results through a Softmax operation, thereby determining the action classification loss function. In some embodiments, a stochastic gradient descent algorithm and a back propagation algorithm are used to optimize network parameters based on the action classification loss function.
[0074] In some embodiments, a boundary probability vector can be obtained based on the time-series boundary measurement vector through a sigmoid operation to determine the time-series boundary loss function. In some embodiments, a stochastic gradient descent algorithm and a backpropagation algorithm are used to optimize network parameters based on the time-series boundary loss function.
[0075] In some embodiments, the temporal action segmentation model performs parameter adjustment according to the intermediate loss function in the aforementioned step S13 and the prediction loss function of the current step until the training is completed.
[0076] In some embodiments, the weights of different loss functions can be set as needed to balance the importance of different loss functions. In some embodiments, the weights of intermediate loss functions and prediction loss functions can be set differently to balance the impact of intermediate loss and final loss on the model.
[0077] In some embodiments, a weighted sum of the intermediate category loss function, the intermediate boundary loss function, the action classification loss function, and the temporal boundary loss function is obtained as the loss function value.
[0078] Among them, the temporal action segmentation model adjusts parameters according to the loss function value. The intermediate loss function includes the intermediate category loss function and the intermediate boundary loss function, and the prediction loss function includes the action classification loss function and the temporal boundary loss function.
[0079] In some embodiments, a flowchart for determining an action category prediction result is shown in FIG4 .
[0080] In step 471, at the beginning of the encoder network, the query feature vector in the class encoder network is used to retrieve key-value pairs in the boundary encoder network to obtain boundary features. In some embodiments, the encoder network includes a class encoder network and a boundary encoder network, and the class encoder network and the boundary encoder network perform feature processing separately and in parallel. The beginning of the encoder network includes the beginning of the class encoder network and the beginning of the boundary encoder network.
[0081] In step 472, at the end of the encoder network, the query feature vector in the boundary encoder network is used to retrieve the key-value feature vector in the boundary encoder to obtain the category feature. In some embodiments, the end of the encoder network includes the end of the category encoder network and the end of the boundary encoder network.
[0082] In step 473, based on the boundary features and category features, dimensionality adjustment is performed to obtain an action category prediction vector and a temporal boundary prediction vector. In some embodiments, the L×D dimensional category features can be adjusted to L×C dimensions using a fully connected layer to obtain the action category prediction vector. In some embodiments, the L×D dimensional boundary features can be adjusted to L×1 dimensions using a fully connected layer to obtain the temporal boundary prediction vector.
[0083] In step 474 , the boundary value is filtered out from the time series boundary prediction vector according to the confidence level, and the action category prediction vector is adjusted according to the boundary value to obtain the action category prediction result.
[0084] In some embodiments, a confidence threshold may be set to filter out boundary values with confidence levels higher than the confidence threshold in the time series boundary prediction vector, and the action category prediction vector may be fine-tuned using the boundary values to obtain an action category prediction result.
[0085] Through the method in the embodiment shown above, it is possible to improve the temporal modeling capability, improve over-segmentation, and improve the accuracy of action segmentation through multi-stage interaction.
[0086] Based on the method in the embodiment shown above, it is possible to extract the features of multimodal data and improve the comprehensiveness of action representation; constrain the information of each modality separately, improve the quality of the features based on the subsequent fusion steps, fully integrate multimodal features, and improve the accuracy of action segmentation.
[0087] FIG5 is a schematic diagram of some embodiments of the temporal action segmentation model disclosed herein.
[0088] The temporal action segmentation model includes a feature extraction module, an intermediate bottleneck module, a dual-stream attention fusion module and a multi-stage interactive dual-branch module.
[0089] 1. Feature extraction module
[0090] In some embodiments, the number of feature extraction modules corresponds to the number of data sources or data categories.
[0091] As shown in Figure 5(a), different network structures are used according to the characteristics of data of different modalities.
[0092] In some embodiments, the inertial data, human key point data, and human bounding box data contain rich information, and the Transformer encoder structure G(·) is used to extract high-dimensional features to obtain the first feature I of each modality. x =G(x), where x can be represented by inertial data, human key point data, or human bounding box data. The first feature is used as the input of the intermediate bottleneck module.
[0093] In some embodiments, the work logs from IoT devices are sparse in time and only related to specific actions. Therefore, as supplementary data, a simple fully connected layer C(·) is used to obtain the corresponding supplementary features I l =C(l), where l represents the work log information from the IoT device.
[0094] 2. Intermediate bottleneck module
[0095] In some embodiments, the number of feature extraction modules corresponds to the number of types of modalities corresponding to the first feature.
[0096] Take the first feature I of a modality as an example, and take L×D dimension I as input, where L is the time length and D is the modality feature dimension. By reducing the feature dimension, the intermediate category loss L is calculated. c inner and the intermediate boundary loss L b inner , and then the enhanced modal feature I′ is obtained by increasing the feature dimension, which is the output of the intermediate bottleneck module. The intermediate bottleneck module includes the intermediate category branch and the intermediate boundary branch.
[0097] In some embodiments, the structure of the intermediate bottleneck module is shown in FIG5( b ).
[0098] A. In the intermediate category branch, given a modal first feature I of L×D dimension, I is input into the intermediate category branch. I is temporally encoded through the long short-term memory network g(·), and then the L×C intermediate category feature I is obtained through a fully connected layer neural network f(·) c =f(g(I)), where C is the number of predicted action categories. Then, the classification probability vector s(I) with L×C dimensions is obtained through softmax operation. c In some embodiments, the intermediate category loss function is calculated according to the following formula (1).
[0099] In some embodiments, the network parameters are optimized using a stochastic gradient descent algorithm and a back-propagation algorithm.
[0100] Furthermore, in the intermediate category branch, the feature dimension is resized from L×C to L×D using the fully connected layer neural network f(·), and the calculated result is added to the original feature input I, as shown in formula (2), to avoid the loss of high-dimensional information caused by dimensionality reduction. I′=f(I c )+I (2)
[0101] In formula (2), I′ is the output of the intermediate category branch and also the input of the intermediate boundary branch.
[0102] B. In the intermediate boundary branch, the intermediate boundary branch receives the output I′ of the intermediate category branch as input, and adjusts the feature vector size from L×D to L×1 through a layer of convolutional neural network p(·), obtaining the feature vector I b =p(I′). Then, the temporal information in the modeling feature is modeled by a multi-layer convolutional neural network p′(·) with residual connections, and the intermediate boundary feature is obtained according to formula (2). b ′=p′(I b )+I b (3)
[0103] For the intermediate boundary feature I b ′, the boundary classification probability vector m(I b ′), the values in the vector are all between 0 and 1. Given a frame's boundary classification probability vector, the closer the value is to 0, the lower the probability that the frame is at the action sequence boundary. Conversely, the closer the value is to 1, the greater the probability that the frame is at the action sequence boundary. The intermediate boundary loss function is calculated according to formula (4).
[0104] In the above formula, y t is the boundary truth value at time t. In some embodiments, the network parameters are optimized using a stochastic gradient descent algorithm and a back propagation algorithm.
[0105] 3. Dual-stream attention fusion module
[0106] The dual-stream attention fusion module takes the enhanced modal feature I′ obtained by the intermediate bottleneck module as input and obtains the comprehensive human motion feature I F The dual-stream attention fusion module consists of a complementary attention fusion module and a motion-spatial attention fusion module.
[0107] The dual-stream attention fusion module structure is shown in Figure 5(c), which includes the motion-spatial attention fusion module and the motion-spatial attention fusion module.
[0108] A. Supplementary attention fusion module
[0109] Through the above-mentioned feature extraction module and intermediate bottleneck module, the feature vectors corresponding to the inertial data, human key point data, human bounding box data and the work log from the IoT device are obtained respectively. The work log from the IoT device provides an accurate representation as a supplementary feature I L , the second feature of the inertial mode is used as the motion flow feature of the human body I M In addition, the second features of the key point modality and the bounding box modality are added as the spatial flow feature I S .
[0110] In some embodiments, feature I will be supplemented L Respectively with motion flow feature I M and spatial flow characteristics I S Add together to obtain the auxiliary enhanced feature vector I M-L and I S-L .
[0111] The auxiliary enhanced feature vector is fed into the multi-head attention layer together with the supplementary features, and the attention enhanced feature vector I is generated using the following formula (5) (6) M-L ′ and I S-L ′:
[0112] In the above formula, d is the normalization constant, T represents the matrix transpose, the auxiliary enhanced feature vector provides the key K and value V, and the supplementary feature provides the query Q.
[0113] B. Motion-Spatial Attention Fusion Module
[0114] The motion features and spatial stream features are modeled through a multi-head cross-attention mechanism. The query feature vector Q of one stream is used to match the key-value feature vector K, V of another stream according to the following formulas (7) and (8):
[0115] In the above formula, I M→S and I S→M is a cross-stream attention feature that encodes the correlation between motion and spatial streams.
[0116] The cross-stream attention features are converted into fine-grained action feature vectors through the frame-by-frame convolution layer c(·), which is called the motion feature vector I M→S ′, spatial eigenvector I S→M In some embodiments, the motion feature vector I can be determined according to formulas (9) and (10): M→S ′, spatial eigenvector I S→M '. I M→S ′=c(I M→S +I M-L ′)+I M→S +I M-L ′ (9) I S→M ′=c(I S→M +I S-L ′)+I S→M +I S-L ′ (10)
[0117] After the supplementary attention fusion module and the motion-space attention fusion module, the supplementary information is well embedded and the information of the two streams of motion and space are fully interacted.
[0118] In some embodiments, the motion feature vector I is added by element-by-element addition. M→S ′, spatial eigenvector I S→M ′, as shown in formula (11), a comprehensive and robust motion representation is obtained. F =I M→S ′+I S→M ′ (11)
[0119] 4. Multi-stage interactive dual-branch module
[0120] The multi-stage interactive dual-branch module includes a category branch and a boundary branch. F The two branches are respectively fed into two branches. The two branches are composed of multiple layers of Transformer encoder layers.
[0121] The first stage of information interaction is at the beginning of the action branch, where the query feature vector Q obtained from the category branch is C To retrieve the key-value pair feature vector K obtained from the boundary branch B 、V B After the multi-head cross attention, the boundary features after the first information interaction are obtained
[0122] The second stage of information interaction is at the end of the two branches, and the boundary branch provides the query feature vector Q′ B , the category branch provides the feature vector K′ of keys and values C , V′ C . According to the multi-head cross attention, the category features after the second information interaction are obtained
[0123] At the end of the category branch, a fully connected layer adjusts the L×D features to L×C to obtain the action category prediction vector X for each frame C At the end of the boundary branch, a fully connected layer adjusts the L×D features to L×1 to obtain the temporal boundary prediction vector X B First, set the threshold a, from X B Filter out high confidence boundary values Use boundary values For the category prediction vector X C Fine-tune as shown in the following formula (12).
[0124] In the above formula, k controls the decay rate of the weight, L is the time length of the segment, and s∈{-1,1} indicates that the feature vectors are aggregated in both the forward and backward directions during the weighting process. This is the final output of the network, which is the frame-level action category prediction result after temporal boundary constraints.
[0125] In some embodiments, a classification probability vector of L×C dimension is obtained by a softmax operation. The action classification loss function is calculated according to formula (13).
[0126] In some embodiments, the network parameters are optimized using a stochastic gradient descent algorithm and a back-propagation algorithm.
[0127] In some embodiments, the boundary probability vector m(X B ), calculate the temporal boundary loss function according to formula (14):
[0128] In some embodiments, y t is the boundary truth value at time t. In some embodiments, the network parameters are optimized using a stochastic gradient descent algorithm and a back propagation algorithm.
[0129] Based on the training method of the temporal action segmentation model shown above, based on the above-mentioned temporal action segmentation model, the features of multimodal data are extracted through the feature extraction module and the intermediate bottleneck module, and the multimodal features are fully fused through the dual-stream attention fusion module to obtain a comprehensive and robust human action representation. The information interaction between action categories and temporal boundaries is modeled through a multi-stage interactive dual-branch module to enhance the network's temporal modeling capability, thereby achieving accurate temporal action segmentation results; a method of introducing category awareness and temporal awareness capabilities for each modal feature, and providing the most effective information for subsequent feature fusion by calculating the intermediate category loss and intermediate boundary loss; through a multi-stage modal fusion method based on the attention mechanism, auxiliary information is combined and the full fusion interaction between motion flow and spatial flow information is promoted, and the attention mechanism is used to model the relationship between each information, thereby improving the action representation capability; through a multi-stage information interaction method of category-boundary, the contextual relationship between action categories and temporal boundaries is represented, thereby improving the temporal modeling capability and effectively improving the over-segmentation problem.
[0130] In some embodiments, during training, sample data is collected, including multimodal data for a time period and action category labels corresponding to each frame. The multimodal data and labels are sequentially traversed. The multimodal data in the sample data is input into a feature extraction module. During training, an intermediate loss function and a predicted loss function are obtained. Based on the loss functions, the network parameters are optimized using a stochastic gradient descent algorithm and a backpropagation algorithm.
[0131] In some embodiments, the multimodal fusion temporal action segmentation model is an end-to-end network, and each module is trained simultaneously. In some embodiments, the model generates four loss functions during training: the intermediate category loss function L c inner , intermediate boundary loss function L b inner , action classification loss function L c final And the temporal boundary loss function L b final , add the loss function according to formula (15). L = λ1L c inner +(1-λ1)L c final +λ2(L b inner +L b final ) (15)
[0132] In the above formula, λ1 and λ2 are weight coefficients, λ1 balances the intermediate loss and the final loss, and λ2 is used to balance the action category loss and the temporal boundary loss, and then the stochastic gradient descent algorithm is used to optimize the total network parameters.
[0133] Through this method, the intermediate loss and final loss, action category loss and temporal boundary loss can be balanced as needed to improve the controllability of the model.
[0134] FIG6 shows a flowchart of some embodiments of the method for temporal action segmentation disclosed herein.
[0135] In step S61, data and supplementary data for each modality of the target are obtained. In some embodiments, the modalities include inertia, key points, and bounding boxes. In some embodiments, the supplementary data includes a work log of the IoT device.
[0136] In step S62, the data for each modality and the supplementary dimension data are input into a temporal action segmentation model to obtain a prediction result for the target action category. In some embodiments, the temporal action segmentation model is generated using any of the temporal action segmentation model training methods mentioned above. In some embodiments, the temporal action segmentation model is shown in Figure 5.
[0137] Based on the method in the embodiment shown above, it is possible to extract features of multimodal data, constrain the information of each modality separately, improve the quality of features based on subsequent fusion steps, fully integrate multimodal features, and improve the accuracy of action segmentation.
[0138] FIG7 is a schematic diagram of some embodiments of the training device for the temporal action segmentation model disclosed herein.
[0139] The feature extraction module 701 can perform high-dimensional feature extraction on each modality in the sample data to obtain the first feature of each modality. In some embodiments, the feature extraction module 701 can perform step S11 above and any method performed by the feature extraction module in the embodiment corresponding to FIG5 .
[0140] The feature constraint module 702 can obtain an intermediate loss function based on the first feature through dimensionality reduction processing, and obtain the second feature of each modality through dimensionality recovery. In some embodiments, the temporal action segmentation model adjusts parameters based on the intermediate loss function until training is completed. In some embodiments, the feature constraint module 702 can perform any of the methods performed by the intermediate bottleneck module in step S13 above, the embodiment shown in Figure 2, and the embodiment corresponding to Figure 5.
[0141] The fusion module 703 can obtain a motion representation by fusing features of different modalities based on the second feature and the supplementary feature. In some embodiments, the fusion module 703 can perform any of the methods performed by the dual-stream attention fusion module in step S15 above, the embodiment shown in FIG3 , and the embodiment corresponding to FIG5 .
[0142] The interactive dual-branch module 704 can process the motion representation through the category encoder network and the boundary encoder network respectively, obtain the action category prediction vector and the temporal boundary prediction vector, determine the action category prediction result, and determine the prediction loss function. In some embodiments, the temporal action segmentation model adjusts parameters according to the prediction loss function until the training is completed. In some embodiments, the interactive dual-branch module 704 can execute any of the methods performed by the multi-stage interactive dual-branch module in step S17 above, the embodiment shown in Figure 4, and the embodiment corresponding to Figure 5.
[0143] Such a training device can extract the features of multimodal data, constrain the information of each modality separately, improve the quality of the features based on which the subsequent fusion steps are based, fully integrate the multimodal features, and improve the accuracy of action segmentation.
[0144] In some embodiments, the training device further includes a loss value determination unit capable of obtaining a weighted sum of an intermediate category loss function, an intermediate boundary loss function, an action classification loss function, and a temporal boundary loss function as a loss function value. In some embodiments, the temporal action segmentation model adjusts parameters based on the loss function value, wherein the intermediate loss function includes an intermediate category loss function and an intermediate boundary loss function, and the prediction loss function includes an action classification loss function and a temporal boundary loss function. In some embodiments, weights can be adjusted as needed.
[0145] Such a training device can balance the intermediate loss and final loss, action category loss and temporal boundary loss as needed, improving the controllability of the model.
[0146] FIG8 is a schematic diagram of some embodiments of the time sequence action segmentation device disclosed herein.
[0147] The data acquisition module 801 can acquire data and supplementary data of each modality of the target.
[0148] Prediction module 802 inputs the data for each modality and the supplementary dimension data into a temporal action segmentation model to obtain a prediction result for the target action category. The temporal action segmentation model can be any of the models mentioned above, trained using any of the training methods or devices for training any of the models mentioned above.
[0149] Such a device can extract features of multimodal data, constrain the information of each modality separately, improve the quality of features based on subsequent fusion steps, fully integrate multimodal features, and improve the accuracy of action segmentation.
[0150] A structural diagram of an embodiment of a data processing device disclosed in the present invention is shown in Figure 9. The data processing device includes a memory 901 and a processor 902. Among them: the memory 901 can be a disk, a flash memory or any other non-volatile storage medium. The memory is used to store the instructions in the training method of the temporal action segmentation model or the corresponding embodiment of the temporal action segmentation method. The processor 902 is coupled to the memory 901 and can be implemented as one or more integrated circuits, such as a microprocessor or a microcontroller. The processor 902 is used to execute the instructions stored in the memory, which can improve the accuracy of action segmentation.
[0151] In one embodiment, as shown in FIG10 , a data processing device 1000 includes a memory 1001 and a processor 1002. The processor 1002 is coupled to the memory 1001 via a BUS 1003. The data processing device 1000 can also be connected to an external storage device 1005 via a storage interface 1004 to access external data, and can also be connected to a network or another computer system (not shown) via a network interface 1006. A detailed description is omitted here.
[0152] In this embodiment, the accuracy of action segmentation can be improved by storing data instructions in a memory and then processing the instructions through a processor.
[0153] In another embodiment, a computer-readable storage medium stores computer program instructions thereon, which, when executed by a processor, implement the training method of a temporal action segmentation model or the steps of the method in the corresponding embodiment of the temporal action segmentation method. It should be understood by those skilled in the art that the embodiments of the present disclosure may be provided as methods, devices, or computer program products. Therefore, the present disclosure may take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present disclosure may take the form of a computer program product implemented on one or more computer-usable non-transient storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0154] The present disclosure is described with reference to the flowcharts and / or block diagrams of the methods, devices (systems) and computer program products according to the embodiments of the present disclosure. It should be understood that each process and / or box in the flowchart and / or block diagram and the combination of the processes and / or boxes in the flowchart and / or block diagram can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, an embedded processor or other programmable data processing device to produce a machine, so that the instructions executed by the processor of the computer or other programmable data processing device generate a device for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0155] These computer program instructions may also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to operate in a specific manner, so that the instructions stored in the computer-readable memory produce a product including an instruction device that implements the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0156] These computer program instructions can also be loaded onto a computer or other programmable data processing device so that a series of operating steps are executed on the computer or other programmable device to produce a computer-implemented process, so that the instructions executed on the computer or other programmable device provide steps for implementing the functions specified in one or more processes in the flowchart and / or one or more boxes in the block diagram.
[0157] The present disclosure has been described in detail so far. To avoid obscuring the concept of the present disclosure, some details known in the art have not been described. Based on the above description, those skilled in the art can fully understand how to implement the technical solutions disclosed herein.
[0158] The methods and apparatus of the present disclosure may be implemented in many ways. For example, the methods and apparatus of the present disclosure may be implemented by software, hardware, firmware, or any combination of software, hardware, and firmware. The above order of steps for the method is for illustration only, and the steps of the method of the present disclosure are not limited to the order specifically described above, unless otherwise specifically stated. In addition, in some embodiments, the present disclosure may also be implemented as programs recorded in a recording medium, which include machine-readable instructions for implementing the methods according to the present disclosure. Therefore, the present disclosure also covers recording media that store programs for executing the methods according to the present disclosure.
[0159] It should be noted that the terms "first", "second", etc. in the specification, claims, and drawings of the present disclosure are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way are interchangeable where appropriate so that the embodiments of the present disclosure described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions, for example, a process, method, system, product, or apparatus that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products, or apparatus.
[0160] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present disclosure and not to limit it. Although the present disclosure has been described in detail with reference to the preferred embodiments, ordinary technicians in the relevant field should understand that the specific implementation methods of the present disclosure can still be modified or some technical features can be replaced by equivalents without departing from the spirit of the technical solutions of the present disclosure, which should all be included in the scope of the technical solutions requested for protection in the present disclosure.
Claims
1. A training method for a temporal action segmentation model, comprising: Performing high-dimensional feature extraction on data of each modality in the sample data to obtain first features of each modality; According to the first features, through dimensionality reduction processing, obtaining an intermediate loss function, and through dimensionality restoration, obtaining second features of each modality; According to the second features and supplementary features, through the fusion of features of different modalities, obtaining a motion representation; Processing the motion representation through a class encoder network and a boundary encoder network respectively to obtain an action class prediction vector and a temporal boundary prediction vector, determining an action class prediction result, and determining a prediction loss function, wherein the temporal action segmentation model adjusts parameters according to the intermediate loss function and the prediction loss function until the training is completed.
2. The training method according to claim 1, wherein The obtaining a motion representation by fusing features of different modalities according to the second features and supplementary features includes: Obtaining an attention-enhanced motion flow feature and an attention-enhanced spatial flow feature according to the second features and supplementary features; Obtaining a cross-flow attention feature through a multi-head cross-attention mechanism; According to the cross-flow attention feature, the attention-enhanced motion flow feature, and the attention-enhanced spatial flow feature, obtaining a motion feature vector and a spatial feature vector; According to the motion feature vector and the spatial feature vector, determining the motion representation.
3. The training method according to claim 1 or 2, wherein The obtaining an intermediate loss function by dimensionality reduction processing according to the first features and obtaining second features of each modality through dimensionality restoration includes: Performing temporal encoding on the first features, determining intermediate class features through a fully connected layer, and determining an intermediate class loss function; Restoring the dimension of the intermediate class features and adding them to the first features to obtain the second features; Performing dimensionality reduction processing and temporal information extraction on the second features to obtain intermediate boundary features, and obtaining an intermediate boundary loss function according to the intermediate boundary features, wherein the intermediate loss function includes the intermediate class loss function and the intermediate boundary loss function.
4. The training method according to any one of claims 1 to 3, wherein, The processing the motion representation through a class encoder network and a boundary encoder network respectively to obtain an action class prediction vector and a temporal boundary prediction vector and determining an action class prediction result includes: At the start of the encoder network, using the query feature vector in the class encoder network to retrieve the key-value pair feature vector in the boundary encoder network to obtain boundary features; At the end of the encoder network, using the query feature vector in the boundary encoder network to retrieve the key-value pair feature vector in the boundary encoder to obtain class features; According to the boundary features and the class features, obtaining an action class prediction vector and a temporal boundary prediction vector through dimensionality adjustment; In the temporal boundary prediction vector, screening out boundary values according to the confidence level, and adjusting the action class prediction vector according to the boundary values to obtain the action class prediction result.
5. The training method according to any one of claims 1 to 4, wherein, The determining the prediction loss function includes: Obtaining an action classification loss function corresponding to the action class prediction result; Obtaining a boundary probability vector according to the temporal boundary prediction vector, and determining a temporal boundary loss function according to the boundary probability vector, Among them, the prediction loss function includes the action classification loss function and the temporal boundary loss function.
6. The training method according to claim 2, wherein The obtaining of the attention-enhanced motion flow feature and the attention-enhanced spatial flow feature according to the second feature and the supplementary feature includes: Obtaining an initial motion flow feature according to the second feature of the inertial modality, and obtaining an initial spatial flow feature according to the second features of the key point modality and the bounding box modality, where the modalities include inertia, key points, and bounding boxes; Adding the initial motion flow feature and the initial spatial flow feature to the supplementary feature respectively to obtain an auxiliary enhanced motion flow feature and an auxiliary enhanced spatial flow feature; According to the auxiliary enhanced motion flow feature and the supplementary feature, using a multi-head attention mechanism to obtain the attention-enhanced motion flow feature and the attention-enhanced spatial flow feature.
7. The training method according to claim 2 or 6, wherein The obtaining of the motion feature vector and the spatial feature vector according to the cross-flow attention feature, the attention-enhanced motion flow feature, and the attention-enhanced spatial flow feature includes: Processing the first cross-flow attention feature and the attention-enhanced motion flow feature through a frame-by-frame convolutional layer to obtain a third feature; Processing the second cross-flow attention feature and the attention-enhanced spatial flow feature through a frame-by-frame convolutional layer to obtain a fourth feature; Adding the third feature, the first cross-flow attention feature, and the attention-enhanced motion flow feature to obtain the motion feature vector; Adding the fourth feature, the second cross-flow attention feature, and the attention-enhanced spatial flow feature to obtain the spatial feature vector, where the first cross-flow attention feature is the cross-flow attention feature obtained by matching the query feature vector of the spatial flow with the key-value pair feature vector of the motion flow, and the second cross-flow attention feature is the cross-flow attention feature obtained by matching the query feature vector of the motion flow with the key-value pair feature vector of the spatial flow.
8. The training method according to claim 2, 6 or 7, wherein The determining of the motion representation according to the motion feature vector and the spatial feature vector includes: Adding the motion feature vector and the spatial feature vector as the motion representation.
9. The training method according to claim 3, wherein, The obtaining of the intermediate boundary feature by performing dimensionality reduction processing and temporal information extraction on the second feature, and obtaining the intermediate boundary loss function according to the intermediate boundary feature includes: Processing the second feature through a one-layer convolutional neural network to obtain a fifth feature; Obtaining the temporal information in the fifth feature through a multi-layer convolutional neural network with residual connections and adding it to the fifth feature to determine the intermediate boundary feature; Obtaining a boundary classification probability vector by passing the intermediate boundary feature through an activation function; Determining the intermediate boundary loss function according to the boundary classification probability vector.
10. The training method according to any one of claims 1 to 9 further includes: Obtaining the weighted sum of the intermediate class loss function, the intermediate boundary loss function, the action classification loss function, and the temporal boundary loss function as the loss function value. Among them, the temporal action segmentation model adjusts parameters according to the loss function value. The intermediate loss function includes the intermediate class loss function and the intermediate boundary loss function. The prediction loss function includes the action classification loss function and the temporal boundary loss function.
11. The training method according to any one of claims 1 to 9, wherein, The supplementary feature is determined according to the working log of the Internet of Things device.
12. A temporal action segmentation method, comprising: Obtaining data of each modality of a target and supplementary data; Inputting the data of each modality and the supplementary dimension data into the temporal action segmentation model to obtain a predicted result of the action category of the target, where the temporal action segmentation model is generated according to the training method of the temporal action segmentation model according to any one of claims 1 to 11.
13. The timing action segmentation method according to claim 12, wherein, The method conforms to at least one of the following: The modality includes inertia, key points, and bounding boxes; or The supplementary data includes the working log of the Internet of Things device.
14. A training device for a temporal action segmentation model, comprising: A feature extraction module configured to perform high-dimensional feature extraction on data of each modality in sample data to obtain a first feature of each modality; A feature constraint module configured to obtain an intermediate loss function through dimensionality reduction processing according to the first feature, and obtain a second feature of each modality through dimension restoration; A fusion module configured to obtain a motion representation through fusion of features of different modalities according to the second feature and supplementary features; and An interactive double-branch module configured to process the motion representation through a class encoder network and a boundary encoder network respectively to obtain an action category prediction vector and a temporal boundary prediction vector, determine a predicted result of the action category, and determine a prediction loss function, where the temporal action segmentation model adjusts parameters according to the intermediate loss function and the prediction loss function until the training is completed.
15. A temporal action segmentation device, comprising: A data acquisition module configured to acquire data of each modality of a target and supplementary data; and A prediction module configured to input the data of each modality and the supplementary dimension data into the temporal action segmentation model to obtain a predicted result of the action category of the target, where the temporal action segmentation model is generated according to the training method of the temporal action segmentation model according to any one of claims 1 to 11.
16. A data processing device, comprising: A memory; and A processor coupled to the memory, the processor being configured to execute the method according to any one of claims 1 to 13 based on instructions stored in the memory.
17. A computer-readable storage medium, on which computer program instructions are stored, and when the instructions are executed by a processor, the steps of the method according to any one of claims 1 to 13 are implemented.
18. A computer program product, comprising a computer program or instructions, and when the computer program or instructions are executed by a processor, the method according to any one of claims 1 to 13 is implemented.
19. A computer program for causing a processor to execute the method according to any one of claims 1 to 13.
Citation Information
Patent Citations
Multi-modal fusion sign language recognition system and method based on graph convolution
CN111259804A
Training method and device of detection segmentation model, electronic equipment and storage medium
CN115249304A
Model training method and device, model identification method and device, processing equipment and storage medium
CN115859112A
Model training method and device, electronic equipment and storage medium
CN116070169A
Time sequence action segmentation method and device, model training method and device and storage medium
CN117876934A
Cited By
Time series data new category discovery method based on multi-modal fusion
CN120726410A
Video conference multi-modal data alignment method and device based on causal mask, equipment and medium
CN120763869A
A video conference multi-modal data alignment method and device based on a causal mask, equipment and medium
CN120763869B
Interactive radiotherapy planning system optimization method
CN120960658A
Weakly supervised group behavior identification method and system based on double-branch space-time motion fusion network
CN121982780A