Action anomaly detection method and system, terminal and storage medium

Through the spatiotemporal graph convolution and spatiotemporal decoupling technology of the action abnormality detection model, combined with pseudo-label generation and comparison learning, the problem of low accuracy of action abnormality detection is solved, and efficient and accurate screening of children with ADHD is achieved.

CN120279601AActive Publication Date: 2025-07-08JIANGXI NORMAL UNIV
View PDF 9 Cites 0 Cited by

Patent Information

Application Number
CN202510756823.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-08
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

In the prior art, the accuracy of motion abnormality detection is low, mainly relying on doctors' subjective judgment and questionnaire surveys, resulting in inaccurate test results.

Method used

The action anomaly detection model is used to perform spatiotemporal graph convolution, extract the spatial-temporal graph features of the sample, and combine the pseudo-label generation and comparison learning of fuzzy label samples to optimize the model parameters until the model converges.

Benefits of technology

It improves the accuracy of motion abnormality detection, can objectively and quickly identify children with ADHD, reduces dependence on subjective judgments, and improves the efficiency and accuracy of detection.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279601A_ABST
    Figure CN120279601A_ABST
Patent Text Reader

Abstract

The invention provides a motion anomaly detection method and system, a terminal and a storage medium, and the method comprises the steps: obtaining sample data, inputting the sample data into a motion anomaly detection model, and carrying out the space-time diagram convolution, and obtaining a sample space-time diagram feature; performing space-time decoupling on the sample space-time diagram feature to obtain a sample decoupling feature, and performing category prediction on the sample decoupling feature to obtain an action prediction result; determining model loss according to the motion prediction result and the sample decoupling features, and performing parameter updating on the motion anomaly detection model according to the model loss until the motion anomaly detection model converges; and inputting to-be-detected data into the converged motion anomaly detection model for motion anomaly detection to obtain a motion anomaly detection result. According to the embodiment of the invention, action anomaly detection can be effectively carried out on the to-be-detected data based on the converged action anomaly detection model, and the accuracy of action anomaly detection is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of action detection, and in particular, to a method, system, terminal and storage medium for detecting abnormal actions. Background Art

[0002] Attention Deficit Hyperactivity Disorder (ADHD), commonly known as hyperactivity disorder, is the most common neurodevelopmental disorder in children and adolescents, and its clinical manifestations are difficulty in concentrating, hyperactivity, impulsiveness, emotional instability, learning difficulties, etc.

[0003] In the prior art, it is mainly generally judged whether there is an abnormal action of ADHD in patients through the subjective judgment of doctors and questionnaires, resulting in low accuracy of abnormal action detection. Summary of the Invention

[0004] The purpose of the embodiments of the present invention is to provide a method, system, terminal and storage medium for detecting abnormal actions to solve the problem of low accuracy of abnormal action detection in the prior art.

[0005] The embodiments of the present invention are implemented as follows. An abnormal action detection method, the method includes: Obtain sample data, and input the sample data into an abnormal action detection model for spatio-temporal graph convolution to obtain sample spatio-temporal graph features; Perform spatio-temporal decoupling on the sample spatio-temporal graph features to obtain sample decoupled features, and perform class prediction on the sample decoupled features to obtain an action prediction result; Determine a model loss according to the action prediction result and the sample decoupled features, and update the parameters of the abnormal action detection model according to the model loss until the abnormal action detection model converges; Input the data to be detected into the converged abnormal action detection model for abnormal action detection to obtain an abnormal action detection result.

[0006] Preferably, inputting the sample data into an abnormal action detection model for spatio-temporal graph convolution to obtain sample spatio-temporal graph features includes: Calculate the neighborhood node indices of the feature nodes in the sample data, determine local neighborhood features according to the neighborhood node indices, and calculate edge convolution features according to the local neighborhood features; Perform graph convolution on the sample data to obtain graph convolution features, and splice the graph convolution features and the edge convolution features to obtain spliced features; Perform representative spatial average pooling processing on the spliced features to obtain average pooling features, and perform hierarchical edge convolution on the average pooling features to obtain an attention map; Combine the attention map with the graph convolutional features to obtain combined features, and sum the combined features in the hierarchical dimension to obtain the sample spatio-temporal graph features.

[0007] Preferably, the formula for calculating the neighborhood node indices of the feature nodes in the sample data includes: where, is the feature node i 's n neighborhood node indices, represents the i th feature of the feature node, represents the j th feature of the feature node, represents the proximity algorithm function, n represents the number of proximities set in the proximity algorithm function; The formula for calculating the edge convolution features based on the local neighborhood features includes: where, represents a linear transformation, represents feature concatenation, represents max pooling of the features in dimension n , represents replicating the features T times, represents the edge convolution features, X is the feature of the feature node, X n is the local neighborhood feature; The formula for performing graph convolution on the sample data includes: where, represents the k th hierarchical graph convolution feature, represents the input sample data, represents pointwise convolution operation, T represents the temporal window size, represents the k th level adjacency matrix, S represents the total edge set of human joint points, s represents the edge subset of human joint points; The formula for performing representative spatial average pooling on the concatenated features includes: where, represents the RSAP function, N k represents the level Hk The number of joints in v represents a joint point, t represents a temporal frame, represents the joint point at the k-th level v On the temporal frame t the eigenvalue, represents the average pooling feature, and max represents the maximum value function; The formula for performing hierarchical edge convolution on the average pooling feature includes: where M represents the attention map, represents hierarchical edge convolution, σ represents the sigmoid activation function, L represents the total set of human joint points at different levels.

[0008] Preferably, decoupling the spatio-temporal features of the sample to obtain sample decoupled features includes: Performing temporal pooling on the spatio-temporal features of the sample to obtain a first pooled feature, and performing convolution on the first pooled feature to obtain sample spatial features; Performing spatial pooling on the spatio-temporal features of the sample to obtain a second pooled feature, and performing convolution on the second pooled feature to obtain sample temporal features; The sample decoupled features include the sample temporal features and the sample spatial features.

[0009] Preferably, obtaining sample data includes: Obtaining fuzzy label samples, and performing sample enhancement on the fuzzy label samples to obtain enhanced samples; Performing label prediction on the fuzzy label samples and the enhanced samples according to the action anomaly detection model to obtain a first predicted label and a second predicted label; Calculating the average value of the first predicted label and the second predicted label to obtain a label average value, and sharpening the label average value to obtain a pseudo-label; Performing label marking on the fuzzy label samples and the enhanced samples according to the pseudo-label, and obtaining clear label samples; Combining the clear label samples, the fuzzy label samples after label marking, and the enhanced samples to obtain the sample data.

[0010] Preferably, the formula for determining the model loss according to the action prediction result and the sample decoupled features includes: where represents the model loss, W CL represents the first weight hyperparameter, W u represents the second weight hyperparameter, represents the cross - entropy loss, represents the mean squared error loss, represents the contrastive loss; where, represents the number of samples of the fuzzy label samples, u represents the fuzzy label samples, q represents the pseudo - label, represents the class probability predicted by the fuzzy label samples in the action prediction result, represents the data of the fuzzy label samples, represents the model parameters of the action anomaly detection model; where, represents the number of samples of the clear label samples, x represents the clear label samples, p represents the true label of the clear label samples, represents the class probability predicted by the clear label samples in the action prediction result, L ′ represents the data of the clear label samples; where, represents the loss weight, represents the sample space feature, represents the sample time feature, represents the contrastive learning loss function, n represents the number of times of contrastive learning construction.

[0011] Preferably, the contrastive learning loss function is: where, represents the prediction probability of sample i for class k in the action prediction result, τ represents the sensitivity hyperparameter, represents the compensation term, represents the penalty term, represents the prototype representation of class k represents the prototype representation of class F ​​i representing a sample i of the sample time feature or the sample space feature; wherein, representing a class k representation of the central feature of a false positive sample, representing a confidence sample set, representing the data volume of the confidence sample set, representing a class k data volume of the set of false positive samples; wherein, representing a class k representation of the central feature of a false negative sample, representing a class k data volume of the set of false negative samples; wherein, F j representing a sample j of the sample time feature or the sample space feature; wherein, α representing a momentum term.

[0012] Another object of the embodiments of the present invention is to provide an action anomaly detection system, the system comprising: a spatio-temporal graph convolution module, configured to obtain sample data and input the sample data into an action anomaly detection model for spatio-temporal graph convolution to obtain sample spatio-temporal graph features; a class prediction module, configured to perform spatio-temporal decoupling on the sample spatio-temporal graph features to obtain sample decoupled features, and perform class prediction on the sample decoupled features to obtain an action prediction result; a model training module, configured to determine a model loss according to the action prediction result and the sample decoupled features, and update parameters of the action anomaly detection model according to the model loss until the action anomaly detection model converges; an action anomaly detection module, configured to input data to be detected into the converged action anomaly detection model for action anomaly detection to obtain an action anomaly detection result.

[0013] In the embodiments of the present invention, by inputting sample data into an action anomaly detection model for spatio-temporal graph convolution, sample spatio-temporal graph features in the sample data can be effectively extracted. By performing spatio-temporal decoupling on the sample spatio-temporal graph features, sample decoupled features can be effectively extracted. By performing class prediction on the sample decoupled features, an action prediction result for the sample data can be effectively obtained. Based on the model loss, the parameters of the action anomaly detection model are updated, so that the converged action anomaly detection model can effectively perform action anomaly detection on the data to be detected, without using subjective judgment and questionnaire surveys for action anomaly detection of attention deficit hyperactivity disorder, improving the accuracy of action anomaly detection. BRIEF DESCRIPTION OF THE DRAWINGS

[0014] Figure 1 is a flowchart of an action anomaly detection method provided by the first embodiment of the present invention; Figure 2 is a schematic diagram of an action anomaly detection model provided by the first embodiment of the present invention; Figure 3 is a schematic diagram of a feature refinement module provided by the first embodiment of the present invention; Figure 4 is a schematic diagram of a feature mixing module provided by the first embodiment of the present invention; Figure 5 is a schematic structural diagram of an action anomaly detection system provided by the second embodiment of the present invention; Figure 6 is a schematic structural diagram of a terminal device provided by the third embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0015] In order to make the objectives, technical solutions and advantages of the present invention clearer, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0016] In order to illustrate the technical solutions described in the present invention, the following will be described through specific embodiments.

[0017] Embodiment 1 Please refer to Figure 1 , which is a flowchart of an action anomaly detection method provided by the first embodiment of the present invention. This action anomaly detection method can be applied to any device or system. This action anomaly detection method includes the steps: Step S10, obtain sample data, and input the sample data into an action anomaly detection model for spatio-temporal graph convolution to obtain sample spatio-temporal graph features; Please refer to Figure 2, the action anomaly detection model adopts a spatio-temporal graph convolutional semi-supervised ADHD recognition model (ADHD Recognition Model, ADHD-RM) based on the temporal characteristics of human skeletons. The structure of the action anomaly detection model consists of ten nested spatio-temporal graph convolutional modules (Spatio Temporal Graph Convolution Module, STGCM), a feature refinement module (Feature Refinement Module, FRM), a feature blend module (Feature Blend Module, FBM), a linear classification layer, and a Softmax activation function.

[0018] Taking the three-dimensional human pose sequence as the model input, in the first 20 iterations of the training of the action anomaly detection model, only the clear label sample data is used to jointly pre-train the spatio-temporal graph convolutional module and the feature refinement module. Due to the label noise problem in the fuzzy label sample data, it is regarded as unlabeled data.

[0019] In this step, spatio-temporal graph convolution is performed on the sample data based on the spatio-temporal graph convolutional module to obtain the sample spatio-temporal graph features. The spatio-temporal graph convolutional module represents the human bone sequence as a graph network composed of several layers of joint nodes and bone edges, and extracts the joint information of time series and space through unified spatio-temporal modeling. In order to enable the model to focus on local and global actions in the action stimulus paradigm, an attention-guided hierarchical aggregation mechanism is combined, and a weighted summation strategy is applied to the outputs of different levels, so that the model can focus on key local or global features, thereby helping the model to distinguish the subtle differences between ADHD subjects and normal subjects under similar action stimuli.

[0020] Optionally, inputting the sample data into the action anomaly detection model for spatio-temporal graph convolution to obtain the sample spatio-temporal graph features includes: Calculating the neighborhood node indices of the feature nodes in the sample data, determining the local neighborhood features according to the neighborhood node indices, and calculating the edge convolution features according to the local neighborhood features; Among them, an edge convolution is used to construct a local neighborhood graph to extract the local neighborhood features in the graph structure. Through the K-Nearest Neighbor (K-NN) algorithm based on the Euclidean distance between two nodes in the feature space, the n adjacent nodes with the smallest distance are selected to obtain the neighborhood node indices. Calculate the neighborhood node indices of each node, so as to obtain the local neighborhood features of each node.

[0021] Performing graph convolution on the sample data to obtain graph convolution features, and splicing the graph convolution features and the edge convolution features to obtain the spliced features; Among them, in the human skeleton structure, the spinal nodes of a person are selected as the central nodes, and the human skeleton structure is unfolded into a hierarchical tree structure to obtain a set of hierarchical nodes. In the human skeleton structure, the central nodes are at the first level, the nodes connected to the central nodes are at the second level, and the nodes connected to the second-level nodes are at the third level. The set of hierarchical nodes contains the hierarchical information of the graph and is represented as an adjacency matrix , N L represents the number of levels, N S is the number of edge subset types defined as 3. Through hierarchical decomposition, the edges between all nodes in the same semantic space are obtained by connecting all nodes in the adjacent-level edge sets, forming a structure of a hierarchical graph with fully connected edges; Adjacency matrix is defined as: Among them, H k represents the k th set of hierarchical nodes at the level, H k+1 represents the k+ 1st set of hierarchical nodes at the level, NL represents the number of level sets, s id , s cp , s cf represent the same-edge subset, centripetal-edge subset, and centrifugal-edge subset respectively, represents the splicing of the three edge subsets. The same-edge subset represents the set of all joint points in the upper and lower levels and the edges connecting to their own joint points. For example, if the first level contains a joint point e1 and the second level contains e2 and e3, then the same-edge subset between the first level and the second level includes the sets of the edges connecting e1, e2, and e3 to their own joint points respectively. The centripetal-edge subset represents the set of directed edges connecting the joint points in the upper level to the joint points in the lower level, and the centrifugal-edge subset represents the set of directed edges connecting the joint points in the lower level to the joint points in the upper level; S represents the total set of the same-edge subset, centripetal-edge subset, and centrifugal-edge subset.

[0022] Perform representative spatial average pooling on the splicing feature to obtain an average pooling feature, and perform hierarchical edge convolution on the average pooling feature to obtain an attention map; Among them, the Attention-Guided Hierarchy Aggregation (A-HA) module is applied to perform Representative Spatial Average Pooling (RSAP) on the concatenated features. The task of RSAP is to extract the temporal frames with the maximum scores from each level, and calculate the weighted values based on the features of these frames to avoid the scaling bias caused by inconsistent numbers of node connection edges; After the RSAP layer, all hierarchical features are regarded as nodes on the graph, and the similarities between levels are learned through hierarchical edge convolution to obtain the attention map; The attention map is combined with the graph convolution features to obtain combined features, and the combined features are summed along the hierarchical dimension to obtain the sample spatio-temporal graph features; Among them, the attention map is multiplied by the graph convolution features and summed along the hierarchical dimension to obtain the sample spatio-temporal graph features. Compared with the traditional temporal convolution method with a fixed convolution kernel size, this step convolves the output features of the graph convolution in the time dimension through convolution kernels of different sizes, so as to extract the temporal features at different time scales, and the output dimension is the input dimension of the next layer of graph convolution. This step does not simply rely on a larger convolution kernel, but combines dilated convolution to expand the receptive field, so that while improving the computational efficiency of the model, it can also extract the temporal sequence with a longer time span. Finally, the sample spatio-temporal graph features contain richer temporal information and can effectively improve the model's ability to extract features from skeleton data.

[0023] Furthermore, the formulas for calculating the neighborhood node indices of the feature nodes in the sample data include: Among them, is the feature node i 's n neighborhood node indices, represents the i th feature of the feature node, represents the j th feature of the feature node, represents the proximity algorithm function, n represents the number of neighbors set in the proximity algorithm function; The formulas for calculating the edge convolution features based on the local neighborhood features include: Among them, represents the linear transformation, represents the feature concatenation, represents the dimension of the featuren perform max pooling on it indicating that the feature is replicated T times indicating the edge convolution feature X is the feature of the feature node X n is the local neighborhood feature; The formula for performing graph convolution on the sample data includes: where represents the k -th level graph convolution feature represents the input sample data represents pointwise convolution operation, T represents the time series window size represents the k -th level adjacency matrix, S represents the total edge set of human joint points, s represents the edge subset of human joint points; The formula for performing representative spatial average pooling on the concatenated feature includes: where represents the RSAP function N k represents the level H k the number of joint points in v represents the joint point t represents the time series frame represents the joint point at the k-th level v at the time series frame t the feature value on represents the average pooling feature, max represents the maximum value function; The formula for performing hierarchical edge convolution on the average pooling feature includes: where M represents the attention map represents hierarchical edge convolution σ represents the sigmoid activation function L represents the hierarchical total set of human joint points

[0024] Furthermore, obtaining sample data includes: Obtain fuzzy label samples, and perform sample enhancement on the fuzzy label samples to obtain enhanced samples; According to the action anomaly detection model, perform label prediction on the fuzzy label samples and the enhanced samples to obtain the first prediction label and the second prediction label; Calculate the average value of the first predicted label and the second predicted label to obtain a label average value, and sharpen the label average value to obtain a pseudo-label; Among them, the blurred label sample is randomly rotated by an angle around the y-axis to obtain an augmented sample. The pseudo-labels of the blurred label sample and its augmented sample are predicted by an action anomaly detection model. After averaging the pseudo-labels of the two generated samples, they are sharpened and used as the common pseudo-label of these two samples. Label the blurred label sample and the augmented sample according to the pseudo-label, and obtain clear label samples; Combine the clear label samples, the blurred label samples and the augmented samples after label marking to obtain the sample data.

[0025] In step S20, perform spatio-temporal decoupling on the sample spatio-temporal graph features to obtain sample decoupled features, and perform class prediction on the sample decoupled features to obtain an action prediction result; Among them, please refer to Figure 3 , the feature refinement module consists of a spatio-temporal decoupling module and a contrast learning module. The spatio-temporal decoupling module extracts temporal features and spatial features through pooling and convolution operations respectively, thereby reducing the interference of similar actions in time and space. The decoupled features can represent independent changes in time and space, improving the fineness of feature expression. At the same time, the contrast learning module optimizes the decoupled feature expression, making the cosine distance of the action features of subjects with ADHD and normal subjects under the same paradigm as large as possible, thereby improving the classification accuracy and enhancing the discriminative ability of the model for difficult samples.

[0026] Optionally, performing spatio-temporal decoupling on the sample spatio-temporal graph features to obtain sample decoupled features includes: Perform temporal pooling on the sample spatio-temporal graph features to obtain a first pooled feature, and perform convolution on the first pooled feature to obtain sample spatial features; Perform spatial pooling on the sample spatio-temporal graph features to obtain a second pooled feature, and perform convolution on the second pooled feature to obtain sample temporal features; The sample decoupled features include the sample temporal features and the sample spatial features.

[0027] In step S30, determine the model loss according to the action prediction result and the sample decoupled features, and update the parameters of the action anomaly detection model according to the model loss until the action anomaly detection model converges; Among them, if a sample is correctly classified, it is considered a confident sample, and the sample space features and sample time features of the confident sample are collected, and the collected sample space features and sample time features are updated to the global feature representation of the corresponding category through exponential moving average, which is called prototype representation.

[0028] In this embodiment, the fuzzy label samples are divided into two categories: False negative (0 Negative, FN): For example, a sample that should belong to the category k , but is misclassified as another category; False positive (0 Positive, FP): For example, a sample that should belong to another category, but is misclassified as the category k .

[0029] Optionally, the formula for determining the model loss according to the action prediction result and the sample decoupling feature includes: Among them, represents the model loss, W CL represents the first weight hyperparameter, W u represents the second weight hyperparameter, represents the cross-entropy loss, represents the mean squared error loss, represents the contrastive loss; Among them, represents the number of samples of the fuzzy label sample, u represents the fuzzy label sample, q represents the pseudo label, represents the class probability predicted by the fuzzy label sample in the action prediction result, represents the data of the fuzzy label sample, represents the model parameters of the action anomaly detection model; Among them, represents the number of samples of the clear label sample, x represents the clear label sample, p represents the true label of the clear label sample, represents the class probability predicted by the clear label sample in the action prediction result, L ′ represents the data of the clear label sample; Among them, represents the loss weight, representing the sample space feature representing the sample time feature representing the contrastive learning loss function n representing the number of times of contrastive learning construction for the spatio-temporal graph convolution intermediate features at different levels in the model, which is set to 4 in this embodiment

[0030] Furthermore, the contrastive learning loss function is as follows: wherein represents, in the action prediction result, the prediction probability of the sample i for the category k τ represents the sensitivity hyperparameter represents the compensation term represents the penalty term represents the prototype representation of the category k represents the prototype representation of the category F i represents the sample time feature or the sample space feature of the sample i wherein represents the central feature representation of the false positive samples of the category k represents the set of confident samples represents the data volume of the set of confident samples represents the set of false positive samples of the category k represents the data volume of the set of false positive samples

[0031] To calibrate the prediction of fuzzy label samples, the confident samples of the category k are used as anchor points, and a compensation term and a penalty term are introduced in the feature space to adjust the feature representations of FN and FP samples for making the FN samples closer to the confident samples of the category k for making the FP samples farther away from the confident samples of the category k

[0032] wherein represents the central feature representation of the false negative samples of the category k represents the set of false negative samples of the category k ​​​​​​​​​​ The data volume representing the set of false negative samples; Wherein, F j represents the sample j of the sample time feature or the sample space feature; Wherein, α represents the momentum term.

[0033] In this embodiment, the loss function encourages the feature vector of the sample F i to be close to the prototype representation of its belonging category k of the prototype representation p k , punishing F i the feature vector of being close to the prototype representation of non-belonging categories . Using the compensation term and the penalty term, extra guidance is provided for the fuzzy label samples. Through the method of contrastive learning, strong supervision is performed on those TP samples with weak confidence by the confidence samples, so as to optimize and calibrate the classification ability of the model.

[0034] Step S40, input the data to be detected into the converged action anomaly detection model for action anomaly detection, and obtain the action anomaly detection result; Wherein, when the data to be detected in the action anomaly detection result is identified as action anomaly data, it is determined that the person to be detected corresponding to the data to be detected may have ADHD, effectively achieving the ADHD screening effect.

[0035] The Azure-Kinect camera array is used for three-dimensional human pose estimation to obtain the data to be detected of the person to be detected. Since limb occlusion may occur when the human body is moving, the pose estimation accuracy will decrease. This embodiment is based on a multi-view scheme of three Azure-Kinect depth camera arrays, which can combine color information and depth information from multiple angles, provide richer depth information and more comprehensive three-dimensional human pose data, and use a checkerboard calibration board to align the spatial coordinate systems of the three depth cameras, and transform the human pose estimation coordinates of the sub-cameras into the main camera coordinate system. Combining the median fusion method of pose coordinates can improve the accuracy and robustness of pose estimation.

[0036] The time alignment method of the Azure-Kinect camera is to connect the synchronization ports of multiple cameras in a daisy-chain manner using a 3.5mm audio cable. The synchronization principle is to calculate the time stamp difference between the master device and the slave device and set a fixed threshold for judgment. If the time stamp difference exceeds the set threshold, the device with the smaller time stamp will reshoot the image to achieve time synchronization again.

[0037] Use a self-made rigid camera calibration board of 60cm * 60cm, where the size of a single checkerboard is 5cm * 5cm. The person to be detected stands in front of the Azure-Kinect depth camera array while holding the calibration board, starts the cameras to shoot synchronously, and slowly moves the calibration board to ensure that each camera's video sequence contains 200 frames of calibration board images.

[0038] Define w as the world coordinate system, with the origin set at the corner point of the upper left corner of the calibration board, and define the camera coordinate system as o. The internal parameter matrix of each camera is respectively , and the internal parameter matrix is responsible for converting the coordinate points in the camera coordinate system to the pixel coordinate system. Select the camera m in the array as the reference camera. In the camera coordinate system, the Euclidean transformation from the sub-camera k to the reference camera m is: Among them, ok represents the camera coordinate system of the sub-camera k , om represents the camera coordinate system of the reference camera m . w represents the world coordinate system. R represents the rotation matrix in the transformation matrix, describes the rotation relationship between the camera coordinate system of the sub-camera k and the camera coordinate system of the reference camera m ; the translation vector t represents the translation amount from the origin of the coordinate system of the sub-camera k to the origin of the coordinate system of the reference camera m . By using the world coordinate system as the intermediate coordinate system, both the rotation relationship and the translation vector are converted to the world coordinate system and then to the camera coordinate system of the reference camera m , and the formula is as follows: Among them, T represents matrix transpose, represents the human body estimation point coordinates of the sub-camera k to the reference camera mThe camera coordinate system transformation matrix, is a series of 3D points in the coordinate system of the checkerboard calibration board, and are the 2D coordinates projected onto the pixel coordinate system. The PnP algorithm calculates the Euclidean transformation based on the correspondence between 3D points and 2D points k . The remaining calculations are the same as the above formula. After calculating the transformation matrix from any camera k to the reference camera m , the spatial alignment of the three cameras can be achieved, and the 3D skeleton information predicted by the sub-camera can be transformed to the coordinate system of the reference camera m .

[0039] In the RGB-D camera array, each camera can estimate the positions of 32 joint points of the human body. For the j-th joint point estimated by the k-th camera, its three-dimensional coordinates can be expressed as: where, j represents the node index among the 32 joint points of the human body, represents the k coordinate position of the x axis of the j-th joint point among the 32 joint points of the human body in the sub-camera represents the k coordinate position of the y axis of the j-th joint point among the 32 joint points of the human body in the sub-camera represents the k coordinate position of the z axis of the j-th joint point among the 32 joint points of the human body in the sub-camera represents the set of three-dimensional coordinate values of the j-th joint point among the 32 joint points of the human body in the sub-camera k .

[0040] To fuse the estimation results of multiple cameras, the three-dimensional coordinate points obtained by the sub-camera k need to be transformed to the coordinate system of the reference camera m . Based on the Euclidean transformation formula between cameras, a point can be transformed from the coordinate system of camera k to the coordinate system of the reference camera m : For each joint point j , the estimation result after being transformed from the sub-camera to the coordinate system of the reference camera m is a set . To obtain the final three-dimensional coordinate estimation, the formula is as follows: where, represents the coordinate of the human body estimation point of the first Azure-Kinect camera in the reference cameram Coordinate representation in the coordinate system Indicates the coordinates of the human body estimation points of the second Azure-Kinect camera in the reference camera m Coordinate representation in the coordinate system Indicates the coordinates of the human body estimation points of the third Azure-Kinect camera in the reference camera m Coordinate representation in the coordinate system Indicates the final estimation after fusing the coordinates of the three Azure-Kinect cameras

[0041] All points in the set are 3D estimations of the same point in the camera coordinate system of the reference camera m. Considering that there will be a certain amount of limb occlusion when the human body is moving, which causes coordinate jitter in the human body pose estimation by the Azure-Kinect camera, a function is defined to take the median value of the (x, y, z) of the joint point estimations of all cameras to obtain a relatively stable and accurate human body pose estimation structure, that is, to obtain the data to be detected

[0042] Please refer to Figure 4 , for the problem of limited data volume in the dataset, this embodiment designs a feature mixing module and adopts a semi-supervised learning strategy for model training. Since the fuzzy label samples may have certain noisy labels, directly relying on the initial labels will have a negative impact on model training. To solve this problem, the model is used to generate pseudo-labels for training based on the prediction results of the fuzzy label samples and their data augmentation samples, rather than using the initial labels. Then, the feature mixing module mixes the high-dimensional features of the clear label samples and the fuzzy label samples by Mixup to mine the potential information in the fuzzy label samples, enabling the model to learn more diverse sample features, which helps the model better adapt to unknown or fuzzy label samples and improves the classification accuracy

[0043] The Mixup method is a data augmentation technique, and its basic idea is to perform linear interpolation on two samples and their corresponding labels according to weights. The formula is as follows Among them, x i and x j are the features of two samples respectively y i and y j are respectively x i , x j the corresponding labels is from BetaMixing coefficients sampled from the distribution, Control the weight distribution of sampling, Indicates based on x i and x j The mixed features after feature mixing, Indicates based on y i and y j The mixed labels after label mixing.

[0044] In the study of pose sequences, since 3D pose data contains rich spatio-temporal information, directly performing the Mixup operation on sequence data may cause the spatio-temporal information of samples to be disordered, resulting in the loss of key features. Therefore, in this embodiment, the Mixup operation is performed in the high-dimensional feature space after feature extraction of clear-label samples and fuzzy-label samples, so as to generate mixed features and corresponding mixed pseudo-labels. The new mixed features and mixed labels are regarded as the feature of new training samples for subsequent classification training.

[0045] Compared with the original Mixup method, this embodiment utilizes the semantic expression ability of hidden layer features and the smooth characteristics of the distribution to generate interpolation samples with more reasonable semantics in the feature space. This can not only generate new samples through interpolation in the feature space, but also directly affect the intermediate representation of the neural network, providing more effective training information for the model, thereby prompting the model to learn more robust feature representations and decision boundaries.

[0046] In this embodiment, clear-label samples, fuzzy-label samples, and their augmented samples are propagated forward through the model. For clear-label sample data, after passing through the intermediate spatio-temporal graph convolution unit, its features are passed into the feature refinement module to decouple the spatio-temporal features, extract temporal features and spatial features respectively. Then, the contrast learning mechanism optimizes the decoupled feature expressions, making the cosine distance of the action features of ADHD subjects and normal subjects in the same paradigm as large as possible, so as to more easily distinguish difficult samples. For fuzzy-label samples, after the fifth layer of spatio-temporal graph convolution, the data features and labels of clear-label samples are mixed with the features and labels of fuzzy-label samples through Mixup to generate the features and labels of new sample data, and they are used for subsequent training. Finally, by calculating the cross-entropy loss, contrast loss between the predicted probability and the label of clear-label samples, and the mean square error loss between the predicted label and the pseudo-label of fuzzy-label samples, backpropagation is performed to update the model parameters.

[0047] In this embodiment, by inputting sample data into an action anomaly detection model for spatio-temporal graph convolution, the sample spatio-temporal graph features in the sample data can be effectively extracted. By performing spatio-temporal decoupling on the sample spatio-temporal graph features, the sample decoupled features can be effectively extracted. By performing class prediction on the sample decoupled features, the action prediction result of the sample data can be effectively obtained. Based on the model loss, the parameters of the action anomaly detection model are updated, so that the converged action anomaly detection model can effectively perform action anomaly detection on the data to be detected, without using subjective judgment and questionnaire surveys for action anomaly detection of ADHD, improving the accuracy of action anomaly detection. By combining a deep learning model with high-precision pose data, children with ADHD can be quickly and objectively identified, and the problems of fuzzy labels and fuzzy behavior characteristics can be effectively addressed. The introduced feature refinement module and contrast learning mechanism enhance the classification ability for fuzzy label samples and improve the recognition accuracy of the model. At the same time, for the occlusion problem in pose data, the proposed preprocessing method significantly improves the reliability of the data through coordinate fusion and coordinate normalization. The ADHD screening based on this embodiment can not only be effectively applied in areas with scarce medical resources, but also improve the efficiency and accuracy of ADHD screening, providing strong technical support for medical institutions.

[0048] Embodiment 2 Please refer to Figure 5 , which is a schematic structural diagram of an action anomaly detection system 100 provided by the second embodiment of the present invention, including: A spatio-temporal graph convolution module 10, configured to obtain sample data, input the sample data into an action anomaly detection model for spatio-temporal graph convolution, and obtain sample spatio-temporal graph features.

[0049] Optionally, the spatio-temporal graph convolution module 10 is further configured to: calculate the neighborhood node indices of the feature nodes in the sample data, determine local neighborhood features according to the neighborhood node indices, and calculate edge convolution features according to the local neighborhood features; Perform graph convolution on the sample data to obtain graph convolution features, splice the graph convolution features and the edge convolution features to obtain spliced features; Perform representative spatial average pooling processing on the spliced features to obtain average pooling features, and perform hierarchical edge convolution on the average pooling features to obtain an attention map; Combine the attention map with the graph convolution features to obtain combined features, and sum the combined features according to the hierarchical dimension to obtain the sample spatio-temporal graph features.

[0050] Furthermore, the spatio-temporal graph convolution module 10 is further configured to: obtain fuzzy label samples, and perform sample enhancement on the fuzzy label samples to obtain enhanced samples; Predict labels for the fuzzy label samples and the augmented samples according to the action anomaly detection model to obtain a first predicted label and a second predicted label; Calculate the average value of the first predicted label and the second predicted label to obtain a label average value, and sharpen the label average value to obtain a pseudo label; Label the fuzzy label samples and the augmented samples according to the pseudo label, and obtain clear label samples; Combine the clear label samples, the fuzzy label samples after label marking, and the augmented samples to obtain the sample data.

[0051] The class prediction module 11 is used to perform spatio-temporal decoupling on the sample spatio-temporal graph features to obtain sample decoupled features, and perform class prediction on the sample decoupled features to obtain an action prediction result.

[0052] Optionally, the class prediction module 11 is further used to: perform temporal pooling on the sample spatio-temporal graph features to obtain a first pooled feature, and perform convolution on the first pooled feature to obtain sample spatial features; Perform spatial pooling on the sample spatio-temporal graph features to obtain a second pooled feature, and perform convolution on the second pooled feature to obtain sample temporal features; The sample decoupled features include the sample temporal features and the sample spatial features.

[0053] The model training module 12 is used to determine a model loss according to the action prediction result and the sample decoupled features, and update the parameters of the action anomaly detection model according to the model loss until the action anomaly detection model converges.

[0054] The action anomaly detection module 13 is used to input the data to be detected into the converged action anomaly detection model for action anomaly detection to obtain an action anomaly detection result.

[0055] In this embodiment, by inputting the sample data into the action anomaly detection model for spatio-temporal graph convolution, the sample spatio-temporal graph features in the sample data can be effectively extracted. By performing spatio-temporal decoupling on the sample spatio-temporal graph features, the sample decoupled features can be effectively extracted. By performing class prediction on the sample decoupled features, the action prediction result of the sample data can be effectively obtained. Based on the model loss, the parameters of the action anomaly detection model are updated, so that the converged action anomaly detection model can effectively perform action anomaly detection on the data to be detected, without using subjective judgment and questionnaire survey methods for action anomaly detection of attention deficit hyperactivity disorder, improving the accuracy of action anomaly detection.

[0056] Embodiment III Figure 6It is a structural block diagram of a terminal device 2 provided in the third embodiment of the present application. As Figure 6 shown, the terminal device 2 in this embodiment includes: a processor 20, a memory 21, and a computer program 22 stored in the memory 21 and executable on the processor 20, such as a program for an action anomaly detection method. When the processor 20 executes the computer program 22, the steps in each embodiment of the above various action anomaly detection methods are implemented.

[0057] Exemplarily, the computer program 22 can be divided into one or more modules. The one or more modules are stored in the memory 21 and executed by the processor 20 to complete the present application. The one or more modules can be a series of computer program instruction segments capable of completing specific functions, and the instruction segments are used to describe the execution process of the computer program 22 in the terminal device 2. The terminal device may include, but is not limited to, the processor 20 and the memory 21.

[0058] The so-called processor 20 may be a central processing unit (CPU), or may also be other general-purpose processors, digital signal processors (DSPs), application specific integrated circuits (ASICs), off-the-shelf programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor may be a microprocessor or the processor may also be any conventional processor, etc.

[0059] The memory 21 may be an internal storage unit of the terminal device 2, such as the hard disk or memory of the terminal device 2. The memory 21 may also be an external storage device of the terminal device 2, such as a plug-in hard disk equipped on the terminal device 2, a smart media card (SMC), a secure digital (SD) card, a flash card, etc. Further, the memory 21 may also include both the internal storage unit of the terminal device 2 and the external storage device. The memory 21 is used to store the computer program and other programs and data required by the terminal device. The memory 21 may also be used to temporarily store data that has been output or will be output.

[0060] In addition, in each embodiment of the present application, each functional module can be integrated in a processing unit, can exist separately as individual units physically, or two or more units can be integrated in one unit. The above integrated unit can be implemented in the form of hardware or in the form of a software functional unit.

[0061] If the integrated module is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Among them, the computer-readable storage medium can be non-volatile or volatile. Based on such an understanding, to implement all or part of the processes in the above-mentioned embodiment methods of the present application, it can also be completed by instructing relevant hardware through a computer program. The computer program can be stored in a computer-readable storage medium. When the computer program is executed by a processor, the steps of the above-mentioned various method embodiments can be implemented. Among them, the computer program includes computer program code, and the computer program code can be in the form of source code, object code, executable file or some intermediate form, etc. The computer-readable storage medium can include: any entity or device capable of carrying computer program code, recording medium, USB flash drive, mobile hard disk, magnetic disk, optical disc, computer memory, read-only memory (ROM, Read-Only Memory), random access memory (RAM, Random Access Memory), electrical carrier signal, telecommunication signal, and software distribution medium, etc. It should be noted that the content included in the computer-readable storage medium can be appropriately increased or decreased according to the requirements of legislation and patent practice in the jurisdiction. For example, in some jurisdictions, according to legislation and patent practice, the computer-readable storage medium does not include electrical carrier signals and telecommunication signals.

[0062] The above-described embodiments are only used to illustrate the technical solutions of the present application, rather than to limit them; although the present application has been described in detail with reference to the foregoing embodiments, those of ordinary skill in the art should understand that: they can still modify the technical solutions recorded in the foregoing embodiments, or perform equivalent replacements on some of the technical features; and these modifications or replacements do not make the essence of the corresponding technical solutions deviate from the spirit and scope of the technical solutions of each embodiment of the present application, and should all be included in the protection scope of the present application.

Claims

1. A method for detecting abnormal actions, characterized in that, The method includes: Obtain sample data, and input the sample data into an action anomaly detection model for spatio-temporal graph convolution to obtain sample spatio-temporal graph features; Perform spatio-temporal decoupling on the sample spatio-temporal graph features to obtain sample decoupled features, and perform class prediction on the sample decoupled features to obtain an action prediction result; Determine a model loss according to the action prediction result and the sample decoupled features, and update the parameters of the action anomaly detection model according to the model loss until the action anomaly detection model converges; Input the data to be detected into the converged action anomaly detection model for action anomaly detection to obtain an action anomaly detection result.

2. The abnormal action detection method according to claim 1, wherein Inputting the sample data into an action anomaly detection model for spatio-temporal graph convolution to obtain sample spatio-temporal graph features includes: Calculate the neighborhood node indices of the feature nodes in the sample data, determine local neighborhood features according to the neighborhood node indices, and calculate edge convolution features according to the local neighborhood features; Perform graph convolution on the sample data to obtain graph convolution features, and splice the graph convolution features and the edge convolution features to obtain spliced features; Perform representative spatial average pooling on the spliced features to obtain average pooling features, and perform hierarchical edge convolution on the average pooling features to obtain an attention graph; Combine the attention graph with the graph convolution features to obtain combined features, and sum the combined features according to the hierarchical dimension to obtain the sample spatio-temporal graph features.

3. The abnormal action detection method according to claim 2, wherein The formula for calculating the neighborhood node indices of the feature nodes in the sample data includes: Among them, is the feature node i 's n neighborhood node index, represents the feature of the i th feature node, represents the feature of the j th feature node, represents the proximity algorithm function, n represents the number of proximities set in the proximity algorithm function; The formula for calculating edge convolution features according to the local neighborhood features includes: Among them, represents a linear transformation, represents feature concatenation, represents performing max pooling on the feature in dimension n and represents replicating the feature T times, represents the edge convolution feature, X is the feature of the feature node, X n is the local neighborhood feature; The formula for performing graph convolution on the sample data includes: Among them, represents the k th hierarchical graph convolutional feature, represents the input sample data, represents the pointwise convolutional operation, T represents the size of the time series window, represents the k th hierarchical adjacency matrix, S represents the total edge set of human joint points, s represents the edge subset of human joint points; The formula for performing representative spatial average pooling on the spliced features includes: Among them, represents the RSAP function, N k represents the hierarchy H k the number of joint points in it, v represents the joint point, t represents the timing frame, represents the joint point of the k-th hierarchy v in the timing frame t the eigenvalue on it, represents the average pooling feature, and max represents the maximum value function; The formula for performing hierarchical edge convolution on the average pooling features includes: Among them, M represents the attention map, represents hierarchical edge convolution, σ represents the sigmoid activation function, L represents the hierarchical total set of human joint points.

4. The abnormal action detection method according to claim 1, wherein Performing spatio-temporal decoupling on the sample spatio-temporal graph features to obtain sample decoupled features includes: Perform temporal pooling on the sample spatio-temporal graph features to obtain first pooling features, and perform convolution on the first pooling features to obtain sample spatial features; Perform spatial pooling on the sample spatio-temporal graph features to obtain second pooling features, and perform convolution on the second pooling features to obtain sample temporal features; The sample decoupled features include the sample temporal features and the sample spatial features.

5. The abnormal action detection method according to claim 4, wherein Obtaining sample data includes: Obtain fuzzy label samples, and perform sample enhancement on the fuzzy label samples to obtain enhanced samples; Perform label prediction on the fuzzy label samples and the enhanced samples according to the action anomaly detection model to obtain a first predicted label and a second predicted label; Calculate the average value of the first predicted label and the second predicted label to obtain a label average value, and sharpen the label average value to obtain pseudo-labels; Perform label marking on the fuzzy label samples and the enhanced samples according to the pseudo-labels, and obtain clear label samples; Combine the clear label samples, the fuzzy label samples and the enhanced samples after label marking to obtain the sample data.

6. The abnormal action detection method according to claim 5, wherein The formula for determining the model loss based on the action prediction result and the sample decoupled feature includes: Among them, represents the model loss, W CL represents the first weight hyperparameter, W u represents the second weight hyperparameter, represents the cross-entropy loss, represents the mean squared error loss, represents the contrastive loss; Among them, represents the number of samples of the fuzzy label sample, u represents the fuzzy label sample, q represents the pseudo label, represents the class probability predicted by the fuzzy label sample in the action prediction result, represents the data of the fuzzy label sample, represents the model parameters of the action anomaly detection model; Among them, represents the number of samples of the clear label samples, x represents the clear label samples, p represents the true label of the clear label samples, represents the class probability predicted by the clear label samples in the action prediction result, L ′ represents the data of the clear label samples; Among them, represents the loss weight, represents the sample space feature, represents the sample time feature, represents the contrastive learning loss function, n represents the number of times of contrastive learning construction.

7. The abnormal action detection method according to claim 6, characterized in that The contrastive learning loss function is: Among them, represents the prediction probability of the sample i for the category k in the action prediction result, τ represents the sensitivity hyperparameter, represents the compensation term, represents the penalty term, represents the prototype representation of the category k , represents the prototype representation of the category , F i represents the sample time feature or the sample space feature of the sample i ; Among them, represents the central feature representation of false positive samples of k the class, represents the confidence sample set, represents the data volume of the confidence sample set, represents the class k of the data volume of the set of false positive samples; Among them, represents the central feature representation of false negative samples of k category represents the data volume of the set of false negative samples of k category Among them, F j represents the sample j of the sample time feature or the sample space feature; Among them, α represents the momentum term.

8. An abnormal action detection system, characterized in that, The system includes: A spatio-temporal graph convolution module, configured to obtain sample data, input the sample data into an action anomaly detection model for spatio-temporal graph convolution, and obtain sample spatio-temporal graph features; A class prediction module, configured to perform spatio-temporal decoupling on the sample spatio-temporal graph features to obtain sample decoupled features, and perform class prediction on the sample decoupled features to obtain an action prediction result; A model training module, configured to determine a model loss according to the action prediction result and the sample decoupled feature, and update parameters of the action anomaly detection model according to the model loss until the action anomaly detection model converges; An action anomaly detection module, configured to input data to be detected into the converged action anomaly detection model for action anomaly detection to obtain an action anomaly detection result.

9. A terminal device, comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, the steps of the method according to any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, the steps of the method according to any one of claims 1 to 7 are implemented.

Citation Information

Patent Citations

  • RNA-Seq sequencing data analysis method

    CN115295083A

  • Distributed traffic flow prediction method and system based on space-time decoupling

    CN115966083A

  • Abnormal behavior recognition method based on space-time diagram convolutional neural network

    CN116959099A

  • Fall detection model training improvement method and fall detection method

    CN117671794A

  • Skeleton human behavior recognition method based on multi-view space-time contrast loss

    CN119049128A