Classification method and system based on skeleton information and action reaction feature data
By combining deep learning models that incorporate skeletal features of limb movements and reaction time features, the problem of low accuracy in depression identification in existing technologies has been solved, achieving more efficient depression identification and assisted diagnosis and treatment.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- QILU UNIVERSITY OF TECHNOLOGY (SHANDONG ACADEMY OF SCIENCES)
- Filing Date
- 2022-10-11
- Publication Date
- 2026-04-28
AI Technical Summary
Existing methods for identifying depression lack objective evaluation criteria. Commonly used models such as LSTM and ResNet18 have low accuracy in recognizing skeletal information and motor response features, leading to misdiagnosis or missed diagnosis.
By combining limb movement skeleton features and movement reaction time features, the collected data is preprocessed using a deep learning model to extract movement reaction time features, which are then input into a recognition model based on movement skeleton sequences for classification.
It improves the accuracy of depression identification, better assists in early identification and treatment, and reflects the clinical manifestations of depression patients, such as reduced activity and slowed thinking and cognitive function.
Smart Images

Figure CN115409072B_ABST
Abstract
Description
Technical Field
[0001] This disclosure relates to the field of data recognition and classification technology, specifically to a classification method and system based on skeleton information and action response feature data. Background Technology
[0002] The statements in this section are merely background information relating to this disclosure and do not necessarily constitute prior art.
[0003] Depression is a group of affective mental disorders characterized by significant and persistent low mood, reduced activity levels, and slowed thinking and cognitive function. People with depression are a group who experience severe psychological distress and negative emotions, generally exhibiting symptoms such as lethargy and sleep deprivation. Those with severe depression may even engage in self-harm or suicidal behavior.
[0004] Compared to the subjective assessment of symptoms in clinical practice, which lacks objective and effective evaluation indicators and is therefore prone to misdiagnosis or missed diagnosis, depression patients often experience reduced activity levels and slowed thinking and cognitive functions. Current methods for identifying depression lack the ability to judge and analyze based on skeletal information of limb movements and characteristics of movement reaction time. Most models, such as LSTM, ResNet18, and Transformer, which only use skeletal sequence data as input, have low recognition accuracy and cannot classify and identify based on skeletal information and characteristic data of movement reaction time. Summary of the Invention
[0005] To address the aforementioned issues, this disclosure proposes a classification method and system based on skeletal information and motor response feature data. By combining limb movement skeletal features and motor response time features, the acquired data is identified, judged, and classified using a deep learning method, providing a convenient and reliable screening tool for the early identification of depression and assisting in clinical diagnosis and treatment.
[0006] According to some embodiments, the present disclosure adopts the following technical solutions:
[0007] Classification methods based on skeleton information and action response feature data include:
[0008] Raw limb data and instruction audio data of the person to be classified and identified under the stimulus task are collected and preprocessed.
[0009] Extract the overall skeleton sequence data from the raw limb data and the audio sequence data from the command audio;
[0010] Align the segments in the skeleton sequence data with the time of the action command issued in the audio sequence data, locate the time period of the action start and the time point of the action command issued in the audio data based on the change of the joint coordinate position in the skeleton sequence segment, and obtain the action reaction time characteristics.
[0011] The action reaction time features and action skeleton sequence data are input into the action skeleton sequence-based recognition model, and the classification results are output.
[0012] According to some embodiments, the present disclosure adopts the following technical solutions:
[0013] Classification systems based on skeleton information and action response feature data include:
[0014] The data acquisition module is used to collect raw limb data and instruction audio data of the person to be classified and identified under the stimulus task, and to perform preprocessing.
[0015] The data extraction module is used to extract the overall skeleton sequence data from the raw limb data and the audio sequence data from the command audio.
[0016] The feature acquisition module is used to align the segments in the skeleton sequence data with the time of the action command issued in the audio sequence data, locate the time period of the start of the action and the time point of the action command issued in the audio data based on the change of the coordinate position of the joints in the skeleton sequence segment, and acquire the action reaction time features.
[0017] The classification module is used to input action reaction time features and action skeleton sequence data into the recognition model based on action skeleton sequence and output classification results.
[0018] Compared with the prior art, the beneficial effects of this disclosure are as follows:
[0019] This disclosure, through preprocessing and analysis of raw data, reveals a difference in reaction time between depressed patients and non-depressed individuals upon hearing instructions. This reflects clinical manifestations in depressed patients, such as reduced activity levels and slowed thinking and cognitive functions.
[0020] Compared to models like LSTM, ResNet18, and Transformer, which only use skeleton sequence data as input, this disclosure incorporates preprocessed reaction time features into the deep learning model through the steps described above. This improves the accuracy of deep learning models in identifying and classifying depression. Adding skeletal information of human limb movements and reaction time features to deep learning models can better assist in the identification and detection of depression. Attached Figure Description
[0021] The accompanying drawings, which form part of this disclosure, are used to provide a further understanding of this disclosure. The illustrative embodiments of this disclosure and their descriptions are used to explain this disclosure and do not constitute an undue limitation of this disclosure.
[0022] Figure 1 This is a schematic diagram illustrating the change of audio intensity over time in this disclosure;
[0023] Figure 2 This is a schematic diagram of the 3D coordinate changes of the hand joints disclosed herein;
[0024] Figure 3 This is a schematic diagram of the model structure disclosed herein;
[0025] Figure 4 This is a schematic diagram of the data identification and classification process disclosed herein. Detailed Implementation
[0026] The present disclosure will be further described below with reference to the accompanying drawings and embodiments.
[0027] It should be noted that the following detailed descriptions are illustrative and intended to provide further explanation of this disclosure. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this disclosure pertains.
[0028] It should be noted that the terminology used herein is for the purpose of describing particular embodiments only and is not intended to limit the exemplary embodiments according to this disclosure. As used herein, the singular form is intended to include the plural form as well, unless the context clearly indicates otherwise. Furthermore, it should be understood that when the terms “comprising” and / or “including” are used in this specification, they indicate the presence of features, steps, operations, devices, components, and / or combinations thereof.
[0029] Example 1
[0030] One embodiment of this disclosure provides a classification method based on skeleton information and action response feature data, including:
[0031] Step 1: Collect raw limb data and instruction audio data of the person to be classified and identified under the stimulus task, and perform preprocessing;
[0032] Step 2: Extract the overall skeleton sequence data from the original limb data and the audio sequence data from the command audio;
[0033] Step 3: Align the segments in the skeleton sequence data with the action command issued in the audio sequence data, locate the time period when the action starts and the time point when the action command is issued in the audio data based on the change in the coordinate position of the joints in the skeleton sequence segment, and obtain the action reaction time characteristics.
[0034] Step 4: Input the action reaction time features and action skeleton sequence data into the recognition model based on the action skeleton sequence, and output the classification results.
[0035] As one example, in step 1, the data sources consisted of two groups: one group comprised 77 patients with depression, and the control group comprised 104 healthy volunteers. The two groups were screened and matched according to information such as age, gender, education level, occupation, and marital status, and the 24-item Hamilton Depression Rating Scale (HAMD) assessed by psychiatrists was strictly used as an important reference standard for data processing.
[0036] During data acquisition, the subjects were given a limb movement stimulation task. The subject's body movements should be as pronounced as possible so that the Kinect device could capture a sufficiently large range of motion. The stimulation task consisted of five kinematic segments: both arms raised and lowered, left arm raised and lowered, right arm raised and lowered, turning left and lowered, and turning right and lowered. After the subject reached a designated position within the Kinect's recording range, an audio command for the movement was played. The stimulation task was controlled by this audio command, and the process of the subject hearing the command and responding was recorded by Kinect and saved to a .xef file.
[0037] Furthermore, in step 2, the overall skeleton sequence data and the audio sequence data from the command audio are extracted from the original limb data. The entire skeleton sequence data and the audio sequence data from the recording site are extracted from the original data recorded by the Kinect V2 device. The joint coordinates in the skeleton sequence segment are composed of seven dimensions representing spatial positional relationships. That is, each human skeleton joint is composed of (x, y, z) and ((Rx, Ry, Rz), Rw). In other words, each human skeleton joint n at kinematic segment time t is composed of seven dimensions representing spatial positional relationships. A quaternion is a form used to represent the rotation of an object in three-dimensional space. A 3D vector (Rx, Ry, Rz) represents the axis of rotation and an angular component, and Rw represents the rotation angle around this axis. It represents the rotation information generated by each skeleton joint at a certain moment.
[0038] Each human skeleton joint is detected at its spatial position (x, y, z), where x, y, and z represent the joint's position coordinates along the x, y, and z axes in a spatial coordinate system with the Kinect acquisition device as the origin. For the target object i to be locked at any time t, the following relationship holds:
[0039]
[0040] For the denoised data, extract its quaternion ((Rx,Ry,Rz),Rw) in spatial coordinate system, as shown in the following formula:
[0041]
[0042] The motion command issued from the audio sequence data is aligned with the segment in the skeleton sequence data. Each frame of the skeleton sequence segment includes time, 3D coordinates of joints, and quaternion information. The process of locating the time period of motion start and the time point of motion command issuance in the audio data based on the changes in the joint coordinate positions in the skeleton sequence segment is as follows: the motion start time point T1 is obtained, and the time point of motion command issuance in the audio information is T2. Then, the motion reaction time T3 = T1 - T2.
[0043] The specific method is as follows: based on the changes in the coordinate positions of joints in the skeleton sequence data (e.g., the changes in the x-coordinates of the experimental subject's two fingertips over time, see...) Figure 2 The time point T1 when the experimental subject begins to perform the action is located, and the time point T2 when the action command is issued in the audio information is located. Then the reaction time T3 = T1 - T2, thereby obtaining the reaction time characteristics of each experimental subject in the five action tasks. Figure 2 Specifically, when the experimental subject remained in its original position without making any movement, the coordinate values of the fingertips remained near fixed values (0.2 and -0.2 in the figure). However, when a movement was made, the coordinate values of the fingertips changed significantly from these initial values. The time point at which the movement began was determined by the difference between the coordinate values at each moment and the initial coordinate values.
[0044] The average reaction time data of the two groups were calculated. The data showed that the patients in the depression group had an average reaction time of about 1 second longer than the normal people in the control group.
[0045] The model construction disclosed herein is as follows:
[0046] Because of its outstanding ability to capture long-term dependencies and interactions, Transformer is particularly attractive for time series modeling and has made exciting progress in a variety of time series applications.
[0047] The 181 sets of samples were divided into training set, validation set and test set in a ratio of 6:2:2.
[0048] The aligned skeleton data is then fed into the Encoder through a positional encoding layer. A multi-head attention mechanism uses multiple attention methods for computation to obtain more layers of information. Next is an Add & Norm residual connection and layer normalization block. This addresses network degradation and laterally normalizes the sequence data. Following this is a feedforward block, a two-layer fully connected neural network. The intermediate layer is ReLU, which not only helps the Attention output extract more abstract information but also filters out invalid information, retaining the more important parts.
[0049] After passing through the Encoder module, the output features are reduced in dimensionality using a linear layer. Then, the reduced skeleton data features are concatenated with the reaction time features from before the classifier, resulting in the final output.
[0050] Example 2
[0051] One embodiment of this disclosure provides a classification system based on skeleton information and action response feature data, including:
[0052] The data acquisition module is used to collect raw limb data and instruction audio data of the person to be classified and identified under the stimulus task, and to perform preprocessing.
[0053] The data extraction module is used to extract the overall skeleton sequence data from the raw limb data and the audio sequence data from the command audio.
[0054] The feature acquisition module is used to align the segments in the skeleton sequence data with the time of the action command issued in the audio sequence data, locate the time period of the start of the action and the time point of the action command issued in the audio data based on the change of the coordinate position of the joints in the skeleton sequence segment, and acquire the action reaction time features.
[0055] The classification module is used to input action reaction time features and action skeleton sequence data into the recognition model based on action skeleton sequence and output classification results.
[0056] Furthermore, the Kinect device was used to collect raw limb data and instruction audio data during the stimulation task.
[0057] The data came from two groups: one group consisted of 77 patients with depression, and the control group consisted of 104 healthy volunteers. The participants were screened and matched according to age, gender, education level, occupation, marital status, etc., and the 24-item Hamilton Depression Rating Scale (HAMD) administered by psychiatrists was strictly used as a key reference standard for data processing.
[0058] During data acquisition, the subjects were given a limb movement stimulation task. The subject's body movements should be as pronounced as possible so that the Kinect device could capture a sufficiently large range of motion. The stimulation task consisted of five kinematic segments: both arms raised and lowered, left arm raised and lowered, right arm raised and lowered, turning left and lowered, and turning right and lowered. After the subject reached a designated position within the Kinect's recording range, an audio command for the movement was played. The stimulation task was controlled by this audio command, and the process of the subject hearing the command and responding was recorded by Kinect and saved to a .xef file.
[0059] Furthermore, in step 2, the overall skeleton sequence data and the audio sequence data in the command audio are extracted from the original limb data. The entire skeleton sequence data and the audio sequence data of the recording site are extracted from the original data recorded by the Kinect V2 device. The joint coordinates in the skeleton sequence segment are composed of seven dimensions of data that describe the spatial position relationship. That is, each human skeleton joint is composed of (x,y,z) and ((Rx,Ry,Rz),Rw). That is, each human skeleton joint n at kinematic segment time t is composed of seven dimensions of data that describe the spatial position relationship.
[0060] The spatial position (x, y, z) of each human skeleton joint is detected. For the target object i to be locked, the following relationship holds at any time t:
[0061]
[0062] For the denoised data, extract its quaternion ((Rx,Ry,Rz),Rw) in spatial coordinate system, as shown in the following formula:
[0063]
[0064] The motion command issued from the audio sequence data is aligned with the segment in the skeleton sequence data. Each frame of the skeleton sequence segment includes time, 3D coordinates of joints, and quaternion information. The process of locating the time period of motion start and the time point of motion command issuance in the audio data based on the changes in the joint coordinate positions in the skeleton sequence segment is as follows: the motion start time point T1 is obtained, and the time point of motion command issuance in the audio information is T2. Then, the motion reaction time T3 = T1 - T2.
[0065] The specific method is as follows: based on the changes in the coordinate positions of joints in the skeleton sequence data (e.g., the changes in the x-coordinates of the experimental subject's two fingertips over time, see...) Figure 2The time point T1 when the experimental subject begins to perform the action is located, and the time point T2 when the action command is issued in the audio information is located. Then the reaction time T3 = T1 - T2, thereby obtaining the reaction time characteristics of each experimental subject in the five action tasks.
[0066] The average reaction time data of the two groups were calculated. The data showed that the patients in the depression group had an average reaction time of about 1 second longer than the normal people in the control group.
[0067] The model construction disclosed herein is as follows:
[0068] Because of its outstanding ability to capture long-term dependencies and interactions, Transformer is particularly attractive for time series modeling and has made exciting progress in a variety of time series applications.
[0069] The 181 sets of samples were divided into training set, validation set and test set in a ratio of 6:2:2.
[0070] The aligned skeleton data is then fed into the Encoder through a positional encoding layer. A multi-head attention mechanism uses multiple attention methods for computation to obtain more layers of information. Next is an Add & Norm residual connection and layer normalization block. This addresses network degradation and laterally normalizes the sequence data. Following this is a feedforward block, a two-layer fully connected neural network. The intermediate layer is ReLU, which not only helps the Attention output extract more abstract information but also filters out invalid information, retaining the more important parts.
[0071] After passing through the Encoder module, the output features are reduced in dimensionality using a linear layer. Then, the reduced skeleton data features are concatenated with the reaction time features from before the classifier, resulting in the final output.
[0072] Example 3
[0073] One embodiment of this disclosure provides a computer-readable storage medium, characterized in that it stores a plurality of instructions adapted for loading and execution by a processor of a terminal device of the classification method steps based on skeleton information and motion response feature data.
[0074] Example 4
[0075] One embodiment of this disclosure provides a terminal device, including a processor and a computer-readable storage medium. The processor is used to implement various instructions; the computer-readable storage medium is used to store multiple instructions, which are adapted to be loaded by the processor and executed by the classification method steps based on skeleton information and motion response feature data.
[0076] This disclosure is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this disclosure. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create a machine for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0077] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0078] While the specific embodiments of this disclosure have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of this disclosure. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of this disclosure are still within the scope of protection of this disclosure.
Claims
1. A classification method based on skeleton information and action reaction feature data, characterized in that, The method comprises the following steps: Collecting original body data and instruction audio data of a person to be classified and identified under a stimulation task, and performing preprocessing; Extracting overall skeleton sequence data in the original body data and audio sequence data in the instruction audio; Aligning a segment in the skeleton sequence data according to a time of issuing an action instruction in the audio sequence data, locating a time period of action start and a time point of issuing the action instruction in the audio data according to a change of a coordinate position of a joint node in the skeleton sequence segment, and obtaining an action reaction time feature; Inputting the action reaction time feature and the action skeleton sequence data into an identification model based on the action skeleton sequence, and outputting a classification result.
2. The classification method based on skeleton information and motion reaction feature data according to claim 1, characterized in that, The stimulation task is five kinematic segments, including double-hand lifting and resetting, left-arm lifting and resetting, right-arm lifting and resetting, left-turning and resetting, and right-turning and resetting.
3. The classification method based on skeleton information and motion reaction feature data according to claim 1, characterized in that, The stimulation task is controlled by the action instruction audio, and a reaction action process performed under the instruction is recorded by Kinect and saved into a file in.xef format.
4. The classification method based on skeleton information and motion reaction feature data according to claim 1, wherein, The joint node coordinates in the skeleton sequence segment are composed of seven-dimensional data representing spatial position relationships.
5. The classification method based on skeleton information and motion response feature data as described in claim 1, characterized in that, The skeleton sequence segment includes time, joint node 3D coordinates, and quaternion information.
6. The classification method based on skeleton information and motion reaction feature data according to claim 1, wherein, The process of locating the time period of action start and the time point of issuing the action instruction in the audio data according to the change of the coordinate position of the joint node in the skeleton sequence segment, and obtaining the action reaction time feature, is as follows: obtaining a time point T1 of action start, a time point T2 of issuing the action instruction in the audio information, and then the action reaction time T3 = T1-T2.
7. A classification system based on skeleton information and motion reaction feature data, characterized in that, The method comprises the following steps: A data acquisition module is configured to collect original body data and instruction audio data of a person to be classified and identified under a stimulation task, and perform preprocessing; A data extraction module is configured to extract overall skeleton sequence data in the original body data and audio sequence data in the instruction audio; A feature acquisition module is configured to align a segment in the skeleton sequence data according to a time of issuing an action instruction in the audio sequence data, locate a time period of action start and a time point of issuing the action instruction in the audio data according to a change of a coordinate position of a joint node in the skeleton sequence segment, and obtain an action reaction time feature; A classification module is configured to input the action reaction time feature and the action skeleton sequence data into an identification model based on the action skeleton sequence, and output a classification result.
8. The skeleton information and motion reaction feature data-based classification system of claim 7, wherein, The original body data and instruction audio data under the stimulation task are collected by using a Kinect device.
9. A computer-readable storage medium, characterized in that, A terminal device has a processor and a computer readable storage medium, and the processor is configured to implement instructions; the computer readable storage medium is configured to store a plurality of instructions, and the instructions are adapted to be loaded and executed by the processor to implement the classification method based on skeleton information and action reaction feature data according to any one of claims 1-6.
10. A terminal device, comprising: A terminal device has a processor and a computer readable storage medium, and the processor is configured to implement instructions; the computer readable storage medium is configured to store a plurality of instructions, and the instructions are adapted to be loaded and executed by the processor to implement the classification method based on skeleton information and action reaction feature data according to any one of claims 1-6.
Citation Information
Patent Citations
Depression recognition method and system based on human skeleton kinematics characteristic information
CN111938670A
Health scale forming method and device, equipment and storage medium
CN113112991A