Gait emotion recognition method, system and device and medium

Through the combination of the gait skeleton framework and the gait emotion framework, the cascading spatio-temporal feature extraction network and emotional denoising network are used to extract and integrate gait features, and the problems of large computing resources and low recognition accuracy in the prior art are solved, and efficient gait emotion recognition is achieved.

CN120148102APending Publication Date: 2025-06-13SOUTH CHINA UNIV OF TECH
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510136212.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-07
Publication Date
2025-06-13

AI Technical Summary

Technical Problem

The network model constructed by the existing gait emotion recognition method is complex and has a large number of parameters, which leads to a large consumption of computing resources. When there are fewer gait data samples, the recognition accuracy is not high.

Method used

The gait skeleton framework and gait emotion framework are adopted to extract networks and emotional denoising networks through cascading spatiotemporal features, extract skeleton characteristics and emotional characteristics, and perform decision-level fusion to reduce the consumption of computing resources and improve recognition accuracy.

Benefits of technology

It effectively reduces the computing resources required for gait emotion recognition and improves the accuracy of gait emotion recognition, especially when there are fewer gait data samples.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148102A_ABST
    Figure CN120148102A_ABST
Patent Text Reader

Abstract

The invention discloses a gait emotion recognition method, system and device and a medium, and the method comprises the steps: obtaining gait sequence data and gait emotion information of the gait sequence data; inputting the gait sequence data into a gait skeleton framework for skeleton feature extraction to obtain enhanced skeleton features and skeleton feature scores of the enhanced skeleton features; inputting the gait emotion information into a gait emotion framework for emotion feature extraction to obtain de-noised emotion features and emotion feature scores of the de-noised emotion features; and according to the enhanced skeleton features and the de-noised emotion features, performing decision-level fusion on the skeleton feature score and the emotion feature score to obtain an emotion recognition result of the gait sequence data. According to the method, computing resources required by gait emotion recognition can be reduced, and the accuracy of gait emotion recognition is effectively improved. The invention relates to the technical field of artificial intelligence.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of artificial intelligence, and in particular to a gait emotion recognition method, system, device and medium. Background Art

[0002] Since gait emotion recognition can realize the recognition of the individual emotion state of users in a non-invasive manner, it has gradually become one of the key contents that people focus on.

[0003] Currently, traditional gait emotion recognition methods usually use spatio-temporal graph convolutional neural networks to jointly capture spatio-temporal features in gait information, and then realize the recognition of gait emotions based on the extracted spatio-temporal features. However, since the network model constructed in this way is relatively complex and has a large number of parameters, it requires a large amount of computing resources to realize gait emotion recognition; moreover, in the case of few gait data samples, the accuracy of gait emotion recognition is not high.

[0004] Therefore, the problems existing in the prior art still need to be solved and optimized urgently. Summary of the Invention

[0005] An object of the present invention is to solve at least to some extent one of the technical problems existing in the related art.

[0006] To this end, an object of an embodiment of the present invention is to provide a gait emotion recognition method, system, device and medium, wherein the method can reduce the computing resources required for gait emotion recognition and effectively improve the accuracy of gait emotion recognition.

[0007] In order to achieve the above technical object, the technical solutions adopted in the embodiments of the present application include:

[0008] In a first aspect, an embodiment of the present application provides a gait emotion recognition method, including:

[0009] Obtain gait sequence data and the gait emotion information of the gait sequence data;

[0010] Input the gait sequence data into a gait skeleton framework for skeleton feature extraction to obtain enhanced skeleton features output by the gait skeleton framework and the skeleton feature scores of the enhanced skeleton features;

[0011] Input the gait emotion information into a gait emotion framework for emotion feature extraction to obtain denoised emotion features output by the gait emotion framework and the emotion feature scores of the denoised emotion features;

[0012] According to the enhanced skeleton features and the denoised emotion features, perform decision-level fusion on the skeleton feature scores and the emotion feature scores to obtain the emotion recognition result of the gait sequence data.

[0013] In addition, according to the method of the above embodiments of the present application, the following additional technical features may also be included:

[0014] Further, in an embodiment of the present application, the gait skeleton framework includes a plurality of cascaded spatio-temporal feature extraction networks, and the spatio-temporal feature extraction network is used to perform the following steps:

[0015] Obtain the skeleton input feature;

[0016] Extract the gait spatial features from the skeleton input feature to obtain the intermediate spatial feature;

[0017] Extract the skeleton temporal attention from the intermediate spatial feature to obtain the skeleton output feature;

[0018] Wherein, the skeleton input feature is the original gait sequence feature of the gait sequence data or the skeleton output feature output by the previous spatio-temporal feature extraction network.

[0019] Further, in an embodiment of the present application, the step of extracting the skeleton temporal attention from the intermediate spatial feature to obtain the skeleton output feature includes:

[0020] Extract the shallow temporal features from the intermediate spatial feature to obtain a shallow spatio-temporal feature map;

[0021] Extract the deep temporal features from the shallow spatio-temporal feature map to obtain a deep spatio-temporal feature map;

[0022] Capture the attention features from the deep spatio-temporal feature map to obtain the skeleton output feature.

[0023] Further, in an embodiment of the present application, the spatio-temporal feature extraction network includes a plurality of parallel temporal depthwise separable convolution modules, and the dilation rate of each temporal depthwise separable convolution module is different. The step of extracting the deep temporal features from the shallow spatio-temporal feature map to obtain the deep spatio-temporal feature map includes:

[0024] Input the shallow spatio-temporal feature map into all the temporal depthwise separable convolution modules for multi-scale temporal feature extraction to obtain a plurality of deep dilated feature maps, and each deep dilated feature map corresponds to one temporal depthwise separable convolution module;

[0025] Fuse the deep dilated feature maps to obtain the deep spatio-temporal feature map.

[0026] Further, in an embodiment of the present application, the step of capturing the attention features from the deep spatio-temporal feature map to obtain the skeleton output feature includes:

[0027] Calculate the attention weights for the deep spatio-temporal feature map to obtain spatio-temporal attention weights corresponding to the deep spatio-temporal feature map;

[0028] Fuse the attention weights to the deep spatio-temporal feature map according to the spatio-temporal attention weights to obtain the skeleton output feature.

[0029] Further, in an embodiment of the present application, the gait emotion framework includes a plurality of cascaded emotion denoising networks, and the emotion denoising network is used to perform the following steps:

[0030] Obtain emotion input features and a preset zero-tensor threshold;

[0031] Extract eigenvalue from the emotion input features to obtain the absolute value of the feature and the global average value of the feature;

[0032] Scale the emotion input features according to the global average value of the feature to obtain emotion-scaled features;

[0033] Perform soft-threshold feature selection on the emotion-scaled features according to the zero-tensor threshold and the absolute value of the feature to obtain emotion soft-threshold features;

[0034] Denoise the emotion input features according to the emotion soft-threshold features to obtain emotion output features;

[0035] Wherein, the emotion input features are the gait emotion features of the gait emotion information or the emotion output features output by the previous emotion denoising network.

[0036] Further, in an embodiment of the present application, the step of performing soft-threshold feature selection on the emotion-scaled features according to the zero-tensor threshold and the absolute value of the feature to obtain emotion soft-threshold features includes:

[0037] Calculate the feature difference of the emotion-scaled features according to the absolute value of the feature to obtain emotion difference features;

[0038] Perform soft-threshold feature selection on the emotion difference features according to the zero-tensor threshold to obtain the emotion soft-threshold features.

[0039] In a second aspect, an embodiment of the present application provides a gait emotion recognition system, including:

[0040] A first processing unit, configured to obtain gait sequence data and the gait emotion information of the gait sequence data;

[0041] A second processing unit, configured to input the gait sequence data into a gait skeleton framework for skeleton feature extraction, to obtain enhanced skeleton features output by the gait skeleton framework, and skeleton feature scores of the enhanced skeleton features;

[0042] A third processing unit, configured to input the gait emotion information into a gait emotion framework for emotion feature extraction, to obtain denoised emotion features output by the gait emotion framework, and emotion feature scores of the denoised emotion features;

[0043] A fourth processing unit, configured to perform decision-level fusion on the skeleton feature scores and the emotion feature scores according to the enhanced skeleton features and the denoised emotion features, to obtain an emotion recognition result of the gait sequence data.

[0044] In a third aspect, an embodiment of the present application further provides an electronic device, including:

[0045] At least one processor;

[0046] At least one memory, configured to store at least one program;

[0047] When the at least one program is executed by the at least one processor, the at least one processor implements the above method.

[0048] In a fourth aspect, an embodiment of the present application further provides a computer-readable storage medium, in which a program executable by a processor is stored, and the program executable by the processor is used to implement the above method when executed by the processor.

[0049] The advantages and beneficial effects of the present application will be partially given in the following description, partially become obvious from the following description, or be understood through the practice of the present application:

[0050] A gait emotion recognition method, system, device and medium disclosed in an embodiment of the present application. The method obtains gait sequence data and gait emotion information of the gait sequence data; inputs the gait sequence data into a gait skeleton framework for skeleton feature extraction to obtain enhanced skeleton features output by the gait skeleton framework and skeleton feature scores of the enhanced skeleton features; inputs the gait emotion information into a gait emotion framework for emotion feature extraction to obtain denoised emotion features output by the gait emotion framework and emotion feature scores of the denoised emotion features; and performs decision-level fusion on the skeleton feature scores and the emotion feature scores according to the enhanced skeleton features and the denoised emotion features to obtain an emotion recognition result of the gait sequence data. By using the gait skeleton framework to extract skeleton features from gait sequence data and using the gait emotion framework to extract emotion features from gait emotion information, and performing decision-level fusion on the feature scores of the skeleton and emotion based on the obtained enhanced skeleton features and denoised emotion features, the method can reduce the computing resources required for gait emotion recognition while effectively improving the accuracy of gait emotion recognition when the number of gait data samples is small. BRIEF DESCRIPTION OF THE DRAWINGS

[0051] In order to more clearly illustrate the technical solutions in the embodiments of the present application or the prior art, the following introduces the accompanying drawings related to the technical solutions in the embodiments of the present application or the prior art. It should be understood that the accompanying drawings in the following introduction are only for conveniently and clearly expressing some embodiments of the technical solutions in the present application. For those skilled in the art, other drawings can be obtained based on these drawings without creative efforts.

[0052] Figure 1 It is a schematic flowchart of a gait emotion recognition method provided by an embodiment of the present application;

[0053] Figure 2 It is a schematic structural framework diagram of a gait emotion recognition system provided by an embodiment of the present application;

[0054] Figure 3 It is a schematic structural diagram of an electronic device provided by an embodiment of the present application. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0055] Embodiments of the present application will be described in detail below. Examples of the embodiments are shown in the accompanying drawings, where like or similar reference numerals denote like or similar elements or elements having the same or similar functions throughout. The embodiments described below by referring to the accompanying drawings are exemplary and are only used to explain the present application and should not be construed as a limitation of the present application. For the step numbers in the following embodiments, they are only set for the convenience of description and explanation, and no limitation is imposed on the order between the steps. The execution order of each step in the embodiments can be adaptively adjusted according to the understanding of those skilled in the art.

[0056] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by those of ordinary skill in the technical field to which this application belongs. The terms used herein are only for the purpose of describing the embodiments of this application and are not intended to limit this application.

[0057] Currently, traditional gait emotion recognition methods usually use spatio-temporal graph convolutional neural networks to jointly capture spatio-temporal features in gait information, and then based on the extracted spatio-temporal features, realize the recognition of gait emotions. However, since the network model constructed in this way is relatively complex and has a large number of parameters, it requires a lot of computing resources to achieve gait emotion recognition; and, although spatio-temporal graph convolutional neural networks can capture spatio-temporal features of gait data to a certain extent, there are still deficiencies in the feature extraction in the time dimension of gait data in this way. It only extracts shallow time features and does not pay attention to deep time features and more important spatio-temporal implicit dependency relationships.

[0058] In addition, there are a small number of existing technologies that use emotional supplementary information of gait data to assist in gait emotion recognition. However, this method only considers information directly related to gait emotion recognition (such as speed, acceleration), etc., and does not consider information not directly related to gait emotion recognition (such as joint distance, angle), and this information not directly related may contain a lot of noise. Moreover, when existing technologies use emotional supplementary information to assist in gait emotion recognition, they mainly splice emotional supplementary features and gait features at the feature level, and then perform emotion recognition based on the features spliced at the feature level. This method does not consider the difference problem of emotional supplementary features and gait features in different feature spaces, and the accuracy of gait emotion recognition is not satisfactory.

[0059] In view of this, embodiments of the present invention provide a gait emotion recognition method, system, device and medium. Among them, this method extracts skeleton features from gait sequence data through a gait skeleton framework. Specifically, by using multiple cascaded spatio-temporal feature extraction networks, each spatio-temporal feature extraction network sequentially performs spatial feature extraction, shallow and deep spatio-temporal feature extraction, and attention feature capture on the skeleton input features. When performing deep spatio-temporal feature extraction, temporal depthwise dilated convolution is used, which can extract deep temporal features in a lightweight manner and adaptively capture spatio-temporal dependencies and important spatio-temporal features based on the attention mechanism, so as to fully consider the complex spatio-temporal dependencies and important spatio-temporal features in each time dimension, which is beneficial to improving the accuracy of gait emotion recognition.

[0060] In addition, this method takes gait emotion information as the input of the gait emotion framework and extracts emotion features through the gait emotion framework. Specifically, through cascaded emotion denoising modules, each emotion denoising module uses emotion soft threshold features to denoise the emotion input features. It can adaptively remove the noise in the non-directly associated information of the emotion input features based on the learnable soft threshold mechanism, and at the same time retain the emotion discrimination characteristics in the non-directly associated information. Therefore, while fully considering the directly associated information of gait emotion recognition, it also fully considers the information non-directly associated with gait emotion recognition, which is beneficial to improving the emotion supplement role played by emotion supplementary information in assisting gait emotion recognition.

[0061] Moreover, this method performs decision-level fusion on the skeleton feature score and the emotion feature score based on the enhanced skeleton features and the denoised emotion features. It can efficiently play the supplementary role of gait emotion information, fully consider the differences between the enhanced skeleton features and the denoised emotion features in different feature spaces, which is beneficial to improving the accuracy of gait emotion recognition.

[0062] A gait emotion recognition method, system, device and medium provided by an embodiment of the present application can be specifically described through the following embodiments. First, a gait emotion recognition method in an embodiment of the present application is described.

[0063] The gait emotion recognition method provided by an embodiment of the present application can be applied to an emotion detection application scenario. In the emotion detection application scenario, the emotion detection service provider can detect the emotions of its user individuals through the method provided by an embodiment of the present application, which is beneficial to reducing the computing resources required for emotion detection and effectively improving the accuracy of emotion detection.

[0064] The gait emotion recognition method provided by the embodiments of the present application can be applied to the application scenario of disease-assisted diagnosis. In the application scenario of disease-assisted diagnosis, the user of the assisted diagnosis service can use the method provided by the embodiments of the present application to perform mental health diagnosis on its user individuals. For example, the method provided by the embodiments of the present application can be used to perform mental health diagnosis on whether a user individual has depression, anxiety symptoms, etc., which is beneficial to improving the efficiency and accuracy of diagnosing mental diseases such as depression.

[0065] Referring to Figure 1 , in the embodiments of the present application, a gait emotion recognition method includes:

[0066] Step 110, obtaining gait sequence data and gait emotion information of the gait sequence data;

[0067] In the embodiments of the present application, the gait sequence data may be a gait skeleton sequence of a user individual, and the gait skeleton sequence may be obtained based on a gait video; alternatively, the gait sequence data may be a feature map corresponding to the gait skeleton sequence, which may be obtained by performing skeleton key point extraction and normalization processing on the gait skeleton sequence based on the OpenPose algorithm to obtain a feature map with dimensions C×T×V. C is the number of channels, which is initially the coordinate dimension of human joints, T is the number of frames corresponding to the gait skeleton sequence, and V is the number of extracted skeleton key points.

[0068] It can be understood that the gait emotion information contains features (such as speed, acceleration, joint distance, etc.) that can directly (strong emotion expression) or indirectly (weak emotion expression) reflect emotions in different dimensions. The gait emotion information can be calculated based on the gait sequence data. For example, the joint distance feature can be calculated based on the key points in the gait sequence data. Specifically, the joint distance feature can be obtained by calculating the coordinate distance between two skeleton joint points. Other features in the gait emotion information can be obtained by simple analogy, and the present application will not elaborate here.

[0069] Step 120, inputting the gait sequence data into a gait skeleton framework for skeleton feature extraction to obtain enhanced skeleton features output by the gait skeleton framework and skeleton feature scores of the enhanced skeleton features;

[0070] In the embodiments of the present application, the gait sequence data can be input into the gait skeleton framework, and the gait skeleton framework enhances and extracts the input gait sequence data to obtain enhanced skeleton features and skeleton feature scores of the enhanced skeleton features. The skeleton feature scores can be obtained based on performing a Softmax operation on the enhanced skeleton features.

[0071] In some embodiments, the gait skeleton framework includes a plurality of cascaded spatio-temporal feature extraction networks, and the spatio-temporal feature extraction network is used to perform the following steps:

[0072] A1. Obtain the skeleton input features;

[0073] A2. Extract the gait spatial features from the skeleton input features to obtain intermediate spatial features;

[0074] In the embodiments of the present application, the gait skeleton framework includes a plurality of cascaded spatio-temporal feature extraction networks. The specific number of cascaded spatio-temporal feature extraction networks can be set according to the actual situation. For example, the number of cascaded spatio-temporal feature extraction networks in the gait skeleton framework can be any one of 3, 4, 5, etc. The present application does not limit the number of spatio-temporal feature extraction networks here.

[0075] Specifically, taking the number of cascaded spatio-temporal feature extraction networks in the gait skeleton framework as 3 as an example in the embodiments of the present application, for the first spatio-temporal feature extraction network in the gait skeleton framework, the skeleton input features of this spatio-temporal feature extraction network can be the original gait sequence features of the gait sequence data, and the original gait sequence features can specifically be the feature map of the gait sequence data; for the second spatio-temporal feature extraction network in the gait skeleton framework, the skeleton input features of this spatio-temporal feature extraction network can be the skeleton output features output by the first spatio-temporal feature extraction network; and for the third spatio-temporal feature extraction network in the gait skeleton framework, the skeleton input features of this spatio-temporal feature extraction network can be the skeleton output features output by the second spatio-temporal feature extraction network, and the skeleton output features of this spatio-temporal feature extraction network can be determined as the enhanced skeleton features output by the gait skeleton framework.

[0076] It can be understood that for a certain spatio-temporal feature extraction network, the gait spatial feature extraction can be based on a graph convolutional network (GCN) to extract the gait features in the space of the skeleton input features input to this spatio-temporal feature extraction network, so as to obtain intermediate spatial features.

[0077] A3. Extract the skeleton temporal attention from the intermediate spatial features to obtain skeleton output features;

[0078] Further, the step A3 of extracting the skeleton temporal attention from the intermediate spatial features to obtain skeleton output features includes:

[0079] A31. Extract shallow temporal features from the intermediate spatial features to obtain a shallow spatio-temporal feature map;

[0080] In the embodiment of the present application, the skeleton temporal attention extraction may be to extract the temporal features at a deep level of the intermediate spatial features, and then capture the complex spatio-temporal implicit dependencies and spatio-temporal features corresponding to the skeleton input features based on the attention mechanism, so as to obtain the skeleton output features output by the spatio-temporal feature extraction network.

[0081] It can be understood that the shallow temporal feature extraction in step A31 may be to apply a layer of dilated convolution to the intermediate spatial features in the temporal dimension, and extract the shallow temporal features of the intermediate spatial features through the dilated convolution in the temporal dimension, so as to obtain a shallow spatio-temporal feature map.

[0082] A32. Perform deep temporal feature extraction on the shallow spatio-temporal feature map to obtain a deep spatio-temporal feature map;

[0083] Further, the spatio-temporal feature extraction network includes a plurality of parallel temporal depth dilated convolution modules, and the dilation rate of each temporal depth dilated convolution module is different. The step A32, performing deep temporal feature extraction on the shallow spatio-temporal feature map to obtain a deep spatio-temporal feature map, includes:

[0084] A321. Input the shallow spatio-temporal feature map into all the temporal depth dilated convolution modules for multi-scale temporal feature extraction to obtain a plurality of deep dilated feature maps, and each deep dilated feature map corresponds to one of the temporal depth dilated convolution modules;

[0085] A323. Perform fusion of the deep dilated feature maps to obtain the deep spatio-temporal feature map.

[0086] In the embodiment of the present application, for a certain spatio-temporal feature extraction network, it includes a plurality of parallel temporal depth dilated convolution modules, and the dilation rate corresponding to each temporal depth dilated convolution module is different. Specifically, the specific number of parallel temporal depth dilated convolution modules in the spatio-temporal feature extraction network can be set according to the actual situation. For example, the number of parallel temporal depth dilated convolution modules in the spatio-temporal feature extraction network can be any one of 2, 3, 5, 7, etc. The present application does not limit the specific number of parallel temporal dilated convolution modules here.

[0087] Specifically, in the embodiment of the present application, taking the number of parallel temporal depthwise dilated convolution modules in the spatio-temporal feature extraction network as 3, the spatio-temporal feature extraction network includes 3 dilated convolutions in parallel with different dilation rates. For example, the dilation rate of the first dilated convolution can be 1, the dilation rate of the second dilated convolution can be 5, and the dilation rate of the third dilated convolution can be 5. Each dilated convolution is used to simulate a temporal convolution kernel of a corresponding size. By respectively inputting the shallow spatio-temporal feature maps into the 3 dilated convolutions with different dilation rates for deep temporal feature extraction, deep dilated feature maps corresponding to each dilated convolution are obtained, and the time dimensions corresponding to each deep dilated feature map are different.

[0088] It can be understood that step A323 can fuse all the deep dilated feature maps in the spatio-temporal feature extraction network to obtain a deep spatio-temporal feature map, and the deep spatio-temporal feature map includes a set of deep temporal features containing different time dimension information in the time dimension.

[0089] A33. Perform attention feature capture on the deep spatio-temporal feature map to obtain the skeleton output feature.

[0090] Further, the step A33, performing attention feature capture on the deep spatio-temporal feature map to obtain the skeleton output feature, includes:

[0091] A331. Calculate the spatio-temporal attention weights corresponding to the deep spatio-temporal feature map to obtain spatio-temporal attention weights corresponding to the deep spatio-temporal feature map;

[0092] A332. Perform attention weight fusion on the deep spatio-temporal feature map according to the spatio-temporal attention weights to obtain the skeleton output feature.

[0093] In the embodiment of the present application, attention feature capture is used to further capture complex spatio-temporal implicit dependency relationships and key spatio-temporal features from the deep spatio-temporal feature map based on the attention mechanism, so as to obtain the skeleton output feature output by the spatio-temporal feature extraction network.

[0094] It can be understood that step A331 can be based on the attention mechanism to obtain the spatio-temporal attention weights corresponding to the deep spatio-temporal feature map. Specifically, operations such as global average pooling and fully connected operations can be sequentially performed on the deep spatio-temporal feature map to obtain the spatio-temporal attention weights in sequence. Also, step A332 can perform attention fusion on the spatio-temporal attention weights and the deep spatio-temporal feature map. Specifically, a Hadamard product operation can be performed on the spatio-temporal attention weights and the deep spatio-temporal feature map to obtain the skeleton output feature containing complex spatio-temporal implicit dependency relationships and key features.

[0095] Step 130: Input the gait emotion information into the gait emotion framework for emotion feature extraction to obtain the denoised emotion features output by the gait emotion framework and the emotion feature scores of the denoised emotion features;

[0096] In the embodiments of the present application, the gait emotion information can be input into the gait emotion framework, and the gait emotion framework performs emotion denoising and extraction on the input gait emotion information, so as to obtain the denoised emotion features and the emotion feature scores of the denoised emotion features. The emotion feature scores can be obtained based on performing a Softmax operation on the denoised emotion features.

[0097] In some embodiments, the gait emotion framework includes a plurality of cascaded emotion denoising networks, and the emotion denoising network is used to perform the following steps:

[0098] B1: Obtain emotion input features and a preset zero tensor threshold;

[0099] B2: Extract eigenvalue of the emotion input features to obtain the absolute value of the feature and the global average value of the feature;

[0100] B3: According to the global average value of the feature, perform feature scaling on the emotion input features to obtain emotion scaled features;

[0101] In the embodiments of the present application, the gait emotion framework includes a plurality of cascaded emotion denoising networks. The specific number of cascaded emotion denoising networks can be set according to actual situations. For example, the number of cascaded emotion denoising networks in the gait emotion framework can be any one of 2, 4, 6, etc. Specifically, for the first emotion denoising network in the gait emotion framework, the input of this emotion denoising network can be gait emotion features, and the gait emotion features can be obtained by performing dimensional mapping on the gait emotion information using a 1×1 convolution; for the second and subsequent emotion denoising networks in the gait emotion framework, their input can be the emotion output features output by the previous emotion denoising network. For example, the emotion input features input to the third emotion denoising network can be the emotion output features output by the second emotion denoising network.

[0102] It can be understood that the zero tensor threshold is a tensor with all component values being zero. Step B2 can be to perform an absolute value operation and a global average pooling operation on the feature map of the emotion input features respectively, so as to obtain the absolute value of the feature and the global average value of the feature of the emotion input features; then, the feature weight scaling in step B3 can first be based on a fully connected operation to calculate the weighted weight coefficients of each feature on the feature map of the emotion input features, and then perform a Hadamard product operation on the global average value of the feature and each weighted weight coefficient, so as to realize the scaling of the weight of each feature of the emotion input features to obtain emotion scaled features.

[0103] B4. Perform soft-threshold feature selection on the emotion scaling features according to the zero tensor threshold and the absolute value of the features, to obtain emotion soft-threshold features;

[0104] Further, the step B4. Perform soft-threshold feature selection on the emotion scaling features according to the zero tensor threshold and the absolute value of the features, to obtain emotion soft-threshold features, includes:

[0105] B41. Calculate the feature difference of the emotion scaling features according to the absolute value of the features, to obtain emotion difference features;

[0106] B42. Perform soft-threshold feature selection on the emotion difference features according to the zero tensor threshold, to obtain the emotion soft-threshold features.

[0107] In the embodiment of the present application, step B41 may first be to calculate the feature difference between the absolute value of the feature and each feature on the feature map of the emotion scaling feature based on the absolute value of the feature, so as to obtain a feature map composed of a number of feature differences, and determine the feature map composed of the number of feature differences as the emotion difference features. Step B42 may first be to compare the size of each feature difference on the feature map of the emotion difference features with the corresponding component value on the zero tensor threshold respectively, then retain the maximum value between the feature difference and the corresponding component value, and then construct a feature map based on all the retained maximum values, and determine the feature map constructed based on all the retained maximum values as the emotion soft-threshold features. Specifically, for a certain feature difference in the emotion difference features, the size of the feature difference may be compared with the corresponding component value. If the feature difference is greater than the component value, the feature difference is determined as the maximum value and retained; or, if the feature difference is less than or equal to the component value, the component value is determined as the maximum value and retained. The same applies to the remaining feature differences and the corresponding component values, and can be simply deduced by analogy.

[0108] It should be noted that since the emotion input features input to each emotion denoising network in the gait emotion framework are not the same, the feature soft thresholds corresponding to each emotion denoising network in the embodiment of the present application are not necessarily the same, and the corresponding emotion soft-threshold features obtained by each emotion denoising network can change adaptively based on the emotion input features.

[0109] B5. Perform feature denoising on the emotion input features according to the emotion soft-threshold features, to obtain emotion output features;

[0110] In an embodiment of the present application, step B5 may be to perform a softening and denoising operation on the features on the feature map of the emotional input features based on each soft threshold feature on the feature map of the emotional soft threshold features, so as to obtain the emotional output features. There are already various specific softening and denoising methods. For example, the emotional soft threshold features and the emotional input features can be multiplied to obtain the emotional output features.

[0111] Step 140: According to the enhanced skeleton features and the denoised emotional features, perform decision-level fusion on the skeleton feature scores and the emotional feature scores to obtain the emotion recognition result of the gait sequence data.

[0112] In an embodiment of the present application, the decision-level fusion in step 140 may be to determine the weighting coefficients of the skeleton feature scores and the weighting coefficients of the emotional feature scores based on the enhanced skeleton features and the denoised emotional features. There are already various ways to obtain the specific weighting coefficients, which will not be elaborated here in the present application; then, based on the skeleton feature scores and their weighting coefficients, and the emotional feature scores and their weighting coefficients, fusion is performed by weighted addition to obtain the final probability distribution result (i.e., the emotion recognition result).

[0113] Next, a gait emotion recognition system proposed according to an embodiment of the present application will be described in detail with reference to the accompanying drawings.

[0114] Refer to Figure 2 , a gait emotion recognition system proposed in an embodiment of the present application includes:

[0115] A first processing unit, configured to obtain gait sequence data and the gait emotion information of the gait sequence data;

[0116] A second processing unit, configured to input the gait sequence data into a gait skeleton framework for skeleton feature extraction to obtain the enhanced skeleton features output by the gait skeleton framework and the skeleton feature scores of the enhanced skeleton features;

[0117] A third processing unit, configured to input the gait emotion information into a gait emotion framework for emotional feature extraction to obtain the denoised emotional features output by the gait emotion framework and the emotional feature scores of the denoised emotional features;

[0118] A fourth processing unit, configured to perform decision-level fusion on the skeleton feature scores and the emotional feature scores according to the enhanced skeleton features and the denoised emotional features to obtain the emotion recognition result of the gait sequence data.

[0119] It can be understood that the content in the above method embodiments is applicable to the system embodiments of the present application. The functions specifically implemented in the system embodiments of the present application are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0120] Referring to Figure 3 , the embodiments of the present application further provide an electronic device, including:

[0121] At least one processor 201;

[0122] At least one memory 202, configured to store at least one program;

[0123] When the at least one program is executed by the at least one processor 201, the at least one processor 201 implements the above method embodiments.

[0124] Similarly, it can be understood that the content in the above method embodiments is applicable to the device embodiments of the present application. The functions specifically implemented in the device embodiments of the present application are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0125] The embodiments of the present application further provide a computer-readable storage medium, in which a program executable by a processor 201 is stored, and the program executable by the processor 201 is used to implement the above method embodiments when executed by the processor 201.

[0126] Similarly, the content in the above method embodiments is applicable to the computer-readable storage medium embodiments of the present application. The functions specifically implemented in the computer-readable storage medium embodiments of the present application are the same as those in the above method embodiments, and the beneficial effects achieved are also the same as those in the above method embodiments.

[0127] In some alternative embodiments, the functions / operations mentioned in the block diagrams may not occur in the order mentioned in the operation diagrams. For example, depending on the functions / operations involved, two consecutive blocks shown may actually be executed substantially simultaneously or the blocks can sometimes be executed in the reverse order. In addition, the embodiments presented and described in the flowcharts of the present application are provided by way of example for the purpose of providing a more comprehensive understanding of the technology. The disclosed methods are not limited to the operations and logical flows presented herein. Alternative embodiments are foreseeable, in which the order of various operations is changed and the sub-operations described as part of a larger operation are executed independently.

[0128] In addition, although the present application has been described in the context of functional modules, it should be understood that, unless otherwise stated to the contrary, one or more of the functions and / or features may be integrated in a single physical device and / or software module, or one or more functions and / or features may be implemented in separate physical devices or software modules. It should also be understood that a detailed discussion of the actual implementation of each module is not necessary for understanding the present application. Rather, given the attributes, functions, and internal relationships of the various functional modules in the devices disclosed herein, the actual implementation of the modules will be understood within the ordinary skills of an engineer. Thus, those skilled in the art can implement the present application as set forth in the claims without undue experimentation. It should also be understood that the specific concepts disclosed are merely illustrative and are not intended to limit the scope of the present application, which is determined by the full scope of the appended claims and their equivalents.

[0129] If a function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of this technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which may be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method according to an embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read-only memories (ROMs), random access memories (RAMs), magnetic disks, or optical discs.

[0130] The logic and / or steps represented in the flowchart or otherwise described herein, for example, can be considered as a predefined sequence of executable instructions for implementing a logical function, and can be specifically implemented in any computer-readable medium for use by an instruction execution system, apparatus, or device (such as a computer-based system, a system including a processor, or other systems that can fetch and execute instructions from the instruction execution system, apparatus, or device), or in conjunction with these instruction execution systems, apparatuses, or devices. For the purposes of this specification, a "computer-readable medium" can be any device that can contain, store, communicate, propagate, or transport a program for use by or in conjunction with an instruction execution system, apparatus, or device.

[0131] More specific examples (non-exhaustive list) of computer-readable media include the following: an electrical connection (electronic device) having one or more wirings, a portable computer diskette (magnetic device), a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber device, and a portable compact disc read-only memory (CDROM). Additionally, the computer-readable media can even be paper or other suitable media on which a program can be printed, since the program can be obtained electronically, for example, by optically scanning the paper or other media, followed by editing, interpretation, or otherwise processing as appropriate, and then stored in a computer memory.

[0132] It should be understood that various parts of the present application can be implemented by hardware, software, firmware, or a combination thereof. In the above-described embodiments, multiple steps or methods can be implemented by software or firmware stored in a memory and executed by a suitable instruction execution system. For example, if implemented in hardware, as in another embodiment, any one or a combination of the following techniques well known in the art can be used: discrete logic circuits having logic gate circuits for implementing logical functions on data signals, application-specific integrated circuits having appropriate combinational logic gate circuits, programmable gate arrays (PGAs), field programmable gate arrays (FPGAs), etc.

[0133] In the foregoing description of the present specification, descriptions with reference to the terms "one embodiment / example", "another embodiment / example", or "certain embodiments / examples", etc. mean that the specific features, structures, materials, or characteristics described in connection with the embodiment or example are included in at least one embodiment or example of the present application. In this specification, the schematic representations of the above terms do not necessarily refer to the same embodiment or example. Moreover, the specific features, structures, materials, or characteristics described can be combined in any one or more embodiments or examples in a suitable manner.

[0134] Although the embodiments of the present application have been shown and described, those of ordinary skill in the art can understand that various changes, modifications, substitutions, and variations can be made to these embodiments without departing from the principles and spirit of the present application, and the scope of the present application is defined by the claims and their equivalents.

[0135] The above has specifically described the preferred embodiments of the present application, but the present application is not limited to the embodiments. Those skilled in the art can make various equivalent deformations or substitutions without departing from the spirit of the present application, and these equivalent deformations or substitutions are all included within the scope defined by the claims of the present application.

Claims

1. A gait emotion recognition method, characterized in that: include: Acquire gait sequence data and gait emotion information of the gait sequence data; Inputting the gait sequence data into a gait skeleton framework to extract skeleton features, thereby obtaining enhanced skeleton features output by the gait skeleton framework and skeleton feature scores of the enhanced skeleton features; Inputting the gait emotion information into a gait emotion framework to extract emotion features, thereby obtaining denoised emotion features output by the gait emotion framework and an emotion feature score of the denoised emotion features; According to the enhanced skeleton features and the denoised emotional features, the skeleton feature scores and the emotional feature scores are fused at a decision level to obtain an emotion recognition result of the gait sequence data.

2. The method according to claim 1, characterized in that The gait skeleton framework includes several cascaded spatiotemporal feature extraction networks, which are used to perform the following steps: Get skeleton input features; Extracting gait space features from the skeleton input features to obtain intermediate space features; Performing skeleton temporal attention extraction on the intermediate spatial features to obtain skeleton output features; The skeleton input feature is the original gait sequence feature of the gait sequence data or the skeleton output feature output by the previous spatiotemporal feature extraction network.

3. The method according to claim 2, characterized in that The step of performing skeleton time attention extraction on the intermediate spatial features to obtain skeleton output features includes: Extracting shallow temporal features from the intermediate spatial features to obtain a shallow spatiotemporal feature map; Performing deep temporal feature extraction on the shallow spatiotemporal feature map to obtain a deep spatiotemporal feature map; Attention features are captured on the deep spatiotemporal feature map to obtain the skeleton output features.

4. The method according to claim 3, characterized in that The spatiotemporal feature extraction network includes a plurality of parallel time-depth hole convolution modules, each of which has a different hole rate, and the shallow spatiotemporal feature map is subjected to deep time feature extraction to obtain a deep spatiotemporal feature map, including: Input the shallow spatiotemporal feature map into all the temporal depth hole convolution modules to extract multi-scale temporal features, and obtain a plurality of deep hole feature maps, each of which corresponds to one temporal depth hole convolution module; All the deep hole feature maps are fused to obtain the deep spatiotemporal feature map.

5. The method according to claim 3, characterized in that: The step of capturing the attention features of the deep spatiotemporal feature map to obtain the skeleton output features includes: Calculating the attention weight of the deep spatiotemporal feature map to obtain the spatiotemporal attention weight corresponding to the deep spatiotemporal feature map; According to the spatiotemporal attention weights, the deep spatiotemporal feature map is fused with attention weights to obtain the skeleton output features.

6. The method according to claim 1, characterized in that The gait emotion framework includes several cascaded emotion denoising networks, which are used to perform the following steps: Get sentiment input features and preset zero tensor threshold; Extracting feature values ​​of the emotion input features to obtain feature absolute values ​​and feature global average values; Performing feature scaling on the emotion input feature according to the feature global average value to obtain an emotion scaling feature; According to the zero tensor threshold and the feature absolute value, performing soft threshold feature selection on the sentiment scaling feature to obtain a sentiment soft threshold feature; According to the emotion soft threshold feature, feature denoising is performed on the emotion input feature to obtain an emotion output feature; Among them, the emotion input feature is the gait emotion feature of the gait emotion information or the emotion output feature output by the previous emotion denoising network.

7. The method according to claim 6, characterized in that The step of performing soft threshold feature selection on the emotion scaling feature according to the zero tensor threshold and the feature absolute value to obtain the emotion soft threshold feature comprises: According to the feature absolute value, performing feature difference calculation on the emotion scaling feature to obtain an emotion difference feature; According to the zero tensor threshold, soft threshold feature selection is performed on the sentiment difference feature to obtain the sentiment soft threshold feature.

8. A gait emotion recognition system, characterized in that: include: A first processing unit, used for acquiring gait sequence data and gait emotion information of the gait sequence data; A second processing unit is used for inputting the gait sequence data into a gait skeleton framework to extract skeleton features, and obtain enhanced skeleton features output by the gait skeleton framework, and skeleton feature scores of the enhanced skeleton features; A third processing unit is used to input the gait emotion information into a gait emotion framework to extract emotion features, obtain denoised emotion features output by the gait emotion framework, and obtain an emotion feature score of the denoised emotion features; The fourth processing unit is used to perform decision-level fusion on the skeleton feature score and the emotion feature score according to the enhanced skeleton feature and the denoised emotion feature to obtain the emotion recognition result of the gait sequence data.

9. An electronic device, characterized in that: include: at least one processor; at least one memory for storing at least one program; When the at least one program is executed by the at least one processor, the at least one processor implements the method according to any one of claims 1 to 7.

10. A computer-readable storage medium storing a program executable by a processor, characterized in that: The program executable by the processor is used to implement the method according to any one of claims 1 to 7 when executed by the processor.

Citation Information

Cited By

  • Emotion information identification method and device based on discrete labels, equipment and storage medium

    CN120954087A

  • Self-supervised gait emotion recognition method based on multi-scale contrast learning and related equipment

    CN122200785A