Head action recognition method and device, electronic equipment and storage medium
By analyzing the data synchronization and timing of different sensor signal channels in the sensor data, high-quality feature extraction is performed, which solves the problem of visual sensor dependence in the existing technology and realizes high-precision head motion recognition on devices without visual sensor equipment.
Patent Information
- Application Number
- CN202511795281.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-01
- Publication Date
- 2026-02-17
AI Technical Summary
Existing head pose estimation methods rely on visual sensors, which limits their applicability to devices without visual sensors and also result in low recognition accuracy.
By analyzing the data synchronization and timing of different sensor signal channels in the sensor data, high-quality feature extraction is performed to achieve head movement recognition.
It improves the accuracy and applicability of head motion recognition, enabling effective application on devices without visual sensors.
Smart Images

Figure CN121542583A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of artificial intelligence technology, and in particular to a head motion recognition method, device, electronic device, and storage medium. Background Technology
[0002] Head pose estimation has wide applications in many fields, such as facial recognition systems, driver monitoring systems, virtual reality, security monitoring systems, and student attention tracking. Currently, most head pose estimation methods are based on acquired facial images, using techniques like face detection and facial landmark detection to identify specific head pose angles. However, this method relies on visual sensor input and cannot be applied to devices without visual sensors, thus limiting its applicability. Summary of the Invention
[0003] This application provides a head motion recognition method, device, electronic device, and storage medium. The technical solution is as follows: On one hand, embodiments of this application provide a head motion recognition method, the method comprising: Acquire sensor data, which is used to characterize the head movement state and includes channel data of multiple sensor signal channels; Based on the data synchronization and data timing of different sensor signal channels in the sensor data, feature extraction is performed on the channel data of the multiple sensor signal channels to obtain the first feature; Head motion recognition is performed based on the first feature to obtain the head motion recognition result.
[0004] On the other hand, embodiments of this application provide a head motion recognition device, the device comprising: The data acquisition module is used to acquire sensor data, which is used to characterize the head movement state and includes channel data of multiple sensor signal channels; The feature extraction module is used to extract features from the channel data of the multiple sensor signal channels based on the data synchronization and data timing of different sensor signal channels in the sensor data, and obtain the first feature; The action recognition module is used to perform head action recognition based on the first feature to obtain the head action recognition result.
[0005] On the other hand, embodiments of this application provide an electronic device, which includes a processor and a memory, wherein the memory stores at least one computer instruction, which is loaded and executed by the processor to implement the head motion recognition method as described above.
[0006] On the other hand, embodiments of this application provide a computer-readable storage medium storing at least one computer instruction, which is executed by a processor to implement the head motion recognition method as described above.
[0007] On the other hand, embodiments of this application provide a computer program product, which includes computer instructions, and when a processor executes the computer instructions, it implements the head motion recognition method as described above.
[0008] In this embodiment, when sensor data representing head movement state and including channel data from multiple sensor signal channels are acquired, features can be extracted from the channel data of multiple sensor signal channels based on the data synchronization and temporal sequence of different sensor signal channels in the sensor data to obtain a first feature. This first feature is then used for head movement recognition to obtain a head movement recognition result. Thus, after acquiring sensor data reflecting head movement state, by specifically analyzing the data correlation and temporal sequence of different sensor signal channels in the sensor data, high-quality feature mining of the channel data from multiple sensor signal channels can be achieved for head movement recognition, improving the accuracy of head movement recognition. Furthermore, since the solution adopted in this application does not rely on the input of a visual sensor, it can be widely applied to various devices (e.g., headphones without a visual sensor). Attached Figure Description
[0009] Figure 1 This is a flowchart of a head motion recognition method provided in an exemplary embodiment of this application; Figure 2 This is a flowchart of a head motion recognition method provided in another exemplary embodiment of this application; Figure 3 This is a flowchart of a head motion recognition method provided in yet another exemplary embodiment of this application; Figure 4 This is a schematic diagram of the architecture of a spatial graph convolutional neural network provided in an exemplary embodiment of this application; Figure 5 This is a schematic diagram of the architecture of a time-graph convolutional neural network provided in an exemplary embodiment of this application; Figure 6 This is a system flowchart of a head motion recognition system provided in an exemplary embodiment of this application; Figure 7 This invention provides a structural block diagram of a head motion recognition device according to an exemplary embodiment of the present application. Figure 8A schematic diagram of the structure of an electronic device provided in an exemplary embodiment of this application is shown. Detailed Implementation
[0010] To make the objectives, technical solutions, and advantages of this application clearer, the embodiments of this application will be described in further detail below with reference to the accompanying drawings.
[0011] In this article, "multiple" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, or B alone. The character " / " generally indicates that the preceding and following related objects have an "or" relationship.
[0012] In related technologies, head pose estimation is mostly achieved by simply analyzing changes in sensor data, and is usually limited to a single type of sensor data. For example, it may rely solely on changes in data measured by a gyroscope, or solely on changes in images output by a visual sensor, or solely on changes in sound measured by a microphone. This coarse data processing method and over-reliance on a single type of sensor data will limit the accuracy of head pose recognition and restrict its application scenarios.
[0013] To address the aforementioned issues, this application provides a head motion recognition method that specifically analyzes the data correlation and temporal sequence of different sensor signal channels in sensor data. This enables in-depth mining of high-quality features from multiple sensor signal channels, allowing for more accurate and stable head motion recognition based on these features, thus improving the accuracy of the recognition results. Compared to general and coarse data processing methods, the approach adopted in this application allows for fine-grained analysis of sensor data at the granularity of specific sensor signal channels, resulting in more accurate head motion recognition. Furthermore, the approach is not limited to a single type of sensor data and can be widely applied to various devices (such as headphones), significantly improving the accuracy and practicality of head motion recognition.
[0014] The head motion recognition method provided in this application can be applied to wearable devices, which may include at least one of the following: headphones (such as in-ear headphones, semi-in-ear headphones, over-ear headphones, etc.), smart glasses, smart earrings or smart collars, augmented reality (AR) glasses, virtual reality (VR) glasses, mixed reality (MR) glasses, smart helmets, etc. This application does not limit the specific type of wearable device.
[0015] Optionally, the wearable device includes sensors, which may include, but are not limited to, at least one of the following: a gyroscope sensor, an accelerometer, an inertial measurement unit (IMU), a pressure sensor, a microphone, and other sensors. The gyroscope sensor can be used to detect the rotational motion of the carrier, such as the angular velocity (rotational speed) or angular displacement (rotation angle) around various axes. The accelerometer can be used to detect the acceleration motion of the carrier, such as the magnitude of acceleration in various directions; when stationary, it can detect the magnitude and direction of gravity. The IMU can be a sensor that detects the acceleration and rotational motion of the carrier through inertial devices, and internally consists of a gyroscope sensor and an accelerometer sensor. Other sensors may include Bluetooth sensors, voice pick-up sensors (VPU), etc., and are not limited here.
[0016] The head motion recognition method provided in this application can also be applied to a terminal, which may include at least one of the following devices: smartphone, tablet computer, personal computer (PC), smartwatch, smart home device (such as smart TV, smart speaker, etc.). This application does not limit the specific type of terminal.
[0017] Optionally, a communication connection can be established between the terminal (such as a mobile phone) and the wearable device (such as headphones) via wired or wireless means. When a communication connection is established, the terminal and the wearable device can also exchange data through the established communication connection, such as transmitting the sensor data collected by the wearable device to the terminal, or transmitting the head motion recognition results determined by the wearable device to the terminal.
[0018] The head motion recognition method provided in this application embodiment can also be executed jointly by a terminal and a wearable device, wherein the terminal and the wearable device establish a communication connection. For example, after the wearable device performs in-depth mining of high-quality features of channel data from multiple sensor signal channels in the sensor data, the terminal performs head motion recognition based on the mined features.
[0019] For ease of description, the head motion recognition method provided in this application embodiment will be described below using an electronic device as the execution subject. The electronic device can be the aforementioned terminal and / or wearable device.
[0020] It is understood that when the electronic device is the aforementioned terminal, the noise reduction method provided in this application embodiment can be implemented by the terminal alone. When the electronic device is the aforementioned wearable device, the noise reduction method provided in this application embodiment can be implemented by the wearable device alone. When the electronic device is both the aforementioned terminal and the aforementioned wearable device, the noise reduction method provided in this application embodiment can be implemented by the terminal and the wearable device through data interaction. This application embodiment does not limit this aspect.
[0021] For ease of explanation, the following embodiments use mobile phones as terminals and earphones as wearable devices for illustrative purposes.
[0022] Please refer to Figure 1 The diagram illustrates a flowchart of a head motion recognition method provided in an exemplary embodiment of this application. The method may include the following steps: Step 101: Acquire sensor data. The sensor data is used to characterize the head movement state and includes channel data of multiple sensor signal channels.
[0023] The aforementioned sensor data can be collected by at least one sensor, which can be located in the same wearable device, such as all sensors in the headphones, or in different wearable devices, such as headphones and smart glasses. No limitation is imposed here.
[0024] It's understandable that when a user wears wearable devices like headphones, their head movements cause the headphones to move as well, resulting in a high degree of overlap between the head's and head's movements. Different head movements also affect the data collected by the sensors within the headphones, causing them to fluctuate. Therefore, the data collected by the headphones' sensors can, to some extent, reflect the headphones' movement, and consequently, the user's head movements.
[0025] Optionally, when the above sensor data is acquired by at least one sensor, the data acquired by each of the at least one sensor can characterize the current head movement state.
[0026] In one possible implementation, at least one of the sensors mentioned above includes an accelerometer. It can be understood that when the user's head performs different movements such as nodding or shaking, the accelerometer in the headphones collects acceleration data along the X, Y, and Z axes. , , The data will show different fluctuations, so the data collected by the accelerometer can characterize the current head movement state.
[0027] In one possible implementation, at least one of the aforementioned sensors includes a gyroscope sensor. It can be understood that when the user's head performs different movements such as nodding or shaking, the gyroscope sensor in the headphones collects angular velocity data on the X, Y, and Z axes. , , The data will also show different fluctuations, so the data collected by the gyroscope sensor can also characterize the current head movement state.
[0028] In one possible implementation, at least one of the sensors mentioned above includes a pressure sensor. It is understood that when a user's head performs different movements such as nodding or shaking, the pressure data collected by the pressure sensor in the headphones will exhibit different fluctuations. Therefore, the data collected by the pressure sensor can also characterize the current head movement state.
[0029] It is understood that the type of sensor in the above-mentioned at least one sensor is not limited in this application embodiment. It can also be other sensors that can sense the head movement state, such as sound sensors (such as microphones), which will not be described in detail here.
[0030] In some embodiments, where sensor data is acquired by a single sensor, that single sensor may have multiple sensor signal channels. That is, the single sensor can acquire channel data from multiple sensor signal channels. For example, the single sensor may be an accelerometer, which can acquire... , , The data from these three sensor signal channels. For example, the single sensor could be a gyroscope sensor, which can collect... , , The data from these three sensor signal channels. For example, the single sensor could be a six-axis IMU composed of an accelerometer and a gyroscope, which can acquire... , , , , The data from these six sensor signal channels. It is understood that the type of a single sensor is not limited in this embodiment; it can also be other sensors capable of acquiring channel data from multiple sensor signal channels, such as a 9-axis IMU composed of an accelerometer, a gyroscope, and a magnetometer, which will not be elaborated upon here.
[0031] In some embodiments, when sensor data is acquired by multiple sensors, each of the multiple sensors may have one or more sensor signal channels, so that the sensor data may consist of channel data from different sensor signal channels of different modes. For example, when the multiple sensors include a six-axis IMU and a pressure sensor, the sensor data acquired by the multiple sensors may consist of channel data from the six sensor signal channels of the six-axis IMU and channel data from a single sensor signal channel of the pressure sensor.
[0032] Step 102: Based on the data synchronization and data timing of different sensor signal channels in the sensor data, feature extraction is performed on the channel data of multiple sensor signal channels to obtain the first feature.
[0033] In this embodiment, when sensor data containing channel data of multiple sensor signal channels is obtained, fine-grained analysis of the sensor data can be performed at the granularity of specific sensor signal channels, based on the data synchronization and temporal sequence of different sensor signal channels. Optionally, since the channel data of each sensor signal channel collected by the sensor in continuous time is continuous temporal data, the sensor signal channel can also be called a temporal channel.
[0034] Specifically, the data synchronization of the sensor signal channels refers to the synchronized changes in channel data between various sensor signal channels during various head movements such as nodding and shaking. The data temporality of the sensor signal channels refers to the changes in channel data over time during various head movements such as nodding and shaking.
[0035] Optionally, the strength of data synchronization can be distinguished according to the level of synchronization. The stronger the synchronous change in the channel data between two sensor signal channels, the higher the synchronization level of the two sensor signal channels; conversely, the weaker the synchronous change in the channel data between the two sensor signal channels, the lower the synchronization level of the two sensor signal channels. Optionally, the synchronization level can have two levels, namely strong synchronization and weak synchronization, or three levels, namely strong synchronization, weak synchronization, and no synchronization. This application does not limit the specific division of synchronization levels in its embodiments.
[0036] In one possible implementation, when the channel data of two sensor signal channels do not change synchronously, these two sensor signal channels can be defined as an asynchronous channel pair, meaning the synchronization level of these two sensor signal channels is very low / none. When the synchronous changes in the channel data of two sensor signal channels are relatively weak, these two sensor signal channels can be defined as a weakly synchronized channel pair, meaning the synchronization level of these two sensor signal channels is low. When the synchronous changes in the channel data of two sensor signal channels are relatively strong, these two sensor signal channels can be defined as a strongly synchronized channel pair, meaning the synchronization level of these two sensor signal channels is high.
[0037] In some embodiments, the first feature extracted from the channel data of multiple sensor signal channels in the sensor data may be a feature that integrates the data synchronization and data temporal sequence of each sensor signal channel in the multiple sensor signal channels. Thus, the extracted first feature has stronger representational power and robustness, enabling in-depth mining of high-quality features in the sensor data.
[0038] Optionally, the first feature can be extracted through various data processing methods, and the embodiments of this application do not limit this.
[0039] In one possible implementation, a strong synchronization channel pair can be pre-set. After extracting the channel data of the strong synchronization channel pair from the channel data of multiple sensor signal channels, the trend of the channel data of the strong synchronization channel pair changing over time can be analyzed to obtain the first feature that integrates the data synchronization and data timing of the sensor signal channels.
[0040] In one possible implementation, the first feature can also be extracted using a neural network model.
[0041] Step 103: Perform head action recognition based on the first feature to obtain the head action recognition result.
[0042] Optionally, the head action recognition result includes the probability of each head action category, which can also be referred to as confidence level, score, etc. The head action categories may include nodding and shaking, or tilting the head up, tilting the head to the left, tilting the head to the right, etc. This application embodiment does not limit the identifiable head action categories. In one possible implementation, in addition to the probability of each head action category, the head action recognition result may also include a label corresponding to each head action category.
[0043] In some embodiments, reference features corresponding to each head movement category can be preset. After feature matching between the first feature and the reference features of each head movement category, the probability of each head movement category can be determined based on the degree of feature matching, wherein the degree of feature matching is positively correlated with the probability.
[0044] In some embodiments, the first feature can be input into the head action classification model to obtain the probability of each head action category output by the head action classification model. Optionally, the head action classification model can consist of fully connected layers and a Softmax activation function.
[0045] In one possible implementation, the head action category with the highest probability can be used as the head action recognition result.
[0046] In one possible application scenario, when a user wears the headphones and makes a nodding (or shaking) motion, the headphones collect real-time sensor data through at least one sensor on its own device. Based on the data synchronization and timing of different sensor signal channels in the sensor data, the headphones can extract features from the channel data of multiple sensor signal channels to obtain a first feature in real time. Then, the headphones perform head movement recognition based on the first feature obtained in real time, and can obtain the current head movement recognition result as nodding (or shaking) in real time. The headphones can then transmit the current head movement recognition result of nodding (or shaking) to the mobile phone.
[0047] In one possible application scenario, when a user wears headphones and makes a nodding (or shaking) motion, the headphones collect real-time sensor data through at least one sensor on its own device. This data is then transmitted to a mobile phone. The mobile phone, based on the data synchronization and timing of different sensor signal channels within the sensor data, extracts features from the channel data of multiple sensor signal channels to obtain a first feature in real time. The mobile phone then performs head movement recognition based on this first feature, resulting in a real-time head movement recognition result of nodding (or shaking). Alternatively, the headphones can extract features from the channel data of multiple sensor signal channels based on the data synchronization and timing of different sensor signal channels, obtain a first feature in real time, and transmit this first feature to the mobile phone. The mobile phone then performs head movement recognition based on this first feature, also resulting in a real-time head movement recognition result of nodding (or shaking).
[0048] In summary, in this embodiment, when sensor data representing head movement state and including channel data from multiple sensor signal channels are obtained, features can be extracted from the channel data of multiple sensor signal channels based on the data synchronization and temporal sequence of different sensor signal channels in the sensor data to obtain a first feature. This first feature is then used for head movement recognition to obtain a head movement recognition result. Thus, by specifically analyzing the data correlation and temporal sequence of different sensor signal channels in the sensor data, high-quality feature mining of channel data from multiple sensor signal channels can be achieved for head movement recognition, improving the accuracy of head movement recognition.
[0049] Feature extraction methods In some embodiments, a Spatial Temporal Graph Convolutional Network (ST-GCN) can be introduced to extract the first feature. Here, the sensor signal channels can be considered as "spatial nodes" of the ST-GCN, the data synchronization between sensor signal channels can be considered as "spatial correlations between spatial nodes," and the data temporality of the sensor signal channels can be considered as the temporal information of the spatial nodes. Thus, the ST-GCN can be used to extract the first feature, which integrates the data synchronization and data temporality of the sensor signal channels.
[0050] In one possible implementation, please refer to Figure 2 The diagram illustrates a flowchart of a head motion recognition method provided in another exemplary embodiment of this application. The method may include the following steps: Step 201: Acquire sensor data. The sensor data is used to characterize the head movement state and includes channel data of multiple sensor signal channels.
[0051] In one possible implementation, sensor data corresponding to each time point within a single time window can be acquired to analyze the sensor data of the local time window and obtain the head action recognition result for that local time window. This single time window can slide along the time dimension to acquire sensor data within different time windows; this time window can also be referred to as a data sliding window.
[0052] Optionally, the time window can be a time window with a first length to acquire sensor data of a fixed length. For example, acquiring 3 seconds of sensor data (3 seconds is 36 frames, assuming a sampling rate of 12Hz). Optionally, the time window can be a time window with a data frame number equal to a first frame number to acquire sensor data of a fixed number of frames. For example, acquiring 64 frames of sensor data.
[0053] Step 202: Input the channel data of multiple sensor signal channels into the feature extraction model to obtain the first feature output by the feature extraction model. The feature extraction model includes a spatial graph convolutional neural network and a temporal graph convolutional neural network. The spatial graph convolutional neural network is used to determine the data synchronization of different sensor signal channels in the sensor data, and the temporal graph convolutional neural network is used to determine the data temporality of different sensor signal channels in the sensor data.
[0054] Optionally, the feature extraction model can be constructed and trained based on a spatiotemporal graph convolutional neural network. This spatiotemporal graph convolutional neural network can include a spatial graph convolutional neural network and a temporal graph convolutional neural network.
[0055] In some embodiments, an input vector can be generated based on channel data from multiple sensor signal channels to be input into a feature extraction model. Optionally, a channel vector corresponding to each of the multiple sensor signal channels can be generated based on channel data from each of the multiple sensor signal channels, and the channel vectors corresponding to each sensor signal channel can be concatenated to obtain a multidimensional vector corresponding to multiple sensor signal channels, thereby generating an input vector based on the multidimensional vector.
[0056] In one possible implementation, the input vector has dimensions [C, V]. Here, C is the feature dimension of the channel vector of a single sensor signal channel, and V is the number of spatial nodes, i.e., the number of multiple sensor signal channels in the sensor data.
[0057] For example, channel data from multiple sensor signal channels, including data acquired by a 6-axis IMU. , , , , Pressure data collected by pressure sensors For example, we can , , , , , There are 7 independent sensor signal channels, i.e., V=7. Since the channel data of each sensor signal channel is a 1-dimensional sensor measurement, C=1. Therefore, the input vector can be represented as [ , , , , , ].
[0058] In one possible implementation, when extracting features from sensor data within a single time window, the dimension of the input vector can be [C, T, V]. Here, T is the number of data frames corresponding to a single time window, for example, 64 frames.
[0059] In one possible implementation, during the training of the feature extraction model, the dimension of the input vector can be [B, C, T, V]. Here, B is the number of samples processed in a batch during a single training iteration of the feature extraction model.
[0060] Step 203: Perform head action recognition based on the first feature to obtain the head action recognition result.
[0061] Optionally, when head action recognition is achieved through a head action classification model, the feature extraction model and the head action classification model can be trained together.
[0062] In one possible implementation, during the training of the feature extraction model and the head action classification model, sample channel data from multiple sensor signal channels within a single time window can be used as input, and the sample head action category corresponding to that single time window can be used as output to jointly train the feature extraction model and the head action classification model. Thus, during online inference, the first feature can be extracted and head action recognition can be achieved through the trained feature extraction model and head action classification model.
[0063] It is understandable that, in practical application scenarios, the solution proposed in this application can define head movements other than nodding and shaking, based on the user's use of electronic devices, to expand the categories of head movements. Therefore, by retraining the feature extraction model and the head movement classification model based on sample data corresponding to the expanded head movement categories, it is possible to support the detection and recognition of more head interaction movements without increasing the amount of computation, further enriching the intelligent interactive experience of electronic devices such as headphones and mobile phones.
[0064] In this embodiment, when sensor data representing head movement state and including channel data from multiple sensor signal channels are acquired, the channel data from these multiple sensor signal channels can be input into a feature extraction model to obtain a first feature output by the model. The feature extraction model includes a spatial graph convolutional neural network (SBR) and a temporal graph convolutional neural network (TBR). The SBR is used to determine the data synchronization of different sensor signal channels in the sensor data, and the TBR is used to determine the data temporality of different sensor signal channels in the sensor data. Then, head movement recognition is performed based on this first feature to obtain the head movement recognition result. Thus, by introducing a spatiotemporal graph convolutional network to deeply mine the data correlation and temporality of different sensor signal channels in the sensor data, high-quality feature mining of channel data from multiple sensor signal channels can be achieved for head movement recognition, improving the accuracy of head movement recognition.
[0065] The order of spatial graph convolutional neural networks and temporal graph convolutional neural networks It is understood that the embodiments of this application do not limit the connection relationship between spatial graph convolutional neural networks and temporal graph convolutional neural networks in the feature extraction model.
[0066] In some embodiments, the output of a spatial graph convolutional neural network can be connected to the input of a temporal graph convolutional neural network. Thus, after receiving channel data from multiple sensor signal channels, the feature extraction model can input this channel data into the spatial graph convolutional neural network to obtain its output. Then, the output of the spatial graph convolutional neural network is input into the temporal graph convolutional neural network to obtain its output. This output is the first feature, which integrates the data synchronization and temporal sequence of the various sensor signal channels.
[0067] In one possible implementation, please refer to Figure 3 The diagram illustrates a flowchart of a head motion recognition method provided in yet another exemplary embodiment of this application. The method may include the following steps: Step 301: Acquire sensor data. The sensor data is used to characterize the head movement state and includes channel data of multiple sensor signal channels.
[0068] In some embodiments, sensor data collected by multiple sensors can be acquired. The sensor data includes channel data of multiple sensor signal channels corresponding to multiple sensors, and each of the multiple sensors corresponds to at least one sensor signal channel.
[0069] For example, the multiple sensors could be a 6-axis IMU in the left earphone, a pressure sensor in the left earphone, and a 6-axis IMU in the right earphone. The multiple sensor signal channels corresponding to these multiple sensors include six sensor signal channels of the 6-axis IMU in the left earphone, one sensor signal channel of the pressure sensor in the left earphone, and six sensor signal channels of the 6-axis IMU in the right earphone.
[0070] In some embodiments, after obtaining sensor data collected by multiple sensors, the sensor data can be preprocessed to obtain sensor data that meets the requirements for subsequent feature extraction. The data preprocessing may include, but is not limited to, at least one of the following: timestamp alignment, data cleaning (outlier removal), and data interpolation (missing value imputation).
[0071] In some embodiments, data preprocessing may further include noise reduction. Optionally, when the sensor data includes data acquired by an IMU, the IMU-acquired data may be denoised to reduce noise and bias in the IMU-acquired data. In one possible implementation, high-frequency noise removal processing of the IMU-acquired data may be performed using a low-pass digital filter.
[0072] In some embodiments, data preprocessing may further include sliding window sampling to extract sensor data within each time window from sensor data collected over a long period of time by sliding a fixed-size time window, so as to achieve feature extraction and head motion recognition of sensor data within a local time window.
[0073] In some embodiments, data preprocessing may further include channel normalization.
[0074] It is understandable that the channel data of different sensor signal channels in different modes have significant numerical differences, even substantial differences in magnitude. These differences can interfere with subsequent feature extraction from the channel data of multiple sensor signal channels. Therefore, we can first normalize the channel data of each sensor signal channel in multiple sensors to obtain normalized channel data for multiple sensor signal channels. Then, based on the data synchronization and temporal sequence of different sensor signal channels in the sensor data, we can extract features from the normalized channel data of multiple sensor signal channels to obtain the first feature.
[0075] Optionally, for any sensor signal channel, a normalization function can be used to normalize the channel data of that sensor signal channel. Optionally, the value range of the normalized channel data of any sensor signal channel is [0, 1].
[0076] Thus, with multimodal sensor data input, normalization of each sensor signal channel can effectively avoid interference caused by excessive numerical differences between different sensor signal channels of different modes, thereby improving the accuracy of feature extraction.
[0077] In one possible implementation, the multiple sensors may include a pressure sensor and a 6-axis inertial sensor IMU. In this case, the multiple sensor signal channels in the sensor data may include a single sensor signal channel corresponding to the pressure sensor and six sensor signal channels corresponding to the 6-axis inertial sensor IMU.
[0078] In this way, by fusing multimodal sensor data such as 6-axis IMU and pressure sensor to achieve head action recognition, more comprehensive information can be obtained. By utilizing multi-dimensional features to recognize head actions, the occurrence of false recognition scenarios can be effectively suppressed, greatly improving the accuracy and robustness of head action recognition. It is suitable for various daily use scenarios and avoids the problems of high false recognition rate and difficulty in adapting to complex dynamic scenarios caused by single-modal sensor data.
[0079] Step 302: Input the channel data of multiple sensor signal channels into the spatial graph convolutional neural network in the feature extraction model to obtain the intermediate features output by the spatial graph convolutional neural network. The intermediate features are used to characterize the data synchronization of different sensor signal channels in the sensor data.
[0080] Optionally, after the spatial graph convolutional neural network treats the sensor signal channels as "spatial nodes" and the data synchronization between sensor signal channels as "spatial correlation between spatial nodes", it can extract correlation features between sensor signal channels from the channel data of multiple sensor signal channels based on the data synchronization between sensor signal channels, i.e., the spatial correlation between spatial nodes. These correlation features are intermediate features used to characterize the channel synchronization correlation characteristics of different sensor signal channels in the sensor data.
[0081] In one possible implementation, in a spatial graph convolutional neural network, for each spatial node, i.e., each sensor signal channel, information can be integrated with that of neighboring nodes to extract the correlation features between spatial nodes. Optionally, neighboring nodes can be sensor signal channels that have data synchronization with themselves, or sensor signal channels that have relatively strong data synchronization with themselves; this is not limited here.
[0082] In some embodiments, channel data from multiple sensor signal channels corresponding to each time point can be input into a spatial graph convolutional neural network in a feature extraction model to obtain intermediate features corresponding to each time point output by the spatial graph convolutional neural network.
[0083] In some embodiments, the channel data of multiple sensor signal channels corresponding to each time point within a single time window can be simultaneously input into the spatial graph convolutional neural network in the feature extraction model, so that the channel data of multiple sensor signal channels corresponding to each time point can be processed by the spatial graph convolutional neural network to obtain the intermediate features corresponding to each time point within a single time window output by the spatial graph convolutional neural network.
[0084] Step 303: Input the intermediate features into the time-map convolutional neural network in the feature extraction model to obtain the first feature output by the time-map convolutional neural network. The first feature is used to characterize the data temporal sequence of different sensor signal channels in the intermediate features.
[0085] Optionally, in a temporal graph convolutional neural network, the temporal features of each sensor signal channel can be extracted from the intermediate features output by the spatial graph convolutional neural network through a sliding convolution operation in the temporal dimension. This temporal feature is the first feature, which reflects the trend of the channel data of each sensor signal channel in the intermediate features changing over time.
[0086] It is understandable that, since the intermediate features contain the data synchronization of different sensor signal channels, i.e., the channel synchronization correlation characteristics, further time-series feature extraction is performed on the intermediate features so that the final first feature integrates data synchronization, i.e., the channel synchronization correlation characteristics, as well as data temporality, i.e., the dynamic characteristics of channel data changing over time.
[0087] In one possible implementation, the intermediate feature may include the association features of strongly synchronized channel pairs, and correspondingly, the first feature may include the temporal co-variations of the channel data of strongly synchronized channel pairs. Optionally, the intermediate feature may also include the association features of weakly synchronized channel pairs, and correspondingly, the first feature may include the temporal co-variations of the channel data of weakly synchronized channel pairs. Optionally, the intermediate feature may also exclude the association features of asynchronous channel pairs, and correspondingly, the first feature may also exclude the temporal co-variations of the channel data of asynchronous channel pairs.
[0088] In some embodiments, during the training of the feature extraction model, sample data corresponding to different head movement scenarios can be used to jointly train the spatial graph convolutional neural network and the temporal graph convolutional neural network in the feature extraction model. This allows the spatial graph convolutional neural network and the temporal graph convolutional neural network in the feature extraction model to deeply mine the synchronous correlation characteristics between channels and the temporal dynamic characteristics of channel data, and learn high-quality features that stably distinguish each head movement category. In this way, the advantages of channel data fusion of multimodal and multi-sensor signal channels are fully utilized, improving the accuracy and robustness of head movement recognition, adapting to various complex dynamic scenarios, and demonstrating high practicality.
[0089] Step 304: Perform head motion recognition based on the first feature to obtain the head motion recognition result.
[0090] Optionally, after processing with a spatial graph convolutional neural network and a temporal graph convolutional neural network sequentially, a first feature that integrates channel synchronization correlation characteristics and temporal dynamic characteristics can be obtained. At this point, pooling can be applied to the first feature to compress its feature dimension. In one possible implementation, when compressing the feature dimension of the first feature, pooling can be performed on both the temporal and spatial dimensions.
[0091] Optionally, after compressing the feature dimension of the first feature, the dimensionality-reduced first feature can be input into a head action classifier / classification layer consisting of a fully connected layer and a Softmax activation function to obtain the probability of each head action category output by the classifier / classification layer.
[0092] In this embodiment, when sensor data representing head movement states and including channel data from multiple sensor signal channels are acquired, the channel data from these multiple sensor signal channels can first be input into a spatial graph convolutional neural network in the feature extraction model to obtain intermediate features output by the spatial graph convolutional neural network. These intermediate features characterize the data synchronization of different sensor signal channels in the sensor data. Then, the intermediate features are input into a temporal graph convolutional neural network in the feature extraction model to obtain a first feature output by the temporal graph convolutional neural network. This first feature characterizes the temporal sequence of data from different sensor signal channels in the intermediate features. Head movement recognition is then performed based on this first feature to obtain the head movement recognition result. Thus, by sequentially performing feature extraction processing using spatial graph convolutional neural networks and temporal graph convolutional neural networks, after deeply mining the synchronous correlation characteristics of sensor signal channels in the sensor data, the temporal coordinated changes of synchronous channel pairs can be further explored, thereby obtaining high-quality features that can stably distinguish head movement categories, achieving high-precision and highly robust head movement recognition.
[0093] In other embodiments, the output of the temporal graph convolutional neural network can be connected to the input of the spatial graph convolutional neural network. Thus, after receiving channel data from multiple sensor signal channels, the feature extraction model can first input the channel data from these multiple sensor signal channels into the temporal graph convolutional neural network to obtain its output. Then, the output of the temporal graph convolutional neural network is input into the spatial graph convolutional neural network to obtain its output. This output can be a second feature that integrates the data synchronization and temporal sequence of each sensor signal channel.
[0094] Optionally, although both the second feature and the first feature integrate the data synchronization and data timing of each sensor signal channel in multiple sensor signal channels, they are actually different features.
[0095] In one possible implementation, channel data from multiple sensor signal channels are input into a time-graph convolutional neural network in a feature extraction model to obtain a temporal feature output by the time-graph convolutional neural network. This temporal feature characterizes the temporal sequence of data from different sensor signal channels in the sensor data, reflecting the trend of channel data changes over time. Then, this temporal feature is input into a spatial graph convolutional neural network in the feature extraction model to obtain a second feature output by the spatial graph convolutional neural network. This second feature represents the data synchronization of different sensor signal channels in the temporal feature, i.e., the channel synchronization correlation characteristic.
[0096] Optionally, after inputting the temporal feature into the spatial graph convolutional neural network in the feature extraction model, the channel features of the multi-sensor signal channels at different time points in the temporal feature can be processed in the spatial graph convolutional neural network to obtain the correlation features between sensor signal channels at different time points in the temporal feature output by the spatial graph convolutional neural network. This correlation feature is the second feature, which reflects the channel synchronous correlation characteristics at different time points in the temporal feature.
[0097] In one possible implementation, the timing feature may include the timing variation of the channel data of each sensor signal channel. Correspondingly, the second feature may include the timing variation of the channel data of strongly synchronized channel pairs in the timing feature, or it may include the timing variation of the channel data of weakly synchronized channel pairs.
[0098] Thus, by performing feature extraction processing on temporal graph convolutional neural networks and spatial graph convolutional neural networks sequentially, after mining the temporal changes of each sensor signal channel in the sensor data, we can continue to mine the temporal coordinated changes of strongly synchronized channel pairs, thereby obtaining high-quality features that can stably distinguish head movement categories, and achieving high-precision and high-robust head movement recognition.
[0099] Feature extraction from spatial graph convolutional neural networks In some embodiments, an adjacency matrix can be introduced to characterize the data synchronization between sensor signal channels, that is, the spatial association between spatial nodes is characterized by an adjacency matrix.
[0100] It is understood that the adjacency matrix in a graph network structure typically represents the adjacency relationship between nodes, and it can aggregate the features of adjacent nodes. However, in this embodiment of the application, the adjacency matrix in ST-GCN represents the data synchronization between sensor signal channels. The stronger the data synchronization between two sensor signal channels, the larger the value in the adjacency matrix, which can aggregate the features of strongly synchronized channel pairs.
[0101] In one possible implementation, an adjacency matrix of multiple sensor signal channels in the sensor data can be determined first. This adjacency matrix is used to characterize the data synchronization degree between any two sensor signal channels. Optionally, this data synchronization degree can refer to the degree / intensity of synchronization change in the channel data, such as strong synchronization or weak synchronization.
[0102] The dimension of the adjacency matrix is related to the number of sensor signal channels. For example, when the sensor data includes 7 sensor signal channels, the dimension of the adjacency matrix is 7*7. The matrix value at coordinate [i, j] in the adjacency matrix represents the degree of synchronization between the channel data of sensor signal channel i and sensor signal channel j, which can reflect the strength of the synchronous changes between the channel data of the sensor signal channels.
[0103] For example, a matrix value of 0.2 at coordinates [2, 3] indicates that the synchronization degree of the channel data of sensor signal channel 2 and sensor signal channel 3 is not high, and sensor signal channel 2 and sensor signal channel 3 can form a weak synchronization channel pair. As another example, a matrix value of 0.7 at coordinates [2, 5] indicates that the synchronization degree of the channel data of sensor signal channel 2 and sensor signal channel 5 is high, and sensor signal channel 2 and sensor signal channel 5 can form a strong synchronization channel pair.
[0104] In some embodiments, after obtaining the adjacency matrix of multiple sensor signal channels in the sensor data, the adjacency matrix and the channel data of the multiple sensor signal channels can be input into the spatial graph convolutional neural network in the feature extraction model to obtain the intermediate features output by the spatial graph convolutional neural network. The core objective of the spatial graph convolutional neural network can be to perform "neighbor aggregation" on the channel data of multiple sensor signal channels based on the adjacency matrix to extract the correlation features between sensor signal channels. These correlation features are used to characterize the channel synchronization correlation characteristics of different sensor signal channels.
[0105] Optionally, in a spatial graph convolutional neural network, the adjacency matrix and the channel data of multiple sensor signal channels can be multiplied and summed to aggregate the features of strongly synchronized channel pairs and obtain aggregated intermediate features.
[0106] In one possible implementation, when extracting features from sensor data across T frames within a single time window, if the sensor data includes channel data from V sensor signal channels, the dimension of the input vector of the spatial graph convolutional neural network can be [V, T]. In the spatial graph convolutional neural network, for the input vector [V, T], the channel data of the V sensor signal channels sampled at each time step (i.e., each time point) along the T dimension can be extracted sequentially (dimension [V, 1]), and the channel data of the V sensor signal channels in dimension [V, V] can be multiplied by the adjacency matrix of dimension [V, V]. This is equivalent to performing a weighted summation through a weighted adjacency matrix to aggregate the features between strongly synchronized channels and obtain intermediate features, which can be feature vectors in dimension [V, 1].
[0107] In this way, by encoding the synchronous changes between various sensor signal channels during various head movements into an adjacency matrix of a spatial graph convolutional network (GCN), the correlation features between different sensor signal channels can be accurately extracted based on this adjacency matrix. This results in the final extracted features aggregating the features of strongly correlated, i.e., strongly synchronously changing sensor signal channels, thereby improving the accuracy of head movement recognition.
[0108] In some embodiments, the adjacency matrix may be determined in real time during online head motion recognition.
[0109] Optionally, the adjacency matrix can be generated in real time using a neural network model. The input to this neural network model can be channel data from multiple sensor signal channels in the sensor data, and the output is the adjacency matrix corresponding to those multiple sensor signal channels. This neural network model is trained jointly with a feature extraction model. Optionally, when head action recognition is achieved through a head action classification model, this neural network model is trained together with both the feature extraction model and the head action classification model. Therefore, based on the currently acquired sensor data, an adjacency matrix suitable for the current sensor data can be generated in real time, improving the accuracy of the adjacency matrix.
[0110] In one possible implementation, the neural network model can be a multilayer perceptron (MLP). Channel data from multiple sensor signal channels in the sensor data can be input into the MLP, thereby obtaining the adjacency matrix of the multiple sensor signal channels in the sensor data output by the MLP. The MLP is trained jointly with the feature extraction model. Optionally, when head action recognition is achieved through a head action classification model, the MLP is trained together with both the feature extraction model and the head action classification model.
[0111] Optionally, a very lightweight MLP can be added for joint training during the joint training of the feature extraction model and the head action classification model. The input to this MLP can be a set of basic statistical features for each sensor signal channel within the current time window, and the output of the MLP can be the adjacency matrix corresponding to the current time window. This output adjacency matrix is then input into the spatial graph convolutional neural network in the feature extraction model to participate in subsequent neighbor aggregation calculations. Optionally, the basic statistical features can include, but are not limited to, at least one of the following statistical features: mean, average, time-domain features, frequency-domain features, correlation features, etc., without limitation here.
[0112] In this way, during online inference, the real-time adjacency matrix corresponding to the sensor data within the current time window can be quickly calculated using the trained ultra-lightweight MLP. This improves the accuracy of the adjacency matrix without significantly increasing the computational load, thereby enhancing the accuracy of feature mining in the spatiotemporal graph convolutional network and achieving high-precision and robust head action recognition.
[0113] In some embodiments, the adjacency matrix may also be predetermined.
[0114] In one possible implementation, a set of coefficients accurately representing the strength of synchronous changes can be generated based on the synchronous changes between sensor signal channels during each head movement and empirical values obtained from data analysis. This allows for the pre-construction of a fixed adjacency matrix, avoiding the need for real-time calculation of synchronous change coefficients during online inference, which would otherwise require more computational power. In other words, the values of each matrix in the adjacency matrix can be set based on empirical values, resulting in an empirical adjacency matrix.
[0115] Let's illustrate the process of generating an empirical adjacency matrix with an example. We can analyze several sets of typical head movement sensor data collected in practice. For instance, we can analyze the channel data of multiple sensor signal channels in the sensor data through various methods such as visualization comparison, data correlation statistics, and traversal. If we find that sensor signal channel 1 and sensor signal channel 2 are phase-synchronized under head movement A, with a high synchronization level (i.e., a strong synchronization channel pair), then based on experience, we can assign a value to the matrix at coordinate [1, 2] in the adjacency matrix, for example, 0.8. For sensor signal channels 1 and 3 that do not change synchronously, with a low / no synchronization level (i.e., an asynchronous channel pair), we can assign a value of 0.1 to the matrix at coordinate [1, 3].
[0116] Optionally, for asynchronous channel pairs with low / no synchronization levels in the adjacency matrix, the corresponding matrix value can be set to non-zero to preserve the association characteristics of weak synchronization channel pairs and avoid information isolation.
[0117] In some embodiments, after generating the empirical adjacency matrix, it can be used as an initial value to define learnable adjacency matrix parameters, and corresponding physical constraints can be added to ensure the rationality of the synchronous change relationship between sensor signal channels. This learnable adjacency matrix is then added to the feature extraction calculation in the subsequent spatial graph convolutional neural network. During the training of the feature extraction model, or during the joint training of the feature extraction model and the head action classification model, the parameters of the adjacency matrix are learned and optimized. Thus, the adjacency matrix corresponding to the trained feature extraction model can be determined as the adjacency matrix for the inference process.
[0118] In this way, by initializing with empirical values and defining a learnable adjacency matrix, and by optimizing the adjacency matrix through training, we can avoid repeatedly calculating the adjacency matrix online, reduce power consumption, and accurately characterize the feature correlation of synchronous changes between sensor signal channels, thereby supporting the spatiotemporal graph convolutional neural network to mine stable features with high discriminative power.
[0119] In one possible implementation, the added physical constraints may include, but are not limited to, at least one of the following: nonnegativity, symmetry, and self-loop = 1. Nonnegativity is used to constrain the strength of synchronous changes in channel data between two sensor signal channels to be non-negative, avoiding the physical contradiction of "negative correlation." Symmetry is used to constrain the strength of synchronous changes in sensor signal channel i relative to sensor signal channel j to be consistent with the strength of synchronous changes in sensor signal channel j relative to sensor signal channel i, to conform to the bidirectional nature of the synchronous change relationship. Self-loop = 1 is used to ensure that each sensor signal channel retains its own characteristics during neighbor aggregation and is not completely covered by its neighbors.
[0120] In this embodiment, by introducing an adjacency matrix to characterize the data synchronization between sensor signal channels, the synchronization correlation features of sensor signal channels in sensor data can be extracted more accurately, which helps to improve the accuracy and robustness of subsequent head motion recognition.
[0121] Feature extraction from time-plot convolutional neural networks In some embodiments, when intermediate features are obtained from the output of a spatial graph convolutional neural network, these intermediate features can be first subjected to dimensionality upscaling to obtain high-dimensional features, thereby enhancing the expressive power of the feature vector. Then, the high-dimensional features obtained after dimensionality upscaling are input into a temporal graph convolutional neural network to obtain the first feature output by the temporal graph convolutional neural network. After dimensionality upscaling, the number of feature dimensions in the high-dimensional features is greater than the number of feature dimensions in the intermediate features.
[0122] Optionally, dimensionality enhancement of intermediate features can be achieved through convolution. In one possible implementation, a convolution kernel can be introduced to enhance the dimensionality of each feature value in the intermediate features. For example, using a 1*32 convolution kernel, when the intermediate features output by the spatial graph convolutional neural network are feature vectors of dimension [V, 1], a 1*32 convolution kernel can be used to perform convolution enhancement on this feature vector of dimension [V, 1]. After the convolution enhancement, the dimension of the high-dimensional feature becomes [V, 32]. This enhances the expressive power of the intermediate features.
[0123] In some embodiments, when acquiring sensor data corresponding to each time point within a single time window, the spatial graph convolutional neural network can output intermediate features corresponding to each time point within the single time window. In this case, the intermediate features corresponding to each time point within the single time window can be concatenated in chronological order to obtain the concatenated features corresponding to the single time window. These concatenated features are then input into the temporal graph convolutional neural network in the feature extraction model to obtain the first feature output by the temporal graph convolutional neural network. This first feature characterizes the changes in channel data of different sensor signal channels within the single time window. Thus, by limiting the amount of data processed by the feature extraction model through a time window, feature extraction and head movement recognition of sensor data within a short period can be achieved, realizing low-computing-power, high-precision head movement recognition.
[0124] In one possible implementation, the intermediate features can be first upscaled, and then the intermediate features from multiple time points can be concatenated. Optionally, when acquiring sensor data corresponding to each time point within a single time window, the spatial graph convolutional neural network can output intermediate features corresponding to each time point within the single time window. In this case, the intermediate features corresponding to each time point within the single time window can first be upscaled to obtain high-dimensional features corresponding to each time point within the single time window. Then, according to chronological order, the high-dimensional features corresponding to each time point within the single time window are concatenated to obtain concatenated features. Finally, the concatenated features are input into the temporal graph convolutional neural network in the feature extraction model to obtain the first feature output by the temporal graph convolutional neural network. This first feature is used to characterize the changes in channel data of different sensor signal channels within a single time window.
[0125] For example, when extracting features from sensor data across T frames within a single time window, if the T intermediate features corresponding to the current time window output by the spatial graph convolutional neural network have dimensions [V, 1], then these T intermediate features can first be upscaled using a 1*32 convolutional kernel to obtain T high-dimensional features with dimensions [V, 32]. Then, along the time dimension, the T high-dimensional features within the current time window are concatenated to obtain concatenated features, which are then input into the temporal graph convolutional neural network.
[0126] In this way, by increasing the dimensionality of intermediate features, the expressive power of intermediate features can be enhanced, thereby enabling the feature extraction model to better learn high-quality features for distinguishing various head movements and improving the accuracy of head movement recognition results.
[0127] In some embodiments, the above-described dimensionality enhancement and feature concatenation are deployed in a spatial graph convolutional neural network, so that the output of the spatial graph convolutional neural network can be the concatenated features. For example, as... Figure 4 As shown, a spatial graph convolutional neural network can include several execution operators, such as strongly synchronized channel feature aggregation, convolutional dimensionality increase, and feature concatenation. Strongly synchronized channel feature aggregation refers to multiplying the adjacency matrix with channel data from multiple sensor signal channels to obtain intermediate features that aggregate the features between strongly synchronized channels. Convolutional dimensionality increase refers to increasing the dimensionality of the intermediate features to higher dimensions to enhance feature expressive power. Feature concatenation refers to concatenating the high-dimensional features to obtain concatenated features. These concatenated features serve as the feature output of the spatial graph convolutional neural network and also as the feature input of the temporal graph convolutional neural network.
[0128] In some embodiments, the dimensionality upscaling and feature concatenation described above can be deployed independently after the spatial graph convolutional neural network, or they can be deployed within the temporal graph convolutional neural network. No limitation is imposed here.
[0129] In some embodiments, a temporal graph convolutional neural network can capture local temporal changes in the time dimension through temporal convolution processing. The core objective of this temporal graph convolutional neural network may be to capture the changing trends of channel characteristics of sensor signal channels over time, such as the temporal coordinated changes of strongly synchronized channel pairs, by sliding convolution in the time dimension.
[0130] Optionally, after concatenating the intermediate features (or high-dimensional features) corresponding to each time point within a single time window to obtain concatenated features, local window features can be extracted from the concatenated features using a convolutional sliding window. These local window features are then subjected to temporal convolution processing to obtain a first feature. This first feature characterizes the changes in channel data of different sensor signal channels within the convolutional sliding window. The convolutional sliding window is used to slide along the time dimension.
[0131] In one possible implementation, temporal convolution processing can be a one-dimensional temporal convolution operation. This one-dimensional temporal convolution operation can be understood as sliding a fixed-size convolutional window through a convolutional kernel along the temporal dimension of the concatenated features. This extracts local window features from the convolutional sliding window, performs weighted summation of these features, and thus captures the local temporal dynamics of the concatenated features.
[0132] For example, when the spliced feature is obtained by splicing T high-dimensional features along the time dimension, the convolution kernel for temporal convolution processing can be a convolution kernel with a time length of K. This allows for the extraction of k high-dimensional features from the spliced feature to extract local temporal dynamic features, thereby obtaining the first feature.
[0133] Considering that during the training of the feature extraction model, as the parameters of the spatial convolutional neural network are updated, the distribution of input features in the temporal convolutional neural network will continuously change, i.e., "covariate shift," and that temporal convolution processes sliding window data in the time dimension, the feature distribution may further fluctuate due to the dynamic nature of time-series data (such as sudden signal changes in a frame), potentially leading to training instability. For example, excessively high feature values may cause the Rectified Linear Unit (ReLU) activation function to enter the "saturation region," with gradients approaching zero (vanishing), resulting in training stagnation. Alternatively, drastic fluctuations in the input feature distribution may require the temporal convolutional neural network to continuously adjust its parameters to adapt to the new distribution, leading to slow training speed, difficulty in convergence, or even oscillations. Therefore, after completing the first feature extraction, the first feature is batch normalized to ensure stable training. Finally, the normalized first feature is input into the ReLU activation function to introduce nonlinearity and further enhance the feature representation capability, resulting in the nonlinearly transformed first feature output by the ReLU activation function.
[0134] For example, such as Figure 5 As shown, a temporal graph convolutional neural network can include several execution operators such as sequentially connected one-dimensional temporal convolution, batch normalization, and the ReLU activation function. The output of the ReLU activation function serves as the feature output of the temporal graph convolutional neural network.
[0135] In some embodiments, after obtaining the first feature after the nonlinear transformation of the ReLU activation function output, pooling can be applied to the first feature in both the time and spatial dimensions to compress its feature dimension. Finally, the dimensionality-reduced first feature is fed into a classifier consisting of a fully connected layer and a Softmax activation function to obtain the probability of each head action category output by the classifier.
[0136] Thus, the solution provided in this application can first aggregate features between strongly synchronized channels through an adjacency matrix in a spatial graph convolutional neural network, and then introduce convolutional dimensionality enhancement processing to further enhance feature representation capabilities. Next, in a temporal graph convolutional neural network, local temporal dynamic features are extracted through one-dimensional temporal convolution in the time dimension, and batch normalization is used to improve training stability. Finally, pooling is used to compress the feature dimension, providing the classifier with the predicted probabilities for each target category. Therefore, through the complementary combination of "explicitly capturing the synchronization characteristics between sensor signal channels using a spatial graph convolutional neural network" and "capturing the short-term temporal characteristics of sensor signal channels using a temporal convolutional neural network," a lightweight head motion recognition method with high real-time performance and high accuracy can be achieved.
[0137] In this embodiment, by sliding the convolutional sliding window in the time dimension to extract intermediate features (or upgraded intermediate features) within the convolutional sliding window and performing temporal convolution processing, the temporal changes of sensor signal channels in sensor data can be extracted more accurately. Specifically, this can include the temporal coordinated changes of strongly synchronized channel pairs. This enables in-depth mining of the correlation features between different sensor signal channels and the local temporal dynamic characteristics of the channel data of sensor signal channels along the time dimension, resulting in the extraction of more accurate and stable high-quality features, thereby improving the accuracy of subsequent head action recognition results.
[0138] Methods for determining head motion recognition results In some embodiments, when head action recognition is performed based on the first feature to obtain the probability of each head action category, the head action category with the highest probability among the probabilities of each head action category can be output as the head action recognition result.
[0139] In some embodiments, the current head action recognition result determined based on the highest probability can be corrected based on historical head action recognition results. This ensures that the current head action recognition result follows a continuous and smooth trend compared to historical head action recognition results, avoiding drastic fluctuations in the action recognition result.
[0140] Optionally, the reasonableness of the current head motion recognition result is then verified based on historical head motion recognition results. If the current head motion recognition result is found to be unreasonable, it is corrected based on historical head motion recognition results. If the current head motion recognition result is found to be reasonable, no correction is made, and the current head motion recognition result can be saved.
[0141] In some embodiments, when head action recognition is performed based on the first feature to obtain the probability of each head action category, the historical probability of each head action category can also be obtained by combining historical data to determine the head action recognition result.
[0142] Optionally, given the probabilities of each head action category in the current output, these probabilities can be cached. Optionally, historical probabilities of each head action category output within a historical time period can be cached, as can the corresponding head action category labels. In one possible implementation, head action category labels within a historical time period and their corresponding confidence windows can also be cached. This confidence window can be a probability range.
[0143] The historical time period can be reasonably set according to the actual situation, and there is no limitation here. For example, it can be the time period corresponding to the first 3 time windows.
[0144] In some embodiments, the historical probabilities of each head action category obtained within a historical time period and the current probabilities of each head action category can be fused to obtain the current fused probabilities of each head action category. This fused probability can also be called the fused confidence level. The current head action recognition result is determined based on the current fused probabilities of each head action category.
[0145] Optionally, a weighted average method (with lower weights for values further back from the current time) can be used to fuse the historical probabilities of each head action category obtained within the historical time period and the current probabilities of each head action category to calculate the fused probability of each head action category at the current time. The head action recognition result can then be determined based on the fused probability.
[0146] In one possible implementation, the head action category with the highest fusion probability among the current head action categories can be output as the current head action recognition result.
[0147] In this way, by fusing the current probabilities of each head action category with the probabilities output over a period of time, we can avoid large anomalies in the current predicted probabilities and effectively suppress the occurrence of misidentification scenarios.
[0148] In one possible implementation, the current head motion recognition result is corrected based on historical head motion recognition results. This ensures that the current head motion recognition result follows a continuous and smooth trend compared to historical results, avoiding drastic fluctuations that could lead to low accuracy and robustness in head motion recognition.
[0149] Optionally, the rationality of the current head action recognition result can be verified based on historical head action recognition results. If the current head action recognition result is found to be unreasonable, it can be corrected based on historical head action recognition results. If the current head action recognition result is found to be reasonable, the head action recognition result determined based on the highest fusion probability is confirmed, no correction is made, and the current head action recognition result can be maintained.
[0150] Optionally, verifying the rationality of the current head action recognition result can be done by verifying the rationality of the user performing the current head action under the head action corresponding to the historical head action recognition result.
[0151] For example, if the head action recognition result for the first three outputs is nodding, and the head action recognition result for the current output is shaking, then the current output head action recognition result is incorrect and needs to be corrected, because according to the interaction habits of head actions, one would not intentionally nod and then immediately shake one's head.
[0152] In this way, by verifying the rationality of the user's current head action under the head action corresponding to the historical head action recognition results, the recognition results of some head actions that violate normal interaction habits can be automatically corrected in a timely manner, effectively suppressing the occurrence of misidentification scenarios, improving the accuracy of head action recognition, and adapting to various daily use scenarios.
[0153] In one possible application scenario, the output head motion recognition results can be adjusted based on the usage scenario. For example, in a certain usage scenario, the interval between two consecutive nodding actions by a user is at least 0.5 seconds, so the output head motion recognition results can be constrained by adding corresponding time interval rules.
[0154] In some embodiments, after determining the head motion recognition result, an interactive operation corresponding to the head motion recognition result can be performed.
[0155] In one possible implementation, if the head motion recognition result is a nodding motion, the interactive operation corresponding to the nodding motion can be executed, such as the interactive operation corresponding to an "agree" or "confirm" type interactive command; if the head motion recognition result is a head shaking motion, the interactive operation corresponding to the head shaking motion can be executed, such as the interactive operation corresponding to a "deny" or "cancel" type interactive command.
[0156] In one possible application scenario, where it is inconvenient for users to use their hands or do not want to pick up their phones to operate them, users can perform simple nodding (or shaking) actions after wearing headphones. This allows the headphones and / or phones to use the solution provided in this application to recognize the user's nodding (or shaking) action with high real-time performance and high accuracy, thereby quickly triggering the headphones / phone to perform high-frequency interactive operations such as answering (or hanging up) incoming calls, playing music and switching tracks (pausing playback).
[0157] Please see Figure 6 The diagram illustrates a system flowchart of a head motion recognition system provided in an exemplary embodiment of this application. The system may include: an input module 601, a preprocessing module 602, a spatiotemporal graph convolution module 603, a correction module 604, and an output module 605.
[0158] The input module 601 may include at least one sensor for collecting sensor data and inputting the collected sensor data into the preprocessing module 602. Figure 6 The example given is of input module 601, which includes a pressure sensor and an inertial sensor (IMU), but this does not constitute a limitation.
[0159] Optionally, the input module 601 can input the channel data of each sensor signal channel collected by different sensors into the preprocessing module 602. For example, the input module 601 can input the channel data of one sensor signal channel collected by a pressure sensor. Input to preprocessing module 602. It can also process channel data from the six sensor signal channels acquired by the inertial sensor IMU. , , , , Input preprocessing module 602.
[0160] The preprocessing module 602 may include multiple preprocessing units for preprocessing the channel data of each mode and each sensor signal channel input by the input module 601, and inputting the preprocessed channel data of each mode and each sensor signal channel into the spatiotemporal graph convolution module 603. Figure 6 The explanation uses the preprocessing module 602, which includes noise reduction, sliding window sampling, and channel normalization, as an example, but this does not constitute a limitation.
[0161] Optionally, the preprocessing performed on the channel data of different sensor signal channels can be different or the same. For example, Figure 6 As shown, the preprocessing module 602 can perform noise reduction processing on the channel data of the six sensor signal channels acquired by the inertial sensor IMU, but it cannot process the channel data of the one sensor signal channel acquired by the pressure sensor. Noise reduction is not required.
[0162] The spatiotemporal graph convolution module 603 can be used to extract features from the preprocessed channel data of each modality and each sensor signal channel input by the preprocessing module 602, so as to extract a first feature that integrates the data synchronization and data temporal sequence of different sensor signal channels. Then, based on the first feature, head action recognition is performed to obtain the predicted probability of each head action category. Then, the predicted probability of each head action category is input into the correction module 604.
[0163] Optionally, such as Figure 6 As shown, the spatiotemporal graph convolution module 603 may include units for spatial graph convolution, temporal graph convolution, feature compression, and classification prediction. The spatial graph convolution unit may include a spatial graph convolutional neural network, used to perform "neighbor aggregation" on the channel features of multiple sensor signal channels at each time step based on the synchronous change relationship (i.e., weighted adjacency matrix) between sensor signal channels, extracting intermediate features. These intermediate features can be understood as the correlation features between sensor signal channels. The temporal graph convolution unit may include a temporal graph convolutional neural network used to slide a convolutional sliding window in the time dimension to capture the changing trends of the channel features of each sensor signal channel within the convolutional sliding window over time, such as the temporal coordinated changes of strongly synchronized channel pairs, ultimately extracting the first feature. The feature compression unit can be used to perform pooling processing on the first feature to reduce its feature dimensionality. Classification prediction can be used to predict the probability of each head action category based on the dimensionality-reduced first feature.
[0164] The correction module 604 caches the probabilities of each head action category input by the spatiotemporal graph convolution module 603 over a period of time. This allows it to determine the fusion probability of each head action category at the current moment based on the historical probabilities of each head action category over a historical time period, and to determine the initial head action category label based on this fusion probability. Then, the initial head action category label at the current moment is validated according to the final output head action category labels over a historical time period. If the validation passes, the initial head action category label at the current moment is used as the final output head action category label and input into the output module 605. If the validation fails, the initial head action category label at the current moment is corrected using the historically output head action category labels, and the corrected head action category label is used as the final output head action category label and input into the output module 605.
[0165] The output module 605 is used to receive the determined head action category label input by the correction module 604 and output the finally determined head action category label, for example, by sending the head action category label to the target device.
[0166] In some embodiments, the system may further include a first encoding module 606, which is used to encode the synchronous change relationships between the various sensor signal channels into an adjacency matrix. Optionally, the first encoding module 606 may be used to generate the corresponding adjacency matrix in real time based on the current channel data of the multiple sensor signal channels.
[0167] like Figure 6 As shown, the first encoding module 606 may include multiple units such as data analysis and a multilayer perceptron. Data analysis can be used to obtain preprocessed channel data of each sensor signal channel for each modality from the preprocessing module 602, and to analyze the basic statistical characteristics of the channel data of each sensor signal channel within the current time window. The multilayer perceptron is used to map and output the corresponding adjacency matrix based on the basic statistical characteristics of each sensor signal channel within the current time window. The first encoding module 606 and the spatiotemporal graph convolution module 603 can be trained together.
[0168] In some embodiments, the system may further include a second encoding module 607, which can also be used to encode the synchronous change relationships between the sensor signal channels into an adjacency matrix. Optionally, the second encoding module 607 is used to initialize the adjacency matrix based on empirical assignment, and then optimize and update the adjacency matrix during the training process of the spatiotemporal graph convolution module 603. Thus, after the spatiotemporal graph convolution module 603 is trained, the corresponding optimized and updated adjacency matrix can be saved so that it can be added to the spatiotemporal graph convolution module 603 for feature extraction calculation during the inference process.
[0169] In one possible application scenario, during the training process of the head motion recognition system, the head motion recognition system may include an input module 601, a preprocessing module 602, a spatiotemporal graph convolution module 603, a correction module 604, an output module 605, and a first encoding module 606; or, the head motion recognition system may include an input module 601, a preprocessing module 602, a spatiotemporal graph convolution module 603, a correction module 604, an output module 605, and a second encoding module 607. Optionally, each module in the head motion recognition system may be deployed on the aforementioned earphone, allowing the earphone to independently implement the training process of the head motion recognition system; or all modules may be deployed on the aforementioned mobile phone, allowing the mobile phone to independently implement the training process of the head motion recognition system; or some modules may be deployed on the earphone and some on the mobile phone, allowing the training process of the head motion recognition system to be achieved through data interaction between the mobile phone and the earphone. This application does not limit the subsystems deployed on the mobile phone or the subsystems deployed on the earphone.
[0170] In one possible application scenario, during the inference process of the head motion recognition system, the head motion recognition system may include an input module 601, a preprocessing module 602, a spatiotemporal graph convolution module 603, a correction module 604, and an output module 605; or, the head motion recognition system may include an input module 601, a preprocessing module 602, a spatiotemporal graph convolution module 603, a correction module 604, an output module 605, and a first encoding module 606. Optionally, each module in the head motion recognition system may be deployed on the aforementioned earphone, allowing the earphone to independently implement the inference process of the head motion recognition system; or all modules may be deployed on the aforementioned mobile phone, allowing the mobile phone to independently implement the inference process of the head motion recognition system; or some modules may be deployed on the earphone and some on the mobile phone, allowing the inference process of the head motion recognition system to be realized through data interaction between the mobile phone and the earphone. This application does not limit the subsystems deployed on the mobile phone or the subsystems deployed on the earphone.
[0171] It should be noted that the above embodiments can be combined to obtain new embodiments, and the embodiments of this application do not limit the combination method.
[0172] Please refer to Figure 7 This illustration shows a structural block diagram of a head motion recognition device provided in an exemplary embodiment of this application. The device includes: The data acquisition module 701 is used to acquire sensor data, which is used to characterize the head movement state and includes channel data of multiple sensor signal channels. The feature extraction module 702 is used to extract features from the channel data of the multiple sensor signal channels based on the data synchronization and data timing of different sensor signal channels in the sensor data, and obtain a first feature; The action recognition module 703 is used to perform head action recognition based on the first feature to obtain the head action recognition result.
[0173] Optionally, the feature extraction module 702 is used for: The channel data of the multiple sensor signal channels are input into the feature extraction model to obtain the first feature output by the feature extraction model. The feature extraction model includes a spatial graph convolutional neural network and a temporal graph convolutional neural network. The spatial graph convolutional neural network is used to determine the data synchronization of different sensor signal channels in the sensor data, and the temporal graph convolutional neural network is used to determine the data temporality of different sensor signal channels in the sensor data.
[0174] Optionally, the feature extraction module 702 is specifically used for: The channel data of the multiple sensor signal channels are input into the spatial graph convolutional neural network in the feature extraction model to obtain the intermediate features output by the spatial graph convolutional neural network. The intermediate features are used to characterize the data synchronization of different sensor signal channels in the sensor data. The intermediate features are input into the time-map convolutional neural network in the feature extraction model to obtain the first feature output by the time-map convolutional neural network. The first feature is used to characterize the data temporal sequence of different sensor signal channels in the intermediate features.
[0175] Optionally, the device further includes a synchronization determination module for: Determine the adjacency matrix of the plurality of sensor signal channels in the sensor data, the adjacency matrix being used to characterize the data synchronization degree between any two sensor signal channels in the plurality of sensor signal channels.
[0176] Optionally, the feature extraction module 702 is specifically used for: The adjacency matrix and the channel data of the multiple sensor signal channels are input into the spatial graph convolutional neural network in the feature extraction model to obtain the intermediate features output by the spatial graph convolutional neural network.
[0177] Optionally, the synchronization determination module is specifically used for: The channel data of the multiple sensor signal channels in the sensor data are input into the multilayer sensing model to obtain the adjacency matrix of the multiple sensor signal channels in the sensor data output by the multilayer sensing model. The multilayer sensing model and the feature extraction model are trained together.
[0178] Optionally, the data acquisition module 701 is used for: Acquire the sensor data corresponding to each time point within a single time window.
[0179] Optionally, the feature extraction module 702 is specifically used for: According to the chronological order, the intermediate features corresponding to each time point in the single time window are spliced together to obtain the spliced features; The concatenated features are input into the time-map convolutional neural network in the feature extraction model to obtain the first feature output by the time-map convolutional neural network.
[0180] Optionally, the feature extraction module 702 is specifically used for: The intermediate features corresponding to each time point in the single time window are subjected to dimensionality upscaling to obtain high-dimensional features corresponding to each time point in the single time window. After dimensionality upscaling, the number of feature dimensions of the high-dimensional features is greater than the number of feature dimensions of the intermediate features. In chronological order, the high-dimensional features corresponding to each time point in a single time window are concatenated to obtain concatenated features.
[0181] Optionally, the feature extraction module 702 is specifically used for: Local window features are extracted from the stitched features using a convolutional sliding window, which is used to slide along the time dimension; The local window features are subjected to temporal convolution processing to obtain the first feature, which is used to characterize the changes in channel data of different sensor signal channels in the sensor data within the convolution sliding window.
[0182] Optionally, the device further includes an adjustment module for: Based on historical head motion recognition results, the current head motion recognition results are corrected.
[0183] Optionally, the adjustment module is specifically used for: Based on the historical head motion recognition results, the rationality of the current head motion recognition results is verified; If the current head motion recognition result is found to be unreasonable, the current head motion recognition result is corrected based on the historical head motion recognition results.
[0184] Optionally, the action recognition module 703 is specifically used for: Based on the first feature, head action recognition is performed to obtain the probability of each head action category; The historical probabilities of each head action category obtained within a historical time period and the current probabilities of each head action category are fused together to obtain the current fused probabilities of each head action category. Based on the fusion probability of each head action category, the current head action recognition result is determined.
[0185] Optionally, the data acquisition module 701 is used for: Acquire sensor data collected by multiple sensors, the sensor data including channel data of multiple sensor signal channels corresponding to the multiple sensors, and each of the multiple sensors corresponds to at least one sensor signal channel.
[0186] Optionally, the feature extraction module 702 is specifically used for: The channel data of each sensor signal channel corresponding to each of the multiple sensors are normalized to obtain the normalized channel data of the multiple sensor signal channels. Based on the data synchronization and data timing of different sensor signal channels in the sensor data, feature extraction is performed on the channel data of multiple normalized sensor signal channels to obtain the first feature.
[0187] Optionally, the plurality of sensors include a pressure sensor and an inertial sensor, and the plurality of sensor signal channels include a single sensor signal channel corresponding to the pressure sensor and six sensor signal channels corresponding to the inertial sensor.
[0188] Optionally, the device also includes a motion interaction module for: If the head action recognition result is a nodding action, the interactive operation corresponding to the nodding action is executed; If the head movement recognition result is a head shaking movement, the interactive operation corresponding to the head shaking movement is executed.
[0189] In summary, in this embodiment, when sensor data representing head movement state and including channel data from multiple sensor signal channels are obtained, features can be extracted from the channel data of multiple sensor signal channels based on the data synchronization and temporal sequence of different sensor signal channels in the sensor data to obtain a first feature. This first feature is then used for head movement recognition to obtain a head movement recognition result. Thus, by specifically analyzing the data correlation and temporal sequence of different sensor signal channels in the sensor data, high-quality feature mining of channel data from multiple sensor signal channels can be achieved for head movement recognition, improving the accuracy of head movement recognition.
[0190] It should be noted that the apparatus provided in the above embodiments is only an example of the division of the above functional modules. In practical applications, the above functions can be assigned to different functional modules as needed, that is, the internal structure of the apparatus can be divided into different functional modules to complete all or part of the functions described above. In addition, the apparatus and method embodiments provided in the above embodiments belong to the same concept, and their implementation process can be found in the method embodiments, which will not be repeated here.
[0191] See Figure 8 , Figure 8 This is a schematic diagram of the structure of an electronic device provided in an exemplary embodiment of this application. The electronic device may include one or more of the following components: a processor 810 and a memory 820. The electronic device may be the aforementioned terminal and / or wearable device.
[0192] Optionally, the processor 810 connects to various parts of the electronic device using various interfaces and lines, and performs various functions and processes data by running or executing instructions, programs, code sets, or instruction sets stored in the memory 820, and by calling data stored in the memory 820. Optionally, the processor 810 can be implemented in at least one hardware form selected from Digital Signal Processing (DSP), Field-Programmable Gate Array (FPGA), and Programmable Logic Array (PLA).
[0193] The processor 810 can integrate one or more of the following: a central processing unit (CPU), a graphics processing unit (GPU), a neural network processing unit (NPU), and a baseband chip. The CPU primarily handles the operating system, user interface, and applications; the GPU is responsible for rendering and drawing the content displayed on the touchscreen; the NPU implements artificial intelligence (AI) functions; and the baseband chip handles wireless communication. It is understood that the baseband chip can also be implemented as a separate chip without being integrated into the processor 810.
[0194] The memory 820 may include random access memory (RAM) or read-only memory (ROM). Optionally, the memory 820 may include a non-transitory computer-readable storage medium. The memory 820 may be used to store instructions, programs, code, code sets, or instruction sets. The memory 820 may include a program storage area and a data storage area, wherein the program storage area may store instructions for implementing an operating system, instructions for at least one function, instructions for implementing the various method embodiments described above, etc.; the data storage area may store data created based on the use of the electronic device, etc.
[0195] In addition, those skilled in the art will understand that the structure of the electronic device shown in the above figures does not constitute a limitation on the electronic device. The electronic device may include more (e.g., microphone, speaker, power supply component, display component, sensor component) or fewer components than shown, or combine certain components, or have different component arrangements.
[0196] This application provides a computer-readable storage medium storing at least one computer instruction, which is executed by a processor to implement the head motion recognition method as described in the above embodiments.
[0197] On the other hand, this application provides a computer program product, which includes computer instructions. When the processor executes the computer instructions, it implements the head motion recognition method as described in the above embodiments.
[0198] Those skilled in the art will recognize that the functions described in the embodiments of this application in one or more of the above examples can be implemented using hardware, software, firmware, or any combination thereof. When implemented using software, these functions can be stored in a computer-readable medium or transmitted as one or more instructions or code on a computer-readable medium. Computer-readable media include computer storage media and communication media, wherein communication media include any medium that facilitates the transfer of a computer program from one place to another. Storage media can be any available medium that can be accessed by a general-purpose or special-purpose computer.
[0199] The above description is merely an optional embodiment of this application and is not intended to limit this application. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A head motion recognition method, characterized in that, The method includes: Acquire sensor data, which is used to characterize the head movement state and includes channel data of multiple sensor signal channels; Based on the data synchronization and data timing of different sensor signal channels in the sensor data, feature extraction is performed on the channel data of the multiple sensor signal channels to obtain the first feature; Head motion recognition is performed based on the first feature to obtain the head motion recognition result.
2. The method according to claim 1, characterized in that, Based on the data synchronization and data timing of different sensor signal channels in the sensor data, feature extraction is performed on the channel data of the multiple sensor signal channels to obtain a first feature, including: The channel data of the multiple sensor signal channels are input into the feature extraction model to obtain the first feature output by the feature extraction model. The feature extraction model includes a spatial graph convolutional neural network and a temporal graph convolutional neural network. The spatial graph convolutional neural network is used to determine the data synchronization of different sensor signal channels in the sensor data, and the temporal graph convolutional neural network is used to determine the data temporality of different sensor signal channels in the sensor data.
3. The method according to claim 2, characterized in that, The step of inputting channel data from the multiple sensor signal channels into a feature extraction model to obtain the first feature output by the feature extraction model includes: The channel data of the multiple sensor signal channels are input into the spatial graph convolutional neural network in the feature extraction model to obtain the intermediate features output by the spatial graph convolutional neural network. The intermediate features are used to characterize the data synchronization of different sensor signal channels in the sensor data. The intermediate features are input into the time-map convolutional neural network in the feature extraction model to obtain the first feature output by the time-map convolutional neural network. The first feature is used to characterize the data temporal sequence of different sensor signal channels in the intermediate features.
4. The method according to claim 3, characterized in that, The method further includes: Determine the adjacency matrix of the plurality of sensor signal channels in the sensor data, wherein the adjacency matrix is used to characterize the data synchronization degree between any two sensor signal channels in the plurality of sensor signal channels; The step of inputting the channel data of the multiple sensor signal channels into the spatial graph convolutional neural network in the feature extraction model to obtain the intermediate features output by the spatial graph convolutional neural network includes: The adjacency matrix and the channel data of the multiple sensor signal channels are input into the spatial graph convolutional neural network in the feature extraction model to obtain the intermediate features output by the spatial graph convolutional neural network.
5. The method according to claim 4, characterized in that, Determining the adjacency matrix of the plurality of sensor signal channels in the sensor data includes: The channel data of the multiple sensor signal channels in the sensor data are input into the multilayer sensing model to obtain the adjacency matrix of the multiple sensor signal channels in the sensor data output by the multilayer sensing model. The multilayer sensing model and the feature extraction model are trained together.
6. The method according to claim 2, characterized in that, The acquisition of sensor data includes: Acquire the sensor data corresponding to each time point in a single time window; The step of inputting the intermediate features into the temporal graph convolutional neural network in the feature extraction model to obtain the first feature output by the temporal graph convolutional neural network includes: According to the chronological order, the intermediate features corresponding to each time point in the single time window are spliced together to obtain the spliced features; The concatenated features are input into the time-map convolutional neural network in the feature extraction model to obtain the first feature output by the time-map convolutional neural network.
7. The method according to claim 6, characterized in that, The step of concatenating the intermediate features corresponding to each time point in a single time window according to chronological order to obtain concatenated features includes: The intermediate features corresponding to each time point in the single time window are subjected to dimensionality upscaling to obtain high-dimensional features corresponding to each time point in the single time window. After dimensionality upscaling, the number of feature dimensions of the high-dimensional features is greater than the number of feature dimensions of the intermediate features. In chronological order, the high-dimensional features corresponding to each time point in a single time window are concatenated to obtain concatenated features.
8. The method according to claim 6, characterized in that, The step of inputting the concatenated features into the time-map convolutional neural network in the feature extraction model to obtain the first feature output by the time-map convolutional neural network includes: Local window features are extracted from the stitched features using a convolutional sliding window, which is used to slide along the time dimension; The local window features are subjected to temporal convolution processing to obtain the first feature, which is used to characterize the changes in channel data of different sensor signal channels in the sensor data within the convolution sliding window.
9. The method according to any one of claims 1-8, characterized in that, The method further includes: Based on historical head motion recognition results, the current head motion recognition results are corrected.
10. The method according to claim 9, characterized in that, Based on historical head motion recognition results, the current head motion recognition results are corrected, including: Based on the historical head motion recognition results, the rationality of the current head motion recognition results is verified; If the current head motion recognition result is found to be unreasonable, the current head motion recognition result is corrected based on the historical head motion recognition results.
11. The method according to any one of claims 1-8, characterized in that, The head action recognition based on the first feature to obtain the head action recognition result includes: Based on the first feature, head action recognition is performed to obtain the probability of each head action category; The historical probabilities of each head action category obtained within a historical time period and the current probabilities of each head action category are fused together to obtain the current fused probabilities of each head action category. Based on the fusion probability of each head action category, the current head action recognition result is determined.
12. The method according to any one of claims 1-8, characterized in that, The acquisition of sensor data includes: Acquire sensor data collected by multiple sensors, wherein the sensor data includes channel data of multiple sensor signal channels corresponding to the multiple sensors, and each of the multiple sensors corresponds to at least one sensor signal channel; Based on the data synchronization and data timing of different sensor signal channels in the sensor data, feature extraction is performed on the channel data of the multiple sensor signal channels to obtain a first feature, including: The channel data of each sensor signal channel corresponding to each of the multiple sensors are normalized to obtain the normalized channel data of the multiple sensor signal channels. Based on the data synchronization and data timing of different sensor signal channels in the sensor data, feature extraction is performed on the channel data of multiple normalized sensor signal channels to obtain the first feature.
13. The method according to claim 12, characterized in that, The plurality of sensors include pressure sensors and inertial sensors, and the plurality of sensor signal channels include a single sensor signal channel corresponding to the pressure sensor and six sensor signal channels corresponding to the inertial sensors.
14. The method according to any one of claims 1-8, characterized in that, The method further includes: If the head action recognition result is a nodding action, the interactive operation corresponding to the nodding action is executed; If the head movement recognition result is a head shaking movement, the interactive operation corresponding to the head shaking movement is executed.
15. A head motion recognition device, characterized in that, The device includes: The data acquisition module is used to acquire sensor data, which is used to characterize the head movement state and includes channel data of multiple sensor signal channels; The feature extraction module is used to extract features from the channel data of the multiple sensor signal channels based on the data synchronization and data timing of different sensor signal channels in the sensor data, and obtain the first feature; The action recognition module is used to perform head action recognition based on the first feature to obtain the head action recognition result.
16. An electronic device, characterized in that, The electronic device includes a processor and a memory, the memory storing at least one computer instruction, which is loaded and executed by the processor to implement the head motion recognition method as described in any one of claims 1 to 14.
17. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores at least one computer instruction, which is executed by a processor to implement the head motion recognition method as described in any one of claims 1 to 14.
18. A computer program product, characterized in that, The computer program product includes computer instructions, and when the processor executes the computer instructions, it implements the head motion recognition method as described in any one of claims 1 to 14.
Citation Information
Patent Citations
Multi-channel action recognition method and device based on adaptive convolutional neural network
CN114417911A
Multi-channel gesture recognition method and device
CN119970010A
Wheeled robot fault diagnosis method and system based on dynamic graph convolutional network
CN120408074A