Non-inductive facial emotion recognition method based on ultrasonic multi-region perception of smart phone
By emitting ultrasonic signals through the smartphone's speaker and microphone, and combining multi-regional acoustic reflection modeling and deep learning models, this technology solves the problems of privacy infringement, high computational cost, and high hardware dependence of existing facial emotion recognition technologies. It achieves high-precision, robust, and seamless facial emotion recognition that is applicable to existing smartphones.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-12-04
- Publication Date
- 2026-03-10
AI Technical Summary
Existing facial emotion recognition technologies suffer from problems such as privacy violations, high computational costs, high hardware dependence, coarse perception granularity, simple signal processing, and insufficient model capabilities in ubiquitous mobile scenarios, making it difficult to achieve high-precision, robust, and seamless facial emotion recognition.
By using the built-in speaker and microphone of a smartphone to transmit and receive ultrasonic signals, and through multi-region acoustic reflection modeling and deep learning models, signal preprocessing, multi-path phase compensation and emotion-related acoustic map construction are performed. Combined with the emotion-acoustic Transformer model, high-precision emotion recognition is achieved.
It achieves seamless, high-precision, and robust facial emotion recognition, is applicable to existing smartphones, can work stably in complex indoor environments, achieves a recognition accuracy of 90.25%, reduces computational overhead, and protects user privacy.
Smart Images

Figure CN121622040A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the technical field of mobile computing, contactless sensing and artificial intelligence, in particular to a method and system for contactless facial emotion recognition with high precision and robustness without visual information, by emitting and receiving ultrasonic signals with the built-in loudspeaker and microphone of a smartphone, through innovative multi-region acoustic reflection modeling, multi-path signal processing and deep learning models. BACKGROUND
[0002] Emotion, as a core component of human psychological activities, its accurate recognition has revolutionary significance for the next generation of human-computer interaction, remote mental health monitoring, personalized education, intelligent entertainment and vehicle safety systems. Facial expression is the most direct and richest natural expression channel of emotion, and its subtle changes can accurately reflect the individual's psychological state. Therefore, automatic facial emotion recognition technology has become a research hotspot in academia and industry.
[0003] Existing facial emotion recognition technology solutions can be mainly summarized into the following three categories, but they all have inherent defects and are difficult to apply in ubiquitous mobile scenarios.
[0004] Computer vision-based method: This method is the most mainstream solution at present. It continuously captures user facial images or video streams through the front camera of the smartphone, and uses deep learning models such as convolutional neural networks to extract expression features (such as action units, geometric features) and classify them. Although this method can achieve high accuracy in controlled lighting environments, its inherent defects are extremely prominent: first, it is extremely dependent on environmental lighting conditions, and its performance drops sharply in dark light, backlight or severe lighting changes; second, continuous video capture constitutes a serious invasion of user privacy, and users' concerns about the collection and use of facial data greatly limit its application scenarios; finally, the video processing computation overhead is large, resulting in high phone power consumption, which cannot achieve long-term and contactless emotion monitoring.
[0005] Physiological signal-based method: This method infers emotional state by collecting electroencephalogram, electrocardiogram, galvanic skin response and other physiological signals. Its advantage is that there is a strong physiological correlation between signals and emotions, and the accuracy is high. However, its biggest disadvantage is that it requires users to wear special sensor devices such as head-mounted EEG devices, heart rate bands or smart bracelets. This not only brings great inconvenience and discomfort to users, disrupting the naturalness of use, but also the devices themselves are expensive, which cannot be popularized among the general public of smartphone users, and are contrary to the vision of "anytime, anywhere" emotion sensing.
[0006] Radio frequency (RF) signal-based methods: In recent years, research has begun to explore using the reflection of wireless RF signals such as Wi-Fi, millimeter waves, and ultra-wideband radar to sense minute facial movements. These methods enable contactless sensing and are unaffected by light. However, their fatal flaw lies in hardware dependence. To achieve sufficiently accurate facial micro-motion sensing, these systems typically require dedicated Wi-Fi network cards with multiple antennas, and expensive millimeter-wave or UWB radar modules—hardware that is not currently standard on most commercial smartphones. Therefore, such solutions cannot be directly deployed on the billions of existing smartphones, severely limiting their application scope.
[0007] With the widespread adoption of smartphones, contactless sensing using their built-in acoustic sensors (speakers and microphones) has become an emerging field. Preliminary research (such as FacER) has attempted to use near-ultrasound for facial expression recognition, demonstrating the feasibility of acoustic sensing. However, these pioneering works face significant technical bottlenecks:
[0008] Coarse perceptual granularity: These devices typically treat the entire face as a uniform reflective surface, failing to distinguish the unique contributions of different facial organs in expression formation. This results in their inability to capture key emotional features formed by the activity of subtle local muscle groups (such as a slightly furrowed brow or a slight upturn of the corners of the mouth), leading to insufficient information acquisition.
[0009] The signal processing is simplistic: it ignores the phase differences and multipath effects that occur when sound waves travel along different paths. These unprocessed signal interferences can "contaminate" effective emotional features, leading to inaccurate feature extraction, especially in complex indoor multipath environments.
[0010] Insufficient model capabilities: They mostly use traditional machine learning models (such as SVM) or simple neural networks, which are difficult to learn and characterize the nonlinear, deep mapping relationship between multi-region, high-dimensional acoustic features and complex emotional states, and have poor adaptability to individual differences and environmental changes.
[0011] In summary, the following key scientific problems urgently need to be solved in the existing technology:
[0012] How can we extract sufficiently rich and nuanced emotional signature information from the acoustic signals of a smartphone without infringing on privacy or adding external devices?
[0013] How to overcome the multipath interference and phase misalignment problems faced by acoustic signals in complex indoor environments, thereby improving signal quality and characteristic reliability?
[0014] How can we build a robust model that can effectively integrate heterogeneous acoustic features from different facial regions and stably map them to multiple emotion categories, while also being robust against individual differences and environmental changes? Summary of the Invention
[0015] The purpose of this invention is to completely overcome the many shortcomings of existing facial emotion recognition technologies and provide a non-intrusive facial emotion recognition method and system based on multi-region ultrasonic sensing on smartphones. The core idea of this method is to transform the smartphone into a high-precision "acoustic radar," using innovative signal processing and deep learning models to "interpret" the user's emotional state from inaudible ultrasonic echoes. This invention aims to achieve a completely non-intrusive (user does not need to change any behavior), absolutely privacy-protecting (no camera used), high-precision (recognizing multiple emotions), strong robustness (adapting to various real-world scenarios), and universal applicability (suitable for existing mainstream smartphones) emotion recognition solution.
[0016] To achieve the above-mentioned objectives, the present invention adopts the following technical solution:
[0017] In a first aspect, the present invention provides a non-sensory facial emotion recognition method based on multi-region ultrasonic sensing of a smartphone, comprising the following steps:
[0018] Step S100: Ultrasonic signal transmission and raw data acquisition.
[0019] The system controls the smartphone's speaker to emit frequency-modulated continuous wave signals, i.e., inaudible linear frequency modulation ultrasound chirp signals, directionally toward the user's face at a preset power and period. Simultaneously, it controls one or more microphones on the smartphone to collect mixed echo signals reflected from the user's face and its surrounding environment at a high sampling rate and convert them into digital signals for storage.
[0020] Step S200: Signal preprocessing and multi-region acoustic reflection modeling.
[0021] This step aims to transform the coarse raw signal into fine features with spatial resolution. Its sub-steps include:
[0022] S210: Signal Denoising and Enhancement. The original echo signal is bandpass filtered to retain the effective components of the ultrasonic frequency band, and an adaptive filtering algorithm is used to initially suppress steady-state noise in the environment.
[0023] S220: Construction of a Multi-Region Reflection Model. Based on facial anatomy, the face is divided into six key reflection regions at the acoustic perception level: the forehead region (covering the forehead and brow bone), the eye region (covering the upper and lower eyelids and periorbital muscles), the nose region (covering the bridge and tip of the nose), the cheek region (covering the cheekbone and soft tissue of the cheek), the mouth region (covering the lips, orbicularis oris muscle, and chin), and the ear region (covering the outer ear). This division ensures that each region is associated with a specific facial muscle group, allowing us to capture the acoustic features generated by changes in facial expressions in greater detail. This method improves the accuracy and robustness of emotion recognition because it enables refined perception and analysis of different facial regions involved in different expressions.
[0024] S230: Regional Energy Feature Extraction. Matched filtering is used to convolve the received echo signal with the known transmitted chirp signal to improve the signal-to-noise ratio and compress the pulse. Subsequently, digital beamforming is employed to spatially separate and extract sub-signal components corresponding to each region from the mixed signal based on the geometric models of the six regions. Finally, the energy integral of each region's sub-signal within a specific time window is calculated as the preliminary acoustic energy feature for that region. This process not only improves signal purity but also effectively captures acoustic energy changes generated by facial muscle activity, providing crucial acoustic information for emotion recognition. Through these two steps, we are able to extract emotion-related acoustic features from ultrasound signals, laying a solid foundation for subsequent emotion recognition.
[0025] Step S300: Multipath phase compensation and construction of emotion-related acoustic maps. The core of this step is to solve the problem of physical distortion during signal propagation and generate high-quality model input features.
[0026] S310: Phase Difference Calculation and Compensation. Identify a reference region (e.g., the forehead area) and calculate the acoustic wave propagation path difference of the other five regions relative to this reference region. According to wave physics, this path difference causes a phase difference in the received signal. The calculation formula is:
[0027]
[0028] Where λ is the wavelength of the ultrasound wave. Subsequently, for each region's sub-signal... Phase correction is performed, and the compensation formula is as follows:
[0029]
[0030] in, Let be the signal amplitude correction coefficient for the i-th region. The original signal, The signal arrival time delay for the i-th region is... This is a noise signal.
[0031] This operation ensures that signal components from different regions but reflecting the same emotional event are aligned in phase, effectively eliminating destructive interference caused by different paths and significantly improving the quality of subsequent time-frequency analysis.
[0032] S320: Construct an emotion-related acoustic map. The signals from all regions, after phase compensation, are superimposed to obtain the total signal. A short-time Fourier transform is performed on the total signal to obtain its high-resolution time-frequency spectrum. Then, a specific time delay window is applied to each region of the time-frequency spectrum. and ultrasonic frequency range A two-dimensional integral is performed to calculate the representative energy value Ei of the region at that moment. Finally, the energy values of all six regions are arranged in a fixed order to form a six-dimensional feature vector, i.e., the emotion-related acoustic spectrum.
[0033]
[0034] in, These represent the acoustic energy values for the forehead, eyes, nose, cheeks, mouth, and ears, respectively. These values are obtained by analyzing the time-spectrum maps of the corresponding regions.
[0035] As a compact and information-rich representation, the vector value distribution of ERAM changes sensitively and discriminatively with changes in contour and reflectivity caused by facial muscle movements.
[0036] Step S400: Emotion Recognition and Classification Based on the Emotion-Acoustic Transformer Model. This step is the brain of the entire method, responsible for learning from and making decisions in the ERAM. The ERAM generated in Step S300 is input into a pre-trained offline emotion-acoustic Transformer model. This model is processed through the following sub-modules:
[0037] S410: Parallel Region Encoding. EA-Former first compresses and maps the ERAM (which can be viewed as a 2D map over multiple chirp cycles) of each region into a high-dimensional feature map by passing it through a weighted 2D convolutional layer. This feature map is then segmented into multiple non-overlapping image patches, and each patch is flattened into a token. To distinguish between different regions and positions within patches, a learnable region ID encoding and a relative position encoding are embedded in each token. All tokens from the six regions together form a long input sequence.
[0038] S420: Local Feature Learning within Regions. The input sequence first enters a region-specific self-attention module consisting of a 4-layer Transformer encoder. This module uses a window self-attention mechanism to restrict each token to interacting only with other tokens within the same region. This allows the model to focus on learning subtle local patterns within each region, such as the specific texture of a wrinkled nose caused by a "disgust" expression, or the unique pattern of drooping eyes caused by "sadness."
[0039] S430: Inter-regional global feature fusion. The sequence then enters an inter-regional global attention module consisting of an 8-layer Transformer encoder. In this module, tokens from all regions undergo global, fully connected attention computation. Rotational position encoding is introduced to better capture the sequential relationships within the sequence. The core function of this module is to learn the co-variation patterns of different facial regions when expressing specific emotions. For example, it can learn that "true happiness" is not only the upward turn of the corners of the mouth (mouth region feature) but must also be accompanied by the contraction of the orbicularis oculi muscle (eye region feature), thus effectively distinguishing between a genuine smile and a fake smile.
[0040] S440: Emotion Classification Decision. The output token sequence, after being processed by multiple Transformer layers, is aggregated into a fixed-size, global facial emotion representation vector through a global average pooling layer. This vector is then passed through a fully connected layer to map the dimensions to the number of target emotion categories (e.g., 9 dimensions), and finally through a Softmax activation function, outputting a probability distribution vector. Each element of this vector represents the confidence probability that the input acoustic signal belongs to the corresponding emotion category. The category with the highest probability is taken as the final emotion recognition result.
[0041] The emotion-acoustic Transformer model includes:
[0042] Parallel Region Encoding Module: Designed to improve computational efficiency on mobile platforms, this module supports six... ERAM inputs perform shared weights Convolution is performed to compress it into a 128-dimensional feature space. Subsequently, each feature map is segmented into 16 non-overlapping regions. The blocks are formed into tokens. To distinguish different regions, each token is appended with an 8-dimensional region ID embedding and a 2-dimensional relative position code. Finally, 96 tokens are concatenated into a sequence carrying region identity and local geometric information, facilitating efficient processing by the subsequent self-attention module.
[0043] Intra-region self-attention module: This module is designed specifically for computational efficiency on mobile platforms. It optimizes computation by performing windowed self-attention within a local window of 16 tokens, focusing on simulating dependencies within a single facial region. This method effectively identifies subtle phase changes within a region, such as cheek hollowing or mouth lifting, while reducing unnecessary redundancy in cross-region computation. Subsequently, a SwiGLU feedforward network performs feature transformation on the attention output, expanding the channels to enhance the model's expressive power. To improve generalization and prevent overfitting, DropPath regularization is introduced. The overall design achieves efficient and robust feature extraction with limited parameters.
[0044] Inter-region global attention module: After intra-region processing, this module integrates 96 tokens from six facial regions into a global sequence to identify cross-regional emotional connections. Through cross-regional relative position encoding (RPE) and rotational position encoding (RoPE), the model can capture synergies and long-range dependencies between different regions. An eight-head self-attention mechanism enables the model to focus on key features in different subspaces, while residual connections ensure the stability of feature transformation, thus providing a rich feature set that integrates local and global information for emotion classification.
[0045] The Emotion Classification Header Module aims to achieve fast and efficient emotion recognition. This module first performs global average pooling on 96 token features, compressing them into a 256-dimensional facial descriptor while retaining key statistical features. Next, a fully connected layer maps the descriptor to a nine-dimensional space, with each dimension corresponding to an emotion category. The Softmax function converts the output into an emotion probability distribution for decision-making, such as selecting the emotion category with the highest probability or threshold-based detection, ensuring that the recognition process is both accurate and suitable for real-time applications.
[0046] Architecturally, the system is built around four core modules: parallel region encoding for initial feature extraction and dimensionality reduction; intra-region self-attention modules for simulating local texture patterns within each region; inter-region global attention modules for capturing global context and collaborative dynamics across regions; and a sentiment classification head for feature aggregation and probability estimation. This tight design results in an end-to-end lightweight framework that effectively integrates local details and global semantic understanding for real-time sentiment inference.
[0047] The emotion recognition method of this invention achieves accurate capture and classification of facial expressions through an emotion-acoustic Transformer model. This method utilizes high-dimensional feature extraction and fine-grained region differentiation to improve recognition accuracy. Simultaneously, by combining local and global feature learning, it effectively captures subtle changes in facial expressions, enhancing the ability to recognize complex emotions. Furthermore, the application of rotational position encoding improves the model's understanding of dynamic changes in expressions, while global average pooling and the Softmax activation function ensure the efficiency and accuracy of the classification process. Overall, this invention outperforms existing technologies in terms of real-time performance, robustness, and generalization ability.
[0048] Secondly, based on the method described in the first aspect, the present invention also provides a non-sensory facial emotion recognition system, which is integrated into the operating system of a smartphone in software form or exists as a standalone application, and includes the following logical modules:
[0049] The ultrasonic signal transceiver control module manages the audio driver, instructing the speaker to emit chirp signals in a specific format and timing, and configuring the microphone gain and sampling rate to receive echoes. This module ensures synchronization between transmission and reception and processes the underlying audio stream data.
[0050] Signal preprocessing and feature engineering module: This module encapsulates all the algorithms of steps S200 and S300 in claim 1. It receives the raw audio data and outputs a standardized, aligned emotion-related acoustic map. This module is the core of the system's signal processing and determines the quality of the input features.
[0051] Emotion Recognition Inference Engine Module: This module incorporates a trained EA-Former model. It receives ERAM features from the previous module, performs forward inference calculations, and outputs the emotion classification result. This module can be optimized for different mobile phone computing capabilities, for example, by using mobile deep learning frameworks for acceleration.
[0052] Context Management and Application Interface Module: Serving as an intelligent bridge between the emotion recognition service and external applications, this module dynamically senses device usage status and user presence by fusing multi-source sensor data in real time. Based on contextual information such as phone posture, distance from the face, and ambient lighting, it intelligently schedules the execution timing and resource allocation of emotion recognition tasks. This module also provides a standardized and secure interface. Through permission verification mechanisms and data anonymization, it securely provides structured recognition results, including emotion type, confidence level, timestamp, and quality indicators, to authorized third-party applications. This ensures system-level sharing of emotion perception capabilities and application-layer innovation while protecting user privacy.
[0053] Compared with existing technologies, the core advantages of this invention are reflected in the following aspects:
[0054] 1. A groundbreaking, seamless privacy protection solution: This invention completely eliminates the need for a camera, employing ultrasonic waves, which are completely inaudible to the human ear, for sensing. While users are engaged in routine activities on their phones (such as browsing the web or watching videos), the system runs seamlessly in the background, realizing the ideal paradigm of "perception as a service." This fundamentally eliminates users' concerns about visual privacy leaks and clears the way for the application of emotion recognition in sensitive scenarios (such as psychological counseling rooms).
[0055] 2. Breakthrough recognition accuracy and robustness:
[0056] Multi-region fine perception: By dividing the human face into six acoustic reflection zones, the present invention is able to capture richer and more subtle emotional features than existing acoustic methods, especially those subtle expressions formed only by local muscle activity.
[0057] Phase compensation improves signal-to-noise ratio: The original multi-path phase compensation mechanism effectively solves the key interference problem in acoustic perception, making the extracted features purer and more reliable, and significantly improving the stability of the system in complex indoor multipath environments.
[0058] Powerful model drives high accuracy: The EA-Former model combines the local perception capability of CNN with the global relationship modeling capability of Transformer, enabling it to deeply understand the complex interaction of multi-region features and the intrinsic connection of emotions, thereby achieving an average recognition accuracy of up to 90.25%, surpassing known emotion recognition solutions based on mobile phone acoustics.
[0059] 3. Excellent universality and practicality:
[0060] Zero hardware requirements: This solution utilizes only the existing standard hardware of smartphones (speakers and microphones) without any external devices, enabling it to be deployed on commercial smartphones at zero cost, and possessing enormous market potential and social value.
[0061] Strong scene adaptability: After rigorous experimental verification, the present invention can maintain excellent performance degradation (the accuracy decrease is generally less than 5%) under various non-ideal conditions such as changes in the distance between the mobile phone and the face (10cm-40cm), user head posture deviation (±30°), and partial facial occlusion, proving its high usability in real life.
[0062] 4. Low power consumption and high efficiency: Through careful optimization of the signal processing flow and EA-Former model (such as using a lightweight attention mechanism), this system greatly reduces computational overhead while ensuring accuracy, making it possible to perform real-time and continuous emotion monitoring on mobile devices, with minimal impact on battery life. Attached Figure Description
[0063] Figure 1 This is an overall flowchart of the non-sensory facial emotion recognition method based on multi-region ultrasonic sensing of a smartphone provided in this embodiment of the invention.
[0064] Figure 2 This is a schematic diagram of multi-region acoustic reflection modeling and phase compensation provided in an embodiment of the present invention, which illustrates the division of six key regions of a face and the principle of phase compensation.
[0065] Figure 3 This is a structural diagram of the Emotion-Acoustic Transformer (EA-Former) model provided in this embodiment of the invention, which details four key components: parallel region encoding, intra-region self-attention computation, inter-region global attention fusion, and emotion classification head.
[0066] Figure 4 This is a performance comparison chart between the present invention and existing methods, provided by an embodiment of the invention. The bar chart shows the comparison results of the recognition accuracy of the present invention method with baseline models such as ResNet-18 and SonicFace.
[0067] Figure 5 This is a diagram showing the impact of different experimental settings on system performance provided in the embodiments of the present invention. Sub-figures (a), (b), and (c) respectively illustrate the system performance under different distances, different user postures, and ablation experimental conditions. Detailed Implementation
[0068] To make the objectives, technical solutions, and advantages of this invention clearer, the invention will be described in detail below with reference to specific embodiments and the accompanying drawings. It should be noted that these embodiments are for illustrative purposes only and are not intended to limit the invention.
[0069] Example 1: Complete emotion recognition process based on the method of the present invention. This example describes in detail the application of the present invention to an emotion monitoring plugin in a smartphone video application.
[0070] Step 1: System initialization and parameter configuration.
[0071] When a user first enables this feature, the system performs a one-time calibration, prompting the user to maintain a neutral expression for a moment to record their baseline acoustic characteristics. Simultaneously, the system loads pre-trained EA-Former model parameters, ultrasound signal parameters, and reference distance parameters required for phase compensation from the configuration file.
[0072] Step 2: Real-time signal acquisition and processing.
[0073] The emotion recognition plugin is activated when a user launches a video application and starts playing it.
[0074] Signal transmission: The system intermittently inserts Chirp ultrasound signals with the above parameters at a rate of 20 times per second into the underlying video audio stream via the audio API. Because the frequency of this signal is higher than the upper limit of human hearing and the energy is low, the user cannot detect it at all.
[0075] Signal reception: The main microphone at the bottom of the phone records audio simultaneously. To prevent interference from the recorded video audio, the system uses adaptive echo cancellation technology to subtract the currently playing video audio signal from the recording, thus retaining only the ultrasonic echo.
[0076] Step 3: Multi-region feature extraction and phase compensation (corresponding to) Figure 1 Steps S200 and S300 in the process.
[0077] The preprocessing module performs bandpass filtering and down-conversion on the received echo signal. Subsequently, the Hilbert transform is used to calculate the analytical form of the signal to facilitate phase processing. Figure 2 The image shows six key reflective areas on the face, which receive reflected signals via smartphone sensors for acoustic analysis used in emotion recognition.
[0078] Multi-region acoustic reflection modeling: The system calls the built-in six-region geometric model of the face, combines the relative attitude of the mobile phone estimated by the current mobile phone gyroscope data, dynamically adjusts the beamforming weight vector, and separates the sub-signals of the six regions from the mixed signal in real time.
[0079] Phase compensation: Using the forehead region as a reference, the path difference in other regions is calculated (e.g., approximate distances estimated based on the built-in face model and distance sensors), and then a formula is used. Perform real-time phase correction.
[0080] Constructing ERAM: Perform STFT on the compensated total signal (window length of 256, overlap of 128), and calculate the energy in the time-frequency interval corresponding to each region to form a new six-dimensional ERAM vector.
[0081] Step 4: Emotion Recognition Reasoning (corresponding to...) Figure 1 Step S400 in the process.
[0082] In this step, we combine the current ERAM vector with the ERAM vectors from several previous time steps to form a time series, which is then input into the Emotion-Acoustic Transformer (EA-Former) model. The structure of the EA-Former model is as follows: Figure 3As shown, it consists of four main modules: parallel region encoding, intra-region self-attention module, inter-region global attention module, and emotion classification head. First, the model maps the ERAM to a high-dimensional feature map using the parallel region encoding module, then segments it into multiple image patches. Each patch is flattened into a token and embedded with a region ID encoding and a relative position encoding. Next, the intra-region self-attention module uses a window self-attention mechanism to restrict each token to interact only with other tokens within the same region, learning subtle local patterns within each region. Then, the inter-region global attention module uses global fully connected attention computation to learn the collaborative change patterns of different facial regions when expressing specific emotions, and introduces rotational position encoding to capture the sequential relationships in the sequence. Finally, the emotion classification head aggregates the token sequence through a global average pooling layer and outputs a 9-dimensional probability vector through a fully connected layer and a Softmax activation function, for example: [Happy: 0.85, Sad: 0.02, ..., Neutral: 0.01].
[0083] Step 5: Output and Application of Results
[0084] The system determines "Happy" as the current emotion based on the probability vector output by the EA-Former model and sends the result [emotion: "Happy", confidence: 0.85, timestamp: ...] to the video application via the application programming interface. The video application can then use this result to record the user's emotional response in the background or intelligently recommend more similar cheerful video content, thus achieving personalized emotional interaction. This method of emotion recognition and application not only enhances the user experience but also provides valuable user feedback to content providers, helping to optimize content strategies and improve user satisfaction.
[0085] Example 2: Software Architecture of a Seamless Facial Emotion Recognition System
[0086] This embodiment describes, from a system implementation perspective, how the present invention is integrated into a mobile phone system-level service.
[0087] The system runs as a background service, EmotionService, on Android or iOS systems.
[0088] AcousticManager (ultrasound signal transceiver control module): Inherited from AudioTrack and AudioRecord, it is responsible for managing the playback and recording of ultrasound signals in a low-latency manner.
[0089] FeatureEngine (Signal Preprocessing and Feature Engineering Module): Implemented in C++ and encapsulated as an NDK library, it contains all digital signal processing algorithms (filtering, matched filtering, beamforming, phase compensation, STFT) to ensure computational efficiency.
[0090] InferenceCore (Emotion Recognition Inference Engine Module): Integrates with mobile machine learning frameworks (such as TensorFlow Lite or PyTorch Mobile) and is responsible for loading and running the optimized EA-Former Lite model.
[0091] ContextAwarenessManager (Context Management Module): Monitors system sensor services to obtain data such as device posture, ambient light, and proximity sensor readings, which are used to dynamically adjust signal processing parameters (e.g., pause emotion recognition when the phone is placed flat on a table).
[0092] EmotionAPI (Application Interface Module): Provides a standard RESTful API or AIDL interface for third-party applications to subscribe to emotion recognition results after requesting permissions.
[0093] Example 3: Robustness Testing and Performance Verification
[0094] This embodiment further illustrates the technical effects of the present invention through experimental data.
[0095] We implemented a UFE prototype system on a Huawei Nova 7 and recruited 20 participants of different genders and ages. We conducted over 50 hours of testing in laboratory and multiple real-world home environments.
[0096] Overall accuracy: In a classification task involving nine emotions (happiness, sadness, surprise, anger, fear, disgust, neutrality, contempt, and shame), the system achieved an average accuracy of 90.25%. Figure 4 The comparison results show that UFE outperforms baseline models such as ResNet-18 and SonicFace.
[0097] Robustness testing:
[0098] Distance robustness: Within the range of 10cm to 40cm, the recognition accuracy gradually decreases from 92.1% to 87.5%, with a variation of less than 5%, indicating that the system has good adaptability to different usage distances (e.g., Figure 5 (a) is shown.
[0099] Posture robustness: When the user's head is horizontally veered from 0° to 30°, the system accuracy remains above 88%. This is thanks to the phase compensation mechanism and the Transformer model's powerful learning ability for feature relationships, which can compensate for signal attenuation in some areas caused by posture changes (such as...). Figure 5 (b) is shown.
[0100] Occlusion robustness: In scenarios where users are wearing ordinary medical masks, the system can still maintain an accuracy of 87.9% by utilizing features of the unoccluded forehead, eyes, and part of the cheek area, demonstrating its practical value.
[0101] The specific embodiments described above further illustrate the purpose, technical solution, and beneficial effects of the present invention. It should be understood that the above descriptions are merely preferred embodiments of the present invention and are not intended to limit the scope of protection of the present invention. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of the present invention should be included within the scope of protection of the present invention.
Claims
1. A method for emotion recognition of a face without awareness based on ultrasonic multi-region sensing of a smart phone, characterized in that, Comprising the following steps: Step S100: ultrasonic signal emission and raw data acquisition, Step S200: signal preprocessing and multi-region acoustic reflection modeling, Step S300: multi-path phase compensation and emotion-related acoustic map construction, Step S400: emotion recognition and classification based on emotion-acoustic Transformer model, input the emotion-related acoustic map into the emotion-acoustic Transformer model for processing, and the model outputs the classification result corresponding to the user's facial emotion category. 2.The method of claim 1, wherein, Step S100 is as follows: control the loudspeaker of the smart phone to emit frequency-modulated continuous wave signals, i.e. inaudible chirp signals, to the user's face at a preset power and period; at the same time, one or more microphones of the smart phone are controlled synchronously to collect mixed echo signals reflected from the user's face and its surrounding environment at a high sampling rate and convert them into digital signals for storage. 3.The method of claim 1, wherein, Step S200: signal preprocessing and multi-region acoustic reflection modeling, Its sub-steps include: S210: signal noise reduction and enhancement, band-pass filtering is performed on the original echo signal to retain the effective components in the ultrasonic frequency band, and an adaptive filtering algorithm is used to preliminarily suppress the steady-state noise in the environment, S220: multi-region reflection model construction, based on the knowledge of human face anatomy, the face is divided into six key reflection regions in the acoustic perception layer, including: forehead area (covering forehead, brow bone area), eye area (covering upper and lower eyelids and eye muscles), nose area (covering nose bridge and nose tip), cheek area (covering zygomatic bone and cheek soft tissue), mouth area (covering lips, orbicularis oris muscle and chin), and ear area (covering external ear), this division ensures that each region is associated with a specific expression muscle group, S230: regional energy feature extraction, the received echo signal is convoluted with the known transmitted chirp signal using a matching filter technique to improve the signal-to-noise ratio and compress the pulse, then, through a digital beamforming algorithm, the sub-signal components corresponding to each region are spatially separated and extracted from the mixed signal according to the geometric model of the six regions, finally, the energy integral of each regional sub-signal in a specific time window is calculated as the preliminary acoustic energy feature of the region. 4.The method of claim 1, wherein, Step S300: multi-path phase compensation and emotion-related acoustic map construction, Specifically, S310: Phase difference calculation and compensation, identify a reference region, calculate the sound wave propagation path difference of other five regions relative to the reference region According to wave physics, this path difference will cause the received signal to produce a phase difference The calculation formula is: where λ is the wavelength of the ultrasound, and then the phase correction is applied to each region's sub-signal with the compensation formula: wherein is a signal amplitude correction coefficient for the i-th region, is an original signal, is a signal arrival time delay for the i-th region, is a noise signal, S320: build emotion-related acoustic map, superimpose all the region signals after phase compensation to get the total signal , perform short-time Fourier transform on the total signal to get its high-resolution time-frequency spectrogram, then, on the time-frequency spectrogram, for each region's specific time delay window and ultrasonic frequency range , perform two-dimensional integration to calculate the representative energy value Ei of the region at that moment, finally, arrange the energy values of all six regions in a fixed order to form a six-dimensional feature vector, i.e. emotion-related acoustic map: wherein, respectively represent the acoustic energy values of the forehead, eye, nose, cheek, mouth and ear regions, which are obtained by analyzing the time-frequency spectrograms of the corresponding regions, ERAM, as a compact and information-rich representation, its vector value distribution will change sensitively and distinguishably with the contour and reflection coefficient changes caused by facial muscle movement.
5. The emotion recognition method based on smartphone ultrasonic multi-region perception according to claim 1, wherein, Step S400: emotion recognition and classification based on emotion-acoustic Transformer model, specifically as follows, S410: Parallel region encoding, EA-Former first compresses and maps the ERAM of each region (which can be regarded as a 2D map under multiple Chirp cycles) into a high-dimensional feature map through a two-dimensional convolution layer with shared weights, then divides the feature map into multiple non-overlapping image blocks and flattens each image block into a token, in order to distinguish different regions and positions within the block, a learnable region ID code and a relative position code are embedded for each token, all tokens of the six regions together form a long input sequence, S420: Local feature learning within the region, the input sequence first enters a self-attention module within the region composed of 4 layers of Transformer encoder, which restricts each token to interact only with other tokens within the same region through window self-attention mechanism, which enables the model to focus on learning subtle local patterns within each region, S430: Global feature fusion between regions, then the sequence enters a global attention module between regions composed of 8 layers of Transformer encoder, in which all tokens of different regions perform global fully connected attention calculation, and rotation position encoding is introduced to better capture the order relationship in the sequence, the core role of this module is to learn the collaborative change pattern of different facial regions when expressing a specific emotion, S440: Emotion classification decision, the output token sequence after multiple layers of Transformer processing is aggregated into a fixed-size global facial emotion representation vector through a global average pooling layer, which is then mapped to the number of target emotion categories through a fully connected layer, and finally a Softmax activation function is used to output a probability distribution vector, each element value of the vector represents the confidence probability of the input acoustic signal belonging to the corresponding emotion category, and the category corresponding to the maximum probability is taken as the final emotion recognition result. 6.The method of claim 1, wherein the method further comprises: determining a face region in the image; and determining a face region in the image. The emotion-acoustic Transformer model includes: Parallel Region Encoding Module: Designed to improve computational efficiency on mobile platforms, this module supports six... ERAM inputs perform shared weights Convolution is performed to compress it into a 128-dimensional feature space. Subsequently, each feature map is segmented into 16 non-overlapping regions. The blocks are formed into tokens. To distinguish different regions, each token is appended with an 8-dimensional region ID embedding and a 2-dimensional relative position code. Finally, 96 tokens are concatenated into a sequence carrying region identity and local geometric information, facilitating efficient processing by the subsequent self-attention module. Intra-region self-attention module: This module is designed for computing efficiency on mobile platforms, optimizing the computing efficiency on mobile platforms, focusing on simulating the dependencies within a single facial region by performing windowed self-attention within a local window of 16 tokens, then a SwiGLU feedforward network is used to convert the attention output into features, expanding the channels to enhance the model's expression ability, and DropPath regularization is introduced to improve generalization and prevent overfitting, Inter-region global attention module: After intra-region processing, this module integrates the 96 tokens of the six facial regions into a global sequence to identify cross-regional emotional connections, through cross-regional relative position encoding (RPE) and rotation position encoding (RoPE), the model can capture the synergy and long-distance dependencies between different regions, eight-headed self-attention mechanism enables the model to focus on key features in different subspaces, and residual connection ensures the stability of feature conversion, thereby providing a rich feature set that integrates local and global information for emotion classification, The emotion classification head module: first, the 96 token features are globally averaged pooled into a 256-dimensional face descriptor, retaining key statistical features. Then, the descriptor is mapped to a 9-dimensional space through a fully connected layer, with each dimension corresponding to an emotion category. The Softmax function converts the output into an emotion probability distribution for decision-making. 7.The smart phone based ultrasonic multi-region sensing based emotion recognition method without explicit face detection according to claim 1, wherein, The emotion categories output by the emotion classification head module include: happy, sad, surprised, angry, scared, disgusted, neutral, contemptuous, and ashamed.
8. A system for implementing the method of any one of claims 1-7, characterized by The system is deployed on a smartphone and includes: An ultrasonic signal transmission and reception control module: used to control the speaker to emit ultrasonic Chirp signals and the microphone to receive reflected echoes; A signal preprocessing and feature engineering module: used to perform multi-region acoustic reflection modeling, phase compensation, and emotion-related acoustic map construction; An emotion recognition inference engine module: containing a pre-trained emotion-acoustic Transformer model, used to perform emotion classification based on the emotion-related acoustic map; An application interface module: used to provide the recognized emotion results to third-party applications.
9. A smartphone, characterized by A computer program is stored in a memory and executed by a processor, and the computer program, when executed by the processor, implements the method of any one of claims 1-7.
10. A computer-readable storage medium having stored thereon a computer program, characterized in that, The computer program, when executed by the processor, implements the steps of the method of any one of claims 1-7.