Facial expression recognition method and system based on sound perception

Through the combination of acoustic signal processing and deep learning network, the invasiveness, light sensitivity and privacy leakage of facial expression recognition system are solved, and low-cost and efficient facial expression recognition is achieved, which is suitable for a variety of application scenarios.

CN120279951AActive Publication Date: 2025-07-08SHANDONG UNIV +1
View PDF 7 Cites 0 Cited by

Patent Information

Application Number
CN202510756910.0
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-09
Publication Date
2025-07-08
Estimated Expiration
2045-06-09

AI Technical Summary

Technical Problem

The existing facial expression recognition system has problems such as strong invasiveness, sensitive lighting and privacy leakage risks, and wireless signal equipment is costly, making it difficult to popularize on a large scale.

Method used

Acoustic signals are used as perception media, and sound reflected signals are captured using microphone arrays, and spectrograms are generated through signal processing, and expression recognition is used using deep learning networks to avoid privacy leakage and light interference from traditional methods.

Benefits of technology

It realizes efficient facial expression recognition without relying on expensive hardware equipment, can effectively identify facial subtle dynamics under the conditions of scarcity of data, avoids the problem of deception by traditional methods, and has broad application prospects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120279951A_ABST
    Figure CN120279951A_ABST
Patent Text Reader

Abstract

The invention belongs to the technical field of facial expression recognition, and particularly relates to a facial expression recognition method and system based on sound perception. According to the method, an acoustic signal is used as a sensing medium, a microphone array is used for capturing a sound reflection signal, a spectrogram containing facial features is generated through signal processing, features are extracted from the spectrogram, facial expressions are classified through a deep learning network, and expression recognition is achieved. According to the facial expression recognition method provided by the invention, the facial expression recognition can be effectively realized by utilizing a commercial loudspeaker and a microphone array without depending on expensive hardware equipment even under the condition that the data volume is relatively small.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The invention belongs to the technical field of facial expression recognition, and in particular relates to a facial expression recognition method and system based on sound perception. Background Art

[0002] With the continuous development of computer technology, users' requirements for applications are also increasing. They not only hope that they can complete the predetermined functions, but also hope that they are more intelligent and user-friendly. By recognizing user emotions, intelligent systems can interact with users more naturally and efficiently, improving the overall experience. Facial expressions, as the most direct way to express emotions, can clearly reflect the psychological and physiological changes of individuals, and have broad application prospects in various human-computer interaction systems. Facial Expression Recognition (FER), as an important research direction in the computer field, has received widespread attention from academia and industry in recent years due to its wide application value in social robots, intelligent medical care, driving fatigue monitoring and many human-computer interaction systems.

[0003] Existing facial expression recognition systems are mainly based on three types of technical solutions: wearable devices, cameras, and wireless signals. Among them, although wearable devices can provide higher recognition accuracy, they require users to wear additional devices for a long time, which is not only invasive, but also may cause discomfort after long-term use, thus affecting the user experience. Many studies have achieved non-contact and accurate monitoring by using cameras to capture facial images or videos. However, such systems are sensitive to lighting conditions and have potential risks of privacy leakage. In recent years, wireless signals (such as WiFi, millimeter waves, and ultrasound) have provided new ideas for facial expression recognition, showing significant advantages in environmental adaptability, non-contact perception, and privacy protection. Despite this, WiFi and millimeter wave technologies are difficult to achieve large-scale popularization due to their high equipment costs. Summary of the invention

[0004] In view of the above technical problems, the present invention provides a facial expression recognition method and system based on sound perception. Acoustic signals are used as the perception medium, and a microphone array is used to capture the sound reflection signal. A spectrogram directly containing facial features is generated through a series of signal processing processes. Then, key features are extracted from the spectrogram, and efficient expression recognition is achieved with the help of a deep learning network. Compared with traditional visual methods, the method provided by the present invention not only effectively avoids the risk of privacy leakage, but also overcomes the interference of ambient light on recognition performance, providing new possibilities for the practical application of facial expression recognition FER technology.

[0005] The present invention is achieved through the following technical solutions: A facial expression recognition method based on sound perception, which uses acoustic signals as the perception medium, captures sound reflection signals by means of a microphone array, generates a spectrogram containing facial features through signal processing, extracts features from the spectrogram, and classifies facial expressions using a deep learning network to achieve expression recognition.

[0006] Further, the method includes the following steps: (1) Signal transmission: Transmit sound signals of spaced FMCW waveforms through a speaker. After the sound signals reach the human face, they are reflected to generate sound reflection signals, and the microphone array receives the sound reflection signals; (2) Signal processing: The signal processing process includes an audio frame segmentation sub-step, a distance measurement sub-step, a spectrogram generation sub-step, and a coordinate system conversion sub-step; The audio frame segmentation sub-step is used to extract the sound signals reflected by the face from the recording of the microphone array; The distance measurement sub-step is used to calculate the distance between the human face and the microphone array; The spectrogram generation sub-step generates a spectrogram containing facial information through the SRP-PHAT algorithm; The coordinate system rotation sub-step is used to convert the spectrogram from spherical coordinates to Cartesian coordinates; (3) Expression classification: Extract features from the spectrogram through a pre-trained deep learning model, input the extracted features into a fully connected deep neural network. The fully connected deep neural network outputs the facial expression classification result of the user by learning the mapping relationship between the features and the expression categories.

[0007] Further, in step (2), the audio frame segmentation sub-step includes: Locate the direct propagation signal from the original recording of the microphone array; Locate the main reflection signal after the direct propagation signal to determine the position of the facial reflection signal; After determining the position of the facial reflection signal from the recording of the microphone array, then segment and extract the sound signals reflected by the face from the recording of the microphone array.

[0008] Further, in step (2), the distance measurement sub-step specifically includes: Use the FMCW radar ranging algorithm to mix the received signal with the transmitted signal to generate an intermediate frequency signal; When measuring the distance, assume that there is a continuous FMCW signal during the blank time period between two FMCW pulses, and search backward step by step through a sliding window until a significant amplitude appears at a certain frequency in the intermediate frequency signal; The distance between the human face and the microphone array dCalculated by the following formula: ; Wherein, N represents the number of sliding windows, represents the maximum ranging distance of a single pulse, represents the distance measured in the current window.

[0009] Furthermore, in step (2), the spectrogram generation sub-step uses the SRP-PHAT algorithm to generate a spectrogram; In the process of generating a spectrogram using the SRP-PHAT algorithm, the calculation range of the azimuth angle is [-180°, 180°], and the calculation range of the elevation angle is: ; Wherein, ; d is the distance between the human face and the microphone array; is the diagonal length of the drawn image; is the minimum elevation angle of the scanning range of the SRP-PHAT in space; The minimum resolution res of the elevation angle is: .

[0010] Furthermore, in step (2), the coordinate system rotation sub-step is used to convert the spectrogram from the spherical coordinate system to the Cartesian coordinate system; Convert the coordinates of the intersection of a certain direction represented by the azimuth angle and the elevation angle in the spherical coordinate system with the projection plane into the specific coordinates in the Cartesian coordinate system, so as to map the azimuth information in the polar coordinate system to the Cartesian coordinate system and generate an image that can clearly display the contour.

[0011] Furthermore, step (3) specifically includes: (3.1) Data preprocessing: Align the spectrogram with the camera image through two-dimensional affine transformation, use the face localization algorithm to determine the position of the face in the camera image, and then obtain the position of the face in the spectrogram; (3.2) Facial expression recognition: Use the pre-trained encoder DINOv2 to encode the face image in the spectrogram; after encoding the image into a feature vector, use a fully connected deep neural network to classify the feature vector and output the facial expression category.

[0012] Furthermore, in step (3.1) of the present invention, aligning the spectrogram with the camera image through two-dimensional affine transformation specifically includes: Use four microphones to align the coordinate system of the microphone array with the camera coordinate system; Combined with the sound source localization method, calculate the positions of the four microphones in the spectrogram; Determine the corresponding positions of four microphones in the camera image through an image matching algorithm; Based on the corresponding points of the four microphones, arbitrarily select any three corresponding points to calculate the affine transformation matrix, and take the average of the four obtained affine transformation matrices to obtain the final transformation matrix.

[0013] A facial expression recognition system based on sound perception, adopting the facial expression method, the system includes: Signal sending module: The signal sending module emits sound signals of intermittent FMCW waveforms through a speaker. After the sound signals reach the human face, they are reflected to generate sound reflection signals, and the microphone array receives the sound reflection signals; Signal processing module: The signal processing module includes an audio frame segmentation sub-module, a distance measurement sub-module, a spectrogram generation sub-module, and a coordinate system conversion sub-module; The audio frame segmentation sub-module is used to extract the sound signals reflected by the face from the recordings of the microphone array; The distance measurement sub-module is used to calculate the distance between the human face and the sensor; The spectrogram generation sub-module generates a spectrogram containing facial information through the SRP-PHAT algorithm; The coordinate system rotation sub-module is used to convert the spectrogram from the spherical coordinate system to the Cartesian coordinate system; Expression classification module: Extract features from the spectrogram through a pre-trained deep learning model, input the extracted features into a fully connected deep neural network. The fully connected deep neural network outputs the facial expression classification result of the user by learning the mapping relationship between the features and the expression categories.

[0014] The beneficial technical effects of the present invention: 1) The facial expression recognition method provided by the present invention utilizes the multipath effect, adopts the SPR-PHAT algorithm to generate a spectrogram containing facial information, extracts facial features from multi-channel sound signals, can effectively and intuitively capture subtle facial dynamics, and can effectively avoid the spoofing problem of traditional image or video recognition methods.

[0015] 2) The facial expression recognition method provided by the present invention can realize facial expression recognition effectively without relying on expensive hardware devices, using commercial speakers and microphone arrays, even in the case of less data volume.

[0016] 3) The facial expression recognition method provided by the present invention combines acoustic imaging with a pre-trained visual model for the first time, and has broad development prospects. Description of the Drawings

[0017] Figure 1 It is a flowchart of a facial expression recognition method based on sound perception in an embodiment of the present invention; Figure 2 It is a physical diagram of an acoustic transceiver in an embodiment of the present invention; Figure 3 The signal positioning result in the embodiment of the present invention; Figure 4 The schematic diagram of converting from spherical coordinate system to Cartesian coordinate system in the embodiment of the present invention; Figure 5 The recognition results of five basic emotion expressions by the facial expression recognition method based on sound perception in the embodiment of the present invention; Figure 6 The change diagram of Loss during the training process of the facial expression recognition method based on sound perception in the embodiment of the present invention; Reference signs: 1. Development board; 2. Driver board; 3. Microphone array; 4. Speaker. Detailed implementation manners

[0018] In order to make the objectives, technical solutions and advantages of the present invention clearer and more understandable, the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present invention and are not used to limit the present invention.

[0019] On the contrary, the present invention covers any alternatives, modifications, equivalent methods and solutions made within the spirit and scope of the present invention defined by the claims. Further, in order to enable the public to have a better understanding of the present invention, some specific details are described in detail in the following detailed description of the present invention. Those skilled in the art can fully understand the present invention without the description of these details.

[0020] In the field of wireless perception, in order to meet the requirements of different devices and application scenarios, it is usually necessary to collect data by oneself, which is not only costly but also cumbersome. Therefore, how to achieve efficient emotion recognition under the conditions of low cost and small amount of data is a huge challenge. The facial expression recognition method provided by the present invention realizes expression recognition under the condition of scarce data.

[0021] Embodiment 1: A facial expression recognition method based on sound perception, the method uses acoustic signals as the perception medium, uses a microphone array to capture sound reflection signals, generates a spectrogram containing facial features through signal processing, extracts features from the spectrogram, and uses a deep learning network to classify facial expressions to achieve expression recognition.

[0022] As Figure 1 shown, the method in this embodiment includes the following steps: (1) Signal transmission: Transmit a sound signal with an FMCW waveform through a speaker, the sound signal is reflected after reaching the human face to generate a sound reflection signal, and the microphone array receives the sound reflection signal; (2)Signal processing: The signal processing process includes an audio frame segmentation sub-step, a distance measurement sub-step, a spectrogram generation sub-step, and a coordinate system conversion sub-step; The audio frame segmentation sub-step is used to extract the sound signal reflected by the face from the recording of the microphone array; The distance measurement sub-step is used to calculate the distance between the human face and the sensor; The spectrogram generation sub-step generates a spectrogram containing facial information through the SRP-PHAT algorithm; The coordinate system rotation sub-step is used to convert the spectrogram from the spherical coordinate system to the Cartesian coordinate system.

[0023] (3)Facial expression classification: Features are extracted from the spectrogram through a pre-trained deep learning model, and the extracted features are input into a fully connected deep neural network. The fully connected deep neural network outputs the facial expression classification result of the user by learning the mapping relationship between the features and the expression categories.

[0024] In step (1) of this embodiment, the length of the sound signal of the spaced FMCW waveform emitted by the speaker is 2 ms for a single pulse duration, and the frequency increases linearly in the range of 14 KHz to 15 KHz; an interval of 8 ms is set between adjacent pulses.

[0025] In the present invention, the method for designing the sound signal is: If the distance from the human face to the microphone array is d, the time required for the signal to return is: ; When d = 0.8 m, t is approximately 0.0047 s, that is, 4.7 ms; when d = 1 m, t is approximately 6 ms.

[0026] To reduce the overlap between the background reflection and the human face reflection signal, the pulse duration is set to 2 ms; to reduce the influence of the previous reflection signal on the next reflection signal, the interval is set to 8 ms; the selection of the frequency band from 14 KHz to 15 KHz is mainly affected by the spectral response of the speaker and the microphone array. Since it is an inexpensive device, the response in the high-frequency band is weak and unstable. Through experiments, 14 - 15 KHz is the frequency band with relatively good effects. Among them, the frequency band from 14 KHz to 15 KHz avoids the common environmental noise frequency bands (usually below 8 KHz), and noise can be removed through a band-pass filter. Also, considering the situation of using inexpensive speakers and microphone arrays, when the signal frequency exceeds 20 KHz for inexpensive speakers and microphone arrays, the power attenuation is obvious and the signal-to-noise ratio decreases, resulting in the inability to separate the human face reflection signal and background noise from the reflection signal. Considering these factors comprehensively, the present invention selects a pulse signal with a length of 2 ms, whose frequency increases linearly within the range of 14 KHz to 15 KHz. Each such signal segment is called a pulse, and an interval of 8 ms is set between adjacent pulses.

[0027] In step (1) of this embodiment, an acoustic transceiver composed of a speaker and a microphone array is used for acoustic sensing; as Figure 2 shown, the acoustic transceiver includes a development board 1, a driver board 2, a microphone array 3, and a speaker 4; Specifically, the size of the speaker is 5 cm × 5 cm, supporting 3.5 mm audio interfaces and USB interfaces, and can be compatible with a variety of sound output devices. The speaker is used for audio signals; the microphone array is composed of 8 microphone modules and is equipped with a driver board and a development board; the 8 microphone modules are arranged in a circular pattern, the array aperture is 42.5 mm, the outer diameter is 102 mm, and the inner diameter is 68.7 mm. All microphone modules are synchronized with the internal clock and can receive sound signals simultaneously. The microphone array is connected to the driver board through a 28P reverse flexible cable and samples the analog signal at a sampling rate of 192 kHz. The driver board is equipped with a 12V2A power interface, and the power indicator light is on when powered on. The development board can control each module of the acoustic transceiver simultaneously; the development board is equipped with a 24GB memory card and can store long-time sound data.

[0028] In step (2) of this embodiment, the direct propagation signal refers to the signal directly propagated from the speaker to the microphone array. Ideally, the direct propagation signal should be exactly the same as the transmitted signal and have the highest intensity.

[0029] The main reflected signal is a mixture of the reflected signals from the main surfaces of the face, such as the cheeks and forehead. Due to the different distances between other surfaces of the face, such as the nose and chin, and the microphone array, the echoes generated by these parts will vary in time, and may arrive at the microphone array earlier or later than the main reflected signal. Therefore, the reflected signal in the face area actually includes the combination of all these reflected signals, capturing all the information of the face.

[0030] To minimize the interference of other reflected signals and effectively extract the target information, it is crucial to accurately segment the reflected signal in the face area.

[0031] In step (2), the audio frame segmentation sub-step includes: Locate the direct propagation signal from the original recording of the microphone array; Locate the main reflected signal after the direct propagation signal, so as to determine the position of the facial reflected signal; After determining the position of the facial reflected signal from the recording of the microphone array, segment and extract the sound signal reflected by the face from the recording of the microphone array.

[0032] An intuitive method to locate the direct propagation signal is to play and record sounds simultaneously; however, since the playback and recording processes need to go through multiple layers of hardware and software processing in the operating system, and the delays of many of these processing links are unpredictable, it is not feasible to use a constant delay to locate the direct propagation signal. Since the direct propagation signal usually has the highest amplitude, the position of the direct propagation signal can be determined by the amplitude of the signal. The present invention uses a method from coarse to fine to locate the signal directly propagated from the speaker to the microphone.

[0033] First, use a sliding window method to find the starting position of the sound in the sound signal. Divide the signal into two adjacent and non-overlapping windows A and B, each window having a length of L. By traversing the signal, each time slide the two windows forward by one sampling point until window B exceeds the total length of the signal. In each sliding step, calculate the energy ratio of window B and window A. By comparing the energy ratios of these two windows, the mutation point of the signal energy can be detected, thus preliminarily determining the starting position of the sound.

[0034] After initially determining the candidate starting positions of the sound signal through the sliding window method, a fine search is further carried out around the candidate points to more accurately locate the starting position of the sound. For each candidate point, a signal segment of length L is intercepted starting from the candidate point, and the intercepted signal is mixed to obtain an intermediate frequency signal. Then, a fast Fourier transform (FFT) is performed on the intermediate frequency signal, and the obtained spectrum is centered to obtain the frequency domain signal. The DC component in the frequency domain signal is extracted, and the point with the largest DC component among all candidate points is selected as the starting position of the sound signal. By analyzing the changes in these DC components, the starting position of the sound signal is determined more precisely. Through the above method, the direct propagation signal found is as Figure 3 shown in the 0.095 - 0.097s part in

[0035] In the case of obtaining the direct propagation signal, the main reflection signal in the signal is identified by calculating the distance between the human face and the microphone array. The obtained main reflection signal is as Figure 3 shown in the 1 - 1.002s part in. In addition, since the human face is a three-dimensional structure rather than a plane, in the present invention, 50 sampling points are extended before and after the main reflection signal, so that the intercepted reflection signal covers the entire facial depth (about 9 cm). The main reflection signal and the extended sampling points will be used as inputs in the next step of signal processing (filtering and spectrogram generation (SRP - PHAT)).

[0036] In step (2) of this embodiment, the distance measurement sub-step specifically includes: Using the FMCW radar ranging algorithm, the received signal is mixed with the transmitted signal to generate an intermediate frequency signal; According to the relationship formula between distance and difference frequency: ; where R represents the target distance, c is the speed of light, Δf is the frequency difference, and μ is the slope. Through this formula, the distance between the target and the sensor can be calculated. Directly using this formula requires the device to usually send continuous FMCW signals. In the present invention, a blank time period is introduced between two FMCW pulses.

[0037] When measuring the distance, assuming that there is a continuous FMCW signal during the blank time period between two FMCW pulses, a sliding window is used to search backward step by step until a significant amplitude appears at a certain frequency in the intermediate frequency signal; The distance between the human face and the microphone array d is calculated by the following formula: ; where N represents the number of sliding windows, represents the maximum ranging distance of a single pulse, Indicates the distance measured in the current window. The distance d Specifically, it is the horizontal distance from the human face to the center of the circular microphone array; In step (2) of this embodiment, the spectrogram generation sub-step uses the SRP-PHAT algorithm to generate a spectrogram;

[0038] Facial expressions are minute and irregular deformations. Sound signals are continuous signals that vary over time and usually exist in the form of sound waves. In a real environment, sound signals are often affected by various noises, echoes, and other environmental interferences, making it more difficult to distinguish weak reflection signals and resulting in the inability to identify subtle changes in the human face. To achieve expression classification, it is necessary to extract the subtle movements of the human face from the sound signals. The facial expression recognition method provided by the present invention scans the face by emitting a specific sound signal through a speaker to generate an acoustic representation that can retain rich facial features, thereby capturing facial dynamic information. When sound waves reach the face, the reflections and scatterings in different regions will vary according to their shapes. The sound signal from one incident direction will be scattered to different reflection angles, generating multiple multipath components, and these multipath signals will ultimately be superimposed on each microphone. The energy of the reflected sound in each direction reflects the three-dimensional surface features of the face. The present invention uses the SRP-PHAT (Steered Response Power with Phase Transform) algorithm to calculate the scattering signal intensity in a specific direction according to the PHAT (Phase Weighted Delay and Sum) algorithm. The signal intensities of each pair of microphones at this position are accumulated to generate a fine-grained facial spectrogram. This spectrogram retains rich facial features and can be effectively used for facial recognition. This process accumulates the signal intensities of all microphones at the facial position.

[0039] In the present invention, only the spectrogram in the projection plane is required, rather than scanning the entire space. Therefore, by dynamically adjusting the scanning angle range and resolution according to the distance from the microphone array to the tester and the size of the drawing range, the calculation amount can be reduced and the efficiency can be improved. If the spatial angle range and resolution in the spectrogram are set too large, blank points may appear on the canvas; while if they are set too small, a large amount of repeated calculations will occur.

[0040] In the process of generating a spectrogram using the SRP-PHAT algorithm, the calculation range of the azimuth angle is [-180°, 180°], and the calculation range of the elevation angle is: ; Wherein, ; d is the distance between the human face and the microphone array; is the diagonal length of the drawn image; is the minimum pitch angle of the scanning range of the SRP-PHAT in space.

[0041] In the current experimental scenario, the plotting range of the azimuth angle θ is [-180°, 180°], so there is no need to dynamically adjust its range. Here, the pitch angle 's influence on the coordinates is mainly considered. When the pitch angle decreases, that is, when the angle between the ray and the z-axis increases, 's slope decreases, and the rate of change of the coordinate value of point M accelerates; in order to ensure that each pixel point in the image is covered, the minimum resolution is required.

[0042] Obviously, for the pixel points farthest along the diagonal direction of the square canvas, the required resolution is the smallest; the minimum resolution res of the pitch angle is: .

[0043] In step (2) of this embodiment, the coordinate system rotor step is used to convert the spectrogram from the spherical coordinate system to the Cartesian coordinate system; Convert the coordinates of the intersection point of a certain direction represented by the direction angle and the pitch angle in the spherical coordinate system with the projection plane into specific coordinates in the Cartesian coordinate system, so as to map the azimuth information in the polar coordinate system to the Cartesian coordinate system and generate an image that can clearly display the contour.

[0044] The spectrogram generated by the spectrogram generation sub-step uses the direction angle and the azimuth angle as coordinates. According to the measured distance from the human face to the microphone array, the coordinate system is converted, and the spectrogram is converted from the polar coordinate system to the Cartesian coordinate system. In the converted spectrogram, the units of the horizontal axis and the vertical axis are centimeters, which can more intuitively display the spatial information of facial expressions.

[0045] To achieve this conversion, first, the relative positions of the Cartesian coordinate system and the spherical coordinate system need to be clarified. In the present invention, the Cartesian coordinate system is obtained by horizontally translating the spherical coordinate system along the positive direction of the Z axis by a known distance. During the conversion process, the x-y plane of the Cartesian coordinate system is regarded as the projection plane. In order to generate the corresponding image on this projection plane, it is necessary to convert the coordinates of the intersection point of a certain direction represented by the direction angle θ and the pitch angle in the spherical coordinate system with this projection plane into specific coordinates in the Cartesian coordinate system. The conversion process is based on geometric relationships. Through the known direction angle, pitch angle, and the distance between the projection plane and the microphone array, the azimuth information in the polar coordinate system can be mapped to the Cartesian coordinate system.

[0046] The relationship between the two coordinate systems is as Figure 4 shown, where the right plane represents the projection plane. O 1 is the coordinate system where the microphone array is located.O 2 is the coordinate system where the imaging plane is located. O 2 The coordinate system is obtained by translating the O 1 coordinate system horizontally by the distance R between the projection plane and the microphone array. Therefore, x 1 the axis is parallel to the x 2 axis, y 1 the axis is parallel to the y 2 axis. The azimuth angle of the ray is θ, and the elevation angle is , and it intersects the plane at point m. Therefore, the line is perpendicular to both the plane and the plane. It can be seen from this that the abscissa and ordinate of point m are equal to the abscissa and ordinate of point n respectively.

[0047] ; ; In triangle , the included angle is a right angle. Therefore, the length of is: ; According to the geometric relationship, the coordinates of point m can be obtained as: ; This coordinate transformation enables the original spatial information based on the direction angle and elevation angle to be accurately projected onto the projection plane, generating an image that can clearly display the contour. In this way, the spatial data generated by beamforming can be visualized in the Cartesian coordinate system, generating an intuitive image that conforms to human spatial perception. Since it is assumed that the distance between the projection plane and the microphone array plane is relatively large, the small distance deviation caused by the facial undulation of the human face is ignored here.

[0048] Step (3) of this embodiment specifically includes: (3.1) Data preprocessing: Align the spectrogram with the camera image through two-dimensional affine transformation, use the face localization algorithm to determine the position of the face in the camera image, and then obtain the position of the face in the spectrogram; (3.2) Facial expression recognition: Encode the image using the pre-trained encoder DINOv2; after encoding the image into a feature vector, use a fully connected deep neural network to classify the feature vector and output the facial expression category.

[0049] In a spectrogram, similar to the imaging principle of a camera, it contains information in all directions within a region. Currently, the spectrogram does not have sufficient information to precisely distinguish a human face from regions very close to the face such as the forehead and neck. In step (3.1) of the present invention, in order to obtain the position of the human face, the pixel points in the generated spectrogram are aligned and matched with the pixel points in the camera photo. This process requires the assistance of camera calibration technology to achieve precise spatial mapping between the spectrogram coordinate system and the camera coordinate system under the same reference framework. In the present invention, the coordinate system where the microphone array is located is set as the world coordinate system. By fixing the relative spatial positions of the camera and the microphone array, it is ensured that the camera pose and position remain constant throughout the experiment, thereby maintaining the stability of the external parameter matrix. At the same time, the distance between the camera and the object being photographed is controlled to be basically the same to ensure the consistency of the camera internal parameters during multiple shootings. The present invention uses an affine transformation to map the x-y coordinates in the world coordinate system to the x-y coordinates in the camera coordinate system, and the present invention sets the z-direction coordinates of all points to a unified value, that is, the z value in the world coordinates corresponding to the point represented by each pixel in the spectrogram is the same. This assumption is based on the fact that the camera and the microphone array are located on the same plane, and the distance between the camera and the human face is much larger than the local concavity and convexity changes on the human face surface. Therefore, the influence of the change in the human face geometry on the z value of each point in the image can be ignored. Based on this, the present invention further assumes that in the camera coordinate system, the depth information of all image points remains consistent, that is, the z value in the camera coordinate system is the same. In the experiment, the camera and the microphone array are fixed in the same plane to ensure the stability of their relative positions and orientations, so that the relative transformation between the world coordinate system and the camera coordinate system only includes rotation around the z axis and translation along the x-y plane.

[0050] The internal parameter matrix satisfies the form of an affine transformation matrix, and the process of projecting from the camera coordinate system to the image plane and sampling can also be regarded as an affine transformation. Therefore, the mapping process from three-dimensional space to the image can be regarded as the multiplication of two affine transformation matrices, and the final result is still an affine transformation matrix. An affine transformation has six degrees of freedom. Each pair of two-dimensional coordinate points can provide two linear equations, and three pairs of two-dimensional points can determine an affine transformation matrix. In order to reduce experimental errors and improve the transformation accuracy, four microphones are used in the present invention to align the coordinate system of the microphone array with the camera coordinate system; Combined with the sound source localization method, calculate the positions of the four microphones in the spectrogram; Determine the corresponding positions of the four microphones in the camera image through an image matching algorithm; Based on the corresponding points of the four microphones, select any three corresponding points from them to calculate the affine transformation matrix, and take the average value of the four obtained affine transformation matrices to obtain the final transformation matrix.

[0051] This step method effectively reduces the deviation caused by single measurement errors, improving the stability and accuracy of the correction algorithm in practical applications. After aligning the spectrogram with the camera image, a mature face localization algorithm is used to determine the position of the face in the camera image, thereby obtaining the position of the face in the spectrogram.

[0052] In step (3.2) of the present invention, the pre-trained model DINOv2 is adopted as the feature extractor. DINOv2 is trained through self-supervised learning on a dataset of approximately 1 billion images, capable of generating highly generalizable feature representations and applicable to various image tasks such as depth detection, image segmentation, and image generation.

[0053] In the present invention, the decoder in the fully connected deep neural network consists of three fully connected layers and two ReLU activation layers alternatingly. This decoder is used to classify the 384-dimensional features output by DINOv2. It consists of three fully connected layers and two ReLU activation layers alternatingly, and the specific structure is as follows: the dimension of the first fully connected layer is 384×256, the second is 256×128, and the third is 128×n, where n represents the number of emotion categories. A ReLU activation layer is connected after each fully connected layer to enhance the network's non-linear expression ability.

[0054] The present invention experimentally evaluates the classification performance of the facial expression recognition method based on sound perception provided by the present invention for five basic emotion expressions (i.e., happy, angry, disgusted, neutral, surprised). The experimental results are as Figure 5 shown. The experimental results show that the method proposed by the present invention can effectively recognize the emotion expressions of users, and the overall average recognition accuracy reaches 74%. The change of Loss during the training process is as Figure 6 shown. It can be seen that as the number of training rounds increases, the Loss value gradually decreases and tends to be stable, indicating that the facial expression recognition method based on sound perception provided by the present invention can effectively learn the features of emotion expressions during the training process and gradually converge to a better state.

[0055] Embodiment 2: A facial expression recognition system based on sound perception, adopting the facial expression method described in Embodiment 1. The system includes: Signal sending module: The signal sending module emits a sound signal with an FMCW waveform through a speaker. After the sound signal reaches the human face and is reflected, a sound reflection signal is generated, and the microphone array receives the sound reflection signal; Signal processing module: The signal processing module includes an audio frame segmentation sub-module, a distance measurement sub-module, a spectrogram generation sub-module, and a coordinate system conversion sub-module; The audio frame segmentation sub-module is used to extract the sound signal reflected by the face from the recording of the microphone array; The distance measurement sub-module is used to calculate the distance between the human face and the sensor; The spectrogram generation sub-module generates a spectrogram containing facial information through the SRP-PHAT algorithm; The coordinate system rotation sub-module is used to convert the spectrogram from the spherical coordinate system to the Cartesian coordinate system.

[0056] Facial expression classification module: Extract features from the spectrogram through a pre-trained deep learning model, input the extracted features into a fully connected deep neural network. The fully connected deep neural network outputs the facial expression classification result of the user by learning the mapping relationship between the features and the expression categories.

[0057] The facial expression recognition system provided by the present invention aims to achieve high-precision facial expression recognition through inexpensive devices in scenarios with scarce data while protecting user privacy.

[0058] The above are only the preferred embodiments of the present invention and are not intended to limit the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principles of the present invention shall be included within the protection scope of the present invention.

Claims

1. A facial expression recognition method based on sound perception, characterized in that, The method uses acoustic signals as the perception medium, captures sound reflection signals using a microphone array, generates a spectrogram containing facial features through signal processing, extracts features from the spectrogram, and classifies facial expressions using a deep learning network to achieve expression recognition.

2. The facial expression recognition method based on sound perception according to claim 1, characterized in that, The method includes the following steps: (1) Signal transmission: Transmit sound signals of an intermittent FMCW waveform through a speaker. After the sound signals reach the human face, they are reflected to generate sound reflection signals, and the microphone array receives the sound reflection signals; (2) Signal processing: The signal processing process includes an audio frame segmentation sub-step, a distance measurement sub-step, a spectrogram generation sub-step, and a coordinate system conversion sub-step; The audio frame segmentation sub-step is used to extract the sound signals reflected by the face from the recording of the microphone array; The distance measurement sub-step is used to calculate the distance between the human face and the microphone array; The spectrogram generation sub-step generates a spectrogram containing facial information through the SRP-PHAT algorithm; The coordinate system rotation sub-step is used to convert the spectrogram from the spherical coordinate system to the Cartesian coordinate system; (3) Expression classification: Extract features from the spectrogram through a pre-trained deep learning model, input the extracted features into a deep neural network. The deep neural network outputs the facial expression classification result of the user by learning the mapping relationship between the features and the expression categories.

3. The facial expression recognition method based on sound perception according to claim 2, wherein In step (2), the audio frame segmentation sub-step includes: Locate the direct propagation signal from the original recording of the microphone array; Locate the main reflection signal after the direct propagation signal to determine the position of the facial reflection signal; After determining the position of the facial reflection signal from the recording of the microphone array, segment and extract the sound signals reflected by the face from the recording of the microphone array.

4. The facial expression recognition method based on sound perception according to claim 2, wherein In step (2), the distance measurement sub-step specifically includes: Use the FMCW radar ranging algorithm to mix the received signal with the transmitted signal to generate an intermediate frequency signal; When measuring the distance, assume that there is a continuous FMCW signal during the blank time period between two FMCW pulses, and gradually search backward through a sliding window until a significant amplitude appears at a certain frequency in the intermediate frequency signal; The distance between the human face and the microphone array d is calculated by the following formula: ; Among them, N represents the number of sliding windows, represents the maximum ranging distance of a single pulse, represents the distance measured in the current window.

5. The method for facial expression recognition based on sound perception according to claim 2, wherein In step (2), the spectrogram generation sub-step uses the SRP-PHAT algorithm to generate a spectrogram; In the process of generating spectrograms using the SRP-PHAT algorithm, the calculation range of the azimuth angle is [-180°, 180°], and the calculation range of the elevation angle is: ; Among them, ; d is the distance between the human face and the microphone array; is the diagonal length of the drawn image; is the minimum value of the pitch angle of the scanning range of the SRP-PHAT in space; The minimum resolution res of the elevation angle is: 。 6. The facial expression recognition method based on sound perception according to claim 2, wherein In step (2), the coordinate system rotation sub-step is used to convert the spectrogram from the spherical coordinate system to the Cartesian coordinate system; Convert the coordinates of the intersection point of a certain direction represented by the azimuth angle and the elevation angle in the spherical coordinate system with the projection plane into specific coordinates in the Cartesian coordinate system, so as to map the azimuth information in the polar coordinate system to the Cartesian coordinate system and generate an image that can clearly display the contour.

7. The facial expression recognition method based on sound perception according to claim 2, characterized in that Step (3) specifically includes: (3.1) Data preprocessing: Align the spectrogram with the camera image through a two-dimensional affine transformation, use a face localization algorithm to determine the position of the face in the camera image, and then obtain the position of the face in the spectrogram; (3.2)Facial expression recognition: Use the pre-trained encoder DINOv2 to encode the face image in the spectrogram; after encoding the image into a feature vector, use a fully-connected deep neural network to classify the feature vector and output the facial expression category.

8. The method for facial expression recognition based on sound perception according to claim 2, characterized in that In step (3.1) of the present invention, the spectrogram is aligned with the camera image through a two-dimensional affine transformation, specifically: Four microphones are used to align the coordinate system of the microphone array with the camera coordinate system; Combined with the sound source localization method, calculate the positions of the four microphones in the spectrogram; Determine the corresponding positions of the four microphones in the camera image through an image matching algorithm; Based on the corresponding points of the four microphones, select any three corresponding points to calculate the affine transformation matrix, and take the average of the four obtained affine transformation matrices to obtain the final transformation matrix.

9. A facial expression recognition system based on sound perception, which adopts the facial expression recognition method described in any one of claims 1-8, characterized in that, The system includes: Signal sending module: The signal sending module emits an intermittent FMCW waveform sound signal through a speaker. After the sound signal reaches the human face and is reflected, a sound reflection signal is generated, and the microphone array receives the sound reflection signal; Signal processing module: The signal processing module includes an audio frame segmentation sub-module, a distance measurement sub-module, a spectrogram generation sub-module, and a coordinate system conversion sub-module; the audio frame segmentation sub-module is used to extract the sound signal reflected by the face from the recording of the microphone array; the distance measurement sub-module is used to calculate the distance between the human face and the sensor; the spectrogram generation sub-module generates a spectrogram containing facial information through the SRP-PHAT algorithm; the coordinate system conversion sub-module is used to convert the spectrogram from the spherical coordinate system to the Cartesian coordinate system; Facial expression classification module: Extract features from the spectrogram through a pre-trained deep learning model, input the extracted features into a fully-connected deep neural network, and the fully-connected deep neural network outputs the facial expression classification result of the user by learning the mapping relationship between the features and the expression categories.

Citation Information

Patent Citations

  • Avatar facial expression and / or speech driven animations

    CN107431635A

  • Emotion analysis method and device, equipment and storage medium

    CN113468983A

  • Facial information acquisition method and electronic equipment

    CN119181126A

  • A face-identifying device

    CN201084168Y

  • Face identification using millimeter-wave radar sensor data

    US20220327863A1