Voice directional enhancement method and near-to-eye display equipment
By waking up the microphone array and camera with a bone conduction microphone, and combining gaze direction detection and voice orientation enhancement methods, the problem of accurately capturing user voice in noisy environments using traditional microphone pickup methods has been solved. This achieves low-power, high-precision directional voice pickup and improves the voice interaction performance of near-eye display devices.
Patent Information
- Application Number
- CN202511420247.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-29
- Publication Date
- 2025-12-12
AI Technical Summary
In noisy environments or multi-party dialogue scenarios, traditional microphone pickup methods struggle to accurately capture the sound sources that users are interested in, affecting the quality of voice interaction and user experience. Furthermore, frequent calibration operations lead to high power consumption and latency issues.
A bone conduction microphone is used to detect sound signals and wake up the microphone array and camera. Combined with gaze direction detection and speech orientation enhancement methods, a mask is constructed through directional and audio features to achieve directional enhancement of the target sound source, reduce the power consumption of unnecessary components, and improve the accuracy of sound pickup.
It achieves low-power, high-precision directional voice pickup, improves the voice interaction performance of near-eye display devices in real and complex environments, reduces latency and power consumption, and improves the efficiency and accuracy of voice interaction.
Smart Images

Figure CN121122301A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of directional sound pickup technology, and in particular to a method for directional speech enhancement and a near-eye display device. Background Technology
[0002] With the rapid development of augmented reality and virtual reality technologies, the demand for near-eye display devices is increasing. Some near-eye display devices utilize voice pickup technology to achieve voice interaction functions. This is especially important in noisy environments or multi-party dialogue scenarios, where it is crucial to accurately capture the voice of the speaker currently conversing with the user.
[0003] However, speech signals are often accompanied by various background noises, such as environmental noise, reverberation, or the voices of other speakers, which can easily cause confusion and affect the directional sound pickup effect. Summary of the Invention
[0004] In view of this, some embodiments of this application provide a speech direction enhancement method and a near-eye display device, which can accurately directionally enhance the target sound source and improve the sound pickup effect.
[0005] In a first aspect, some embodiments of this application provide a speech direction enhancement method, including: acquiring the current frame and gaze direction of mixed speech, wherein the mixed speech is acquired by a microphone array; constructing directional features based on the current frame and gaze direction of the mixed speech, and constructing audio features based on the current frame of the mixed speech; determining a current frame mask based on the directional features and audio features; and determining the target enhanced speech of the current frame based on the current frame mask, the current frame of the mixed speech, and multiple historical frames of the mixed speech.
[0006] In this embodiment, since the directional features are constructed from the current frame of the mixed speech and the gaze direction, these features are used to determine the current frame mask. This ensures that the current frame mask is generated under the guidance of the gaze direction, which facilitates accurate separation of the target speech from the mixed speech. Furthermore, the target enhanced speech for the current frame is determined based on the current frame mask, the current frame of the mixed speech, and multiple historical frames of the mixed speech. That is, by combining the current frame mask and the current frame of the mixed speech to distinguish between noise and target speech in the current frame, and referring to multiple historical frames, target speech enhancement is performed frame by frame based on the overall features of the mixed speech. This effectively improves the accuracy of the target enhanced speech in the current frame, thereby improving the accuracy of directional sound pickup.
[0007] In one or more embodiments, the aforementioned construction of directional features based on the current frame of mixed speech and gaze direction includes: obtaining the actual phase difference between every two speech channels in the current frame of mixed speech, and obtaining the theoretical phase difference between every two speech channels under the gaze direction; and constructing directional features based on the actual phase difference between every two speech channels and the theoretical phase difference between every two speech channels.
[0008] In this embodiment, directional features are constructed based on the actual phase difference between every two speech channels and the theoretical phase difference between every two speech channels. Since the directional features reflect the gaze direction and the audio features reflect the energy distribution of the mixed speech in the current frame, the current frame mask can be accurately determined based on the directional features and the audio features.
[0009] In one or more embodiments, the aforementioned acquisition of the theoretical phase difference between every two speech channels in the gaze direction includes: determining the beam steering vector of each speech channel based on the gaze direction; and determining the theoretical phase difference between every two speech channels in the gaze direction based on the beam steering vectors of every two speech channels.
[0010] In this embodiment, since the beam steering vector describes the relative phase relationship when sound waves arrive at each microphone in the array from a specific direction, the theoretical phase difference between each pair of voice channels in the gaze direction can be accurately determined based on the beam steering vector of each pair of voice channels.
[0011] In one or more embodiments, the aforementioned construction of directional features based on the actual phase difference between each pair of speech channels and the theoretical phase difference between each pair of speech channels includes: calculating the difference between the actual phase difference and the theoretical phase difference between each pair of speech channels to obtain multiple phase difference distances; and normalizing the multiple phase difference distances to obtain directional features. In this embodiment, the directional features constructed in the above manner reflect the degree to which the target sound source is close to the gaze direction. The directional features are used as input to a mask model to calculate the current frame mask, allowing the mask model to learn the degree to which the target sound source is close to the gaze direction, thus guiding the mask model to output an accurate current frame mask.
[0012] In one or more embodiments, the aforementioned determination of the target enhanced speech for the current frame based on the current frame mask, the current frame of the mixed speech, and multiple historical frames of the mixed speech includes: Based on the current frame mask and the current frame of the mixed speech, determine the noise speech spectrum corresponding to the current frame of the mixed speech; based on the noise speech spectrum corresponding to the current frame of the mixed speech, the noise speech spectrum corresponding to multiple historical frames of the mixed speech, and the gaze direction, determine the weight of the current frame; use the weight of the current frame to perform weighted processing on the current frame of the mixed speech and multiple historical frames of the mixed speech to obtain the target enhanced speech of the current frame.
[0013] In this embodiment, considering the noise information corresponding to the current frame of the mixed speech and the noise information corresponding to multiple historical frames of the mixed speech, and combining the gaze direction, the weight of the current frame is determined to make the weight of the current frame accurate. The current frame weight is used to weight the current frame of the mixed speech and multiple historical frames of the mixed speech, taking into account the preceding information, so that the target enhanced speech of the current frame is more accurate.
[0014] In one or more embodiments, the aforementioned determination of the current frame weight based on the noise speech spectrum corresponding to the current frame of the mixed speech, the noise speech spectrum corresponding to multiple historical frames of the mixed speech, and the gaze direction includes: calculating a noise covariance matrix using the noise speech spectrum corresponding to the current frame of the mixed speech and the noise speech spectrum corresponding to multiple historical frames of the mixed speech; determining the beam steering vector of each speech channel in the current frame of the mixed speech based on the gaze direction; and determining the current frame weight based on the noise covariance matrix and the beam steering vector of each speech channel in the current frame of the mixed speech.
[0015] In this embodiment, the noise covariance matrix includes the noise information statistics of the current frame and multiple historical frames. The beam steering vector of each speech channel in the current frame of the mixed speech describes the relative phase relationship when the sound source arrives at each microphone in the microphone array from the gaze direction. The weight of the current frame is determined based on both, so that the weight of the current frame is determined after considering the noise information statistics of the current frame and multiple historical frames as well as the relative phase relationship when the sound source arrives at each microphone in the microphone array from the gaze direction, which has high accuracy.
[0016] In one or more embodiments, the aforementioned determination of the current frame weight based on the noise covariance matrix and the beam steering vectors of each speech channel in the current frame of the mixed speech includes: normalizing the product of the inverse of the noise covariance matrix and the beam steering vectors of each speech channel in the current frame of the mixed speech to obtain the current frame weight.
[0017] In this embodiment, by multiplying the inverse of the noise covariance matrix with the vector matrix formed by the beam steering vectors of each speech channel, the noise covariance matrix is applied to the beam steering vectors of each speech channel. Then, normalization is performed, which helps to ensure unity gain.
[0018] In one or more embodiments, the method further includes: caching processing data corresponding to the current frame of the mixed speech, so that the processing data can be applied to speech direction enhancement of the next frame of the mixed speech, wherein the processing data includes a noise covariance matrix. In this embodiment, the processing data corresponding to each frame is cached after processing. When performing speech direction enhancement on the next frame of the mixed speech, the processing data can be called without repeated calculation, thus realizing streaming inference and real-time caching. Therefore, it can effectively avoid recalculating the entire segment, improve computational efficiency, facilitate low-latency continuous audio processing, and is suitable for streaming speech enhancement scenarios.
[0019] Secondly, some embodiments of this application provide a near-eye display device, including: at least one processor, and a memory communicatively connected to the at least one processor, wherein the memory stores instructions executable by the at least one processor, the instructions being executed by the at least one processor to enable the at least one processor to perform the method in the first aspect.
[0020] Thirdly, some embodiments of this application provide a computer-readable storage medium storing computer-executable instructions for causing a computer device to perform the method as described in the first aspect.
[0021] The above description is only an overview of the technical solution of this application. In order to better understand the technical means of this application and to implement it in accordance with the contents of the specification, and to make the above and other objects, features and advantages of this application more obvious and understandable, the following are specific embodiments of this application. Attached Figure Description
[0022] One or more embodiments are illustrated by way of example with reference numerals in the accompanying drawings. These illustrations do not constitute a limitation on the embodiments. Elements with the same reference numerals in the drawings are denoted as similar elements. Unless otherwise stated, the figures in the drawings are not to be limited by scale.
[0023] Figure 1 This is a schematic diagram of the structure of the near-eye display device 100 in some embodiments of this application; Figure 2 This is a flowchart illustrating the method for detecting gaze direction in some embodiments of this application; Figure 3 This is a schematic diagram of a sub-process of the method for detecting gaze direction in some embodiments of this application; Figure 4 This is a schematic diagram of a sub-process of the method for detecting gaze direction in some embodiments of this application; Figure 5 This is a schematic diagram of a sub-process of the method for detecting gaze direction in some embodiments of this application; Figure 6 This is a schematic diagram of the gaze direction detection model in some embodiments of this application; Figure 7 This is a schematic diagram of the structure of the near-eye display device 200 in some embodiments of this application; Figure 8 This is a flowchart illustrating the speech direction enhancement method in some embodiments of this application; Figure 9 This is a schematic diagram of the sub-processes of the speech direction enhancement method in some embodiments of this application; Figure 10 This is a schematic diagram of another sub-process of the speech direction enhancement method in some embodiments of this application; Figure 11 This is a schematic diagram of another sub-process of the speech direction enhancement method in some embodiments of this application; Figure 12 This is a schematic diagram of another sub-process of the speech direction enhancement method in some embodiments of this application; Figure 13 This is a schematic diagram of the structure of the near-eye display device 300 in some embodiments of this application; Figure 14 This is a schematic diagram illustrating the process of directional sound pickup in a near-eye display device in some embodiments of this application. Detailed Implementation
[0024] The present application will now be described in detail with reference to specific embodiments. These embodiments will help those skilled in the art to further understand the present application, but do not limit the present application in any way. It should be noted that those skilled in the art can make several modifications and improvements without departing from the concept of the present application. These all fall within the protection scope of the present application.
[0025] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0026] It should be noted that, unless there is a conflict, the various features in the embodiments of this application can be combined with each other, all of which are within the protection scope of this application. Furthermore, although functional modules are divided in the device schematic diagram and a logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than the module division in the device or the order in the flowchart. In addition, the terms "first," "second," and "third" used herein do not limit the data or execution order, but only distinguish identical or similar items with essentially the same function and effect.
[0027] Unless otherwise defined, all technical and scientific terms used in this specification have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used in this specification is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The term "and / or" as used in this specification includes any and all combinations of one or more of the associated listed items.
[0028] Furthermore, the technical features involved in the various embodiments of this application described below can be combined with each other as long as they do not conflict with each other.
[0029] With the rapid development of augmented reality and wearable devices, near-eye display devices with voice interaction capabilities are gaining popularity. In noisy environments or multi-party conversations, traditional microphone pickup methods struggle to accurately capture the user's preferred sound source, impacting voice interaction quality and user experience. Near-eye display devices can be AR glasses, AR headsets, VR glasses, or VR headsets, etc. This application does not impose any limitations on the form of the near-eye display device; any wearable device that enables near-field visual presentation is acceptable.
[0030] To capture the sounds that users are focused on, some known solutions, based on the assumption of a fixed head position, use a static mapping matrix to map eye image information into a gaze direction vector to capture sound in the gaze direction. This mapping matrix needs to be obtained through initial or multiple calibrations after the user wears the device. However, once the head posture changes, the original calibration parameters become invalid, causing a drift error in the gaze direction estimation.
[0031] Other solutions known to the inventors of this application employ an inertial measurement unit (IMU) to detect head posture in order to improve sound pickup accuracy, triggering calibration when the head posture changes. However, frequent recalibration interrupts the audio data stream, causing stuttering. Consequently, it is difficult to simultaneously meet the requirements of low latency and high accuracy in real-time voice interaction.
[0032] In addition, the above solution relies on the IMU to detect head posture in real time, which frequently triggers calibration and results in high power consumption.
[0033] Other solutions known to the inventors of this application use the gaze direction as the input to the beamforming module for directional sound pickup. However, gaze direction drift error is further propagated to the beamforming module, causing the main lobe of the beam to deviate from the actual direction of the sound source of interest, severely affecting the signal-to-noise ratio and quality of the picked-up speech.
[0034] Based on this, some embodiments of this application provide a near-eye display device that can achieve highly robust directional voice pickup with low power consumption, which is beneficial for achieving efficient and accurate voice interaction.
[0035] Other embodiments of this application provide a method for detecting gaze direction, which can achieve high-precision estimation of gaze direction.
[0036] Other embodiments of this application provide a speech direction enhancement method that can accurately enhance the target sound source based on the gaze direction described above, thereby achieving directional sound pickup.
[0037] The aforementioned methods for detecting gaze direction and enhancing speech direction are applied to near-eye display devices, enabling them to achieve low-power, high-precision directional sound pickup. This improves the voice interaction performance of near-eye display devices in real-world, complex environments.
[0038] To facilitate understanding, the hardware structure of the near-eye display device in some embodiments of this application will be introduced first.
[0039] Near-eye display devices Some embodiments of this application provide a near-eye display device; please refer to [link / reference]. Figure 1 The near-eye display device 100 includes a frame 10, a bone conduction microphone 20, a camera 30, a microphone array 40, and a controller (not shown).
[0040] The frame 10 includes a frame 11 and temples 12. The temples 12 are connected to the frame 11. For example, one end of the temples 12 can be rotatably connected to one end of the frame 11 via a connecting shaft, so that the temples 12 can switch between an unfolded state and a folded state.
[0041] The eyeglass frame 11 includes a left frame body 111 and a right frame body 112, which are connected by a connecting part 113. The left frame body 111 and the right frame body 112 respectively house lenses. The left frame body 111 corresponds to the wearer's left eye, and the right frame body 112 corresponds to the wearer's right eye. The connecting part 113 between the left frame body 111 and the right frame body 112 corresponds to the wearer's bridge of the nose.
[0042] A left nose pad 51 is provided on the left frame 111, close to the connecting part 113; a right nose pad 52 is provided on the right frame 112, also close to the connecting part 113. That is, the left nose pad 51 and right nose pad 52 are located on both sides of the connecting part 113. When the near-eye display device 100 is worn, the left nose pad 51 and right nose pad 52 fit snugly against the wearer's nose bridge to ensure the wearing stability of the near-eye display device 100.
[0043] A bone conduction microphone 20 is mounted on the frame 11. The bone conduction microphone 20 is a sound pickup device that captures sound signals through bone vibrations. In some embodiments, the bone conduction microphone 20 has a size of less than 6mm x 6mm and a thickness of no more than 2.5mm. The bone conduction microphone 20 is configured to detect sound signals. When the bone conduction microphone 20 detects a sound signal, it indicates that the wearer is making a sound or speaking, and the near-eye display device 100 needs to enter a voice interaction state.
[0044] A camera 30 is mounted on the frame 11 and is configured to capture images of the wearer's eyes. In some embodiments, there are two cameras 30. One camera 30 is mounted on the left frame 111 and is referred to as the left camera; the other camera 30 is mounted on the right frame 112 and is referred to as the right camera. The fields of view of the two cameras 30 cover both eyes of the wearer, which is beneficial for capturing images of the eyes.
[0045] A microphone array 40 is mounted on the temple 12, for example, at one end of the temple 12 near the frame 11. In some embodiments, the microphone array 40 includes four microphones arranged in a matrix array or a circular array. The microphone array 40 is configured to acquire audio data. This audio data is audio data from the environment in which the near-eye display device 100 is located.
[0046] The controller is configured to run a directional sound pickup model. This model is used for directional sound pickup based on eye images and audio data. Specifically, the controller's memory stores the directional sound pickup model, and the controller's processor invokes this model to perform directional sound pickup based on eye images and audio data. In some embodiments, the controller and other circuitry are encapsulated within the temple 12.
[0047] The near-eye display device 100 is configured to wake up the microphone array 40, camera 30, and controller in response to the bone conduction microphone 20 detecting a sound signal, thereby enabling directional sound pickup of audio data. In some embodiments, after detecting a sound signal, the bone conduction microphone 20 sends a signal to the controller. Upon receiving the signal, the controller performs a wake-up action, waking up the microphone array 40, camera 30, and directional sound pickup model. The microphone array 40 collects audio data and sends it to the controller, the camera 30 collects eye images and sends them to the controller, and the processor in the controller calls the directional sound pickup model to perform directional sound pickup based on the eye images and audio data.
[0048] In some embodiments, the near-eye display device 100 further includes other slave controllers (not shown), defining a controller configured with a directional sound pickup model as the master controller. After the bone conduction microphone 20 detects a sound signal, it sends a signal to the slave controller. Upon receiving the signal, the slave controller wakes up the microphone array 40, the camera 30, and the master controller.
[0049] It is worth noting that the embodiments of this application do not impose any limitations on the wake-up methods of the microphone array 40, the camera 30, and the controller.
[0050] When the aforementioned near-eye display device 100 is worn, the bone conduction microphone 20 is fitted against the wearer. Since the bone conduction microphone collects sound signals transmitted through bones or human tissue, when the bone conduction microphone detects a sound signal, it indicates that the wearer is making a sound or speaking, and the near-eye display device 100 needs to enter the voice interaction state. At this time, the microphone array 40, camera 30, and controller are activated. The camera 30 collects images of the wearer's eyes, the microphone array 40 collects audio data, and the controller runs a directional sound pickup model, which performs directional sound pickup based on the eye images and audio data.
[0051] In this embodiment, the bone conduction microphone 20 continuously listens for sound signals from the wearer. Only after the bone conduction microphone 20 detects a sound signal are the microphone array 40, camera 30, and controller activated to collaboratively perform directional sound pickup. Compared to scenarios where multiple or all of the bone conduction microphone 20, microphone array 40, camera 30, and controller are continuously operating, in this embodiment only the bone conduction microphone 20 is continuously listening, while other components operate intermittently. This results in lower power consumption for the near-eye display device 100 during use, thereby enabling dynamic switching from low-power listening to directional sound pickup. Combined with a directional sound pickup model, this achieves highly robust directional voice pickup with lower power consumption, facilitating efficient and accurate voice interaction.
[0052] In some embodiments, the near-eye display device 100 is also configured to be in a sleep state in response to the bone conduction microphone 20 not detecting a sound signal for a first duration, the microphone array 40, the camera 30, and the controller.
[0053] If the bone conduction microphone 20 does not detect a sound signal for a specified duration, it indicates that the wearer has not engaged in voice interaction with others for the duration, and the voice interaction may end. The first duration can be 30 seconds or 1 minute. Those skilled in the art can set the first duration according to the actual scenario; this application does not impose any limitations on the first duration.
[0054] In some embodiments, if the bone conduction microphone 20 does not detect a sound signal for a first duration, it sends a signal to the controller. Upon receiving the signal, the controller puts the microphone array 40, camera 30, and directional sound pickup model into a sleep state. Thus, the near-eye display device 100 is in a low-power monitoring state. In some embodiments, if the bone conduction microphone 20 does not detect a sound signal for a first duration, the controller puts the microphone array 40, camera 30, and main controller into a sleep state.
[0055] In this embodiment, if the bone conduction microphone 20 does not detect a sound signal for a first duration, it indicates that voice interaction has stopped. The microphone array 40, camera 30, and controller enter a sleep state, while the bone conduction microphone 20 continues to listen, maintaining a low power consumption state.
[0056] In some embodiments, if the microphone array 40 does not collect audio data for a second duration, the microphone array 40, camera 30, and controller are put into a sleep state.
[0057] If microphone array 40 fails to collect audio data for a second duration, it indicates that the wearer's environment has remained silent for the second duration, and the voice interaction may end. This second duration can be 1 minute or 2 minutes. Those skilled in the art can set the second duration according to the actual scenario; this application does not impose any limitations on the second duration.
[0058] Thus, if the microphone array 40 fails to collect audio data for a second duration, it indicates that voice interaction has stopped. The microphone array 40, camera 30, and controller enter a sleep state, while the bone conduction microphone 20 continues to monitor and maintain a low power consumption state.
[0059] In some embodiments, the bone conduction microphone 20 is mounted on the nose pads 51 / 52 of the eyeglass frame 11, and the bone conduction microphone 20 is configured to fit snugly against the wearer's nose bridge. For example... Figure 1 As shown, the bone conduction microphone 20 is mounted on the left nose pad 51 and / or the right nose pad 52. Exemplarily, the bone conduction microphone 20 is mounted on the surface of the left nose pad 51 facing the wearer's nose bridge. This facilitates the accurate detection of sound signals from the wearer by the bone conduction microphone 20.
[0060] In the above solution, since the bone conduction microphone 20 is installed on the nose pads 51 / 52 of the frame 11, when the near-eye display device 100 is worn, the bone conduction microphone 20 fits snugly against the wearer's nose bridge, stably acquiring the wearer's bone vibration signals, thereby accurately detecting sound signals. This allows subsequent wake-up actions to be performed promptly, reducing delays and facilitating real-time interaction.
[0061] In some embodiments, the bone conduction microphone 20 is bonded to the nose pads 51 / 52 of the eyeglass frame 11. For example, the bone conduction microphone is bonded to the left nose pad 51 and / or the right nose pad 52 using silicone adhesive.
[0062] In the above solution, an adhesive is used to bond the bone conduction microphone 20 to the nose pads 51 / 52, ensuring a stable contact coupling between the bone conduction microphone 20 and the wearer's nose bridge, which is beneficial for accurately detecting whether the wearer is speaking. Furthermore, using adhesive to fix the bone conduction microphone 20 results in a simple and stable structure.
[0063] In some embodiments, the bone conduction microphone 20 is mechanically mounted to the nose pad 51 / 52. Exemplarily, the mechanical structure may be a U-shaped spring, the open end of which conforms to the contour of the nose pad 51 / 52, and the closed end which is integrally formed with the bone conduction microphone 20; or, the bone conduction microphone 20 may be welded to the closed end of the U-shaped spring.
[0064] In some embodiments, the surface of the bone conduction microphone 20 is covered with a flexible silicone sleeve. This flexible silicone sleeve is elastic and adaptable to the shape of the nose bridge of different face shapes. Thus, on the one hand, the flexible silicone sleeve protects the bone conduction microphone; on the other hand, compared to the bone conduction microphone 20 directly contacting the nose bridge, the elastic flexible silicone sleeve can press firmly against the nose bridge, adapting to different face shapes and improving the stability of the contact coupling.
[0065] In some embodiments, such as Figure 1 As shown, the camera 30 includes a left camera 31 and a right camera 32. The left camera 31 is located below the left nose pad 51, and the right camera 32 is located below the right nose pad 52. It can be understood that when the near-eye display device 100 is worn, the nose pads 51 / 52 are generally located below the eyes, and the camera 30 is located below the nose pads 51 / 52, so that the camera 30 will not obstruct the eyes and will not affect the wearer's vision.
[0066] In this embodiment, two cameras 30 cover the wearer's eyes, which facilitates accurate tracking of the wearer's gaze direction and enables the pickup of sound in the direction of the wearer's gaze, i.e., directional sound pickup. In addition, the cameras 30 are positioned below the nose pads 51 / 52, so as not to interfere with the wearer's field of vision.
[0067] In some embodiments, the left camera 31 is tilted upwards and configured to face the wearer's right eye; the right camera 32 is tilted upwards and configured to face the wearer's left eye. That is, the left camera 31 is used to capture an image of the right eye, and the right camera 32 is used to capture an image of the left eye.
[0068] In this way, the two cameras 31 / 32 capture images in an interleaved manner, and the two cameras do not obstruct the lens, making the structure of the near-eye display device 100 compact.
[0069] In some embodiments, the angle between the central axis of the left camera 31 and the plane containing the temple and the angle between the central axis of the right camera 32 and the plane containing the temple are both α, where 30° ≤ α ≤ 60°. Optionally, 40° ≤ α ≤ 50°. In some embodiments, α can be 30°, 34°, 38°, 40°, 46°, 50°, 53°, 57°, 60°, or a range of any two of these values.
[0070] Not intending to be limited by any theory or explanation, the inventors unexpectedly discovered that when 30°≤α≤60°, the camera 30 can be positioned well to capture eye images. When α is less than 30° or greater than 60°, the camera 30 cannot capture eye images or the captured eye images are of poor quality.
[0071] This angle design allows the two cameras 30 to face the wearer's eye area, covering both eyes without interfering with the wearer's vision, while also giving the cameras 30 good stability and concealment.
[0072] In some embodiments, both the left camera 31 and the right camera 32 are lensless cameras. The lensless camera modulates light from different directions using an optical mask, processing the received signals into a superimposed coded image of multiple light channels. Because the lensless camera records a coded image, the near-eye display device 100 can protect the wearer's privacy.
[0073] In summary, the near-eye display device 100 in this embodiment introduces a bone conduction microphone 20 to detect sound signals. Since the bone conduction microphone 20 collects sound signals transmitted through bones or human tissue, when the bone conduction microphone 20 detects a sound signal, it indicates that the wearer is making a sound or speaking, and the near-eye display device 100 needs to enter a voice interaction state. At this time, the microphone array 40, camera 30, and controller are activated. The camera 30 collects images of the wearer's eyes, the microphone array 40 collects audio data, and the controller runs a directional sound pickup model. The directional sound pickup model performs directional sound pickup based on the eye images and audio data.
[0074] The bone conduction microphone 20 continuously listens for voice signals from the wearer. Only after the bone conduction microphone 20 detects a voice signal are the microphone array 40, camera 30, and controller activated to collaboratively perform directional sound pickup. Compared to scenarios where multiple or all of the bone conduction microphone 20, microphone array 40, camera 30, and controller are continuously operating, in this embodiment only the bone conduction microphone 20 is continuously listening, while other components operate intermittently. This results in lower power consumption for the near-eye display device 100 during use, enabling dynamic switching from low-power listening to directional sound pickup. Combined with a directional sound pickup model, this achieves highly robust directional voice pickup with lower power consumption, facilitating efficient and accurate voice interaction.
[0075] It is understood that a directional sound pickup model is an algorithmic model that picks up sound from a specific direction. Those skilled in the art can program and design directional sound pickup models and deploy them in near-eye display devices. In some embodiments, the directional sound pickup model includes a gaze direction determination module and a speech direction enhancement module, achieving eye-tracking-based directional sound pickup by directionally enhancing the target sound source in the gaze direction.
[0076] The following section will first introduce the method for detecting gaze direction in the gaze direction determination module.
[0077] Methods for detecting gaze direction Some embodiments of this application provide a method for detecting gaze direction, the method comprising acquiring an eye image, the eye image being an coded image; reconstructing the eye image to obtain a reconstructed image; and inputting the reconstructed image into a pre-trained gaze direction detection model to obtain the gaze direction.
[0078] In this embodiment, the eye image is an coded image. Image reconstruction is performed on the coded image to restore eye features, enabling the gaze direction detection model to recognize eye features in the reconstructed image and thus accurately detect the gaze direction based on these features. Furthermore, the coded image effectively protects user privacy, thereby improving the accuracy of gaze direction detection while protecting the wearer's privacy.
[0079] In some embodiments of this application, the method for detecting gaze direction is applied to a near-eye display device. The near-eye display device includes a processor and a memory, the method for detecting gaze direction is stored in the memory, and the processor executes the method for detecting gaze direction to realize gaze direction detection.
[0080] See Figure 2 , Figure 2 A flowchart illustrating a method for detecting gaze direction provided in an embodiment of this application is shown below. Figure 2 As shown, method A100 includes the following steps: A10, acquire the eye image, which is an encoded image.
[0081] The eye image can be acquired by the near-eye display device described in the above embodiments. The wearer wears the near-eye display device, and a camera within the device acquires an image of the wearer's eyes. This eye image is an encoded image, meaning a digital image that has undergone compression and encoding. Encoded images protect the wearer's privacy and save storage space.
[0082] A20 performs image reconstruction on the eye image to obtain the reconstructed image.
[0083] Since the aforementioned eye images are multi-channel coded images, a reconstructed image for gaze estimation needs to be generated through image reconstruction. The reconstructed image reflects eye features, thus enabling the gaze direction detection model to identify eye features in the reconstructed image and accurately detect the gaze direction based on these features.
[0084] In some embodiments, please refer to Figure 3 Step A20 specifically includes: A21 decomposes the eye image to obtain the high-frequency feature map and low-frequency feature map of the eye image.
[0085] Among them, the high-frequency feature map of an eye image refers to the feature map that reflects the rapid and subtle changes in details, edges, textures, or noise in the eye image. The low-frequency feature map of an eye image refers to the feature map that reflects the overall smoothness of large areas of color, gradations of brightness and darkness, contours, etc. in the eye image, including the structure and main features of the image.
[0086] In this step, a decomposition algorithm can be used to decompose the eye image and extract high-frequency and low-frequency feature maps. In some embodiments, please refer to... Figure 4 Step A21 specifically includes: A211 uses a low-pass filter to perform convolution processing on each row of the eye image to obtain a low-frequency feature row map.
[0087] The low-pass filter is a filtering function that preserves the low-frequency components of the signal, where the low-frequency components correspond to slowly changing parts of the signal, such as smooth regions in an eye image. In some embodiments, the low-pass filter may be a Haar filter.
[0088] A low-pass filter performs a convolution operation on each row of the eye image, assigning a convolution weight coefficient to each pixel. This results in pixels corresponding to low-frequency features having larger convolution weight coefficients, while pixels corresponding to high-frequency features have smaller convolution weight coefficients. This preserves low-frequency features and attenuates high-frequency features in the eye image, resulting in a low-frequency feature row map. The low-frequency components recorded in the low-frequency feature row map correspond to features such as the scleral background color, smooth iris regions, and eyelid contours in the eye image. Those skilled in the art will understand that the specific calculation process of the low-pass filter can be found in related technologies and will not be described in detail here.
[0089] A212 uses a high-pass filter to convolve each row of the eye image to obtain a high-frequency feature row map.
[0090] A high-pass filter is a filtering function that extracts the high-frequency components of a signal. These high-frequency components correspond to rapidly changing parts of the signal, such as edges or textures in an eye image. In some embodiments, the high-pass filter may be a Daubechies filter.
[0091] The high-pass filter performs a convolution operation on each row of the eye image, assigning a convolution weight coefficient to each pixel in the eye image. This results in pixels corresponding to high-frequency features having larger convolution weight coefficients, while pixels corresponding to low-frequency features have smaller convolution weight coefficients. Thus, it preserves high-frequency features in the eye image while attenuating low-frequency features, resulting in a high-frequency feature row map. Those skilled in the art will understand that the specific calculation process of the high-pass filter can be found in related technologies and will not be described in detail here.
[0092] A213 samples each row of the low-frequency feature row plot at intervals, and samples each row of the high-frequency feature row plot at intervals.
[0093] In some embodiments, "sampling each row of the low-frequency feature row plot at intervals" specifically includes: retaining the even-numbered columns in the low-frequency feature row plot and removing the odd-numbered columns to obtain the sampled low-frequency feature row plot. "Sampling each row of the high-frequency feature row plot at intervals" specifically includes: retaining the even-numbered columns in the high-frequency feature row plot and removing the odd-numbered columns to obtain the sampled high-frequency feature row plot.
[0094] In some embodiments, "sampling each row of the low-frequency feature row plot at intervals" specifically includes: retaining the odd columns in the low-frequency feature row plot and removing the even columns to obtain the sampled low-frequency feature row plot. "Sampling each row of the high-frequency feature row plot at intervals" specifically includes: retaining the odd columns in the high-frequency feature row plot and removing the even columns to obtain the sampled high-frequency feature row plot.
[0095] A214 uses a low-pass filter to convolve each column of the sampled low-frequency feature map to obtain a low-frequency feature map, which reflects the low-frequency features of the eye image.
[0096] The low-pass filter performs a convolution operation on each column of the sampled low-frequency feature map, assigning a convolution weight coefficient to each pixel in the sampled low-frequency feature map. This results in pixels corresponding to low-frequency features having larger convolution weight coefficients, while pixels corresponding to high-frequency features have smaller convolution weight coefficients. This preserves the low-frequency features in the sampled low-frequency feature map while attenuating the high-frequency features, thus obtaining the low-frequency feature map. Therefore, the low-frequency feature map is obtained by low-pass filtering the eye image in both row and column directions, effectively preserving low-frequency features while removing high-frequency features.
[0097] A215 uses a low-pass filter to convolve each column of the sampled high-frequency feature map to obtain a horizontal high-frequency feature map, which reflects the detailed features of the eye image in the horizontal direction.
[0098] The low-pass filter performs convolution processing on each column of the sampled high-frequency feature row map, assigning a convolution weight coefficient to each pixel in the sampled high-frequency feature row map. This results in pixels corresponding to low-frequency features having larger convolution weight coefficients, while pixels corresponding to high-frequency features have smaller convolution weight coefficients. This preserves low-frequency features in the column direction of the sampled high-frequency feature row map while attenuating high-frequency features in the column direction. Therefore, the horizontal high-frequency feature map is obtained by performing high-pass filtering on the rows and low-pass filtering on the columns of the eye image. The horizontal high-frequency feature map effectively preserves the detailed features in the horizontal direction of the eye image, i.e., the vertical edge features of the eye image.
[0099] A216 uses a high-pass filter to convolve each column of the sampled low-frequency feature map to obtain a vertical high-frequency feature map, which reflects the detailed features of the eye image in the vertical direction.
[0100] The high-pass filter performs a convolution operation on each column of the sampled low-frequency feature row map, assigning a convolution weight coefficient to each pixel in the sampled low-frequency feature row map. This results in pixels corresponding to high-frequency features having larger convolution weight coefficients and pixels corresponding to low-frequency features having smaller convolution weight coefficients, thus preserving high-frequency features in the column direction of the sampled low-frequency feature row map while attenuating low-frequency features in the column direction. Consequently, the vertical high-frequency feature map is obtained by performing low-pass filtering on the rows and high-pass filtering on the columns of the eye image. The vertical high-frequency feature map effectively preserves the detailed features in the vertical direction of the eye image, i.e., the horizontal edge features of the eye image.
[0101] A217 uses a high-pass filter to convolve each column of the sampled high-frequency feature map to obtain a diagonal high-frequency feature map, which reflects the detailed features in the diagonal direction of the eye image.
[0102] The high-pass filter performs a convolution operation on each column of the sampled high-frequency feature row map, assigning a convolution weight coefficient to each pixel in the sampled high-frequency feature row map. This results in pixels corresponding to high-frequency features having larger convolution weight coefficients, while pixels corresponding to low-frequency features have smaller convolution weight coefficients. This preserves high-frequency features in the column direction of the sampled high-frequency feature row map while attenuating low-frequency features in the column direction. Consequently, the diagonal high-frequency feature map is obtained by performing high-pass filtering on both the row and column directions of the eye image. The diagonal high-frequency feature map effectively preserves the detailed features in the diagonal direction of the eye image, i.e., the pixel edge features in the diagonal direction of the eye image.
[0103] It is understandable that the above decomposition method can be expressed by the following formula:
[0104] Where Y represents the eye image, It is a low-frequency characteristic. High-frequency characteristics The decomposition function is used to process the eye image, as described in steps A211 to A217 above.
[0105] In this embodiment, low-frequency feature maps, horizontal high-frequency feature maps, vertical high-frequency feature maps, and diagonal high-frequency feature maps are obtained through steps A211 to A217 described above. The low-frequency features of the eye image include the low-frequency feature map, and the high-frequency features of the eye image include the horizontal high-frequency feature map, vertical high-frequency feature map, and diagonal high-frequency feature map. This ensures that the high-frequency features include high-frequency detail features from multiple directions, making the high-frequency features more comprehensive and beneficial for the subsequent gaze direction detection model to determine the gaze direction by recognizing these high-frequency detail features.
[0106] A22 performs noise reduction on high-frequency features, and the low-frequency features and the noise-reduced high-frequency features form a reconstructed image.
[0107] Understandably, since high-frequency features include rapidly changing and subtle features such as details, edges, textures, or noise, they may interfere with gaze direction detection. Therefore, noise reduction processing is performed on high-frequency features to remove noise from them.
[0108] In some embodiments, please refer to Figure 5 Step A22 includes the following steps: A221. If the convolution weight coefficient corresponding to a pixel in the high-frequency feature map is less than the threshold, then the convolution weight coefficient is set to zero.
[0109] A222: If the convolution weight coefficient corresponding to a pixel in the high-frequency feature map is greater than or equal to the threshold, then the convolution weight coefficient is reduced.
[0110] The threshold is a pre-set threshold for the convolution weight coefficients. In some embodiments, this threshold can be determined based on the pixel features of the eye image.
[0111] As described above, when performing convolution processing, the filter assigns convolution weight coefficients to pixels. To preserve detailed features and remove noise, convolution weight coefficients in the high-frequency feature map that are less than a threshold are set to zero. Pixels with convolution weight coefficients less than the threshold are considered noise pixels. Setting the convolution weight coefficients of these noise pixels to zero eliminates them, thus removing noise. Convolution weight coefficients in the high-frequency feature map that are greater than or equal to the threshold are reduced to suppress the amplitude of outlier noise and avoid edge breaks in high-frequency features due to zeroing. Reduction refers to decreasing the absolute value of convolution weight coefficients greater than or equal to the threshold; for example, a convolution weight coefficient of -5 is reduced by -2. This reduces the absolute value of all non-zero convolution weight coefficients in the high-frequency feature map, thereby reducing the convolution weight coefficients of outlier noise pixels and suppressing the amplitude of outlier noise. Furthermore, reducing the absolute value of all non-zero convolution weight coefficients in the high-frequency feature map, rather than setting them to zero, effectively avoids edge breaks in high-frequency features.
[0112] In some embodiments, steps A221 and A222 described above can be characterized by the following formula:
[0113] in, The first high-frequency feature Line 1 The convolution weights of the column pixels, The high-frequency features after noise reduction processing Line 1 The convolution weights of the column pixels, T is the threshold, and the sign() function is used to extract them. The positive and negative signs.
[0114] For example, =-5, T=3, =2, because 2>0, so max(0,2)=2, sing( )=-1, -1*2=-2.
[0115] The low-frequency features and the denoised high-frequency features form the reconstructed image, which is:
[0116] in, To reconstruct the image, It is a low-frequency characteristic. The high-frequency features are denoised. The "+" operator indicates that the pixel values at the same pixel location are added channel by channel. For example, at the pixel position in row i and column j, the pixel values of the R channel are added, the pixel values of the G channel are added, and the pixel values of the B channel are added. The denoised high-frequency and low-frequency features have the same size. The pixel value at the pixel position in row i and column j of the denoised high-frequency features is added channel by channel to the pixel value at the pixel position in row i and column j of the low-frequency features to obtain the pixel value at the pixel position in row i and column j of the reconstructed image. In this way, the low-frequency features reflecting the overall contour and the high-frequency features reflecting clean details are fused into the reconstructed image.
[0117] In this embodiment, the eye image is first decomposed to obtain high-frequency and low-frequency features. Then, the high-frequency features are denoised, and the denoised high-frequency and low-frequency features form a reconstructed image. This results in a reconstructed image with less noise, highlighting the eye features, which is beneficial for the gaze direction detection model to recognize eye features and output an accurate gaze direction.
[0118] A30 inputs the reconstructed image into a pre-trained gaze direction detection model to obtain the gaze direction.
[0119] The gaze direction detection model is a pre-trained neural network model. In some embodiments, the gaze direction detection network is iteratively trained using several reconstructed images labeled with gaze directions until the gaze direction detection network converges, thus obtaining the gaze direction detection model.
[0120] The gaze direction detection network is a pre-configured neural network with neural network components (convolutional layers, deconvolutional layers, pooling layers, or activation functions, etc.). The basic structure and principles of neural networks are well-known in the field of machine learning and will not be described in detail here.
[0121] Here, several reconstructed images are used to iteratively train the gaze direction detection network. The model parameters at convergence are then configured as those of the gaze direction detection network to obtain the gaze direction detection model. Those skilled in the art will understand that the model training process can refer to relevant techniques; however, the specific training and parameter tuning process will not be described in detail here.
[0122] In some embodiments, the gaze direction detection model includes a feature extraction module and a gaze regression module. For example... Figure 6 As shown, from the input to the output of the gaze direction detection model, it includes a feature extraction module and a gaze regression module.
[0123] In some embodiments, the feature extraction module includes a cascaded histogram equalization layer, a normalization layer, and a convolutional network. The feature extraction module is configured to extract features from the input image and output a feature map. The reconstructed image is input to the feature extraction module, first undergoes histogram equalization, then normalization, and finally is input to the convolutional network for feature extraction, outputting a feature map.
[0124] Histogram equalization layers are used to perform histogram equalization on the input reconstructed image, increasing the gray-level spacing or making the gray levels more uniform, thereby increasing gray-level contrast and making the image clearer. Normalization layers are used to normalize the input image to simplify the computation of image reconstruction.
[0125] Convolutional networks, consisting of multiple convolutional layers, are used to extract features from images after histogram equalization and normalization, enabling the extraction of detailed features. This improves the gaze direction detection model's ability to perceive image details.
[0126] The gaze regression module is configured to map the feature map output by the feature extraction module to the gaze direction. Specifically, the gaze regression module processes the feature map using the following formula: , ,
[0127] in, For feature maps, This represents the weight vector corresponding to each feature in the gaze direction detection model. for The dimension of , where d is the feature dimension. This represents the bias corresponding to each feature in the gaze direction detection model. For the dimension of b, This is the gaze direction vector.
[0128] In this embodiment, the histogram equalization layer increases grayscale contrast, making the image clearer; the normalization layer simplifies computation. The convolutional network extracts features from the image after histogram equalization and normalization, enabling the extraction of detailed features. Therefore, the gaze direction detection model's ability to perceive image details is improved, thus increasing detection accuracy.
[0129] In some embodiments, the loss function configured for training the gaze direction detection network includes a cosine similarity loss function. The cosine similarity loss function is shown in the following formula: =1-cos(θ) in, For gaze direction prediction labels, θ represents the true label for the gaze direction, and θ is the angle between the predicted label for the gaze direction and the true label for the gaze direction.
[0130] In this embodiment, compared to the Euclidean distance loss function, the cosine similarity loss function is more consistent with the geometric features of the gaze direction, which is beneficial to improving the accuracy of the gaze direction detection model.
[0131] In some embodiments, the method A100 further includes: The A40 uses an adaptive filter to enhance the details of the reconstructed image.
[0132] In this embodiment, an adaptive filter is introduced to improve the local details of the reconstructed image.
[0133] The adaptive filter's processing of the reconstructed image can be expressed by the following formula:
[0134] in,
[0135] in, It is the pixel mean of the local window corresponding to the pixel position in the i-th row and j-th column of the reconstructed image. It is the pixel value of the pixel in the i-th row and j-th column of the reconstructed image before filtering. It is the pixel value of the pixel in the i-th row and j-th column of the reconstructed image after filtering. It is the dynamic weighting factor corresponding to the pixel position in the i-th row and j-th column of the reconstructed image, which is determined by the local gradient. and variance control.
[0136] The local gradient at the pixel position in the i-th row and j-th column is used to measure whether the region where the pixel is located is a smooth region or a detailed region; The larger the value, the more drastic the pixel value change. This parameter is used to suppress smooth features in edge regions to preserve detailed features. The pixel variance within the local window corresponding to the pixel position in the i-th row and j-th column is used to measure texture complexity. The larger the pixel variance, the more noise or texture there is. This is a gradient weight adjustment factor that controls the degree of influence of the gradient in weight calculation. The larger the value, the more the adaptive filter emphasizes edge features; This is the variance weight adjustment factor, used to control the degree of influence of variance in weight calculation. The larger the value, the more the adaptive filter emphasizes reducing smoothing features in complex areas such as texture or noise.
[0137] After the reconstructed image is processed by the aforementioned adaptive filter, the image noise is further reduced and the detailed features such as the edge texture of the pupil region are enhanced, which is beneficial for subsequent gaze direction estimation.
[0138] In this embodiment, step A30 specifically includes: A31 inputs the reconstructed image after detail enhancement into a pre-trained gaze direction detection model to obtain the gaze direction.
[0139] In this embodiment, by performing detail enhancement processing on the reconstructed image, the gaze direction detection model is able to identify detailed features, thereby improving the detection accuracy of the gaze direction detection model.
[0140] In summary, the embodiments of this application reconstruct eye images (coded images) to restore eye features, enabling the gaze direction detection model to recognize eye features in the reconstructed image and thus accurately detect the gaze direction based on these features. Furthermore, the coded image effectively protects the wearer's privacy, thereby improving the gaze direction detection accuracy while protecting the wearer's privacy. In some embodiments, the eye image is decomposed, and the obtained high-frequency features are denoised. The low-frequency features and the denoised high-frequency features form the reconstructed image. This results in a reconstructed image with less noise, highlighting eye features and facilitating feature recognition by the gaze direction detection model, thereby improving the model's detection accuracy. In some embodiments, an adaptive filter is used to filter the reconstructed image, further reducing image noise and enhancing detailed features. Using the filtered reconstructed image as input data for the gaze direction detection model further improves the model's detection accuracy.
[0141] The method for detecting gaze direction provided in this application embodiment can be implemented by various types of near-eye display devices, such as AR glasses or VR glasses.
[0142] See Figure 7 , Figure 7 This is a schematic diagram of the structure of a near-eye display device provided in an embodiment of this application. The near-eye display device 200 includes a processor 201 and a memory 202. The memory 202 is connected to the processor 201, for example, via a bus.
[0143] Processor 201 is configured to support the near-eye display device 200 in performing the corresponding functions in the methods described in the above-described method embodiments. Processor 201 may be a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof. The aforementioned hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The aforementioned PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0144] Memory 202 is used to store program code, etc. Memory 202 may include volatile memory (VM), such as random access memory (RAM); memory 202 may also include non-volatile memory (NVM), such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); memory 202 may also include combinations of the above types of memory.
[0145] The memory 202 is used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the method for detecting gaze direction in the embodiments of this application. The processor 201 executes various functional applications and data processing of the method for detecting gaze direction by running the non-volatile software programs, instructions, and modules stored in the memory 202, thereby realizing the function of the method for detecting gaze direction provided in the above method embodiments.
[0146] Memory 202 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and applications required for at least one function. In some embodiments, the memory may include memory remotely configured relative to the processor. Examples of the aforementioned networks include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0147] The one or more modules are stored in the memory. When executed by the one or more processors, they perform the method for detecting gaze direction in any of the above method embodiments, for example, the method steps described in the above method embodiments.
[0148] This application also provides a computer-readable storage medium storing a computer program, the computer program including program instructions, which, when executed by a computer, cause the computer to perform the method for detecting gaze direction as described in the foregoing embodiments.
[0149] The gaze direction reflects the direction the wearer focuses on, and the sound source in the gaze direction is also the sound source the wearer focuses on. After obtaining the gaze direction, the speech direction enhancement module in the directional pickup model captures the target sound source in the gaze direction and performs directional enhancement using the speech direction enhancement method in this embodiment to achieve directional pickup. The speech direction enhancement method applied to the speech direction enhancement module in this embodiment is described below.
[0150] Speech-oriented enhancement methods The speech direction enhancement method provided in some embodiments of this application first obtains the current frame and gaze direction of mixed speech, which is acquired by a microphone array. Directional features are constructed based on the current frame and gaze direction of the mixed speech, and audio features are constructed based on the current frame of the mixed speech. A current frame mask is determined based on the directional features and audio features. The target speech to be enhanced in the current frame is determined based on the current frame mask, the current frame of the mixed speech, and multiple historical frames of the mixed speech.
[0151] In this embodiment, since the directional features are constructed from the current frame of the mixed speech and the gaze direction, these features are used to determine the current frame mask. This ensures that the current frame mask is generated under the guidance of the gaze direction, which facilitates accurate separation of the target speech from the mixed speech. Furthermore, the target enhanced speech for the current frame is determined based on the current frame mask, the current frame of the mixed speech, and multiple historical frames of the mixed speech. That is, by combining the current frame mask and the current frame of the mixed speech to distinguish between noise and target speech in the current frame, and referring to multiple historical frames, target speech enhancement is performed frame by frame based on the overall features of the mixed speech. This effectively improves the accuracy of the target enhanced speech in the current frame, thereby improving the accuracy of directional sound pickup.
[0152] The speech-oriented enhancement method is applied to near-eye devices. These devices execute this method to achieve directional sound pickup.
[0153] See Figure 8 , Figure 8 This is a flowchart illustrating a speech direction enhancement method provided in an embodiment of this application. Figure 8As shown, method S100 includes the following steps: S10, obtain the current frame and gaze direction of the mixed speech.
[0154] The mixed speech is acquired by a microphone array, such as the microphone array of the near-eye display device in the above embodiment. The microphone array includes four microphone units, resulting in four channels of audio data for the mixed speech. After acquisition, the mixed speech is divided into multiple audio frames based on time. For example, each 25ms segment of the mixed speech is Fourier transformed to obtain one audio frame, resulting in a 25ms frame length. For a total duration of 1 second of mixed speech, dividing it into 25ms frames yields 40 audio frames.
[0155] The near-eye display device performs frame-by-frame speech-oriented enhancement on the mixed speech, where the current frame refers to the audio frame being processed. The gaze direction is the direction of the wearer's eye focus. In some embodiments, the gaze direction can be obtained based on eye image detection using the gaze direction detection method described in the above embodiments. For details, please refer to the embodiments of the gaze direction detection method described above; they will not be repeated here.
[0156] S20: Construct directional features based on the current frame of mixed speech and gaze direction, and construct audio features based on the current frame of mixed speech.
[0157] Here, directional features refer to features that incorporate the gaze direction. In some embodiments, the gaze direction can be extracted using a convolutional layer module to obtain directional features. In some embodiments, the gaze direction can be fused with the beam steering vectors of each speech channel in the current frame of the mixed speech to obtain directional features.
[0158] Audio features encompass time-domain features, frequency-domain features, time-frequency-domain features, and auditory perception features, including features such as logarithmic power spectrum, phase difference, or spectrum. In some embodiments, audio features can be extracted from the current frame of the mixed speech.
[0159] In some embodiments, please refer to Figure 9 The aforementioned step S20 specifically includes: S21, obtain the actual phase difference between every two speech channels in the current frame of the mixed speech, and obtain the theoretical phase difference between every two speech channels in the gaze direction.
[0160] Since the microphone array includes M microphones, the mixed speech has M speech channels, with one microphone corresponding to one speech channel. The current frame of the mixed speech has M speech channels. It is understandable that, due to the different positions and directions of each microphone relative to the sound source, the time and phase of the sound arriving at each microphone from the same sound source are also different. Therefore, there is a phase difference between different speech channels in the current frame of the mixed speech.
[0161] In some embodiments, the actual phase difference between any two speech channels in the current frame of the mixed speech is calculated using the following formula.
[0162]
[0163] in, It is the actual phase difference between the m-th and n-th speech channels. It is the complex spectrum of the m-th speech channel at the f-th frequency point in the t-th frame. yes The conjugate of the complex number, symbol It is complex number multiplication, symbol The phase angle of a complex number.
[0164] In some embodiments, the actual phase difference between each pair of voice channels can be calculated using the above formula. For example, taking M as 4, the actual phase difference between each pair of voice channels includes: the actual phase difference between the first and second voice channels, the actual phase difference between the first and third voice channels, the actual phase difference between the first and fourth voice channels, the actual phase difference between the second and third voice channels, the actual phase difference between the second and fourth voice channels, and the actual phase difference between the third and fourth voice channels.
[0165] It is understandable that any two speech channels also have a theoretical phase difference in the gaze direction. The theoretical phase difference between two speech channels in the gaze direction is the theoretical phase difference between the two speech channels in the gaze direction.
[0166] In some embodiments, the aforementioned “obtaining the theoretical phase difference between every two speech channels in the gaze direction” includes: determining the beam steering vector of each speech channel based on the gaze direction; and determining the aforementioned theoretical phase difference between every two speech channels in the gaze direction based on the beam steering vectors of every two speech channels.
[0167] The beam steering vector describes the relative phase relationship of sound waves as they arrive at each microphone in the array from a specific direction. Therefore, the beam steering vector indicates how to adjust the phase of the speech signal in each speech channel so that speech signals from the gaze direction can be added in phase during superposition, facilitating subsequent enhancement of the sound from that gaze direction, while speech signals from other directions may be suppressed due to out-of-phase conditions.
[0168] The direction of gaze is represented by the vector g. In some embodiments, the direction of gaze will be... The projection is the target horizontal angle θ∈[0°, 180°]. The beam steering vector is constructed using the following formula. Each voice channel corresponds to a beam steering vector.
[0169]
[0170] Where k is the imaginary unit and p is the frequency of the speech signal at the current moment; This indicates that the first voice channel is for the direction from which the voice comes. The time delay of sound waves, This indicates that the (M-1)th voice channel is related to the direction. The time delay of the sound waves. Those skilled in the art will understand that the microphone that first captures the speech signal is the reference microphone. The distances between the other microphones and the reference microphone are calculated, and the distance divided by the speed of sound yields the corresponding time delay, for example... This is the result of dividing the distance between the first microphone and the reference microphone by the speed of sound. This is the result of dividing the distance between the (M-1)th microphone and the reference microphone by the speed of sound.
[0171] Then, the theoretical phase difference is determined based on the beam steering vector of each pair of voice channels. In some embodiments, the theoretical phase difference is... ,in, It is the beam steering vector of the m-th voice channel. It is the complex conjugate of the beam steering vector of the nth voice channel.
[0172] In this embodiment, since the beam steering vector describes the relative phase relationship when sound waves arrive at each microphone in the array from a specific direction, the theoretical phase difference between each pair of voice channels in the gaze direction can be accurately determined based on the beam steering vector of each pair of voice channels.
[0173] S22, constructs directional features based on the actual phase difference between every two speech channels and the theoretical phase difference between every two speech channels.
[0174] Understandably, every two speech channels correspond to one actual phase difference and one theoretical phase difference. For example, the m-th speech channel and the n-th speech channel have one actual phase difference. and a theoretical phase difference .
[0175] In some embodiments, please refer to Figure 10 The above step S22 specifically includes: S221, calculate the difference between the actual phase difference and the theoretical phase difference between every two speech channels to obtain multiple phase difference distances.
[0176] S222 normalizes multiple phase difference distances to obtain directional features.
[0177] The phase difference distance reflects how close the target sound source is to the direction of gaze. Understandably, the smaller the phase difference distance, the more likely the target sound source is in the direction of gaze. In other words, the closer the actual phase difference is to the theoretical phase difference, the more likely the target sound source is in the direction of gaze.
[0178] To facilitate the processing of the phase difference distance input mask model, in this embodiment, each phase difference distance is normalized and mapped to the range [-1, 1].
[0179] In some embodiments, the directional features are calculated using the following formula. ;
[0180] in, It is the actual phase difference between the m-th and n-th speech channels. It is the theoretical phase difference between the m-th and n-th speech channels, where M is the number of speech channels. As a directional feature, This includes the phase difference distance between every two voice channels in the M voice channels.
[0181] In this embodiment, directional features are constructed in the manner described above, enabling these features to reflect the degree to which the target sound source is close to the gaze direction. Using these directional features to guide the generation of the current frame mask ensures that the degree to which the target sound source is close to the gaze direction is taken into account during mask generation, thus improving the accuracy of the current frame mask.
[0182] In some embodiments, the logarithmic power spectrum of the current frame of the mixed speech is extracted as an audio feature.
[0183] Logarithmic power spectrum represents the power spectrum of an audio signal in logarithmic form, presenting a linear relationship between frequency and intensity, which facilitates the analysis of the energy distribution of the signal in different frequency bands.
[0184] Here, the logarithmic power spectrum of the current frame of the mixed speech is used as an audio feature, which facilitates the mask model to learn the energy distribution of the current frame of the mixed speech in different frequency bands and calculate and output an accurate current frame mask.
[0185] In this embodiment, directional features are constructed based on the actual phase difference between every two speech channels and the theoretical phase difference between every two speech channels, and the logarithmic power spectrum of the current frame of the extracted mixed speech is used as audio features. Since the directional features reflect the gaze direction and the audio features reflect the energy distribution of the current frame of the mixed speech, the directional features and audio features are used to guide the generation of the current frame mask, making the current frame mask accurate.
[0186] S30, determine the current frame mask based on directional and audio features.
[0187] Since the directional features are constructed from the current frame of the mixed speech and the gaze direction, the directional features are used to determine the current frame mask. Thus, the current frame mask is generated under the guidance of the gaze direction, which is beneficial for accurately separating the target speech from the mixed speech.
[0188] In some embodiments, directional features and audio features are input into a pre-trained mask model, enabling the gaze direction to be learned by the mask model. The gaze direction guides the mask model to output an accurate current frame mask. Those skilled in the art will understand that the current frame mask is a matrix of the same size as the spectrum of the current frame of the mixed speech, with weight values between [0,1]. The current frame mask assigns a weight value to each time-frequency unit of the current frame of the mixed speech; a weight value close to 1 indicates that the corresponding time-frequency unit is retained, while a weight value close to 0 indicates that the corresponding time-frequency unit is suppressed.
[0189] The mask model is a model obtained by training a neural network using training data. In some embodiments, the neural network is a lightweight neural network with an encoder and a decoder. In some embodiments, the training data includes several sets of directional features and audio features, each set of directional features and audio features labeled with a real mask.
[0190] Therefore, when applying the mask model for mask estimation, the above-mentioned directional features and audio features can be input into the mask model to obtain the mask for the current frame.
[0191] S40, determine the target enhanced speech for the current frame based on the current frame mask, the current frame of the mixed speech, and multiple historical frames of the mixed speech.
[0192] As mentioned above, the current frame mask is a matrix of the same size as the spectrum of the current frame of the mixed speech, with weight values between [0,1]. The current frame mask assigns a weight value to each time-frequency unit of the current frame of the mixed speech. If the weight value is close to 1, the corresponding time-frequency unit is retained; if the weight value is close to 0, the corresponding time-frequency unit is suppressed.
[0193] Understandably, the current frame mask is used to weight the mixed speech in the current frame, preserving the target speech and suppressing noise, thus allowing the target speech to be extracted. The current frame mask is then inverted, i.e., subtracted from a matrix with all values of 1, to obtain the inverse mask. This inverse mask is then used to weight the mixed speech in the current frame, preserving noise and suppressing the target speech, thus allowing the noise to be extracted. In this way, the target speech and noisy speech are separated from the mixed speech in the current frame.
[0194] Multiple historical frames of the mixed speech reflect its overall characteristics. In this embodiment, the target speech for enhancement in the current frame is determined based on the current frame mask, the current frame of the mixed speech, and multiple historical frames of the mixed speech. That is, by combining the current frame mask and the current frame of the mixed speech to distinguish between noise and target speech in the current frame, and referring to multiple historical frames, target speech enhancement is performed frame by frame based on the overall characteristics of the mixed speech. This can effectively improve the accuracy of target speech enhancement in the current frame, thereby improving the accuracy of directional sound pickup.
[0195] In some embodiments, please refer to Figure 11 The aforementioned step S40 specifically includes: S41, Based on the current frame mask and the current frame of the mixed speech, determine the noise speech spectrum corresponding to the current frame of the mixed speech.
[0196] In some embodiments, a current frame mask is used to process the current frame of the mixed speech to obtain the target speech spectrum corresponding to the current frame of the mixed speech. The target speech spectrum of the current frame of the mixed speech can be expressed by the following formula:
[0197] in, It is the current frame of the mixed speech. It is the mask corresponding to the f-th frequency point in the t-th frame of the current frame of the mixed speech. yes The conjugate of complex numbers, It is the target speech spectrum of the current frame of the mixed speech.
[0198] In some embodiments, the current frame mask Perform an inversion operation to obtain the reverse mask. The conjugate complex number of the anti-mask is used to weight the current frame of the mixed speech to obtain the noisy speech spectrum. The noisy speech spectrum of the current frame of the mixed speech can be expressed by the following formula:
[0199] in, It is the current frame of the mixed speech. It is the inverse mask of the current frame mask. It is the conjugate complex number of the anti-mask. It is the noise speech spectrum of the current frame of the mixed speech.
[0200] S42, determine the weight of the current frame based on the gaze direction, the noise speech spectrum corresponding to the current frame of the mixed speech, and the sum of the noise speech spectra corresponding to multiple historical frames of the mixed speech.
[0201] The noise-speech spectrum corresponding to the current frame of the mixed speech reflects the noise information of the current frame, and the noise-speech spectra corresponding to multiple historical frames of the mixed speech reflect the noise information of the multiple historical frames. The weight of the current frame is determined by combining the noise information of the current frame and the noise information of the multiple historical frames. In some embodiments, the noise information of the current frame and the noise information of the multiple historical frames can be statistically analyzed, and the weight of the current frame can be determined based on the statistical results. For example, positions with less noise information in the current frame of the mixed speech are assigned larger weights, while positions with more noise information are assigned smaller weights, in order to reduce noise interference.
[0202] In some embodiments, please refer to Figure 12 The aforementioned step S42 specifically includes: S421, calculate the noise covariance matrix using the noise speech spectrum corresponding to the current frame of the mixed speech and the noise speech spectrum corresponding to multiple historical frames of the mixed speech.
[0203] The noise covariance matrix reflects the statistical results of noise information. Here, the noise covariance matrix includes not only the statistical results of noise information in the current frame, but also the statistical results of noise information in multiple historical frames.
[0204] In some embodiments, the noise covariance matrix is calculated using the following formula:
[0205] Where t is the frame number of the audio frame, t=T is the current frame, and t=1 to t=T-1 are multiple historical frames. It is the noise speech spectrum corresponding to the t-th frame. yes The conjugate transpose of . Let be the noise covariance matrix.
[0206] S422, determine the beam steering vector of each speech channel in the current frame of the mixed speech based on the gaze direction.
[0207] Referring to the description of step S21 above, the beam steering vector describes the relative phase relationship of sound waves arriving at each microphone in the array from a specific direction. The gaze direction is represented by vector g. In some embodiments, the direction of gaze will be... The projection is the target horizontal angle θ∈[0°, 180°]. The beam steering vector is constructed using the following formula. Each voice channel corresponds to a beam steering vector.
[0208]
[0209] Where k is the imaginary unit, and p is the current frequency. This indicates that the first voice channel is for the direction from which the voice comes. The time delay of sound waves, This indicates that the (M-1)th voice channel is related to the direction. The time delay of sound waves.
[0210] S423, determine the weight of the current frame based on the noise covariance matrix and the beam steering vector of each speech channel in the current frame of the mixed speech.
[0211] The noise covariance matrix includes the statistical results of noise information from the current frame and multiple historical frames. The beam steering vectors of each speech channel in the current frame of the mixed speech describe the relative phase relationship of the sound source as it arrives at each microphone in the microphone array from the gaze direction. The weight of the current frame is determined based on both of these factors. In other words, the weight of the current frame is determined after considering the statistical results of noise information from the current frame and multiple historical frames, as well as the relative phase relationship of the sound source as it arrives at each microphone in the microphone array from the gaze direction, thus possessing high accuracy.
[0212] In some embodiments, the aforementioned step S423 specifically includes: normalizing the product of the inverse of the noise covariance matrix and the beam steering vector of each speech channel in the current frame of the mixed speech to obtain the current frame weight.
[0213] Taking M speech channels as an example, the beam steering vector for each of the M speech channels is a vector matrix. This vector matrix is multiplied by the inverse of the noise covariance matrix, and then normalized to obtain the weight of the current frame.
[0214] In some embodiments, the weight of the current frame can be calculated using the following formula: ,in
[0215] in, The inverse of the noise covariance matrix. Let M be the beam steering vectors for the M voice channels. yes The conjugate transpose of . This is the weight of the current frame.
[0216] In this embodiment, the inverse of the noise covariance matrix is multiplied by a vector matrix formed by the beam steering vectors of each speech channel. This allows the noise covariance matrix to act on the beam steering vectors of each speech channel. Normalization is then performed to ensure that the current frame weight is between 0 and 1. Therefore, after the mixed speech is weighted by this current frame weight, the signal amplitude remains unchanged and is neither amplified nor attenuated. This ensures unity gain for the target-enhanced speech in the gaze direction of the current frame. Thus, the target-enhanced speech in the gaze direction of the current frame retains its original amplitude and sound quality after noise suppression, achieving both distortion-free performance and enhanced noise reduction.
[0217] S43, the current frame of the mixed speech and multiple historical frames of the mixed speech are weighted using the current frame weight to obtain the target enhanced speech of the current frame.
[0218] In some embodiments, the current frame of the mixed speech and multiple historical frames are concatenated to obtain joint speech. Then, the joint speech is weighted using the weight of the current frame to obtain the target enhanced speech of the current frame.
[0219] In some embodiments, the target enhanced speech of the current frame can be represented by the following formula:
[0220] in,
[0221] For the current frame of mixed speech, The previous frame for the current frame of the mixed speech. For the current frame containing mixed speech, the preceding T-1 frames. For joint speech of consecutive T frames, Weight of the current frame The conjugate transpose of . Enhance the speech of the target in the current frame.
[0222] In this embodiment, the current frame and multiple historical frames of the mixed speech are weighted using the current frame weight, taking into account prior information, to make the target enhanced speech obtained in the current frame more accurate. Furthermore, the current frame weight... After the aforementioned noise suppression and normalized gain design, the joint speech is weighted so that the target speech of the current frame in the gaze direction is preserved and amplified while the noise is weakened, thus obtaining the target enhanced speech of the current frame.
[0223] In some embodiments, the method S100 further includes: S50, the processing data corresponding to the current frame of the mixed speech is cached so that the processing data can be applied to the speech direction enhancement of the next frame of the mixed speech. The processing data includes the current frame mask, the noise speech spectrum corresponding to the current frame, the noise covariance matrix, or the current frame weights.
[0224] Understandably, near-eye display devices process mixed speech frame by frame, caching the processing data for each frame for inference in subsequent frames. This processing data includes the current frame mask, the noise covariance matrix, and the current frame weights. For example, when processing the next frame, the noise covariance matrix of the current frame is called to calculate the noise covariance matrix for the next frame.
[0225] In this embodiment, the processing data for each frame is cached, thus achieving streaming inference and real-time caching. This effectively avoids recalculating the entire segment and improves computational efficiency. Consequently, it facilitates low-latency continuous audio processing, making it suitable for streaming speech enhancement scenarios.
[0226] In summary, since the directional features are constructed from the current frame of the mixed speech and the gaze direction, they are used to determine the current frame mask. This ensures that the current frame mask is generated under the guidance of the gaze direction, which is beneficial for accurately separating the target speech from the mixed speech. Furthermore, the target enhanced speech for the current frame is determined based on the current frame mask, the current frame of the mixed speech, and multiple historical frames of the mixed speech. That is, by combining the current frame mask and the current frame of the mixed speech to distinguish between noise and target speech in the current frame, and referring to multiple historical frames, target speech enhancement is performed frame by frame based on the overall features of the mixed speech, effectively improving the accuracy of the target enhanced speech in the current frame, thereby improving the accuracy of directional sound pickup. In some embodiments, directional features are constructed based on the actual phase difference between every two speech channels and the theoretical phase difference between every two speech channels under the gaze direction, so that the directional features can reflect the degree to which the target sound source is close to the gaze direction. Using directional features to guide the generation of the current frame mask ensures that the degree to which the target sound source is close to the gaze direction is considered when generating the current frame mask, which is beneficial for improving the accuracy of the current frame mask. In some embodiments, the noise information of the current frame and multiple historical frames are combined to determine the current frame weight, enabling the current frame weight to effectively remove noise and reduce interference. In some embodiments, the current frame weight is used to weight the current frame of the mixed speech and multiple historical frames of the mixed speech, taking into account the preceding information, so that the target enhanced speech of the current frame is more accurate. In some embodiments, the processing data corresponding to each frame is cached after processing, thus realizing streaming inference and real-time caching, which can effectively avoid recalculating the entire segment and improve computational efficiency. Therefore, it is beneficial to achieve low-latency continuous audio processing, which is suitable for streaming speech enhancement scenarios.
[0227] The speech orientation enhancement method provided in this application embodiment can be used by various types of near-eye display devices with computing power.
[0228] See Figure 13 , Figure 13 This is a schematic diagram of a near-eye display device provided in an embodiment of this application. The electronic device 300 includes a processor 301 and a memory 302. The memory 302 is connected to the processor 301, for example, via a bus.
[0229] Processor 301 is configured to support the electronic device 300 in performing the corresponding functions in the methods described in the above method embodiments. Processor 301 may be a central processing unit (CPU), a network processor (NP), a hardware chip, or any combination thereof. The aforementioned hardware chip may be an application-specific integrated circuit (ASIC), a programmable logic device (PLD), or a combination thereof. The aforementioned PLD may be a complex programmable logic device (CPLD), a field-programmable gate array (FPGA), a generic array logic (GAL), or any combination thereof.
[0230] Memory 302 is used to store program code, etc. Memory 302 may include volatile memory (VM), such as random access memory (RAM); memory 302 may also include non-volatile memory (NVM), such as read-only memory (ROM), flash memory, hard disk drive (HDD), or solid-state drive (SSD); memory 302 may also include combinations of the above types of memory.
[0231] The memory 302 is used to store non-volatile software programs, non-volatile computer-executable programs, and modules, such as the program instructions / modules corresponding to the speech orientation enhancement method in the embodiments of this application. The processor 301 executes various functional applications and data processing of the speech orientation enhancement method by running the non-volatile software programs, instructions, and modules stored in the memory 201, thereby realizing the speech orientation enhancement function provided in the above method embodiments.
[0232] Memory 302 may include a program storage area and a data storage area, wherein the program storage area may store the operating system and application programs required for at least one function. In some embodiments, the memory may include memory remotely configured relative to the processor. Examples of the networks described above include, but are not limited to, the Internet, corporate intranets, local area networks, mobile communication networks, and combinations thereof.
[0233] The one or more modules are stored in the memory, and when executed by the one or more processors, they perform the speech orientation enhancement function in any of the above method embodiments, for example, by performing the method steps described in the above method embodiments.
[0234] This application also provides a computer-readable storage medium storing a computer program, the computer program including program instructions, which, when executed by a computer, cause the computer to perform the speech orientation enhancement method as described in the foregoing embodiments.
[0235] Please see Figure 14 , Figure 14 This is a schematic diagram illustrating the directional sound pickup process of a near-eye display device in some embodiments of this application. The near-eye display device includes a bone conduction microphone, a camera, a microphone array, and a controller. The directional sound pickup process performed by the near-eye display device includes the following steps: a) Sound signals were detected by the bone conduction microphone in the near-eye display device.
[0236] (b) In the near-eye display device, the microphone array, camera, and controller are activated. The microphone array captures mixed speech and sends it to the controller, and the camera captures eye images and sends them to the controller.
[0237] c) The controller receives mixed speech and eye images.
[0238] d) The controller performs image reconstruction on the eye image to obtain the reconstructed image.
[0239] The image reconstruction process can refer to the specific implementation process of step A20 in the above-described method embodiment for detecting gaze direction, and will not be repeated here.
[0240] e) The controller calls the stored gaze direction detection model, inputs the reconstructed image into the gaze direction detection model for detection, and obtains the gaze direction.
[0241] The gaze direction detection model is a pre-trained neural network model. For details on using the gaze direction detection model to detect the gaze direction, please refer to the specific implementation process of step A30 in the above-mentioned gaze direction detection method embodiment, which will not be repeated here.
[0242] f) The controller constructs directional features based on the current frame of the mixed speech and the gaze direction, and constructs audio features based on the current frame of the mixed speech.
[0243] The specific methods for constructing directional features and audio features can be found in the specific implementation process of step S20 in the above speech orientation enhancement method embodiment, and will not be repeated here.
[0244] g) The controller determines the current frame mask based on directional and audio features.
[0245] h) Determine the target enhanced speech for the current frame based on the current frame mask, the current frame of the mixed speech, and multiple historical frames of the mixed speech.
[0246] The specific process of step h) can be referred to the specific implementation process of step S40 in the above speech direction enhancement method embodiment, and will not be repeated here.
[0247] In this embodiment, the bone conduction microphone continuously listens for sound signals from the wearer. Only after the bone conduction microphone detects a sound signal are the microphone array, camera, and controller activated to collaboratively perform directional sound pickup. Compared to scenarios where multiple or all of the bone conduction microphone, microphone array, camera, and controller are continuously operating, in this embodiment only the bone conduction microphone is continuously listening, while other components operate intermittently. This results in lower power consumption of the near-eye display device during use, thereby enabling dynamic switching from low-power listening to directional sound pickup. Combined with the gaze direction detection method and speech direction enhancement method in this embodiment, highly robust directional speech pickup is achieved with lower power consumption, which is beneficial for achieving efficient and accurate voice interaction.
[0248] Through the above description of the embodiments, those skilled in the art can clearly understand that each embodiment can be implemented using software and a general-purpose hardware platform, or of course, using hardware. Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.
[0249] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and not to limit them; under the concept of this application, the technical features of the above embodiments or different embodiments can also be combined, the steps can be implemented in any order, and there are many other variations of different aspects of this application as described above, which are not provided in detail for the sake of brevity; although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the foregoing embodiments, or make equivalent substitutions for some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A speech direction enhancement method, characterized in that, include: The current frame and gaze direction of the mixed speech are obtained, which is acquired by a microphone array; A directional feature is constructed based on the current frame of the mixed speech and the gaze direction, and an audio feature is constructed based on the current frame of the mixed speech. The current frame mask is determined based on the directional features and the audio features; The target enhanced speech of the current frame is determined based on the current frame mask, the current frame of the mixed speech, and multiple historical frames of the mixed speech.
2. The method according to claim 1, characterized in that, The construction of directional features based on the current frame of the mixed speech and the gaze direction includes: Obtain the actual phase difference between every two speech channels in the current frame of the mixed speech, and obtain the theoretical phase difference between every two speech channels in the gaze direction; The directional feature is constructed based on the actual phase difference between any two of the said voice channels and the theoretical phase difference between any two of the said voice channels.
3. The method according to claim 2, characterized in that, The step of obtaining the theoretical phase difference between every two speech channels in the gaze direction includes: The beam steering vector for each of the voice channels is determined based on the gaze direction; The theoretical phase difference between each pair of the two voice channels is determined based on the beam steering vector of each pair of the two voice channels in the gaze direction.
4. The method according to claim 2, characterized in that, The construction of the directional feature based on the actual phase difference between every two speech channels and the theoretical phase difference between every two speech channels includes: Calculate the difference between the actual phase difference and the theoretical phase difference between every two of the aforementioned speech channels to obtain multiple phase difference distances; The direction features are obtained by normalizing the multiple phase difference distances.
5. The method according to claim 1, characterized in that, The step of determining the target enhanced speech for the current frame based on the current frame mask, the current frame of the mixed speech, and multiple historical frames of the mixed speech includes: Based on the current frame mask and the current frame of the mixed speech, determine the noise speech spectrum corresponding to the current frame of the mixed speech; The weight of the current frame is determined based on the noise speech spectrum corresponding to the current frame of the mixed speech, the noise speech spectrum corresponding to multiple historical frames of the mixed speech, and the gaze direction. The current frame of the mixed speech and multiple historical frames of the mixed speech are weighted using the current frame weight to obtain the target enhanced speech of the current frame.
6. The method according to claim 5, characterized in that, The step of determining the current frame weight based on the noise speech spectrum corresponding to the current frame of the mixed speech, the noise speech spectrum corresponding to multiple historical frames of the mixed speech, and the gaze direction includes: The noise covariance matrix is calculated using the noise speech spectrum corresponding to the current frame of the mixed speech and the noise speech spectrum corresponding to multiple historical frames of the mixed speech. The beam steering vector of each speech channel in the current frame of the mixed speech is determined based on the gaze direction. The weights of the current frame are determined based on the noise covariance matrix and the beam steering vectors of each speech channel in the current frame of the mixed speech.
7. The method according to claim 6, characterized in that, The step of determining the current frame weight based on the noise covariance matrix and the beam steering vectors of each speech channel in the current frame of the mixed speech includes: The product of the inverse of the noise covariance matrix and the beam steering vector of each speech channel in the current frame of the mixed speech is normalized to obtain the current frame weight.
8. The method according to claim 6, characterized in that, The method further includes: The processing data corresponding to the current frame of the mixed speech is cached so that the processing data can be applied to the speech direction enhancement of the next frame of the mixed speech, wherein the processing data includes the noise covariance matrix.
9. A near-eye display device, characterized in that, include: At least one processor, and The memory communicatively connected to the at least one processor, wherein, The memory stores instructions executable by the at least one processor, which, when executed by the at least one processor, enables the at least one processor to perform the method according to any one of claims 1-8.
10. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions for causing a computer device to perform the method as described in any one of claims 1-8.
Citation Information
Patent Citations
Directional pickup method based on dual microphone array and computing device
CN108269582A
Voice enhancing method, voice enhancing device, voice enhancing equipment and storage medium
CN110085246A
Speech enhancement method and device, equipment, storage medium and program
CN113223552A
VR large-space active noise reduction and directional speech enhancement system
CN120340516A
Beamforming device
US20240365072A1
Cited By
User input audio low-distortion processing method for multi-source noise environment
CN121617411A