Head-mounted device and audio data processing method
By combining the sound pickup component and eye-tracking component of the head-mounted device, the device locks onto the wearer's gaze target and uses voiceprint and shape features to separate the target audio data from mixed audio data, solving the problem of poor audio quality in noisy environments and improving the user experience.
Patent Information
- Application Number
- CN202610518275.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-04-17
- Publication Date
- 2026-08-25
AI Technical Summary
In noisy environments, head-mounted devices struggle to accurately lock onto the wearer's gaze, resulting in poor audio data quality and impacting the user experience.
It employs a combination of a microphone pickup component and an eye-tracking component. The microphone array collects mixed audio data, the eye-tracking component determines the wearer's gaze direction, the control unit determines the target object based on the gaze direction and environmental data, and extracts the target audio data from the mixed audio data through voiceprint and shape features.
It effectively increases the difficulty for wearers to distinguish the speech of specific individuals, and improves the quality of audio data and user experience.
Smart Images

Figure CN122633018A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of smart wearable device technology, and more specifically, to a head-mounted device and an audio data processing method. Background Technology
[0002] With the development of technology, more and more wearable smart devices are entering people's daily lives. Head-mounted devices are a common type of wearable smart device, which can include, for example, virtual reality (VR) glasses, augmented reality (AR) glasses, mixed reality (MR) glasses, or displayless AI glasses, etc.
[0003] Some head-mounted devices have audio capture capabilities, allowing them to collect audio data from the wearer's environment. However, if the environment is noisy and multiple people are speaking simultaneously, the audio data quality is usually poor, making it difficult for the wearer to distinguish specific speakers and thus reducing the user experience. Summary of the Invention
[0004] In view of this, embodiments of the present invention provide a head-mounted device and an audio data processing method, which, after locking onto the wearer's gaze target, separates the audio data of a specific object from the mixed audio data based on the characteristic data of the gaze target, reducing the difficulty for the wearer to distinguish the speech of a specific object, thereby effectively improving the wearer's user experience.
[0005] In a first aspect, embodiments of the present invention provide a head-mounted device, the device comprising: The pickup component, including a microphone array, is configured to acquire mixed audio data; The eye-tracking component is configured to determine the wearer's gaze direction; The control unit is configured to determine a target object based on the line-of-sight and environmental data, extract feature data of the target object, and extract target audio data of the target object from the mixed audio data based on the feature data, so as to perform directional tracking of the target object, wherein the feature data includes at least one of the voiceprint features and shape features of the target object.
[0006] Optionally, the control unit is further configured to: The object of gaze is determined based on the direction of the gaze and the environmental data; In response to the duration of the gaze direction pointing to the same gaze object satisfying a first duration condition, or upon receiving an object confirmation instruction, the corresponding gaze object is identified as the target object.
[0007] Optionally, the control unit is further configured to: In response to receiving an object change instruction, or if the duration for which the voiceprint feature is not extracted meets the second duration condition, the target object is re-determined.
[0008] Optionally, the control unit is further configured to: The real-time orientation of the target object is determined based on the feature data; Beamforming is performed on the mixed audio data according to the real-time direction to obtain the target audio data.
[0009] Optionally, the control unit is further configured to: The directions of arrival of multiple sound sources, including the target object, are determined based on the mixed audio data. The real-time direction is determined from each of the directions of arrival based on the characteristic data.
[0010] Optionally, the control unit is further configured to: The real-time distance between the wearer and the target object is determined, and the mixed audio data is beamformed based on the real-time direction and the real-time distance to obtain the target audio data.
[0011] Optionally, the control unit is further configured to: The mixed audio data is subjected to signal separation processing to obtain multiple audio data. The target audio data is determined from each of the audio data based on the feature data.
[0012] Optionally, the control unit is further configured to: The target audio data is translated to obtain the corresponding translation result; The device also includes: The output component is configured to output the translation result.
[0013] Secondly, embodiments of the present invention provide an audio data processing method, the method comprising: To determine the wearer's line of sight; The target object is determined based on the line of sight and environmental data; Acquire feature data of the target object, wherein the feature data includes at least one of the voiceprint features and shape features of the target object; Based on the feature data, target audio data of the target object is extracted from the collected mixed audio data in order to perform targeted tracking of the target object.
[0014] Thirdly, embodiments of the present invention provide a computer-readable storage medium storing computer program instructions thereon, wherein the computer program, when executed by a processor, implements the method described in the second aspect.
[0015] Fourthly, embodiments of the present invention provide a computer program product including instructions that, when executed on a head-mounted device, cause the head-mounted device to perform the method described in the second aspect.
[0016] The head-mounted device of this invention includes a sound pickup component, an eye-tracking component, and a control unit. The sound pickup component is used to collect mixed audio data, the eye-tracking component is used to determine the wearer's gaze direction, and the control unit is used to determine a target object based on the gaze direction and environmental data, extract feature data of the target object, and extract the audio data of the target object from the mixed audio data based on the feature data, thereby achieving directional tracking of the target object. After locking onto the wearer's gaze target, this invention separates the audio data of the specific object from the mixed audio data based on the feature data of the gaze target, reducing the difficulty for the wearer to distinguish the speech of a specific object, thus effectively improving the wearer's user experience. Attached Figure Description
[0017] The above and other objects, features and advantages of the present invention will become clearer from the following description of embodiments of the invention with reference to the accompanying drawings, in which: Figure 1 This is a structural block diagram of a head-mounted device according to an embodiment of the present invention; Figure 2 This is a flowchart of an audio data processing method according to an embodiment of the present invention; Figure 3 This is a flowchart of an audio data processing method according to an embodiment of the present invention; Figure 4 This is a flowchart of an audio data processing method according to an embodiment of the present invention; Figure 5 This is a flowchart of an audio data processing method according to an embodiment of the present invention; Figure 6 This is a schematic diagram illustrating an application scenario of the audio data processing method according to an embodiment of the present invention; Figure 7 This is a flowchart of an audio data processing method according to an embodiment of the present invention; Figure 8 This is a schematic diagram of an audio data processing device according to an embodiment of the present invention. Detailed Implementation
[0018] The present application is described below based on embodiments, but it is not limited to these embodiments. In the detailed description of the present application below, certain specific details are described in detail. Those skilled in the art can fully understand the present application without these details. To avoid obscuring the substance of the present application, well-known methods, processes, flows, elements, and circuits are not described in detail.
[0019] Furthermore, those skilled in the art should understand that the accompanying drawings provided herein are for illustrative purposes only and are not necessarily drawn to scale.
[0020] Unless the context explicitly requires it, words such as "including" or "contains" throughout the application should be interpreted as including rather than exclusive or exhaustive; that is, meaning "including but not limited to".
[0021] In the description of this application, it should be understood that the terms "first," "second," etc., are used for descriptive purposes only and should not be construed as indicating or implying relative importance. Furthermore, in the description of this application, unless otherwise stated, "a plurality of" means two or more.
[0022] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties. Furthermore, the collection, use and processing of the relevant data must comply with the relevant laws, regulations and standards of the relevant countries and regions, and corresponding operation entry points are provided for users to choose to authorize or refuse.
[0023] In noisy environments with multiple people speaking simultaneously, the quality of audio data captured by head-mounted devices is typically poor. To improve audio quality, existing technologies usually rely on beamforming algorithms to extract audio data of specific objects from the raw audio data. However, the directional enhancement effect of beamforming algorithms depends on the accuracy of the pickup direction, and existing technologies struggle to accurately determine the wearer's gaze direction. This makes it difficult to accurately determine the directional pickup direction of the head-mounted device, and consequently, to accurately extract audio data of specific objects. Therefore, the improvement in audio quality is limited.
[0024] To address the aforementioned problems, embodiments of the present invention propose a head-mounted device and an audio data processing method. Figure 1 This is a structural block diagram of a head-mounted device according to an embodiment of the present invention. Figure 1 As shown, the head-mounted device of this embodiment includes a sound pickup component 11, an eye-tracking component 12, and a control unit 13.
[0025] The sound pickup component 11 mainly includes multiple microphones, which can be arranged in an array to collect mixed audio data. In some embodiments, the microphones can be analog microphones, such as dynamic microphones, electret microphones, micro-electro-mechanical system (MEMS) analog microphones, etc.; in some embodiments, the microphones can be digital microphones, such as MEMS digital microphones, etc.
[0026] The eye-tracking component 12 is used to determine the wearer's gaze direction and may include optical components, a visual sensor system, and a processing unit. The optical components include an infrared light source system, a light transmission system, and an infrared bandpass filter. The infrared light source system can be at least one of an infrared light-emitting diode (LED) array, an infrared laser diode (LD) array, or a vertical-cavity surface-emitting laser (VCSEL) array. The infrared light-emitting devices can be arranged in a circumferential ring, multi-quadrant symmetrical, or other manner. The light transmission system may include a lens array and a waveguide component. The lens array can be at least one of a planar lens array, a microlens array, or a planar convex lens array. The waveguide component can be a waveguide sheet made with a transparent substrate such as glass or polymer, with diffraction gratings or geometric reflection structures integrated internally or on its surface. The infrared bandpass filter is used to filter out visible light interference. The visual sensor system can be an infrared camera array. The infrared camera can be a monocular (multi-view) camera component composed of a global shutter infrared complementary metal-oxide-semiconductor (CMOS) image sensor, a rolling shutter infrared CMOS sensor, or a high frame rate infrared array image sensor. The processing unit can control the LED array to illuminate the wearer's eyeball from different angles to form corneal reflection points, and control the light to be transmitted to the system to perform multi-beam parallel decomposition or light field homogenization processing on the incident light. Then, it controls the vision sensor system to acquire eye images, and then calculates the vector relationship between the pupil center and the reflection point according to the pupil center-corneal reflection method. After image processing algorithms, such as ORB (Oriented Fast and Rotated BRIEF) feature point detection, corneal reflection interference removal, and lens distortion correction, the eye optical features are accurately located. Finally, the eye vector is mapped to the actual gaze direction through the eye calibration model.
[0027] The control unit 13 is electrically connected to the microphone assembly 11 and the eye-tracking assembly 12, enabling control of these components. The control unit 13 may include a processor for controlling the overall operation of the head-mounted device and may include one or more processing units. For example, the processor may include at least one of the following: a central processing unit (CPU), an application processor (AP), a modem processor, a graphics processing unit (GPU), an image signal processor (ISP), a video processing unit (VPU), a video codec, a digital signal processor (DSP), a baseband processor, and a neural network processing unit (NPU). Different processing units may be independent devices or integrated into one or more processors. The controller can generate operation control signals based on instruction opcodes and timing signals to control the acquisition and execution of instructions. The processor may also include a memory for storing instructions and data. The processor executes various functional applications and data processing of the head-mounted device by running the instructions stored in the memory. In some embodiments, the memory is a cache memory.
[0028] In this embodiment of the invention, the control unit 13 is configured to determine a target object based on the wearer's gaze direction and environmental data, extract feature data of the target object, and extract target audio data of the target object from mixed audio data based on the feature data, so as to perform directional tracking of the target object. The feature data includes at least one of the target object's voiceprint features and shape features.
[0029] In an optional implementation of this embodiment, the control unit 13 is further configured to determine the object of gaze based on the wearer's gaze direction and environmental data, and to determine the corresponding object of gaze as the target object in response to the wearer's gaze direction pointing to the same object for a duration that meets a first duration condition, or to receive an object confirmation instruction.
[0030] In an optional implementation of this embodiment, the control unit 13 is further configured to re-determine the target object in response to receiving an object change instruction or when the duration for which the voiceprint features of the target object are not extracted meets a second duration condition.
[0031] In an optional implementation of this embodiment, the control unit 13 is further configured to determine the real-time direction of the target object based on the feature data of the target object, and to perform beamforming processing on the mixed audio data based on the real-time direction of the target object to obtain the target audio data.
[0032] In an optional implementation of this embodiment, the control unit 13 is further configured to determine the directions of arrival (DOA) of multiple sound sources based on the mixed audio data, and to determine the real-time direction from each DOA based on the characteristic data of the target object. The multiple sound sources include the target object.
[0033] In one optional implementation of this embodiment, the control unit 13 is further configured to determine the real-time distance between the wearer and the target object, and to perform beamforming processing on the mixed audio data based on the real-time direction and real-time distance of the target object to obtain the target audio data.
[0034] In an optional implementation of this embodiment, the control unit 13 is further configured to perform signal separation processing on the mixed audio data, acquire multiple audio data, and determine target audio data from each audio data based on the feature data of the target object.
[0035] In one optional implementation of this embodiment, the head-mounted device further includes an output component. The control unit 13 is further configured to perform translation processing on the target audio data to obtain the corresponding translation result; the output component is configured to output the translation result.
[0036] In this embodiment of the invention, after locking onto the wearer's gaze target, the audio data of the specific object is separated from the mixed audio data based on the characteristic data of the gaze target, reducing the difficulty for the wearer to distinguish the speech of the specific object, and thus effectively improving the wearer's user experience.
[0037] The following describes the method through examples. Figure 2 This is a flowchart of an audio data processing method according to an embodiment of the present invention. Figure 2 As shown, the method in this embodiment includes the following steps: Step S100: Obtain the wearer's gaze direction.
[0038] In real-world situations, if the wearer is in a noisy environment, they will typically focus their attention on the target in order to hear their words clearly, for example, by looking directly at the target. Therefore, in this step, the head-mounted device can obtain the wearer's gaze direction through eye-tracking components.
[0039] Optionally, to reduce the power consumption and improve the battery life of the head-mounted device, the device can obtain the wearer's gaze direction through an eye-tracking component after receiving an eye-tracking command. In one optional implementation, if the head-mounted device includes a voice interaction component, the wearer can issue eye-tracking commands through the voice interaction component. The head-mounted device may include at least one of a voice interaction component and a command trigger button. The voice interaction component may include a microphone pickup unit, a voice recognition unit, and an intent recognition unit. The microphone pickup unit can be at least one microphone in a microphone array for collecting voice information emitted by the wearer; the voice recognition unit performs voice recognition on the voice information collected by the microphone pickup unit to determine the text content of the voice information; the intent recognition unit performs intent recognition on the text content of the voice information to determine the intent information of the voice information. Then, the voice interaction component can send the intent information to the head-mounted device. Head-mounted devices can pre-store the mapping relationship between various intent information and commands. Therefore, in this case, after receiving the intent information of the voice information, the head-mounted device can determine whether the intent information is used to represent eye tracking. If so, the head-mounted device can determine that the wearer has issued an eye tracking command.
[0040] In one alternative implementation, if the head-mounted device includes a command trigger button, the wearer can issue eye-tracking commands via the command trigger button. The command trigger button can be a mechanical switch or a pressure sensing unit; this embodiment is not limited in this regard. The command trigger button can convert the wearer's trigger operation (such as a pressing operation) into a trigger signal and send the trigger signal to the head-mounted device. The head-mounted device can pre-store the mapping relationship between each trigger signal and command. Therefore, in this case, after receiving a trigger signal, the head-mounted device can determine whether the trigger signal corresponds to an eye-tracking command. If so, the head-mounted device can determine that the wearer has issued an eye-tracking command.
[0041] Depending on the hardware structure of the head-mounted device, it can also determine whether it has received an eye-tracking command in other ways. For example, if the head-mounted device has a touchpad, such as using the temples as a touchpad, the wearer can perform specific gestures using the touchpad. If the head-mounted device recognizes the specific gesture used to trigger magnetic calibration, it can determine that the wearer has issued an eye-tracking command.
[0042] Step S200: Determine the target object based on the line of sight and environmental data.
[0043] While acquiring the wearer's gaze direction, the head-mounted device can also collect environmental data of the wearer's surroundings through visual sensors. Based on the wearer's gaze direction and the environmental data, the device can then identify the object the wearer is looking at as the target object. Target object identification can be achieved in various ways, such as the method described in "Using point cloud data to improve three-dimensional gaze estimation, Haofei Wang, Marco Antonelli, Bertram E. Shi, 39th Annual International Conference of the IEEE Engineering in Medicine and Biology Society, EMBC'17, 2017, DOI:10.1111 / CGF.70144," which will not be elaborated upon in this embodiment.
[0044] In practical applications, the object of a wearer's listening is usually not frequently changed. In noisy environments, to hear the target's speech clearly, the wearer typically maintains focus on the target for a period of time. Therefore, in some embodiments, the head-mounted device can determine whether the duration of the wearer's gaze on the target meets a first duration condition after determining the wearer's gaze direction and environmental data of the wearer's environment. If the duration of the wearer's gaze on the target meets the first duration condition, the head-mounted device can identify the target as the target. The first duration condition can be that the duration of the wearer's gaze on the same target is higher than (or not lower than) a certain duration. The first duration can be set according to actual needs, for example, it can be determined based on the continuous gaze duration of each tester in a noisy environment without changing the target, or it can be determined based on the wearer's historical continuous gaze duration in a noisy environment without changing the target.
[0045] In some embodiments, the wearer can also trigger the object confirmation command in various ways. After determining the object the wearer is gazing at, if the head-mounted device receives an object confirmation command, it can designate that object as the target object. The way the wearer triggers the head-mounted device to receive the object confirmation command can be referred to the way the head-mounted device receives the eye-tracking command, and will not be repeated here.
[0046] In some embodiments, if the head-mounted device receives an object change instruction, or if the duration for which the voiceprint features of the current target object are not extracted meets the second duration condition, it indicates that the wearer is no longer paying attention to the current target object. Therefore, the head-mounted device can re-determine the target object using the above-described method. The second duration condition can be that the duration for which the voiceprint features of the current target object are not extracted is higher than (or not lower than) the second duration, and the second duration can also be set according to actual needs. The way the wearer triggers the head-mounted device to receive an object change instruction can refer to the way the head-mounted device receives an eye-tracking instruction.
[0047] Step S300: Obtain the feature data of the target object.
[0048] In this embodiment, the feature data of the target object can be either the voiceprint feature or the shape feature of the target object. Therefore, in this step, the head-mounted device can acquire feature data in different ways depending on the type of feature data.
[0049] When the feature data is voiceprint features, head-mounted devices can extract the voiceprint features of the target object from the acquired mixed audio data in various ways. In some optional implementations, the head-mounted device can separate the target audio data of the target object from the mixed audio data and extract the voiceprint features of the target object from the target audio data of the target object according to the voiceprint feature extraction algorithm.
[0050] Furthermore, head-mounted devices can separate target audio data for a specific object in various ways. For example, a head-mounted device can use the wearer's gaze direction as the main lobe direction of the beam and perform spatial filtering on the mixed audio data based on beamforming algorithms to obtain the target audio data. As another example, a head-mounted device can use blind source signal separation algorithms, such as Independent Component Analysis (ICA) and Independent Vector Analysis (IVA), or deep learning algorithms used for blind source signal separation, such as Multi-Conv-TasNet and Multi-DPRNN, to process the mixed audio data into independent audio data for each sound source. It can then calculate the direction of arrival (DOA) of each audio data based on the time difference of arrival (TDOA), and identify the audio data whose DOA matches the wearer's gaze direction as the target audio data.
[0051] When the feature data is shape features, the head-mounted device can extract the shape features of the target object from the acquired environmental images in various ways. In some optional implementations, the head-mounted device can determine the shape region of the target object from the environmental image and extract the target object's shape features from the shape region using shape feature extraction algorithms, such as Histogram of Oriented Gradient (HOG) and Fourier descriptors. Optionally, to further improve the accuracy of directional tracking, when the shape feature is facial features, the head-mounted device can determine the facial region of the target object from the environmental image and extract the target object's facial features from the facial region using facial feature extraction algorithms.
[0052] Furthermore, head-mounted devices can determine the shape region of a target object in various ways. For example, a head-mounted device can identify the target environment image as an environmental image that matches the wearer's gaze direction, and then perform image segmentation processing on the target environment image based on image segmentation algorithms, such as threshold segmentation algorithms, edge detection segmentation algorithms, region segmentation algorithms, etc., or deep learning models used to achieve image segmentation, such as semantic segmentation models (including fully convolutional networks, Convolutional Networks for Biomedical Image Segmentation (U-Net), A Deep Convolutional Encoder-Decoder Architecture for Image Segmentation (SegNet), etc.), instance segmentation models (including Mask R-CNN, Segmenting Objects by Locations (SOLO) series models, etc.), and panoramic segmentation models (including Panoptic FPN, Unified Panoptic Segmentation Network (UPSNet), etc.) to obtain the shape region of the target object. For example, a head-mounted device can identify the target environment image as an environmental image that matches the wearer's gaze direction, and determine the three-dimensional spatial bounding box of the target object based on an object detection algorithm or a deep learning model used to achieve object detection, and then determine the shape region of the target object based on the three-dimensional spatial bounding box of the target object.
[0053] In some alternative implementations, the head-mounted device can also extract the shape features of each object from the environmental image based on the shape feature extraction algorithm, and determine the shape features of the object that matches the wearer's gaze direction as the shape features of the target object.
[0054] Alternatively, other methods may be used to determine the feature data of the target object in this embodiment, and this embodiment does not limit this.
[0055] Step S400: Extract the target audio data of the target object from the acquired mixed audio data based on the feature data.
[0056] In this embodiment, after acquiring the feature data of the target object, the head-mounted device can perform directional tracking based on the feature data of the target object, so as to reduce the possibility of audio quality degradation caused by the wearer's line of sight deviating from the target object for a short time.
[0057] In this step, the head-mounted device can extract the target audio data of the target object in various ways based on the target object's feature data. Figure 3 This is a flowchart of an audio data processing method according to an embodiment of the present invention. Figure 3 As shown, in an optional implementation of this embodiment, step S400 may include the following steps: Step S410: Perform signal separation processing on the mixed audio data to obtain multiple audio data.
[0058] In this embodiment, the head-mounted device can continuously acquire mixed audio data through a sound pickup component. After acquiring the audio data, the head-mounted device can perform signal separation processing on the acquired mixed audio data to obtain independent audio data for each sound source.
[0059] Taking the signal separation processing of mixed audio data using Multi-Conv-TasNet as an example, Multi-Conv-TasNet is a fully convolutional temporal audio separation network, mainly consisting of an encoder, a separation module, and a decoder. The encoder transforms the temporal waveform of the mixed audio data acquired by the microphone array into a high-dimensional feature representation and fuses multi-channel spatial features. The separation module, composed of a Temporal Convolutional Network (TCN) and a mask generator, transforms the high-dimensional feature representation into mask features corresponding to each sound source to achieve blind source signal separation. The decoder maps the mask features corresponding to each sound source to a temporal signal through deconvolution to reconstruct the temporal waveform corresponding to each sound source. Therefore, when performing signal separation processing of mixed audio data using Multi-Conv-TasNet, head-mounted devices can use the mixed audio data as input to Multi-Conv-TasNet to obtain the audio data corresponding to each sound source.
[0060] Step S420: Determine the target audio data from each audio data based on the feature data.
[0061] In this step, when the feature data of the head-mounted device is voiceprint features, the head-mounted device can extract the corresponding voiceprint features from the audio data corresponding to each sound source according to the voiceprint feature extraction algorithm, and determine the similarity between the voiceprint features corresponding to each sound source and the voiceprint features of the target object. Then, the audio data whose similarity meets the preset similarity conditions is determined as the target audio data.
[0062] The preset similarity conditions can be determined based on the similarity calculation method used. For example, if the similarity calculation method is cosine similarity, then for any object's voiceprint feature, the closer the cosine similarity corresponding to the voiceprint feature is to 1, the higher the probability that the object is the target object. The closer the cosine similarity corresponding to the voiceprint feature is to -1, the lower the probability that the object is the target object. Therefore, the preset similarity conditions can be determined as the similarity being higher than (or not lower than) a preset threshold, and the cosine similarity ranking being the first highest.
[0063] In this way, head-mounted devices can improve audio quality simply by separating the target audio data of the object being gazed at from the mixed audio data. This effectively reduces the complexity of audio data processing and improves audio data processing efficiency.
[0064] Figure 4 This is a flowchart of an audio data processing method according to an embodiment of the present invention. Figure 4 As shown, in another optional implementation of this embodiment, step S400 may include the following steps: Step S410': Determine the real-time orientation of the target object based on the feature data.
[0065] In this step, the head-mounted device can determine the real-time orientation of the target object in various ways. In some embodiments, the head-mounted device can determine the real-time orientation of the target object based on its voiceprint characteristics. Figure 5 This is a flowchart of an audio data processing method according to an embodiment of the present invention. Figure 5 As shown, in some embodiments, step S410' may include the following steps: Step S411': Determine the directions of arrival of multiple sound sources based on the mixed audio data.
[0066] In this step, the head-mounted device can also perform signal separation processing on the mixed audio data, determine the audio data corresponding to each sound source, and calculate the direction of arrival (DOA) of each sound source based on the arrival time difference of the audio data corresponding to each sound source. Optionally, the DOA of the i-th sound source... This can be expressed by the following formula:
[0067] Where c represents the speed of sound in air. This represents the time difference between the arrival of the sound from the i-th sound source at the j-th and k-th microphones in the microphone array. This represents the distance between the j-th microphone and the k-th microphone.
[0068] Step S412': Determine the real-time direction from each direction of arrival based on the feature data.
[0069] In this step, the head-mounted device can determine the target audio data of the target object based on the voiceprint characteristics, and determine the direction of arrival corresponding to the target audio data as the real-time direction of the target object. The target audio data of the target object can be determined in the manner described in step S420, and will not be repeated here.
[0070] In some embodiments, the head-mounted device can determine the real-time orientation of the target object based on its shape features. Optionally, the head-mounted device can perform image recognition on the environmental image based on the shape features of the target object to determine the position of the target object in the environmental image, and then determine the real-time orientation of the target object based on the intrinsic parameters of the image sensor, the mounting angle, and the position of the target object in the environmental image. The intrinsic parameters of the image sensor can be pre-calibrated or determined using methods such as OpenCV; the image sensor is typically fixedly mounted on the head-mounted device, so the mounting angle of the image sensor can be determined by an attitude sensing component or also using methods such as OpenCV.
[0071] Step S420': Beamforming processing is performed on the mixed audio data according to the real-time direction to obtain the target audio data.
[0072] After determining the real-time orientation of the target object, in this step, the head-mounted device can use the real-time orientation of the target object as the main lobe orientation of the beam, and perform spatial filtering on the mixed audio data based on the beamforming algorithm to obtain the target audio data.
[0073] In some embodiments, the head-mounted device can calculate the sound propagation delay compensation amount of the target object's sound to each microphone in the microphone array relative to the reference microphone (i.e., the predefined reference microphone) based on the real-time direction of the target object, and determine the total delay of each microphone based on the sound propagation delay compensation amount corresponding to each microphone and the pre-determined delay calibration amount. The total delay is converted into a phase factor, and then the frequency domain fixed weight of each audio channel in the mixed audio data is determined based on the pre-determined amplitude weighting coefficient. Thus, the audio signals of each audio channel are transformed in the frequency domain and then weighted and summed. Finally, the weighted sum is transformed in the time domain to realize the spatial filtering of the mixed audio data.
[0074] In some optional implementations, to further improve audio quality, the head-mounted device can also determine the real-time distance between the wearer and the target object. The real-time distance between the wearer and the target object can be determined in several ways. For example, the head-mounted device can use eye-tracking components to determine the wearer's convergence angle and interpupillary distance, and then calculate the distance between the wearer and the target object based on these parameters; alternatively, the distance can be determined using images of the target environment and parameters from image sensors. Taking the calculation of the distance between the wearer and the target object based on convergence as an example, the real-time distance between the wearer and the target object... This can be expressed by the following formula:
[0075] in, Indicates the wearer's interpupillary distance. It indicates the radius of the wearer.
[0076] After determining the real-time distance between the wearer and the target object, the head-mounted device can determine the beamforming focusing distance, pointing angle, and gain parameters based on the real-time distance and the real-time direction of the target object. Based on the focusing distance, pointing angle, and gain parameters, the near-field spatial range centered on the target object is determined as the effective sound pickup range. Then, the target sound component generated in the mixed audio data within the effective sound pickup range is enhanced, while the interfering sound component generated in other ranges is suppressed and attenuated, thereby obtaining the target audio data.
[0077] In this way, head-mounted devices can separate the target audio data of the object being gazed at by the wearer from the mixed audio data through spatial filtering, allowing the wearer to receive a purer human voice, thus further improving audio quality and enhancing the user experience.
[0078] Figure 6 This is a schematic diagram illustrating an application scenario of the audio data processing method according to an embodiment of the present invention. For example... Figure 6 As shown, after the head-mounted device determines that the wearer's gaze is object B based on the device's gaze direction and environmental data, it can designate object B as the target object when the duration of the wearer's gaze on object B meets a first duration condition or when an object confirmation command is received. The device then extracts the target audio data of object B from the collected mixed audio data based on object B's characteristic data for directional tracking. If the wearer's gaze temporarily changes from object B to object A, the duration of the wearer's gaze on object A does not meet the first duration condition, and no object confirmation command is received. Therefore, the head-mounted device will still perform directional tracking of object B.
[0079] Figure 7 This is a flowchart of an audio data processing method according to an embodiment of the present invention. Figure 7As shown, in some embodiments, this embodiment may further include the following steps: Step S500: Translate the target audio data to obtain the corresponding translation result.
[0080] In some embodiments, the head-mounted device can also translate the target audio data to obtain the corresponding translation result.
[0081] Optionally, the head-mounted device can perform Automatic Speech Recognition (ASR) processing on the target audio data to obtain the corresponding speech-recognized text, and translate the speech-recognized text into target text written in the target language. The target text can be specified by the wearer in various ways. Optionally, after obtaining the target text corresponding to the target audio data, the head-mounted device can also perform speech synthesis processing on the target text to obtain the corresponding target speech.
[0082] Alternatively, the head-mounted device may also use other methods to acquire the target text or target speech corresponding to the target audio data.
[0083] Step S600: Output the translation result.
[0084] In this step, the head-mounted device can output the translation results in various ways to synchronize the translation results with the wearer. For example, if the output component is a display component, the head-mounted device can display the translation results through the display component; if the output component is an audio playback component, the head-mounted device can broadcast the translation results through the audio playback component.
[0085] The head-mounted device of this invention includes a sound pickup component, an eye-tracking component, and a control unit. The sound pickup component is used to collect mixed audio data, the eye-tracking component is used to determine the wearer's gaze direction, and the control unit is used to determine a target object based on the gaze direction and environmental data, extract feature data of the target object, and extract the audio data of the target object from the mixed audio data based on the feature data, thereby achieving directional tracking of the target object. After locking onto the wearer's gaze target, this invention separates the audio data of the specific object from the mixed audio data based on the feature data of the gaze target, reducing the difficulty for the wearer to distinguish the speech of a specific object, thus effectively improving the wearer's user experience.
[0086] Figure 8 This is a schematic diagram of an audio data processing apparatus according to an embodiment of the present invention. Figure 8 As shown, the audio data processing device in this embodiment includes a direction determination unit 801, an object determination unit 802, a feature acquisition unit 803, and an audio extraction unit 804.
[0087] The direction determination unit 801 is used to obtain the wearer's gaze direction; the object determination unit 802 is used to determine the target object based on the gaze direction and environmental data; the feature acquisition unit 803 is used to acquire the feature data of the target object, the feature data including at least one of the target object's voiceprint features and shape features; and the audio extraction unit 804 is used to extract the target audio data of the target object from the collected mixed audio data based on the feature data, so as to perform directional tracking of the target object.
[0088] The head-mounted device of this invention includes a sound pickup component, an eye-tracking component, and a control unit. The sound pickup component is used to collect mixed audio data, the eye-tracking component is used to determine the wearer's gaze direction, and the control unit is used to determine a target object based on the gaze direction and environmental data, extract feature data of the target object, and extract the audio data of the target object from the mixed audio data based on the feature data, thereby achieving directional tracking of the target object. After locking onto the wearer's gaze target, this invention separates the audio data of the specific object from the mixed audio data based on the feature data of the gaze target, reducing the difficulty for the wearer to distinguish the speech of a specific object, thus effectively improving the wearer's user experience.
[0089] Those skilled in the art will understand that embodiments of this application can be provided as methods, apparatus (devices), or computer program products. Therefore, this application can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, this application can take the form of a computer program product embodied on one or more computer-readable storage media (including but not limited to disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0090] These computer program instructions may be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including an instruction means, the implementation process of which is described in the instruction means. Figure 1 The function specified in one or more processes.
[0091] These computer program instructions may also be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing device to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing device, produce instructions for implementing processes. Figure 1 A device for a function specified in one or more processes.
[0092] Another embodiment of the present invention relates to a non-volatile storage medium for storing a computer-readable program, which is used by a computer to execute some or all of the above-described method embodiments.
[0093] That is, those skilled in the art will understand that all or part of the steps in the methods of the above embodiments can be implemented by specifying relevant hardware through a program. This program is stored in a storage medium and includes several instructions to cause a device (which may be a microcontroller, chip, etc.) or processor to execute all or part of the steps of the methods described in the embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as a USB flash drive, a portable hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk. In one application scenario, the non-volatile storage medium storing the above-mentioned computer program product can be part of the control unit of a head-mounted device.
[0094] The above description is merely a preferred embodiment of this application and is not intended to limit this application. Various modifications and variations can be made to this application by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this application should be included within the protection scope of this application.
Claims
1. A head-mounted device, characterized in that, The device includes: The pickup component, including a microphone array, is configured to acquire mixed audio data; The eye-tracking component is configured to determine the wearer's gaze direction; The control unit is configured to determine a target object based on the line-of-sight and environmental data, extract feature data of the target object, and extract target audio data of the target object from the mixed audio data based on the feature data, so as to perform directional tracking of the target object, wherein the feature data includes at least one of the voiceprint features and shape features of the target object.
2. The device according to claim 1, characterized in that, The control unit is further configured to: The object of gaze is determined based on the direction of the gaze and the environmental data; In response to the duration of the gaze direction pointing to the same gaze object satisfying a first duration condition, or upon receiving an object confirmation instruction, the corresponding gaze object is identified as the target object.
3. The device according to claim 2, characterized in that, The control unit is further configured to: In response to receiving an object change instruction, or if the duration for which the voiceprint feature is not extracted meets the second duration condition, the target object is re-determined.
4. The device according to claim 1, characterized in that, The control unit is further configured to: The real-time orientation of the target object is determined based on the feature data; Beamforming is performed on the mixed audio data according to the real-time direction to obtain the target audio data.
5. The device according to claim 4, characterized in that, The control unit is further configured to: The directions of arrival of multiple sound sources, including the target object, are determined based on the mixed audio data. The real-time direction is determined from each of the directions of arrival based on the characteristic data.
6. The device according to claim 4, characterized in that, The control unit is further configured to: The real-time distance between the wearer and the target object is determined, and the mixed audio data is beamformed based on the real-time direction and the real-time distance to obtain the target audio data.
7. The device according to claim 1, characterized in that, The control unit is further configured to: The mixed audio data is subjected to signal separation processing to obtain multiple audio data. The target audio data is determined from each of the audio data based on the feature data.
8. The device according to claim 1, characterized in that, The control unit is also configured to: The target audio data is translated to obtain the corresponding translation result; The device also includes: The output component is configured to output the translation result.
9. An audio data processing method, characterized in that, The method includes: To determine the wearer's line of sight; The target object is determined based on the line of sight and environmental data; Acquire feature data of the target object, wherein the feature data includes at least one of the voiceprint features and shape features of the target object; Based on the feature data, target audio data of the target object is extracted from the collected mixed audio data in order to perform targeted tracking of the target object.
10. A computer-readable storage medium storing computer program instructions thereon, characterized in that, When the computer program is executed by a processor, it implements the method as described in claim 9.
11. A computer program product, comprising instructions, characterized in that, When the instructions are executed on the head-mounted device, the head-mounted device performs the method as described in claim 9.