Method and device for realizing directional unvarnished transmission according to voice environment and voice interaction equipment
By combining microphone array and gyroscope detection, directional transparent transmission of wireless headphones is achieved, solving the problem of manual mode switching in existing technologies and improving the convenience and clarity of voice communication.
Patent Information
- Application Number
- CN202510657737.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-21
- Publication Date
- 2025-09-05
AI Technical Summary
The transparent transmission function of existing wireless headphones requires manual mode switching and cannot intelligently identify the voice environment, resulting in inconvenience in voice communication and serious interference from environmental noise.
The microphone array is used to directionally identify human voices, and the gyroscope and call microphone are combined to detect the user's status. It automatically determines the voice communication and forms a directional receiving field to directionally amplify the voice of the communicator.
It realizes automatic recognition and directional transparent transmission of voice communication, improves the convenience of use and the clarity of voice communication, and reduces environmental noise interference.
Smart Images

Figure CN120602832A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of audio algorithm technology, and in particular to a method, device and voice interaction equipment for achieving directional transparent transmission according to a voice environment. Background Art
[0002] Wireless headphones such as TWS (True Wireless Stereo) and OWS (Open Wireless Stereo) typically feature both ANC (Active Noise Cancellation) and transparent transmission. In ANC mode, the headphones suppress ambient noise through active noise reduction technology, providing a quiet listening environment. However, when the user needs to communicate with others, the ANC function will also suppress the human voice as noise, resulting in the user being unable to clearly hear the other party's communication content and having to manually switch to transparent transmission mode.
[0003] The current transparent transmission function of TWS / OWS headphones amplifies ambient noise to prevent interference with communication. This traditional transparent transmission method has several shortcomings. First, the transparent transmission function requires the user to manually switch modes, which cannot achieve intelligent operation and reduces convenience. Second, the transparent transmission mode amplifies all ambient noise, including non-human noise such as wind and traffic noise. These noises interfere with voice communication and affect the clarity and comfort of communication. In addition, the traditional transparent transmission mode lacks precision and cannot specifically identify and amplify human voices, resulting in poor voice communication.
[0004] The problems of traditional transparent transmission are particularly prominent in complex acoustic environments, such as noisy streets, restaurants, or conference venues. Ambient noise is equally amplified and mixed with the human voice, making communication difficult. Furthermore, since it cannot distinguish sound sources from different directions, users may be distracted by sounds from multiple directions, making it difficult to focus on the specific person they are communicating with.
[0005] Microphone array beamforming, a technology that uses multiple microphones to create a directional sound field, has been widely used in speech recognition, noise suppression, and sound source localization. However, in voice communication scenarios, existing technology cannot automatically identify and amplify voices from specific directions, failing to meet users' needs for high-quality voice communication without removing their headphones. Summary of the Invention
[0006] The purpose of the present invention is to provide a method, device and voice interaction device for achieving directional transparent transmission according to the voice environment, so as to solve the technical problem that the existing voice interaction devices need to manually switch modes during use and cannot intelligently identify the voice environment.
[0007] To solve the above technical problems, the present invention provides a method for achieving directional transparent transmission according to the voice environment, comprising:
[0008] Directional recognition of human voices within a set distance around the user through a microphone array;
[0009] Get user status information;
[0010] determining whether the user is conducting voice communication based on the human voice and the status information;
[0011] When it is determined that the user is engaged in voice communication, a directional sound receiving field is formed to directionally amplify the voice of the communicator.
[0012] Optionally, obtaining user status information includes:
[0013] Detecting the user's head or limb movements; and / or
[0014] Detect user's voice feedback.
[0015] Optionally, detecting the movement of the user's head or limbs includes detecting the movement direction and amplitude of the user's head or limbs.
[0016] Optionally, the directionally identifying human voices within a set distance around the user by using a microphone array includes:
[0017] Beamforming technology is used to capture human voices in a directionally controlled manner while excluding ambient noise and interference from other directions.
[0018] Optionally, determining whether the user is conducting voice communication includes:
[0019] Comprehensively analyze the user's head movements and voice signals to determine whether the user is engaged in voice communication.
[0020] Optionally, the comprehensive analysis of the user's head movements and voice signals includes:
[0021] Based on whether the user is facing the sound source, whether he or she nods, and whether there is voice feedback information, it is comprehensively judged whether the user is engaged in voice communication.
[0022] The present invention also provides a device for implementing directional transparent transmission according to a voice environment, comprising:
[0023] Microphone array, used for directionally identifying human voices within a set distance around the user;
[0024] A status detection unit, used to obtain user status information;
[0025] The processor is used to determine whether the user is conducting voice communication based on the human voice and the status information; and when it is determined that the user is conducting voice communication, form a directional sound receiving field and directionally amplify the human voice of the communicator.
[0026] Optionally, the state detection unit includes:
[0027] A gyroscope to detect movements of the user's head or limbs; and / or
[0028] The call microphone is used to detect the user's voice feedback.
[0029] Optionally, the microphone array captures human voice in a directionally manner while excluding ambient noise and interference from other directions.
[0030] The present invention also provides a voice interaction device, comprising the apparatus for achieving directional transparent transmission according to the voice environment as described above.
[0031] Compared with the prior art, the present invention has at least the following technical effects:
[0032] The technical solution provided by this invention can automatically identify the voice communication scenario. Based on the user's status information and the surrounding human voice information, it intelligently determines and automatically switches to directional transparent transmission mode, eliminating the need for manual operation by the user and improving user convenience. At the same time, by directional sound reception and amplifying the communicator's voice, it effectively filters out ambient noise, improving the clarity and comfort of voice communication.
[0033] In addition, the present invention achieves precise directional sound reception through the beamforming technology of the microphone array, and combines the multi-dimensional information fusion of the gyroscope detecting the user's head movement and the call microphone detecting voice feedback, which greatly improves the accuracy of voice scene recognition and reduces the misjudgment rate, enabling the system to more intelligently adapt to the voice interaction needs in various complex environments. It is particularly suitable for voice interaction devices such as wireless headphones and smart speakers. BRIEF DESCRIPTION OF THE DRAWINGS
[0034] Figure 1 This is a block diagram of a method for implementing directional transparent transmission according to a voice environment in an embodiment of the present invention;
[0035] Figure 2 The figure is a schematic diagram of the structure of a device for achieving directional transparent transmission according to the voice environment in an embodiment of the present invention.
[0036] In the figure, 10, microphone array; 20, state detection unit; 30, processor; 21, gyroscope; 22, call microphone. DETAILED DESCRIPTION
[0037] The following describes a method, apparatus, and voice interaction device for implementing directional transparent transmission based on a voice environment, in conjunction with schematic diagrams. These diagrams illustrate preferred embodiments of the present invention. It should be understood that those skilled in the art may modify the invention described herein while still achieving the beneficial effects of the invention. Therefore, the following description should be understood as generally known to those skilled in the art and not as a limitation of the present invention.
[0038] The following paragraphs describe the present invention in more detail by way of example with reference to the accompanying drawings. The advantages and features of the present invention will become more apparent from the following description. It should be noted that the drawings are greatly simplified and not to exact scale, and are provided solely for the purpose of assisting in the description of the embodiments of the present invention.
[0039] Example 1
[0040] This embodiment provides a method for achieving directional transparent transmission according to the voice environment. Figure 1 ,include:
[0041] S1: Directedly identifying human voices within a set distance around the user through the microphone array 10.
[0042] In this step, the microphone array 10 uses beamforming technology to capture the human voice within a set distance around the user.
[0043] In a specific example, the set distance can be within 2-3 meters around the user, and the microphone array 10 can include at least 4 microphone units to form a circular or linear array. The microphone array 10 can form one or more beams through beamforming technology, each beam pointing in a different direction, thereby capturing sound signals from different directions. Through beamforming technology, the microphone array 10 can effectively capture human voices in a directionally controlled manner while eliminating ambient noise and interference from other directions.
[0044] S2: Get user status information.
[0045] In this step, the status information includes the user's head or body movement information and / or voice feedback information.
[0046] Specifically, the movement of the user's head or limbs can be detected by the gyroscope 21 , and / or the user's voice feedback can be detected by the call microphone 22 .
[0047] The gyroscope 21 detects the movement of the user's head or limbs, including detecting the movement direction and amplitude of the user's head or limbs.
[0048] In a specific example, the gyroscope 21 can detect the user's head rotation angle, speed, acceleration and other data to determine whether the user is facing the sound source. The communication microphone 22 can detect whether the user has issued a voice signal to further confirm whether the user has the intention to communicate.
[0049] S3: Determine whether the user is performing voice communication based on the human voice and the status information.
[0050] In this step, the user's head movements and voice signals are comprehensively analyzed to determine whether the user is conducting voice communication.
[0051] Specifically, it is possible to comprehensively judge whether the user is engaging in voice communication based on information such as whether the user is facing the direction of the sound source, whether the user has body movements such as nodding, and whether there is voice feedback.
[0052] In a specific example, a judgment threshold can be set. For example, when the user's head turning angle exceeds a set value (such as 30 degrees) and lasts for more than a set time (such as 2 seconds), it is determined that the user is paying attention to the sound source in that direction; when the user detects voice feedback through the call microphone 22, it is further confirmed that the user is conducting voice communication.
[0053] S4: When it is determined that the user is in the process of voice communication, a directional sound receiving field is formed to directionally amplify the voice of the communicator.
[0054] In this step, when it is determined that the user is conducting voice communication, the system automatically adjusts the beam direction and gain of the microphone array 10 to form a directional sound receiving field and directionally amplify the voice of the communicator.
[0055] Specifically, the beam direction of the microphone array 10 can be adjusted according to the sound source direction information so that it accurately points to the position of the communicator; at the same time, the signal gain in this direction is increased, the clarity of the human voice is improved, and the interference of environmental noise is reduced.
[0056] In a specific example, different gain levels can be set, such as low, medium, and high, and the gain level can be automatically adjusted according to the needs of the communication scenario to obtain optimal speech clarity and comfort.
[0057] It should be noted that in steps S1-S3, if the human voice within the set distance around the user is not recognized, or the user status information is not obtained, or it is determined that the user is not communicating by voice, the process returns to step S1 to directionally identify the human voice within the set distance around the user through the microphone array 10.
[0058] The beneficial effects of this embodiment are: directional sound reception is achieved through the beamforming technology of the microphone array 10, and automatic recognition of voice communication scenes is achieved in combination with the gyroscope 21 and the call microphone 22, without the need for the user to manually switch modes, thereby improving ease of use; at the same time, the human voice is directionally amplified to avoid interference from environmental noise, thereby improving the clarity and comfort of voice communication.
[0059] Example 2
[0060] This embodiment provides a device for achieving directional transparent transmission according to the voice environment. Figure 2 ,include:
[0061] The microphone array 10 is used for directionally identifying human voices within a set distance around the user;
[0062] A status detection unit 20 is used to obtain user status information;
[0063] The processor 30 is configured to determine whether the user is engaging in voice communication based on the human voice and the status information; and when it is determined that the user is engaging in voice communication, form a directional sound receiving field and directionally amplify the human voice of the communicator.
[0064] The microphone array 10 captures the human voice in a directionally manner through beamforming technology while eliminating ambient noise and interference from other directions.
[0065] In a specific example, the microphone array 10 may include at least 4 microphone units to form a ring or linear array.
[0066] The microphone unit can be an omnidirectional microphone or a directional microphone, and a beamforming function is achieved through a specific arrangement and signal processing algorithm.
[0067] The state detection unit 20 includes a gyroscope 21 for detecting the movement of the user's head or limbs; and / or a call microphone 22 for detecting the user's voice feedback.
[0068] The gyroscope 21 can detect data such as the user's head rotation angle, speed, and acceleration, and then determine whether the user is facing the direction of the sound source.
[0069] The communication microphone 22 can detect whether the user sends a voice signal, and further confirm whether the user has the intention to communicate.
[0070] The processor 30 is electrically connected to the microphone array 10 and the state detection unit 20 , receives the human voice signal collected by the microphone array 10 and the state information collected by the state detection unit 20 , and performs processing and analysis.
[0071] In a specific example, the processor 30 may be a digital signal processor (DSP), a microcontroller unit (MCU), an application processor (AP), or the like.
[0072] In another specific example, the processor 30 can run a specific algorithm to comprehensively analyze the user's head movements and voice signals to determine whether the user is conducting voice communication; when it is determined that the user is conducting voice communication, the microphone array 10 is controlled to adjust the beam direction and gain to form a directional receiving field and directionally amplify the communicator's voice.
[0073] In addition, in a specific example, the device may also include a memory (not shown in the figure) for storing preset algorithm parameters and program codes; and an audio processing unit (not shown in the figure) for further processing the amplified human voice signal, such as noise reduction, equalization, etc., to improve the sound quality.
[0074] The device of this embodiment provides a structural basis for achieving directional transparent transmission according to the voice environment. It realizes directional sound reception through microphone array beamforming technology, and combines the gyroscope and call microphone to realize automatic recognition of voice communication scenes. There is no need for users to manually switch modes, which improves the convenience of use. At the same time, it amplifies the human voice in a direction to avoid interference from environmental noise, thereby improving the clarity and comfort of voice communication.
[0075] Example 3
[0076] This embodiment provides a voice interaction device, which may include the apparatus described in the second embodiment.
[0077] In this embodiment, the voice interaction device may be a wireless headset or a smart speaker.
[0078] When the voice interaction device is a wireless headset, the microphone array 10 can be set on the housing of the headset, the gyroscope 21 can be built into the headset, and the call microphone 22 can be set below the headset to better capture the user's voice.
[0079] When the voice interaction device is a smart speaker, the microphone array 10 can be set on the top or side of the speaker to form a 360-degree circular array to capture sounds from all directions.
[0080] In addition to the device described in Example 2, the voice interaction device may also include conventional components such as a speaker, a battery, and a communication module. The speaker is used to play the processed sound signal; the battery provides power to the device; and the communication module is used to achieve wireless connection with other devices, such as Bluetooth connection or Wi-Fi connection.
[0081] In a specific example, when the voice interaction device is a TWS (True Wireless Stereo) headset, the left and right earbuds can exchange data and work together through wireless communication. For example, the microphone arrays 10 of the left and right earbuds can form a larger virtual array to improve beamforming accuracy; and the data from the gyroscopes 21 of the left and right earbuds can be combined and analyzed to more accurately determine the user's head movements.
[0082] The voice interaction device of this embodiment applies directional transparent transmission technology to actual products. Through the coordinated operation of microphone array 10, gyroscope 21, and call microphone 22, intelligent directional transparent transmission is achieved, eliminating the need for users to manually switch modes, improving the product's ease of use and user experience. At the same time, the voice is amplified in a targeted manner, avoiding interference from ambient noise, and improving the clarity and comfort of voice communication. Especially in noisy environments, the voice interaction device of the present invention can effectively extract and amplify human voices from a specific direction, allowing users to clearly hear the voice of the communicator without being disturbed by ambient noise.
[0083] In summary, the method, device, and voice interaction device for achieving directional transparent transmission based on the voice environment provided by the present invention solve the technical problem in the prior art that the device needs to manually switch modes in voice communication scenarios and cannot directionally amplify human voices through the organic combination of microphone array beamforming technology, state detection, and intelligent judgment algorithms. The solution of the present invention does not require the user to manually switch modes, intelligently identifies the voice communication scenario, automatically forms a directional receiving field, and amplifies the voice of the communicator, while effectively suppressing environmental noise interference, significantly improving the convenience and clarity of voice communication, and providing users with a more comfortable voice communication experience.
[0084] Obviously, those skilled in the art may make various changes and modifications to the present invention without departing from the spirit and scope of the present invention. Thus, if such changes and modifications fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. A method for achieving directional transparent transmission according to a voice environment, characterized in that: include: Directional recognition of human voices within a set distance around the user through a microphone array; Get user status information; determining whether the user is conducting voice communication based on the human voice and the status information; When it is determined that the user is engaged in voice communication, a directional sound receiving field is formed to directionally amplify the voice of the communicator.
2. The method for achieving directional transparent transmission according to the voice environment according to claim 1, characterized in that: The obtaining of user status information includes: Detecting the user's head or limb movements; and / or Detect user's voice feedback.
3. The method for achieving directional transparent transmission according to the voice environment according to claim 2, characterized in that: The detecting of the movement of the user's head or limbs includes detecting the movement direction and amplitude of the user's head or limbs.
4. The method for achieving directional transparent transmission according to a voice environment according to claim 1, wherein: The method of directionally identifying human voices within a set distance around the user by using a microphone array includes: Beamforming technology is used to capture human voices in a directionally controlled manner while excluding ambient noise and interference from other directions.
5. The method for achieving directional transparent transmission according to a voice environment according to claim 1, wherein: Determining whether the user is conducting voice communication includes: Comprehensively analyze the user's head movements and voice signals to determine whether the user is engaged in voice communication.
6. The method for achieving directional transparent transmission according to the voice environment according to claim 5, characterized in that: The comprehensive analysis of the user's head movements and voice signals includes: Based on whether the user is facing the sound source, whether he or she nods, and whether there is voice feedback information, it is comprehensively judged whether the user is engaged in voice communication.
7. A device for achieving directional transparent transmission according to a voice environment, characterized in that: include: Microphone array, used for directionally identifying human voices within a set distance around the user; A status detection unit, used to obtain user status information; a processor, configured to determine whether the user is performing voice communication based on the human voice and the status information; And when it is determined that the user is conducting voice communication, a directional sound receiving field is formed to directionally amplify the voice of the communicator.
8. The device for achieving directional transparent transmission according to the voice environment according to claim 7, characterized in that: The state detection unit includes: A gyroscope to detect movements of the user's head or limbs; and / or The call microphone is used to detect the user's voice feedback.
9. The device for achieving directional transparent transmission according to a voice environment according to claim 7, characterized in that: The microphone array captures the human voice in a directionally oriented manner while excluding ambient noise and interference from other directions.
10. A voice interaction device, comprising the apparatus for achieving directional transparent transmission according to a voice environment as claimed in any one of claims 7 to 9.