AR glasses that can distinguish sound sources

By using a dual-microphone assembly and an audio processor in AR glasses to calculate the angle of the sound source, the voices of the wearer and the person being communicated with are distinguished, solving the problem of inaccurate recognition caused by environmental interference in existing technologies. This achieves efficient and flexible voice interaction, suitable for a variety of application scenarios.

CN115220230BActive Publication Date: 2026-04-14CETHIK GRP
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-06-20
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing AR glasses are easily affected by external environmental interference during voice recognition, making it difficult to accurately locate the wearer's voice information. They also consume a lot of computing power and memory, limiting their application scenarios.

Method used

It employs a dual-microphone assembly and an audio processor to calculate the sound source angle, distinguishing the voice of the wearer from that of the person being communicated with. It also uses a wireless communication module to analyze voice information with external devices, reducing environmental noise interference and optimizing microphone positions to improve recognition accuracy.

Benefits of technology

It achieves accurate voice recognition of the wearer in complex environments, reduces recognition errors, improves user experience, expands application scenarios, avoids computational and memory consumption, and is suitable for various environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115220230B_ABST
    Figure CN115220230B_ABST
Patent Text Reader

Abstract

The application discloses AR glasses capable of distinguishing sound sources, which comprise a frame, a leg, a glasses display screen, a mainboard, a first microphone assembly, a second microphone assembly, a wireless communication module, a first audio processor and a second audio processor, the first microphone assembly comprises a first microphone and a second microphone which are arranged side by side in the same leg, the second microphone assembly comprises a third microphone and a fourth microphone which are symmetrically arranged at the bottom of the frame, the mainboard, the audio processors and the wireless communication module are built in the frame, the first microphone assembly is electrically connected with the first audio processor, the second microphone assembly is electrically connected with the second audio processor, the glasses display screen, the audio processors and the wireless communication module are all electrically connected with the mainboard, and the sound source body can be distinguished and responded by calculating the angle of the sound source during work. The device can distinguish the sound source body, reduce environmental sound interference, facilitate voice recognition interaction, avoid occupying a large amount of computing power and memory, is less affected by the environment and is flexible to use.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention belongs to the field of AR glasses technology, specifically relating to an AR glasses capable of distinguishing sound sources. Background Technology

[0002] AR glasses are a new product of modern technology. Existing voice recognition technology simply uses a single microphone or dual microphones to uniformly pick up sound signals for recognition, or uniformly picks up sound signals and uses camera image acquisition for directional recognition.

[0003] One approach involves using a single or dual microphone to uniformly pick up sound signals for recognition. This involves converting the obtained speech into text using TTS (Text-to-Speech) technology, then processing it with algorithms to perform subsequent operations, such as displaying the converted text on the AR glasses screen or extracting keywords from the text to determine the next step, like playing a video or face-to-face translation. However, this approach has drawbacks: the picked-up sound has significant noise, is easily affected by external environmental interference, requires software algorithm correction, consumes computing power and memory, and also contains a lot of useless speech information, making it difficult to accurately locate the wearer's output speech and hindering flexible and rapid application. Another approach combines uniform sound pickup with camera image acquisition for directional recognition. For example, Chinese patent CN110188179 B discloses a voice directional recognition interaction method, including the following steps: picking up and recognizing sound signals directly in front to obtain speech-text content; acquiring the speech-text content; acquiring a face image that simultaneously meets the image acquisition angle and acquisition distance; and determining whether to respond based on the speech-text content and the face image; wherein the image acquisition angle is 60-70 degrees and the acquisition distance is less than or equal to 1 meter. The disadvantages are that the camera used has a small image acquisition angle, generally within 70 degrees, and the acquisition distance is often within 1 meter. These two conditions are too strict on the target of directional voice recognition, which greatly limits the actual application scenarios. It also cannot accurately locate and capture the wearer's voice information as subsequent voice commands. Summary of the Invention

[0004] The purpose of this invention is to address the above-mentioned problems by proposing an AR glasses that can distinguish sound sources. This AR glasses can distinguish the main sound source, reduce interference from ambient sounds, facilitate voice recognition interaction, avoid consuming a lot of computing power and memory, are less affected by the environment, and are flexible in application.

[0005] To achieve the above objectives, the technical solution adopted by the present invention is as follows:

[0006] This invention proposes an AR glasses system capable of distinguishing sound sources, comprising a frame, temples, and a display screen, and further comprising a motherboard, a first microphone assembly, a second microphone assembly, a wireless communication module, a first audio processor, and a second audio processor, wherein:

[0007] The first microphone assembly includes a first microphone and a second microphone, which are arranged side by side in a horizontal direction on the outside of the same temple, with the second microphone positioned close to the glasses display screen.

[0008] The second microphone assembly includes a third microphone and a fourth microphone, which are symmetrically positioned at the bottom of the frame.

[0009] The motherboard, the first audio processor, the second audio processor, and the wireless communication module are built into the frame. The first microphone and the second microphone are both electrically connected to the first audio processor, and the third microphone and the fourth microphone are both electrically connected to the second audio processor. The glasses display, the first audio processor, the second audio processor, and the wireless communication module are all electrically connected to the motherboard.

[0010] In operation, the first and second microphones collect sound source information and transmit it to the first audio processor. The first audio processor calculates the first sound source angle based on the received sound source information. The first sound source angle is the angle between the sound source emission point and the lines connecting the first and second microphones, respectively. The third and fourth microphones collect sound source information and transmit it to the second audio processor. The second audio processor calculates the second sound source angle based on the received sound source information. The second sound source angle is the angle between the line connecting the sound source emission point and the third microphone and the plane of symmetry of the frame. The motherboard performs the following operations:

[0011] When the first sound source angle is greater than or equal to the first preset angle, the sound source information is considered to come from the wearer of the glasses and is responded to as a command word or wake-up word. The system then determines whether the second sound source angle is greater than the second preset angle. If so, the sound source information is considered to come from the communication object. The main board receives the corresponding sound source information and sends it to an external device for analysis via the wireless communication module. The analyzed sound source information is then sent back to the main board and projected onto the glasses display screen. Otherwise, the sound source information is considered to come from the wearer of the glasses and no response is made.

[0012] Preferably, both the first microphone assembly and the second microphone assembly are MIC linear array digital silicon microphones.

[0013] Preferably, the distance between the first microphone and the second microphone is 80mm to 100mm, and the distance between the second microphone and the front wall of the glasses display screen is 20mm to 24mm.

[0014] Preferably, the distance between the third microphone and the fourth microphone is less than or equal to 10mm.

[0015] Preferably, the first preset angle is 30° to 90°, and the second preset angle is 0° to 20°.

[0016] Preferably, the first preset angle is 54° to 58°, and the second preset angle is 2° to 8°.

[0017] Preferably, the audio processor is model ZL38063.

[0018] Preferably, the parsed sound source information is text information or graphic information.

[0019] Compared with the prior art, the beneficial effects of the present invention are as follows:

[0020] 1) By optimizing the position of the microphone and working with the audio processor to calculate the angle of the sound source, the AR glasses can distinguish the main body of the sound source, greatly reducing the interference of ambient sound. They can accurately identify the voice information of the wearer for voice recognition interaction, reduce errors, improve the user experience, and avoid consuming a lot of computing power and memory. Compared with existing technologies, they are less affected by the environment (such as the collection angle, collection distance, light, etc.) and are more flexible in application.

[0021] 2) By distinguishing the voices of the wearer and the person being communicated with, the system enables targeted voice wake-up of the wearer and voice-text interaction with the person being communicated with. It can selectively receive the voice of the person being communicated with and combine it with external devices (such as servers, mobile phones, etc.) to achieve cloud recognition and analysis, converting it into text or graphics for display on the glasses' screen. It has a wide range of applications. Attached Figure Description

[0022] Figure 1 This is a schematic diagram of the structure of the AR glasses of the present invention that can distinguish sound sources;

[0023] Figure 2 This is a side view of a human wearing AR glasses according to the present invention;

[0024] Figure 3 This is a front view of a human wearing the AR glasses of the present invention.

[0025] Figure 4 This is a circuit diagram of the AR glasses of the present invention that can distinguish sound sources.

[0026] Explanation of reference numerals in the attached diagram: 1. Frame; 2. Temple; 3. Display screen; 11. Third microphone; 12. Fourth microphone; 21. First microphone; 22. Second microphone. Detailed Implementation

[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of this application.

[0028] It should be noted that when a component is referred to as being "connected" to another component, it can be directly connected to the other component or there may be an intervening component. Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein in the specification of this application is for the purpose of describing particular embodiments only and is not intended to limit the application.

[0029] like Figure 1-4 As shown, an AR glasses capable of distinguishing sound sources includes a frame 1, temples 2, and a display screen 3. It also includes a motherboard, a first microphone assembly, a second microphone assembly, a wireless communication module, a first audio processor, and a second audio processor, wherein:

[0030] The first microphone assembly includes a first microphone 21 and a second microphone 22. The first microphone 21 and the second microphone 22 are arranged side by side in the horizontal direction on the outside of the same temple 2, and the second microphone 22 is positioned close to the glasses display screen 3.

[0031] The second microphone assembly includes a third microphone 11 and a fourth microphone 12, which are symmetrically arranged at the bottom of the frame 1.

[0032] The motherboard, the first audio processor, the second audio processor, and the wireless communication module are built into the frame 1. The first microphone 21 and the second microphone 22 are both electrically connected to the first audio processor. The third microphone 11 and the fourth microphone 12 are both electrically connected to the second audio processor. The glasses display screen 3, the first audio processor, the second audio processor, and the wireless communication module are all electrically connected to the motherboard.

[0033] In operation, the first microphone 21 and the second microphone 22 collect sound source information and transmit it to the first audio processor. The first audio processor calculates the first sound source angle based on the received sound source information. The first sound source angle is the angle between the sound source emission point and the lines connecting the first microphone 21 and the second microphone 22. The third microphone 11 and the fourth microphone 12 collect sound source information and transmit it to the second audio processor. The second audio processor calculates the second sound source angle based on the received sound source information. The second sound source angle is the angle between the line connecting the sound source emission point and the third microphone 11 and the plane of symmetry of the frame 1. The motherboard performs the following operations:

[0034] When the first sound source angle is greater than or equal to the first preset angle, the sound source information is considered to come from the glasses wearer and is responded to as a command word or wake-up word. The system then determines whether the second sound source angle is greater than the second preset angle. If so, the sound source information is considered to come from the communication object. The main board receives the corresponding sound source information and sends it to an external device for analysis via the wireless communication module. The analyzed sound source information is then sent back to the main board and projected onto the glasses display screen 3. Otherwise, the sound source information is considered to come from the glasses wearer and no response is made.

[0035] In this application, the orientation of the AR glasses is defined based on the wearing state of the human body; that is, when the AR glasses are in a horizontal position, the line of sight is considered "front," and the back is considered "back." For example... Figure 1 As shown, the first microphone 21 and the second microphone 22 are horizontally arranged side by side on the outside of the right temple 2, and the third microphone 11 and the fourth microphone 12 are symmetrically arranged at the bottom of the frame 1, preferably directly above the bridge of the nose. Figure 2 As shown, the sound source is point A.

[0036] like Figure 4 As shown, during operation, the system receives sound source information from all directions through the first and second microphone components. The picked-up sound signals are then transmitted to the corresponding audio processors (such as the ASR audio processor) to calculate the sound source angle. Selective sound source collection is performed; for example, the motherboard only begins processing the audio information of the wearer and the person being communicated with when the wearer's voice signal is identified as a command word or wake-up word. This processing may include translating the audio information of the person being communicated with, or directly responding to commands such as answering a phone call, playing a video, or turning on the camera. Otherwise, no audio information processing is performed. The first microphone component is primarily responsible for calculating the sound source angle to determine whether the sound source is the wearer of the glasses. The second microphone component further calculates the angle to determine whether to respond to the wearer's voice commands and process the audio information. Each audio processor can also improve the accuracy of sound source localization by taking the average of multiple samples of the sound source angle.

[0037] Each audio processor acquires sound source information via the SPI serial peripheral interface, processes it on the motherboard (e.g., performing calibration, packaging, etc.), and then wirelessly transmits it to external devices via wireless communication modules (e.g., 5G / WIFI modules). External devices can then identify and analyze the information via cloud-based solutions such as Alibaba Cloud, iFlytek, and Speechocean. The resulting text and image information is then transmitted to the MIPI interface glasses display screen 3 for display. This wireless data transmission (including text, video, and images) effectively solves the cumbersome problem of wired data transmission in AR glasses, making operation convenient.

[0038] When a person wears AR glasses, the system can accurately locate and capture the wearer's voice information as subsequent voice commands. By distinguishing between the wearer's and the person being communicated with, it can directionally enable voice wake-up for the wearer and voice-text interaction for the person being communicated with. It can selectively receive the voice of the person being communicated with and combine it with external devices (such as servers, mobile phones, etc.) for cloud-based recognition and analysis, converting it into text or graphics for display on the glasses' display screen 3. Through selective voice discrimination combined with cloud processing, it can be expanded to many practical scenarios. It can effectively recognize the voices of outsiders and the wearer, as well as environmental sounds, enabling effective communication with users interacting in front of them and accurately recognizing the wearer's voice commands to control the AR glasses.

[0039] This AR glasses, by optimizing microphone placement and using an audio processor to calculate the sound source angle, can distinguish the main sound source, greatly reducing interference from ambient sounds. It can accurately identify the wearer's voice, facilitating voice recognition interaction, reducing voice command errors, improving the user experience, and avoiding excessive computing power and memory consumption. Compared to existing technologies, this AR glasses is not affected by factors such as acquisition angle, acquisition distance, or ambient environment, ensuring accurate identification of the wearer's voice commands. Unlike traditional single-microphone or dual-microphone systems that rely on a single microphone for sound signal pickup, this AR glasses avoids situations where a person with a cold might be incorrectly rejected. Furthermore, it avoids the influence of sound sample quality, emotion, background noise, and changes in sound over time, and is not interfered with by other surrounding voices. Its flexible applications include cultural tourism, industry, gaming, education, healthcare, aviation, and urban security, and it can be used in various scenarios, such as AR and VR.

[0040] In one embodiment, both the first microphone assembly and the second microphone assembly are MIC linear array digital silicon microphones.

[0041] In one embodiment, the distance between the first microphone 21 and the second microphone 22 is 80mm to 100mm, and the distance between the second microphone 22 and the front wall of the glasses display screen 3 is 20mm to 24mm. The first microphone 21 and the second microphone 22 are at the same horizontal position, and the center distance is preferably 100mm, and the distance between the second microphone 22 and the front wall of the glasses display screen 3 is preferably 20mm.

[0042] In one embodiment, the distance between the third microphone 11 and the fourth microphone 12 is less than or equal to 10 mm.

[0043] In one embodiment, the first preset angle is 30° to 90°, and the second preset angle is 0° to 20°.

[0044] In one embodiment, the first preset angle is 54° to 58°, and the second preset angle is 2° to 8°.

[0045] like Figure 2 As shown, this is a side view of the device worn by the user. Point A is the sound source. Preferably, the angle between point A and the lines connecting the first microphone 21 and the second microphone 22 is 54° to 58°. Figure 3 As shown, this is a front view of the human body wearing the glasses. The third microphone 11 and the fourth microphone 12 are located at the bottom of the frame 1, and the angle between the line connecting point A and the third microphone 11 or the fourth microphone 12 and the plane of symmetry of the frame 1 is 2° to 8°.

[0046] In one embodiment, the audio processor is a ZL38063. Using the ZL38063 audio processor as the voice recognition chip, it not only calculates the sound source angle but also provides echo cancellation (AEC) technology, improving the user experience. It should be noted that the audio processor can be selected according to actual needs.

[0047] In one embodiment, the parsed sound source information is text information or graphic information.

[0048] How do these AR glasses work?

[0049] After the device is activated, the first microphone 21 and the second microphone 22 collect sound source information and transmit it to the first audio processor. The first audio processor calculates the first sound source angle based on the received sound source information. The third microphone 11 and the fourth microphone 12 collect sound source information and transmit it to the second audio processor. The second audio processor calculates the second sound source angle based on the received sound source information. The mainboard performs the following operations: when the first sound source angle is greater than or equal to 30°, generally between 54° and 58°, the sound source is considered to be from the wearer of the glasses. When the first sound source angle is less than 30°, it further determines whether the second sound source angle is greater than 20°, generally between 2° and 8°. If so, the sound source is considered to be from the person being communicated with. The mainboard receives the corresponding sound source information and sends it to an external device for parsing via the wireless communication module. The parsed text or graphic information is returned to the mainboard and projected onto the glasses display screen 3. Otherwise, the sound source is considered to be from the wearer of the glasses, and no response is made.

[0050] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0051] The embodiments described above are merely specific and detailed examples of the embodiments described in this application, and should not be construed as limiting the scope of the patent application. It should be noted that those skilled in the art can make various modifications and improvements without departing from the concept of this application, and these all fall within the scope of protection of this application. Therefore, the scope of protection of this patent application should be determined by the appended claims.

Claims

1. An AR glasses capable of distinguishing sound sources, comprising a frame, temples, and a display screen, characterized in that: The AR glasses capable of distinguishing sound sources further include a motherboard, a first microphone assembly, a second microphone assembly, a wireless communication module, a first audio processor, and a second audio processor, wherein: The first microphone assembly includes a first microphone and a second microphone, which are arranged side by side in a horizontal direction on the outside of the same temple, with the second microphone positioned close to the display screen of the glasses. The second microphone assembly includes a third microphone and a fourth microphone, which are symmetrically disposed at the bottom of the frame; The motherboard, the first audio processor, the second audio processor, and the wireless communication module are built into the frame. The first microphone and the second microphone are both electrically connected to the first audio processor, and the third microphone and the fourth microphone are both electrically connected to the second audio processor. The glasses display screen, the first audio processor, the second audio processor, and the wireless communication module are all electrically connected to the motherboard. In operation, the first and second microphones collect sound source information and transmit it to the first audio processor. The first audio processor calculates a first sound source angle based on the received sound source information. The first sound source angle is the angle between the sound source emission point and the lines connecting the first and second microphones. The third and fourth microphones collect sound source information and transmit it to the second audio processor. The second audio processor calculates a second sound source angle based on the received sound source information. The second sound source angle is the angle between the line connecting the sound source emission point and the third microphone and the plane of symmetry of the frame. The motherboard performs the following operations: When the first sound source angle is greater than or equal to the first preset angle, the sound source information is considered to come from the wearer of the glasses and is responded to as a command word or wake-up word. It is then determined whether the second sound source angle is greater than the second preset angle. If so, the sound source information is considered to come from the communication object. The motherboard receives the corresponding sound source information and sends it to an external device for parsing through the wireless communication module. The parsed sound source information is then transmitted back to the motherboard and projected onto the glasses display screen. Otherwise, the sound source information is considered to come from the wearer of the glasses and no response is made.

2. The AR glasses capable of distinguishing sound sources as described in claim 1, characterized in that: Both the first microphone assembly and the second microphone assembly are MIC linear array digital silicon microphones.

3. The AR glasses capable of distinguishing sound sources as described in claim 1, characterized in that: The distance between the first microphone and the second microphone is 80mm to 100mm, and the distance between the second microphone and the front wall of the glasses display screen is 20mm to 24mm.

4. The AR glasses capable of distinguishing sound sources as described in claim 1, characterized in that: The distance between the third and fourth microphones is less than or equal to 10mm.

5. The AR glasses capable of distinguishing sound sources as described in claim 1, characterized in that: The first preset angle is 30° to 90°, and the second preset angle is 0° to 20°.

6. The AR glasses capable of distinguishing sound sources as described in claim 5, characterized in that: The first preset angle is 54° to 58°, and the second preset angle is 2° to 8°.

7. The AR glasses capable of distinguishing sound sources as described in claim 1, characterized in that: The audio processor is model ZL38063.

8. The AR glasses capable of distinguishing sound sources as described in claim 1, characterized in that: The parsed sound source information is either text information or graphic information.

Citation Information

Patent Citations

  • Voice orientation recognition interaction methods, devices, equipment and media

    CN110188179B

  • Sound source positioning method and device and computer storage medium

    CN112466325A

  • A electronic system for sound localization

    CN208334627U