AR (Augmented Reality) glasses capable of identifying environment sound and carrying out multi-mode prompt
The AR eyeglass system addresses noisy output by using location-based noise filtering and multi-modal prompts to enhance clarity and accuracy of audio and visual information for the hearing impaired.
Patent Information
- Application Number
- CN202510612321.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-13
- Publication Date
- 2025-07-15
AI Technical Summary
The microphones of existing AR glasses extract too much sound information, resulting in messy output information, making it difficult for people with hearing impairments to obtain information accurately.
By setting the positioning component to determine the frame position, the information extraction component determines the noise frequency based on the position and filters out the noise, and uses the multimodal prompt component to output accurate information, including a text display and a sound outputter.
It improves the accuracy of information extraction and the accuracy of noise removal, making prompt information easier to distinguish, and optimizes the performance of AR glasses.
Smart Images

Figure CN120318998A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of intelligent devices, and particularly to an AR glasses that can identify ambient sounds and give multimodal prompts. Background Art
[0002] A hearing aid is a device used to help people with hearing impairments hear sounds clearly. Hearing aids can be customized according to individual usage needs. Common hearing aids generally include in-ear hearing aids, behind-the-ear hearing aids, and glasses-type hearing aids. Among them, glasses-type hearing aids generally consist of a frame and a hearing aid, and the hearing aid is usually installed on the temple of the frame. The glasses of glasses-type hearing aids are usually AR glasses.
[0003] The working principle of glasses-type hearing aids usually involves using a microphone to extract sound information in the usage scenario, and then outputting the sound information using a hearing aid or / and a display, so that people with hearing impairments can obtain the sound information extracted by the microphone in a timely manner. However, in the existing AR glasses, there is too much noise in the sound information extracted by the microphone, resulting in the information output by the hearing aid or / and the display being rather messy, which is not conducive to people with hearing impairments to accurately obtain information. For example, when there are many people and vehicles in the current scenario, the microphone may extract the sound information of multiple people and vehicles, resulting in very messy information output by the hearing aid or / and the display.
[0004] Therefore, there is a technical problem that the prompt information of the existing AR glasses is messy and not conducive to people with hearing impairments to distinguish. Summary of the Invention
[0005] An AR glasses that can identify ambient sounds and give multimodal prompts provided by the present invention solves the technical problem that the prompt information of the existing AR glasses is messy and not conducive to people with hearing impairments to distinguish.
[0006] Some implementation schemes for solving the above technical problems include: An AR glasses that can identify ambient sounds and give multimodal prompts, comprising a frame; Lenses, which are arranged on the frame; An information extraction component, which is installed on the frame and extracts ambient sound information; A prompt component, which is installed on the frame and emits prompt information; A positioning component, which is installed on the frame, and moreover, the positioning component determines the position information of the frame; And a data processing component, and the information extraction component, the positioning component, and the prompt component are all in communication with the data processing component; Among them, the AR glasses emit prompt information according to the following steps: The positioning component positions the location where the spectacle frame is located; Determine the noise frequency in the current scene according to the location where the spectacle frame is located; According to the determined noise frequency, the information extraction component extracts the sound information in the current scene and filters out the noise in the current scene to obtain the prompt information data; The prompt component outputs the prompt information data.
[0007] Preferably, the prompt component includes a text display and a sound output device, and the prompt component issues prompt information through the text display and the sound output device respectively.
[0008] Preferably, the sound output device includes an in-ear speaker and a bone conduction vibration player.
[0009] Preferably, the text display displays the prompt information data in text form on the lens.
[0010] Preferably, the prompt component outputting the prompt information data includes the following steps: The positioning component determines whether the current scene is indoors or outdoors; Determine the prompt type of the prompt component according to the positioning information determined by the positioning component. When the positioning component determines that the spectacle frame is indoors, the text display and the sound output device output prompt information at the same time. When the positioning component determines that the spectacle frame is outdoors, the text display does not output prompt information, and the sound output device outputs prompt information.
[0011] Preferably, when the positioning component determines that the spectacle frame is indoors, the text display and the sound output device output prompt information at the same time. The text display receives the prompt information text data, and the words of different parts of speech in the prompt information text data are displayed in different formats.
[0012] Preferably, the determining the noise frequency in the current scene according to the location where the spectacle frame is located includes the following steps: Extract the key sound frequencies in the current scene, where the key sound frequencies refer to the sound frequencies that will inevitably appear in the current scene; Determine the noise frequency in the current scene according to the key sound frequencies in the current scene. Among them, except for the key sound frequencies in the current scene, other sound frequencies are noise frequencies.
[0013] Preferably, the AR glasses further include an electronic fence, and the electronic fence communicates with the positioning component. Among them, the electronic fence is used to determine the current location where the spectacle frame is located. When the spectacle frame is within the area enclosed by the electronic fence, the spectacle frame is indoors.
[0014] Preferably, the sound outputter is a multi-channel sound outputter, and sounds of different frequencies are output through different channels.
[0015] Preferably, the information extractor includes at least two microphones, and the information extractor tracks the sound source direction in real time.
[0016] Compared with the prior art, the present invention has the following advantages: By setting the positioning component, the positioning component determines the position where the spectacle frame is located. The information extraction component can determine the noise frequency in the current scene according to the position where the spectacle frame is located, so that the sound extracted by the information extraction component is more accurate. Moreover, when filtering out the noise in the information extracted by the information extraction component, the noise removal is more accurate, effectively improving the accuracy of the prompt component's prompt, making the information prompted by the prompt component easier to distinguish, and optimizing the use performance of the AR glasses.
[0017] Determine the noise frequency in the current scene according to the position of the spectacle frame. Different scenes have different noise frequencies. Using the position where the spectacle frame is located to determine the noise frequency, the noise frequency convention is more accurate, and further making the prompt information issued by the prompt component more accurate. BRIEF DESCRIPTION OF THE DRAWINGS
[0018] For purposes of explanation, several embodiments of the technology of the present invention are set forth in the following drawings. The following drawings are incorporated into this text and form a part of the specific embodiments. In some cases, well-known structures and components are shown in block diagram form to avoid obscuring the concepts of the subject technology of the present invention.
[0019] Figure 1 It is a structural block diagram of the present invention.
[0020] Figure 2 It is a working flowchart of the present invention.
[0021] As shown in the figure: 1. Information extraction component, 2. Prompt component, 3. Positioning component, 4. Data processing component, 6. Electronic fence. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0022] The specific embodiments shown below are intended as descriptions of various configurations of the subject technology of the present invention, and are not intended to represent the only configurations in which the subject technology of the present invention can be practiced. The specific embodiments include specific details intended to provide a thorough understanding of the subject technology of the present invention. However, it will be clear and obvious to those skilled in the art that the subject technology of the present invention is not limited to the specific details shown herein, and can be practiced without these specific details.
[0023] It will be understood that, in this document, relational terms such as "first" and "second" are intended to distinguish one entity or operation from another entity or operation, and are not intended to imply any actual relationship or order between these entities or operations.
[0024] The term "comprising", "including" or any other variation thereof is intended to cover non-exclusive inclusion, such that a process, method, article or apparatus comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or apparatus. Without further limitation, an element defined by the phrase "comprising a..." does not exclude the presence of additional identical elements in the process, method, article or apparatus comprising the element.
[0025] Referring to Figures 1 to 2 as shown, an AR glasses for identifying environmental sounds and performing multimodal prompts, comprising a frame; lenses, the lenses being provided on the frame; an information extraction component 1, the information component being installed on the frame and extracting environmental sound information; a prompt component 2, the prompt component 2 being installed on the frame and emitting prompt information; a positioning component 3, the positioning component 3 being installed on the frame, and the positioning component 3 determining the position information of the frame; and a data processing component 4, the information extraction component 1, the positioning component 3 and the prompt component 2 all communicating with the data processing component 4; wherein, the AR glasses emit prompt information according to the following steps: the positioning component 3 locates the position where the frame is located; determine the noise frequency in the current scene according to the position where the frame is located; according to the determined noise frequency, the information extraction component 1 extracts the sound information in the current scene and filters out the noise in the current scene to obtain prompt information data; the prompt component 2 outputs the prompt information data.
[0026] In some embodiments, the prompt component 2 includes a text display and a sound output device, and the prompt component 2 emits prompt information through the text display and the sound output device respectively.
[0027] It can be understood that the lens is a transparent lens, and the lens has a display function, and the text display displays the prompt information on the lens in text form.
[0028] In some embodiments, the information extraction component 1 may be a sound extractor such as a microphone. Among them, the information extraction component 1 may have a pre-data processing function and can perform noise reduction processing on the extracted information. For example, noise in the extracted information can be directly filtered out.
[0029] In some embodiments, the data processing component 4 may be a cloud server. Alternatively, the data processing component 4 may also be an electronic component with data processing capabilities such as a data processing chip.
[0030] In some embodiments, the sound outputter includes an in-ear speaker and a bone conduction vibration player.
[0031] In some embodiments, the text displayer displays the prompt information data in text form on the lens.
[0032] In some embodiments, the prompt component 2 outputs the prompt information data including the following steps: The positioning component 3 determines whether the current scene is indoors or outdoors; Based on the positioning information determined by the positioning component 3, the prompt type of the prompt component 2 is determined. When the positioning component 3 determines that the frame is indoors, the text displayer and the sound outputter output prompt information simultaneously. When the positioning component 3 determines that the frame is outdoors, the text displayer does not output prompt information, and the sound outputter outputs prompt information.
[0033] In some embodiments, when the positioning component 3 determines that the frame is indoors, the text displayer and the sound outputter output prompt information simultaneously. The text displayer receives the prompt information text data, and words of different parts of speech in the prompt information text data are displayed in different formats.
[0034] In some embodiments, determining the noise frequency in the current scene according to the position of the frame includes the following steps: Extract the key sound frequencies in the current scene, where the key sound frequencies refer to the sound frequencies that will necessarily appear in the current scene; Based on the key sound frequencies in the current scene, determine the noise frequency in the current scene. Among them, except for the key sound frequencies in the current scene, other sound frequencies are noise frequencies.
[0035] In some embodiments, the AR glasses further include an electronic fence 6, and the electronic fence 6 communicates with the positioning component 3. Among them, the electronic fence 6 is used to determine the current position of the frame. When the frame is within the area enclosed by the electronic fence 6, the frame is indoors.
[0036] In some embodiments, the sound outputter is a multi-channel sound outputter, and sounds of different frequencies are output through different channels.
[0037] In some embodiments, the information extractor includes at least two microphones, and the information extractor tracks the sound source direction in real time.
[0038] Understandably, the information extractor tracks the sound source direction in real time. For example, the microphone can have a rotation function. Or, there can be multiple microphones, and the multiple microphones face different directions to extract sounds from different directions.
[0039] In some embodiments, the sound outputter can have the following functions: Noise reduction function: A high-quality sound outputter has strong noise reduction ability, such as adaptive noise reduction technology, which can automatically adjust according to the environment, effectively reduce background noise, and improve speech clarity.
[0040] Sound quality performance: The sound quality of the sound outputter is natural, clear, and soft, with low distortion, and can accurately restore sounds, allowing the wearer to obtain a good auditory experience.
[0041] Multi-channel dynamic compression: The sound outputter usually divides the sound into multiple independent channels according to frequency, and adjusts the gain and compression ratio of each channel separately to achieve "amplifying soft sounds and limiting loud sounds", thereby improving the speech recognition rate.
[0042] The more channels the sound outputter has, the finer the adjustment of sounds at different frequencies, which can better match the individual's hearing condition, and improve the wearing comfort and listening effect.
[0043] The frequency response range of the sound outputter should cover most of the speech frequency range. For example, the low frequency can reach about 200 Hz, and the high frequency can reach about 8000 Hz to ensure the accuracy of sound restoration. Maximum sound output and sound gain: The maximum sound output and sound gain of the sound outputter should be below the patient's discomfort threshold to ensure that the hearing is not damaged.
[0044] Equivalent input noise level and total harmonic distortion: The equivalent input noise level and total harmonic distortion of the sound outputter should be controlled within a certain range to ensure operation in a low-noise environment and reduce background noise interference.
[0045] The information extraction component 1 can perform adaptive beamforming: The dual microphone array tracks the sound source direction in real time, enhances the forward gain, suppresses the backward noise, and improves the signal-to-noise ratio.
[0046] In some embodiments, the data processing component 4 has AI scene recognition: It automatically recognizes different environments (such as quiet, meeting, restaurant, outdoor, music, vehicle-mounted, etc.) through machine learning, switches the best parameter combination within 0.2 seconds, optimizes the auditory experience. The data processing component 4 uses an advanced chip with a fast operation speed, can process sound signals more precisely, and improves the overall performance of the sound outputter.
[0047] The technical solution of the present invention will be further introduced below in combination with specific application scenarios: In a specific scenario, an electronic fence 6 can be set. For example, an electronic fence 6 can be set in the environment where hearing-impaired people work.
[0048] When a hearing-impaired person enters the area enclosed by the electronic fence 6, the positioning component 3 and the electronic fence 6 simultaneously determine the position of the spectacle frame, that is, the spectacle frame enters the environment where the hearing-impaired person works. At this time, the data processing component 4 and the information extraction component 1 can determine the noise information in the working environment. For example, when the working environment of the hearing-impaired person is close to a road, both the information extraction component 1 and the data processing component 4 can directly identify the sounds emitted by vehicles and other means of transportation on the road, as well as wild animals such as birds, as noise frequencies. At this time, when the hearing-impaired person uses the AR glasses in the working environment, the information extraction component 1 only extracts human voices. Or, in order to further improve the accuracy of the voices extracted by the information extraction component 1, only human voices within a certain range can be extracted.
[0049] When the positioning component 3 and the electronic fence determine that the spectacle frame leaves the working environment of the hearing-impaired person, and the positioning component 3 determines that the position of the spectacle frame is outdoors, considering the safety needs of the hearing-impaired person outside, at this time, the information extraction component 1 and the data processing component 4 can determine human voices as noise information, so that the hearing-impaired person can more accurately obtain the sounds emitted by vehicles, etc., and prevent the hearing-impaired person from being hit by vehicles.
[0050] The above application examples are only exemplary applications of sounds. The hearing-impaired person can determine the noise frequency according to different environmental requirements. That is, after determining the sound frequency that will definitely appear in the current scenario, the noise frequency can be determined, thereby improving the correctness of noise removal and making the prompts of the prompt component 2 more accurate.
[0051] The above has introduced the technical solution of the present invention theme and the corresponding details. It can be understood that the above introduction is only some implementation schemes of the technical solution of the present invention theme, and some details can also be omitted during its specific implementation.
[0052] In addition, in some implementation schemes of the above invention, it is possible to combine multiple implementation schemes. Due to space limitations, various combination schemes are not listed one by one. Those skilled in the art can freely combine and implement the above implementation schemes according to needs during specific implementation to obtain a better application experience.
[0053] When implementing the technical solution of the present invention theme, those skilled in the art can obtain other detailed configurations or drawings according to the technical solution of the present invention theme and the drawings. Obviously, without departing from the technical solution of the present invention theme, these details still fall within the scope covered by the technical solution of the present invention theme.
Claims
1. An AR glasses that recognizes environmental sounds and gives multimodal prompts, characterized in that: It includes a frame; lenses, which are arranged on the frame; an information extraction component (1), which is installed on the frame and extracts ambient sound information; a prompt component (2), which is installed on the frame and emits prompt information; a positioning component (3), which is installed on the frame, and the positioning component (3) determines the position information of the frame; and a data processing component (4), and the information extraction component (1), the positioning component (3) and the prompt component (2) are all in communication with the data processing component (4); wherein, the AR glasses emit prompt information according to the following steps: the positioning component (3) locates the position where the frame is located; determines the noise frequency in the current scene according to the position where the frame is located; according to the determined noise frequency, the information extraction component (1) extracts the sound information in the current scene and filters out the noise in the current scene to obtain prompt information data; the prompt component (2) outputs the prompt information data.
2. The AR glasses for recognizing environmental sounds and performing multimodal prompts according to claim 1, characterized in that: The prompt component (2) includes a text display and a sound outputter, and the prompt component (2) emits prompt information through the text display and the sound outputter respectively.
3. The AR glasses for recognizing environmental sounds and performing multimodal prompts according to claim 2, characterized in that: The sound outputter includes an in-ear speaker and a bone conduction vibration player.
4. The AR glasses for recognizing environmental sounds and performing multimodal prompts according to any one of claims 1 to 3, characterized in that: The text display displays the prompt information data in text form on the lens.
5. The AR glasses for identifying environmental sounds and performing multimodal prompts according to claim 2, characterized in that: The output of the prompt information data by the prompt component (2) includes the following steps: the positioning component (3) determines whether the current scene is indoors or outdoors; determines the prompt type of the prompt component (2) according to the positioning information determined by the positioning component (3). When the positioning component (3) determines that the frame is indoors, the text display and the sound outputter output prompt information at the same time. When the positioning component (3) determines that the frame is outdoors, the text display does not output prompt information, and the sound outputter outputs prompt information.
6. The AR glasses for recognizing environmental sounds and performing multimodal prompts according to claim 5, characterized in that: When the positioning component (3) determines that the frame is indoors, the text display and the sound outputter output prompt information at the same time. The text display receives the prompt information text data, and the words of different parts of speech in the prompt information text data are displayed in different formats.
7. The AR glasses for identifying environmental sounds and performing multimodal prompts according to claim 1, characterized in that: The determination of the noise frequency in the current scene according to the position where the frame is located includes the following steps: extracting the key sound frequency in the current scene, where the key sound frequency refers to the sound frequency that will necessarily appear in the current scene; determining the noise frequency in the current scene according to the key sound frequency in the current scene, where, except for the key sound frequency in the current scene, other sound frequencies are noise frequencies.
8. The AR glasses for identifying environmental sounds and performing multimodal prompts according to claim 1, characterized in that: The AR glasses further include an electronic fence (6), and the electronic fence (6) is in communication with the positioning component (3), where the electronic fence (6) is used to determine the position where the current frame is located. When the frame is within the area enclosed by the electronic fence (6), the frame is indoors.
9. The AR glasses for identifying environmental sounds and performing multimodal prompts according to claim 2, wherein: The sound outputter is a multi-channel sound outputter, and sounds of different frequencies are output through different channels.
10. The AR glasses for recognizing environmental sounds and performing multimodal prompts according to claim 1, characterized in that: The information extractor includes at least two microphones, and the information extractor tracks the sound source direction in real time.