A sound early warning method based on augmented reality glasses and mobile terminal cooperation

CN122511297APending Publication Date: 2026-08-04ZHIZHENRUIHONG ARTIFICIAL INTELLIGENCE TECHNOLOGY (SHANGHAI) CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ZHIZHENRUIHONG ARTIFICIAL INTELLIGENCE TECHNOLOGY (SHANGHAI) CO LTD
Filing Date
2026-06-10
Publication Date
2026-08-04

AI Technical Summary

Technical Problem

若在眼镜端实时运行声音识别和定位模型,会导致功耗显著上升,续航时间缩短至1-2小时,无法满足全天候使用需求

Benefits of technology

[0010] Compared with the prior art, the sound warning method provided by the present disclosure has the following beneficial effects.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122511297A_ABST
    Figure CN122511297A_ABST
Patent Text Reader

Abstract

A sound warning method based on the collaboration between augmented reality glasses and a mobile terminal is disclosed. The system includes: augmented reality glasses for collecting ambient audio and presenting warning information in a pop-up window; and a mobile terminal connected to the augmented reality glasses via Bluetooth, used to call a sound classification model to identify the sound type, call a self-learning sound distance determination model based on energy attenuation characteristics and frequency distortion characteristics to determine the direction and distance of the sound source, and send a pop-up window command to the augmented reality glasses after stabilization processing through a de-jitter state machine. This disclosure reduces the energy consumption of the augmented reality glasses by migrating the computational task to the mobile terminal; utilizes a self-learning distance determination model to achieve device ranging; allows users to download and update the model via a mobile APP to customize the recognition type; effectively suppresses false alarms, and provides accurate and real-time ambient sound warnings for hearing-impaired individuals and users in complex environments.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments disclosed herein relate to the fields of sound signal processing and augmented reality technology, and in particular to a sound warning method based on the collaboration between augmented reality glasses and a mobile terminal. Background Technology

[0002] With the maturation of augmented reality (AR) technology, AR glasses are increasingly being applied in fields such as hearing assistance, communication in noisy environments, and intelligent early warning systems. However, the existing technology still has the following shortcomings.

[0003] First, augmented reality glasses are limited by size and battery capacity, resulting in limited local computing power. Running sound recognition and localization models in real time on the glasses would significantly increase power consumption, reducing battery life to 1-2 hours, which is insufficient for all-day use. Second, most existing sound source localization technologies rely on multi-microphone array time difference of arrival (TDOA) or angle of arrival (DOA) algorithms, requiring at least two spatially separated sound source devices or fixed base stations, making them unsuitable for single-device mobile scenarios like augmented reality glasses. Third, traditional sound recognition models cannot be updated once deployed, preventing users from customizing the types of sounds they need to recognize based on their specific life scenarios (such as needing to monitor a newborn's cries at home or having specific warning needs in their community). Finally, existing warning systems directly output the recognition results, lacking effective filtering of transient noise, which can easily lead to repeated pop-ups and false alarms, negatively impacting the user experience.

[0004] Therefore, how to achieve real-time recognition of ambient sounds, single-device distance estimation, user-customized updates, and false alarm suppression while ensuring low power consumption of augmented reality glasses is a technical problem that urgently needs to be solved in this field. Summary of the Invention

[0005] This disclosure provides a sound warning method based on the collaboration between augmented reality glasses and a mobile terminal. This method can offload computing tasks to the mobile terminal, use a self-learning distance determination model to achieve single-device distance measurement, support dynamic model updates, and reduce false alarms through a jitter-reducing state machine, thereby improving the practicality of augmented reality glasses and user experience. Technical solution

[0006] According to a first aspect of this disclosure, a sound warning method based on the collaboration between augmented reality glasses and a mobile terminal is provided, comprising: augmented reality glasses for collecting ambient audio and presenting warning information in the form of a pop-up window; a mobile terminal connected to the augmented reality glasses via Bluetooth for calling a sound classification model to determine the sound type, calling a self-learning sound distance determination model based on energy / frequency characteristics to determine the direction and distance of the sound source, and sending a pop-up window command to the augmented reality glasses after processing by a de-jitter state machine; the sound classification model can be updated by downloading from the network of the mobile terminal.

[0007] According to a second aspect of this disclosure, a sound warning method based on the collaboration between augmented reality glasses and a mobile terminal is provided, comprising: an audio acquisition and transmission step, a sound type identification step, a sound source direction and distance determination step, a shake-reduction and stability processing step, and a warning generation and display step.

[0008] According to a third aspect of this disclosure, an electronic device is provided, including a processor and a memory, wherein the processor performs the methods described above.

[0009] According to a fourth aspect of this disclosure, a non-transitory computer-readable storage medium is provided. Beneficial effects

[0010] Compared with the prior art, the sound warning method provided by the present disclosure has the following beneficial effects.

[0011] (1) Low power consumption: Augmented reality glasses only perform audio acquisition and Bluetooth transmission, without performing local calculations. According to actual tests, the battery life of glasses can be increased from 1.5 hours to more than 6 hours.

[0012] (2) Single-device passive ranging: Single-device distance estimation is achieved by using a self-learning model based on energy attenuation and frequency distortion, without relying on multiple base stations or sound source devices, and is suitable for mobile scenarios.

[0013] (3) Customizable: Users can select and download the required sound types through a mobile APP. The model supports incremental updates to meet personalized early warning needs.

[0014] (4) False alarm suppression: By effectively filtering out instantaneous false triggers caused by environmental noise through the de-jitter state machine, the false alarm rate is reduced by about 81%, improving the accuracy of early warning.

[0015] (5) Hearing impairment assistance: Display direction and distance information on augmented reality glasses in the form of pop-up text, etc., to help people with hearing impairments obtain key sound information in noisy environments or during telephone calls, thereby improving their quality of life. Attached Figure Description

[0016] To more clearly illustrate the technical solutions in the embodiments of this disclosure, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this disclosure. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Figure 1 is a flowchart illustrating a sound warning method based on the collaboration between augmented reality glasses and a mobile terminal disclosed herein.

[0017] Figure 2 This is a schematic diagram of a sound warning method provided in one embodiment of the present disclosure.

[0018] Figure 3 This is a schematic flowchart of a sound warning method provided in one embodiment of the present disclosure.

[0019] Figure 4 A schematic diagram illustrating the training and inference principles of a self-learning distance determination model in a sound warning method provided for another embodiment of this disclosure.

[0020] Figure 5 The working state transition diagram of the debouncing state machine provided in one embodiment of this disclosure.

[0021] Figure 6 This is a schematic diagram illustrating the pop-up warning effect of augmented reality glasses according to one embodiment of the present disclosure. Detailed Implementation

[0022] The technical solutions of the embodiments of this disclosure will be clearly described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this disclosure, and not all embodiments. Based on the embodiments of this disclosure, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this disclosure.

[0023] One embodiment of this disclosure provides a sound warning method, such as... Figure 1 As shown, the system includes augmented reality glasses 110, a mobile terminal 120, and a cloud server 130. The mobile terminal 120 has an application (APP) installed to accompany the augmented reality glasses 110. The augmented reality glasses 110 are wirelessly connected to the mobile terminal 120 via Bluetooth, and the mobile terminal 120 is connected to the cloud server 130 via WiFi or a cellular network.

[0024] The augmented reality glasses 110 are equipped with a linear microphone array, for example, four MEMS microphones are arranged along the upper edge of the frame or on both temples, for collecting ambient audio. Unlike traditional solutions, the augmented reality glasses 110 in this disclosure does not perform any audio recognition or positioning calculations, but only performs audio acquisition, Bluetooth transmission, and pop-up display to minimize power consumption.

[0025] The mobile terminal 120 stores pre-trained sound classification and sound distance determination models. Upon receiving audio data from the augmented reality glasses 110, the mobile terminal 120 first extracts 40-dimensional log-Mel spectrum features and inputs them into the sound classification model trained based on CRNN for type discrimination. If the sound is identified as a sound category requiring warning, such as car horns, baby crying, or fire alarms, the sound distance determination model is further invoked. This model takes the energy attenuation rate of the audio and the power spectral density of a specific frequency band as input and outputs an estimated distance (in meters). Simultaneously, it uses the dual-microphone energy difference (ILD) or phase difference (IPD) to estimate the direction of the sound source (left / right / front / back, etc.). Afterward, the discrimination result enters the de-jitter state machine: a time window T and a confirmation threshold K are set. Only when the same sound category appears 2-4 times consecutively within 2 seconds is it confirmed as a valid event; otherwise, it is considered transient noise and discarded. For valid events, the mobile terminal 120 generates a pop-up command, such as in JSON format: {"category":"car_horn", "direction":"left_rear", "distance":12.5}. This command is sent to the augmented reality glasses 110 via Bluetooth. The augmented reality glasses 110 renders the command as a text pop-up and displays it at the appropriate position within the lenses' field of view.

[0026] In addition, users can select the type of sound they need to be alerted via an app on mobile terminal 120, such as selecting options like "baby crying" or "fire alarm". The app downloads the corresponding model update data from cloud server 130, supports incremental updates, and replaces the local model, thereby achieving self-customized expansion.

[0027] One embodiment of this disclosure provides a sound warning method applied to the aforementioned system. For example... Figure 2 As shown, the method includes the following steps.

[0028] Step S201: Audio Acquisition and Transmission. The augmented reality glasses acquire ambient audio in real time through their microphone array, with a sampling rate of 16kHz, and package it into a frame every 200ms, which is then transmitted to the mobile terminal via Bluetooth 5.3.

[0029] Step S202: Sound Type Recognition. After receiving the audio frame, the mobile terminal extracts the log-Melogram features and inputs them into the sound classification model. This model is based on a convolutional recurrent neural network, containing 3 convolutional layers and 2 LSTM layers. It outputs the probability of each category; when the probability of the "car horn" category is >0.85, it is considered valid.

[0030] Step S203: Sound Source Direction and Distance Determination. The mobile terminal extracts the RMS energy value of the audio frame, representing energy attenuation and frequency band energy distribution, and frequency distortion, and inputs it into the sound distance determination model. This model is a support vector regression model, pre-trained with 100 sets of samples collected at different distances (1m, 5m, 10m, 20m, 30m) for each sound type, establishing an energy-distance regression curve. During inference, the estimated distance is output based on the current energy value. Direction is determined based on the energy ratio of the left and right microphones: left energy - right energy > 3dB is determined as "left", otherwise "right". Combining the front and rear microphone groups can further distinguish "left front" / "left rear", etc.

[0031] Step S204: Debouncing Stability Processing. The outputs of steps S202 and S203 are fed into the debouncing state machine. For example... Figure 4 As shown, the state machine starts from the "idle" state. When a sound category is detected for the first time, it enters the "detecting" state and starts the timer. Within the time window T, the counter is incremented by 1 for each detection of the same category. If the counter reaches the threshold K, it enters the "trigger warning" state. If the time window expires and the counter has not reached the threshold, it returns to "idle" and the counter is cleared.

[0032] Step S205: Warning Generation and Display. After the warning is triggered, the mobile terminal generates a pop-up command and sends it to the augmented reality glasses. The glasses render the command as a text pop-up window, such as "There is a vehicle horn approximately 12 meters to your left rear, please be careful," which is displayed for 2 seconds and then automatically fades out.

[0033] In one specific embodiment, a hearing-impaired user wears the augmented reality glasses disclosed herein and connects them to a mobile phone while walking on the street. When a car horn sounds 15 meters to the right rear, the glasses' microphone array captures the audio, the mobile phone identifies it as a "car horn," the distance determination model outputs 15.2 meters, and the direction is determined to be "right rear." After the anti-shake state machine confirms three times consecutively within one second, a yellow text message pops up on the glasses lens: "Vehicle horn approximately 15 meters to the right rear, please be careful." The user sees the prompt and promptly avoids the vehicle.

[0034] In another embodiment, a user with an infant at home downloads an additional "baby crying" recognition model via a mobile app. When the baby cries in another room, the glasses recognize the crying and display "Baby crying approximately 5 meters to the left front," while the phone vibrates to alert the user and help them check on the baby promptly.

[0035] In another embodiment, while the user is on a phone call, the system disclosed herein can simultaneously process ambient sound and phone audio. A mobile app monitors the call status, and when a call is in progress, automatically transcribes the other party's speech into text and displays it on the glasses, providing call caption assistance.

[0036] One embodiment of this disclosure also provides a sound warning method. The device can be deployed on a mobile terminal and includes: an audio receiving module, a sound classification module, a distance determination module, a de-jitter module, and a command sending module. The functions of each module correspond to the steps of the method described above, and will not be repeated here.

[0037] This disclosure also provides an electronic device including a processor and a memory, the processor performing the methods described above. The electronic device may be augmented reality glasses or a mobile terminal.

[0038] This disclosure also provides a non-transitory computer-readable storage medium having computer-executable instructions stored thereon, which, when executed by a processor, implement the above-described method.

[0039] It is understood that the specific examples in this document are merely to help those skilled in the art better understand the embodiments of this disclosure, and are not intended to limit the scope of the invention. The various embodiments described in this disclosure can be implemented individually or in combination.

[0040] The above description is merely a specific embodiment of this disclosure, but the scope of protection of this invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this disclosure should be included within the scope of protection of this disclosure. Therefore, the scope of protection of this invention should be determined by the scope of the claims.

Claims

1. A sound warning method based on the collaboration between augmented reality glasses and a mobile terminal, characterized in that, include: Augmented reality glasses are used to collect ambient audio signals, process the ambient audio signals into frames, and transmit them to a mobile terminal via wireless communication; wherein, the augmented reality glasses do not perform sound recognition and sound source localization calculations, but only retain the functions of audio acquisition, basic preprocessing, and warning information display; A mobile terminal, which is communicatively connected to the augmented reality glasses, is used to receive the framed environmental audio signal and process the environmental audio signal based on a preset stored sound processing model. The mobile terminal includes a processor and a memory, and the memory stores a sound classification model, a sound source distance estimation model, and a sound source direction estimation model. The processor is configured to perform the following processing steps: An acoustic feature vector is extracted from the received ambient audio signal. The acoustic feature vector includes at least spectral energy distribution features and energy attenuation features. Based on the sound classification model, the acoustic feature vector is used to identify the sound category, and the corresponding sound category and confidence level are obtained. When the recognition result meets the preset category conditions, the sound source distance estimation model and the sound source direction estimation model are invoked to perform joint reasoning on the acoustic feature vector and output the distance information and direction information of the sound source relative to the augmented reality glasses. The sound source distance estimation model establishes a distance mapping function based on the acoustic propagation loss constraint relationship. The consistency of the sound category, distance information and direction information is determined based on the time continuity constraint mechanism. The stability of the continuous output results is evaluated within a preset time window. When the same sound category and its corresponding spatial location information meet the consistency threshold condition within the time window, it is determined to be a valid sound source event. The effective sound source event is converted into a warning command and sent to the augmented reality glasses via wireless communication, so as to control the augmented reality glasses to output the corresponding warning information in a visual pop-up window. The sound classification model supports downloading and updating via the network, and performs model replacement or incremental updates on the mobile terminal side to expand the set of recognizable sound categories.

2. The system according to claim 1, characterized in that, The sound distance determination model is a self-learning distance determination model. Its training process includes: collecting audio samples at different distances for each identifiable sound type, extracting the energy attenuation features and frequency distortion features of the samples, and training a regression model to establish a mapping relationship between energy / frequency features and distance.

3. The system according to claim 1, characterized in that, The de-shaking state machine is configured to: set a time window T and an acknowledgment threshold K; when the same sound category is received continuously for K times within the time window T, it is determined to be a valid event and a pop-up warning is triggered. Otherwise, it will be considered transient noise and will not trigger a pop-up warning.

4. The system according to claim 1, characterized in that, The energy consumption optimization strategy for the augmented reality glasses is as follows: the augmented reality glasses do not perform any sound recognition or distance calculation functions, but only perform audio acquisition, Bluetooth transmission and pop-up display functions; all sound classification models and distance determination models are stored and run on the mobile terminal.

5. The system according to claim 1, characterized in that, The update mechanism of the sound classification model includes: the user selects the type to be identified from the list of selectable warning sound types through the application interface of the mobile terminal; the mobile terminal downloads the corresponding type of model update data from the server via the Internet; the updated sound classification model replaces the original model, realizing the user's customized warning type expansion.

6. The system according to claim 1, characterized in that, The warning information format in the pop-up command is: sound category identifier + direction indicator + distance estimate, for example, "Vehicle horn detected approximately 12 meters to the left rear, please be aware".

7. A sound warning method based on the collaboration between augmented reality glasses and a mobile terminal, applied to the system described in any one of claims 1 to 6, characterized in that, Includes the following steps: Step S1, Audio Acquisition and Transmission: The augmented reality glasses acquire environmental audio data in real time through their microphone array and send the acquired audio data to the mobile terminal via Bluetooth communication; Step S2, Sound Type Recognition: The mobile terminal receives the audio data, calls the pre-trained sound classification model stored in the local memory, performs sound type identification on the audio data, and outputs one or more sound categories and their corresponding confidence scores; Step S3, Sound Source Direction and Distance Determination: The mobile terminal calls the pre-trained sound distance determination model stored in the local memory, extracts the energy attenuation characteristics and frequency distortion characteristics of the audio data, and determines the azimuth and distance of the sound source relative to the augmented reality glasses based on the output of the sound distance determination model. Step S4, De-jitter Stability Processing: Input the discrimination results output from Step S2 and Step S3 into the de-jitter state machine, count the continuously output discrimination results within a preset time window, and when the cumulative count of the same sound category reaches a preset threshold, the sound event is confirmed as a valid event; otherwise, it is considered invalid noise. Step S5, Warning Generation and Display: For a confirmed valid judgment result, the mobile terminal generates a pop-up command containing sound category, azimuth angle and distance information, and sends it to the augmented reality glasses via Bluetooth. The display screen of the augmented reality glasses presents the warning information in the form of text and numbers.

8. The method according to claim 7, characterized in that, The input features of the sound distance determination model in step S3 include: the energy attenuation rate of the audio signal, the power spectral density of a specific frequency band, and a library of energy-distance regression coefficients for various sound types.

9. The method according to claim 7, characterized in that, In step S4, the preset time window for the de-jitter state machine is 0.5 to 2 seconds, and the preset threshold number of times is 2 to 4.

10. The method according to claim 7, characterized in that, Also includes: The mobile terminal receives the user's model update instruction through its application interface, downloads the model parameters for the new sound type from the cloud server, and incrementally updates the sound classification model.

11. The method according to claim 7, characterized in that, Also includes: The augmented reality glasses use beamforming technology through their linear microphone array to directionally amplify ambient audio in a specific direction and send the enhanced audio to the mobile terminal.

12. The method according to claim 11, characterized in that, Also includes: The augmented reality glasses utilize a deep learning model or a machine learning model to adjust beamforming parameters, wherein the beamforming parameter adjustment includes adaptively changing the weighting coefficients of the microphone array based on the environmental signal-to-noise ratio.

13. The method according to claim 7, characterized in that, Also includes: The augmented reality glasses receive subtitle adjustment instructions sent by the mobile terminal and adjust the position, size, or color of the pop-up subtitles according to the subtitle adjustment instructions.

14. An electronic device, characterized in that, The electronic device includes: processor; Memory used to store the processor's executable instructions; The processor is used to execute the sound warning method as described in claim 7.

15. The electronic device according to claim 14, characterized in that, The electronic device includes augmented reality glasses.

16. The electronic device according to claim 15, characterized in that, The augmented reality glasses include a linear microphone array comprising multiple microphone sensors distributed along a straight line.

17. The electronic device according to claim 15, characterized in that, The augmented reality glasses include those worn by people with hearing impairments.

18. An electronic device, characterized in that, The electronic device includes: processor; Memory used to store the processor's executable instructions; The processor is used to perform the mobile terminal-side operation in the method of claim 10.

19. A sound warning system, characterized in that, include: Augmented reality glasses, mobile devices, and cloud servers. The augmented reality glasses are used to collect ambient audio and send it to the mobile terminal via Bluetooth. The mobile terminal is used to call the locally stored sound classification model to identify the type of audio, call the locally stored sound distance determination model to determine the direction and distance of the sound source, and generate a pop-up command to send to the augmented reality glasses after performing stability judgment through the anti-shake state machine. The mobile terminal is also used to download updated data of the sound classification model from the cloud server via the Internet to expand the recognition types; The augmented reality glasses are used to display the warning information in the pop-up command via subtitles.

20. A non-transitory computer-readable storage medium having computer-executable instructions stored thereon, characterized in that, When the computer-executable instructions are executed by a processor, they implement the method described in any one of claims 7 to 13.