An AI intelligent ambient sound synchronization system and method for over-ear wearable devices
By integrating an external microphone, audio processing module, and speaker into an over-ear wearable device, and combining noise suppression and sound event detection technologies, the problem of over-ear devices being unable to distinguish between harmful noise and valuable sound has been solved, achieving intelligent pass-through and enhanced security.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SUZHOU AIDOMUKE INTELLIGENT TECHNOLOGY CO LTD
- Filing Date
- 2026-03-17
- Publication Date
- 2026-06-16
AI Technical Summary
Existing over-ear wearable devices, while providing an immersive audio experience, cannot intelligently distinguish between harmful noise and valuable sounds, causing critical information to be drowned out by useless noise, affecting user safety and communication.
It employs an external microphone, audio processing module, control unit, and built-in speaker, combined with noise suppression, sound event detection and enhancement units. It uses a lightweight neural network and adaptive filter to identify and enhance valuable sounds, and outputs clear ambient sound after mixing.
It achieves intelligent pass-through of valuable sounds, improves user security and information transmission efficiency, provides a better user experience, is compatible with a variety of over-ear devices, and has a high degree of customizability and personalized control.
Smart Images

Figure CN122227133A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of audio processing and wearable device technology, specifically an AI intelligent ambient sound synchronization system and method for over-ear wearable devices. Background Technology
[0002] Wearable devices that cover the ears, such as noise-canceling headphones, safety helmets, and augmented reality / virtual reality (AR / VR) headsets, generally employ physical soundproofing structures and active noise cancellation technology to isolate external ambient noise in order to provide users with an immersive audio experience. However, this complete auditory isolation presents significant safety hazards and inconveniences in many everyday and professional scenarios: pedestrians wearing noise-canceling headphones may not be able to hear vehicle horns or bicycle bells; industrial workers wearing protective earmuffs may miss abnormal equipment sounds or colleagues' emergency calls; office workers using noise-canceling headphones may not be able to detect fire alarms or colleagues approaching. All of these situations can adversely affect user safety or normal communication.
[0003] Existing technologies include "transparency modes" or "ambient sound modes," but these typically simply pick up external sounds with a microphone, amplify them, and play them to the user without intelligent sound filtering. This results in users hearing unprocessed ambient sound containing a large amount of useless noise, leading to a poor user experience in noisy environments. Crucially, the truly important sound information is drowned out by the noise, failing to effectively transmit environmental information. Therefore, there is an urgent need in this field for a technical solution that can intelligently distinguish between "harmful noise" and "valuable sound," and deliver the latter clearly and customizablely to address the shortcomings of existing technologies. Summary of the Invention
[0004] The purpose of this invention is to provide an AI intelligent ambient sound synchronization system and method for ear-covering wearable devices, so as to solve the problems mentioned in the background art.
[0005] To achieve the above objectives, the present invention provides the following technical solution: An AI-powered intelligent ambient sound synchronization system for over-ear wearable devices includes an external microphone, an audio processing module, a control unit, a mixing module, and a built-in speaker. The external microphone is used to collect ambient sound. The audio processing module includes a noise suppression unit and a sound event detection and enhancement unit. The noise suppression unit is used to filter out broadband steady-state noise in ambient sound, and the sound event detection and enhancement unit is used to identify valuable sound events in ambient sound and enhance them. The control unit is used to coordinate the working sequence of the external microphone, audio processing module and mixing module, and to store preset valuable sound event categories and response strategies. The mixing module is used to mix the processed ambient sound output by the audio processing module with the built-in audio as needed; The built-in speaker is used to play the mixed audio output from the mixing module.
[0006] As a further aspect of the present invention: the external microphone is a microphone array that supports beamforming technology and is capable of directional sound pickup.
[0007] As a further aspect of the present invention, the noise suppression unit employs at least one noise reduction algorithm, such as spectral subtraction or adaptive filtering, and can achieve precise noise suppression by combining a preset noise model database or signal correlation analysis.
[0008] As a further aspect of the present invention: the sound event detection and enhancement unit includes a feature extraction subunit, a classification and recognition subunit, and a selective enhancement subunit; The feature extraction subunit is used to extract at least one of the following: Mel frequency cepstral coefficients, time-domain peak features, and modulation frequency features of the audio signal; The classification and recognition subunit adopts any one of the following: lightweight neural network model, support vector machine, and convolutional neural network, and is trained with scene sound data of a preset duration. The selective enhancement subunit is used to boost the gain and optimize the frequency band of the identified valuable sound events.
[0009] As a further aspect of the present invention: the control unit supports users to manually switch working modes and is connected to sensors or indicator lights. The sensors are used to acquire external parameters of the device's usage scenario, and the indicator lights are used to provide visual reminders when valuable sound events are detected.
[0010] As a further aspect of the present invention: the mixing module allows users to adjust the mixing ratio of processed ambient sound and built-in audio, and presets high-priority, valuable sound event response rules: When a high-priority, valuable sound event is detected, the built-in audio volume is automatically adjusted or the built-in audio is muted.
[0011] An AI intelligent ambient sound synchronization method for over-ear wearable devices, applied to the AI intelligent ambient sound synchronization system for over-ear wearable devices as described above, includes the following steps: S1: External microphone collects ambient sound signals in real time; S2: The audio processing module filters out broadband steady-state noise in the ambient sound through the noise suppression unit, and then identifies valuable sound events and performs enhancement processing through the sound event detection and enhancement unit to obtain the processed ambient sound. S3: The control unit coordinates the mixing module to mix the processed ambient sound with the built-in audio according to the preset response strategy; S4: Built-in speaker plays the mixed audio signal.
[0012] As a further aspect of the present invention: the valuable sound events include at least one of traffic warning sounds, equipment malfunction sounds, voice commands, and emergency alarm sounds; The classification and recognition subunit trains a model based on scene sound data of a preset duration to achieve accurate recognition of the valuable sound events.
[0013] As a further aspect of the present invention: the mixing ratio of the mixing module can be customized and adjusted by the user, and when a high-priority valuable sound event is detected, a preset response rule is automatically executed to reduce the built-in audio volume or mute the built-in audio.
[0014] As a further aspect of the present invention: the control unit acquires external parameters through connected sensors and dynamically adjusts the gain amplitude and mixing ratio of the processed ambient sound. The external parameters include at least one of vehicle speed and usage scenario type.
[0015] Compared with the prior art, the beneficial effects of the present invention are: 1. Wide range of applications: It can be widely used in various over-ear devices such as noise-canceling headphones, industrial safety helmets, motorcycle helmets and AR / VR devices, and has strong versatility; 2. Intelligent selective pass-through: Unlike a simple "pass-through mode", this invention can intelligently filter out useless noise and only amplify the specific sounds that the user needs to hear, resulting in high information transmission efficiency and a better user experience; 3. Proactive Security: By identifying key sounds such as alarms and shouts, and amplifying or prioritizing their playback, the system transforms passive reception into proactive alerts, greatly enhancing user security in various scenarios. 4. Highly customizable: Users or equipment manufacturers can define the categories of "valuable sounds" that need to be listened to and set corresponding response strategies according to specific application scenarios (such as urban commuting, industrial workshops and outdoor sports); 5. Intelligent noise reduction: Through advanced audio processing algorithms, it selectively eliminates noise such as wind noise, rather than simply blocking or retaining it all, ensuring that the transmitted ambient sound is clear and useful. 6. Excellent user experience: Users can customize the loudness and mixing ratio of ambient sounds, achieving a personalized balance between safety and entertainment; the automatic gain control function ensures optimal performance at different speeds. 7. High structural integration: All components can be integrated into the existing structure of smart wearable devices without compromising the integrity and security of the device. Attached Figure Description
[0016] Figure 1 This is a schematic diagram of the connection structure of the AI intelligent ambient sound synchronization system for an over-ear wearable device in an embodiment of the present invention.
[0017] Figure 2 This is a schematic diagram of the audio signal processing link principle of the AI intelligent ambient sound synchronization system for ear-covering wearable devices in an embodiment of the present invention.
[0018] Figure 3 This is a schematic diagram of the process steps of the AI intelligent ambient sound synchronization method for ear-covering wearable devices in an embodiment of the present invention.
[0019] Figure 4 This is a frequency-phase characteristic curve of the core acoustic component of the system in this embodiment of the invention; Wherein, A is the amplitude-frequency response curve of the system's core acoustic components (external microphone, built-in speaker) under a sound pressure level input of 94dB, and B is the phase-frequency response curve of the corresponding acoustic components and audio processing unit.
[0020] In the diagram: 1-External microphone, 2-Noise suppression unit. Detailed Implementation
[0021] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0022] The specific implementation of the present invention will be described in detail below with reference to specific embodiments.
[0023] Please see Figures 1-4 The present invention provides an AI intelligent ambient sound synchronization system and method for over-ear wearable devices. Through intelligent audio processing technology, it achieves a balance between immersive experience and environmental safety perception for over-ear wearable devices, and can flexibly adapt to different types of over-ear devices. The following three typical embodiments are described in detail.
[0024] Example 1: Application in smart noise-canceling headphones This embodiment integrates the intelligent ambient sound synchronization system of the present invention into a pair of over-ear active noise-canceling headphones, suitable for urban commuting and daily travel scenarios. The specific configuration and implementation process are as follows: I. Specific Configuration of System Components 1. External Microphone 1: A dual-microphone array is set on each of the left and right sides of the earphone shell. It adopts an omnidirectional microphone, and its core parameters are as follows: The sensitivity range is -39dB to -37dB (@1kHz ref 1V / Pa), the operating voltage is 3.6V, the frequency response is referenced to the frequency response curve corresponding to the sensitivity @1kHz, the signal-to-noise ratio is 66dB (A-weighted), and the total harmonic distortion (THD) is ≤0.1% (Nom, Rload>2kΩ).
[0025] This dual-microphone array supports beamforming technology, enabling directional sound pickup, improving the targeting of ambient sound acquisition, and reducing interference from noise in irrelevant directions.
[0026] 2. Audio processing module: Integrated into the headphone main control board, its core components include: Noise Suppression Unit 2: Employs spectral subtraction (a noise suppression algorithm that estimates the noise spectrum and subtracts it from the original signal spectrum to suppress persistent broadband steady-state noise such as traffic background noise and wind noise in urban commuting scenarios. It does not require additional motion sensors and directly constructs a noise model through signal feature analysis, ensuring real-time noise reduction.
[0027] Sound event detection and enhancement unit: The feature extraction subunit extracts the Mel-frequency cepstral coefficients (MFCC, a commonly used acoustic feature in audio signal feature extraction) of the audio signal; the classification and recognition subunit adopts a lightweight neural network model, which has been trained with no less than 2000 hours of urban traffic scene sound data, specifically for recognizing two types of high-priority sound events: "car horn" and "emergency vehicle alarm"; the selective enhancement subunit performs gain enhancement (preset enhancement amplitude is 10dB) and frequency band sharpening on the identified target sound events, focusing on the key frequency band of 2kHz to 4kHz, to further improve the recognizability of the sound.
[0028] 3. Control Unit: Employs a low-power microprocessor to coordinate the working sequence of the microphone array, audio processing module, and mixing module; stores preset sound event categories and response strategies (such as volume adjustment rules when high-priority sound events are triggered); and allows users to manually switch working modes via buttons on the side of the headphones.
[0029] 4. Mixing Module: Connects to the built-in audio source of the headphones (such as music from a Bluetooth-connected mobile phone) and the output of the audio processing module. Users can set the mixing ratio of ambient sound and built-in audio through the control unit (adjustable range is 1:9 to 9:1). At the same time, it presets high-priority sound event response rules: when a car horn or emergency vehicle alarm is detected, the built-in audio volume is automatically reduced by 50% to ensure that ambient sound can be clearly transmitted.
[0030] 5. Built-in speaker: It adopts a high-fidelity dynamic speaker with a frequency response range of 100Hz to 8kHz, which can adapt to the playback needs of processed ambient sound and built-in audio, ensuring the clarity and naturalness of the sound output.
[0031] II. Specific Process of Method Implementation Step S1: When a user is walking on the street wearing headphones and listening to music, this system is activated. The dual microphone arrays on the left and right sides collect ambient sound signals in real time. The directional sound pickup function reduces the interference of noise reflected from buildings behind, ensuring that the collected signals are mainly traffic-related sounds from the front and sides.
[0032] Step S2: The noise suppression unit 2 processes the acquired raw signal through spectral subtraction to quickly attenuate continuous traffic background noise (such as tire noise and engine noise of vehicles) and wind noise, filter out low-frequency useless noise below 100Hz, and retain the effective audio frequency range of 100Hz to 8kHz to avoid noise drowning out key information.
[0033] Step S3: The sound event detection and enhancement unit performs real-time analysis on the noise-reduced signal and quickly matches sound features through a lightweight neural network model. When an ambulance approaches from behind, the features of its alarm sound are successfully matched with the preset "emergency vehicle alarm sound" features in the model, thus completing event recognition.
[0034] Step S4: The selective enhancement subunit immediately boosts the alarm sound signal by 10dB and sharpens the frequency band from 2kHz to 4kHz, making the alarm sound stand out more in the mixed audio.
[0035] Step S5: The mixing module automatically reduces the music volume by 50% according to preset rules, and mixes the enhanced alarm sound with the reduced music volume to ensure that the two sound signals do not interfere with each other.
[0036] Step S6: The mixed audio signal is output through the built-in speaker, allowing users to clearly hear the alarm sound, be alert in time, and take evasive action. At the same time, there is no need to completely turn off the music, achieving a balance between immersive music experience and traffic environment safety perception, significantly improving the safety and user experience in urban commuting scenarios.
[0037] Example 2: Application of industrial safety earmuffs This embodiment integrates the system into protective earmuffs for industrial workshops, adapting to scenarios such as factory production and equipment operation. It focuses on solving the problem of identifying and transmitting critical sounds (equipment malfunction sounds, command sounds, and alarm sounds) under strong noise conditions in industrial environments. The specific configuration and implementation process are as follows: I. Specific Configuration of System Components 1. External microphone 1: A dual microphone array is set at the front of the earcup shell. The microphone parameters are the same as in Example 1. It adopts an omnidirectional design to ensure that it can comprehensively collect sound signals from all directions in the workshop and adapt to the complex sound source distribution scenario in the industrial workshop.
[0038] 2. Audio Processing Module: Core configuration optimized for high-noise industrial environments. Noise Suppression Unit 2: Employs an adaptive filter. The reference signal originates from a pre-set factory machine noise model database (containing steady-state noise characteristics of common machine tools and fans). Acquisition method: The module is placed in the field for intelligent self-learning, acquiring noise models from the normal working environment and extracting key features for storage within the system. During operation, the system monitors the noise models in the scene in real time. If sound features outside the noise model appear, the extra-feature noise is processed and simultaneously played back at the speaker. It also integrates a VU-0005 high-pass filter with an attenuation slope of 24dB / octave, possessing phase linear Bessel characteristics. Through multi-stage circuitry, it achieves an actual attenuation slope of 12dB / octave, specifically filtering out low-frequency noise below 100Hz (such as equipment vibration noise and workshop floor resonance noise), preventing interference from low-frequency noise on effective signals such as human voices and abnormal equipment sounds. In addition, the unit also uses the UVR5AI model, which has been trained for no less than 2,000 hours of noise composition analysis in different industrial scenarios. It can intelligently identify the noise source audio track and accurately separate noise from valuable sounds (human voices and abnormal equipment sounds), further improving the noise reduction effect.
[0039] Sound event detection and enhancement unit: The feature extraction subunit extracts Mel-frequency cepstral coefficients (MFCC) and temporal peak features of the sound; the classification and recognition subunit uses support vector machine (SVM, a commonly used machine learning classification model in this field) to pre-train and recognize three types of valuable sound events: "abnormal machine friction sound", "voice command from the safety supervisor" and "plant-wide emergency alarm". Among them, the voice command recognition supports accurate recognition of preset keywords such as "stop" and "attention"; the selective enhancement subunit performs gain enhancement (15dB) and frequency band optimization (enhancing the 100Hz-800Hz frequency band of human voice and the 1kHz-5kHz frequency band of abnormal equipment sound) on the recognized valuable sounds.
[0040] 3. Control Unit: It adopts an industrial-grade low-power processor with anti-electromagnetic interference capability. In addition to coordinating the work of each module, it is also connected to the LED indicator on the outside of the earcup. When a valuable sound event is detected, the LED indicator flashes synchronously (red corresponds to an emergency alarm, yellow corresponds to an abnormal sound of the device, and green corresponds to a voice command), realizing a dual reminder of hearing and vision.
[0041] 4. Mixing module: Allows users to adjust the mixing ratio of ambient sound and built-in audio (such as work instruction broadcasts and audiobooks) via the knob on the earcups. The preset rules are: when the keywords "plant-wide emergency alarm" or "stop" are detected, the built-in audio is immediately muted and only the processed ambient sound is played; when "abnormal machine friction sound" is detected, the volume of the built-in audio is reduced by 70%.
[0042] 5. Built-in speaker: The speaker features a sweat-proof and dust-proof design with a frequency response range of 80Hz to 10kHz, making it suitable for industrial environments and ensuring clear and stable sound output.
[0043] II. Specific Process of Method Implementation Step S1: When workers are working in the workshop wearing protective earmuffs, the external microphone array 1 continuously collects sound signals in the workshop, including complex sounds such as machine noise, equipment operation noise and people talking.
[0044] Step S2: The noise suppression unit 2 uses an adaptive filter to match the factory machine noise model, suppressing continuous machine noise. At the same time, the VU-0005 high-pass filter filters out low-frequency vibration noise below 100Hz, and the UVR5AI model further separates noise from effective signals, greatly reducing interference from useless noise and ensuring the accuracy of subsequent sound recognition.
[0045] Step S3: The sound event detection and enhancement unit performs feature extraction and classification on the noise-reduced signal. When a nearby device emits an abnormal high-frequency friction sound, its time-domain peak feature and 1kHz~5kHz frequency band feature match the preset feature of "abnormal machine friction sound" to complete the event recognition. When the safety supervisor issues a "stop" command, the keyword feature is accurately recognized.
[0046] Step S4: The selective enhancement subunit performs a 15dB gain boost and corresponding frequency band optimization on the abnormal friction sound and the "stop" command sound respectively to ensure that the sound is clearly distinguishable.
[0047] Step S5: When an abnormal friction sound is detected, the mixing module reduces the built-in audio volume by 70% and mixes the enhanced abnormal friction sound; when a "stop" command is detected, the mixing module immediately mutes the built-in audio and only outputs the command sound.
[0048] Step S6: The mixed audio signal is played through the built-in speaker, and the LED indicator flashes accordingly. Workers can quickly detect equipment abnormalities and troubleshoot faults, or respond to the "stop" command in time to avoid production safety accidents.
[0049] This solution addresses the shortcomings of traditional industrial protective earmuffs that isolate all sound. While effectively protecting against strong industrial noise, it ensures the accurate transmission of critical safety information and work instructions, thereby improving the safety and practicality of industrial production.
[0050] Example 3: Application in motorcycle full-face helmets This embodiment integrates the system into a motorcycle full-face helmet, adapting to riding scenarios and focusing on solving problems such as wind noise suppression and key traffic sound recognition during riding. The specific configuration and implementation process are as follows: I. Specific Configuration of System Components 1. External Microphone 1: A microphone array consisting of two microphones is installed at the chin of the helmet. The position is optimized by wind tunnel design to reduce the impact of direct wind pressure on microphone acquisition during riding and reduce direct interference from wind noise. The microphone parameters are consistent with the previous two embodiments, supporting beamforming technology to directionally acquire traffic sound signals (such as car horns and pedestrian shouts) from the front and sides.
[0051] 2. Audio Processing Module: The main control board is integrated into the helmet liner, and its core includes a dedicated DSP chip, optimized for cycling scenarios. Noise Suppression Unit 2 employs a two-step wind noise suppression strategy. The first step calculates the correlation between the two microphone signals to identify and initially suppress incoherent wind noise signals. The second step uses spectral subtraction to subtract wind noise components from the signal spectrum based on the established wind noise model (containing wind noise characteristics at different vehicle speeds), further reducing wind noise interference. Simultaneously, a bandpass filter with a center frequency of 2.5kHz is integrated to specifically enhance the recognition of critical traffic sounds such as car horns. Furthermore, this unit also uses the UVR5AI model, trained for at least 2000 hours of noise composition analysis at different vehicle speeds (0–120 km / h), enabling it to intelligently distinguish between wind noise, traffic noise, and valuable sounds.
[0052] Sound event detection and enhancement unit: The feature extraction subunit extracts Mel-frequency cepstral coefficients (MFCC) and modulation frequency features of the sound; the classification and recognition subunit uses a convolutional neural network (CNN, a commonly used deep learning classification model in this field) to pre-train and identify three valuable sound events: "car horn sound", "bicycle bell sound" and "pedestrian shouting sound"; the selective enhancement subunit enhances the gain of the identified target sound (the enhancement magnitude can be dynamically adjusted according to the vehicle speed, ranging from 5dB to 15dB).
[0053] 3. Control Unit: Employs a low-power ARM Cortex-M series processor connected to the inertial measurement unit (IMU) inside the helmet. The IMU acquires vehicle speed estimates, which are then used to dynamically adjust system parameters (such as automatic gain and wind noise model matching). Users can switch between three operating modes via buttons on the side of the helmet: "Full Entertainment Mode" (ambient sound off, only built-in audio plays), "Safety Priority Mode" (ambient sound gain increased by 5dB, mixing ratio of ambient sound:built-in audio = 8:2), and "Automatic Mode" (dynamically adjusts the mixing ratio and ambient sound gain based on vehicle speed).
[0054] 4. Mixing Module: Connects to the rider's phone via Bluetooth to acquire music and navigation audio. The default mixing ratio is ambient sound: built-in audio = 7:3. Preset rules are as follows: when a high-priority sound event is detected (such as a car horn at close range), the built-in audio volume is automatically reduced by 60%; when the vehicle speed exceeds 60km / h, the ambient sound gain is automatically increased by 3dB; when the vehicle speed is below 20km / h, the ambient sound gain is appropriately reduced to avoid excessive ambient sound affecting the built-in audio experience.
[0055] 5. Built-in speaker: Located next to the rider's ear, connected to a power amplifier, with a frequency response range covering 100Hz to 10kHz, ensuring that navigation voice, music and processed ambient sound can be played clearly and naturally.
[0056] II. Specific Process of Method Implementation Step S1: When the rider is wearing a full-face helmet, the microphone array at the chin collects ambient sounds in real time. The wind tunnel-optimized installation position effectively reduces wind pressure interference with the collection, and the beamforming technology of the dual microphone array improves the targeting of traffic-related sound collection.
[0057] Step S2: Noise suppression unit 2 first uses dual-microphone signal correlation analysis to initially suppress incoherent wind noise, and then uses spectral subtraction combined with a wind noise model to further remove wind noise components at different vehicle speeds. A bandpass filter enhances the car horn sound in the 2.5kHz frequency band, and the UVR5AI model separates noise from valuable sounds to ensure that the noise-reduced signal is dominated by key traffic sounds.
[0058] Step S3: The sound event detection and enhancement unit extracts and classifies the signal features to determine in real time whether there are valuable sound events such as car horns, bicycle bells, and pedestrian shouts.
[0059] Step S4: After identifying the target sound event, the selective enhancement subunit adjusts the gain amplitude according to the current vehicle speed (e.g., the gain is increased by 12dB when the vehicle speed is 80km / h) and optimizes the corresponding frequency band to improve sound clarity.
[0060] Step S5: The control unit obtains the vehicle speed through the IMU. When the vehicle speed exceeds 60km / h, the automatic gain control unit increases the ambient sound volume by 3dB. The mixing module mixes the processed ambient sound with the mobile phone music / navigation audio at the default ratio of 7:3. When a car horn is detected at close range, the built-in audio volume is immediately reduced by 60% to ensure that the horn sound is clearly transmitted.
[0061] Step S6: The mixed signal is driven by the power amplifier to play through the built-in speaker, allowing the rider to clearly perceive key sounds of the surrounding traffic environment without affecting the listening of music and navigation. Dynamic parameter adjustments at different vehicle speeds ensure the best experience throughout the journey.
[0062] All components are integrated inside the helmet, without compromising its integrity and safety, achieving an effective balance between cycling entertainment and traffic safety.
[0063] It should be noted that, in this invention, although the specification describes the embodiments, not every embodiment contains only one independent technical solution. This way of describing the specification is only for clarity. Those skilled in the art should regard the specification as a whole. The technical solutions in each embodiment can also be appropriately combined to form other embodiments that can be understood by those skilled in the art.
Claims
1. An AI-powered intelligent ambient sound synchronization system for over-ear wearable devices, characterized in that, It includes an external microphone, an audio processing module, a control unit, a mixing module, and a built-in speaker; The external microphone is used to collect ambient sound. The audio processing module includes a noise suppression unit and a sound event detection and enhancement unit. The noise suppression unit is used to filter out broadband steady-state noise in ambient sound, and the sound event detection and enhancement unit is used to identify valuable sound events in ambient sound and enhance them. The control unit is used to coordinate the working sequence of the external microphone, audio processing module and mixing module, and to store preset valuable sound event categories and response strategies. The mixing module is used to mix the processed ambient sound output by the audio processing module with the built-in audio as needed; The built-in speaker is used to play the mixed audio output from the mixing module.
2. The AI intelligent ambient sound synchronization system for over-ear wearable devices according to claim 1, characterized in that, The external microphone is a microphone array that supports beamforming technology and can achieve directional sound pickup.
3. The AI intelligent ambient sound synchronization system for over-ear wearable devices according to claim 1, characterized in that, The noise suppression unit employs at least one noise reduction algorithm, such as spectral subtraction or adaptive filtering, and can achieve precise noise suppression by combining a preset noise model database or signal correlation analysis.
4. The AI intelligent ambient sound synchronization system for over-ear wearable devices according to claim 1, characterized in that, The sound event detection and enhancement unit includes a feature extraction subunit, a classification and recognition subunit, and a selective enhancement subunit; The feature extraction subunit is used to extract at least one of the following: Mel frequency cepstral coefficients, time-domain peak features, and modulation frequency features of the audio signal; The classification and recognition subunit adopts any one of the following: lightweight neural network model, support vector machine, and convolutional neural network, and is trained with scene sound data of a preset duration. The selective enhancement subunit is used to boost the gain and optimize the frequency band of the identified valuable sound events.
5. The AI intelligent ambient sound synchronization system for over-ear wearable devices according to claim 1, characterized in that, The control unit supports manual switching of the working mode by the user and is connected to sensors or indicator lights. The sensors are used to acquire external parameters of the device's usage scenario, and the indicator lights are used to provide visual alerts when valuable sound events are detected.
6. The AI intelligent ambient sound synchronization system for over-ear wearable devices according to claim 1, characterized in that, The mixing module allows users to adjust the mixing ratio of processed ambient sound and built-in audio, and presets high-priority, valuable sound event response rules: When a high-priority, valuable sound event is detected, the built-in audio volume is automatically adjusted or the built-in audio is muted.
7. An AI intelligent ambient sound synchronization method for over-ear wearable devices, applied to the AI intelligent ambient sound synchronization system for over-ear wearable devices as described in any one of claims 1-6, characterized in that, Includes the following steps: S1: External microphone collects ambient sound signals in real time; S2: The audio processing module filters out broadband steady-state noise in the ambient sound through the noise suppression unit, and then identifies valuable sound events and performs enhancement processing through the sound event detection and enhancement unit to obtain the processed ambient sound. S3: The control unit coordinates the mixing module to mix the processed ambient sound with the built-in audio according to the preset response strategy; S4: Built-in speaker plays the mixed audio signal.
8. The AI intelligent ambient sound synchronization method for over-ear wearable devices according to claim 7, characterized in that, In step S2, the valuable sound events include at least one of traffic warning sounds, equipment malfunction sounds, voice commands, and emergency alarm sounds; The classification and recognition subunit of the sound event detection and enhancement unit trains a model based on scene sound data of a preset duration to achieve accurate recognition of the valuable sound events.
9. The AI intelligent ambient sound synchronization method for over-ear wearable devices according to claim 7, characterized in that, In step S3, the mixing ratio of the mixing module can be customized by the user, and when a high-priority valuable sound event is detected, a preset response rule is automatically executed to reduce the built-in audio volume or mute the built-in audio.
10. The AI intelligent ambient sound synchronization method for over-ear wearable devices according to claim 7, characterized in that, In step S3, the control unit acquires external parameters through connected sensors and dynamically adjusts the gain amplitude and mixing ratio of the processed ambient sound. The external parameters include at least one of vehicle speed and usage scenario type.