Listening assistance signal enhancement method and device, intelligent glasses and storage medium

By extracting and enhancing speech signals using a microphone array and generating a cancellation signal to synthesize an auxiliary hearing signal, the problem of auditory confusion in open sound fields is solved, achieving a clear speech enhancement effect in open sound fields.

CN121838787APending Publication Date: 2026-04-10GOERTEK INC
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-23
Publication Date
2026-04-10

AI Technical Summary

Technical Problem

In an open-field soundscape shaped like glasses, the wearer simultaneously hears the original sound transmitted directly through the air and the enhanced sound processed and played by the device, resulting in auditory confusion and a poor experience.

Method used

The speech signal is extracted and enhanced by a microphone array, a cancellation signal is generated, and an enhanced auxiliary hearing signal is synthesized to suppress the direct transmission of speech through the air, thereby enhancing the human ear's hearing ability.

Benefits of technology

In an open sound field, the target speech can be heard clearly, background noise and airborne sound interference can be suppressed, and the listening experience can be improved.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121838787A_ABST
    Figure CN121838787A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of voice enhancement, and discloses a hearing aid signal enhancement method and device, intelligent glasses and a storage medium, and the method comprises the steps: extracting a voice signal from a sound signal of a current environment, and carrying out the enhancement of the voice signal, and obtaining an enhanced signal; generating an offset signal corresponding to the voice signal; and synthesizing the enhanced signal and the offset signal to obtain an enhanced hearing-aiding signal, and outputting the enhanced hearing-aiding signal. According to the method, the environment sound is collected, the voice signal of the dialogue person is extracted and enhanced in a targeted manner, and the offset signal is generated to suppress the dialogue voice which is directly propagated in the air, so that the offset effect of the dialogue voice which is propagated in the air can be formed in the ear area of the user, and an auxiliary hearing enhancement function is realized under the condition of an open sound field.
Need to check novelty before this filing date? Find Prior Art

Description

TECHNICAL FIELD

[0001] The present application relates to the technical field of voice enhancement, in particular to an auxiliary hearing signal enhancement method and device, smart glasses and a storage medium. BACKGROUND

[0002] In an open sound field in the form of glasses, using a high-performance microphone to pick up the voice of a person speaking on the other side and amplify it and then play it from the loudspeaker of the glasses will simultaneously hear the original voice directly transmitted by the air and the enhanced voice processed and played by the device, resulting in auditory confusion and poor experience. SUMMARY

[0003] The main purpose of the present application is to provide an auxiliary hearing signal enhancement method, device, smart glasses and storage medium, aiming to solve the technical problem that in an open sound field in the form of glasses, the original voice directly transmitted by the air and the enhanced voice processed and played by the device will be heard simultaneously.

[0004] To achieve the above purpose, the present application provides an auxiliary hearing signal enhancement method, which comprises: extracting a voice signal from a sound signal of the current environment, and enhancing the voice signal to obtain an enhanced signal; generating a cancellation signal corresponding to the voice signal; synthesizing the enhanced signal and the cancellation signal to obtain an enhanced auxiliary hearing signal, and outputting the enhanced auxiliary hearing signal.

[0005] Optionally, the step of extracting a voice signal from a sound signal of the current environment comprises: determining filter coefficients of adaptive beamforming corresponding to a sound pickup unit based on a target direction, the target direction being a gaze direction of a user; filtering the sound signal based on the filter coefficients to obtain a receiving beam pointing to the target direction; determining a voice signal of the sound signal in the target direction according to the receiving beam.

[0006] Optionally, before the step of determining filter coefficients of adaptive beamforming corresponding to a sound pickup unit based on a target direction, the method further comprises: collecting an environmental image and obtaining eye movement data of the user; determining a current gaze point of the user based on the eye movement data and the environmental image; setting the direction corresponding to the current gaze point as the target direction.

[0007] Optionally, the step of generating a cancellation signal corresponding to the voice signal comprises: generating a voice cancellation signal corresponding to the voice signal; Generate a noise cancellation signal corresponding to the sound signal; A cancellation signal is generated based on the speech cancellation signal and the noise cancellation signal.

[0008] Optionally, the step of generating a cancellation signal corresponding to the speech signal includes: Predict the transfer function of the speech signal as it reaches the user's ear region based on the acoustic model; The speech signal is processed based on the transfer function to obtain the predicted original sound components; A cancellation signal with the opposite phase to the original sound component is generated.

[0009] Optionally, after the step of generating a cancellation signal with the opposite phase to the original sound component, the method further includes: The actual speech signal generated by the speech output unit in the user's ear region is acquired. The actual speech signal is compared with the sound signal to obtain the residual error signal; The transfer function is updated based on the residual error signal.

[0010] Optionally, the step of synthesizing the enhanced signal and the canceled signal to obtain the enhanced auxiliary hearing signal, and outputting the enhanced auxiliary hearing signal, includes: The enhanced signal and the canceled signal are combined to generate an enhanced hearing aid signal; When the voice output unit outputs the enhanced hearing aid signal, the target signal of the current environment is collected through the sound pickup unit; Determine the azimuth information and first sharpness index corresponding to the target signal; If the angular deviation between the target direction and the azimuth information is greater than a preset deviation threshold, and the first clarity index is greater than the second clarity index of the sound signal, then the target direction is updated based on the azimuth information, and an enhanced hearing aid signal corresponding to the target signal is generated according to the updated target direction.

[0011] Furthermore, to achieve the above objectives, this application also proposes an auxiliary hearing signal enhancement device, which includes: The signal enhancement module is used to extract speech signals from the sound signals of the current environment and enhance the speech signals to obtain an enhanced signal; A signal cancellation module is used to generate a cancellation signal corresponding to the speech signal; The signal output module is used to synthesize the enhanced signal and the canceled signal to obtain the enhanced hearing aid signal, and output the enhanced hearing aid signal.

[0012] In addition, to achieve the above objectives, this application also proposes a smart glasses, the smart glasses comprising: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the hearing enhancement method described above.

[0013] In addition, to achieve the above objectives, this application also proposes a storage medium, which is a computer-readable storage medium, on which a computer program is stored, and when the computer program is executed by a processor, it implements the steps of the hearing enhancement method described above.

[0014] This application discloses a method for extracting speech signals from ambient sound signals, enhancing the speech signals to obtain enhanced signals, generating cancellation signals corresponding to the speech signals, synthesizing the enhanced signals and the cancellation signals to obtain enhanced hearing aids, and outputting the enhanced hearing aids. By collecting ambient sound, selectively extracting and enhancing the speech signals of the speakers, and simultaneously generating cancellation signals to suppress the direct airborne propagation of the speech, a cancellation effect against airborne sound can be achieved in the user's ear region, ensuring the realization of hearing aid enhancement functionality under open sound field conditions. Attached Figure Description

[0015] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.

[0016] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, for those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating the first embodiment of the hearing enhancement method of this application; Figure 2 This is a flowchart illustrating the second embodiment of the hearing enhancement method of this application; Figure 3 This is a flowchart illustrating the third embodiment of the hearing enhancement method of this application; Figure 4 This is a flowchart illustrating the fourth embodiment of the hearing enhancement method of this application; Figure 5 This is a schematic diagram of the module structure of the hearing aid signal enhancement device according to an embodiment of this application; Figure 6This is a schematic diagram of the device structure of the hardware operating environment involved in the smart glasses in this application embodiment.

[0018] The purpose, features, and advantages of this application will be further explained in conjunction with the embodiments and with reference to the accompanying drawings. Detailed Implementation

[0019] It should be understood that the specific embodiments described herein are merely illustrative of the technical solutions of this application and are not intended to limit this application.

[0020] To better understand the technical solution of this application, a detailed description will be provided below in conjunction with the accompanying drawings and specific implementation methods.

[0021] Current in-ear headphones can make it difficult to hear the other person when communicating at a distance or in noisy environments. However, if a high-performance microphone is used to pick up the speaker's voice, amplify it, and play it through the headphones' speakers, the problem of not being able to hear the other person can be solved.

[0022] Applying this function directly to glasses would present certain problems. Glasses are designed for an open sound field, meaning that while the human ear may not be able to hear the exact content, it can still pick up snippets of conversation. Under these conditions, if the sound picked up by the microphone is played directly from the speaker, the wearer will hear two sounds: the sound that travels directly through the air to the ear and the sound that has been processed and played back by the microphone. There will also be a delay between these two sounds.

[0023] Therefore, this application provides a method for enhancing auxiliary hearing signals in the form of glasses. Under the open sound field conditions of glasses, the airborne sound of the speaker opposite is eliminated, while only the sound picked up and enhanced by the microphone is retained, thereby enhancing the hearing ability of the human ear.

[0024] It should be noted that the executing entity in this embodiment can be a computing service device with data processing, network communication, and program execution functions, such as a computer, or an electronic device capable of performing the above functions. The following description uses smart glasses as an example to illustrate this embodiment and the subsequent embodiments.

[0025] Based on this, embodiments of this application provide a method for enhancing auxiliary hearing signals, referring to... Figure 1 , Figure 1 This is a flowchart illustrating the first embodiment of the hearing enhancement method of this application.

[0026] In this embodiment, the auxiliary hearing signal enhancement method includes: Step S10: Extract speech signals from the sound signals of the current environment and enhance the speech signals to obtain enhanced signals.

[0027] It should be noted that the current ambient sound signal is a multi-channel raw audio signal collected by a microphone array symmetrically arranged on the smart glasses. This raw audio signal is then converted into a digital audio stream after analog-to-digital conversion and preprocessing. The signal contains a mixture of speech from different directions, ambient noise, and other acoustic components. The speech signal refers to the human voice signal propagating in the current environment along the target direction (the direction the smart glasses wearer is looking), containing the speaker's semantic information. To enhance the speech signal from a specific direction, it is necessary to separate and amplify the human voice from this mixed signal.

[0028] Understandably, before step S10 is executed, the smart glasses first continuously collect ambient sound through the microphone array and determine the spatial pointing parameters of beamforming based on the target direction (usually the wearer's current gaze direction). If the smart glasses are in hearing enhancement mode, the speech enhancement processing link can be activated, the beamforming algorithm module can be called, and the initial beam weights or filter coefficients corresponding to the current target direction can be loaded to enhance the speech signal in the sound signal.

[0029] Specifically, the multi-microphone signals are first time-domain aligned and framed. Then, spatial filtering is applied to each channel signal using predetermined beamforming filter coefficients to form a receiving beam with a main lobe pointing towards the target and side lobes suppressed. The beamout signal is the target speech signal initially separated from environmental reverberation and noise. Next, the speech signal undergoes further noise reduction, gain control, and sound quality optimization. For example, residual noise is suppressed through spectral subtraction, Wiener filtering, or deep learning noise reduction models, and the speech gain is dynamically adjusted based on the signal-to-noise ratio, ultimately outputting a clear and highly intelligible enhanced signal.

[0030] For example, in a noisy public setting, a user wearing smart glasses converses with other users. The smart glasses, using built-in sensors, determine that the user's gaze is 30 degrees to the right of directly in front. It then uses beamforming coefficients matched to this direction to filter the audio signal collected by a four-microphone array, extracting the speech of other users in that direction. After noise reduction and appropriate gain, the output is an enhanced signal, allowing the user to clearly hear what the other person is saying, while ambient noise and other people's conversations are significantly suppressed. This scenario is merely an example; in practice, it can be adaptively adjusted according to the acoustic environment and user needs.

[0031] Step S20: Generate a cancellation signal corresponding to the speech signal.

[0032] It should be noted that the cancellation signal is used to acoustically cancel the reverse sound wave signal of the spoken dialogue that travels directly through the air to the human ear. Its generation depends on the ANC (Active Noise Cancellation) processing link. The ANC link works in parallel with the speech enhancement link to ensure that only the enhanced target speech is output in the end.

[0033] Understandably, when the smart glasses enter the auxiliary hearing enhancement mode, the ANC link continuously receives the raw ambient sound signal from the microphone array. If the voice enhancement link has output an enhanced signal that meets the signal-to-noise ratio requirements, the ANC link will simultaneously initiate the generation process of the cancellation signal. Optionally, the smart glasses can also decide whether to enable or adjust the strength of the ANC function based on the ambient noise intensity or user settings.

[0034] In one example, to achieve noise suppression while avoiding excessive cancellation or interference with the target speech components, step S20 may include: Generate a speech cancellation signal corresponding to the speech signal; generate a noise cancellation signal corresponding to the sound signal; generate a cancellation signal based on the speech cancellation signal and the noise cancellation signal.

[0035] It should be noted that cancellation signals can also be used to cancel environmental noise that travels directly through the air to the ear. Speech cancellation signals are inverse sound waves generated from the speaker's speech components extracted from the environment, used to cancel the portion of that speech that travels directly through the air to the ear, thus eliminating sound ghosting. Noise cancellation signals are inverse sound waves generated from the background noise components in the environment other than the target speech, used to suppress direct interference from environmental noise.

[0036] Understandably, after smart glasses acquire the sound signals of the current environment through the microphone array, they simultaneously activate the voice enhancement processing link and the active noise cancellation (ANC) processing link. This means that while the voice enhancement link performs beamforming to extract and enhance the target voice, the ANC link also begins to perform environmental acoustic analysis based on the same or multiple original sound signals and executes a path cancellation decision.

[0037] As an optional implementation, the smart glasses use the signal-to-noise ratio (SNR) of the sound signal or the residual intensity of the speaker's voice as a decision metric and compare it with a preset threshold. For example, if the calculated SNR of the current sound signal is 10dB, which is lower than the preset threshold of 15dB, it is determined that noise suppression is insufficient and a noise cancellation signal needs to be generated. Simultaneously, if the detected speech signal decibel level is high, it is determined that a speech cancellation signal needs to be generated. Subsequently, the ANC link generates two anti-phase sound waves according to the acoustic model and performs weighted synthesis in the frequency domain to form the final cancellation signal.

[0038] In another optional implementation, the smart glasses adaptively adjust the threshold for triggering split-path cancellation based on the complexity of the ambient sound (such as the number of sound sources and the characteristics of the noise spectrum). For example, in a relatively quiet meeting room scenario (low complexity), the smart glasses set a higher signal-to-noise ratio threshold, such as 20dB, and only activate noise cancellation when noise has a significant impact; while in a noisy street scenario (high complexity), the threshold is lowered (such as 8dB), and noise cancellation is activated more aggressively to ensure listening quality. Simultaneously, the activation of the voice cancellation signal is dynamically determined based on the stability of the target voice. If the voice direction remains stable, its cancellation gain is reduced to avoid overprocessing. For example, when a user is talking to a friend in a restaurant, the smart glasses first extract and enhance the friend's voice. At the same time, the ANC link analyzes the ambient sound. The friend's voice is relatively strong, so a precise voice cancellation signal is generated; the restaurant's background noise (such as the sound of cutlery and conversation) has a wide spectrum, so a wideband noise cancellation signal is generated. The two signals are combined and played through a speaker, effectively canceling their respective original sound components, allowing the user to hear only the clearly enhanced friend's voice without ghosting or background noise interference.

[0039] Step S30: Combine the enhanced signal and the canceled signal to obtain the enhanced hearing aid signal, and output the enhanced hearing aid signal.

[0040] It should be noted that the enhanced auxiliary hearing signal is the final output signal obtained by precisely synthesizing the enhanced speech signal from the target direction and the cancellation signal used to cancel out the original ambient noise in the time domain. Its purpose is to play a single, fused acoustic signal through the speakers on the temples of the smart glasses, so that the wearer can only perceive the clearly enhanced target speech in an open sound field, no longer disturbed by airborne sounds and background noise, thus achieving a hearing enhancement effect.

[0041] Understandably, the enhanced signal output from the voice enhancement link and the canceled signal output from the ANC link can be synchronously transmitted to the signal synthesis module after generation. This module first performs time synchronization alignment on the two signals to compensate for minor asynchrony that may be introduced by differences in the internal algorithm delays of the two links. After alignment, the two signals are superimposed in the time domain to generate a preliminary auxiliary hearing signal. To optimize the listening experience, the synthesis module can also perform dynamic gain control or amplitude limiting on the synthesized signal based on the energy or clarity of the enhanced signal to prevent output overload or distortion. Finally, the processed digital auxiliary hearing signal is converted into an analog signal by a digital-to-analog converter and then driven by a power amplifier to play synchronously on the speakers on both sides of the smart glasses.

[0042] In this embodiment, ambient sound is accurately collected through a microphone array, and speech in the wearer's gaze direction is extracted and enhanced. At the same time, a cancellation signal is generated to suppress conversational speech, which can create a cancellation effect on airborne sound in the human ear area, ensuring the hearing enhancement function is achieved under open sound field conditions.

[0043] Reference Figure 2 , Figure 2 This is a flowchart illustrating the second embodiment of the hearing enhancement method of this application. Based on the first embodiment described above, a second embodiment of the hearing enhancement method of this application is proposed. In the second embodiment, step S10 includes: Step S101: Based on the target direction, determine the filter coefficients of the adaptive beamforming corresponding to the pickup unit, wherein the target direction is the user's gaze direction.

[0044] It should be noted that the target direction refers to the convergence direction of the wearer's gaze, i.e., the spatial location currently being focused on. The sound pickup unit specifically refers to the microphone array deployed on the smart glasses. Adaptive beamforming is a signal processing technique that adjusts the weights (i.e., filter coefficients) of the signals received by each microphone in the array in real time, thereby forming a directional receiving beam in the acoustic space. This generates high gain in the target direction to enhance the sound coming from that direction while suppressing interference noise from other directions.

[0045] It should be understood that the sound pickup unit can be an array of an even number of microphones symmetrically distributed, and at least one set of the microphones is arranged on the left and right temples of the smart glasses, with each set of microphones arranged along the front-back direction of the smart glasses.

[0046] Understandably, based on the precise three-dimensional position of each microphone in the microphone array and an estimation of the current acoustic environment, an array manifold matrix is ​​established. This matrix describes the relative delay or phase difference when plane waves from different spatial directions arrive at each microphone in the array. Then, the calculated target direction vector is substituted into the matrix to calculate the corresponding steering vector. The adaptive beamforming algorithm uses this steering vector as a constraint to minimize the total output power of the array while ensuring the target direction signal passes through without distortion. This suppresses interference from all non-target directions, thus solving for a set of optimal complex weights, i.e., filter coefficients. The filter coefficients determine how the amplitude and phase of the signal from each microphone are adjusted and then summed to form a beam pointing towards the target direction.

[0047] It's important to understand that the target's direction may change its absolute position in the global coordinate system due to the user's head rotation. Therefore, in practical implementations, smart glasses typically use the head coordinate system as a reference and integrate inertial measurement unit data to stably map the gaze direction to a coordinate system with the head as the origin, thereby ensuring the stability and intuitiveness of the beam pointing.

[0048] It should be understood that, to balance performance and computational complexity, smart glasses can pre-set multiple beam coefficient codebooks for typical directions. Once the target direction is determined, a set of approximate filter coefficients is quickly obtained as initial values ​​through table lookup or interpolation, and then fine-tuned. For example, the codebook stores nominal coefficients for directions such as directly in front, 30 degrees to the left / right, and 60 degrees to the left / right. When the smart glasses determine that the user's gaze direction is approximately 25 degrees to the right front, they take the coefficients from the directly in front and the 30-degree to the right and perform weighted interpolation to obtain the initial coefficients suitable for the 25-degree direction. Then, they perform rapid adaptive convergence based on a small amount of real-time collected data to cope with subtle changes in individual head movements or environmental reflections.

[0049] In one example, a user in a meeting room looks at colleague A who is speaking. The eye-tracking module determines the gaze direction to be 35 degrees to the right front. The smart glasses immediately use this direction as the target direction and, based on a four-microphone array model, calculate a new set of filter coefficients using a beamforming algorithm. After applying these coefficients, the main lobe of the microphone array's receiving beam is precisely aligned to the 35-degree direction, significantly enhancing colleague A's speech, while interference such as air conditioning noise from the left window and whispers from others on the right are suppressed by the beam's sidelobe nulls or low-gain regions. These coefficients remain in effect until the user turns their gaze to another speaker (such as colleague B to the left front). The smart glasses then recalculate and update the filter coefficients based on the new target direction, achieving a smooth switch in auditory focus.

[0050] Furthermore, in order to directly reflect the wearer's gaze intention, the actual direction corresponding to the gaze point is located by combining the environmental image, making the determination of the target direction more in line with the user's communication needs, without the need for manual adjustment of direction parameters. Before step S101, the method further includes: Acquire environmental images and obtain the user's eye movement data; determine the user's current gaze point based on the eye movement data and the environmental images; set the direction corresponding to the current gaze point as the target direction.

[0051] It should be noted that the environmental image refers to the real-time scene captured by the front-facing camera of the smart glasses, including spatial information such as communication objects and environmental objects, providing a scene basis for gaze point positioning; eye movement data is a parameter reflecting the wearer's eye movement state collected by the eye movement sensor built into the smart glasses, including pupil position, gaze angle, eye movement trajectory, etc.; the current gaze point is determined based on the matching analysis of eye movement data and environmental image to determine the specific location where the wearer's gaze is focused in the actual scene.

[0052] Understandably, after a user issues an activation command for the hearing enhancement function via voice, gesture, or physical button, the smart glasses respond to the function call command and generate a corresponding processing task call request based on the main control computing module to initiate the gaze point analysis module. Simultaneously, the aligned environmental image frame and eye-tracking data are transmitted to this module for fusion analysis to determine the target direction reflecting the user's immediate communication intent.

[0053] Specifically, the eye-tracking data is first filtered and calibrated to remove noise introduced by blinking or brief gaze drift, and the stable gaze point coordinates in the current environmental image's two-dimensional coordinate system are calculated. Then, combining the camera's intrinsic parameters and posture information, this two-dimensional gaze point is mapped to a three-dimensional head space coordinate system with the smart glasses as the origin, generating a spatial direction vector pointing towards the user's point of focus.

[0054] Step S102: Filter the sound signal based on the filter coefficients to obtain a receiving beam pointing towards the target direction.

[0055] It should be noted that the receiving beam is a spatially directional synthesized signal formed by applying specific weights of filter coefficients to the original sound signals collected by each channel of the microphone array and then synthesizing them.

[0056] Specifically, the multi-microphone signals are synchronized and preprocessed, including framing, windowing, and time-frequency transformation, converting the signals to frequency domain sub-bands that are easier to process. Then, in each frequency sub-band, the calculated complex filter coefficients are spatially filtered with the spectral data of each microphone in the corresponding sub-band. For example, beamforming is used to calculate the inner product of the filter coefficient vector and the microphone spectral vector; the result is the beam output spectrum pointing towards the target in that frequency sub-band. Finally, the beam output spectra of all frequency sub-bands are inversely transformed to synthesize the received beam signal in the time domain. In this signal, the sound components from the target direction are enhanced by in-phase superposition, while sound components from other directions are significantly suppressed due to out-of-phase cancellation or low weighting.

[0057] For example, the user locks their gaze onto speaker A. The smart glasses process the mixed signal (including A's voice, air conditioning noise, other people's whispers, keyboard sounds, etc.) captured by the four microphones at that moment, frame by frame, and apply this set of coefficients to multiple frequency bands for weighted summation. In the final synthesized received beam signal, the speech component from A's direction (e.g., directly in front) is significantly enhanced, and its volume may be increased by more than 10dB; while the low-frequency noise from the side air conditioning vent and the conversations of other people behind are suppressed by more than 15dB.

[0058] Step S103: Determine the speech signal of the sound signal in the target direction based on the received beam.

[0059] Understandably, speech activity detection technology can distinguish the received beam into speech segments containing human voices and non-speech segments containing only environmental noise. Then, further noise reduction processing (such as spectral subtraction or Wiener filtering) is applied to the identified speech segments to filter out residual environmental noise in the same direction as the target speech, as well as a small amount of lateral interference that cannot be completely suppressed by the beam. Finally, a target-direction speech signal with a significantly improved signal-to-noise ratio is output.

[0060] In this embodiment, the optimal filter coefficients are dynamically calculated and applied according to the preset target direction to enhance the speech signal from the user's gaze direction, while suppressing interference noise and reverberation from other directions to the greatest extent.

[0061] Reference Figure 3 , Figure 3 This is a flowchart illustrating the third embodiment of the hearing enhancement method of this application. Based on the first embodiment described above, the third embodiment of the hearing enhancement method of this application is proposed. In the third embodiment, step S20 includes: Step S201: Predict the transfer function of the speech signal reaching the user's ear region based on the acoustic model.

[0062] It should be noted that the transfer function is a mathematical model describing the propagation of sound from the speaker of smart glasses to the wearer's ear canal. It quantifies the frequency response changes (such as the amplification or attenuation of certain frequencies) and phase shifts experienced by the sound signal along this propagation path. The acoustic model is a mathematical or data-driven model that is pre-stored or built in real time in smart glasses, representing a specific acoustic path, and depends on the physical structure of the smart glasses.

[0063] Understandably, precisely measuring the personalized transfer function for each pair of glasses and each user is costly and impractical. Therefore, this embodiment can use a parameterized model or adaptive model pre-trained based on typical wearing scenarios as a foundation. Since the transfer function will change slightly with the tightness and angle of the glasses, the model needs to be able to fine-tune during operation to achieve the best cancellation effect. However, this is usually only done as needed after the active noise cancellation function is enabled, in order to balance effect and power consumption.

[0064] Understandably, after the smart glasses invoke the acoustic model, they take the current speech signal to be canceled as input, calculate the spectral changes and time delays the speech signal should experience as it propagates from the glasses reference point (usually the speaker diaphragm position) to a preset human ear reference point, and output a complex-form prediction transfer function. The function can be represented as gain and phase values ​​at a series of frequency points.

[0065] For example, a user wears smart glasses and activates the ANC (Auxiliary Hearing Enhancement) function to converse quietly with a friend across from them. While generating the enhanced speech signal, the smart glasses also need to generate a cancellation signal to eliminate the portion of the friend's original speech directly reaching the user's ear. At this time, the ANC link calls the acoustic model, inputting the extracted spectral features of the friend's speech. Based on the current wearing condition of the glasses (e.g., determined to be a standard fit), the acoustic model predicts the transfer function of the speech signal propagating from the temple speaker to the user's ear canal entrance. This transfer function may have a slight attenuation in the main human voice frequency band of 1kHz-3kHz and introduce a fixed delay of approximately 0.1 milliseconds. This predicted transfer function will be used in the next step to generate an accurate inverse signal. The parameters mentioned above are for illustrative purposes only. Specific algorithms for predicting the transfer function using the acoustic model (such as boundary element method simulation, measured data fitting, neural network regression, etc.) are existing technologies in this field, and this application does not limit their specific implementation.

[0066] Step S202: Process the speech signal based on the transfer function to obtain the predicted original sound components.

[0067] Step S203: Generate a cancellation signal that is out of phase with the original sound component.

[0068] It's important to note that the predicted original sound component refers to the theoretical acoustic waveform of the original environmental sound signal (especially speech from the target direction), calculated by the smart glasses based on its transfer function, as it propagates through the air and eventually reaches the user's eardrum. The cancellation signal is a reverse sound wave generated to acoustically eliminate the predicted original sound component. It has the same amplitude as the predicted component but is completely opposite in phase. When these two signals are superimposed in space, according to the principle of destructive interference of sound waves, their combined sound pressure will tend to zero, thus achieving a silent effect.

[0069] Understandably, the smart glasses perform a forward propagation computation. It will predict the resulting transfer function (e.g., representing the complex response in the frequency domain). ) Applied to the speech signal to be canceled The calculation process is essentially frequency domain multiplication. .here, This is the predicted spectrum of the original sound component that will appear in the ear canal. This calculation takes into account all frequency-selective attenuation caused by the propagation path (e.g., high frequencies attenuate faster in air) and phase delay. Subsequently, for Performing an inverse Fourier transform yields the predicted original sound components in the time domain. .

[0070] Furthermore, the predicted original sound components Perform a phase reversal. In the time domain, this typically means multiplying the signal waveform by -1, i.e. The generated This constitutes the initial cancellation signal. However, in practical systems, to compensate for the additional distortion (which can be understood as another transfer function) introduced by the new path from generating the cancellation signal to its playback through the speaker and then propagating to the ear canal, smart glasses often... It then undergoes further processing via a pre-equalization filter. The design goal of this filter is to make... After being played through the speaker and propagated a second time, the sound component at the ear canal can exactly match the original sound component. To achieve equal amplitude and opposite phase.

[0071] Furthermore, in order to adaptively compensate for transfer function deviations caused by individual wearing differences, glasses slippage, or environmental changes, step S203 further includes: The actual speech signal generated by the speech output unit in the user's ear region is acquired; the actual speech signal is compared with the sound signal to obtain a residual error signal; and the transfer function is updated based on the residual error signal.

[0072] It should be noted that the voice output unit refers to the speaker on the smart glasses used to play voice. The actual voice signal refers to the acoustic signal collected by a microphone placed close to the user's ear canal, which includes the original ambient sound and the auxiliary hearing signal played by the speaker. The residual error signal is the difference between the expected silent state (ideal cancellation) of the smart glasses and the actual measured sound pressure level, reflecting the imperfection of the current cancellation effect, caused by factors such as environmental changes and model mismatch.

[0073] Specifically, while outputting a cancellation signal through a speaker, the smart glasses continuously collect actual acoustic signals near the user's ear through a microphone. After preprocessing, the actual acoustic signal is compared with a sound signal that has been appropriately delayed and aligned in the time or frequency domain. The goal of the comparison is to evaluate the remaining acoustic energy after actual cancellation; the difference is the residual error signal. This error signal is fed back to the transfer function estimation module. Then, using the error signal as a performance indicator, iterative calculations dynamically adjust the model parameters used to predict the transfer function, so that the model's predicted output (i.e., the predicted original sound component) gets closer and closer to reality. This allows the cancellation signal generated based on this model to more effectively minimize the energy of the error signal. Through this continuous adaptive learning, the smart glasses can gradually approximate and track the actual acoustic path characteristics of the current user and the current wearing state.

[0074] In this embodiment, the active noise cancellation control link can accurately estimate the effect of ambient sound propagating to the human ear and generate an inverse sound wave with the same amplitude but opposite phase to it for acoustic cancellation, thereby achieving effective suppression of ambient sound propagating directly to the human ear through the air.

[0075] Reference Figure 4 , Figure 4 This is a flowchart illustrating the fourth embodiment of the hearing enhancement method of this application. Based on the above embodiments, the fourth embodiment of the hearing enhancement method of this application is proposed. In the fourth embodiment, step S30 includes: Step S301: Combine the enhanced signal and the canceled signal to generate an enhanced hearing aid signal.

[0076] Step S302: When the voice output unit outputs the enhanced hearing aid signal, the target signal of the current environment is collected through the sound pickup unit.

[0077] It should be noted that when the enhanced hearing signal is played through the voice output unit (i.e., speaker) of the smart glasses, the smart glasses simultaneously collect the sound of the current environment again through the microphone array (pickup unit). The signal collected here is defined as the target signal.

[0078] Understandably, in open-back hearing aid systems, the user's acoustic environment, head posture, and the way the glasses are worn can all change in real time. If smart glasses only perform signal processing and output based on the initial state, they cannot cope with these dynamic changes, which may lead to a decrease in enhancement effect (such as the sound source deviating from the main lobe of the beam) or the introduction of new interference (such as insufficient cancellation causing echoes).

[0079] Step S303: Determine the azimuth information and the first sharpness index corresponding to the target signal.

[0080] It should be noted that the directional information reflects the direction of the dominant sound source in the igniter environment after being amplified and output by the smart glasses, such as the direction of the next speaker. The First Clarity Index is a quantitative metric used to measure the speech intelligibility of the target signal.

[0081] It should be understood that in dynamic environments, the user's gaze target may move, and environmental noise may change abruptly. Smart glasses must be able to perceive the current output effect. Orientation information is used to determine whether the direction of enhancement is correct or if the correct direction is needed in the next moment. If the orientation information deviates significantly from the user's intention (target direction), it indicates that the target sound source may have been deviated from, or the direction of the speech signal in the current environment may have changed. The first clarity index is used to determine whether the target signal is clear. If the clarity index is too low, it indicates insufficient noise suppression or that the speech enhancement algorithm needs adjustment. Conversely, a high clarity index may indicate that the speech information at this moment is more important.

[0082] Step S304: If the angular deviation between the target direction and the azimuth information is greater than a preset deviation threshold, and the first clarity index is greater than the second clarity index of the sound signal, then the target direction is updated based on the azimuth information, and an enhanced hearing aid signal corresponding to the target signal is generated according to the updated target direction.

[0083] It should be noted that the preset deviation threshold is a pre-set angular tolerance value (e.g., 15 degrees), used to determine whether the estimated sound source location has significantly deviated from the target direction currently locked by the system. The second clarity index is a clarity index calculated by analyzing the sound signal (i.e., the original ambient sound without enhancement) in the target direction, representing the intelligibility benchmark of the sound signal in that direction before the smart glasses' enhancement.

[0084] Understandably, if the angular deviation between the target direction and the azimuth information exceeds a preset deviation threshold, this indicates that the dominant sound source direction in the environment has significantly deviated from the focus of the system's current service. This could be due to speaker movement, failure to recalibrate after the user turns their head, or the appearance of a new, stronger sound source to focus on in the environment. When the first clarity index is greater than the second clarity index of the sound signal, the target is switched to ensure that the smart glasses can provide clearer speech from the new sound source direction than from the original direction.

[0085] For example, in a meeting, the user initially focuses on and listens to colleague A speaking, with the target direction being directly in front. At this moment, colleague B begins speaking from 30 degrees to the left and front. The smart glasses detect that the actual dominant sound source location has changed to 28 degrees to the left and front, a deviation of 28 degrees from the current target direction, exceeding the preset threshold of 15 degrees; and the speech clarity index (first clarity index) in the new direction is 80, while the clarity index of the original sound in the original direction (second clarity index) has dropped to 20 because A has stopped speaking. The smart glasses then update the target direction to 28 degrees to the left and front and immediately regenerate an enhanced post-auditory signal for colleague B's speech and output it to the user. Without any manual operation, the user's auditory focus automatically and smoothly switches from colleague A to colleague B, maintaining the continuity and clarity of the conversation.

[0086] In this embodiment, when the target direction shifts or a clearer voice signal appears, the system can automatically update the target direction and regenerate the auxiliary hearing signal, thus avoiding the problem of decreased voice clarity in the fixed output mode.

[0087] It should be noted that the above examples are only for understanding this application and do not constitute a limitation on the hearing enhancement method of this application. Any simple modifications based on this technical concept are within the protection scope of this application.

[0088] This application also provides a hearing aid signal enhancement device; please refer to [reference needed]. Figure 5 The hearing aid signal enhancement device includes: The signal enhancement module 10 is used to extract speech signals from the sound signals of the current environment and enhance the speech signals to obtain an enhanced signal; Signal cancellation module 20 is used to generate a cancellation signal corresponding to the speech signal; The signal output module 30 is used to synthesize the enhanced signal and the canceled signal to obtain the enhanced hearing aid signal, and output the enhanced hearing aid signal.

[0089] The hearing enhancement device provided in this application, employing the hearing enhancement method described in the above embodiments, can solve the technical problem that, in an open sound field resembling eyeglasses, one can simultaneously hear the original sound transmitted directly through the air and the enhanced sound processed and played by the device. Compared with the prior art, the beneficial effects of the hearing enhancement device provided in this application are the same as those of the hearing enhancement method provided in the above embodiments, and other technical features in the hearing enhancement device are the same as those disclosed in the methods of the above embodiments, and will not be repeated here.

[0090] This application provides a smart glasses, which includes: at least one processor; and a memory communicatively connected to the at least one processor; wherein the memory stores instructions executable by the at least one processor, which are executed by the at least one processor to enable the at least one processor to perform the hearing enhancement method in Embodiment 1 above.

[0091] The following is for reference. Figure 6 It shows a structural schematic diagram suitable for implementing smart glasses in the embodiments of this application. Figure 6 The smart glasses shown are merely an example and should not be construed as limiting the functionality and scope of use of the embodiments of this application.

[0092] like Figure 6 As shown, the smart glasses may include a processing device 1001 (e.g., a central processing unit, a graphics processing unit, etc.), which can perform various appropriate actions and processes according to a program stored in a read-only memory 1002 or a program loaded from a storage device 1003 into a random access memory 1004. The random access memory 1004 also stores various programs and data required for the operation of the smart glasses. The processing device 1001, the read-only memory 1002, and the random access memory 1004 are interconnected via a bus 1005. An input / output interface 1006 is also connected to the bus. Typically, the following systems can be connected to the input / output interface 1006: an input device 1007 including, for example, a touchscreen, touchpad, keyboard, mouse, image sensor, microphone, accelerometer, gyroscope, etc.; an output device 1008 including, for example, a liquid crystal display (LCD), speaker, vibrator, etc.; a storage device 1003 including, for example, magnetic tape, hard disk, etc.; and a communication device 1009. The communication device 1009 allows the smart glasses to communicate wirelessly or wiredly with other devices to exchange data. While the diagram shows smart glasses with various systems, it should be understood that implementing or having all of the systems shown is not required. More or fewer systems may be implemented alternatively.

[0093] Specifically, according to the embodiments disclosed in this application, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments disclosed in this application include a computer program product comprising a computer program carried on a computer-readable medium, the computer program containing program code for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via a communication device, or installed from storage device 1003, or installed from read-only memory 1002. When the computer program is executed by processing device 1001, it performs the functions defined in the methods of the embodiments disclosed in this application.

[0094] The smart glasses provided in this application, employing the auxiliary hearing signal enhancement method described in the above embodiments, can solve the technical problem of simultaneously hearing the original sound transmitted directly through the air and the enhanced sound processed and played by the device in an open sound field shaped like glasses. Compared with the prior art, the beneficial effects of the smart glasses provided in this application are the same as those of the auxiliary hearing signal enhancement method provided in the above embodiments, and other technical features of the smart glasses are the same as those disclosed in the previous embodiment method, and will not be repeated here.

[0095] It should be understood that the various parts disclosed in this application can be implemented using hardware, software, firmware, or a combination thereof. In the description of the above embodiments, specific features, structures, materials, or characteristics can be combined in any suitable manner in one or more embodiments or examples.

[0096] The above description is merely a specific embodiment of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

[0097] This application provides a computer-readable storage medium having computer-readable program instructions (i.e., a computer program) stored thereon, the computer-readable program instructions being used to perform the hearing enhancement method in the above embodiments.

[0098] The computer-readable storage medium provided in this application may be, for example, a USB flash drive, but is not limited to, electrical, magnetic, optical, electromagnetic, infrared, or semiconductor systems, devices, or any combination thereof. More specific examples of computer-readable storage media may include, but are not limited to: electrical connections having one or more wires, portable computer disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fibers, portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination thereof. In this embodiment, the computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, system, or device. The program code contained on the computer-readable storage medium may be transmitted using any suitable medium, including but not limited to: wires, optical cables, RF (Radio Frequency), etc., or any suitable combination thereof.

[0099] The aforementioned computer-readable storage medium may be included in the smart glasses; or it may exist independently and not assembled into the smart glasses.

[0100] The aforementioned computer-readable storage medium carries one or more programs, which, when executed by the smart glasses, cause the smart glasses to perform the hearing enhancement method described above.

[0101] Computer program code for performing the operations of this application can be written in one or more programming languages ​​or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, and C++, and conventional procedural programming languages ​​such as the "C" language or similar programming languages. The program code can be executed entirely on the user's computer, partially on the user's computer, as a standalone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving remote computers, the remote computer can be connected to the user's computer via any type of network—including a Local Area Network (LAN) or a Wide Area Network (WAN)—or can be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0102] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of this application. In this regard, each block in a flowchart or block diagram may represent a module, segment, or portion of code containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer instructions.

[0103] The modules described in the embodiments of this application can be implemented in software or hardware. The names of the modules do not necessarily limit the functionality of the unit itself.

[0104] The readable storage medium provided in this application is a computer-readable storage medium that stores computer-readable program instructions (i.e., a computer program) for executing the above-described auxiliary hearing signal enhancement method. This solves the technical problem that, in an open sound field resembling eyeglasses, both the original sound transmitted directly through the air and the enhanced sound processed and played by the device are heard simultaneously. Compared to the prior art, the beneficial effects of the computer-readable storage medium provided in this application are the same as those of the auxiliary hearing signal enhancement method provided in the above embodiments, and will not be repeated here.

[0105] The above description is only a part of the embodiments of this application and does not limit the scope of this application. All equivalent structural transformations made under the technical concept of this application and using the content of this application specification and drawings, or direct / indirect applications in other related technical fields, are included within the protection scope of this application.

Claims

1. A method for enhancing auxiliary hearing signals, characterized in that, Applied to smart glasses, the method includes: Extract speech signals from the sound signals of the current environment, and enhance the speech signals to obtain an enhanced signal; Generate a cancellation signal corresponding to the speech signal; The enhanced signal and the canceled signal are synthesized to obtain an enhanced hearing aid signal, and the enhanced hearing aid signal is output.

2. The method as described in claim 1, characterized in that, The step of extracting speech signals from the sound signals of the current environment includes: Based on the target direction, the filter coefficients of the adaptive beamforming corresponding to the pickup unit are determined, wherein the target direction is the user's gaze direction; The sound signal is filtered based on the filter coefficients to obtain a receiving beam pointing towards the target direction; The speech signal of the sound signal in the target direction is determined based on the received beam.

3. The method as described in claim 2, characterized in that, Before the step of determining the adaptive beamforming filter coefficients corresponding to the pickup unit based on the target direction, the method further includes: Acquire environmental images and obtain the user's eye movement data; Based on the eye-tracking data and the environmental image, the user's current gaze point is determined; Set the direction corresponding to the current gaze point as the target direction.

4. The method according to any one of claims 1 to 3, characterized in that, The step of generating a cancellation signal corresponding to the speech signal includes: Generate a speech cancellation signal corresponding to the speech signal; Generate a noise cancellation signal corresponding to the sound signal; A cancellation signal is generated based on the speech cancellation signal and the noise cancellation signal.

5. The method according to any one of claims 1 to 3, characterized in that, The step of generating a cancellation signal corresponding to the speech signal includes: Predict the transfer function of the speech signal as it reaches the user's ear region based on the acoustic model; The speech signal is processed based on the transfer function to obtain the predicted original sound components; A cancellation signal with the opposite phase to the original sound component is generated.

6. The method as described in claim 5, characterized in that, Following the step of generating a cancellation signal that is out of phase with the original sound component, the method further includes: The actual speech signal generated by the speech output unit in the user's ear region is acquired. The actual speech signal is compared with the sound signal to obtain the residual error signal; The transfer function is updated based on the residual error signal.

7. The method according to any one of claims 1 to 3, characterized in that, The step of synthesizing the enhanced signal and the canceled signal to obtain the enhanced auxiliary hearing signal, and outputting the enhanced auxiliary hearing signal, includes: The enhanced signal and the canceled signal are combined to generate an enhanced hearing aid signal; When the voice output unit outputs the enhanced hearing aid signal, the target signal of the current environment is collected through the sound pickup unit; Determine the azimuth information and first sharpness index corresponding to the target signal; If the angular deviation between the target direction and the azimuth information is greater than a preset deviation threshold, and the first clarity index is greater than the second clarity index of the sound signal, then the target direction is updated based on the azimuth information, and an enhanced hearing aid signal corresponding to the target signal is generated according to the updated target direction.

8. A hearing aid signal enhancement device, characterized in that, The device includes: The signal enhancement module is used to extract speech signals from the sound signals of the current environment and enhance the speech signals to obtain an enhanced signal; A signal cancellation module is used to generate a cancellation signal corresponding to the speech signal; The signal output module is used to synthesize the enhanced signal and the canceled signal to obtain the enhanced hearing aid signal, and output the enhanced hearing aid signal.

9. A type of smart glasses, characterized in that, The smart glasses include: a memory, a processor, and a computer program stored in the memory and executable on the processor, the computer program being configured to implement the steps of the hearing enhancement method as described in any one of claims 1 to 7.

10. A storage medium, characterized in that, The storage medium is a computer-readable storage medium, and a computer program is stored on the storage medium. When the computer program is executed by a processor, it implements the steps of the hearing enhancement method as described in any one of claims 1 to 7.