Complex sound scene target voice selectivity enhancement method and system based on eye movement guidance

By integrating a miniature eye-tracking module with a ring microphone array, and combining eye movement intentions with acoustic signals, the audio enhancement strategy is dynamically adjusted. This solves the problem that traditional hearing aids cannot dynamically respond to the objects of attention of hearing-impaired children. It achieves high-precision sound source localization and speech intelligibility improvement in complex acoustic environments, thereby enhancing the social participation ability of hearing-impaired children.

CN121963736AInactive Publication Date: 2026-05-01NANJING TECHN COLLEGE OF SPECIAL EDUCATION
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
NANJING TECHN COLLEGE OF SPECIAL EDUCATION
Filing Date
2026-02-03
Publication Date
2026-05-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Traditional hearing aids cannot dynamically respond to changes in the instantaneous focus of hearing-impaired children. Existing multimodal assistive technologies lack robustness in complex acoustic environments and lack human-machine collaborative feedback mechanisms, resulting in limited speech intelligibility and social participation abilities for hearing-impaired children in real-world scenarios.

Method used

By integrating a miniature eye-tracking module with a ring microphone array, combining eye movement intentions and acoustic signals, dynamically adjusting audio enhancement strategies, and introducing visual confidence scores, speech intelligibility assessments, and tactile feedback mechanisms, a low-latency human-machine collaborative enhancement system is formed.

Benefits of technology

It achieves high-precision sound source localization in complex acoustic environments, dynamically adjusts enhancement strategies, improves speech intelligibility and social participation in hearing-impaired children, and meets the ergonomic wearing needs of children.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121963736A_ABST
    Figure CN121963736A_ABST
Patent Text Reader

Abstract

The invention discloses a complex sound scene target voice selectivity enhancement method and system based on eye movement guidance, and the method comprises the steps: synchronously collecting a gazing direction and a multi-channel audio signal through a miniature eye movement tracking module and an annular quaternary microphone array, and constructing a dynamic sound source positioning mechanism based on eye movement guidance; based on visual confidence score adaptive switching or fusion beam forming and deep learning enhancement strategies, voice intelligibility evaluation and bone conduction tactile feedback closed loops are introduced, and children are actively guided to adjust gazing behaviors. The system comprises an eye movement tracking module, a sound source positioning module, an adaptive enhancement module, a feedback prompt module and the like, and is integrally light, low in delay and high in comfort. According to the invention, through multi-mode fusion and closed-loop interaction, the selective perception ability and communication efficiency of hearing-impaired children to the voice of the target speaker in a complex sound scene are improved.
Need to check novelty before this filing date? Find Prior Art

Description

A method and system for selectively enhancing target speech in complex soundscapes based on eye-tracking guidance Technical Field

[0001] This invention relates to the field of artificial intelligence and intelligent human-computer interaction technology, specifically to a method and system for selectively enhancing speech in complex soundscapes based on eye-tracking guidance. Background Technology

[0002] The rapid development of artificial intelligence and intelligent human-computer interaction technologies has provided new possibilities for improving the speech perception ability of hearing-impaired individuals in complex acoustic environments. Hearing-impaired children often face difficulties in effectively extracting target speech from multiple speakers simultaneously in multi-person interaction scenarios such as classrooms and streets. Traditional hearing aids mainly rely on fixed beamforming or noise suppression strategies based on signal-to-noise ratio. Their enhancement logic is disconnected from the user's actual attentional intent and cannot dynamically respond to changes in the instantaneous focus of hearing-impaired children. In recent years, multimodal assistive technologies have begun to explore incorporating visual cues into the speech enhancement process, but existing solutions still have limitations: some methods rely on lip-reading video for audiovisual fusion, but in practical use, they are easily affected by side profiles, occlusion, or insufficient lighting, leading to feature extraction failure; some methods, while using cameras to locate speakers and guide beamforming, do not combine user gaze behavior, often misidentifying unfocused speakers as target sound sources. Research on speech enhancement technologies based on eye-tracking or visual attention modeling is mostly confined to laboratory environments, lacking systematic integrated designs for child users. On the one hand, eye-tracking modules are not yet deeply integrated with hearing aid hardware, resulting in insufficient wearing comfort and long-term stability. On the other hand, existing modal fusion strategies mostly employ static weights or post-fusion mechanisms, failing to dynamically adjust audio processing intensity based on visual input quality (such as gaze clarity and head movement interference), leading to decreased enhancement robustness. The overall system architecture is not optimized for children's use scenarios, exhibiting significant shortcomings in end-to-end latency control, wearing comfort, and hardware miniaturization. Furthermore, the systems generally lack human-computer collaborative feedback mechanisms, failing to guide children to actively correct their gaze behavior through tactile or visual means when the enhancement result is inconsistent with the user's cognitive intent. These issues collectively restrict the speech intelligibility and social participation abilities of hearing-impaired children in real, complex acoustic environments. Therefore, there is an urgent need for a soundscape-selective enhancement method and system that uses eye tracking as an intent proxy, is lightweight, has low latency, and possesses closed-loop feedback capabilities. Summary of the Invention

[0003] To address the shortcomings of existing technologies, this invention provides a method and system for selectively enhancing speech in complex soundscapes based on eye-tracking guidance, which can effectively solve the problems mentioned in the background technology.

[0004] To achieve the above objectives, this invention provides a method for selectively enhancing speech in complex soundscapes based on eye-tracking guidance, comprising the following steps:

[0005] Step 1: Acquire user eye movement data. The eye movement data is collected by a miniature eye-tracking module integrated into the inner side of the temple of the hearing aid, including pupil center coordinates, spatial azimuth of fixation point, and fixation duration.

[0006] Step 2: Synchronously acquire multi-channel audio signals. The multi-channel audio signals are acquired through a ring microphone array set in the ear hook of the hearing aid. The timestamps of the multi-channel audio signals are aligned with the eye movement data.

[0007] Step 3: Calculate the initial spatial orientation estimate of the target sound source based on the spatial azimuth angle of the gaze point, and perform sound source localization correction in combination with multi-channel audio signals to generate a dynamically updated target sound source direction vector;

[0008] Step 4: Perform short-time Fourier transform on the multi-channel audio signal to obtain the time-frequency domain complex spectrum matrix, and construct a steering vector based on the target sound source direction vector to generate a preliminary enhanced speech signal;

[0009] Step 5: Perform quality assessment on the eye movement data. The quality assessment includes determining whether the fixation point is stable, whether it is located in the effective face area, and whether the ambient lighting meets the visual feature extraction conditions, and generating a visual confidence score.

[0010] Step 6: Dynamically adjust the audio enhancement strategy based on the visual confidence score. When the visual confidence score > the first preset threshold, high-gain beamforming is used to suppress non-target direction noise, and the preliminary enhanced speech signal generated in Step 4 is output as the final enhanced speech signal. When the visual confidence score < the second preset threshold, switch to the omnidirectional enhancement mode based on speech activity detection, enable the general speech enhancement model based on deep learning, weight the time-frequency domain complex spectrum matrix of Step 4, generate the general model output signal, and use it as the final enhanced speech signal. When the second preset threshold ≤ the visual confidence score ≤ the first preset threshold, a weighted fusion strategy is used to weightedly fuse the preliminary enhanced speech signal with the signal output by the general speech enhancement model to generate the final enhanced speech signal.

[0011] Step 7: Evaluate the speech intelligibility of the unique final enhanced speech signal generated in Step 6. The speech intelligibility evaluation is calculated based on the dynamic rate of change of the Mel frequency cepstral coefficients and the fundamental frequency stability. If the speech intelligibility evaluation result is lower than the third preset threshold, a feedback mechanism is triggered to transmit a tactile cue signal to the user, guiding them to refocus on the target speaker's face. After the user corrects their gaze, Steps 1-6 are executed again. If the evaluation result meets the standard, proceed directly to Step 8.

[0012] Step 8: Output the final enhanced speech signal after the above processing to the user through the miniature speaker of the hearing aid.

[0013] To optimize the above technical solution, the specific measures also include:

[0014] Furthermore, the miniature eye-tracking module includes an infrared light-emitting diode and a complementary metal-oxide-semiconductor image sensor, with a center wavelength of a preset wavelength, a radiation intensity of a preset intensity, a half-intensity angle of a preset angle range, an image sensor resolution of a preset resolution, a frame rate of a preset frame rate, and is connected to the main control chip via a flexible printed circuit board. The overall thickness does not exceed a preset thickness, and the weight is less than a preset weight.

[0015] Furthermore, the circular four-element microphone array is evenly distributed around the circumference, with a pre-set circumference diameter. Each microphone unit is an omnidirectional condenser microphone with a predetermined sensitivity, a signal-to-noise ratio greater than a preset signal-to-noise ratio threshold, and a preset sampling frequency. After pre-amplification and anti-aliasing filtering by the analog front-end circuit, it is input to the digital signal processor with a predetermined precision via an analog-to-digital converter.

[0016] Furthermore, the sound source localization correction process includes: firstly, using a generalized cross-correlation phase transform algorithm to calculate the time delay estimate between each microphone pair, and constructing an initial sound source direction candidate set; projecting the spatial azimuth angle of the gaze point onto the horizontal plane, constructing a visual guidance sector with a width of a preset angle centered on it, retaining only the sound source direction candidates located within the visual guidance sector, and if the number of candidates is zero, expanding the sector angle to a larger preset angle range; and selecting the candidate direction with the highest energy as the corrected target sound source direction vector.

[0017] Furthermore, the spatial azimuth angle of the gaze point is projected onto the horizontal plane, and a visual guidance sector with a width of ±15 degrees is constructed with it as the center. Only the sound source direction candidates located within the visual guidance sector are retained. If the number of candidates is zero, the sector angle is expanded to ±30 degrees.

[0018] Furthermore, the visual confidence score is obtained by weighted summation of three factors: the normalized value of ambient light intensity, the normalized reciprocal of the Euclidean distance from the fixation point to the center of the nearest detected face, and the ratio of the current fixation duration to a preset reference time. The sum of the weights of these three factors is 1, and the visual confidence score ranges from [0, 1]. The visual confidence score formula is as follows: ,in, This is the normalized value of ambient light intensity, ranging from 0 to 1; The normalized reciprocal of the Euclidean distance from the gaze point to the center of the nearest detected face is 0 to 1. It is the ratio of the current gaze duration to the preset reference time, with an upper limit of 1; the coefficients α, β, and γ are preset weight coefficients, and satisfy α+β+γ=1.

[0019] Furthermore, the general speech enhancement model based on deep learning is a lightweight convolutional recurrent neural network, containing multiple convolutional layers, multiple bidirectional gated recurrent unit layers, and a fully connected output layer. The total number of parameters is less than a predetermined size, and the inference latency is less than a predetermined latency threshold. It runs on an embedded neural network accelerator, with the input being the logarithmic power spectrum of noisy speech and the output being an ideal ratio mask IRM(f,t). The general speech enhancement model based on deep learning weights the multi-channel audio signal X(f,t) after Fourier transform: the output signal... , This represents the average value of the complex spectrum matrix in the time-frequency domain of the multi-channel audio signal.

[0020] Furthermore, the speech intelligibility assessment process includes: extracting the first few dimensions of the Mel frequency cepstral coefficients of the enhanced speech signal, calculating the first-order difference standard deviation between adjacent frames as a dynamic rate of change index; estimating the fundamental frequency trajectory using the autocorrelation function method, calculating the standard deviation of the fundamental frequency within several consecutive frames to obtain a fundamental frequency stability index; and finally, the intelligibility score is a weighted sum of the dynamic rate of change index and the fundamental frequency stability index.

[0021] Furthermore, a trigger feedback mechanism is implemented to transmit tactile cues to the user through bone conduction vibration. The frequency, amplitude, and duration of each trigger are all set to fixed parameters, and the interval between two triggers is no less than a preset time interval.

[0022] This invention also provides a system for selectively enhancing speech in complex soundscapes based on eye-tracking guidance, used to implement a method for selectively enhancing speech in complex soundscapes. The system includes: a miniature eye-tracking module for acquiring the user's eye movement data; a circular four-element microphone array for simultaneously acquiring multi-channel audio signals; a sound source localization and correction module for generating a dynamically updated target sound source direction vector based on the gaze point spatial azimuth angle of the eye movement data and the multi-channel audio signals; a beamforming module for generating a preliminary enhanced speech signal based on the target sound source direction vector; a quality assessment module for assessing the quality of the eye movement data and generating a visual confidence score; an adaptive enhancement strategy control module for dynamically adjusting the audio enhancement strategy based on the visual confidence score; a speech intelligibility assessment module for assessing the intelligibility of the final enhanced speech signal; a feedback module for triggering a tactile cue signal when the intelligibility assessment result is below a third preset threshold; and an audio output module for outputting the final enhanced speech signal to the user.

[0023] Furthermore, the system also includes a power management module, which is powered by a lithium polymer battery with a predetermined capacity, supports continuous operation for no less than a predetermined duration, and has a low battery warning function. When the remaining battery power is lower than a preset battery threshold, it automatically reduces the power consumption of non-core modules.

[0024] Furthermore, the entire system is encapsulated in an ergonomic shell that conforms to the shape of a child's head, with a total weight not exceeding a predetermined weight. The left and right ear devices achieve data synchronization through near-field magnetic induction technology, with a synchronization error less than a predetermined time error.

[0025] The present invention also provides an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor, when executing the computer program, implements the above-described method for selective enhancement of speech in complex soundscapes based on eye-tracking guidance.

[0026] The present invention also provides a computer-readable storage medium comprising a stored computer program, wherein, when the computer program is executed, it controls the device on which the computer-readable storage medium is located to perform an eye-tracking-guided method for selectively enhancing speech in complex soundscapes.

[0027] The beneficial effects of this invention are as follows: By integrating a miniature eye-tracking module with a hearing aid, this invention achieves an active target speech selection mechanism based on eye movement intent. It effectively overcomes the technical deficiency of traditional hearing aids in sensing the user's cognitive focus. The proposed dynamic sound source localization correction method effectively integrates visual priors and acoustic cues, maintaining high localization accuracy even under adverse conditions such as side profiles, occlusion, or low light. The visual confidence-driven adaptive enhancement strategy abandons rigid fixed fusion modes, intelligently switching or hybrid enhancement algorithms based on actual visual input quality, balancing robustness and performance. Low-latency speech intelligibility assessment and bone conduction feedback closed-loop endow the system with self-diagnosis and user guidance capabilities. Tactile cues correct gaze behavior, effectively solving the enhancement failure problem caused by user attention drift and improving the effectiveness of human-computer interaction. Lightweight hardware and optimized algorithm design achieve end-to-end low-latency processing, meeting the ergonomic wearing needs of children and all-weather usage scenarios. Attached Figure Description

[0028] Figure 1 is a flowchart illustrating the method for selective speech enhancement of complex soundscape targets based on eye-tracking guidance proposed in this invention.

[0029] Figure 2 is a schematic diagram of the structure of the eye-tracking-guided target speech selective enhancement system for complex soundscapes proposed in this invention. Detailed Implementation

[0030] The invention will now be described in further detail with reference to the accompanying drawings.

[0031] It should be noted that the terms such as "upper", "lower", "left", "right", "front", and "back" used in the invention are only for clarity of description and are not intended to limit the scope of the invention. Changes or adjustments to their relative relationships, without substantially altering the technical content, should also be considered within the scope of the invention.

[0032] The embodiments described in this invention are merely some, not all, of the embodiments of this invention. All other embodiments obtained by those skilled in the art based on the embodiments of this invention without inventive effort are within the scope of protection of this invention.

[0033] Example 1, as shown in Figure 1, discloses an eye-tracking-guided method for selectively enhancing target speech in complex soundscapes. This method addresses the difficulty hearing-impaired children face in perceiving target speech in typical social scenarios such as classrooms and streets where multiple people speak simultaneously. The method deeply integrates a miniature eye-tracking module with a circular four-element microphone array into the hearing aid, constructing a dynamic sound source localization mechanism based on eye-tracking signals as cognitive intent. Building upon this, it introduces an adaptive audio enhancement strategy control logic driven by visual confidence. Finally, through low-latency speech intelligibility assessment and bone conduction tactile feedback closed-loop, it forms a human-machine collaborative enhancement system with self-healing performance capabilities. The entire system adheres to end-to-end processing latency of less than 90 milliseconds, a total weight of no more than 28 grams, and a continuous working time of no less than 8 hours, ensuring its suitability for the daily wear needs of hearing-impaired children.

[0034] In one implementation, step 1 involves acquiring eye movement data of the hearing-impaired child. This process is performed by a miniature eye-tracking module integrated into the inner side of the temple of the hearing aid. This module consists of an infrared light-emitting diode with a center wavelength of 850 nm and a complementary metal-oxide-semiconductor (CMOS) image sensor with a resolution of 640×480 pixels. The infrared light source has a radiation intensity of 12 mW / sr, with a maximum of 20 mW / sr, and a half-intensity angle of ±15 degrees (in one implementation, the models selected are ROHM SML-P14RW and Osram SFH4030B). It employs a pulse-driven mode (frequency 200 Hz, duty cycle 8%), meeting the requirements of photobiological safety risk level 1 in GB / T20145-2006 (corresponding to IEC62471). The image sensor has a frame rate of 120 Hz, meeting the dynamic response requirements of real-time eye tracking. The module is no more than 3.5 mm thick and weighs less than 1.2 grams. It is connected to the main control chip via a flexible printed circuit board, ensuring long-term wearing comfort and not interfering with children's head movements. The miniature eye-tracking module also integrates a closed-loop circuit for monitoring light source intensity. It collects the actual radiation intensity of the infrared light source in real time. When the intensity exceeds the 20 mW / sr threshold, it immediately triggers power-off protection. After restarting, it defaults to a safe level of 10 mW / sr, further enhancing the safety for children.

[0035] During operation, an infrared light source illuminates the surface of the eyeball, forming a stable reflective spot on the cornea. The image sensor captures a sequence of images containing the pupil outline and the corneal reflective point at a rate of 120 frames per second. The main control chip runs a real-time pupil center localization algorithm, employing a geometric model based on ellipse fitting or a weighted centroid method to accurately extract the pupil center coordinates from each frame of the image. Subsequently, using a pre-calibrated head-eye geometry model (established through an offline calibration procedure, containing a fixed spatial relationship between the center of the eye socket, the center of the pupil, and the head coordinate system), the two-dimensional image coordinates are converted into a spatial three-dimensional gaze direction vector. This vector is projected onto the horizontal plane to obtain the spatial azimuth of the gaze point. The unit is degrees, and the range is [-180°, 180°]. Simultaneously, the system records the start timestamp of the current gaze event. With end timestamp Calculate fixation duration The unit is milliseconds. All eye movement data, including pupil center coordinates ( ), Spatial azimuth of the gaze point fixation duration — Each one is accompanied by a high-precision timestamp with a time resolution better than 0.5 milliseconds to ensure strict synchronization with subsequent audio signals.

[0036] Step 2: Synchronous acquisition of multi-channel audio signals. This process is accomplished by a circular quad microphone array located in the ear hook of the hearing aid. This array consists of four omnidirectional condenser microphone units, evenly distributed in a 35mm diameter circle. Each microphone has a sensitivity of -38 dB (0 dB reference 1 V / Pa), a signal-to-noise ratio greater than 65 dB, and a sampling frequency of 16 kHz. After the ambient sound signals are synchronously picked up by the four microphones, they first enter the analog front-end circuit. This circuit includes a preamplifier providing a fixed gain of 20 dB and an integrated 8 kHz cutoff Butterworth anti-aliasing low-pass filter to effectively suppress high-frequency noise above the Nyquist frequency. The filtered analog signal is then input to a 24-bit Σ-Δ analog-to-digital converter (model SGM58601) for quantization in pulse density modulation format. The digital signal processor has an internal hardware timestamp unit, whose clock source shares the same 10 MHz temperature-compensated crystal oscillator with the eye-tracking module, ensuring that the time alignment error between the two data streams is less than 0.1 milliseconds. Multi-channel audio signals are buffered in a double buffer using a block processing approach. Each block has 512 sampling points, corresponding to a 32-millisecond time window, with a frame shift of 256 points (i.e., an overlap rate of 50%), to balance real-time performance and spectral resolution. Each audio block carries a timestamp that is strictly aligned with the eye-tracking data, providing a precise time reference for subsequent multimodal fusion.

[0037] Step 3: Generate dynamically updated target sound source direction vectors. This process consists of two stages: initial estimation and acoustic correction. The initial estimation stage directly uses the gaze point spatial azimuth angle obtained in Step 1. This serves as a rough location of the target sound source. The calibration stage utilizes acoustic information to optimize the initial estimate, improving localization robustness. First, the four-channel audio signals acquired in step 2 are combined in pairs to form a coarse location. =6 microphone pairs. For the signal of each microphone pair and Perform the Generalized Cross-Correlation Phase Transform (GCC-PHAT) algorithm to calculate its cross-power spectrum. ,in for Perform a Fourier transform; then calculate the cross-correlation function after the phase transform. ,in Indicates the inverse Fourier transform; take The time delay corresponding to the peak position This serves as the time delay estimate for the microphone pair. Based on the known geometric layout of the microphone array (circumference radius r = 17.5 mm), geometric relationships are used... Six sets of time delay estimates The mapping is divided into six candidate sound source directions, among which Let be the candidate angles for the direction of the sound source, in degrees, with a range of [-180°, 180°]. Let be the azimuth angle of the k-th microphone, k = 1, 2, ..., 6, and c be the speed of sound (343 m / s). Then, the system will assign the spatial azimuth angle of the gaze point. Projecting onto a horizontal plane, a visual guidance sector with a width of ±15 degrees is constructed centered on it. -15°, [+15°], only candidate sound source directions located within this sector are retained. If the number of candidates after filtering is zero (indicating a serious conflict between acoustic localization and visual guidance, possibly due to strong noise or reverberation), the sector is expanded to ±30 degrees and filtered again. Finally, among the retained candidate directions, the total energy of the four-channel microphone signal corresponding to each candidate direction is calculated. And select the candidate direction with the highest energy. This serves as the corrected target sound source direction vector. This vector is in unit two-dimensional vector form. This indicates that it is used for subsequent beamforming calculations.

[0038] Step 4: Generate preliminary enhanced speech signal. The system performs Short Time Fourier Transform (STFT) on the four-channel audio signals, using a Hamming window with a window length N of 512 points and a frame shift of 256 points, to obtain the four-channel time-frequency domain complex spectrum matrix. Where 257 is the number of frequency bins (corresponding to 0 to 8 kHz), and N is the number of frames. Based on the target sound source direction vector generated in step 3. Based on the geometric layout of the microphone array, the steering vector corresponding to each frequency bin f is calculated. The m-th element of the guiding vector. The sampling frequency is (Unit: Hz), FFT window length: N, Let be the propagation delay of the m-th microphone relative to the reference microphone (usually the array center), determined by geometric relationships. Sure, Let be the distance from the m-th microphone to the reference point. Let be the azimuth angle of the m-th microphone. Then, the Minimum Variance Distortionless Response (MVDR) beamforming algorithm is employed, with its weight vector... The following optimization problem is obtained at each frequency bin f: ;in The noise covariance matrix is... Online estimation is performed during periods of speech inactivity (determined by the speech activity detection module), and an exponential smoothing factor is used. =0.95 was updated. The time-spectrum of the initial enhanced speech signal was analyzed. ,in ∈ , ∈ Finally, regarding Overlapping addition (OLA) synthesis was performed to obtain a preliminary time-domain enhanced speech signal. With an overlap rate of 50%, inter-frame stitching distortion is eliminated.

[0039] Step 5: Generate a visual confidence score. The reliability of the current visual input for sound source localization is comprehensively evaluated and calculated using a weighted average of three dimensions. The first dimension is the normalized value of ambient light intensity. Its value is obtained by linearly mapping the feedback value of the image sensor's automatic exposure control (AEC) from the eye-tracking module, and the mapping formula is: ,in and These are the preset minimum and maximum effective exposure thresholds, respectively, to ensure... The value ranges from [0,1]. A value closer to 1 indicates more abundant lighting, which is more conducive to visual feature extraction. The second dimension is... This is the normalized reciprocal of the Euclidean distance from the gaze point to the center of the nearest detected face. The system incorporates a lightweight face detector (based on the MobileNetV2 backbone network, with fewer than 50,000 parameters), running once every 10 frames (approximately 83 milliseconds) and outputting the center coordinates of all face bounding boxes in the image. , ). The coordinates of the current gaze point ( , Calculate the Euclidean distance between the face centers and all faces. Take the minimum value Then through the formula Calculation, where To preset the maximum effective distance, corresponding to 40% of the image diagonal length, ensure... ∈[0,1]. The third dimension is That is, the ratio of the current gaze duration to 500 milliseconds. The upper limit is 1, reflecting fixation stability. Finally, the visual confidence score... ,in =0.4、 =0.4、 =0.2, and satisfies α+β+γ=1. This score is calculated in real time after each eye-tracking event update, and its value ranges from [0,1].

[0040] Step 6: Dynamically adjust the audio enhancement strategy. The adaptive enhancement strategy adjustment module adjusts according to... Three different strategies are applied to the numerical range to generate a unique final enhanced speech signal. The first preset threshold is set to 0.7, and the second preset threshold is set to 0.3 (the preset thresholds are based on the optimal threshold range obtained from statistical analysis of 100 sets of experimental data). When When the value is greater than 0.7, the system determines that the visual information is highly reliable and adopts a high-gain beamforming strategy, that is, directly outputs the preliminary enhanced speech signal generated in step 4. As It also activates the strong noise suppression module, applying an additional 20 dB attenuation to the spectral components in non-target directions.

[0041] when When the value is less than 0.3, the system determines that visual information is unavailable (e.g., insufficient lighting, absence of face regions, gaze drift) and switches to an omnidirectional enhancement mode based on speech activity detection. In this mode, beamforming is disabled, and a general speech enhancement model based on deep learning is enabled instead. This model is a lightweight convolutional recurrent neural network, containing three convolutional layers (kernel sizes of 7, 5, and 3, channel numbers of 32, 64, and 128, stride of 1, and same padding), two bidirectional gated recurrent unit (BiGRU) layers (with 128 hidden units), and a fully connected output layer. The total number of model parameters is less than 200,000, and it runs on an embedded neural network accelerator (NPU) with an inference latency of less than 25 milliseconds. The training process of a general speech enhancement model based on deep learning is as follows: Training data preparation: The TIMIT subset of children's speech is used as clean speech. The NOISEX-92 noisy dataset is used, and noisy speech samples are generated by mixing with a signal-to-noise ratio of -5dB to 15dB. The samples are resampled to 16kHz, and the total number of training samples is ≥100,000. The frame segmentation parameters are the same as in step 4 (window length 512 points, frame shift 256 points). Feature and label generation: The log power spectrum (LPS) of the noisy speech is extracted as the input feature of the model, and the corresponding clean speech is used as the input feature. With noise Extract LPS separately, according to the formula The true IRM was calculated as the training label, and the label was normalized to the [0,1] interval using min-max. Model training: The Adam optimizer was used, with mean squared error (MSE) as the loss function, and training was conducted for 30 epochs with a batch size of 32. Convolutional and fully connected layers were initialized using Xavier, and a dropout layer (rate=0.2) was added to prevent overfitting. Model optimization: After training, the optimal model was selected through a validation set, and INT8 quantization compression was used to ensure that the inference latency was ≤25 milliseconds. The SNR improvement on the test set was ≥8dB, and the STOI was ≥0.85, meeting the deployment requirements of embedded devices. The model input was the log power spectrum of a single-channel noisy speech (averaged across four channels). Generation process: Take the four-channel time-frequency domain complex spectrum matrix from step 4. The arithmetic mean is calculated according to the channel dimension to obtain the single-channel complex spectrum. ,right Take the modulus to obtain the amplitude spectrum. Calculate the square of the amplitude spectrum to obtain the power spectrum. Perform a natural logarithmic transform on the power spectrum to obtain the logarithmic power spectrum. = The model output is an Ideal Ratio Mask (IRM). This is used to weight the time-frequency domain complex spectrum matrix in step 4, while preserving the phase. The original phase: ,in, The average value of the complex spectrum matrix in the time-frequency domain of the four-channel audio signal. yes The phase term is used to obtain the output signal of the general model. China as .

[0042] When 0.3≤ When the value is ≤0.7, the system adopts a weighted fusion strategy to initially enhance the speech signal. Output signal of the general model Weighted fusion ultimately enhances the speech signal. , where the fusion weight λ=( -0.3) / 0.4, to achieve a smooth transition from omnidirectional mode to directional mode.

[0043] Step 7, trigger the feedback mechanism. The speech intelligibility assessment module evaluates the unique response generated in Step 6. Perform real-time analysis to assess whether its understandability meets usage requirements. First, extract... The first 13 dimensions of the Mel-frequency cepstral coefficients (MFCCs) were used, with a frame length of 25 milliseconds (400 points) and a frame shift of 10 milliseconds (160 points) to obtain the MFCC sequence. Calculate the first-order difference between adjacent frames. The dynamic feature sequence is obtained, and then the standard deviation of the sequence is calculated. M=, as a dynamic rate of change indicator Secondly, the fundamental frequency trajectory F0(t) is estimated using the autocorrelation function method. The F0 values ​​of 10 consecutive frames are selected, and their standard deviation is calculated. As a fundamental frequency stability indicator Ensure P∈[0,1]. The final intelligibility score C=0.6·M+0.4·P. The third preset threshold is set to 0.5. If C>0.5, it indicates that the intelligibility of the enhanced speech meets the standard, and proceed directly to step 8; if C<0.5, it indicates that the enhanced speech lacks sufficient dynamic fluctuations or the fundamental frequency is unstable, resulting in low intelligibility, and the system triggers the feedback mechanism. The bone conduction vibration unit then starts, generating a sinusoidal vibration with a frequency of 200 Hz and an amplitude of 10 micrometers, lasting for 300 milliseconds. This vibration is transmitted directly to the inner ear through the skull, forming a clear and identifiable tactile cue, guiding the hearing-impaired child to realize that the current gaze may be deviating from the target, and prompting them to refocus on the target speaker's face. To prevent frequent interference, the system sets a cooling-off period, with an interval of no less than 2 seconds between two feedback triggers; after the user corrects their gaze, the system re-executes steps 1-6 to generate a new... And conduct another comprehensibility assessment.

[0044] Step 8: Output the enhanced speech signal. The final enhanced speech signal obtained after all the above processing steps. The signal is fed into the audio output module. This module contains a 24-bit digital-to-analog converter (16 kHz sampling rate), a 15 mW Class-D power amplifier, and a miniature balanced armature speaker with a 32 ohm impedance. The signal is converted into sound waves by the speaker and then inserted into the ear canal of a hearing-impaired child through a custom silicone earmold, completing the entire enhancement process. The system's end-to-end processing latency is the sum of the latency of each module: eye-tracking processing latency of 15 ms, audio acquisition and preprocessing latency of 5 ms, sound source localization and beamforming latency of 25 ms, adaptive strategy adjustment latency of 10 ms, and intelligibility assessment latency of 10 ms, totaling 65 ms, far below the 90 ms real-time requirement.

[0045] Example 2, building upon Example 1, introduces a gaze direction compensation mechanism based on inertial measurement unit (IMU) assistance for use by hearing-impaired children in outdoor bright light or rapid head-turning scenarios, to improve the robustness of eye-tracking-acoustic fusion positioning. The core difference in this example lies in the expansion of steps 1 and 3, while the remaining steps remain the same.

[0046] In step 1, in addition to the miniature eye-tracking module, the hearing aid also integrates a triaxial microelectromechanical system (MEMS) inertial measurement unit (IMU) on the inner side of the temple, containing an accelerometer and a gyroscope, with a sampling frequency of 200 Hz. This IMU is used to monitor head angular velocity in real time. With linear acceleration When the detected head angular velocity exceeds a preset threshold (e.g., 50 degrees / second), the system determines that the user is in a state of rapid head rotation. At this point, the spatial azimuth of the fixation point calculated solely based on eye-tracking data... Distortions can occur due to head movements. Therefore, the system incorporates head movement compensation: the relative eye movement angle output by the eye-tracking module is converted into a more accurate representation of the original image. Absolute head orientation angle relative to IMU output By merging, the absolute spatial gaze direction can be obtained. = + .in, The absolute gaze direction is obtained by integrating the gyroscope angular velocity and performing zero-bias correction using the accelerometer. This absolute gaze direction is then input as a new initial estimate in step 3.

[0047] In step 3, the sound source localization correction process is adjusted accordingly. The center of the visually guided sector no longer uses the original... Instead, it uses the IMU-compensated version. Furthermore, the sector width is dynamically adjusted according to the head's movement: when the head is stationary ( When the speed is <10 degrees / second, the sector width remains ±15 degrees; when the head rotates slowly (10 ≤ When the speed is <50 degrees / second, the sector expands to ±20 degrees; when the head rotates rapidly ( When the speed is ≥50 degrees / second, the sector expands to ±30 degrees. This dynamic sector mechanism effectively addresses the visual-acoustic spatial alignment error caused by head movement, significantly improving the success rate of sound source localization in dynamic scenes. Experiments show that in running or head-turning scenarios, the target sound source localization accuracy of this embodiment is 22% higher than that of Embodiment 1.

[0048] As shown in Figure 2, the present invention also provides a speech selective enhancement system for complex soundscape targets based on eye-tracking guidance, which is used to implement the above-mentioned speech selective enhancement method for complex soundscape targets. The system includes: a miniature eye-tracking module, as described above, integrated into the inner side of the temple of the hearing aid to acquire the user's eye movement data; a ring-shaped four-element microphone array, embedded in the ear hook of the hearing aid with a fixed physical layout, for synchronously acquiring multi-channel audio signals; a sound source localization and correction module, implemented by a dedicated instruction set accelerator in the main control chip, for generating a dynamically updated target sound source direction vector based on the gaze point spatial azimuth angle of the eye movement data and the multi-channel audio signals; a beamforming module, running on the vector operation unit of the digital signal processor, for generating a preliminary enhanced speech signal based on the target sound source direction vector; a quality assessment module, executed by the general-purpose computing core of the main control chip, for assessing the quality of the eye movement data and generating a visual confidence score; an adaptive enhancement strategy control module, for dynamically adjusting the audio enhancement strategy based on the visual confidence score; a speech intelligibility assessment module, for assessing the intelligibility of the final enhanced speech signal; a feedback module, for triggering a tactile cue signal and driving the bone conduction vibration unit when the intelligibility assessment result is lower than a third preset threshold; and an audio output module, for outputting the final enhanced speech signal to the user.

[0049] The entire system is encapsulated in an ergonomic shell designed to fit a child's head, with a total weight of no more than 28 grams. The left and right earbuds achieve data synchronization via near-field magnetic induction technology, with a synchronization error of less than 0.5 milliseconds. The system also includes a power management module powered by a 120mAh lithium polymer battery with a predetermined capacity, supporting over eight hours of continuous operation. It features a low-battery warning function and enters a low-power mode when the battery level drops below 15%, disabling non-core functions to extend usage time.

[0050] The present invention also discloses an electronic device, including a processor, a memory, and a computer program stored in the memory and configured to be executed by the processor, wherein the processor executes the computer program to implement the above-described method for selective enhancement of speech in complex soundscapes based on eye-tracking guidance.

[0051] The present invention also discloses a computer-readable storage medium, including a stored computer program, wherein, when the computer program is running, it controls the device where the computer-readable storage medium is located to execute the above-described method for selective enhancement of speech in complex soundscapes based on eye-tracking guidance.

[0052] The above are merely preferred embodiments of the present invention. The scope of protection of the present invention is not limited to the above embodiments. All technical solutions falling within the scope of the present invention's concept are within the scope of protection of the present invention. It should be noted that for those skilled in the art, any improvements and modifications made without departing from the principles of the present invention should be considered within the scope of protection of the present invention.

Claims

1. A method for selective speech enhancement of complex soundscape targets based on eye-tracking guidance, characterized in that, Includes the following steps: Step 1: Acquire user eye movement data. This data is collected through a miniature eye-tracking module integrated into the inner side of the temple of the hearing aid, including pupil center coordinates, spatial azimuth of the fixation point, and fixation duration. Step 2: Synchronously acquire multi-channel audio signals. These signals are acquired through a ring microphone array located in the ear hook of the hearing aid. The timestamps of the multi-channel audio signals are aligned with the eye movement data. Step 3: Calculate the initial spatial azimuth estimate of the target sound source based on the spatial azimuth of the fixation point. Combine this with the multi-channel audio signals for sound source localization correction, generating a dynamically updated target sound source direction vector. Step 4: Perform a short-time Fourier transform on the multi-channel audio signals to obtain a time-frequency domain complex spectrum matrix. Construct a steering vector based on the target sound source direction vector to generate a preliminary enhanced speech signal. Step 5: Perform a quality assessment on the eye movement data. This assessment includes determining whether the fixation point is stable, whether it is located within a valid face area, and whether the ambient lighting meets the visual feature extraction conditions, generating a visual confidence score. Step 6: Dynamically adjust the audio enhancement strategy according to the visual confidence score. When the visual confidence score is greater than the first preset threshold, high-gain beamforming is used and non-target direction noise is suppressed. The preliminary enhanced speech signal generated in step 4 is output as the final enhanced speech signal. When the visual confidence score is less than the second preset threshold, switch to the omnidirectional enhancement mode based on speech activity detection, enable the general speech enhancement model based on deep learning, weight the time-frequency domain complex spectrum matrix of step 4, generate the general model output signal and use it as the final enhanced speech signal. When the second preset threshold ≤ visual confidence score ≤ first preset threshold, a weighted fusion strategy is adopted to weightedly fuse the preliminary enhanced speech signal with the signal output by the general speech enhancement model to generate the final enhanced speech signal; Step 7: The speech intelligibility of the unique final enhanced speech signal generated in Step 6 is evaluated. The speech intelligibility evaluation is calculated based on the dynamic change rate of the Mel frequency cepstral coefficients and the fundamental frequency stability; If the speech intelligibility evaluation result is lower than the third preset threshold, a feedback mechanism is triggered to transmit a tactile cue signal to the user, guiding them to refocus on the target speaker's face. After the user corrects their gaze, Steps 1-6 are executed again; If the evaluation result meets the standard, the process proceeds directly to Step 8; Step 8: The final enhanced speech signal processed as described above is output to the user through the miniature speaker of the hearing aid device.

2. The method for selectively enhancing the target speech in complex soundscapes according to claim 1, characterized in that, The miniature eye-tracking module includes an infrared light-emitting diode and a complementary metal-oxide-semiconductor image sensor, which are connected to the main control chip via a flexible printed circuit board.

3. The method for selectively enhancing the target speech in complex soundscapes according to claim 1, characterized in that, The circular four-element microphone array is evenly distributed around the circumference, and each microphone unit is an omnidirectional condenser microphone.

4. The method for selectively enhancing the target speech in complex soundscapes according to claim 1, characterized in that, The sound source localization correction process includes: calculating the time delay estimate between each microphone pair and constructing an initial sound source direction candidate set; projecting the spatial azimuth angle of the gaze point onto the horizontal plane, constructing a visual guidance sector with a width of a preset angle centered on it, retaining only the sound source direction candidates located within the visual guidance sector, and if the number of candidates is zero, expanding the sector angle to a larger preset angle range; and selecting the candidate direction with the highest energy as the corrected target sound source direction vector.

5. The method for selectively enhancing the target speech in complex soundscapes according to claim 4, characterized in that, Project the spatial azimuth of the gaze point onto the horizontal plane, and construct a visual guidance sector with a width of ±15 degrees centered on it. Only retain the candidate sound source direction located within the visual guidance sector. If the number of candidates is zero, expand the sector angle to ±30 degrees.

6. The method for selectively enhancing the target speech in complex soundscapes according to claim 1, characterized in that, The visual confidence score is obtained by weighted summing of three factors: the normalized value of ambient light intensity, the normalized reciprocal of the Euclidean distance from the fixation point to the center of the nearest detected face, and the ratio of the current fixation duration to a preset reference time. The sum of the weights of these three factors is 1, and the visual confidence score ranges from [0, 1]. The visual confidence score formula is as follows: ,in, For visual confidence, This is the normalized value of ambient light intensity, ranging from 0 to 1; The normalized reciprocal of the Euclidean distance from the gaze point to the center of the nearest detected face is 0 to 1. It is the ratio of the current gaze duration to the preset reference time, with an upper limit of 1; the coefficients α, β, and γ are preset weight coefficients, and satisfy α+β+γ=1.

7. The method for selectively enhancing the target speech in complex soundscapes according to claim 1, characterized in that, The general speech enhancement model based on deep learning is a lightweight convolutional recurrent neural network, containing multiple convolutional layers, multiple bidirectional gated recurrent unit layers, and a fully connected output layer. The input is the logarithmic power spectrum of the noisy speech, and the output is an ideal ratio mask IRM(f,t), used to enhance the multi-channel audio signal after Fourier transform. Weighting: Output signal , This represents the average value of the complex spectrum matrix in the time-frequency domain of the multi-channel audio signal.

8. The method for selectively enhancing the target speech in complex soundscapes according to claim 1, characterized in that, The speech intelligibility assessment process includes: extracting the first few dimensions of the Mel frequency cepstral coefficients of the enhanced speech signal, calculating the first-order difference standard deviation between adjacent frames as a dynamic rate of change index; estimating the fundamental frequency trajectory using the autocorrelation function method, calculating the standard deviation of the fundamental frequency within several consecutive frames to obtain a fundamental frequency stability index; and finally, the intelligibility score is a weighted sum of the dynamic rate of change index and the fundamental frequency stability index.

9. The method for selectively enhancing the target speech in complex soundscapes according to claim 1, characterized in that, The trigger feedback mechanism transmits tactile cues to the user through bone conduction vibration. The frequency, amplitude, and duration of each trigger are all set to fixed parameters, and the interval between two triggers is no less than a preset time interval.

10. A speech selective enhancement system for complex soundscape targets based on eye-tracking guidance, characterized in that, A method for selectively enhancing target speech in complex soundscapes as described in any one of claims 1-9, comprising: a miniature eye-tracking module for acquiring user eye movement data; a circular four-element microphone array for simultaneously acquiring multi-channel audio signals; a sound source localization and correction module for generating a dynamically updated target sound source direction vector based on the gaze point spatial azimuth angle of the eye movement data and the multi-channel audio signals; a beamforming module for generating a preliminary enhanced speech signal based on the target sound source direction vector; a quality assessment module for assessing the quality of the eye movement data and generating a visual confidence score; an adaptive enhancement strategy adjustment module for dynamically adjusting the audio enhancement strategy based on the visual confidence score; a speech intelligibility assessment module for assessing the intelligibility of the final enhanced speech signal; a feedback module for triggering a tactile cue signal when the intelligibility assessment result is lower than a third preset threshold; and an audio output module for outputting the final enhanced speech signal to the user.