BOSSA-based deaf-mute multi-mode dynamic response interaction method, system and device and medium
By using a BOSSA-based multimodal dynamic response interaction method, combining sound, dynamic sign language, and environmental data, and simulating the physiological characteristics of the human ear for filtering and sound source tracking, the problem of deaf-mute people locating the speaker in complex environments is solved. This achieves high-precision sound source recognition and multimodal feedback, thus improving the interactive experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHAANXI YIZUYUAN REGENERATIVE MEDICINE CO LTD
- Filing Date
- 2026-01-30
- Publication Date
- 2026-05-01
AI Technical Summary
Deaf and mute individuals often struggle to accurately locate the speaker in complex environments. Existing sound source tracking methods are ill-equipped to handle vertical displacement in three-dimensional space, leading to interruptions in interaction. Furthermore, particle filtering algorithms suffer from issues such as random initialization, computational redundancy, and particle degradation.
A multimodal dynamic response interaction method based on BOSSA is adopted. By acquiring sound, dynamic sign language and environmental data, the physiological characteristics of cochlear hair cells in the human ear are simulated for filtering. The sound classification is performed using Gammatone filter bank and Gaussian mixture model. The sound source is tracked by combining bio-inspired spatial tuning neuron clusters, generating spatial activation map and outputting multimodal feedback.
It improves the accuracy of sound source recognition and the naturalness of interaction, enhances the ability of deaf and mute people to perceive the surrounding acoustic environment, solves the problems of large sound source tracking range and difficulty in vertical tracking, and significantly improves the interactive experience and communication efficiency.
Smart Images

Figure CN121963768A_ABST
Abstract
Description
A multimodal dynamic response interaction method, system, device, and medium for deaf and mute users based on Bossa A. Technical Field
[0001] This invention relates to the field of voice interaction, and more particularly to a multimodal dynamic response interaction method, system, device and medium for deaf and mute people based on BOSAS. Background Technology
[0002] Human sound perception relies heavily on cochlear hair cells, with high and low frequencies handled by hair cells in different areas of the cochlea. Because high-frequency hair cells are more fragile, various types of hearing loss often manifest first as high-frequency hearing loss. In human speech, vowel energy is concentrated in the low-frequency range of 500Hz–2kHz, while the key consonants that determine semantic clarity are mainly distributed in the high-frequency range of 2–4kHz. Therefore, many deaf-mute individuals with residual hearing, although able to perceive sound, struggle to understand meaning due to the absence of high-frequency consonants, forming the core obstacle of "hearing but not understanding."
[0003] Furthermore, the daily interactions of deaf and mute individuals are characterized by significant mobility; actions such as a speaker getting up or going up and down stairs cause dynamic changes in the sound source in three-dimensional space. Traditional sound source tracking methods are mostly based on two-dimensional horizontal plane modeling, which struggles to handle vertical displacement, easily leading to sound source loss and interruption of interaction. Existing particle filtering algorithms suffer from problems such as random initialization, computational redundancy, and particle degradation, making it difficult to meet the requirements of real-time, stable, and omnidirectional sound source tracking. Therefore, how to enable deaf and mute individuals to more accurately locate speakers in complex environments is an urgent problem to be solved. Summary of the Invention
[0004] This invention provides a BOSSA-based multimodal dynamic response interaction method, system, device, and medium for deaf and mute individuals, solving the technical problem in the prior art where deaf and mute individuals have difficulty accurately locating the speaker in complex environments.
[0005] In a first aspect, the present invention provides a BOSSA-based multimodal dynamic response interaction method for deaf and mute individuals, comprising:
[0006] The system acquires and preprocesses multimodal data to obtain multimodal data, which includes sound data, dynamic sign language data, and environmental data. A filtered signal is obtained based on the sound data and the physiological characteristics of the human ear, including the high-frequency sensitivity and low-frequency tolerance characteristics of cochlear hair cells. The category determination result of the sound data is obtained based on the filtered signal. A spatial activation map is generated based on the category determination result. The optimal sound source tracking result is output based on the spatial activation map. Finally, a multimodal feedback result is output based on the optimal sound source tracking result and the category determination result.
[0007] Furthermore, based on the sound data and the physiological characteristics of the human ear, a filtered signal is obtained, including: designing a filter bank based on the physiological characteristics of the human ear and the hearing needs of deaf and mute individuals. The filter bank consists of 55 Gammatone filters covering a frequency band of 20Hz-20kHz; inputting the sound data into the filter bank, and completing frequency decomposition through parallel filtering and envelope extraction to obtain the narrowband output signal corresponding to each Gammatone filter; filtering and suppressing the 55 narrowband output signals to obtain the filtered signal.
[0008] Furthermore, the classification result of the sound data is obtained based on the filtered signal, including: extracting key identifier features based on the filtered signal and constructing feature vectors; constructing a sound object feature library and training the Gaussian mixture model with the sound object feature library; inputting the feature vectors into the trained Gaussian mixture model and determining the sound object category of the sound data through the likelihood value; and determining the classification result of the sound data based on the sound object priority, classification confidence, and sound object category.
[0009] Furthermore, a spatial activation map is generated based on the category determination results, including: extracting binaural cues from each frequency channel of the Gammatone filter based on a bio-inspired spatially tuned neuron cluster architecture; fusing the binaural cues based on a nonlinear multiplication mechanism to obtain the spatial orientation response intensity of the STN cluster; optimizing the spatial orientation response intensity of the STN cluster based on sound object priority, classification confidence, and category determination results; and generating a spatial activation map based on the optimized spatial orientation response intensity of the STN cluster through an inhibitory side connection network.
[0010] Furthermore, the optimal sound source tracking result is output based on the spatial activation map, including: extracting the horizontal azimuth angle corresponding to the peak in the spatial activation map and generating the corresponding initial particle set; updating the weight of each particle in the initial particle set based on the spatial activation map as the observation basis; determining the adaptive motion model based on the motion state of the sound source; obtaining the optimal sound source tracking result in two-dimensional space under the adaptive motion model, or, under the adaptive motion model, increasing the state vector of the particles and the observation dimension to obtain the optimal sound source tracking result in three-dimensional space.
[0011] Furthermore, based on the optimal sound source tracking results and category determination results, multimodal feedback results are output, including: based on the spatiotemporal correlation verification results, sound sources that meet the correlation conditions are determined as interactive correlation sound sources and bound to the corresponding dynamic sign language in the verification; the interactive experience is optimized based on the optimal sound source tracking results, user habits, and scene characteristics.
[0012] Secondly, this invention provides a multimodal dynamic response interaction system for deaf and mute individuals based on BOSAS, comprising: a multimodal data acquisition and preprocessing module for acquiring and preprocessing multimodal data to obtain multimodal data, wherein the multimodal data includes sound data, dynamic sign language data, and environmental data; a BOSAS-customized sound processing module, comprising a cochlear filtering unit, a sound object modeling unit, and a spatial tuning and localization unit, wherein the cochlear filtering unit is used to obtain a filtered signal based on the sound data and the physiological characteristics of the human ear; the sound object modeling unit is used to obtain a category determination result of the sound data based on the filtered signal; the spatial tuning and localization unit is used to generate a spatial activation map based on the category determination result; a dynamic sound source tracking module is used to output the optimal sound source tracking result based on the spatial activation map; and a multimodal intelligent interaction control module is used to output multimodal feedback results based on the optimal sound source tracking result and the category determination result.
[0013] Thirdly, the present invention provides an electronic device, comprising: a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to cause the electronic device to perform the BOSSA-based multimodal dynamic response interaction method for deaf and mute users as provided in the first aspect.
[0014] Fourthly, the present invention provides a non-transitory computer-readable storage medium, wherein when the instructions in the storage medium are executed by the processor of an electronic device, the electronic device is able to execute the BOSSA-based multimodal dynamic response interaction method for deaf and mute individuals as provided in the first aspect.
[0015] The present invention provides one or more technical solutions that have at least the following technical effects or advantages: Based on the BOSAS framework, the present invention effectively improves the naturalness and real-time performance of information interaction by fusing sound, dynamic sign language, and environmental data. The present invention improves the accuracy of sound source identification by simulating the physiological characteristics of human cochlear hair cells—sensitivity to high frequencies and tolerance to low frequencies—through biomimetic filtering of sound signals. It utilizes 55 narrowband signals to accurately determine sound categories and generates spatial activation maps to achieve high-precision sound source tracking. Combining sound source location and category information, it can output multimodal feedback (such as visualization or vibration cues) adapted to the perceptual characteristics of deaf and mute users, significantly enhancing their ability to perceive the surrounding acoustic environment, and possessing good practicality and inclusiveness.
[0016] This invention solves the problems of large sound source tracking range, particle degradation, and difficulty in vertical tracking in existing technologies by organically combining a bio-inspired sound separation algorithm with particle filtering. It enables deaf and mute people to accurately locate the speaker in complex environments, provides a stable foundation for sign language-sound source binding and multimodal feedback, and significantly improves the interactive experience and communication efficiency. Attached Figure Description
[0017] To more clearly illustrate the technical solutions in the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 is a flowchart illustrating the BOSSA-based multimodal dynamic response interaction method for deaf and mute individuals provided by the present invention; Figure 2 is a flowchart illustrating the dynamic sound source tracking provided by the present invention; Figure 3 is a structural diagram illustrating the BOSSA-based multimodal dynamic response interaction system for deaf and mute individuals provided by the present invention. Detailed Implementation
[0019] This invention provides a BOSSA-based multimodal dynamic response interaction method for deaf and mute individuals, which solves the technical problem in the prior art where deaf and mute individuals have difficulty accurately locating the speaker in complex environments.
[0020] The technical solution of this invention aims to solve the above-mentioned technical problems. The overall approach is as follows: A multimodal dynamic response interaction method for deaf and mute individuals based on Bossa A, comprising: acquiring multimodal data to be processed and preprocessing it to obtain multimodal data, wherein the multimodal data includes sound data, dynamic sign language data, and environmental data; obtaining a filtered signal based on the sound data and the physiological characteristics of the human ear, wherein the physiological characteristics of the human ear include the high-frequency sensitivity characteristics and low-frequency tolerance characteristics of the cochlear hair cells; obtaining a category determination result of the sound data based on the filtered signal; generating a spatial activation map based on the category determination result; outputting the optimal sound source tracking result based on the spatial activation map; and outputting a multimodal feedback result based on the optimal sound source tracking result and the category determination result.
[0021] To better understand the above technical solutions, the following will provide a detailed explanation of the technical solutions in conjunction with the accompanying drawings and specific implementation methods.
[0022] First, it should be clarified that the term "and / or" in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.
[0023] Explanation of technical terms: BOSSA (Biologically Oriented Sound Segregation Algorithm) is a framework that simulates the neural mechanisms of the human brain to analyze spatial cues of sound and achieve sound source separation and optimization.
[0024] The present invention provides a multimodal dynamic response interaction method for deaf and mute people based on BOSSA, as shown in Figure 1, including steps S11-S16: Step S11, acquire the multimodal data to be processed and preprocess it to obtain multimodal data, wherein the multimodal data includes sound data, dynamic sign language data and environmental data.
[0025] The multimodal data to be processed is acquired and preprocessed to obtain multimodal data, including S111-S1114: S111 is based on a dual-channel omnidirectional MEMS silicon microphone array to acquire sound data.
[0026] Sound data can be acquired through a dual-channel omnidirectional MEMS silicon microphone array deployed in the left and right ear canals, with the microphone spacing fixed at 2cm to simulate the interaural distance of the human ear; the acquisition end is equipped with a preamplifier circuit to amplify μV-level signals to mV-level (amplification gain 40dB) to prevent weak consonant signals in the 2kHz-4kHz range from being drowned out by noise; ultra-high frequency electromagnetic interference is filtered by an RC low-pass filter (cutoff frequency 15kHz), and sampling distortion is eliminated by an 8th-order FIR anti-aliasing filter (cutoff frequency 20kHz); the acquisition parameters are set to a sampling rate of 44.1kHz (covering the 20Hz-20kHz range audible to the human ear), 16-bit bit depth, and a signal-to-noise ratio of ≥60dB.
[0027] Adaptive Wiener filtering is used to remove some steady-state background noise, the audio amplitude is linearly mapped to the [-1, 1] interval, a 1ms precision timestamp is added for precise alignment with other modal data, and standardized dual-channel raw sound data is output.
[0028] The S112 uses a wide-angle camera and a six-axis attitude sensor to collect dynamic sign language data.
[0029] Dynamic sign language data can be collected through the fusion of a wide-angle camera and a six-axis attitude sensor. An external wearable camera (90° field of view, 1280×720 resolution, 25fps frame rate) focuses on the hand area, and the attitude sensor simultaneously captures the dynamics of the hand and the posture of the arm, collecting the three-dimensional coordinates of 21 hand joints and the angle data of the arm, and generating binary signs of hand occlusion or out of frame in real time.
[0030] Accurate separation of the hand region can be achieved by introducing a lightweight U-Net semantic segmentation algorithm. The joint trajectory is extracted based on the MediaPipe optimization model to generate a sign language coordinate sequence. Arm posture feature labels are extracted and associated with occlusion markers to ensure data temporal consistency and correlation, and standardized dynamic sign language data with synchronized timestamps are output.
[0031] S113 is a miniature noise sensor based on microphone array co-integration that collects environmental data, including environmental noise.
[0032] Environmental data can be collected using a miniature noise sensor integrated with a microphone array (sampling rate 5Hz, accuracy ±1dB), and simultaneously collected using a light sensor deployed in conjunction with a sign language camera (light sampling rate 5Hz). Hardware clock synchronization ensures consistency with the audio and sign language data timing. After three rounds of moving average noise reduction to filter out transient fluctuations, noise intensity and light intensity are uniformly mapped to the [0, 1] interval, generating standardized environmental parameters.
[0033] S114 aligns the sound data, dynamic sign language data, and environmental data according to timestamps to obtain multimodal data.
[0034] The collected and processed sound data, dynamic sign language data, and environmental data are aligned according to timestamps.
[0035] Step S12: Obtain a filtered signal based on the sound data and the physiological characteristics of the human ear, including the high-frequency sensitivity and low-frequency tolerance characteristics of the cochlear hair cells.
[0036] The filtered signal is obtained based on the sound data and the physiological characteristics of the human ear, including S121-S123: S121, based on the physiological characteristics of the human ear and the hearing needs of deaf and mute people, a filter bank is designed. The filter bank consists of 55 Gammatone filters and covers the frequency band of 20Hz-20kHz.
[0037] To accurately address the semantic comprehension difficulties faced by deaf-mute individuals with residual hearing due to high-frequency hearing loss, and to adapt to the sound characteristics of different frequency bands and the perceptual differences of deaf-mute individuals, the core parameters of the filters (center frequency, bandwidth, order, and gain) are configured differently according to the functional positioning of each frequency band, as follows: 8 channels are set in the 20Hz-500Hz low-frequency range, with a center frequency interval of 60Hz, a bandwidth of 55Hz, an order of 4, and an initial gain of 0dB; this frequency band is mainly composed of low-frequency noises such as footsteps and refrigerator humming, containing only low-frequency components of vowels such as "o" (with limited impact on semantic comprehension); the fewer number of filters and the fixed narrow bandwidth can reduce noise redundancy while avoiding interference with high-frequency signal processing.
[0038] The 500Hz-2kHz mid-low frequency range has 12 channels, with the center frequency interval reduced to 125Hz to increase density. The bandwidth is 80Hz, the order is 4th, and the initial gain is 0dB. This range is the main distribution frequency band of core vowels such as "a" and "e", which are relatively complete for deaf and mute people. The 12-channel filter can balance signal integrity and computing resources, preserving vowel features (laying the foundation for semantic understanding) without excessive subdivision leading to redundancy.
[0039] The high-frequency core range of 2kHz-4kHz has 20 channels with a center frequency interval of 100Hz, a bandwidth narrowed to 30Hz, an order increased to 6th order, and an initial gain of +2dB. This range is the core distribution area of high-frequency consonants such as "s", "sh", and "f" (key to semantic discrimination), but deaf people have the weakest perception of this frequency band. The combination of multiple channels, narrow bandwidth (improving resolution), high order (enhancing anti-interference), and positive gain (signal pre-enhancement) can specifically solve the pain point of "semantic ambiguity caused by consonant absence".
[0040] Ten channels are set in the mid-to-high frequency range of 4kHz-8kHz, with the center frequency spacing extended to 400Hz to reduce density, bandwidth of 150Hz, order of 4th order, and initial gain of +1dB. This range includes some frequency components of voiceless consonants such as "k" and "t", which supplements the semantic integrity. Low-amplitude positive gain can enhance high-frequency details and reduce information loss while controlling the amount of computation.
[0041] The 8kHz-20kHz ultra-high frequency range has only 5 channels, with a center frequency spacing of 2400Hz, a bandwidth of 300Hz, a fourth-order frequency, and an initial gain of -3dB. This frequency band is dominated by useless noise such as sharp tinnitus and electronic interference, which is meaningless for identifying key sounds. The low-density configuration and negative gain can suppress noise, reduce signal processing redundancy, and reduce potential hearing discomfort for deaf and mute people.
[0042] Time-domain transfer function of Gammatone filter for:
[0043] Where A is the gain (in dB), set according to the initial value mentioned above; The filter order determines the steepness of the frequency response; This is a bandwidth parameter, which is positively correlated with the frequency band bandwidth. The center frequency (in Hz) is set according to the frequency band interval; The initial phase is set to 0. For time, and .
[0044] A high-frequency hearing data input interface is reserved, supporting users to import pure-tone audiometry data, including hearing loss values in core frequency bands such as 2kHz, 3kHz, and 4kHz. (Unit: dB). In the pre-processing stage of filtering, a dynamic mapping between the loss value and gain compensation is established, specifically as follows: If the hearing loss value at a certain frequency is... Then the gain of the corresponding filter is adjusted to be (A coefficient of 0.8 is used to avoid overcompensation that could lead to signal distortion.) For example, if hearing loss is 15dB in the 4kHz band, the filter gain for that band is increased by 12dB. The maximum compensation gain is ≤30dB (to avoid the risk of physical damage) to ensure that the intensity of the high-frequency signal after compensation matches the level of normal human hearing.
[0045] In addition, to accommodate individual perceptual differences among deaf and mute individuals, a "frequency band number adjustment channel" is reserved in the hardware or algorithm layer of the filter bank, supporting dynamic increases or decreases of the number of filters within a single frequency band by ±5 (the total number of channels is controlled between 45 and 65, i.e., 55 ± 10, to avoid exceeding the computing power limit of the wearable device). Specific examples are as follows: If users report weak perception of key sounds such as horns in the 3kHz-4kHz frequency band, this sub-range (the 3kHz-4kHz part of the original 20 channels) can be increased from 10 channels to 15 channels, and the center frequency interval can be reduced to 67Hz (1kHz range is evenly divided into 15 channels), further improving the resolution of this frequency band; in the reverse adjustment scenario, if users are sensitive to ultra-high frequency noise in the 8kHz-20kHz range, this frequency band can be reduced from 5 channels to 3 channels, and the center frequency interval can be expanded to 6000Hz, reducing redundant processing of invalid signals and reducing auditory interference.
[0046] S122 inputs the sound data into the filter bank, and completes frequency decomposition through parallel filtering and envelope extraction to obtain the narrowband output signal corresponding to each Gammatone filter.
[0047] Obtain audio data from the left and right channels ( , The input filters (55 Gammatone filters) are simultaneously processed to reduce signal delay and meet real-time interactive requirements.
[0048] Each Gammatone filter retains only the signal within the corresponding center frequency ± bandwidth range, outputting 55 narrowband signals.
[0049] For example, if the input mixed sound contains dialogue (including 2800Hz "sh") and background music (including 1000Hz low frequency), the 2800Hz filter outputs the narrowband signal corresponding to "sh", and the 1000Hz filter outputs the low frequency signal of the background music, thus achieving frequency layering.
[0050] The signal envelope reflects the intensity variation of sound in the corresponding frequency range. The envelope of dialogue fluctuates periodically, while the envelope of noise changes randomly and irregularly, providing a basis for judgment in subsequent invalid frequency band suppression.
[0051] Based on the Hilbert transform, the signal envelope of each narrowband output signal is extracted, including:
[0052] in, This is the narrowband output signal of a single-channel Gammatone filter. For Hilbert transform, For the actual part, The virtual part, This is the signal envelope of a single-channel Gammatone filter.
[0053] S123 filters and suppresses the 55 narrowband output signals to obtain the filtered signal.
[0054] For low-frequency narrowband output signals in the 20Hz-500Hz range, set the envelope strength threshold. Intensity threshold The signal type can be further subdivided based on one-third of the average envelope intensity in the key frequency band (2kHz-4kHz) and envelope characteristics. Specifically, if the percentage of envelope fluctuations without repeating patterns within 100ms is >85% (e.g., current noise, continuous wind noise), and the intensity remains below [a certain value] for 100ms... If the noise level is low, it is considered invalid noise, and the narrowband output signal gain of that path is reduced to -15dB (deep suppression); if the time interval deviation between two consecutive envelope peaks is <30% (low-frequency repetition, such as footsteps or knocking), even if the intensity is lower than... Only reduce the gain to -5dB (mild suppression, preserving environmental safety perception value); if the envelope variability consistency is >75% within 3 consecutive cycles (periodic variability, such as low-frequency vowels like "o" and "u"), and the intensity is ≥ If the original gain is retained, the basic semantic information of the vowels will be preserved.
[0055] For narrowband output signals in the 8kHz-20kHz ultra-high frequency band, this band is dominated by random noise (such as electronic interference, with a completely disordered envelope), but may contain sudden critical sounds (such as sharp alarms, emergency braking sounds, with a pulse-like periodic envelope). The intensity ratio is set accordingly. Intensity percentage This can be the ratio of the peak intensity of the actual envelope of the narrowband signal to the average envelope intensity of the 2kHz-4kHz key frequency band. It can be specifically processed based on envelope characteristics as follows: If the proportion of envelope fluctuations without repeating patterns within 100ms is >85%, and the intensity proportion... <1 / 2 (e.g., continuous electronic interference), reduce the initial gain from -3dB to -10dB for deep noise suppression; if the time interval deviation of three consecutive envelope peaks is <20%, and the duration of a single peak is <50ms (pulse-like periodicity, such as a sudden alarm), even if the intensity ratio is <20%, the noise will be suppressed. <1 / 2, still retain the original -3dB gain to ensure that no danger signals are missed.
[0056] For narrowband output signals in the 2kHz-4kHz high frequency range, this frequency band is the core area of high-frequency consonants such as "s" and "sh" (the envelope has a short, high-frequency periodic rhythm). The gain can be dynamically optimized according to the ambient noise, as follows: Identify "consonant signal segments" by the envelope rhythm (e.g., if the envelope has more than 3 short peaks within 100ms, it is determined to contain consonants); calculate the associated envelope and signal-to-noise ratio (SNR) for the consonant signal segments to avoid noise interference during non-consonant periods; if the consonant segment SNR < 5dB (e.g., in a noisy restaurant scene), increase the gain by an additional 3dB on top of the base +2dB gain (total gain +5dB) to strengthen the short envelope peaks of the consonants and prevent them from being masked by noise; if the consonant segment SNR > 15dB (e.g., in a quiet home scene), maintain the base gain of +2dB to avoid over-amplification that could distort the consonant envelope (e.g., the long periodic envelope of "sh" is truncated).
[0057] The final result is 55 narrowband output signals (i.e. filtered signals) after screening and suppression, in which the consonant signals (with clear envelope features) in the core frequency band of 2kHz-4kHz are mainly preserved and enhanced.
[0058] Step S13: Obtain the category determination result of the sound data based on the filtered signal.
[0059] The category determination results of the sound data are obtained based on the filtered signal, including S131-S134: S131, key identification features are extracted based on the filtered signal, and feature vectors are constructed.
[0060] Five key identifying features of the sound object are extracted from the filtered signal to form a 5-dimensional feature vector. ,in, For transpose; The center frequency is used to distinguish between high-frequency signals and mid-to-low-frequency signals; This refers to the duration, used to distinguish between continuous and burst signals; It is the spatial azimuth angle, used to distinguish between fixed-direction signals and signals without a fixed direction; This represents the signal energy in the time domain, used to distinguish between strong and weak signals. High-frequency pulse density is used to distinguish key sound objects (such as alarms which have fixed pulses). Stable values) and non-critical sound objects (such as dialogue, noise) (Values are irregular).
[0061] The features are calculated as follows: ,in, It is frequency The corresponding frequency domain signal energy is obtained by decomposing the signal frequency components using Fourier transform; ,in, For the first time the signal envelope is greater than the energy determination threshold And three consecutive frames higher than Take the moment when the threshold is first exceeded; for The signal envelope is first less than The moment; ,in, 1 represents the time difference between the two microphones. The speed of sound is 340 m / s. The spacing between wheat grains is 2cm; ,in, For a moment The corresponding signal envelope energy is calculated using the sliding window smoothing method; ,in, The number of high-frequency pulses within 200ms. For window length, , The window length is designed to ensure real-time capture of pulse characteristics without increasing computational load or device power consumption due to excessive window length.
[0062] When the feature vector Corresponding signal energy satisfy At this time, it is a weak signal (such as a whispered conversation at a distance, or a weak alarm through a wall), among which The average energy of all signal frames within the last 3 seconds, with each frame corresponding to a 10ms audio signal, is the feature vector for weak signals. For smoothing, the formula is:
[0063] in, The smoothed feature vector , It is the average of the feature vectors of the first 3 frames, each frame is 10ms, covering the historical features of the first 30ms; by fusing the features of the current frame and the historical frames, the feature jump caused by energy fluctuations of weak signals is weakened, making the subsequent likelihood value calculation more stable.
[0064] S132, construct a sound object feature library, and train the Gaussian mixture model using the sound object feature library.
[0065] To accurately capture the key sound information needed by deaf and mute individuals in their daily lives and avoid ineffective interference, the sound object feature library is constructed based on five core sound objects: safety warnings, public transportation, social interaction, daily environmental prompts, and environmental noise. Specifically: Safety warnings are emergency alerts directly related to personal safety, including emergency alarms, vehicle warnings, and hazardous environment alerts. Emergency alarms include continuous beeping of fire alarms and pulse alerts from earthquake early warning systems. Vehicle warnings include short car horns and alternating high and low frequency sirens from ambulances, fire trucks, and police cars. Hazardous environment alerts include periodic beeping announcements of "Please be careful" at construction sites and aggressive barking from large dogs. Public transportation focuses on essential travel scenario alerts, including public transportation announcements, road interaction prompts, and hub broadcasts. Public transportation announcements include bus announcements of "XX station has arrived" and the "beep" warning sound of subway doors opening and closing. Road interaction prompts include rapid "beep" sounds from pedestrian green lights at zebra crossings and traffic light switching sounds. Hub announcements include announcements such as "The train bound for XX is about to arrive" at train stations / airports; social interaction announcements focus on targeted sounds related to interpersonal communication, including dialogues and calls, and social scene prompts; among them, dialogues and calls include natural conversations between different genders / ages, and the sound of one's own name being called; social scene prompts include light / heavy / continuous knocking sounds, and electronic / mechanical doorbell sounds; daily environment prompts revolve around equipment / facilities necessary for daily convenience, including home appliance prompts and public facility prompts; among them, home appliance prompts include the "ding" sound of a microwave oven finishing, the whistle of a kettle boiling, and the sound of a washing machine finishing; public facility prompts include the "ding" sound of an elevator arriving at its stop, the sound of a toilet flushing, and the sound of an automatic door sensor motor; environmental noise specifically refers to interference sounds without key informational value, including natural background sounds, household noise, and easily confused interference sounds; among them, natural background sounds include the sound of wind and rain; household noise includes distant unrelated conversations and the sound of home appliances running; easily confused interference sounds include toy alarm sounds.
[0066] Feature vector samples of the above five types of sound objects (covering variations under different distances and noise environments) were collected, and each group of samples was labeled with a category as a training label; The training data for the sound object (which has been filtered and labeled) is denoted as ,in For the first The number of samples in each class All are 5-dimensional feature vectors .
[0067] A Gaussian Mixture Model (GMM) is used to fit the feature distribution patterns of different sound objects. An independent GMM model is trained for each type of sound object, denoted as . ( ), each Include Gaussian components ( (Set separately according to the complexity of category features). The probability density function is:
[0068] in, The input feature vector, For the first The mean vector of a Gaussian distribution, For the first The covariance matrix of a Gaussian distribution; For the first The weights of a Gaussian distribution ; The probability density function of a single Gaussian distribution is expressed by the formula:
[0069] in, For the feature vector dimension, For transpose, , Let be the determinant and inverse of the covariance matrix.
[0070] Training via Expectation Maximization (EM) algorithm The model's parameters enable it to accurately distinguish the categories of sound objects (i.e., the five categories mentioned above).
[0071] Initialize weights mean From training data Random selection 1 sample is used as the initial mean, covariance use Initialize the global covariance matrix; calculate based on the current parameters. Each sample belong Inner The posterior probability of each component as follows:
[0072] in, Indicates the first The class of Each sample belongs to The One portion, For probability; based on ,renew parameters , , Maximize the expected likelihood value.
[0073] Repeat the above calculation steps of "maximizing the expected likelihood value based on the current parameters" when The parameter change is less than the preset threshold (e.g.) At this point, the parameters of the Gaussian mixture model are stable, and training is complete.
[0074] S133: Input the feature vector into the trained Gaussian mixture model, and determine the sound object category of the sound data through the likelihood value.
[0075] Regarding the priority of sound objects: In the process of classifying sound objects, safety warnings and public travel are prioritized as high-priority key sound objects (corresponding to sound object priority 1-2), social interactions and daily environmental prompts are classified as ordinary sound objects (corresponding to sound object priority 3-4), and environmental noise is classified as low-priority interference sound objects (corresponding to sound object priority 5). The core identification features of key sound objects are pre-stored to provide a benchmark for rapid matching and response.
[0076] For example: emergency alarm sound, Corresponding to typical high-frequency pulse frequency bands. It possesses stable duration characteristics, and the average environmental energy is denoted as... , Energy levels are significantly higher than the environmental average. It is characterized by high pulse density and a regular pulse interval of 0.5s; car horns, For sharp, sudden high-frequency bands, Corresponding to the duration of a single short vocalization, It exhibits sudden high-energy characteristics. It features high pulse density and high-frequency repetition with ≥2 consecutive triggers within 1 second; special vehicles sound their horns. It covers alternating high and low frequency bands (such as the alternating "beep-du" sound of an ambulance). To maintain the duration of the warning, It exhibits strong energy penetration. It features medium to high pulse density and a periodic switching pattern of "high frequency - low frequency" (period 1 second); used for bus and subway station announcements. Covering the core frequency bands for clear voice communication. The duration of a single complete sentence's audio. It exhibits a moderate and stable energy state. It features low pulse density and includes pre-stored keywords such as "XX station has arrived" or "Please exit from the rear door"; it serves as a prompt for subway and bus door opening and closing. A short, high-frequency prompt tone. This is the total duration of the warning signal for a single door opening / closing operation. It exhibits moderate energy. It features extremely high pulse density and "closing countdown" accelerated repetition (such as doubling the frequency in the last 3 seconds).
[0077] In real-time acquired environmental signals (specifically, sound data), pure sound can be separated after Gammatone filtering, and feature vectors can be extracted. Adaptively select the feature vector type for different signal intensities, weak signals ( Using smoothed feature vectors To improve stability, ordinary signals ( Directly using the original feature vector To preserve detailed information.
[0078] The basic threshold for determining useful signals can be set as follows: Meanwhile, to avoid weak signals being misjudged as noise due to low energy, the useful signal determination threshold is lowered in weak signal scenarios. To enhance the priority of key sound objects, weight coefficients are set for the GMM model of key sound objects. Normal sound object settings The system calls the GMM models corresponding to the five core sound objects, adapts to the computing power of portable devices, and calculates the feature vectors for each. Likelihood values in each model And correct it to In weak signal scenarios, if the global maximum corrected likelihood value of the 5 types of sound objects is ≥ If the signal is positive, it is considered a valid signal and proceeds to the subsequent sound object classification process (i.e., determining the sound object category); otherwise, it is considered invalid noise and is filtered out directly. In ordinary signal scenarios, if the global maximum corrected likelihood value is ≥ If the signal is positive, it is considered a valid signal and proceeds to the subsequent sound object classification process (i.e., determining the sound object category); otherwise, it is considered invalid noise and is filtered out directly.
[0079] The category corresponding to the global maximum likelihood value can be taken as the candidate classification result (i.e., the category of the sound object). .
[0080] S134. Determine the category determination result of the sound data based on the priority of the sound object, the classification confidence, and the category of the sound object.
[0081] For all class likelihood values, use Function normalization yields classification confidence. (Value range 0) 1) The calculation is as follows:
[0082] Confidence is used to determine candidate classification results The reliability of the information.
[0083] Based on sound object priority and candidate classification results (i.e., the category of the sound object) and classification credibility are used to determine the category judgment result of the sound data, and to perform differential judgment and response triggering, as follows: If Key sound objects (sound object priority 1-2) and their corresponding credibility. Key threshold ,in If the model is validated using a test set of samples (requiring an accuracy rate ≥ 95% and a false positive rate ≤ 3%), it is determined to be a "clearly identified key sound object." Through optimization methods such as simplifying the computational logic and reusing pre-stored parameters, the real-time signal category modeling can be completed within 5ms, triggering the highest priority multimodal warning. For key sound objects (sound object priority 1-2) but (lower than) If the sound is identified as a "suspected key sound object," a medium-priority alert will be triggered, and the suspected category will be marked; if It is a regular sound object (sound object priority 3-4) and has a corresponding credibility level. ordinary threshold ,in The model is validated using a test set of samples (requiring both an effective classification accuracy of ≥80% and a low-reliability classification filtering rate of ≥90% to balance the identification of useful information and the control of invalid interference). Dialogue triggers regular multimodal feedback, and daily environmental prompts output corresponding feedback according to preset rules.
[0084] like For ambient noise (sound object priority 5) or Noise filtering threshold (Low-confidence classification is considered an uncertain result), which does not trigger additional modal feedback and avoids invalid interference; among which Verification using a test set confirmed that the low-intensity noise filtration rate is ≥95%.
[0085] In addition, individual hearing profiles can be established for deaf and mute individuals to achieve personalized gain compensation and accurate modeling of weak signals, ensuring that different individuals have consistent recognition accuracy and perceptual experience for various sound objects.
[0086] Specifically, individual hearing profiles record personalized data such as the user's sensitivity thresholds to various sound objects, weak signal recognition preferences, and signal loss at different frequencies. This data directly guides the cochlear filter unit's weak signal gain compensation and signal-to-noise ratio (SNR) threshold adjustment to suit the perceptual needs of different users. The adjustment rule is that the more severe the signal loss at a frequency, the lower the SNR threshold for fundamental tone extraction. For example, when the signal loss in the 4kHz band is 15dB, the SNR threshold is lowered from the usual 5dB to 2dB to ensure that key features such as weak consonants are not lost.
[0087] Based on the personalized preprocessed signal and combined with the weak signal identification preferences in the individual's hearing profile, precise modeling and optimization of weak signals are carried out. The weak signal type is determined by integrating three key features: pitch characteristics, signal duration, and spatial location, with the judgment criteria adapted to the user's personalized needs. For example, for users sensitive to high-frequency perception, the pitch characteristic intensity threshold can be relaxed; for users who need to capture short conversations, the duration threshold can be lowered from 3 seconds to 2 seconds. Even if the pitch characteristic is weak, if the signal meets the individualized duration threshold and the spatial location is fixed (such as someone speaking across from you), it is still judged as "weak conversation sound" and included in the modeling scope, avoiding misjudgment as invalid noise due to general screening rules.
[0088] Step S14: Generate a spatial activation map based on the category determination results.
[0089] As shown in Figure 2, a spatial activation map is generated based on the category determination result, including S141-S144: S141, based on the bio-inspired spatial tuning neuron cluster architecture, extracts binaural cues from each frequency channel in the Gammatone filter.
[0090] Spatial tuning neurons (STNs) are neuronal models that simulate the human auditory system's sensitivity to sounds from specific spatial orientations. A bio-inspired STN cluster is constructed, with each STN corresponding to a unique horizontal azimuth angle. (0°) (360°), for example, 360 STNs, with one STN corresponding to each 1°, or 64 STNs evenly covering 0°. 360° resolution is optimized through interpolation; horizontal target matching parameters are configured for each STN, including the interaural timing difference (ITD) and interaural level difference (ILD), where the target ITD is denoted as... The target ILD is denoted as The parameter values can be calibrated based on measured head-related transfer function (HRTF) data to ensure that they match the acoustic characteristics of the corresponding azimuth angle, providing a benchmark for subsequent response intensity calculations.
[0091] Based on the left and right channel signals of each frequency band, the actual ITD and ILD (i.e., binaural cues) corresponding to each frequency band can be extracted; the constructed horizontal STN cluster can be invoked, and the horizontal azimuth angle bound to each STN can be determined. First, the actual parameters and the target parameters are calculated separately using an exponential function. =Horizontal ITD, The matching degree of ITD (Level ILD) is as follows: ITD matching degree =
[0092] ILD matching degree =
[0093] in, =0.1ms =1dB is the optimal parameter calibrated using HRTF measured data. The matching degree of ITD and ILD is designed based on the nonlinear tuning characteristics of the human auditory system, taking into account both positioning accuracy and anti-interference ability.
[0094] S142, based on the nonlinear multiplication mechanism, fuses the binaural cues to obtain the spatial orientation response intensity of the STN cluster.
[0095] The response intensity of the STN in that frequency band is obtained by performing a nonlinear multiplicative combination of the two types of matching degrees in the same frequency band (i.e., the spatial azimuth response intensity of the STN cluster) as follows: Response intensity = ITD matching degree × ILD matching degree; if multiple frequency bands are involved, the response intensity of all frequency bands is averaged and normalized to finally obtain the spatial azimuth response intensity of the STN (range 0). (1, 1 represents a perfect match), to complete the preliminary response strength calculation for all STNs (i.e., STN clusters).
[0096] S143 optimizes the spatial orientation response intensity of the STN cluster based on sound object priority, classification confidence, and category determination results.
[0097] After the initial calculation of the spatial orientation response intensity of the STN cluster is completed, the "sound object priority, classification confidence level, and category determination result" are received and output; for those with a classification confidence level ≥ the corresponding priority threshold (key sound objects corresponding to...). Corresponding to ordinary sound objects The sound data (priority 1-4) is enhanced in a priority-oriented manner, with weighted enhancement only applied to associated STN clusters within ±30° of the response peak location. The upper limit of the weighted response intensity is clamped to 1.0 to avoid numerical overflow. Irrelevant noise (sound data with priority 5 classified as environmental noise, and classification confidence < 0.0) is not considered. (Low-reliability, uncertain sound data) triggers dual suppression.
[0098] First pass Filter low-intensity responses (set to 0 directly if below the threshold); suppress strong interference noise above the threshold with an attenuation coefficient of 0.3; if irrelevant noise exists in the sound data of the same spatial location, force the response intensity of the irrelevant noise to be reduced to less than 10% of the response intensity of the (non-noise) sound data to avoid noise masking useful sound information; perform optimization on the spatial location response intensity of the STN cluster. Interval normalization ensures the consistency and validity of response data.
[0099] S144, based on the optimized STN cluster spatial orientation response intensity, generates a spatial activation map through an inhibitory side-connection network.
[0100] Based on the optimized spatial orientation response intensity of the STN cluster, a global competitive suppression of the response intensity of all STNs is performed through an inhibitory side-connection network. This further enhances the response peak of the STN in the target direction and deeply suppresses the sidelobe response of the STN in the non-target direction (avoiding "spatial leakage"). Finally, the final response intensity of each STN (range 0) is obtained. 1, 1 means a perfect match.
[0101] Based on biological auditory mechanisms, "horizontal azimuth angle" is generated. A two-dimensional spatial activation diagram of “× response intensity”.
[0102] With horizontal azimuth (0°) (360°) is the horizontal axis, and the final STN response intensity (0) is the vertical axis. 1) Using the vertical axis, map the "azimuth-response intensity" data one by one to form a two-dimensional matrix, then convert it into a recognizable standardized format (such as JSON), containing a "horizontal azimuth coordinate array" and a "corresponding response intensity array"; to balance basic use and full-space perception needs, a vertical elevation angle is added to the horizontal STN cluster. (-90°) (90°), construct a three-dimensional spatial coordinate system For each coordinate Configure vertical target parameters ( =Vertical ITD, =Vertical ILD), and horizontal parameter ( This forms a three-dimensional matching parameter set; each STN corresponds to a unique parameter set. Orientation, generating " × The spatial activation map (e.g., a 360×180×1 matrix) with "× response intensity" is directly output after standardization; the two-dimensional or three-dimensional spatial activation map is output in real time at a frequency of 10ms / frame, providing observational basis for particle initialization and weight update.
[0103] Step S15: Output the optimal sound source tracking result based on the spatial activation map.
[0104] The optimal sound source tracking result is output based on the spatial activation map, including S151-S154: S151, extract the horizontal azimuth angle corresponding to the peak in the spatial activation map, and generate the corresponding initial particle set.
[0105] Receive two-dimensional spatial activation map (including the horizontal azimuth of the sound source) (Response intensity), extract the horizontal azimuth angle corresponding to the peak value of the activation map. As the initial localization reference for the sound source; around An initial particle set is generated, and the number of particles in the initial particle set is dynamically adjusted based on the peak intensity of the activation map. 300 particles are generated when the peak intensity is ≥0.8, and 600 particles are generated when the peak intensity is <0.5, achieving a balance between tracking accuracy and computational power consumption. The particle positions are randomly distributed within ±15° around the peak intensity according to a Gaussian distribution to avoid blind searching. Each particle carries a two-dimensional state vector. ,in The horizontal azimuth angle (0°) corresponding to the particle 360° The corresponding angular velocity of the particle (° / frame) supports horizontal two-dimensional sound source tracking.
[0106] S152 updates the weights of each particle in the initial particle set based on the spatial activation graph.
[0107] The matching degree between each particle and the spatial activation map is calculated as a weight. The weight calculation formula is as follows:
[0108] in, The weight of the particle; This is the adjustment coefficient (set to 0.1). The current azimuth angle of the particle The corresponding activation map response intensity, Set the peak azimuth angle of the activation map for the current frame; normalize the weights of all particles to the interval [0, 1] to ensure that the total weights are 1.
[0109] In addition, a roulette wheel resampling method is used to retain the top 80% of particles by weight, eliminate low-weight particles, and add a small number of new particles (in... Randomly generated within a range of ±15°, with a total quantity maintained at 300-600, to avoid particle degradation.
[0110] S153, determine the appropriate motion model based on the motion state of the sound source.
[0111] Based on the peak azimuth angle of the spatial activation map of 5 consecutive frames Calculate the position change between adjacent frames If the change in position is stable ( (≤3° / frame), enabling the uniform motion model, predicting the current frame state based on the particle state vector of the previous frame, calculated as follows:
[0112]
[0113] in, The horizontal azimuth angle of the particles in the current frame. This refers to the horizontal azimuth angle of the particles in the previous frame. The horizontal angular velocity of the particle in the current frame. The horizontal angular velocity of the particles in the previous frame. The frame interval is fixed at 10ms. To simulate the uncertainty of sound source motion using small random noise (±0.5° / frame), and to avoid slow particle set convergence; if the variation fluctuates greatly ( >3° / frame), with uniform acceleration model enabled, the calculation is as follows:
[0114]
[0115]
[0116] in, The horizontal angular acceleration of the particles in the current frame. This represents the horizontal angular acceleration of the particles in the previous frame. The angular acceleration is random noise (±0.2° / frame²), and the meanings of the other parameters are the same as those of the uniform velocity model.
[0117] S154, under the adapted motion model, obtain the optimal sound source tracking result in two-dimensional space, or, under the adapted motion model, increase the state vector of the particle and the observation dimension to obtain the optimal sound source tracking result in three-dimensional space.
[0118] [Obtaining the optimal sound source tracking result in two-dimensional space] When the sound data is obscured by strong noise (no obvious peak in the spatial activation map), the observation data update is paused, and the particle state is predicted solely by the motion model. The tracking result is continuously output until the spatial activation map reappears with a peak, at which point the weight update is resumed. The sound source tracking position, i.e., the horizontal azimuth angle, of the current frame is calculated by weighted averaging of all particle states. With horizontal angular velocity The update rate for this result is 10 milliseconds per frame, and the output... This represents the optimal sound source tracking result in two-dimensional space at the current moment.
[0119] [Obtain the optimal sound source tracking result in three-dimensional space] Receive spatial activation map (including horizontal azimuth angle) Vertical elevation angle (Activation intensity), extracting the peak position of the activation map. As the initial location coordinates of the sound source; surrounding An initial particle set is generated within a horizontal ±15° and vertical ±10° range. Each particle carries three-dimensional spatial positioning and two-dimensional sound source motion state information. ,in The horizontal angular velocity, The vertical angular velocity; based on the adapted motion model, predict the state of each particle at the next moment. .
[0120] Particle weights and particle positions relative to the peak position of the activation map The Euclidean distance is negatively correlated with the activation intensity at the corresponding location, and the weighting formula is as follows:
[0121] in, For the first The weight of each particle; For adjustment coefficients, For particle position The corresponding spatial activation map activation intensity; based on the weighted average of particle states, the three-dimensional position estimate of the sound source in the current frame is calculated. and its horizontal and vertical speeds This enables full-space three-dimensional sound source tracking. The final output state vector This is the optimal sound source tracking result in three-dimensional space at the current moment.
[0122] Step S16: Output multimodal feedback results based on the optimal sound source tracking results and category determination results.
[0123] Specifically, this includes: outputting multimodal feedback that integrates visual and tactile feedback based on the optimal sound source tracking results and category determination results.
[0124] Visual feedback conveys information through a comprehensive multi-dimensional feature set. Subtitle content matches the semantics of the audio object (e.g., real-time dialogue transcription, alarm display "Emergency Alarm"), and color is strongly correlated with priority (red = priority 1-2, yellow = suspected key, white = priority 3-4). Dynamically, the flashing frequency of borders or indicators increases with the level of urgency. Simultaneously, directional arrows are superimposed on the screen, appearing as horizontal fan-shaped guidance in 2D tracking scenarios and providing three-dimensional guidance including height information in 3D scenarios. Tactile feedback quantifies information through the intensity and orientation of directional vibrations. Strong vibrations (amplitude ≥ 0.5mm) correspond to clearly defined key audio objects, weak vibrations (amplitude ≤ 0.2mm) correspond to ordinary audio objects, and moderate vibrations (amplitude 0.3-0.4mm) correspond to suspected key audio objects. Combining optimal sound source tracking results, it outputs comprehensive and precise tactile cues from four-way directional vibrations (2D) to vibrations including vertical vibrations (3D). Based on standardized environmental data (noise intensity, light intensity) and binary markers indicating hand occlusion or out-of-frame movement, the proportion of visual and tactile feedback is dynamically adjusted. When the noise intensity is greater than 60dB, visual feedback is enhanced, such as increasing the interface brightness by 30% and enlarging the interactive font by 20%. When the lighting is too dim (the threshold can be set according to the actual scene), or when the hand occludes or the proportion of the text outside the frame exceeds 30%, tactile feedback is enhanced, such as increasing the intensity of tactile vibration by 20% and extending the feedback duration by 1 second. When encountering both noise and occlusion interference, redundant feedback such as highlighting and flashing, centering the text, and strong tactile vibration is triggered to ensure the effectiveness of information transmission in complex interaction scenarios.
[0125] By combining the optimal sound source tracking results with the category determination results, the tracking strategy is dynamically adjusted and an adaptation warning is triggered. Specifically, when a clearly defined key sound object is detected (priority 1-2 and classification confidence ≥ 1), the tracking strategy is dynamically adjusted and an adaptation warning is triggered. When a sound source is detected (such as an alarm or car horn), it is immediately set to the highest tracking priority, interrupting the tracking of current non-critical sound objects (priority 3-5) and switching to dedicated directional tracking of that sound source. Strong feedback is activated to ensure rapid delivery of emergency information. For suspected critical sound objects (priority 1-2 but classification confidence < 0.05), the tracking is stopped. If the target is a sound object, a medium-priority alert will be activated. While maintaining the existing tracking task, the system will monitor changes in its characteristics (such as energy intensity and pulse pattern) in parallel and provide a medium-level alert to avoid tracking switching interference or omission of key information due to misjudgment. For ordinary sound objects (priority 3-4 and classification confidence ≥ 100%), a medium-priority alert will be activated. For sounds like dialogue or knocking, it will trigger regular feedback matching its type, such as dialogue transcription or icon hints, without interfering with other tracking tasks; for ambient noise (priority 5) or low-confidence uncertain sound objects (classification confidence < This does not trigger redundant feedback, but only maintains basic background monitoring, reducing invalid interference.
[0126] Based on the optimal sound source tracking result and the category determination result, output multimodal feedback result, including: based on the spatiotemporal correlation verification result, determine the sound source that meets the correlation condition as the interactive correlation sound source, and bind it with the corresponding dynamic sign language in the verification; optimize the interactive experience based on the optimal sound source tracking result, user habits and scene characteristics.
[0128] The system can perform spatiotemporal correlation verification between synchronized sound source coordinates and sign language coordinates. When the spatial distance between the two is less than 0.5m and the timestamp difference is less than 100ms, it is determined to be an "interactively associated sound source" and its priority is automatically increased by 1 level (the maximum priority is 1). After binding, the system focuses on strengthening the multimodal feedback of associated sound sources, such as prioritizing the display of subtitles, highlighting directional arrows, and delaying human voice dialogue subtitles by ≤100ms. At the same time, the feedback intensity of non-associated sound sources is weakened. If the sign language action stops for more than 3 seconds or the spatial distance between the two is greater than 1m, the binding is automatically debound and the original priority of the sound source is restored, thus improving the relevance of the feedback.
[0129] In multi-speaker scenarios, each speaker is assigned a unique visual identifier for intuitive differentiation. For example, speaker 1 is matched with a blue #28 subtitle and a blue-clad male 3D sign language character A, while speaker 2 is matched with a green #28 subtitle and a green-clad female 3D sign language character B. The sign language character's movements are synchronized with the speaker's speech rate in real time to enhance identity recognition. All speaker tags are equipped with weak vibration feedback to ensure accurate interaction.
[0130] Multimodal feedback is used to transmit the sound source tracking status in real time, allowing users to intuitively perceive the system's operating status. Specifically: when the sound source is stably tracked (successfully matched for 3 consecutive frames), the wearable device vibrates slightly once and displays the "Sound locked" subtitle, with no additional redundant feedback to avoid interference; when the sound source is masked or not matched (matching fails for 2 consecutive frames), the wearable device vibrates twice consecutively and displays the light gray subtitle "Relocating sound," clearly indicating that the system is in normal working condition; after relocking the sound source, the wearable device vibrates slightly once and the subtitle flashes twice, indicating "Tracking restored," enhancing the user's perception of system status switching to avoid misjudging device malfunction.
[0131] By learning user behavior preferences and scenario data, the interaction strategy is dynamically optimized, specifically as follows: A built-in common sound source memory function automatically stores feature labels and user history feedback configurations (such as default subtitle size for dialogues and default vibration intensity for knocking sounds) for 50 high-frequency target sound sources (e.g., family conversations, common appliance notifications). These configurations are directly called upon during subsequent recognition, reducing repetitive operations. User interaction data (e.g., manual adjustment of feedback intensity) is collected to train a preference model, adaptively adjusting visual subtitle size and tactile vibration parameters, and dynamically adjusting various thresholds and response intensity weighting coefficients to improve classification and localization accuracy. The usage scenario is determined by combining posture and environmental data. The system enhances tactile feedback and simplifies visual presentation when walking, and enhances visual presentation and reduces tactile interference when stationary, continuously improving the personalized experience. Combining posture data (acceleration, angular velocity) and environmental data (light intensity, background noise level) collected by the wearable device's built-in sensors, it automatically determines the usage scenario: in walking or noisy environments, it enhances tactile feedback (increases vibration intensity by 20%) and simplifies visual presentation (retaining only the core arrows and icons, hiding redundant text); in stationary or quiet environments, it enhances visual feedback (increases subtitle brightness by 30%) and reduces tactile interference (reduces vibration intensity by 30%), achieving dynamic adaptation to the scenario-based experience.
[0132] In summary, this invention, based on the BOSAS framework, effectively enhances the naturalness and real-time performance of information interaction by fusing sound, dynamic sign language, and environmental data. By simulating the physiological characteristics of human cochlear hair cells—sensitivity to high frequencies and tolerance to low frequencies—this invention performs biomimetic filtering on sound signals to improve sound source recognition accuracy. It utilizes 55 narrowband signals to accurately determine sound categories and generates spatial activation maps for high-precision sound source tracking. Combining sound source location and category information, it can output multimodal feedback (such as visual or vibration cues) adapted to the perceptual characteristics of deaf and mute users, significantly enhancing their ability to perceive the surrounding acoustic environment, demonstrating good practicality and inclusiveness.
[0133] This invention solves the problems of large sound source tracking range, particle degradation, and difficulty in vertical tracking in existing technologies by organically combining a bio-inspired sound separation algorithm with particle filtering. It enables deaf and mute people to accurately locate the speaker in complex environments, provides a stable foundation for sign language-sound source binding and multimodal feedback, and significantly improves the interactive experience and communication efficiency.
[0134] Based on the same inventive concept, this invention provides a BOSSA-based multimodal dynamic response interactive system for deaf and mute individuals, as shown in Figure 3. The system includes: a multimodal data acquisition and preprocessing module for acquiring and preprocessing multimodal data to obtain multimodal data, which includes sound data, dynamic sign language data, and environmental data; a BOSSA-customized sound processing module, comprising a cochlear filtering unit, a sound object modeling unit, and a spatial tuning and localization unit. The cochlear filtering unit obtains a filtered signal based on the sound data and the physiological characteristics of the human ear; the sound object modeling unit obtains a category determination result for the sound data based on the filtered signal; the spatial tuning and localization unit generates a spatial activation map based on the category determination result; a dynamic sound source tracking module outputs the optimal sound source tracking result based on the spatial activation map; and a multimodal intelligent interactive control module outputs multimodal feedback results based on the optimal sound source tracking result and the category determination result.
[0135] Based on the same inventive concept, the present invention provides an electronic device, including: a memory and a processor, wherein the memory is used to store a computer program, and the processor runs the computer program to enable the electronic device to perform the BOSSA-based multimodal dynamic response interaction method for deaf and mute people as provided above.
[0136] Based on the same inventive concept, the present invention provides a non-transitory computer-readable storage medium, which, when the instructions in the storage medium are executed by the processor of an electronic device, enables the electronic device to execute the BOSSA-based multimodal dynamic response interaction method for deaf and mute users as described above.
[0137] Since the electronic device described in this embodiment is an electronic device used to implement the information processing method in the embodiments of the present invention, those skilled in the art can understand the specific implementation methods and various variations of the electronic device in this embodiment based on the information processing method described in the embodiments of the present invention. Therefore, how the electronic device implements the method in the embodiments of the present invention will not be described in detail here. Any electronic device used by those skilled in the art to implement the information processing method in the embodiments of the present invention falls within the scope of protection of the present invention.
[0138] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0139] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, create means for implementing the functions specified in one or more blocks of the flowchart illustrations and / or one or more blocks of the block diagrams.
[0140] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means that implement the functions specified in one or more flowcharts and / or one or more block diagrams.
[0141] These computer program instructions may also be loaded onto a computer or other programmable data processing apparatus to cause a series of operational steps to be performed on the computer or other programmable apparatus to produce a computer-implemented process, such that the instructions, which execute on the computer or other programmable apparatus, provide steps for implementing the functions specified in one or more flowcharts and / or one or more block diagrams.
[0142] Although preferred embodiments of the invention have been described, those skilled in the art, upon learning the basic inventive concept, can make other changes and modifications to these embodiments. Therefore, the appended claims are intended to be interpreted as including both the preferred embodiments and all changes and modifications falling within the scope of the invention.
[0143] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of the claims of this invention and their equivalents, this invention also intends to include these modifications and variations.
Claims
1. A multimodal dynamic response interaction method for deaf and mute people based on Bossa, characterized in that, include: The process involves acquiring and preprocessing multimodal data to obtain multimodal data, which includes sound data, dynamic sign language data, and environmental data. A filtered signal is then obtained based on the sound data and the physiological characteristics of the human ear, including the high-frequency sensitivity and low-frequency tolerance characteristics of the cochlear hair cells. The category determination result of the sound data is obtained based on the filtered signal; A spatial activation map is generated based on the category determination result; the optimal sound source tracking result is output based on the spatial activation map; and a multimodal feedback result is output based on the optimal sound source tracking result and the category determination result.
2. The BOSSA-based multimodal dynamic response interaction method for deaf and mute individuals as described in claim 1, characterized in that, The process involves obtaining a filtered signal based on sound data and the physiological characteristics of the human ear, including: designing a filter bank based on the physiological characteristics of the human ear and the hearing needs of deaf people; the filter bank consists of 55 Gammatone filters covering a frequency band of 20Hz-20kHz; inputting sound data into the filter bank; performing frequency decomposition through parallel filtering and envelope extraction to obtain the narrowband output signal corresponding to each Gammatone filter; and filtering and suppressing the 55 narrowband output signals to obtain the filtered signal.
3. The BOSSA-based multimodal dynamic response interaction method for deaf and mute individuals as described in claim 2, characterized in that, Obtaining the category determination result of the sound data based on the filtered signal includes: extracting key identifier features based on the filtered signal and constructing a feature vector; constructing a sound object feature library and training a Gaussian mixture model with the sound object feature library; inputting the feature vector into the trained Gaussian mixture model and determining the sound object category of the sound data through the likelihood value; and determining the category determination result of the sound data based on the sound object priority, classification confidence, and sound object category.
4. The BOSSA-based multimodal dynamic response interaction method for deaf and mute individuals as described in claim 3, characterized in that, The process of generating a spatial activation map based on the category determination results includes: extracting binaural cues from each frequency channel of the Gammatone filter based on a bio-inspired spatially tuned neuron cluster architecture; fusing the binaural cues using a nonlinear multiplication mechanism to obtain the spatial orientation response intensity of the STN cluster; optimizing the spatial orientation response intensity of the STN cluster based on sound object priority, classification confidence, and category determination results; and generating a spatial activation map based on the optimized spatial orientation response intensity of the STN cluster through an inhibitory side connection network.
5. The BOSSA-based multimodal dynamic response interaction method for deaf and mute individuals as described in claim 4, characterized in that, The optimal sound source tracking result is output based on the spatial activation map, including: extracting the horizontal azimuth angle corresponding to the peak value in the spatial activation map and generating the corresponding initial particle set; updating the weight of each particle in the initial particle set based on the spatial activation map as the observation basis; determining the adaptive motion model according to the motion state of the sound source; obtaining the optimal sound source tracking result in two-dimensional space under the adaptive motion model, or, under the adaptive motion model, increasing the state vector of the particles and the observation dimension to obtain the optimal sound source tracking result in three-dimensional space.
6. The BOSSA-based multimodal dynamic response interaction method for deaf and mute individuals as described in claim 5, characterized in that, Based on the optimal sound source tracking result and the category determination result, output multimodal feedback result, including: based on the spatiotemporal correlation verification result, determine the sound source that meets the correlation condition as the interactive correlation sound source, and bind it with the corresponding dynamic sign language in the verification; optimize the interactive experience based on the optimal sound source tracking result, user habits and scene characteristics.
7. A multimodal dynamic response interaction system for deaf and mute people based on Bossa, characterized in that, The BOSSA-based multimodal dynamic response interaction method for deaf and mute individuals, applicable to any one of claims 1-6, comprises: a multimodal data acquisition and preprocessing module, used to acquire and preprocess multimodal data to be processed, thereby obtaining multimodal data, wherein the multimodal data includes sound data, dynamic sign language data, and environmental data; a BOSSA-customized sound processing module, including a cochlear filtering unit, a sound object modeling unit, and a spatial tuning and positioning unit, wherein the cochlear filtering unit is used to obtain a filtered signal based on the sound data and the physiological characteristics of the human ear; the sound object modeling unit is used to obtain a category determination result of the sound data based on the filtered signal; the spatial tuning and positioning unit is used to generate a spatial activation map based on the category determination result; a dynamic sound source tracking module, used to output the optimal sound source tracking result based on the spatial activation map; and a multimodal intelligent interaction control module, used to output a multimodal feedback result based on the optimal sound source tracking result and the category determination result.
8. An electronic device, characterized in that, include: The device includes a memory and a processor, wherein the memory stores a computer program and the processor runs the computer program to enable the electronic device to perform the BOSSA-based multimodal dynamic response interaction method for deaf and mute users as described in any one of claims 1-6.
9. A non-transitory computer-readable storage medium, wherein instructions in the storage medium, when executed by a processor of an electronic device, enable the electronic device to perform the BOSSA-based multimodal dynamic response interaction method for deaf and mute individuals as described in any one of claims 1 to 6.