A training method and system for an panoramic sound spatial sound recognition model
By simulating multiple scenarios in a spatial classroom, collecting and reconstructing sound information, combining multi-layer networks and head-related transfer functions, and monitoring behavioral responses, the problem of low sound recognition accuracy in complex environments for visually impaired people is solved, and their ability to adapt to life is improved.
Patent Information
- Application Number
- CN202510175276.1
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-18
- Publication Date
- 2025-09-19
- Estimated Expiration
- 2045-02-18
AI Technical Summary
Existing panoramic sound recognition technology has extremely complex background noise, where multiple sound sources overlap or the ambient noise interference is too strong. This reduces the sound recognition accuracy of visually impaired people in complex environments, affecting their quality of life and adaptability.
Build a spatial classroom, collect sound direction, position and spectrum information through multi-scenario simulation exercises, construct a panoramic sound recognition model, combine multi-layer network structure and head-related transfer function for audio reconstruction, monitor behavioral responses, calculate environmental sensitivity and adjust training.
It improves the sound recognition accuracy and spatial perception ability of visually impaired people in complex environments, and enhances their adaptability and independence in daily life.
Smart Images

Figure CN120032644B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of hearing aid training, and in particular to a training method and system for an panoramic sound space sound recognition model. Background Art
[0002] For people with visual impairments, sound is an important way to communicate with the outside world. Through sound, they can identify key environmental information, such as traffic signs, speech, and warning sounds. This sound perception not only aids in daily navigation but also helps them make safe decisions and avoid potential dangers.
[0003] Traditional assistance methods for people with visual impairments include Braille, voice guidance, and assistive devices. Braille, as a tactile method, helps blind people read and write. Voice guidance technology uses smart devices to provide users with directions, locations, and environmental information, helping them to travel independently. Assistive devices, such as guide dogs and canes, use touch and hearing to guide people around obstacles. While these methods can effectively compensate for vision loss, in complex or dynamic environments, blind people may face insufficient information or untimely environmental feedback, impacting their quality of life.
[0004] Currently, panoramic sound recognition technology has been applied to assistance systems for people with disabilities. Combined with multi-microphone arrays and sound source localization technology, the system can identify and enhance specific sounds, such as human speech or warning signals, in complex environments, reducing the interference of ambient noise. In addition, deep learning technology is used for sound classification and noise cancellation, improving recognition accuracy. However, while these technologies can enhance the sound recognition ability of people with disabilities in noisy environments, they still face the problem of multiple overlapping sound sources or excessive ambient noise interference in extremely complex background noise, resulting in reduced recognition accuracy, which in turn makes it impossible to improve the adaptability of people with disabilities in society.
[0005] To this end, a training method and system for an panoramic sound spatial sound recognition model are proposed. Summary of the Invention
[0006] The present invention aims to provide a training method and system for an Atmos spatial sound recognition model, designed to improve the comprehensive adaptability of individuals with disabilities in daily life and learning. First, a spatial classroom is constructed, which includes multiple scenarios encountered by individuals with visual impairments in real life. Simulation drills are conducted with individuals in the classroom to obtain simulation drill data, which is then preprocessed. Second, an Atmos spatial sound recognition model is constructed, which includes an audio information acquisition process and a spatial audio reconstruction process. The audio information acquisition process generates sound direction information, sound source location information, and sound spectrum information from the simulation drill data. The spatial audio reconstruction process reconstructs audio based on this sound direction, sound source location, and sound spectrum information to obtain a scene audio generation result. Then, in the classroom, individuals with disabilities respond to the scene audio generation result. The behavioral response includes the individual's reaction time and behavioral actions to the scene audio generation result. Based on the behavioral response, environmental sensitivity is calculated. Finally, an environmental sensitivity threshold is set. If the environmental sensitivity threshold is not greater than the environmental sensitivity threshold, the corresponding scene is re-simulated in the classroom, and the Atmos spatial sound recognition model is retrained.
[0007] To achieve the above object, the present invention provides the following technical solutions:
[0008] A method and system for training an ambient sound spatial sound recognition model, comprising:
[0009] A space classroom is constructed, wherein the space classroom includes multiple scenarios that visually impaired persons encounter in real life; a simulation exercise is conducted on the disabled persons in the space classroom to obtain simulation exercise data; and data preprocessing is performed on the simulation exercise data;
[0010] Constructing an ambient sound spatial sound recognition model, including an audio information collection process and a spatial audio reconstruction process; the audio information collection process generates sound direction information, sound source location information, and sound spectrum information in the simulation rehearsal data; the spatial audio reconstruction process performs audio reconstruction based on the sound direction information, the sound source location information, and the sound spectrum information to obtain a scene audio generation result;
[0011] In the spatial classroom, the disabled person makes a behavioral response according to the scene audio generation result; the behavioral response includes the disabled person's reaction time and behavioral action to the scene audio generation result; and based on the behavioral response, the environmental sensitivity is calculated;
[0012] An environmental sensitivity threshold is set. If the environmental sensitivity threshold is not greater than the environmental sensitivity threshold, the corresponding scene is re-simulated in the spatial classroom and the panoramic sound recognition model is retrained.
[0013] Furthermore, the multiple scenarios are simulated using venues, props and sound equipment; the simulation rehearsal data is the environmental sound that can be received by the human ear and is formed by various sounds emitted in the multiple scenarios, including the audio information required by the disabled person and unnecessary noise.
[0014] Furthermore, the audio information collection process generates the sound direction information, sound source location information and sound spectrum information in the simulation exercise data, and the specific steps include:
[0015] Preprocessing the simulation exercise data;
[0016] The simulation exercise data is input into a direction timing extraction path, a position timing extraction path, and a spectrum timing extraction path respectively; the direction timing extraction path includes: using multiple sensors to receive sound signals and record signal arrival times, calculating the time difference of the signal arrival times, and using the position difference of the multiple sensors to calculate the initial direction of the sound; based on the initial direction, simulating the direction change of the sound source through a rotation matrix to obtain the direction angle that changes with time, and summing the direction angles to obtain the sound direction information;
[0017] The position time series extraction path includes: using the propagation time of the sound signals received by the multiple sensors to calculate the three-dimensional position coordinates of the sound source using multi-point positioning technology; based on the three-dimensional position coordinates, using a second-order dynamic equation to generate the sound source position information;
[0018] The spectrum time series extraction path includes: designing an adaptive filter to automatically adjust the bandwidth according to the intensity and spectral characteristics of the sound signal to capture the frequency information in the sound signal in real time; using a high-order spectrum analysis method to capture the dynamic changes in the frequency information and generate the sound spectrum information.
[0019] Furthermore, the spatial audio reconstruction process performs audio reconstruction based on the sound direction information, the sound source position information, and the sound spectrum information to obtain a scene audio generation result, and the specific steps include:
[0020] Inputting the sound direction information, the sound source location information and the sound spectrum information into a multi-layer network structure to extract audio features;
[0021] Based on the audio features, synthesizing the sound using a head-related transfer function to obtain reconstructed audio;
[0022] Audio rendering and binaural rendering are performed on the reconstructed audio to obtain the scene audio generation result.
[0023] Furthermore, the method for monitoring the disabled person's reaction time and behavioral actions to the scene audio generation result is as follows: the reaction time is obtained by monitoring the disabled person's heartbeat and muscle activity changes, and the formula is: reaction time = reaction start time - sound trigger time; using the disabled person's muscle activity changes to obtain the disabled person's behavioral actions, including two aspects: first, determining whether the disabled person's action corresponds to the sound source; second, determining the disabled person's action amplitude;
[0024] Based on the behavioral response, the environmental sensitivity is calculated, and the specific process is as follows:
[0025] Assigning a weight factor to each of the behavioral actions to indicate its contribution to the environmental sensitivity;
[0026] A matching index is defined, where if the matching index is equal to 1, it indicates that the action perfectly matches the direction of the sound, and if the matching index is equal to 0, it indicates that the action does not match the direction of the sound at all; the matching index is calculated based on the actual direction of the action of the disabled person and the sound;
[0027] defining a motion range index, and calculating the motion range index by the head turning angle and walking distance of the disabled person;
[0028] The environmental sensitivity is calculated by comprehensively considering the reaction time, the matching index, the movement amplitude index, and the corresponding weight factors.
[0029] A training system for an ambient sound spatial sound recognition model, comprising:
[0030] A simulation data collection unit is provided to construct a space classroom, wherein the space classroom includes multiple scenarios that visually impaired persons encounter in real life; simulates the disabled persons in the space classroom to obtain simulation data; and pre-processes the simulation data.
[0031] A sound recognition model construction unit constructs a panoramic sound spatial sound recognition model, including an audio information collection process and a spatial audio reconstruction process; the audio information collection process generates sound direction information, sound source location information, and sound spectrum information in the simulation rehearsal data; the spatial audio reconstruction process performs audio reconstruction based on the sound direction information, the sound source location information, and the sound spectrum information to obtain a scene audio generation result;
[0032] a sound recognition result evaluation unit, wherein, in the spatial classroom, the disabled person makes a behavioral response according to the scene audio generation result; the behavioral response includes the disabled person's reaction time and behavioral action to the scene audio generation result; and based on the behavioral response, calculates environmental sensitivity;
[0033] The sound recognition model retraining unit sets an environmental sensitivity threshold. If the environmental sensitivity threshold is not greater than the environmental sensitivity threshold, the corresponding scene is re-simulated in the spatial classroom to retrain the panoramic sound spatial sound recognition model.
[0034] Furthermore, the multiple scenarios are simulated using venues, props and sound equipment; the simulation rehearsal data is the environmental sound that can be received by the human ear and is formed by various sounds emitted in the multiple scenarios, including the audio information required by the disabled person and unnecessary noise.
[0035] Furthermore, the audio information collection process generates the sound direction information, sound source location information and sound spectrum information in the simulation exercise data, and the specific steps include:
[0036] Preprocessing the simulation exercise data;
[0037] The simulation exercise data is input into a direction timing extraction path, a position timing extraction path, and a spectrum timing extraction path respectively; the direction timing extraction path includes: using multiple sensors to receive sound signals and record signal arrival times, calculating the time difference of the signal arrival times, and using the position difference of the multiple sensors to calculate the initial direction of the sound; based on the initial direction, simulating the direction change of the sound source through a rotation matrix to obtain the direction angle that changes with time, and summing the direction angles to obtain the sound direction information;
[0038] The position time series extraction path includes: using the propagation time of the sound signals received by the multiple sensors to calculate the three-dimensional position coordinates of the sound source using multi-point positioning technology; based on the three-dimensional position coordinates, using a second-order dynamic equation to generate the sound source position information;
[0039] The spectrum time series extraction path includes: designing an adaptive filter to automatically adjust the bandwidth according to the intensity and spectral characteristics of the sound signal to capture the frequency information in the sound signal in real time; using a high-order spectrum analysis method to capture the dynamic changes in the frequency information and generate the sound spectrum information.
[0040] Furthermore, the spatial audio reconstruction process performs audio reconstruction based on the sound direction information, the sound source position information, and the sound spectrum information to obtain a scene audio generation result, and the specific steps include:
[0041] Inputting the sound direction information, the sound source location information and the sound spectrum information into a multi-layer network structure to extract audio features;
[0042] Based on the audio features, synthesizing the sound using a head-related transfer function to obtain reconstructed audio;
[0043] Audio rendering and binaural rendering are performed on the reconstructed audio to obtain the scene audio generation result.
[0044] Furthermore, the method for monitoring the disabled person's reaction time and behavioral actions to the scene audio generation result is as follows: the reaction time is obtained by monitoring the disabled person's heartbeat and muscle activity changes, and the formula is: reaction time = reaction start time - sound trigger time; using the disabled person's muscle activity changes to obtain the disabled person's behavioral actions, including two aspects: first, determining whether the disabled person's action corresponds to the sound source; second, determining the disabled person's action amplitude;
[0045] Based on the behavioral response, the environmental sensitivity is calculated, and the specific process is as follows:
[0046] Assigning a weight factor to each of the behavioral actions to indicate its contribution to the environmental sensitivity;
[0047] A matching index is defined, where if the matching index is equal to 1, it indicates that the action perfectly matches the direction of the sound, and if the matching index is equal to 0, it indicates that the action does not match the direction of the sound at all; the matching index is calculated based on the actual direction of the action of the disabled person and the sound;
[0048] defining a motion range index, and calculating the motion range index by the head turning angle and walking distance of the disabled person;
[0049] The environmental sensitivity is calculated by comprehensively considering the reaction time, the matching index, the movement amplitude index, and the corresponding weight factors.
[0050] Compared with the prior art, the present invention has the following beneficial effects:
[0051] 1. The audio information acquisition process of this invention accurately extracts the spatial information and spectral characteristics of sound, providing high-quality data support for the training of panoramic sound recognition models. By using multiple sensors to calculate the arrival time difference and position difference of sound signals, the direction and location of the sound can be accurately determined. Adaptive filters and high-order spectrum analysis methods can capture frequency changes in sound signals and provide dynamic spectral information. This information can effectively reconstruct complex spatial sound environments, helping visually impaired people better perceive and understand their surroundings.
[0052] 2. The spatial audio reconstruction process of the present invention can efficiently and accurately reconstruct a real three-dimensional audio environment, providing an immersive auditory experience for people with disabilities. By inputting sound direction, sound source location, and spectrum information into a multi-layer network for feature extraction, and combining it with the head-related transfer function for sound synthesis, it is possible to simulate the propagation and changes of sound in space and reproduce complex sound scenes. Audio rendering and binaural rendering further enhance the sense of space and positioning accuracy, allowing people with disabilities to accurately perceive the location and direction of sound sources in the environment. This process improves the spatial perception ability of people with disabilities, enabling them to better identify their surroundings in daily life.
[0053] 3. The present invention provides a quantitative basis for assessing the environmental sensitivity of people with disabilities by accurately monitoring their reaction time and behavioral movements. By monitoring reaction time through changes in heart rate and muscle activity, combined with movement amplitude and matching indicators, a comprehensive analysis of the responsiveness of people with disabilities to environmental audio can be made. Multi-dimensional data collection and analysis methods help assess the adaptability of people with disabilities in auditory spatial perception and optimize training programs for different scenarios. Ultimately, by calculating environmental sensitivity, model training can be adjusted in real time to improve the adaptability of people with disabilities in real life, and enhance their independence and quality of life. BRIEF DESCRIPTION OF THE DRAWINGS
[0054] Figure 1 This is a flow chart of a method for training an ambient sound spatial sound recognition model according to the present invention;
[0055] Figure 2 This is a model flow chart of the panoramic sound spatial sound recognition model of the present invention;
[0056] Figure 3 This is a system structure diagram of a training system for an panoramic sound spatial sound recognition model of the present invention. DETAILED DESCRIPTION
[0057] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts are within the scope of protection of the present invention.
[0058] To improve the comprehensive adaptability of people with disabilities in their daily lives and studies, the present invention provides a method and system for training an ambient sound spatial sound recognition model. To illustrate the effectiveness of the present invention, the following examples will be used to illustrate the effectiveness of the present invention.
[0059] Example 1
[0060] A rehabilitation hospital in City A has launched a training method for panoramic sound recognition models for visually impaired people, helping them to conduct adaptive training for sensory compensation. Figure 1 As shown, a flowchart of a method for training an panoramic spatial sound recognition model is shown.
[0061] Reference Figure 1 In step S01, a venue is set up to form a space classroom, which includes multiple scenarios that visually impaired people encounter in real life; simulation exercises are conducted on the disabled people in the space classroom to obtain simulation exercise data; and data preprocessing is performed on the simulation exercise data.
[0062] Specifically, the multiple scenarios are simulated using venues, props and sound equipment; the simulation rehearsal data is the environmental sound that can be received by the human ear and is formed by various sounds emitted in the multiple scenarios, including the audio information required by the disabled person and unnecessary noise.
[0063] Furthermore, the preprocessing steps include: removing noise from the audio data collected from the simulation drills, using filters or noise suppression algorithms to remove background noise and non-target sounds to ensure that the collected sound signals are clear and effective; and amplifying the data by simulating different scene changes, for example: adjusting the duration of the sound, changing the speed or frequency of the audio, adding different background noises, etc., to generate more diverse training samples.
[0064] By building multiple scenarios and simulating the actual life situations of people with disabilities, we can generate simulation data containing the required audio information and noise, which helps to realistically restore the sounds in the environment, improve the accuracy and pertinence of model training, and optimize the adaptive training of people with disabilities.
[0065] Further, refer to Figure 1 Step S02 is to construct an Atmos spatial sound recognition model, including an audio information collection process and a spatial audio reconstruction process; the audio information collection process generates sound direction information, sound source location information and sound spectrum information in the simulation rehearsal data; the spatial audio reconstruction process reconstructs the audio based on the sound direction information, sound source location information and sound spectrum information to obtain the scene audio generation result. Figure 2 , showing the model flow chart of the panoramic sound spatial sound recognition model.
[0066] Specifically, the audio information collection process generates the sound direction information, sound source location information and sound spectrum information in the simulation exercise data, and the specific steps include:
[0067] Preprocessing the simulation exercise data;
[0068] Inputting the simulation exercise data into the direction time series extraction path, the position time series extraction path and the spectrum time series extraction path respectively;
[0069] The directional timing extraction path includes: using multiple sensors to receive sound signals and recording the signal arrival time. In this embodiment, two sensors S1 and S2 are used as an example, and the positions of the two sensors are P1 and P2 respectively. The time when the sound is emitted from the source point S and arrives at the two sensors is t1 and t2 respectively. The time difference of the signal arrival time Δt = t2-t1 is calculated, and the position difference d of the sensors is used to calculate the time difference of the signal arrival time. 12 =||P2-P1||, calculate the initial direction of the sound Where c is the speed of sound. Based on the initial direction θ0, a rotation matrix is used to simulate the direction change of the sound source, obtaining a direction angle that varies over time. The rotation rate of the rotation matrix is denoted as ω. The direction angle of the sound source at time t is θ(t) = θ0 + ωt. In practice, multiple sensors are used, and multiple direction angles are calculated. These multiple direction angles are aggregated to obtain the sound direction information.
[0070] Furthermore, the position time series extraction path includes: using the propagation time of the sound signals received by the multiple sensors to calculate the three-dimensional position coordinates of the sound source using multi-point positioning technology. Specifically, assuming there are N sensors S i (1,2,...,N), the positions are P i (1,2,...,N), the location of the sound source is X=(x,y,z), and the time difference of sound propagation is Δt i , the speed of sound is c, then the positioning equation of each sensor is ||XP i ||=c·Δt i This equation can be solved by the least squares method. Based on the three-dimensional position coordinates, the sound source position information is generated using a second-order dynamic equation, which is described as follows:
[0071]
[0072] Here, X0 represents the initial position, V0 represents the initial velocity, and A represents the acceleration. These parameters can be continuously updated based on real-time sensor data to generate dynamic position information.
[0073] Furthermore, the spectrum time series extraction path includes: designing an adaptive filter, the transfer function of the filter is H(f, t) which changes with time t and frequency f. In this embodiment, the spectrum X(f, t) of the signal is obtained by Fourier transform, and the output of the filter can be expressed as Y(f, t) = X(f, t) · H(f, t), wherein H(f, t) is dynamically adjusted according to the real-time signal characteristics and bandwidth requirements to capture the frequency information in the sound signal in real time. Furthermore, a high-order spectrum analysis method is used to capture the dynamic changes in the frequency information. In this embodiment, the Hilbert transform is used to extract the instantaneous value of the frequency:
[0074] X inst (f, t) = H{X(f, t)};
[0075] Where H represents the Hilbert transform. The change in instantaneous frequency can be expressed by the following formula:
[0076]
[0077] By using this method, the fluctuation pattern of the spectrum over time can be extracted to generate the sound spectrum information.
[0078] The audio information collection process achieves comprehensive capture of ambient audio by precisely extracting the direction, position, and spectral characteristics of sound. This not only accurately locates the location and dynamics of sound sources, but also captures subtle variations in sound frequency, providing high-quality input data for subsequent spatial audio reconstruction. Through multi-sensor collaboration and advanced analysis methods, the accuracy and robustness of sound recognition models can be effectively improved, helping people with disabilities better understand and adapt to ambient audio information and enhance their daily lives.
[0079] Furthermore, the spatial audio reconstruction process performs audio reconstruction based on the sound direction information, the sound source position information, and the sound spectrum information to obtain a scene audio generation result, and the specific steps include:
[0080] The sound direction information θ(t), the sound source position information X(t) and the sound spectrum information X inst (f, t) is input into the multi-layer network structure to extract audio features. The specific structure of the multi-layer network structure is:
[0081] Input layer: input sound direction information θ(t), sound source location information X(t) and sound spectrum information X inst (f,t);
[0082] Convolutional layer: convolutions the input features to generate filtered feature maps for detecting specific patterns, temporal changes, and directional features in the spectrum;
[0083] Pooling layer: downsamples the features extracted by the convolutional layer to reduce the dimension of the feature map while retaining the most important feature information;
[0084] Fully connected layer: converts the spatial audio features extracted by convolution and pooling operations into a fixed-dimensional feature vector for subsequent audio synthesis and reconstruction;
[0085] Time series modeling layer: Introducing a long short-term memory network, which retains and updates long-term memory in audio signals through a gating mechanism, models the temporal dependencies of sounds, and feeds this back into the model for subsequent audio reconstruction.
[0086] Output layer: A fully connected layer that generates the final audio features, which will be used for subsequent audio synthesis.
[0087] Furthermore, based on the audio features, a head-related transfer function is used to synthesize sound to obtain reconstructed audio; audio rendering and binaural rendering are performed on the reconstructed audio to obtain the scene audio generation result.
[0088] The spatial audio reconstruction process achieves a high degree of fidelity to real-world audio by precisely integrating information about sound direction, source location, and spectrum. This process simulates the sound perception of individuals with disabilities in complex environments, enhancing their spatial awareness of their surroundings. Through audio reconstruction and binaural rendering, individuals with disabilities can receive clearer and more accurate auditory feedback, improving their adaptability in daily life and learning, thereby enhancing their environmental awareness and reaction speed, and ultimately improving their ability to live independently and interact socially.
[0089] Further, refer to Figure 1 In step S03, in the spatial classroom, the disabled person makes a behavioral response according to the scene audio generation result; the behavioral response includes the disabled person's reaction time and behavioral actions to the scene audio generation result; and based on the behavioral response, the environmental sensitivity is calculated.
[0090] Specifically, the method for monitoring the disabled person's reaction time and behavior to the scene audio generation result is as follows: the reaction time is obtained by monitoring the disabled person's heartbeat and muscle activity changes, and the formula is: reaction time = reaction start time - sound trigger time; using the disabled person's muscle activity changes to obtain the disabled person's behavior, including two aspects: first, determining whether the disabled person's action corresponds to the sound source; second, determining the disabled person's action amplitude;
[0091] Based on the behavioral response, the environmental sensitivity is calculated, and the specific process is as follows:
[0092] Assign a weight factor to each of the behavioral actions to indicate their contribution to the environmental sensitivity; the weight factor for whether the action corresponds to the sound source is 0.8, and the weight factor for the action amplitude is 1.2;
[0093] A matching index is defined. If the matching index is equal to 1, it means that the action perfectly matches the direction of the sound. If the matching index is equal to 0, it means that the action does not match the direction of the sound at all. The matching index is calculated based on the actual direction of the action of the disabled person and the sound. The calculation formula is: Among them, angle represents the angle between the sound direction and the actual action direction, angle threshold Indicates the set threshold angle. This formula calculates the matching degree based on the angular difference between the direction of the sound and the direction of the disabled person's movement. If the angular difference is less than a certain threshold (for example, 10°), the matching degree is considered 1. Otherwise, the matching degree is reduced according to the angular difference.
[0094] Defining the range of motion index I magnitude , reflecting the intensity of the movement, and calculating the movement amplitude index by the head turning angle and walking distance of the disabled person's movement. The formula is The maximum possible amplitude is set according to the experimental scenario and can be the maximum head turning angle or the maximum step length.
[0095] The reaction time t react , the matching index match and the movement amplitude index I magnitude , and its corresponding weight factor, calculate the environment sensitivity, the formula is:
[0096]
[0097] According to this formula, shorter reaction times indicate higher environmental sensitivity. Furthermore, the closer the match between movement and sound source, and the larger the movement amplitude, the more sensitive the person with disabilities is to sound perception and reaction, thus positively enhancing environmental sensitivity.
[0098] By monitoring the reaction time and behavioral movements of individuals with disabilities, we can effectively assess their ability to perceive and respond to scene audio generation results. Reaction time monitoring, achieved through changes in heart rate and muscle activity, accurately reflects the speed of their auditory response. Behavioral analysis comprehensively assesses their spatial cognition by examining whether their movements match the sound source and the amplitude of their movements. This data can be used to quantify the environmental sensitivity of individuals with disabilities, providing a basis for training and model optimization, improving their adaptability in real life and ensuring more accurate and efficient training methods.
[0099] Further, refer to Figure 1 In step S04, the environmental sensitivity threshold is set to 0.8. If the threshold is not exceeded, the corresponding scene is re-simulated in the spatial classroom and the panoramic sound recognition model is retrained. Table 1 shows the experimental data on the responses of people with different disabilities in several scenarios.
[0100] Table 1 Experimental data display
[0101]
[0102] As can be seen from Table 1, based on the value of environmental sensitivity, if the environmental sensitivity of a disabled person is lower than 0.8, the model will be re-simulated and trained in this scenario.
[0103] A rehabilitation hospital in City A uses a training method for an immersive spatial sound recognition model to effectively improve the adaptability of individuals with disabilities to their daily environments. By building a spatial classroom simulating diverse scenarios and combining sound direction, location, and spectral data, the model accurately reconstructs the sound environment, helping individuals with disabilities navigate and make decisions in complex scenarios. Furthermore, by providing feedback on behavioral responses, the model can timely assess their environmental sensitivity and adjust training content, thereby further enhancing their spatial perception and responsiveness, strengthening their ability to care for themselves, and enhancing their sense of security.
[0104] Example 2
[0105] In order to help visually impaired people with disabilities better plan their routes and avoid danger, a company that manufactures guide equipment uses a training system for an ambient sound spatial sound recognition model to simulate the sound environment in diverse urban scenes. This allows people with disabilities to receive real-time sound prompts from their surroundings through functions such as audio recognition and sound direction perception. Figure 3 As shown, a system structure diagram of a training system for an panoramic spatial sound recognition model is shown.
[0106] Specifically, the multiple scenarios are simulated using venues, props and sound equipment; the simulation rehearsal data is the environmental sound that can be received by the human ear and is formed by various sounds emitted in the multiple scenarios, including the audio information required by the disabled person and unnecessary noise.
[0107] Furthermore, the preprocessing steps include: removing noise from the audio data collected from the simulation drills, using filters or noise suppression algorithms to remove background noise and non-target sounds to ensure that the collected sound signals are clear and effective; and amplifying the data by simulating different scene changes, for example: adjusting the duration of the sound, changing the speed or frequency of the audio, adding different background noises, etc., to generate more diverse training samples.
[0108] The system includes a simulation data acquisition unit, a venue is set up to form a spatial classroom, and the spatial classroom includes multiple scenarios encountered by visually impaired people in real life; the disabled people are simulated in the spatial classroom to obtain simulation data; and the simulation data is preprocessed.
[0109] Specifically, the audio information collection process generates the sound direction information, sound source location information and sound spectrum information in the simulation exercise data, and the specific steps include:
[0110] Preprocessing the simulation exercise data;
[0111] Inputting the simulation exercise data into the direction time series extraction path, the position time series extraction path and the spectrum time series extraction path respectively;
[0112] The directional timing extraction path includes: using multiple sensors to receive sound signals and recording the signal arrival time. In this embodiment, two sensors S1 and S2 are used as an example, and the positions of the two sensors are P1 and P2 respectively. The time when the sound is emitted from the source point S and arrives at the two sensors is t1 and t2 respectively. The time difference of the signal arrival time Δt = t2-t1 is calculated, and the position difference d of the sensors is used to calculate the time difference of the signal arrival time. 12 =||P2-P1||, calculate the initial direction of the sound Where c is the speed of sound. Based on the initial direction θ0, a rotation matrix is used to simulate the direction change of the sound source, obtaining a direction angle that varies over time. The rotation rate of the rotation matrix is denoted as ω. The direction angle of the sound source at time t is θ(t) = θ0 + ωt. In practice, multiple sensors are used, and multiple direction angles are calculated. These multiple direction angles are aggregated to obtain the sound direction information.
[0113] Furthermore, the position time series extraction path includes: using the propagation time of the sound signals received by the multiple sensors to calculate the three-dimensional position coordinates of the sound source using multi-point positioning technology. Specifically, assuming there are N sensors S i (1,2,…,N), positions are P i (1,2,…,N), the location of the sound source is X=(x,y,z), and the time difference of sound propagation is Δt i , the speed of sound is c, then the positioning equation of each sensor is ||XP i ||=c·Δt i This equation can be solved by the least squares method. Based on the three-dimensional position coordinates, the sound source position information is generated using a second-order dynamic equation, which is described as follows:
[0114]
[0115] Here, X0 represents the initial position, V0 represents the initial velocity, and A represents the acceleration. These parameters can be continuously updated based on real-time sensor data to generate dynamic position information.
[0116] Furthermore, the spectrum time series extraction path includes: designing an adaptive filter, the transfer function of the filter is H(f, t) which changes with time t and frequency f. In this embodiment, the spectrum X(f, t) of the signal is obtained by Fourier transform, and the output of the filter can be expressed as Y(f, t) = X(f, t) · H(f, t), wherein H(f, t) is dynamically adjusted according to the real-time signal characteristics and bandwidth requirements to capture the frequency information in the sound signal in real time. Furthermore, a high-order spectrum analysis method is used to capture the dynamic changes in the frequency information. In this embodiment, the Hilbert transform is used to extract the instantaneous value of the frequency:
[0117] X inst (f, t) = H{X(f, t)};
[0118] Where H represents the Hilbert transform. The change in instantaneous frequency can be expressed by the following formula:
[0119]
[0120] By using this method, the fluctuation pattern of the spectrum over time can be extracted to generate the sound spectrum information.
[0121] Furthermore, the spatial audio reconstruction process performs audio reconstruction based on the sound direction information, the sound source position information, and the sound spectrum information to obtain a scene audio generation result, and the specific steps include:
[0122] The sound direction information θ(t), the sound source position information X(t) and the sound spectrum information X inst (f, t) is input into the multi-layer network structure to extract audio features. The specific structure of the multi-layer network structure is:
[0123] Input layer: input sound direction information θ(t), sound source location information X(t) and sound spectrum information X inst (f,t);
[0124] Convolutional layer: convolutions the input features to generate filtered feature maps for detecting specific patterns, temporal changes, and directional features in the spectrum;
[0125] Pooling layer: downsamples the features extracted by the convolutional layer to reduce the dimension of the feature map while retaining the most important feature information;
[0126] Fully connected layer: converts the spatial audio features extracted by convolution and pooling operations into a fixed-dimensional feature vector for subsequent audio synthesis and reconstruction;
[0127] Time series modeling layer: Introducing a long short-term memory network, which retains and updates long-term memory in audio signals through a gating mechanism, models the temporal dependencies of sounds, and feeds this back into the model for subsequent audio reconstruction.
[0128] Output layer: A fully connected layer that generates the final audio features, which will be used for subsequent audio synthesis.
[0129] Furthermore, based on the audio features, a head-related transfer function is used to synthesize sound to obtain reconstructed audio; audio rendering and binaural rendering are performed on the reconstructed audio to obtain the scene audio generation result.
[0130] Furthermore, the system also includes a sound recognition model construction unit, which constructs a panoramic sound space sound recognition model, including an audio information collection process and a spatial audio reconstruction process; the audio information collection process generates sound direction information, sound source position information and sound spectrum information in the simulation rehearsal data; the spatial audio reconstruction process performs audio reconstruction based on the sound direction information, the sound source position information and the sound spectrum information to obtain a scene audio generation result.
[0131] Specifically, the method for monitoring the disabled person's reaction time and behavior to the scene audio generation result is as follows: the reaction time is obtained by monitoring the disabled person's heartbeat and muscle activity changes, and the formula is: reaction time = reaction start time - sound trigger time; using the disabled person's muscle activity changes to obtain the disabled person's behavior, including two aspects: first, determining whether the disabled person's action corresponds to the sound source; second, determining the disabled person's action amplitude;
[0132] Based on the behavioral response, the environmental sensitivity is calculated, and the specific process is as follows:
[0133] Assign a weight factor to each of the behavioral actions to indicate their contribution to the environmental sensitivity; the weight factor for whether the action corresponds to the sound source is 0.8, and the weight factor for the action amplitude is 1.2;
[0134] A matching index is defined. If the matching index is equal to 1, it means that the action perfectly matches the direction of the sound. If the matching index is equal to 0, it means that the action does not match the direction of the sound at all. The matching index is calculated based on the actual direction of the action of the disabled person and the sound. The calculation formula is: Among them, angle represents the angle between the sound direction and the actual action direction, angle thresholdIndicates the set threshold angle. This formula calculates the matching degree based on the angular difference between the direction of the sound and the direction of the disabled person's movement. If the angular difference is less than a certain threshold (for example, 10°), the matching degree is considered 1. Otherwise, the matching degree is reduced according to the angular difference.
[0135] Defining the range of motion index I magnitude , reflecting the intensity of the movement, and calculating the movement amplitude index by the head turning angle and walking distance of the disabled person's movement. The formula is The maximum possible amplitude is set according to the experimental scenario and can be the maximum head turning angle or the maximum step length.
[0136] The reaction time t react , the matching index match and the movement amplitude index I magnitude , and its corresponding weight factor, calculate the environment sensitivity, the formula is:
[0137]
[0138] According to this formula, shorter reaction times indicate higher environmental sensitivity. Furthermore, the closer the match between movement and sound source, and the larger the movement amplitude, the more sensitive the person with disabilities is to sound perception and reaction, thus positively enhancing environmental sensitivity.
[0139] Furthermore, the system also includes a sound recognition result evaluation unit. In the spatial classroom, the disabled person makes a behavioral response based on the scene audio generation result; the behavioral response includes the reaction time and behavioral actions of the disabled person to the scene audio generation result; based on the behavioral response, the environmental sensitivity is calculated.
[0140] Furthermore, the system also includes a sound recognition model retraining unit that sets an environmental sensitivity threshold. If the environmental sensitivity threshold is not exceeded, the corresponding scene is re-simulated in the spatial classroom to retrain the panoramic sound recognition model. Table 2 shows experimental data on the responses of people with different disabilities in several scenarios.
[0141] Table 2 Experimental data display
[0142]
[0143] As can be seen from Table 2, based on the value of environmental sensitivity, if the environmental sensitivity of a disabled person is lower than 0.9, the model will be re-simulated and trained in this scenario.
[0144] While embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions, and variations may be made to these embodiments without departing from the principles and spirit of the invention, and that the scope of the invention is defined by the appended claims and their equivalents.
Claims
1. A method for training an panoramic sound recognition model, characterized in that: include: Building a space classroom, which includes multiple scenarios that visually impaired people encounter in real life; The multi-faceted scenarios are simulated using venues, props, and sound equipment; a simulated exercise is conducted on the disabled person in the space classroom to obtain simulated exercise data; the simulated exercise data is ambient sound that can be heard by the human ear, formed by various sounds emitted in the multi-faceted scenarios, including audio information required by the disabled person and unnecessary noise; performing data preprocessing on the simulation drill data; Construct an Atmos spatial sound recognition model, including the audio information acquisition process and the spatial audio reconstruction process; The audio information collection process generates sound direction information, sound source location information and sound spectrum information in the simulation exercise data; The spatial audio reconstruction process performs audio reconstruction based on the sound direction information, the sound source position information and the sound spectrum information to obtain a scene audio generation result; In the spatial classroom, the disabled person makes a behavioral response based on the scene audio generation result; the behavioral response includes the disabled person's reaction time and behavioral actions to the scene audio generation result; based on the behavioral response, the environmental sensitivity is calculated; the reaction time is obtained by monitoring the changes in the disabled person's heartbeat and muscle activity, and the formula is: reaction time = reaction start time - sound trigger time; Utilizing the changes in the disabled person's muscle activity to obtain the disabled person's behavioral movements includes two aspects: first, determining whether the disabled person's movements correspond to the sound source; second, determining the amplitude of the disabled person's movements; Based on the behavioral response, the environmental sensitivity is calculated, and the specific process is as follows: Assigning a weight factor to each of the behavioral actions to indicate its contribution to the environmental sensitivity; A matching index is defined, where if the matching index is equal to 1, it indicates that the action perfectly matches the direction of the sound, and if the matching index is equal to 0, it indicates that the action does not match the direction of the sound at all; the matching index is calculated based on the actual direction of the action of the disabled person and the sound; defining a motion range index, and calculating the motion range index by the head turning angle and walking distance of the disabled person; Calculating the environmental sensitivity by comprehensively considering the reaction time, the matching index, the movement amplitude index, and the corresponding weight factors; An environmental sensitivity threshold is set. If the environmental sensitivity threshold is not greater than the environmental sensitivity threshold, the corresponding scene is re-simulated in the spatial classroom and the panoramic sound recognition model is retrained.
2. The method for training an ambient sound spatial sound recognition model according to claim 1, wherein: The audio information collection process generates sound direction information, sound source location information, and sound spectrum information in the simulation exercise data, and the specific steps include: Preprocessing the simulation exercise data; The simulation exercise data is input into a direction timing extraction path, a position timing extraction path, and a spectrum timing extraction path respectively; the direction timing extraction path includes: using multiple sensors to receive sound signals and record signal arrival times, calculating the time difference of the signal arrival times, and using the position difference of the multiple sensors to calculate the initial direction of the sound; based on the initial direction, simulating the direction change of the sound source through a rotation matrix to obtain the direction angle that changes with time, and summing the direction angles to obtain the sound direction information; The position time series extraction path includes: using the propagation time of the sound signals received by the multiple sensors to calculate the three-dimensional position coordinates of the sound source using multi-point positioning technology; based on the three-dimensional position coordinates, using a second-order dynamic equation to generate the sound source position information; The spectrum time series extraction path includes: designing an adaptive filter to automatically adjust the bandwidth according to the intensity and spectral characteristics of the sound signal to capture the frequency information in the sound signal in real time; using a high-order spectrum analysis method to capture the dynamic changes in the frequency information and generate the sound spectrum information.
3. The method for training an ambient sound spatial sound recognition model according to claim 1, wherein: The spatial audio reconstruction process performs audio reconstruction based on the sound direction information, the sound source position information, and the sound spectrum information to obtain a scene audio generation result, and specifically includes the following steps: Inputting the sound direction information, the sound source location information and the sound spectrum information into a multi-layer network structure to extract audio features; Based on the audio features, synthesizing the sound using a head-related transfer function to obtain reconstructed audio; Audio rendering and binaural rendering are performed on the reconstructed audio to obtain the scene audio generation result.
4. A training system for an ambient sound spatial sound recognition model, characterized in that: include: A simulated data collection unit is built to form a space classroom, which includes multiple scenarios that visually impaired people encounter in real life; The multi-faceted scenarios are simulated using venues, props, and sound equipment; a simulated drill is conducted on the disabled person in the space classroom to obtain simulated drill data; the simulated drill data is ambient sound that can be heard by the human ear, formed by various sounds emitted in the multi-faceted scenarios, including audio information required by the disabled person and unnecessary noise; and data preprocessing is performed on the simulated drill data; The sound recognition model construction unit constructs an panoramic sound spatial sound recognition model, including the audio information collection process and the spatial audio reconstruction process; The audio information collection process generates sound direction information, sound source location information and sound spectrum information in the simulation exercise data; The spatial audio reconstruction process performs audio reconstruction based on the sound direction information, the sound source position information and the sound spectrum information to obtain a scene audio generation result; A sound recognition result evaluation unit, wherein in the spatial classroom, the disabled person makes a behavioral response based on the scene audio generation result; the behavioral response includes the disabled person's reaction time and behavioral actions to the scene audio generation result; based on the behavioral response, environmental sensitivity is calculated; the reaction time is obtained by monitoring changes in the disabled person's heartbeat and muscle activity, and the formula is: reaction time = reaction start time - sound trigger time; Utilizing the changes in the disabled person's muscle activity to obtain the disabled person's behavioral movements includes two aspects: first, determining whether the disabled person's movements correspond to the sound source; second, determining the amplitude of the disabled person's movements; Based on the behavioral response, the environmental sensitivity is calculated, and the specific process is as follows: Assigning a weight factor to each of the behavioral actions to indicate its contribution to the environmental sensitivity; A matching index is defined, where if the matching index is equal to 1, it indicates that the action perfectly matches the direction of the sound, and if the matching index is equal to 0, it indicates that the action does not match the direction of the sound at all; the matching index is calculated based on the actual direction of the action of the disabled person and the sound; defining a motion range index, and calculating the motion range index by the head turning angle and walking distance of the disabled person; Calculating the environmental sensitivity by comprehensively considering the reaction time, the matching index, the movement amplitude index, and the corresponding weight factors; The sound recognition model retraining unit sets an environmental sensitivity threshold. If the environmental sensitivity threshold is not greater than the environmental sensitivity threshold, the corresponding scene is re-simulated in the spatial classroom to retrain the panoramic sound spatial sound recognition model.
5. The training system for an ambient sound spatial sound recognition model according to claim 4, characterized in that: The audio information collection process generates sound direction information, sound source location information, and sound spectrum information in the simulation exercise data, and the specific steps include: Preprocessing the simulation exercise data; The simulation exercise data is input into a direction timing extraction path, a position timing extraction path, and a spectrum timing extraction path respectively; the direction timing extraction path includes: using multiple sensors to receive sound signals and record signal arrival times, calculating the time difference of the signal arrival times, and using the position difference of the multiple sensors to calculate the initial direction of the sound; based on the initial direction, simulating the direction change of the sound source through a rotation matrix to obtain the direction angle that changes with time, and summing the direction angles to obtain the sound direction information; The position time series extraction path includes: using the propagation time of the sound signals received by the multiple sensors to calculate the three-dimensional position coordinates of the sound source using multi-point positioning technology; based on the three-dimensional position coordinates, using a second-order dynamic equation to generate the sound source position information; The spectrum time series extraction path includes: designing an adaptive filter to automatically adjust the bandwidth according to the intensity and spectral characteristics of the sound signal to capture the frequency information in the sound signal in real time; using a high-order spectrum analysis method to capture the dynamic changes in the frequency information and generate the sound spectrum information.
6. The training system for an ambient sound spatial sound recognition model according to claim 4, characterized in that: The spatial audio reconstruction process performs audio reconstruction based on the sound direction information, the sound source position information, and the sound spectrum information to obtain a scene audio generation result, and specifically includes the following steps: Inputting the sound direction information, the sound source location information and the sound spectrum information into a multi-layer network structure to extract audio features; Based on the audio features, synthesizing the sound using a head-related transfer function to obtain reconstructed audio; Audio rendering and binaural rendering are performed on the reconstructed audio to obtain the scene audio generation result.
Citation Information
Patent Citations
Spatial perception training method and device based on auditory information
CN118454057A
KR20220082440A