Training method and system of panoramic sound space sound recognition model

By simulating multiple scenes in the spatial classroom and constructing a panoramic sound recognition model, the problem of low sound recognition accuracy for people with disabilities in complex noise environments is solved, and their adaptability and quality of life are improved.

CN120032644AActive Publication Date: 2025-05-23GUANG DONG ULTRA PICTURES CULTURE COMM CO LTD
View PDF 13 Cites 0 Cited by

Patent Information

Application Number
CN202510175276.1
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-18
Publication Date
2025-05-23
Estimated Expiration
2045-02-18

AI Technical Summary

Technical Problem

In extremely complex background noise, when multiple sound sources overlap or environmental noise interference is too large, the sound recognition accuracy of people with disabilities is reduced, affecting their ability to adapt to society.

Method used

By building a spatial classroom, we simulate the diverse scenarios encountered by people with disabilities in real life, collect and preprocess audio data, and build a panoramic sound recognition model, including audio information acquisition and spatial audio reconstruction processes, calculate environmental sensitivity and adjust model training according to threshold values.

Benefits of technology

It improves the comprehensive adaptability of people with disabilities in daily life and learning, enhances their spatial perception of complex environments, and improves the quality of life and independence.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120032644A_ABST
    Figure CN120032644A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of hearing aided training, in particular to a training method and system of a panoramic sound space sound recognition model. The method comprises the following steps: firstly, building a site to form a space classroom including multi-element scenes in actual life; and performing simulation drilling on the disabled in the space class to obtain simulation drilling data. Secondly, constructing a panoramic sound space sound recognition model, including an audio information acquisition process and a space audio reconstruction process; then, in a space classroom, the disabled person makes a behavior response according to a scene audio generation result; based on the behavioral response, environmental sensitivity is calculated. And finally, setting an environment sensitivity threshold value, if the environment sensitivity threshold value is not greater than the environment sensitivity threshold value, re-simulating the corresponding scene in the space classroom, and re-training the panoramic sound space sound recognition model. According to the invention, the comprehensive adaptive capacity of the disabled in daily life and learning can be improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of hearing assistance training, and in particular, to a training method and system for a panoramic sound spatial sound recognition model. Background Art

[0002] For visually impaired disabled people, sound is one of the important ways to communicate with the outside world. Through sound, disabled people can identify key information in the environment, such as traffic signals, people's speech, warning sounds, etc. This sound perception not only helps with navigation in daily life but also helps disabled people make safe decisions and avoid potential dangers.

[0003] For visually impaired disabled people, traditional assistance methods mainly include Braille, voice guidance, and blind assistance devices. Braille, as a tactile method, helps blind people with reading and writing; voice guidance technology provides users with direction, location, and environmental information through intelligent devices to help blind people travel independently; blind assistance tools, such as guide dogs and white canes, use touch and hearing to guide disabled people to avoid obstacles. Although these methods can effectively compensate for the lack of vision, in complex environments or dynamic scenarios, blind people may face problems such as insufficient information or untimely environmental feedback, which affects their quality of life.

[0004] Currently, panoramic sound spatial sound recognition technology has been applied to disabled people's assistance systems. Combining multi-microphone arrays and sound source localization technology, the system can identify and enhance specific sounds, such as people's speech or warning signals, in complex environments and reduce the interference of environmental noise. In addition, deep learning technology is used for sound classification and noise cancellation, improving the accuracy of recognition. However, although these technologies can enhance the sound recognition ability of disabled people in noisy environments, in extremely complex background noise, they still face problems such as multiple sound sources overlapping or excessive environmental noise interference, resulting in a decrease in recognition accuracy, which in turn leads to an inability to improve the adaptability of disabled people in society.

[0005] Therefore, a training method and system for a panoramic sound spatial sound recognition model are proposed. Summary of the Invention

[0006] The purpose of the present invention is to provide a training method and system for an panoramic sound space sound recognition model, which is used to improve the comprehensive adaptability of disabled people in daily life and learning. First, a venue is set up to form a spatial classroom, which includes multiple scenes encountered by disabled people with visual impairments in real life; simulated exercises are conducted on disabled people in the spatial classroom to obtain simulated exercise data; and data preprocessing is performed on the simulated exercise data. Secondly, a panoramic sound space sound recognition model is constructed, including an audio information collection process and a spatial audio reconstruction process; the audio information collection process generates sound direction information, sound source location information and sound spectrum information in the simulated exercise data; the spatial audio reconstruction process performs audio reconstruction based on the sound direction information, sound source location information and sound spectrum information to obtain the scene audio generation result. Then, in the spatial classroom, the disabled person makes a behavioral response according to the scene audio generation result; the behavioral response includes the reaction time and behavioral action of the disabled person to the scene audio generation result; based on the behavioral response, the environmental sensitivity is calculated. Finally, the environmental sensitivity threshold is set. If it is not greater than the environmental sensitivity threshold, the corresponding scene is re-simulated in the spatial classroom, and the panoramic sound space sound recognition model is re-trained.

[0007] To achieve the above object, the present invention provides the following technical solutions:

[0008] A training method and system for an ambient sound spatial sound recognition model, comprising:

[0009] A venue is set up to form a space classroom, wherein the space classroom includes multiple scenarios that a visually impaired person encounters in real life; a simulation exercise is conducted on the disabled person in the space classroom to obtain simulation exercise data; and data preprocessing is performed on the simulation exercise data;

[0010] Constructing an ambient sound spatial sound recognition model, including an audio information collection process and a spatial audio reconstruction process; the audio information collection process generates sound direction information, sound source location information and sound spectrum information in the simulation exercise data; the spatial audio reconstruction process performs audio reconstruction based on the sound direction information, the sound source location information and the sound spectrum information to obtain a scene audio generation result;

[0011] In the space classroom, the disabled person makes a behavioral response according to the scene audio generation result; the behavioral response includes the disabled person's reaction time and behavioral action to the scene audio generation result; based on the behavioral response, the environmental sensitivity is calculated;

[0012] An environmental sensitivity threshold is set. If it is not greater than the environmental sensitivity threshold, the corresponding scene is re-simulated in the spatial classroom and the panoramic sound spatial sound recognition model is re-trained.

[0013] Furthermore, the multiple scenarios are simulated using venues, props and sound equipment; the simulation rehearsal data is environmental sound that can be received by human ears and is formed by various sounds emitted in the multiple scenarios, including audio information required by the disabled person and unnecessary noise.

[0014] Furthermore, the audio information collection process generates the sound direction information, sound source location information and sound spectrum information in the simulation exercise data, and the specific steps include:

[0015] Preprocessing the simulation exercise data;

[0016] The simulation exercise data are input into the direction timing extraction path, the position timing extraction path and the spectrum timing extraction path respectively; the direction timing extraction path comprises: using multiple sensors to receive sound signals and record the signal arrival time, calculating the time difference of the signal arrival time, and using the position difference of multiple sensors to calculate the initial direction of the sound; based on the initial direction, simulating the direction change of the sound source through the rotation matrix to obtain the direction angle that changes with time, and summarizing the direction angle to obtain the sound direction information;

[0017] The position timing extraction path includes: calculating the three-dimensional position coordinates of the sound source by using the propagation time of the sound signals received by the multiple sensors using the multi-point positioning technology; based on the three-dimensional position coordinates, generating the sound source position information using the second-order dynamic equation;

[0018] The spectrum time series extraction path includes: designing an adaptive filter to automatically adjust the bandwidth according to the intensity and spectrum characteristics of the sound signal to capture the frequency information in the sound signal in real time; using a high-order spectrum analysis method to capture the dynamic changes in the frequency information to generate the sound spectrum information.

[0019] Furthermore, the spatial audio reconstruction process performs audio reconstruction based on the sound direction information, the sound source position information and the sound spectrum information to obtain a scene audio generation result, and the specific steps include:

[0020] Input the sound direction information, the sound source location information and the sound spectrum information into a multi-layer network structure to extract audio features;

[0021] Based on the audio features, synthesizing the sound using a head-related transfer function to obtain reconstructed audio;

[0022] Audio rendering and binaural rendering are performed on the reconstructed audio to obtain the scene audio generation result.

[0023] Furthermore, the method for monitoring the reaction time and behavior of the disabled person to the scene audio generation result is as follows: the reaction time is obtained by monitoring the heartbeat and muscle activity changes of the disabled person, and the formula is: reaction time = reaction start time - sound trigger time; using the muscle activity changes of the disabled person, the behavior of the disabled person is obtained, including two aspects: first, judging whether the action of the disabled person corresponds to the sound source; second, judging the action amplitude of the disabled person;

[0024] Based on the behavioral response, the environmental sensitivity is calculated, and the specific process is as follows:

[0025] Assigning a weight factor to each of the behavioral actions to indicate the contribution to the environmental sensitivity;

[0026] A matching index is defined, where if the matching index is equal to 1, it means that the action perfectly matches the direction of the sound, and if the matching index is equal to 0, it means that the action does not match the direction of the sound at all; and the matching index is calculated according to the action of the disabled person and the actual direction of the sound;

[0027] defining a movement range index, and calculating the movement range index according to the head turning angle and walking distance of the disabled person;

[0028] The environmental sensitivity is calculated by comprehensively considering the reaction time, the matching index, the motion amplitude index, and the corresponding weight factors.

[0029] A training system for an ambient sound spatial sound recognition model, comprising:

[0030] A simulation data collection unit is provided to build a space classroom, wherein the space classroom includes multiple scenarios that a visually impaired person encounters in real life; a simulation exercise is performed on the disabled person in the space classroom to obtain simulation exercise data; and data preprocessing is performed on the simulation exercise data;

[0031] A sound recognition model construction unit constructs a panoramic sound space sound recognition model, including an audio information collection process and a spatial audio reconstruction process; the audio information collection process generates sound direction information, sound source location information and sound spectrum information in the simulation exercise data; the spatial audio reconstruction process performs audio reconstruction based on the sound direction information, the sound source location information and the sound spectrum information to obtain a scene audio generation result;

[0032] A sound recognition result evaluation unit, in the space classroom, the disabled person makes a behavioral response according to the scene audio generation result; the behavioral response includes the disabled person's reaction time and behavioral action to the scene audio generation result; based on the behavioral response, the environmental sensitivity is calculated;

[0033] The sound recognition model retraining unit sets an environmental sensitivity threshold. If the environmental sensitivity threshold is not greater than the environmental sensitivity threshold, the corresponding scene is re-simulated in the spatial classroom to retrain the panoramic sound spatial sound recognition model.

[0034] Furthermore, the multiple scenarios are simulated using venues, props and sound equipment; the simulation rehearsal data is environmental sound that can be received by human ears and is formed by various sounds emitted in the multiple scenarios, including audio information required by the disabled person and unnecessary noise.

[0035] Furthermore, the audio information collection process generates the sound direction information, sound source location information and sound spectrum information in the simulation exercise data, and the specific steps include:

[0036] Preprocessing the simulation exercise data;

[0037] The simulation exercise data are input into the direction timing extraction path, the position timing extraction path and the spectrum timing extraction path respectively; the direction timing extraction path includes: using multiple sensors to receive sound signals and record the signal arrival time, calculating the time difference of the signal arrival time, and using the position difference of multiple sensors to calculate the initial direction of the sound; based on the initial direction, the direction change of the sound source is simulated by a rotation matrix to obtain the direction angle that changes with time, and the direction angle is summarized to obtain the sound direction information;

[0038] The position timing extraction path includes: calculating the three-dimensional position coordinates of the sound source by using the propagation time of the sound signals received by the multiple sensors using the multi-point positioning technology; based on the three-dimensional position coordinates, generating the sound source position information using the second-order dynamic equation;

[0039] The spectrum time series extraction path includes: designing an adaptive filter to automatically adjust the bandwidth according to the intensity and spectrum characteristics of the sound signal to capture the frequency information in the sound signal in real time; using a high-order spectrum analysis method to capture the dynamic changes in the frequency information to generate the sound spectrum information.

[0040] Furthermore, the spatial audio reconstruction process performs audio reconstruction based on the sound direction information, the sound source position information and the sound spectrum information to obtain a scene audio generation result, and the specific steps include:

[0041] Input the sound direction information, the sound source location information and the sound spectrum information into a multi-layer network structure to extract audio features;

[0042] Based on the audio features, synthesizing the sound using a head-related transfer function to obtain reconstructed audio;

[0043] Audio rendering and binaural rendering are performed on the reconstructed audio to obtain the scene audio generation result.

[0044] Furthermore, the method for monitoring the reaction time and behavior of the disabled person to the scene audio generation result is as follows: the reaction time is obtained by monitoring the heartbeat and muscle activity changes of the disabled person, and the formula is: reaction time = reaction start time - sound trigger time; using the muscle activity changes of the disabled person, the behavior of the disabled person is obtained, including two aspects: first, judging whether the action of the disabled person corresponds to the sound source; second, judging the action amplitude of the disabled person;

[0045] Based on the behavioral response, the environmental sensitivity is calculated, and the specific process is as follows:

[0046] Assigning a weight factor to each of the behavioral actions to indicate the contribution to the environmental sensitivity;

[0047] A matching index is defined, where if the matching index is equal to 1, it means that the action perfectly matches the direction of the sound, and if the matching index is equal to 0, it means that the action does not match the direction of the sound at all; and the matching index is calculated according to the action of the disabled person and the actual direction of the sound;

[0048] defining a movement range index, and calculating the movement range index according to the head turning angle and walking distance of the disabled person;

[0049] The environmental sensitivity is calculated by comprehensively considering the reaction time, the matching index, the motion amplitude index, and the corresponding weight factors.

[0050] Compared with the prior art, the present invention has the following beneficial effects:

[0051] 1. The audio information collection process of the present invention can accurately extract the spatial information and spectral characteristics of the sound, providing high-quality data support for the training of the panoramic sound spatial sound recognition model. Through the collaborative work of multiple sensors, the arrival time difference and position difference of the sound signal are calculated, and the direction and position information of the sound can be accurately obtained; the adaptive filter and high-order spectrum analysis method can capture the frequency changes in the sound signal and provide dynamic spectrum information. This information can effectively reconstruct the complex spatial sound environment and help the visually impaired better perceive and understand the surrounding environment.

[0052] 2. The spatial audio reconstruction process of the present invention can efficiently and accurately reconstruct a real three-dimensional audio environment, providing an immersive auditory experience for people with disabilities. By inputting the sound direction, sound source location and spectrum information into a multi-layer network for feature extraction, and combining the head-related transfer function for sound synthesis, it is possible to simulate the propagation and change of sound in space and reproduce complex sound scenes. Audio rendering and binaural rendering further enhance the sense of space and positioning accuracy, allowing people with disabilities to accurately perceive the location and direction of sound sources in the environment. This process enhances the spatial perception ability of people with disabilities, enabling them to better identify their surroundings in daily life.

[0053] 3. The present invention provides a quantitative basis for evaluating the environmental sensitivity of disabled persons by accurately monitoring their reaction time and behavioral actions. By monitoring the reaction time through changes in heartbeat and muscle activity, combined with movement amplitude and matching indicators, the responsiveness of disabled persons to environmental audio can be comprehensively analyzed. Multi-dimensional data collection and analysis methods help evaluate the adaptability of disabled persons in auditory spatial perception and optimize training programs for different scenarios. Finally, by calculating environmental sensitivity, model training can be adjusted in real time to enhance the adaptability of disabled persons in real life, improve independence and quality of life. BRIEF DESCRIPTION OF THE DRAWINGS

[0054] Figure 1 A method flow chart of a method for training a panoramic sound spatial sound recognition model of the present invention;

[0055] Figure 2 A model flow chart of the panoramic sound spatial sound recognition model of the present invention;

[0056] Figure 3 The system structure diagram of a training system for an panoramic sound spatial sound recognition model of the present invention. DETAILED DESCRIPTION

[0057] The following will be combined with the drawings in the embodiments of the present invention to clearly and completely describe the technical solutions in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0058] In order to improve the comprehensive adaptability of disabled people in daily life and learning, the present invention provides a training method and system for an panoramic sound space sound recognition model. In order to illustrate the role of the present invention, the effectiveness of the present invention will be described from the following examples.

[0059] Embodiment 1

[0060] A rehabilitation hospital in City A has launched a training method for panoramic sound recognition models for visually impaired people, helping them to conduct adaptive training for sensory compensation. Figure 1 As shown, a flowchart of a method for training an ambient sound spatial sound recognition model is shown.

[0061] Reference Figure 1 In step S01, a venue is built to form a space classroom, where the space classroom includes multiple scenarios that visually impaired persons encounter in real life; simulated drills are conducted on the disabled persons in the space classroom to obtain simulated drill data; and data preprocessing is performed on the simulated drill data.

[0062] Specifically, the multiple scenarios are simulated using venues, props and sound equipment; the simulation rehearsal data is environmental sound that can be received by the human ear and is formed by various sounds emitted in the multiple scenarios, including audio information required by the disabled person and unnecessary noise.

[0063] Furthermore, the preprocessing steps include: removing noise from the audio data collected from the simulation exercises, using filters or noise suppression algorithms to remove background noise and non-target sounds, and ensuring that the collected sound signals are clear and effective; amplifying the data by simulating different scene changes, for example: adjusting the duration of the sound, changing the speed or frequency of the audio, adding different background noises, etc., to generate more diverse training samples.

[0064] By building multiple scenarios and simulating the actual life situations of people with disabilities, we can generate simulation data containing the required audio information and noise, which helps to truly restore the sounds in the environment, improve the accuracy and pertinence of model training, and optimize the adaptive training of people with disabilities.

[0065] Further, refer to Figure 1 Step S02 of the invention is to construct a panoramic sound spatial sound recognition model, including an audio information collection process and a spatial audio reconstruction process; the audio information collection process generates sound direction information, sound source location information and sound spectrum information in the simulation exercise data; the spatial audio reconstruction process reconstructs the audio based on the sound direction information, sound source location information and sound spectrum information to obtain a scene audio generation result. Figure 2 , showing the model flow chart of the panoramic sound spatial sound recognition model.

[0066] Specifically, the audio information collection process generates the sound direction information, sound source location information and sound spectrum information in the simulation exercise data, and the specific steps include:

[0067] Preprocessing the simulation exercise data;

[0068] Inputting the simulation exercise data into the direction timing extraction path, the position timing extraction path and the spectrum timing extraction path respectively;

[0069] The direction timing extraction path includes: using multiple sensors to receive sound signals and record the signal arrival time. In this embodiment, two sensors S are used. 1 and S 2 As an example, the positions of the two sensors are P 1 and P 2 The time it takes for the sound to be emitted from the source point S and reach the two sensors is t 1 and t 2 . Calculate the time difference of the signal arrival time Δt=t 2 -t 1 , and using the position difference d of the sensor 12 =||P 2 -P 1 ||, calculate the initial direction of the sound Where c is the speed of sound; based on the initial direction θ 0 , the direction change of the sound source is simulated by the rotation matrix, and the direction angle that changes with time is obtained. The rotation rate of the rotation matrix is ​​recorded as ω, and the direction angle of the sound source at time t is θ(t) = θ 0 In actual situations, there are multiple sensors and multiple calculated direction angles, and the multiple direction angles are summarized to obtain the sound direction information.

[0070] Furthermore, the position time series extraction path includes: using the propagation time of the sound signals received by the multiple sensors to calculate the three-dimensional position coordinates of the sound source using multi-point positioning technology. Specifically, assuming there are N sensors S i (1,2,...,N), the positions are P i (1,2,...,N), the location of the sound source is X=(x,y,z), and the time difference of sound propagation is Δt i , the speed of sound is c, then the positioning equation of each sensor is ||XP i ||=c·Δt i This equation can be solved by the least square method. Based on the three-dimensional position coordinates, the sound source position information is generated using a second-order dynamic equation, which is described as follows:

[0071]

[0072] Among them, X 0 represents the initial position, V 0 represents the initial velocity, and A represents the acceleration. According to the real-time sensor data, these parameters can be continuously updated to generate dynamic position information.

[0073] Furthermore, the spectrum time series extraction path includes: designing an adaptive filter, the transfer function of the filter is H(f, t) which changes with time t and frequency f. In this embodiment, the spectrum X(f, t) of the signal is obtained by Fourier transform, and the output of the filter can be expressed as Y(f, t) = X(f, t) · H(f, t), wherein H(f, t) is dynamically adjusted according to the real-time signal characteristics and bandwidth requirements to capture the frequency information in the sound signal in real time. Furthermore, a high-order spectrum analysis method is used to capture the dynamic changes in the frequency information. In this embodiment, the Hilbert transform is used to extract the instantaneous value of the frequency:

[0074] X inst (f, t) = H{X(f, t)};

[0075] Wherein, H represents Hilbert transform. Wherein, the change of instantaneous frequency can be expressed by the following formula:

[0076]

[0077] Through this method, the fluctuation pattern of the spectrum over time can be extracted to generate the sound spectrum information.

[0078] The audio information collection process achieves comprehensive capture of environmental audio by accurately extracting the direction, position and spectral characteristics of the sound. This can not only accurately locate the position and dynamic changes of the sound source, but also capture subtle changes in the sound frequency, providing high-quality input data for subsequent spatial audio reconstruction. Through multi-sensor collaboration and high-order analysis methods, the accuracy and robustness of the sound recognition model can be effectively improved, thereby helping people with disabilities better understand and adapt to environmental audio information and improve their life experience.

[0079] Furthermore, the spatial audio reconstruction process performs audio reconstruction based on the sound direction information, the sound source position information and the sound spectrum information to obtain a scene audio generation result, and the specific steps include:

[0080] The sound direction information θ(t), the sound source location information X(t) and the sound spectrum information X inst (f, t) is input into the multi-layer network structure to extract audio features. The specific structure of the multi-layer network structure is:

[0081] Input layer: input sound direction information θ(t), sound source location information X(t) and sound spectrum information X inst (f,t);

[0082] Convolutional layer: convolutions the input features to generate filtered feature maps for detecting specific patterns, temporal changes, and directional features in the spectrum;

[0083] Pooling layer: downsamples the features extracted by the convolutional layer to reduce the dimension of the feature map while retaining the most important feature information;

[0084] Fully connected layer: converts the spatial audio features extracted by convolution and pooling operations into a feature vector of fixed dimension for subsequent audio synthesis and reconstruction;

[0085] Time series modeling layer: Long short-term memory network is introduced to retain and update long-term memory in audio signals through gating mechanism, model the temporal dependency of sound, and feed it back into the model for subsequent audio reconstruction;

[0086] Output layer: It is a fully connected layer that generates the final audio features, which will be used for subsequent audio synthesis.

[0087] Furthermore, based on the audio features, the sound is synthesized using a head-related transfer function to obtain reconstructed audio; audio rendering and binaural rendering are performed on the reconstructed audio to obtain the scene audio generation result.

[0088] The spatial audio reconstruction process achieves a high degree of restoration of real-scene audio by accurately integrating sound direction information, sound source location information, and sound spectrum information. This process can simulate the sound perception of people with disabilities in complex environments and enhance their spatial perception of the surrounding environment. Through audio reconstruction and binaural rendering, people with disabilities can obtain clearer and more accurate auditory feedback, improve their adaptability in daily life and learning, thereby improving their environmental perception and reaction speed, and further improving their ability to live independently and interact with society.

[0089] Further, refer to Figure 1 In step S03, in the space classroom, the disabled person makes a behavioral response according to the scene audio generation result; the behavioral response includes the disabled person's reaction time and behavioral actions to the scene audio generation result; based on the behavioral response, the environmental sensitivity is calculated.

[0090] Specifically, the method for monitoring the reaction time and behavior of the disabled person to the scene audio generation result is as follows: the reaction time is obtained by monitoring the heartbeat and muscle activity changes of the disabled person, and the formula is: reaction time = reaction start time - sound trigger time; using the muscle activity changes of the disabled person, the behavior of the disabled person is obtained, including two aspects: first, judging whether the action of the disabled person corresponds to the sound source; second, judging the action amplitude of the disabled person;

[0091] Based on the behavioral response, the environmental sensitivity is calculated, and the specific process is as follows:

[0092] Assign a weight factor to each of the behavioral actions to indicate the contribution to the environmental sensitivity; the weight factor of whether the action corresponds to the sound source is 0.8, and the weight factor of the action amplitude is 1.2;

[0093] A matching index is defined. If the matching index is equal to 1, it means that the action perfectly matches the sound direction. If the matching index is equal to 0, it means that the action does not match the sound direction at all. According to the actual direction of the action and sound of the disabled person, the matching index is calculated. The calculation formula is: Among them, angle represents the angle between the sound direction and the actual action direction, angle threshold Indicates the set threshold angle. This formula calculates the matching degree based on the angle difference between the direction of the sound and the direction of the disabled person's movement. If the angle difference is less than a certain threshold (for example, 10°), the matching degree is considered to be 1, otherwise the matching degree is attenuated according to the angle difference.

[0094] Defining the range of motion index I magnitude , reflects the intensity of the action, and the action amplitude index is calculated by the head turning angle and walking distance of the disabled person's action. The formula is The maximum possible amplitude is set according to the experimental scenario, which can be the maximum head turning angle or the maximum step length.

[0095] The reaction time t react , the matching index match and the action amplitude index I magnitude , and its corresponding weight factor, calculate the environment sensitivity Environment Sensitivity, the formula is:

[0096]

[0097] According to this formula, the shorter the reaction time, the higher the environmental sensitivity. In addition, the higher the matching degree between the action and the sound source and the larger the movement amplitude, the more sensitive the disabled person is to the sound perception and reaction, thus having a positive effect on improving environmental sensitivity.

[0098] By monitoring the reaction time and behavioral actions of people with disabilities, their perception and response ability to the scene audio generation results can be effectively evaluated. The monitoring of reaction time is achieved through changes in heartbeat and muscle activity, which can accurately reflect the speed of their auditory response. The analysis of behavioral actions starts from whether the action matches the sound source and the amplitude of the action, and comprehensively evaluates their spatial cognitive ability. Through these data, the environmental sensitivity of people with disabilities can be quantified, which in turn provides a basis for training and model optimization, improves their adaptability in real life, and ensures that the training method is more accurate and efficient.

[0099] Further, refer to Figure 1 In step S04, the environmental sensitivity threshold is set to 0.8. If it is not greater than the environmental sensitivity threshold, the corresponding scene is re-simulated in the space classroom and the panoramic sound recognition model is re-trained. As shown in Table 1, the experimental data of the responses of people with different disabilities in several scenarios are shown.

[0100] Table 1 Experimental data display

[0101]

[0102] As can be seen from Table 1, based on the value of environmental sensitivity, if the environmental sensitivity of a disabled person is lower than 0.8, the model will be re-simulated and trained in this scenario.

[0103] A training method for an all-around sound spatial sound recognition model used by a rehabilitation hospital in City A can effectively improve the environmental adaptability of people with disabilities in their daily lives. By building a spatial classroom that simulates multiple scenes and combining sound direction, location information, and spectrum data, the model can accurately reconstruct the sound environment and help people with disabilities orient and judge in complex scenes. At the same time, by providing feedback on behavioral responses, their sensitivity to the environment is timely evaluated and the training content is adjusted, so as to better improve the spatial perception and action response capabilities of people with disabilities, and enhance their ability to take care of themselves and their sense of security.

[0104] Embodiment 2

[0105] In order to help visually impaired people with disabilities better plan their routes and avoid danger, a company producing guide equipment used a training system for panoramic sound recognition models to simulate the sound environment in a variety of urban scenes, so that people with disabilities can get real-time sound prompts from the surrounding environment through functions such as audio recognition and sound direction perception. Figure 3 As shown, a system structure diagram of a training system for an panoramic spatial sound recognition model is shown.

[0106] Specifically, the multiple scenarios are simulated using venues, props and sound equipment; the simulation rehearsal data is environmental sound that can be received by the human ear and is formed by various sounds emitted in the multiple scenarios, including audio information required by the disabled person and unnecessary noise.

[0107] Furthermore, the preprocessing steps include: removing noise from the audio data collected from the simulation exercises, using filters or noise suppression algorithms to remove background noise and non-target sounds, and ensuring that the collected sound signals are clear and effective; amplifying the data by simulating different scene changes, for example: adjusting the duration of the sound, changing the speed or frequency of the audio, adding different background noises, etc., to generate more diverse training samples.

[0108] The system includes a simulation data acquisition unit, a venue is built to form a space classroom, and the space classroom includes multiple scenarios encountered by disabled people with visual impairments in real life; simulated exercises are performed on the disabled people in the space classroom to obtain simulated exercise data; and data preprocessing is performed on the simulated exercise data.

[0109] Specifically, the audio information collection process generates the sound direction information, sound source location information and sound spectrum information in the simulation exercise data, and the specific steps include:

[0110] Preprocessing the simulation exercise data;

[0111] Inputting the simulation exercise data into the direction timing extraction path, the position timing extraction path and the spectrum timing extraction path respectively;

[0112] The direction timing extraction path includes: using multiple sensors to receive sound signals and record the signal arrival time. In this embodiment, two sensors S are used. 1 and S 2 As an example, the positions of the two sensors are P 1 and P 2 The time it takes for the sound to be emitted from the source point S and reach the two sensors is t 1 and t 2 . Calculate the time difference of the signal arrival time Δt=t 2 -t 1 , and using the position difference d of the sensor 12 =||P 2 -P 1 ||, calculate the initial direction of the sound Where c is the speed of sound; based on the initial direction θ 0 , the direction change of the sound source is simulated by the rotation matrix, and the direction angle that changes with time is obtained. The rotation rate of the rotation matrix is ​​recorded as ω, and the direction angle of the sound source at time t is θ(t) = θ 0 In actual situations, there are multiple sensors and multiple calculated direction angles, and the multiple direction angles are summarized to obtain the sound direction information.

[0113] Furthermore, the position time series extraction path includes: using the propagation time of the sound signals received by the multiple sensors to calculate the three-dimensional position coordinates of the sound source using multi-point positioning technology. Specifically, assuming there are N sensors S i (1,2,…,N), the positions are P i (1,2,…,N), the location of the sound source is X=(x,y,z), and the time difference of sound propagation is Δt i, the speed of sound is c, then the positioning equation of each sensor is ||XP i ||=c·Δt i This equation can be solved by the least square method. Based on the three-dimensional position coordinates, the sound source position information is generated using a second-order dynamic equation, which is described as follows:

[0114]

[0115] Among them, X 0 represents the initial position, V 0 represents the initial velocity, and A represents the acceleration. According to the real-time sensor data, these parameters can be continuously updated to generate dynamic position information.

[0116] Furthermore, the spectrum time series extraction path includes: designing an adaptive filter, the transfer function of the filter is H(f, t) which changes with time t and frequency f. In this embodiment, the spectrum X(f, t) of the signal is obtained by Fourier transform, and the output of the filter can be expressed as Y(f, t) = X(f, t) · H(f, t), wherein H(f, t) is dynamically adjusted according to the real-time signal characteristics and bandwidth requirements to capture the frequency information in the sound signal in real time. Furthermore, a high-order spectrum analysis method is used to capture the dynamic changes in the frequency information. In this embodiment, the Hilbert transform is used to extract the instantaneous value of the frequency:

[0117] X inst (f, t) = H{X(f, t)};

[0118] Wherein, H represents Hilbert transform. Wherein, the change of instantaneous frequency can be expressed by the following formula:

[0119]

[0120] Through this method, the fluctuation pattern of the spectrum over time can be extracted to generate the sound spectrum information.

[0121] Furthermore, the spatial audio reconstruction process performs audio reconstruction based on the sound direction information, the sound source position information and the sound spectrum information to obtain a scene audio generation result, and the specific steps include:

[0122] The sound direction information θ(t), the sound source location information X(t) and the sound spectrum information X inst (f, t) is input into the multi-layer network structure to extract audio features. The specific structure of the multi-layer network structure is:

[0123] Input layer: input sound direction information θ(t), sound source location information X(t) and sound spectrum information X inst (f,t);

[0124] Convolutional layer: convolutions the input features to generate filtered feature maps for detecting specific patterns, temporal changes, and directional features in the spectrum;

[0125] Pooling layer: downsamples the features extracted by the convolutional layer to reduce the dimension of the feature map while retaining the most important feature information;

[0126] Fully connected layer: converts the spatial audio features extracted by convolution and pooling operations into a feature vector of fixed dimension for subsequent audio synthesis and reconstruction;

[0127] Time series modeling layer: Long short-term memory network is introduced to retain and update long-term memory in audio signals through gating mechanism, model the temporal dependency of sound, and feed it back into the model for subsequent audio reconstruction;

[0128] Output layer: It is a fully connected layer that generates the final audio features, which will be used for subsequent audio synthesis.

[0129] Furthermore, based on the audio features, the sound is synthesized using a head-related transfer function to obtain reconstructed audio; audio rendering and binaural rendering are performed on the reconstructed audio to obtain the scene audio generation result.

[0130] Furthermore, the system also includes a sound recognition model construction unit, which constructs a panoramic sound space sound recognition model, including an audio information collection process and a spatial audio reconstruction process; the audio information collection process generates sound direction information, sound source position information and sound spectrum information in the simulation rehearsal data; the spatial audio reconstruction process performs audio reconstruction based on the sound direction information, the sound source position information and the sound spectrum information to obtain a scene audio generation result.

[0131] Specifically, the method for monitoring the reaction time and behavior of the disabled person to the scene audio generation result is as follows: the reaction time is obtained by monitoring the heartbeat and muscle activity changes of the disabled person, and the formula is: reaction time = reaction start time - sound trigger time; using the muscle activity changes of the disabled person, the behavior of the disabled person is obtained, including two aspects: first, judging whether the action of the disabled person corresponds to the sound source; second, judging the action amplitude of the disabled person;

[0132] Based on the behavioral response, the environmental sensitivity is calculated, and the specific process is as follows:

[0133] Assign a weight factor to each of the behavioral actions to indicate the contribution to the environmental sensitivity; the weight factor of whether the action corresponds to the sound source is 0.8, and the weight factor of the action amplitude is 1.2;

[0134] A matching index is defined. If the matching index is equal to 1, it means that the action perfectly matches the sound direction. If the matching index is equal to 0, it means that the action does not match the sound direction at all. According to the actual direction of the action and sound of the disabled person, the matching index is calculated. The calculation formula is: Among them, angle represents the angle between the sound direction and the actual action direction, angle threshold Indicates the set threshold angle. This formula calculates the matching degree based on the angle difference between the direction of the sound and the direction of the disabled person's movement. If the angle difference is less than a certain threshold (for example, 10°), the matching degree is considered to be 1, otherwise the matching degree is attenuated according to the angle difference.

[0135] Defining the range of motion index I magnitude , reflects the intensity of the action, and the action amplitude index is calculated by the head turning angle and walking distance of the disabled person's action. The formula is The maximum possible amplitude is set according to the experimental scenario, which can be the maximum head turning angle or the maximum step length.

[0136] The reaction time t react , the matching index match and the action amplitude index I magnitude , and its corresponding weight factor, calculate the environment sensitivity Environment Sensitivity, the formula is:

[0137]

[0138] According to this formula, the shorter the reaction time, the higher the environmental sensitivity. In addition, the higher the matching degree between the action and the sound source and the larger the movement amplitude, the more sensitive the disabled person is to the sound perception and reaction, thus having a positive effect on improving environmental sensitivity.

[0139] Furthermore, the system also includes a sound recognition result evaluation unit. In the spatial classroom, the disabled person makes a behavioral response based on the scene audio generation result; the behavioral response includes the reaction time and behavioral actions of the disabled person to the scene audio generation result; based on the behavioral response, the environmental sensitivity is calculated.

[0140] Furthermore, the system also includes a sound recognition model retraining unit, which sets an environmental sensitivity threshold. If the environmental sensitivity threshold is not greater than the environmental sensitivity threshold, the corresponding scene is re-simulated in the spatial classroom to retrain the panoramic sound spatial sound recognition model. As shown in Table 2, the experimental data of the reactions of people with different disabilities in several scenarios are shown.

[0141] Table 2 Experimental data display

[0142]

[0143] As can be seen from Table 2, based on the value of environmental sensitivity, if the environmental sensitivity of a disabled person is lower than 0.9, the model will be re-simulated and trained in this scenario.

[0144] Although embodiments of the present invention have been shown and described, it will be appreciated by those skilled in the art that various changes, modifications, substitutions and variations may be made to the embodiments without departing from the principles and spirit of the present invention, and that the scope of the present invention is defined by the appended claims and their equivalents.

Claims

1. A method for training an panoramic sound recognition model, characterized in that: include: Building a venue to form a space classroom, wherein the space classroom includes multiple scenarios that visually impaired people encounter in real life; Conducting simulation exercises on the disabled person in the space classroom to obtain simulation exercise data; Performing data preprocessing on the simulation exercise data; Construct an Atmos spatial sound recognition model, including the audio information collection process and the spatial audio reconstruction process; The audio information collection process generates sound direction information, sound source location information and sound spectrum information in the simulation exercise data; The spatial audio reconstruction process performs audio reconstruction based on the sound direction information, the sound source position information and the sound spectrum information to obtain a scene audio generation result; In the space classroom, the disabled person makes a behavioral response according to the scene audio generation result; the behavioral response includes the disabled person's reaction time and behavioral action to the scene audio generation result; based on the behavioral response, the environmental sensitivity is calculated; An environmental sensitivity threshold is set. If it is not greater than the environmental sensitivity threshold, the corresponding scene is re-simulated in the spatial classroom and the panoramic sound spatial sound recognition model is re-trained.

2. The training method of a panoramic sound spatial sound recognition model according to claim 1, characterized in that: The multiple scenarios are simulated using venues, props and sound equipment; the simulation rehearsal data are environmental sounds that can be received by human ears and are formed by various sounds emitted in the multiple scenarios, including audio information required by the disabled person and unnecessary noise.

3. The training method of a panoramic sound spatial sound recognition model according to claim 1, characterized in that: The audio information collection process generates the sound direction information, sound source location information and sound spectrum information in the simulation exercise data, and the specific steps include: Preprocessing the simulation exercise data; The simulation exercise data are input into the direction timing extraction path, the position timing extraction path and the spectrum timing extraction path respectively; the direction timing extraction path includes: using multiple sensors to receive sound signals and record the signal arrival time, calculating the time difference of the signal arrival time, and using the position difference of multiple sensors to calculate the initial direction of the sound; based on the initial direction, the direction change of the sound source is simulated by a rotation matrix to obtain the direction angle that changes with time, and the direction angle is summarized to obtain the sound direction information; The position timing extraction path includes: calculating the three-dimensional position coordinates of the sound source by using the propagation time of the sound signals received by the multiple sensors using the multi-point positioning technology; based on the three-dimensional position coordinates, generating the sound source position information using the second-order dynamic equation; The spectrum time series extraction path includes: designing an adaptive filter to automatically adjust the bandwidth according to the intensity and spectrum characteristics of the sound signal to capture the frequency information in the sound signal in real time; using a high-order spectrum analysis method to capture the dynamic changes in the frequency information to generate the sound spectrum information.

4. The method for training an ambient sound spatial sound recognition model according to claim 1, characterized in that: The spatial audio reconstruction process performs audio reconstruction based on the sound direction information, the sound source position information and the sound spectrum information to obtain a scene audio generation result, and the specific steps include: Input the sound direction information, the sound source location information and the sound spectrum information into a multi-layer network structure to extract audio features; Based on the audio features, synthesizing the sound using a head-related transfer function to obtain reconstructed audio; Audio rendering and binaural rendering are performed on the reconstructed audio to obtain the scene audio generation result.

5. The method for training an ambient sound spatial sound recognition model according to claim 1, characterized in that: The method for monitoring the reaction time and behavior of the disabled person to the scene audio generation result is as follows: the reaction time is obtained by monitoring the heartbeat and muscle activity changes of the disabled person, and the formula is: reaction time = reaction start time - sound trigger time; Using the changes in the disabled person's muscle activity to obtain the disabled person's behavior and actions includes two aspects: first, determining whether the disabled person's actions correspond to the sound source; second, determining the disabled person's action amplitude; Based on the behavioral response, the environmental sensitivity is calculated, and the specific process is as follows: Assigning a weight factor to each of the behavioral actions to indicate the contribution to the environmental sensitivity; A matching index is defined, where if the matching index is equal to 1, it means that the action perfectly matches the direction of the sound, and if the matching index is equal to 0, it means that the action does not match the direction of the sound at all; and the matching index is calculated according to the action of the disabled person and the actual direction of the sound; defining a movement range index, and calculating the movement range index according to the head turning angle and walking distance of the disabled person; The environmental sensitivity is calculated by comprehensively considering the reaction time, the matching index, the motion amplitude index, and the corresponding weight factors.

6. A training system for an ambient sound spatial sound recognition model, characterized in that: include: A simulated data collection unit is constructed to form a space classroom, wherein the space classroom includes multiple scenarios that visually impaired people encounter in real life; Conducting simulation exercises on the disabled person in the space classroom to obtain simulation exercise data; Performing data preprocessing on the simulation exercise data; A sound recognition model building unit, which builds a panoramic sound space sound recognition model, including an audio information collection process and a spatial audio reconstruction process; The audio information collection process generates sound direction information, sound source location information and sound spectrum information in the simulation exercise data; The spatial audio reconstruction process performs audio reconstruction based on the sound direction information, the sound source position information and the sound spectrum information to obtain a scene audio generation result; A sound recognition result evaluation unit, in the space classroom, the disabled person makes a behavioral response according to the scene audio generation result; the behavioral response includes the disabled person's reaction time and behavioral action to the scene audio generation result; based on the behavioral response, the environmental sensitivity is calculated; The sound recognition model retraining unit sets an environmental sensitivity threshold. If the environmental sensitivity threshold is not greater than the environmental sensitivity threshold, the corresponding scene is re-simulated in the spatial classroom to retrain the panoramic sound spatial sound recognition model.

7. The training system of the panoramic sound spatial sound recognition model according to claim 6, characterized in that: The multiple scenarios are simulated using venues, props and sound equipment; the simulation rehearsal data are environmental sounds that can be received by human ears and are formed by various sounds emitted in the multiple scenarios, including audio information required by the disabled person and unnecessary noise.

8. The training system of the panoramic sound spatial sound recognition model according to claim 6, characterized in that: The audio information collection process generates the sound direction information, sound source location information and sound spectrum information in the simulation exercise data, and the specific steps include: Preprocessing the simulation exercise data; The simulation exercise data are input into the direction timing extraction path, the position timing extraction path and the spectrum timing extraction path respectively; the direction timing extraction path includes: using multiple sensors to receive sound signals and record the signal arrival time, calculating the time difference of the signal arrival time, and using the position difference of multiple sensors to calculate the initial direction of the sound; based on the initial direction, the direction change of the sound source is simulated by a rotation matrix to obtain the direction angle that changes with time, and the direction angle is summarized to obtain the sound direction information; The position timing extraction path includes: calculating the three-dimensional position coordinates of the sound source by using the propagation time of the sound signals received by the multiple sensors using the multi-point positioning technology; based on the three-dimensional position coordinates, generating the sound source position information using the second-order dynamic equation; The spectrum time series extraction path includes: designing an adaptive filter to automatically adjust the bandwidth according to the intensity and spectrum characteristics of the sound signal to capture the frequency information in the sound signal in real time; using a high-order spectrum analysis method to capture the dynamic changes in the frequency information to generate the sound spectrum information.

9. The training system of the panoramic sound spatial sound recognition model according to claim 6, characterized in that: The spatial audio reconstruction process performs audio reconstruction based on the sound direction information, the sound source position information and the sound spectrum information to obtain a scene audio generation result, and the specific steps include: Input the sound direction information, the sound source location information and the sound spectrum information into a multi-layer network structure to extract audio features; Based on the audio features, synthesizing the sound using a head-related transfer function to obtain reconstructed audio; Audio rendering and binaural rendering are performed on the reconstructed audio to obtain the scene audio generation result.

10. The training system of the panoramic sound spatial sound recognition model according to claim 6, characterized in that: The method for monitoring the reaction time and behavior of the disabled person to the scene audio generation result is as follows: the reaction time is obtained by monitoring the heartbeat and muscle activity changes of the disabled person, and the formula is: reaction time = reaction start time - sound trigger time; Using the changes in the disabled person's muscle activity to obtain the disabled person's behavior and actions includes two aspects: first, determining whether the disabled person's actions correspond to the sound source; second, determining the disabled person's action amplitude; Based on the behavioral response, the environmental sensitivity is calculated, and the specific process is as follows: Assigning a weight factor to each of the behavioral actions to indicate the contribution to the environmental sensitivity; A matching index is defined, where if the matching index is equal to 1, it means that the action perfectly matches the direction of the sound, and if the matching index is equal to 0, it means that the action does not match the direction of the sound at all; and the matching index is calculated according to the action of the disabled person and the actual direction of the sound; defining a movement range index, and calculating the movement range index according to the head turning angle and walking distance of the disabled person; The environmental sensitivity is calculated by comprehensively considering the reaction time, the matching index, the motion amplitude index, and the corresponding weight factors.

Citation Information

Patent Citations

  • Method and device for sensing surrounding environment, terminal and storage medium

    CN112927718A

  • Method and apparatus for sound source location detection

    CN113056925A

  • Perception system based on auditory sense and use method thereof

    CN113196390A

  • Construction method of sound classification model, and sound classification method and system

    CN114255783A

  • Sound visualization method and device, equipment, storage medium and program product

    CN116013352A