AI and multi-microphone array-based three-dimensional sound event accurate identification and positioning method

By using a method based on AI and multi-microphone arrays, we collect and process three-dimensional sound signals, build and optimize a three-dimensional sound detection model, and solve the problem of insufficient ability to capture the multi-dimensional distribution characteristics of the three-dimensional sound field. This achieves high-precision three-dimensional sound event recognition and positioning, and enhances the real-time adaptability and robustness of the model.

CN120428168BActive Publication Date: 2025-10-14BEIJING YUANZHI DIGITAL INFORMATION TECHNOLOGY CO LTD

Patent Information

Application Number
CN202510577584.7
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-05-07
Publication Date
2025-10-14
Estimated Expiration
2045-05-07

AI Technical Summary

Technical Problem

The existing technology is unable to capture the multi-dimensional distribution characteristics of the three-dimensional sound field, resulting in reduced classification and positioning accuracy, limited model generalization ability caused by single data training and static environment assumptions, and insufficient real-time adaptability under dynamic noise interference and long-term performance degradation.

Method used

Through a method based on AI and multi-microphone arrays, multi-channel sound signals are collected, frame segmentation and windowing processing and Mel feature extraction are performed to build a three-dimensional sound detection model. Multi-angle virtual sound field data is generated through simulation algorithms for joint training, and dynamic calibration is performed in combination with environmental noise feedback to generate a real-time calibrated three-dimensional sound detection model.

Benefits of technology

It improves the ability to capture the multi-dimensional distribution characteristics of the three-dimensional sound field, enhances the generalization ability and real-time adaptability of the model, solves the problem of reduced classification and positioning accuracy, and ensures high-precision recognition and positioning in dynamic noise environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120428168B_ABST
    Figure CN120428168B_ABST
Patent Text Reader

Abstract

The application relates to the technical field of acoustic signal processing. A three-dimensional sound event accurate identification and positioning method based on AI and a multi-microphone array is provided, which comprises the following steps: acquiring a multi-channel sound signal; performing frame windowing processing on the multi-channel sound signal to obtain time-frequency domain information; generating mel feature information; performing time-space joint feature extraction to generate an initial event category and a three-dimensional direction code; generating multi-angle virtual sound field data by combining the attenuation relationship between the sound source direction and the microphone angle through a simulation algorithm; inputting the real-time collected multi-channel sound signal into an optimized three-dimensional sound detection model for reasoning, and outputting an optimized event category and a three-dimensional direction code; performing dynamic calibration to generate a real-time calibrated three-dimensional sound detection model, so as to solve the problems of insufficient multi-dimensional distribution characteristic capturing capability of a three-dimensional sound field, limited model generalization capability caused by single data training and static environment assumption, and insufficient real-time adaptability under dynamic noise interference and long-term performance attenuation.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of acoustic signal processing technology, and in particular to a method for accurately identifying and locating three-dimensional sound events based on AI and a multi-microphone array. Background Art

[0002] With the rapid development of artificial intelligence and the Internet of Things (IoT), sound event recognition and localization in three-dimensional space are playing an increasingly important role in intelligent security, robotics interaction, virtual reality, and other fields. Accurate sound event detection and directional localization are core technologies for environmental perception, target tracking, and human-computer interaction.

[0003] However, related technologies have the problem of insufficient ability to capture the multi-dimensional distribution characteristics of three-dimensional sound fields, making it difficult to maintain high-precision classification and positioning in complex scenarios; the technical framework relies on single-domain data training and static environment assumptions, which makes it difficult to cope with dynamic noise interference and performance degradation in long-term deployment, limiting the model's generalization ability and real-time adaptability. Summary of the Invention

[0004] Based on this, it is necessary to provide a method for accurately identifying and positioning three-dimensional sound events based on AI and multi-microphone arrays to address the above-mentioned technical problems, so as to solve the problems of reduced classification and positioning accuracy caused by insufficient ability to capture the multi-dimensional distribution characteristics of the three-dimensional sound field in related technologies, limited model generalization ability caused by single data training and static environment assumptions, and insufficient real-time adaptability under dynamic noise interference and long-term performance degradation.

[0005] In the first aspect, the present application provides a method for accurately identifying and locating three-dimensional sound events based on AI and a multi-microphone array, the method comprising:

[0006] The sound events in the three-dimensional space are collected by a preset microphone array system to obtain multi-channel sound signals;

[0007] Perform frame division and windowing processing on the multi-channel sound signal to obtain time-frequency domain information; perform Mel feature extraction based on the time-frequency domain information to generate Mel feature information;

[0008] A 3D sound detection model is constructed based on Mel feature information. The 3D sound detection model performs spatiotemporal joint feature extraction on the Mel feature information to generate the initial event category and 3D direction code.

[0009] Based on single-microphone recording data, a simulation algorithm is used to combine the attenuation relationship between the sound source direction and the microphone angle to generate multi-angle virtual sound field data. This multi-angle virtual sound field data and real data are input into a 3D sound detection model for joint training to generate an optimized 3D sound detection model.

[0010] The multi-channel sound signals collected in real time are input into the optimized 3D sound detection model for inference, and the optimized event categories and 3D direction codes are output. The optimized 3D sound detection model is dynamically calibrated based on environmental noise feedback to generate a real-time calibrated 3D sound detection model.

[0011] Furthermore, based on the single-microphone recording data, a simulation algorithm is used to combine the attenuation relationship between the sound source direction and the microphone angle to generate multi-angle virtual sound field data, including:

[0012] Model the geometric relationship between the incident direction of the sound source and the angle between the microphones for single-microphone recording data, and generate a sound source direction-attenuation relationship mapping table;

[0013] Based on the sound source direction-attenuation relationship mapping table, the amplitude of the single-channel recording signal is nonlinearly and dynamically adjusted to generate a multi-angle virtual sound field signal;

[0014] The multi-channel signal spatial distribution of the multi-angle virtual sound field signal is reconstructed to generate multi-angle virtual sound field data that matches the physical layout of the target microphone array.

[0015] Furthermore, the multi-channel signal spatial distribution of the multi-angle virtual sound field signal is reconstructed to generate multi-angle virtual sound field data that matches the physical layout of the target microphone array, including:

[0016] Based on the physical layout of the target microphone array, the spatial relative parameters between the microphone units are calculated to generate the array spatial distribution model;

[0017] According to the array spatial distribution model, the multi-angle single-channel virtual signals are distributed to each microphone channel according to the spatial attenuation law of the sound wave propagation path to generate the initial multi-channel virtual signal;

[0018] The inter-channel phase consistency of the initial multi-channel virtual signal is adjusted to eliminate the spatial interference error in the virtual sound field data, and generate multi-angle virtual sound field data that matches the physical layout of the target microphone array.

[0019] Furthermore, the amplitude of the single-channel recording signal is nonlinearly and dynamically adjusted based on the sound source direction-attenuation relationship mapping table to generate a multi-angle virtual sound field signal, including:

[0020] Based on the sound source direction-attenuation relationship mapping table, the attenuation weight parameters corresponding to different sound source directions are extracted to generate a set of directional attenuation weight parameters;

[0021] According to the directional attenuation weight parameter set, the amplitude of each time frame in the single-channel recording signal is non-uniformly weighted to generate a directional correlation basic signal;

[0022] Based on the sound field superposition effect of sound waves between adjacent directions, multi-angle energy diffusion compensation is performed on the direction-related basic signal to generate a multi-angle virtual sound field signal.

[0023] Further, the multi-angle virtual sound field data and the real data are input into a three-dimensional sound detection model for joint training to generate an optimized three-dimensional sound detection model, including:

[0024] The multi-angle virtual sound field data and the real data are mixed and enhanced in time-frequency features to generate a mixed training data set;

[0025] Based on the mixed training data set, the spatio-temporal feature extraction layer of the three-dimensional sound detection model is trained in layers to generate a preliminary optimized spatio-temporal joint feature extraction layer;

[0026] Based on the preliminary optimized spatio-temporal joint feature extraction layer, the direction regression branch of the three-dimensional sound detection model is jointly optimized through a direction perception loss function to generate an optimized three-dimensional sound detection model.

[0027] Further, the multi-angle virtual sound field data and the real data are mixed and enhanced in time-frequency features to generate a mixed training data set, including:

[0028] The time-domain waveforms of the multi-angle virtual sound field data and the real data are timestamped and aligned to generate time-domain aligned virtual-real mixed signals;

[0029] Based on the time-domain aligned virtual-real mixed signals, their frequency energy distribution features are extracted and nonlinearly fused to generate frequency domain fusion features;

[0030] The frequency domain fusion features are introduced into a random time-frequency mask to generate a mixed training data set.

[0031] Further, based on environmental noise feedback, the optimized three-dimensional sound detection model is dynamically calibrated to generate a real-time calibrated three-dimensional sound detection model, including:

[0032] The real-time environmental noise is analyzed by feedback to extract the time-frequency distribution features of the noise signal to generate a noise level parameter;

[0033] Based on the noise level parameter, the feature extraction weight of the optimized three-dimensional sound detection model is directionally perceived and dynamically adjusted to generate a noise suppression weight matrix;

[0034] The noise suppression weight matrix is fused into the optimized three-dimensional sound detection model through an incremental learning strategy to generate a real-time calibrated three-dimensional sound detection model.

[0035] In a second aspect, the present application also provides an AI and multi-microphone array based three-dimensional sound event accurate identification and positioning system, which comprises:

[0036] A multi-channel signal acquisition module is used to collect sound events in three-dimensional space through a preset microphone array system to obtain multi-channel sound signals;

[0037] Mel feature extraction module, used to perform frame and window processing on multi-channel sound signals to obtain time-frequency domain information; Mel feature extraction is performed based on the time-frequency domain information to generate Mel feature information;

[0038] The spatiotemporal feature extraction module is used to build a 3D sound detection model based on the Mel feature information, perform spatiotemporal joint feature extraction on the Mel feature information through the 3D sound detection model, and generate the initial event category and 3D direction code;

[0039] The data training and inference module is used to generate multi-angle virtual sound field data based on single-microphone recording data by combining the attenuation relationship between the sound source direction and the microphone angle through a simulation algorithm. The multi-angle virtual sound field data and real data are input into the 3D sound detection model for joint training to generate an optimized 3D sound detection model.

[0040] The dynamic calibration module is used to input the multi-channel sound signals collected in real time into the optimized 3D sound detection model for inference, and output the optimized event category and 3D direction code; the optimized 3D sound detection model is dynamically calibrated based on environmental noise feedback to generate a real-time calibrated 3D sound detection model.

[0041] In a third aspect, the present application further provides a computer device comprising a memory and a processor, wherein the memory stores a computer program, and wherein the processor implements the steps of any method in the first aspect of the present application when executing the computer program.

[0042] In a fourth aspect, the present application further provides a computer-readable storage medium having a computer program stored thereon, characterized in that when the computer program is executed by a processor, the steps of any method in the first aspect of the present application are implemented.

[0043] The technical scheme provided in the application has the following technical effects: by providing an AI and multi-microphone array-based three-dimensional sound event accurate identification and positioning method, the following is performed: a multi-channel sound signal is obtained by collecting sound events in a three-dimensional space through a preset microphone array system; the multi-channel sound signal is subjected to frame windowing processing to obtain time-frequency domain information; Mel feature extraction is performed based on the time-frequency domain information to generate Mel feature information; a three-dimensional sound detection model is constructed based on the Mel feature information, time-space joint feature extraction is performed on the Mel feature information through the three-dimensional sound detection model to generate an initial event category and three-dimensional direction encoding; multi-angle virtual sound field data is generated by combining the attenuation relationship between the sound source direction and the microphone angle based on single microphone recording data through a simulation algorithm; the multi-angle virtual sound field data and real data are input into the three-dimensional sound detection model for joint training to generate an optimized three-dimensional sound detection model; the real-time collected multi-channel sound signal is input into the optimized three-dimensional sound detection model for inference to output an optimized event category and three-dimensional direction encoding; the optimized three-dimensional sound detection model is dynamically calibrated based on environmental noise feedback to generate a real-time calibrated three-dimensional sound detection model, so as to solve the problems of classification and positioning precision decline caused by insufficient multi-dimensional distribution characteristic capture capability of a three-dimensional sound field in related technologies, limited model generalization capability caused by single data training and static environment assumption, and insufficient real-time adaptability under dynamic noise interference and long-term performance decay. BRIEF DESCRIPTION OF DRAWINGS

[0044] In order to more clearly illustrate the technical scheme in the embodiments of the present application or the related art, the drawings needed to be used in the embodiments or the related art description will be briefly introduced. Obviously, the drawings in the following description are only some embodiments of the present application, and other drawings can be obtained by those skilled in the art without creative labor.

[0045] Figure 1 A flow chart of the AI and multi-microphone array-based three-dimensional sound event accurate identification and positioning method in an embodiment of the present application;

[0046] Figure 2 A structural diagram of the AI and multi-microphone array-based three-dimensional sound event accurate identification and positioning system in an embodiment of the present application. DETAILED DESCRIPTION

[0047] In order to make the above-mentioned purposes, features and advantages of the present application more obvious and easy to understand, the specific embodiments of the present application are described in detail below with reference to the drawings. In the following description, a large number of specific details are set forth in order to provide a thorough understanding of the present application. However, the present application can be implemented in many other ways different from those described herein, and those skilled in the art can make similar improvements without departing from the scope of the application. Therefore, the present application is not limited to the specific embodiments disclosed below.

[0048] As shown in Figure 1 The present application provides an AI-based and multi-microphone array three-dimensional sound event accurate identification and positioning method, which comprises:

[0049] S101: Collecting sound events in a three-dimensional space through a preset microphone array system to obtain multi-channel sound signals.

[0050] Specifically, according to the characteristics of the application scenario and the target sound source, a preset array system composed of multiple microphones is set up. The microphone array can adopt linear, circular, spherical or other geometric layouts to adapt to different three-dimensional space sound field distribution requirements. When a sound event occurs in a three-dimensional space, sound waves will propagate from the sound source to each microphone in the microphone array. Due to the time difference and phase difference of sound waves reaching different microphones, each microphone will receive slightly different sound signals. Each microphone converts the received sound waves into electrical signals and converts the analog signals into digital signals through an analog-to-digital converter (ADC). The above-mentioned digital signals are collected and stored in real time to form a multi-channel sound signal dataset. In order to ensure the consistency of the multi-channel signals, the collected signals need to be time-synchronized. In addition, the signals can also be preprocessed, such as filtering to remove noise, gain adjustment, etc., to improve signal quality. The preprocessed multi-channel sound signals are integrated into a unified dataset for subsequent feature extraction and analysis. The signal of each channel contains information about the sound source position and characteristics, which will be used for subsequent sound event identification and positioning. Through the above steps, the preset microphone array system can effectively collect sound events in a three-dimensional space and generate multi-channel sound signals to provide basic data support for subsequent sound event identification and positioning.

[0051] S102: Frame windowing processing is performed on the multi-channel sound signals to obtain time-frequency domain information; Mel feature extraction is performed based on the time-frequency domain information to generate Mel feature information.

[0052] Specifically, the continuous multi-channel sound signal is divided into multiple short-time frames, usually with a frame length of 20-40 milliseconds and a frame shift of 10-20 milliseconds. The purpose of framing is to convert the non-stationary audio signal into a series of short-time stationary signal segments for subsequent processing. A window function (such as Hamming window, Hanning window, etc.) is applied to each short-time frame to reduce spectral leakage. The window function reduces signal discontinuity at the frame boundary by smoothing the start and end points of the frame, thereby improving the accuracy of spectral analysis. The windowed short-time frame is subjected to FFT processing to convert the time-domain signal to the frequency-domain signal. The frequency spectrum data output by FFT contains the frequency components and amplitude information of the signal. The power spectrum of the FFT result is calculated, i.e. the square of the amplitude of each frequency component. The power spectrum reflects the energy distribution of the signal at different frequencies. A set of mel frequency filters is designed to convert the linear frequency scale to the mel frequency scale. The mel frequency scale is more consistent with the perceptual characteristics of the human auditory system and can better capture the characteristics of sound. The power spectrum is passed through the mel filter bank to calculate the energy output of each filter. The above energy values constitute the mel spectrum. The mel spectrum is subjected to DCT processing to extract the mel frequency cepstral coefficients (MFCC). DCT can convert the mel spectrum into a set of cepstral coefficients, of which the first few dimensions contain the main features of the sound. Through the above steps, the mel feature information is extracted from the multi-channel sound signal, providing key feature representation for subsequent sound event recognition and positioning. In addition, by performing frequency domain feature extraction (such as mel spectrum and mel cepstral coefficient calculation) on the time domain signal, the spectral energy distribution and time-varying characteristics of the sound can be effectively captured.

[0053] S103: Construct a three-dimensional sound detection model based on the mel feature information, and perform spatio-temporal joint feature extraction on the mel feature information through the three-dimensional sound detection model to generate an initial event category and a three-dimensional direction code.

[0054] Specifically, a suitable deep learning architecture (such as CNN, LSTM, or their combination) is selected to build the 3D sound detection model. The design of the model needs to consider the ability to process mel feature information and have spatio-temporal joint feature extraction. The extracted mel feature information is used as the input of the model. The mel feature information contains the energy distribution of the sound event in different mel frequency bands, which can effectively represent the spectral characteristics of the sound. The time series processing layer (such as LSTM) is used to capture the change pattern of mel feature on the time axis and extract the time sequence feature of the sound event. The convolution layer or other spatial feature extraction method is used to analyze the distribution pattern of mel feature on different frequency bands and extract the spatial feature of the sound event. The extracted time feature and spatial feature are fused to form the spatio-temporal joint feature. The above content is realized through connection layer, attention mechanism or other feature fusion technology to fully utilize the time sequence and spatial information of the sound event. Based on the fused spatio-temporal joint feature, the sound event is classified or regressed through the full connection layer or other classifier to generate the initial event category. At the same time, the regression branch or special direction encoding module in the model is used to calculate the 3D direction information of the sound event according to the spatio-temporal joint feature to generate the 3D direction encoding. Through the above steps, the 3D sound detection model based on mel feature information can effectively identify and locate the sound event, providing a basis for subsequent analysis and application.

[0055] S104: Based on the single microphone recording data, the multi-angle virtual sound field data is generated by combining the attenuation relationship of the sound source direction and the microphone angle through the simulation algorithm; the multi-angle virtual sound field data and the real data are input into the 3D sound detection model for joint training to generate an optimized 3D sound detection model.

[0056] Specifically, by recording sound events with a single microphone, the original single-channel sound signal is obtained. A geometric relationship model of the angle between the sound source direction and the microphone is established to determine the attenuation law of sound in different directions. According to the sound source direction and attenuation relationship model, the single-channel recording signal is processed to simulate sound signals in different sound source directions. By adjusting the amplitude, phase and other parameters of the signal, multi-angle virtual sound field data is generated to simulate sound sources in different directions. The generated multi-angle virtual sound field data is fused with the real multi-channel sound signal data to ensure the consistency and compatibility of the data. A three-dimensional sound detection model capable of processing multi-channel sound signals is designed and constructed, which should have the ability of spatio-temporal joint feature extraction. The fused virtual and real data are input into the three-dimensional sound detection model. Through the training process of the model, the parameters of the model are optimized so that it can learn the features in the virtual data and the real data at the same time. After multiple iterations of training, the performance of the model in sound event recognition and positioning is verified to ensure the accuracy and robustness of the model. Then the three-dimensional sound detection model optimized by joint training is obtained, which has better performance and adaptability in processing complex sound environments.

[0057] S105: input the real-time collected multi-channel sound signal into the optimized three-dimensional sound detection model for inference, output the optimized event category and three-dimensional direction code; based on the environmental noise feedback, dynamically calibrate the optimized three-dimensional sound detection model to generate a real-time calibrated three-dimensional sound detection model.

[0058] Specifically, the real-time multi-channel sound signal collected serves as input, containing information about sound events in the current environment. This real-time signal is fed into an optimized 3D sound detection model. The model analyzes the input signal using its joint spatiotemporal feature extraction mechanism. The model outputs optimized event categories (i.e., the identified sound event types) and 3D directional codes (i.e., the locations of the sound events in 3D space). Simultaneously, the noise in the real-time environment is monitored and analyzed to extract its time-frequency distribution characteristics. Key characteristic parameters, such as noise level and spectral distribution, are extracted from the noise signal. These parameters serve as feedback to evaluate the performance of the current model in noisy environments. Based on this extracted noise feedback, the 3D sound detection model is dynamically calibrated. The model's feature extraction weights and directional regression parameters are adjusted to adapt to the current noise environment. The calibrated parameters are then updated to the model using an incremental learning strategy. After dynamic calibration, a real-time calibrated 3D sound detection model is generated, which exhibits improved robustness and accuracy in the current noise environment. During model operation, the ambient noise is continuously monitored and dynamically calibrated to ensure the model remains optimized. Through the above steps, it is possible to process multi-channel sound signals in real time, accurately identify and locate sound events, and continuously optimize model performance through environmental noise feedback to adapt to dynamically changing environmental conditions.

[0059] The technical solution provided by the present application includes the following technical effects: by providing a method for accurately identifying and locating three-dimensional sound events based on AI and a multi-microphone array, including: collecting sound events in three-dimensional space through a preset microphone array system to obtain multi-channel sound signals; performing frame and window processing on the multi-channel sound signals to obtain time-frequency domain information; performing Mel feature extraction based on the time-frequency domain information to generate Mel feature information; constructing a three-dimensional sound detection model based on the Mel feature information, performing spatiotemporal joint feature extraction on the Mel feature information through the three-dimensional sound detection model to generate initial event categories and three-dimensional direction codes; based on single-microphone recording data, generating a Multi-angle virtual sound field data; input the multi-angle virtual sound field data and real data into the three-dimensional sound detection model for joint training to generate an optimized three-dimensional sound detection model; input the multi-channel sound signal collected in real time into the optimized three-dimensional sound detection model for inference, and output the optimized event category and three-dimensional direction code; dynamically calibrate the optimized three-dimensional sound detection model based on environmental noise feedback to generate a real-time calibrated three-dimensional sound detection model to solve the problems in related technologies such as the lack of ability to capture the multi-dimensional distribution characteristics of the three-dimensional sound field, resulting in reduced classification and positioning accuracy, limited model generalization ability caused by single data training and static environment assumptions, and insufficient real-time adaptability under dynamic noise interference and long-term performance degradation.

[0060] Further, based on single microphone recording data, a multi-angle virtual sound field data is generated by simulating the attenuation relationship of the angle between the sound source direction and the microphone, including:

[0061] Modeling the geometric relationship between the sound source incident direction and the angle between the microphone for the single microphone recording data, generating a sound source direction-attenuation relationship mapping table;

[0062] Based on the sound source direction-attenuation relationship mapping table, the amplitude of the single-channel recording signal is nonlinearly and dynamically adjusted to generate a multi-angle virtual sound field signal;

[0063] The multi-angle virtual sound field signal is reconstructed to generate a multi-angle virtual sound field data matching the physical layout of the target microphone array.

[0064] Specifically, a single microphone is used to record a sound event to obtain an original single-channel sound signal as the basis data for subsequent processing. The geometric relationship between the sound source incident direction and the microphone is analyzed to determine the angle between the sound source and the microphone in different directions. A corresponding relationship model of the sound source direction and the angle is established to generate a sound source direction-attenuation relationship mapping table. The mapping table describes the attenuation law of the sound signal when the sound source is in different directions. According to the sound source direction-attenuation relationship mapping table, the amplitude of the single-channel recording signal is nonlinearly and dynamically adjusted. By adjusting the amplitude, the sound signal intensity change under different sound source directions is simulated to generate a multi-angle virtual sound field signal. The generated multi-angle virtual sound field signal is reconstructed to generate a multi-angle virtual sound field data matching the physical layout of the target microphone array. Through the above steps, the multi-angle virtual sound field data can be generated based on the single microphone recording data, providing diversified data support for the training of the three-dimensional sound detection model, and enhancing the recognition and positioning ability of the model for different sound source directions.

[0065] Further, the multi-angle virtual sound field signal is reconstructed to generate a multi-angle virtual sound field data matching the physical layout of the target microphone array, including:

[0066] Based on the physical layout of the target microphone array, the spatial relative parameters between each microphone unit are calculated to generate an array spatial distribution model;

[0067] According to the array spatial distribution model, the multi-angle single-channel virtual signal is distributed to each microphone channel according to the spatial attenuation law of the sound wave propagation path to generate an initial multi-channel virtual signal;

[0068] The inter-channel phase consistency of the initial multi-channel virtual signal is adjusted, the spatial interference error in the virtual sound field data is eliminated, and multi-angle virtual sound field data matching the physical layout of the target microphone array is generated.

[0069] Specifically, according to the geometric layout of the target microphone array (such as linear, circular, spherical, etc.), the spatial relative parameters between each microphone unit are calculated, including distance, angle, direction, etc. Based on the calculated spatial relative parameters, a mathematical model describing the spatial distribution of the microphone array is generated. The model is used to guide the subsequent signal distribution and adjustment. The multi-angle single-channel virtual sound field signal is distributed to each channel of the target microphone array according to the spatial attenuation law of the sound wave propagation path. According to the sound source direction and array layout, the signal amplitude of each channel is adjusted to simulate the attenuation characteristics of sound waves propagating in different directions. Through the above distribution process, the initial multi-channel virtual signal is generated, and the signal of each channel reflects the acoustic characteristics of the sound source in the corresponding direction. The phase of the initial multi-channel virtual signal is analyzed to ensure the phase consistency of the signals of each channel. By adjusting the phase difference between the channels, the spatial interference error in the virtual sound field data is eliminated, and the spatial distribution of the signal is more accurate. After amplitude adjustment and phase consistency calibration, multi-angle virtual sound field data matching the physical layout of the target microphone array is generated, ensuring that it is consistent with the real sound field in terms of spatial distribution. Through the above steps, the multi-angle virtual sound field signal can be reconstructed into a multi-channel signal matching the physical layout of the target microphone array, providing high-quality virtual data support for subsequent sound event recognition and positioning.

[0070] Further, based on the sound source direction-attenuation relationship mapping table, the amplitude of the single-channel recording signal is nonlinearly dynamically adjusted to generate a multi-angle virtual sound field signal, including:

[0071] Based on the sound source direction-attenuation relationship mapping table, the attenuation weight parameters corresponding to different sound source directions are extracted to generate a set of direction attenuation weight parameters;

[0072] According to the set of direction attenuation weight parameters, the amplitudes of each time frame in the single-channel recording signal are non-uniformly weighted and adjusted to generate a direction-related basic signal;

[0073] Based on the sound field superposition effect of sound waves between adjacent directions, the multi-angle energy diffusion compensation is performed on the direction-related basic signal to generate a multi-angle virtual sound field signal.

[0074] Specifically, according to the sound source direction-attenuation relationship mapping table, the attenuation weight parameters corresponding to different sound source directions are extracted to generate a set of direction attenuation weight parameters. The mapping table describes the attenuation law of the sound signal when the sound source is in different directions. The attenuation weight parameters of different directions are extracted from the mapping table to form a parameter set. The set includes the attenuation weight corresponding to each direction, which is used for subsequent amplitude adjustment. According to the set of direction attenuation weight parameters, the amplitudes of each time frame in the single-channel recording signal are non-uniformly weighted and adjusted. For each time frame, according to its corresponding sound source direction, the corresponding attenuation weight is applied to adjust the amplitude of the time frame. This non-uniform weighting adjustment can simulate the amplitude change of sound sources in different directions to generate a direction-related basic signal. Considering the sound field superposition effect between adjacent directions, the multi-angle energy diffusion compensation is performed on the direction-related basic signal. By analyzing the energy distribution of the sound field in adjacent directions, the signal is adjusted so that the generated virtual sound field signal can better simulate the diffusion characteristics of sound waves in the real sound field. The compensation can ensure that the energy distribution of the virtual sound field signal in different directions is more accurate, and a multi-angle virtual sound field signal is generated. Through the above steps, based on the sound source direction-attenuation relationship mapping table, the single-channel recording signal is non-linearly and dynamically adjusted to generate a multi-angle virtual sound field signal, which provides diversified data support for subsequent sound event recognition and positioning.

[0075] Further, the multi-angle virtual sound field data and the real data are input into a three-dimensional sound detection model for joint training to generate an optimized three-dimensional sound detection model, including:

[0076] The multi-angle virtual sound field data and the real data are mixed and enhanced in time-frequency features to generate a mixed training data set;

[0077] Based on the mixed training data set, the spatio-temporal feature extraction layer of the three-dimensional sound detection model is trained layer by layer to generate a preliminarily optimized spatio-temporal joint feature extraction layer;

[0078] Based on the preliminarily optimized spatio-temporal joint feature extraction layer, the direction regression branch of the three-dimensional sound detection model is jointly optimized through a direction perception loss function to generate an optimized three-dimensional sound detection model.

[0079] Specifically, the multi-angle virtual sound field data is mixed and enhanced with real data in time-frequency features to generate a mixed training dataset. This step enhances the model's adaptability to different acoustic environments by combining the time-frequency features of virtual and real data. Based on the mixed training dataset, the spatio-temporal feature extraction layer of the three-dimensional sound detection model is trained in a hierarchical adversarial manner. Hierarchical adversarial training is a technique to enhance model robustness by introducing adversarial noise at different levels, enabling the model to better handle complex and variable acoustic signals. Through hierarchical adversarial training, a preliminary optimized spatio-temporal joint feature extraction layer is generated. This step optimizes the model's feature extraction capability in both time and spatial dimensions, enabling it to more accurately capture the spatio-temporal features of sound events. Based on the preliminary optimized spatio-temporal joint feature extraction layer, the direction regression branch of the three-dimensional sound detection model is jointly optimized through a direction perception loss function. The direction perception loss function is used to optimize the model's prediction ability for sound direction, ensuring that the model can accurately locate the direction of sound events. After hierarchical adversarial training and joint optimization of the direction regression branch, an optimized three-dimensional sound detection model is generated. This model has higher accuracy and robustness in handling complex acoustic environments, and can effectively identify and locate sound events in three-dimensional space. Through the above steps, the three-dimensional sound detection model can fully utilize the mixed training of virtual and real data to improve its performance and adaptability in practical applications.

[0080] Further, the multi-angle virtual sound field data and the real data are mixed and enhanced in time-frequency features to generate a mixed training dataset, comprising:

[0081] The time-domain waveforms of the multi-angle virtual sound field data and the real data are timestamped and aligned to generate time-domain aligned virtual-real mixed signals;

[0082] Based on the time-domain aligned virtual-real mixed signals, their frequency energy distribution features are extracted and nonlinearly fused to generate frequency domain fused features;

[0083] Random time-frequency mask perturbation is introduced to the frequency domain fused features to generate a mixed training dataset.

[0084] Specifically, the multi-angle virtual sound field data is time-stamped aligned with the time-domain waveform of the real data, ensuring consistency of both on the time axis. This step generates a time-synchronized virtual data set and a real data set by adjusting the timing reference of the virtual data and the real data, making their timestamps strictly match but maintaining signal independence. Based on the aligned independent data sets, the frequency energy distribution features of the virtual data and the real data are extracted respectively, and a mixed training data set is generated through feature-level nonlinear fusion (such as feature matrix splicing or weighted superposition). This step converts the time-domain signal to the frequency domain signal through frequency domain transformation (such as Fast Fourier Transform), and calculates the energy of each frequency component. The extracted frequency energy distribution features are nonlinearly fused. Nonlinear fusion methods (such as weighted average, exponential fusion, etc.) can better combine the frequency domain features of virtual and real data to generate frequency domain fusion features. Random time-frequency mask perturbation is introduced to the frequency domain fusion features. This step randomly adds masks in the frequency domain to simulate noise and interference in the real environment, increasing the diversity and robustness of the data. Through the above processing, a mixed training data set is generated, which combines the time-frequency features of virtual and real data and enhances the generalization ability of the model through random time-frequency mask perturbation. Through the above steps, multi-angle virtual sound field data and real data can be effectively mixed and enhanced in time-frequency features to generate high-quality mixed training data sets, providing strong support for the training of three-dimensional sound detection models.

[0085] Further, based on the feedback of environmental noise, the optimized three-dimensional sound detection model is dynamically calibrated to generate a real-time calibrated three-dimensional sound detection model, including:

[0086] Feedback analysis is performed on the real-time environmental noise to extract the time-frequency distribution features of the noise signal and generate a noise level parameter;

[0087] Based on the noise level parameter, the feature extraction weight of the optimized three-dimensional sound detection model is dynamically adjusted in the direction of perception to generate a noise suppression weight matrix;

[0088] The noise suppression weight matrix is fused into the optimized three-dimensional sound detection model through an incremental learning strategy to generate a real-time calibrated three-dimensional sound detection model.

[0089] Specifically, feedback analysis is performed on the real-time environmental noise, the time-frequency distribution characteristics of the noise signal are extracted, and the noise level parameters are generated. This step obtains the distribution of noise at different frequencies through spectrum analysis and energy calculation, providing a basis for subsequent model calibration. Based on the noise level parameters, the feature extraction weights of the optimized three-dimensional sound detection model are dynamically adjusted with direction perception to generate a noise suppression weight matrix. This step adjusts the weights of the model so that it can better suppress the influence of noise during the feature extraction process and enhance the recognition ability of the target sound. The noise suppression weight matrix is ​​integrated into the optimized three-dimensional sound detection model through an incremental learning strategy to generate a real-time calibrated three-dimensional sound detection model. This step gradually integrates the new weight matrix into the model through online learning, ensuring that the model can quickly adapt to the new noise environment while maintaining its original performance. Through the above steps, the three-dimensional sound detection model can adapt to changes in environmental noise in real time, improving its robustness and accuracy in complex acoustic environments.

[0090] It should be understood that, although the various steps in the flowcharts involved in the various embodiments described above are displayed in sequence according to the instructions of the arrows, these steps are not necessarily executed in sequence in the order indicated by the arrows. Unless otherwise specified herein, there is no strict order restriction on the execution of these steps, and these steps can be executed in other orders. Moreover, at least a portion of the steps in the flowcharts involved in the various embodiments described above can include multiple steps or multiple stages, and these steps or stages are not necessarily executed and completed at the same time, but can be executed at different times, and the execution order of these steps or stages is not necessarily to be carried out in sequence, but can be executed in turn or alternately with other steps or at least a portion of steps or stages in other steps.

[0091] In one embodiment, if Figure 2 As shown, the present application also provides a three-dimensional sound event precise identification and positioning system 200 based on AI and a multi-microphone array, and the system 200 includes:

[0092] The multi-channel signal acquisition module 201 is used to acquire sound events in a three-dimensional space through a preset microphone array system to obtain a multi-channel sound signal;

[0093] Mel feature extraction module 202 is used to perform frame segmentation and windowing processing on the multi-channel sound signal to obtain time-frequency domain information; perform Mel feature extraction based on the time-frequency domain information to generate Mel feature information;

[0094] The spatiotemporal feature extraction module 203 is used to construct a 3D sound detection model based on the Mel feature information, perform spatiotemporal joint feature extraction on the Mel feature information through the 3D sound detection model, and generate an initial event category and a 3D direction code;

[0095] The data training and inference module 204 is used to generate multi-angle virtual sound field data based on single-microphone recording data by combining the attenuation relationship between the sound source direction and the microphone angle through a simulation algorithm; the multi-angle virtual sound field data and real data are input into the 3D sound detection model for joint training to generate an optimized 3D sound detection model;

[0096] The dynamic calibration module 205 is used to input the multi-channel sound signals collected in real time into the optimized 3D sound detection model for inference, and output the optimized event category and 3D direction code; dynamically calibrate the optimized 3D sound detection model based on environmental noise feedback to generate a real-time calibrated 3D sound detection model.

[0097] Specifically, the multi-channel signal acquisition module 201 collects sound events in three-dimensional space through a preset microphone array system to obtain multi-channel sound signals. The microphone array system can capture sound signals from different directions and provide basic data for subsequent processing. The Mel feature extraction module 202 performs frame and window processing on the collected multi-channel sound signals to obtain time-frequency domain information. Mel feature extraction is performed based on the time-frequency domain information to generate Mel feature information. This step converts the sound signal into Mel-frequency cepstral coefficients (MFCC) through a Mel-frequency filter bank and discrete cosine transform to characterize the spectral characteristics of the sound. The spatiotemporal feature extraction module 203 constructs a three-dimensional sound detection model based on the Mel feature information, and uses the three-dimensional sound detection model to perform spatiotemporal joint feature extraction on the Mel feature information to generate the initial event category and three-dimensional direction encoding. The model uses deep learning technologies such as convolutional neural networks (CNN) and long short-term memory networks (LSTM) to extract the temporal and spatial features of sound events respectively, and fuse them to identify the event category and locate its three-dimensional direction.

[0098] The data training and inference module 204 generates multi-angle virtual sound field data based on single microphone recording data by simulating the attenuation relationship of sound source direction and microphone angle through an algorithm. The multi-angle virtual sound field data and real data are input into a three-dimensional sound detection model for joint training to generate an optimized three-dimensional sound detection model. This step simulates sound signals of different sound source directions to enhance the model's adaptability to different acoustic environments. The dynamic calibration module 205 inputs real-time collected multi-channel sound signals into the optimized three-dimensional sound detection model for inference, and outputs optimized event categories and three-dimensional direction encodings. Based on environmental noise feedback, the optimized three-dimensional sound detection model is dynamically calibrated to generate a real-time calibrated three-dimensional sound detection model. This step monitors environmental noise and adjusts model parameters to ensure the accuracy and robustness of the model in real-time applications. Through the above steps, sound events in a three-dimensional space can be effectively collected, processed and analyzed to achieve accurate recognition and positioning.

[0099] The data training and inference module 204 is further configured to:

[0100] model the geometric relationship between the sound source incident direction and the microphone angle for the single microphone recording data, and generate a sound source direction-attenuation relationship mapping table;

[0101] perform nonlinear dynamic adjustment on the amplitude of the single-channel recording signal based on the sound source direction-attenuation relationship mapping table, and generate multi-angle virtual sound field signals;

[0102] perform multi-channel signal spatial distribution reconstruction on the multi-angle virtual sound field signals to generate multi-angle virtual sound field data matching the physical layout of the target microphone array.

[0103] The data training and inference module 204 is further configured to:

[0104] calculate the spatial relative parameters between each microphone unit based on the physical layout of the target microphone array, and generate an array spatial distribution model;

[0105] According to the array spatial distribution model, the multi-angle single-channel virtual signals are distributed to each microphone channel according to the spatial attenuation law of the sound wave propagation path to generate initial multi-channel virtual signals;

[0106] perform inter-channel phase consistency adjustment on the initial multi-channel virtual signals to eliminate spatial interference errors in the virtual sound field data, and generate multi-angle virtual sound field data matching the physical layout of the target microphone array.

[0107] The data training and inference module 204 is further configured to:

[0108] extract attenuation weight parameters corresponding to different sound source directions based on the sound source direction-attenuation relationship mapping table to generate a set of direction attenuation weight parameters;

[0109] According to the directional attenuation weight parameter set, the amplitude of each time frame in the single-channel recording signal is non-uniformly weighted and adjusted to generate a directional correlation basis signal;

[0110] Based on the sound field superposition effect of sound waves between adjacent directions, the multi-angle energy diffusion compensation is performed on the directional correlation basis signal to generate a multi-angle virtual sound field signal.

[0111] The data training and inference module 204 is further used for:

[0112] The time-frequency feature mixing enhancement is performed on the multi-angle virtual sound field data and the real data to generate a mixed training data set;

[0113] Based on the mixed training data set, the spatio-temporal feature extraction layer of the three-dimensional sound detection model is trained in layers to generate a preliminary optimized spatio-temporal joint feature extraction layer;

[0114] Based on the preliminary optimized spatio-temporal joint feature extraction layer, the directional regression branch of the three-dimensional sound detection model is jointly optimized through a directional perception loss function to generate an optimized three-dimensional sound detection model.

[0115] The data training and inference module 204 is further used for:

[0116] The time stamp alignment processing is performed on the time domain waveforms of the multi-angle virtual sound field data and the real data to generate a time domain aligned virtual-real mixed signal;

[0117] Based on the time domain aligned virtual-real mixed signal, its frequency energy distribution features are extracted and non-linearly fused to generate frequency domain fusion features;

[0118] The frequency domain fusion features are introduced into a random time-frequency mask to generate a mixed training data set.

[0119] The dynamic calibration module 205 is further used for:

[0120] The real-time environmental noise is analyzed in feedback to extract the time-frequency distribution features of the noise signal and generate a noise level parameter;

[0121] Based on the noise level parameter, the feature extraction weight of the optimized three-dimensional sound detection model is dynamically adjusted in directional perception to generate a noise suppression weight matrix;

[0122] The noise suppression weight matrix is fused into the optimized three-dimensional sound detection model through an incremental learning strategy to generate a real-time calibrated three-dimensional sound detection model.

[0123] In one embodiment, the present application further provides a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor implements the steps in the above-mentioned method embodiments when executing the computer program.

[0124] In one embodiment, the present application further provides a computer-readable storage medium having a computer program stored thereon, which implements the steps in the above-mentioned method embodiments when the computer program is executed by a processor.

[0125] In one embodiment, (1) sound signal acquisition and preprocessing: a microphone array system is used to acquire sound events in three-dimensional space, and Mel-frequency cepstral coefficients (MFCC) are used to extract features and obtain intensity information of the sound waveform. (2) three-dimensional sound detection model construction: based on multi-channel three-dimensional sound signals, an AI model with directional recognition capability is established through a deep learning algorithm, which can input multi-channel sound data at one time and output event categories and corresponding direction information. (3) directional positioning expansion and simulation algorithm: different directional positioning algorithms are designed for the sound positioning requirements in the horizontal and vertical directions, and a single-microphone recording simulation algorithm based on the sound positioning direction and the microphone angle attenuation is proposed, which can generate multi-angle sound data sets.

[0126] In one embodiment, (1) Sound signal acquisition: A preset microphone array system is used to acquire sound events and record multi-channel sound signals. (2) Signal preprocessing: Time-frequency analysis and intensity information extraction of sound waveforms are performed using MFCC feature extraction technology. (3) Model training and inference: Based on the designed three-dimensional sound detection model, the model is trained using a training dataset to complete event classification and directional positioning functions. (4) Directional positioning simulation: For single-microphone recording data, the sound positioning direction and the microphone angle attenuation relationship are combined to generate a multi-angle sound dataset.

[0127] In one embodiment, horizontal directional sound event detection is achieved by using a three-microphone array, with the three microphones positioned at a 120-degree angle to each other for directional sound pickup. To capture 360-degree sound events and achieve stereo directional sound event detection, four microphones are used, with an additional microphone facing vertically upward. The sound channel waveform is processed using Mel-frequency cepstral coefficients (MFCC) (or Mel-spectrogram system) to obtain sound feature information and sound intensity information. A multi-channel 3D sound detection model is established and trained to input multi-channel sound information at once. The model outputs the event category and the direction of the sound event. The AI ​​model is characterized by the first layer corresponding to the number of microphones, with input of 3 or 4 channels and output of the sound time type and direction. The output direction of the AI ​​model inference is: for horizontal positioning, taking the first horizontal microphone at 0 degrees as an example, the positioning output is [horizontal (sin, cos) value]. For omnidirectional positioning, taking the first horizontal microphone at 0 degrees as an example, the positioning output is [horizontal angle (cos, sin), vertical angle (cos, sin)]. Horizontal positioning is extended to stereo positioning, and data collection, data processing, AI models, and hardware equipment are all continuous and easy to expand.

[0128] For the device embodiments, since they basically correspond to the method embodiments, the relevant parts can be referred to the partial description of the method embodiments. The device embodiments described above are merely illustrative, wherein the components described as separate parts may or may not be physically separated, and the parts displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules can be selected according to actual needs to achieve the purpose of the disclosed solution. A person of ordinary skill in the art can understand and implement it without expending creative work.

[0129] The above-described embodiments merely represent several implementation methods of the embodiments of the present application. While the descriptions are relatively specific and detailed, they should not be construed as limiting the scope of the patent application. It should be noted that a person skilled in the art may make various modifications and improvements without departing from the concept of the embodiments of the present application, and these modifications and improvements fall within the scope of protection of the embodiments of the present application.

Claims

1. A method for accurately identifying and locating three-dimensional sound events based on AI and a multi-microphone array, characterized by: The method comprises: The sound events in the three-dimensional space are collected by a preset microphone array system to obtain multi-channel sound signals; Performing frame and window processing on the multi-channel sound signal to obtain time-frequency domain information; performing Mel feature extraction based on the time-frequency domain information to generate Mel feature information; Building a 3D sound detection model based on the Mel feature information, performing spatiotemporal joint feature extraction on the Mel feature information using the 3D sound detection model to generate an initial event category and a 3D direction code; Based on single-microphone recording data, a simulation algorithm is used to combine the attenuation relationship between the sound source direction and the microphone angle to generate multi-angle virtual sound field data; the multi-angle virtual sound field data and real data are input into the 3D sound detection model for joint training to generate an optimized 3D sound detection model; The multi-channel sound signal collected in real time is input into the optimized three-dimensional sound detection model for inference, and the optimized event category and three-dimensional direction code are output; the optimized three-dimensional sound detection model is dynamically calibrated based on environmental noise feedback to generate a real-time calibrated three-dimensional sound detection model.

2. The method for accurately identifying and locating three-dimensional sound events based on AI and a multi-microphone array according to claim 1, characterized in that: The method generates multi-angle virtual sound field data based on single microphone recording data by combining the attenuation relationship between the sound source direction and the microphone angle through a simulation algorithm, including: Modeling the geometric relationship between the incident direction of the sound source and the angle between the microphones for the single-microphone recording data to generate a sound source direction-attenuation relationship mapping table; Based on the sound source direction-attenuation relationship mapping table, the amplitude of the single-channel recording signal is nonlinearly and dynamically adjusted to generate a multi-angle virtual sound field signal; Multi-channel signal spatial distribution reconstruction is performed on the multi-angle virtual sound field signal to generate the multi-angle virtual sound field data that matches the physical layout of the target microphone array.

3. The method for accurately identifying and locating three-dimensional sound events based on AI and a multi-microphone array according to claim 2, characterized in that: The reconstructing the multi-channel signal spatial distribution of the multi-angle virtual sound field signal to generate the multi-angle virtual sound field data matching the physical layout of the target microphone array includes: Based on the physical layout of the target microphone array, calculating the spatial relative parameters between the microphone units to generate an array spatial distribution model; According to the array spatial distribution model, the multi-angle virtual sound field signal is distributed to each microphone channel according to the spatial attenuation law of the sound wave propagation path to generate an initial multi-channel virtual signal; Inter-channel phase consistency adjustment is performed on the initial multi-channel virtual signal to eliminate spatial interference errors in the virtual sound field data, and to generate the multi-angle virtual sound field data that matches the physical layout of the target microphone array.

4. The method for accurately identifying and locating three-dimensional sound events based on AI and a multi-microphone array according to claim 2, characterized in that: The nonlinear dynamic adjustment of the amplitude of the single-channel recording signal based on the sound source direction-attenuation relationship mapping table to generate a multi-angle virtual sound field signal includes: Based on the sound source direction-attenuation relationship mapping table, extracting attenuation weight parameters corresponding to different sound source directions to generate a directional attenuation weight parameter set; According to the directional attenuation weight parameter set, the amplitude of each time frame in the single-channel recording signal is non-uniformly weighted adjusted to generate a direction-related basic signal; Based on the sound field superposition effect of sound waves in adjacent directions, multi-angle energy diffusion compensation is performed on the direction-related basic signal to generate the multi-angle virtual sound field signal.

5. The method for accurately identifying and locating three-dimensional sound events based on AI and a multi-microphone array according to claim 1, characterized in that: The step of inputting the multi-angle virtual sound field data and the real data into the three-dimensional sound detection model for joint training to generate an optimized three-dimensional sound detection model includes: Performing mixed time-frequency feature enhancement on the multi-angle virtual sound field data and real data to generate a mixed training data set; Based on the mixed training data set, performing hierarchical training on the spatiotemporal feature extraction layer of the three-dimensional sound detection model to generate a preliminarily optimized spatiotemporal joint feature extraction layer; Based on the preliminarily optimized spatiotemporal joint feature extraction layer, the directional regression branch of the 3D sound detection model is jointly optimized by using a directional perception loss function to generate the optimized 3D sound detection model.

6. The method for accurately identifying and locating three-dimensional sound events based on AI and a multi-microphone array according to claim 5, characterized in that: The step of performing time-frequency feature hybrid enhancement on the multi-angle virtual sound field data and real data to generate a hybrid training data set includes: Performing time stamp alignment processing on the time domain waveforms of the multi-angle virtual sound field data and the real data to generate a time domain aligned virtual-real mixed signal; Based on the time-domain aligned virtual-real mixed signal, extracting its frequency-domain energy distribution characteristics and performing nonlinear fusion to generate frequency-domain fusion features; A random time-frequency mask perturbation is introduced into the frequency domain fusion feature to generate the mixed training data set.

7. The method for accurately identifying and locating three-dimensional sound events based on AI and a multi-microphone array according to claim 1, characterized in that: The dynamically calibrating the optimized three-dimensional sound detection model based on environmental noise feedback to generate a real-time calibrated three-dimensional sound detection model includes: Conduct feedback analysis on real-time environmental noise, extract the time-frequency distribution characteristics of the noise signal, and generate noise level parameters; Based on the noise level parameter, dynamically adjust the feature extraction weights of the optimized three-dimensional sound detection model based on direction perception to generate a noise suppression weight matrix; The noise suppression weight matrix is ​​fused into the optimized three-dimensional sound detection model through an incremental learning strategy to generate the real-time calibrated three-dimensional sound detection model.

8. A three-dimensional sound event precise identification and positioning system based on AI and multi-microphone array, characterized by: The system comprises: A multi-channel signal acquisition module is used to collect sound events in three-dimensional space through a preset microphone array system to obtain multi-channel sound signals; A Mel feature extraction module is used to perform frame segmentation and windowing processing on the multi-channel sound signal to obtain time-frequency domain information; perform Mel feature extraction based on the time-frequency domain information to generate Mel feature information; a spatiotemporal feature extraction module, configured to construct a 3D sound detection model based on the Mel feature information, perform spatiotemporal joint feature extraction on the Mel feature information using the 3D sound detection model, and generate an initial event category and a 3D direction code; A data training and inference module is used to generate multi-angle virtual sound field data based on single-microphone recording data by combining the attenuation relationship between the sound source direction and the microphone angle through a simulation algorithm; the multi-angle virtual sound field data and real data are input into the 3D sound detection model for joint training to generate an optimized 3D sound detection model; The dynamic calibration module is configured to input the multi-channel sound signal collected in real time into the optimized 3D sound detection model for inference, and output optimized event categories and 3D direction codes; dynamically calibrate the optimized 3D sound detection model based on ambient noise feedback to generate a real-time calibrated 3D sound detection model.

9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, wherein: When the processor executes the computer program, the steps of the method for accurately identifying and locating three-dimensional sound events based on AI and a multi-microphone array are implemented as described in any one of claims 1 to 7.

10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by a processor, the steps of the method for accurately identifying and locating three-dimensional sound events based on AI and a multi-microphone array are implemented as described in any one of claims 1 to 7.

Citation Information

Patent Citations

  • Sound source positioning method and system based on polyhedral microphone array

    CN114167356A

  • Sound event positioning and recognition method of residual module based on fusion channel attention mechanism

    CN116631386A

Cited By

  • Interference modeling and multi-source cooperative suppression method and system for hybrid dynamic sound source

    CN121838786A