A deep learning-based sound source position imaging method

By processing and meshing the microphone array signals, and combining them with a convolutional neural network that integrates prior information, the device dependence and noise adaptability problems of existing sound source localization methods are solved, achieving efficient and intuitive sound source location imaging.

CN115375920BActive Publication Date: 2025-11-11CHINA AGRI UNIV
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202210373852.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2022-04-11
Publication Date
2025-11-11
Estimated Expiration
2042-04-11

AI Technical Summary

Technical Problem

Existing sound source localization methods are highly dependent on sound acquisition equipment, have low adaptability to environmental noise, have complex input structures for deep learning networks, and produce sound source location images that are not intuitive.

Method used

By framing, windowing, and filtering the signals collected by the microphone array, a sound image is synthesized. A convolutional neural network that integrates prior information is built using a grid-divided sound source plane as input to perform sound source location localization and imaging. The confidence level is then used to form a sound source location imaging map.

Benefits of technology

It reduces reliance on sound acquisition equipment, enhances adaptability to environmental noise, improves the intuitiveness of sound source location imaging and network training efficiency, and reduces hardware requirements and time complexity.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN115375920B_ABST
    Figure CN115375920B_ABST
Patent Text Reader

Abstract

This invention discloses a deep learning-based sound source location imaging method. The invention learns the spatial location of sound sources from different environments using deep learning and forms a sound source location imaging map. It constructs a dataset using acoustic images (which have lower memory footprint and complexity) as input to the deep learning network, and uses the grid numbers within the sound source plane as the output. A convolutional neural network fusing prior information is established based on the characteristics of the acoustic images. The trained network obtains the confidence level of the presence of a sound source in each grid corresponding to a number within the sound source plane, forming the sound source location imaging map. This invention overcomes the problems of common sound source localization methods being highly dependent on sound acquisition equipment and having poor adaptability to environmental noise, as well as the complex input structure and large number of parameters in some deep learning-based sound source location imaging methods, resulting in poor intuitiveness of the sound source location imaging map.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of sound source location imaging technology, and in particular to a sound source location imaging method based on deep learning. Background Technology

[0002] With the rapid development of technologies such as big data and artificial intelligence, the demand for sound source localization imaging methods is increasing, and their application scenarios are becoming more widespread. Currently, microphones are often used as pickups, and multiple microphones are arranged into arrays of a specific shape to filter the sound in the spatial domain, finally obtaining the relative position of the sound in a plane or space. Microphone array-based sound source localization technology plays a crucial role in scenarios such as speech enhancement, noise detection, and smart farms.

[0003] Currently, sound source localization methods can be categorized into two types based on their localization principles: those based on physical sound field modeling and those based on data-driven modeling. The former mainly includes methods based on time delay estimation, high-resolution spectral estimation, and controllable beamforming, while the latter primarily uses deep learning-based sound source localization methods.

[0004] Sound source localization based on controllable beamforming is a common method for sound source localization and imaging. It involves scanning the time delay difference between each point on the sound source plane and each microphone, calculating the power of the controllable response, and drawing a heatmap based on the power difference at each scanned point. This heatmap is then used as the sound source location image, with the scanned point with the highest power identified as the sound source location. This sound source location imaging method requires calculating the time delay difference for each microphone, which places high demands on the performance of the sensor and data acquisition card. Scanning each point leads to a relatively high time complexity for the algorithm. Furthermore, the algorithm needs to determine the sound source frequency before localization, making it highly susceptible to environmental noise.

[0005] Deep learning-based sound source localization methods are data-driven sound source location estimation methods that require a large amount of data to train the network from input to output. The input and output data determine the type of model. When the output data is coordinate points, the sound source localization model is a regression model, which often leads to difficulty in convergence and high requirements for sampled data. Currently, sound source location estimation models based on classification models mostly use cross-correlation functions, time delays between signals, and spectrograms as model inputs. This method of constructing model input data increases the complexity of the algorithm, but directly using the voltage signal obtained from the microphone array as the model input will result in low model accuracy.

[0006] Deep learning-based sound source localization methods are also influenced by deep network structures. Currently, sound source localization models often employ network structures commonly used in image classification, such as AlexNet, GoogleNet, and ResNet. Networks with more parameters have higher hardware requirements, while networks with fewer parameters have relatively lower accuracy. Summary of the Invention

[0007] The purpose of this invention is to provide a sound source location imaging method based on deep learning, in order to solve the problems mentioned in the background art, such as the strong dependence of existing methods on sound acquisition equipment, low adaptability to environmental noise, and the relatively complex structure of the neural network input of some deep learning-based sound source location imaging methods, resulting in poor intuitiveness of the sound source location imaging map.

[0008] To achieve the above objectives, the embodiments of the present invention provide the following technical solutions:

[0009] This invention provides a complete method for sound source location imaging based on deep learning. During the deep learning network training phase, a microphone array acquires sound signals. The raw signals are then segmented, windowed, and filtered. The filtered array signals are synthesized into the input of a deep convolutional neural network, i.e., a sound image. The plane containing the sound source is then divided into grids to determine the sound source's location number. This number is used as the output to train the network. The network's prediction accuracy and loss function value are then analyzed. After reducing the size to a suitable range, the network is used for sound source localization and imaging, where p(x) i Let q(x) represent the true probability distribution. i The ) represents the predicted probability distribution. In sound source localization and imaging, the array signal acquired by the microphone array is filtered to form an acoustic image. A well-trained network is then used for classification to obtain the classification confidence score and location number. Finally, the confidence score is used as a heat value to form a sound source location imaging map. Confidence score Where x i ,x j This is the output value of the last fully connected layer in the network.

[0010] The method includes the following steps:

[0011] 1) Based on the requirements of sound source location imaging, determine the accuracy of sound source imaging, thereby dividing the limited sound source plane into grids and numbering the resulting grids sequentially.

[0012] 2) Place the sound source at any position in the grid divided in step 1), sample it in different ambient noise, and after sampling all the grids generated in step 1), perform filtering preprocessing on the array sound signal.

[0013] 3) Obtain the preprocessed array acoustic signal as in step 2), and perform framing and windowing processing on it. The frame shift between every two frames is half the frame length, and a rectangular window is used in the windowing process.

[0014] 4) Synthesize an acoustic image from each frame of array acoustic signal obtained in step 3) after framing and windowing. Define y i ,i=1,,…,m, where m is the number of elements in the microphone array, y i Given an array acoustic signal for a given frame, take the values ​​of the first p samples of the first channel of this frame, i.e., y1(1:p), where p = m. 2 As the first row of matrix Q, we also take the values ​​of the first p samples of the second channel of this frame, i.e., y2(1:p), as the second row of matrix pic. After taking the first p samples of all m channels of this frame, we continue to group them into subsequent rows of matrix Q, with p samples per channel. Figure 4 The sampling process and results are as follows when p = 64 and m = 8; this forms a matrix Q of size p × p. Normalizing the value of Q to between 0 and 255, it can be represented as... Where Q′ represents the normalized matrix, uint8 indicates rounding to the nearest integer and adjusting the data structure to 8 bits, and max and min represent the signs of finding the maximum and minimum values. At this point, matrix Q′ can form a grayscale image, which is positioned as a sound image.

[0015] 5) Using the sound image obtained in step 4) as input and the location number obtained in step 1) as output, construct a convolutional neural network (CNN) as a classifier; the CNN network structure includes convolutional layers, pooling layers, and fully connected layers, and obtains the sound source location number through the SoftMax function; the model structure is as follows. Figure 3 As shown. Since the acoustic image obtained by the method in step 4) has extremely strong texture features, its gradient histogram (HOG) feature vector is calculated as prior information for the model and concatenated with the feature vector obtained from the CNN network in the first fully connected layer; finally, it passes through two fully connected layers and the SoftMax function, i.e. The sound source location number is obtained, where C is the length of the last fully connected layer vector, and f y The output value of the fully connected layer. This is the sum of the output values ​​of the fully connected layer.

[0016] 6) Divide the acoustic images and their corresponding location numbers into training and test sets. Use the training set to train the network obtained in step 5). The loss function value generally shows a gradual decreasing trend. After multiple iterations, the loss value basically no longer changes. On the test set, stop training when the network's prediction accuracy meets the requirements for sound source localization, and save the network parameters of the last iteration.

[0017] 7) Obtain a mature convolutional neural network that integrates prior information according to the method in step 6), use a microphone array to collect sound source signals, repeat steps 2 to 4) on the collected array sound signals to obtain sound images, and obtain the location number and confidence level of the output sound source through the mature convolutional neural network that integrates prior information.

[0018] 8) The confidence level obtained in step 7) represents the probability that a sound source exists in the grid divided in step 1). Create a grayscale image corresponding to the sound source plane grid, with a pixel value of 255 and a pixel size equal to the number of horizontal and vertical grids in the sound source plane. Multiply the confidence level corresponding to each grid by the pixel value of the grayscale image to generate the sound source location image.

[0019] Compared with the prior art, the beneficial effects of the present invention are:

[0020] 1) Divide the sound source plane into a grid and use a deep learning network to determine whether there is a sound source in the grid. This grid division method avoids the convergence problem that occurs in general deep learning regression networks, making deep learning networks usable for sound source location imaging.

[0021] 2) Acquiring sound signals under different signal-to-noise ratio environments enriches the diversity of training data, improves the detection robustness of deep learning networks, reduces dependence on sound acquisition equipment, and enhances adaptability to noise in the environment.

[0022] 3) Median filtering of the array signal eliminates the impact of noise from audio acquisition devices such as microphones and data acquisition cards on the array signal, thus improving the reliability of the dataset.

[0023] 4) Compared to commonly used methods such as cross-correlation functions and spectrograms, acoustic-image synthesis methods have lower time and space complexity, and acoustic images occupy less hardware memory. Acoustic images have strong texture characteristics, and they fully consider the time delay difference of sound signals between channels and the energy differences contained in the channels, which is beneficial for deep learning networks to mine data from them.

[0024] 5) Convolutional neural networks that integrate prior information have fewer parameters, as shown in Table 1, and exhibit higher convergence efficiency during network training. Figure 5 As shown in Table 1 and 6. Figure 5 In 6, AlexNet, GoogleNet, and ResNet50 represent the three deep learning classification networks AlexNet, GoogleNet, and ResNet50, respectively; hog_fc represents a multilayer perceptron network using HOG as feature vectors; and hog_cnn is the convolutional neural network that integrates prior information, as mentioned in this invention. Figure 5The vertical axis acc represents the accuracy of the network during the iteration process. Figure 6 In this context, `loss` represents the loss function value of the network during the iterative process. Combined with... Figure 5 According to Table 6 and Table 1, it can be concluded that convolutional neural networks that integrate prior information have faster convergence speed and accuracy while having a lower number of parameters.

[0025] Network Name Parameters AlexNet 14712576 GoogleNet 10500080 ResNet50 23639168 hog_fc 13948320 hog_cnn 14553952

[0026] 6) The sound source location imaging map uses the size of the pixel value to represent the probability of the existence of a sound source in the grid on the sound source plane, which improves the intuitiveness of the sound source location imaging. Attached Figure Description

[0027] To more clearly illustrate the technical solutions in the embodiments of the present invention or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on the provided drawings without creative effort.

[0028] Figure 1 This is a schematic diagram of sound source location imaging based on a microphone array;

[0029] Figure 2 This is a basic block diagram of a deep learning-based sound source location imaging method.

[0030] Figure 3 This is a diagram of the convolutional neural network structure that incorporates prior information used in this invention.

[0031] Figure 4 The method for forming acoustic images used in this invention;

[0032] Figure 5 The classification accuracy of various deep learning networks during the iteration process;

[0033] Figure 6 This represents the loss function values ​​of various deep learning networks during the iteration process.

[0034] In the diagram: 1 microphone, 1 microphone array, 3 sound sources, 4 array plane, 5 sound source plane. Detailed Implementation

[0035] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0036] In this application, the terms "comprising," "including," or any other variations thereof are intended to cover a non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitation, an element defined by the phrase "comprising one..." does not exclude the presence of other identical elements in the process, method, article, or apparatus that includes said element.

[0037] The specific embodiments of the present invention will be described in more detail below with reference to the accompanying drawings. The present invention addresses the shortcomings of traditional sound source localization and imaging methods by proposing a sound source location imaging method based on deep learning. Figure 2 The diagram shows the basic block diagram of a sound source location imaging method based on deep learning. The specific implementation steps of this invention include grid division and numbering of the sound source plane, data acquisition and array signal preprocessing, audio signal framing, sound image synthesis, convolutional neural network fusing prior information, and sound source location imaging.

[0038] The specific implementation process of each step is as follows:

[0039] 1. Grid division and numbering of the sound source plane

[0040] In this embodiment, the positional relationship between the microphone array and the sound source plane is as follows: Figure 1 As shown, the center of the microphone array plane (4) is aligned with the center of the sound source plane (5) to determine the sound source imaging range and the distance between the sound source plane and the array plane. The sound source plane is divided into a grid with consistent length and width, and each grid is numbered, denoted as b. i Each grid cell has one and only one location number.

[0041] 2. Data Acquisition and Array Signal Preprocessing

[0042] In this embodiment, the sound source localization deep learning network has a sufficient training set during the training phase. In environments with signal-to-noise ratios (SNR) of 5dB, 10dB, and 20dB, the sound source in each grid is repeatedly sampled n times using a microphone array.

[0043] In this embodiment, an eight-element ring microphone array is used for sound source localization and imaging. The array sound signal collected by the array is defined as M, which is an 8×k matrix, where k is the number of sampling points of the data collected by each microphone.

[0044] In this embodiment, median filtering refers to replacing the value of any point in a digital sequence with the median of its surrounding points, which can be expressed mathematically as: Y i =Med(fi-v ,…,f i ,…,f i+v ), i∈Z, v=(m-1) / 2, Med is the sign of the average, f i The values ​​in the sequence are m, where m is the size of the filter window, and Y is... i This is the filtered value.

[0045] 3. Audio signal framing

[0046] In this embodiment, the octa-microphone array collects sound signals through the microphones, which are then processed by a data acquisition card to obtain discrete array signals. The array signals are preprocessed using the method described in step 2 of this embodiment, with 512 sampling points taken as one frame. The frame shift is 256 sampling points. Each channel of each frame's array signal can be represented as x. i ,i=1,…,8.

[0047] 4. Sound and image synthesis

[0048] In this embodiment, an eight-element microphone array collects sound signals, and the array sound signals are obtained after being segmented into frames through steps 1 to 3 in this embodiment. i The steps for i = 1, ..., 8 are consistent with step 3 of the embodiment. Take the values ​​of the first 64 sampling points of the first channel of a frame, i.e., x1(1:64), as the first row of matrix pic. Similarly, take the values ​​of the first 64 sampling points of the second channel of a frame, i.e., x2(1:64), as the second row of matrix pic. After taking the first 64 sampling points of all 8 channels of this frame, continue to group them into subsequent rows of matrix pic, with 64 sampling points per channel. The sampling process and results are as follows... Figure 4 As shown. This forms a 64×64 matrix pic. Normalizing the values ​​of pic to between 0 and 255, it can be represented as... Where p represents the normalized matrix, uint8 indicates rounding to the nearest integer and adjusting the data structure to 8 bits, and max and min represent the signs of finding the maximum and minimum values. In this case, matrix p is a grayscale image, which is positioned as a sound image.

[0049] 5. Convolutional Neural Networks that Integrate Prior Information

[0050] In this embodiment, the network input is an audio image, and the network output is the location number of the sound source. The Convolutional Neural Network (CNN), commonly used in image classification, is used as a Multilayer Perceptron (MLP). The CNN network structure includes convolutional layers, pooling layers, and fully connected layers, and the sound source location number is obtained through the SoftMax function.

[0051] In this embodiment, the convolutional layer extracts image features through matrix convolution operations. The convolution operation formula is as follows: Where K is an n×n matrix and x is an m×m matrix. In a CNN network, K is defined as the convolution kernel, which traverses the feature map with a certain stride and outputs the convolved feature matrix.

[0052] In this embodiment, the specific operations of the pooling layer are basically the same as those of the convolutional layer. The difference is that the convolutional kernel of the pooling layer only takes the maximum value, average value, etc. at the corresponding position (max pooling, average pooling), that is, the operation rules between matrices are different, and it does not undergo modification through backpropagation.

[0053] In this embodiment, the fully connected layer acts as a classifier in the entire CNN. The fully connected layer is implemented by convolutional operations. When the convolutional kernel size is 1×1, it transforms the pooling layer into a feature vector of size h×w, where h and w are the height and width of the pooling layer, respectively.

[0054] In this embodiment, the CNN model uses two convolutions and pooling operations, and activates it with the Leaky ReLU activation function, i.e., a = g(x) = max(0, 0.01z), to fully extract the features of the sound image. The CNN feature vector is then obtained through a fully connected layer. The sound image obtained in this embodiment 4 has extremely strong texture features. A Histogram of Oriented Gradients (HOG) feature vector is concatenated in the first fully connected layer as prior information for the network. Finally, the sound source location number is obtained through two fully connected layers and the SoftMax function. The value obtained after the SoftMax function is applied to the fully connected layer is used as the confidence value for each sound source location number, denoted as conf. i .

[0055] 6. Sound source location imaging map

[0056] In this embodiment, a grayscale image is constructed with a pixel size consistent with the number of grids in the sound source plane. The pixel size of the grayscale image corresponds to the number of grids in both the horizontal and vertical dimensions of the sound source plane, and the grayscale value of each pixel is gray. i =255*conf i conf i This is consistent with the confidence value obtained in step 5 of this embodiment. The grayscale image is an imaging map of the sound source location.

[0057] Although specific embodiments and accompanying drawings of the present invention have been disclosed for illustrative purposes to help understand the content of the present invention and implement it accordingly, those skilled in the art will understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features therein; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention. Therefore, the present invention should not be limited to the content disclosed in the preferred embodiments and accompanying drawings.

[0058] The various embodiments in this specification are described in a progressive manner. Similar or identical parts between embodiments can be referred to mutually. Each embodiment focuses on describing the differences from other embodiments. In particular, for system or system embodiments, since they are basically similar to method embodiments, the description is relatively simple, and relevant parts can be referred to the descriptions in the method embodiments. The systems and system embodiments described above are merely illustrative. The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to achieve the purpose of this embodiment according to actual needs. Those skilled in the art can understand and implement this without creative effort.

[0059] Those skilled in the art will further recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, computer software, or a combination of both. To clearly illustrate the interchangeability of hardware and software, the components and steps of the various examples have been generally described in terms of functionality in the foregoing description. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.

[0060] The above description of the disclosed embodiments enables those skilled in the art to make or use the invention. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the invention. Therefore, the invention is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.

Claims

1. A sound source location imaging method based on deep learning, characterized in that, Includes the following steps: Step 1: Determine the accuracy requirements for imaging the sound source location in the spatial plane, and then divide the limited sound source plane into grids and number the resulting grids sequentially. Step 2: Place the sound source at any position in all the grids within the sound source plane, play sound from the sound source in different ambient noise conditions and sample it with a microphone array, and record and store the array sound signal with a signal acquisition device; Step 3: Perform preprocessing on the audio signal, including filtering, framing, and windowing. Step 4: Combine the filtered array signals of each frame into a sound image; Step 5: Using the sound image as the input to the deep learning network and the position number of the sound source in the sound source plane as the output of the deep learning network, a convolutional neural network that integrates prior information is built. Step 6: Divide the sound image and its corresponding location number into a training set and a test set. Use the training set to train the convolutional neural network that integrates prior information. The loss function value generally shows a gradual decreasing trend. After multiple iterations, the loss value basically no longer changes. On the test set, stop training when the network's prediction accuracy reaches the requirements for sound source localization, and save the network parameters of the last iteration. Step 7: Construct a convolutional neural network that fuses prior information and can be used for sound source location imaging based on the network parameters of the last iteration; The sound source signal is collected using a microphone array. Steps 2 to 4 are repeated on the collected array sound signal to obtain a sound image. The location number and confidence level of the output sound source are obtained by using a convolutional neural network that fuses prior information. Step 8: Construct a matrix with a pixel size consistent with the number of grid cells in the sound source plane, where the matrix size matches the number of grid cells in both the horizontal and vertical dimensions of the sound source plane. The value g of each point in the matrix is... i = uint8(255*c i ), c i This is the confidence level value corresponding to each number. uint8 means rounding to the nearest integer and adjusting the data structure to 8 bits. At this point, the matrix can be converted into a grayscale image, also known as a sound source location imaging image.

2. The sound source location imaging method based on deep learning according to claim 1, characterized in that: In step 3, during frame division, the frame shift between two frames is half the frame length, and the multiple channel signals obtained by the microphone array are treated as a whole and simultaneously divided into frames and windowed. During filtering, median filtering is performed on each channel of a frame.

3. The sound source location imaging method based on deep learning according to claim 1, characterized in that: The method for synthesizing the audio-visual image in step 4 is as follows: the data of each channel of a frame is represented as f. i ,i=1,…,mic,where mic is the number of elements in the microphone array; take the values ​​of the first q sampling points of the first channel of a frame, i.e., f1(1:q), where q=mic 2 As the first row of matrix A, the values ​​of the first q samples of the second channel of this frame, i.e., f2(1:q), are taken as the second row of matrix A. After taking the first q samples of each of the mic channels of this frame, the values ​​are grouped into subsequent rows of matrix A with q samples per channel, until a complete frame length is taken. At this point, the frame length is q. 2 / mic; This forms a matrix A of size q×q. Normalizing the values ​​of A to between 0 and 255 can be represented as: Where A' represents the normalized matrix, and max and min represent the signs for finding the maximum and minimum values; at this time, matrix A' can form a grayscale image, which is positioned as a sound image.

4. The sound source location imaging method based on deep learning according to claim 1, characterized in that: Step 5 involves a convolutional neural network structure that integrates prior information, comprising convolutional layers, pooling layers, and fully connected layers. The sound source location is obtained using the SoftMax function. The sound image possesses strong texture features; its gradient histogram (HOG) feature vector is calculated and concatenated as prior information to the feature vector obtained from the CNN network. Finally, two fully connected layers are used, and the SoftMax function is applied to the final layer. The position number of the sound source in the sound source plane is obtained, where C is the length of the last fully connected layer vector, and f y The output value of the fully connected layer. This is the sum of the output values ​​of the fully connected layer.

Citation Information

Patent Citations

  • Binaural sound source location method based on convolutional neural network

    CN109164415A

  • Underwater sound source positioning method based on deep learning

    CN109993280A