An Adaptive Sound Pickup Method and System Based on Multi-Source Sound Localization

By performing feature extraction and sound source positioning model processing on multi-sound source sound data, accurately positioning the target sound source and adjusting the parameters of sound picking equipment, the problem of inaccurate sound source positioning and sound confusion in the existing technology is solved, and a more accurate and fast sound picking effect is achieved.

CN119094937BActive Publication Date: 2025-06-20GUANGZHOU BAOLUN ELECTRONICS CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202411277624.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2024-09-12
Publication Date
2025-06-20
Estimated Expiration
2044-09-12

AI Technical Summary

Technical Problem

The existing multi-microphone sound pickup technology is not accurate enough in complex environments, making it difficult to distinguish multiple sound sources, resulting in sound confusion and signal overlap.

Method used

By extracting the sound data of multiple sound sources, a comprehensive feature vector set is obtained, and a preset sound source positioning model is input for processing, to obtain the target sound source characteristics. According to the target sound source characteristics, the optimal sound pickup device parameters are predicted in the current environment, and the sound pickup device parameters are adjusted to achieve more accurate and fast sound pickup.

Benefits of technology

It realizes accurate positioning of the target sound source when multiple sound sources are sounded simultaneously, improves the accuracy and speed of sound pickup, and reduces sound confusion and signal overlap problems.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119094937B_ABST
    Figure CN119094937B_ABST
Patent Text Reader

Abstract

The present application discloses an adaptive sound pickup method and system based on multi-source localization. The method includes: extracting features from the collected multi-source sound generation data to obtain a comprehensive feature vector set; inputting the comprehensive feature vector set into a preset sound source localization model for processing to obtain target sound source features; according to the target sound source features, obtaining the optimal sound pickup device parameters in the current environment through a preset sound pickup strategy prediction model; and setting the adjustable parameters of the sound pickup device according to the optimal sound pickup device parameters to perform sound pickup. The present application accurately locates the target sound source when multiple sound sources are generating sound simultaneously, and sets the relevant parameters of the sound pickup device according to the information of the target sound source, so as to achieve more accurate and rapid sound pickup.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application belongs to the technical field of microphone sound pickup, and particularly relates to an adaptive sound pickup method and system based on multi-source localization. Background Art

[0002] In current multi-microphone sound pickup technologies, the accuracy of sound source localization is often limited by various factors. For environmental interference, existing processing methods may only perform compensation based on some basic acoustic models. In complex and changing actual environments, this compensation effect is limited, resulting in inaccurate sound source localization and the need to use more complex sound source separation and tracking technologies.

[0003] In the case of multiple sound sources existing simultaneously, existing separation and tracking technologies mainly rely on energy or spectral features for simple separation, making it difficult to distinguish the main sound sources and causing problems of sound confusion. In actual complex sound scenarios, fixed screening means based on simple thresholds or rules are usually adopted, lacking the ability to learn and adapt to complex sound features and unable to adjust in real time following the dynamic changes of sounds. When the sound intensity, frequency, etc. change rapidly, inaccurate screening results occur, and it is impossible to capture and process effective audio in a timely manner. Summary of the Invention

[0004] This application proposes an adaptive sound pickup method and system based on multi-source localization, which can accurately locate the target sound source when multiple sound sources sound simultaneously and achieve more accurate and rapid sound pickup.

[0005] The first aspect of this application provides an adaptive sound pickup method based on multi-source localization, and the method includes:

[0006] Extract features from the collected multi-source sound data to obtain a comprehensive feature vector set;

[0007] Input the comprehensive feature vector set into a preset sound source localization model for processing to obtain target sound source features;

[0008] According to the target sound source features, obtain the optimal sound pickup device parameters in the current environment through a preset sound pickup strategy prediction model;

[0009] Set the adjustable parameters of the sound pickup device according to the optimal sound pickup device parameters to perform sound pickup.

[0010] The above solution first extracts sound source features from multi-source sound data to obtain a comprehensive feature vector set, and then obtains target sound source features through a sound source localization model, realizing the distinction and localization of the target sound source and solving the problem of sound confusion. Then, according to the target sound source features, the optimal sound pickup device parameters in the current environment are obtained, enabling the sound pickup device to more accurately and quickly capture the sound signal emitted by the target sound source.

[0011] In a possible implementation method of the first aspect, feature extraction is performed on the collected multi-source sound data to obtain a comprehensive feature vector set, specifically:

[0012] Extract the sound signal, environmental information, and sound source position information from the multi-source sound data;

[0013] Perform feature extraction on the sound signal to obtain sound features;

[0014] Perform feature extraction on the sound source position information to obtain sound source spatial features and sound source temporal features;

[0015] Based on the same time dimension, align and splice the sound features, sound source spatial features, sound source temporal features, and environmental information in multiple dimensions in sequence to obtain a comprehensive feature vector set.

[0016] The above scheme respectively extracts the sound signal, environmental information, and sound source position information, and accordingly obtains the sound features, and the features of the sound source in time and space. Then, these features are aligned and spliced in time to provide data support for subsequent sound source localization.

[0017] In a possible implementation method of the first aspect, feature extraction is performed on the sound signal to obtain sound features, specifically:

[0018] According to a preset time length, divide the sound signal into several frame signals;

[0019] Convert the frame signal into a Mel spectrogram, and extract the first threshold number of first coefficients through the Mel spectrogram to obtain the energy distribution of the sound signal within a preset frequency range;

[0020] According to the frame signal, calculate the short-time energy of the sound signal to obtain the intensity change of the sound signal within the time length;

[0021] Calculate the zero-crossing rate of each frame signal to obtain the frequency change of the sound signal within the time length;

[0022] According to the first coefficient, short-time energy, and zero-crossing rate, obtain sound features.

[0023] The above scheme respectively uses three different algorithms to quantify the changes in the energy, intensity, and frequency of the sound signal over time, so as to obtain the features of the sound signal. Moreover, based on the sound features, it can provide data support for subsequent sound source localization

[0024] In a possible implementation method of the first aspect, the short-time energy, specifically:

[0025] The calculation formula of the short-time energy is:

[0026]

[0027] In the formula, x() is the said sound signal, N is the frame length of the frame signal, n is the serial number of the frame signal, m is the index variable of the sample points within the current frame signal, and E(n) is the short-time energy of the nth sample point in the frame signal.

[0028] In a possible implementation method of the first aspect, feature extraction is performed on the sound source position information to obtain a sound source spatial feature and a sound source temporal feature. Specifically:

[0029] According to the time difference for each sound source to reach the sound pickup device, a sound source spatial feature is extracted from the sound source position information;

[0030] According to the start time and duration of each sound source, a sound source temporal feature is extracted from the sound source position information.

[0031] The sound source spatial feature and the sound source temporal feature obtained in the above solution can accurately distinguish each sound source when multiple sound sources occur simultaneously and have the same sound characteristics, providing data support for sound source localization.

[0032] In a possible implementation method of the first aspect, the comprehensive feature vector set is input into a preset sound source localization model for processing to obtain a target sound source feature. Specifically:

[0033] Based on the comprehensive feature vector set, sound source localization information, as well as the sound source distortion degree and the sound source confusion degree of the target sound source, are obtained through the sound source localization model;

[0034] According to the sound source localization information, a target signal corresponding to the target sound source is extracted from the multi-sound source sound data, and the target signal is optimized according to the sound source distortion degree and the sound source confusion degree of the target sound source to obtain a target sound source feature.

[0035] In the above solution, first, sound source localization information, as well as the sound source distortion degree and the sound source confusion degree affecting sound source localization, are obtained through the sound source localization model. Then, the signal of the target sound source is extracted based on the sound source localization information, and then the signal is further optimized using the sound source distortion degree and the sound source confusion degree of the target sound source to reduce the error of the target sound source feature and effectively and accurately locate the target sound source.

[0036] In a possible implementation method of the first aspect, the sound source localization model is specifically:

[0037] Among them, the loss function of the sound source localization model has the following expression:

[0038] L total =

[0039] Wlocation L location +W distortion L distortion +W classification L classification ;

[0040] Wherein, L total is the loss function of the sound source localization model, W location is the weight of the sound source localization loss, L location is the sound source localization loss, W distortion is the weight of the sound source distortion degree loss, L distortion is the sound source distortion degree loss, W classification is the weight of the sound source confusion degree loss, L classification is the sound source confusion degree loss.

[0041] In a possible implementation method of the first aspect, according to the target sound source characteristics, the optimal pick-up device parameters in the current environment are obtained through a preset pick-up strategy prediction model, specifically:

[0042] According to the target sound source characteristics, the pick-up strategy prediction model screens out the optimal pick-up strategy that meets the current environment from a preset pick-up strategy library;

[0043] According to the hardware parameters of the pick-up device and the optimal pick-up strategy, the optimal pick-up device parameters are determined.

[0044] The above solution selects the optimal pick-up strategy that can most accurately pick up the target sound source from the existing pick-up strategy library based on the target sound source characteristics, and this optimal pick-up strategy can also adapt to the current environment and minimize the impact of the environment on pick-up.

[0045] In a possible implementation method of the first aspect, according to the optimal pick-up device parameters, the adjustable parameters of the pick-up device are set for pick-up, specifically:

[0046] According to the optimal pick-up device parameters, the pick-up gain parameter, frequency response parameter, sampling rate and noise threshold value of the pick-up device are set.

[0047] In the second aspect of the present application, an adaptive pick-up system based on multi-source localization is provided, and the system includes: a data preprocessing module, a sound source localization module, an optimal pick-up parameter acquisition module, and a pick-up device setting module;

[0048] Among them, the data preprocessing module is used to extract features from the collected multi-source sound data to obtain a comprehensive feature vector set;

[0049] The sound source localization module is used to input the comprehensive feature vector set into a preset sound source localization model for processing, and obtain sound source localization information and sound source error information;

[0050] The optimal pickup parameter acquisition module is used to obtain the optimal pickup device parameters in the current environment through a preset pickup strategy prediction model according to the sound source localization information and the sound source error information;

[0051] The pickup device setting module is used to set a pickup device to pick up sound according to the optimal pickup device parameters. Description of the Drawings

[0052] In order to more clearly illustrate the technical solutions of the present application, the accompanying drawings required for implementation will be briefly introduced below. Obviously, the accompanying drawings in the following description are only some implementations of the present application. For those of ordinary skill in the art, other accompanying drawings can be obtained based on these drawings without creative efforts.

[0053] Figure 1 It is a specific flowchart of an adaptive sound pickup method based on multi - sound - source localization provided by an embodiment of the present application;

[0054] Figure 2 It is a structural diagram of an adaptive sound pickup system based on multi - sound - source localization provided by an embodiment of the present application. Specific Embodiments

[0055] The technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all of the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present application without creative efforts fall within the scope of protection of the present application.

[0056] It should be understood that the step numbers used in the text are only for convenient description and are not used as a limitation on the execution order of the steps.

[0057] First Embodiment

[0058] Existing multi-source localization methods rely solely on simple geometric acoustic models and cannot accurately simulate the complex sound propagation in the actual environment. In a complex acoustic environment, there are likely to be large deviations in source localization. Moreover, when multiple sound sources emit sounds simultaneously and their sound characteristics are similar, it is difficult to accurately distinguish and track each sound source, resulting in chaotic source localization, inability to accurately locate the source position, and problems such as sound confusion, signal overlap, and sound distortion. Therefore, how to adjust the pickup strategy according to the dynamic changes of sound, cope with the diversity and complexity of the sound scene, and achieve more accurate multi-source localization is an important research direction in the embodiments of this application.

[0059] As Figure 1 shown, Figure 1 FIG. is a schematic flow chart of a specific adaptive pickup method based on multi-source localization provided by an embodiment of this application. The adaptive pickup method based on multi-source localization in this embodiment includes steps S1 to S4, which are described in detail as follows:

[0060] Step S1: Extract features from the collected multi-source sound data to obtain a comprehensive feature vector set.

[0061] In the embodiments of this application, first, a large amount of multi-source sound data of multiple sound sources emitting sounds simultaneously is collected through a microphone array, and detailed environmental parameter annotations and source position information corresponding to the environment are marked on the multi-source sound data.

[0062] Among them, the marked multi-source sound data includes data such as environmental temperature, environmental humidity, room size, relevant influencing materials (such as wall material, floor material, and ceiling material, etc.), time difference of the sound source reaching different microphones, intensity difference of the sound source reaching different microphones, and quantization value of the distortion degree.

[0063] Preprocess the marked multi-source sound data, check and remove outliers and missing values in the data, and then extract features from the preprocessed multi-source sound data to obtain sound signals, environmental information, and source position information respectively. Among them, the sound signal is the original sound signal collected by the microphone array.

[0064] Next, analyze the sound signal to obtain the changes in the energy, intensity, and frequency of the sound signal over time, so as to obtain the characteristics of the sound signal. Three main methods are used for feature extraction of the sound signal: Mel spectrum method, calculation of short-time energy, and calculation of zero-crossing rate. And before performing feature extraction, the sound signal also needs to be divided into several frame signals according to a preset time length.

[0065] The first method for extracting features of a sound signal is the Mel spectrum method, which is mainly used to determine the energy distribution of the sound signal within a preset frequency range and can effectively describe the key information of the sound. The specific process is as follows: First, each frame signal is optimized using an application window function to reduce the discontinuity at both ends of each frame signal. Then, the windowed frame signal is subjected to a fast Fourier transform to convert the time-domain signal into a frequency-domain signal to obtain the spectrum. Then, the linear frequency scale in the spectrum is converted into a Mel frequency scale, and a set of triangular filters evenly distributed on the Mel frequency scale are used to filter the spectrum, resulting in the Mel spectrum. Next, the logarithm of the Mel spectrum is taken to simulate the non-linear perception of sound intensity by the human ear, obtaining the logarithmic Mel spectrum. Finally, the discrete cosine transform is performed on the logarithmic Mel spectrum to obtain a number of Mel frequency cepstral coefficients, and the first threshold number of coefficients are extracted from the Mel frequency cepstral coefficients as the sound features.

[0066] It is worth mentioning that, generally, the first 12 to 16 coefficients in the Mel frequency cepstral coefficients can contain the main features and most of the energy of the sound signal, can effectively describe the key information of the sound, and retaining fewer coefficients can reduce the data dimension and computational complexity without excessive loss of information. Therefore, in the embodiments of this application, at most the first 16 Mel frequency cepstral coefficients are taken as the sound features. In addition, these Mel frequency cepstral coefficients are sorted from large to small according to the energy magnitude.

[0067] Optionally, in the case where the scenario is relatively simple, such as in a small meeting room, the first 12 Mel frequency cepstral coefficients can be taken. The 12 Mel frequency cepstral coefficients can better capture the key features, accurately identify them, and reduce the calculation time. In a complex scenario, such as an abnormal multilingual academic conference, the speaker may speak with various accents and speaking speeds. To more accurately capture these complex speech features and improve the recognition accuracy, 16 Mel frequency cepstral coefficients are selected.

[0068] Exemplarily, the application window function used in the embodiments of this application is the Hamming window.

[0069] The second method for extracting features of a sound signal is to calculate the short-time energy of the sound signal to obtain the intensity change of the sound signal within a certain time length. The specific calculation formula is as follows:

[0070]

[0071] In the formula, x() is the sound signal, N is the frame length of the frame signal, n is the serial number of the frame signal, m is the index variable of the sample points within the current frame signal, and E(n) is the short-time energy of the nth sample point in the frame signal.

[0072] Among them, short-time energy is an independent voice feature parameter, mainly used in aspects such as monitoring the start and end of voice signals and intensity evaluation. By traversing from 0 to N-1, the traversal of a frame signal is completed, enabling the formula to sum the squares of each sample point within the current frame and calculate the short-time energy of this frame.

[0073] The third method for extracting features of voice signals is to calculate the zero-crossing rate of each frame signal to obtain the frequency change of the voice signal within a certain time length. The specific process is as follows: For each frame signal, count the number of times the frame signal crosses the zero level from positive to negative or from negative to positive. That is, for a frame signal, if the signs of two adjacent sampling points are different, it is considered that a zero-crossing has occurred. The zero-crossing rate can be expressed as the ratio of the number of zero-crossings to the frame length. By calculating the zero-crossing rate, information about the frequency change of the voice signal and the distinction between voiceless and voiced sounds can be obtained.

[0074] Based on the above three voice feature extraction methods, the voice features corresponding to the voice signal can be obtained.

[0075] Then, feature extraction is performed on the sound source position information to obtain the sound source spatial feature and the sound source time feature. Specifically: According to the time difference when each sound source reaches the microphone array, the sound source spatial feature is extracted from the sound source position information. According to the start time and duration of each sound source, the sound source time feature is extracted from the sound source position information.

[0076] After obtaining the voice features, the sound source spatial features, the sound source time features, and the environmental information, these data are first aligned in time and then aligned in other data dimensions. After obtaining the aligned data, based on the same time dimension, these different types of feature data are spliced into several unified vectors according to the alignment result to obtain a comprehensive feature vector set.

[0077] Exemplarily, the voice feature is a vector of length n, the sound source spatial feature is a vector of length m, the sound source time feature is a vector of length p, and the environmental information is a vector of length q. These data are sequentially spliced into a comprehensive feature vector of length n+m+p+q.

[0078] Step S2, input the comprehensive feature vector set into a preset sound source localization model for processing to obtain the target sound source features.

[0079] In the embodiment of the present application, first inputting the comprehensive feature vector set into the trained sound source localization model can obtain the sound source localization information, as well as the sound source distortion degree and the sound source confusion degree of multiple sound sources. Among them, an additional fully connected layer is added after the TCN layer of the sound source localization model to achieve multiple data outputs of the model.

[0080] Among them, the loss function of the sound source localization model is expressed as:

[0081] L total =W location L location +W distortion L distortion +W classification L classification ;

[0082] In the formula, L total is the loss function of the sound source localization model, W location is the weight of the sound source localization loss, L location is the sound source localization loss, W distortion is the weight of the sound source distortion degree loss, L distortion is the sound source distortion degree loss, W classification is the weight of the sound source confusion degree loss, L classification is the sound source confusion degree loss.

[0083]

[0084] In the formula, N is the total number of samples, x pred,i , y pred,i , z pred,i are the position coordinates of the sound source predicted by the model in the x, y, and z directions in the i-th sample, respectively, and x true,i , y true,i , z true,i are the true position coordinates of the sound source in the x, y, and z directions in the i-th sample, respectively.

[0085]

[0086] In the formula, distortion pred,i is the degree of sound source distortion predicted by the model in the i-th sample, and distortion true,i is the true value of the degree of sound source distortion in the i-th sample.

[0087]

[0088] In the formula, C is the total number of categories classified by the model, y true,i,c is the true probability that the i-th sample belongs to category c. If the sample belongs to category c, this value is 1; otherwise, it is 0; y pred,i,c is the probability that the model predicts that the i-th sample belongs to category c.

[0089] Finally, based on the output result of the model, the sound source localization information, the sound source distortion degree, and the sound source confusion degree of the target sound source are obtained. Then, according to the sound source localization information, the target signal corresponding to the target sound source is extracted from the multi-source sound data, and the target signal is optimized according to the sound source distortion degree and the sound source confusion degree of the target sound source to obtain the target sound source feature.

[0090] Optionally, in the embodiment of the present application, the beamforming technology is used to extract the sound source localization information of the target sound source from the sound source localization information output by the model, then the fast Fourier transform is performed on the sound source localization information of the target sound source to obtain the corresponding target spectrum, and the target spectrum is used to normalize the sound source distortion degree and the sound source confusion degree to obtain the sound source distortion degree and the sound source confusion degree of the target sound source.

[0091] In addition, the embodiment of the present application collects a large amount of historical multi-source sound data under different environmental conditions, and details the corresponding environmental parameters and the corresponding sound source position information for each data sample, and converts the historical multi-source sound data into a historical feature vector set. The historical feature vector set is divided into a training set, a validation set, and a test set according to the ratio of 8:1:1. The training set is input into the initial time-domain convolutional model and trained using the ReLu activation function. The validation set and the test set are used to evaluate the model performance and optimize the model according to the training results to obtain a sound source localization model that can meet the prediction accuracy.

[0092] Step S3, according to the target sound source feature, the optimal pick-up device parameters in the current environment are obtained through a preset pick-up strategy prediction model.

[0093] In the embodiment of the present application, the target sound source feature is first preprocessed, and then the preprocessed target sound source feature is input into the pick-up strategy prediction model. According to the provided pick-up strategy library, the optimal pick-up strategy that meets the current environment is selected from it.

[0094] Among them, the data in the pick-up strategy library is the optimal pick-up strategy made according to the relevant historical data extraction or the professional knowledge and experience of relevant professionals for the relevant situations of different sound source localization information, distortion degree quantization values, and sound source confusion degrees.

[0095] Optionally, the embodiment of the present application collects a large amount of historical sound source features, divides the historical sound source features into a training set, a test set, and a validation set according to the ratio of 7:2:1 to train the random forest model, and obtains the pick-up strategy prediction model.

[0096] Finally, based on the hardware parameters of the pick-up device and the optimal pick-up strategy, the optimal pick-up device parameters are determined.

[0097] Step S4: Set the adjustable parameters of the sound pickup device according to the optimal sound pickup device parameters for sound pickup.

[0098] In the embodiment of the present application, according to the obtained optimal sound pickup device parameters, adjust the sound pickup gain parameter, frequency response parameter, sampling rate, noise gate threshold, compression ratio, equalizer parameter, etc. of the sound pickup device to achieve the adaptive sound pickup of the sound pickup device, accurately locate the target sound source when multiple sound sources sound simultaneously, and achieve more accurate and rapid sound pickup.

[0099] Implementing the embodiment of the present application has the following beneficial effects:

[0100] In the embodiment of the present application, the source feature extraction is performed from the multi-source sound data to obtain a comprehensive feature vector set. Among them, the sound feature, the features of the sound source in space and time, and the environmental information are respectively extracted, and the source features are determined according to these data. When multiple sound sources occur simultaneously and the sound features are the same, each sound source is accurately distinguished, providing data support for sound source localization. Then, the target sound source features are obtained through the sound source localization model, and the sound source localization information obtained through the prediction and correction of the sound source distortion degree and the sound source confusion degree is used to realize the distinction and localization of the target sound source, solving the problem of sound confusion. Then, according to the target sound source features, the optimal sound pickup device parameters in the current environment are obtained, so that the sound pickup device can capture the sound signal emitted by the target sound source more accurately and quickly.

[0101] Second Embodiment

[0102] Furthermore, in order to execute the adaptive sound pickup system based on multi-source localization corresponding to the above method embodiment to achieve the corresponding functions and technical effects, Figure 2 A structural diagram of an adaptive sound pickup system based on multi-source localization is provided. For the convenience of description, only the parts related to this embodiment are shown. The adaptive sound pickup system based on multi-source localization provided by the embodiment of the present application includes:

[0103] A data preprocessing module 201, configured to perform feature extraction on the collected multi-source sound data to obtain a comprehensive feature vector set.

[0104] In the embodiment of the present application, first, the sound signal, environmental information, and sound source position information are respectively extracted from the multi-source sound data, and then feature extraction is performed on the sound signal to obtain sound features; feature extraction is performed on the sound source position information to obtain sound source spatial features and sound source time features. Finally, based on the same time dimension, the sound features, sound source spatial features, sound source time features, and environmental information are aligned and spliced in multiple dimensions in sequence to obtain a comprehensive feature vector set, providing data support for subsequent sound source localization.

[0105] The sound source localization module 202 is configured to input the comprehensive feature vector set into a preset sound source localization model for processing, so as to obtain sound source localization information and sound source error information.

[0106] In the embodiment of the present application, the comprehensive feature vector set is input into the sound source localization model for data prediction to obtain sound source localization information, the degree of sound source distortion, and the degree of sound source confusion. Then, the localization information of the target sound source is separated from the sound source localization information, and based on the localization information, the degree of sound source distortion and the degree of sound source confusion of the target sound source are determined. Finally, the target signal corresponding to the target sound source is extracted from the multi-source sound data, and the target signal is optimized according to the degree of sound source distortion and the degree of sound source confusion of the target sound source to obtain the target sound source feature.

[0107] The optimal pickup parameter acquisition module 203 is configured to obtain the optimal pickup device parameters in the current environment through a preset pickup strategy prediction model according to the sound source localization information and the sound source error information.

[0108] In the embodiment of the present application, according to the target sound source feature, the pickup strategy prediction model screens out the optimal pickup strategy that meets the current environment from a preset pickup strategy library; and determines the optimal pickup device parameters according to the hardware parameters of the pickup device and the optimal pickup strategy.

[0109] Among them, the pickup strategy prediction model is constructed by collecting a large number of historical sound source features and dividing the historical sound source features into a training set, a test set, and a validation set according to a ratio of 7:2:1 to train a random forest model.

[0110] The pickup device setting module 204 is configured to set the pickup device to pick up sound according to the optimal pickup device parameters.

[0111] In the embodiment of the present application, according to the obtained optimal pickup device parameters, the pickup gain parameter, frequency response parameter, sampling rate, noise gate threshold, compression ratio, equalizer parameter, etc. of the pickup device are adjusted to achieve the adaptive pickup of the pickup device, accurately locate the target sound source when multiple sound sources sound simultaneously, and achieve more accurate and rapid pickup.

[0112] In some embodiments, the data preprocessing module 201 further includes:

[0113] A voice feature extraction unit, configured to segment the voice signal into a plurality of frame signals according to a preset time length; convert the frame signals into Mel spectrograms, extract the first threshold number of first coefficients through the Mel spectrograms, so as to obtain the energy distribution of the voice signal within a preset frequency range; calculate the short-time energy of the voice signal according to the frame signals, so as to obtain the intensity change of the voice signal within the time length; calculate the zero-crossing rate of each frame signal, so as to obtain the frequency change of the voice signal within the time length; and obtain voice features according to the first coefficients, short-time energy and zero-crossing rate.

[0114] Specifically, it mainly analyzes the voice signal to obtain the changes of the energy, intensity and frequency of the voice signal over time, so as to obtain the features of the voice signal. Three methods are mainly used for the feature extraction of the voice signal: the Mel spectrogram method, calculating short-time energy, and calculating the zero-crossing rate. Before performing feature extraction, the voice signal needs to be divided into a plurality of frame signals according to a preset time length.

[0115] The first method for the feature extraction of the voice signal is the Mel spectrogram method, which is mainly used to determine the energy distribution of the voice signal within a preset frequency range and can effectively describe the key information of the voice. The specific process is as follows: First, each frame signal is optimized by applying a window function to reduce the discontinuity at both ends of each frame signal. Then, the windowed frame signals are subjected to a fast Fourier transform to convert the time-domain signal into a frequency-domain signal to obtain a spectrogram. Then, the linear frequency scale in the spectrogram is converted into a Mel frequency scale, and the spectrogram is filtered using a set of triangular filters evenly distributed on the Mel frequency scale to obtain a Mel spectrogram. Next, the logarithm of the Mel spectrogram is taken to simulate the non-linear perception of the human auditory system to the voice intensity, obtaining a logarithmic Mel spectrogram. Finally, the logarithmic Mel spectrogram is subjected to a discrete cosine transform to obtain a plurality of Mel frequency cepstral coefficients, and the first threshold number of coefficients are extracted from the Mel frequency cepstral coefficients as voice features.

[0116] It is worth mentioning that generally, the first 12 to 16 coefficients in the Mel frequency cepstral coefficients can contain the main features and most of the energy of the voice signal, can effectively describe the key information of the voice, and retaining fewer coefficients can reduce the data dimension and computational complexity without excessive loss of information. Therefore, in the embodiments of the present application, at most the first 16 Mel frequency cepstral coefficients are taken as voice features. In addition, these Mel frequency cepstral coefficients are sorted from large to small according to the energy magnitude.

[0117] Optionally, in the case of a relatively simple scenario, such as when used in a small meeting room, the first 12 Mel-frequency cepstral coefficients can be taken. The 12 Mel-frequency cepstral coefficients can better capture key features, accurately identify them, and reduce calculation time. In a complex scenario, such as an abnormal multilingual academic conference, the speaker may speak with various accents and speech rates. To more accurately capture these complex speech features and improve the recognition accuracy, 16 Mel-frequency cepstral coefficients are selected.

[0118] Exemplarily, the application window function used in the embodiments of the present application is the Hamming window.

[0119] The second method for extracting features of the sound signal is to calculate the short-time energy of the sound signal to obtain the intensity change of the sound signal within a certain time length. The specific calculation formula is as follows:

[0120]

[0121] In the formula, x() is the sound signal, N is the frame length of the frame signal, n is the serial number of the frame signal, m is the index variable of the sample point in the current frame signal, and E(n) is the short-time energy of the nth sample point in the frame signal.

[0122] Among them, the short-time energy is an independent sound feature parameter, mainly used in aspects such as the monitoring of the start and end of the sound signal and the intensity evaluation. By traversing from 0 to N-1, the traversal of a frame signal is completed, enabling the formula to sum the squares of each sample point in the current frame and calculate the short-time energy of this frame.

[0123] The third method for extracting features of the sound signal is to calculate the zero-crossing rate of each frame signal to obtain the frequency change of the sound signal within a certain time length. The specific process is as follows: For each frame signal, count the number of times the frame signal crosses the zero level from positive to negative or from negative to positive. That is, for a frame signal, if the signs of two adjacent sampling points are different, it is considered that a zero crossing has occurred. The zero-crossing rate can be expressed as the ratio of the number of zero crossings to the frame length. By calculating the zero-crossing rate, information about the frequency change of the sound signal and the distinction between voiceless and voiced sounds can be obtained.

[0124] Based on the above three methods for extracting sound features, the sound features corresponding to the sound signal can be obtained.

[0125] The sound source feature extraction unit is used to extract the sound source spatial features from the sound source position information according to the time difference of each sound source reaching the sound pickup device; and extract the sound source time features from the sound source position information according to the start time and duration of each sound source.

[0126] Extract the sound source spatial features from the sound source position information according to the time difference of each sound source reaching the microphone array. Extract the sound source temporal features from the sound source position information according to the start time and duration of each sound source.

[0127] Specifically, after obtaining the sound features, sound source spatial features, sound source temporal features, and environmental information, align these data in time first, and then align them in other data dimensions. After obtaining the aligned data, based on the same time dimension, splice these different types of feature data into several unified vectors according to the alignment result to obtain a comprehensive feature vector set.

[0128] Exemplarily, the sound feature is a vector of length n, the sound source spatial feature is a vector of length m, the sound source temporal feature is a vector of length p, and the environmental information is a vector of length q. Splice these data in sequence to form a comprehensive feature vector of length n + m + p + q.

[0129] In some embodiments, the sound source localization module 202 further includes:

[0130] The sound source localization information of the target sound source can be extracted from the sound source localization information output by the model using beamforming technology, and then the fast Fourier transform is performed on the sound source localization information of the target sound source to obtain the corresponding target spectrum, and the target spectrum is used to normalize the sound source distortion degree and the sound source confusion degree to obtain the sound source distortion degree and the sound source confusion degree of the target sound source.

[0131] Implementing the embodiments of the present application has the following beneficial effects:

[0132] In the embodiments of the present application, relevant environmental data and sound source data are first obtained in an environment of multiple scenarios, and environmental information such as temperature, humidity, spatial size, and material is comprehensively considered. The distortion degree and confusion degree in the relevant environment are predicted through a sound source localization model to verify the predicted sound source localization, reduce the prediction error, and output accurate sound source localization information. Then, based on the sound pickup strategy prediction model and the preset sound pickup strategy library, the best sound pickup screening conditions are accurately predicted to achieve adaptive real-time adjustment, accurately locate the target sound source when multiple sound sources are sounding simultaneously, and achieve more accurate and rapid sound pickup.

[0133] The above specific embodiments have further detailed the purpose, technical solution, and beneficial effects of the present application. It should be understood that the above is only the specific embodiments of the present application and is not used to limit the protection scope of the present application. It is particularly pointed out that for those skilled in the art, any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present application shall be included in the protection scope of the present application.

Claims

1. An adaptive sound pickup method based on multi-sound source localization, characterized in that: include: The collected multi-source sound data are subjected to feature extraction to obtain a comprehensive feature vector set, specifically: extracting sound signals, environmental information and sound source location information from the multi-source sound data; extracting features from the sound signals to obtain sound features; extracting features from the sound source location information to obtain sound source spatial features and sound source temporal features; Based on the same time dimension, the sound features, sound source spatial features, sound source temporal features and environmental information are sequentially aligned and spliced ​​in multiple dimensions to obtain a comprehensive feature vector set; The process of acquiring the sound feature is as follows: according to a preset time length, the sound signal is divided into a plurality of frame signals; the frame signal is converted into a Mel spectrum, and a first threshold number of first coefficients are extracted through the Mel spectrum to obtain the energy distribution of the sound signal within a preset frequency range; wherein the Mel spectrum is logarithmically and discrete cosine transformed in sequence to obtain a plurality of Mel-frequency cepstral coefficients, and the first first threshold number of coefficients are selected from the Mel-frequency cepstral coefficients as the first coefficient; according to the frame signal, the short-time energy of the sound signal is calculated to obtain the intensity change of the sound signal within the time length; the zero-crossing rate of each frame signal is calculated to obtain the frequency change of the sound signal within the time length; the sound feature is obtained according to the first coefficient, the short-time energy and the zero-crossing rate; The calculation formula of the short-time energy is: Wherein, x() is the sound signal, N is the frame length of the frame signal, n is the sequence number of the frame signal, m is the index variable of the sample point in the current frame signal, and E(n) is the short-time energy of the nth sample point in the frame signal; Inputting the comprehensive feature vector set into a preset sound source localization model for processing to obtain the target sound source features; According to the characteristics of the target sound source, the optimal sound pickup device parameters in the current environment are obtained through the preset sound pickup strategy prediction model; According to the optimal sound pickup device parameters, adjustable parameters of the sound pickup device are set to pick up sound.

2. The adaptive sound pickup method based on multi-sound source localization according to claim 1, characterized in that: The feature extraction of the sound source position information to obtain the sound source spatial feature and the sound source temporal feature is specifically as follows: Extracting the sound source spatial features from the sound source position information according to the time difference between each sound source reaching the sound pickup device; According to the start time and duration of each sound source, a sound source time feature is extracted from the sound source position information.

3. The adaptive sound pickup method based on multi-sound source localization according to claim 1, characterized in that: The comprehensive feature vector set is input into a preset sound source localization model for processing to obtain the target sound source features, specifically: Based on the comprehensive feature vector set, the sound source localization information, as well as the sound source distortion degree and the sound source confusion degree of the target sound source are obtained through the sound source localization model; According to the sound source localization information, a target signal corresponding to a target sound source is extracted from the multi-sound source sound data, and the target signal is optimized according to the sound source distortion degree and the sound source confusion degree of the target sound source to obtain the target sound source feature.

4. The adaptive sound pickup method based on multi-sound source localization according to claim 3, characterized in that: The sound source localization model is specifically: Among them, the loss function of the sound source localization model is expressed as: L total = W location L location +W distortion L distortion +W classification L classification ; Where, L total is the loss function of the sound source localization model, W location is the weight of the sound source localization loss, L location is the sound source localization loss, W distortion is the weight of the sound source distortion loss, L distortion is the distortion loss of the sound source, W classification is the weight of the loss of sound source confusion, L classification It is the loss of sound source confusion.

5. The adaptive sound pickup method based on multi-sound source localization according to claim 1, characterized in that: According to the target sound source characteristics, the optimal sound pickup device parameters in the current environment are obtained through a preset sound pickup strategy prediction model, specifically: According to the characteristics of the target sound source, the sound pickup strategy prediction model selects the optimal sound pickup strategy that meets the current environment from a preset sound pickup strategy library; The optimal sound pickup device parameters are determined according to the hardware parameters of the sound pickup device and the optimal sound pickup strategy.

6. The adaptive sound pickup method based on multi-sound source localization according to claim 1, characterized in that: The step of setting the adjustable parameters of the sound pickup device to pick up sound according to the optimal sound pickup device parameters is specifically as follows: According to the optimal sound pickup device parameters, the sound pickup gain parameters, frequency response parameters, sampling rate and noise threshold value of the sound pickup device are set.

7. An adaptive sound pickup system based on multi-sound source localization, characterized in that: include: Data preprocessing module, sound source localization module, optimal sound pickup parameter acquisition module and sound pickup device setting module; The data preprocessing module is used to perform feature extraction on the collected multi-source sound data to obtain a comprehensive feature vector set, specifically: extracting sound signals, environmental information and sound source position information from the multi-source sound data; performing feature extraction on the sound signals to obtain sound features; performing feature extraction on the sound source position information to obtain sound source spatial features and sound source temporal features; based on the same time dimension, aligning and splicing the sound features, sound source spatial features, sound source temporal features and environmental information in multiple dimensions in turn to obtain a comprehensive feature vector set; The process of acquiring the sound feature is as follows: according to a preset time length, the sound signal is divided into a plurality of frame signals; the frame signal is converted into a Mel spectrum, and a first threshold number of first coefficients are extracted through the Mel spectrum to obtain the energy distribution of the sound signal within a preset frequency range; wherein the Mel spectrum is logarithmically and discrete cosine transformed in sequence to obtain a plurality of Mel-frequency cepstral coefficients, and the first first threshold number of coefficients are selected from the Mel-frequency cepstral coefficients as the first coefficient; according to the frame signal, the short-time energy of the sound signal is calculated to obtain the intensity change of the sound signal within the time length; the zero-crossing rate of each frame signal is calculated to obtain the frequency change of the sound signal within the time length; the sound feature is obtained according to the first coefficient, the short-time energy and the zero-crossing rate; The calculation formula of the short-time energy is: Wherein, x() is the sound signal, N is the frame length of the frame signal, n is the sequence number of the frame signal, m is the index variable of the sample point in the current frame signal, and E(n) is the short-time energy of the nth sample point in the frame signal; The sound source localization module is used to input the comprehensive feature vector set into a preset sound source localization model for processing to obtain sound source localization information and sound source error information; The optimal sound pickup parameter acquisition module is used to obtain the optimal sound pickup device parameters in the current environment through a preset sound pickup strategy prediction model according to the sound source positioning information and the sound source error information; The sound pickup device setting module is used to set the sound pickup device to pick up sound according to the optimal sound pickup device parameters.

Citation Information

Patent Citations

  • Self-adaptive pickup method with voiceprint recognition

    CN117877491A