A self-calibrating acoustic sensor network method and system
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-03-29
- Publication Date
- 2026-08-14
AI Technical Summary
[0006]因此,本发明解决的技术问题是:本发明是解决深度卷积网络在时频域中的大多数应用都是根据梅尔频率频谱图或恒定Q小波频谱图的点对数进行操作的,没有任何形式的实例归一化的问题
[0051]本发明的有益效果:本发明所提出的一种自校正声学传感器网络方法不仅能改善传感器的覆盖范围,还能减少传感器对虚假变化因素的依赖,可以提高传感器网络的可靠性和准确性,增强对真实环境变化的感知能力,提高对真实变化的敏感性。
Smart Images

Figure CN118430574B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of acoustic sensor network technology, specifically to a self-calibrating acoustic sensor network method and system. Background Technology
[0002] The proliferation of machine learning models for wireless sensor networks has brought new opportunities to environmental acoustics. However, these models are prone to statistical biases, such as those caused by unforeseen changes in hardware or the environment. Mitigating these biases becomes even more difficult in the context of supervised learning due to the wide coverage area. Far-field recognition is a challenging task not only because sound waves gradually attenuate as they propagate away from the source, but also because sound waves undergo absorption and reverberation depending on the propagation medium. Furthermore, the propagation of the spectrum itself is affected by meteorological variables such as temperature, pressure, and humidity. In acoustic sensor networks, ensuring the robustness of machine learning systems and preventing missed or false detections requires distributed signal processing techniques that take these varying factors into account. Auditory neurophysiology provides useful domain-specific knowledge for far-field acoustic event detection problems, which has the potential to be applied to acoustic engineering.
[0003] Speech recognition involves higher cognitive processes related to language ability, while meaningless stimuli such as dynamic ripples are less likely to produce individual learning effects; therefore, they reveal early stages of our auditory system. At the cochlear level, two functional elements can explain the human listener's ability to identify distant sounds: the bandpass selectivity of the stereocilia in the inner hair cells, i.e., pitch; and the loudness adaptation of the outer hair cells, i.e., electromigration.
[0004] While tones frequently appear in the form of time-frequency decompositions in machine hearing, electrodynamics does not have such a well-rounded computational equivalent. For example, most applications of deep convolutional networks (convnets) in the time-frequency domain operate on logarithms of Mel frequency spectrograms or constant Q wavelet spectrograms, without any form of instance normalization. Summary of the Invention
[0005] In view of the above-mentioned problems, the present invention is proposed.
[0006] Therefore, the technical problem solved by this invention is that most applications of deep convolutional networks in the time-frequency domain operate based on the logarithm of the Mel frequency spectrum or the constant Q wavelet spectrum, without any form of instance normalization.
[0007] To solve the above-mentioned technical problems, the present invention provides the following technical solution:
[0008] In a first aspect, the present invention provides a self-calibrating acoustic sensor network method, comprising:
[0009] The time-frequency signal of mono audio is acquired, and the time-frequency signal is used to extract features to obtain Mel-frequency cepstral coefficients. The Mel-frequency cepstral coefficients are then subjected to a first preprocessing step to obtain the active channel and sensing channel of the Mel-frequency cepstral coefficients.
[0010] The active channel and sensing channel of the Mel-frequency cepstral coefficients are combined to form a dual-channel time-frequency map. The dual-channel time-frequency map is then subjected to a second preprocessing to obtain a first dual-channel time-frequency map.
[0011] The time-frequency graph of the first dual channel is used as the input data of the neural network, and a neural network classification model is obtained through training.
[0012] In a preferred embodiment of the self-calibrating acoustic sensor network method of the present invention, the following steps are performed: Mel-spectral coefficients are obtained by feature extraction from the time-frequency signal; and the active channel and sensing channel of the Mel-spectral coefficients are obtained by performing a first preprocessing on the Mel-spectral coefficients, including...
[0013] Apply a low-pass filter φ with an attenuation constant equal to T T (t), to obtain a smooth time-frequency representation,
[0014] M(t,f)=(E * φ T (t,f)
[0015] φ is expressed using a first-order autoregressive model. T (t) is defined as an infinite impulse response filter, expressed as,
[0016] M(t,f)=sE(t,f)+(1-s)M(t-1,f)
[0017] The normalized transform spectrum is obtained, represented as follows:
[0018]
[0019] Where E(t,f) is the filterbank energy of each time-frequency block, M(t,f) is the intermediate smoothing energy, t is time, f is frequency, * is convolution, s is the smoothing coefficient, and δ, ε, a, and r are preset parameters.
[0020] As a preferred embodiment of the self-calibrating acoustic sensor network method described in this invention, the active channel of the Mel-frequency cepstral coefficients includes,
[0021] The active channels of the Mel-frequency cepstral coefficients are normalized and expressed as follows:
[0022]
[0023] Where, max f NORMAL(t,f) represents the maximum normalized value at all frequencies f at time point t, min f NORMAL(t,f) is the minimum normalized value at all frequencies f at time point t, median t′ This represents the median over all time frames t′.
[0024] As a preferred embodiment of the self-calibrating acoustic sensor network method described in this invention, wherein:
[0025] The sensing channels for Mel-Cepstral coefficients include,
[0026] The sensor channels of the Mel-frequency cepstral coefficients are normalized and expressed as follows:
[0027]
[0028] Where Flux is the sensing channel, τ is the time factor, s is the smoothing coefficient, t is time, and f is the frequency.
[0029] In a preferred embodiment of the self-calibrating acoustic sensor network method described in this invention, the active channel and sensing channel of the Mel-frequency cepstral coefficients are combined to form a dual-channel time-frequency diagram, including:
[0030] Mel-spectral coefficients are obtained by performing Mel-spectral transform on mono audio data;
[0031] The activity characteristics of the Mel-Cepstral Coefficients are obtained by squaring the Mel-Cepstral Coefficients within each time window and summing all the coefficients; the perceptual characteristics of the Mel-Cepstral Coefficients are obtained by calculating the differences in the Mel-Cepstral Coefficients between adjacent time windows.
[0032] The activity and sensing characteristics of the Mel cepstral coefficients are standardized by dividing the activity characteristic value at each time point by the maximum value of the activity characteristic and the sensing characteristic value at each time point by the maximum value of the sensing characteristic, to obtain the standardized NORMAL time-frequency diagram, which is the time-frequency diagram of the activity channel and the sensing channel of the Mel cepstral coefficients.
[0033] The standardized activity features and perception features are combined to form a dual-channel time-frequency map.
[0034] In a preferred embodiment of the self-calibrating acoustic sensor network method of the present invention, the time-frequency graph of the dual-channel signal is subjected to a second preprocessing to obtain a first dual-channel time-frequency graph, including:
[0035] The dual-channel time-frequency graph is converted into a grayscale image, and the grayscale image is segmented into multiple overlapping blocks.
[0036] Calculate the histogram of each pixel value block, and calculate the constraint factor based on the histogram to perform histogram equalization on each pixel value block;
[0037] The small blocks that have undergone histogram equalization are recombined into a complete CLAHE image to obtain the first dual-channel time-frequency map.
[0038] In a preferred embodiment of the self-calibrating acoustic sensor network method described in this invention, the time-frequency graph of the first dual-channel circuit is used as input data for the neural network, and a neural network classification model is obtained through training, including:
[0039] The time-frequency graph of the first dual-channel array is used as input data for the neural network. Through training, a neural network classification model is obtained, including...
[0040] The output of the acoustic feature extractor is processed using Conv1D+ReLU+BN;
[0041] The feature maps processed by Conv1D+ReLU+BN are input into the SE-Res2Net module to fuse the output feature maps of the SE-Res2Net module.
[0042] The output feature map of the SE-Res2Net module is input into the ASP layer block, and the output features of the ASP layer block are adjusted in dimension through a fully connected layer to obtain the speaker embedding code.
[0043] In a second aspect, the present invention provides a system for a self-calibrating acoustic sensor network, comprising:
[0044] The preprocessing module is used to acquire the time-frequency signal of mono audio, extract features from the time-frequency signal to obtain Mel-frequency cepstral coefficients, and obtain the active channel and sensing channel of the Mel-frequency cepstral coefficients by performing a first preprocessing on the Mel-frequency cepstral coefficients.
[0045] The channel combination module is used to combine the active channel and the sensing channel of the Mel-frequency cepstral coefficients to form a dual-channel time-frequency map, and to perform a second preprocessing on the dual-channel time-frequency map to obtain a first dual-channel time-frequency map;
[0046] The model building module is used to take the time-frequency graph of the first dual channel as input data for the neural network and obtain a neural network classification model through training.
[0047] Thirdly, the present invention provides a computing device, comprising:
[0048] Memory and processor;
[0049] The memory is used to store computer-executable instructions, and the processor is used to execute the computer-executable instructions, which, when executed by the processor, implement the steps of the self-calibrating acoustic sensor network method.
[0050] Fourthly, the present invention provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, implement the steps of the self-calibrating acoustic sensor network method.
[0051] The beneficial effects of this invention are as follows: The self-calibrating acoustic sensor network method proposed in this invention can not only improve the coverage of the sensor, but also reduce the sensor's dependence on spurious changes, thereby improving the reliability and accuracy of the sensor network, enhancing its ability to perceive changes in the real environment, and increasing its sensitivity to real changes. Attached Figure Description
[0052] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein:
[0053] Figure 1 This is an overall flowchart of a self-calibrating acoustic sensor network method provided by the present invention. Detailed Implementation
[0054] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0055] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0056] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0057] This invention is described in detail with reference to the schematic diagrams. When detailing the embodiments of this invention, for ease of explanation, the cross-sectional views illustrating the device structure may be partially enlarged, not adhering to the usual scale. Furthermore, the schematic diagrams are merely examples and should not be construed as limiting the scope of protection of this invention. In actual fabrication, the three-dimensional spatial dimensions of length, width, and depth should be included.
[0058] Furthermore, in the description of this invention, it should be noted that the terms "upper," "lower," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used solely for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In addition, the terms "first," "second," or "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0059] Unless otherwise explicitly specified and limited, the terms "installation," "connection," and "joining" in this invention should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; similarly, they can refer to mechanical connections, electrical connections, or direct connections, or indirect connections through an intermediate medium, or internal connections between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0060] Example 1
[0061] Reference Figure 1 As one embodiment of the present invention, a self-calibrating acoustic sensor network method is provided, comprising:
[0062] S100: Acquire the time-frequency signal of mono audio, extract features from the time-frequency signal to obtain Mel-frequency cepstral coefficients, and obtain the active channel and sensing channel of the Mel-frequency cepstral coefficients by performing a first preprocessing on the Mel-frequency cepstral coefficients;
[0063] Furthermore, a low-pass filter φ with an attenuation constant equal to T is applied. T (t), to obtain a smooth time-frequency representation,
[0064] M(t,f)=(E * φ T (t,f)
[0065] φ is expressed using a first-order autoregressive model. T (t) is defined as an infinite impulse response filter, expressed as,
[0066] M(t,f)=sE(t,f)+(1-s)M(t-1,f)
[0067] The normalized transform spectrum is obtained, represented as follows:
[0068]
[0069] Where E(t,f) is the filterbank energy of each time-frequency block, M(t,f) is the intermediate smoothing energy, t is time, f is frequency, * is convolution, s is the smoothing coefficient, and δ, ε, a, and r are preset parameters.
[0070] It should be noted that the variable f can correspond to a frequency in Hertz, such as the complex modulus of the short-term Fourier transform; it can also correspond to a frequency in Myers, such as the Myer frequency spectrum; and it can also correspond to a frequency in semitones, such as the constant Q tone spectrum. * represents convolution and is implicit in the "channel" f at different frequencies.
[0071] s is the smoothing coefficient, and δ, ε, a, and r are preset parameters. These parameter values are determined through experience and can be set as follows: s = 0.025, a = 0.98, δ = 2, r = 0.5, ε = 0.000001.
[0072] It should also be noted that by applying a low-pass filter φ with an attenuation constant equal to T... T (t) can smooth the time-frequency representation, reduce noise and details in the high-frequency part, make the energy distribution smoother, help extract more significant time-frequency features, and reduce unnecessary fluctuations and abrupt changes. A first-order autoregressive model is used to analyze the filter φ. T The model defines a smoothed energy M(t-1,f) and uses the smoothed energy M(t,f) from the previous time step to perform a weighted average of the smoothed energy M(t,f) from the previous time step. This autoregressive model introduces temporal continuity, making the changes in smoothed energy more gradual and helping to eliminate sudden noise or anomalous changes.
[0073] By normalizing the spectrum, the smoothed energy M(t,f) is compared with the filterbank energy E(t,f) and then normalized. Normalization maps the energy values to a uniform range, amplifies low-energy components, and compresses high-energy components, helping to balance energy differences between different time-frequency blocks and improve the perception of the overall energy distribution. Pre-set parameters δ, ε, a, and r can be adjusted according to specific application scenarios and requirements. By appropriately setting these parameters, the normalization process can be flexibly controlled to achieve better normalization results.
[0074] Furthermore, the Mel-frequency cepstral coefficients' active channels (MFCC-Activity) are normalized and expressed as follows:
[0075]
[0076] Where, max f NORMAL(t,f) represents the maximum normalized value at all frequencies f at time point t, min f NORMAL(t,f) is the minimum normalized value at all frequencies f at time point t, median t′ This represents the median over all time frames t′.
[0077] The sensing channel (MFCC-Flux) of the Mel-frequency cepstral coefficients is normalized and expressed as follows:
[0078]
[0079] Where Flux is the sensing channel, τ is the time factor, s is the smoothing coefficient, t is time, and f is the frequency.
[0080] It should be noted that the analogy between NORMAL and the spectral flux in the limiting case is as follows: ε→0, a→0, r→0... The sensing channel of NORMAL can be obtained through the limiting case of NORMAL0(t,f).
[0081] It should also be noted that by calculating the maximum and minimum normalized values at all frequencies f at time point t, and the median at all time frames t′, the active channel of the Mel-frequency cepstral coefficients can be normalized. This normalizes the energy range of the active channel to a unified interval, allowing for comparison and analysis of energy values at different time points. The normalized active channel better reflects energy changes at different time points, helping to identify important signal characteristics.
[0082] By calculating the ratio of the energy value at all frequencies f at time point t to the sum of the energy values at the previous moment corresponding to the smoothing coefficient s multiplied by the time factor τ, and taking the maximum value and applying the logarithmic function, the sensing channel of the Mel-Cepstral Coefficient can be normalized. This can balance the energy distribution of the sensing channel, highlight important energy peaks and suppress unnecessary fluctuations. The normalized sensing channel can better highlight the dynamic characteristics of the signal and improve the sensitivity to signal changes.
[0083] Normalization enhances the contrast between different frequencies and time points in the Mel-Frequency Cepstral Coefficient (MFCC), making important energy changes more apparent. This helps extract key features from the signal and reduces the impact of noise and unnecessary details on the analysis results. Normalized MFCCs better reflect the frequency characteristics and time-domain dynamics of the signal, providing a more accurate feature representation.
[0084] S200: Combine the active channel and sensing channel of the Mel-frequency cepstral coefficients to form a dual-channel time-frequency map, and perform a second preprocessing on the dual-channel time-frequency map to obtain a first dual-channel time-frequency map;
[0085] Furthermore, Mel-frequency cepstral transform is performed on the mono audio data to obtain the Mel-frequency cepstral coefficients;
[0086] The activity characteristics of the Mel-Cepstral Coefficients are obtained by squaring the Mel-Cepstral Coefficients within each time window and summing all the coefficients; the perceptual characteristics of the Mel-Cepstral Coefficients are obtained by calculating the differences in the Mel-Cepstral Coefficients between adjacent time windows.
[0087] The activity and sensing characteristics of the Mel cepstral coefficients are standardized by dividing the activity characteristic value at each time point by the maximum value of the activity characteristic and the sensing characteristic value at each time point by the maximum value of the sensing characteristic, to obtain the standardized NORMAL time-frequency diagram, which is the time-frequency diagram of the activity channel and the sensing channel of the Mel cepstral coefficients.
[0088] The standardized activity features and perception features are combined to form a dual-channel time-frequency map.
[0089] It should be noted that the MFCC time-frequency diagram is generated by using the normal operating time of the hydropower station as the total segmentation length, segmenting the sound data of the hydro-generator according to the preset window length and preset movement length to obtain multiple data windows; and performing MFCC transformation on each data window based on the preset first transformation parameters to generate a spectrogram, i.e., a time-frequency diagram.
[0090] It should also be noted that standardizing the Activity and Flux features involves dividing the Activity value at each time point by the maximum value of Activity, normalizing its range to between 0 and 1. Similarly, dividing the Flux value at each time point by the maximum value of Flux, normalizing its range to between 0 and 1.
[0091] The dual-channel time-frequency plots of MFCC-Activity and MFCC-Flux combine the activity level and rate of change information of MFCC coefficients. Traditional MFCC time-frequency plots only provide spectral information, while adding Activity and Flux features can more comprehensively reflect the dynamic characteristics of the audio signal. Such dual-channel time-frequency plots can better capture the time and frequency domain information of the audio signal, providing richer feature representations. The dual-channel time-frequency plots of MFCC-Activity and MFCC-Flux provide more contextual information and dynamic feature change information. Such feature representations are highly beneficial for speech recognition and speaker recognition tasks. Activity level information can help identify active speech regions, which is helpful for speech detection and speech segmentation tasks, while the rate of change information can provide more speaker-specific features, which is helpful for speaker recognition and speech verification tasks.
[0092] Furthermore, the dual-channel time-frequency image is converted into a grayscale image, and the grayscale image is segmented into multiple overlapping small blocks;
[0093] Calculate the histogram of each pixel value block, and calculate the constraint factor based on the histogram to perform histogram equalization on each pixel value block;
[0094] The small blocks that have undergone histogram equalization are recombined into a complete CLAHE image to obtain the first dual-channel time-frequency map.
[0095] It should be noted that the time-frequency image is converted into a grayscale image by averaging or weighted averaging the values of the active and sensing channels, resulting in a grayscale matrix of shape (T,F). The grayscale image is then segmented into multiple overlapping blocks. The segmentation method is controlled by setting the block size and overlap. The image is divided into blocks of size (12,12) with a stride of (7,7). Each block is extracted using a loop and histogram equalization is applied to it. CLAHE objects are created using functions from the OpenCV library and applied to each block, resulting in histogram equalized blocks. These blocks are then recombined into a complete CLAHE image. The complete CLAHE image is the result of histogram equalization of the first dual-channel time-frequency image.
[0096] It should also be noted that histogram equalization can enhance image contrast by redistributing pixel values. Applying the CLAHE algorithm to time-frequency graphs can make the features of different frequencies and time periods more obvious and prominent, which helps to improve the visualization effect of the image. It also makes the difference between the active channel and the sensing channel of the Mel-Cepstral Coefficient more explicit, making the features of the active channel and the sensing channel of the Mel-Cepstral Coefficient more obvious and recognizable. This helps to further analyze and interpret the features of acoustic signals and improve the performance of classification and recognition tasks.
[0097] S300: The time-frequency graph of the first dual channel is used as the input data of the neural network, and a neural network classification model is obtained through training;
[0098] Furthermore, the output of the acoustic feature extractor is processed using Conv1D+ReLU+BN.
[0099] The feature maps processed by Conv1D+ReLU+BN are input into the SE-Res2Net module to fuse the output feature maps of the SE-Res2Net module.
[0100] The output feature map of the SE-Res2Net module is input into the ASP layer block, and the output features of the ASP layer block are adjusted in dimension through a fully connected layer to obtain the speaker embedding code.
[0101] It should be noted that neural networks can use ECAPA-TDNN and other technologies for voiceprint modeling. Taking ECAPA-TDNN as an example, it is an improvement on the traditional x-vector architecture, which places greater emphasis on local multi-scale feature representation and global multi-level feature fusion.
[0102] The entire ECAPA-TDNN network structure processes the output of the acoustic feature extractor through a single layer of Conv1D+ReLU+BN (equivalent to one TDNN module); it then fuses the output feature maps from three SE-Res2Net modules and feeds them into an ASP (attention statistics pooling) layer; finally, it connects to a fully connected layer to adjust the dimensions, resulting in a 192-dimensional speaker embedding code. Cosine similarity is used as the backend similarity discriminator. When discrepancies arise, the one with the higher similarity score is used.
[0103] It should also be noted that, assuming the output of the acoustic feature extractor is a tensor, the output of the acoustic feature extractor is input into a Conv1D layer. This layer has a ReLU activation function and a Batch Normalization (BN) operation, the purpose of which is to introduce local contextual information in the temporal dimension. Hyperparameters such as the number of filters, filter size, and stride of the Conv1D can be defined. The feature map processed by Conv1D+ReLU+BN is input into three SE-Res2Net modules, each with its own feature map. The purpose of the ASP layer block is to perform attention-weighted and statistical convergence of features. Attention mechanisms can be used to weight features to highlight important information. The fully connected layer converts the input feature map into a 192-dimensional embedding code.
[0104] The above is an illustrative scheme of a self-calibrating acoustic sensor network method according to this embodiment. It should be noted that the technical solution of this self-calibrating acoustic sensor network device belongs to the same concept as the technical solution of the self-calibrating acoustic sensor network method described above. Details not described in detail in the technical solution of the self-calibrating acoustic sensor network device in this embodiment can be found in the description of the technical solution of the self-calibrating acoustic sensor network method described above.
[0105] The apparatus for a self-calibrating acoustic sensor network in this embodiment includes:
[0106] The preprocessing module is used to acquire the time-frequency signal of mono audio, extract features from the time-frequency signal to obtain Mel-frequency cepstral coefficients, and obtain the active channel and sensing channel of the Mel-frequency cepstral coefficients by performing a first preprocessing on the Mel-frequency cepstral coefficients.
[0107] The channel combination module is used to combine the active channel and the sensing channel of the Mel-frequency cepstral coefficients to form a dual-channel time-frequency map, and to perform a second preprocessing on the dual-channel time-frequency map to obtain a first dual-channel time-frequency map;
[0108] The model building module is used to take the time-frequency graph of the first dual channel as input data for the neural network and obtain a neural network classification model through training.
[0109] This embodiment also provides a computing device suitable for self-calibrating acoustic sensor networks, including:
[0110] The system includes a memory and a processor; the memory stores computer-executable instructions, and the processor executes the computer-executable instructions to implement the self-calibrating acoustic sensor network method proposed in the above embodiments.
[0111] This embodiment also provides a storage medium on which a computer program is stored, which, when executed by a processor, implements the method for implementing a self-calibrating acoustic sensor network as proposed in the above embodiments.
[0112] The storage medium proposed in this embodiment and the method for realizing a self-calibrating acoustic sensor network proposed in the above embodiments belong to the same inventive concept. Technical details not described in detail in this embodiment can be found in the above embodiments, and this embodiment has the same beneficial effects as the above embodiments.
[0113] Based on the above description of the implementation methods, those skilled in the art can clearly understand that the present invention can be implemented using software and necessary general-purpose hardware, and of course, it can also be implemented using hardware, but in many cases the former is a better implementation method. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a computer-readable storage medium, such as a computer floppy disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk, or optical disk, etc., including several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute the methods of the various embodiments of the present invention.
[0114] Example 2
[0115] Referring to Table 1, this embodiment differs from the first embodiment in that it provides a self-calibrating acoustic sensor network method to verify and explain the technical effects used in this method.
[0116] ECAPA-TDNN is used for voiceprint modeling, which is an improvement on the traditional x-vector architecture, and places greater emphasis on local multi-scale feature representation and global multi-level feature fusion.
[0117] Using ECAPA-TDNN as the voiceprint embedding code extractor and cosine similarity as the back-end similarity discriminator, four anomaly detection systems were constructed, each with MFCC (Mel-Cepstral Coefficients), MFCC-Activity (active channel of Mel-Cepstral Coefficients), MFCC-Flux (perceptual channel of Mel-Cepstral Coefficients), and MFCC-Activity-Flux (fusion of active and perceptual channels of Mel-Cepstral Coefficients) as acoustic feature extractors.
[0118] Sounds emitted from different locations were detected in the noisy environment of the turbine room of a hydroelectric generator. These sounds were recorded over five days by nine sensors. The distance between the sensors and the sound sources was estimated using acoustic beamforming. The dataset contains 40,000 segments, each lasting 3.015 seconds. These segments were annotated by a technician and are part of a large continuous recording set totaling 1,000 hours.
[0119] Randomly extract segments from the dataset and extract fixed segments with a length of 3.015s. The dimension of the input acoustic features is set to 80, the window length is 25ms, and the frame shift is 10ms. The fixed parameter values of the MFCC-Activity module and the initial values of the LearnablePCEN parameters are set to α = 0.98, δ = 2.0, r = 0.5, and ε = 0.025. SoftMax (m = 0.2, s = 30) is used as the backend classifier of the model. The Adam optimizer is used for training, the batch size is set to 300, and the initial learning rate is set to 0.001. At the same time, the interval learning rate adjustment strategy is adopted, the step size is set to 1, and the learning rate of each round is learning rate of each round = learning rate of the previous round × 0.97 (8). Audio features are extracted from the first fully connected layer after the ASP pooling layer, and finally a 512-dimensional speaker feature vector is output. The performance of the speaker recognition system was evaluated by calculating the equal error rate (EEER) and the minimum detection cost function (minDCF), with a risk factor of 1.0.
[0120] The performance evaluation results of the speaker recognition system in the domain environment are shown in Table 1.
[0121] Feature extraction <![CDATA[E EER / %]]> MFCC 3.96 MFCC-Activity-Flux 3.63 MFCC-Flux 3.85 MFCC-Activity 3.91
[0122] In terms of EEER, MFCC-Activity-Flux reduced performance by 8.35%, 5.73%, and 7.18% compared to MFCC-Activity, MFCC-Flux-SV, and MFCC, respectively, indicating that MFCC-Activity-Flux provides better performance for abnormal systems.
[0123] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A self-calibrating acoustic sensor network method, characterized in that, include: The time-frequency signal of mono audio is acquired, and the time-frequency signal is used to extract features to obtain Mel-frequency cepstral coefficients. The Mel-frequency cepstral coefficients are then subjected to a first preprocessing step to obtain the active channel and sensing channel of the Mel-frequency cepstral coefficients. The active channel and sensing channel of the Mel-frequency cepstral coefficients are combined to form a dual-channel time-frequency map. The dual-channel time-frequency map is then subjected to a second preprocessing to obtain a first dual-channel time-frequency map. Among them, the active channels of the Mel-Cepstral Coefficients include... The active channels of the Mel-frequency cepstral coefficients are normalized and expressed as follows: in, For time points All frequencies The maximum value of the normalized value, For time points All frequencies The minimum value of the normalized value on the , Indicates all time frames the median on; The sensing channels for Mel-Cepstral coefficients include, The sensor channels of the Mel-frequency cepstral coefficients are normalized and expressed as follows: in, For sensing channels, As a time factor, For smoothing coefficients, For time, For frequency; The active channel and sensing channel of the Mel-frequency cepstral coefficients are combined to form a dual-channel time-frequency diagram, including: Mel-spectral coefficients are obtained by performing Mel-spectral transform on mono audio data; The activity characteristics of the Mel-Cepstral Coefficients are obtained by squaring the Mel-Cepstral Coefficients within each time window and summing all the coefficients; the perceptual characteristics of the Mel-Cepstral Coefficients are obtained by calculating the differences in the Mel-Cepstral Coefficients between adjacent time windows. The activity and sensing characteristics of the Mel cepstral coefficients are standardized by dividing the activity characteristic value at each time point by the maximum value of the activity characteristic and the sensing characteristic value at each time point by the maximum value of the sensing characteristic, to obtain the standardized NORMAL time-frequency diagram, which is the time-frequency diagram of the activity channel and the sensing channel of the Mel cepstral coefficients. The standardized activity features and perception features are combined to form a dual-channel time-frequency map; The time-frequency graph of the dual channels is subjected to a second preprocessing step to obtain the time-frequency graph of the first dual channel, including: The dual-channel time-frequency graph is converted into a grayscale image, and the grayscale image is segmented into multiple overlapping blocks. Calculate the histogram of each pixel value block, and calculate the constraint factor based on the histogram to perform histogram equalization on each pixel value block; The small blocks that have undergone histogram equalization are recombined into a complete CLAHE image to obtain the first dual-channel time-frequency map. The time-frequency graph of the first dual channel is used as the input data of the neural network, and a neural network classification model is obtained through training. The neural network adopts the ECAPA-TDNN neural network model.
2. The self-calibrating acoustic sensor network method as described in claim 1, characterized in that, Mel-frequency cepstral coefficients are obtained by feature extraction from the time-frequency signal. Through a first preprocessing step, the active channel and sensing channel of the Mel-frequency cepstral coefficients are obtained, including... Apply a low-pass filter with an attenuation constant equal to T ,get: Using a first-order autoregressive model Defined as an infinite impulse response filter, denoted as, The normalized transform spectrum is obtained, represented as follows: in, The filterbank energy for each time-frequency block. For intermediate smooth energy, For time, For frequency, For convolution, For smoothing coefficients, , , , These are pre-set parameters.
3. The self-calibrating acoustic sensor network method as described in claim 2, characterized in that, The time-frequency graph of the first dual-channel array is used as input data for the neural network. Through training, a neural network classification model is obtained, including... The output of the acoustic feature extractor is processed using Conv1D+ReLU+BN; The feature maps processed by Conv1D+ReLU+BN are input into the SE-Res2Net module to fuse the output feature maps of the SE-Res2Net module. The output feature map of the SE-Res2Net module is input into the ASP layer block, and the output features of the ASP layer block are adjusted in dimension through a fully connected layer to obtain the speaker embedding code.
4. A self-calibrating acoustic sensor network system, employing the self-calibrating acoustic sensor network method as described in any one of claims 1 to 3, characterized in that, include: The preprocessing module is used to acquire the time-frequency signal of mono audio, extract features from the time-frequency signal to obtain Mel-frequency cepstral coefficients, and obtain the active channel and sensing channel of the Mel-frequency cepstral coefficients by performing a first preprocessing on the Mel-frequency cepstral coefficients. The channel combination module is used to combine the active channel and the sensing channel of the Mel-frequency cepstral coefficients to form a dual-channel time-frequency map, and to perform a second preprocessing on the dual-channel time-frequency map to obtain a first dual-channel time-frequency map; The model building module is used to take the time-frequency graph of the first dual channel as input data for the neural network and obtain a neural network classification model through training.
5. An electronic device, characterized in that, The device includes: processor; Memory used to store processor-executable instructions; The processor is configured to invoke instructions stored in the memory to execute the method described in any one of claims 1 to 3.
6. A computer-readable storage medium having computer program instructions stored thereon, characterized in that, When the computer program instructions are executed by the processor, they implement the method described in any one of claims 1 to 3.
Citation Information
Patent Citations
Speaker clustering method based on multi-scale channel separation convolution feature extraction
CN115101076A
Speech emotion recognition method and system based on time-frequency feature double fusion
CN116884441A