Audio recognition and noise reduction processing method and system for power grid regulation and control workbench environment

By adopting a comprehensive method of endpoint detection, mixed audio separation and voiceprint recognition in the grid control workbench environment, the problem of poor results in traditional audio acquisition in complex environments is solved, and high-accurate audio acquisition and noise removal effects are achieved.

CN120148543AActive Publication Date: 2025-06-13STATE GRID FUJIAN ELECTRIC POWER CO LTD +3

Patent Information

Application Number
CN202510486589.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-18
Publication Date
2025-06-13
Estimated Expiration
2045-04-18

AI Technical Summary

Technical Problem

Traditional audio acquisition methods cannot effectively obtain clean audio in complex power grid control workbench environments, especially when there are multiple interference sources such as telephone ringtones, alarm ringtones, communication between people and equipment noise, and cannot effectively deal with multiple rounds of interactive power scheduling tasks.

Method used

A comprehensive discriminant audio acquisition method with fusion endpoint detection, mixed audio separation, and voiceprint recognition authentication is adopted. By obtaining audio features in different working scenarios, cross-domain fusion training of speaker models is carried out, voice signals are collected in real time for endpoint detection and weighting, vocal fragments are identified and separated, target speaker vocals are extracted, and the speaker model is updated.

Benefits of technology

Effectively remove noise in a complex grid control workbench environment, improve the accuracy and recognition rate of audio acquisition, and accurately identify and extract target speaker fragments, adapt to different work scenarios and environmental changes.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120148543A_ABST
    Figure CN120148543A_ABST
Patent Text Reader

Abstract

The invention discloses an audio recognition and noise reduction processing method and system for a power grid regulation and control workbench environment, and the method comprises the steps: obtaining pure audio signals of all power grid regulation and control workers in different working scenes and different background noises in a dispatching desk, carrying out the cross-domain fusion of two features of a time domain and a frequency domain, and taking the two features as a training set to train a speaker model; voice signals are collected in real time, and pure noise and silence in the voice signals are removed; according to the noise level estimation and the noise power, carrying out weighting processing on the voice signal after the pure noise and the mute are removed; recognizing and separating human voice fragments by adopting a continuous down-sampling and resampling algorithm with multi-resolution characteristics; judging whether the human voice fragment contains a target speaker or not through a speaker model; and generating a plurality of voice segments based on the human voice segments, and extracting target speaking human voice through clustering and a speaker model. The method adapts to environmental changes of regulation and control dispatching desks at all levels, environmental noise can be remarkably reduced, and accurate recognition and extraction of human voices are achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of speech recognition, and more specifically, relates to an audio recognition noise reduction processing method and system for the environment of a power grid dispatching workbench. Background Art

[0002] Due to the existence of various interference sources in the dispatching desk environment, such as telephone rings, alarm bells, conversations of others, equipment noises, etc., traditional audio acquisition methods cannot obtain clean audio in such a complex environment. In addition, power dispatching tasks usually involve multiple rounds of interaction and require continuous audio acquisition, which makes it impossible for traditional endpoint detection processing methods to effectively cope with.

[0003] The existing technologies for voice recognition include: The patent with the authorization announcement number CN109686377B discloses an audio recognition method, device, and computer-readable storage medium. The method includes: obtaining a voiceprint vectorization model; obtaining multiple different first audio files of the same target speaker; vectorizing each first audio file using the voiceprint vectorization model; determining the central vector of the voiceprint vectors of the multiple different first audio files in at least one way, and respectively determining a similarity acceptance range using each central vector; obtaining the voiceprint vector of the audio file to be recognized using the voiceprint vectorization model, and calculating the similarity between the voiceprint vector of the audio file to be recognized and each central vector; for each central vector, determining whether the voiceprint vector of the audio file to be recognized is within the similarity acceptance range determined using it, and determining whether the audio file to be recognized belongs to the target speaker according to the judgment result. The patent with the publication number CN116312570A discloses a voice noise reduction method, device, equipment, and medium based on voiceprint recognition. The method includes: obtaining the voiceprint template information of a specified person and the scene audio including the voice of the specified person; performing voice separation on the scene audio to obtain multiple person audios corresponding to multiple single persons respectively, and the person audio includes scene noise; determining the specified person audio corresponding to the specified person by matching the person audio and the voiceprint template information; performing noise reduction processing on the specified person audio to obtain the target audio. By performing voice separation on the scene audio and matching the audio corresponding to the specified person in the person audios corresponding to multiple single persons, the audio corresponding to the specified person can be obtained, so that in the scenario voice with multiple speakers, other audios in the voice except the target speaker can be regarded as noise, and the voice of the target speaker can be retained. However, the above patents cannot obtain clean audio in a complex environment with various interference sources, such as telephone rings, alarm bells, conversations of others, equipment noises, etc. Summary of the Invention

[0004] To address the deficiencies in the existing technology, the present invention provides a comprehensive discriminative audio acquisition method that integrates endpoint detection, mixed audio separation, and voiceprint recognition and authentication.

[0005] The present invention adopts the following technical solutions.

[0006] A first aspect of the present invention proposes an audio recognition and noise reduction processing method for the environment of a power grid control workbench, characterized by including: Obtain the pure audio signals of all power grid control staff in different working scenarios and different background noises in the dispatching console. Extract the time-domain waveform features and frequency-domain features of the pure audio signals from different time scales, and fuse the two features across domains to train a speaker model as a training set; Real-time collect voice signals, perform endpoint detection processing on the voice signals to remove pure noise and silence in the voice signals; According to the noise level estimation and the noise power of different background noises in the dispatching console, calculate the weights of each time frame and different frequencies of the voice signal after removing pure noise and silence, so as to perform weighted processing on the voice signal after removing pure noise and silence; Adopt a pre-trained continuous downsampling and resampling algorithm with multi-resolution features for the weighted voice signal to identify and separate the human voice segments, and the human voice segments include other people's voices and the target speaker's voice; Determine whether the human voice segment contains the target speaker through the speaker model; If the human voice segment contains the target speaker, then extract the target speaker's voice through region proposal, clustering, and the speaker model, establish a voiceprint mapping table for the working scenario period of the staff, construct a comprehensive loss function based on this table, and use it and the extracted target speaker's voice to update the speaker model.

[0007] Preferably, different working scenarios of power grid control staff include: monitoring and dispatching, fault handling, and safety analysis; Different background noises in the dispatching console include transformer noise, alarm noise, telephone ringtone noise, server noise, and personnel noise; The server noise includes the noise of air conditioners, fans, and power supplies, and the personnel noise includes the sounds of personnel walking, flipping through documents, and typing on keyboards.

[0008] Preferably, extracting the time-domain waveform features and frequency-domain features of the pure audio signal from different time scales specifically includes: The audio signal is divided into several long-time window signals with a set duration T1 as one frame; each long-time window signal is divided into several medium-time window signals with a set duration T2 as one frame; each medium-time window signal is divided into several short-time window signals with a set duration T3 as one frame; calculate the waveform skewness of all long-time window signals, the energy envelope of medium-time window signals, and the zero-crossing rate of short-time window signals as time-domain waveform features; Extract the Mel spectrum and phase derivative spectrum of these audio signals as frequency-domain features.

[0009] Preferably, the two types of features are cross-domain fused and used as a training set to train a speaker model, specifically: The speaker model is a convolutional neural network; The cross-domain fusion of the two types of features is to calculate the attention weights of time-domain waveform features to frequency-domain features and the attention weights of frequency-domain features to time-domain waveform features respectively; When calculating the attention weights of time-domain waveform features to frequency-domain features, first obtain a query vector that maps time-domain waveform features to a set dimension, and key vectors and value vectors that map frequency-domain features to a set dimension, and use them to calculate the attention weights of time-domain waveform features to frequency-domain features; When calculating the attention weights of frequency-domain features to time-domain waveform features, first obtain a query vector that maps frequency-domain features to a set dimension, and key vectors and value vectors that map time-domain waveform features to a set dimension, and use them to calculate the attention weights of time-domain waveform features to frequency-domain features; multiply the time-domain waveform features by the attention weights of time-domain waveform features to frequency-domain features, multiply the frequency-domain features by the attention weights of frequency-domain features to time-domain waveform features, and splice and fuse the two calculation results into features.

[0010] Preferably, the process of endpoint detection processing is specifically: Step 1: Collect voice signals through a microphone. After collecting voice signals within a set time period, input the collected voice signals into a voice activity detection algorithm model for judgment, and output the judged voice signal type; Step 2: If the voice type is a non-voice signal, it is judged that it is pure noise or silence at this time, and the voice collection of the microphone is stopped; if it is a voice signal, go to Step 3 to continue using the microphone to collect voice signals; Step 3: During the voice collection process, sequentially input the voice signals collected in the next time period through the voice activity detection algorithm to judge whether a non-voice signal is detected. If not, continue to collect signals. If detected, go to Step 4; Step 4: Continue to collect the voice signals collected within the set M time periods. The voice signals collected in each time period are input into the voice activity detection algorithm for detection. If a voice signal is detected, the task remains in the voice process and returns to Step 3. If there is no voice signal throughout the M time periods, it is considered that the speech has ended, and the microphone collection is stopped.

[0011] Preferably, the voice activity detection algorithm is specifically: After pre-emphasizing, framing, and windowing the voice signal, it is converted into a frequency-domain signal through Mel Frequency Cepstral Coefficients (MFCC). The frequency-domain signal is divided into six sub-bands according to frequency. For each sub-band, its energy is calculated. The probability density function of the trained Gaussian model is used to operate on the energy of these six sub-bands and the weighted sum of the energies of the six sub-bands respectively to obtain the corresponding log-likelihood ratio. Among them, the log-likelihood functions calculated from the energies of the six sub-bands are all local log-likelihood functions, and the weighted sum of the energies of the six sub-bands calculates the global log-likelihood ratio. When making a voice decision, all local and global log-likelihood ratios are judged. If any local or global log-likelihood ratio exceeds the set log-likelihood ratio threshold, it is considered that there is voice. Among them, the weights and thresholds of the weighted energies of the six sub-bands are obtained through training. After each detection, the algorithm self-learns and updates the mean and variance parameters of the Gaussian model.

[0012] Preferably, the calculation of the weights of each time frame and different frequencies of the voice signal after removing pure noise and silence according to the noise level estimation and the noise power of different background noises in the dispatching console is specifically: The voice signal after removing pure noise and silence and different background noises in the dispatching console are framed with a set duration T4 as one frame and decomposed into several frequency points through short-time Fourier transform. Calculate the noise level estimation of each frame of the voice signal after removing pure noise and silence and the noise power of different background noises in the dispatching console; Calculate the th i frequency point corresponding to the weight of the

[0013] The formula is: , , are all regularization coefficients; is the set noise level estimation threshold, is the maximum noise level estimation, is the sum of the noise powers of the background noises that dominate at different frequencies. , , , , respectively represent the noise powers of transformer noise, alarm noise, telephone ringtone noise, server noise, and personnel noise.

[0014] Preferably, calculate the noise level estimate of the speech signal after removing pure noise and silence for each frame and the noise powers of different background noises in the dispatching console, specifically: The noise level estimate of the frame is calculated as follows:

[0015] where is the total number of frequency points of the short-time Fourier transform; is the frequency corresponding to the i th frequency point of the th frame; The noise powers of different background noises in the dispatching console are calculated by calculating the noise level estimate of each frame of different background noises in the dispatching console and taking the average value as its noise power.

[0016] Preferably, the target speaker voice is extracted through region proposal, clustering, and speaker model, specifically: Use the region proposal algorithm to generate multiple non-overlapping regions of speech segments based on the voice segment. This algorithm is a fully convolutional network model that uses a 3*3 sliding window to slide the voice segment to generate multiple speech segments; Use clustering to form multiple clusters from the speech segments of multiple non-overlapping regions. Input each speech segment in each cluster into the speaker model for recognition, count the number of speech segments containing the target speaker in each cluster, and integrate all the speech segments of the cluster with the largest number into a speech signal, which is the extracted target speaker voice.

[0017] Preferably, the voiceprint mapping table for the working scenario time period of the staff is established, and based on this table, a comprehensive loss function is constructed. Using it and the extracted target speaker voice, the speaker model is updated, specifically: According to the duty schedules and historical work record information of all grid dispatching staff, establish a voiceprint mapping table for the working scenario time period of the staff. This table establishes the mapping relationship between the most frequent working scenario in each working period of each grid dispatching staff, the corresponding pure audio signal in the working scenario, the cross-domain fusion feature vector, and the short-time Fourier transform complex spectrum. Extract the time-domain waveform features and frequency-domain features of the target speaker's voice from different time scales respectively, and cross-domain fuse them into a feature vector ; According to the voiceprint mapping table of the staff's working scenario period, obtain the feature vector of the pure audio signal of the target speaker in the current working scenario and the short-time Fourier transform complex spectrum ; The comprehensive loss function is:

[0018] wherein, 、 are set parameters, and their values are all in the interval of [0,1], is the norm, is transpose of, is the short-time Fourier transform complex spectrum of the target speaker's voice

[0019] The second aspect of the present invention provides an audio recognition and noise reduction processing system for the power grid control workbench environment using the audio recognition and noise reduction method described in the first aspect of the present invention, including: a speaker model construction module, an endpoint detection module, a separated voice segment module, a speech recognition module, and a target speaker voice extraction module, characterized in that: Speaker model construction module: Obtain the pure audio signals of all power grid control staff in different working scenarios and different background noises in the dispatching desk, extract the time-domain waveform features and frequency-domain features of the pure audio signals from different time scales respectively, and use the cross-domain fusion of the two features as the training set to train the speaker model; Endpoint detection module: Real-time collect voice signals, perform endpoint detection processing on the voice signals, and remove pure noise and silence in the voice signals; Separate voice segment module: Estimate the noise level and calculate the weights of each time frame and different frequencies of the voice signal after removing pure noise and silence according to the noise power of different background noises in the dispatching desk, so as to perform weighted processing on the voice signal after removing pure noise and silence; Use the pre-trained continuous downsampling and resampling algorithm of multi-resolution features for the weighted voice signal to identify and separate voice segments, and the voice segments include other people's voices and the target speaker's voice; Speech recognition module: Judge whether the voice segment contains the target speaker through the speaker model; Target speaker voice extraction module: If the voice segment contains the target speaker, the target speaker voice is extracted through region proposal, clustering, and speaker model. A voiceprint mapping table for the working scenario period of the staff is established, and a comprehensive loss function is constructed based on this table. Using this function and the extracted target speaker voice, the speaker model is updated.

[0020] The beneficial effects of the present invention are as follows. Compared with the prior art, the present invention takes into account the differences in the working state voices of different staff members at the control console, as well as various background noises of different frequency bands at the control console. The time-domain waveform features and frequency-domain features of these audio signals are extracted from different time scales, and the two features are fused as the training set to train the speaker model. At the same time, the time-domain and frequency-domain features of the speech signal are utilized, making the model recognition more accurate; endpoint detection processing is performed through the voice activity detection algorithm; and in order to adapt to the environmental changes of control and dispatching consoles at all levels, weights of different frequencies in different time frames are calculated through noise level estimation and noise power weighting, enhancing the noise characteristics and significantly reducing environmental noise, especially the high-frequency noise in the main frequency band; target speaker speech extraction based on the speaker model and clustering is adopted to achieve accurate recognition and extraction of human voices. By establishing a voiceprint mapping table for the working scenario period of the staff and self-learning and updating according to the comprehensive loss function in this table, adaptability to the environment is realized, and the target speaker segment can be accurately recognized and extracted, especially the recognition rate for the control console environment is greatly improved. BRIEF DESCRIPTION OF THE DRAWINGS

[0021] Figure 1 It is a flowchart of the method of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0022] To make the objectives, technical solutions, and advantages of the present invention clearer, the technical solutions of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. The embodiments described in this application are only a part of the embodiments of the present invention, rather than all embodiments. Based on the spirit of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present invention.

[0023] As Figure 1 shown, Embodiment 1 of the present invention proposes an audio recognition and noise reduction processing method for the environment of a power grid control workbench, which is characterized by including: Obtain the pure audio signals of all power grid control staff in different working scenarios and different background noises in the dispatching console. Extract the time-domain waveform features and frequency-domain features of the pure audio signals from different time scales respectively. After cross-domain fusion of the two features, use them as the training set to train the speaker model. It should be noted that the features obtained after cross-domain fusion of each pure audio signal of the same power grid control staff in different working scenarios are used as a training sample of the power grid control staff.

[0024] Collect voice signals in real time, perform endpoint detection on the voice signals, and remove pure noise and silence in the voice signals. Estimate the noise level and the noise power of different background noises in the dispatching console, calculate the weights of each time frame and different frequencies of the voice signal after removing pure noise and silence, so as to perform weighted processing on the voice signal after removing pure noise and silence. Adopt the continuous downsampling and resampling algorithms of pre-trained multi-resolution features for the weighted voice signal to identify and separate the human voice segments, where the human voice segments include the voices of others and the target speaker's voice; use the speaker model to judge whether the human voice segment contains the target speaker. If the human voice segment contains the target speaker, then extract the target speaker's voice through region proposal, clustering and the speaker model, establish a voiceprint mapping table for the working scenario period of the staff, construct a comprehensive loss function according to the table, and use it and the extracted target speaker's voice to update the speaker model.

[0025] Preferably, different working scenarios of power grid control staff include: monitoring and dispatching, fault handling, safety analysis; Different background noises in the dispatching console include transformer noise, alarm noise, telephone ring noise, server noise, personnel noise; The server noise includes the noise of air conditioners, fans, and power supplies, and the personnel noise includes the sounds of personnel walking, flipping through documents, and typing on keyboards.

[0026] Preferably, extract the time-domain waveform features and frequency-domain features of these audio signals from different time scales, specifically: Divide the audio signal into several long-time window signals with a set duration T1 as one frame; divide each long-time window signal into several medium-time window signals with a set duration T2 as one frame; divide each medium-time window signal into several short-time window signals with a set duration T3 as one frame; calculate the waveform skewness of all long-time window signals, the energy envelope of medium-time window signals, and the zero-crossing rate of short-time window signals as time-domain waveform features; The formula is as follows:

[0027] Among them, When the independent variable of the function is greater than 0, output 1; when it is less than 0, output -1; when it is equal to 0, output 0; is the zero-crossing rate of the short-time window signal of the is the energy envelope of the medium-time window signal of the is the waveform skewness of the long-time window signal of the is the b th a sampling value of the corresponding time window signal of the N is the total number of sampling values included in a frame of short-time window signal, n is the n th sampling value in a frame of short-time window signal; T2 / T3 is the number of short-time window signals included in the medium-time window signal, and k is the kth short-time window signal included in the medium-time window signal at this time; is the set of all sampling values included in the long-time window signal of the and are respectively the mean and standard deviation of all sampling values in is the expectation obtained by subtracting the cube of the mean from all sampling values included in the long-time window signal of the Extract the Mel spectrogram and phase derivative spectrum of these audio signals as frequency domain features.

[0028] Preferably, the speaker model is specifically: The speaker model is a convolutional neural network; when performing cross-domain attention fusion, calculate the fusion feature after multi-head attention weighting of the waveform time feature to the frequency domain feature and the fusion feature after multi-head attention weighting of the frequency domain feature to the waveform time feature, and splice the two fusion features; When calculating the fusion feature after multi-head attention weighting of feature A to feature B, first obtain the query vector that maps feature A to a set dimension, and the key vector that maps feature B to a set dimension and the value vector ; Calculate the query matrix of feature A to feature B,

[0029] Calculate the multi-head attention weight of feature A to feature B, and the formula is:

[0030] where, a vector composed of feature A; a vector composed of feature B, is the dimension of.

[0031] Preferably, the process of endpoint detection processing is specifically as follows: Step 1: Collect the voice signal through a microphone. After collecting the voice signal within a set time period, input the collected voice signal into the voice activity detection algorithm model for judgment, and output the type of the judged voice signal; Step 2: If the voice type is a non-voice signal, it is judged that it is pure noise or silence at this time, and the voice collection of the microphone is stopped; if it is a voice signal, go to Step 3 and continue to use the microphone to collect the voice signal; Step 3: During the voice collection process, sequentially pass the voice signal collected in the next time period through the voice activity detection algorithm to judge whether a non-voice signal is detected. If no non-voice signal is detected, continue to collect the signal. If a non-voice signal is detected, go to Step 4; Step 4: Continue to collect the voice signals collected within the set M time periods. The voice signals collected in each time period are input into the voice activity detection algorithm for detection; if a voice signal is detected, the task is still in the voice process, and return to Step 3; if there is no voice signal during the M time periods, it is considered that the speech ends, and the microphone collection is stopped.

[0032] Preferably, the voice activity detection algorithm is specifically as follows: After pre-emphasizing, framing, and windowing the voice signal, convert it into a frequency-domain signal through Mel Frequency Cepstral Coefficients (MFCC); Divide the frequency-domain signal into six sub-bands according to frequency. The frequency ranges of these sub-bands are 80Hz - 250Hz, 250Hz - 500Hz, 500Hz - 1K, 1K - 2K, 2K - 3K, and 3K - 4K respectively; for each sub-band, calculate its energy; use the probability density function of the trained Gaussian model to perform operations on the energy of these six sub-bands and the weighted sum of the energies of the six sub-bands respectively to obtain the corresponding log-likelihood ratio; among them, the log-likelihood functions calculated from the energies of the six sub-bands are all local log-likelihood functions, and the log-likelihood ratio calculated from the weighted sum of the energies of the six sub-bands is the global log-likelihood ratio; When performing voice decision-making, judge all local and global log-likelihood ratios. If any local or global log-likelihood ratio exceeds the set log-likelihood ratio threshold, it is considered that there is voice; among them, the weights and thresholds of the weighted energies of the six sub-bands are obtained through training; after each detection, the algorithm self-learns and updates the mean and variance parameters of the Gaussian model.

[0033] Preferably, the noise power of different background noises in the noise level estimation and the dispatching console is calculated, and the weights of each time frame and different frequencies of the speech signal after removing pure noise and silence are calculated, specifically as follows: The speech signal after removing pure noise and silence and different background noises in the dispatching console are framed with a set duration T4 as one frame, and decomposed into a number of frequency points through short-time Fourier transform; Calculate the noise level estimation of the speech signal after removing pure noise and silence for each frame and the noise power of different background noises in the dispatching console; Calculate the th i frequency point of the th frame of the speech signal after removing pure noise and silence The formula for the weight

[0034] where , , are all regularization coefficients; is the set noise level estimation threshold, is the maximum noise level estimation, is the sum of the noise powers of the background noises dominant at different frequencies, , , , , respectively represent the noise powers of transformer noise, alarm noise, telephone ringtone noise, server noise, and personnel noise.

[0035] Calculate the noise level estimation of the speech signal after removing pure noise and silence for each frame and the noise power of different background noises in the dispatching console, specifically as follows: The noise level estimation of the th frame

[0036] where is the total number of frequency points of the short-time Fourier transform; is the th i frequency point of the th frame The noise power of different background noises in the dispatching console is to calculate the noise level estimation of each frame of different background noises in the dispatching console and take its average value as its noise power.

[0037] Preferably, the target speaker voice is extracted through region proposal, clustering, and a speaker model, specifically as follows: The region proposal algorithm is used to generate speech segments of multiple non-overlapping regions based on the voice segment. This algorithm is a fully convolutional network model that uses a 3*3 sliding window to slide the voice segment to generate multiple speech segments. Clustering is used to form multiple clusters from the speech segments of multiple non-overlapping regions. Each speech segment in each cluster is input into the speaker model for recognition. The number of speech segments of the target speaker contained in each cluster is counted, and all the speech segments of the cluster with the largest number are integrated into a speech signal, which is the extracted target speaker voice.

[0038] Preferably, the staff work scenario time period voiceprint mapping table is established, and based on this table, a comprehensive loss function is constructed. Using it and the extracted target speaker voice, the speaker model is updated, specifically as follows: According to the duty schedules and historical work record information of all power grid dispatching staff, a staff work scenario time period voiceprint mapping table is established. This table establishes the mapping relationship between the most frequent work scenarios in each work period of each power grid dispatching staff, the corresponding cross-domain fusion feature vectors of the pure audio signals under the corresponding work scenarios, and the short-time Fourier transform complex spectra. The time-domain waveform features and frequency-domain features of the target speaker voice are extracted from different time scales respectively, and they are cross-domain fused into a feature vector. According to the staff work scenario time period voiceprint mapping table, the feature vector of the pure audio signal of the target speaker in the current work scenario is obtained. And the short-time Fourier transform complex spectrum. ; The comprehensive loss function is:

[0039] Among them, , are set parameters, and their values are all in the interval [0,1]. is the norm. is transpose of, is the short-time Fourier transform complex spectrum of the target speaker voice.

[0040] Embodiment 2 of the present invention provides an audio recognition and noise reduction processing system for the power grid dispatching workbench environment using the audio recognition and noise reduction method described in Embodiment 1 of the present invention, including: a speaker model construction module, an endpoint detection module, a voice segment separation module, a speech recognition module, and a target speaker voice extraction module, characterized in that: Speaker model construction module: Obtain the pure audio signals of all power grid dispatching staff in different working scenarios and different background noises in the dispatching console. Extract the time-domain waveform features and frequency-domain features of the pure audio signals from different time scales respectively. After cross-domain fusion of the two features, use them as the training set to train the speaker model; Endpoint detection module: Collect voice signals in real time, perform endpoint detection processing on the voice signals, and remove pure noise and silence in the voice signals; Separation of human voice segment module: Estimate the noise level and the noise power of different background noises in the dispatching console, calculate the weights of each time frame and different frequencies of the voice signal after removing pure noise and silence, so as to perform weighted processing on the voice signal after removing pure noise and silence; Use the pre-trained continuous downsampling and resampling algorithm of multi-resolution features for the weighted voice signal to identify and separate human voice segments, and the human voice segments include other people's voices and the target speaker's voice; Speech recognition module: Determine whether the human voice segment contains the target speaker through the speaker model; Target speaker voice extraction module: If the human voice segment contains the target speaker, then through region proposal, clustering and the speaker model, extract the target speaker voice, establish a voiceprint mapping table for the working scenario period of the staff, construct a comprehensive loss function according to this table, and use it and the extracted target speaker voice to update the speaker model.

[0041] The present disclosure may be a system, a method, and / or a computer program product. The computer program product may include a computer-readable storage medium having computer-readable program instructions thereon for causing a processor to implement various aspects of the present disclosure.

[0042] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those of ordinary skill in the art should understand that: It is still possible to modify the specific implementation manners of the present invention or make equivalent replacements, and any modification or equivalent replacement that does not depart from the spirit and scope of the present invention shall be covered by the protection scope of the claims of the present invention.

Claims

1. An audio recognition and noise reduction processing method for a power grid control workbench environment, characterized in that: include: Obtain the pure audio signals of all power grid control staff in different working scenarios and different background noises in the dispatching desk, extract the time domain waveform features and frequency domain features of all pure audio signals from different time scales, and use the two features as training sets to train the speaker model after cross-domain fusion; collect voice signals in real time, perform endpoint detection on the voice signals, and remove pure noise and silence in the voice signals; According to the noise level estimation and the noise power of different background noises in the dispatching station, the weights of each time frame and different frequencies of the speech signal after pure noise and silence are removed are calculated, so as to perform weighted processing on the speech signal after pure noise and silence are removed; Using a pre-trained continuous downsampling and resampling algorithm of multi-resolution features on the weighted speech signal, identifying and separating vocal segments, wherein the vocal segments include side person voices and target speaking voices; Using the speaker model, determine whether the vocal segment contains the target speaker; If the vocal segment contains the target speaker, the target speaker’s voice is extracted through region proposal, clustering and speaker model, and a voiceprint mapping table for the staff’s work scene and time period is established. A comprehensive loss function is constructed based on the table, and the speaker model is updated using it and the extracted target speaker’s voice.

2. The method for audio recognition and noise reduction in a power grid control workbench environment according to claim 1, characterized in that: The different work scenarios of power grid control staff include: monitoring and dispatching, fault handling, and safety analysis; Different background noises in the dispatch desk include transformer noise, alarm noise, phone ringing noise, server noise, and personnel noise; The server noise includes the noise of air conditioners, fans, and power supplies, and the personnel noise includes the sounds of people walking, flipping through documents, and typing on keyboards.

3. The method for audio recognition and noise reduction in a power grid control workbench environment according to claim 2, characterized in that: Extract the time domain waveform features and frequency domain features of pure audio signals from different time scales, specifically: The audio signal is divided into a plurality of long-time window signals with a set time length T1 as a frame; each long-time window signal is divided into a plurality of medium-time window signals with a set time length T2 as a frame; each medium-time window signal is divided into a plurality of short-time window signals with a set time length T3 as a frame; the waveform skewness of all long-time window signals, the energy envelope of the medium-time window signal and the zero-crossing rate of the short-time window signal are calculated as time domain waveform features; The Mel spectrum and phase derivative spectrum of these audio signals are extracted as frequency domain features.

4. The method for audio recognition and noise reduction in a power grid control workbench environment according to claim 3, characterized in that: The two features are fused across domains and used as a training set to train the speaker model, specifically: The speaker model is a convolutional neural network; The cross-domain fusion of the two features is to respectively calculate the attention weight of the time domain waveform feature to the frequency domain feature, and the attention weight of the frequency domain feature to the time domain waveform feature; When calculating the attention weight of the time domain waveform feature to the frequency domain feature, first obtain a query vector that maps the time domain waveform feature to a set dimension, and a key vector and a value vector that map the frequency domain feature to a set dimension, and use them to calculate the attention weight of the time domain waveform feature to the frequency domain feature; When calculating the attention weight of the frequency domain feature to the time domain waveform feature, first obtain a query vector that maps the frequency domain feature to a set dimension, and a key vector and a value vector that map the time domain waveform feature to a set dimension, and use them to calculate the attention weight of the time domain waveform feature to the frequency domain feature; Multiply the time domain waveform feature by the attention weight of the time domain waveform feature to the frequency domain feature, multiply the frequency domain feature by the attention weight of the frequency domain feature to the time domain waveform feature, and concatenate the two calculation results to form the fused feature.

5. The method for audio recognition and noise reduction in a power grid control workbench environment according to claim 1, characterized in that: The specific process of endpoint detection processing is as follows: Step 1: Collect voice signals through a microphone. After collecting voice signals within a set time period, input the collected voice signals into a voice activity detection algorithm model for judgment, and output the judged voice signal type; Step 2: If the voice type is a non-voice signal, it is determined to be pure noise or silence at this time, and the voice collection of the microphone is stopped; if it is a voice signal, go to step 3 and continue to use the microphone to collect the voice signal; Step 3: During the voice collection process, the voice signals collected in the next time period are sequentially passed through the voice activity detection algorithm to determine whether a non-voice signal is detected. If no non-voice signal is detected, the signal collection continues. If a non-voice signal is detected, the process proceeds to step 4. Step 4: Continue to collect voice signals collected within the set M time periods. The voice signals collected in each time period are input into the voice activity detection algorithm for detection; if a voice signal is detected, the task is still in the voice process and returns to step 3; if there is no voice signal within M time periods, it is considered that the speech is over and the microphone collection is stopped.

6. The method for audio recognition and noise reduction in a power grid control workbench environment according to claim 5, characterized in that: The voice activity detection algorithm is specifically: After the speech signal is pre-processed by pre-emphasis, framing and windowing, it is converted into a frequency domain signal through Mel-frequency cepstral coefficients (MFCC). The frequency domain signal is divided into six sub-bands according to the frequency; for each sub-band, its energy is calculated; the probability density function of the trained Gaussian model is used to calculate the energy of the six sub-bands and the weighted sum of the energy of the six sub-bands, and the corresponding log-likelihood ratio is obtained; among which, the energy of the six sub-bands is calculated as the local log-likelihood function, and the weighted sum of the energy of the six sub-bands is calculated as the global log-likelihood ratio; When making speech judgments, all local and global log-likelihood ratios are judged. If any local or global log-likelihood ratio exceeds the set log-likelihood ratio threshold, it is considered that speech exists. The weights and thresholds of the six sub-band energy weightings are obtained through training. After each detection, the algorithm performs self-learning updates on the mean and variance parameters of the Gaussian model.

7. The method for audio recognition and noise reduction in a power grid control workbench environment according to claim 2, characterized in that: The weights of each time frame and different frequencies of the speech signal after removing pure noise and silence are calculated based on the noise level estimation and the noise power of different background noises in the dispatching station, specifically: The voice signal after pure noise and silence are removed, and the different background noises in the dispatching console are divided into frames with a set duration T4 as one frame, and decomposed into several frequency points through short-time Fourier transform; Calculate the noise level estimate of the speech signal after removing pure noise and silence for each frame and noise power of different background noises in the dispatch desk; Calculate the speech signal after removing pure noise and silence Frame No. i The frequency corresponding to the frequency point Weight The formula is: in, , , are regularization coefficients; is the noise level estimation threshold set, is the maximum noise level estimate, is the sum of the noise powers of the dominant background noise at different frequencies, , , , , They represent the noise power of transformer noise, alarm noise, telephone ringing noise, server noise, and personnel noise respectively.

8. The method for audio recognition and noise reduction in a power grid control workbench environment according to claim 7, characterized in that: Calculate the noise level estimate of the speech signal after removing pure noise and silence for each frame The noise power of the background noise different from that in the dispatch station is: No. Noise level estimation of a frame The calculation formula is: in, is the total number of frequency points of short-time Fourier transform; For the Frame No. i The frequency corresponding to the frequency point The complex spectrum value of ; The noise power of different background noises in the dispatching station is to calculate the noise level estimation of each frame of different background noises in the dispatching station, and calculate the average value as its noise power.

9. The method for audio recognition and noise reduction in a power grid control workbench environment according to claim 1, characterized in that: The target speaker voice is extracted through region proposal, clustering and speaker model, specifically: A region proposal algorithm is used to generate multiple non-overlapping speech segments based on the vocal segments. The algorithm is a fully convolutional network model that uses a 3*3 sliding window to slide the vocal segments to generate multiple speech segments. Clustering is used to form multiple clusters from speech clips in multiple non-overlapping areas. Each speech clip in each cluster is input into the speaker model for recognition. The number of speech clips of the target speaker contained in each cluster is counted, and all speech clips of the cluster with the largest number are integrated into a speech signal, which is the extracted voice of the target speaker.

10. The method for audio recognition and noise reduction in a power grid control workbench environment according to claim 2, characterized in that: The voiceprint mapping table of the staff's working scene and time period is established, and a comprehensive loss function is constructed according to the table, and the speaker model is updated by using the comprehensive loss function and the extracted target speaker voice, specifically: According to the duty schedule and historical work record information of all power grid control staff, a voiceprint mapping table for staff working scenes and periods is established. This table establishes the mapping relationship between the most common working scenes in each working period of each power grid control staff, the cross-domain fused feature vector corresponding to the pure audio signal in the corresponding working scene, and the short-time Fourier transform complex spectrum; The time domain waveform features and frequency domain features of the target speaker are extracted from different time scales respectively, and then fused across domains into a feature vector. ; According to the staff working scene period voiceprint mapping table, obtain the characteristic vector of the pure audio signal of the target speaker in the working scene of the current period and the short-time Fourier transform complex spectrum ; The comprehensive loss function is: in, , are the set parameters, and their values ​​are all in the range of [0,1]. is the norm, for The transpose of is the short-time Fourier transform complex spectrum of the target speaker's voice.

11. An audio recognition and noise reduction processing system in a power grid control workbench environment based on the audio recognition and noise reduction method according to any one of claims 1 to 10, comprising: Speaker model building module, endpoint detection module, voice segment separation module, speech recognition module, target speaker voice extraction module, characterized by: Speaker model construction module: obtain the pure audio signals of all power grid control staff in different working scenarios and different background noises in the dispatching desk, extract the time domain waveform features and frequency domain features of the pure audio signals from different time scales, and use the two features as the training set to train the speaker model after cross-domain fusion; Endpoint detection module: collects voice signals in real time, performs endpoint detection on the voice signals, and removes pure noise and silence from the voice signals; Separation module of human voice segments: according to the noise level estimation and the noise power of different background noises in the dispatching station, the weights of each time frame and different frequencies of the speech signal after the pure noise and silence are removed are calculated, so as to perform weighted processing on the speech signal after the pure noise and silence are removed; the continuous downsampling and resampling algorithm of the pre-trained multi-resolution features is used for the weighted speech signal to identify and separate the human voice segments, which include the voices of the bystanders and the target speaker; Speech recognition module: determines whether the vocal segment contains the target speaker through the speaker model; Target speaker voice extraction module: If the voice segment contains the target speaker, the target speaker voice is extracted through region proposal, clustering and speaker model, and a voiceprint mapping table for the staff's work scene and time period is established. A comprehensive loss function is constructed based on the table, and the speaker model is updated using it and the extracted target speaker voice.

Citation Information

Patent Citations

  • Audio recognition method and apparatus, computer-readable storage medium

    CN109686377B

  • Voice noise reduction method and device based on voiceprint recognition, equipment and medium

    CN116312570A

  • Voice denoising method based on audio recognition

    CN101404160A

  • Target speaker real-time voice information extraction method based on voiceprint features

    CN115240688A

  • Speech recognition hybrid model construction method and system for power grid equipment monitoring

    CN119380714A

Cited By

  • Intelligent ward intercom calling system with background noise suppression function

    CN120877758A

  • Power grid dispatching telephone audio corpus processing and model adaptive iterative training method and system

    CN122201258A