Audio processor-based voice data processing method
By analyzing ambient background noise and real-time user speech, and combining this with convolutional neural networks to recognize user voices, the problem of low background noise filtering efficiency is solved, achieving efficient audio data processing and clean audio acquisition.
Patent Information
- Application Number
- CN202510157592.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-02-13
- Publication Date
- 2025-12-09
- Estimated Expiration
- 2045-02-13
AI Technical Summary
The lack of simple and effective methods to filter out background noise in existing technologies leads to low efficiency in voice data processing.
By continuously acquiring ambient background noise, dividing it into equal-length segments, and analyzing the similarity of waveforms, the identified background noise is determined. This is then filtered out by combining real-time voice input from the user, and a convolutional neural network is used to recognize the user's voice to eliminate background noise.
It effectively processes audio data, enhances the user's voice interaction experience, and yields clean audio data.
Smart Images

Figure CN119993187B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The application belongs to the technical field of speech data processing, and particularly relates to a speech data processing method based on an audio processor. BACKGROUND
[0002] A patent with the publication number CN105869645A discloses a speech data processing method and device. The method comprises the following steps: obtaining an I-Vector vector of each speech sample in a plurality of speech samples, and determining a target seed sample in the plurality of speech samples; respectively calculating a cosine distance between the I-Vector vector of the target seed sample and an I-Vector vector of a target remaining speech sample, the target remaining speech sample being a speech sample other than the target seed sample in the plurality of speech samples; and filtering a target speech sample from the plurality of speech samples or the target remaining speech sample according to the cosine distance, the cosine distance between the I-Vector vector of the target speech sample and the I-Vector vector of the target seed sample being higher than a first predetermined threshold. The application solves the technical problem of low speech data cleaning efficiency caused by the fact that the related art cannot clean the speech data by using a manual labeling method.
[0003] However, for speech data processing, there is a lack of a relatively simple method for filtering out background noise or distinguishing and reasonably filtering out background noise. Therefore, a technical solution is provided. SUMMARY
[0004] The present application aims to at least solve one of the technical problems existing in the prior art. To this end, the present application provides a speech data processing method based on an audio processor, which specifically comprises the following steps:
[0005] Continuously obtaining background noise of an environment, and retaining a period of background noise before the appearance of a target signal from the background noise, the target signal being used to indicate a user's demand for speech input;
[0006] Dividing the background noise into a plurality of segments of equal length, and determining a specified background noise or generating a synchronous filtering signal according to the discrete degree and mean value of the similarity data of the waveform graph of each environmental sound and other waveform graphs;
[0007] Using the sound collected in a space where the user inputs real-time speech as the specified background noise when the synchronous filtering signal is generated;
[0008] Then, deleting the part of the user's input real-time speech that belongs to the specified background noise, obtaining processed data, and marking the processed data as noise-reduced speech data.
[0009] Further, after obtaining the noise-reduced speech data, if the user does not input new speech signals within a period of time, the generation of the target signal is re-detected.
[0010] Further, the specific duration of a period is set by an administrator.
[0011] Further, the way of determining the authorized background sound or generating the synchronous filtering signal according to several environmental sounds is as follows:
[0012] All environmental sounds and their corresponding waveform graphs are obtained;
[0013] One of all the environmental sounds is selected in turn and marked as a selected environmental sound. When a selected environmental sound is obtained, the similarity of the waveform graphs of other environmental sounds to the waveform graph of the selected environmental sound is calculated, obtaining several similarities. The mean value of the similarities and a stability value representing the dispersion degree of the similarities are calculated;
[0014] Then, the mean value is proportional to the stability value, and the mean value and the stability value are multiplied by a corresponding weight value and added to obtain a value marked as a sure selection value;
[0015] The sure selection values of all selected environmental sounds are obtained, and the selected environmental sound with the highest sure selection value is marked as the authorized background sound.
[0016] Further, before marking the selected environmental sound with the highest sure selection value as the authorized background sound, the selected environmental sound needs to be screened. The specific screening method is:
[0017] All selected environmental sounds with a mean value exceeding a set value X1 and a stability value lower than a set value X2 are screened out.
[0018] Further, when no selected environmental sound can be screened out, a synchronous filtering signal is generated.
[0019] Further, the specific way of determining the authorized background sound when the synchronous filtering signal is generated is:
[0020] The speech collection device farthest from the aforementioned speech collection device collecting the target signal in the same environment is automatically obtained to collect real-time spatial sound, which is marked as the authorized background sound.
[0021] Further, the specific way of determining the authorized background sound when the synchronous filtering signal is generated is:
[0022] The speech collection device farthest from the aforementioned speech collection device collecting the target signal in the same environment is automatically obtained to collect real-time spatial sound, which is marked as the authorized background sound.
[0023] Then the user voice in the space sound is suppressed to the minimum or eliminated, and the remaining is as the authorized background sound.
[0024] Further, the user voice is identified through an intelligent identification model, and the intelligent identification model is a convolutional neural network (CNN) model.
[0025] Further, the specific training method of the intelligent identification model is as follows:
[0026] A plurality of user voices without interference are obtained from the corresponding environment as training data, and other user voices and arbitrary voices are obtained as verification data.
[0027] The plurality of training data are input into the model for training, and after the training is completed, the verification data are input into the model for verification of the accuracy rate. When the accuracy rate exceeds a set value B1, it is indicated that the model training is completed, otherwise, the plurality of training data are reacquired, and the parameters in the model are adjusted until the accuracy rate exceeds the set value B1.
[0028] The training of the intelligent identification model is completed.
[0029] Compared with the prior art, the present application has the following beneficial effects:
[0030] The present application can effectively process audio data and obtain a good pure audio data, and can improve the experience of user voice interaction, by continuously obtaining the background sound of the environment, retaining the background sound of one period before the target signal appears from the background sound, extracting the authorized background sound from the background sound, or using the sound collected in the space when the user inputs real-time voice as the authorized background sound, and then processing the audio data. BRIEF DESCRIPTION OF DRAWINGS
[0031] Figure 1 The system block diagram of the present application is shown in the figure. DETAILED DESCRIPTION
[0032] The technical solutions of the present application will be described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present application, but not all. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative labor are within the scope of protection of the present application.
[0033] Please refer to Figure 1 The present application provides a voice data processing method based on an audio processor.
[0034] As an embodiment of the present application, the method specifically includes the following steps:
[0035] Step one: before processing data, in the case of no human voice, analyze the background sound and extract several environmental sounds from the background sound, in the specific way:
[0036] First, set the audio acquisition device to keep the time to obtain the sound of the environment, and mark it as background sound;
[0037] Continuously acquire background sound, and then set a period according to T1 time; T1 is a preset value, which is set by the administrator in advance;
[0038] When the audio duration of the background sound exceeds two periods, if no target signal is detected at the latest time, delete all voice data one period ago; Then continuously acquire the background sound, if the audio duration of the background sound acquired subsequently exceeds two periods, and no target signal is detected, automatically delete all background sounds one period ago from the nearest time, until the target signal is detected;
[0039] The target signal is generally the wake-up voice corresponding to the device, of course, it can also be other wake-up signals, indicating that the user needs to perform voice interaction or other voice input behavior;
[0040] After detecting the target signal, the background sound one period before the target signal is generated is automatically divided into several environmental sounds with the same duration; The specific number is set by the administrator, and is guaranteed to be no less than ten;
[0041] Step two: analyze the similarity of the obtained several environmental sounds, and the specific way of similarity analysis is:
[0042] First, obtain all environmental sounds and their corresponding waveform graphs, and then select an environmental sound and mark it as the selected environmental sound, obtain its waveform graph, and then obtain the similarity of the waveform graphs of the other environmental sounds and the selected environmental sound, to obtain several similarities Di, i = 1,..., n;
[0043] Then obtain the mean value of Di, mark it as P, and calculate the stability value W of all Di using the formula, the specific calculation formula is:
[0044]
[0045] Then select the next environmental sound in order, mark it as the selected environmental sound, calculate the mean value and stability value of the similarity of the waveform graphs of the remaining other environmental sounds; Repeat to obtain the mean value and stability value corresponding to when all environmental sounds are marked as the selected environmental sound;
[0046] All selected environmental sounds with mean value exceeding a set value X1 and stable value below a set value X2 are screened out, and then the selected value corresponding to all screened selected environmental sounds is calculated according to the formula, and the specific formula is:
[0047] Selected value = 0.63*mean value + 0.37 / stable value
[0048] The selected environmental sound with the highest selected value is marked as the approved background sound;
[0049] If no selected environmental sound is screened out, a synchronous filtering signal is generated at this time;
[0050] Step three: after obtaining the approved background sound, the real-time voice entered by the user is automatically obtained after the user enters the target signal;
[0051] Then the approved background sound in the real-time voice is eliminated to obtain the processed data, which is marked as the noise reduction voice data;
[0052] The basic processing of the voice data is completed;
[0053] Step four: if the synchronous filtering signal is generated, the voice collection device farthest from the voice collection device collecting the target signal in the same environment is automatically obtained to collect real-time spatial sound, which is marked as the approved background sound;
[0054] Then the approved background sound in the real-time voice is eliminated to obtain the processed data, which is marked as the noise reduction voice data;
[0055] The basic processing of the voice data is completed;
[0056] Step five: after the processing of the voice data is completed, it is continuously monitored whether the user is still continuously entering voice, if no new voice signal entered by the user is detected within T1 time, that is, a period of time, then jump to step one for reprocessing;
[0057] As an embodiment of the present application, the embodiment is implemented on the basis of embodiment one, and the difference lies in that, in the embodiment, the processing method when the synchronous filtering signal is generated is different, and the specific method in the embodiment is:
[0058] If the synchronous filtering signal is generated, the voice collection device farthest from the voice collection device collecting the target signal in the same environment is automatically obtained to collect real-time spatial sound, which is marked as the spatial sound;
[0059] Then the user sound in the spatial sound is suppressed to the lowest or eliminated, and the remaining is marked as the approved background sound, and then the approved background sound in the real-time voice is eliminated to obtain the processed data, which is marked as the noise reduction voice data;
[0060] Basic processing of voice data is completed;
[0061] The user's voice is identified by an intelligent recognition model, which is determined by the following method:
[0062] A convolutional neural network (CNN) is selected as the basic model. A number of interference-free user voice data is obtained from the corresponding environment as training data, and other user voices and arbitrary sound mixed with a set number of user voices are obtained as verification data.
[0063] The training data is first identified. First, the original audio signal is converted into a frequency spectrum or mel frequency cepstral coefficient (MFCC) to serve as the input of the CNN.
[0064] The audio data is enhanced, such as adding noise and changing the volume, to improve the generalization ability of the model.
[0065] The user selects a suitable CNN architecture for audio processing, including convolutional layers, pooling layers, and fully connected layers. A pre-trained CNN model, such as VGG19, is used to improve the accuracy of voice recognition through transfer learning. The convolutional layer of the CNN automatically extracts local features in the audio signal, such as frequency and time changes. Through the pooling layer and the fully connected layer, the CNN can integrate local features to form a global feature representation. Cross-entropy loss is selected to guide the training process of the CNN. Gradient descent or other optimization algorithms are used to adjust the weights of the CNN to minimize the loss function.
[0066] Then, a number of training data are input into the model for training. After training is completed, the verification data are input into the model to verify the accuracy. When the accuracy exceeds a set value B1, the model training is complete. Otherwise, a number of training data are reacquired, and the parameters in the model are adjusted until the accuracy exceeds the set value B1.
[0067] The training of the intelligent recognition model is completed.
[0068] Then, the user's voice is extracted from the spatial sound using the intelligent model.
[0069] Some of the data in the above formula are dimensionless values calculated by removing the dimension. The formula is obtained by software simulation of a large amount of collected data to obtain a formula closest to the real situation. The preset parameters and preset thresholds in the formula are set by a person skilled in the art according to the actual situation or obtained by a large amount of data simulation.
[0070] The above examples are only used to illustrate the technical method of the present application but not limit the present application. Although the present application is described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present application can be modified or equivalently replaced without departing from the spirit and scope of the technical method of the present application.
Claims
1. A voice data processing method based on an audio processor, characterized by, The method specifically comprises the following steps: Continuously acquire the background sound of the environment, and retain the background sound before the appearance of the target signal from the background sound, the target signal being used to indicate the user's demand for voice input; Divide the background sound into several segments of environment sound with equal length, determine the authorized background sound or generate the synchronous filtering signal according to the discrete degree and mean value of the similarity data of the waveform graph of each environment sound and other waveform graphs; In the process of generating the synchronous filtering signal, the sound collected in the space where the user inputs real-time voice is used as the authorized background sound; Then, the part of the real-time voice input by the user that belongs to the authorized background sound is deleted to obtain processed data, which is marked as noise-reduced voice data; The way of determining the authorized background sound or generating the synchronous filtering signal according to several segments of environment sound is as follows: Acquire all environment sounds and their corresponding waveform graphs; Select one from all environment sounds in turn and mark it as the selected environment sound, and then calculate the similarity of the waveform graph of the selected environment sound and the waveform graph of other environment sounds to obtain several similarities, and then calculate the mean value of the similarities and the stability value representing the discrete degree of the similarities; Then, multiply the mean value and the stability value by a corresponding weight value in a manner that is proportional to the mean value and inversely proportional to the stability value, and then add them to obtain a value marked as the selected value; Obtain the selected values of all selected environment sounds, and mark the selected environment sound with the highest selected value as the authorized background sound.
2. The voice data processing method based on an audio processor according to claim 1, wherein, After obtaining the noise-reduced voice data, if the user does not input new voice signals within one period of time, the generation of the target signal is re-detected.
3. The voice data processing method based on an audio processor according to claim 1, wherein, The specific length of one period is set by an administrator.
4. The voice data processing method based on an audio processor according to claim 1, wherein, Before marking the selected environment sound with the highest selected value as the authorized background sound, the selected environment sound needs to be screened, and the specific screening method is as follows: Screen all selected environment sounds with a mean value higher than a set value X1 and a stability value lower than a set value X2.
5. The audio processor-based voice data processing method of claim 4, wherein, When no selected environment sound can be screened, the synchronous filtering signal is generated.
6. The voice data processing method based on an audio processor according to claim 1, wherein, The specific way of determining the authorized background sound when the synchronous filtering signal is generated is as follows: The voice collection device farthest from the voice collection device of the preceding target signal in the same environment is automatically acquired to collect real-time space sound, which is used as the authorized background sound.
7. The audio processor-based voice data processing method of claim 1, wherein, The specific way of determining the authorized background sound when the synchronous filtering signal is generated is as follows: The voice collection device farthest from the voice collection device of the preceding target signal in the same environment is automatically acquired to collect real-time space sound, which is marked as the space sound; Then, the user sound in the space sound is suppressed to the lowest or eliminated, and the remaining part is used as the authorized background sound.
8. The audio processor-based voice data processing method of claim 7, wherein, The user sound is identified by an intelligent identification model, and the intelligent identification model is a convolutional neural network (CNN) model.
9. The voice data processing method based on an audio processor according to claim 1, wherein, The specific training method of the intelligent identification model is as follows: Acquire several data of user sound without interference in the corresponding environment as training data, and then acquire other user sound and arbitrary sound mixed with a set number of user sound as verification data; The model is trained by inputting a plurality of training data, and after the training is completed, the verification data is input into the model to verify the accuracy, and when the accuracy exceeds a set value B1, it indicates that the model training is completed, otherwise a plurality of training data are reacquired, and the parameters in the model are adjusted until the accuracy exceeds the set value B1. The training of the intelligent identification model is completed.
Citation Information
Patent Citations
Voice data processing method and device
CN105869645A
Voice noise reduction method and system for interaction, electronic equipment and storage medium
CN115376538A
Methods, apparatus, and non-transitory computer readable medium for audio processing
US20230080446A1