Voice data processing method based on audio processor

By analyzing the waveform similarity and discretization of the ambient sound, determining the background sound, and eliminating the corresponding background sound when the user enters the voice, the problem of low background sound filtering efficiency in the prior art is solved, and effective noise reduction of voice data and improved user experience is achieved.

CN119993187AActive Publication Date: 2025-05-13SUZHOU LEADER INTELLIGENT TECH CO LTD
View PDF 6 Cites 0 Cited by

Patent Information

Application Number
CN202510157592.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-02-13
Publication Date
2025-05-13
Estimated Expiration
2045-02-13

AI Technical Summary

Technical Problem

The prior art is difficult to effectively filter out background sounds, resulting in low efficiency in voice data processing.

Method used

By continuously obtaining the background sound of the environment, dividing it into equal-time environmental sound segments, analyzing the similarity and discretization of the waveform graph, determining the background sound, and eliminating the corresponding background sound when the user enters the voice, obtaining noise-reducing voice data.

Benefits of technology

It realizes effective noise reduction of voice data, improves the user's voice interaction experience, and the method is simple and easy to use.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119993187A_ABST
    Figure CN119993187A_ABST
Patent Text Reader

Abstract

The invention discloses a voice data processing method based on an audio processor, and relates to the technical field of voice data processing, and the method comprises the steps: continuously obtaining the background sound of an environment, and reserving the background sound of a period before a downloading target signal appears from the background sound, extracting a checking background sound from the background sound or taking the sound collected in the space when the user records the real-time voice as the checking background sound; according to the invention, the audio data can be effectively processed in a mode of processing the audio data, so that good pure audio data can be obtained, and the experience feeling of voice interaction of a user can be improved; the method is simple, effective, easy and practical.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of voice data processing, and in particular relates to a voice data processing method based on an audio processor. Background Art

[0002] The patent with publication number CN105869645A discloses a method and device for processing speech data. The method includes: obtaining the I-Vector vector of each speech sample in a plurality of speech samples, and determining a target seed sample in a plurality of speech samples; respectively calculating the cosine distance between the I-Vector vector of the target seed sample and the I-Vector vector of the target remaining speech sample, the target remaining speech sample being the speech sample other than the target seed sample in the plurality of speech samples; filtering the target speech sample from the plurality of speech samples or the target remaining speech samples at least according to the cosine distance, and the cosine distance between the I-Vector vector of the target speech sample and the I-Vector vector of the target seed sample is higher than a first predetermined threshold. The present invention solves the technical problem that the related art cannot use manual annotation methods to clean speech data, resulting in low speech data cleaning efficiency.

[0003] However, for voice data processing, there is a lack of a simple method for filtering out background sounds, or distinguishing background sounds and filtering them out reasonably. Based on this, a technical solution is provided. Summary of the invention

[0004] The present invention aims to solve at least one of the technical problems existing in the prior art; to this end, the present invention proposes a voice data processing method based on an audio processor, and the method specifically comprises the following steps:

[0005] Continuously obtain the background sound of the environment, and retain the background sound of one period before the target signal appears from the background sound. The target signal is used to indicate that the user has a need for voice recording;

[0006] The background sound is divided into several segments of ambient sound of equal duration, and the discrete degree and mean value of the data of the similarity between the waveform of each ambient sound and other waveforms are used to determine the background sound or generate a synchronous filtering signal;

[0007] When generating the synchronous filtering signal, the sound collected from the space where the user is recording the real-time voice is used as the verified background sound;

[0008] Then, the part of the real-time voice recorded by the user that belongs to the approved background sound is deleted to obtain processed data, which is marked as noise reduction voice data.

[0009] Furthermore, after the noise reduction voice data is obtained, if the user does not record a new voice signal within a period of time, the generation of the target signal is re-detected.

[0010] Furthermore, the specific duration of a cycle is set by an administrator.

[0011] Furthermore, the method of determining the background sound or generating the synchronous filtering signal according to the plurality of environmental sounds is as follows:

[0012] Get all the ambient sounds and their corresponding waveforms;

[0013] One is selected from all the ambient sounds in turn and marked as the selected ambient sound. For each selected ambient sound, the similarity between the waveforms of other ambient sounds and the waveform of the selected ambient sound is calculated to obtain several similarities. The mean value of the similarities and the stable value representing the degree of dispersion of the similarities are first calculated.

[0014] Then, in a manner proportional to the mean and inversely proportional to the stable value, the mean and stable value are multiplied by a corresponding weight value and then added together, and the resulting value is marked as the selected value;

[0015] The confirmation values ​​of all the selected ambient sounds are obtained, and the selected ambient sound with the highest confirmation value is marked as the approved background sound.

[0016] Furthermore, before marking the selected ambient sound with the highest confirmation value as the approved background sound, the selected ambient sound needs to be screened. The specific screening method is as follows:

[0017] Filter out all selected ambient sounds whose average value exceeds the set value X1 and whose stable value is lower than the set value X2.

[0018] Further, when any selected ambient sound cannot be screened out, a synchronous filtering signal is generated.

[0019] Furthermore, the specific method of determining the background sound when generating the synchronous filtering signal is as follows:

[0020] The voice collection device that is in the same environment and farthest from the voice collection device for collecting the target signal mentioned above will be automatically acquired to collect real-time spatial sound and used as the verified background sound.

[0021] Furthermore, the specific method of determining the background sound when generating the synchronous filtering signal is as follows:

[0022] The real-time spatial sound will be automatically collected by the voice collection device that is in the same environment and farthest from the voice collection device for collecting the target signal, and marked as spatial sound;

[0023] The user's voice in the spatial sound is then suppressed to a minimum or eliminated, and the remaining sound is used as the approved background sound.

[0024] Furthermore, the user's voice is recognized through an intelligent recognition model, which is a convolutional neural network (CNN) model.

[0025] Furthermore, the specific training method of the intelligent recognition model is:

[0026] Acquire a number of non-interfering user voice data from the corresponding environment as training data, and then acquire the voices of other users and a set number of user voices mixed with any voice as verification data;

[0027] Input a number of training data into the model for training. After the training is completed, input the verification data into the model to verify its accuracy. When the accuracy exceeds the set value B1, it means that the model training is completed. Otherwise, reacquire a number of training data and adjust the parameters in the model until the accuracy exceeds the set value B1.

[0028] Complete the training of the intelligent recognition model.

[0029] Compared with the prior art, the present invention has the following beneficial effects:

[0030] The present invention continuously acquires the background sound of the environment, retains the background sound of one period before the target signal appears, and then extracts the verified background sound from the background sound or uses the sound collected in the space where the user is when recording real-time voice as the verified background sound; and then performs audio data processing, which can effectively process the audio data and obtain good pure audio data, thereby improving the user's voice interaction experience; the present invention is simple, effective and easy to use. BRIEF DESCRIPTION OF THE DRAWINGS

[0031] Figure 1 It is a system block diagram of the present invention. DETAILED DESCRIPTION

[0032] The technical solution of the present invention will be clearly and completely described below in conjunction with the embodiments. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without creative work are within the scope of protection of the present invention.

[0033] See also Figure 1 , the present application provides a voice data processing method based on an audio processor;

[0034] As an embodiment 1 of the present application, the method specifically includes the following steps:

[0035] Step 1: Before processing the data, analyze the background sound without human voice, and extract several segments of ambient sound from the background sound. The specific method is as follows:

[0036] First, set the audio acquisition device to always acquire the sound of the environment and mark it as background sound;

[0037] The background sound is continuously acquired and then set as a cycle according to the T1 time; T1 is a preset value, which is pre-set by the administrator;

[0038] When the duration of the background sound obtained exceeds two cycles, if the target signal is not detected at the latest moment, all the voice data before one cycle will be deleted; if the background sound is continuously obtained, if the duration of the subsequent background sound obtained exceeds two cycles, if the target signal is not detected, all the background sound before the closest cycle will be automatically deleted until the target signal is detected;

[0039] The target signal is generally the wake-up voice of the corresponding device. Of course, it can also be other wake-up signals, indicating that the user needs to perform voice interaction or other voice input actions;

[0040] After detecting the target signal, the background sound of a period before the target signal is generated will be automatically divided into several segments of ambient sound with the same duration; the specific number of segments is set by the administrator and is guaranteed to be no less than ten segments;

[0041] Step 2: Perform similarity analysis on the obtained several environmental sounds. The specific method of similarity analysis is as follows:

[0042] First, all ambient sounds and their corresponding waveforms are obtained, where the waveform is a waveform of the sound; then, one ambient sound is selected, marked as the selected ambient sound, and its waveform is obtained; then, similarities between other ambient sounds and the waveform of the selected ambient sound are obtained, and several similarities D i, i=1, ..., n are obtained;

[0043] Then get the mean value of Di, mark it as P, and use the formula to calculate the stable value W of all Di. The specific calculation formula is:

[0044]

[0045] Then, the next ambient sound is selected in sequence, marked as the selected ambient sound, and the mean and stable values ​​of the similarities of the waveforms of the other ambient sounds are calculated; and the process is repeated to obtain the corresponding mean and stable values ​​when all ambient sounds are marked as the selected ambient sounds;

[0046] All selected ambient sounds whose average value exceeds the set value X1 and whose stable value is lower than the set value X2 are screened out, and then the corresponding confirmed values ​​of all the screened selected ambient sounds are calculated according to the formula. The specific formula is:

[0047] Selected value = 0.63*mean + 0.37 / stable value;

[0048] Mark the selected ambient sound with the highest confirmation value as the approved background sound;

[0049] If any of the selected ambient sounds cannot be filtered out, a synchronous filtering signal will be generated at this time;

[0050] Step 3: After the background sound is verified, the real-time voice recorded by the user will be automatically obtained after the user enters the target signal;

[0051] Then, the approved background sound in the real-time speech is eliminated to obtain processed data, which is marked as noise-reduced speech data;

[0052] Complete basic processing of voice data;

[0053] Step 4: If a synchronous filtering signal is generated, the voice collection device that is in the same environment and farthest from the voice collection device that collects the target signal will be automatically acquired to collect real-time spatial sound, which will be used as the background sound for verification;

[0054] Then, the approved background sound in the real-time speech is eliminated to obtain processed data, which is marked as noise-reduced speech data;

[0055] Complete basic processing of voice data;

[0056] Step 5: After completing the processing of the voice data, continue to monitor whether the user is still recording voice. If no new voice signal from the user is detected within T1 time, that is, within one cycle time, jump to step 1 and reprocess;

[0057] As the second embodiment of the present application, this embodiment is implemented on the basis of the first embodiment, and the difference is that, in this embodiment, the processing method when generating the synchronous filtering signal is different. The specific method in this embodiment is:

[0058] If a synchronous filtering signal is generated, the voice collection device that is in the same environment and farthest from the voice collection device that collects the target signal will be automatically acquired to collect the real-time spatial sound and mark it as spatial sound;

[0059] Then, the user voice in the spatial sound is suppressed to the minimum or eliminated, and the remaining is used as the approved background sound. Then, the approved background sound in the real-time voice is eliminated to obtain the processed data, which is marked as noise-reduced voice data;

[0060] Complete basic processing of voice data;

[0061] The user's voice is identified and determined by the intelligent recognition model, which is determined in the following way:

[0062] The convolutional neural network (CNN) is selected as the basic model. Several user voices without interference are obtained from the corresponding environment as training data. Then the voices of other users and a set number of user voices mixed with arbitrary voices are obtained as verification data.

[0063] First, the training data is recognized. The original audio signal is first converted into a spectrogram or Mel-frequency cepstral coefficient MFCC so as to serve as the input of CNN.

[0064] Enhance the audio data, such as adding noise, changing the volume, etc., to improve the generalization ability of the model;

[0065] The user selects a CNN architecture suitable for audio processing, including convolutional layers, pooling layers, and fully connected layers. Use pre-trained CNN models, such as VGG19, to improve the accuracy of sound recognition through transfer learning. Use the convolutional layers of CNN to automatically extract local features in audio signals, such as changes in frequency and time. Through pooling layers and fully connected layers, CNN can integrate local features to form a global feature representation. Select cross-entropy loss to guide the CNN training process. Use gradient descent or other optimization algorithms to adjust the weights of CNN to minimize the loss function.

[0066] Then, several training data are input into the model for training. After the training is completed, the verification data is input into the model to verify its accuracy. When the accuracy exceeds the set value B1, it means that the model training is completed. Otherwise, several training data are obtained again and the parameters in the model are adjusted until the accuracy exceeds the set value B1.

[0067] Complete the training of intelligent recognition model;

[0068] Then, an intelligent model is used to extract the user’s voice from the spatial sound.

[0069] Some of the data in the above formula are calculated by removing the dimensions and taking their numerical values. The formula is a formula that is closest to the actual situation obtained by software simulation of a large amount of collected data; the preset parameters and preset thresholds in the formula are set by technical personnel in this field according to actual conditions or obtained through simulation of a large amount of data.

[0070] The above embodiments are only used to illustrate the technical method of the present invention rather than to limit it. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical method of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical method of the present invention.

Claims

1. A voice data processing method based on an audio processor, characterized in that: The method specifically comprises the following steps: Continuously obtain the background sound of the environment, and retain the background sound of one period before the target signal appears from the background sound. The target signal is used to indicate that the user has a need for voice recording; The background sound is divided into several segments of ambient sound of equal duration, and the discrete degree and mean value of the data of the similarity between the waveform of each ambient sound and other waveforms are used to determine the background sound or generate a synchronous filtering signal; When generating the synchronous filtering signal, the sound collected from the space where the user is recording the real-time voice is used as the verified background sound; Then, the part of the real-time voice recorded by the user that belongs to the approved background sound is deleted to obtain processed data, which is marked as noise reduction voice data.

2. The method for processing speech data based on an audio processor according to claim 1, characterized in that: After obtaining the noise reduction voice data, if the user does not record a new voice signal within a period of time, the generation of the target signal is re-detected.

3. The method for processing speech data based on an audio processor according to claim 1, characterized in that: The specific duration of a cycle is set by the administrator.

4. The method for processing speech data based on an audio processor according to claim 1, characterized in that: The method of determining the background sound or generating a synchronous filtering signal based on several segments of ambient sound is as follows: Get all the ambient sounds and their corresponding waveforms; One is selected from all the ambient sounds in turn and marked as the selected ambient sound. For each selected ambient sound, the similarity between the waveforms of other ambient sounds and the waveform of the selected ambient sound is calculated to obtain several similarities. The mean value of the similarities and the stable value representing the degree of dispersion of the similarities are first calculated. Then, in a manner proportional to the mean and inversely proportional to the stable value, the mean and stable value are multiplied by a corresponding weight value and then added together, and the resulting value is marked as the selected value; The confirmation values ​​of all the selected ambient sounds are obtained, and the selected ambient sound with the highest confirmation value is marked as the approved background sound.

5. The method for processing speech data based on an audio processor according to claim 4, characterized in that: Before marking the selected ambient sound with the highest confirmation value as the approved background sound, the selected ambient sound needs to be screened. The specific screening method is as follows: Filter out all selected ambient sounds whose average value exceeds the set value X1 and whose stable value is lower than the set value X2.

6. The method for processing speech data based on an audio processor according to claim 5, characterized in that: When any selected ambient sound cannot be filtered out, a synchronous filtering signal is generated.

7. The method for processing speech data based on an audio processor according to claim 1, characterized in that: The specific method of determining the background sound when generating the synchronous filtering signal is as follows: The voice collection device that is in the same environment and farthest from the voice collection device for collecting the target signal mentioned above will be automatically acquired to collect real-time spatial sound and used as the verified background sound.

8. The method for processing speech data based on an audio processor according to claim 1, characterized in that: The specific method of determining the background sound when generating the synchronous filtering signal is as follows: The real-time spatial sound will be automatically collected by the voice collection device that is in the same environment and farthest from the voice collection device for collecting the target signal, and marked as spatial sound; The user's voice in the spatial sound is then suppressed to a minimum or eliminated, and the remaining sound is used as the approved background sound.

9. The method for processing speech data based on an audio processor according to claim 8, characterized in that: The user's voice is recognized through an intelligent recognition model, which is a convolutional neural network (CNN) model.

10. The method for processing speech data based on an audio processor according to claim 1, characterized in that: The specific training method of the intelligent recognition model is: Acquire a number of non-interfering user voice data from the corresponding environment as training data, and then acquire the voices of other users and a set number of user voices mixed with any voice as verification data; Input a number of training data into the model for training. After the training is completed, input the verification data into the model to verify its accuracy. When the accuracy exceeds the set value B1, it means that the model training is completed. Otherwise, reacquire a number of training data and adjust the parameters in the model until the accuracy exceeds the set value B1. Complete the training of the intelligent recognition model.

Citation Information

Patent Citations

  • Voice data processing method and device

    CN105869645A

  • Voice noise reduction method and system for interaction, electronic equipment and storage medium

    CN115376538A

  • Outdoor noise and dust environment monitoring method and system

    CN116958133A

  • Apparatus and Method For Feature Compensation Using Weighted Auto-Regressive Moving Average Filter and Global Cepstral Mean and Variance Normalization

    KR1020120077527A

  • Speech and Noise Models for Speech Recognition

    US20110307253A1