Speech Recognition Method, Apparatus, Computer Device, Storage Medium, and Program Product
By segmenting and feature processing of the original speech data, combined with modal decomposition technology, the problem of inaccurate speech recognition in the prior art is solved, and the accuracy and security of speech recognition are improved.
Patent Information
- Application Number
- CN202210194783.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-03-01
- Publication Date
- 2025-06-03
- Estimated Expiration
- 2042-03-01
AI Technical Summary
There are false speech such as virtual voice and imitation speech in existing voiceprint recognition applications, resulting in reduced accuracy and security of speech recognition. How to accurately identify synthesized speech or convert speech has become an urgent problem.
By performing segmented processing and feature processing on the original speech data, candidate speech data are obtained, and then modal decomposition is performed to extract modal components with high independence, convert them into speech data to be recognized, and input them into the preset speech recognition model for recognition and verification.
It improves the accuracy and detection accuracy of speech recognition, enhances the recognition ability of false speech, and thus improves the security of the authentication system of voiceprint recognition function.
Smart Images

Figure CN114664313B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, and in particular, to a voice recognition method, apparatus, computer device, storage medium, and computer program product. Background Art
[0002] In recent years, with the increasing maturity of voiceprint recognition technology, authentication systems with voiceprint recognition functions have been widely used in various scenarios, such as voice assistants, online banking, etc.
[0003] However, with the development of robotics and virtual technologies, there are many deceptive and aggressive voices in existing applications, such as synthetic voices, converted voices, etc., such as synthesized voices, converted voices, etc., which reduce the accuracy of voice recognition, and further reduce the security of the authentication system using the voiceprint recognition function.
[0004] Based on this, how to accurately identify synthetic voices or converted voices has become an urgent technical problem in current voiceprint recognition applications. Summary of the Invention
[0005] Based on this, in view of the above technical problems, it is necessary to provide a voice recognition method, apparatus, computer device, storage medium, and computer program product that can accurately identify false voices such as synthetic voices or converted voices.
[0006] In a first aspect, this application provides a voice recognition method, which includes:
[0007] Perform segmentation processing on the original voice data to obtain a plurality of sub-voice data, and perform feature processing on each sub-voice data to obtain candidate voice data;
[0008] Perform modal decomposition processing on the candidate voice data to obtain the voice data to be recognized;
[0009] Input the voice data to be recognized into a preset voice recognition model to obtain a voice recognition result; the voice recognition result is used to indicate whether the voice data to be recognized is real voice data.
[0010] In one optional embodiment, performing feature processing on each sub-voice data to obtain candidate voice data includes:
[0011] Perform voice signal feature extraction processing on each sub-voice data to determine the feature degree of each sub-voice data;
[0012] Perform feature synthesis processing on the sub-voice data with a feature degree greater than a preset threshold to obtain candidate voice data.
[0013] In one optional embodiment, the method further includes:
[0014] Denoise each sub-audio data to obtain the denoised sub-audio data;
[0015] Perform feature processing on each sub-audio data to obtain candidate audio data, including:
[0016] Perform feature processing on each denoised sub-audio data to obtain candidate audio data.
[0017] In one optional embodiment, perform modal decomposition processing on the candidate audio data to obtain the audio data to be recognized, including:
[0018] Perform modal decomposition processing on the candidate audio data to obtain multiple modal components;
[0019] Determine the modal components that meet the preset independence requirement as candidate modal components;
[0020] Convert the candidate modal components into the audio data to be recognized.
[0021] In one optional embodiment, the method further includes:
[0022] Perform speech segmentation on the original audio data to obtain the silent segments and the voiced segments in the original audio data;
[0023] Merge the voiced segments to obtain the merged audio data;
[0024] Perform segmentation processing on the original audio data to obtain multiple sub-audio data, including:
[0025] Perform segmentation processing on the merged audio data to obtain multiple sub-audio data.
[0026] In one optional embodiment, the method further includes:
[0027] Perform speech segmentation on the candidate audio data to obtain the silent segments and the voiced segments in the candidate audio data;
[0028] Merge the voiced segments to obtain the merged audio data;
[0029] Perform modal decomposition processing on the candidate audio data to obtain the audio data to be recognized, including:
[0030] Perform modal decomposition processing on the merged audio data to obtain the audio data to be recognized.
[0031] In one optional embodiment, the method further includes:
[0032] Perform speech segmentation on the audio data to be recognized to obtain the silent segments and the voiced segments in the audio data to be recognized;
[0033] Merge the voice segments to obtain merged voice data;
[0034] Input the voice data to be recognized into a preset voice recognition model to obtain a voice recognition result, including:
[0035] Input the merged voice data into a preset voice recognition model to obtain a voice recognition result.
[0036] In one optional embodiment, input the voice data to be recognized into a preset voice recognition model to obtain a voice recognition result, including:
[0037] Input the voice data to be recognized into a preset voice recognition model to obtain the credibility of the voice data to be recognized;
[0038] If the credibility is greater than or equal to a preset credibility threshold, determine that the voice recognition result of the voice data to be recognized is real voice data;
[0039] If the credibility is less than the preset credibility threshold, determine that the voice recognition result of the voice data to be recognized is false voice data.
[0040] In one optional embodiment, the method further includes:
[0041] Output an alarm message when the voice recognition result of the voice data to be recognized is false voice data.
[0042] In a second aspect, a voice recognition device is provided. The device includes:
[0043] A first processing module, configured to segment the original voice data to obtain multiple sub-voice data, and perform feature processing on each sub-voice data to obtain candidate voice data;
[0044] A second processing module, configured to perform modal decomposition processing on the candidate voice data to obtain the voice data to be recognized;
[0045] A recognition module, configured to input the voice data to be recognized into a preset voice recognition model to obtain a voice recognition result; the voice recognition result is used to indicate whether the voice data to be recognized is real voice data.
[0046] In a third aspect, the present application further provides a computer device. The computer device includes a memory and a processor. The memory stores a computer program, and when the processor executes the computer program, the method provided in the first aspect is implemented.
[0047] Fourthly, the present application also provides a computer-readable storage medium. The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method provided in the first aspect is implemented.
[0048] Fifthly, the present application also provides a computer program product. The computer program product includes a computer program, and when the computer program is executed by a processor, the method provided in the first aspect is implemented.
[0049] For the above voice recognition method, device, computer device, storage medium and computer program product, the computer device performs segmentation processing on the original voice data to obtain multiple sub-voice data, performs feature processing on each sub-voice data to obtain candidate voice data, performs modal decomposition processing on the candidate voice data to obtain the voice data to be recognized, and inputs the voice data to be recognized into a preset voice recognition model to obtain a voice recognition result; the voice recognition result is used to indicate whether the voice data to be recognized is real voice data. In this solution, the computer device performs segmentation processing on the original voice data to refine the processing granularity of the original voice data. Based on the sub-voice data obtained by segmentation for feature processing, the accuracy of feature processing can be improved, thereby obtaining relatively accurate candidate voice data. On this basis, modal decomposition processing is performed on the candidate voice data. Modal decomposition processing can extract data with relatively high modal component independence from the candidate voice data, thereby further improving the accuracy of the obtained voice data to be recognized. Based on the two voice data processing, when the voice data to be recognized is input into the voice recognition model, the obtained voice recognition result is more accurate, and the detection accuracy of real voice data recognition is improved. Description of the Drawings
[0050] Figure 1 It is an application environment diagram of the voice recognition method in an embodiment;
[0051] Figure 2 It is a schematic flowchart of the voice recognition method in an embodiment;
[0052] Figure 3 It is a schematic flowchart of the voice recognition method in another embodiment;
[0053] Figure 4 It is a schematic flowchart of the voice recognition method in another embodiment;
[0054] Figure 5 It is a schematic flowchart of the voice recognition method in another embodiment;
[0055] Figure 6 It is a schematic flowchart of the voice recognition method in another embodiment;
[0056] Figure 7Schematic flowchart of a voice recognition method in another embodiment;
[0057] Figure 8 Schematic flowchart of a voice recognition method in another embodiment;
[0058] Figure 9 Schematic flowchart of a voice recognition method in another embodiment;
[0059] Figure 10 Block diagram of the structure of a voice recognition device in one embodiment;
[0060] Figure 11 Block diagram of the structure of a voice recognition device in another embodiment;
[0061] Figure 12 Block diagram of the structure of a voice recognition device in another embodiment. Detailed implementation manners
[0062] In order to make the objectives, technical solutions and advantages of the present application clearer and more understandable, the present application will be further described in detail below with reference to the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are only used to explain the present application and are not used to limit the present application.
[0063] The voice recognition method provided by the embodiments of the present application can be applied to an application environment as shown in Figure 1 In one embodiment, a computer device is provided. The computer device can be a terminal or a server, and its internal structure diagram can be as shown in Figure 1 The computer device includes a processor, a memory, a communication interface, a display screen, and an input device connected through a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and the computer program in the non-volatile storage medium. The communication interface of the computer device is used to communicate with an external terminal in a wired or wireless manner, and the wireless manner can be implemented through WIFI, a mobile cellular network, NFC (Near Field Communication) or other technologies. When the computer program is executed by the processor, a voice recognition method is implemented. The display screen of the computer device can be a liquid crystal display screen or an electronic ink display screen, and the input device of the computer device can be a touch layer covered on the display screen, or a button, a trackball or a touchpad provided on the housing of the computer device, or an external keyboard, touchpad or mouse, etc.
[0064] Those skilled in the art can understand that Figure 1The structure shown is only a block diagram of some structures related to the solution of this application, and does not constitute a limitation on the computer device to which the solution of this application is applied. The specific computer device may include more or fewer components than those shown in the figure, or combine some components, or have a different component layout.
[0065] In one embodiment, as Figure 2 shown, a speech recognition method is provided. Taking the computer device in Figure 1 as an example, the method includes the following steps:
[0066] Step 201: Segment the original speech data to obtain multiple sub-speech data, and perform feature processing on each sub-speech data to obtain candidate speech data.
[0067] Among them, the original speech data can be data received from a terminal or a third party, or speech data input by a user, or speech data obtained from a database. After the computer device obtains the original speech data, considering that the original speech data is large in quantity, the original speech data can be segmented. Exemplarily, the computer device can segment the original speech data according to the number of frames. For example, every 10 frames form a segment to form multiple sub-speech data; or the computer device can also segment the original speech data according to the duration. For example, every 30s forms a segment to form multiple sub-speech data; or the computer device can also identify the speech pause part in the original speech data and perform segment cutting at the pause part to form multiple sub-speech data. The purpose of segmenting the original speech data is to refine the processing granularity of the data to make the processing result more accurate. In this embodiment, no limitation is imposed on the principle based on which the computer device segments the original speech data.
[0068] After the computer device obtains multiple sub-speech data, the computer device performs feature processing on each sub-speech data respectively. Optionally, the feature processing includes feature extraction, feature synthesis, and other processing. The computer device performs feature extraction processing on each sub-speech data. Among them, there are various feature extraction methods. Exemplarily, the computer device can input each sub-speech data into a feature extraction model to obtain the speech features corresponding to each sub-speech data; or, the computer device can also perform wavelet decomposition on each sub-speech data. Exemplarily, the computer device performs wavelet decomposition processing on each frame of signal under the small basis wave db1 / db2 / db3 to obtain the wavelet decomposition coefficients corresponding to each sub-speech data. Here, the wavelet decomposition coefficients can be used to represent the speech features of each sub-speech data. After obtaining the speech features corresponding to each sub-speech data, the computer device can also perform feature screening on these speech features, so as to determine the speech features that meet the feature conditions for synthesis to obtain candidate speech data. Exemplarily, the feature conditions can be that the speech features match the calibrated features; or, the wavelet decomposition coefficients are greater than a preset threshold, etc. This embodiment does not make any limitations in this regard.
[0069] Step 202: Perform modal decomposition processing on the candidate speech data to obtain the speech data to be recognized.
[0070] In this embodiment, since the candidate speech data may be fake speech data obtained through synthesis or conversion, and this fake speech data is a non-linear time series signal. For this type of data, the computer device can perform modal decomposition processing on the candidate speech data after the feature processing in the above step 201, analyze the signal modal components in the candidate speech data, and determine the speech data to be recognized. Exemplarily, the computer device can perform modal decomposition processing on the candidate speech data through an empirical mode decomposition algorithm to obtain multiple signal modal components. Or, the computer device can also perform modal decomposition processing on the candidate speech data through a dynamic mode decomposition algorithm to obtain multiple signal modal components. This embodiment does not make any limitations on the modal decomposition method.
[0071] Considering that the candidate speech data includes a mixed signal source, the computer device can also perform independent component analysis on the obtained multiple signal modal components. Exemplarily, the computer device can determine the modal components with stronger independence in the signal modal components through an independent component analysis algorithm (Independent Component Analysis, ICA), so as to perform signal conversion on the modal components with stronger independence to obtain the speech data to be recognized.
[0072] Step 203: Input the speech data to be recognized into a preset speech recognition model to obtain a speech recognition result; the speech recognition result is used to indicate whether the speech data to be recognized is real speech data.
[0073] Among them, the preset speech recognition model is a pre-trained speech recognition model. Optionally, the samples for training the model can be a collection of false speech data samples. Taking the real speech data samples as a reference, according to the similarity between the false speech data samples and the real speech data samples, the recognition results of the false speech data samples are determined. Based on the recognition results, the parameters of the speech recognition model are adjusted until the preset number of iterations is reached, and the trained speech recognition model is obtained.
[0074] Based on the trained speech recognition model, the computer device inputs the speech data to be recognized into the speech recognition model to obtain the speech recognition result corresponding to the speech data to be recognized. Exemplarily, when the computer device inputs the speech data to be recognized into the speech recognition model, the credibility of the speech data to be recognized can be obtained. The credibility is used to represent the probability that the speech data to be recognized is real speech data. For example, if the credibility is 80%, it means that the probability that the speech data to be recognized is real speech data is 80%. If a credibility threshold of 70% is set and the credibility of the speech data to be recognized is greater than the credibility threshold, the recognition result of the speech data to be recognized is determined to be real speech data. Optionally, when the computer device inputs the speech data to be recognized into the speech recognition model, multiple feature credibilities of the speech data to be recognized can be obtained. The computer device can determine the recognition result of the speech data to be recognized based on the multiple feature credibilities. For example, if the weight distribution of the multiple feature credibilities conforms to the preset distribution rule, the recognition result of the speech data to be recognized is determined to be real speech data. This embodiment does not make any limitations in this regard.
[0075] In the above speech recognition method, the computer device performs segmentation processing on the original speech data to obtain multiple sub-speech data, performs feature processing on each sub-speech data to obtain candidate speech data, performs modal decomposition processing on the candidate speech data to obtain the speech data to be recognized, inputs the speech data to be recognized into the preset speech recognition model to obtain the speech recognition result; the speech recognition result is used to indicate whether the speech data to be recognized is real speech data. In this solution, the computer device performs segmentation processing on the original speech data to refine the processing granularity of the original speech data. Based on the sub-speech data obtained by segmentation for feature processing, the accuracy of feature processing can be improved, thereby obtaining relatively accurate candidate speech data. On this basis, modal decomposition processing is performed on the candidate speech data. Modal decomposition processing can extract data with higher modal component independence from the candidate speech data, thereby further improving the accuracy of the obtained speech data to be recognized. Based on the two speech data processing, when the speech data to be recognized is input into the speech recognition model, the obtained speech recognition result is more accurate, and the detection accuracy of real speech data recognition is improved.
[0076] The computer device performs feature processing on each sub-audio data after the processing granularity is refined, and the obtained candidate audio data is data with a higher degree of effective audio data. In one optional embodiment, as Figure 3 shown, the above step 201 performs feature processing on each sub-audio data to obtain candidate audio data, including:
[0077] Step 301, performing speech signal feature extraction processing on each sub-audio data to determine the feature degree of each sub-audio data.
[0078] In this embodiment, the computer device can perform speech signal feature extraction processing on each sub-audio data based on a wavelet decomposition function. Exemplarily, the computer device sequentially performs wavelet decomposition processing on each sub-audio data under the small fundamental wave db3 to obtain the wavelet decomposition coefficients of each sub-audio data. The wavelet decomposition coefficients are used to indicate the feature degree of the speech features of each sub-audio data, and this feature degree can also be understood as the weight of the speech features.
[0079] Step 302, performing feature synthesis processing on the sub-audio data with a feature degree greater than a preset threshold to obtain candidate audio data.
[0080] Among them, the speech data segments that obtain the wavelet decomposition coefficients of the noisy speech signal are synthesized into a complete speech data information chain to improve the efficiency of wavelet reconstruction of the wavelet decomposition coefficients of the noisy speech signal.
[0081] In this embodiment, after the computer device obtains the feature degrees of each sub-audio data, it screens the sub-audio data with a feature degree greater than a preset threshold, that is, based on the wavelet decomposition coefficients of each sub-audio data, it determines the sub-audio data with a wavelet decomposition coefficient greater than the preset threshold, performs wavelet reconstruction on these sub-audio data, and synthesizes a speech data information chain as candidate audio data. Taking the sub-audio data with a feature degree greater than the preset threshold means that the weight of its speech features is higher and more accurate.
[0082] In this embodiment, the computer device processes the sub-audio data through feature extraction and feature synthesis processing, and there is more valid data in the obtained candidate audio data, and the candidate audio data is more accurate.
[0083] To further improve the processing accuracy of the processed data, in one optional embodiment, the method further includes:
[0084] Performing denoising processing on each sub-audio data to obtain the denoised sub-audio data.
[0085] In this embodiment, the computer device can perform denoising on each sub-voice data. Exemplarily, the denoising algorithm includes a least mean square (LMS) adaptive filtering algorithm, a recursive least squares RLS adaptive filtering algorithm, a transform domain adaptive filtering algorithm, an affine projection algorithm, a conjugate gradient algorithm, and an adaptive filtering algorithm based on subband decomposition. The computer device can perform denoising on each sub-voice data according to any denoising algorithm to obtain denoised sub-voice data. Alternatively, the computer device can also determine the voice signal threshold T according to the Neyman-Pearson criterion, determine the signal whose signal value in each sub-voice data is greater than the threshold T as noise and remove it, and obtain the denoised sub-voice data. Exemplarily, the signal value can be a signal frequency or a signal amplitude, and correspondingly, the voice signal threshold T can be a frequency T value or an amplitude T value, which is not limited in this embodiment.
[0086] Then, feature processing is performed on each sub-speech data to obtain candidate speech data, which includes:
[0087] Feature processing is performed on each denoised sub-speech data to obtain candidate speech data.
[0088] In this embodiment, after obtaining the denoised sub-speech data, the computer device performs the feature processing of steps 301-302 to obtain candidate speech data, which will not be described in detail in this embodiment.
[0089] Optionally, the computer device may perform denoising on the candidate voice data after obtaining the candidate voice data; or, the computer device may first perform denoising on the original voice data, and then perform segmentation processing on the denoised voice data. This embodiment does not limit the execution order and number of executions of the denoising process. Ideally, the more denoising times are performed, the more accurate the voice data obtained is.
[0090] The computer device performs secondary data processing, that is, modal decomposition processing on the candidate voice data, with the purpose of analyzing independent signals in the voice signal. In one optional embodiment, as Figure 4 As shown, the candidate speech data is subjected to modal decomposition processing to obtain speech data to be recognized, including:
[0091] Step 401, performing modal decomposition processing on the candidate speech data to obtain multiple modal components.
[0092] Since the candidate speech data is a non-linear time series signal, for this type of data, the computer device can perform modal decomposition processing on the candidate speech data, analyze the signal modal components at each layer in the candidate speech data, and determine the speech data to be recognized. Exemplarily, the computer device can perform modal decomposition processing on the candidate speech data through an empirical mode decomposition algorithm to obtain multiple signal modal components.
[0093] Step 402, determine the modal components that meet the preset independence requirement as candidate modal components.
[0094] In this embodiment, considering that the candidate speech data includes a mixed signal source, the computer device can also perform independent component analysis on the obtained multiple signal modal components. Exemplarily, the computer device can use the autocorrelation criterion as a benchmark and determine the modal components in the signal modal components whose independence meets the preset independence requirement through the independent component analysis algorithm ICA. Exemplarily, the computer device can analyze according to the peak value, envelope spectrum, and kurtosis of the signal modal components, and determine that the signal modal components with independence greater than the threshold are candidate modal components.
[0095] Step 403, convert the candidate modal components into the speech data to be recognized.
[0096] In this embodiment, after the computer device obtains the independence analysis result of the modal components based on the ICA algorithm and determines the candidate modal components, the computer device can use the FastICA algorithm to perform signal conversion on the selected candidate modal components and convert them into the speech data to be recognized.
[0097] In this embodiment, since there is a non-linear relationship between the noise-free speech and the noisy speech in the candidate speech data, performing modal decomposition processing on the candidate speech data makes the accuracy higher in subsequent speech processing, realizes the separation of the independent source signals with independence from the candidate speech data information, facilitates the peak analysis, envelope spectrum, and kurtosis analysis of the separation result, and the obtained speech data to be recognized has high signal quality, effectively improving the accuracy of synthetic speech detection.
[0098] To further improve the processing efficiency of the speech data, the original speech data can also be filtered for silent segments. In one optional embodiment, as Figure 5 shown, the method further includes:
[0099] Step 501, perform speech segmentation on the original speech data to obtain the silent segments and the voiced segments in the original speech data.
[0100] In this embodiment, after obtaining the original voice data, the computer device can detect the original voice data to determine the voiced segments and unvoiced segments in the original voice data. Optionally, since there is no voice signal in the unvoiced segments, the computer device can determine the unvoiced segments in the original voice data by calculating the zero-crossing rate of the original voice data, so as to obtain the unvoiced segments and voiced segments in the original voice data. Alternatively, the computer device can also distinguish the unvoiced segments through spectrograms, audio signals, etc., so as to obtain the unvoiced segments and voiced segments in the original voice data.
[0101] Step 502: Merge the voiced segments to obtain merged voice data.
[0102] In this embodiment, after obtaining the voiced segments in the original voice data, the computer device merges the voiced segments in a certain order to obtain merged voice data. Here, the certain order can be the default order or the time frame order, and this embodiment does not make any limitations in this regard.
[0103] Step 503: Segment the merged voice data to obtain multiple sub-voice data.
[0104] In this embodiment, after obtaining the merged voice data, the computer device segments the merged voice data. Optionally, the segmentation process is similar to that in step 201 and will not be elaborated here.
[0105] In this embodiment, the computer device eliminates the unvoiced segments from the original voice data, reducing the data processing volume of subsequent voice data segmentation.
[0106] Alternatively, the filtering operation of the unvoiced segments can also be performed after obtaining the candidate voice data. In one optional embodiment, as Figure 6 shown, the method further includes:
[0107] Step 601: Segment the candidate voice data to obtain the unvoiced segments and voiced segments in the candidate voice data.
[0108] In this embodiment, after obtaining the candidate voice data, the computer device can detect the candidate voice data to determine the voiced segments and unvoiced segments in the candidate voice data. Optionally, since there is no voice signal in the unvoiced segments, the computer device can determine the unvoiced segments in the candidate voice data by calculating the zero-crossing rate of the candidate voice data, so as to obtain the unvoiced segments and voiced segments in the candidate voice data. Alternatively, the computer device can also distinguish the unvoiced segments through spectrograms, audio signals, etc., so as to obtain the unvoiced segments and voiced segments in the candidate voice data.
[0109] Step 602: Merge the voiced segments to obtain merged voice data.
[0110] In this embodiment, similar to step 502, details are not described herein.
[0111] Step 603: Perform modal decomposition processing on the merged voice data to obtain the voice data to be recognized.
[0112] In this embodiment, after the computer device obtains the merged voice data, it performs modal decomposition processing on the merged voice data. Optionally, the segmentation processing is similar to that in step 202, and details are not described herein.
[0113] In this embodiment, the computer device removes the silent segments from the candidate voice data, reducing the data processing volume of subsequent modal decomposition processing.
[0114] Alternatively, the filtering operation of the silent segments can also be performed after obtaining the voice data to be recognized. In one optional embodiment, as Figure 7 shown, the method further includes:
[0115] Step 701: Perform voice segmentation on the voice data to be recognized to obtain the silent segments and the voiced segments in the voice data to be recognized.
[0116] In this embodiment, after the computer device obtains the voice data to be recognized, it can detect the voice data to be recognized to determine the voiced segments and the silent segments in the voice data to be recognized. Optionally, since there is no voice signal in the silent segments, the computer device can determine the silent segments in the voice data to be recognized by calculating the zero-crossing rate of the voice data to be recognized, so as to obtain the silent segments and the voiced segments in the voice data to be recognized. Alternatively, the computer device can also distinguish the silent segments through spectrograms, audio signals, etc., so as to obtain the silent segments and the voiced segments in the voice data to be recognized.
[0117] Step 702: Merge the voiced segments to obtain the merged voice data.
[0118] In this embodiment, similar to step 502, details are not described herein.
[0119] Step 703: Input the merged voice data into a preset voice recognition model to obtain a voice recognition result.
[0120] In this embodiment, after the computer device obtains the merged voice data, it inputs the merged voice data into a preset voice recognition model. Optionally, the segmentation processing is similar to that in step 203, and details are not described herein.
[0121] In this embodiment, the computer device removes the silent segments from the voice data to be recognized, reducing the data processing volume of subsequent voice recognition.
[0122] After obtaining the speech data to be recognized, the computer device determines whether the speech data to be recognized is real speech data or fake speech data based on a preset speech recognition model. In one optional embodiment, as Figure 8 shown, the speech data to be recognized is input into the preset speech recognition model to obtain a speech recognition result, including:
[0123] Step 801: Input the speech data to be recognized into the preset speech recognition model to obtain the credibility of the speech data to be recognized.
[0124] Among them, the preset speech recognition model is the model obtained through training described in step 201, and this speech recognition model has a high recognition accuracy for synthesized speech. In this embodiment, the computer device uses the speech data to be recognized as the input data of the preset speech recognition model to obtain the credibility of the speech data to be recognized. This credibility can be a percentage, representing the probability value that the speech data to be recognized is real speech data. For example, the credibility of the speech data to be recognized is 80%.
[0125] Step 802: If the credibility is greater than or equal to the preset credibility threshold, determine that the speech recognition result is that the speech data to be recognized is real speech data.
[0126] In this embodiment, the preset credibility threshold is a value set according to the actual situation. Exemplarily, this credibility threshold can be 80%. If the credibility of the speech data to be recognized is greater than or equal to this credibility threshold, taking the example in step 801 to illustrate, the credibility of the speech data to be recognized is 80%, equal to the preset credibility threshold, then determine that the speech data to be recognized is real speech data.
[0127] Step 803: If the credibility is less than the preset credibility threshold, determine that the speech recognition result is that the speech data to be recognized is fake speech data.
[0128] In this embodiment, if the credibility of the speech data to be recognized is less than this credibility threshold, exemplarily, this credibility threshold can be 80%, and the credibility of the speech data to be recognized is 40%, less than the preset credibility threshold, then determine that the speech data to be recognized is fake speech data, that is, the speech data to be recognized is more likely to be synthesized speech data or converted speech data, and the speech data to be recognized is less likely to be real speech data.
[0129] In this embodiment, the computer device determines the speech recognition result of the speech data to be recognized based on the trained speech recognition model, thereby determining whether the speech data to be recognized is fake speech data, so as to achieve accurate recognition of fake speech data.
[0130] In the case where the voice data to be recognized is determined to be false voice data, in one optional embodiment, the method further includes:
[0131] When the voice recognition result is that the voice data to be recognized is false voice data, an alarm message is output.
[0132] In this embodiment, in order to indicate to the user or staff that the voice data to be recognized is false voice data, when the computer device determines that the voice data to be recognized is false voice data, it can output an alarm message. Exemplarily, the computer device can display a reminder message on the display interface, such as in the form of a pop-up window or by opening the display interface; or, the computer device can also output a prompt voice through a buzzer or an alarm voice; or, the computer device can also send an alarm message to the relevant staff. The computer device can perform one or more of the above to output an alarm message.
[0133] In addition, when the voice recognition result is that the voice data to be recognized is true voice data, the computer device can also delete the true voice data to release storage resources.
[0134] In this embodiment, when the computer device determines that the voice data to be recognized is false voice data and outputs an alarm message, it can prompt and remind the staff, further improving the security of data recognition. When the computer device determines that the voice data to be recognized is true voice data, the to-be-recognized data is deleted in a timely manner to release computing resources and storage resources.
[0135] To better illustrate the above method, as Figure 9 shown, this embodiment provides a voice recognition method (illustrated by filtering out silent segments of candidate voice data), which specifically includes:
[0136] S101. Perform segmentation processing on the original voice data to obtain a plurality of sub-voice data;
[0137] S102. Perform denoising processing on each sub-voice data to obtain the denoised sub-voice data;
[0138] S103. Perform voice signal feature extraction processing on each denoised sub-voice data to determine the feature degree of each sub-voice data;
[0139] S104. Perform feature synthesis processing on the sub-voice data with a feature degree greater than a preset threshold to obtain candidate voice data;
[0140] S105. Perform voice segmentation on the candidate voice data to obtain silent segments and voiced segments in the candidate voice data;
[0141] S106. Merge the audible segments to obtain merged speech data;
[0142] S107. Perform modal decomposition processing on the merged speech data to obtain multiple modal components;
[0143] S108. Determine the modal components that meet the preset independence requirements as candidate modal components;
[0144] S109. Convert the candidate modal components into speech data to be recognized;
[0145] S110. Input the speech data to be recognized into a preset speech recognition model to obtain the credibility of the speech data to be recognized;
[0146] S111. If the credibility is greater than or equal to the preset credibility threshold, determine that the speech recognition result of the speech data to be recognized is real speech data;
[0147] S112. In the case where the speech recognition result of the speech data to be recognized is real speech data, delete the real speech data;
[0148] S113. If the credibility is less than the preset credibility threshold, determine that the speech recognition result of the speech data to be recognized is false speech data;
[0149] S114. In the case where the speech recognition result of the speech data to be recognized is false speech data, output an alarm message.
[0150] In this embodiment, the computer device performs segmentation processing on the original speech data, refines the processing granularity of the original speech data, and performs feature processing based on the sub-speech data obtained by segmentation, which can improve the accuracy of feature processing, thereby obtaining relatively accurate candidate speech data. On this basis, modal decomposition processing is performed on the candidate speech data. Modal decomposition processing can extract data with relatively high modal component independence from the candidate speech data, thereby further improving the accuracy of the speech data to be recognized. Based on the two-time speech data processing, inputting the speech data to be recognized into the speech recognition model, the obtained speech recognition result is more accurate, and the detection accuracy of real speech data recognition is improved.
[0151] The speech recognition method provided in the above embodiment has the same implementation principle and technical effect as the above method embodiment, and will not be elaborated here.
[0152] It should be understood that although the steps in the flowcharts involved in the above-described embodiments are sequentially shown according to the indications of the arrows, these steps are not necessarily executed sequentially in the order indicated by the arrows. Unless there is a clear indication in this article, the execution of these steps has no strict order restriction, and these steps can be executed in other orders. Moreover, at least a part of the steps in the flowcharts involved in the above-described embodiments may include multiple steps or multiple stages. These steps or stages are not necessarily executed at the same moment, but can be executed at different moments. The execution order of these steps or stages is not necessarily sequential, but can be executed alternately or in turn with at least a part of other steps or steps or stages in other steps.
[0153] Based on the same inventive concept, an embodiment of the present application further provides a speech recognition device for implementing the above-mentioned speech recognition method. The solution provided by this device to solve the problem is similar to the solution described in the above method. Therefore, the specific limitations in one or more embodiments of the speech recognition device provided below can refer to the limitations on the speech recognition method in the above text, and will not be repeated here.
[0154] In one embodiment, as Figure 10 shown, a speech recognition device is provided, including:
[0155] A first processing module 01, configured to segment the original speech data to obtain a plurality of sub-speech data, and perform feature processing on each sub-speech data to obtain candidate speech data;
[0156] A second processing module 02, configured to perform modal decomposition processing on the candidate speech data to obtain the speech data to be recognized;
[0157] A recognition module 03, configured to input the speech data to be recognized into a preset speech recognition model to obtain a speech recognition result; the speech recognition result is used to indicate whether the speech data to be recognized is real speech data.
[0158] In one of the optional embodiments, the first processing module 01 is configured to perform speech signal feature extraction processing on each sub-speech data to determine the feature degree of each sub-speech data; perform feature synthesis processing on the sub-speech data with a feature degree greater than a preset threshold to obtain candidate speech data.
[0159] In one of the optional embodiments, the first processing module 01 is further configured to perform denoising processing on each sub-speech data to obtain the denoised sub-speech data; perform feature processing on each denoised sub-speech data to obtain candidate speech data.
[0160] In one optional embodiment, the second processing module 02 is configured to perform modal decomposition processing on the candidate speech data to obtain a plurality of modal components; determine the modal components that meet the preset independence requirement as candidate modal components; and convert the candidate modal components into speech data to be recognized.
[0161] In one optional embodiment, as Figure 11 shown, the speech recognition device further includes a third processing module 04;
[0162] The third processing module 04 is configured to perform speech segmentation on the original speech data to obtain silent segments and voiced segments in the original speech data; merge the voiced segments to obtain merged speech data; the first processing module 01 is further configured to perform segmentation processing on the merged speech data to obtain a plurality of sub-speech data.
[0163] In one optional embodiment, the third processing module 04 is further configured to perform speech segmentation on the candidate speech data to obtain silent segments and voiced segments in the candidate speech data; merge the voiced segments to obtain merged speech data; the second processing module 02 is configured to perform modal decomposition processing on the merged speech data to obtain speech data to be recognized.
[0164] In one optional embodiment, the third processing module 04 is further configured to perform speech segmentation on the speech data to be recognized to obtain silent segments and voiced segments in the speech data to be recognized; merge the voiced segments to obtain merged speech data; the recognition module 03 is configured to input the merged speech data into a preset speech recognition model to obtain a speech recognition result.
[0165] In one optional embodiment, the recognition module 03 is configured to input the speech data to be recognized into a preset speech recognition model to obtain the credibility of the speech data to be recognized; if the credibility is greater than or equal to a preset credibility threshold, determine that the speech recognition result of the speech data to be recognized is real speech data; if the credibility is less than the preset credibility threshold, determine that the speech recognition result of the speech data to be recognized is false speech data.
[0166] In one optional embodiment, as Figure 12 shown, the speech recognition device further includes an output module 05, configured to output an alarm message when the speech recognition result of the speech data to be recognized is false speech data.
[0167] Each module in the above speech recognition device can be implemented in whole or in part by software, hardware, and their combination. The above modules can be embedded in the processor of the computer device in hardware form or be independent of it, or be stored in the memory of the computer device in software form, so that the processor can call and execute the operations corresponding to the above respective modules.
[0168] In one embodiment, a computer device is provided, including a memory and a processor. A computer program is stored in the memory. When the processor executes the computer program, the following steps are implemented:
[0169] Segment the original voice data to obtain multiple sub-voice data, and perform feature processing on each sub-voice data to obtain candidate voice data;
[0170] Perform modal decomposition processing on the candidate voice data to obtain the voice data to be recognized;
[0171] Input the voice data to be recognized into a preset voice recognition model to obtain a voice recognition result; the voice recognition result is used to indicate whether the voice data to be recognized is real voice data.
[0172] For the computer device provided in the above embodiment, its implementation principle and technical effect are similar to those of the above method embodiment, and will not be elaborated here.
[0173] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:
[0174] Segment the original voice data to obtain multiple sub-voice data, and perform feature processing on each sub-voice data to obtain candidate voice data;
[0175] Perform modal decomposition processing on the candidate voice data to obtain the voice data to be recognized;
[0176] Input the voice data to be recognized into a preset voice recognition model to obtain a voice recognition result; the voice recognition result is used to indicate whether the voice data to be recognized is real voice data.
[0177] For the computer-readable storage medium provided in the above embodiment, its implementation principle and technical effect are similar to those of the above method embodiment, and will not be elaborated here.
[0178] In one embodiment, a computer program product is provided, including a computer program. When the computer program is executed by a processor, the following steps are implemented:
[0179] Segment the original voice data to obtain multiple sub-voice data, and perform feature processing on each sub-voice data to obtain candidate voice data;
[0180] Perform modal decomposition processing on the candidate voice data to obtain the voice data to be recognized;
[0181] Input the speech data to be recognized into a preset speech recognition model to obtain a speech recognition result; the speech recognition result is used to indicate whether the speech data to be recognized is real speech data.
[0182] For the computer program product provided in the above embodiments, its implementation principle and technical effects are similar to those of the above method embodiments, and will not be elaborated here.
[0183] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data for analysis, stored data, displayed data, etc.) involved in this application are all information and data that have been authorized by the user or fully authorized by all parties.
[0184] Those of ordinary skill in the art can understand that all or part of the processes in the methods of the above embodiments can be completed by instructing relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the above method embodiments. Among them, any reference to a memory, database, or other medium used in the embodiments provided in this application can include at least one of non-volatile and volatile memories. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetoresistive random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can be in various forms, such as static random access memory (SRAM) or dynamic random access memory (DRAM), etc. The databases involved in the embodiments provided in this application can include at least one of relational databases and non-relational databases. Non-relational databases can include distributed databases based on blockchain, etc., and are not limited thereto. The processors involved in the embodiments provided in this application can be general-purpose processors, central processors, graphics processors, digital signal processors, programmable logic devices, data processing logics based on quantum computing, etc., and are not limited thereto.
[0185] The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity of description, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, it should be considered as the scope described in this specification.
[0186] The above-described embodiments merely represent several implementation manners of the present application. The description thereof is relatively specific and detailed, but it should not be construed as a limitation on the scope of the patent of the present application. It should be noted that for those of ordinary skill in the art, without departing from the concept of the present application, several modifications and improvements can still be made, and these all belong to the protection scope of the present application. Therefore, the protection scope of the present application shall be subject to the appended claims.
Claims
1. A speech recognition method, characterized in that, the method includes: segmenting the original speech data according to a preset number of frames or a preset duration to obtain a plurality of sub-speech data, and performing feature extraction processing and feature synthesis processing on each of the sub-speech data to obtain candidate speech data; the feature extraction processing includes inputting each of the sub-speech data into a feature extraction model to obtain the speech features corresponding to each of the sub-speech data, or performing wavelet decomposition on each of the sub-speech data to obtain the wavelet decomposition coefficients of each of the sub-speech data; performing modal decomposition processing on the candidate speech data through an empirical mode decomposition algorithm to obtain a plurality of modal components; performing independent component analysis on the plurality of modal components, and determining the modal components that meet the preset independence requirement as candidate modal components; converting the candidate modal components into speech data to be recognized; the independent component analysis includes analyzing the peak value, envelope spectrum and kurtosis of the modal components, or the independent component analysis includes analyzing the modal components through an independent component analysis algorithm ICA; inputting the speech data to be recognized into a preset speech recognition model to obtain a speech recognition result; the speech recognition result is used to indicate whether the speech data to be recognized is real speech data.
2. The method according to claim 1, characterized in that, the performing feature processing on each of the sub-speech data to obtain candidate speech data includes: performing speech signal feature extraction processing on each of the sub-speech data to determine the feature degree of each of the sub-speech data; performing feature synthesis processing on the sub-speech data with a feature degree greater than a preset threshold to obtain the candidate speech data.
3. The method according to claim 1 or 2, characterized in that, the method further includes: performing denoising processing on each of the sub-speech data to obtain the sub-speech data after denoising; the performing feature processing on each of the sub-speech data to obtain candidate speech data includes: performing feature processing on each of the sub-speech data after denoising to obtain the candidate speech data.
4. The method according to claim 1, characterized in that, the method further includes: performing speech segmentation on the original speech data to obtain silent segments and voiced segments in the original speech data; merging the voiced segments to obtain merged speech data; the segmenting the original speech data according to a preset number of frames or a preset duration to obtain a plurality of sub-speech data includes: segmenting the merged speech data to obtain a plurality of the sub-speech data.
5. The method according to claim 1, characterized in that, the method further includes: performing speech segmentation on the candidate speech data to obtain silent segments and voiced segments in the candidate speech data; merging the voiced segments to obtain merged speech data; the performing modal decomposition processing on the candidate speech data to obtain the speech data to be recognized includes: performing modal decomposition processing on the merged speech data to obtain the speech data to be recognized.
6. The method according to claim 1, characterized in that, the method further includes: Perform speech segmentation on the speech data to be recognized to obtain silent segments and audible segments in the speech data to be recognized; Merge the audible segments to obtain merged speech data; The step of inputting the speech data to be recognized into a preset speech recognition model to obtain a speech recognition result includes: Input the merged speech data into the preset speech recognition model to obtain the speech recognition result.
7. The method according to claim 1, wherein, the step of inputting the speech data to be recognized into a preset speech recognition model to obtain a speech recognition result includes: Input the speech data to be recognized into a preset speech recognition model to obtain the credibility of the speech data to be recognized; If the credibility is greater than or equal to a preset credibility threshold, determine that the speech recognition result is that the speech data to be recognized is real speech data; If the credibility is less than the preset credibility threshold, determine that the speech recognition result is that the speech data to be recognized is false speech data.
8. The method according to claim 7, wherein, the method further includes: Output an alarm message when the speech recognition result is that the speech data to be recognized is false speech data.
9. A speech recognition device, wherein, the device includes: A first processing module, configured to segment the original speech data according to a preset number of frames or a preset duration to obtain a plurality of sub-speech data, and perform feature extraction processing and feature synthesis processing on each of the sub-speech data to obtain candidate speech data; the feature extraction processing includes inputting each of the sub-speech data into a feature extraction model to obtain speech features corresponding to each of the sub-speech data, or performing wavelet decomposition on each of the sub-speech data to obtain wavelet decomposition coefficients of each of the sub-speech data; A second processing module, configured to perform modal decomposition processing on the candidate speech data through an empirical mode decomposition algorithm to obtain a plurality of modal components; perform independent component analysis on the plurality of modal components, and determine modal components that meet the preset independence requirement as candidate modal components; convert the candidate modal components into speech data to be recognized; the independent component analysis includes analyzing the peak value, envelope spectrum, and kurtosis of the modal components, or the independent component analysis includes analyzing the modal components through an independent component analysis algorithm ICA; A recognition module, configured to input the speech data to be recognized into a preset speech recognition model to obtain a speech recognition result; the speech recognition result is used to indicate whether the speech data to be recognized is real speech data.
10. A computer device, including a memory and a processor, the memory stores a computer program, wherein, when the processor executes the computer program, the steps of the method according to any one of claims 1 to 8 are implemented.
11. A computer-readable storage medium, on which a computer program is stored, wherein, when the computer program is executed by a processor, the steps of the method according to any one of claims 1 to 8 are implemented.
12. A computer program product comprising a computer program, wherein, when the computer program is executed by a processor, it implements the steps of the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
Voiceprint recognition method, device and equipment and computer readable storage medium
CN110364169A
Method and device for extracting voiced segment from audio file, equipment and storage medium
CN110910863A
Voiceprint recognition method based on voice noise reduction and related device
CN111108554A
Conference recording method and device based on voice recognition, equipment and storage medium
CN113870892A