Speech recognition method and device applied to noise environment, equipment and medium

By extracting and separating speech data in a noisy environment and fusing the denoised speech data in the semantic encoding and decoding process, the problem of low speech recognition accuracy in noisy environments is solved, and higher speech recognition reliability is achieved.

CN120612929APending Publication Date: 2025-09-09CHENGDU XIAOCHANG TECH CO LTD
View PDF 5 Cites 0 Cited by

Patent Information

Application Number
CN202510848566.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-24
Publication Date
2025-09-09

AI Technical Summary

Technical Problem

In a noisy environment, the accuracy of speech recognition is relatively low, and it is difficult to effectively recognize speech data.

Method used

By acquiring target speech data from a target environment and pre-configured standard noise data, the system extracts the speech data to be separated at the level of latent semantic information. Based on this information, it performs denoising and separation processing to generate denoised speech data. The denoised speech data is then integrated into the semantic encoding and decoding process of the speech data to form the target speech decoding features, ultimately enabling speech recognition.

Benefits of technology

It improves the accuracy of speech recognition in noisy environments, avoids the loss of semantic information caused by directly using denoised speech data, and ensures the reliability of speech recognition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120612929A_ABST
    Figure CN120612929A_ABST
Patent Text Reader

Abstract

The invention provides a speech recognition method and device applied to a noise environment, equipment and a medium, and relates to the technical field of computers. In the application, firstly, target voice data in a target environment is acquired, and standard noise data configured for the target environment in advance is acquired; secondly, extracting to-be-separated voice data having correlation with the standard noise data in a potential semantic information level from the target voice data; then, based on the to-be-separated voice data, performing denoising separation processing on the target voice data to form denoised voice data; further, in the semantic coding and decoding process of the target voice data, fusing the de-noised voice data to obtain target voice decoding features; and finally, performing speech recognition based on the target semantic decoding features, and outputting a target speech recognition result. Based on the above content, the problem that in the prior art, the accuracy of a speech recognition result in a noisy environment is relatively low can be solved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present application relates to the field of computer technology, and in particular to a method and apparatus, device, and medium for speech recognition in a noisy environment. Background Art

[0002] In recent years, the application of large AI models has become increasingly widespread, especially since the release of Deepseek's large model. It has penetrated multiple fields, including cloud computing, communications, healthcare, and audio-visual entertainment. When integrated with large AI models, voice recognition can take on the roles of intelligent assistant, customer service representative, and waiter, significantly lowering the user barrier to entry, improving product usability, and significantly increasing product competitiveness. For example, the AI ​​intelligent voice assistant is a comprehensive AI intelligent voice assistant for Yinchuang's full line of hardware products. It serves user needs such as requesting songs, watching movies, placing orders, controlling devices, and troubleshooting common problems, fulfilling the roles of intelligent assistant, customer service representative, and waiter. However, noise interference is often present during use (for example, in a karaoke setting, the noise level is typically 80-100dB; in a home living room, the noise level is typically 60-80dB; and in outdoor settings, the noise level is typically 40-60dB). This can lead to relatively low accuracy during speech recognition. Summary of the Invention

[0003] In view of this, the purpose of the present application is to provide a method and apparatus, device and medium for speech recognition in a noisy environment, so as to improve the problem in the prior art of relatively low accuracy of speech recognition results in a noisy environment.

[0004] To achieve the above objectives, this application adopts the following technical solutions: A speech recognition method applied in a noisy environment, comprising: Acquire target speech data in a target environment, and acquire standard noise data pre-configured for the target environment, wherein the target speech data belongs to speech data to be recognized, and the standard noise data includes at least one type of noise; Extracting speech data to be separated from the target speech data, which has correlation with the standard noise data at the level of latent semantic information; Performing denoising and separation processing on the target voice data based on the voice data to be separated to form denoised voice data; In the semantic encoding and decoding process of the target speech data, the denoised speech data is integrated to obtain target speech decoding features; Speech recognition is performed based on the target semantic decoding features, and a target speech recognition result is output.

[0005] In a preferred embodiment of the present application, in the above-mentioned speech recognition method applied in a noisy environment, the step of extracting the speech data to be separated that is correlated with the standard noise data at the level of latent semantic information from the target speech data includes: Performing feature encoding on the target speech data and the standard noise data respectively to form target speech features corresponding to the target speech data and standard noise features corresponding to the standard noise data, wherein the target speech features and the standard noise features belong to latent semantic features of the same dimension, which belongs to the time domain or the frequency domain; Performing multi-level forward diffusion on the target speech feature and the standard noise feature respectively to form a plurality of forward diffusion feature combinations, wherein each of the forward diffusion feature combinations includes a speech forward diffusion feature corresponding to the target speech feature and a noise forward diffusion feature corresponding to the standard noise feature; respectively determining a feature correlation between the speech forward diffusion feature and the noise forward diffusion feature in each of the forward diffusion feature combinations; A forward diffusion feature combination having a maximum feature correlation is determined, and the speech forward diffusion features in the forward diffusion feature combination are subjected to at least one backward diffusion, and the obtained backward diffusion features are feature decoded to form speech data to be separated, wherein the number of the at least one backward diffusion is equal to the number of forward diffusions corresponding to the forward diffusion feature combination having the maximum feature correlation.

[0006] In a preferred embodiment of the present application, in the above-mentioned speech recognition method applied in a noisy environment, the step of performing multi-level forward diffusion on the target speech features and the standard noise features to form multiple forward diffusion feature combinations includes: Based on the standard noise feature, performing attention processing on the target speech feature to form a speech attention feature corresponding to the target speech feature; Multi-level forward diffusion is performed on the speech attention feature and the standard noise feature respectively to form multiple forward diffusion feature combinations.

[0007] In a preferred embodiment of the present application, in the above-mentioned speech recognition method applied in a noisy environment, the step of performing multi-level forward diffusion on the target speech features and the standard noise features to form multiple forward diffusion feature combinations includes: Based on the standard noise feature, performing attention processing on the target speech feature to form a speech attention feature corresponding to the target speech feature, and combining the speech attention feature and the standard noise feature as a forward diffusion feature; Performing at least one level of compression operation on the target speech feature and the standard noise feature respectively to achieve forward diffusion, thereby forming at least one compressed feature combination, wherein each compressed feature combination includes a compressed speech feature corresponding to the target speech feature and a compressed noise feature corresponding to the standard noise feature; For each of the compressed feature combinations, attention processing is performed on the compressed speech features included in the compressed feature combination based on the compressed noise features included in the compressed feature combination to form a speech attention feature corresponding to the compressed speech feature, and the speech attention feature and the compressed noise feature are used as a forward diffusion feature combination.

[0008] In a preferred embodiment of the present application, in the above-mentioned method for speech recognition in a noisy environment, the step of fusing the denoised speech data to obtain target speech decoding features during the semantic encoding and decoding process of the target speech data includes: Performing feature encoding on the target speech data to form target speech coding features, and performing feature encoding on the denoised speech data to form denoised speech coding features; In the process of performing multiple levels of forward diffusion processing on the target speech coding features, the denoised speech coding features are gradually fused to form multiple levels of forward fused speech features, wherein the object of a subsequent forward diffusion processing includes the forward fused speech features of the previous level and the denoised speech coding features; In the process of performing multiple levels of backward diffusion processing on the multiple levels of fused speech features, the denoised speech coding features are gradually fused to form multiple levels of backward fused speech features, wherein the objects of the subsequent backward diffusion processing include the backward fused speech features of the previous level, the forward fused speech features of the corresponding level, and the denoised speech coding features, and the manner in which the denoised speech coding features are fused in the forward diffusion processing is different from the manner in which the denoised speech coding features are fused in the backward diffusion processing; Based on the backward fusion speech features of the last level, the target speech decoding features are determined.

[0009] In a preferred embodiment of the present invention, in the above-mentioned speech recognition method for a noisy environment, the step of gradually fusing the denoised speech coding features to form multiple levels of forward fused speech features during the process of performing multiple levels of forward diffusion processing on the target speech coding features includes: In the forward diffusion process of the first level, attention processing is performed on the target speech coding features based on the denoised speech coding features to form the forward fusion speech features of the first level; In the forward diffusion process of the i-th level, the forward fusion speech features of the i-1-th level and the denoised speech coding features are respectively compressed, and based on the compressed denoised speech coding features, the compressed forward fusion speech features are subjected to attention processing to form the forward fusion speech features of the i-th level, where i is greater than 1.

[0010] In a preferred embodiment of the present invention, in the above-mentioned speech recognition method for a noisy environment, the step of gradually fusing the denoised speech coding features to form multiple levels of backward fused speech features during the process of performing multiple levels of backward diffusion processing on the multiple levels of fused speech features includes: In the backward diffusion process of the first level, linear mapping and nonlinear activation operations are performed on the denoised speech coding features to form a target parameter distribution, and each parameter in the target parameter distribution is used as a weight coefficient of a parameter at a corresponding position in the fused speech features of the last level to update the fused speech features of the last level to form the backward fused speech features of the first level; During the forward diffusion process at the jth level, the backward fusion speech features at the j-1th level are expanded, and the expanded backward fusion speech features and the forward fusion speech features at the corresponding level are added to form a bidirectional fusion speech feature, and the denoising speech coding features are linearly mapped and nonlinearly activated to form a target parameter distribution, and each parameter in the target parameter distribution is used as a weight coefficient of the parameter at the corresponding position in the bidirectional fusion speech feature to update the bidirectional fusion speech feature to form a backward fusion speech feature at the jth level, where j is greater than 1.

[0011] The present application also provides a speech recognition device for use in a noisy environment, comprising: a data acquisition module, configured to acquire target speech data in a target environment and acquire standard noise data pre-configured for the target environment, wherein the target speech data belongs to speech data to be recognized and the standard noise data includes at least one type of noise; A data extraction module is used to extract the speech data to be separated that is correlated with the standard noise data at the level of latent semantic information from the target speech data; A speech denoising module, configured to perform denoising and separation processing on the target speech data based on the speech data to be separated to form denoised speech data; A speech encoding and decoding module, configured to fuse the denoised speech data during the semantic encoding and decoding process of the target speech data to obtain target speech decoding features; The speech recognition module is used to perform speech recognition based on the target semantic decoding features and output a target speech recognition result.

[0012] Based on the above, the present application further provides an electronic device, including: memory for storing computer programs; The processor connected to the memory is used to execute the computer program stored in the memory to implement the above-mentioned speech recognition method applied in a noisy environment.

[0013] On the basis of the above, the present application further provides a computer-readable storage medium, in which a computer program is stored. When the computer program is run, each step of the above-mentioned method for speech recognition in a noisy environment is executed.

[0014] The present application provides a speech recognition method, apparatus, device, and medium for use in a noisy environment. First, target speech data in a target environment is obtained, and standard noise data pre-configured for the target environment is obtained. Second, speech data to be separated that has a correlation with the standard noise data at the level of potential semantic information is extracted from the target speech data. Then, denoising and separation processing is performed on the target speech data based on the speech data to be separated to form denoised speech data. Furthermore, in the semantic encoding and decoding process of the target speech data, the denoised speech data is integrated to obtain target speech decoding features. Finally, speech recognition is performed based on the target semantic decoding features, and the target speech recognition result is output. Based on the above content, when extracting noise data (i.e., the speech data to be separated) from the target speech data, the trained neural network model is not directly used for noise prediction, but the standard noise data configured in the corresponding environment is used, so that the reliability of the extracted noise data is higher. In this way, the reliability of the denoised speech data formed can be further guaranteed. On this basis, since speech recognition is not performed directly on the denoised speech data, but it is further combined with the target speech data, the semantic representation accuracy of the target speech decoding features formed can be higher, avoiding the problem of direct use of the denoised speech data resulting in the loss of some speech semantic information. Accordingly, the reliability of speech recognition can be further improved, thereby improving the problem of relatively low accuracy of speech recognition results in noisy environments in the existing technology. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to make the above-mentioned objects, features and advantages of the present application more obvious and easy to understand, preferred embodiments are given below and described in detail with reference to the accompanying drawings.

[0016] Figure 1 This is a structural block diagram of an electronic device provided in an embodiment of the present application.

[0017] Figure 2 This is a first schematic diagram of a speech recognition method applied in a noisy environment provided in an embodiment of the present application.

[0018] Figure 3 This is a second schematic diagram of a speech recognition method applied in a noisy environment provided in an embodiment of the present application.

[0019] Figure 4 A schematic diagram of fusing denoised speech data into the semantic encoding and decoding process of target speech data provided in an embodiment of the present application.

[0020] Figure 5 A block diagram of a speech recognition device for use in a noisy environment provided in an embodiment of the present application. DETAILED DESCRIPTION

[0021] To make the objectives, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Generally, the components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations.

[0022] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application for protection, but merely represents selected embodiments of the present application. All other embodiments obtained by persons of ordinary skill in the art based on the embodiments in the present application without creative work are within the scope of protection of the present application.

[0023] like Figure 1 As shown, an embodiment of the present application provides an electronic device, wherein the electronic device may include a memory, a processor, and a speech recognition device for use in a noisy environment.

[0024] In detail, the memory and the processor are electrically connected directly or indirectly to realize data transmission or interaction. For example, the memory and the processor can be electrically connected through one or more communication buses or signal lines. The speech recognition device for use in a noisy environment includes at least one software function module stored in the memory in the form of software or firmware. The processor is used to execute the executable computer program stored in the memory, for example, the software function module and computer program included in the speech recognition device for use in a noisy environment, so as to realize the speech recognition method for use in a noisy environment provided in the embodiment of the present application.

[0025] Optionally, the memory may be, but is not limited to, random access memory (RAM), read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), etc.

[0026] Furthermore, the processor may be a general-purpose processor, including a central processing unit (CPU), a network processor (NP), a system on chip (SoC), etc.; it may also be a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components.

[0027] I understand. Figure 1 The structure shown is only for illustration, and the electronic device may also include Figure 1 More or fewer components than shown, or with Figure 1 The different configurations shown, for example, may also include a communication unit for exchanging information with other devices (such as a voice acquisition device, etc.).

[0028] Combine Figure 2 and Figure 3 The embodiment of the present application also provides a method for speech recognition in a noisy environment that can be applied to the above electronic device. The method steps defined in the process related to the speech recognition method in a noisy environment can be implemented by the electronic device. Figure 2 The specific process shown is explained in detail.

[0029] Step S110 , obtaining target speech data in a target environment, and obtaining standard noise data pre-configured for the target environment.

[0030] In an embodiment of the present application, the electronic device can obtain target voice data in a target environment and obtain standard noise data pre-configured for the target environment. The target voice data is the voice data to be recognized, and the standard noise data includes at least one type of noise. For example, a KTV environment, a home environment, and an outdoor environment generally have different noises. Therefore, by configuring the corresponding standard noise data, the reliability of subsequent noise separation can be improved.

[0031] Step S120 , extracting the speech data to be separated from the target speech data, which is correlated with the standard noise data at the level of latent semantic information.

[0032] In an embodiment of the present application, after acquiring the target speech data and the standard noise data, the electronic device may extract, from the target speech data, speech data to be separated that is correlated with the standard noise data at the level of latent semantic information. Since the standard noise data is noise, extracting speech data to be separated that is correlated with the standard noise data from the target speech data actually means extracting noise from the target speech data, i.e., the speech data to be separated is the extracted noise.

[0033] Step S130 : performing denoising and separation processing on the target voice data based on the voice data to be separated to form denoised voice data.

[0034] In an embodiment of the present application, after extracting the to-be-separated voice data, the electronic device may perform denoising and separation processing on the target voice data based on the to-be-separated voice data to form denoised voice data. For example, the to-be-separated voice data and the target voice data are both time-domain data. Therefore, the target voice data and the to-be-separated voice data may be subtracted from each other to denoise the target voice data, thereby obtaining denoised voice data after denoising, i.e., denoised voice data with no noise or relatively reduced noise content.

[0035] Step S140 , in the semantic encoding and decoding process of the target speech data, the denoised speech data is integrated to obtain target speech decoding features.

[0036] In an embodiment of the present application, after forming the denoised voice data, the electronic device can fuse the denoised voice data in the semantic encoding and decoding process of the target voice data to obtain the target voice decoding features. That is to say, at the actual processing level, there may be a certain error between the extracted voice data to be separated and the actual noise. Therefore, if voice recognition is performed directly based on the denoised voice data, it may be difficult to effectively guarantee the reliability of voice recognition. That is to say, a certain degree of effective semantic information may be lost in the denoised voice data. In this case, it is necessary to further combine the original target voice data so that while ensuring that the semantic information is not lost, the noise can be suppressed or removed to a certain extent. That is, in the semantic encoding and decoding process of the target voice data, the denoised voice data is fused to obtain the target voice decoding features.

[0037] Step S150: performing speech recognition based on the target semantic decoding features and outputting a target speech recognition result.

[0038] In an embodiment of the present application, after forming the target semantic decoding feature, the electronic device can perform speech recognition based on the target semantic decoding feature and output the target speech recognition result. For example, the target semantic decoding feature can be fully connected to obtain the corresponding fully connected feature, wherein the size of the fully connected feature is the same as the number of speech recognition types, and the type can be "pause", "play", "previous song", "next song", "accompaniment", etc. In this way, the fully connected feature can be further activated (which can be achieved through functions such as softmax) to obtain the corresponding probability distribution, and then the type corresponding to the probability with the maximum value is determined from the probability distribution as the target speech recognition result. It can be understood that in other embodiments, the target semantic decoding feature can also be processed by a trained generative model to generate corresponding instructions and other data, such as recommending songs, lowering the volume, etc.

[0039] Based on the above content, when extracting noise data (i.e., the speech data to be separated) from the target speech data, the trained neural network model is not directly used for noise prediction, but the standard noise data configured in the corresponding environment is used, so that the reliability of the extracted noise data is higher. In this way, the reliability of the denoised speech data formed can be further guaranteed. On this basis, since speech recognition is not performed directly on the denoised speech data, but it is further combined with the target speech data, the semantic representation accuracy of the target speech decoding features formed can be higher, avoiding the problem of direct use of the denoised speech data resulting in the loss of some speech semantic information. Accordingly, the reliability of speech recognition can be further improved, thereby improving the problem of relatively low accuracy of speech recognition results in noisy environments in the existing technology.

[0040] First, it should be noted that for step S120, the specific method of extracting the speech data to be separated that is correlated with the standard noise data at the level of latent semantic information from the target speech data is not limited and can be selected according to actual needs.

[0041] For example, in an alternative embodiment, the target speech data and the standard noise data may be feature-encoded separately to generate target speech features corresponding to the target speech data and standard noise features corresponding to the standard noise data. Then, based on the target speech features, cross-attention processing may be performed on the standard noise features to obtain the speech data to be separated, thereby improving the efficiency of feature extraction.

[0042] For example, in another alternative embodiment, in order to improve the reliability of feature extraction, that is, to ensure that effective noise is extracted from the target speech data as much as possible, the above-mentioned step S120 may further include step S121, step S122, step S123 and step S124, and the specific content of each step is described below.

[0043] Step S121 , feature encoding is performed on the target speech data and the standard noise data respectively to form target speech features corresponding to the target speech data and standard noise features corresponding to the standard noise data.

[0044] In an embodiment of the present application, the target speech data and the standard noise data can be feature-encoded respectively to form target speech features corresponding to the target speech data and standard noise features corresponding to the standard noise data. The target speech features and the standard noise features belong to the same dimension of latent semantic features, which belongs to the time domain or the frequency domain. That is to say, the target speech data in the time domain and the standard noise data in the time domain can be feature-encoded respectively to form corresponding target speech features and corresponding standard noise features. The target speech data and the standard noise data can also be feature-encoded on the corresponding spectrograms after being converted to the frequency domain respectively, thereby forming target speech features and standard noise features. The specific method of feature encoding the data in the time domain and the frequency domain is not limited, and any existing technology can be used. For example, the spectrogram can be convolved (as well as pooled, fully connected, etc.) to form corresponding features.

[0045] Step S122 , performing multi-level forward diffusion on the target speech feature and the standard noise feature respectively to form a plurality of forward diffusion feature combinations.

[0046] In an embodiment of the present application, after forming the target speech feature and the standard noise feature, the target speech feature and the standard noise feature can be subjected to multi-level forward diffusion respectively to form multiple forward diffusion feature combinations. Each forward diffusion feature combination includes a speech forward diffusion feature corresponding to the target speech feature and a noise forward diffusion feature corresponding to the standard noise feature, and the speech forward diffusion feature and the noise forward diffusion feature are formed based on the same level of forward diffusion, that is, the feature sizes are guaranteed to be the same.

[0047] Step S123 , determining the feature correlation between the speech forward diffusion feature and the noise forward diffusion feature in each of the forward diffusion feature combinations.

[0048] In an embodiment of the present application, after the forward diffusion feature combination is formed, the feature correlation between the speech forward diffusion feature and the noise forward diffusion feature in each forward diffusion feature combination can be determined respectively, such as cosine similarity.

[0049] Step S124: determine a forward diffusion feature combination with a maximum feature correlation, perform at least one backward diffusion on the speech forward diffusion features in the forward diffusion feature combination, and perform feature decoding on the obtained backward diffusion features to form speech data to be separated.

[0050] In an embodiment of the present application, after determining the feature correlation, a forward diffusion feature combination with a maximum feature correlation, that is, a forward diffusion feature combination with the highest correlation, can be determined, and the speech forward diffusion features in the forward diffusion feature combination are subjected to at least one backward diffusion, and the obtained backward diffusion features are feature decoded to form the speech data to be separated. The number of backward diffusions is equal to the number of forward diffusions corresponding to the forward diffusion feature combination with the maximum feature correlation. For example, if the forward diffusion feature combination with the maximum feature correlation is formed based on the first forward diffusion, the speech forward diffusion features are subjected to one backward diffusion; if the forward diffusion feature combination with the maximum feature correlation is formed based on the second forward diffusion, the speech forward diffusion features are subjected to two backward diffusions, and so on. In addition, considering that features are compressed (i.e., downsampled) during the forward diffusion process to gradually reduce their size, features can be expanded (e.g., upsampled) during the backward diffusion process to gradually restore their size to the target speech feature's size (the same as the standard noise feature's size). During feature decoding, the backward diffusion features can be mapped to the size corresponding to the spectrograms of the target speech data and the standard noise data, thereby forming a decoded spectrogram corresponding to the backward diffusion features. This decoded spectrogram is then inverse Fourier transformed to form the speech data to be separated, i.e., converting from the frequency domain to the time domain.

[0051] It is understandable that, in the above-mentioned step S122, the specific manner of performing multi-level forward diffusion on the target speech feature and the standard noise feature is not limited and can be selected according to actual needs.

[0052] For example, in an alternative embodiment, in order to reduce the amount of calculation and improve the calculation efficiency, the above-mentioned step S122 may further include step S122a and step S122b, and the specific content of each step is described as follows.

[0053] Step S122a: performing attention processing on the target speech feature based on the standard noise feature to form a speech attention feature corresponding to the target speech feature.

[0054] In an embodiment of the present application, attention processing (cross-attention processing, wherein the query feature corresponds to the standard noise feature, and the key feature and the value feature correspond to the target speech feature) can be performed on the target speech feature based on the standard noise feature to form a speech attention feature corresponding to the target speech feature, that is, semantic features that have a correlation relationship with the standard noise feature are mined from the target speech feature.

[0055] Step S122b: Perform multi-level forward diffusion on the speech attention feature and the standard noise feature respectively to form multiple forward diffusion feature combinations.

[0056] In an embodiment of the present application, multi-level forward diffusion is performed on the speech attention feature and the standard noise feature respectively to form multiple forward diffusion feature combinations. Among them, the object of the latter forward diffusion is the previous forward diffusion feature combination. For example, by performing convolution, pooling and full connection processing on the speech attention feature, the speech forward diffusion feature in the first forward diffusion feature combination can be obtained; and by performing convolution, pooling and full connection processing on the standard noise feature, the noise forward diffusion feature in the first forward diffusion feature combination can be obtained.

[0057] For another example, in another alternative implementation, in order to more reliably determine the voice data to be separated, the above-mentioned step S122 may further include step S122c, step S122d and step S122e, and the specific content of each step is described as follows.

[0058] Step S122c: performing attention processing on the target speech feature based on the standard noise feature to form a speech attention feature corresponding to the target speech feature, and combining the speech attention feature and the standard noise feature as a forward diffusion feature.

[0059] In an embodiment of the present application, attention processing (i.e., cross-attention processing) can be performed on the target speech feature based on the standard noise feature to form a speech attention feature corresponding to the target speech feature, and the speech attention feature and the standard noise feature can be combined as a forward diffusion feature.

[0060] Step S122d: performing at least one level of compression operation on the target speech feature and the standard noise feature respectively to achieve forward diffusion and form at least one compression feature combination.

[0061] In an embodiment of the present application, at least one level of compression (i.e., downsampling) can be performed on each of the target speech feature and the standard noise feature to achieve forward diffusion, thereby forming at least one compressed feature combination. Each compressed feature combination includes a compressed speech feature corresponding to the target speech feature and a compressed noise feature corresponding to the standard noise feature.

[0062] Step S122e, for each of the compressed feature combinations, based on the compressed noise features included in the compressed feature combination, attention processing is performed on the compressed speech features included in the compressed feature combination to form a speech attention feature corresponding to the compressed speech feature, and the speech attention feature and the compressed noise feature are used as a forward diffusion feature combination.

[0063] In an embodiment of the present application, after forming the compressed feature combinations, for each of the compressed feature combinations, attention processing (i.e., cross-attention processing) can be performed on the compressed speech features included in the compressed feature combination based on the compressed noise features included in the compressed feature combination to form a speech attention feature corresponding to the compressed speech feature, and the speech attention feature and the compressed noise feature are used as a forward diffusion feature combination. In other words, each forward diffusion feature combination includes features obtained by the cross-attention processing.

[0064] Secondly, it should be noted that, with respect to step S140 , the specific manner of fusing the denoised speech data during the semantic encoding and decoding process of the target speech data is not limited and can be selected according to actual needs.

[0065] For example, in an alternative embodiment, to improve computational efficiency and reduce computational overhead, the target speech data may be feature-encoded to form target speech coding features, and the denoised speech data may be feature-encoded to form denoised speech coding features. The target speech coding features and the denoised speech coding features may then be concatenated to obtain concatenated features, which may then be feature-decoded to form target speech decoding features.

[0066] For example, in another alternative embodiment, in order to achieve effective fusion of the denoised speech data so that the semantic representation accuracy of the formed target speech decoding features can be higher, the above-mentioned step S140 can further include step S141, step S142, step S143 and step S144, and the specific content of each step is described below.

[0067] Step S141 , feature encoding is performed on the target speech data to form target speech coding features, and feature encoding is performed on the denoised speech data to form denoised speech coding features.

[0068] In an embodiment of the present application, the target speech data can be feature-encoded to form target speech coding features, and the denoised speech data can be feature-encoded to form denoised speech coding features. For example, the target speech data can be converted into the frequency domain to obtain a corresponding spectrogram, and then the spectrogram can be subjected to convolution or other processing to achieve feature coding and form target speech coding features. Alternatively, the denoised speech data can be converted into the frequency domain to obtain a corresponding spectrogram, and then the spectrogram can be subjected to convolution or other processing to achieve feature coding and form denoised speech coding features.

[0069] Step S142: In the process of performing multiple levels of forward diffusion processing on the target speech coding features, the denoised speech coding features are gradually fused to form multiple levels of forward fused speech features.

[0070] In an embodiment of the present application, after forming the target speech coding features and the denoised speech coding features, the denoised speech coding features can be gradually fused during the process of performing multiple levels of forward diffusion processing on the target speech coding features to form multiple levels of forward fused speech features. The objects of the subsequent forward diffusion processing include the forward fused speech features of the previous level and the denoised speech coding features. In this way, by performing the fusion of the denoised speech coding features during the multiple levels of forward diffusion processing, the fusion effect of the denoised speech coding features can be improved, that is, the fusion is more complete.

[0071] Step S143 , in the process of performing multiple-level backward diffusion processing on the multiple-level fused speech features, the denoised speech coding features are gradually fused to form multiple-level backward fused speech features.

[0072] In an embodiment of the present application, after forming the multiple levels of fused speech features, the denoised speech coding features can be gradually fused in the process of performing multiple levels of backward diffusion processing on the multiple levels of fused speech features to form multiple levels of backward fused speech features. Among them, the objects of the latter backward diffusion processing include the backward fused speech features of the previous level, the forward fused speech features of the corresponding level, and the denoised speech coding features, and the way of fusing the denoised speech coding features in the forward diffusion processing is different from the way of fusing the denoised speech coding features in the backward diffusion processing. That is to say, the denoised speech coding features will be fused in the process of forward diffusion processing and backward diffusion processing, so that the fusion of the denoised speech coding features is more complete, that is, the semantic information after denoising can be fully utilized.

[0073] Step S144: Determine target speech decoding features based on the backward fusion speech features of the last level.

[0074] In the embodiment of the present application, after obtaining the last level of backward fusion speech features, the target speech decoding features can be determined based on the last level of backward fusion speech features. For example, the last level of backward fusion speech features can be used as the target speech decoding features.

[0075] It is understood that, in the above step S142, the specific manner of forming multiple levels of forward fusion speech features is not limited. For example, in an alternative embodiment, in order to achieve reliable fusion of denoised speech coding features in each level of forward diffusion processing, the above step S142 may further include step S142a and step S142b. The specific contents of each step are as follows (combined with Figure 4 shown).

[0076] Step S142a: In the forward diffusion process of the first level, attention processing is performed on the target speech coding features based on the denoised speech coding features to form the forward fusion speech features of the first level.

[0077] In the embodiment of the present application, in the forward diffusion process of the first level, the target speech coding feature is subjected to attention processing (i.e., cross attention processing) based on the denoised speech coding feature to form the forward fusion speech feature of the first level. Step S142b, in the forward diffusion process of the i-th level, compression operations are performed on the forward fusion speech features and the denoised speech coding features of the i-1-th level respectively, and based on the compressed denoised speech coding features, attention processing is performed on the compressed forward fusion speech features to form the forward fusion speech features of the i-th level.

[0078] In an embodiment of the present application, after forming the forward fusion speech features of the first level, in the forward diffusion process of the i-th level, the forward fusion speech features of the i-1th level and the denoising speech coding features are respectively compressed (i.e., downsampling operation, and the two compressed features have the same feature size), and based on the compressed denoising speech coding features, the compressed forward fusion speech features are subjected to attention processing to form the forward fusion speech features of the i-th level, where i is greater than 1, i.e., an integer greater than or equal to 2.

[0079] It is understandable that, in the above-mentioned step S143, the specific manner of forming multiple levels of backward fusion speech features is not limited. For example, in an alternative embodiment, in order to achieve reliable fusion of denoised speech coding features in each level of forward diffusion processing, so that backward fusion speech features that can effectively represent global semantic information can be obtained, the above-mentioned step S143 can further include step S143a and step S143b. The specific contents of each step are as follows (further combined with Figure 4 shown).

[0080] Step S143a, in the backward diffusion process of the first level, linear mapping and nonlinear activation operations are performed on the denoised speech coding features to form a target parameter distribution, and each parameter in the target parameter distribution is used as a weight coefficient of the parameter at the corresponding position in the fused speech features of the last level to update the fused speech features of the last level to form the backward fused speech features of the first level.

[0081] In an embodiment of the present application, during the backward diffusion process of the first level, the denoised speech coding features are subjected to linear mapping (to capture linear relationships, such as by performing full connection processing) and nonlinear activation operations (to capture nonlinear relationships, which can be specifically achieved through a sigmiod function) to form a target parameter distribution, and each parameter in the target parameter distribution is used as a weight coefficient of the parameter at the corresponding position in the fused speech features of the last level to update the fused speech features of the last level (for example, a bitwise multiplication operation can be performed) to form the backward fused speech features of the first level.

[0082] Step S143b, in the forward diffusion process of the j-th level, the backward fusion speech features of the j-1-th level are expanded, and the expanded backward fusion speech features and the forward fusion speech features of the corresponding level are added to form a bidirectional fusion speech feature, and the denoising speech coding features are linearly mapped and nonlinearly activated to form a target parameter distribution, and each parameter in the target parameter distribution is used as a weight coefficient of the parameter at the corresponding position in the bidirectional fusion speech feature to update the bidirectional fusion speech feature to form the backward fusion speech feature of the j-th level.

[0083] In an embodiment of the present application, after forming the backward fusion speech features of the first level, the backward fusion speech features of the j-1th level can be expanded during the forward diffusion process of the jth level, and the expanded backward fusion speech features and the forward fusion speech features of the corresponding level (the expansion operation can make the sizes of the two features the same, thereby facilitating the addition operation) are added to form a bidirectional fusion speech feature, and the denoising speech coding features are linearly mapped and nonlinearly activated to form a target parameter distribution, and each parameter in the target parameter distribution is used as a weight coefficient of the parameter at the corresponding position in the bidirectional fusion speech feature to update the bidirectional fusion speech feature (for example, a bitwise multiplication operation can be performed) to form a backward fusion speech feature of the jth level, where j is greater than 1, that is, an integer greater than or equal to 2.

[0084] It should be noted here that compared to linear mapping, nonlinear activation operations, and bitwise multiplication operations, the computational complexity of cross-attention processing is greater. Therefore, cross-attention processing is not used to achieve the fusion of different features during the backward diffusion process. The purpose of using cross-attention processing to achieve the fusion of different features during the forward diffusion process is that, during the forward diffusion process, considering the large difference between the target speech coding features and the denoised speech coding features, the combination of the above-mentioned linear mapping, nonlinear activation operations, and bitwise multiplication operations has a strong dependence on the features used as weight coefficients, which can easily lead to poor fusion effects (poor stability of feature fusion). However, during the backward diffusion process, due to the high-precision fusion in the forward diffusion process, the features used as weight coefficients are more stable, thereby reducing computational overhead while also ensuring the reliability of fusion.

[0085] In addition, it should be noted that in the process of forward diffusion processing and backward diffusion processing, using different fusion methods can also facilitate the capture of different complex semantic relationships, thereby further improving the semantic representation ability of the obtained candidate fused speech features.

[0086] Combine Figure 5 The present application also provides a speech recognition device for use in a noisy environment, which can be applied to the above-mentioned electronic device. The speech recognition device for use in a noisy environment may include a data acquisition module, a data extraction module, a speech denoising module, a speech encoding and decoding module, and a speech recognition module.

[0087] In detail, the data acquisition module can be used to acquire target speech data in a target environment and acquire standard noise data pre-configured for the target environment, wherein the target speech data belongs to speech data to be recognized and the standard noise data includes at least one type of noise. In the embodiment of the present application, the data acquisition module can be used to execute Figure 2 As shown in step S110, for the relevant content of the data acquisition module, reference may be made to the above description of step S110.

[0088] In detail, the data extraction module can be used to extract the speech data to be separated that is correlated with the standard noise data at the level of potential semantic information from the target speech data. Figure 2 As shown in step S120, for the relevant content of the data extraction module, reference can be made to the above description of step S120.

[0089] In detail, the speech denoising module can be used to perform denoising and separation processing on the target speech data based on the speech data to be separated to form denoised speech data. In the embodiment of the present application, the speech denoising module can be used to perform Figure 2 As shown in step S130, for the relevant content of the speech denoising module, reference can be made to the above description of step S130.

[0090] In detail, the speech encoding and decoding module can be used to fuse the denoised speech data in the semantic encoding and decoding process of the target speech data to obtain the target speech decoding features. Figure 2 As shown in step S140, for the relevant content of the speech encoding and decoding module, reference can be made to the above description of step S140.

[0091] In detail, the speech recognition module can be used to perform speech recognition based on the target semantic decoding features and output the target speech recognition result. Figure 2 As shown in step S150, for the relevant content of the speech recognition module, reference can be made to the above description of step S150.

[0092] In an embodiment of the present application, corresponding to the above-mentioned speech recognition method applied to the electronic device in a noisy environment, a computer-readable storage medium is also provided, in which a computer program is stored. When the computer program is run, the various steps of the speech recognition method applied to the noisy environment are executed.

[0093] The steps executed when the aforementioned computer program is running will not be described in detail here. Please refer to the above explanation of the speech recognition method applied in a noisy environment.

[0094] In summary, the speech recognition method, apparatus, device and medium provided in the present application for application in a noisy environment, first, obtain the target speech data in the target environment, and obtain the standard noise data pre-configured for the target environment; secondly, extract the speech data to be separated that has a correlation with the standard noise data at the level of potential semantic information from the target speech data; then, perform denoising and separation processing on the target speech data based on the speech data to be separated to form denoised speech data; further, in the semantic encoding and decoding process of the target speech data, fuse the denoised speech data to obtain the target speech decoding features; finally, perform speech recognition based on the target semantic decoding features, and output the target speech recognition results. Based on the above content, when extracting noise data (i.e., the speech data to be separated) from the target speech data, the trained neural network model is not directly used for noise prediction, but the standard noise data configured in the corresponding environment is used, so that the reliability of the extracted noise data is higher. In this way, the reliability of the denoised speech data formed can be further guaranteed. On this basis, since speech recognition is not performed directly on the denoised speech data, but it is further combined with the target speech data, the semantic representation accuracy of the target speech decoding features formed can be higher, avoiding the problem of direct use of the denoised speech data resulting in the loss of some speech semantic information. Accordingly, the reliability of speech recognition can be further improved, thereby improving the problem of relatively low accuracy of speech recognition results in noisy environments in the existing technology.

[0095] In the several embodiments provided in the embodiments of the present application, it should be understood that the disclosed devices and methods can also be implemented in other ways. The device and method embodiments described above are merely schematic. For example, the flowcharts and block diagrams in the accompanying drawings show the possible architectures, functions and operations of the devices, methods and computer program products according to the multiple embodiments of the present application. In this regard, each box in the flowchart or block diagram can represent a module, a program segment or a part of the code, and the module, program segment or a part of the code contains one or more executable instructions for implementing the specified logical function. It should also be noted that in some alternative implementations, the functions marked in the box can also occur in an order different from that marked in the accompanying drawings. For example, two consecutive boxes can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flowchart, and the combination of boxes in the block diagram and / or flowchart, can be implemented with a dedicated hardware-based system that performs the specified function or action, or can be implemented with a combination of dedicated hardware and computer instructions.

[0096] In addition, the functional modules in each embodiment of the present application can be integrated together to form an independent part, or each module can exist independently, or two or more modules can be integrated to form an independent part.

[0097] If the functions are implemented in the form of software modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the portion that contributes to the prior art, or the portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, electronic device, or network device, etc.) to perform all or part of the steps of the methods described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program code, such as a USB flash drive, a mobile hard drive, a read-only memory (ROM), a random access memory (RAM), a magnetic disk, or an optical disk. It should be noted that, in this document, the terms "comprise," "include," or any other variant thereof are intended to cover non-exclusive inclusion, so that a process, method, article, or device that includes a series of elements includes not only those elements, but also other elements not explicitly listed, or also includes elements inherent to such process, method, article, or device. Without further constraints, an element defined by the phrase "comprises a..." does not preclude the existence of additional identical elements in the process, method, article or apparatus that includes the element.

[0098] The above description is merely a preferred embodiment of the present application and is not intended to limit the present application. Various modifications and variations are possible for those skilled in the art. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.

Claims

1. A speech recognition method applied in a noisy environment, characterized in that: include: Acquire target speech data in a target environment, and acquire standard noise data pre-configured for the target environment, wherein the target speech data belongs to speech data to be recognized, and the standard noise data includes at least one type of noise; Extracting speech data to be separated from the target speech data, which has correlation with the standard noise data at the level of latent semantic information; Performing denoising and separation processing on the target voice data based on the voice data to be separated to form denoised voice data; In the semantic encoding and decoding process of the target speech data, the denoised speech data is integrated to obtain target speech decoding features; Speech recognition is performed based on the target semantic decoding features, and a target speech recognition result is output.

2. The method for speech recognition in a noisy environment according to claim 1, wherein: The step of extracting the speech data to be separated that is correlated with the standard noise data at the level of latent semantic information from the target speech data comprises: Performing feature encoding on the target speech data and the standard noise data respectively to form target speech features corresponding to the target speech data and standard noise features corresponding to the standard noise data, wherein the target speech features and the standard noise features belong to latent semantic features of the same dimension, which belongs to the time domain or the frequency domain; Performing multi-level forward diffusion on the target speech feature and the standard noise feature respectively to form a plurality of forward diffusion feature combinations, wherein each of the forward diffusion feature combinations includes a speech forward diffusion feature corresponding to the target speech feature and a noise forward diffusion feature corresponding to the standard noise feature; respectively determining a feature correlation between the speech forward diffusion feature and the noise forward diffusion feature in each of the forward diffusion feature combinations; A forward diffusion feature combination having a maximum feature correlation is determined, and the speech forward diffusion features in the forward diffusion feature combination are subjected to at least one backward diffusion, and the obtained backward diffusion features are feature decoded to form speech data to be separated, wherein the number of the at least one backward diffusion is equal to the number of forward diffusions corresponding to the forward diffusion feature combination having the maximum feature correlation.

3. The method for speech recognition in a noisy environment according to claim 2, wherein: The step of performing multi-level forward diffusion on the target speech feature and the standard noise feature to form multiple forward diffusion feature combinations includes: Based on the standard noise feature, performing attention processing on the target speech feature to form a speech attention feature corresponding to the target speech feature; Multi-level forward diffusion is performed on the speech attention feature and the standard noise feature respectively to form multiple forward diffusion feature combinations.

4. The method for speech recognition in a noisy environment according to claim 2, wherein: The step of performing multi-level forward diffusion on the target speech feature and the standard noise feature to form multiple forward diffusion feature combinations includes: Based on the standard noise feature, performing attention processing on the target speech feature to form a speech attention feature corresponding to the target speech feature, and combining the speech attention feature and the standard noise feature as a forward diffusion feature; Performing at least one level of compression operation on the target speech feature and the standard noise feature respectively to achieve forward diffusion, thereby forming at least one compressed feature combination, wherein each compressed feature combination includes a compressed speech feature corresponding to the target speech feature and a compressed noise feature corresponding to the standard noise feature; For each of the compressed feature combinations, attention processing is performed on the compressed speech features included in the compressed feature combination based on the compressed noise features included in the compressed feature combination to form a speech attention feature corresponding to the compressed speech feature, and the speech attention feature and the compressed noise feature are used as a forward diffusion feature combination.

5. The method for speech recognition in a noisy environment according to any one of claims 1 to 4, characterized in that: The step of fusing the denoised speech data to obtain target speech decoding features during the semantic encoding and decoding process of the target speech data comprises: Performing feature encoding on the target speech data to form target speech coding features, and performing feature encoding on the denoised speech data to form denoised speech coding features; In the process of performing multiple levels of forward diffusion processing on the target speech coding features, the denoised speech coding features are gradually fused to form multiple levels of forward fused speech features, wherein the object of a subsequent forward diffusion processing includes the forward fused speech features of the previous level and the denoised speech coding features; In the process of performing multiple levels of backward diffusion processing on the multiple levels of fused speech features, the denoised speech coding features are gradually fused to form multiple levels of backward fused speech features, wherein the objects of the subsequent backward diffusion processing include the backward fused speech features of the previous level, the forward fused speech features of the corresponding level, and the denoised speech coding features, and the manner in which the denoised speech coding features are fused in the forward diffusion processing is different from the manner in which the denoised speech coding features are fused in the backward diffusion processing; Based on the backward fusion speech features of the last level, the target speech decoding features are determined.

6. The method for speech recognition in a noisy environment according to claim 5, characterized in that: The step of gradually fusing the denoised speech coding features to form multiple levels of forward fused speech features during the process of performing multiple levels of forward diffusion processing on the target speech coding features comprises: In the forward diffusion process of the first level, attention processing is performed on the target speech coding features based on the denoised speech coding features to form the forward fusion speech features of the first level; In the forward diffusion process of the i-th level, the forward fusion speech features of the i-1-th level and the denoised speech coding features are respectively compressed, and based on the compressed denoised speech coding features, the compressed forward fusion speech features are subjected to attention processing to form the forward fusion speech features of the i-th level, where i is greater than 1.

7. The method for speech recognition in a noisy environment according to claim 5, wherein: The step of gradually fusing the denoised speech coding features to form multiple levels of backward fused speech features during the process of performing multiple levels of backward diffusion processing on the multiple levels of fused speech features comprises: In the backward diffusion process of the first level, linear mapping and nonlinear activation operations are performed on the denoised speech coding features to form a target parameter distribution, and each parameter in the target parameter distribution is used as a weight coefficient of a parameter at a corresponding position in the fused speech features of the last level to update the fused speech features of the last level to form the backward fused speech features of the first level; During the forward diffusion process at the jth level, the backward fusion speech features at the j-1th level are expanded, and the expanded backward fusion speech features and the forward fusion speech features at the corresponding level are added to form a bidirectional fusion speech feature, and the denoising speech coding features are linearly mapped and nonlinearly activated to form a target parameter distribution, and each parameter in the target parameter distribution is used as a weight coefficient of the parameter at the corresponding position in the bidirectional fusion speech feature to update the bidirectional fusion speech feature to form a backward fusion speech feature at the jth level, where j is greater than 1.

8. A speech recognition device for use in a noisy environment, characterized in that: include: a data acquisition module, configured to acquire target speech data in a target environment and acquire standard noise data pre-configured for the target environment, wherein the target speech data belongs to speech data to be recognized and the standard noise data includes at least one type of noise; A data extraction module is used to extract the speech data to be separated that is correlated with the standard noise data at the level of latent semantic information from the target speech data; A speech denoising module, configured to perform denoising and separation processing on the target speech data based on the speech data to be separated to form denoised speech data; A speech encoding and decoding module, configured to fuse the denoised speech data during the semantic encoding and decoding process of the target speech data to obtain target speech decoding features; The speech recognition module is used to perform speech recognition based on the target semantic decoding features and output a target speech recognition result.

9. An electronic device, characterized in that: include: memory for storing computer programs; A processor connected to the memory is used to execute the computer program stored in the memory to implement the speech recognition method applied in a noisy environment as described in any one of claims 1 to 7.

10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when running, executes the method for speech recognition in a noisy environment according to any one of claims 1 to 7.

Citation Information

Patent Citations

  • Voice recognizing result processing method and device, computer device and storage medium

    CN110459224A

  • Speech feature processing method and device, electronic equipment and storage medium

    CN112735397A

  • Noise processing method and system based on artificial intelligence

    CN119694327A

  • Speech recognition apparatus, method and computer program product

    US20050010406A1

  • Speech recognition method and device based on artificial intelligence

    US20180330726A1