A signal processing method, device and medium

CN122575401APending Publication Date: 2026-08-14SAIC GM WULING AUTOMOBILE CO LTD +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-05-13
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0007]本申请提供一种信号处理方法、设备与介质,有利于解决语音信号的处理效果与运算负荷难以协调的问题

Benefits of technology

根据所述T帧初始语音子帧,确定所述T帧初始语音子帧中每一帧对应的多个语音频率子带的频域特征;

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122575401A_ABST
    Figure CN122575401A_ABST
Patent Text Reader

Abstract

This application provides a signal processing method, apparatus, and medium, relating to the field of signal processing technology, which helps to solve the problem of balancing the processing effect and computational load of speech signals. The method includes: determining the frequency domain characteristics of multiple speech frequency sub-bands based on an initial speech signal; determining the frequency domain characteristics corresponding to the initial speech signal based on the frequency domain characteristics of the multiple speech frequency sub-bands; and determining the target speech signal based on the frequency domain characteristics corresponding to the initial speech signal and a preset speech signal processing method. A higher frequency point distribution density is set in the relatively lower frequency speech frequency sub-bands to effectively extract low-frequency speech information and ensure the processing effect of the speech signal; a lower frequency point distribution density is set in the relatively higher frequency speech frequency sub-bands to reduce the number of frequency points and the computational load in subsequent processing, thereby reducing the computational load of the apparatus and effectively balancing the processing effect of the speech signal and the computational load of the apparatus.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of signal processing technology, and in particular to a signal processing method, apparatus and medium. Background Technology

[0002] With the development of intelligent acoustic interaction, remote auditory communication, and wearable hearing aids, speech signal processing technology has been widely applied in various scenarios such as voice communication, voice interaction, and voice enhancement. In practical applications, using speech signal processing technology to process the initial speech signal can improve the quality of the speech signal, extract effective information from the speech signal, and improve the playback or recognition effect of the speech signal to meet the needs of various application scenarios.

[0003] When performing speech signal processing, it is usually necessary to perform time-frequency conversion on the initial speech signal, converting the initial speech signal from the time domain to the frequency domain, and determining the frequency domain characteristics of the initial speech signal so as to analyze and process the frequency characteristics of the speech signal in the future.

[0004] In related technologies, Fourier transform can be used to process the initial speech signal, extract feature values ​​of multiple frequency points evenly distributed at fixed intervals in the frequency domain, and determine the frequency domain features corresponding to the initial speech signal.

[0005] It is understandable that the spacing or density of frequency points directly affects the processing performance of the voice signal and the computational load on the equipment. A higher frequency point density leads to a larger total number of frequency points, increasing the computational load for subsequent processing, resulting in a heavier equipment load, longer processing time, and lower processing efficiency. Conversely, a lower frequency point density, while reducing the computational load, results in an insufficient number of frequency points, failing to capture sufficient effective information and thus leading to poor voice signal processing performance. Therefore, how to balance the processing performance of the voice signal with the computational load on the equipment has become a pressing technical problem to be solved.

[0006] It should be noted that the information disclosed in the background section of this application is intended only to enhance the understanding of the general background of this application, and should not be construed as an admission or in any way implying that the information constitutes prior art known to those skilled in the art. Summary of the Invention

[0007] This application provides a signal processing method, device, and medium that helps to solve the problem of difficulty in coordinating the processing effect and computational load of speech signals.

[0008] In a first aspect, embodiments of this application provide a signal processing method, including: Based on the initial speech signal, the frequency domain characteristics of multiple speech frequency sub-bands are determined. The frequency domain characteristics of any one of the speech frequency sub-bands include the feature values ​​of multiple frequency points. The multiple frequency points are evenly distributed within the speech frequency sub-band. Among the multiple speech frequency sub-bands, the speech frequency sub-bands with relatively lower frequencies have a relatively larger frequency point distribution density. Based on the frequency domain characteristics of the multiple speech frequency sub-bands, determine the frequency domain characteristics corresponding to the initial speech signal; The target speech signal is determined based on the frequency domain characteristics corresponding to the initial speech signal and the preset speech signal processing method.

[0009] In this embodiment, based on the initial speech signal, the frequency domain characteristics of multiple speech frequency sub-bands are determined. A larger frequency point distribution density is set in the speech frequency sub-band with relatively lower frequencies, which can effectively extract low-frequency speech information from the initial speech signal and, to a certain extent, guarantee or even improve the processing effect of the speech signal. A smaller frequency point distribution density is set in the speech frequency sub-band with relatively higher frequencies, which reduces the number of frequency points, thereby reducing the amount of computation in the subsequent preset speech signal processing process, reducing the computing load of the device, and effectively coordinating the processing effect of the speech signal and the computing load of the device.

[0010] In some possible implementations, determining the frequency domain characteristics of multiple speech frequency sub-bands based on the initial speech signal includes: According to the formula: ; Determine the frequency domain characteristics of multiple speech frequency sub-bands; in, The first audio subframe in the t-th frame is the first audio subframe in the t-th frame. One sampling point; The frequency domain features of the i-th speech frequency sub-band corresponding to the initial speech sub-frame of frame t; The frequency corresponding to the i-th speech frequency sub-band; The length of the wavelet kernel corresponding to the i-th speech frequency sub-band; Let be the wavelet kernel of the i-th speech frequency sub-band.

[0011] In the embodiments of this application, wavelet kernels of different lengths and parameters are used to match speech signals of different frequency ranges, thereby achieving independent filtering and feature extraction of different speech frequency sub-bands, which can improve the accuracy of frequency domain feature extraction.

[0012] In some possible implementations, the frequency domain features corresponding to the initial speech signal include feature values ​​at M1 frequency points, where M1 > 2. The step of determining the target speech signal based on the frequency domain features corresponding to the initial speech signal and a preset speech signal processing method includes: N target frequency points are determined from the M1 frequency points corresponding to the initial speech signal. For any one of the N target frequency points, the extended feature value of the target frequency point is determined based on the feature value of the target frequency point and the feature value of at least one frequency point adjacent to the target frequency point, where M1≥N≥1. The target speech signal is determined based on the extended feature values ​​of the N target frequency points and the feature values ​​of the M1-N remaining frequency points, wherein the remaining frequency points are the frequency points other than the N target frequency points among the M1 frequency points.

[0013] In this embodiment, a target frequency point is selected from the frequency domain features corresponding to the initial speech signal. The extended feature value of the target frequency point is determined by the feature values ​​of the target frequency point and its adjacent frequency points. This effectively supplements the frequency domain correlation information between adjacent frequency points, makes full use of the harmonic characteristics and frequency domain context information of the speech signal, and improves the detail performance and fidelity of the speech signal.

[0014] In some possible implementations, determining the extended feature value of the arbitrary target frequency point based on the feature value of the arbitrary target frequency point and the feature value of at least one frequency point adjacent to the arbitrary target frequency point includes: Based on the feature value of any target frequency point and the feature values ​​of k1 associated frequency points, the extended feature value of the target frequency point is determined. The associated frequency points are frequency points other than the target frequency point within a first preset window range. The first preset window range includes the target frequency point and at least one frequency point adjacent to the target frequency point, where k1≥1.

[0015] In this embodiment, k1 associated frequency points are determined through a first preset window range, and then the extended feature value of any target frequency point is determined. This can further expand the frequency domain receptive field corresponding to the target frequency point and capture frequency domain association information over a longer distance. This enriches the information content of the extended feature value and improves the robustness of speech signal processing in complex noise environments.

[0016] In some possible implementations, determining the extended feature value of the arbitrary target frequency point based on the feature value of the arbitrary target frequency point and the feature values ​​of k1 associated frequency points includes: Based on the feature value of any target frequency point, the feature values ​​of k1 associated frequency points, and the weight corresponding to each feature value, the extended feature value of the target frequency point is determined.

[0017] In this embodiment, the degree of information contribution of different feature values ​​to the extended feature values ​​can be differentiated according to the weight corresponding to each feature value, thereby improving the accuracy of the extended feature value expression and further enhancing the robustness of speech signal processing in complex noise environments.

[0018] In some possible implementations, determining the frequency domain characteristics of multiple speech frequency sub-bands based on the initial speech signal includes: Based on the initial speech signal, determine the initial speech subframe of frame T, where T≥2; Based on the initial speech subframe of the T-frame, determine the frequency domain characteristics of multiple speech frequency subbands corresponding to each frame in the initial speech subframe of the T-frame; The step of determining the frequency domain features corresponding to the initial speech signal based on the frequency domain features of the plurality of speech frequency sub-bands includes: Based on the frequency domain characteristics of multiple speech frequency sub-bands corresponding to each frame in the initial speech sub-frame of the T-frame, a first time-frequency matrix corresponding to the initial speech signal is determined, and the elements in the first time-frequency matrix are feature values.

[0019] In this embodiment of the application, by performing frame-by-frame processing on the initial speech signal and constructing a first time-frequency matrix, the temporal continuous information of the speech signal is introduced, which to a certain extent avoids the problems of temporal feature breakage and unnatural speech connection.

[0020] In some possible implementations, determining the first time-frequency matrix corresponding to the initial speech signal based on the frequency domain characteristics of multiple speech frequency sub-bands corresponding to each frame in the initial speech sub-frame of the T-frame includes: Based on the frequency domain energy of each frame in the initial speech subframe of the T-frame, the attention weight corresponding to each frame in the initial speech subframe of the T-frame is determined. Based on the frequency domain features of multiple speech frequency sub-bands corresponding to each frame in the initial speech sub-frame of the T-frame, and the attention weight corresponding to each frame, the first time-frequency matrix corresponding to the initial speech signal is determined.

[0021] In this embodiment, attention weights are matched to different time frames based on the energy distribution of the speech signal. In non-stationary noise scenarios, the focus is on time segments with significant speech energy, thus suppressing transient noise and further improving the quality of speech signal processing.

[0022] In some possible implementations, determining the target speech signal based on the frequency domain features corresponding to the initial speech signal and a preset speech signal processing method includes: The first time-frequency matrix corresponding to the initial speech signal is input into the encoder to determine the second time-frequency matrix corresponding to the initial speech signal. The elements in the second time-frequency matrix are semantic feature vectors, and the dimension of the semantic feature vectors is greater than the dimension of the feature values. The second time-frequency matrix corresponding to the initial speech signal is input into the decoder to determine the target speech signal.

[0023] In this embodiment of the application, by mapping low-dimensional feature values ​​to high-dimensional semantic feature vectors through an encoder, deep time-frequency correlation features and global semantic features in the initial speech signal can be extracted, thereby improving the processing effect and robustness of the speech signal.

[0024] In some possible implementations, for any initial speech subframe in the second time-frequency matrix corresponding to the initial speech signal, the second feature sequence corresponding to the arbitrary initial speech subframe is determined based on the first feature sequence corresponding to the arbitrary initial speech subframe and the first feature sequences corresponding to k2 associated speech subframes. The third time-frequency matrix corresponding to the initial speech signal is determined based on the second feature sequence corresponding to each initial speech subframe; The third time-frequency matrix corresponding to the initial speech signal is input into the decoder to determine the target speech signal; Wherein, the first feature sequence includes all elements corresponding to the initial speech subframe of any given frame in the second time-frequency matrix; the associated speech subframe is the frame other than the initial speech subframe of any given frame within the second preset window range, the second preset window range includes the initial speech subframe of any given frame and at least one initial speech subframe adjacent to the initial speech subframe of any given frame, k2≥1.

[0025] In this embodiment, by combining the first feature sequence corresponding to the associated speech subframe, the second feature sequence corresponding to any initial speech subframe is determined, and then a third time-frequency matrix is ​​constructed. This can fully utilize the temporal correlation information of the speech, supplement speech details, and to a certain extent suppress speech distortion caused by short-term burst noise and inter-frame jitter, thereby improving the processing effect of the speech signal.

[0026] In some possible implementations, before determining the second feature sequence corresponding to the arbitrary initial speech subframe, the method further includes: Based on the dimension of the semantic feature vector, the first feature sequence corresponding to each frame in the second time-frequency matrix is ​​decomposed into multiple sub-feature sequences; The step of determining the second feature sequence corresponding to the arbitrary initial speech subframe based on the first feature sequence corresponding to the arbitrary initial speech subframe and the first feature sequences corresponding to k2 associated speech subframes includes: Based on the multiple sub-feature sequences corresponding to any initial speech sub-frame and the multiple sub-feature sequences corresponding to k2 associated speech sub-frames, the second feature sequence corresponding to any initial speech sub-frame is determined.

[0027] In this embodiment of the application, the first feature sequence corresponding to each frame in the second time-frequency matrix is ​​decomposed into multiple sub-feature sequences and then time-series fusion is performed. This can significantly reduce the number of parameters and the amount of computation, reduce the computing load of the device, and better coordinate the relationship between the speech signal processing effect and the computing load of the device, so that the method can be better adapted to electronic devices with limited computing power.

[0028] Secondly, embodiments of this application provide an electronic device, including: processor; Memory; And a computer program, wherein the computer program is stored in the memory, the computer program including instructions that, when executed by the processor, cause the electronic device to perform the method described in any one of the first aspects.

[0029] Thirdly, embodiments of this application provide a computer-readable storage medium including a stored program, wherein, when the program is executed, it controls the device where the computer-readable storage medium is located to perform the method described in any one of the first aspects. Attached Figure Description

[0030] To more clearly illustrate the technical solutions of the embodiments of this application, the drawings used in the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0031] Figure 1 A schematic flowchart of a signal processing method provided in an embodiment of this application; Figure 2 A schematic diagram of a frequency point provided in an embodiment of this application; Figure 3 A schematic flowchart illustrating another signal processing method provided in an embodiment of this application; Figure 4 A flowchart illustrating another speech signal processing method provided in an embodiment of this application; Figure 5 A flowchart illustrating another speech signal processing method provided in an embodiment of this application; Figure 6 A flowchart illustrating another speech signal processing method provided in an embodiment of this application; Figure 7 A flowchart illustrating another speech signal processing method provided in an embodiment of this application; Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Detailed Implementation

[0032] To better understand the technical solution of this application, the embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0033] It should be understood that the described embodiments are merely some, not all, of the embodiments in this application. All other embodiments obtained by those skilled in the art based on the embodiments in this application without inventive effort are within the scope of protection of this application.

[0034] The terminology used in the embodiments of this application is for the purpose of describing particular embodiments only and is not intended to be limiting of this application. The singular forms “a,” “the,” and “the” used in the embodiments of this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0035] It should be understood that the term "and / or" used in this article is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A existing alone, A and B existing simultaneously, or B existing alone. Additionally, the character " / " in this article generally indicates that the preceding and following related objects have an "or" relationship.

[0036] With the development of intelligent acoustic interaction, remote auditory communication, and wearable hearing aids, speech signal processing technology has been widely applied in various scenarios such as voice communication, voice interaction, and voice enhancement. In practical applications, using speech signal processing technology to process the initial speech signal can improve the quality of the speech signal, extract effective information from the speech signal, and improve the playback or recognition effect of the speech signal to meet the needs of various application scenarios.

[0037] When performing speech signal processing, it is usually necessary to perform time-frequency conversion on the initial speech signal, converting the initial speech signal from the time domain to the frequency domain, and determining the frequency domain characteristics of the initial speech signal so as to analyze and process the frequency characteristics of the speech signal in the future.

[0038] In related technologies, Fourier transform can be used to process the initial speech signal, extract feature values ​​of multiple frequency points evenly distributed at fixed intervals in the frequency domain, and determine the frequency domain features corresponding to the initial speech signal.

[0039] It is understandable that the spacing or density of frequency points directly affects the processing performance of the voice signal and the computational load on the equipment. A higher frequency point density leads to a larger total number of frequency points, increasing the computational load for subsequent processing, resulting in a heavier equipment load, longer processing time, and lower processing efficiency. Conversely, a lower frequency point density, while reducing the computational load, results in an insufficient number of frequency points, failing to capture sufficient effective information and thus leading to poor voice signal processing performance. Therefore, how to balance the processing performance of the voice signal with the computational load on the equipment has become a pressing technical problem to be solved.

[0040] In view of this, embodiments of this application provide a signal processing method that helps to solve the problem of difficulty in coordinating the processing effect and computational load of speech signals. The embodiments of this application will be described in detail below with reference to the accompanying drawings.

[0041] See Figure 1 This is a flowchart illustrating a signal processing method provided in an embodiment of this application, as shown below. Figure 1 As shown, the method specifically includes the following steps.

[0042] S101: Determine the frequency domain characteristics of multiple speech frequency sub-bands based on the initial speech signal.

[0043] In this embodiment, multiple voice frequency sub-bands can be determined based on a preset voice frequency range. The preset voice frequency range is typically the frequency interval corresponding to the valid information in the initial voice signal. Those skilled in the art can set specific preset voice frequency ranges according to actual needs; however, this embodiment does not specifically limit the preset voice frequency range.

[0044] In some possible implementations, a preset speech frequency range is divided into multiple speech frequency sub-bands, each corresponding to the complete preset speech frequency range and not overlapping with the others. For example, the preset speech frequency range can be evenly divided into 5 speech frequency sub-bands, and these 5 sub-bands are not overlapping with each other.

[0045] Of course, those skilled in the art can set the size, number, and position of other multiple voice frequency sub-bands according to actual needs. In this embodiment, the size, number, and position of the multiple voice frequency sub-bands are not specifically limited.

[0046] It can be understood that the frequency corresponding to a speech frequency sub-band is used to characterize the position of that speech frequency sub-band within a preset speech frequency range. For example, the frequency corresponding to a speech frequency sub-band can be the midpoint frequency of that speech frequency sub-band.

[0047] Furthermore, the frequency domain features of a speech frequency sub-band include feature values ​​of multiple frequency points within the speech frequency sub-band. Specifically, the feature values ​​of a frequency point include at least one of amplitude, real part, and imaginary part. Within a single speech frequency sub-band, multiple frequency points are evenly distributed. Between multiple speech frequency sub-bands, speech frequency sub-bands with relatively lower frequencies have a relatively larger frequency point distribution density, while speech frequency sub-bands with relatively higher frequencies have a relatively smaller frequency point distribution density.

[0048] See Figure 2 This is a schematic diagram of a frequency point provided in an embodiment of this application, such as... Figure 2 As shown in Figure A, a preset speech frequency range can be evenly divided into 5 speech frequency sub-bands, and these 5 sub-bands do not overlap. Within a single speech frequency sub-band, multiple frequency points are evenly distributed. Between the multiple speech frequency sub-bands, the corresponding speech frequency sub-bands with relatively lower frequencies have a relatively larger frequency point distribution density, while the corresponding speech frequency sub-bands with relatively higher frequencies have a relatively smaller frequency point distribution density.

[0049] Based on the initial speech signal, the frequency domain features corresponding to each speech frequency sub-band are extracted.

[0050] In some possible implementations, frequency domain feature extraction is performed on multiple speech frequency sub-bands. Specifically, according to the formula: ; Determine the frequency domain characteristics of multiple speech frequency sub-bands; in, The first audio subframe in the t-th frame is the first audio subframe in the t-th frame. One sampling point; The frequency domain features of the i-th speech frequency sub-band corresponding to the initial speech sub-frame of frame t; The frequency corresponding to the i-th speech frequency sub-band; The length of the wavelet kernel corresponding to the i-th speech frequency sub-band; Let be the wavelet kernel of the i-th speech frequency sub-band.

[0051] It is understood that in the embodiments of this application, different speech frequency sub-bands may correspond to wavelet kernels of different lengths and parameters.

[0052] Specifically, the initial speech signal can be segmented into frames to obtain multiple initial speech subframes. For each initial speech subframe, its frequency domain features corresponding to each speech frequency sub-band are extracted. For any speech frequency sub-band, the time-domain sampling points of the initial speech subframe are weighted and summed with the wavelet kernel corresponding to that speech frequency sub-band to obtain the frequency domain features of the initial speech subframe corresponding to that speech frequency sub-band. This process further determines the frequency domain features of all speech frequency sub-bands corresponding to the initial speech subframe, ultimately achieving the extraction of frequency domain features from the initial speech subframe.

[0053] In the embodiments of this application, wavelet kernels of different lengths and parameters are used to match speech signals of different frequency ranges, thereby achieving independent filtering and feature extraction of different speech frequency sub-bands, which can improve the accuracy of frequency domain feature extraction.

[0054] S102: Determine the frequency domain characteristics of the initial speech signal based on the frequency domain characteristics of multiple speech frequency sub-bands.

[0055] It is understandable that multiple speech frequency sub-bands correspond to different frequency ranges. Based on the frequency domain characteristics of multiple speech frequency sub-bands, the frequency domain characteristics of the initial speech signal can be completely reconstructed. Typically, the frequency domain characteristics of multiple independent speech frequency sub-bands can be arranged and spliced ​​in an orderly manner to form the frequency domain characteristics corresponding to the initial speech signal.

[0056] Of course, those skilled in the art can also process the frequency domain features of multiple speech frequency sub-bands according to actual needs, and then determine the frequency domain features corresponding to the initial speech signal.

[0057] In some possible implementations, compressed frequency domain features of multiple speech frequency sub-bands are determined based on their frequency domain characteristics; and frequency domain features corresponding to the initial speech signal are determined based on these compressed frequency domain features. The number of frequency points in the compressed frequency domain features of the speech frequency sub-bands is less than the number of frequency points in the corresponding frequency domain features.

[0058] It is understood that those skilled in the art can, according to actual needs, select a portion of frequency points from the frequency domain features of the speech frequency sub-bands according to a preset filtering strategy to form corresponding compressed frequency domain features. They can also determine the corresponding compressed frequency domain features based on a preset fusion strategy and the frequency domain features of the speech frequency sub-bands. In the embodiments of this application, the frequency domain features of multiple speech frequency sub-bands can be further processed to determine compressed frequency domain features of multiple speech frequency sub-bands with fewer frequency points, further reducing the computational load.

[0059] S103: Determine the target speech signal based on the frequency domain characteristics corresponding to the initial speech signal and the preset speech signal processing method.

[0060] In this embodiment, the preset speech signal processing method can be used as a speech processing strategy to perform corresponding processing based on the frequency domain features corresponding to the initial speech signal. It can typically be used to achieve speech enhancement, noise suppression, frequency domain correction, etc. For example, the preset speech signal processing method may include at least one of frequency domain filtering, feature compensation, and interference removal. Those skilled in the art can select or combine different types of processing strategies according to the processing needs of actual business scenarios. This embodiment does not limit the specific type or execution logic of the preset speech signal processing method.

[0061] In this way, based on the frequency domain characteristics of the initial speech signal and the preset speech signal processing method, the target speech signal that meets the actual needs is determined and output.

[0062] It should be noted that the effective core information of a speech signal is often concentrated in a relatively low frequency range, while the effective speech information contained in the relatively high frequency range is relatively limited. This embodiment, through a differentiated frequency point layout, deploys frequency points with a higher distribution density within the low-frequency speech frequency sub-band, which can fully collect key low-frequency speech information and ensure the basic effect of subsequent speech signal processing. The high-frequency speech frequency sub-band is matched with a smaller frequency point distribution density, which can reasonably reduce the total number of frequency points and reduce the computational load in the frequency domain feature calculation and data processing process. In this embodiment, effective speech information can be well preserved, ensuring the processing quality of the target speech signal to a certain extent, while reducing invalid data, lowering the computational load during continuous device operation, improving the overall efficiency of the speech signal processing flow, and thus achieving a reasonable coordination between the speech signal processing effect and the device's computational load.

[0063] To address the challenge of balancing processing efficiency and computational load in speech signal processing, other methods can be used to perform time-frequency conversion on the initial speech signal. Accordingly, this application also provides another speech signal processing method. See [link to relevant documentation]. Figure 3 This is a flowchart illustrating a signal processing method provided in an embodiment of this application, as shown below. Figure 3 As shown, the method specifically includes the following steps.

[0064] S301: Determine the frequency domain characteristics corresponding to the initial speech signal based on the initial speech signal.

[0065] As mentioned above, the preset speech frequency range is typically the frequency interval corresponding to the effective information in the initial speech signal. Those skilled in the art can set specific preset speech frequency ranges according to actual needs; however, this application does not specifically limit the preset speech frequency range in its embodiments.

[0066] In this embodiment, M1 frequency points can be determined based on a preset speech frequency range, where M1 > 2. Within the preset speech frequency range, the interval between two adjacent frequency points is positively correlated with the frequency value. It can be understood that the interval between adjacent frequency points is negatively correlated with the frequency point distribution density; therefore, as the frequency value increases, the frequency point distribution density gradually decreases. Those skilled in the art can set the specific value of M1 and the correspondence between the frequency point interval and the frequency value according to actual needs; this embodiment does not impose specific limitations on this.

[0067] See Figure 2 This is a schematic diagram of a frequency point provided in an embodiment of this application, such as... Figure 2 As shown in B, M1 frequency points are continuously distributed within a preset speech frequency range. Within the preset speech frequency range, the interval between two adjacent frequency points is positively correlated with the frequency value. That is, the lower the frequency value, the smaller the interval between adjacent frequency points and the greater the local distribution density of the frequency points; the higher the frequency value, the larger the interval between adjacent frequency points and the smaller the local distribution density of the frequency points.

[0068] Based on the initial speech signal, feature values ​​corresponding to M1 frequency points are extracted. Specifically, the feature values ​​of each frequency point include at least one of amplitude, real part, and imaginary part. Based on the feature values ​​of the M1 frequency points, the frequency domain features corresponding to the initial speech signal are determined.

[0069] S103: Determine the target speech signal based on the frequency domain characteristics corresponding to the initial speech signal and the preset speech signal processing method.

[0070] For details on this step, please refer to the description above. For the sake of brevity, it will not be repeated here.

[0071] It should be noted that the effective core information of a speech signal is often concentrated in a relatively low frequency range, while the effective speech information contained in a relatively high frequency range is relatively limited.

[0072] In this embodiment, the interval between two adjacent frequency points is positively correlated with the frequency value. A larger number of frequency points are deployed in the lower frequency domain, which effectively extracts low-frequency speech information from the initial speech signal, ensuring or even improving the speech signal processing effect to a certain extent. Simultaneously, a smaller number of frequency points are deployed in the higher frequency domain, reducing the total number of frequency points and thus reducing the computational load in subsequent preset speech signal processing, lowering the device's computational load, and effectively coordinating the speech signal processing effect with the device's computational load.

[0073] In some possible implementations, M2 high-frequency points are determined from the M1 frequency points corresponding to the initial speech signal. Compressed high-frequency domain features are determined based on the feature values ​​of the M2 high-frequency points. The frequencies of the high-frequency points are greater than or equal to a preset frequency threshold. The compressed high-frequency domain features include the feature values ​​of M3 frequency points, where M1 > M2 ≥ 2 and M2 > M3 ≥ 1. Based on the compressed high-frequency domain features and the feature values ​​of the frequency points other than the M2 high-frequency points among the M1 frequency points, the splicing frequency domain features corresponding to the initial speech signal are determined. Finally, the target speech signal is determined based on the splicing frequency domain features corresponding to the initial speech signal and a preset speech signal processing method.

[0074] It is understood that a preset frequency threshold can be used to determine the boundary of a high-frequency region. Those skilled in the art can set specific preset frequency thresholds according to actual needs; however, the preset frequency thresholds are not specifically limited in the embodiments of this application.

[0075] Based on the feature values ​​of M2 high-frequency points in the high-frequency region and a preset compression fusion strategy, compressed high-frequency domain features are determined, including feature values ​​of M3 frequency points. Wherein, M1 > M2 ≥ 2, and M2 > M3 ≥ 1. Those skilled in the art can set the specific values ​​of M2 and M3 and the compression method of the high-frequency features according to actual needs; this application does not impose specific limitations on these settings.

[0076] Based on the compressed high-frequency domain features and the feature values ​​of the frequency points other than the M2 high-frequency points out of the M1 frequency points, the spliced ​​frequency domain features corresponding to the initial speech signal are determined. Typically, the original feature values ​​of the low-frequency region are sequentially spliced ​​with the compressed high-frequency domain features in ascending order of frequency to obtain the spliced ​​frequency domain features corresponding to the initial speech signal.

[0077] The target speech signal is determined based on the splicing frequency domain characteristics corresponding to the initial speech signal and the preset speech signal processing method.

[0078] In this embodiment, M2 high-frequency points are determined based on a preset frequency threshold, and compressed high-frequency domain features are determined based on the feature values ​​of the M2 high-frequency points. The frequency domain features with less effective information in the high-frequency region are compressed, which further reduces the amount of computation in the subsequent preset speech signal processing process. Under the premise of ensuring the speech processing effect, the computing load of the device is reduced.

[0079] In the signal processing section of related technologies, each frequency point is usually processed independently, ignoring the frequency domain correlation information between adjacent frequency points. This fails to fully utilize the harmonic characteristics and frequency domain context information of the speech signal, which can easily lead to problems such as loss of speech details and distortion, thereby affecting the processing effect of the speech signal.

[0080] To address the aforementioned technical problems, this application also provides a specific method for processing preset speech signals. In this embodiment, the frequency domain features corresponding to the initial speech signal include feature values ​​at M1 frequency points. See [link to relevant documentation]. Figure 4 This is a flowchart illustrating another speech signal processing method provided in an embodiment of this application, as shown below. Figure 4 As shown, step S103 specifically includes the following steps.

[0081] S401: Determine N target frequency points from the M1 frequency points corresponding to the initial speech signal. For any one of the N target frequency points, determine the extended feature value of any one target frequency point based on the feature value of the target frequency point and the feature value of at least one frequency point adjacent to the target frequency point, where M1≥N≥1.

[0082] In this embodiment, the target frequency point is the frequency point that needs to undergo feature expansion processing. Those skilled in the art can set the number, distribution, and specific filtering conditions of the target frequency points according to actual needs. For example, all M1 frequency points can be determined as target frequency points, or the remaining frequency points (excluding the first and last) among the M1 frequency points can be selected as target frequency points, or frequency points within a certain frequency range can be selected as target frequency points. This embodiment does not specifically limit the method of determining the target frequency point.

[0083] In an ordered sequence of frequency points corresponding to frequency domain features, the frequency points adjacent to the target frequency point typically include the preceding and following frequency points. At least one frequency point adjacent to any target frequency point can be the preceding or following frequency point, or it can include both the preceding and following frequency points.

[0084] In some possible implementations, based on the feature value of any target frequency point and the feature values ​​of k1 associated frequency points, the extended feature value of any target frequency point is determined. The associated frequency points are frequency points other than any target frequency point within a first preset window range. The first preset window range includes any target frequency point and at least one frequency point adjacent to any target frequency point, where k1≥1.

[0085] It is understood that by adjusting the size of the first preset window and the relative positional relationship between the target frequency point and the range of the first preset window, the associated frequency points corresponding to the target frequency point can be flexibly selected. For example, in an ordered sequence of frequency points corresponding to frequency domain features, the associated frequency points corresponding to the target frequency point may include the frequency point following the target frequency point and the frequency point immediately after that; it may also include the frequency point preceding the target frequency point, the frequency point following the target frequency point, and the frequency point immediately after that. Of course, those skilled in the art can set the size of the first preset window range and the relative positional relationship between the target frequency point and the first preset window range according to actual needs.

[0086] This allows for a further expansion of the receptive field in the frequency domain corresponding to the target frequency point, capturing frequency domain correlation information over longer distances, fully utilizing the long-range harmonic characteristics and frequency band coupling characteristics of speech signals, enriching and expanding the information content of feature values, and improving the robustness of speech signal processing in complex noise environments.

[0087] It should be noted that the dimension of the extended eigenvalue is usually the same as the dimension of the eigenvalue. For example, if the original eigenvalue is a three-dimensional vector containing amplitude, real part, and imaginary part, then the extended eigenvalue is also a three-dimensional vector of the same dimension; if the original eigenvalue is a one-dimensional eigenvalue obtained by fusing amplitude, real part, and imaginary part, then the extended eigenvalue is also a one-dimensional eigenvalue of the same dimension.

[0088] When determining the extended feature value, the feature value of the target frequency point can be concatenated with the feature value of at least one adjacent frequency point to determine the concatenation feature of the target frequency point; based on the concatenation feature of the target frequency point and the preset mapping processing strategy, the extended feature value corresponding to the target frequency point can be determined.

[0089] In some possible implementations, the extended feature value of any target frequency point is determined based on the feature value of any target frequency point, the feature values ​​of k1 associated frequency points, and the weight corresponding to each feature value.

[0090] In this embodiment, corresponding weights can be assigned to the feature values ​​of the target frequency point and the feature values ​​of k1 associated frequency points. The weight corresponding to each feature value can be used to characterize the contribution of the feature value in the extended feature value. Then, based on the feature value of any target frequency point, the feature values ​​of k1 associated frequency points, and the weight corresponding to each feature value, the extended feature value of any target frequency point is determined.

[0091] Those skilled in the art can set the weights corresponding to each feature value according to actual needs. In this application embodiment, the method of weight allocation is not specifically limited. For example, the weight corresponding to each feature value can be determined automatically by the neural network; or, for different associated frequency points, different weights can be set according to their distance from the target frequency point, with the weight of the associated frequency point being higher the closer it is to the target frequency point.

[0092] This allows for differentiated information contribution of different feature values ​​to extended feature values, improving the accuracy of extended feature value representation and further enhancing the robustness of speech signal processing in complex noise environments.

[0093] S402: Determine the target speech signal based on the extended feature values ​​of N target frequency points and the feature values ​​of M1-N remaining frequency points.

[0094] The remaining frequency points are the frequency points other than the N target frequency points out of the M1 frequency points. It should be noted that when N=M1, all frequency points are target frequency points, and the target speech signal is determined based on the extended feature values ​​of the M1 target frequency points (and the feature values ​​of the 0 remaining frequency points).

[0095] Understandably, based on frequency magnitude, the extended feature values ​​of N target frequency points can be systematically integrated with the original feature values ​​of M1-N remaining frequency points, and then inverse time-frequency transform processing can be performed to determine the target speech signal in the time domain. It should be noted that inverse time-frequency transform processing corresponds to time-frequency conversion methods.

[0096] In this embodiment, a target frequency point is selected from the frequency domain features corresponding to the initial speech signal. The extended feature value of the target frequency point is determined by the feature values ​​of the target frequency point and its adjacent frequency points. This effectively supplements the frequency domain correlation information between adjacent frequency points, makes full use of the harmonic characteristics and frequency domain context information of the speech signal, and improves the detail performance and fidelity of the speech signal.

[0097] Based on the concept of using multiple speech frequency sub-bands to perform time-frequency conversion on the initial speech signal, embodiments of this application also provide a speech signal processing method that performs frame-based processing of the initial speech signal, see [link to relevant documentation]. Figure 5 This is a flowchart illustrating another speech signal processing method provided in an embodiment of this application, as shown below. Figure 5 As shown, the method specifically includes the following steps.

[0098] S1011: Determine the initial speech subframe of frame T based on the initial speech signal. T≥2; The initial speech signal is a continuous time-domain signal. In actual processing, the initial speech signal is usually divided into frames according to a preset frame length and frame shift to obtain multiple initial speech subframes. Those skilled in the art can set the specific frame length, frame shift, and total number T of initial speech subframes according to actual needs; however, this embodiment does not impose specific limitations on these settings.

[0099] S1012: Based on the initial speech subframe of T frame, determine the frequency domain characteristics of multiple speech frequency subbands corresponding to each frame in the initial speech subframe of T frame.

[0100] For each initial speech subframe, frequency domain feature extraction processing of multiple speech frequency sub-bands is performed. Specific details regarding the embodiments of this application can be found above; for the sake of brevity, they will not be repeated here.

[0101] S1021: Determine the first time-frequency matrix corresponding to the initial speech signal based on the frequency domain characteristics of multiple speech frequency sub-bands corresponding to each frame in the initial speech sub-frame of T frame.

[0102] It can be understood that the frequency points corresponding to the initial speech subframe of each frame are arranged in ascending order of frequency to obtain the frequency domain feature sequence corresponding to a single frame. Then, the frequency domain feature sequences corresponding to the initial speech subframes of T frames are arranged in temporal order to construct the first time-frequency matrix corresponding to the initial speech signal. The elements in the first time-frequency matrix are the feature values ​​of the frequency points at the corresponding positions.

[0103] In some possible implementations, the attention weight corresponding to each frame in the initial speech subframe of the T-frame is determined based on the frequency domain energy of each frame in the initial speech subframe of the T-frame; and the first time-frequency matrix corresponding to the initial speech signal is determined based on the frequency domain characteristics of multiple speech frequency sub-bands corresponding to each frame in the initial speech subframe of the T-frame and the attention weight corresponding to each frame.

[0104] It is understood that the attention weight corresponding to each frame in the initial speech subframe of the T-frame can be determined based on the frequency domain energy of each frame, and then the first time-frequency matrix corresponding to the initial speech signal can be determined by combining the frequency domain features of multiple speech frequency sub-bands corresponding to each frame.

[0105] In some possible implementations, according to the formula: ; Determine the frequency domain energy of each initial speech subframe, where, Let X(t,f) be the frequency domain energy corresponding to the initial speech subframe of frame t, F be the total number of frequency points in multiple speech frequency sub-bands, and X(t,f) be the feature value of the f-th frequency point in the initial speech subframe of frame t. According to the formula: ; Determine the attention weights for each frame in the initial speech subframe of T-frame; in, Use the Sigmoid activation function; This is the first learnable parameter; This is the second learnable parameter; Let be the frequency domain energy corresponding to the initial speech subframe of frame t; represents the attention weight corresponding to the initial speech subframe of frame t.

[0106] Specifically, the attention weight corresponding to each frame is multiplied point-by-point by the frequency domain features of all speech frequency sub-bands in that frame to obtain a weighted single-frame frequency domain feature sequence. Then, the weighted frequency domain feature sequences of T frames are arranged in temporal order to construct the first time-frequency matrix corresponding to the initial speech signal.

[0107] In this way, attention weights can be matched to different time frames according to the energy distribution of the speech signal, focusing on time segments with significant speech energy in non-stationary noise scenarios, suppressing transient noise, and further improving the quality of speech signal processing.

[0108] S1031: Input the first time-frequency matrix corresponding to the initial speech signal into the encoder to determine the second time-frequency matrix corresponding to the initial speech signal.

[0109] In this embodiment, the encoder is used to perform feature upscaling and deep feature extraction on the input first time-frequency matrix. After processing by the encoder, the feature value at each position in the first time-frequency matrix is ​​mapped to a high-dimensional semantic feature vector, thereby obtaining the second time-frequency matrix corresponding to the initial speech signal.

[0110] It is understood that the elements in the second time-frequency matrix are semantic feature vectors, and the dimension of the semantic feature vectors is greater than the dimension of the feature values. For example, the elements in the first time-frequency matrix are one-dimensional scalar feature values ​​determined according to the frequency amplitude, real part and imaginary part of the corresponding position. After the encoder performs feature dimensionality upscaling and deep feature extraction, the elements in the determined second time-frequency matrix are 16-dimensional speech feature vectors.

[0111] Of course, those skilled in the art can set the network structure of the encoder and the specific dimensions of the semantic feature vector according to actual needs, and this application does not impose specific limitations on this.

[0112] S1032: Input the second time-frequency matrix corresponding to the initial speech signal into the decoder to determine the target speech signal.

[0113] It is understood that the decoder is used to perform feature dimensionality reduction and inverse time-frequency transformation on the input second time-frequency matrix, restoring the high-dimensional semantic feature vector to the time domain, and thus determining the target speech signal. Those skilled in the art can set the network structure of the decoder according to actual needs, and this application embodiment does not make specific limitations on this.

[0114] In some possible implementations, the decoder employs a reverse mapping structure symmetrical to the encoder, progressively recovering the frequency resolution of abstract features through multiple layers of transposed convolutions. At the decoder's output layer, sub-band gain masks corresponding to each speech frequency sub-band are predicted. These sub-band gain masks are then multiplied point-by-point with the frequency domain features of the speech frequency sub-bands corresponding to the initial speech signal to obtain the feature values ​​of the enhanced speech frequency sub-bands. Finally, an inverse wavelet transform is performed to project the feature values ​​of the multi-resolution speech frequency sub-bands back into the one-dimensional time domain, directly yielding the reconstructed target speech signal.

[0115] By using the above-mentioned decoding method, speech enhancement and reconstruction can be completed directly in the wavelet domain, avoiding information loss during time-frequency conversion and improving the reconstruction accuracy of the target speech signal.

[0116] In this embodiment, by performing frame-based processing on the initial speech signal and constructing a first time-frequency matrix, the temporal continuity information of the speech signal is introduced, which to some extent avoids the problems of temporal feature breaks and unnatural speech transitions. By mapping low-dimensional feature values ​​to high-dimensional semantic feature vectors through an encoder, deep time-frequency correlation features and global semantic features in the initial speech signal can be extracted, improving the processing effect and robustness of the speech signal.

[0117] In practical applications, the decoding of the first time-frequency matrix corresponding to the initial speech signal is usually performed independently on a per-frame basis. This approach easily overlooks the temporal correlation information of continuous speech, leading to problems such as loss of speech details and distortion, thus affecting the processing effect of the speech signal. See also Figure 6 This is a flowchart illustrating another speech signal processing method provided in an embodiment of this application, as shown below. Figure 6 As shown, step S1032 specifically includes the following steps.

[0118] S601: For any initial speech subframe in the second time-frequency matrix corresponding to the initial speech signal, determine the second feature sequence corresponding to the initial speech subframe based on the first feature sequence corresponding to the initial speech subframe and the first feature sequences corresponding to k2 associated speech subframes.

[0119] In this embodiment, the first feature sequence is an ordered set of all semantic feature vectors corresponding to a single initial speech subframe in the second time-frequency matrix. Based on the first feature sequence corresponding to any initial speech subframe and the first feature sequences corresponding to k2 associated speech subframes, the second feature sequence corresponding to any initial speech subframe is determined.

[0120] It is understood that by adjusting the size of the second preset window and the relative positional relationship between any initial speech subframe and the range of the second preset window, the associated speech subframe corresponding to any initial speech subframe can be flexibly selected. Those skilled in the art can set the size and coverage of the second preset window according to actual needs, and this application embodiment does not impose specific limitations on this.

[0121] The first feature sequence corresponding to the current frame is fused with the first feature sequences corresponding to k2 associated speech subframes to obtain the second feature sequence corresponding to the current frame.

[0122] In some possible implementations, based on the dimension of the semantic feature vector, the first feature sequence corresponding to each frame in the second time-frequency matrix is ​​decomposed into multiple sub-feature sequences. The second feature sequence corresponding to any initial speech sub-frame is determined according to the multiple sub-feature sequences corresponding to any initial speech sub-frame and the multiple sub-feature sequences corresponding to k2 associated speech sub-frames.

[0123] It is understandable that the elements in the first feature sequence are multi-dimensional semantic feature vectors. Based on the dimension of the semantic feature vector of the second time-frequency matrix, the first feature sequence corresponding to each frame in the second time-frequency matrix can be decomposed into multiple sub-feature sequences. For example, if the semantic feature vector has a 16-dimensional dimension, the first feature sequence corresponding to each frame in the second time-frequency matrix can be decomposed into 8 sub-feature sequences. The elements in the first sub-feature sequence are a two-dimensional feature vector composed of the first and second dimensions of the semantic feature vector; the elements in the second sub-feature sequence are a two-dimensional feature vector composed of the third and fourth dimensions of the semantic feature vector, and so on.

[0124] Those skilled in the art can set the number of sub-feature sequences according to actual needs, and no specific limitation is made in this embodiment.

[0125] Temporal feature fusion processing is performed on each sub-feature sequence. For any given sub-feature sequence, the sub-feature sequence corresponding to any initial speech sub-frame is extracted, and simultaneously, the sub-feature sequences of the same dimension corresponding to k2 associated speech sub-frames are extracted and fused to obtain the corresponding fused sub-feature sequence.

[0126] All fused sub-feature sequences are concatenated according to the order in which they were divided to obtain the second feature sequence corresponding to any initial speech sub-frame. The dimension of the concatenated second feature sequence is consistent with the dimension of the original first feature sequence.

[0127] Understandably, the computational complexity of determining the second feature sequence corresponding to any initial speech subframe, based on the first feature sequence corresponding to any initial speech subframe and the first feature sequences corresponding to k2 associated speech subframes, is typically O(C). 2 When the dimension C of the semantic feature vector is large, the computational load is large, making it difficult to deploy on resource-constrained devices.

[0128] In this embodiment, each sub-feature sequence processes features in only C / G dimensions, where G is the number of sub-feature sequences, and the computational complexity of a single sub-feature sequence is O((C / G)). 2 The overall computational complexity is reduced from O(C) to O(C). 2 ) decreases to O(G (C / G) 2 )=O(C 2 Therefore, when G takes a large value, the computational complexity can be significantly reduced, thereby reducing the computational load on the device.

[0129] In this embodiment of the application, the first feature sequence corresponding to each frame in the second time-frequency matrix is ​​decomposed into multiple sub-feature sequences and then time-series fusion is performed. This can significantly reduce the number of parameters and the amount of computation, reduce the computing load of the device, and better coordinate the relationship between the speech signal processing effect and the computing load of the device, so that the method can be better adapted to electronic devices with limited computing power.

[0130] S602: Determine the third time-frequency matrix corresponding to the initial speech signal based on the second feature sequence corresponding to each initial speech subframe.

[0131] The second feature sequences corresponding to all initial speech subframes are arranged in chronological order to construct the third time-frequency matrix corresponding to the initial speech signal. The dimension of the third time-frequency matrix is ​​consistent with that of the second time-frequency matrix.

[0132] S603: Input the third time-frequency matrix corresponding to the initial speech signal into the decoder to determine the target speech signal.

[0133] In this embodiment, by combining the first feature sequence corresponding to the associated speech subframe, the second feature sequence corresponding to any initial speech subframe is determined, and then a third time-frequency matrix is ​​constructed. This can fully utilize the temporal correlation information of the speech, supplement speech details, and to a certain extent suppress speech distortion caused by short-term burst noise and inter-frame jitter, thereby improving the processing effect of the speech signal.

[0134] Meanwhile, based on the concept of unevenly distributed frequency points, this application also provides another speech signal processing method for processing the initial speech signal in frames, see [link to relevant documentation]. Figure 7 This is a flowchart illustrating another speech signal processing method provided in an embodiment of this application, as shown below. Figure 7 As shown, the method specifically includes the following steps.

[0135] S701: Determine the initial speech subframe of frame T based on the initial speech signal, where T≥2.

[0136] For details on this step, please refer to the relevant description in S1011 above. For the sake of brevity, it will not be repeated here.

[0137] S702: Based on the initial speech subframe of T frame, determine the frequency domain features corresponding to each frame in the initial speech subframe of T frame.

[0138] The frequency domain features corresponding to the initial speech subframe include feature values ​​of M1 frequency points. The M1 frequency points are distributed within a preset speech frequency range. Within the preset speech frequency range, the interval between two adjacent frequency points is positively correlated with the frequency value, and M1 > 2.

[0139] For each initial speech subframe, frequency domain feature extraction is performed. For details of this step, please refer to the previous description in S301; for the sake of brevity, it will not be repeated here.

[0140] S703: Determine the fourth time-frequency matrix corresponding to the initial speech signal based on the frequency domain characteristics of each frame in the initial speech subframe of T frame.

[0141] It can be understood that the feature values ​​of all frequency points corresponding to the initial speech subframe of each frame are arranged in ascending order of frequency to obtain the frequency domain feature sequence corresponding to a single frame. Then, the frequency domain feature sequences corresponding to the initial speech subframes of T frames are arranged in temporal order to construct the fourth time-frequency matrix corresponding to the initial speech signal. The elements in the fourth time-frequency matrix are the feature values ​​of the frequency points at the corresponding positions.

[0142] In some possible implementations, the fourth time-frequency matrix corresponding to the initial speech signal is determined based on the frequency domain features corresponding to each frame in the initial speech subframe of the T-frame and the attention weight corresponding to each frame.

[0143] In some possible implementations, according to the formula: ; Determine the frequency domain energy of each initial speech subframe, where, Let X(t,f) be the frequency domain energy corresponding to the initial speech subframe of frame t, F be the total number of frequency points in multiple speech frequency sub-bands, and X(t,f) be the feature value of the f-th frequency point in the initial speech subframe of frame t. According to the formula: ; Determine the attention weights for each frame in the initial speech subframe of T-frame; in, Use the Sigmoid activation function; This is the first learnable parameter; This is the second learnable parameter; Let be the frequency domain energy corresponding to the initial speech subframe of frame t; represents the attention weight corresponding to the initial speech subframe of frame t.

[0144] Specifically, the attention weight corresponding to each frame is multiplied point-by-point with the frequency domain feature corresponding to each frame to obtain the weighted frequency domain feature of each frame. Then, the weighted frequency domain feature sequence of T frames is arranged in temporal order to construct the fourth time-frequency matrix corresponding to the initial speech signal.

[0145] In this way, attention weights can be matched to different time frames according to the energy distribution of the speech signal, focusing on time segments with significant speech energy in non-stationary noise scenarios, suppressing transient noise, and further improving the quality of speech signal processing.

[0146] S704: Input the fourth time-frequency matrix corresponding to the initial speech signal into the encoder to determine the fifth time-frequency matrix corresponding to the initial speech signal.

[0147] Among them, the elements in the fifth time-frequency matrix are semantic feature vectors, and the dimension of the semantic feature vectors is greater than the dimension of the eigenvalues.

[0148] In this embodiment, the encoder is used to perform feature upscaling and deep feature extraction on the input fourth time-frequency matrix. After processing by the encoder, the feature value at each position in the fourth time-frequency matrix is ​​mapped to a high-dimensional semantic feature vector, thereby obtaining the fifth time-frequency matrix corresponding to the initial speech signal.

[0149] It is understandable that the elements in the fifth time-frequency matrix are semantic feature vectors, and the dimension of the semantic feature vectors is greater than the dimension of the feature values. For example, the elements in the fourth time-frequency matrix are one-dimensional scalar feature values ​​determined according to the frequency amplitude, real part and imaginary part of the corresponding position. After the encoder performs feature dimensionality upscaling and deep feature extraction, the elements in the fifth time-frequency matrix are determined to be 16-dimensional speech feature vectors.

[0150] Of course, those skilled in the art can set the network structure of the encoder and the specific dimensions of the semantic feature vector according to actual needs, and this application does not impose specific limitations on this.

[0151] S705: Input the fifth time-frequency matrix corresponding to the initial speech signal into the decoder to determine the target speech signal.

[0152] Understandably, the decoder performs feature dimensionality reduction and inverse time-frequency transformation on the input fifth time-frequency matrix, restoring the high-dimensional semantic feature vector to the time domain, thereby determining the target speech signal. Those skilled in the art can configure the decoder's network structure according to actual needs; this application does not impose specific limitations on this aspect.

[0153] In this embodiment, by performing frame-based processing on the initial speech signal and constructing a fourth time-frequency matrix, the temporal continuity information of the speech signal is introduced, which to some extent avoids the problems of temporal feature breaks and unnatural speech transitions. By mapping low-dimensional feature values ​​to high-dimensional semantic feature vectors through an encoder, deep time-frequency correlation features and global semantic features in the initial speech signal can be extracted, improving the processing effect and robustness of the speech signal.

[0154] In practical applications, the decoding of the fourth time-frequency matrix corresponding to the initial speech signal is usually performed independently on a single initial speech subframe basis. This processing method easily ignores the temporal correlation information of continuous speech, which can easily lead to problems such as loss of speech details and distortion, thereby affecting the processing effect of the speech signal.

[0155] In some possible implementations, for any initial speech subframe in the fifth time-frequency matrix corresponding to the initial speech signal, the fourth feature sequence corresponding to the initial speech subframe is determined based on the third feature sequence corresponding to the initial speech subframe and the third feature sequences corresponding to k2 associated speech subframes. The sixth time-frequency matrix corresponding to the initial speech signal is determined based on the fourth feature sequence corresponding to each initial speech subframe; The sixth time-frequency matrix corresponding to the initial speech signal is input into the decoder to determine the target speech signal; The third feature sequence includes all elements corresponding to any initial speech subframe in the fifth time-frequency matrix; the associated speech subframe is a frame other than the initial speech subframe within a second preset window range, the second preset window range includes the initial speech subframe and at least one initial speech subframe adjacent to the initial speech subframe, k2≥1.

[0156] For details regarding the embodiments of this application, please refer to the description of steps S601-S603 above. For the sake of brevity, these details will not be repeated here.

[0157] In this embodiment, the fourth feature sequence corresponding to any initial speech subframe is determined by combining the third feature sequence corresponding to the associated speech subframe, thereby constructing the sixth time-frequency matrix. This fully utilizes the temporal correlation information of the speech, supplements speech details, and to a certain extent suppresses speech distortion caused by short-term burst noise and inter-frame jitter, improving the processing effect of the speech signal.

[0158] In some possible implementations, based on the dimension of the semantic feature vector, the third feature sequence corresponding to each frame in the fifth time-frequency matrix is ​​decomposed into multiple sub-feature sequences: the fourth feature sequence corresponding to any initial speech sub-frame is determined according to the multiple sub-feature sequences corresponding to any initial speech sub-frame and the multiple sub-feature sequences corresponding to k2 associated speech sub-frames.

[0159] For details regarding the embodiments of this application, please refer to the description of step S601 above. For the sake of brevity, it will not be repeated here.

[0160] In this embodiment of the application, the third feature sequence corresponding to each frame in the fifth time-frequency matrix is ​​decomposed into multiple sub-feature sequences and then fused in time sequence. This can significantly reduce the number of parameters and the amount of computation, reduce the computing load of the device, and better coordinate the relationship between the speech signal processing effect and the computing load of the device, so that the method can be better adapted to electronic devices with limited computing power.

[0161] It is understandable that encoder networks typically correspond to speech signal processing models. During the training of these models, relying solely on a single optimization objective often fails to simultaneously consider both subjective listening quality and objective acoustic metrics. For example, simply minimizing the temporal mean square error may numerically achieve a smaller reconstruction bias, but it often results in deficiencies in speech intelligibility and phase consistency. To address these issues, this application also provides a method for training a speech signal processing model.

[0162] In this embodiment, a multi-domain collaborative optimization objective function is used to train the encoder network or speech signal processing model. The multi-domain collaborative optimization objective function is used to impose constraints in the time domain, amplitude spectrum domain, and complex spectrum domain to control the reconstruction quality of the speech signal.

[0163] Specifically, the time-domain portion uses the scale-invariant signal-to-noise ratio (SNR) metric, which is defined as: ; Where s is the reference target speech signal; The target speech signal output by the model; This is the energy normalization coefficient. Those skilled in the art can adjust the calculation method of this index according to actual needs, and there are no specific limitations on this in the embodiments of this application.

[0164] In the amplitude spectrum domain, a weighted amplitude error term is introduced, which is defined as: ; in, The amplitude spectrum of the target speech signal is used as a reference. γ is the amplitude spectrum of the target speech signal output by the model; T is the total number of frames in the speech signal; and F is the total number of frequency points. Those skilled in the art can set the specific value of the amplitude compression index according to actual needs, but this application does not impose specific limitations on this.

[0165] Ultimately, the objective function for multi-domain collaborative optimization is defined as a weighted sum of multiple error terms, and it is defined as follows: ; in, , and These are the first, second, and third weighting coefficients, respectively. For the complex spectral domain error term; L SI SNR This refers to the time-domain error term. Those skilled in the art can adjust the values ​​of each weight coefficient according to the actual application scenario, and can also add or delete corresponding error terms according to actual needs. This application embodiment does not make specific limitations in this regard.

[0166] Training the speech signal processing model using the above method allows for simultaneous constraints on the speech signal reconstruction process in the time domain, amplitude spectrum domain, and complex spectrum domain, overcoming the limitations of a single optimization objective. This approach effectively improves the subjective listening quality of the speech signal while ensuring the objective accuracy of speech signal reconstruction, enabling the trained model to better adapt to different application scenarios.

[0167] In some possible implementations, a dynamic weight adaptation mechanism is introduced during the training phase to further improve the robustness of the model.

[0168] The first, second, and third weighting coefficients are adaptively generated by a lightweight parameter generation submodule based on the statistical characteristics of the current input initial speech signal. These statistical characteristics may include signal-to-noise ratio, spectral sparsity, etc. Those skilled in the art can configure the structure of the parameter generation submodule and the composition of the environment state vector according to actual needs; this embodiment does not impose specific limitations on these aspects.

[0169] By training the model using the above method, the weights of each optimization objective can be dynamically adjusted based on the actual environmental characteristics of the input speech signal. For example, under low signal-to-noise ratio conditions, the weights can be automatically increased. The weights are adjusted to enhance the model's noise reduction performance; under high signal-to-noise ratio conditions, the performance can be automatically improved. and The weights are adjusted to achieve higher speech quality. This approach enables the model to achieve better processing results in different noise environments, further improving the adaptability and robustness of the speech signal processing model.

[0170] Corresponding to the above embodiments, this application also provides an electronic device, including: a processor; a memory; and a computer program, wherein the computer program is stored in the memory, and when the computer program is executed by the processor, the electronic device performs any one of the methods described in the method embodiments.

[0171] Corresponding to the above embodiments, this application also provides an electronic device.

[0172] See Figure 8 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 8 As shown, the electronic device 800 may include a processor 801, a memory 802, and a communication unit 803. These components communicate via one or more buses. Those skilled in the art will understand that the electronic device structure shown in the figure does not constitute a limitation on the embodiments of this application. It may be a bus topology or a star topology, and may include more or fewer components than shown, or combine certain components, or have different component arrangements.

[0173] The communication unit 803 is used to establish a communication channel, thereby enabling the electronic device to communicate with other devices.

[0174] The processor 801 serves as the control center of the electronic device, connecting various parts of the device via various interfaces and lines. It executes software programs and / or modules stored in the memory 802, and calls data stored in the memory to perform various functions and / or process data. The processor can be composed of integrated circuits (ICs), such as a single packaged IC or multiple packaged ICs with the same or different functions connected together. For example, the processor 801 may consist only of a central processing unit (CPU). In this embodiment, the CPU may have a single processing core or include multiple processing cores.

[0175] Memory 802 is used to store the execution instructions of processor 801. Memory 802 can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk or optical disk.

[0176] When the execution instructions in memory 802 are executed by processor 801, the electronic device 800 is able to perform some or all of the steps in the above method embodiments. Corresponding to the above embodiments, this application also provides a computer-readable storage medium, wherein the computer-readable storage medium may store a program, and when the program runs, it can control the device where the computer-readable storage medium is located to perform some or all of the steps in the above method embodiments. Specifically, the computer-readable storage medium may be a magnetic disk, optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0177] For details regarding the embodiments of this application, please refer to the description of the above method embodiments. For the sake of brevity, these details will not be repeated here.

[0178] Corresponding to the above embodiments, this application also provides a computer-readable storage medium, wherein the computer-readable storage medium may store a program, and when the program runs, it can control the device where the computer-readable storage medium is located to execute some or all of the steps in the above method embodiments. In specific implementation, the computer-readable storage medium may be a magnetic disk, an optical disk, read-only memory (ROM), or random access memory (RAM), etc.

[0179] For details regarding the embodiments of this application, please refer to the description of the above method embodiments. For the sake of brevity, these details will not be repeated here.

[0180] In this application embodiment, "at least one" refers to one or more, and "more than one" refers to two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent the existence of A alone, the simultaneous existence of A and B, or the existence of B alone. A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" and similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, and c can represent: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.

[0181] Those skilled in the art will recognize that the units and algorithm steps described in the embodiments disclosed herein can be implemented using electronic hardware, computer software, or a combination of electronic hardware and software. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0182] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the above-described apparatus, controller, and computer storage medium can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.

[0183] In the several embodiments provided in this application, any function, if implemented as a software functional unit and sold or used as an independent product, can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0184] The above description is merely a specific embodiment of this application. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the protection scope of this application. The protection scope of this application shall be determined by the scope of the appended claims.

Claims

1. A signal processing method, characterized in that, include: Based on the initial speech signal, the frequency domain characteristics of multiple speech frequency sub-bands are determined. The frequency domain characteristics of any one of the speech frequency sub-bands include the feature values ​​of multiple frequency points. The multiple frequency points are evenly distributed within the speech frequency sub-band. Among the multiple speech frequency sub-bands, the speech frequency sub-bands with relatively lower frequencies have a relatively larger frequency point distribution density. Based on the frequency domain characteristics of the multiple speech frequency sub-bands, determine the frequency domain characteristics corresponding to the initial speech signal; The target speech signal is determined based on the frequency domain characteristics corresponding to the initial speech signal and the preset speech signal processing method.

2. The method according to claim 1, characterized in that, The step of determining the frequency domain characteristics of multiple speech frequency sub-bands based on the initial speech signal includes: According to the formula: ; Determine the frequency domain characteristics of multiple speech frequency sub-bands; in, The first audio subframe in the t-th frame is the first audio subframe in the t-th frame. One sampling point; The frequency domain features of the i-th speech frequency sub-band corresponding to the initial speech sub-frame of frame t; The frequency corresponding to the i-th speech frequency sub-band; The length of the wavelet kernel corresponding to the i-th speech frequency sub-band; Let be the wavelet kernel of the i-th speech frequency sub-band.

3. The method according to claim 1, characterized in that, The frequency domain features corresponding to the initial speech signal include feature values ​​at M1 frequency points, where M1 > 2. Determining the target speech signal based on the frequency domain features corresponding to the initial speech signal and a preset speech signal processing method includes: N target frequency points are determined from the M1 frequency points corresponding to the initial speech signal. For any one of the N target frequency points, the extended feature value of the target frequency point is determined based on the feature value of the target frequency point and the feature value of at least one frequency point adjacent to the target frequency point, where M1≥N≥1. The target speech signal is determined based on the extended feature values ​​of the N target frequency points and the feature values ​​of the M1-N remaining frequency points, wherein the remaining frequency points are the frequency points other than the N target frequency points among the M1 frequency points.

4. The method according to claim 3, characterized in that, The step of determining the extended feature value of any target frequency point based on the feature value of any target frequency point and the feature value of at least one frequency point adjacent to the target frequency point includes: Based on the feature value of any target frequency point and the feature values ​​of k1 associated frequency points, the extended feature value of the target frequency point is determined. The associated frequency points are frequency points other than the target frequency point within a first preset window range. The first preset window range includes the target frequency point and at least one frequency point adjacent to the target frequency point, where k1≥1.

5. The method according to claim 4, characterized in that, The step of determining the extended feature value of any target frequency point based on the feature value of any target frequency point and the feature values ​​of k1 associated frequency points includes: Based on the feature value of any target frequency point, the feature values ​​of k1 associated frequency points, and the weight corresponding to each feature value, the extended feature value of the target frequency point is determined.

6. The method according to claim 1, characterized in that, The step of determining the frequency domain characteristics of multiple speech frequency sub-bands based on the initial speech signal includes: Based on the initial speech signal, determine the initial speech subframe of frame T, where T≥2; Based on the initial speech subframe of the T-frame, determine the frequency domain characteristics of multiple speech frequency subbands corresponding to each frame in the initial speech subframe of the T-frame; The step of determining the frequency domain features corresponding to the initial speech signal based on the frequency domain features of the plurality of speech frequency sub-bands includes: Based on the frequency domain characteristics of multiple speech frequency sub-bands corresponding to each frame in the initial speech sub-frame of the T-frame, a first time-frequency matrix corresponding to the initial speech signal is determined, and the elements in the first time-frequency matrix are feature values.

7. The method according to claim 6, characterized in that, The step of determining the first time-frequency matrix corresponding to the initial speech signal based on the frequency domain characteristics of multiple speech frequency sub-bands corresponding to each frame in the initial speech sub-frame of the T-frame includes: Based on the frequency domain energy of each frame in the initial speech subframe of the T-frame, the attention weight corresponding to each frame in the initial speech subframe of the T-frame is determined. Based on the frequency domain features of multiple speech frequency sub-bands corresponding to each frame in the initial speech sub-frame of the T-frame, and the attention weight corresponding to each frame, the first time-frequency matrix corresponding to the initial speech signal is determined.

8. The method according to claim 6, characterized in that, The step of determining the target speech signal based on the frequency domain features corresponding to the initial speech signal and a preset speech signal processing method includes: The first time-frequency matrix corresponding to the initial speech signal is input into the encoder to determine the second time-frequency matrix corresponding to the initial speech signal. The elements in the second time-frequency matrix are semantic feature vectors, and the dimension of the semantic feature vectors is greater than the dimension of the feature values. The second time-frequency matrix corresponding to the initial speech signal is input into the decoder to determine the target speech signal.

9. The method according to claim 8, characterized in that, The step of inputting the second time-frequency matrix corresponding to the initial speech signal into the decoder to determine the target speech signal includes: For any initial speech subframe in the second time-frequency matrix corresponding to the initial speech signal, the second feature sequence corresponding to the initial speech subframe is determined based on the first feature sequence corresponding to the initial speech subframe and the first feature sequences corresponding to k2 associated speech subframes. The third time-frequency matrix corresponding to the initial speech signal is determined based on the second feature sequence corresponding to each initial speech subframe; The third time-frequency matrix corresponding to the initial speech signal is input into the decoder to determine the target speech signal; Wherein, the first feature sequence includes all elements corresponding to the initial speech subframe of any given frame in the second time-frequency matrix; the associated speech subframe is the frame other than the initial speech subframe of any given frame within the second preset window range, the second preset window range includes the initial speech subframe of any given frame and at least one initial speech subframe adjacent to the initial speech subframe of any given frame, k2≥1.

10. The method according to claim 9, characterized in that, Before determining the second feature sequence corresponding to any initial speech subframe, the method further includes: Based on the dimension of the semantic feature vector, the first feature sequence corresponding to each frame in the second time-frequency matrix is ​​decomposed into multiple sub-feature sequences; The step of determining the second feature sequence corresponding to the arbitrary initial speech subframe based on the first feature sequence corresponding to the arbitrary initial speech subframe and the first feature sequences corresponding to k2 associated speech subframes includes: Based on the multiple sub-feature sequences corresponding to any initial speech sub-frame and the multiple sub-feature sequences corresponding to k2 associated speech sub-frames, the second feature sequence corresponding to any initial speech sub-frame is determined.

11. An electronic device, characterized in that, include: processor; Memory; And a computer program, wherein the computer program is stored in the memory, the computer program including instructions that, when executed by the processor, cause the electronic device to perform the method of any one of claims 1 to 10.

12. A computer-readable storage medium, characterized in that, The computer-readable storage medium includes a stored program, wherein, when the program is executed, it controls the device on which the computer-readable storage medium is located to perform the method according to any one of claims 1 to 10.