Speech enhancement method, device, equipment, medium and product

By introducing a combined model of multi-core inverted residual module, convolutional multi-focus attention module and grouped attention gate module into banking and financial instruments, the problem of low speech recognition accuracy of banking and financial instruments in complex acoustic environments is solved, and the speech quality is enhanced and the recognition rate is improved.

CN120748419APending Publication Date: 2025-10-03AGRICULTURAL BANK OF CHINA
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202510988942.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-17
Publication Date
2025-10-03

AI Technical Summary

Technical Problem

The voice recognition accuracy of banking and financial equipment in complex acoustic environments is low, resulting in degraded voice quality and an inability to meet business processing needs.

Method used

A speech enhancement method is adopted to obtain the speech information spectrum of the user input, and use the combined model of multi-core inverted residual module, convolutional multi-focus attention module and grouped attention gate module to perform speech feature extraction and noise suppression to achieve the quality enhancement of speech signal.

Benefits of technology

It significantly improves the accuracy of speech recognition, reduces model overhead, and adapts to the hardware resource limitations of banking and financial equipment.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120748419A_ABST
    Figure CN120748419A_ABST
Patent Text Reader

Abstract

The invention discloses a speech enhancement method, device and equipment, a medium and a product. The method comprises the following steps: acquiring a to-be-processed spectrogram corresponding to voice information input by a user; the spectrogram to be processed is input into a target speech enhancement model to obtain a target spectrogram, and the target speech enhancement model comprises a plurality of convolution multi-focus attention modules, a plurality of multi-kernel inverted residual modules and a plurality of grouping attention gate modules. The multi-core inverted residual module comprises a multi-core channel-by-channel convolution layer, a dimension reduction convolution layer, a plurality of batch standardization layers, a dimension raising convolution layer and a first activation layer; and converting the target spectrogram to obtain enhanced voice information, through the technical scheme of the invention, the quality of the input voice can be enhanced, and then the accuracy of voice recognition is significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present invention relate to the field of computer technology, and in particular to a speech enhancement method, apparatus, device, medium, and product. Background Art

[0002] With the continuous development of banking and financial equipment, voice interaction has become a key component in improving customer service experience and transaction efficiency. To provide customers with convenient operation, many banking and financial equipment are now equipped with built-in microphones, allowing customers to complete various operations through voice commands. However, the actual acoustic environment of banking operations is relatively complex. The voices of customers communicating with tellers and other customers' conversations during transactions can add up to a noisy environment, significantly reducing the quality of voice captured by the financial equipment's built-in microphones. This results in low voice recognition accuracy, making it unable to meet business processing requirements. Summary of the Invention

[0003] Embodiments of the present invention provide a speech enhancement method, apparatus, device, medium, and product to enhance the quality of input speech, thereby significantly improving the accuracy of speech recognition.

[0004] According to one aspect of the present invention, there is provided a method for speech enhancement, comprising:

[0005] Obtain the to-be-processed spectrogram corresponding to the voice information input by the user;

[0006] Inputting the to-be-processed spectrogram into a target speech enhancement model to obtain a target spectrogram, wherein the target speech enhancement model comprises: a plurality of convolutional multi-focus attention modules, a plurality of multi-core inverted residual modules and a plurality of grouped attention gate modules, the multi-core inverted residual module comprises: a multi-core channel-by-channel convolution layer, a dimensionality reduction convolution layer, a plurality of batch normalization layers, a dimensionality increase convolution layer and a first activation layer, the convolutional multi-focus attention module comprises: an adaptive average pooling layer, an adaptive maximum pooling layer, a dimensionality reduction convolution layer, a dimensionality increase convolution layer, a large kernel convolution layer, a first activation layer, a plurality of second activation layers, and two parallel channel pooling layers, and the grouped attention gate module comprises: a plurality of grouped convolution layers, a plurality of batch normalization layers, a first activation layer, a second activation layer and a first convolution layer;

[0007] The target spectrogram is converted to obtain enhanced speech information.

[0008] According to another aspect of the present invention, a speech enhancement device is provided, the speech enhancement device comprising:

[0009] An acquisition module is used to obtain a to-be-processed spectrogram corresponding to the voice information input by the user;

[0010] An enhancement module, configured to input the to-be-processed spectrogram into a target speech enhancement model to obtain a target spectrogram, wherein the target speech enhancement model comprises: a plurality of convolutional multi-focus attention modules, a plurality of multi-core inverted residual modules, and a plurality of grouped attention gate modules, the multi-core inverted residual module comprising: a multi-core channel-by-channel convolution layer, a dimensionality reduction convolution layer, a plurality of batch normalization layers, a dimensionality increase convolution layer, and a first activation layer, the convolutional multi-focus attention module comprising: an adaptive average pooling layer, an adaptive maximum pooling layer, a dimensionality reduction convolution layer, a dimensionality increase convolution layer, a large kernel convolution layer, a first activation layer, a plurality of second activation layers, and two parallel channel pooling layers, and the grouped attention gate module comprising: a plurality of grouped convolution layers, a plurality of batch normalization layers, a first activation layer, a second activation layer, and a first convolution layer;

[0011] The conversion module is used to convert the target spectrogram to obtain enhanced speech information.

[0012] According to another aspect of the present invention, an electronic device is provided, comprising:

[0013] at least one processor; and

[0014] a memory communicatively connected to the at least one processor; wherein,

[0015] The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the speech enhancement method described in any embodiment of the present invention.

[0016] According to another aspect of the present invention, a computer-readable storage medium is provided, wherein the computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the speech enhancement method according to any embodiment of the present invention when executed.

[0017] According to another aspect of the present invention, a computer program product is provided. When the computer program is executed by a processor, the computer program implements the speech enhancement method as described in any one of the embodiments of the present invention.

[0018] The embodiment of the present invention first obtains the to-be-processed spectrogram corresponding to the voice information input by the user; then inputs the to-be-processed spectrogram into the target speech enhancement model to obtain the target spectrogram, wherein the target speech enhancement model includes: multiple convolutional multi-focus attention modules, multiple multi-core inverted residual modules and multiple grouped attention gate modules, the multi-core inverted residual module includes: a multi-core channel-by-channel convolution layer, a dimensionality reduction convolution layer, multiple batch normalization layers, a dimensionality increase convolution layer and a first activation layer, the convolutional multi-focus attention module includes: an adaptive average pooling layer, an adaptive maximum pooling layer, a dimensionality reduction convolution layer, a dimensionality increase convolution layer, a large kernel convolution layer, a first activation layer, multiple second activation layers, two parallel channel pooling layers, the grouped attention gate module includes: multiple grouped convolution layers, multiple batch normalization layers, a first activation layer, a second activation layer and a first convolution layer; finally, the target spectrogram is converted to obtain the enhanced speech information. The introduction of a multi-core inverted residual module enables comprehensive and detailed extraction of speech features. The introduction of a convolutional multi-focus attention module highlights key channel and spatial features, precisely focusing on important areas in the speech signal and effectively establishing contextual interactions and spatial relationships between speech features. The introduction of a grouped attention gate module optimizes the flow of information within the network, strengthens the connections between features at different levels, and suppresses the expression of noise and irrelevant features. This enhances the quality of input speech while reducing model overhead and achieving a lightweight model.

[0019] It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present invention, nor is it intended to limit the scope of the present invention. Other features of the present invention will become readily understood through the following description. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present invention and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.

[0021] Figure 1 is a flow chart of a speech enhancement method according to an embodiment of the present invention;

[0022] Figure 2 is a structural diagram of a multi-core inverted residual module in an embodiment of the present invention;

[0023] Figure 3 is a schematic structural diagram of a multi-core channel-by-channel convolutional layer in an embodiment of the present invention;

[0024] Figure 4is a structural diagram of a convolutional multi-focus attention module in an embodiment of the present invention;

[0025] Figure 5 is a structural diagram of a multi-core inverted residual attention module in an embodiment of the present invention;

[0026] Figure 6 is a structural diagram of a group attention gate module in an embodiment of the present invention;

[0027] Figure 7 is a schematic structural diagram of a target speech enhancement model in an embodiment of the present invention;

[0028] Figure 8 is a structural diagram of a speech enhancement device according to an embodiment of the present invention;

[0029] Figure 9 It is a structural diagram of an electronic device in an embodiment of the present invention. DETAILED DESCRIPTION

[0030] In order to enable those skilled in the art to better understand the solutions of the present invention, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the embodiments described are only part of the embodiments of the present invention, not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by ordinary technicians in this field without making creative efforts should fall within the scope of protection of the present invention.

[0031] It should be noted that the terms "first", "second", etc. in the description and claims of the present invention and the above-mentioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the numbers used in this way can be interchanged where appropriate, so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein. In addition, the terms "including" and "having" and any variations thereof are intended to cover non-exclusive inclusions. For example, a process, method, system, product or device that includes a series of steps or units is not necessarily limited to those steps or units clearly listed, but may include other steps or units that are not clearly listed or inherent to these processes, methods, products or devices.

[0032] It is understandable that before using the technical solutions disclosed in the various embodiments of this disclosure, the type, scope of use, usage scenarios, etc. of the personal information involved in this disclosure should be informed to the user and the user's authorization should be obtained in an appropriate manner in accordance with relevant laws and regulations.

[0033] Example 1

[0034] Figure 1 This is a flow chart of a speech enhancement method provided by an embodiment of the present invention. This embodiment is applicable to speech enhancement. The method can be performed by a speech enhancement device in an embodiment of the present invention. The device can be implemented in software and / or hardware. Figure 1 As shown, the method specifically includes the following steps:

[0035] S110, obtaining a to-be-processed spectrogram corresponding to the voice information input by the user.

[0036] In this embodiment, a method for obtaining the to-be-processed spectrogram corresponding to the voice information input by the user may be: performing a short-time Fourier transform on the voice information input by the user to obtain the to-be-processed spectrogram corresponding to the voice information input by the user.

[0037] S120, inputting the to-be-processed spectrogram into a target speech enhancement model to obtain a target spectrogram.

[0038] In this embodiment, the target speech enhancement model includes: multiple convolutional multi-focus attention modules, multiple multi-core inverted residual modules and multiple grouped attention gate modules, the multi-core inverted residual module includes: a multi-core channel-by-channel convolution layer, a dimensionality reduction convolution layer, multiple batch normalization layers, a dimensionality increase convolution layer and a first activation layer, the convolutional multi-focus attention module includes: an adaptive average pooling layer, an adaptive maximum pooling layer, a dimensionality reduction convolution layer, a dimensionality increase convolution layer, a large kernel convolution layer, a first activation layer, multiple second activation layers, two parallel channel pooling layers, and the grouped attention gate module includes: multiple grouped convolution layers, multiple batch normalization layers, a first activation layer, a second activation layer and a first convolution layer.

[0039] In this embodiment, the activation function corresponding to the first activation layer may be a ReLU6 activation function, and the activation function corresponding to the second activation layer may be a Sigmoid activation function. The dimensionality reduction convolution layer may be a 1×1 convolution layer, the dimensionality increase convolution layer may be a 1×1 convolution layer, and the large kernel convolution layer may be a 7×7 convolution layer. The two parallel channel pooling layers may be a channel average pooling layer and a channel maximum pooling layer, respectively. The group convolution layer may be a 3×3 group convolution layer. The first convolution layer is a 1×1 convolution layer.

[0040] In this embodiment, the multi-core channel-by-channel convolution layer includes: multiple parallel first combination layers, each first combination layer includes: a channel-by-channel convolution layer, a batch normalization layer and a first activation layer from input to output, and after adding the features of the outputs of the first activation layers in each first combination layer, the channels are shuffled to obtain the output feature map of the multi-core channel-by-channel convolution layer, wherein the kernel sizes of the channel-by-channel convolution layers in each first combination layer are different.

[0041] In this embodiment, the target speech enhancement model may include: 5 convolutional multi-focus attention modules, 10 multi-core inverted residual modules, and 4 grouped attention gate modules. The target speech enhancement model may also include: multiple dimensionality-increasing convolutional layers, multiple dimensionality-reducing convolutional layers, multiple maximum pooling layers, and multiple linear interpolation layers, wherein the spatial maximum pooling layer is a 2×2 maximum pooling layer.

[0042] Optionally, the multi-core inverted residual module includes, from input to output, in sequence: a dimensionality reduction convolution layer, a batch normalization layer, a first activation layer, a multi-core channel-by-channel convolution layer, a dimensionality increase convolution layer, and a batch normalization layer. The input feature map of the multi-core inverted residual module and the output feature map of the batch normalization layer are feature added to obtain the output feature map of the multi-core inverted residual module. The multi-core channel-by-channel convolution layer includes: multiple parallel first combination layers, each first combination layer includes, from input to output, a channel-by-channel convolution layer, a batch normalization layer, and a first activation layer. After feature addition of the outputs of the first activation layers in each first combination layer, channel shuffling is performed to obtain the output feature map of the multi-core channel-by-channel convolution layer, wherein the kernel sizes of the channel-by-channel convolution layers in each first combination layer are different.

[0043] In this embodiment, the kernel sizes of the channel-by-channel convolutional layers in each first combination layer are 1×1, 3×3, and 5×5, respectively.

[0044] In a specific example, the structural diagram of the multi-core inverted residual module is as follows: Figure 2 As shown in Figure 1, the multi-core inverted residual module consists of a cascade of multi-core channel-by-channel convolutional layer, a 1×1 convolutional layer, a batch normalization layer, and the first activation layer.

[0045] The structural diagram of the multi-core channel-by-channel convolution layer is as follows Figure 3 As shown, the input feature map is first convolved using multiple channel-by-channel convolution layers with different kernel sizes, and then batch normalization and ReLU6 activation processing are performed in sequence. The results are then subjected to feature addition operations, and the channel shuffling operation is used to ensure the flow of information between channels, and finally the feature map is output. In the multi-core channel-by-channel convolution layer, the channel-by-channel convolution layer operation with a large kernel helps to capture the contextual information of a larger area, while the operation with a small kernel focuses more on local details. Therefore, the multi-core channel-by-channel convolution layer can learn input features of different scales. In the embodiment of the present invention, the kernel sizes of the channel-by-channel convolution layers in the multi-core channel-by-channel convolution layer are 1×1, 3×3, and 5×5, respectively. Therefore, the calculation formula for the output MKDC(x) of the multi-core channel-by-channel convolution layer is shown as follows:

[0046] MKDC(x)=CS(∑ k∈K DWCB(x)

[0047] DWCB k(x) = ReLU6(BN(DWC k (x)));

[0048] Among them, x is the input of the multi-core channel-by-channel convolution module, CS(·) represents the channel shuffle operation, BN(·) represents the batch normalization operation, and DWC k (·) represents a channel-by-channel convolution operation with a kernel size of k×k. In the embodiment of the present invention, K={1, 3, 5}.

[0049] The multi-core inverted residual module first uses 1×1 convolution to expand the number of channels of the input feature map (the value of the expansion factor in the embodiment of the present invention is 2), so that the subsequent network can capture features in more channel dimensions. Next, the multi-core inverted residual module uses batch normalization and ReLU6 activation function to enhance the training stability of the network and its ability to express complex features. Then, the multi-core channel-by-channel convolution layer in the multi-core inverted residual module captures contextual information of different scales and enhances the understanding of the input feature map by using multiple convolutions of different kernel sizes and other operations. Finally, the 1×1 convolution and batch normalization operations of the multi-core inverted residual module restore the number of channels of the feature map, and add the features to the input feature map before outputting it. Therefore, the calculation formula for the output MKIR(x) of the multi-core inverted residual module is shown as follows:

[0050] MKIR(x)=BN(PWC2(MKDC(ReLU6(BN(PWC1(x))))))+x;

[0051] PWC1(·) is a 1×1 convolution for dimensionality reduction, and PWC2(·) is a 1×1 convolution for dimensionality increase. They respectively represent two 1×1 convolution operations in the multi-core inverted residual module, and BN is a batch normalization operation.

[0052] Optionally, the convolutional multi-focus attention module includes, from input to output, two parallel second combination layers, a second activation layer, two parallel channel pooling layers, a large kernel convolution layer and a second activation layer. Each second combination layer includes, from input to output, an adaptive pooling layer, a dimensionality reduction convolution layer, a first activation layer and a dimensionality increase convolution layer. The adaptive pooling layers in each second combination layer are respectively: an adaptive average pooling layer and an adaptive maximum pooling layer. The output feature maps of the adaptive average pooling layer and the adaptive maximum pooling layer are added together to serve as the input feature map of the second activation layer. The output feature map of the second activation layer is dot-multiplied with the input feature map of the convolutional multi-focus attention module to serve as the input feature map of the two parallel channel pooling layers. The two parallel channel pooling layers are respectively channel average pooling layers and channel maximum pooling layers. The output feature maps of the two parallel channel pooling layers are connected in series to serve as the input feature map of the large kernel convolution layer. The output feature map of the second activation layer is dot-multiplied with the input feature map of the two parallel channel pooling layers to serve as the output feature map of the convolutional multi-focus attention module.

[0053] In this embodiment, the structural diagram of the convolutional multi-focus attention module is as follows: Figure 4 As shown. The convolutional multi-focus attention module enhances the important channel and spatial features of the input feature map respectively. The convolutional multi-focus attention module first performs adaptive maximum pooling and adaptive average pooling operations on the input feature map respectively. Adaptive maximum pooling can highlight the important features on each channel in the spectrogram, and the adaptive average pooling operation can retain the overall feature information of each channel. Then, the convolutional multi-focus attention module uses 1×1 convolution for dimensionality reduction, and then uses the activation function ReLU to introduce nonlinearity, and then uses 1×1 convolution for dimensionality increase. After adding the features of the output obtained by the above operations, the channel attention weight is generated by the Sigmoid activation function. Finally, the generated channel attention weight is multiplied by the input feature map to enhance the important channel features and suppress the unimportant channel features. Therefore, the calculation formula of the output result CA(x) after channel feature enhancement is shown as follows:

[0054]

[0055] Among them, AMP(·) represents the adaptive maximum pooling operation, and AAP(·) represents the adaptive average pooling operation.

[0056] Similar to obtaining channel attention weights, the convolutional multi-focus attention module continues to perform channel average pooling and channel maximum pooling on the input feature map after channel feature enhancement, and then performs a large kernel convolution operation (the convolution kernel size is 7×7 in the embodiment of the present invention) to capture a wider range of spatial contextual relationships. The spatial attention weights are then generated through the Sigmoid activation function. Finally, the generated spatial attention weights are dot-multiplied with the input feature map after channel feature enhancement to further enhance important spatial features. Therefore, the calculation formula for the output SA(x) of the convolutional multi-focus attention module is shown as follows:

[0057]

[0058] Among them, LKC(·) represents the large kernel convolution operation, Channel max (·) represents the channel maximum pooling operation, Channel avg (·) represents the channel-wise average pooling operation.

[0059] The convolutional multi-focus attention module can adaptively adjust the degree of attention to different channels and spaces of the input features, thereby effectively extracting key information in the input feature map.

[0060] In this embodiment, the convolutional multi-focus attention module and the multi-core inverted residual module can also be cascaded to form a multi-core inverted residual attention module. The structural diagram of the multi-core inverted residual attention module is as follows: Figure 5 shown.

[0061] The Multi-core Inverted Residual Attention Module combines the attention mechanism of the Convolutional Multi-focus Attention Module with the Multi-core Inverted Residual Module to conduct deeper mining and integration of the input feature maps. The Multi-core Inverted Residual Attention Module can learn the dependencies between channels and spaces, thereby generating more representative feature representations. The formula for calculating the output MKIRA(x) of the Multi-core Inverted Residual Attention Module is shown below:

[0062] MKIRA(x)=MKIR(CMFA(x)).

[0063] Optionally, the grouped attention gate module includes, from input to output direction, two parallel third combination layers, a first activation layer, a dimensionality-increasing convolution layer, a batch normalization layer, and a second activation layer. Each of the third combination layers includes a grouped convolution layer and a batch normalization layer. After feature addition of the output feature maps of the two parallel third combination layers, the input feature map of the first activation layer is obtained. After dot multiplication of the attention coefficient output by the second activation layer and the input feature map of the grouped attention gate module, the output feature map of the grouped attention gate module is obtained.

[0064] In this embodiment, the structural diagram of the group attention gate module is as follows: Figure 6 As shown. The grouped attention gate module first performs 3×3 grouped convolution operations on the input feature map and the gate signal for feature extraction, and adds and fuses the batch-normalized feature maps. Next, the grouped attention gate module performs ReLU function activation, 1×1 convolution, batch normalization, and Sigmoid function activation operations on the fused feature map to obtain the attention coefficient. Finally, the grouped attention gate module performs a dot product of the attention coefficient and the input feature map to enhance the features of important areas and suppress the features of unimportant areas before outputting them to the subsequent network. Therefore, the calculation formula for the output GAG(g,x) of the grouped attention gate module is shown as follows:

[0065]

[0066] Here, g represents the gating signal and GConv(·) represents the grouped convolution operation.

[0067] The grouped attention gate module uses the gating signal of high-resolution features to guide the information flow in low-resolution feature maps, enabling the network to more effectively integrate features at different levels, improve the distinguishability of features, suppress the interference of noise and irrelevant features, and improve the ability to understand and extract the overall information of input features.

[0068] Optionally, the multiple multi-core inverted residual modules are respectively: a first multi-core inverted residual module, a second multi-core inverted residual module, a third multi-core inverted residual module, a fourth multi-core inverted residual module, a fifth multi-core inverted residual module, a sixth multi-core inverted residual module, a seventh multi-core inverted residual module, an eighth multi-core inverted residual module, a ninth multi-core inverted residual module and a tenth multi-core inverted residual module, the multiple convolutional multi-focus attention modules are respectively: a first convolutional multi-focus attention module, a second convolutional multi-focus attention module, a third convolutional multi-focus attention module, a fourth convolutional multi-focus attention module and a fifth convolutional multi-focus attention module, the multiple group attention gate modules are respectively: a first group attention gate module, a second group attention gate module, a third group attention gate module and a fourth group attention gate module;

[0069] Inputting the to-be-processed spectrogram into a target speech enhancement model to obtain a target spectrogram comprises:

[0070] Inputting the to-be-processed spectrogram into a first multi-core inverted residual module, a dimension-raising convolutional layer, and a spatial maximum pooling layer in sequence to obtain a first feature map;

[0071] Inputting the first feature map into a second multi-core inverted residual module, a dimension-increasing convolutional layer, and a spatial maximum pooling layer in sequence to obtain a second feature map;

[0072] Inputting the second feature map into a third multi-core inverted residual module, a dimension-increasing convolutional layer, and a spatial maximum pooling layer in sequence to obtain a third feature map;

[0073] Inputting the third feature map into a fourth multi-core inverted residual module, a dimension-raising convolutional layer, and a spatial maximum pooling layer in sequence to obtain a fourth feature map;

[0074] Inputting the fourth feature map into a fifth multi-core inverted residual module, a dimension-raising convolutional layer, and a spatial maximum pooling layer in sequence to obtain a fifth feature map;

[0075] Input the fifth feature map into the first convolutional multi-focus attention module, the sixth multi-core inverted residual module, the dimension reduction convolution layer and the linear interpolation layer in sequence to obtain the sixth feature map;

[0076] Inputting the sixth feature map and the fourth feature map into a first group attention gate module to obtain a seventh feature map, and performing feature addition on the seventh feature map and the sixth feature map to obtain an eighth feature map;

[0077] Inputting the eighth feature map into the second convolutional multi-focus attention module, the seventh multi-core inverted residual module, the dimensionality reduction convolution layer and the linear interpolation layer in sequence to obtain a ninth feature map;

[0078] Inputting the ninth feature map and the third feature map into a second grouping attention gate module to obtain a tenth feature map, and performing feature addition on the tenth feature map and the ninth feature map to obtain an eleventh feature map;

[0079] Inputting the eleventh feature map into the third convolutional multi-focus attention module, the eighth multi-core inverted residual module, the dimensionality reduction convolution layer and the linear interpolation layer in sequence to obtain the eleventh feature map;

[0080] Inputting the eleventh feature map and the second feature map into a third grouping attention gate module to obtain a twelfth feature map, and performing feature addition on the twelfth feature map and the eleventh feature map to obtain a thirteenth feature map;

[0081] Inputting the thirteenth feature map into the fourth convolutional multi-focus attention module, the ninth multi-core inverted residual module, the dimensionality reduction convolution layer and the linear interpolation layer in sequence to obtain a fourteenth feature map;

[0082] Inputting the fourteenth feature map and the first feature map into a fourth grouping attention gate module to obtain a fifteenth feature map, and performing feature addition on the fifteenth feature map and the fourteenth feature map to obtain a sixteenth feature map;

[0083] The sixteenth feature map is sequentially input into the fifth convolutional multi-focus attention module, the tenth multi-core inverted residual module, the dimensionality reduction convolution layer and the linear interpolation layer to obtain the target spectrum map.

[0084] It should be noted that the target speech enhancement model is Figure 7 As shown, the target speech enhancement model can be a U-shaped network model based on lightweight attention convolution. In this embodiment, the input of the left 3×3 grouped convolution layer of the first grouped attention gate module is the sixth feature map, and the input of the right 3×3 grouped convolution layer is the fourth feature map. The output feature map of the second activation layer is dot-multiplied with the fourth feature map to serve as the output feature map of the first grouped attention gate module.

[0085] In the embodiment of the present invention, compared with the traditional U-type network, the U-type network based on lightweight attention convolution uses a cascade of multi-core inverted residual modules, 1×1 convolution and 2×2 maximum pooling layers to replace the traditional coding blocks. Therefore, in the encoding stage, the input I of the n+1th coding block (the n+1th coding block includes: the n+1th multi-core inverted residual module, the dimension-increasing convolution layer and the spatial maximum pooling layer) is n+1 The calculation formula is as follows:

[0086] I n+1 =MaxPool(PWC2(MKIR(I n ))),1<=n<=4;

[0087] Among them, MaxPool(·) represents the spatial maximum pooling operation.

[0088] In the decoding stage, the U-shaped network based on lightweight attention convolution uses a multi-core inverted residual attention module to replace the traditional decoding block, and uses a grouped attention gate module to process the high-resolution features from the encoding block and the low-resolution output features of the previous layer decoding block, and then uses them as the input to the next layer decoding block. Therefore, in the decoding stage, the input O of the nth decoding block (the nth decoding block includes: multi-core inverted residual module, convolution multi-focus attention module, grouped attention gate module, dimensionality reduction convolution layer and linear interpolation layer) is n The calculation formula is as follows:

[0089] O1=MaxPool(PWC2(MKIR(I5)));

[0090] O n+1 =GAG(I Lay-n ,Up(PWC1(MKIRA(O n ))))+Up(PWC1(MKIRA(O n ))),1<=n<=4,Lay=6;

[0091] Where Up(·) represents the bilinear interpolation operation.

[0092] Finally, the U-network based on lightweight attention convolution uses 1×1 convolution and bilinear interpolation operations to restore the feature map output by the decoding block to the enhanced speech spectrum features.

[0093] During encoding, the encoding layer of the U-shaped network, based on lightweight attention convolution, can simultaneously capture features at different scales, improving the ability to extract complex speech features. Furthermore, the multi-core inverted residual module employs a lightweight convolutional design, significantly reducing the number of parameters and computational complexity. During decoding, the multi-core inverted residual attention module, combined with a convolutional multi-focus attention mechanism, focuses on key features, enhancing feature discrimination. The grouped attention gate module, through a unique gating mechanism, optimizes information flow and strengthens the connection between features at different levels.

[0094] Optionally, the training process of the target speech enhancement model includes:

[0095] Acquire a training sample set, wherein the training sample set includes: speech information after adding noise and speech information before adding noise;

[0096] Inputting the spectrogram corresponding to the speech information after adding noise in the training sample set into the to-be-trained model to obtain a predicted spectrogram;

[0097] The parameters of the model to be trained are trained according to the difference between the predicted spectrogram and the spectrogram corresponding to the speech information before adding noise to obtain a target speech enhancement model.

[0098] In a specific example, a speech signal is a continuous analog signal that needs to be framed and sampled for digital signal processing and analysis. In this embodiment of the present invention, the speech in the training set (including the speech before and after noise addition, where the speech before noise addition is the clean speech) is framed and sampled to obtain clean speech and noisy speech. The clean speech is the real speech collected in a quiet environment, and the noisy speech is the speech after noise addition.

[0099] The embodiment of the present invention performs short-time Fourier transform on the framed and sampled speech to obtain a spectrogram of clean speech and a spectrogram of noisy speech.

[0100] Train and test the U-net based on lightweight attention convolution. The training steps are as follows:

[0101] Step 1: Frame, sample, and perform short-time Fourier transform on the clean speech and the noisy speech to obtain the spectrogram of the clean speech and the spectrogram of the noisy speech.

[0102] Step 2: Build a model to be trained, which can be a U-shaped network based on lightweight attention convolution.

[0103] Step 3: Network training: Input the spectrum features obtained in step 1 into the network model in step 2 to start neural network training.

[0104] The test steps are as follows:

[0105] Step 1: Collect noisy speech in real scenarios, extract its spectral features, and obtain a spectrogram;

[0106] Step 2: Input the spectrogram into the trained target speech enhancement model to obtain the enhanced spectrogram, and convert the enhanced spectrogram into speech;

[0107] Step 3: Calculate the perceived speech quality evaluation, short-term objective intelligibility, and segmented signal-to-noise ratio of the enhanced speech to evaluate the enhancement performance of the model.

[0108] In this embodiment, the core function of Perceptual Evaluation of Speech Quality (PESQ) is to assess the overall quality of the speech signal, especially for speech affected by coding, transmission, or noise. The core function of Short-Time Objective Intelligibility (STOI) is to assess the intelligibility of the speech signal (i.e., whether the human ear can understand the speech content), especially for speech in noisy, reverberant, or distorted environments. The core function of Segmental Signal-to-Noise Rati (SegSNR) is to assess the local degree of noise interference in the speech signal.

[0109] S130, converting the target spectrogram to obtain enhanced speech information.

[0110] In order to adapt to the complex noise environment at the end of banking and financial equipment and improve the enhancement effect of the model, the embodiment of the present invention introduces multi-core inverted residual and multi-core inverted residual attention modules in the encoding and decoding layers of the traditional U-shaped network. The multi-core inverted residual module introduced in the embodiment of the present invention uses convolution kernels of different sizes to capture speech features of different scales at the same time. Small convolution kernels focus on subtle local features of the speech signal, and large convolution kernels capture a wider range of contextual information, thereby achieving a comprehensive and detailed extraction of speech features. The multi-core inverted residual attention module further combines the convolutional multi-focus attention mechanism to highlight the features of key channels and spaces, accurately focus on important areas in the speech signal, and effectively establish contextual interactions and spatial relationships between speech features. In addition, the embodiment of the present invention also uses a grouped attention gate module to optimize the jump connection method of the traditional U-shaped network. The grouped attention gate module introduced in the embodiment of the present invention can optimize the flow of information in the network, strengthen the connection between features at different levels, and suppress the expression of noise and irrelevant features. The optimized skip connection scheme of the present invention allows the decoding block to obtain more effective input features during operation, deeply fuse features at different levels, effectively alleviate the problem of detailed information being ignored during decoding, and achieve efficient aggregation of semantic features at different scales.

[0111] In addition, considering the hardware resource limitations of banking and financial equipment, the embodiment of the present invention uses multi-core lightweight convolution to achieve lightweight models. During the encoding process, the multi-core inverted residual module introduced in the embodiment of the present invention uses channel-by-channel convolution, multi-size convolution kernel convolution, and channel shuffling to reduce the computational overhead of the model, thereby significantly reducing the number of model parameters. During the decoding process, the convolutional multi-focus attention module also uses channel dimensionality reduction operations to ensure that the main convolution calculation process is performed in a low dimension, thereby reducing the additional overhead brought by the attention module.

[0112] Example 2

[0113] Figure 8 This is a schematic diagram of the structure of a speech enhancement device provided by an embodiment of the present invention. This embodiment is applicable to speech enhancement. The device can be implemented in software and / or hardware. The device can be integrated into any device that provides speech enhancement function, such as Figure 8 As shown, the speech enhancement device specifically includes: an acquisition module 810, an enhancement module 820 and a conversion module 830.

[0114] The acquisition module is used to obtain the to-be-processed spectrogram corresponding to the voice information input by the user;

[0115] An enhancement module, configured to input the to-be-processed spectrogram into a target speech enhancement model to obtain a target spectrogram, wherein the target speech enhancement model comprises: a plurality of convolutional multi-focus attention modules, a plurality of multi-core inverted residual modules, and a plurality of grouped attention gate modules, the multi-core inverted residual module comprising: a multi-core channel-by-channel convolution layer, a dimensionality reduction convolution layer, a plurality of batch normalization layers, a dimensionality increase convolution layer, and a first activation layer, the convolutional multi-focus attention module comprising: an adaptive average pooling layer, an adaptive maximum pooling layer, a dimensionality reduction convolution layer, a dimensionality increase convolution layer, a large kernel convolution layer, a first activation layer, a plurality of second activation layers, and two parallel channel pooling layers, and the grouped attention gate module comprising: a plurality of grouped convolution layers, a plurality of batch normalization layers, a first activation layer, a second activation layer, and a first convolution layer;

[0116] The conversion module is used to convert the target spectrogram to obtain enhanced speech information.

[0117] The above-mentioned product can execute the method provided by any embodiment of the present invention, and has the corresponding functional modules and beneficial effects of the execution method.

[0118] Example 3

[0119] Figure 9 A schematic diagram of the structure of an electronic device 10 that can be used to implement an embodiment of the present invention is shown. The electronic device is intended to represent various forms of digital computers, such as laptop computers, desktop computers, workstations, personal digital assistants, servers, blade servers, mainframe computers, and other suitable computers. The electronic device can also represent various forms of mobile devices, such as personal digital assistants, cellular phones, smart phones, wearable devices (such as helmets, glasses, watches, etc.) and other similar computing devices. The components shown herein, their connections and relationships, and their functions are merely examples and are not intended to limit the implementation of the present invention described and / or claimed herein.

[0120] like Figure 9 As shown, the electronic device 10 includes at least one processor 11 and a memory, such as a read-only memory (ROM) 12, a random access memory (RAM) 13, etc., which is communicatively connected to the at least one processor 11. The memory stores a computer program that can be executed by the at least one processor. The processor 11 can perform various appropriate actions and processes according to the computer program stored in the read-only memory (ROM) 12 or the computer program loaded from the storage unit 18 into the random access memory (RAM) 13. Various programs and data required for the operation of the electronic device 10 can also be stored in the RAM 13. The processor 11, ROM 12, and RAM 13 are connected to each other via a bus 14. An input / output (I / O) interface 15 is also connected to the bus 14.

[0121] Multiple components in the electronic device 10 are connected to the I / O interface 15, including an input unit 16, such as a keyboard, a mouse, etc.; an output unit 17, such as various types of displays, speakers, etc.; a storage unit 18, such as a magnetic disk, an optical disk, etc.; and a communication unit 19, such as a network card, a modem, a wireless communication transceiver, etc. The communication unit 19 allows the electronic device 10 to exchange information / data with other devices via a computer network such as the Internet and / or various telecommunication networks.

[0122] The processor 11 can be any general-purpose and / or specialized processing component with processing and computing capabilities. Some examples of the processor 11 include, but are not limited to, a central processing unit (CPU), a graphics processing unit (GPU), various specialized artificial intelligence (AI) computing chips, various processors running machine learning model algorithms, a digital signal processor (DSP), and any appropriate processor, controller, microcontroller, etc. The processor 11 executes the various methods and processes described above, such as the speech enhancement method.

[0123] In some embodiments, the speech enhancement method may be implemented as a computer program tangibly embodied in a computer-readable storage medium, such as the storage unit 18. In some embodiments, part or all of the computer program may be loaded and / or installed on the electronic device 10 via the ROM 12 and / or the communication unit 19. When the computer program is loaded into the RAM 13 and executed by the processor 11, one or more steps of the speech enhancement method described above may be performed. Alternatively, in other embodiments, the processor 11 may be configured to perform the speech enhancement method in any other suitable manner (e.g., by means of firmware).

[0124] Various embodiments of the systems and techniques described herein can be implemented in digital electronic circuit systems, integrated circuit systems, field programmable gate arrays (FPGAs), application specific integrated circuits (ASICs), application specific standard products (ASSPs), system-on-chip systems (SOCs), programmable logic devices (CPLDs), computer hardware, firmware, software, and / or combinations thereof. These various embodiments can include being implemented in one or more computer programs that are executable and / or interpreted on a programmable system that includes at least one programmable processor, which can be a special purpose or general purpose programmable processor that can receive data and instructions from a storage system, at least one input device, and at least one output device, and transmit data and instructions to the storage system, the at least one input device, and the at least one output device.

[0125] Computer programs for implementing the methods of the present invention may be written in any combination of one or more programming languages. These computer programs may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when the computer program is executed by the processor, the functions / operations specified in the flowcharts and / or block diagrams are implemented. The computer program may be executed entirely on the machine, partially on the machine, as a stand-alone software package, partially on the machine and partially on a remote machine, or entirely on a remote machine or server.

[0126] In the context of the present invention, computer-readable storage media can be tangible media that can contain or store a computer program for use with an instruction execution system, device or equipment or used in combination with an instruction execution system, device or equipment. Computer-readable storage media can include but are not limited to electronic, magnetic, optical, electromagnetic, infrared or semiconductor systems, devices or equipment, or any suitable combination of the foregoing. Alternatively, computer-readable storage media can be machine-readable signal media. More specific examples of machine-readable storage media can include electrical connections based on one or more lines, portable computer disks, hard disks, random access memories (RAM), read-only memories (ROM), erasable programmable read-only memories (EPROM or flash memory), optical fibers, portable compact disk read-only memories (CD-ROM), optical storage devices, magnetic storage devices, or any suitable combination of the foregoing.

[0127] To provide interaction with a user, the systems and techniques described herein can be implemented on an electronic device having: a display device (e.g., a CRT (cathode ray tube) or LCD (liquid crystal display) monitor) for displaying information to the user; and a keyboard and pointing device (e.g., a mouse or trackball) through which the user can provide input to the electronic device. Other types of devices can also be used to provide interaction with the user; for example, the feedback provided to the user can be any form of sensory feedback (e.g., visual feedback, auditory feedback, or tactile feedback); and input from the user can be received in any form (including acoustic input, voice input, or tactile input).

[0128] The systems and techniques described herein can be implemented in a computing system that includes back-end components (e.g., as a data server), or a computing system that includes middleware components (e.g., an application server), or a computing system that includes front-end components (e.g., a user computer with a graphical user interface or web browser through which a user can interact with implementations of the systems and techniques described herein), or a computing system that includes any combination of such back-end components, middleware components, or front-end components. The components of the system can be interconnected by any form or medium of digital data communication (e.g., a communication network). Examples of communication networks include: a local area network (LAN), a wide area network (WAN), a blockchain network, and the Internet.

[0129] A computing system may include clients and servers. The clients and servers are typically remote from each other and typically interact via a communication network. This client-server relationship arises through computer programs running on the respective computers, creating a client-server relationship. The server may be a cloud server, also known as a cloud computing server or cloud host. This server is a hosting product within the cloud computing service ecosystem that addresses the management difficulties and limited scalability of traditional physical hosting and VPS services.

[0130] It should be understood that the various forms of the processes shown above can be used to reorder, add, or delete steps. For example, the steps described in the present invention can be performed in parallel, sequentially, or in a different order, as long as the desired results of the technical solution of the present invention can be achieved. This is not limited herein.

[0131] An embodiment of the present invention further provides a computer program product, including a computer program, which implements the speech enhancement method according to any embodiment of the present invention when executed by a processor.

[0132] The computer program product may be implemented by writing computer program code for performing the operations of the present invention in one or more programming languages, or a combination thereof, including object-oriented programming languages ​​such as Java, Smalltalk, C++, and conventional procedural programming languages ​​such as "C" or similar programming languages. The program code may be executed entirely on the user's computer, partially on the user's computer, as a stand-alone software package, partially on the user's computer and partially on a remote computer, or entirely on a remote computer or server. In cases involving a remote computer, the remote computer may be connected to the user's computer via any type of network, including a local area network (LAN) or a wide area network (WAN), or may be connected to an external computer (e.g., via the Internet using an Internet service provider).

[0133] The above specific embodiments do not limit the scope of protection of the present invention. Those skilled in the art will appreciate that various modifications, combinations, sub-combinations, and substitutions may be made based on design requirements and other factors. Any modifications, equivalent substitutions, and improvements made within the spirit and principles of the present invention are intended to be included within the scope of protection of the present invention.

Claims

1. A speech enhancement method, characterized in that: include: Obtain the to-be-processed spectrogram corresponding to the voice information input by the user; Inputting the to-be-processed spectrogram into a target speech enhancement model to obtain a target spectrogram, wherein the target speech enhancement model comprises: a plurality of convolutional multi-focus attention modules, a plurality of multi-core inverted residual modules and a plurality of grouped attention gate modules, the multi-core inverted residual module comprises: a multi-core channel-by-channel convolution layer, a dimensionality reduction convolution layer, a plurality of batch normalization layers, a dimensionality increase convolution layer and a first activation layer, the convolutional multi-focus attention module comprises: an adaptive average pooling layer, an adaptive maximum pooling layer, a dimensionality reduction convolution layer, a dimensionality increase convolution layer, a large kernel convolution layer, a first activation layer, a plurality of second activation layers, and two parallel channel pooling layers, and the grouped attention gate module comprises: a plurality of grouped convolution layers, a plurality of batch normalization layers, a first activation layer, a second activation layer and a first convolution layer; The target spectrogram is converted to obtain enhanced speech information.

2. The method according to claim 1, characterized in that The multi-core inverted residual module includes, from input to output, a dimensionality reduction convolution layer, a batch normalization layer, a first activation layer, a multi-core channel-by-channel convolution layer, a dimensionality increase convolution layer, and a batch normalization layer. The input feature map of the multi-core inverted residual module and the output feature map of the batch normalization layer are feature added to obtain the output feature map of the multi-core inverted residual module. The multi-core channel-by-channel convolution layer includes: multiple parallel first combination layers, each first combination layer includes, from input to output, a channel-by-channel convolution layer, a batch normalization layer, and a first activation layer. After feature addition of the output of the first activation layer in each first combination layer, channel shuffling is performed to obtain the output feature map of the multi-core channel-by-channel convolution layer, wherein the kernel size of the channel-by-channel convolution layer in each first combination layer is different.

3. The method according to claim 1, characterized in that The convolutional multi-focus attention module includes, from input to output, two parallel second combination layers, a second activation layer, two parallel channel pooling layers, a large kernel convolution layer and a second activation layer. Each second combination layer includes, from input to output, an adaptive pooling layer, a dimensionality reduction convolution layer, a first activation layer and a dimensionality increase convolution layer. The adaptive pooling layers in each second combination layer are respectively: an adaptive average pooling layer and an adaptive maximum pooling layer. The output feature maps of the adaptive average pooling layer and the adaptive maximum pooling layer are added as the input feature maps of the second activation layer. The output feature map of the second activation layer is dot-multiplied with the input feature map of the convolutional multi-focus attention module to serve as the input feature map of the two parallel channel pooling layers. The two parallel channel pooling layers are respectively the channel average pooling layer and the channel maximum pooling layer. The output feature maps of the two parallel channel pooling layers are connected in series to serve as the input feature map of the large kernel convolution layer. The output feature map of the second activation layer is dot-multiplied with the input feature map of the two parallel channel pooling layers to serve as the output feature map of the convolutional multi-focus attention module.

4. The method according to claim 1, wherein The grouped attention gate module includes, from input to output, two parallel third combination layers, a first activation layer, a dimensionality-increasing convolution layer, a batch normalization layer, and a second activation layer. Each of the third combination layers includes a grouped convolution layer and a batch normalization layer. The output feature maps of the two parallel third combination layers are added to obtain the input feature map of the first activation layer. The attention coefficient output by the second activation layer is dot-multiplied by the input feature map of the grouped attention gate module to obtain the output feature map of the grouped attention gate module.

5. The method according to claim 1, wherein The multiple multi-core inverted residual modules are respectively: a first multi-core inverted residual module, a second multi-core inverted residual module, a third multi-core inverted residual module, a fourth multi-core inverted residual module, a fifth multi-core inverted residual module, a sixth multi-core inverted residual module, a seventh multi-core inverted residual module, an eighth multi-core inverted residual module, a ninth multi-core inverted residual module and a tenth multi-core inverted residual module, the multiple convolutional multi-focus attention modules are respectively: a first convolutional multi-focus attention module, a second convolutional multi-focus attention module, a third convolutional multi-focus attention module, a fourth convolutional multi-focus attention module and a fifth convolutional multi-focus attention module, the multiple group attention gate modules are respectively: a first group attention gate module, a second group attention gate module, a third group attention gate module and a fourth group attention gate module; Inputting the to-be-processed spectrogram into a target speech enhancement model to obtain a target spectrogram comprises: Inputting the to-be-processed spectrogram into a first multi-core inverted residual module, a dimension-raising convolutional layer, and a maximum pooling layer in sequence to obtain a first feature map; Inputting the first feature map into a second multi-core inverted residual module, a dimension-increasing convolutional layer, and a maximum pooling layer in sequence to obtain a second feature map; Inputting the second feature map into a third multi-core inverted residual module, a dimension-increasing convolutional layer, and a maximum pooling layer in sequence to obtain a third feature map; Inputting the third feature map into a fourth multi-core inverted residual module, a dimension-raising convolutional layer, and a maximum pooling layer in sequence to obtain a fourth feature map; Inputting the fourth feature map into a fifth multi-core inverted residual module, a dimension-raising convolutional layer, and a maximum pooling layer in sequence to obtain a fifth feature map; Input the fifth feature map into the first convolutional multi-focus attention module, the sixth multi-core inverted residual module, the dimension reduction convolution layer and the linear interpolation layer in sequence to obtain the sixth feature map; Inputting the sixth feature map and the fourth feature map into a first group attention gate module to obtain a seventh feature map, and performing feature addition on the seventh feature map and the sixth feature map to obtain an eighth feature map; Inputting the eighth feature map into the second convolutional multi-focus attention module, the seventh multi-core inverted residual module, the dimensionality reduction convolution layer and the linear interpolation layer in sequence to obtain a ninth feature map; Inputting the ninth feature map and the third feature map into a second grouping attention gate module to obtain a tenth feature map, and performing feature addition on the tenth feature map and the ninth feature map to obtain an eleventh feature map; Inputting the eleventh feature map into the third convolutional multi-focus attention module, the eighth multi-core inverted residual module, the dimensionality reduction convolution layer and the linear interpolation layer in sequence to obtain the eleventh feature map; Inputting the eleventh feature map and the second feature map into a third grouping attention gate module to obtain a twelfth feature map, and performing feature addition on the twelfth feature map and the eleventh feature map to obtain a thirteenth feature map; Inputting the thirteenth feature map into the fourth convolutional multi-focus attention module, the ninth multi-core inverted residual module, the dimensionality reduction convolution layer and the linear interpolation layer in sequence to obtain a fourteenth feature map; Inputting the fourteenth feature map and the first feature map into a fourth grouping attention gate module to obtain a fifteenth feature map, and performing feature addition on the fifteenth feature map and the fourteenth feature map to obtain a sixteenth feature map; The sixteenth feature map is sequentially input into the fifth convolutional multi-focus attention module, the tenth multi-core inverted residual module, the dimensionality reduction convolution layer and the linear interpolation layer to obtain the target spectrum map.

6. The method according to claim 1, characterized in that The training process of the target speech enhancement model includes: Acquire a training sample set, wherein the training sample set includes: speech information after adding noise and speech information before adding noise; Inputting the spectrogram corresponding to the speech information after adding noise in the training sample set into the to-be-trained model to obtain a predicted spectrogram; The parameters of the model to be trained are trained according to the difference between the predicted spectrogram and the spectrogram corresponding to the speech information before adding noise to obtain a target speech enhancement model.

7. A speech enhancement device, characterized in that: include: An acquisition module is used to obtain a to-be-processed spectrogram corresponding to the voice information input by the user; An enhancement module, configured to input the to-be-processed spectrogram into a target speech enhancement model to obtain a target spectrogram, wherein the target speech enhancement model comprises: a plurality of convolutional multi-focus attention modules, a plurality of multi-core inverted residual modules, and a plurality of grouped attention gate modules, the multi-core inverted residual module comprising: a multi-core channel-by-channel convolution layer, a dimensionality reduction convolution layer, a plurality of batch normalization layers, a dimensionality increase convolution layer, and a first activation layer, the convolutional multi-focus attention module comprising: an adaptive average pooling layer, an adaptive maximum pooling layer, a dimensionality reduction convolution layer, a dimensionality increase convolution layer, a large kernel convolution layer, a first activation layer, a plurality of second activation layers, and two parallel channel pooling layers, and the grouped attention gate module comprising: a plurality of grouped convolution layers, a plurality of batch normalization layers, a first activation layer, a second activation layer, and a first convolution layer; The conversion module is used to convert the target spectrogram to obtain enhanced speech information.

8. An electronic device, characterized in that: The electronic device comprises: at least one processor; and a memory communicatively connected to the at least one processor; wherein, The memory stores a computer program that can be executed by the at least one processor. The computer program is executed by the at least one processor to enable the at least one processor to perform the speech enhancement method according to any one of claims 1 to 6.

9. A computer-readable storage medium, characterized in that The computer-readable storage medium stores computer instructions, and the computer instructions are used to enable a processor to implement the speech enhancement method according to any one of claims 1 to 6 when executed.

10. A computer program product, characterized in that The computer program product comprises a computer program, which, when executed by a processor, implements the speech enhancement method according to any one of claims 1 to 6.