Voice signal processing method, device, equipment and medium
By performing convolutional mixing and multi-layer perception processing on the voice signal, the number of parameters of the speech recognition model is reduced, and the problems of slow response speed and poor user experience in the prior art are solved, thereby achieving faster speech recognition response and better user experience.
Patent Information
- Application Number
- CN202210560595.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2022-05-20
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2042-05-20
AI Technical Summary
The existing voice wake-up technology has the problems of large model parameters, the need to be configured in the cloud, resulting in slow response speed and poor user experience.
By extracting speech features from the speech signal, performing convolution mixing processing and multi-layer perception-based mixing processing, the amount of parameters of the speech recognition model can be reduced, so that it can be configured on the terminal device.
It realizes faster voice recognition response, improves user experience, and reduces the amount of parameters of the voice recognition model, so that it can be better configured on terminal devices.
Smart Images

Figure CN115035887B_ABST
Abstract
Description
Technical Field
[0001] The present disclosure generally relates to the field of natural language processing, and specifically to a method, device, equipment and medium for processing a speech signal. Background Art
[0002] With the development of hardware technologies such as artificial intelligence algorithms and AI chips, smart devices have been widely used in daily life. Such as smart home voice control systems, smart speakers, smart conference systems, etc. Voice interaction is widely used in smart devices and is becoming increasingly mature. In traditional voice interaction scenarios, the device can only be interacted with by clicking a button, such as pressing the recording button, to wake up the device. In order to further improve the human-computer interaction experience, voice wake-up technology came into being.
[0003] There are currently three main methods of voice wake-up: wake-up technology based on template matching; wake-up technology based on hidden Markov models; and wake-up technology based on deep learning. Among them, the most widely used is the voice recognition wake-up method based on deep learning. However, the relevant wake-up technologies all have a large number of model parameters and need to be configured in the cloud for calculation, resulting in slow wake-up response speed and poor user experience. Summary of the invention
[0004] In view of the above-mentioned defects or deficiencies in the prior art, it is desirable to provide a method, apparatus, device and medium for processing a voice signal, which can provide a faster voice recognition response and improve the user experience.
[0005] In a first aspect, an embodiment of the present application provides a method for processing a speech signal, comprising:
[0006] Acquire speech signals collected from the environment;
[0007] Extracting speech features from the speech signal to obtain speech features corresponding to the speech signal;
[0008] Performing convolution mixing processing on the speech features to obtain shallow speech recognition features;
[0009] Performing a hybrid process based on multi-layer perception on the shallow speech recognition features to obtain deep speech recognition features;
[0010] Obtaining a recognition result of the speech signal according to the deep speech recognition feature;
[0011] According to the recognition result, a response strategy corresponding to the recognition result is executed.
[0012] In some embodiments, the performing convolution mixing processing on the speech features to obtain shallow speech recognition features includes:
[0013] The speech features are subjected to convolution mixing processing using a convolution mixing model.
[0014] Among them, the convolution mixing model includes a spatial position convolution mixing module and a channel position convolution mixing module, and the mixing result of the spatial position convolution mixing module and the input of the spatial position convolution mixing module are input into the channel position convolution mixing module through a residual connection.
[0015] In some embodiments, the spatial position convolution mixing module includes a depth-separable convolution layer, a first excitation function layer and a first normalization layer, and the channel position convolution mixing module includes a point-by-point convolution layer, a second excitation function layer and a second normalization layer.
[0016] In some embodiments, the shallow speech recognition features are subjected to hybrid processing based on multi-layer perception to obtain deep speech recognition features, including:
[0017] A multi-layer perceptual hybrid model is used to perform convolutional hybrid processing on the shallow speech recognition features.
[0018] Among them, the multi-layer perceptual mixing model includes a space-perceptual mixing module and a channel-perceptual mixing module. The space-perceptual mixing module performs space-perceptual mixing on the transposed feature information and transposes the mixing result again and inputs it into the channel-perceptual mixing module for channel-perceptual mixing.
[0019] In some embodiments, the spatial-aware mixing module includes a first fully connected layer, a third activation function layer, and a second fully connected layer, and the channel-aware mixing module includes a third fully connected layer, a fourth activation function layer, and a fourth fully connected layer.
[0020] In some embodiments, before performing convolution mixing processing on the speech features to obtain shallow speech recognition features, the method further includes:
[0021] Downsampling the speech features using a feature embedding module;
[0022] Among them, the feature embedding module includes a feature embedding layer, a fifth activation function layer and a third normalization layer.
[0023] In some embodiments, extracting speech features from the speech signal to obtain speech features corresponding to the speech signal includes:
[0024] Mel-frequency cepstral coefficients are used to extract speech features from the speech signal to obtain speech features corresponding to the speech signal.
[0025] In a second aspect, an embodiment of the present application provides a speech signal processing device, including:
[0026] An acquisition module, used to acquire the collected voice signal;
[0027] A speech feature extraction module, used to extract speech features from the speech signal to obtain speech features corresponding to the speech signal;
[0028] A convolutional mixing model is used to perform convolutional mixing processing on the speech features to obtain shallow speech recognition features;
[0029] A multi-layer perception hybrid model is used to perform a hybrid process based on multi-layer perception on the shallow speech recognition features to obtain deep speech recognition features;
[0030] A classification module, used to obtain a recognition result of the speech signal according to the deep speech recognition feature;
[0031] An execution module is used to execute a response strategy corresponding to the identification result according to the identification result.
[0032] In a third aspect, an embodiment of the present application provides an electronic device, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor implements the method described in the embodiment of the present application when executing the program.
[0033] In a fourth aspect, an embodiment of the present application provides a computer-readable storage medium having a computer program stored thereon, which, when executed by a processor, implements the method described in the embodiment of the present application.
[0034] The voice signal processing method proposed in the embodiment of the present application can reduce the number of parameters of the voice recognition model while effectively recognizing voice commands by performing convolution mixing operations and multi-perception-based mixing processing on the voice features extracted from the voice signal, so that the processing device for executing the voice signal can be better configured on the terminal device, thereby improving the response efficiency of the terminal device to voice commands.
[0035] Additional aspects and advantages of the present invention will be given in part in the following description and in part will be obvious from the following description, or will be learned through practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS
[0036] Other features, objects and advantages of the present application will become more apparent by reading the detailed description of non-limiting embodiments made with reference to the following drawings:
[0037] Figure 1 For the system architecture in the related technology;
[0038] Figure 2 This is an application scenario diagram of the speech signal processing method proposed in the embodiment of the present application;
[0039] Figure 3 A flowchart of a method for processing a speech signal in one embodiment of the present application;
[0040] Figure 4 A schematic diagram of the structure of a convolutional mixing model proposed in one embodiment of the present application;
[0041] Figure 5 A schematic diagram of the structure of a multi-layer perception hybrid model highlighted in one embodiment of the present application;
[0042] Figure 6 A flowchart of a method for processing a speech signal in another embodiment of the present application is shown in FIG.
[0043] Figure 7 For Figure 6 The corresponding model structure diagram;
[0044] Figure 8 A block diagram of a speech signal processing device in one embodiment of the present application;
[0045] Fig. 9 A block diagram of a speech signal processing device in another embodiment of the present application
[0046] Fig.10 A schematic diagram of the structure of a computer system of an electronic device or server suitable for implementing an embodiment of the present application is shown. DETAILED DESCRIPTION
[0047] The present application is further described in detail below in conjunction with the accompanying drawings and embodiments. It is to be understood that the specific embodiments described herein are only used to explain the relevant invention, rather than to limit the invention. It is also necessary to explain that, for ease of description, only the parts related to the invention are shown in the accompanying drawings.
[0048] It should be noted that, in the absence of conflict, the embodiments and features in the embodiments of the present application can be combined with each other. The present application will be described in detail below with reference to the accompanying drawings and in combination with the embodiments.
[0049] With the development of hardware technologies such as artificial intelligence algorithms and AI chips, smart devices have been widely used in daily life. Such as smart home voice control systems, smart speakers, smart conference systems, etc. These smart voice control systems need to be awakened before executing voice control strategies, that is, wake up the device. In related technologies, the initial wake-up method was to manually press the recording button for voice input. Later, wake-up technologies based on deep learning models gradually took shape, forming Figure 1The system architecture shown is that the client collects the user's voice information in real time, and then sends it to the cloud server through the wireless network. The voice information is analyzed by the deep learning model configured on the cloud server to obtain the voice recognition result or the voice control instruction based on the voice recognition result. The cloud server returns the voice recognition result or voice control instruction to the client, and the client executes the corresponding response strategy according to the returned voice recognition result or voice control instruction, such as responding to the user in response to the wake-up sentence, or executing the control plan corresponding to the control instruction. It can be seen that since the terminal cannot perform voice recognition by itself, the voice recognition operation cannot be realized when the network is not smooth, which affects the user's experience of using the voice control device.
[0050] Based on this, the present application proposes a method for processing voice signals, which has a smaller amount of model parameters and can be configured in a terminal device to ensure the user's voice control experience in an offline state.
[0051] Figure 2 This is an application scenario diagram of the speech signal processing method proposed in the embodiment of the present application. Figure 1 The application scenario includes a sound sampling device 20 and a speech signal processing device 10.
[0052] Among them, the voice collection device 20 and the voice signal processing device 10 are jointly arranged on the terminal device, and the terminal device can be an intelligent voice control device, such as an intelligent voice control device, a home appliance with voice control function, a smart phone, a tablet computer, a laptop computer, a wearable device, etc.
[0053] The sound collection device 20 is a device for collecting voice data, such as a microphone array. The voice signal processing device 10 is connected to the sound collection device 20 and is used to execute the voice signal processing method proposed in the present application to recognize the voice signal collected by the sound collection device 20 to obtain a recognition result and control the voice control device to execute a response strategy corresponding to the recognition result.
[0054] Figure 3 The figure is a flowchart of a method for processing a speech signal in one embodiment of the present application.
[0055] like Figure 3 As shown, the method for processing a speech signal proposed in the embodiment of the present application comprises the following steps:
[0056] Step 301: Acquire a speech signal collected from an environment.
[0057] The environment is composed of various natural factors. The environment can include real environment and virtual environment. The real environment exists in real life, and the virtual environment is obtained by simulating the real environment.
[0058] That is to say, the voice signal collected by the sound collection device from the environment can be directly obtained from the sound collection device, and the voice signal processed by the sound collection device based on the voice signal collected from the environment can also be obtained from the sound collection device. Among them, the processing of the voice signal by the sound collection device includes but is not limited to sound source localization, enhanced voice, voice endpoint detection, etc.
[0059] Step 302: extract speech features from the speech signal to obtain speech features corresponding to the speech signal.
[0060] Optionally, the speech features of the speech signal are extracted using Mel-frequency cepstral coefficients to obtain speech features corresponding to the speech signal.
[0061] That is to say, the speech features obtained by performing speech feature extraction on the speech signal are Mel features of the speech.
[0062] Specifically, the voice signal is framed to obtain multiple audio frames. Among them, the frame processing refers to dividing the voice signal into voice signal segments of fixed size, and each segment of the voice signal is called a frame, and the general frame length is 10-30ms. When performing frame processing, the overlapping segmentation method can be adopted, and the ratio of the frame shift to the frame length ranges from 0 to 1 / 2, wherein the frame shift is the difference between the previous frame and the next frame. By utilizing the short-term stability of the signal, a smooth transition between frames is achieved to maintain its continuity, and at the same time, the problem of information omission caused by the boundary of the time window can be avoided. In an embodiment of the present application, the frame length is 25ms, the frame shift is 10ms, the sampling frequency of the voice signal is 16000 / s, and the voice signal length is 1s.
[0063] Optionally, before the voice signal is framed, pre-emphasis processing may be performed on the voice signal to enhance the high-frequency signal in the voice signal, and windowing processing may be performed on the voice signal after the voice signal is framed to eliminate signal discontinuity that may be caused at both ends of each frame.
[0064] Then, each audio frame is subjected to Fourier transform to obtain the spectrum information corresponding to each audio. Among them, Fourier transform is used to convert the time domain signal into the frequency domain signal, and Fourier transform can adopt the fast Fourier transform method. The spectrum information corresponding to each audio frame is filtered by Mel filter to obtain the spectrum characteristics of each audio frame. Among them, the Mel filter can be a triangular filter bank. Filtering by Mel filter can make the obtained spectrum characteristics more consistent with the auditory characteristics of the human ear.
[0065] Optionally, logarithmic operations are performed on all filter outputs in the triangular filter bank, and then discrete cosine transform (DCT) is performed to finally obtain an MFCC (Mel Frequency Cepstrum Coefficient, MFCC) feature vector. In the embodiment of the present application, the length of the MFCC feature vector is 40, that is, in the embodiment of the present application, after the speech feature extraction is performed on the speech signal, a 98×40 two-dimensional feature vector is obtained, which is similar to an image with a length of 98, a width of 40, and a channel number of 1.
[0066] Step 303, performing convolution mixing processing on the speech features to obtain shallow speech recognition features.
[0067] In one or more embodiments, a convolutional mixing model is used to perform convolutional mixing processing on speech features.
[0068] Among them, the convolution mixing model includes a spatial position convolution mixing module and a channel position convolution mixing module. The mixing result of the spatial position convolution mixing module and the input of the spatial position convolution mixing module are input into the channel position convolution mixing module through a residual connection.
[0069] That is to say, the embodiment of the present application first uses the spatial position convolution mixing module in the convolution mixing model to mix the spatial position features in the speech features, and then uses the channel position convolution mixing module in the convolution mixing module to mix the channel position features in the speech features.
[0070] Furthermore, if Figure 4 As shown, the spatial position convolution module may include a depth-separable convolution layer, a first excitation function layer and a first normalization layer, and the channel position convolution mixing module includes a point-by-point convolution layer, a second excitation function layer and a second normalization layer. Among them, the first excitation function layer and the second excitation function layer can both use the GLU function as the excitation function, which can increase the nonlinearity of the matrix and further facilitate the extraction of correlations in different dimensions.
[0071] Optionally, the convolutional mixed model may be a ConvMixer model, which can realize speech recognition based on speech features and has a smaller number of model parameters than a traditional full convolutional model, and can be better configured on a terminal device.
[0072] Optionally, when performing convolution mixing processing on speech features, multiple convolution mixing models can be set continuously according to the computing power of the terminal or the requirements of speech recognition, and this application does not make specific limitations here.
[0073] Step 304, performing hybrid processing based on multi-layer perception on the shallow speech features to obtain deep speech recognition features.
[0074] In one or more embodiments, a multi-layer perceptual hybrid model may be used to perform convolutional hybrid processing on shallow speech recognition features.
[0075] Among them, the multi-layer perception mixing model includes a space perception mixing module and a channel perception mixing module. The space perception mixing module performs space perception mixing on the transposed feature information and transposes the mixing result again and inputs it into the channel mixing perception model for channel perception mixing.
[0076] That is to say, the spatial-aware mixing module has a transposition layer, which is used to transpose the feature matrix, that is, after converting the row features into column features, the spatial features are mixed based on the column features, and then the transposition layer is used again to transpose the mixed column features back to row features, so as to facilitate the mixing of channel features based on the row features.
[0077] Specifically, if Figure 5 As shown, the multi-layer perceptual hybrid model includes a fourth normalization layer, a spatial perceptual hybrid module (token-mixing MLP), a transposition layer, a fifth normalization layer, and a channel-perceptual hybrid module (channel-mixing MLP). The shallow speech recognition features are subjected to the fourth normalization layer to obtain channel-based row vector features. The transposition layer transposes the row vector features into column vector features, which are then input into the spatial perceptual hybrid layer for mixing to obtain mixed column vector features. The transposition layer then transposes the mixed column vector features to obtain row vector features. The row vector features are input into the fifth normalization layer for normalization processing. The normalized features are input into the channel-perceptual hybrid layer for mixing to obtain deep speech recognition features. Among them, a skip connection is used between the fifth normalization layer and the transposition layer, and a residual connection is used in the channel-perceptual hybrid module.
[0078] The spatial perception hybrid module includes a first fully connected layer, a third activation function layer, and a second fully connected layer, and the channel perception hybrid module includes a third fully connected layer, a fourth activation function layer, and a fourth fully connected layer. The third activation function layer and the fourth excitation function layer can both use the GLU function as an excitation function.
[0079] It should be understood that in the embodiment of the present application, when feature mixing is performed through the spatial perception mixing module, all columns share parameters in the spatial perception mixing module, and when feature mixing is performed through the channel perception mixing module, all rows share parameters in the column spatial perception mixing module. Alternating execution of the two types of perception mixing modules can promote information interaction between the two dimensions.
[0080] Optionally, when performing hybrid processing based on multi-layer perception on shallow speech recognition features, multiple hybrid models based on multi-layer perception can be set continuously according to the computing power of the terminal or the requirements of speech recognition. This application does not make specific limitations here.
[0081] It should also be understood that by performing convolution mixing processing on shallow speech features through a multi-layer perceptual hybrid model to obtain deep speech recognition features, the spatial mixing effect and channel mixing effect between speech features can be further improved, thereby effectively improving the expression effect of speech features, and further improving the speech recognition effect based on speech features.
[0082] Step 305: Obtain the recognition result of the speech signal according to the deep speech recognition feature.
[0083] It should be noted that the present application classifies speech recognition features through a classifier to obtain recognition results of speech signals.
[0084] Specifically, the speech recognition features are input into the 2D average pooling layer, and then through the fully connected layer using the softmax activation function, an N+2 classification output result can be obtained. Wherein, N is the number of wake-up words or command words obtained by the speech signal processing device through training, 2 represents silence and unknown, the wake-up words include but are not limited to "XiaoXiaoX", "Hello XiaoX", etc., and the command words include but are not limited to "turn on the air conditioner", "increase the volume", "turn up the brightness", etc.
[0085] In one or more embodiments, the model used in the present application to process speech features can be converted into tflite format after training, so that it can be directly run on smart terminals such as smartphones and tablets through Java or C++ language.
[0086] Step 306: Execute a response strategy corresponding to the recognition result according to the recognition result.
[0087] For example, when the recognition result is a wake-up word, the intelligent voice control device is controlled to respond, such as "Xiao X is here", "Are you here", etc. When the recognition result is a command word, the corresponding intelligent terminal is controlled to execute the control command. For example, when the command word is "turn on the air conditioner", the air conditioner can be controlled to turn on and execute the preset or last set control strategy. When the command word is "increase the volume", the player currently in the playing state is controlled to increase the volume by one level. When the command word is "increase the brightness", the lighting device currently in the lighting state is controlled to increase the brightness by one level.
[0088] In one or more embodiments, before performing convolution mixing processing on the speech features to obtain shallow speech features, it also includes: using a feature embedding module to downsample the speech features.
[0089] Among them, the feature embedding module includes a feature embedding layer, a fifth activation function layer and a third normalization layer.
[0090] Among them, the feature embedding layer is a 2D convolution layer with a convolution kernel of 64 and a channel number of 2. Therefore, the feature input to the convolutional hybrid model after processing by the feature embedding module is a tensor of [batch size, 49, 20, 64]. The feature embedding operation performed by the feature embedding module can complete all downsampling processes of the neural network, effectively reducing the resolution of the image, increasing the receptive field, and facilitating the convolutional hybrid model and the hybrid model based on multi-layer perception to find more distant spatial information.
[0091] As a specific example, Figure 6 and Figure 7 As shown, the method for processing a speech signal comprises the following steps:
[0092] Step 601: Acquire a voice signal collected from an environment.
[0093] Step 602: extract features of the speech signal using Mel-frequency cepstral coefficients to obtain speech features.
[0094] Step 603: input the speech features into a feature embedding module to obtain downsampled speech features.
[0095] Step 604, input the downsampled speech features into multiple consecutive convolutional mixed models to obtain shallow speech recognition features.
[0096] Step 605, input the shallow speech recognition features into the multi-layer perceptual hybrid model to obtain the deep speech recognition features.
[0097] Step 606, input the deep speech recognition features into the average pooling layer and the fully connected layer in sequence to obtain the speech recognition result.
[0098] Step 607: Execute the response strategy corresponding to the speech recognition result.
[0099] Furthermore, the present application also uses the speech signal processing method proposed in the present application to verify its effectiveness.
[0100] Specifically, the speech signal processing method proposed in the embodiment of the present application is tested using the Google Speech Commands V2 (GSC-V2) dataset. GSC-V2 contains 105,829 command words, 2,618 speakers, and 35 command words, including 'down', 'go', 'left', 'no', 'off', 'on', 'right', 'stop', 'up', and 'yes'. Convolutional mixed models of different sizes are used for experiments, such as the model depth is 8, the hidden layer size is 64, called ConvMlp-Mixer-S, the model depth is 12, the hidden layer size is 64, called ConvMlp-Mixer-M, the model depth is 12, the hidden layer size is 128, called ConvMlp-Mixer-L. The adamw optimizer is used in the experiment, the warmup epoch is 10, the number of iterations is 25000, the learning rate is 0.02, and the batch size is 256. The results are shown in Table 1.
[0101] Table 1
[0102]
[0103]
[0104] It can be seen that the ConvMlp-Mixer-S model corresponding to the speech signal processing method proposed in the embodiment of the present application uses only 96k parameters to achieve an accuracy of 96.24, while the ConvMlp-Mixer-L uses 0.299M parameters to achieve an accuracy of 97.77. Compared with the MLP-based model and the tansformer-based model, the effect is better and the model parameters are smaller.
[0105] To sum up, the voice signal processing method proposed in the embodiment of the present application, by performing convolution mixing operations and multi-perception-based mixing processing on the voice features extracted from the voice signal, can reduce the number of parameters of the voice recognition model while effectively recognizing voice commands, so that the processing device for executing the voice signal can be better configured on the terminal device, thereby improving the response efficiency of the terminal device to voice commands.
[0106] It should be noted that although the operations of the method of the present invention are described in a particular order in the drawings, this does not require or imply that the operations must be performed in this particular order or that all illustrated operations must be performed to achieve desired results.
[0107] Figure 8 The block diagram of a speech signal processing device in one embodiment of the present application is shown.
[0108] like Figure 8 As shown, the speech signal processing device 10 of the embodiment of the present application includes:
[0109] An acquisition module 11 is used to acquire the collected voice signal;
[0110] A speech feature extraction module 12 is used to extract speech features from the speech signal to obtain speech features corresponding to the speech signal;
[0111] A convolutional mixing model 13, used for performing convolutional mixing processing on the speech features to obtain shallow speech recognition features;
[0112] A multi-layer perception hybrid model 14, used for performing a hybrid process based on multi-layer perception on the shallow speech recognition features to obtain deep speech recognition features;
[0113] A classification module 15, configured to obtain a recognition result of the speech signal according to the deep speech recognition feature;
[0114] The execution module 16 is used to execute a response strategy corresponding to the identification result according to the identification result.
[0115] In some embodiments, the convolution mixing model 13 includes a spatial position convolution mixing module and a channel position convolution mixing module, and the mixing result of the spatial position convolution mixing module and the input of the spatial position convolution mixing module are input into the channel position convolution mixing module through a residual connection.
[0116] In some embodiments, the spatial position convolution mixing module includes a depth-separable convolution layer, a first excitation function layer and a first normalization layer, and the channel position convolution mixing module includes a point-by-point convolution layer, a second excitation function layer and a second normalization layer.
[0117] In some embodiments, the multi-layer perceptual mixing model 14 includes a spatial perceptual mixing module and a channel perceptual mixing module. The spatial perceptual mixing module performs spatial perceptual mixing on the transposed feature information and transposes the mixing result again and inputs it into the channel perceptual mixing module for channel perceptual mixing.
[0118] In some embodiments, the spatial-aware mixing module includes a first fully connected layer, a third activation function layer, and a second fully connected layer, and the channel-aware mixing module includes a third fully connected layer, a fourth activation function layer, and a fourth fully connected layer.
[0119] In some embodiments, Fig. 9 As shown, the speech signal processing device 10 also includes: a feature embedding module 17, the feature embedding module 17 is used to downsample the speech features, wherein the feature embedding module 17 includes a feature embedding layer, a fifth activation function layer and a third normalization layer.
[0120] In some embodiments, the speech feature extraction module 12 is used to extract speech features from the speech signal using Mel-frequency cepstral coefficients to obtain speech features corresponding to the speech signal.
[0121] It should be understood that the units or modules described in the device 10 are similar to those described in the reference Figure 2 The various steps in the described method correspond to each other. Therefore, the operations and features described above for the method are also applicable to the device 10 and the units contained therein, and will not be repeated here. The device 10 can be pre-implemented in a browser or other security application of an electronic device, or loaded into a browser or its security application of an electronic device by downloading or the like. The corresponding units in the device 10 can cooperate with the units in the electronic device to implement the solution of the embodiment of the present application.
[0122] For the several modules or units mentioned in the above detailed description, this division is not mandatory. In fact, according to the embodiments of the present disclosure, the features and functions of two or more modules or units described above can be embodied in one module or unit. On the contrary, the features and functions of one module or unit described above can be further divided into multiple modules or units to be embodied.
[0123] To sum up, the speech signal processing device proposed in the embodiment of the present application can reduce the number of parameters of the speech recognition model while effectively recognizing voice commands by performing convolution mixing operations and multi-perception-based mixing processing on the speech features extracted from the speech signal, so that the processing device for executing the speech signal can be better configured on the terminal device, thereby improving the response efficiency of the terminal device to voice commands.
[0124] Reference below Fig.10 , Fig.10 A schematic diagram of the structure of a computer system of an electronic device or server suitable for implementing an embodiment of the present application is shown.
[0125] like Fig.10 As shown, the computer system includes a central processing unit (CPU) 1001, which can perform various appropriate actions and processes according to the program stored in the read-only memory (ROM) 1002 or the program loaded from the storage part 1008 into the random access memory (RAM) 1003. In the RAM 1003, various programs and data required for the operation instructions of the system are also stored. The CPU 1001, the ROM 1002, and the RAM 1003 are connected to each other through a bus 1004. An input / output (I / O) interface 1005 is also connected to the bus 1004.
[0126] The following components are connected to the I / O interface 1005: an input section 1006 including a keyboard, a mouse, etc.; an output section 1007 including a cathode ray tube (CRT), a liquid crystal display (LCD), etc., and a speaker, etc.; a storage section 1008 including a hard disk, etc.; and a communication section 1009 including a network interface card such as a LAN card, a modem, etc. The communication section 1009 performs communication processing via a network such as the Internet. A drive 1010 is also connected to the I / O interface 1005 as needed. A removable medium 1011, such as a magnetic disk, an optical disk, a magneto-optical disk, a semiconductor memory, etc., is installed on the drive 1010 as needed, so that a computer program read therefrom is installed into the storage section 1008 as needed.
[0127] In particular, according to an embodiment of the present application, the above reference flow chart Figure 2 The described process can be implemented as a computer software program. For example, an embodiment of the present application includes a computer program product, which includes a computer program carried on a computer-readable medium, and the computer program includes a program code for executing the method shown in the flow chart. In such an embodiment, the computer program includes a program code for executing the method shown in the flow chart. In such an embodiment, the computer program can be downloaded and installed from the network through the communication part 1009, and / or installed from the removable medium 1011. When the computer program is executed by the central processing unit (CPU) 1001, the above-mentioned functions defined in the system of the present application are executed.
[0128] It should be noted that the computer-readable medium shown in the present application may be a computer-readable signal medium or a computer-readable storage medium or any combination of the above two. The computer-readable storage medium may be, for example, but not limited to, an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, device or device, or any combination of the above. More specific examples of computer-readable storage media may include, but are not limited to: an electrical connection with one or more wires, a portable computer disk, a hard disk, a random access memory (RAM), a read-only memory (ROM), an erasable programmable read-only memory (EPROM or flash memory), an optical fiber, a portable compact disk read-only memory (CD-ROM), an optical storage device, a magnetic storage device, or any suitable combination of the above. In the present application, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in combination with an instruction execution system, device or device. In the present application, a computer-readable signal medium may include a data signal propagated in a baseband or as part of a carrier wave, which carries a computer-readable program code. This propagated data signal may take a variety of forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination of the above. Computer-readable signal media may also be any computer-readable medium such as a computer-readable storage medium that can send, propagate or transmit a program for use by or in conjunction with an instruction execution system, device or device. The program code contained on the computer-readable medium may be transmitted using any appropriate medium, including but not limited to: wireless, wire, optical cable, RF, etc., or any suitable combination of the above.
[0129] The flow chart and block diagram in the accompanying drawings illustrate the possible architecture, functions and operating instructions of the system, method and computer program product according to various embodiments of the present application. In this regard, each box in the flow chart or block diagram can represent a module, a program segment or a part of a code, and the aforementioned module, program segment or a part of a code contains one or more executable instructions for realizing the specified logical function. It should also be noted that in some implementations as replacements, the functions marked in the box can also occur in a sequence different from that marked in the accompanying drawings. For example, the boxes represented by two connections can actually be executed substantially in parallel, and they can sometimes be executed in the opposite order, depending on the functions involved. It should also be noted that each box in the block diagram and / or flow chart, and the combination of the boxes in the block diagram and / or flow chart can be implemented with a dedicated hardware-based system that performs a specified function or operating instruction, or can be implemented with a combination of dedicated hardware and computer instructions.
[0130] The units or modules involved in the embodiments described in the present application may be implemented by software or hardware. The units or modules described may also be arranged in a processor. For example, they may be described as: a processor includes an acquisition module, a speech feature extraction module, a convolutional hybrid model, a multi-layer perceptual hybrid model, a classification module, and an execution module. The names of these units or modules do not, in some cases, constitute limitations on the units or modules themselves. For example, the acquisition module may also be described as "acquiring a collected speech signal."
[0131] As another aspect, the present application further provides a computer-readable storage medium, which may be included in the electronic device described in the above embodiment, or may exist independently without being assembled into the electronic device. The above computer-readable storage medium stores one or more programs, and when the above programs are used by one or more processors to execute the method for processing speech signals described in the present application.
[0132] The above description is only a preferred embodiment of the present application and an explanation of the technical principles used. Those skilled in the art should understand that the scope of disclosure involved in the present application is not limited to the technical solution formed by a specific combination of the above technical features, but should also cover other technical solutions formed by any combination of the above technical features or their equivalent features without departing from the aforementioned disclosed concept. For example, the above features are replaced with the technical features with similar functions disclosed in this application (but not limited to) by each other to form a technical solution.
Claims
1. A method for processing a speech signal, characterized in that: include: Acquire speech signals collected from the environment; Extracting speech features from the speech signal to obtain speech features corresponding to the speech signal; Performing convolution mixing processing on the speech features to obtain shallow speech recognition features; Performing a hybrid process based on multi-layer perception on the shallow speech recognition features to obtain deep speech recognition features; Obtaining a recognition result of the speech signal according to the deep speech recognition feature; According to the recognition result, executing a response strategy corresponding to the recognition result; The convolution mixing process is performed on the speech features to obtain shallow speech recognition features, including: The speech features are subjected to convolution mixing processing using a convolution mixing model. The convolutional mixing model includes a spatial position convolutional mixing module and a channel position convolutional mixing module, and the mixing result of the spatial position convolutional mixing module and the input of the spatial position convolutional mixing module are input into the channel position convolutional mixing module through a residual connection; The shallow speech recognition features are subjected to a mixed processing based on multi-layer perception to obtain deep speech recognition features, including: A multi-layer perceptual hybrid model is used to perform convolutional hybrid processing on the shallow speech recognition features. Among them, the multi-layer perceptual mixing model includes a space-perceptual mixing module and a channel-perceptual mixing module. The space-perceptual mixing module performs space-perceptual mixing on the transposed feature information and transposes the mixing result again and inputs it into the channel-perceptual mixing module for channel-perceptual mixing.
2. The method according to claim 1, characterized in that The spatial position convolution mixing module includes a depth-separable convolution layer, a first excitation function layer and a first normalization layer, and the channel position convolution mixing module includes a point-by-point convolution layer, a second excitation function layer and a second normalization layer.
3. The method according to claim 1, characterized in that The spatial perception hybrid module includes a first fully connected layer, a third activation function layer and a second fully connected layer, and the channel perception hybrid module includes a third fully connected layer, a fourth activation function layer and a fourth fully connected layer.
4. The method according to claim 1, characterized in that: Before performing convolution mixing processing on the speech features to obtain shallow speech recognition features, the method further includes: Downsampling the speech features using a feature embedding module; Among them, the feature embedding module includes a feature embedding layer, a fifth activation function layer and a third normalization layer.
5. The method according to claim 1, characterized in that The extracting speech features from the speech signal to obtain speech features corresponding to the speech signal includes: Mel-frequency cepstral coefficients are used to extract speech features from the speech signal to obtain speech features corresponding to the speech signal.
6. A speech signal processing device, characterized in that: include: An acquisition module, used to acquire the collected voice signal; A speech feature extraction module, used to extract speech features from the speech signal to obtain speech features corresponding to the speech signal; A convolutional mixing model is used to perform convolutional mixing processing on the speech features to obtain shallow speech recognition features; A multi-layer perception hybrid model is used to perform a hybrid process based on multi-layer perception on the shallow speech recognition features to obtain deep speech recognition features; A classification module, used to obtain a recognition result of the speech signal according to the deep speech recognition feature; An execution module, configured to execute a response strategy corresponding to the identification result according to the identification result; The convolutional mixing model is specifically used to perform convolutional mixing processing on the speech features using the convolutional mixing model. The convolutional mixing model includes a spatial position convolutional mixing module and a channel position convolutional mixing module, and the mixing result of the spatial position convolutional mixing module and the input of the spatial position convolutional mixing module are input into the channel position convolutional mixing module through a residual connection; The multi-layer perceptual hybrid model is specifically used to perform convolutional hybrid processing on the shallow speech recognition features using the multi-layer perceptual hybrid model. Among them, the multi-layer perceptual mixing model includes a space-perceptual mixing module and a channel-perceptual mixing module. The space-perceptual mixing module performs space-perceptual mixing on the transposed feature information and transposes the mixing result again and inputs it into the channel-perceptual mixing module for channel-perceptual mixing.
7. An electronic device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that: When the processor executes the program, it implements the method for processing a speech signal as described in any one of claims 1 to 5.
8. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the program is executed by a processor, the method for processing a speech signal as described in any one of claims 1 to 5 is implemented.
Citation Information
Patent Citations
Speech recognition method and device, equipment and storage medium
CN114333782A