A method for optimizing robot speech electrical signal processing
By proposing a new processing model in the robotic voice electrical signal processing system, using convolutional layer, activation function, parallel processing module and pooling operations, the existing system's problems of inaccurate processing and poor personalization adaptability in complex environments are solved, and more efficient and accurate voice signal processing is achieved.
Patent Information
- Application Number
- CN202411735411.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-29
- Publication Date
- 2025-05-16
- Estimated Expiration
- 2044-11-29
AI Technical Summary
Existing robotic voice and electrical signal processing systems are difficult to achieve real-time and accurate voice signal processing in complex environments, and are difficult to adapt to the personalized voice characteristics and environmental changes of different users.
A robot voice electrical signal processing model is proposed. The spatial characteristics of voice electrical signals are initially extracted through convolutional layers and activation functions. The overall processing module and the local processing module process data in parallel to fuse multi-scale information, and the extraction module integrates global and local feature information, and uses global average pooling and global maximum pooling to reduce information loss and model parameters.
It improves the accuracy and real-time nature of speech signal processing, enhances the model's expression ability and adaptability to complex environments, reduces the computational complexity and model parameters, and improves training and inference efficiency.
Smart Images

Figure CN119541491B_ABST
Abstract
Description
Technical Field
[0001] The invention belongs to the technical field of signal processing, and in particular relates to a method for optimizing robot speech electrical signal processing. Background Art
[0002] Robot speech electrical signal processing is one of the key technologies for realizing human-computer interaction. It can help robots recognize and execute voice commands. However, in practical applications, speech electrical signal processing faces many challenges. First, robots work in complex and changing environments and are often disturbed by environmental noise, such as wind noise, mechanical noise, etc. External noise can significantly reduce the signal-to-noise ratio of speech signals, greatly reducing the accuracy of speech recognition. This problem is particularly evident in outdoor or industrial scenarios. Therefore, it is difficult for existing speech processing systems to stably and accurately obtain clear speech signals.
[0003] Although a variety of speech signal processing algorithms have been used to reduce noise and enhance speech signals, they are insufficient in terms of real-time performance and accuracy in complex environments. Traditional speech signal processing mostly relies on static feature extraction and regular models, which are difficult to adapt to complex sound changes. In addition, real-time processing is a key requirement in robot voice interaction. However, in order to improve accuracy, existing methods often require higher computing resources, which increases processing delays, resulting in the system response speed being unable to meet the needs of real-time interaction, affecting the robot's operating efficiency and user experience.
[0004] In addition, there are problems with the personalization and adaptability of robot voice processing. The voice characteristics, accents, and speaking speeds of different users vary greatly, and the voice characteristics of the same user in different situations may also change. Current voice processing technology shows limited flexibility when dealing with individual differences, and it is difficult to dynamically adjust the algorithm to adapt to diverse voice inputs. Summary of the invention
[0005] The invention provides a method for optimizing robot speech electric signal processing, aiming at proposing a robot speech electric signal processing model, wherein a preliminary extraction module performs preliminary processing on the robot speech electric signal through a convolution layer, effectively extracts spatial features in the speech electric signal, selects a trainable activation function to process the output of the convolution layer, introduces nonlinear transformation, and further enhances the expression ability of the model, an overall processing module and a local processing module can further extract features of the robot speech electric signal, the overall processing module and the local processing module process input data in a parallel manner, and fuse multi-scale information, so that feature representation is richer and more comprehensive, and an integrated extraction module adds the outputs of the overall processing module and the local processing module element by element, fuses global and local feature information, and is helpful to construct a more comprehensive and detailed feature representation, improves classification performance, and the parallel operation of global average pooling and global maximum pooling ensures that global information and local important features can be captured, thereby reducing information loss, and can effectively reduce the number of model parameters, reduce calculation complexity, and improve training and reasoning efficiency.
[0006] In order to achieve the above-mentioned purpose, the present invention provides the following technical solution: a method for optimizing robot speech electrical signal processing, the specific steps are as follows.
[0007] S1. Collect robot voice electrical signal data and pre-process it.
[0008] S2. Build a preliminary extraction module, including convolutional layer, activation function, and maximum pooling layer.
[0009] S3. Build the overall processing module, including convolution layer, pooling layer, and activation function.
[0010] S4. Construct a local processing module, including a local processing block.
[0011] S5. Construct an integration extraction module, including global maximum pooling and global average pooling.
[0012] S6. Construct a robot speech electrical signal processing model, including input, preliminary extraction module, overall processing module, local processing module, integrated extraction module and output.
[0013] S7. Training and detection of the robot speech electric signal processing model. Use the robot speech electric signal data set to train the robot speech electric signal processing model. After the training is completed, obtain the robot speech electric signal to be detected, input it into the robot speech electric signal processing model, and obtain the detection result of the input robot speech electric signal.
[0014] Preferably, in step S1, the voice command of the robot is obtained and converted into an electrical signal, the electrical signal is labeled, and a bandpass filter is used to perform data preprocessing operations on the electrical signal to remove noise interference in the robot voice electrical signal.
[0015] Preferably, in step S2, for the preliminary extraction module, the input is the robot voice electrical signal ,in, represents the number of samples, Represents the number of channels. The robot voice electrical signal first passes through the convolution layer, then is processed by the activation function, and then passes through the maximum pooling layer to obtain the final output. , ,in, represents maximum pooling, represent Activation function, Represents a convolutional layer.
[0016] Preferably, in step S2, for the preliminary extraction module, the robot voice electrical signal is preliminarily processed through the convolution layer, which can effectively extract the spatial features in the signal. The weight of the convolution layer is trainable, which means that it can be adaptively adjusted through training data to extract the most discriminative features. The output of the convolution layer is processed using an activation function, which can introduce nonlinear transformations. The selection of a trainable activation function further enhances the expressive power of the model, can better fit complex data, and can extract more discriminative features in the initial stage, thereby improving classification performance. This is particularly important for the detection and classification tasks of robot voice electrical signals, because the characteristics of voice electrical signals are susceptible to noise interference, and accurate feature extraction directly affects the final classification effect.
[0017] Preferably, in step S3, for the overall processing module, the input is , From , after a convolutional layer, we get , ,in, represents the convolutional layer, followed by a parallel dual branch, where one branch includes a maximum pooling layer and a depth-wise separable convolution, and the output is , ,in, stands for depthwise separable convolution, represents the maximum pooling layer, and the other branch includes the average pooling layer and the depth-wise separable convolution, and the output is , ,in, represents the average pooling layer, and Add element by element and pass the activation function to get , ,in, represents element-by-element addition, represent Activation function, and Perform element-by-element multiplication and feed the result into the activation function to get the final output , ,in, represents element-wise multiplication, represent Activation function.
[0018] Preferably, in step S3, for the overall processing module, the depthwise separable convolution greatly reduces the amount of parameters and calculations compared to the standard convolution, and performs spatial convolution and channel convolution separately, which can not only maintain a high feature extraction capability, but also significantly reduce the computational complexity and improve processing efficiency. Through parallel maximum pooling and average pooling branches, information of different scales can be extracted. Maximum pooling retains the most significant features, and average pooling captures the overall trend. The two pooled features are processed by depthwise separable convolution and then added element by element, which can fuse multi-scale information and make the feature representation richer and more comprehensive. After feature fusion, nonlinear transformation is performed through activation function, which can enhance the expression ability of the model. Through element-by-element addition and activation operations, the feature expression is further enriched. Subsequently, element-by-element multiplication operations and activation function processing are performed, which increases the nonlinear combination between features and improves the model's ability to capture complex patterns.
[0019] Preferably, in step S4, the local processing module is composed of four local processing blocks, and for each local processing block, the input is , the input first passes through a convolutional layer, then passes through the activation function and convolutional layer, and the output is the same as the input Perform element-by-element addition operations and finally obtain the output of the local processing block , ,in, represents the convolutional layer, represents element-by-element addition, represent The activation function takes the input After being processed by four local processing blocks, the output of the final local processing module is obtained. .
[0020] Preferably, in step S4, the local processing module is composed of four local processing blocks, each of which contains two convolutional layers and an activation function, which can extract more complex and discriminative features in the local range. By introducing residual connections, gradients are allowed to propagate directly through jump connections, which alleviates the gradient vanishing problem in deep networks and improves the stability and effect of training. Since the local processing module uses four local processing blocks, the local processing blocks have a higher risk of overfitting. Therefore, a random inactivation layer is added after the activation function to randomly inactivate some neurons to prevent overfitting in the local processing blocks, thereby improving the generalization ability of the model and enhancing the robustness of the model to the input data. Although each local processing block has a high feature extraction capability, the risk of overfitting due to high complexity is avoided by adding random inactivation layers and residual connections. The entire local processing module can avoid overfitting while maintaining high performance, thereby improving the overall processing effect.
[0021] Preferably, in step S5, for the integrated extraction module, the output of the overall processing module Output of the local processing module Doing element-by-element addition, we get , ,in, represents element-by-element addition, then After average pooling and maximum pooling are performed in parallel, the results are concatenated and then input into the fully connected layer to obtain the final output. , ,in, represents the fully connected layer, Represents a splicing operation, represents global average pooling, Represents global maximum pooling.
[0022] Preferably, in step S5, for the integrated extraction module, by adding the outputs of the overall processing module and the local processing module element by element, the global and local feature information is integrated, which helps to build a more comprehensive and detailed feature representation and improve the classification performance. The use of global average pooling and global maximum pooling can capture information of different scales. Average pooling captures the overall trend, while maximum pooling captures the most significant features. The parallel operation of global average pooling and global maximum pooling ensures that both global information and local important features can be captured, thereby reducing information loss. The splicing operation can fuse information of different scales, provide richer and more useful feature representations, and enhance the discrimination ability and robustness of the model. Splicing the pooled results and inputting them into the fully connected layer can effectively reduce the number of model parameters, reduce computational complexity, and improve the efficiency of training and reasoning. At the same time, classification through the fully connected layer can make full use of high-level feature representations and improve classification accuracy.
[0023] Preferably, in step S6, for the robot speech electric signal processing model, the robot speech electric signal data is first input into the model, the speech electric signal is first processed by the preliminary extraction module, and then the processing results are respectively input into the overall processing module and the local processing module for parallel processing, and the robot speech electric signal features are further extracted. Finally, the results of the parallel processing are input into the integrated extraction module, and the final robot speech electric signal processing model is obtained after the processing is completed.
[0024] Compared with the prior art, the present invention has the following technical effects: the technical solution provided by the present invention proposes a robot speech electric signal processing model, wherein the preliminary extraction module performs preliminary processing on the robot speech electric signal through a convolution layer, effectively extracts the spatial features in the speech electric signal, selects a trainable activation function to process the output of the convolution layer, introduces nonlinear transformation, and further enhances the expression ability of the model; the overall processing module and the local processing module can further extract the features of the robot speech electric signal; the overall processing module and the local processing module process the input data in a parallel manner, and fuse multi-scale information to make the feature representation richer and more comprehensive; the integrated extraction module adds the outputs of the overall processing module and the local processing module element by element, fuses the global and local feature information, and helps to construct a more comprehensive and detailed feature representation, improves the classification performance, and the parallel operation of the global average pooling and the global maximum pooling ensures that the global information and the local important features can be captured, thereby reducing information loss, and at the same time can effectively reduce the number of model parameters, reduce the computational complexity, and improve the efficiency of training and reasoning. BRIEF DESCRIPTION OF THE DRAWINGS
[0025] Figure 1 The present invention provides a flowchart of a method for processing robot voice electrical signals.
[0026] Figure 2 It is a structural diagram of the preliminary extraction module provided by the present invention.
[0027] Figure 3 It is a structural diagram of the overall processing module provided by the present invention.
[0028] Figure 4 It is a structural diagram of the local processing block provided by the present invention.
[0029] Figure 5 It is a structural diagram of the local processing module provided by the present invention.
[0030] Figure 6 It is a structural diagram of the integrated extraction module provided by the present invention.
[0031] Figure 7 It is an effect diagram of the training and testing process of the robot speech electrical signal processing model provided by the present invention. DETAILED DESCRIPTION
[0032] The present invention aims to propose a method for optimizing robot speech electric signal processing, and proposes a robot speech electric signal processing model, wherein a preliminary extraction module performs preliminary processing on the robot speech electric signal through a convolution layer, effectively extracts spatial features in the speech electric signal, selects a trainable activation function to process the output of the convolution layer, introduces nonlinear transformation, and further enhances the expression ability of the model, and the overall processing module and the local processing module can further extract the features of the robot speech electric signal, and the overall processing module and the local processing module process the input data in a parallel manner, and fuse multi-scale information, so that the feature representation is richer and more comprehensive, and the integrated extraction module adds the outputs of the overall processing module and the local processing module element by element, fuses the global and local feature information, and helps to construct a more comprehensive and detailed feature representation, improves the classification performance, and the parallel operation of the global average pooling and the global maximum pooling ensures that the global information and the local important features can be captured, thereby reducing the information loss, and at the same time can effectively reduce the number of model parameters, reduce the computational complexity, and improve the efficiency of training and reasoning.
[0033] See also Figure 1 As shown, a method for optimizing robot speech electrical signal processing in an embodiment of the present application, the specific steps are as follows.
[0034] S1. Collect robot voice electrical signal data and pre-process it.
[0035] Furthermore, in step S1, the voice command of the robot is obtained and converted into an electrical signal, the electrical signal is labeled, and a bandpass filter is used to perform data preprocessing operations on the electrical signal to remove noise interference in the robot voice electrical signal.
[0036] S2, build a preliminary extraction module, including a convolutional layer, Activation function and a max pooling layer.
[0037] Furthermore, in step S2, for the preliminary extraction module, its structure is as follows Figure 2 As shown, the input is the robot voice electrical signal ,in, represents the number of samples, Represents the number of channels. The robot voice electrical signal first passes through the convolution layer, then is processed by the activation function, and then passes through the maximum pooling layer to obtain the final output. , ,in, represents maximum pooling, represent Activation function, Represents a convolutional layer.
[0038] S3, build the overall processing module, including a convolution layer, a maximum pooling layer, an average pooling layer, two depth-separable convolutions, Activation function, activation function and a dropout layer.
[0039] Further, in step S3, for the overall processing module, its structure is as follows: Figure 3 As shown, the input is , From , after a convolutional layer, we get , ,in, represents the convolutional layer, followed by a parallel dual branch, where one branch includes a maximum pooling layer and a depth-wise separable convolution, and the output is , ,in, stands for depthwise separable convolution, represents the maximum pooling layer, and the other branch includes the average pooling layer and the depth-wise separable convolution, and the output is , ,in, represents the average pooling layer, and Add them element by element, and reduce overfitting through random inactivation layer, and then pass through activation function to get , ,in, represents element-by-element addition, represent Activation function, and Perform element-by-element multiplication and feed the result into the activation function to get the final output , ,in, represents element-wise multiplication, represent Activation function.
[0040] S4, build a local processing module, including four local processing blocks, each of which includes two convolutional layers, activation function and a dropout layer.
[0041] Further, in step S4, for the local processing module, its structure is as follows: Figure 5 As shown, the local processing module is composed of four local processing blocks. For each local processing block, its structure is as follows Figure 4 As shown, the input is , the input first passes through a convolutional layer, then passes through the activation function, random inactivation layer, and convolutional layer in sequence, and the output is the same as the input Perform element-by-element addition operations and finally obtain the output of the local processing block , ,in, represents the convolutional layer, represents element-by-element addition, represent The activation function takes the input After being processed by four local processing blocks, the output of the final local processing module is obtained. .
[0042] S5. Construct an integration extraction module, which includes a global maximum pooling and a global average pooling.
[0043] Further, in step S5, for the integration and extraction module, its structure is as follows: Figure 6 As shown, the output of the overall processing module Output from local processing module By adding the elements one by one, we get , ,in, represents element-by-element addition, then After average pooling and maximum pooling are performed in parallel, the results are concatenated and then input into the fully connected layer to obtain the final output. , ,in, represents the fully connected layer, Represents a splicing operation, represents global average pooling, Represents global maximum pooling.
[0044] S6. Construct a robot speech electrical signal processing model, including input, preliminary extraction module, overall processing module, local processing module, integrated extraction module and output.
[0045] Furthermore, in step S6, for the robot speech electric signal processing model, the robot speech electric signal data is first input into the model, the speech electric signal is first processed by the preliminary extraction module, and then the processing results are respectively input into the overall processing module and the local processing module for parallel processing, and the robot speech electric signal features are further extracted. Finally, the results of the parallel processing are input into the integrated extraction module, and the final robot speech electric signal processing model is obtained after the processing is completed.
[0046] S7. Training and detection of the robot speech electric signal processing model. Use the robot speech electric signal data set to train the robot speech electric signal processing model. After the training is completed, obtain the robot speech electric signal to be detected, input it into the robot speech electric signal processing model, and obtain the detection result of the input robot speech electric signal.
[0047] Furthermore, in step S6, for the robot voice electrical signal processing model, it is written based on Python language and uses the TensorFlow framework. The robot voice electrical signal is input into the robot voice electrical signal processing model. The SGD optimizer is used, the learning rate is set to 0.001, the training batch size is 60, the training data set is repeated 100 times, the input robot voice electrical signal is trained, and the training accuracy is output to obtain the final trained robot voice electrical signal processing model.
[0048] Furthermore, in step S7, the results of the robot speech electrical signal processing model training and testing process are as follows: Figure 7 As shown, from the comparison of the training and testing effects, it can be seen that the accuracy rate increases continuously during the training process, and finally gradually stabilizes and the accuracy rate can reach above 0.9. The accuracy rate of the test process is lower than that of the training process, but it can still achieve an accuracy rate equivalent to that of the training process. For the robot voice electrical signal, the robot voice electrical signal processing model is used to detect the robot voice electrical signal, which has a reliable detection effect, verifying the effectiveness of the model proposed by this method.
[0049] The above are only preferred embodiments of the present invention. It should be pointed out that a person skilled in the art can make several modifications and improvements without departing from the inventive concept of the present invention, and these all fall within the protection scope of the present invention.
Claims
1. A method for optimizing robot speech electrical signal processing, characterized in that: The specific steps include: S1, collect robot voice electrical signal data and pre-process it; S2, build a preliminary extraction module, including convolutional layer, activation function, and maximum pooling layer; S3, build the overall processing module, including convolution layer, pooling layer, activation function, for the overall processing module, Input is , From , after a convolutional layer, we get , ,in, represents the convolutional layer, followed by a parallel dual branch, where one branch includes a maximum pooling layer and a depth-wise separable convolution, and the output is , ,in, stands for depthwise separable convolution, represents the maximum pooling layer, and the other branch includes the average pooling layer and the depth-wise separable convolution, and the output is , ,in, represents the average pooling layer, and Add element by element and pass the activation function to get , ,in, represents element-by-element addition, represent Activation function, and Perform element-by-element multiplication and feed the result into the activation function to get the final output , ,in, represents element-wise multiplication, represent Activation function; S4, constructing a local processing module, including a local processing block; S5, build the integration and extraction module, including the pooling layer; S6. Construct a robot speech electrical signal processing model, including input, preliminary extraction module, overall processing module, local processing module, integrated extraction module and output; S7. Training and detection of the robot speech electric signal processing model. Use the robot speech electric signal data set to train the robot speech electric signal processing model. After the training is completed, obtain the robot speech electric signal to be detected, input it into the robot speech electric signal processing model, and obtain the detection result of the input robot speech electric signal.
2. A method for optimizing robot speech electrical signal processing according to claim 1, characterized in that: In the step S1, the voice command of the robot is obtained and converted into an electrical signal, the electrical signal is labeled, and a bandpass filter is used to perform data preprocessing operations on the electrical signal to remove noise interference in the robot voice electrical signal.
3. A method for optimizing robot speech electrical signal processing according to claim 2, characterized in that: In the step S2, for the preliminary extraction module, the input is the robot voice electrical signal ,in, represents the number of samples, Represents the number of channels. The robot voice electrical signal first passes through the convolution layer, then is processed by the activation function, and then passes through the maximum pooling layer to obtain the final output. , ,in, represents maximum pooling, represent Activation function, Represents a convolutional layer.
4. A method for optimizing robot speech electrical signal processing according to claim 3, characterized in that: In the step S4, the local processing module is composed of four local processing blocks. For each local processing block, the input is , the input first passes through a convolutional layer, then passes through the activation function and convolutional layer, and the output is the same as the input Finally, the output of the local processing block is obtained. , ,in, represents the convolutional layer, represents element-by-element addition, represent The activation function takes the input After being processed by four local processing blocks, the output of the final local processing module is obtained. .
5. A method for optimizing robot speech electrical signal processing according to claim 4, characterized in that: In the step S5, the output of the overall processing module is processed by the integrated extraction module. Output of the local processing module Doing element-by-element addition, we get , ,in, represents element-by-element addition, then After average pooling and maximum pooling are performed in parallel, the results are concatenated and then input into the fully connected layer to obtain the final output. , ,in, represents the fully connected layer, Represents a splicing operation, represents global average pooling, Represents global maximum pooling.
6. A method for optimizing robot speech electrical signal processing according to claim 5, characterized in that: In the step S6, for the robot speech electric signal processing model, the robot speech electric signal data is first input into the model, the speech electric signal is first processed by the preliminary extraction module, and then the processing results are respectively input into the overall processing module and the local processing module for parallel processing, and the robot speech electric signal features are further extracted. Finally, the results of the parallel processing are input into the integrated extraction module, and the final robot speech electric signal processing model is obtained after the processing is completed.
Citation Information
Patent Citations
Speech emotion classification method and device, equipment and storage medium
CN116312644A