Method for determining attitude of disconnecting switch and electronic device
By introducing a global average pooling layer and an improved Cswish activation function in the voiceprint monitoring model, combining the depth-separable convolution and attention mechanism, the problem of difficult to balance the speed and accuracy of the isolation switch attitude monitoring in the prior art is solved, and fast and accurate attitude recognition is achieved.
Patent Information
- Application Number
- CN202411462553.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-10-18
- Publication Date
- 2025-07-29
- Estimated Expiration
- 2044-10-18
AI Technical Summary
In the prior art, when the voiceprint recognition model performs attitude monitoring of the isolating switch, it is difficult to ensure a faster monitoring speed and a higher monitoring accuracy.
A voiceprint monitoring model is adopted, which includes the first SC attention mechanism layer and a fully connected layer. Using the global average pooling layer and the improved Cswish activation function, the spatial characteristics and channel characteristics of the speech Mel spectrogram are extracted, combined with the depth separation convolution and attention mechanism, and fast and accurate pose recognition is performed.
The recognition accuracy of the isolating switch soundprint monitoring model is improved, and the training convergence speed of the model is accelerated, achieving fast and accurate attitude monitoring.
Smart Images

Figure CN119296572B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the technical field of deep learning, and in particular, to a method for determining the posture of a disconnector and an electronic device. Background Art
[0002] A disconnector is a circuit that disconnects a circuit without load current, so that the equipment under repair has an obvious disconnection point from the power supply, thus ensuring the personal safety of maintenance personnel. The disconnector does not have a special arc extinguishing device, so it cannot cut off the load current and short-circuit current. Therefore, the operation of the disconnector must be carried out when the circuit breaker is disconnected.
[0003] Voiceprint recognition technology is a technology that uses voice features for identity verification or voice analysis. It is mainly based on the unique attributes in the voice signal, such as the frequency range, rhythm, and fluctuation pattern of the voice. These characteristics can be used to distinguish different speakers or voice sources and are widely used in security authentication, intelligent assistants, forensic medicine, and other fields.
[0004] The technology of voiceprint recognition has also been extended to the condition monitoring of specific industries and equipment, such as the fault diagnosis and operation status monitoring of industrial equipment. By collecting the sounds emitted during the operation of the equipment and analyzing these sound data to monitor the working status of the equipment or predict the occurrence of faults.
[0005] During the power grid operation process, it is necessary to disconnect and close the disconnector for the next operation. Therefore, it is necessary to judge the posture of the disconnector. And in this process, it is necessary for manual workers to judge on-site in the power grid operation area whether the current posture of the disconnector is in the closed posture, which has the disadvantages of low efficiency and low reliability. By analyzing the sounds emitted by the disconnector in the closed or open state and using deep learning technology to establish a voiceprint recognition model, the automatic recognition of its operation state can be realized. The application of this technology can greatly improve the operation and maintenance efficiency and safety of the power system, especially has important value in remote control operation and automatic monitoring systems. By accurately analyzing the voiceprint of the disconnector, the dependence on manual visual inspection can be reduced, and the accuracy and efficiency of recognition can be improved.
[0006] In the prior art, the running speed of the voiceprint recognition model is slow. Some methods use complex network training, resulting in a high running time cost. Some methods have a fast running speed, but the final accuracy is not high. In practical applications, a high requirement for accuracy is needed. Summary of the Invention
[0007] The main purpose of the present application is to provide a method for determining the posture of a disconnector and an electronic device, so as to at least solve the problem that it is difficult to ensure both a fast monitoring speed and a high monitoring accuracy when the voiceprint recognition model monitors the posture of the disconnector in the prior art.
[0008] To achieve the above object, according to one aspect of the present application, a method for determining the attitude of a disconnector is provided, including: obtaining the sound emitted by the disconnector at the current moment to obtain voice data, and converting the voice data into a corresponding Mel spectrogram to obtain a voice Mel spectrogram; inputting the voice Mel spectrogram into a voiceprint monitoring model so that the voiceprint monitoring model analyzes the voice Mel spectrogram to obtain a monitoring result, where the monitoring result is a result indicating that the attitude of the disconnector at the current moment is a closed attitude or an open attitude. Among them, the voiceprint monitoring model includes a first SC (Spatial and Channel) attention mechanism layer and a fully connected layer. The first SC attention mechanism layer includes a global average pooling layer and a first Cswish layer. The output end of the global average pooling layer is connected to the input end of the first Cswish activation layer. The global average pooling layer is used to extract the spatial features and channel features of the voice Mel spectrogram. The activation function of the first Cswish activation layer is a function obtained by improving the Swish function using exponential growth. The output end of the first Cswish activation layer is connected to the input end of the fully connected layer, and the fully connected layer is used to output the monitoring result.
[0009] Optionally, the voiceprint monitoring model further includes a first convolutional layer and a first MobileNetV2 layer. The output end of the first convolutional layer is connected to the input end of the first MobileNetV2 layer. The output end of the first MobileNetV2 layer is connected to the input end of the first SC attention mechanism layer. The first SC attention mechanism layer further includes a second convolutional layer, a first batch normalization layer, a ReLU layer, and a third convolutional layer connected in sequence. The input end of the second convolutional layer is connected to the output end of the global average pooling layer. The output end of the third convolutional layer is connected to the input end of the first Cswish activation layer. The first SC attention mechanism layer further includes a product layer. The input ends of the product layer are respectively connected to the output end of the first Cswish activation layer and the output end of the first convolutional layer. The output end of the product layer is connected to the input end of the fully connected layer. Input the voice mel spectrogram into the voiceprint monitoring model so that the voiceprint monitoring model analyzes the voice mel spectrogram to obtain a monitoring result, including: input the voice mel spectrogram into the first convolutional layer so that the first convolutional layer extracts local features of the voice mel spectrogram to obtain a first feature map; input the first feature map into the first MobileNetV2 layer so that the first MobileNetV2 layer extracts features from the first feature map through depthwise separable convolution to obtain a second feature map; input the second feature map into the global average pooling layer so that the global average pooling layer performs global average pooling on the second feature map to obtain a matrix z, where, Q2 represents the second feature map, C represents the number of channels in the second feature map, F represents the number of frequency components in the second feature map, T represents the number of time components in the second feature map, c represents the serial number of the channel, t represents the serial number of the time component, and f represents the serial number of the frequency component; input the matrix z into the interaction layer so that the interaction layer processes the matrix z to obtain a target matrix, where the interaction layer includes the second convolutional layer, the first batch normalization layer, the ReLU layer, and the third convolutional layer connected in sequence; input the target matrix into the first Cswish activation layer so that the first Cswish activation layer performs nonlinear activation on the target matrix to obtain an attention weight, where the activation function of the first Cswish activation layer is f(x) = m(max(a(e x-1), x)) + nx · sigmoid(x), where α is a hyperparameter, m is a preset first adaptive parameter, n is a preset second adaptive parameter, x is the input of the first Cswish activation layer, and f(x) is the output of the first Cswish activation layer; input the attention weight and the second feature map into the product layer, so that the product layer multiplies the attention weight and the second feature map element by element to obtain a third feature map; input the third feature map into the fully connected layer to obtain the monitoring result.
[0010] Optionally, the voiceprint monitoring model further includes a second MobileNetV2 layer, a second SC attention mechanism layer, a first linear layer, and a CFB (Convolutional Feature Block) combination layer connected in sequence. The CFB combination layer includes multiple CFB modules connected in sequence. The input end of the second MobileNetV2 layer is connected to the output end of the first SC attention mechanism layer. The voiceprint monitoring model further includes a splicing layer. The input ends of the splicing layer are respectively connected to the output ends of each CFB module. The output end of the splicing layer is connected to the input end of the fully connected layer. The internal structure of the second SC attention mechanism layer is the same as that of the first SC attention mechanism layer. Inputting the third feature map into the fully connected layer to obtain the monitoring result includes: inputting the third feature map into the second MobileNetV2 layer, so that the second MobileNetV2 layer extracts features from the third feature map through depthwise separable convolution to obtain a fourth feature map; inputting the fourth feature map into the second SC attention mechanism layer to obtain a fifth feature map; inputting the fifth feature map into the first linear layer to obtain a sixth feature map; inputting the sixth feature map into the CFB combination layer, and obtaining the feature maps output by each CFB module in the CFB combination layer, and inputting each feature map into the splicing layer, so that the splicing layer performs a splicing operation on the multiple feature maps to obtain a seventh feature map, where the CFB combination layer is used to extract the global features and local features of the sixth feature map; inputting the seventh feature map into the fully connected layer to obtain the monitoring result.
[0011] Optionally, a plurality of the CFB modules are arranged in sequence, the sixth feature map is input into the CFB combination layer, and the feature maps output by each of the CFB modules in the CFB combination layer are obtained, including: a first input step of inputting the sixth feature map into the first CFB module so that the first CFB module extracts the global feature and the local feature of the sixth feature map to obtain the first feature map; a second input step of inputting the current feature map into the current CFB module so that the current CFB module extracts the global feature and the local feature of the current feature map to obtain the target feature map, where the current CFB module is any one of the remaining CFB modules in the CFB combination layer except the first CFB module, and the current feature map is the feature map output by the CFB module connected to the input end of the current CFB module; the second input step is repeated until the feature map output by the last CFB module is obtained.
[0012] Optionally, the internal structures of all the CFB modules are the same. The CFB module includes a first LFFB (Lightweight FeedForward Block) module, a CVAAttention (Convolutional Vision Attention) module, a ConvB (Convolutional Block) module, a second LFFB module, and a first-layer normalization module, which are connected in sequence. Inputting the sixth feature map into the first CFB module to enable the first CFB module to extract the global and local features of the sixth feature map and obtain the first feature map includes: inputting the sixth feature map into the first LFFB module to enable the first LFFB module to perform feature extraction on the sixth feature map through lightweight convolution to obtain a first sub-feature map; adding the sixth feature map to half of the first sub-feature map to obtain a second sub-feature map; inputting the second sub-feature map into the CVAAttention module to enable the CVAAttention module to extract the global and local features of the second sub-feature map to obtain a third sub-feature map; adding the third sub-feature map and the second sub-feature map to obtain a fourth sub-feature map; inputting the fourth sub-feature map into the ConvB module to enable the ConvB module to extract the deep features of the fourth sub-feature map to obtain a fifth sub-feature map; adding the fifth sub-feature map and the fourth sub-feature map to obtain a sixth sub-feature map; inputting the sixth sub-feature map into the second LFFB module to enable the second LFFB module to perform feature extraction on the sixth sub-feature map through lightweight convolution to obtain a seventh sub-feature map; adding the sixth sub-feature map to half of the seventh sub-feature map to obtain an eighth sub-feature map; inputting the eighth sub-feature map into the first-layer normalization module to enable the first-layer normalization module to perform layer normalization processing on the eighth sub-feature map to obtain the first feature map.
[0013] Optionally, the internal structures of the first LFFB module and the second LFFB module are the same. The first LFFB module includes a first combination layer, a second combination layer, and a third combination layer connected in sequence. The first combination layer includes a first Conv1D layer, a first GELU (Gaussian Error Linear Unit) layer, and a second batch normalization layer connected in sequence. The second combination layer includes a first DWConv1D layer, a second GELU layer, and a third batch normalization layer connected in sequence. The third combination layer includes a second Conv1D layer, a third GELU layer, and a fourth batch normalization layer connected in sequence. Input the sixth feature map into the first LFFB module so that the first LFFB module extracts features from the sixth feature map through lightweight convolution to obtain a first sub-feature map, including: input the sixth feature map into the first combination layer so that the first combination layer performs one-dimensional convolution processing, activation processing, and batch normalization processing on the sixth feature map to obtain a first output feature map; input the first output feature map into the second combination layer so that the second combination layer performs one-dimensional lightweight convolution processing, activation processing, and batch normalization processing on the first output feature map to obtain a second output feature map; input the second output feature map into the third combination layer so that the third combination layer performs one-dimensional convolution processing, activation processing, and batch normalization processing on the second output feature map to obtain the first sub-feature map.
[0014] Optionally, the CVAAttention module includes a fourth convolutional layer, a fifth convolutional layer, a sixth convolutional layer, a seventh convolutional layer, a first Softmax layer, and a second Softmax layer. The second sub-feature map is input into the CVAAttention module so that the CVAAttention module extracts the global features and local features of the second sub-feature map to obtain a third sub-feature map, including: inputting the second sub-feature map into the fourth convolutional layer to obtain a third output feature map, and performing dimensionality reduction on the third output feature map to obtain a fourth output feature map; inputting the second sub-feature map into the fifth convolutional layer to obtain a fifth output feature map, and inputting the fifth output feature map into the first Softmax layer so that the first Softmax layer transposes the fifth output feature map to obtain a sixth output feature map; inputting the second sub-feature map into the sixth convolutional layer to obtain a seventh output feature map, and inputting the seventh output feature map into the second Softmax layer so that the second Softmax layer transposes the seventh output feature map to obtain an eighth output feature map; multiplying the fourth output feature map and the sixth output feature map in matrix form to obtain a ninth output feature map, and multiplying the ninth output feature map and the eighth output feature map in matrix form to obtain a tenth output feature map; inputting the tenth output feature map into the seventh convolutional layer to obtain the third sub-feature map.
[0015] Optionally, the ConvB module includes a second layer normalization module, a first PWConv module, a gated linear unit, a second DWConv1D layer, a fifth batch normalization layer, a second Cswish activation layer, a second PWConv module, and a Dropout layer connected in sequence. The activation function of the second Cswish activation layer is the same as that of the first Cswish activation layer. The fourth sub-feature map is input into the ConvB module so that the ConvB module extracts the deep features of the fourth sub-feature map to obtain a fifth sub-feature map, including: inputting the fourth sub-feature map into the second layer normalization module so that the second layer normalization module performs layer normalization processing on the fourth sub-feature map to obtain a first intermediate feature map; inputting the first intermediate feature map into the first PWConv module so that the first PWConv module performs element-wise convolution on the first intermediate feature map to obtain a second intermediate feature map; inputting the second intermediate feature map into the gated linear unit to obtain a third intermediate feature map; inputting the third intermediate feature map into the second DWConv1D layer so that the second DWConv1D layer performs one-dimensional lightweight convolution processing on the third intermediate feature map to obtain a fourth intermediate feature map; inputting the fourth intermediate feature map into the fifth batch normalization layer so that the fifth batch normalization layer performs batch normalization processing on the fourth intermediate feature map to obtain a fifth intermediate feature map; inputting the fifth intermediate feature map into the second Cswish activation layer so that the second Cswish activation layer performs activation processing on the fourth intermediate feature map to obtain a sixth intermediate feature map; inputting the sixth intermediate feature map into the second PWConv module so that the second PWConv module performs element-wise convolution on the sixth intermediate feature map to obtain a seventh intermediate feature map; inputting the seventh intermediate feature map into the Dropout layer so that the Dropout layer performs regularization processing on the seventh intermediate feature map to obtain the fifth sub-feature map.
[0016] Optionally, the voiceprint monitoring model further includes a third layer normalization module, an AttentiveStatistic Pooling layer, a sixth batch normalization layer, and a second linear layer connected in sequence. The output end of the second linear layer is connected to the input end of the fully connected layer, and the input end of the third layer normalization module is connected to the output end of the splicing layer. Inputting the seventh feature map into the fully connected layer to obtain the monitoring result includes: inputting the seventh feature map into the third layer normalization module so that the third layer normalization module performs layer normalization processing on the seventh feature map to obtain an eighth feature map; inputting the eighth feature map into the AttentiveStatistic Pooling layer so that the Attentive Statistic Pooling layer captures the statistical features and dynamic changes of the eighth feature map through statistical pooling and attention mechanism to obtain a ninth feature map; inputting the ninth feature map into the sixth batch normalization layer so that the sixth batch normalization layer performs batch normalization processing on the ninth feature map to obtain a tenth feature map; inputting the tenth feature map into the second linear layer to obtain an embedded code; inputting the embedded code into the fully connected layer to obtain the monitoring result.
[0017] According to another aspect of the present application, there is provided an electronic device, including: one or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and configured to be executed by the one or more processors, and the one or more programs include those for executing any one of the methods for determining the posture of the disconnecting switch.
[0018] Applying the technical solution of the present application, first, obtain the sound emitted by the disconnector at the current moment to obtain voice data, and convert the voice data into a corresponding Mel spectrogram to obtain a voice Mel spectrogram. Finally, input the voice Mel spectrogram into the voiceprint monitoring model so that the voiceprint monitoring model analyzes the voice Mel spectrogram to obtain a monitoring result indicating that the posture of the disconnector at the current moment is a closed posture or an open posture. Among them, the voiceprint monitoring model includes a first SC attention mechanism layer and a fully connected layer. The first SC attention mechanism layer includes a global average pooling layer and a first Cswish layer. The global average pooling layer is used to extract the spatial features and channel features of the voice Mel spectrogram. The activation function of the first Cswish activation layer is a function obtained by improving the Swish function using exponential growth. Compared with the problem that it is difficult to ensure both a relatively fast monitoring speed and a relatively high monitoring accuracy when the voiceprint recognition model in the prior art monitors the posture of the disconnector, in the voiceprint monitoring model of the present application, the activation function of the first Cswish activation layer helps to avoid the problem of gradient disappearance by introducing the non-linear part of exponential growth, enabling the model to better learn deep features during the training process, thereby improving the recognition accuracy of the disconnector voiceprint monitoring model. In addition, since the activation function of the first Cswish activation layer can maintain a large negative gradient value, it can accelerate the propagation of the gradient during the backpropagation process, making the model training converge faster, thereby improving the running speed of the model. BRIEF DESCRIPTION OF THE DRAWINGS
[0019] The accompanying drawings forming a part of this application are used to provide a further understanding of this application. The schematic embodiments of this application and their descriptions are used to explain this application and do not constitute an improper limitation of this application. In the drawings:
[0020] Figure 1 The hardware structure block diagram of a mobile terminal showing a method for determining the posture of a disconnector provided in an embodiment of the present application is shown;
[0021] Figure 2 The flowchart showing a method for determining the posture of a disconnector provided in an embodiment of the present application is shown;
[0022] Figure 3 The structural schematic diagram of a specific voiceprint monitoring model provided in an embodiment of the present application is shown;
[0023] Figure 4 The structural schematic diagram of a first SC attention mechanism layer provided in an embodiment of the present application is shown;
[0024] Figure 5 The structural schematic diagram of a CFB module provided in an embodiment of the present application is shown;
[0025] Figure 6 The figure shows a schematic structural diagram of a first LFFB module provided according to an embodiment of the present application;
[0026] Figure 7 The figure shows a schematic structural diagram of a CVA Attention module provided according to an embodiment of the present application;
[0027] Figure 8 The figure shows a schematic structural diagram of a ConvB module provided according to an embodiment of the present application.
[0028] Among them, the above-mentioned drawings include the following reference numerals:
[0029] 102, processor; 104, memory; 106, transmission device; 108, input / output device. Detailed implementation manners
[0030] It should be noted that, without conflict, the embodiments in the present application and the features in the embodiments may be combined with each other. The present application will be described in detail below with reference to the drawings and in combination with the embodiments.
[0031] In order to enable those skilled in the art to better understand the solution of the present application, the technical solutions in the embodiments of the present application will be clearly and completely described below with reference to the drawings in the embodiments of the present application. Obviously, the described embodiments are only a part of the embodiments of the present application, rather than all the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those of ordinary skill in the art without creative efforts shall fall within the protection scope of the present application.
[0032] It should be noted that the terms "first", "second", etc. in the specification and claims of the present application and the above-mentioned drawings are used to distinguish similar objects, and do not necessarily need to be used to describe a specific order or sequence. It should be understood that such used data may be interchanged under appropriate circumstances so as to describe the embodiments of the present application herein. In addition, the terms "comprising" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product or device including a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these process, method, product or device.
[0033] As introduced in the background art, when the prior art voiceprint recognition model monitors the posture of the disconnecting switch, it is difficult to ensure both a relatively fast monitoring speed and a relatively high monitoring accuracy. To solve the above problems, the embodiments of the present application provide a method for determining the posture of a disconnecting switch and an electronic device.
[0034] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention.
[0035] The method embodiments provided in the embodiments of this application can be executed on a mobile terminal, a computer terminal, or a similar computing device. Taking running on a mobile terminal as an example, Figure 1 is a hardware structure block diagram of a mobile terminal for a method of determining the attitude of a disconnect switch according to an embodiment of the present invention. As Figure 1 shown, the mobile terminal may include one or more ( Figure 1 only one is shown in the figure) processors 102 (the processor 102 may include, but is not limited to, a processing device such as a microprocessor MCU or a programmable logic device FPGA) and a memory 104 for storing data. Among them, the above-mentioned mobile terminal may further include a transmission device 106 for communication functions and an input / output device 108. Those of ordinary skill in the art can understand that Figure 1 the structure shown is only schematic and does not limit the structure of the above-mentioned mobile terminal. For example, the mobile terminal may further include more or fewer components than Figure 1 shown in the figure, or have a different configuration from Figure 1 shown in the figure.
[0036] The memory 104 can be used to store computer programs. For example, software programs and modules of application software, such as the computer program corresponding to the method of determining the attitude of the disconnect switch in the embodiments of the present invention. The processor 102 executes various functional applications and data processing by running the computer program stored in the memory 104, that is, implements the above method. The memory 104 may include a high-speed random access memory, and may also include a non-volatile memory, such as one or more magnetic storage devices, flash memories, or other non-volatile solid-state memories. In some instances, the memory 104 may further include a memory remotely disposed relative to the processor 102, and these remote memories may be connected to the mobile terminal through a network. Examples of the above networks include, but are not limited to, the Internet, an enterprise intranet, a local area network, a mobile communication network, and combinations thereof. The transmission device 106 is used to receive or send data via a network. Specific examples of the above network may include a wireless network provided by a communication provider of the mobile terminal. In one instance, the transmission device 106 includes a network adapter (abbreviated as NIC), which can be connected to other network devices through a base station and thus communicate with the Internet. In one instance, the transmission device 106 may be a radio frequency (abbreviated as RF) module, which is used to communicate with the Internet wirelessly.
[0037] In this embodiment, a method for determining the attitude of a disconnecting switch operating on a mobile terminal, a computer terminal, or a similar computing device is provided. It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0038] Figure 2 is a flowchart of a method for determining the attitude of a disconnecting switch according to an embodiment of the present application. As Figure 2 shown, the method includes the following steps:
[0039] Step S201, obtain the sound emitted by the disconnecting switch at the current moment, obtain voice data, and convert the above voice data into a corresponding Mel spectrogram to obtain a voice Mel spectrogram;
[0040] Step S202, input the above voice Mel spectrogram into the voiceprint monitoring model, so that the voiceprint monitoring model analyzes the above voice Mel spectrogram to obtain a monitoring result. The above monitoring result is a result indicating that the attitude of the disconnecting switch at the current moment is a closed attitude or an open attitude. Among them, the above voiceprint monitoring model includes a first SC attention mechanism layer and a fully connected layer. The first SC attention mechanism layer includes a global average pooling layer and a first Cswish layer. The output end of the global average pooling layer is connected to the input end of the first Cswish activation layer. The global average pooling layer is used to extract the spatial features and channel features of the voice Mel spectrogram. The activation function of the first Cswish activation layer is a function obtained by improving the Swish function using exponential growth. The output end of the first Cswish activation layer is connected to the input end of the fully connected layer. The fully connected layer is used to output the above monitoring result.
[0041] Through the above embodiments, first, the sound emitted by the disconnector at the current moment is acquired to obtain voice data, and the voice data is converted into a corresponding Mel spectrogram to obtain a voice Mel spectrogram. Finally, the voice Mel spectrogram is input into the voiceprint monitoring model so that the voiceprint monitoring model analyzes the voice Mel spectrogram to obtain a monitoring result indicating that the posture of the disconnector at the current moment is a closed posture or an open posture. Among them, the voiceprint monitoring model includes a first SC attention mechanism layer and a fully connected layer. The first SC attention mechanism layer includes a global average pooling layer and a first Cswish layer. The global average pooling layer is used to extract the spatial features and channel features of the voice Mel spectrogram. The activation function of the first Cswish activation layer is a function obtained by improving the Swish function using exponential growth. Compared with the problem in the prior art that it is difficult to ensure both a relatively fast monitoring speed and a relatively high monitoring accuracy when the voiceprint recognition model monitors the posture of the disconnector, in the voiceprint monitoring model of the present application, the activation function of the first Cswish activation layer helps to avoid the problem of gradient disappearance by introducing the non-linear part of exponential growth, enabling the model to better learn deep features during the training process, thereby improving the recognition accuracy of the disconnector voiceprint monitoring model. In addition, since the activation function of the first Cswish activation layer can maintain a large negative gradient value, it can accelerate the propagation of the gradient during the backpropagation process, making the model training converge faster, thereby improving the running speed of the model.
[0042] Specifically, to acquire the sound emitted by the disconnector at the current moment to obtain voice data and convert the above voice data into a corresponding Mel spectrogram to obtain a voice Mel spectrogram, the specific steps can be as follows: Use a microphone in the substation to record and collect the sound (i.e., voice segment) of the disconnector. Each voice segment is 10 seconds (when the voice segment is less than 10 seconds, the voice segment is repeatedly spliced until the voice length reaches 10 seconds) to obtain the initial voice data. Then, noise is added to 20% of the initial voice data, echo is added to 30% of the initial voice data, and both noise and echo are added to 5% of the initial voice data (i.e., the proportion of the initial voice data containing both noise and echo, the initial voice data containing noise, the initial voice data containing echo, and the original initial voice data is 1:4:6:9) to obtain the voice data. Then, a series of processes such as pre-emphasis, framing, windowing, fast Fourier transform, and Mel filter bank filtering are performed on the voice data to obtain the voice Mel spectrogram.
[0043] Specifically, converting voice data into a Mel spectrogram is a common operation in the field of audio processing. The voice Mel spectrogram is obtained by performing a fast Fourier transform (FFT) on the sound generated during the operation of the disconnector and mapping it to the Mel scale to capture the basic texture and features of the sound.
[0044] In an alternative solution, as Figure 3 and Figure 4 shown, the above voiceprint monitoring model further includes a first convolutional layer and a first MobileNetV2 layer. The output end of the first convolutional layer is connected to the input end of the first MobileNetV2 layer. The output end of the first MobileNetV2 layer is connected to the input end of the first SC attention mechanism layer. The first SC attention mechanism layer further includes a second convolutional layer, a first batch normalization layer, a ReLU layer, and a third convolutional layer connected in sequence. The input end of the second convolutional layer is connected to the output end of the global average pooling layer. The output end of the third convolutional layer is connected to the input end of the first Cswish activation layer. The first SC attention mechanism layer further includes a product layer. The input ends of the product layer are respectively connected to the output end of the first Cswish activation layer and the output end of the first convolutional layer. The output end of the product layer is connected to the input end of the fully connected layer. Input the above voice mel spectrogram into the voiceprint monitoring model so that the voiceprint monitoring model analyzes the above voice mel spectrogram to obtain a monitoring result, including: input the above voice mel spectrogram into the first convolutional layer so that the first convolutional layer extracts local features of the above voice mel spectrogram to obtain a first feature map Q1; input the first feature map Q1 into the above first MobileNetV2 layer so that the first MobileNetV2 layer extracts features of the first feature map through depthwise separable convolution to obtain a second feature map Q2; input the second feature map Q2 into the above global average pooling layer so that the global average pooling layer performs global average pooling on the second feature map to obtain a matrix z, where Q2 represents the second feature map, C represents the number of channels in the second feature map, F represents the number of frequency components in the second feature map, T represents the number of time components in the second feature map, c represents the serial number of the channel, t represents the serial number of the time component, and f represents the serial number of the frequency component; input the matrix z into the interaction layer so that the interaction layer processes the matrix z to obtain a target matrix, where the interaction layer includes the above second convolutional layer, the above first batch normalization layer, the above ReLU layer, and the above third convolutional layer connected in sequence; input the target matrix into the above first Cswish activation layer so that the first Cswish activation layer performs non-linear activation on the target matrix to obtain an attention weight, where the activation function of the above first Cswish activation layer is f(x) = m(max(α(e x-1), x)) + nx · sigmoid(x), where α is a hyperparameter, m is a preset first adaptive parameter, n is a preset second adaptive parameter, x is the input of the first Cswish activation layer, and f(x) is the output of the first Cswish activation layer; input the above attention weight and the second feature map into the above product layer, so that the product layer multiplies the attention weight and the second feature map element by element to obtain a third feature map Q3; input the third feature map Q3 into the above fully connected layer to obtain the above monitoring result.
[0045] In the above embodiments, through the combination of the first convolutional layer and the first MobileNetV2 layer, the model can effectively extract the local features and deep features of the voice mel spectrogram; the global average pooling layer can capture the global information of the second feature map; the interaction layer can further process and refine the global information; the first Cswish activation layer helps the model focus on the most important part of the input data; using the product layer to multiply the attention weight and the second feature map element by element, this operation actually combines the attention mechanism with the original feature map, enhancing the model's ability to distinguish features; the fully connected layer receives the third feature map as input and outputs the final monitoring result. This step usually involves classification or regression tasks to determine the posture of the disconnector, thus further ensuring a high accuracy.
[0046] Specifically, the first convolutional layer, the second convolutional layer, and the third convolutional layer are all common convolutional layers in the field of deep learning. The first MobileNetV2 layer is a common MobileNetV2 in the field of deep learning. The first batch normalization layer is a common batch normalization operation in the field of deep learning. The ReLU layer is a common ReLU in the field of deep learning. The fully connected layer is a common fully connected layer in the field of deep learning. Specifically, MobileNetV2 is a popular lightweight deep learning model. MobileNetV2 uses depthwise separable convolutions as its core building blocks. Through its innovative architecture design, MobileNetV2 significantly reduces the computational burden while maintaining the model performance.
[0047] Specifically, by processing the voice mel spectrogram through the first convolutional layer, the local features of the sound can be effectively extracted; by processing the first feature map through the first MobileNetV2 layer, the computational resources and parameter efficiency are optimized through depthwise separable convolutions. It helps extract important disconnector sound features while keeping the voiceprint monitoring model lightweight; by processing the second feature map through the first SC attention mechanism layer, the model's ability to identify key features is enhanced through global information focusing, thus obtaining a more representative third feature map.
[0048] Specifically, global average pooling is performed on the input second feature map using a global average pooling layer, and each channel and frequency are retained during the pooling process. An interaction layer is used to capture cross-channel interactions. Specifically, the formula for the attention weight is: ω = σ(Conv2(ReLU(BN(Conv1(z)))), where σ(·) is the Cswish function (i.e., the activation function of the first Cswish activation layer), BN represents batch normalization, and ω ∈ R C×F represents the attention weights calculated at different frequency and channel positions; the second convolutional layer and the third convolutional layer calculate the attention weights by aggregating adjacent channels and the frequencies of matrix z, and two Conv layers are stacked to enhance the learning ability. In the embodiments of the present application, the attention weight ω is the output f(x) of the first Cswish activation layer. Specifically, the hyperparameter α is a positive number, m + n = 1, and the magnitudes of m and n can be adjusted according to the training effect of the voiceprint monitoring model. Specifically, the activation function of the first Cswish activation layer introduces a non-linear part with exponential growth, which helps to avoid the problem of gradient disappearance and maintain a large negative gradient value.
[0049] According to some exemplary embodiments of the present application, such as Figure 3As shown, the above voiceprint monitoring model further includes a second MobileNetV2 layer, a second SC attention mechanism layer, a first linear layer, and a CFB combination layer connected in sequence. The above CFB combination layer includes multiple CFB modules connected in sequence. The input end of the second MobileNetV2 layer is connected to the output end of the first SC attention mechanism layer. The voiceprint monitoring model further includes a splicing layer. The input ends of the splicing layer are respectively connected to the output ends of each of the above CFB modules. The output end of the splicing layer is connected to the input end of the fully connected layer. The internal structure of the second SC attention mechanism layer is the same as that of the first SC attention mechanism layer. Inputting the third feature map Q3 into the fully connected layer to obtain the above monitoring result, including: inputting the third feature map Q3 into the second MobileNetV2 layer so that the second MobileNetV2 layer extracts features from the third feature map Q3 through depthwise separable convolution to obtain a fourth feature map Q4; inputting the fourth feature map Q4 into the second SC attention mechanism layer to obtain a fifth feature map Q5; inputting the fifth feature map Q5 into the first linear layer to obtain a sixth feature map Q6; inputting the sixth feature map Q6 into the CFB combination layer, and obtaining the feature maps output by each of the above CFB modules in the CFB combination layer, and inputting each of the above feature maps into the splicing layer so that the splicing layer performs a splicing operation on the multiple above feature maps to obtain a seventh feature map Q7, where the CFB combination layer is used to extract the global features and local features of the sixth feature map; inputting the seventh feature map Q7 into the fully connected layer to obtain the above monitoring result.
[0050] In the above embodiment, through the depthwise separable convolution of the second MobileNetV2 layer, the model can extract deeper features from the third feature map to form the fourth feature map, which further enhances the model's ability to capture the sound features of the disconnector and helps to further improve the accuracy of the monitoring result; the use of the second SC attention mechanism layer enables the model to pay more attention to the features in the fourth feature map that are more important for the disconnector posture judgment. In this way, the model can ignore some unimportant information and further improve the monitoring accuracy; through the multiple CFB modules of the CFB combination layer, the model can extract global features and local features at the same time, which helps the model to understand the input sound data from different perspectives and enhances the generalization ability of the model; the splicing layer splices the feature maps output by multiple CFB modules in the CFB combination layer to form the seventh feature map. This splicing operation helps to integrate the features extracted by different CFB modules, making the final monitoring result more comprehensive, thus further ensuring a relatively high monitoring accuracy.
[0051] Specifically, the second MobileNetV2 layer is the common MobileNetV2 in the field of deep learning, the first linear layer is the common linear layer in the field of deep learning, and the concatenation operation is the common Concat operation in the field of deep learning.
[0052] Specifically, the depthwise separable convolution of the second MobileNetV2 layer is used to further optimize and refine the features.
[0053] Specifically, as Figure 3 shown, the CFB combination layer includes L stacked CFB modules. In the embodiments of the present application, L is 4.
[0054] In some other alternative solutions, multiple CFB modules are arranged in sequence, the sixth feature map is input into the CFB combination layer, and the feature maps output by each CFB module in the CFB combination layer are obtained, including: a first input step of inputting the sixth feature map into the first CFB module to enable the first CFB module to extract the global features and local features of the sixth feature map to obtain the first feature map; a second input step of inputting the current feature map into the current CFB module to enable the current CFB module to extract the global features and local features of the current feature map to obtain the target feature map, where the current CFB module is any one of the remaining CFB modules in the CFB combination layer except the first CFB module, and the current feature map is the feature map output by the CFB module connected to the input end of the current CFB module; and repeating the second input step until the feature map output by the last CFB module is obtained.
[0055] In the above embodiments, first, the sixth feature map is input into the first CFB module, which is responsible for extracting global and local features and generating the first feature map. This step lays the foundation for subsequent processing. In the second input step, the current feature map (i.e., the output of the previous CFB module) is input into the current CFB module, and this process is repeated in each remaining CFB module in the CFB combination layer. This iterative process allows each CFB module to further optimize and refine feature extraction based on the previous module, thereby obtaining a higher-quality target feature map. By repeating the second input step until the last CFB module in the CFB combination layer is processed, this method ensures that the feature map can be fully processed when passing through each module. The finally output feature map contains the cumulative effect of all intermediate processing steps from the initial feature map to the final feature map. That is, through this hierarchical and iterative feature extraction process, the monitoring accuracy of the voiceprint monitoring model for the disconnector posture can be further improved because each CFB module can capture and strengthen features at different levels, making the final monitoring result more reliable.
[0056] Specifically, the sixth feature map is input into the stacked CFB modules. After passing through the first CFB module, the first feature map is output, and then the first feature map is input into the next CFB module, and so on. Each time a CFB module is passed through, the feature map is updated once. The feature maps output by each CFB module are subjected to a Concat operation to improve its ability to extract deep features at different layers. Connecting the feature maps of different layers can improve the performance of the voiceprint monitoring model.
[0057] In some alternative solutions, the internal structures of all the above CFB modules are the same. As Figure 5 shown, the above CFB module includes a first LFFB module, a CVAAttention module, a ConvB module, a second LFFB module, and a first layer normalization module connected in sequence. The above sixth feature map Q6 is input into the first of the above CFB modules to enable the first of the above CFB modules to extract the global and local features of the above sixth feature map and obtain the first of the above feature maps, including: inputting the above sixth feature map Q6 into the above first LFFB module to enable the above first LFFB module to perform feature extraction on the above sixth feature map Q6 through lightweight convolution to obtain a first sub-feature map Q 61 ; adding the above sixth feature map Q6 to half of the above first sub-feature map Q 61 to obtain a second sub-feature map Q 62 ; inputting the above second sub-feature map Q 62 into the above CVAAttention module to enable the above CVAAttention module to extract the above second sub-feature map Q62 The global features and local features to obtain the third sub-feature map Q 63 ; The above-mentioned third sub-feature map Q 63 and the above-mentioned second sub-feature map Q 62 are added together to obtain the fourth sub-feature map Q 64 ; The above-mentioned fourth sub-feature map Q 64 is input into the above-mentioned ConvB module so that the above-mentioned ConvB module extracts the deep features of the above-mentioned fourth sub-feature map Q 64 to obtain the fifth sub-feature map Q 65 ; The above-mentioned fifth sub-feature map Q 65 is added to the above-mentioned fourth sub-feature map Q 64 to obtain the sixth sub-feature map Q 66 ; The above-mentioned sixth sub-feature map Q 66 is input into the above-mentioned second LFFB module so that the above-mentioned second LFFB module extracts the features of the above-mentioned sixth sub-feature map Q 66 through lightweight convolution to obtain the seventh sub-feature map Q 67 ; The above-mentioned sixth sub-feature map Q 66 is added to half of the above-mentioned seventh sub-feature map Q 67 to obtain the eighth sub-feature map Q 68 ; The above-mentioned eighth sub-feature map Q 68 is input into the above-mentioned first layer normalization module so that the above-mentioned first layer normalization module performs layer normalization on the above-mentioned eighth sub-feature map Q 68 to obtain the first above-mentioned feature map Q 69 .
[0058] In the above embodiments, the CFB module combines the advantages of the Transformer and CNN (Convolutional Neural Network) models, captures content-based global features, and performs local feature extraction through multiple convolutional pooling operations, further ensuring that the monitoring results are relatively accurate, and the lightweight convolution of the first LFFB module and the CVAAttention module further ensures a relatively fast monitoring speed.
[0059] Specifically, the first layer normalization module is a common layer normalization operation in the field of deep learning.
[0060] In some exemplary embodiments, the internal structures of the above-mentioned first LFFB module and the above-mentioned second LFFB module are the same, such as Figure 6As shown in the figure, the above-mentioned first LFFB module includes a first combination layer, a second combination layer, and a third combination layer connected in sequence. The above-mentioned first combination layer includes a first Conv1D layer, a first GELU layer, and a second batch normalization layer connected in sequence. The above-mentioned second combination layer includes a first DWConv1D layer, a second GELU layer, and a third batch normalization layer connected in sequence. The above-mentioned third combination layer includes a second Conv1D layer, a third GELU layer, and a fourth batch normalization layer connected in sequence. Input the above-mentioned sixth feature map Q6 into the above-mentioned first LFFB module so that the above-mentioned first LFFB module performs feature extraction on the above-mentioned sixth feature map Q6 through lightweight convolution to obtain a first sub-feature map, including: input the above-mentioned sixth feature map Q6 into the above-mentioned first combination layer so that the above-mentioned first combination layer performs one-dimensional convolution processing, activation processing, and batch normalization processing on the above-mentioned sixth feature map Q6 to obtain a first output feature map; input the above-mentioned first output feature map into the above-mentioned second combination layer so that the above-mentioned second combination layer performs one-dimensional lightweight convolution processing, activation processing, and batch normalization processing on the above-mentioned first output feature map to obtain a second output feature map; input the above-mentioned second output feature map into the above-mentioned third combination layer so that the above-mentioned third combination layer performs one-dimensional convolution processing, activation processing, and batch normalization processing on the above-mentioned second output feature map to obtain the above-mentioned first sub-feature map Q 61 .
[0061] In the above embodiment, through the lightweight convolution design of the first LFFB module and the second LFFB module, efficient feature extraction of the sixth feature map is realized. This design reduces the amount of calculation, further improves the processing speed, and enables the model to respond more quickly.
[0062] Specifically, the first Conv1D layer and the second Conv1D layer are common one-dimensional convolution layers in the field of deep learning; the first GELU layer, the second GELU layer, and the third GELU layer are common Gaussian error linear units (i.e., Gaussian error linear unit activation functions) in the field of deep learning; the second batch normalization layer, the third batch normalization layer, and the fourth batch normalization layer are common batch normalization operations in the field of deep learning; the first DWConv1D layer is a common one-dimensional depth convolution layer in the field of deep learning.
[0063] Specifically, first, the sixth feature map is input into the first Conv1D layer for width convolution processing, then processed through the first GELU layer and the second batch normalization layer (i.e., first processed with the GELU activation function and then through batch normalization), then through the first DWConv1D layer for lightweight convolution, then through the second GELU layer and the third batch normalization layer, then through the second Conv1D layer for width convolution processing, then through the third GELU layer and the fourth batch normalization layer to obtain the first sub-feature map. The first LFFB module uses lightweight convolution, significantly reducing the number of parameters.
[0064] In some other exemplary embodiments, such as Figure 7 shown, the above CVAAttention module includes a fourth convolutional layer, a fifth convolutional layer, a sixth convolutional layer, a seventh convolutional layer, a first Softmax layer, and a second Softmax layer. Input the above second sub-feature map Q 62 into the above CVAAttention module so that the above CVAAttention module extracts the global and local features of the above second sub-feature map Q 62 to obtain a third sub-feature map Q 63 , including: inputting the above second sub-feature map Q 62 into the above fourth convolutional layer to obtain a third output feature map, and reducing the dimension of the above third output feature map to obtain a fourth output feature map; inputting the above second sub-feature map Q 62 into the above fifth convolutional layer to obtain a fifth output feature map, and inputting the above fifth output feature map into the above first Softmax layer so that the above first Softmax layer transposes the above fifth output feature map to obtain a sixth output feature map; inputting the above second sub-feature map Q 62 into the above sixth convolutional layer to obtain a seventh output feature map, and inputting the above seventh output feature map into the above second Softmax layer so that the above second Softmax layer transposes the above seventh output feature map to obtain an eighth output feature map; multiplying the above fourth output feature map and the above sixth output feature map in matrix to obtain a ninth output feature map, and multiplying the above ninth output feature map and the above eighth output feature map in matrix to obtain a tenth output feature map; inputting the above tenth output feature map into the above seventh convolutional layer to obtain the above third sub-feature map Q 63 .
[0065] In the above embodiments, through the processing of the CVAAttention module, global features and local features can be simultaneously extracted from the second sub-feature map. This feature extraction method helps to capture richer information, thereby further improving the accuracy of the model in judging the posture of the disconnector. The above CVAttention module is used to achieve the balance of the receptive field, considering both global and local information, and at the same time reducing the interference of irrelevant information.
[0066] Specifically, the fourth convolutional layer, the fifth convolutional layer, the sixth convolutional layer, and the seventh convolutional layer are common convolutional layers in the field of deep learning; the first Softmax layer and the second Softmax layer are common Softmax layers in the field of deep learning.
[0067] Specifically, the CVAAttention module introduces a channel attention mechanism, with a simple model and fast running speed.
[0068] In other embodiments, as Figure 8 shown, the above ConvB module includes a second normalization module, a first PWConv module, a gated linear unit, a second DWConv1D layer, a fifth batch normalization layer, a second Cswish activation layer, a second PWConv module, and a Dropout layer connected in sequence. The activation function of the above second Cswish activation layer is the same as the activation function of the above first Cswish activation layer. Input the above fourth sub-feature map Q 64 into the above ConvB module so that the above ConvB module extracts the deep features of the above fourth sub-feature map Q 64 to obtain a fifth sub-feature map, including: inputting the above fourth sub-feature map Q 64 into the above second normalization module so that the above second normalization module normalizes the above fourth sub-feature map Q 64Perform layer normalization processing to obtain a first intermediate feature map; input the above first intermediate feature map into the above first PWConv module so that the above first PWConv module performs element-wise convolution on the above first intermediate feature map to obtain a second intermediate feature map; input the above second intermediate feature map into the above gated linear unit to obtain a third intermediate feature map; input the above third intermediate feature map into the above second DWConv1D layer so that the above second DWConv1D layer performs one-dimensional lightweight convolution processing on the above third intermediate feature map to obtain a fourth intermediate feature map; input the above fourth intermediate feature map into the above fifth batch normalization layer so that the above fifth batch normalization layer performs batch normalization processing on the above fourth intermediate feature map to obtain a fifth intermediate feature map; input the above fifth intermediate feature map into the above second Cswish activation layer so that the above second Cswish activation layer performs activation processing on the above fourth intermediate feature map to obtain a sixth intermediate feature map; input the above sixth intermediate feature map into the above second PWConv module so that the above second PWConv module performs element-wise convolution on the above sixth intermediate feature map to obtain a seventh intermediate feature map; input the above seventh intermediate feature map into the above Dropout layer so that the above Dropout layer performs regularization processing on the above seventh intermediate feature map to obtain the above fifth sub-feature map Q 65 。
[0069] In the above embodiment, through the multi-layer processing of the ConvB module, deeper feature information can be extracted from the fourth sub-feature map. These information are crucial for the subsequent analysis of the monitoring results, further ensuring that the monitoring results are relatively accurate.
[0070] Specifically, the second layer normalization module is a common layer normalization operation in the field of deep learning; the first PWConv module and the second PWConv module are common element-wise convolutions (i.e., point convolutions) in the field of deep learning; the gated linear unit is a common gated linear unit GLU (Gated Linear Unit) in the field of deep learning; the second DWConv1D layer is a common one-dimensional depth convolution layer in the field of deep learning; the fifth batch normalization layer is a common batch normalization operation in the field of deep learning; the Dropout layer is a common regularization operation in the field of deep learning.
[0071] In other embodiments, such as Figure 3As shown, the above voiceprint monitoring model further includes a third layer normalization module, an Attentive Statistic Pooling layer, a sixth batch normalization layer, and a second linear layer connected in sequence. The output end of the second linear layer is connected to the input end of the fully connected layer. The input end of the third layer normalization module is connected to the output end of the splicing layer. The seventh feature map Q7 is input into the fully connected layer to obtain the above monitoring result, including: inputting the seventh feature map Q7 into the third layer normalization module so that the third layer normalization module performs layer normalization processing on the seventh feature map to obtain an eighth feature map Q8; inputting the eighth feature map Q8 into the Attentive Statistic Pooling layer so that the Attentive Statistic Pooling layer captures the statistical features and dynamic changes of the eighth feature map Q8 through statistical pooling and attention mechanism to obtain a ninth feature map Q9; inputting the ninth feature map Q9 into the sixth batch normalization layer so that the sixth batch normalization layer performs batch normalization processing on the ninth feature map Q9 to obtain a tenth feature map Q 10 ; inputting the tenth feature map Q 10 into the second linear layer to obtain an embedded code Q 11 ; inputting the embedded code Q 11 into the fully connected layer to obtain the above monitoring result.
[0072] In the above embodiment, through the sequential processing of the third layer normalization module, the Attentive Statistic Pooling layer, the sixth batch normalization layer, and the second linear layer, the model can more effectively capture and represent the statistical features and dynamic changes of the input feature map. This processing method helps to extract richer and more discriminative features, thereby further improving the recognition accuracy of the model for the posture of the disconnector.
[0073] Specifically, the embedded code is converted into the final classification output through the fully connected layer, and it is predicted whether the disconnector is in the closed posture or the open posture according to the extracted features.
[0074] Specifically, the voiceprint monitoring model in this application is a model obtained after training using a sample voice data set. The sample voice data set is collected by collecting voice segments when the historical disconnector is in the closed position and the open position. The specific training process is as follows: First, preprocess the sample voice data set to obtain sample voice Mel spectrogram data. Then, divide the sample voice Mel spectrogram data into a training set, a validation set, and a test set according to a certain ratio. The training set, the validation set, and the test set are mutually exclusive data. The training set accounts for 70% of the sample voice Mel spectrogram data, and the validation set and the test set each account for 15% of the sample voice Mel spectrogram data. Then, use the training set, the validation set, and the test set to train and validate the untrained voiceprint monitoring model. Set the margin of AM-softmax to 0.2 and the scale to 30, which is used as the recognition loss. The loss is optimized by the AdamW optimizer, the initial learning rate is 0.001, and the weight decay is 10 -7 ; For the first 2000 steps, the learning rate is warmed up, and the learning rate is halved every 4 epochs. The batch size for training is 152, and the trained voiceprint monitoring model is obtained.
[0075] It should be noted that the steps shown in the flowchart of the accompanying drawings can be executed in a computer system such as a set of computer-executable instructions. And although the logical order is shown in the flowchart, in some cases, the steps shown or described can be executed in a different order than here.
[0076] The embodiment of this application also provides an electronic device, including: one or more processors, a memory, and one or more programs. Among them, the above one or more programs are stored in the above memory and are configured to be executed by the above one or more processors. The above one or more programs include a method for determining the posture of any of the above disconnectors.
[0077] Obviously, those skilled in the art should understand that the above-mentioned modules or steps of the present invention can be implemented by a general-purpose computing device. They can be concentrated on a single computing device or distributed on a network composed of multiple computing devices. They can be implemented by program codes executable by the computing device. Thus, they can be stored in a storage device and executed by the computing device. And in some cases, the steps shown or described can be executed in a different order than here, or they can be separately made into individual integrated circuit modules, or multiple modules or steps among them can be made into a single integrated circuit module to implement. In this way, the present invention is not limited to any specific combination of hardware and software.
[0078] Those skilled in the art should understand that the embodiments of the present application can be provided as a method, a system, or a computer program product. Therefore, the present application can take the form of a complete hardware embodiment, a complete software embodiment, or an embodiment combining software and hardware aspects. Moreover, the present application can take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to disk memory, CD-ROM, optical memory, etc.) that contain computer-usable program code.
[0079] The present application is described with reference to the flowcharts and / or block diagrams of methods, apparatuses (systems), and computer program products according to the embodiments of the present application. It should be understood that each flow and / or block in the flowchart and / or block diagram, and the combination of flows and / or blocks in the flowchart and / or block diagram, can be realized by computer program instructions. These computer program instructions can be provided to the processor of a general-purpose computer, a special-purpose computer, an embedded processor, or other programmable data processing devices to generate a machine, such that the instructions executed by the processor of the computer or other programmable data processing devices generate means for realizing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0080] These computer program instructions can also be stored in a computer-readable memory that can direct a computer or other programmable data processing device to work in a specific manner, such that the instructions stored in the computer-readable memory generate a manufactured article including instruction means that realizes the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0081] These computer program instructions can also be loaded onto a computer or other programmable data processing device, such that a series of operation steps are executed on the computer or other programmable device to generate a computer-implemented process, and thus the instructions executed on the computer or other programmable device provide steps for realizing the functions specified in Figure 1 one or more flows and / or blocks Figure 1 one or more blocks.
[0082] In a typical configuration, a computing device includes one or more processors (CPUs), an input / output interface, a network interface, and memory.
[0083] The memory may include non-permanent memory in the computer-readable medium, in the form of random access memory (RAM) and / or non-volatile memory, such as read-only memory (ROM) or flash memory (flash RAM). The memory is an example of a computer-readable medium.
[0084] A computer-readable medium includes both permanent and non-permanent, removable and non-removable media and can implement information storage by any method or technology. The information can be computer-readable instructions, data structures, program modules, or other data. Examples of computer storage media include, but are not limited to, phase change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other memory technologies, compact disc read-only memory (CD-ROM), digital versatile disc (DVD) or other optical storage, magnetic cassettes, tape disk storage or other magnetic storage devices, or any other non-transitory medium that can be used to store information that can be accessed by a computing device. As defined herein, a computer-readable medium does not include transitory computer-readable media such as modulated data signals and carrier waves.
[0085] It should also be noted that the term "comprising", "including" or any other variant thereof is intended to cover non-exclusive inclusion, such that a process, method, article or device comprising a series of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such process, method, article or device. Without further limitation, an element defined by the statement "comprising an..." does not exclude the presence of additional identical elements in the process, method, article or device comprising the element.
[0086] From the above description, it can be seen that the above embodiments of the present application achieve the following technical effects:
[0087] In the method for determining the attitude of the disconnector in this application, first, the sound emitted by the disconnector at the current moment is obtained to obtain voice data, and the voice data is converted into the corresponding Mel spectrogram to obtain the voice Mel spectrogram. Finally, the voice Mel spectrogram is input into the voiceprint monitoring model so that the voiceprint monitoring model analyzes the voice Mel spectrogram to obtain a monitoring result indicating that the attitude of the disconnector at the current moment is a closed attitude or an open attitude. Among them, the voiceprint monitoring model includes a first SC attention mechanism layer and a fully connected layer. The first SC attention mechanism layer includes a global average pooling layer and a first Cswish layer. The global average pooling layer is used to extract the spatial features and channel features of the voice Mel spectrogram. The activation function of the first Cswish activation layer is a function obtained by improving the Swish function using exponential growth. Compared with the problem that it is difficult to ensure both a relatively fast monitoring speed and a relatively high monitoring accuracy when the voiceprint recognition model in the prior art monitors the attitude of the disconnector, in the voiceprint monitoring model of this application, the activation function of the first Cswish activation layer helps to avoid the problem of gradient disappearance by introducing the non-linear part of exponential growth, enabling the model to better learn deep features during the training process, thereby improving the recognition accuracy of the disconnector voiceprint monitoring model. In addition, since the activation function of the first Cswish activation layer can maintain a relatively large negative gradient value, it can accelerate the propagation of the gradient during the backpropagation process, making the model training converge faster, thereby improving the running speed of the model.
[0088] The above are only the preferred embodiments of this application and are not used to limit this application. For those skilled in the art, this application can have various changes and modifications. Any modification, equivalent replacement, improvement, etc. made within the spirit and principle of this application shall be included within the protection scope of this application.
Claims
1. A method for determining the attitude of a disconnector, characterized in that Including: Obtain the sound emitted by the disconnector at the current moment to obtain voice data, and convert the voice data into a corresponding Mel spectrogram to obtain a voice Mel spectrogram; Input the voice Mel spectrogram into the voiceprint monitoring model, so that the voiceprint monitoring model analyzes the voice Mel spectrogram to obtain a monitoring result, where the monitoring result is a result indicating that the posture of the disconnector at the current moment is a closed posture or an open posture. Among them, the voiceprint monitoring model includes a first SC attention mechanism layer and a fully connected layer. The first SC attention mechanism layer includes a global average pooling layer and a first Cswish activation layer. The output end of the global average pooling layer is connected to the input end of the first Cswish activation layer. The global average pooling layer is used to extract the spatial features and channel features of the voice Mel spectrogram. The activation function of the first Cswish activation layer is a function obtained by improving the Swish function using exponential growth. The output end of the first Cswish activation layer is connected to the input end of the fully connected layer. The fully connected layer is used to output the monitoring result. Input the voice Mel spectrogram into the voiceprint monitoring model, so that the voiceprint monitoring model analyzes the voice Mel spectrogram to obtain a monitoring result, including: Input the voice Mel spectrogram into the first convolutional layer, so that the first convolutional layer extracts the local features of the voice Mel spectrogram to obtain a first feature map; Input the first feature map into the first MobileNetV2 layer, so that the first MobileNetV2 layer extracts features from the first feature map through depthwise separable convolution to obtain a second feature map; Input the second feature map into the global average pooling layer, so that the global average pooling layer performs global average pooling on the second feature map to obtain a matrix z, where, , , Q2 represents the second feature map, C represents the number of channels in the second feature map, F represents the number of frequency components in the second feature map, T represents the number of time components in the second feature map, c represents the serial number of the channel, t represents the serial number of the time component, and f represents the serial number of the frequency component; Input the matrix z into the interaction layer, so that the interaction layer processes the matrix z to obtain a target matrix, where the interaction layer includes a second convolutional layer, a first batch normalization layer, a ReLU layer, and a third convolutional layer connected in sequence; Input the target matrix into the first Cswish activation layer, so that the first Cswish activation layer performs nonlinear activation on the target matrix to obtain an attention weight; Input the attention weight and the second feature map into the product layer, so that the product layer multiplies the attention weight and the second feature map element by element to obtain a third feature map; Input the third feature map into the fully connected layer to obtain the monitoring result.
2. The method for determining the attitude of the disconnecting switch according to claim 1, wherein The voiceprint monitoring model further includes the first convolutional layer and the first MobileNetV2 layer. The output end of the first convolutional layer is connected to the input end of the first MobileNetV2 layer. The output end of the first MobileNetV2 layer is connected to the input end of the first SC attention mechanism layer. The first SC attention mechanism layer further includes the second convolutional layer, the first batch normalization layer, the ReLU layer, and the third convolutional layer that are connected in sequence. The input end of the second convolutional layer is connected to the output end of the global average pooling layer. The output end of the third convolutional layer is connected to the input end of the first Cswish activation layer. The first SC attention mechanism layer further includes the product layer. The input ends of the product layer are respectively connected to the output end of the first Cswish activation layer and the output end of the first convolutional layer. The output end of the product layer is connected to the input end of the fully connected layer. The activation function of the first Cswish activation layer is , is a hyperparameter, m is a preset first adaptive parameter, n is a preset second adaptive parameter, x is the input of the first Cswish activation layer, and f(x) is the output of the first Cswish activation layer.
3. The method for determining the attitude of the disconnector according to claim 2, wherein The voiceprint monitoring model further includes a second MobileNetV2 layer, a second SC attention mechanism layer, a first linear layer, and a CFB combination layer connected in sequence. The CFB combination layer includes a plurality of CFB modules connected in sequence. The input end of the second MobileNetV2 layer is connected to the output end of the first SC attention mechanism layer. The voiceprint monitoring model further includes a splicing layer. The input ends of the splicing layer are respectively connected to the output ends of the CFB modules. The output end of the splicing layer is connected to the input end of the fully connected layer. The internal structures of the second SC attention mechanism layer and the first SC attention mechanism layer are the same. Inputting the third feature map into the fully connected layer to obtain the monitoring result, including: Inputting the third feature map into the second MobileNetV2 layer, so that the second MobileNetV2 layer extracts features of the third feature map through depthwise separable convolution to obtain a fourth feature map; Inputting the fourth feature map into the second SC attention mechanism layer to obtain a fifth feature map; Inputting the fifth feature map into the first linear layer to obtain a sixth feature map; Inputting the sixth feature map into the CFB combination layer, and obtaining the feature maps output by the CFB modules in the CFB combination layer, and inputting the feature maps into the splicing layer, so that the splicing layer performs a splicing operation on the plurality of feature maps to obtain a seventh feature map, where the CFB combination layer is used to extract the global features and local features of the sixth feature map; Inputting the seventh feature map into the fully connected layer to obtain the monitoring result.
4. The method for determining the attitude of the disconnecting switch according to claim 3, characterized in that The plurality of CFB modules are arranged in order. Inputting the sixth feature map into the CFB combination layer, and obtaining the feature maps output by the CFB modules in the CFB combination layer, including: A first input step of inputting the sixth feature map into the first CFB module, so that the first CFB module extracts the global features and local features of the sixth feature map to obtain the first feature map; A second input step of inputting the current feature map into the current CFB module, so that the current CFB module extracts the global features and local features of the current feature map to obtain a target feature map, where the current CFB module is any one of the remaining CFB modules in the CFB combination layer except the first CFB module, and the current feature map is the feature map output by the CFB module connected to the input end of the current CFB module; Repeating the second input step until the feature map output by the last CFB module is obtained.
5. The method for determining the attitude of the disconnector according to claim 4, characterized in that All the internal structures of the CFB modules are the same. The CFB module includes a first LFFB module, a CVA Attention module, a ConvB module, a second LFFB module, and a first layer normalization module connected in sequence. The sixth feature map is input into the first CFB module to enable the first CFB module to extract the global features and local features of the sixth feature map, obtaining the first feature map, including: Input the sixth feature map into the first LFFB module to enable the first LFFB module to perform feature extraction on the sixth feature map through lightweight convolution, obtaining a first sub-feature map; Add the sixth feature map to half of the first sub-feature map to obtain a second sub-feature map; Input the second sub-feature map into the CVA Attention module to enable the CVA Attention module to extract the global features and local features of the second sub-feature map, obtaining a third sub-feature map; Add the third sub-feature map to the second sub-feature map to obtain a fourth sub-feature map; Input the fourth sub-feature map into the ConvB module to enable the ConvB module to extract the deep features of the fourth sub-feature map, obtaining a fifth sub-feature map; Add the fifth sub-feature map to the fourth sub-feature map to obtain a sixth sub-feature map; Input the sixth sub-feature map into the second LFFB module to enable the second LFFB module to perform feature extraction on the sixth sub-feature map through lightweight convolution, obtaining a seventh sub-feature map; Add the sixth sub-feature map to half of the seventh sub-feature map to obtain an eighth sub-feature map; Input the eighth sub-feature map into the first layer normalization module to enable the first layer normalization module to perform layer normalization processing on the eighth sub-feature map, obtaining the first feature map.
6. The method for determining the attitude of the disconnector according to claim 5, characterized in that, The internal structures of the first LFFB module and the second LFFB module are the same. The first LFFB module includes a first combination layer, a second combination layer, and a third combination layer connected in sequence. The first combination layer includes a first Conv1D layer, a first GELU layer, and a second batch normalization layer connected in sequence. The second combination layer includes a first DWConv1D layer, a second GELU layer, and a third batch normalization layer connected in sequence. The third combination layer includes a second Conv1D layer, a third GELU layer, and a fourth batch normalization layer connected in sequence. Input the sixth feature map into the first LFFB module to enable the first LFFB module to perform feature extraction on the sixth feature map through lightweight convolution, obtaining a first sub-feature map, including: Input the sixth feature map into the first combination layer to enable the first combination layer to perform one-dimensional convolution processing, activation processing, and batch normalization processing on the sixth feature map, obtaining a first output feature map; Input the first output feature map into the second combination layer, so that the second combination layer performs one-dimensional lightweight convolution processing, activation processing, and batch normalization processing on the first output feature map to obtain a second output feature map; Input the second output feature map into the third combination layer, so that the third combination layer performs one-dimensional convolution processing, activation processing, and batch normalization processing on the second output feature map to obtain the first sub-feature map.
7. The method for determining the attitude of the disconnector according to claim 5, characterized in that, The CVAAttention module includes a fourth convolutional layer, a fifth convolutional layer, a sixth convolutional layer, a seventh convolutional layer, a first Softmax layer, and a second Softmax layer. Input the second sub-feature map into the CVAAttention module, so that the CVAAttention module extracts the global feature and local feature of the second sub-feature map to obtain a third sub-feature map, including: Input the second sub-feature map into the fourth convolutional layer to obtain a third output feature map, and perform dimensionality reduction on the third output feature map to obtain a fourth output feature map; Input the second sub-feature map into the fifth convolutional layer to obtain a fifth output feature map, and input the fifth output feature map into the first Softmax layer, so that the first Softmax layer transposes the fifth output feature map to obtain a sixth output feature map; Input the second sub-feature map into the sixth convolutional layer to obtain a seventh output feature map, and input the seventh output feature map into the second Softmax layer, so that the second Softmax layer transposes the seventh output feature map to obtain an eighth output feature map; Perform matrix multiplication on the fourth output feature map and the sixth output feature map to obtain a ninth output feature map, and perform matrix multiplication on the ninth output feature map and the eighth output feature map to obtain a tenth output feature map; Input the tenth output feature map into the seventh convolutional layer to obtain the third sub-feature map.
8. The method for determining the attitude of the disconnector according to claim 5, characterized in that, The ConvB module includes a second layer normalization module, a first PWConv module, a gated linear unit, a second DWConv1D layer, a fifth batch normalization layer, a second Cswish activation layer, a second PWConv module, and a Dropout layer connected in sequence. The activation function of the second Cswish activation layer is the same as the activation function of the first Cswish activation layer. Input the fourth sub-feature map into the ConvB module, so that the ConvB module extracts the deep feature of the fourth sub-feature map to obtain a fifth sub-feature map, including: Input the fourth sub-feature map into the second layer normalization module, so that the second layer normalization module performs layer normalization processing on the fourth sub-feature map to obtain a first intermediate feature map; Input the first intermediate feature map into the first PWConv module, so that the first PWConv module performs element-wise convolution on the first intermediate feature map to obtain a second intermediate feature map; Input the second intermediate feature map into the gated linear unit to obtain a third intermediate feature map; Input the third intermediate feature map into the second 1D DWConv layer so that the second 1D DWConv layer performs one-dimensional lightweight convolution processing on the third intermediate feature map to obtain a fourth intermediate feature map; Input the fourth intermediate feature map into the fifth batch normalization layer so that the fifth batch normalization layer performs batch normalization processing on the fourth intermediate feature map to obtain a fifth intermediate feature map; Input the fifth intermediate feature map into the second Cswish activation layer so that the second Cswish activation layer performs activation processing on the fourth intermediate feature map to obtain a sixth intermediate feature map; Input the sixth intermediate feature map into the second PWConv module so that the second PWConv module performs element-wise convolution on the sixth intermediate feature map to obtain a seventh intermediate feature map; Input the seventh intermediate feature map into the Dropout layer so that the Dropout layer performs regularization processing on the seventh intermediate feature map to obtain the fifth sub-feature map.
9. The method for determining the attitude of the disconnector according to claim 3, characterized in that, The voiceprint monitoring model further includes a third layer normalization module, an Attentive Statistic Pooling layer, a sixth batch normalization layer, and a second linear layer connected in sequence. The output end of the second linear layer is connected to the input end of the fully connected layer, and the input end of the third layer normalization module is connected to the output end of the splicing layer. Input the seventh feature map into the fully connected layer to obtain the monitoring result, including: Input the seventh feature map into the third layer normalization module so that the third layer normalization module performs layer normalization processing on the seventh feature map to obtain an eighth feature map; Input the eighth feature map into the Attentive Statistic Pooling layer so that the Attentive Statistic Pooling layer captures the statistical features and dynamic changes of the eighth feature map through statistical pooling and attention mechanism to obtain a ninth feature map; Input the ninth feature map into the sixth batch normalization layer so that the sixth batch normalization layer performs batch normalization processing on the ninth feature map to obtain a tenth feature map; Input the tenth feature map into the second linear layer to obtain an embedded code; Input the embedded code into the fully connected layer to obtain the monitoring result.
10. An electronic device, characterized in that, Including: One or more processors, a memory, and one or more programs, wherein the one or more programs are stored in the memory and are configured to be executed by the one or more processors. The one or more programs include a method for determining the posture of the disconnecting switch according to any one of claims 1 to 9.
Citation Information
Patent Citations
Electric power communication system voiceprint recognition method and system based on EDRSN
CN116580714A
Power plant equipment state auditory monitoring method fused with frequency band self-downward attention mechanism
CN116825131A