Cable fault voiceprint monitoring method and system based on deep learning

Through the improvement of the Mel frequency cepspectrum sparseness and ResNet model, combined with the cross entropy loss function and the Adam optimization algorithm, the problem of identifying susceptible interference and high computational volume in cable fault monitoring is solved, and efficient and accurate soundprint monitoring of cable faults is achieved.

CN120236607AInactive Publication Date: 2025-07-01STATE GRID ZHEJIANG ELECTRIC POWER COMPANY TAIZHOU POWER SUPPLY

Patent Information

Application Number
CN202510715232.3
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-30
Publication Date
2025-07-01
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

In the monitoring of cable faults, the identification is susceptible to interference, the calculation is large and the solution is complex, making it difficult to efficiently and accurately extract cable fault characteristics, reducing the robustness in complex environments.

Method used

The feature extraction is performed using the sparseness of the Mel frequency cepspectrum, and the last layer of the ResNet model is replaced with a fully connected layer. The number of neurons matches the number of cable fault types. Combined with the cross entropy loss function and the Adam optimization algorithm, the model training process is optimized.

Benefits of technology

It improves robustness in complex environments, reduces the computational volume and solution complexity, and achieves efficient and accurate monitoring and early warning of cable fault soundprints.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120236607A_ABST
    Figure CN120236607A_ABST
Patent Text Reader

Abstract

The invention discloses a cable fault voiceprint monitoring method and system based on deep learning, and the method comprises the steps: obtaining sound data during the operation of a cable, carrying out the preprocessing, obtaining voiceprint data, carrying out the feature extraction through Mel-frequency cepstrum sparsity, replacing the last layer of a ResNet model with a new full-connection layer, and carrying out the detection of the voiceprint data. The number of neurons of the ResNet model is matched with the number of cable fault types needing to be classified, so that the ResNet model originally used for an image recognition task is completely suitable for voiceprint monitoring, the step of image conversion is not needed any more, robustness in a complex environment is improved, assistance of other models is not needed any more, and the calculated amount and the scheme complexity are reduced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of data processing, and particularly relates to a cable fault acoustic fingerprint monitoring method and system based on deep learning. Background Art

[0002] At present, significant progress has been made in the fields of speech and image processing with deep learning technology. Among them, acoustic fingerprint recognition, as a key technology, has made major breakthroughs in fields such as speech recognition and identity verification. Applying deep learning technology to acoustic fingerprint recognition in cable fault monitoring is expected to improve the accurate monitoring of cable status.

[0003] An acoustic fingerprint is an acoustic wave signal with certain information content collected by an acoustic measurement sensor. As an important part of power transmission, the acoustic fingerprint of a cable also has unique characteristics. The cable acoustic fingerprint contains a certain periodicity and stability, reflecting the status and specific information of the cable during operation. For example, there are obvious differences in the sounds between a cable operating normally and a cable with a fault. When a cable is operating normally, its sound exhibits stable characteristics with a certain frequency and amplitude. When a cable fails, the generated sound changes and loses the stability during normal operation. There are also differences in the sounds generated by different types of cable faults. If a high-frequency sharp "hissing" sound is emitted, it may indicate a cable insulation fault; if a sound similar to "crackling" or "exploding" is emitted, it may indicate a current short circuit in the cable, and if a sound similar to "impact" or "click" occurs, it may indicate that the cable has been mechanically impacted.

[0004] For example, the invention application with the publication (announcement) number CN114859269A discloses a cable fault diagnosis method based on acoustic fingerprint recognition technology. The collected signal is preprocessed, and a spectrogram is generated using the preprocessed signal. Then, the ResNet-18 is pre-trained using the ImageNet dataset to improve the model's feature extraction ability. Then, to make the model input unified to conform to the data format of ImageNet, the generated spectrogram set is converted into an RGB image as a training sample. Transfer learning is used to connect the ResNet-18 feature extraction layer to the deepRNN network to capture the information in the time series on the basis of obtaining the image features and establish a long-term dependence relationship in the time series. The signal converted into an RGB image is used as the training set and test set to train the model. Finally, the softmax layer is used by the model to judge whether there is a cable fault sound.

[0005] However, since ResNet was originally used for image recognition tasks and voiceprint recognition does not involve images, this solution generates spectrograms from voice signals and converts them into RGB images to adapt to ResNet. This process is prone to include some redundant information that is not directly related to cable fault recognition, reducing the robustness of this solution in complex environments. At the same time, converting to images also increases the computational complexity and the degree of complexity of the solution.

[0006] Therefore, in the process of cable fault monitoring, how to efficiently and accurately extract features directly related to cable faults, improve the robustness in complex environments, and minimize unnecessary computational complexity is a technical problem that urgently needs to be solved at present. Summary of the Invention

[0007] Aiming at the problems of existing technologies, such as vulnerable recognition to interference, large computational complexity, and complex solutions, the present invention provides a method and system for cable fault voiceprint monitoring based on deep learning. It uses Mel Frequency Cepstral Coefficients (MFCC) sparsity for feature extraction, and replaces the last layer of the ResNet model with a new fully connected layer, the number of neurons of which matches the number of cable fault types to be classified, making the ResNet model originally used for image recognition tasks fully applicable to voiceprint monitoring. Since there is no longer a need for the step of converting images, the robustness in complex environments is improved, and there is no longer a need for other model assistance, reducing the computational complexity and the degree of complexity of the solution.

[0008] The following is the technical solution of the present invention.

[0009] A method for cable fault voiceprint monitoring based on deep learning, comprising the following steps: S1: Obtain the sound data during cable operation and perform preprocessing to obtain voiceprint data; S2: Use Mel Frequency Cepstral Coefficients (MFCC) sparsity to extract features from the obtained voiceprint data, convert the voiceprint data from the time domain to the frequency domain to obtain feature vectors; S3: Input the feature vectors into a pre-trained ResNet model for pattern recognition, wherein the last layer of the ResNet model is replaced by a new fully connected layer, the number of neurons of which matches the number of cable fault types to be classified; S4: Judge and output the current operating state of the cable according to the pattern recognition result of the deep learning network.

[0010] In the present invention, Mel-frequency cepstral sparsity is used for feature extraction, and the last layer of the ResNet model is replaced with a new fully connected layer, the number of neurons of which matches the number of cable fault types to be classified, so that the ResNet model originally used for image recognition tasks is fully applicable to voiceprint monitoring. Since the step of converting images is no longer required, the robustness in complex environments is improved, and no other model assistance is required, reducing the computational amount and the complexity of the solution.

[0011] Preferably, S2: The obtained voiceprint data is subjected to feature extraction using Mel-frequency cepstral sparsity, and the voiceprint data is converted from the time domain to the frequency domain to obtain a feature vector, including: Define the frequency unit as Mel frequency, and convert the signal frequency of the voiceprint data to Mel frequency based on the conversion relationship; Perform DFT transformation on each frame of the voiceprint data under Mel frequency, so that the voiceprint data is converted from the time domain to the frequency domain, then perform power spectrum calculation and filter the power spectrum through a Mel filter bank to obtain a series of energy values of Mel frequency bands; Perform DCT transformation on a series of energy values of Mel frequency bands, concentrate the energy of the signal in the frequency domain, compress the signal while extracting features, remove redundant signals and reduce the data dimension to obtain MFCC feature parameters; Add information of the previous and next frames to the MFCC feature parameters to obtain a feature vector.

[0012] Preferably, the adding information of the previous and next frames to the MFCC feature parameters to obtain a feature vector includes: Add information of the previous and next frames to the MFCC feature parameters through second-order difference, and the expression is: ; where is the second-order difference value, and are the results of the previous and next DCT transformations respectively, that is, the MFCC feature parameters, n is the sampling point sequence value, and N is the total number of sampling points.

[0013] Preferably, in S3, the pre-training process of the ResNet model includes: Build the main structure of the ResNet model, including multiple residual blocks. Each residual block contains a Relu activation function, two convolutional layers and a skip connection. The skip connection allows the network to directly transmit information to subsequent layers, avoiding the disappearance of information in the network; Use historical data as training samples, use the cross-entropy loss function to train the model, and use the Adam optimization algorithm to optimize the model, and finally obtain the pre-trained ResNet model.

[0014] Preferably, in the ResNet model, the output of the residual block after convolution operation includes: Assume the input is , and the expected output is H( ), then: ; The data transfer of each residual unit in the residual block is: ; ; where F() is the residual function, representing the learned residual; f() is the Relu activation function; is the result after internal calculation of the residual unit, , respectively represent the inputs of the 1st, th residual units, W1 represents the weight parameter corresponding to the residual function, and the learned features from the shallow layer 1 to the deep layer L are obtained according to the following formula: ; where , respectively represent the inputs of the th, Lth residual units; In the ResNet model, the network is mainly divided into several stages, each stage contains multiple residual blocks. In each stage, the size of the feature map is reduced by half, and the number of channels is doubled. Features are extracted in this way of gradually increasing the number of channels and reducing the size of the feature map.

[0015] Preferably, the cross-entropy loss function is used to train the model, including: calculating the difference between the predicted probability of each sample and the true label using the cross-entropy loss function, and taking the average value to measure the difference between the model output and the true label, so as to prompt the model to learn to correctly classify sound events. During the process of minimizing the cross-entropy loss function, by adjusting the parameters of the model, the model can more accurately predict the category of sound events.

[0016] Preferably, the Adam optimization algorithm is used to optimize the model, including: Initializing the model parameters, and at each time step, calculating the gradient of the current iteration, that is, the partial derivative of the loss function with respect to the model parameters; Then calculating the first-order momentum term and the second-order momentum term based on the gradient and the exponential decay rate respectively; Performing bias correction to obtain the corrected first-order momentum term and second-order momentum term; Updating the value of the model parameters based on the corrected first-order momentum term, second-order momentum term and learning rate.

[0017] Preferably, the exponential decay rates of both the first-order momentum term and the second-order momentum term are both 0.5.

[0018] The present invention also provides a cable fault acoustic fingerprint monitoring system based on deep learning. When the system runs, it executes the above-mentioned cable fault acoustic fingerprint monitoring method based on deep learning.

[0019] The present invention also provides an electronic device, including a memory and a processor. A computer program is stored in the memory. When the processor calls the computer program in the memory, the steps of the above-mentioned cable fault acoustic fingerprint monitoring method based on deep learning are implemented.

[0020] The present invention also provides a storage medium. Computer-executable instructions are stored in the storage medium. When the computer-executable instructions are loaded and executed by a processor, the steps of the above-mentioned cable fault acoustic fingerprint monitoring method based on deep learning are implemented.

[0021] The substantial effects of the present invention include: The present invention extracts features through Mel Frequency Cepstral Coefficients (MFCC), effectively converts the acoustic fingerprint signal from the time domain to the frequency domain, and extracts the feature vectors crucial for cable fault recognition. This process not only removes redundant information but also significantly improves the sensitivity and recognition accuracy of cable fault features.

[0022] The ResNet model originally used for image recognition is innovatively applied to the field of acoustic fingerprint monitoring. By replacing the last layer with a fully connected layer matching the number of fault types, efficient pattern recognition of cable fault acoustic fingerprints is achieved. The residual learning mechanism of ResNet effectively avoids the problem of gradient disappearance, improving the training efficiency and recognition accuracy of the model.

[0023] Combined with the cross-entropy loss function and the Adam optimization algorithm, the present invention optimizes the training process of the model, enabling the model to converge to the optimal solution more quickly and having stronger generalization ability. This not only enhances the robustness of the model in complex environments but also ensures the effective recognition of the model for unknown cable fault types.

[0024] Therefore, through innovative technical means and methods, the present invention realizes the efficient and accurate monitoring and early warning of cable fault acoustic fingerprints, and solves the problems of susceptibility to interference in recognition, large computational amount, and complex solution existing in the prior art. BRIEF DESCRIPTION OF THE DRAWINGS

[0025] Figure 1 is a flowchart of an embodiment of the present invention; Figure 2 is a system block diagram of an embodiment of the present invention. DETAILED DESCRIPTION OF THE INVENTION

[0026] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the following will, in conjunction with the embodiments, clearly and completely describe the present technical solution. Obviously, the described embodiments are only a part rather than all of the embodiments of the present invention. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the scope of protection of the present invention.

[0027] It should be understood that in various embodiments of the present invention, the sequence numbers of the processes do not imply the order of execution, and the order of execution of the processes should be determined based on their functions and internal logic, and should not constitute any limitation to the implementation process of the embodiments of the present invention.

[0028] It should be understood that in the present invention, "including" and "having" and any variations thereof are intended to cover non-exclusive inclusion. For example, a process, method, system, product, or device that includes a series of steps or units does not necessarily have to be limited to those steps or units clearly listed, but may include other steps or units not clearly listed or inherent to these processes, methods, products, or devices.

[0029] It should be understood that in the present invention, "a plurality of" means two or more. "And / or" is merely a description of the association relationship of associated objects, indicating that three relationships may exist. For example, A and / or B may represent: A exists alone, A and B exist simultaneously, and B exists alone. The character " / " generally indicates that the associated objects before and after are in an "or" relationship. "Including A, B, and C" and "including A, B, C" mean that all of A, B, and C are included. "Including A, B, or C" means including any one of A, B, and C. "Including A, B, and / or C" means including any one, any two, or all three of A, B, and C.

[0030] The following will detail the technical solution of the present invention with specific embodiments. The embodiments can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments.

[0031] Embodiment 1: It should be noted that before implementing this embodiment, first build the operation environment for model operation. Use a computer as a server, build a cable sound data management database on the server, create data tables according to different cable sound data and fault types to save data, and transmit the cable sound fault data back to the server through communication and upload it to the server database. Then prepare a workbench with a Windows operating system locally, download the data on the server to the local workbench, and build a fault classification model using the Python language based on the dimension and length of the data. Build an automated script to classify and preprocess the data to form a training data set (the first training set), a validation data set (the second training set), and a test data set, and create a data input channel. Then use deep learning training to detect abnormal data.

[0032] A method for monitoring cable fault sound patterns based on deep learning according to this embodiment is as Figure 1 shown, and includes the following steps: S1: Obtain the sound data during cable operation and perform preprocessing to obtain sound pattern data.

[0033] Use a sound collector to obtain the sound data during cable operation, and construct a detection network with input channels and output channels corresponding to the cable data dimensions. Perform linear normalization on the data obtained using the sound, and use the following data: ; where represents the data after linear normalization for each dimension data, represents the original data for each dimension data, represents the maximum value in each dimension data, represents the minimum value in each dimension data. After linear normalization, the data for each dimension is mapped to the [0, 1] unified measurement space.

[0034] S2: Use Mel-frequency cepstrum sparsity to extract features from the obtained sound pattern data, convert the sound pattern data from the time domain to the frequency domain, and obtain feature vectors.

[0035] Define the frequency unit as Mel frequency, and convert the signal frequency of the sound pattern data to Mel frequency based on the conversion relationship; Perform DFT transformation on each frame of the sound pattern data at Mel frequency, so that the sound pattern data is converted from the time domain to the frequency domain. Subsequently, perform power spectrum calculation and filter the power spectrum through a Mel filter bank to obtain a series of energy values of Mel frequency bands; where, perform DFT on each frame of the signal: ; where h is the window function of the signal at point N, k is the DFT wrapping frequency, and then the power spectrum is calculated.

[0036] ; where N is the maximum value of the sampling points, is the power spectrum.

[0037] Create a Mel filter bank: ; where is the filter bank, is the frequency value of the sampling point.

[0038] Perform DCT transformation on the energy values of a series of Mel frequency bands, concentrate the energy of the signal in the frequency domain, compress the signal while extracting features, remove redundant signals and reduce the data dimension to obtain MFCC feature parameters; According to the definition of the cepstrum, perform the inverse Fourier transform and filter with a low-pass filter to obtain the final low-frequency signal. DCT is similar to DFT, the difference is that it only operates on the real part and does not involve complex operations. After the signal is transformed by DCT, the energy will be concentrated in the low-frequency part: ; where N represents the maximum value of the sampling points, are the MFCC feature parameters obtained after the discrete cosine transform.

[0039] Add information of the previous and next frames to the MFCC feature parameters through second-order difference, and the expression is: ; where is the second-order difference value, and are the results of the previous and next DCT transformations respectively, that is, the MFCC feature parameters, n is the sampling point sequence value, and N is the total number of sampling points.

[0040] S3: Input the feature vector into the pre-trained ResNet model for pattern recognition, where the last layer of the ResNet model is replaced by a new fully connected layer, and the number of its neurons matches the number of cable fault types to be classified.

[0041] Among them, the pre-training process of the ResNet model includes: Build the main structure of the ResNet model, including multiple residual blocks. Each residual block contains a Relu activation function, two convolutional layers and a skip connection. The skip connection allows the network to directly transmit information to the subsequent layers to avoid the disappearance of information in the network; Using historical data as training samples, a cross-entropy loss function is used to train the model, and the Adam optimization algorithm is used to optimize the model, and finally a pre-trained ResNet model is obtained.

[0042] Preferably, in the ResNet model, the output of the residual block after convolution operation includes: Assume the input is , and the expected output is H( ), then: ; The data transfer of each residual unit in the residual block is: ; ; where F() is the residual function, representing the learned residual; f() is the Relu activation function; is the result after internal calculation of the residual unit, , respectively represent the inputs of the 1st, th residual units, W1 represents the weight parameter corresponding to the residual function, and the learned features from the shallow layer 1 to the deep layer L are obtained according to the following formula: ; where , respectively represent the inputs of the th, Lth residual units; In the ResNet model, the network is mainly divided into several stages, each stage contains multiple residual blocks. In each stage, the size of the feature map is halved and the number of channels is doubled. Features are extracted in this way of gradually increasing the number of channels and decreasing the size of the feature map.

[0043] In this embodiment, a cross-entropy loss function is used to train the model, including: calculating the difference between the predicted probability of each sample and the true label using the cross-entropy loss function, and taking the average value to measure the difference between the model output and the true label, so as to prompt the model to learn to correctly classify sound events. In the process of minimizing the cross-entropy loss function, by adjusting the parameters of the model, the model can more accurately predict the category of sound events.

[0044] The specific formula is as follows: ; where, represents the true label, represents the predicted probability of the model. log represents the natural logarithm. The process of minimizing the cross-entropy loss function is to adjust the parameters of the model so that the model can more accurately predict the category of sound events.

[0045] In addition, this embodiment uses the Adam optimization algorithm to optimize the model, including: Initialize the model parameters. At each time step, calculate the gradient of the current iteration, that is, the partial derivative of the loss function with respect to the model parameters; After that, calculate the first-order momentum term and the second-order momentum term based on the gradient and the exponential decay rate respectively; Perform bias correction to obtain the corrected first-order momentum term and second-order momentum term; Update the values of the model parameters based on the corrected first-order momentum term, second-order momentum term, and learning rate.

[0046] Specifically, Adam optimization first calculates the gradient of the current iteration, that is, the partial derivative of the loss function with respect to the model parameters. The cross-entropy gradient value at the r-th iteration is: ; After that, introduce the first-order momentum term and the second-order momentum term , and the formula is: ; ; where the parameters , represent the exponential decay rates of the first-moment estimate and the second-moment estimate respectively. In the experiment, is set to 0.5, is set to 0.5. In the case of the initial time and a very small decay rate, the moment estimate value will tend to 0. The introduction of these two terms enables the Adam algorithm to adaptively adjust the learning rate, has better adaptability to the gradient changes of different parameters and at different time steps, and can converge to the optimal solution more effectively during the training process. However, due to possible estimation biases at the initial time and when the decay rate is small, bias correction is required to obtain more accurate first-order and second-order momentum terms.

[0047] ; ; At each iteration step, the value of the parameter needs to be updated, and the update expression of is: where η is the learning rate, representing the magnitude of the effective step size in the parameter space, represents a constant parameter.

[0048] S4: Judge and output the current operating state of the cable according to the pattern recognition result of the deep learning network.

[0049] In this embodiment, Mel-frequency cepstrum sparsity is used for feature extraction, and the last layer of the ResNet model is replaced with a new fully connected layer, the number of neurons of which matches the number of cable fault types to be classified, so that the ResNet model originally used for image recognition tasks is fully applicable to voiceprint monitoring. Since there is no longer a need for the step of converting images, the robustness in complex environments is improved, and no other model assistance is required, reducing the computational amount and the complexity of the solution.

[0050] Embodiment 2: A cable fault voiceprint monitoring system based on deep learning is used to execute the above-mentioned cable fault voiceprint monitoring method based on deep learning. As Figure 2 shown, this system includes: Voiceprint acquisition module: used to acquire the voiceprint signal of the target cable; Fault identification module: used to preprocess the acquired voiceprint signal to obtain voiceprint signal data, acquire the voiceprint signal data corresponding to the target cable and perform feature extraction, and identify the target cable fault type through a deep learning network model; Fault warning module: used to send an alarm to the monitoring center according to the judgment result.

[0051] Among them, the fault warning module includes: a WeChat mini-program for cable fault voiceprint monitoring, which receives audio in real time and displays the analysis results of the cable fault voiceprint monitoring system on the audio file.

[0052] In this embodiment, a WeChat mini-program is written in the WeChat developer tool environment. The mini-program connects to the hardware Bluetooth module through the Bluetooth interface and receives the audio collected by the cable fault voiceprint monitoring system in real time.

[0053] Unpack the received data, declare the received data as a global variable, and then convert it into a JSON object for display on the interface.

[0054] The WeChat mini-program mainly includes two parts: Bluetooth data reception and data processing and its visualization. The main process is that the WeChat mini-program visualization platform includes data upload and result analysis. The mini-program connects to the hardware Bluetooth module through the Bluetooth interface, receives audio in real time, and displays the analysis results of the audio file on the mini-program interface.

[0055] Among them, the Bluetooth receiving data first calls the wx.openBluetoothAdapter method to obtain the Bluetooth adapter object. Through the adapter object, the applet can complete the functions of searching, connecting and reading Bluetooth devices. Secondly, the startBluetoothDevicesDiscovery method is called in the Bluetooth adapter object to search for Bluetooth devices. As a result, the applet can see the device name, device ID and other information of nearby Bluetooth devices on the interface, so as to find the hardware Bluetooth module. After that, the createBLEConnection method is called in the Bluetooth adapter object, so that the applet can connect to the specified Bluetooth device and obtain the service list object contained in the connected Bluetooth module. Call the getBLEDeviceCharacteristics method to obtain the characteristic value of the Bluetooth device. Finally, the applet monitors the specified attributes of the specific characteristic value of the Bluetooth device. When the hardware device uploads data, the applet can use this characteristic value to return to the data uploaded by the hardware, which is convenient for subsequent data processing.

[0056] Since the hardware packages and sends the results in the format of a string, it cannot be directly displayed in the WeChat applet, so the received data must be processed so that the applet can display the analysis results. First, the applet receives the data uploaded by the hardware through the Bluetooth interface; secondly, the received data is declared as a global variable so that the data can be called in different interfaces of the applet to prepare for subsequent display; finally, in the js file of the display interface, the JSON.parse function is called to parse the received data into a JSON object to display the analysis results.

[0057] The cable fault voiceprint monitoring system based on deep learning research of the present invention uses Mel frequency cepstral coefficients to extract features of sound signals. MFCC selectively extracts features related to cable faults in voiceprint signals, while ignoring irrelevant voiceprint features. Redundant information is removed to extract the most important features of cable fault sound. The influence of noise on features is effectively suppressed, and the robustness to noise is improved. By performing discrete cosine transform (DCT) on the spectrum, the information in the frequency domain is converted to cepstral coefficients, thereby greatly reducing the dimension of the features, reducing the amount of calculation and storage space. At the same time, an 18-layer ResNet network is selected to train the extracted feature vectors, avoiding the gradient explosion problem in deep learning and improving the recognition efficiency. Finally, the cross entropy loss function is used to measure the difference between the model output and the true label, prompting the model to learn to correctly classify sound events. In addition, a WeChat applet is written in the WeChat developer tool environment, the analysis results sent by the hardware are received via Bluetooth, and the analysis results are visualized on the WeChat applet interface to display the fault condition of the cable, thereby achieving the purpose of early warning.

[0058] This embodiment is based on deep learning technology. It uses Mel Frequency Cepstral Coefficients (MFCC) to extract features from sound signals. MFCC selectively extracts the features related to cable faults in the voiceprint signals while ignoring the irrelevant voiceprint features. It removes redundant information and extracts the most important features of the cable fault sound. It effectively suppresses the influence of noise on the features and improves the robustness to noise. By performing Discrete Cosine Transform (DCT) on the spectrum, the information in the frequency domain is converted into cepstral coefficients, thus greatly reducing the dimension of the features, decreasing the computational amount and storage space. At the same time, the ResNet network is selected to train the extracted feature vectors, avoiding the problem of gradient explosion in deep learning and improving the recognition efficiency. Finally, the cross-entropy loss function is used to measure the difference between the model output and the true label, prompting the model to learn to correctly classify sound events. In addition, a WeChat mini-program is written in the WeChat developer tool environment. It receives the analysis results sent by the hardware through Bluetooth and visualizes the analysis results on the WeChat mini-program interface to display the fault situation of the cable, thus achieving the purpose of early warning.

[0059] In addition, this embodiment also provides an electronic device, including a memory and a processor. A computer program is stored in the memory. When the processor calls the computer program in the memory, the steps of the above-mentioned method for monitoring cable fault voiceprint based on deep learning are implemented.

[0060] This embodiment also provides a storage medium. A computer executable instruction is stored in the storage medium. When the computer executable instruction is loaded and executed by the processor, the steps of the above-mentioned method for monitoring cable fault voiceprint based on deep learning are implemented.

[0061] Therefore, generally speaking, the substantial effects of the above embodiments include: Through feature extraction using Mel Frequency Cepstral Coefficients (MFCC), the voiceprint signal is effectively converted from the time domain to the frequency domain, and the feature vectors crucial for cable fault recognition are extracted. This process not only removes redundant information but also significantly improves the sensitivity and recognition accuracy of cable fault features.

[0062] The ResNet model originally used for image recognition is innovatively applied to the field of voiceprint monitoring. By replacing the last layer with a fully connected layer matching the number of fault types, efficient pattern recognition of cable fault voiceprints is achieved. The residual learning mechanism of ResNet effectively avoids the problem of gradient vanishing, improving the training efficiency and recognition accuracy of the model.

[0063] Combined with the cross-entropy loss function and the Adam optimization algorithm, the present invention optimizes the training process of the model, enabling the model to converge to the optimal solution more quickly and possess stronger generalization ability. This not only enhances the robustness of the model in complex environments but also ensures the effective identification of unknown cable fault types by the model.

[0064] Through innovative technical means and methods, efficient and accurate monitoring and early warning of cable fault sound patterns are achieved, solving the problems of susceptibility to interference in identification, large computational volume, and complex solutions existing in the prior art.

[0065] From the description of the above embodiments, those skilled in the art can understand that, for the convenience and conciseness of description, only the above division of each functional module is used as an example. In actual applications, the above functions can be allocated to different functional modules according to needs, that is, the internal structure of a specific device is divided into different functional modules to complete all or part of the functions described above.

[0066] In the embodiments provided in the present application, it should be understood that the disclosed structure and method can be implemented in other ways. For example, the embodiments of the structure described above are only illustrative. For example, the division of modules or units is only a logical functional division. In actual implementation, there can be other division methods. For example, multiple units or components can be combined or integrated into another structure, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of structures or units can be in electrical, mechanical, or other forms.

[0067] The units described as separate components may or may not be physically separated. The components displayed as units can be one physical unit or multiple physical units, that is, they can be located in one place or distributed to multiple different places. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0068] In addition, each functional unit in the embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit. The above integrated units can be implemented in the form of hardware or in the form of software functional units.

[0069] When the integrated unit is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a readable storage medium. Based on this understanding, the technical solution of the embodiments of the present application, in essence, or the part that contributes to the prior art, or all or part of this technical solution, can be embodied in the form of a software product. This software product is stored in a storage medium and includes several instructions for causing a device (which can be a single-chip microcomputer, a chip, etc.) or a processor to execute all or part of the steps of the methods of the various embodiments of the present application. The aforementioned storage medium includes: various media that can store program codes, such as USB flash drives, mobile hard disks, read only memory (ROM), random access memory (RAM), magnetic disks, or optical discs.

[0070] The above content is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art within the technical scope disclosed by the present application can easily think of changes or substitutions, which should all be covered within the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A cable fault acoustic fingerprint monitoring method based on deep learning, characterized in that, It includes the following steps: S1: Obtain the sound data during cable operation and perform preprocessing to obtain voiceprint data; S2: Use Mel-frequency cepstral sparsity to extract features from the obtained voiceprint data, convert the voiceprint data from the time domain to the frequency domain, and obtain feature vectors; S3: Input the feature vectors into a pre-trained ResNet model for pattern recognition, where the last layer of the ResNet model is replaced by a new fully connected layer, and the number of its neurons matches the number of cable fault types to be classified; S4: Judge and output the current operating state of the cable according to the pattern recognition result of the deep learning network.

2. The method for monitoring cable fault sound patterns based on deep learning according to claim 1, wherein S2: Use Mel-frequency cepstral sparsity to extract features from the obtained voiceprint data, convert the voiceprint data from the time domain to the frequency domain, and obtain feature vectors, including: Define the frequency unit as Mel-frequency, and convert the signal frequency of the voiceprint data to Mel-frequency based on the conversion relationship; Perform DFT transformation on each frame of the voiceprint data under Mel-frequency, so that the voiceprint data is converted from the time domain to the frequency domain, then calculate the power spectrum and filter the power spectrum through a Mel filter bank to obtain a series of energy values of Mel frequency bands; Perform DCT transformation on a series of energy values of Mel frequency bands, concentrate the energy of the signal in the frequency domain, compress the signal while extracting features, remove redundant signals and reduce the data dimension to obtain MFCC feature parameters; Add the information of the previous and next frames to the MFCC feature parameters to obtain feature vectors.

3. The method for monitoring cable fault sound patterns based on deep learning according to claim 2, wherein, The adding the information of the previous and next frames to the MFCC feature parameters to obtain feature vectors includes: Add the information of the previous and next frames to the MFCC feature parameters through second-order difference, and the expression is: ; wherein is the second-order difference value, and are the results of the DCT transformations before and after respectively, that is, the MFCC feature parameters, n is the sampling point sequence value, and N is the total number of sampling points.

4. A method for monitoring cable fault sound patterns based on deep learning according to claim 1, characterized in that, In S3, the pre-training process of the ResNet model includes: Build the main structure of the ResNet model, including multiple residual blocks. Each residual block contains a Relu activation function, two convolutional layers and a skip connection. The skip connection allows the network to directly transmit information to subsequent layers to avoid the disappearance of information in the network; Use historical data as training samples, adopt the cross-entropy loss function to train the model, and adopt the Adam optimization algorithm to optimize the model, and finally obtain the pre-trained ResNet model.

5. The method for monitoring cable fault sound patterns based on deep learning according to claim 4, characterized in that, In the ResNet model, the output of the residual block after convolution operation includes: Assume the input is , and the expected output is H( ), then: ; The data transfer of each residual unit in the residual block is: ; ; where F() is the residual function representing the learned residual; f() is the Relu activation function; is the result after the internal calculation of the residual unit, and represent the inputs of the first and the second residual units respectively. W1 represents the weight parameter corresponding to the residual function. The learned features from the shallow layer 1 to the deep layer L are obtained according to the following formula: ; Among them and respectively represent the inputs of the -th and the L-th residual units; In the ResNet model, the network is mainly divided into several stages. Each stage contains multiple residual blocks. In each stage, the size of the feature map is reduced by half, and the number of channels is doubled. Features are extracted in this way of gradually increasing the number of channels and reducing the size of the feature map.

6. A method for monitoring cable fault acoustic fingerprints based on deep learning according to claim 4, characterized in that The adopting the cross-entropy loss function to train the model includes: calculating the difference between the predicted probability and the true label of each sample by using the cross-entropy loss function, and obtaining the average value, so as to measure the difference between the model output and the true label, and prompting the model to learn to correctly classify sound events. In the process of minimizing the cross-entropy loss function, by adjusting the parameters of the model, the model can more accurately predict the category of sound events.

7. A method for monitoring cable fault sound patterns based on deep learning according to claim 4, characterized in that The adopting the Adam optimization algorithm to optimize the model includes: Initialize the model parameters and calculate the gradient of the current iteration at each time step, that is, the partial derivative of the loss function with respect to the model parameters; Then the first-order momentum term and the second-order momentum term are calculated based on the gradient and the exponential decay rate respectively; Perform deviation correction to obtain the corrected first-order momentum term and second-order momentum term; Update the values ​​of model parameters based on the modified first-order momentum term, second-order momentum term and learning rate.

8. A method for monitoring cable fault sound patterns based on deep learning according to claim 7, characterized in that, The exponential decay rates of the first-order momentum term and the second-order momentum term are both 0.

5.

9. A cable fault acoustic fingerprint monitoring system based on deep learning, characterized in that, When the system is running, a cable fault voiceprint monitoring method based on deep learning as described in any one of claims 1 to 8 is executed.

10. An electronic device, characterized in that, It comprises a memory and a processor, wherein the memory stores a computer program, and when the processor calls the computer program in the memory, it implements the steps of a cable fault voiceprint monitoring method based on deep learning as described in any one of claims 1 to 8.

Citation Information

Patent Citations

  • Cable fault diagnosis method based on voiceprint recognition technology

    CN114859269A

  • High-voltage circuit breaker mechanical fault voiceprint recognition method based on fusion feature and residual neural network

    CN118609592A

  • Cable defect identification method and device based on multiple physical quantities

    CN119226866A

  • Transformer abnormal sound detection method, system and device in complex environment and storage medium

    CN119323972A

Cited By

  • Server hardware fault pre-detection system and monitoring method based on voiceprint recognition

    CN120950288A

  • Cable fault voiceprint monitoring and early warning method and system based on Bluetooth transmission

    CN121811922A