Weakly labeled data based abnormal detection method for voiceprint of power cable joint in power distribution network
By collecting and analyzing the acoustic signature data of cable joints in the distribution network, performing weak marker factor analysis and acoustic signature augmentation, and training the anomaly detector, the problem of low accuracy in acoustic signature anomaly detection of cable joints in the distribution network is solved, and real-time and accurate detection of the cable joint status is achieved.
Patent Information
- Application Number
- CN202511622023.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-11-07
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2045-11-07
AI Technical Summary
In existing technologies, the sound signature detection method for distribution network cable joints has low accuracy due to weak marker data, making it difficult to identify early hidden faults and prone to misjudgment or missed detection.
By collecting operational monitoring data of cable joints in the power distribution network, extracting sample acoustic signature feature sets, performing weak labeling factor analysis, configuring acoustic signature extension length, expanding sample acoustic signature feature sets, training cable joint anomaly detectors, and using machine learning models such as LSTM to capture the evolution trend of cable joint status and output anomaly rate.
It enables real-time and accurate detection of abnormalities in distribution network cable joints, improving detection accuracy and stability, and providing reliable decision-making basis for operation and maintenance.
Smart Images

Figure CN121096377B_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of cable anomaly detection, and in particular to a power distribution network cable joint voiceprint anomaly detection method based on weakly labeled data. BACKGROUND
[0002] The power distribution network cable joint is a key connecting component of the power system, and its operating state directly determines the stability of the power supply of the power distribution network. Once the power distribution network cable joint has faults such as joint loosening and insulation aging, it is easy to cause local power outages or even large-area power outages, which seriously affects industrial production and people's lives.
[0003] However, in actual voiceprint detection scenarios, the voiceprint features of the cable joint are easily disturbed by environmental noise, equipment errors, etc., resulting in a common "weakly labeled" problem in the labeled data. For example, the voiceprint features of early faults of the cable joint are very small different from the normal operating voiceprint, and a single voiceprint feature may correspond to a fuzzy state of a high probability of abnormality and a small probability of normality, which cannot be accurately defined by an absolute label. Such weakly labeled data can greatly reduce the accuracy of traditional detection methods, making it difficult to identify early hidden faults and easily leading to misjudgment / omission.
[0004] Therefore, there is an urgent need for a power distribution network cable joint voiceprint anomaly detection method that can specifically handle weakly labeled data to break through the technical bottleneck of traditional detection methods and provide reliable technical support for anomaly detection and fault handling of power distribution network cable joints. SUMMARY
[0005] The present application provides a power distribution network cable joint voiceprint anomaly detection method based on weakly labeled data to solve the technical problem of low accuracy of existing power distribution network cable joint voiceprint anomaly detection.
[0006] The technical solution of the present application to solve the above technical problem is as follows:
[0007] The present application provides a power distribution network cable joint voiceprint anomaly detection method based on weakly labeled data, comprising:
[0008] According to the operating monitoring data of the power distribution network cable joint, a sample voiceprint data set is collected, a sample detection result set is collected, and a plurality of sample voiceprint feature sets are extracted;
[0009] Weakly labeled factor analysis is performed on the plurality of sample voiceprint feature sets to obtain a plurality of weakly labeled factors, the sample detection result set is labeled, a plurality of voiceprint extension lengths are configured according to the plurality of weakly labeled factors, the plurality of sample voiceprint feature sets and the sample detection result set are extended to obtain a plurality of sample voiceprint feature sequence sets, and a cable joint anomaly detector is trained in combination with the labeled sample detection result set;
[0010] The system collects current voiceprint data, analyzes it to obtain real-time weak marker factors, expands it to obtain a real-time voiceprint feature sequence set, inputs it into the cable joint anomaly detector, and outputs the cable joint anomaly rate as the detection result.
[0011] The beneficial effects of this invention are:
[0012] Compared to existing technologies, this application first collects sample acoustic fingerprint datasets and sample detection result sets based on the operational monitoring data of cable joints in the distribution network, and extracts multiple sample acoustic fingerprint feature sets, providing a reliable data foundation for subsequent weak labeling analysis, data expansion, and model training. Secondly, weak labeling factor analysis is performed on the multiple sample acoustic fingerprint feature sets to obtain multiple weak labeling factors. The sample detection result sets are then labeled, and multiple acoustic fingerprint expansion lengths are configured based on these weak labeling factors to expand the multiple sample acoustic fingerprint feature sets and sample detection result sets, obtaining multiple sample acoustic fingerprint feature sequence sets. Combined with the labeled sample detection result sets, a cable joint anomaly detector is trained, compensating for the deficiencies of low-reliability data. The cable joint anomaly detector can accurately capture the evolution trend of cable joint status from normal to abnormal, providing reliable model support for subsequent anomaly detection of cable joints in the distribution network. Finally, the current voiceprint data is collected, the real-time weak marker factor is analyzed, and the real-time voiceprint feature sequence set is expanded. This set is then input into the cable joint anomaly detector, and the cable joint anomaly rate is output as the detection result. This enables an accurate assessment of the current cable joint anomaly risk and can serve as a basis for maintenance personnel's maintenance decisions.
[0013] Through the above technical solution, this application first performs weak labeling factor analysis on the voiceprint features, transforming the originally ambiguous "high probability of abnormality, low probability of normal" weak labeling state into quantifiable weak labeling factors. Based on this, the voiceprint extension length is dynamically configured, supplementing low-confidence samples with sufficient temporal features to enhance information effectiveness, while avoiding increased training burden due to redundant data for high-confidence samples, thus achieving targeted data extension. Simultaneously, the trained cable joint anomaly detector can capture the dynamic trend of cable joint state evolution, accurately outputting detection results including the anomaly rate based on voiceprint features, thereby achieving real-time and accurate detection of abnormal states of distribution network cable joints. This improves the accuracy and stability of cable joint anomaly detection, providing a reliable decision-making basis for distribution network operation and maintenance. Attached Figure Description
[0014] Figure 1 A schematic flowchart illustrating the acoustic anomaly detection method for distribution network cable joints based on weak marker data provided by the present invention;
[0015] Figure 2 This is a flowchart illustrating the process of constructing the cable joint anomaly detector in the distribution network cable joint acoustic anomaly detection method based on weakly marked data provided by the present invention. Detailed Implementation
[0016] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0017] In the description of this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, a feature defined as "first" or "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0018] In the description of this invention, the term "for example" is used to mean "used as an example, illustration, or description." Any embodiment described as "for example" in this invention is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed herein.
[0019] Examples, such as Figure 1 As shown, this embodiment of the invention provides a method for detecting acoustic anomalies in distribution network cable joints based on weak marker data, including:
[0020] S10: Based on the operation monitoring data of the distribution network cable joints, collect sample voiceprint datasets, collect sample detection result sets, and extract multiple sample voiceprint feature sets.
[0021] Traditional methods for detecting abnormal acoustic signatures in distribution network cable joints often employ a single acoustic signature feature plus a fixed threshold for judgment. For example, anomalies are determined solely by whether the energy value exceeds a preset threshold. This method fails to consider the ambiguity of acoustic signature features due to environmental noise and equipment errors, and also ignores the characteristic differences of different fault types. This can easily lead to misjudgments or omissions of faults such as minor loosening or early insulation damage.
[0022] Meanwhile, the raw voiceprint data collected is a continuous audio signal, which not only lacks a unified time dimension, but also cannot be directly read and analyzed by the algorithm, further resulting in low detection accuracy and poor adaptability, making it difficult to meet the needs of distribution network operation and maintenance for accurate fault identification.
[0023] To address the aforementioned issues, this application collects sample acoustic signature datasets and sample detection result sets based on operational monitoring data of cable joints in the power distribution network, and extracts multiple sample acoustic signature feature sets. Each sample acoustic signature feature set includes at least energy and zero-crossing rate.
[0024] Specifically, step S10 in the method includes:
[0025] During the operation of cable joints in the power distribution network, acoustic fingerprint data of a preset length is collected to obtain a sample acoustic fingerprint dataset.
[0026] Detect the abnormal state of the cable joint under each sample of acoustic fingerprint data to obtain a sample detection result set;
[0027] Voiceprint features are extracted from each sample voiceprint data in the sample voiceprint dataset to obtain multiple sample voiceprint feature sets.
[0028] In this embodiment, during the operation of the power distribution network cable joint, a preset length of acoustic fingerprint data is collected from the operation monitoring data of the power distribution network cable joint to obtain a sample acoustic fingerprint dataset. The preset length is a fixed acoustic fingerprint acquisition duration, such as 50ms. The purpose of the preset length is to ensure that the time dimension of all sample acoustic fingerprint data is completely consistent, avoiding incomparability of indicators during subsequent feature extraction due to differences in the duration of single data segments. Those skilled in the art can flexibly adjust this duration according to actual needs. For example, it can be set to 100ms for high-voltage cable joints to capture richer discharge sound characteristics, and to 30ms for low-voltage joints to improve acquisition efficiency.
[0029] Among them, soundprint data refers to the audio signals generated by cable joints in the power distribution network during operation, which can reflect their working status. Soundprint data is directly related to the health of cable joints. For example, the stable current humming sound of a normal cable joint, the intermittent metallic friction sound of a slightly loose joint, and the weak discharge "buzzing" sound of a joint with damaged insulation.
[0030] For example, if the preset length is 50ms, combined with a fixed time interval (such as once per second), acoustic data of cable joints in the power distribution network are continuously collected during operation. After the collection is completed, all acoustic data segments are sorted and integrated according to the collection time to form a sample acoustic data set. The sample acoustic data set is a series of audio segments labeled with the collection time that can truly reflect the operating status of the cable joints, providing reliable raw data for subsequent feature extraction and anomaly analysis.
[0031] Secondly, the actual abnormal state of the cable joint under each sample of acoustic fingerprint data is detected to obtain a "normal" or "abnormal" detection result, which is then labeled with a binary label to obtain a sample detection result set. Common "abnormal" states include: local overheating of the cable joint due to excessive contact resistance, mechanical friction caused by loose conductors or connectors, partial discharge caused by aging / damage of the insulation layer, and poor contact caused by external dust / moisture intrusion.
[0032] For example, the true abnormal state of cable joints under sample acoustic fingerprint data can be obtained through manual inspection and sensor-assisted verification. For instance, technicians can use infrared thermometers to check for overheating, partial discharge detectors to locate insulation faults, and visual and tactile checks for looseness. Simultaneously, temperature sensors, vibration sensors, and partial discharge sensors can be installed near the cable joints to record the true state of the cable joints at the time of sample acoustic fingerprint data acquisition, helping to confirm the abnormal state of the cable joints under each sample acoustic fingerprint data. Then, binary labels are used to mark the determined "normal" or "abnormal" states, for example, labeling the normal state as "0" and the abnormal state as "1," so that each sample acoustic fingerprint data is bound to a unique binary label, forming a sample detection result set. This provides basic data reflecting the true state for subsequent weak label analysis and model training.
[0033] Finally, since the sample voiceprint data is a continuous audio waveform, such as a sound wave diagram, the computer cannot directly analyze the waveform pattern. Therefore, it is necessary to extract voiceprint features from each sample voiceprint data in the sample voiceprint dataset. From each sample voiceprint data, voiceprint features that can reflect the state of the cable joint are extracted to obtain multiple sample voiceprint feature sets. Among them, the voiceprint features are extracted from the voiceprint data and can quantify indicators that reflect the true state of the cable joint, such as energy, zero-crossing rate, and amplitude envelope statistics: the energy of a normal cable joint is stable within a specific range, while when the joint is loose and causes friction or discharge, the energy will fluctuate abnormally or even increase significantly; the number of times the sound wave crosses the horizontal axis, the value is stable during normal operation, and abnormal vibration will cause the waveform to jitter at high frequencies, causing the zero-crossing rate to increase significantly; the amplitude envelope statistics include the amplitude mean, variance, etc., which can characterize the variation law of sound amplitude.
[0034] For example, voiceprint features can be extracted using specialized signal processing algorithms. For instance, a short-time Fourier transform can be used to convert the voiceprint data from the time domain to the frequency domain, and spectral features can be extracted using Mel-frequency cepstral coefficients (MFCC). Then, a sliding window can be used to calculate the energy, zero-crossing rate, and amplitude envelope statistics for each voiceprint segment. For example, after extracting a 50ms sample voiceprint data, a set of specific feature values is obtained: energy 55dB, zero-crossing rate 130 times / 50ms, and mean amplitude envelope of 28 (normalized value). Integrating the feature values of all sample voiceprint data forms multiple sample voiceprint feature sets.
[0035] It should be noted that the short-time Fourier transform, Mel-frequency cepstral coefficients (MFCC), and sliding window techniques involved in the aforementioned voiceprint feature extraction process are all existing and mature technologies widely used in the field of audio signal processing. Their principles and implementation methods are clearly described in publicly available technical literature, industry standards, or open-source tool libraries. Those skilled in the art can obtain and apply these technologies through various conventional means, such as referring to theoretical methods in professional textbooks like "Digital Signal Processing" and calling readily available functions from open-source tools like Python's Librosa library and MATLAB's signal processing toolbox, without requiring additional research and development. Therefore, this application will not elaborate on the specific implementation details of these existing technologies.
[0036] In summary, compared to existing technologies, this application collects sample acoustic signature datasets and sample detection result sets based on the operational monitoring data of distribution network cable joints, and extracts multiple sample acoustic signature feature sets. In this way, it obtains raw signals reflecting the true state of distribution network cable joints and transforms them into calculable quantitative indicators, providing a reliable data foundation for subsequent weak labeling analysis, data expansion, and model training.
[0037] S20: Perform weak labeling factor analysis on multiple sample voiceprint feature sets to obtain multiple weak labeling factors, label the sample detection result set, configure multiple voiceprint expansion lengths based on multiple weak labeling factors, expand multiple sample voiceprint feature sets and sample detection result sets to obtain multiple sample voiceprint feature sequence sets, and train the cable joint anomaly detector by combining the labeled sample detection result sets.
[0038] The acoustic signature of cable joints in power distribution networks is easily affected by environmental noise and equipment acquisition errors, resulting in a large amount of weakly labeled data. This means that the same acoustic signature may correspond to both normal and abnormal states, making it impossible to define it accurately with absolute labels. However, traditional methods often directly use this weakly labeled data, leading to significant errors in the detection results of acoustic signature anomalies in power distribution network cable joints.
[0039] Meanwhile, due to the varying reliability of sample detection results, directly training the model using the initial sample voiceprint feature set and sample detection result set will result in significant drawbacks: samples with low reliability lack continuous features in the time dimension, making it difficult to provide effective state association information, leading to insufficient learning of such samples by the model; samples with high reliability are forcibly included in the unified training even though no additional data is required, generating redundant data and increasing the computational burden, ultimately resulting in inconsistent training data quality. This leads to inaccurate association between the features learned by the model and abnormal states, resulting in underfitting of low reliability samples and overfitting of high reliability samples, causing the output anomaly detection results to have large deviations and failing to meet the requirements of power distribution network operation and maintenance for fault identification accuracy and stability.
[0040] To address the aforementioned issues, this application performs weak labeling factor analysis on multiple sample voiceprint feature sets to obtain multiple weak labeling factors, labels the sample detection result set, configures multiple voiceprint extension lengths based on the multiple weak labeling factors, expands the multiple sample voiceprint feature sets and sample detection result sets, obtains multiple sample voiceprint feature sequence sets, and trains a cable joint anomaly detector by combining the labeled sample detection result sets.
[0041] Specifically, step S20 in the method includes:
[0042] Within the multiple sample voiceprint feature sets and sample detection result sets, select the first sample voiceprint feature set and the first sample detection result, analyze to obtain the first feature weak labeling factor set, calculate the first weak labeling factor, and label the first sample detection result.
[0043] Further analysis revealed several weakly labeled factors;
[0044] Configure multiple voiceprint extension lengths based on multiple weak marker factors.
[0045] In this embodiment, firstly, from multiple sample voiceprint feature sets and sample detection result sets, a sample voiceprint feature set is randomly selected, and the corresponding sample detection result is obtained as the first sample voiceprint feature set and the first sample detection result. Then, the first feature weak labeling factor set is obtained through analysis, the first weak labeling factor is calculated, and the first sample detection result is labeled. In this way, the sample detection result, which originally only contained the qualitative result of "normal / abnormal", is supplemented with quantitative abnormality rate information, so that the subsequent model training can obtain more refined state association basis, rather than relying on a single binary judgment, thereby improving the model's learning accuracy of the matching relationship between sample voiceprint features and actual state.
[0046] Secondly, repeat the above operations, continue to select sample voiceprint feature sets and sample detection results, analyze and obtain multiple weak labeling factors, form a complete probabilistic labeling system, and solve the problem that traditional absolute labels cannot reflect feature ambiguity.
[0047] Finally, multiple voiceprint augmentation lengths are configured based on multiple weak labeling factors. In this way, the reliability of the sample detection results is quantified by the weak labeling factors. Samples with low reliability (small weak labeling factors) need to be configured with longer voiceprint augmentation lengths to augment more data and use multi-time-segment features to compensate for the deficiencies of single-segment data. Conversely, samples with high reliability (large weak labeling factors) can have their voiceprint augmentation lengths appropriately shortened to achieve a balance between accuracy and efficiency.
[0048] Furthermore, the phrase "analyzing to obtain the first feature weak labeling factor set, calculating the first weak labeling factor, and labeling the first sample detection result" includes:
[0049] Extract the same sample voiceprint features of each first sample voiceprint feature set in multiple sample voiceprint feature sets to obtain multiple first sample voiceprint feature sets of the same type.
[0050] Obtain multiple first-class sample detection result sets corresponding to multiple first-class sample voiceprint feature sets, calculate the proportion of multiple first-class sample detection result sets that are the same as the first sample detection result, obtain multiple first feature weak labeling factors, and calculate the mean to obtain the first weak labeling factor.
[0051] The first weak labeling factor is used to label the first sample detection result to obtain the labeled first sample detection result, wherein the labeled first sample detection result includes the anomaly rate.
[0052] In this embodiment, firstly, for each first sample voiceprint feature in the first sample voiceprint feature set, the same sample voiceprint features within multiple sample voiceprint feature sets are extracted. Then, they are clustered according to the voiceprint feature type to obtain multiple first similar sample voiceprint feature sets. For example, for each first sample voiceprint feature in the first sample voiceprint feature set, such as energy 55dB, zero-crossing rate 130 times / 50ms, and amplitude envelope mean 28, all sample voiceprint features with energy of 55dB are extracted from multiple sample voiceprint feature sets to form a first similar sample voiceprint feature set, all sample voiceprint features with a zero-crossing rate of 130 times / 50ms are extracted to form a first similar sample voiceprint feature set, and all sample voiceprint features with an amplitude envelope mean of 28 are extracted to form a first similar sample voiceprint feature set. In this way, a total of three first similar sample voiceprint feature sets are obtained.
[0053] Secondly, multiple first-type sample detection result sets corresponding to multiple first-type sample voiceprint feature sets are obtained. The proportion of multiple first-type sample detection result sets that are the same as the first sample detection result is calculated to obtain multiple first feature weak labeling factors. The average value is then calculated to obtain the first weak labeling factor. For example, using the data from the previous example, if the first sample detection result is "0" (representing a normal state), the proportion of detection results of "0" in the three first-type sample voiceprint feature sets is calculated separately. For example, the proportion of detection results of "0" in the first-type sample voiceprint feature set with an energy of 55dB is 60%, the proportion of detection results of "0" in the first-type sample voiceprint feature set with a zero-crossing rate of 130 times / 50ms is 70%, and the proportion of detection results of "0" in the first-type sample voiceprint feature set with an amplitude envelope mean of 28 is 80%. The average value of the three is then calculated to obtain the first weak labeling factor = (60% + 70% + 80%) / 3 = 70%. The first weak labeling factor can comprehensively reflect the reliability of the first sample detection result.
[0054] Finally, the first weak labeling factor is used to label the first sample detection result, obtaining the labeled first sample detection result, which includes the anomaly rate. For example, if the first weak labeling factor is 70%, when the first sample detection result is "1" (representing an abnormal state), the first weak labeling factor of 70% is directly labeled as the anomaly rate of the labeled first sample detection result; conversely, when the first sample detection result is "0" (representing a normal state), 1-70%=30% is labeled as the anomaly rate of the labeled first sample detection result. This labeling method preserves the ambiguity of the detection results, facilitating subsequent model learning of more refined association state features.
[0055] Furthermore, the phrase "configuring multiple voiceprint augmentation lengths based on multiple weak marker factors" includes:
[0056] Obtain the preset voiceprint extension length;
[0057] The ratio of the mean of multiple weak marker factors to each weak marker factor is calculated, and the preset voiceprint amplification length is adjusted and configured as multiple voiceprint amplification lengths.
[0058] In this embodiment, a preset acoustic signature expansion length is first obtained. This preset acoustic signature expansion length is the initial base expansion time set for all samples, serving as a benchmark value for subsequent dynamic adjustments. The preset acoustic signature expansion length can be dynamically determined according to the actual application scenario: for example, for high-voltage cable joints with more complex fault acoustic signature characteristics and longer durations, the preset acoustic signature expansion length can be set to 100ms to cover more feature details; for low-voltage cable joints with simpler fault acoustic signatures, the preset acoustic signature expansion length can be set to 50ms to improve processing efficiency. Those skilled in the art can set a preset acoustic signature expansion length that fits the actual detection requirements according to the actual application scenario.
[0059] Secondly, the ratio of the mean of multiple weak labeling factors to each weak labeling factor is calculated, and the preset voiceprint amplification length is adjusted and configured to have multiple voiceprint amplification lengths. The magnitude of the weak labeling factor directly reflects the reliability of the sample detection results: a larger weak labeling factor indicates higher consistency with the current sample detection results within the same sample detection result set, and thus higher reliability; conversely, a smaller weak labeling factor indicates lower reliability.
[0060] The ratio of the mean of weak marker factors to each weak marker factor determines the direction of adjustment for the voiceprint extension length: the larger the weak marker factor, the smaller the ratio of the mean of weak marker factors to each weak marker factor, which means that the required voiceprint extension length is shorter; conversely, the smaller the weak marker factor, the larger the ratio of the mean of weak marker factors to each weak marker factor, which means that the required voiceprint extension length is longer.
[0061] For example, first calculate the mean of all weakly labeled factors, such as 80%. Then, calculate the ratio of the mean of each weakly labeled factor to the mean of each weakly labeled factor. For example, if a weakly labeled factor is 70%, the ratio = 80% / 70% = 1.14. Then, calculate the product of the ratio and the preset voiceprint amplification length, which is the voiceprint amplification length. For example, if the preset voiceprint amplification length is 100ms, the voiceprint amplification length = 100ms × 1.14 = 114ms. This indicates that the reliability of the current sample detection result is lower than the mean, and a longer voiceprint amplification length is needed to improve reliability. In this way, samples with low reliability (small weakly labeled factors) can reduce misjudgments through multi-time-segment feature cross-validation by using a longer voiceprint amplification length, while samples with high reliability (large weakly labeled factors) can avoid invalid data dragging down processing efficiency, achieving a balance between accuracy and efficiency.
[0062] Furthermore, such as Figure 2 As shown, step S20 in the method further includes:
[0063] According to multiple voiceprint extension lengths, the length of each sample voiceprint data in the sample voiceprint dataset is extended and supplementary voiceprint data is collected to obtain a sample voiceprint data sequence set.
[0064] Based on the sample voiceprint data sequence set, multiple sample voiceprint feature sequence sets are extracted;
[0065] A cable joint anomaly detector was constructed based on machine learning.
[0066] The cable joint anomaly detector is trained and tested using the multiple sample voiceprint feature sequence sets and the labeled sample detection result set to obtain the cable joint anomaly detector.
[0067] In this embodiment, each sample voiceprint data in the sample voiceprint dataset is first length-extended and supplemented according to multiple voiceprint extension lengths to obtain a sample voiceprint data sequence set. For example, if the sample voiceprint data collection rule is a preset length of 50ms and one collection per second, when the corresponding voiceprint extension length is 114ms, the voiceprint data of the 114ms preceding the 50ms voiceprint data is collected, forming a sample voiceprint data sequence with a total length of 164ms. The same operation is performed on each sample voiceprint data according to its corresponding voiceprint extension length to obtain the sample voiceprint data sequence set. Compared to the initial sample voiceprint dataset, the sample voiceprint data sequence set contains more continuous voiceprint features, such as the gradual change in voiceprint energy from slight loosening to significant friction in a distribution network cable joint. These continuous voiceprint features can more realistically and accurately reflect the evolution trajectory of the cable joint state, providing a reliable data foundation for subsequent extraction of time-series feature sequences and training of cable joint anomaly detectors.
[0068] Secondly, based on the sample voiceprint data sequence set, voiceprint features are extracted using the same method as in step S10 to obtain multiple sample voiceprint feature sequence sets. For example, each sample voiceprint data sequence is segmented using a sliding window: for example, with a preset length of 50ms as the window size and a step size of 10ms, the 164ms sample voiceprint data sequence is segmented into 12 overlapping 50ms windows. Then, voiceprint features (such as energy, zero-crossing rate, and amplitude envelope statistics) are extracted for each window. Each window obtains a set of feature values, and the 12 sets of feature values are arranged according to voiceprint feature type and time order to form multiple sample voiceprint feature sequences, which serve as the sample voiceprint feature sequence set for the current sample voiceprint data sequence. For example, the sample voiceprint feature sequence set includes: energy sequence [55, 58, 62, ..., 70], zero-crossing rate sequence [130, 135, 140, ..., 150], and amplitude envelope statistics sequence [28, 27, 29, ..., 26]. Multiple sample voiceprint feature sequence sets can capture the changing trend of voiceprint features over time, thus providing a basis for identifying early faults.
[0069] Secondly, a cable joint anomaly detector is constructed based on machine learning. For example, LSTM (Long Short-Term Memory) networks are specifically designed for processing time-series data and can effectively capture the correlation patterns that change over time in voiceprint feature sequences. Therefore, a cable joint anomaly detector can be constructed using an LSTM architecture, mainly composed of an input layer, an LSTM layer, a fully connected layer, and an output layer. Specifically:
[0070] The input layer receives structured input from a set of sample voiceprint feature sequences, with the input dimension matching the length and number of features in the feature sequence. For example, if each sample voiceprint feature sequence contains 12 time steps, and each time step contains 3 voiceprint features (energy, zero-crossing rate, and mean amplitude envelope), then the input layer dimension is (12, 3), ensuring that the model can fully receive the temporal distribution and feature dimension information of the sample voiceprint feature sequence set.
[0071] The LSTM layer contains 1-2 LSTM units to extract the temporal dependencies of the sample voiceprint feature sequences. For example, the first layer has 64 LSTM units, which selectively retain key historical features and discard noise information through a gating mechanism; the second layer has 32 LSTM units to further compress the feature dimension and strengthen the capture of long-term dependencies, such as the periodic fluctuation of a certain voiceprint feature within 10 time steps. Additionally, a Dropout layer (e.g., dropout=0.2) can be added after each layer to prevent overfitting.
[0072] The fully connected layer transforms the temporal features output by the LSTM layer into a one-dimensional vector, integrating key information through nonlinear transformation. For example, a fully connected layer with 16 neurons can be used with the ReLU activation function to enhance the model's ability to express nonlinear features, mapping high-dimensional temporal features into more compact abstract features.
[0073] The output layer contains one neuron and uses the Sigmoid activation function to compress the output value to the 0-1 range, which directly corresponds to the abnormality rate of the cable joint.
[0074] Finally, using multiple sample acoustic signature sequence sets and labeled sample detection result sets, the cable joint anomaly detector is trained and tested under supervision to obtain the cable joint anomaly detector. The cable joint anomaly detector trained with multiple sample acoustic signature sequence sets and labeled sample detection result sets can receive and parse the input acoustic signature sequence set and output reliable detection results including the anomaly rate.
[0075] Specifically, the step of "using the multiple sample voiceprint feature sequence sets and the labeled sample detection result set to perform supervised training and testing on the cable joint anomaly detector" includes:
[0076] The cable joint anomaly detector is trained under supervision using the multiple sample voiceprint feature sequence sets and the labeled sample detection result set. The deviation between the output of the cable joint anomaly detector and the labeled sample detection result is calculated using a loss function as the loss.
[0077] Adjust the cable joint anomaly detector according to the loss, and perform iterative training for parameter tuning;
[0078] Conduct a test; if the accuracy rate is satisfactory, the training is complete.
[0079] For example, the supervised training and testing process can be carried out through the following technical path: 1. Data preparation: Divide multiple sample voiceprint feature sequence sets and labeled sample detection result sets into training set, validation set, and test set in a ratio of 7:1.5:1.5. The training set is used for model parameter learning, the validation set is used for hyperparameter adjustment and overfitting monitoring, and the test set is used for final evaluation of the model's generalization ability. 2. Model training: Using the sample voiceprint feature sequences in the training set as input features and the corresponding labeled sample detection result sets as supervision labels, iteratively train the cable joint anomaly detector. During training, the mean squared error of MSE is used as the loss function, and the Adam optimizer (learning rate 0.001) is used to dynamically adjust the parameter update step size. After each round of training, the model performance is evaluated using the validation set. If the validation set loss does not decrease for 5 consecutive rounds, training is stopped to avoid model overfitting. The model's generalization ability is then evaluated using the test set. When the mean squared error of MSE in the test set meets the preset accuracy requirements and the prediction accuracy is greater than or equal to 95%, the model is considered to have converged, and the trained cable joint anomaly detector is obtained.
[0080] In summary, compared to existing technologies, this application performs weak labeling factor analysis on multiple sample voiceprint feature sets to obtain multiple weak labeling factors. The sample detection result sets are then labeled, and multiple voiceprint expansion lengths are configured based on these weak labeling factors. This expands the multiple sample voiceprint feature sets and sample detection result sets, resulting in multiple sample voiceprint feature sequence sets. Combined with the labeled sample detection result sets, a cable joint anomaly detector is trained. Thus, by quantifying the reliability of sample detection results through weak labeling factors, samples of different confidence levels are dynamically expanded according to the voiceprint expansion length. Sufficient temporal features are added to low-confidence samples to enhance information effectiveness, while avoiding reduced training efficiency due to redundant data for high-confidence samples. The cable joint anomaly detector trained in this way can accurately capture the evolution trend of cable joint status from normal to abnormal, providing reliable model support for subsequent anomaly detection of distribution network cable joints.
[0081] S30: Collect the current voiceprint data, analyze it to obtain the real-time weak marker factor, expand it to obtain the real-time voiceprint feature sequence set, input it into the cable joint anomaly detector, and output the cable joint anomaly rate as the detection result.
[0082] The aforementioned steps, based on techniques such as weak marker factor analysis and acoustic signature feature augmentation, construct a cable joint anomaly detector, which can be used to detect faults based on the real-time acoustic signature features of cable joints.
[0083] To address the aforementioned issues, this application collects current voiceprint data, analyzes it to obtain a real-time weak marker factor, expands it to obtain a real-time voiceprint feature sequence set, inputs it into the cable joint anomaly detector, and outputs the cable joint anomaly rate as the detection result.
[0084] Specifically, step S30 in the method includes:
[0085] Collect current voiceprint data and extract real-time voiceprint feature sets;
[0086] Based on the real-time voiceprint feature set and multiple sample voiceprint feature sets, the real-time weak labeling factor is obtained through analysis.
[0087] Based on the real-time weak labeling factor, configure the voiceprint expansion length, expand the real-time voiceprint feature set, and obtain a real-time voiceprint feature sequence set;
[0088] The real-time voiceprint feature sequence set is input into the cable joint anomaly detector, and the cable joint anomaly rate is output as the detection result.
[0089] In this embodiment, the current voiceprint data is first collected using the same method as step S10, and a real-time voiceprint feature set is extracted. During the voiceprint data collection and voiceprint feature extraction process, it is necessary to ensure that the current voiceprint data and the sample voiceprint data are compatible in terms of format and duration. At the same time, the feature dimensions and physical meanings of the real-time voiceprint feature set and the sample voiceprint feature set are completely consistent to avoid incomparability of features due to differences in collection parameters.
[0090] Secondly, based on the same logic as step S20, a real-time weak labeling factor is obtained by analyzing the real-time voiceprint feature set and multiple sample voiceprint feature sets. For example, the same sample voiceprint features for each voiceprint feature in the real-time voiceprint feature set are extracted from multiple sample voiceprint feature sets to obtain multiple similar sample voiceprint feature sets. Multiple similar sample detection result sets are obtained corresponding to these sets. The proportion of detection results of "1" (representing an abnormal state) in the multiple similar sample detection result sets is calculated to obtain multiple feature weak labeling factors. The average of these factors is then calculated to obtain the real-time weak labeling factor. The proportion of detection results of "1" (representing an abnormal state) in the multiple similar sample detection result sets is calculated because the core objective of this application is cable joint anomaly detection. To prioritize avoiding the risk of missed anomaly detection, the real-time weak labeling factor is prioritized to focus on the correlation between real-time voiceprint features and abnormal states, providing a basis for subsequent dynamic configuration of voiceprint extension length. The real-time weak labeling factor reflects the reliability of the real-time voiceprint features. The larger the real-time weak labeling factor, the higher the matching degree between the real-time voiceprint features and the sample voiceprint features, i.e., the higher the reliability.
[0091] Next, based on the same logic as step S20, the voiceprint expansion length is configured according to the real-time weak labeling factor, and the real-time voiceprint feature set is expanded to obtain a real-time voiceprint feature sequence set. For example, using the average values of multiple weak labeling factors from step S20, the ratio of the average values of multiple weak labeling factors to the real-time weak labeling factor is calculated. Based on this, the preset voiceprint expansion length is adjusted and configured as the voiceprint expansion length. Then, according to the voiceprint expansion length, the current voiceprint data is length-expanded and supplemented with additional voiceprint data. The voiceprint features of the supplemented voiceprint data are extracted to obtain the real-time voiceprint feature sequence set.
[0092] Thus, by configuring the voiceprint extension length according to the real-time weak labeling factor (reflecting the credibility of real-time voiceprint features), sufficient temporal information can be supplemented for real-time voiceprint features with low credibility, while avoiding the increase in processing delay due to redundant data for real-time voiceprint features with high credibility. Ultimately, the information density and structure of the real-time voiceprint feature sequence set are precisely adapted to the input requirements of the cable joint anomaly detector, providing reliable data support for the output anomaly rate detection results.
[0093] Finally, the real-time voiceprint feature sequence set is input into the pre-trained cable joint anomaly detector, and the output is the cable joint anomaly rate, which serves as the detection result. The cable joint anomaly rate directly reflects the real-time anomaly risk level of the cable joint; the higher the rate, the greater the likelihood of an abnormal state.
[0094] Specifically, the step of "inputting the real-time voiceprint feature sequence set into the cable joint anomaly detector and outputting the cable joint anomaly rate as the detection result" includes:
[0095] The real-time acoustic signature sequence set is input into the cable joint anomaly detector, and the cable joint anomaly rate is obtained as the output.
[0096] Obtain the joint abnormality rate threshold, determine whether the cable joint abnormality rate is greater than or equal to the joint abnormality rate threshold, and obtain the detection result.
[0097] In this embodiment, the real-time voiceprint feature sequence set is first input into a pre-trained cable joint anomaly detector, and the cable joint anomaly rate is output. For example, after inputting the real-time voiceprint feature sequence set into the pre-trained cable joint anomaly detector, the output cable joint anomaly rate is 75%. This cable joint anomaly rate can intuitively and quantitatively reflect the current risk level of the cable joint.
[0098] Secondly, the abnormality rate threshold of the cable joint is obtained, and it is determined whether the abnormality rate of the cable joint is greater than or equal to the threshold to obtain the detection result. The abnormality rate threshold is a critical value used to determine whether a cable joint needs to trigger an early warning, such as 60%. Those skilled in the art can dynamically adjust it according to the actual scenario. For example, for old cable joints with an operating life of over 10 years, the abnormality rate threshold can be lowered to 50% to strengthen early warning, while for newly built cable joints, it can be raised to 70% to reduce false alarms.
[0099] If the cable joint anomaly rate exceeds the joint anomaly rate threshold, it indicates that the current cable joint anomaly risk has reached the warning level requiring intervention, and there may be early faults or hidden defects. The warning mechanism should be triggered immediately to prompt maintenance personnel to conduct targeted investigations based on the site environment. Conversely, if the cable joint anomaly rate does not exceed the joint anomaly rate threshold, it is determined that the cable joint is currently in normal operation and can be monitored according to the regular cycle.
[0100] In summary, compared to existing technologies, this application collects current voiceprint data, analyzes it to obtain real-time weak labeling factors, expands it to obtain a real-time voiceprint feature sequence set, inputs it into the cable joint anomaly detector, and outputs the cable joint anomaly rate as the detection result. Thus, by dynamically expanding the real-time voiceprint feature sequence set through real-time weak labeling factors, a quantified cable joint anomaly rate is ultimately output, enabling accurate assessment of the current cable joint anomaly risk and serving as a basis for maintenance personnel's operational decisions.
[0101] In summary, the embodiments of this application have at least the following technical effects:
[0102] Compared to existing technologies, this application first collects sample acoustic signature datasets and sample detection result sets based on the operational monitoring data of cable joints in the distribution network, and extracts multiple sample acoustic signature feature sets. In this way, raw signals reflecting the true state of cable joints in the distribution network are obtained and transformed into calculable quantitative indicators, providing a reliable data foundation for subsequent weak labeling analysis, data expansion, and model training.
[0103] Secondly, this application performs weak labeling factor analysis on multiple sample voiceprint feature sets to obtain multiple weak labeling factors. The sample detection result sets are then labeled, and multiple voiceprint expansion lengths are configured based on these weak labeling factors. This expands the multiple sample voiceprint feature sets and sample detection result sets, resulting in multiple sample voiceprint feature sequence sets. Combined with the labeled sample detection result sets, a cable joint anomaly detector is trained. Thus, by quantifying the reliability of sample detection results through weak labeling factors, samples of different confidence levels are dynamically expanded according to the voiceprint expansion length. This provides sufficient temporal features to enhance the effectiveness of information for low-confidence samples, while avoiding reduced training efficiency due to redundant data for high-confidence samples. The cable joint anomaly detector trained in this way can accurately capture the evolution trend of cable joint status from normal to abnormal, providing reliable model support for subsequent anomaly detection of distribution network cable joints.
[0104] Finally, this application collects current voiceprint data, analyzes it to obtain a real-time weak labeling factor, expands it to obtain a real-time voiceprint feature sequence set, inputs it into the cable joint anomaly detector, and outputs the cable joint anomaly rate as the detection result. Thus, by dynamically expanding the real-time voiceprint feature sequence set through the real-time weak labeling factor, a quantified cable joint anomaly rate is finally output, enabling accurate judgment of the current cable joint anomaly risk, which can serve as a basis for maintenance personnel's operational decisions.
[0105] Through the above technical solution, this application first performs weak labeling factor analysis on the voiceprint features, transforming the originally ambiguous "high probability of abnormality, low probability of normal" weak labeling state into quantifiable weak labeling factors. Based on this, the voiceprint extension length is dynamically configured, supplementing low-confidence samples with sufficient temporal features to enhance information effectiveness, while avoiding increased training burden due to redundant data for high-confidence samples, thus achieving targeted data extension. Simultaneously, the trained cable joint anomaly detector can capture the dynamic trend of cable joint state evolution, accurately outputting detection results including the anomaly rate based on voiceprint features, thereby achieving real-time and accurate detection of abnormal states of distribution network cable joints. This improves the accuracy and stability of cable joint anomaly detection, providing a reliable decision-making basis for distribution network operation and maintenance.
[0106] It should be noted that the descriptions of each embodiment in the above embodiments have different focuses. For parts that are not described in detail in a certain embodiment, please refer to the relevant descriptions in other embodiments.
[0107] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0108] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded computer, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.
[0109] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.
[0110] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.
[0111] Although preferred embodiments of the invention have been described, those skilled in the art, once they have learned the basic inventive concept, can make other changes and modifications to these embodiments.
[0112] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of this invention and its equivalents, this invention also intends to include these modifications and variations.
Claims
1. A method for detecting acoustic anomalies in distribution network cable joints based on weakly marked data, characterized in that, The method includes: Based on the operational monitoring data of cable joints in the power distribution network, a sample acoustic signature dataset was collected, along with a sample detection result set. Multiple sample acoustic signature feature sets were then extracted, including: During the operation of cable joints in the power distribution network, acoustic fingerprint data of a preset length is collected to obtain a sample acoustic fingerprint dataset. The abnormal state of the cable joint is detected under the acoustic fingerprint data of each sample, and a sample detection result set is obtained. Each sample detection result includes abnormal or normal. Voiceprint features are extracted from each sample voiceprint data in the sample voiceprint dataset to obtain multiple sample voiceprint feature sets. Weak labeling factor analysis is performed on multiple sample voiceprint feature sets to obtain multiple weak labeling factors. The sample detection result set is labeled, and multiple voiceprint extension lengths are configured based on the multiple weak labeling factors. Multiple sample voiceprint feature sets and sample detection result sets are then extended to obtain multiple sample voiceprint feature sequence sets. Combined with the labeled sample detection result sets, a cable joint anomaly detector is trained, including: Within the multiple sample voiceprint feature sets and sample detection result sets, select the first sample voiceprint feature set and the first sample detection result, analyze to obtain the first feature weak labeling factor set, calculate the first weak labeling factor, and label the first sample detection result. Further analysis revealed several weakly labeled factors; Configure multiple voiceprint extension lengths based on multiple weak marker factors; The system collects current voiceprint data, analyzes it to obtain real-time weak marker factors, expands it to obtain a real-time voiceprint feature sequence set, inputs it into the cable joint anomaly detector, and outputs the cable joint anomaly rate as the detection result, including: Collect current voiceprint data and extract real-time voiceprint feature sets; Based on the real-time voiceprint feature set and multiple sample voiceprint feature sets, the real-time weak labeling factor is obtained through analysis. Based on the real-time weak labeling factor, configure the voiceprint expansion length, expand the real-time voiceprint feature set, and obtain a real-time voiceprint feature sequence set; The real-time voiceprint feature sequence set is input into the cable joint anomaly detector, and the cable joint anomaly rate is output as the detection result.
2. The method for detecting acoustic anomalies in distribution network cable joints based on weakly marked data according to claim 1, characterized in that, Each sample voiceprint feature set includes at least energy and zero-crossing rate.
3. The method for detecting acoustic anomalies in distribution network cable joints based on weakly marked data according to claim 1, characterized in that, The analysis yields the first set of weak labeling factors, the calculation of the first weak labeling factors, and the annotation of the first sample detection results, including: Extract the same sample voiceprint features of each first sample voiceprint feature set in multiple sample voiceprint feature sets to obtain multiple first sample voiceprint feature sets of the same type. Obtain multiple first-class sample detection result sets corresponding to multiple first-class sample voiceprint feature sets, calculate the proportion of multiple first-class sample detection result sets that are the same as the first sample detection result, obtain multiple first feature weak labeling factors, and calculate the mean to obtain the first weak labeling factor. The first weak labeling factor is used to label the first sample detection result to obtain the labeled first sample detection result, wherein the labeled first sample detection result includes the anomaly rate.
4. The method for detecting acoustic anomalies in distribution network cable joints based on weakly marked data according to claim 1, characterized in that, Multiple voiceprint augmentation lengths are configured based on multiple weak marker factors, including: Obtain the preset voiceprint extension length; The ratio of the mean of multiple weak marker factors to each weak marker factor is calculated, and the preset voiceprint amplification length is adjusted and configured as multiple voiceprint amplification lengths.
5. The method for detecting acoustic anomalies in distribution network cable joints based on weakly marked data according to claim 1, characterized in that, Multiple sample voiceprint feature sets and sample detection result sets are expanded to obtain multiple sample voiceprint feature sequence sets. Combined with the labeled sample detection result sets, a cable joint anomaly detector is trained, including: According to multiple voiceprint extension lengths, the length of each sample voiceprint data in the sample voiceprint dataset is extended and supplementary voiceprint data is collected to obtain a sample voiceprint data sequence set. Based on the sample voiceprint data sequence set, multiple sample voiceprint feature sequence sets are extracted; A cable joint anomaly detector was constructed based on machine learning. The cable joint anomaly detector is trained and tested using the multiple sample voiceprint feature sequence sets and the labeled sample detection result set to obtain the cable joint anomaly detector.
6. The method for detecting acoustic anomalies in distribution network cable joints based on weakly marked data according to claim 5, characterized in that, The cable joint anomaly detector is subjected to supervised training and testing using the multiple sample voiceprint feature sequence sets and the labeled sample detection result set, including: The cable joint anomaly detector is trained under supervision using the multiple sample voiceprint feature sequence sets and the labeled sample detection result set. The deviation between the output of the cable joint anomaly detector and the labeled sample detection result is calculated using a loss function as the loss. Adjust the cable joint anomaly detector according to the loss, and perform iterative training for parameter tuning; Conduct a test; if the accuracy rate is satisfactory, the training is complete.
7. The method for detecting acoustic anomalies in distribution network cable joints based on weakly marked data according to claim 1, characterized in that, The real-time acoustic signature sequence set is input into the cable joint anomaly detector, and the cable joint anomaly rate is output as the detection result, including: The real-time acoustic signature sequence set is input into the cable joint anomaly detector, and the cable joint anomaly rate is obtained as the output. Obtain the joint abnormality rate threshold, determine whether the cable joint abnormality rate is greater than or equal to the joint abnormality rate threshold, and obtain the detection result.
Citation Information
Patent Citations
Power transformation equipment anomaly detection method and device based on voiceprint
CN119360890A
Transformer voiceprint anomaly detection method based on generative adversarial network
CN119559970A