Model training method, anomaly detection method, device, equipment, medium and product
By extracting Mel spectrum and Mel cepstral coefficient features to train a neural network model, the problem of insufficient accuracy in the abnormal sound recognition model of transmission lines was solved, and efficient detection of abnormal hazards in different scenarios was achieved.
Patent Information
- Application Number
- CN202410517437.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2024-04-26
- Publication Date
- 2025-10-31
AI Technical Summary
In existing technologies, abnormal sound recognition models for power transmission lines determine the presence of potential fault sounds by identifying the amplitude and frequency of the sound. However, this method has a relatively limited recognition dimension and poor accuracy.
By extracting Mel spectrum features and Mel cepstral coefficient features from audio data of power transmission line scenarios, a corresponding neural network model is trained. A third neural network model is then trained by combining the model accuracy and the two features to predict whether there are abnormal or potential hazards on power transmission lines.
It improves the accuracy of identifying abnormal and potentially dangerous sounds, is applicable in different application scenarios, and enhances the model's recognition performance in noisy environments.
Smart Images

Figure CN120877772A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of power transmission line detection technology, and in particular to a model training method, anomaly detection method, device, equipment, medium and product. Background Technology
[0002] Transmission lines are power lines that transmit electrical energy from power plants (including hydropower stations, thermal power plants, nuclear power plants, wind farms, and photovoltaic power plants) to power load centers and connect different power systems. To ensure the safety of transmission lines, abnormal and potentially hazardous sounds along the transmission corridor can be monitored, identified, and warned of. In the context of transmission lines, in addition to frequently occurring hazardous sounds such as birds and construction machinery, there are also some less frequent hazards, such as short-circuit discharge, lightning strikes, and corona discharge.
[0003] In existing technologies, a sound recognition device is fixedly installed on the power transmission line to monitor the ambient sound around the power transmission line. When an abnormal sound is detected, it is recorded, and the corresponding audio file is generated and uploaded to the monitoring device. The monitoring device calls the built-in sound recognition model to identify the type of sound. If the sound is identified as a potential hazard or fault sound, an early warning message is generated.
[0004] However, the aforementioned sound recognition model determines the presence of potential malfunctions by identifying the amplitude and frequency of the sound. This approach has a limited dimension for identifying potential malfunctions and is therefore less accurate. Summary of the Invention
[0005] This application provides a model training method, anomaly detection method, device, equipment, medium, and product to address the problem that existing models have a relatively singular dimension in identifying hazardous sounds and poor accuracy.
[0006] Firstly, this application provides a model training method, the method comprising:
[0007] Acquire audio data of the scene where the transmission line is located, and extract Mel spectrum features and Mel cepstral coefficient features from the audio data;
[0008] A first neural network model is trained based on the Mel spectrum features, and the corresponding first model accuracy is determined. A second neural network model is trained based on the Mel cepstral coefficient features, and the corresponding second model accuracy is determined.
[0009] A third neural network model is trained based on the accuracy of the first model, the accuracy of the second model, the Mel spectrum features, and the Mel cepstral coefficient features. The third neural network model is used to predict whether there are abnormal or potential hazards on the transmission line.
[0010] Optionally, training a third neural network model based on the first model accuracy, the second model accuracy, the Mel spectral features, and the Mel cepstral coefficient features includes:
[0011] The first weight parameter and the second weight parameter are determined based on the accuracy of the first model and the accuracy of the second model, respectively.
[0012] Using a predefined algorithm, the first weight parameter, the second weight parameter, the Mel spectrum feature, and the Mel cepstral coefficient feature are calculated to obtain the hybrid voiceprint feature;
[0013] A third neural network model is trained based on the hybrid voiceprint features.
[0014] Optionally, the formula corresponding to the predefined algorithm is:
[0015] F = α*f1 + β*f2
[0016] Where F represents the mixed voiceprint feature; f1 represents the Mel spectrum feature; and f2 represents the Mel cepstral coefficient feature. Indicates the first weight parameter; S1 represents the second weight parameter; S2 represents the first model accuracy; S3 represents the second model accuracy.
[0017] Optionally, the third neural network model is a time-delay neural network model based on deep learning. Training the third neural network model based on the hybrid voiceprint features includes:
[0018] Identify at least one type of mixed voiceprint features;
[0019] The at least one type of mixed voiceprint features and the type of audio data are input into the third neural network model for training, resulting in a trained third neural network model.
[0020] Optionally, acquire audio data of the scene where the transmission line is located, and extract Mel-spectral features and Mel-cepstral coefficient features from the audio data, including:
[0021] Collect at least one type of audio data within a preset range of the power transmission line, and construct a database based on the at least one type of audio data;
[0022] For each type, the audio data is preprocessed to obtain multiple audio frames;
[0023] Mel-spectral features of the multiple audio frames are extracted respectively;
[0024] The Mel spectrum features are processed to obtain the Mel cepstral coefficient features.
[0025] Optionally, a database is constructed based on the at least one type of audio data, including:
[0026] Obtain annotation information for the at least one type of audio data, the annotation information indicating the type and clarity score of the audio data; determine the first type of audio data based on the annotation information.
[0027] Based on the first audio data, decibel value adjustment processing is performed to obtain the second audio data, and based on the first audio data, data augmentation processing is performed to obtain the third audio data;
[0028] The audio data in the database is constructed based on the first audio data, the second audio data, and / or the third audio data.
[0029] Optionally, determining the first audio data of the at least one type based on the annotation information includes:
[0030] For each type of audio data, a clarity score is assigned, and a weighting coefficient for the audio data is determined based on the clarity score.
[0031] The audio score is obtained by calculating the clarity score and the weighting coefficient using a weighted algorithm.
[0032] Determine whether the audio score is greater than the first threshold;
[0033] If so, then the audio data is determined to be the first audio data.
[0034] Optionally, the second audio data is obtained by adjusting the decibel value based on the first audio data, including:
[0035] Acquire sound data simulating at least one of the aforementioned types of audio data;
[0036] For each type of sound data, determine whether the sound data exceeds a second threshold;
[0037] If so, the average decibel value of the first audio data of the aforementioned type is calculated, and the sound data is adjusted based on the average decibel value to obtain the second audio data.
[0038] Optionally, based on the first audio data, data augmentation processing is performed to obtain third audio data, including:
[0039] The data augmentation strategy is determined based on the application scenario requirements; the data augmentation strategy includes at least one of the following: random velocity perturbation strategy, random displacement perturbation strategy, random aliasing strategy, and random cropping and splicing strategy;
[0040] The priority of the data augmentation strategy is determined, and based on the priority, the data augmentation strategy is used to sequentially augment the first audio data to obtain the third audio data.
[0041] Optionally, both the first neural network model and the second neural network model are time-delay neural network models based on deep learning. The first neural network model is trained based on the Mel-spectral characteristics, and the corresponding first model accuracy is determined. The second neural network model is trained based on the Mel-cephalic coefficient characteristics, and the corresponding second model accuracy is determined, including:
[0042] The audio data in the database is divided into a training set and a test set according to a specific ratio;
[0043] The Mel-spectral features extracted from the training set and the type of audio data are input into the first neural network model for training to obtain the trained first neural network model; the Mel-cephalic coefficient features extracted from the training set and the type of audio data are input into the second neural network model for training to obtain the trained second neural network model.
[0044] The trained first neural network model and the trained second neural network model are tested based on the test set to determine the accuracy of the first model and the accuracy of the second model.
[0045] Optionally, the abnormal and potential hazards include: short-circuit discharge sound of transmission lines, lightning strike sound, corona discharge sound, and vibration sound of towers or conductors.
[0046] Secondly, this application provides an anomaly detection method, the method comprising:
[0047] Acquire audio data of the scene where the target transmission line is located;
[0048] The audio data is processed based on a third neural network model to obtain a prediction result; the prediction result is used to indicate whether there are any abnormal or potential hazards on the target transmission line.
[0049] The third neural network model is obtained by training a first model accuracy, a second model accuracy, Mel spectrum features, and Mel cepstral coefficient features. The first model accuracy is determined by training the first neural network model based on the Mel spectrum features. The second model accuracy is determined by training the second neural network model based on the Mel cepstral coefficient features. The Mel spectrum features and the Mel cepstral coefficient features are extracted from pre-acquired audio data.
[0050] Thirdly, this application provides a model training apparatus, the apparatus comprising:
[0051] The first acquisition module is used to acquire audio data of the scene where the transmission line is located, and extract Mel spectrum features and Mel cepstral coefficient features from the audio data;
[0052] The determination module is used to train a first neural network model based on the Mel spectrum features and determine the corresponding first model accuracy, and to train a second neural network model based on the Mel cepstral coefficient features and determine the corresponding second model accuracy.
[0053] The training module is used to train a third neural network model based on the accuracy of the first model, the accuracy of the second model, the Mel spectrum features, and the Mel cepstral coefficient features. The third neural network model is used to predict whether there are abnormal hidden sounds on the transmission line.
[0054] Fourthly, this application provides an anomaly detection device, the device comprising:
[0055] The second acquisition module is used to acquire audio data based on the scene where the target transmission line is located;
[0056] The processing module is used to process the audio data based on a third neural network model to obtain a prediction result; the prediction result is used to indicate whether there are any abnormal or potential hazards on the target transmission line.
[0057] The third neural network model is obtained by training a first model accuracy, a second model accuracy, Mel spectrum features, and Mel cepstral coefficient features. The first model accuracy is determined by training the first neural network model based on the Mel spectrum features. The second model accuracy is determined by training the second neural network model based on the Mel cepstral coefficient features. The Mel spectrum features and the Mel cepstral coefficient features are extracted from pre-acquired audio data.
[0058] Fifthly, this application provides an electronic device, including: a processor, and a memory communicatively connected to the processor;
[0059] The memory stores computer-executed instructions;
[0060] The processor executes computer execution instructions stored in the memory to implement the method as described in any one of the first and second aspects.
[0061] In a sixth aspect, this application provides a computer-readable storage medium storing computer-executable instructions that, when executed by a processor, are used to implement the method as described in any one of the first and second aspects.
[0062] In a seventh aspect, this application provides a computer program product, including a computer program that, when executed by a processor, implements the method as described in any one of the first and second aspects.
[0063] In summary, this application provides a model training method, anomaly detection method, apparatus, device, medium, and product. It can train a corresponding neural network model by extracting Mel-frequency spectral features and Mel-frequency cepstral coefficient features from audio data of a transmission line location, thereby determining the model accuracy. Furthermore, based on the model accuracy and the Mel-frequency spectral features and Mel-frequency cepstral coefficient features, the neural network model is trained to effectively detect potential anomalies in transmission lines. Since Mel-frequency spectral features can describe the energy distribution of audio signals in time and Mel-scale frequencies, they can capture the frequency domain content and temporal dynamics of audio signals, enabling the effective detection of potential anomalies in transmission lines. The representation of the symbol is converted into a more perceptible domain. Mel-frequency cepstral coefficients can effectively represent the short-time energy and frequency distribution of audio signals. At the same time, Mel-frequency cepstral coefficients have good robustness to noise and distortion, and can maintain good recognition performance in noisy environments. This neural network model can combine the advantages of both sound features, which can further improve the performance of detecting abnormal and potentially dangerous sounds. Since the training of this neural network model also requires the addition of model accuracy, different application scenarios and different amounts of audio data affect the numerical value of the model accuracy, so that the trained neural network model can be applied to different application scenarios, thereby improving the accuracy of identifying potentially dangerous sounds. Attached Figure Description
[0064] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0065] Figure 1 This is a schematic diagram of an application scenario provided by an embodiment of this application;
[0066] Figure 2 A schematic flowchart illustrating a model training method provided in an embodiment of this application;
[0067] Figure 3 A schematic diagram of the feature extraction process provided in the embodiments of this application;
[0068] Figure 4 This is a schematic diagram of a TDNN feature context feature extraction structure provided in an embodiment of this application;
[0069] Figure 5 A schematic diagram illustrating a process for determining hybrid voiceprint features, provided in an embodiment of this application;
[0070] Figure 6A flowchart illustrating an anomaly detection method provided in an embodiment of this application;
[0071] Figure 7 A schematic diagram of a process for detecting abnormal sounds in a power transmission line, provided as an embodiment of this application;
[0072] Figure 8 This is a schematic diagram of the structure of a model training device provided in an embodiment of this application;
[0073] Figure 9 This is a schematic diagram of the structure of an anomaly detection device provided in an embodiment of this application;
[0074] Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0075] The accompanying drawings have illustrated specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to specific embodiments. Detailed Implementation
[0076] To facilitate a clear description of the technical solutions in the embodiments of this application, the terms "first" and "second" are used in the embodiments of this application to distinguish identical or similar items with essentially the same function and purpose. For example, "first device" and "second device" are merely used to distinguish different devices and do not limit their order of execution. Those skilled in the art will understand that the terms "first" and "second" do not limit the quantity or execution order, and that "first" and "second" do not necessarily imply that they are different.
[0077] It should be noted that, in this application, the terms "exemplary" or "for example" are used to indicate that something is being described as an example, illustration, or illustration. Any embodiment or design described as "exemplary" or "for example" in this application should not be construed as being more preferred or advantageous than other embodiments or design solutions. Specifically, the use of terms such as "exemplary" or "for example" is intended to present the relevant concepts in a concrete manner.
[0078] In this application, "at least one" means one or more, and "more than one" means two or more. "And / or" describes the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can mean: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. "At least one of the following" or similar expressions refer to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or multiple.
[0079] For fault detection such as abnormal short circuits and corona discharges, short circuit discharge detection can be performed using instantaneous current values, current change rates, or true RMS values; short circuit discharge can also be performed by detecting voltage and phase angle changes during a short circuit; and short circuit discharge can be performed using machine vision detection methods. However, since abnormal short circuit discharges are usually accompanied by sound, light, and heat characteristics, voiceprints can be introduced into the substation field for fault identification. It is understandable that this voiceprint feature was initially used primarily for speaker identification.
[0080] In one possible implementation, a sound recognition device is fixedly installed on the power transmission line to monitor the ambient sound around the power transmission line. When an abnormal sound is detected, it is recorded, and the corresponding audio file is generated and uploaded to the monitoring device. The monitoring device calls the built-in sound recognition model to identify the type of sound. If the sound is identified as a potential hazard or fault sound, an early warning message is generated.
[0081] However, the aforementioned sound recognition model determines the presence of potential malfunctions by identifying the amplitude and frequency of the sound. This approach has a limited dimension for identifying potential malfunctions and is therefore less accurate.
[0082] In another possible implementation, taking transformer detection as an example, a deep neural network is used to extract the transformer acoustic signature features. Then, an optimized multivariate regression model is used to predict the predicted value of the transformer acoustic signature features corresponding to the real-time operating condition environmental factors data. The similarity between the real-time transformer acoustic signature features and the predicted value of the transformer acoustic signature features is calculated to determine whether the transformer is abnormal.
[0083] However, due to the dispersed and uncertain nature of abnormalities in transmission lines, such as abnormal short-circuit discharges, the above detection methods have limitations and cannot quickly and timely capture the sounds of potential abnormalities in transmission lines, resulting in poor accuracy in identifying potential hazards.
[0084] To address the aforementioned issues, this application provides a model training method. This method extracts Mel-spectral characteristics and Mel-cepstral coefficients from audio data of the transmission line location to train a corresponding neural network model, thereby determining the model's accuracy. Furthermore, based on the model accuracy and the Mel-spectral and Mel-cepstral coefficient characteristics, the neural network model is trained to effectively detect potential anomalies in transmission lines. Mel-spectral characteristics describe the energy distribution of audio signals over time and Mel-scale frequencies, capturing the frequency domain content and temporal dynamics of audio signals, converting the audio signal representation into a more easily perceptible domain. Mel-cepstral coefficients effectively represent the short-time energy and frequency distribution of audio signals and exhibit good robustness to noise and distortion, maintaining good recognition performance even in noisy environments. This neural network model combines the advantages of both sound features, further improving the performance of detecting potential anomalies. Since training this neural network model also requires the inclusion of model accuracy, different application scenarios and varying amounts of audio data affect the model's accuracy, allowing the trained neural network model to be applicable to different application scenarios, thus improving the accuracy of identifying potential hazard sounds.
[0085] For example, Figure 1 This is a schematic diagram of an application scenario provided in an embodiment of this application, such as... Figure 1 As shown, this application scenario can be applied to the safety inspection of transmission lines. The application scenario includes: tower mounting equipment 101 and data processing system 102; the tower mounting equipment 101 is used to collect audio data of the environment around the transmission line.
[0086] After acquiring audio data, the tower-mounted device 101 can transmit the audio data to the data processing system 102 for processing. Specifically, the data processing system 102 extracts Mel spectrum features and Mel cepstral coefficient features from the audio data. Further, it trains a first neural network based on the Mel spectrum features and determines the accuracy of the first model. It trains a second neural network model based on the Mel cepstral coefficient features and determines the accuracy of the second model. Further, it trains a third neural network model based on the accuracy of the first model, the accuracy of the second model, the Mel spectrum features, and the Mel cepstral coefficient features, so that the third neural network model can identify whether there are any abnormal or potentially dangerous sounds.
[0087] Optionally, when it is necessary to detect safety hazards of power transmission lines in a certain scenario, the audio data of the scenario can be processed using a third neural network model to detect whether there are abnormal hazard sounds on the power transmission lines in that scenario.
[0088] It should be noted that the data processing system 102 can be used to detect safety hazards in power transmission lines, or it can be used with other equipment loaded with the operating logic of the third neural network model to detect safety hazards in power transmission lines. This application embodiment does not specifically limit this.
[0089] Among them, the first neural network model, the second neural network model, and the third neural network model are all pre-trained models, that is, optimized models.
[0090] It should be noted that the embodiments of this application do not specifically limit the number of tower-mounted devices 101 or the towers on which they are distributed. For example, two or more tower-mounted devices 101 can be deployed on one tower to improve the accuracy of audio data collection, or one tower-mounted device 101 can be deployed on each tower within a preset range to increase the range of audio data collection.
[0091] It is understood that the data processing system 102 may also be a server, a server cluster, or an electronic device that performs the above processing. The embodiments of this application do not specifically limit the executing entity of the model training method.
[0092] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will be described below with reference to the accompanying drawings.
[0093] Figure 2 This is a flowchart illustrating a model training method provided in an embodiment of this application. This model training method can be applied to the aforementioned data processing system, such as... Figure 2 As shown, the model training method includes the following steps:
[0094] S201. Obtain audio data of the scene where the transmission line is located, and extract Mel spectrum features and Mel cepstral coefficient features from the audio data.
[0095] In this embodiment of the application, Mel spectral features can refer to Mel spectral graph features. The extraction process of Mel spectral features includes steps such as preprocessing, Fourier transform, and Mel filter bank. The extraction process of MFCC features includes steps such as preprocessing, Fourier transform, Mel filter bank, logarithmic operation, discrete cosine transform, and dynamic feature extraction.
[0096] Optionally, preprocessing may include pre-emphasis processing, frame splitting processing, and windowing processing. The specific process of preprocessing is not limited in the embodiments of this application. For example, preprocessing may also include wavelet transform processing, adaptive filtering processing, and spectral smoothing processing.
[0097] To better suit real-world applications, the noise portion of the original audio data can be preserved, thereby effectively improving the model's robustness and applicability in complex real-world scenarios. Therefore, preprocessing does not include noise reduction but mainly focuses on pre-emphasis processing to increase the high-frequency resolution of the speech. Assuming the nth sampling point of the audio data is x[n], the pre-emphasis formula can be expressed as follows:
[0098] y[n]=x[n]-ωx[n-1],ω=0.97(ω∈(0.9,1))
[0099] Wherein, ω is the pre-emphasis coefficient. By pre-emphasizing the audio data, the loss of the high-frequency part of the audio data during the propagation process can be compensated to a certain extent, the information of the vocal channel can be protected, and the signal-to-noise ratio of the audio data can be improved.
[0100] Because audio signals in audio data have short-term stationary characteristics, frame segmentation and windowing are necessary to facilitate subsequent Fourier transform. Specifically, the audio signal can be divided into several segments according to time, with each segment becoming a frame. For example, the frame length can be set to 25ms and the frame shift step size to 15ms, meaning there is a 10ms overlap between consecutive frames, ensuring a smooth transition between frames and preventing information leakage. Windowing involves multiplying each frame of audio signal by a specific window function, ensuring the global continuity of the segmented audio signal and facilitating subsequent frequency domain analysis.
[0101] For example, Figure 3 This is a schematic diagram of the feature extraction process provided in the embodiments of this application, such as... Figure 3 As shown, the audio signal is pre-emphasized, framed, and windowed to obtain multiple audio frames. Then, a Fourier transform is performed on each audio frame. The amplitude of the Fourier transformed audio frame is then taken, squared, and divided by the set number of Fourier points to obtain the power spectrum.
[0102] Furthermore, the power spectrum obtained above is processed through a Mel filter bank to obtain the Mel spectrum characteristics.
[0103] In this Mel filter bank, the starting point of each filter is at the midpoint of the previous filter. Optionally, the number of filters can be set to 64. Furthermore, in actual operation, each filter can be multiplied by the power spectrum to obtain the energy in that frequency band, which is the Mel spectral characteristic.
[0104] It should be noted that the first part of the MFCC extraction process is the same as the Mel spectral feature extraction process. Therefore, after obtaining the Mel spectral features, the logarithm of the Mel spectral features obtained after processing by the Mel filter bank can be taken. After the logarithmic operation, a discrete cosine transform is performed, and then the dynamic features are calculated to obtain the MFCC.
[0105] It should be noted that the audio data obtained in this step may include at least one type of hazard sound data in the environment surrounding the transmission line. This application embodiment does not specifically limit the amount of audio data. It may be audio data of the scene where the transmission line is located within a certain period of time, or audio data in a database of abnormal hazard sounds of the transmission line constructed by data augmentation and expansion methods.
[0106] S202. Train a first neural network model based on the Mel spectrum features and determine the corresponding first model accuracy; train a second neural network model based on the Mel cepstral coefficient features and determine the corresponding second model accuracy.
[0107] In this embodiment, model accuracy can refer to the accuracy and reliability of machine learning or models when predicting or classifying data. It can be measured by different metrics, such as accuracy, precision, recall, and F1 score. This embodiment does not specifically limit these metrics, and they can be selected and calculated according to specific application scenarios and user needs.
[0108] Optionally, both the first neural network model and the second neural network model are time-delay neural network (TDNN) models based on deep learning, or they can be other types of neural network models based on deep learning. This application does not specifically limit them, such as convolutional neural network models based on deep learning. It is understood that the first neural network model and the second neural network model can be the same neural network model or different neural network models.
[0109] TDNN is a feedforward neural network with time delay characteristics, capable of processing time-series data. This TDNN consists of convolutional layers, batch normalization layers, and fully connected layers. It uses the ReLU activation function to introduce non-linearity, ensuring sufficient feature representation capability. Since TDNN differs from models using only one frame of features in that it can contain temporal information from multiple frames, this application uses a time-delayed neural network, effectively delaying the weights. This significantly reduces the overall weights of the model and facilitates training.
[0110] For example, taking the extraction of feature representations from 11 frames with 5 frames each of context as an example, the features of the 11 frames can usually be directly concatenated to form an 11*D feature (D is the feature dimension of a frame), and then the model can learn the 11*D feature mapping. However, TDNN processes it differently. TDNN processes a narrower temporal context than 11 frames in the initial layer, and then feeds the processed temporal context into a deeper network to achieve feature extraction from the 11-frame feature context.
[0111] For example, Figure 4 This is a schematic diagram of a TDNN feature context feature extraction structure provided in an embodiment of this application, as shown below. Figure 4 As shown, the time resolution of the bottom layer (layer 1) is 5, the time resolution of the second layer (layer 2) is 8, the time resolution of the third layer (layer 3) is 14, and the time resolution of the top layer (layer 4) is 23.
[0112] S203. A third neural network model is trained based on the accuracy of the first model, the accuracy of the second model, the Mel spectrum features, and the Mel cepstral coefficient features. The third neural network model is used to predict whether there are abnormal hidden danger sounds on the transmission line.
[0113] In this embodiment of the application, abnormal hazard sounds can refer to the sounds corresponding to safety hazards in transmission lines, such as the sounds of short-circuit discharge caused by faults or damage to equipment such as insulators, switches, and transformers. This embodiment of the application does not specifically limit the specific types of abnormal hazard sounds.
[0114] Optionally, abnormal and potential hazards may include: short-circuit discharge sounds of transmission lines, lightning strike sounds, corona discharge sounds, and vibration sounds of towers or conductors.
[0115] The causes of discharge sounds from short circuits in transmission lines may be equipment malfunctions or external factors, such as wind, lightning strikes, or animal contact, which can lead to short circuits and discharges. This application does not specifically limit the causes of discharge sounds. Corona discharge sounds are usually caused by corona discharge phenomena resulting from high-voltage electric fields on the transmission line, producing a weak humming or popping sound, indicating an electrical problem in the transmission line. Vibration sounds from towers or conductors may be caused by loose towers or conductors on the transmission line or by external forces, producing friction or collision sounds, indicating a safety hazard in the transmission line. Alternatively, towers may vibrate under wind, producing low-frequency humming or vibration sounds, indicating a problem with the tower structure.
[0116] Therefore, after the above-mentioned abnormal and potentially dangerous sounds are generated, timely inspection and maintenance are necessary to improve the safety of the power transmission line. Thus, it is very important to accurately identify the audio data of the scene where the power transmission line is located.
[0117] This application, by identifying various potential hazards and sounds on power transmission lines, can detect potential faults and problems early, take maintenance and repair measures in advance, avoid accidents, and ensure the safe operation of power transmission lines.
[0118] In this step, the Mel spectrogram feature is a time-frequency representation of the sound signal, describing the energy distribution of the sound signal in time and Mel-scale frequency. The time dimension corresponds to the frame index, and the frequency dimension represents the Mel frequency band. Thus, the Mel spectrogram feature can capture the frequency domain content and temporal dynamics of the signal, transforming the sound signal representation into a more perceptible domain, thereby improving algorithm performance. MFCC is a cepstral coefficient calculated at the Mel scale, which can effectively represent the short-time energy and frequency distribution of the sound signal. MFCC also exhibits good robustness to noise and distortion, maintaining good recognition performance in noisy environments. Therefore, this application utilizes these two sound features to train a third neural network model. By combining the advantages of both sound features, the model's performance in detecting abnormal and potentially dangerous sounds can be improved, thereby enhancing the accuracy of the model in identifying such sounds.
[0119] Optionally, a predefined algorithm is used to calculate the accuracy of the first model, the accuracy of the second model, the Mel spectrum features, and the Mel cepstral coefficient features to obtain the hybrid voiceprint features. Further, a third neural network model is trained based on the hybrid voiceprint features. The predefined algorithm can be a weighted fusion algorithm, a feature splicing algorithm, a feature cross algorithm, etc., and this application embodiment does not specifically limit it.
[0120] Optionally, during the training of the third neural network model, feature fusion can be performed on the accuracy of the first model, the accuracy of the second model, the Mel spectrum features, and the Mel cepstral coefficient features based on the third neural network model to obtain mixed voiceprint features. Furthermore, the third neural network model can be trained based on these mixed voiceprint features.
[0121] Therefore, the process of determining the hybrid voiceprint features can be implemented before or during the training of the third neural network model, and this application does not specifically limit this.
[0122] Therefore, in training the model, this application effectively combines two types of voiceprint features with different focuses: Mel-spectral features and Mel-cepstral coefficient features. This reduces the limitations of a single feature, improves the robustness and generalization ability of the trained model, and reduces the model's sensitivity to noise and interference by fusing these two voiceprint features. This improves the model's accuracy in identifying abnormal and potential hazards in the transmission line environment, thereby enhancing the safety level of power grid equipment operation. Mel-spectral features and Mel-cepstral coefficient features each have unique advantages in describing sound signals. They capture the spectral information and spectral envelope information of the sound signal, respectively. By combining them with the model's accuracy, the expressive power of the features can be improved, better capturing the characteristics and patterns of the sound signal and better describing its features.
[0123] Optionally, training a third neural network model based on the first model accuracy, the second model accuracy, the Mel spectral features, and the Mel cepstral coefficient features includes:
[0124] The first weight parameter and the second weight parameter are determined based on the accuracy of the first model and the accuracy of the second model, respectively.
[0125] Using a predefined algorithm, the first weight parameter, the second weight parameter, the Mel spectrum feature, and the Mel cepstral coefficient feature are calculated to obtain the hybrid voiceprint feature;
[0126] A third neural network model is trained based on the hybrid voiceprint features.
[0127] In this embodiment, the hybrid voiceprint feature is based on two underlying basic voiceprint features: one is the Mel spectrogram feature, and the other is the Mel cepstral coefficient feature. The two features are combined into a hybrid voiceprint feature through a specific method.
[0128] In this step, when the predefined algorithm is a weighted fusion algorithm, the first weight parameter and the second weight parameter can be determined by the first model accuracy and the second model accuracy, respectively. Then, the mixed voiceprint feature is calculated by weighted summation, that is, mixed voiceprint feature = first weight parameter * Mel spectrum feature + second weight parameter * Mel cepstral coefficient feature.
[0129] For example, Figure 5 This is a schematic diagram of a process for determining hybrid voiceprint features provided in an embodiment of this application, such as... Figure 5 As shown, the sound signal in the audio data is preprocessed, and then Mel spectrum features and Mel cepstral coefficient features are extracted. A first neural network model is trained based on the Mel spectrum features to determine the accuracy of the first model. A second neural network model is trained based on the Mel cepstral coefficient features to determine the accuracy of the second model. Furthermore, based on the accuracy of the first model and the accuracy of the second model, the mixing weight parameters, namely the first weight parameter and the second weight parameter, are determined. Then, the mixed voiceprint features are calculated by weighted summation.
[0130] It should be noted that the blending weights will change depending on the audio data, so the blending weights can be updated. For example, when the amount of audio data changes, the blending weights will also change.
[0131] Optionally, the formula corresponding to the predefined algorithm can be:
[0132] F = α*f1 + β*f2
[0133] Where F represents the mixed voiceprint feature; f1 represents the Mel spectrum feature; and f2 represents the Mel cepstral coefficient feature. Indicates the first weight parameter; S1 represents the second weight parameter; S2 represents the first model accuracy; S3 represents the second model accuracy.
[0134] It should be noted that the type or size of the acquired audio data may vary in different application scenarios. Therefore, α and β may also be different. The specific values of α and β are not limited in the embodiments of this application, but are determined based on the application scenario.
[0135] In this way, after obtaining the weight parameters of the two features, the hybrid voiceprint features can be constructed. Based on the above formula, the hybrid voiceprint features are determined, which greatly improves the calculation speed.
[0136] Optionally, when the first weight parameter is greater than the second weight parameter, a multiple can be assigned to the first weight parameter, such as mixed voiceprint feature = 2 * first weight parameter * Mel spectrum feature + second weight parameter * Mel cepstral coefficient feature. Correspondingly, when the first weight parameter is less than the second weight parameter, a multiple can be assigned to the second weight parameter. The embodiments of this application do not specifically limit the value of the multiple.
[0137] Optionally, the predefined algorithm can also be a weighted subtraction algorithm, or an algorithm that takes the maximum value after weighting. This application does not specifically limit the predefined algorithm.
[0138] Furthermore, after obtaining the mixed voiceprint features, a third neural network model can be trained based on the mixed voiceprint features.
[0139] Therefore, the embodiments of this application can calculate mixed voiceprint features based on weight parameters, Mel spectrum features, and Mel cepstral coefficient features. This has unique advantages in capturing different aspects of the sound signal. By adding the limitation of mixed weight parameters, Mel spectrum features and Mel cepstral coefficient features can be effectively fused, which can more comprehensively describe the sound signal and improve the expressive power of voiceprint features. The mixed voiceprint features are calculated using a predefined algorithm, which improves the flexibility in determining the mixed voiceprint features. Furthermore, the mixed voiceprint features are used to train a third neural network model, thereby improving the accuracy of the third neural network model in identifying abnormal and potentially dangerous sounds.
[0140] Optionally, acquire audio data of the scene where the transmission line is located, and extract Mel-spectral features and Mel-cepstral coefficient features from the audio data, including:
[0141] Collect at least one type of audio data within a preset range of the power transmission line, and construct a database based on the at least one type of audio data;
[0142] For each type, the audio data is preprocessed to obtain multiple audio frames;
[0143] Mel-spectral features of the multiple audio frames are extracted respectively;
[0144] The Mel spectrum features are processed to obtain the Mel cepstral coefficient features.
[0145] In this embodiment of the application, the preset range refers to a certain range corresponding to the surrounding environment of the transmission line, such as the surrounding environment of the transmission line with a radius of 200m centered on a certain tower. This embodiment of the application does not specifically limit the size of the preset range.
[0146] For example, audio data of the surrounding environment of a real transmission line is collected. This audio data is collected continuously by the tower-mounted equipment and uploaded to a storage server in 10-second increments, and then a database is built based on this audio data.
[0147] Optionally, the audio data in the database may include at least one type of hazard sound data collected from the actual environment surrounding the transmission line, at least one type of hazard sound data generated by simulating hazards, and at least one type of hazard sound data obtained by data enhancement and expansion based on the actual hazard sound data. This application embodiment does not specifically limit the amount and content of the audio data in the database.
[0148] In this step, for each audio frame, Mel spectral features are extracted and processed to obtain Mel cepstral coefficient features. The corresponding processing procedure is similar to the process described in S201 above. For details, please refer to the description in S201, which will not be repeated here. The process of processing Mel spectral features includes logarithmic operation, discrete cosine transform, and dynamic feature extraction.
[0149] It should be noted that the embodiments of this application do not specifically limit the type of audio data, which can be at least one of the following: short-circuit discharge sound of transmission lines, lightning strike sound, corona discharge sound, and vibration sound of towers or conductors. It is understood that the type of audio data can also include normal sound types.
[0150] Therefore, the embodiments of this application utilize at least one type of audio data to construct a database, which can provide more data samples and diversity, helping the model to learn and generalize better, thereby covering a wider range of scenarios and situations. This application can also obtain Mel cepstral coefficient features based on the processing of Mel spectral features, simplifying the Mel cepstral coefficient feature extraction process to capture the spectral features of audio signals and ensuring the convenience of feature extraction.
[0151] Optionally, a database is constructed based on the at least one type of audio data, including:
[0152] Obtain annotation information for the at least one type of audio data, the annotation information indicating the type and clarity score of the audio data; determine the first type of audio data based on the annotation information.
[0153] Based on the first audio data, decibel value adjustment processing is performed to obtain the second audio data, and based on the first audio data, data augmentation processing is performed to obtain the third audio data;
[0154] The audio data in the database is constructed based on the first audio data, the second audio data, and / or the third audio data. In this embodiment, the annotation information is obtained by judging the type and clarity of the audio data based on a multi-person annotation and scoring mechanism. For example, multiple people are assigned to annotate and score each audio data, resulting in multiple annotation files. The annotation files contain category labels for all hidden dangers in the transmission line environment. If the audio data contains a clear sound of a hidden danger and the category of the hidden danger is 100% confirmed, it is marked as 1 under that category label. If the sound is weak but the category of the hidden danger can be confirmed, it is marked as 2 under that category label. If the sound is suspected but the category of the hidden danger cannot be determined, it can be assigned a value between 5 and 9 according to the sound strength and specific circumstances. Audio data without hidden dangers is not labeled with any value.
[0155] In each annotation file, the value of the category label is the clarity score.
[0156] Due to the uncertainty of the multi-person annotation and scoring mechanism, there may be multiple different annotation files for the same hazard category. In order to ensure the quality of audio data, it is necessary to filter the annotation files to select the first audio data belonging to the current hazard category.
[0157] Optionally, determining the first audio data of the at least one type based on the annotation information includes:
[0158] For each type of audio data, a clarity score is assigned, and a weighting coefficient for the audio data is determined based on the clarity score.
[0159] The audio score is obtained by calculating the clarity score and the weighting coefficient using a weighted algorithm.
[0160] Determine whether the audio score is greater than the first threshold;
[0161] If so, then the audio data is determined to be the first audio data.
[0162] In this embodiment, the first threshold is a preset threshold set in advance to determine whether the audio data belongs to the current hazard category. It can be set based on experimental data or by the user. This embodiment does not specifically limit the size of the first threshold.
[0163] In this step, for audio data of the same type, based on the multiple annotation files corresponding to the audio data, the values under the category labels of each annotation file are obtained. Each value is multiplied by a pre-set weight coefficient, and the calculated audio score exceeds the first threshold. If the audio data belongs to the current hazard category, the weight score is determined by the user in advance. For example, if a user's judgment is more accurate, the weight corresponding to the value under the category label of the annotation file judged by that user will be greater.
[0164] In this way, we can reasonably filter out the more accurate audio data for each type, that is, the real hidden danger sounds.
[0165] After obtaining accurate first audio data, the first audio data can be simulated and collected according to the characteristics of abnormal hazard sounds, and then the decibel value can be adjusted to obtain second audio data. Optionally, the first audio data can also be expanded using data augmentation methods to obtain third audio data. Then, a database of abnormal hazard sounds of transmission lines can be constructed based on the first audio data, the second audio data, and / or the third audio data.
[0166] Therefore, the embodiments of this application can construct a database based on audio data of a determined and accurate category, ensuring the rationality of the audio data in the database, and effectively compensate for the small sample problem in different scenarios through numerical adjustment of abnormal and potential sound and data amplification methods.
[0167] Optionally, the second audio data is obtained by adjusting the decibel value based on the first audio data, including:
[0168] Acquire sound data simulating at least one of the aforementioned types of audio data;
[0169] For each type of sound data, determine whether the sound data exceeds a second threshold;
[0170] If so, the average decibel value of the first audio data of the aforementioned type is calculated, and the sound data is adjusted based on the average decibel value to obtain the second audio data.
[0171] In this embodiment of the application, the sound data of simulated audio data refers to the sound data obtained by randomly placing a sound source similar to the first audio data in different locations in a real application scenario and collecting the sound using a device.
[0172] In this step, for a certain type, strong change detection is performed on all collected sound data. When the audio of the sound data exceeds the second threshold, it is considered that the clarity of the sound containing the potential hazard and the sound information contained therein meet the requirements of high-quality sound data. Furthermore, based on the average decibel value calculated from the real collected potential hazard sound (first audio data), the sound data collected from the simulated sound source is randomly adjusted within a specific range to obtain the second audio data.
[0173] It should be noted that the embodiments of this application do not specifically limit the size of the second threshold or the value of the decibel adjustment, which can be determined based on the actual application scenario.
[0174] Thus, the embodiments of this application can expand the audio data in the database by simulating the sound of potential hazards. By adjusting the simulated sound data using the mean of the first audio data, the second audio data is obtained, making the second audio data closer to the real sound of potential hazards, thus ensuring the authenticity and accuracy of the audio data in the database.
[0175] Optionally, based on the first audio data, data augmentation processing is performed to obtain third audio data, including:
[0176] The data augmentation strategy is determined based on the application scenario requirements; the data augmentation strategy includes at least one of the following: random velocity perturbation strategy, random displacement perturbation strategy, random aliasing strategy, and random cropping and splicing strategy;
[0177] The priority of the data augmentation strategy is determined, and based on the priority, the data augmentation strategy is used to sequentially augment the first audio data to obtain the third audio data.
[0178] In this embodiment, the random speed perturbation strategy refers to a strategy that changes the frequency of an audio signal by altering its playback speed, thereby changing the pitch and tone of the audio. The random speed perturbation strategy can simulate audio input at different speeds, helping the model better adapt to audio data at different speeds.
[0179] The speed perturbation value ranges from 0.9 to 1.1. Since the audio data is perturbed by randomly selecting perturbation values within a certain range, the playback speed will increase or decrease depending on the adjustment.
[0180] Random displacement perturbation strategy refers to random time shifting of audio data, that is, the strategy of randomly moving the audio signal on the time axis. Random displacement perturbation strategy can simulate audio input at different time points and increase the diversity of data.
[0181] For example, the first audio data can be randomly shifted forward or backward by a few milliseconds along the time axis, and the shift value can be set to a range of -5ms to 5ms.
[0182] Random aliasing strategies can refer to mixing two or more audio signals. This can be done by simple superposition or by mixing through weighted averaging, or by mixing audio signals with different types of noise signals. Random aliasing strategies can simulate the simultaneous presence of multiple sounds and the audio characteristics of audio data in noisy environments, increasing the diversity of the data.
[0183] For example, the first audio data and the simulated sound data can be overlaid with background sounds or similar sounds that frequently occur in the power transmission line scenario to increase the complexity of the data.
[0184] Random cropping and splicing strategy refers to randomly selecting two audio frames from the same type of audio data, randomly initializing a cropping and splicing point, cropping the two audio frames according to the cropping and splicing point, and then cross-sponging the cropped segments of the two audio frames to form two new audio data. Random cropping and splicing strategy can generate more diverse audio data.
[0185] It should be noted that the priority of the data expansion strategy is set in advance. It can be set by the user in advance or according to the needs of the application scenario. This application embodiment does not make specific limitations on this.
[0186] For example, an audio data to be enhanced is input, such as a first audio data. At least one of the data augmentation strategies is randomly selected to form a random combination to augment 15 kinds of sound data. During the augmentation, the execution order of the data augmentation strategies is determined based on the priority of each data augmentation strategy in the random combination. The data augmentation is performed in sequence, and the augmented third audio data is output.
[0187] For example, the random combination includes random velocity perturbation strategy, random displacement perturbation strategy, random aliasing strategy, and random pruning and splicing strategy. Based on priority, the execution order is determined as random velocity perturbation strategy, random displacement perturbation strategy, random aliasing strategy, and random pruning and splicing strategy. Then, the first audio data is augmented sequentially according to these strategies to obtain the third audio data. It should be noted that this application also focuses on collecting and labeling other types of sounds present around the transmission line, mainly as negative samples for the model to learn, further improving the model's generalization ability. Optionally, all audio data in the database are in single-channel format, with an audio duration of 10 seconds and a sampling rate of 44100Hz.
[0188] Understandably, the above process can effectively expand audio data, such as expanding audio data collected from real-world scenarios to more than 10 hours.
[0189] In this way, this application can augment the first audio data based on different data augmentation strategies, which can effectively expand the audio data and generate more diverse training samples, thereby improving the flexibility of data augmentation.
[0190] Optionally, both the first neural network model and the second neural network model are time-delay neural network models based on deep learning. The first neural network model is trained based on the Mel-spectral characteristics, and the corresponding first model accuracy is determined. The second neural network model is trained based on the Mel-cephalic coefficient characteristics, and the corresponding second model accuracy is determined, including:
[0191] The audio data in the database is divided into a training set and a test set according to a specific ratio;
[0192] The Mel-spectral features extracted from the training set and the type of audio data are input into the first neural network model for training to obtain the trained first neural network model; the Mel-cephalic coefficient features extracted from the training set and the type of audio data are input into the second neural network model for training to obtain the trained second neural network model.
[0193] The trained first neural network model and the trained second neural network model are tested based on the test set to determine the accuracy of the first model and the accuracy of the second model.
[0194] This application constructs an audio data database and divides it into a training set and a test set according to a specific ratio. Then, it extracts Mel spectrum features and Mel cepstral coefficient features according to a process similar to S201 above. The specific ratio can be a random ratio or a user-defined ratio. This application does not specifically limit this ratio. The training set is used to train the first neural network model and the second neural network model. The test set is used to calculate the model accuracy corresponding to the trained first neural network model and the second neural network model, which can be represented by an accuracy score.
[0195] For example, the Mel spectrum features extracted from the training set and the type of audio data are input into the first time-delay neural network model for training, and continuously iterated and optimized until the loss function value of the first time-delay neural network model meets the preset conditions. If it is less than the preset value, then the trained first time-delay neural network model is obtained.
[0196] For example, taking the first TDNN model as input to Tx64-dimensional Mel spectrum features, the first TDNN model contains multiple hidden layers, each containing multiple neurons. Each neuron can receive input data from the current time and several previous time steps. After linear transformation of weights and biases, it undergoes nonlinear transformation through activation functions. Then, it is trained and optimized for 100 rounds using the Adam optimizer and backpropagation algorithm to output the trained first TDNN model.
[0197] It is understandable that the process of training the second neural network model using Mel-Cepstral Coefficients features and the type of audio data is similar to the process of training the first neural network model. For details, please refer to the description of the above embodiments, which will not be repeated here.
[0198] Furthermore, after the two models have been trained, each model is tested using a test set to obtain the accuracy score for each model.
[0199] It should be noted that during the extraction of Mel spectral features and Mel cepstral coefficient features from the training and testing sets, the feature dimensions during the extraction process need to be controlled to ensure that the feature dimensions of the two different features are consistent. When using the extracted features to train the first TDNN model and the second TDNN model of the same specifications, it is necessary to ensure that the model training parameters such as the optimizer, learning rate, and number of iterations are consistent.
[0200] Thus, by training a time-delay neural network model, this embodiment of the application can improve the model's performance, generalization ability, and robustness, thereby better adapting to different audio processing tasks and application scenarios. Furthermore, by using Mel spectral features and Mel cepstral coefficient features in combination with the type of audio data as input, the first TDNN model and the second TDNN model are trained respectively, improving the training accuracy of the models and thus helping to obtain reasonable accuracy for the first and second models.
[0201] Optionally, the third neural network model is a time-delay neural network model based on deep learning. Training the third neural network model based on the hybrid voiceprint features includes:
[0202] Identify at least one type of mixed voiceprint features;
[0203] The at least one type of mixed voiceprint features and the type of audio data are input into the third neural network model for training, resulting in a trained third neural network model.
[0204] For example, after training the first TDNN model and the second TDNN model, the trained model is tested using a test set to obtain the first model accuracy corresponding to the first TDNN model and the second model accuracy corresponding to the second TDNN model.
[0205] Furthermore, based on the model accuracy of the two models, the feature mixing weights required to construct the hybrid voiceprint features are calculated, i.e. After obtaining the two feature weight parameters, the hybrid voiceprint features can be constructed.
[0206] Furthermore, a third TDNN model is trained using at least one type of mixed voiceprint features and audio data type to obtain a trained third TDNN model, which is used for the detection of abnormal and potentially dangerous sounds.
[0207] It should be noted that the mixing weight parameters involved in the above-mentioned mixed feature construction process will change as the amount of training data accumulates. Therefore, when the audio data accumulates to a certain extent, the mixing weight parameters also need to be updated.
[0208] Since the data in the training set includes at least one type of audio data, at least one type of Mel spectral feature and at least one type of Mel cepstral coefficient feature can also be extracted from the training set. Accordingly, at least one type of mixed voiceprint feature can be determined, that is, the type of audio data corresponds to the type of mixed voiceprint feature.
[0209] Thus, the time-delay neural network construction method based on mixed voiceprint features proposed in this application combines at least one type of mixed voiceprint features and the type of audio data to train a third TDNN model of deep learning, so as to realize the continuous and effective detection of abnormal hazards in transmission lines through the mixed voiceprint features of sound. This can improve the model's recognition accuracy of abnormal hazard sounds in transmission lines and provide technical support for abnormal detection of transmission lines.
[0210] For example, Figure 6 This is a flowchart illustrating an anomaly detection method provided in an embodiment of this application, as shown below. Figure 6 As shown, this anomaly detection method can be applied to the aforementioned data processing system, and can also be deployed on other devices, such as... Figure 6 As shown, the anomaly detection method includes the following steps:
[0211] S601. Obtain audio data of the scene where the target transmission line is located;
[0212] S602. The audio data is processed based on a third neural network model to obtain a prediction result; the prediction result is used to indicate whether there are any abnormal or potential hazards on the target transmission line.
[0213] The third neural network model is obtained by training a first model accuracy, a second model accuracy, Mel spectrum features, and Mel cepstral coefficient features. The first model accuracy is determined by training the first neural network model based on the Mel spectrum features. The second model accuracy is determined by training the second neural network model based on the Mel cepstral coefficient features. The Mel spectrum features and the Mel cepstral coefficient features are extracted from pre-acquired audio data.
[0214] In this embodiment of the application, the scene where the target transmission line is located is the scene where the transmission line to be detected is located. This embodiment of the application does not specifically limit the scene where the target transmission line is located.
[0215] In this step, the audio data is processed based on the third neural network model. This can involve extracting the target Mel spectral features and target Mel cepstral coefficient features from the audio data, and obtaining the pre-determined first and second weight parameters. Further, the target mixed voiceprint features are calculated using the formula: first weight parameter * target Mel spectral features + second weight parameter * target Mel cepstral coefficient features. The target mixed voiceprint features are then input into the third neural network model to obtain the prediction result. The first and second weight parameters are determined by the accuracy of the first and second models, as detailed in the above embodiment, and will not be repeated here.
[0216] Optionally, processing the audio data based on the third neural network model can also involve extracting the target Mel spectrum features and target Mel cepstral coefficient features from the audio data, obtaining the pre-determined first model accuracy and second model accuracy, and inputting the first model accuracy, second model accuracy, Mel spectrum features, and Mel cepstral coefficient features into the third neural network model to obtain the prediction results.
[0217] For example, Figure 7 This is a flowchart illustrating an abnormal sound detection process for power transmission lines, as provided in an embodiment of this application. Figure 7 As shown, taking the third neural network model as the TDNN model as an example, after the device continuously collects audio data (sound signals) around the transmission line, a series of preprocessing steps are performed on the collected sound signals. Then, Mel spectrum features and Mel cepstral coefficient features are extracted respectively. Next, the mixed voiceprint feature determination method in the above embodiment is used to obtain the mixed voiceprint features input to the TDNN model. The mixed voiceprint features are input into the TDNN model to obtain the prediction result. The prediction result determines whether the current sound is an abnormal hazard sound. When it is determined to be an abnormal hazard sound, an alarm can be issued; otherwise, no measures are taken.
[0218] Optionally, after obtaining the prediction results, the accuracy of the prediction results can be manually verified on-site. If so, the data in the database can be updated based on the audio data of the scene where the target transmission line is located and the prediction results, thereby continuously enriching the database.
[0219] Since the third neural network model can effectively combine two different voiceprint features and better describe the characteristics of sound signals by overcoming the limitations of model accuracy, it is applicable to different application scenarios. Therefore, the embodiments of this application can use the third neural network model to accurately identify abnormal and potentially dangerous sounds.
[0220] In the foregoing embodiments, the model training method and anomaly detection method provided in the embodiments of this application have been described. To implement the functions of the methods provided in the embodiments of this application, the electronic device serving as the execution subject may include hardware structures and / or software modules, implementing the above functions in the form of hardware structures, software modules, or a combination of hardware structures and software modules. Whether a particular function is executed in the form of hardware structures, software modules, or a combination of hardware structures and software modules depends on the specific application and design constraints of the technical solution.
[0221] For example, Figure 8 This is a schematic diagram of the structure of a model training device provided in an embodiment of this application, as shown below. Figure 8 As shown, the device 800 includes: a first acquisition module 801, used to acquire audio data of the scene where the transmission line is located, and extract Mel spectrum features and Mel cepstral coefficient features from the audio data;
[0222] The determination module 802 is used to train a first neural network model based on the Mel spectrum features and determine the corresponding first model accuracy, and to train a second neural network model based on the Mel cepstral coefficient features and determine the corresponding second model accuracy.
[0223] Training module 803 is used to train a third neural network model based on the accuracy of the first model, the accuracy of the second model, the Mel spectrum features, and the Mel cepstral coefficient features. The third neural network model is used to predict whether there are abnormal hidden sounds on the transmission line.
[0224] Optionally, the training module 803 includes a determination unit, a calculation unit, and a training unit;
[0225] The determining unit is configured to determine the first weight parameter and the second weight parameter based on the first model accuracy and the second model accuracy, respectively.
[0226] The calculation unit is used to calculate the first weight parameter, the second weight parameter, the Mel spectrum feature, and the Mel cepstral coefficient feature using a predefined algorithm to obtain the mixed voiceprint feature;
[0227] The training unit is used to train a third neural network model based on the hybrid voiceprint features.
[0228] Optionally, the formula corresponding to the predefined algorithm is:
[0229] F = α*f1 + β*f2
[0230] Where F represents the mixed voiceprint feature; f1 represents the Mel spectrum feature; and f2 represents the Mel cepstral coefficient feature. Indicates the first weight parameter; S1 represents the second weight parameter; S2 represents the first model accuracy; S3 represents the second model accuracy.
[0231] Optionally, the training unit is specifically used for:
[0232] Identify at least one type of mixed voiceprint features;
[0233] The at least one type of mixed voiceprint features and the type of audio data are input into the third neural network model for training, resulting in a trained third neural network model.
[0234] Optionally, the first acquisition module 801 includes a construction unit, a preprocessing unit, an extraction unit, and a first processing unit;
[0235] The construction unit is used to collect at least one type of audio data within a preset range of the transmission line and construct a database based on the at least one type of audio data.
[0236] The preprocessing unit is used to preprocess the audio data for each type to obtain multiple audio frames;
[0237] The extraction unit is used to extract the Mel-spectral features of the multiple audio frames respectively;
[0238] The first processing unit is used to process the Mel spectrum features to obtain Mel cepstral coefficient features.
[0239] Optionally, the construction unit includes an acquisition unit, a second processing unit, and a construction subunit;
[0240] The acquisition unit is configured to acquire annotation information of the at least one type of audio data, wherein the annotation information is used to indicate the type and clarity score of the audio data; and to determine the first type of audio data based on the annotation information.
[0241] The second processing unit is used to perform decibel adjustment processing based on the first audio data to obtain second audio data, and to perform data augmentation processing based on the first audio data to obtain third audio data;
[0242] The construction subunit is used to construct audio data in the database based on the first audio data, the second audio data, and / or the third audio data.
[0243] Optionally, the acquisition unit is specifically used for:
[0244] For each type of audio data, a clarity score is assigned, and a weighting coefficient for the audio data is determined based on the clarity score.
[0245] The audio score is obtained by calculating the clarity score and the weighting coefficient using a weighted algorithm.
[0246] Determine whether the audio score is greater than the first threshold;
[0247] If so, then the audio data is determined to be the first audio data.
[0248] Optionally, the second processing unit is specifically used for:
[0249] Acquire sound data simulating at least one of the aforementioned types of audio data;
[0250] For each type of sound data, determine whether the sound data exceeds a second threshold;
[0251] If so, the average decibel value of the first audio data of the aforementioned type is calculated, and the sound data is adjusted based on the average decibel value to obtain the second audio data.
[0252] Optionally, the processing subunit is specifically used for:
[0253] The data augmentation strategy is determined based on the application scenario requirements; the data augmentation strategy includes at least one of the following: random velocity perturbation strategy, random displacement perturbation strategy, random aliasing strategy, and random cropping and splicing strategy;
[0254] The priority of the data augmentation strategy is determined, and based on the priority, the data augmentation strategy is used to sequentially augment the first audio data to obtain the third audio data.
[0255] Optionally, both the first neural network model and the second neural network model are time-delay neural network models based on deep learning, and the determining module 802 is specifically used for:
[0256] The audio data in the database is divided into a training set and a test set according to a specific ratio;
[0257] The Mel-spectral features extracted from the training set and the type of audio data are input into the first neural network model for training to obtain the trained first neural network model; the Mel-cephalic coefficient features extracted from the training set and the type of audio data are input into the second neural network model for training to obtain the trained second neural network model.
[0258] The trained first neural network model and the trained second neural network model are tested based on the test set to determine the accuracy of the first model and the accuracy of the second model.
[0259] Optionally, the abnormal and potential hazards include: short-circuit discharge sound of transmission lines, lightning strike sound, corona discharge sound, and vibration sound of towers or conductors.
[0260] It should be noted that the specific implementation principle and effect of the above model training device can be found in the relevant description and effect of the above embodiments, and will not be elaborated further here.
[0261] For example, Figure 9 This is a schematic diagram of the structure of an anomaly detection device provided in an embodiment of this application, as shown below. Figure 9 As shown, the device 900 includes: a second acquisition module 901, used to acquire audio data of the scene where the target transmission line is located;
[0262] Processing module 902 is used to process the audio data based on a third neural network model to obtain a prediction result; the prediction result is used to indicate whether there are any abnormal or potential hazards on the target transmission line.
[0263] The third neural network model is obtained by training a first model accuracy, a second model accuracy, Mel spectrum features, and Mel cepstral coefficient features. The first model accuracy is determined by training the first neural network model based on the Mel spectrum features. The second model accuracy is determined by training the second neural network model based on the Mel cepstral coefficient features. The Mel spectrum features and the Mel cepstral coefficient features are extracted from pre-acquired audio data.
[0264] It should be noted that the specific implementation principle and effect of the above-mentioned anomaly detection device can be found in the relevant description and effect of the above embodiments, and will not be elaborated further here.
[0265] This application also provides a schematic diagram of the structure of an electronic device. Figure 10 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application, such as... Figure 10 As shown, the electronic device may include: a processor 1001 and a memory 1002 communicatively connected to the processor; the memory 1002 stores a computer program; the processor 1001 executes the computer program stored in the memory 1002, causing the processor 1001 to perform the method described in any of the above embodiments.
[0266] The memory 1002 and the processor 1001 can be connected via the bus 1003.
[0267] This application also provides a computer-readable storage medium storing computer program execution instructions, which, when executed by a processor, are used to implement the methods described in any of the foregoing embodiments of this application.
[0268] This application also provides a chip for executing instructions, which is used to perform the methods described in any of the foregoing embodiments executed by an electronic device as described in any of the foregoing embodiments of this application.
[0269] This application also provides a computer program product, which includes a computer program that, when executed by a processor, can implement the methods described in any of the foregoing embodiments executed by an electronic device as described in any of the foregoing embodiments of this application.
[0270] In the several embodiments provided in this application, it should be understood that the disclosed apparatus and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of modules is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple modules or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or modules may be electrical, mechanical, or other forms.
[0271] The modules described as separate components may or may not be physically separate. The components shown as modules may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the modules can be selected to implement the solution of this embodiment according to actual needs.
[0272] Furthermore, the functional modules in the various embodiments of this application can be integrated into one processing unit, or each module can exist physically separately, or two or more modules can be integrated into one unit. The unit composed of the above modules can be implemented in hardware or in the form of hardware plus software functional units.
[0273] The integrated modules implemented as software functional modules described above can be stored in a computer-readable storage medium. These software functional modules, stored in a storage medium, include several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) or processor to execute some steps of the methods described in the various embodiments of this application.
[0274] It should be understood that the aforementioned processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. A general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in the application can be directly manifested as being executed by a hardware processor, or executed by a combination of hardware and software modules within the processor.
[0275] The memory may include high-speed random access memory (RAM) or non-volatile memory (NVM), such as at least one disk storage device, and may also be a USB flash drive, external hard drive, read-only memory, disk or optical disc, etc.
[0276] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0277] The aforementioned storage media can be implemented from any type of volatile or non-volatile storage device or a combination thereof, such as Static Random-Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read-Only Memory (EPROM), Programmable Read-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The storage media can be any available medium accessible to general-purpose or special-purpose computers.
[0278] An exemplary storage medium is coupled to a processor, enabling the processor to read information from and write information to the storage medium. Alternatively, the storage medium can be an integral part of the processor. Both the processor and the storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and storage medium can exist as discrete components in an electronic device or host device.
[0279] It should be noted that, for the sake of simplicity, the foregoing method embodiments are all described as a series of actions. However, those skilled in the art should understand that this application is not limited to the described order of actions, as some steps may be performed in other orders or simultaneously according to this application. Furthermore, those skilled in the art should also understand that the embodiments described in the specification are all optional embodiments, and the actions and modules involved are not necessarily essential to this application.
[0280] It should be further noted that although the steps in the flowchart are shown sequentially according to the arrows, these steps are not necessarily executed in the order indicated by the arrows. Unless explicitly stated herein, there is no strict order restriction on the execution of these steps, and they can be executed in other orders. Moreover, at least some steps in the flowchart may include multiple sub-steps or multiple stages. These sub-steps or stages are not necessarily completed at the same time, but can be executed at different times. The execution order of these sub-steps or stages is not necessarily sequential, but can be performed alternately or in turn with other steps or at least some of the sub-steps or stages of other steps.
[0281] In the above embodiments, the descriptions of each embodiment have their own emphasis. For parts not described in detail in a certain embodiment, please refer to the relevant descriptions of other embodiments. The technical features of the above embodiments can be combined arbitrarily. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as the combination of these technical features does not contradict each other, it should be considered within the scope of this specification.
[0282] Other embodiments of this application will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This application is intended to cover any variations, uses, or adaptations of this application that follow the general principles of this application and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this application are indicated by the claims.
[0283] The above description is merely a specific implementation of the embodiments of this application, but the protection scope of the embodiments of this application is not limited thereto. Any changes or substitutions within the technical scope disclosed in the embodiments of this application should be covered within the protection scope of the embodiments of this application. Therefore, the protection scope of the embodiments of this application should be determined by the protection scope of the claims.
Claims
1. A model training method, characterized in that, The method includes: Acquire audio data of the scene where the transmission line is located, and extract Mel spectrum features and Mel cepstral coefficient features from the audio data; A first neural network model is trained based on the Mel spectrum features, and the corresponding first model accuracy is determined. A second neural network model is trained based on the Mel cepstral coefficient features, and the corresponding second model accuracy is determined. A third neural network model is trained based on the accuracy of the first model, the accuracy of the second model, the Mel spectrum features, and the Mel cepstral coefficient features. The third neural network model is used to predict whether there are abnormal or potential hazards on the transmission line.
2. The method according to claim 1, characterized in that, Training a third neural network model based on the accuracy of the first model, the accuracy of the second model, the Mel spectrum features, and the Mel cepstral coefficient features includes: The first weight parameter and the second weight parameter are determined based on the accuracy of the first model and the accuracy of the second model, respectively. Using a predefined algorithm, the first weight parameter, the second weight parameter, the Mel spectrum feature, and the Mel cepstral coefficient feature are calculated to obtain the hybrid voiceprint feature; A third neural network model is trained based on the hybrid voiceprint features.
3. The method according to claim 2, characterized in that, The formula corresponding to the predefined algorithm is: F = α*f1 + β*f2 Where F represents the mixed voiceprint feature; f1 represents the Mel spectrum feature; and f2 represents the Mel cepstral coefficient feature. Indicates the first weight parameter; S1 represents the second weight parameter; S2 represents the first model accuracy; S3 represents the second model accuracy.
4. The method according to claim 2, characterized in that, The third neural network model is a time-delay neural network model based on deep learning. Training the third neural network model based on the hybrid voiceprint features includes: Identify at least one type of mixed voiceprint features; The at least one type of mixed voiceprint features and the type of audio data are input into the third neural network model for training, resulting in a trained third neural network model.
5. The method according to claim 1, characterized in that, Acquire audio data of the scene where the transmission line is located, and extract Mel-spectral features and Mel-cepstral coefficient features from the audio data, including: Collect at least one type of audio data within a preset range of the power transmission line, and construct a database based on the at least one type of audio data; For each type, the audio data is preprocessed to obtain multiple audio frames; Mel-spectral features of the multiple audio frames are extracted respectively; The Mel spectrum features are processed to obtain the Mel cepstral coefficient features.
6. The method according to claim 5, characterized in that, Constructing a database based on at least one type of audio data, including: Obtain annotation information for at least one type of audio data, the annotation information being used to indicate the type and clarity score of the audio data; determine first audio data of at least one type based on the annotation information; Based on the first audio data, decibel value adjustment processing is performed to obtain the second audio data, and based on the first audio data, data augmentation processing is performed to obtain the third audio data; The audio data in the database is constructed based on the first audio data, the second audio data, and / or the third audio data.
7. The method according to claim 6, characterized in that, Determining the first audio data of at least one type based on the annotation information includes: For each type of audio data, a clarity score is assigned, and a weighting coefficient for the audio data is determined based on the clarity score. The audio score is obtained by calculating the clarity score and the weighting coefficient using a weighted algorithm. Determine whether the audio score is greater than the first threshold; If so, then the audio data is determined to be the first audio data.
8. The method according to claim 6, characterized in that, Based on the first audio data, decibel adjustment processing is performed to obtain the second audio data, including: Acquire sound data simulating at least one of the aforementioned types of audio data; For each type of sound data, determine whether the sound data exceeds a second threshold; If so, the average decibel value of the first audio data of the aforementioned type is calculated, and the sound data is adjusted based on the average decibel value to obtain the second audio data.
9. The method according to claim 6, characterized in that, Based on the first audio data, data augmentation processing is performed to obtain the third audio data, including: The data augmentation strategy is determined based on the application scenario requirements; the data augmentation strategy includes at least one of the following: random velocity perturbation strategy, random displacement perturbation strategy, random aliasing strategy, and random cropping and splicing strategy; The priority of the data augmentation strategy is determined, and based on the priority, the data augmentation strategy is used to sequentially augment the first audio data to obtain the third audio data.
10. The method according to claim 5, characterized in that, Both the first neural network model and the second neural network model are time-delay neural network models based on deep learning. The first neural network model is trained based on the Mel-frequency spectral features, and the corresponding first model accuracy is determined. The second neural network model is trained based on the Mel-frequency cepstral coefficient features, and the corresponding second model accuracy is determined, including: The audio data in the database is divided into a training set and a test set according to a specific ratio; The Mel-spectral features extracted from the training set and the type of audio data are input into the first neural network model for training to obtain the trained first neural network model; the Mel-cephalic coefficient features extracted from the training set and the type of audio data are input into the second neural network model for training to obtain the trained second neural network model. The trained first neural network model and the trained second neural network model are tested based on the test set to determine the accuracy of the first model and the accuracy of the second model.
11. The method according to any one of claims 1-10, characterized in that, The abnormal and potentially hazardous sounds include: short-circuit discharge sounds from transmission lines, lightning strike sounds, corona discharge sounds, and vibration sounds from towers or conductors.
12. An anomaly detection method, characterized in that, The method includes: Acquire audio data of the scene where the target transmission line is located; The audio data is processed based on a third neural network model to obtain a prediction result; the prediction result is used to indicate whether there are any abnormal or potential hazards on the target transmission line. The third neural network model is obtained by training a first model accuracy, a second model accuracy, Mel spectrum features, and Mel cepstral coefficient features. The first model accuracy is determined by training the first neural network model based on the Mel spectrum features. The second model accuracy is determined by training the second neural network model based on the Mel cepstral coefficient features. The Mel spectrum features and the Mel cepstral coefficient features are extracted from pre-acquired audio data.
13. A model training device, characterized in that, The device includes: The first acquisition module is used to acquire audio data of the scene where the transmission line is located, and extract Mel spectrum features and Mel cepstral coefficient features from the audio data; The determination module is used to train a first neural network model based on the Mel spectrum features and determine the corresponding first model accuracy, and to train a second neural network model based on the Mel cepstral coefficient features and determine the corresponding second model accuracy. The training module is used to train a third neural network model based on the accuracy of the first model, the accuracy of the second model, the Mel spectrum features, and the Mel cepstral coefficient features. The third neural network model is used to predict whether there are abnormal hidden sounds on the transmission line.
14. An anomaly detection device, characterized in that, The device includes: The second acquisition module is used to acquire audio data based on the scene where the target transmission line is located; The processing module is used to process the audio data based on a third neural network model to obtain a prediction result; the prediction result is used to indicate whether there are any abnormal or potential hazards on the target transmission line. The third neural network model is obtained by training a first model accuracy, a second model accuracy, Mel spectrum features, and Mel cepstral coefficient features. The first model accuracy is determined by training the first neural network model based on the Mel spectrum features. The second model accuracy is determined by training the second neural network model based on the Mel cepstral coefficient features. The Mel spectrum features and the Mel cepstral coefficient features are extracted from pre-acquired audio data.
15. An electronic device, characterized in that, include: A processor, and a memory communicatively connected to the processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory to implement the method as described in any one of claims 1-12.
16. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-12.
17. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method as described in any one of claims 1-12.