Method and System for Abnormal Monitoring of Tower Bolts Based on Deep Learning and Voiceprint
By denoising the tower monitoring audio and extracting the bolt voiceprint feature, combined with the pre-trained anomaly classification model, the problem of insufficient voiceprint data processing in the existing technology is solved, and efficient and intelligent monitoring of the bolt status of the tower is achieved.
Patent Information
- Application Number
- CN202411604664.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-11-12
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2044-11-12
AI Technical Summary
The prior art is relatively traditional in processing and feature extraction of voiceprint data, and it is impossible to fully explore deep features and potential information, which affects the accuracy and reliability of the model.
By obtaining the original monitoring audio of the tower, using the pre-trained denoising model for processing, the bolt soundprint to be analyzed is extracted, and loaded into the pre-trained bolt soundprint exception classification model for classification processing, so as to achieve intelligent and efficient monitoring of the bolt status of the tower.
Accurate and intelligent monitoring of the status of tower bolts is achieved, and the monitoring efficiency and reliability are improved.
Smart Images

Figure CN119152859B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and more particularly, to a method and system for abnormal monitoring of tower bolts based on deep learning and voiceprint. Background Art
[0002] With the continuous advancement of infrastructure construction, towers, as key support structures, are widely used in fields such as communication and power. The stability and safety of towers are of utmost importance, and bolts, as key components connecting various parts of the tower, directly affect the overall stability of the tower. However, traditional bolt condition monitoring methods often rely on manual inspections, which are inefficient and vulnerable to human factors.
[0003] The prior art CN118503671A discloses a method for positioning and maintaining the loosening position of tower bolts based on deep learning and voiceprint, including: First, obtaining the current feedback voiceprint data of tower bolts through a preset sensor. Then, calling a pre-trained deep learning model to process the voiceprint data to obtain the loosening position of the current tower bolts. Next, determining the target tower bolts corresponding to the current loosening position and obtaining the feedback voiceprint data of the target bolts during the overall monitoring period, thereby calculating the structural damage risk coefficient of the tower structure where the target bolts are located. Finally, based on the structural damage risk coefficient, determining the maintenance strategy for the target tower bolts. However, in terms of the processing and feature extraction of voiceprint data in the above prior application patent solution, the processing steps of the prior art are relatively traditional and conventional, and may not be able to fully extract the deep features and potential information in the voiceprint data, affecting the accuracy and reliability of the model; for the feedback voiceprint data, a series of complex calculations and conversions are required to obtain the structural damage risk coefficient, and the model is complex and has poor generalization ability. Summary of the Invention
[0004] The purpose of the present invention is to provide a method and system for abnormal monitoring of tower bolts based on deep learning and voiceprint.
[0005] In a first aspect, an embodiment of the present invention provides a method for abnormal monitoring of tower bolts based on deep learning and voiceprint, including:
[0006] Obtaining the original tower monitoring audio;
[0007] Calling a pre-trained tower monitoring audio denoising model to process the original tower monitoring audio to obtain the target tower monitoring audio;
[0008] Obtaining the voiceprint of the bolt to be analyzed in the target tower monitoring audio;
[0009] Load the soundprint of the bolt to be analyzed into a pre-trained bolt soundprint anomaly classification model for soundprint classification processing to obtain the target soundprint anomaly classification result of the target iron tower monitoring audio.
[0010] In a second aspect, an embodiment of the present invention provides a server system, including a server, and the server is configured to execute the method described in the first aspect.
[0011] Compared with the prior art, the beneficial effects provided by the present invention include: adopting a method and system for abnormal monitoring of iron tower bolts based on deep learning and soundprint disclosed by the present invention, by acquiring the original monitoring audio of the iron tower and using a pre-trained denoising model to process the audio to obtain a clearer target audio. Subsequently, extract the soundprint of the bolt to be analyzed from the target audio and load it into a pre-trained bolt soundprint anomaly classification model for classification processing. Finally, this method can accurately obtain the bolt soundprint anomaly classification result in the iron tower monitoring audio, realizing intelligent and efficient monitoring of the state of iron tower bolts. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings required to be used in the embodiments will be briefly introduced below. It should be understood that the following drawings only show some embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can be obtained based on these drawings without creative efforts.
[0013] Figure 1 It is a schematic flowchart of the steps of the method for abnormal monitoring of iron tower bolts based on deep learning and soundprint provided by the embodiment of the present invention; DETAILED DESCRIPTION OF THE EMBODIMENTS
[0014] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the drawings in the embodiments of the present invention. Obviously, the described embodiments are some, but not all, of the embodiments of the present invention. Usually, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0015] The following will describe in detail the specific embodiments of the present invention with reference to the drawings.
[0016] To solve the technical problems in the foregoing background art, Figure 1 It is a flowchart of the method for abnormal monitoring of iron tower bolts based on deep learning and soundprint provided by the embodiment of the present disclosure. The method for abnormal monitoring of iron tower bolts based on deep learning and soundprint will be introduced in detail below.
[0017] Step S201: Obtain the original audio for tower monitoring;
[0018] Step S202: Call the pre-trained audio denoising model for tower monitoring to process the original audio for tower monitoring, and obtain the target audio for tower monitoring;
[0019] Step S203: Obtain the bolt sound pattern to be analyzed of the target audio for tower monitoring;
[0020] Step S204: Load the bolt sound pattern to be analyzed into the pre-trained bolt sound pattern anomaly classification model for sound pattern classification processing, and obtain the target sound pattern anomaly classification result of the target audio for tower monitoring.
[0021] In an embodiment of the present invention, exemplarily, the server receives the original monitored audio data from the audio acquisition device installed on the iron tower. For example, in a certain monitoring task, the server automatically downloads and saves the audio file of the iron tower environment of the day, named "20230425_tower_audio.wav", from the audio acquisition device at a fixed time every day, such as 3 am. After obtaining the original audio file, the server automatically calls a pre-trained audio denoising model. This model has been trained with a large amount of iron tower audio data containing environmental noise and can effectively identify and remove wind noise, current noise, and other background noises. For example, the server performs denoising processing on the "20230425_tower_audio.wav" file and outputs the denoised audio file "20230425_tower_audio_denoised.wav". Next, the server runs a voiceprint extraction algorithm to extract the unique voiceprint features of the bolts from the denoised audio. These features may include the sound waveforms with specific frequencies, amplitudes, and durations generated when the bolts are loose. For example, from the "20230425_tower_audio_denoised.wav" file, the server successfully extracts a piece of voiceprint feature data suspected of bolt loosening and saves it as "20230425_bolt_soundprint.dat". Finally, the server loads the extracted bolt voiceprint data into a trained bolt voiceprint anomaly classification model. This model can identify the normal and abnormal voiceprints of bolts and classify the abnormal voiceprints in more detail, such as bolt loosening, fracture, etc. In an embodiment of the present invention, the bolt voiceprint anomaly classification model analyzes the "20230425_bolt_soundprint.dat" file and outputs a classification result: "bolt loosening". Corresponding response operations are preset in the server for different classification results. According to the result of "bolt loosening", the server can automatically issue an alarm to notify the maintenance personnel to check and tighten the bolts on the iron tower in a timely manner. In an embodiment of the present invention, the bolt voiceprint anomaly classification model is obtained in the following manner.
[0022] Obtain the current bolt voiceprint instance of the target iron tower monitoring audio instance; the current bolt voiceprint instance is configured with a current voiceprint anomaly target value;
[0023] Construct an integrated voiceprint classification model and a target voiceprint classification model according to the basic voiceprint classification model; the basic voiceprint classification model is obtained by loading the basic bolt voiceprint instance of the target iron tower monitoring audio instance into a preset classification model for bolt voiceprint anomaly classification training; the basic bolt voiceprint instance is configured with an initial target iron tower monitoring audio instance category label; the initial target iron tower monitoring audio instance category label is different from the current voiceprint anomaly target value;
[0024] Load the current bolt voiceprint instance into the integrated voiceprint classification model for voiceprint classification processing to obtain a first sample voiceprint type recognition result;
[0025] Integrate the first sample voiceprint type recognition result and the current voiceprint anomaly target value to obtain a sample voiceprint anomaly target value;
[0026] Load the current bolt voiceprint instance into the target voiceprint classification model for voiceprint classification processing to obtain a second sample voiceprint type recognition result;
[0027] Execute a training process on the target voiceprint classification model according to the deviation between the second sample voiceprint type recognition result and the sample voiceprint anomaly target value. The trained target voiceprint classification model is used to determine a bolt voiceprint anomaly classification model, and the bolt voiceprint anomaly classification model is used to determine the type of the current voiceprint anomaly target value for a bolt voiceprint instance.
[0028] In an embodiment of the present invention, exemplarily, the server selects a specific audio instance from the processed iron tower monitoring audio database, such as "20230425_tower_audio_sample.wav", and extracts the current bolt soundprint instance from it, marked as "bolt_soundprint_sample". This bolt soundprint instance is configured with a current soundprint anomaly target value, such as "bolt loose". First, the server uses a preset classification model to train a large number of basic bolt soundprint instances, which are all configured with initial target iron tower monitoring audio instance category labels, such as "normal", "slightly loose", etc., and these labels are different from the current soundprint anomaly target value "bolt loose". Through this process, the server obtains a basic soundprint classification model. Then, based on this basic model, the server constructs an integrated soundprint classification model and a target soundprint classification model. The server loads "bolt_soundprint_sample" into the integrated soundprint classification model for processing and obtains the first sample soundprint type recognition result. For example, the model determines that this soundprint is "possibly loose". The server integrates the recognition result of "possibly loose" with the current soundprint anomaly target value "bolt loose" to obtain a sample soundprint anomaly target value, such as "highly suspected loose". Next, the server loads "bolt_soundprint_sample" into the target soundprint classification model for processing and obtains the second sample soundprint type recognition result. For example, the model determines that this soundprint is "slightly loose". The server compares the deviation between "slightly loose" and "highly suspected loose" and accordingly adjusts the training of the target soundprint classification model to improve its recognition accuracy of the degree of bolt looseness. After multiple iterative trainings and optimizations, the server obtains a more accurate bolt soundprint anomaly classification model. This model can more accurately determine the type for the current soundprint anomaly target value (such as "bolt loose") of the bolt soundprint instance.
[0029] In an embodiment of the present invention, after obtaining the current bolt soundprint instance of the target iron tower monitoring audio instance, the embodiment of the present invention further provides the following implementation manners.
[0030] Perform data cleaning processing on the current bolt soundprint instance to obtain a transitional bolt soundprint instance, and the data cleaning processing is used to change the composition of the current bolt soundprint instance;
[0031] Load the transitional bolt soundprint instance into the integrated soundprint classification model and the target soundprint classification model respectively to obtain a first error parameter;
[0032] The training process for the target soundprint classification model according to the deviation between the second sample soundprint type recognition result and the sample soundprint anomaly target value includes:
[0033] Determine a second error parameter according to the deviation between the second sample voiceprint type recognition result and the sample voiceprint anomaly target value;
[0034] Determine a target error parameter according to the first error parameter and the second error parameter;
[0035] Adjust the structural parameters of the target voiceprint classification model according to the target error parameter.
[0036] In an embodiment of the present invention, exemplarily, after obtaining the current bolt voiceprint instance "bolt_soundprint_sample", the server first performs data cleaning processing on it. For example, the server will remove background noise, standardize the volume, trim the silent part, etc., to change the composition of the voiceprint instance, and obtain a cleaned transitional bolt voiceprint instance, named "cleaned_bolt_soundprint_sample". Next, the server loads "cleaned_bolt_soundprint_sample" into the integrated voiceprint classification model and the target voiceprint classification model respectively for processing. Both models will output the classification results of the voiceprint instance. The server compares the classification results of the integrated voiceprint classification model and the target voiceprint classification model for "cleaned_bolt_soundprint_sample", calculates the difference between the two, and obtains the first error parameter. For example, if the integrated model identifies as "possibly loose" while the target model identifies as "slightly loose", the first error parameter reflects the difference degree between these two classifications. The server determines the second error parameter according to the deviation between the second sample voiceprint type recognition result (such as "slightly loose") and the sample voiceprint anomaly target value (such as "highly suspected loose"). This parameter reflects the accuracy of the target voiceprint classification model when processing a specific voiceprint instance. The server combines the first error parameter and the second error parameter to determine a comprehensive target error parameter. This target error parameter more comprehensively reflects the performance of the target voiceprint classification model when processing voiceprint instances similar to "cleaned_bolt_soundprint_sample". Based on the target error parameter, the server adjusts the structural parameters of the target voiceprint classification model. For example, the server will adjust the weights, biases or other hyperparameters of the model to improve the classification accuracy of the model for future bolt voiceprint instances. Through such an iterative optimization process, the server can continuously improve the performance of the bolt voiceprint anomaly classification model.
[0037] In an embodiment of the present invention, the step of loading the transitional bolt voiceprint instance into the integrated voiceprint classification model and the target voiceprint classification model respectively to obtain the first error parameter can be implemented through the following examples.
[0038] Load the transitional bolt soundprint instance into the integrated soundprint classification model for soundprint classification processing to obtain a third sample soundprint type recognition result;
[0039] Load the transitional bolt soundprint instance into the target soundprint classification model for soundprint classification processing to obtain a fourth sample soundprint type recognition result;
[0040] Determine the first error parameter according to the deviation between the fourth sample soundprint type recognition result and the third sample soundprint type recognition result.
[0041] In an embodiment of the present invention, by way of example, the server loads the "cleaned_bolt_soundprint_sample" transitional bolt soundprint instance after data cleaning processing into the integrated soundprint classification model. This model is composed of multiple basic classifiers and can comprehensively analyze the input soundprint features and give a classification result. For example, after analyzing the "cleaned_bolt_soundprint_sample", the third sample soundprint type recognition result given by the integrated soundprint classification model is "suspected medium looseness". Subsequently, the server loads the same "cleaned_bolt_soundprint_sample" transitional bolt soundprint instance into the target soundprint classification model. This model is a classifier dedicated to accurately identifying specific bolt soundprint anomalies. For example, after analyzing the "cleaned_bolt_soundprint_sample", the fourth sample soundprint type recognition result given by the target soundprint classification model is "medium looseness". The server will then compare the recognition results given by the integrated soundprint classification model and the target soundprint classification model, namely "suspected medium looseness" and "medium looseness", and calculate the deviation between the two. This deviation reflects the classification consistency of the two models when processing the same input. Based on this deviation, the server determines the first error parameter. For example, if the two recognition results are exactly the same, the first error parameter may be zero; if there are differences, the first error parameter will be a positive value used to quantify the degree of this difference. This error parameter will be used in the subsequent model optimization and adjustment process to improve the accuracy and reliability of bolt soundprint anomaly classification.
[0042] In an embodiment of the present invention, the following implementation manners are further provided.
[0043] In the case where the training period is the first preset period, adjust the integrated soundprint classification model according to the structural parameters of the target soundprint classification model to obtain the bolt soundprint anomaly classification model, where the training period is the period of training the target soundprint classification model using the current bolt soundprint instance.
[0044] In an embodiment of the present invention, exemplarily, when the server trains the target voiceprint classification model using the current bolt voiceprint instance "bolt_soundprint_sample", a training cycle is set. This cycle is called the first preset cycle, which determines the duration and intensity of model training. For example, the server sets the first preset cycle to 100 training iterations. In each iteration, the server uses "bolt_soundprint_sample" and its associated voiceprint anomaly target values to train the target voiceprint classification model, continuously adjusting the model's parameters to improve its ability to identify bolt voiceprint anomalies. When the training reaches the first preset cycle, that is, after 100 iterations are completed, the server adjusts the integrated voiceprint classification model according to the current structural parameters of the target voiceprint classification model. Specifically, the server analyzes the features and classification logic learned by the target voiceprint classification model during the training process and integrates this information into the integrated voiceprint classification model. For example, if the target voiceprint classification model discovers during training that certain specific frequency components in "bolt_soundprint_sample" are crucial for identifying bolt loosening, the server will add these features to the recognition logic of the integrated voiceprint classification model. Through such adjustments, the integrated voiceprint classification model can absorb the learning results of the target voiceprint classification model within a specific training cycle, thereby enhancing its recognition performance and generalization ability for bolt voiceprint anomalies. Finally, the adjusted integrated voiceprint classification model will be used as the bolt voiceprint anomaly classification model to more accurately identify and process bolt voiceprint anomalies in the tower monitoring audio.
[0045] In an embodiment of the present invention, in the case where the training cycle is the first preset cycle, adjusting the integrated voiceprint classification model according to the structural parameters of the target voiceprint classification model to obtain the bolt voiceprint anomaly classification model can be implemented through the following examples.
[0046] Adjust the integrated voiceprint classification model according to the structural parameters of the target voiceprint classification model to obtain a candidate integrated voiceprint classification model;
[0047] During the training process when the training cycle is the first preset cycle, optimize the candidate integrated voiceprint classification model according to the structural parameters of the target voiceprint classification model to obtain the bolt voiceprint anomaly classification model.
[0048] In an embodiment of the present invention, exemplarily, the server first analyzes the structural parameters of the target voiceprint classification model after being trained for a first preset period. These structural parameters may include the weights, biases, connection methods between layers, and other hyperparameters of the model. For example, the target voiceprint classification model may have learned some feature extraction methods and classification boundaries that are particularly effective for identifying the voiceprints of bolt loosening. Based on these analyses, the server adjusts the integrated voiceprint classification model. The adjustment methods may include adding or deleting certain feature extraction layers, modifying the connection weights, adjusting the classification threshold, etc. These adjustments are aimed at enabling the integrated voiceprint classification model to absorb the advantages of the target voiceprint classification model. After these adjustments, the server obtains a candidate integrated voiceprint classification model. In the subsequent training process, the server continues to optimize the candidate integrated voiceprint classification model using bolt voiceprint instances such as "bolt_soundprint_sample". This optimization process may include multiple aspects, such as further adjusting the weights and biases of the model, optimizing the layer structure of the model, improving the activation function, etc. For example, the server may find that the candidate integrated voiceprint classification model makes misclassifications when identifying certain specific types of bolt loosening voiceprints. To address this issue, the server can fine-tune the relevant layers in the model to improve the recognition accuracy of the model for such voiceprints. After multiple iterative optimizations within the first preset period, the server finally obtains a bolt voiceprint anomaly classification model with more excellent performance. This model not only inherits the advantages of the target voiceprint classification model but also improves the generalization ability and stability through the method of ensemble learning. In this way, the server can use this model to more accurately identify bolt voiceprint anomalies in the tower monitoring audio.
[0049] In an embodiment of the present invention, during the training process when the training period is the first preset period, the candidate integrated voiceprint classification model is optimized according to the structural parameters of the target voiceprint classification model to obtain the bolt voiceprint anomaly classification model, which can be implemented through the following examples.
[0050] When the training period is the first preset period, during the training period of a second preset period, the candidate integrated voiceprint classification model is adjusted according to the structural parameters of the target voiceprint classification model;
[0051] When the preset training termination condition is reached, the candidate integrated voiceprint classification model is determined as the bolt voiceprint anomaly classification model.
[0052] In an embodiment of the present invention, exemplarily, during the process of model training by the server, when the training reaches the first preset period, a new training stage will be entered, which is called the second preset period. In this stage, the server will continuously monitor the structural parameters and performance of the target voiceprint classification model, and further adjust the candidate integrated voiceprint classification model based on this. For example, after the end of the first preset period, the target voiceprint classification model shows a high recognition rate for the voiceprints of bolt looseness at certain specific frequencies. After analyzing the structural parameters of this target voiceprint classification model, the server finds that it uses a special feature extraction method. Therefore, during the training process of the second preset period, the server will apply this method to the candidate integrated voiceprint classification model to improve its recognition ability for the voiceprints of bolt looseness at specific frequencies. At the same time, the server will also perform corresponding optimization and adjustment on the candidate integrated voiceprint classification model according to other excellent characteristics of the target voiceprint classification model, such as anti-noise performance and recognition speed. During the training process of the second preset period, the server will continuously monitor the performance of the candidate integrated voiceprint classification model and compare it with the preset training termination conditions. These training termination conditions can include the recognition accuracy of the model, the training duration, the convergence of the loss function, etc. For example, the server sets that when the recognition accuracy of the candidate integrated voiceprint classification model no longer significantly improves in consecutive multiple training iterations, or the training duration reaches a certain preset upper limit, it is considered that the training termination condition is reached. At this time, the server will determine the current candidate integrated voiceprint classification model as the final bolt voiceprint anomaly classification model. This final bolt voiceprint anomaly classification model not only inherits the advantages of the target voiceprint classification model, but also improves the recognition performance and stability through ensemble learning and multi-period training optimization. The server can use this model to more accurately identify and classify the bolt voiceprint anomalies in the tower monitoring audio.
[0053] In an embodiment of the present invention, in the case of the training cycle of the second preset period, according to the structural parameters of the target voiceprint classification model, the adjustment of the candidate integrated voiceprint classification model can be executed through the following examples.
[0054] In the case of the training cycle of the second preset period, obtain the current structural parameters of the target voiceprint classification model;
[0055] Process the current structural parameters according to the state space model to obtain the target structural parameters;
[0056] Adjust the candidate integrated voiceprint classification model according to the target structural parameters.
[0057] In an embodiment of the present invention, exemplarily, during the training process of the second preset period, the server will first obtain the current structural parameters of the target voiceprint classification model. These structural parameters may include the weight matrix, bias vector, convolution kernel parameters, etc. of the model, which jointly determine the recognition performance and feature extraction ability of the model. For example, the server can extract these structural parameters by accessing the internal state of the target voiceprint classification model or using specialized tools. For example, the server obtains a set of weight matrix and bias vector, which represent the feature transformation and decision boundary of the model at different levels. Next, the server will use a state space model to process the currently obtained structural parameters. The state space model can help the server understand and predict the variation law of the model parameters over time, so as to extract more stable and representative target structural parameters. For example, the server can construct a Kalman filter or hidden Markov model as the state space model, and use these models to perform filtering, smoothing or prediction operations on the current structural parameters. Through the processing of the state space model, the server can remove noise and outliers, and obtain more robust and reliable target structural parameters. Finally, the server will adjust the candidate integrated voiceprint classification model according to the obtained target structural parameters. The purpose of the adjustment is to enable the candidate integrated voiceprint classification model to better learn and simulate the excellent characteristics of the target voiceprint classification model, so as to improve its own recognition performance and generalization ability. For example, the server can apply the weight matrix and bias vector in the target structural parameters to the corresponding levels of the candidate integrated voiceprint classification model, or redesign the network structure and connection mode of the candidate integrated voiceprint classification model according to the target structural parameters. Through these adjustments, the candidate integrated voiceprint classification model can more effectively extract the features of the bolt voiceprint and make accurate classification judgments. In summary, by obtaining the current structural parameters of the target voiceprint classification model, processing them using the state space model, and adjusting the candidate integrated voiceprint classification model according to the target structural parameters and other steps, the server can continuously improve the performance of the bolt voiceprint anomaly classification model to more accurately identify and process the bolt voiceprint anomalies in the tower monitoring audio.
[0058] In an embodiment of the present invention, the process of processing the current structural parameters according to the state space model to obtain the target structural parameters can be implemented through the following examples.
[0059] Determine the first influence factor of the integrated voiceprint classification model and the second influence factor of the target voiceprint classification model according to the training period corresponding to the current structural parameters; the first influence factor has a negative feedback relationship with the training period, and the second influence factor has a positive feedback relationship with the training period;
[0060] Process the current structural parameters according to the first influence factor, the second influence factor and the state space model to obtain the target structural parameters.
[0061] In an embodiment of the present invention, exemplarily, during the process of model training by the server, two influencing factors are determined for each training cycle: a first influencing factor for the integrated voiceprint classification model and a second influencing factor for the target voiceprint classification model. These two influencing factors are used to adjust the weights of model parameter updates to ensure the stability and convergence of the training process. For example, when the current training cycle is the 10th cycle of the second preset cycle, the server determines the first influencing factor and the second influencing factor according to a preset functional relationship. Since there is a negative feedback relationship between the first influencing factor and the training cycle, as the training cycle increases, the value of the first influencing factor gradually decreases. On the contrary, there is a positive feedback relationship between the second influencing factor and the training cycle, and as the training cycle increases, its value gradually increases. Specifically, if the initial first influencing factor is set to 0.8 and the second influencing factor is set to 0.2, as the training cycle progresses, the first influencing factor will gradually decrease to 0.6, 0.4, etc., while the second influencing factor will gradually increase to 0.4, 0.6, etc. After determining the first influencing factor and the second influencing factor, the server processes the current structural parameters using these two influencing factors and the state space model. The state space model can help the server predict and estimate the optimal value of the model parameters based on historical information and current information. For example, the server uses a Kalman filter as the state space model, which can update the state of the model by combining the currently observed structural parameters and the previous estimated values. During the update process, the first influencing factor acts on the original parameters of the integrated voiceprint classification model, gradually reducing its influence; while the second influencing factor enhances the influence of the current structural parameters of the target voiceprint classification model. In this way, the server can gradually incorporate the excellent characteristics of the target voiceprint classification model into the integrated voiceprint classification model while maintaining the stability and generalization ability of the integrated model. Finally, the obtained target structural parameters after processing will be a parameter set that combines the advantages of multiple models, which is used to guide the subsequent training and tuning processes. In summary, by determining the first influencing factor and the second influencing factor and processing the current structural parameters using these influencing factors and the state space model, the server can more effectively perform model fusion and parameter optimization, thereby improving the performance of the bolt voiceprint anomaly classification model.
[0062] In an embodiment of the present invention, the training method of the basic voiceprint classification model can be implemented through the following example.
[0063] Obtain the basic bolt voiceprint instance of the target tower monitoring audio instance;
[0064] Load the basic bolt voiceprint instance into the preset classification model for voiceprint classification processing to obtain an initial sample voiceprint type recognition result;
[0065] Determine a sample error parameter according to the deviation between the initial sample voiceprint type recognition result and the class label of the initial target iron tower monitoring audio instance;
[0066] Adjust the structural parameters of the preset classification model according to the sample error parameter, and determine the preset classification model that has completed training as the basic voiceprint classification model.
[0067] In an embodiment of the present invention, by way of example, the server first extracts basic bolt voiceprint instances from the target iron tower monitoring audio instances. These instances may include voiceprint features in the case of bolt loosening, fracture, or other abnormalities. The server can intercept these representative voiceprint segments through professional audio processing software or custom algorithms. For example, the server can accurately intercept the sound wave segments with specific frequencies and amplitudes generated during bolt loosening from a long-time iron tower monitoring audio. This segment is the basic bolt voiceprint instance. Next, the server loads the extracted basic bolt voiceprint instances into a preset classification model for voiceprint classification processing. This preset classification model can be a deep learning network, such as a convolutional neural network (CNN) or a recurrent neural network (RNN), which has been preliminarily trained for audio classification tasks. For example, the server inputs the intercepted bolt loosening voiceprint instance into the preset CNN model, and the model will attempt to identify which type (such as loosening, fracture, etc.) this voiceprint instance belongs to and output an initial sample voiceprint type recognition result. After obtaining the initial sample voiceprint type recognition result, the server will compare it with the class label of the initial target iron tower monitoring audio instance to calculate the deviation between the two. This deviation is called the sample error parameter, which reflects the degree of non-conformity between the model prediction result and the actual label. For example, if the model incorrectly identifies the voiceprint instance of bolt loosening as bolt fracture, the sample error parameter will be relatively large. According to the calculated sample error parameter, the server will adjust the structural parameters of the preset classification model to reduce the error and improve the recognition accuracy of the model. The adjustment methods may include modifying network weights, increasing or decreasing the number of network layers, changing activation functions, etc. For example, the server will update the weight matrix of the CNN model through the backpropagation algorithm so that the model can more accurately identify when encountering a similar bolt loosening voiceprint instance next time. After multiple iterations of training and adjustment, when the performance of the preset classification model reaches a certain preset standard (such as the recognition accuracy exceeds 90%), the server will determine it as the final basic voiceprint classification model. This model will be used for subsequent bolt voiceprint anomaly classification tasks. In summary, by steps such as obtaining basic bolt voiceprint instances, loading them into a preset classification model for processing, determining sample error parameters, and adjusting model structural parameters, the server can train a high-performance basic voiceprint classification model to provide strong support for subsequent iron tower bolt voiceprint anomaly classification.
[0068] In addition, in the "category labels of the initial target tower detection audio instances", defining or annotating the category labels of voiceprint instances in different situations is a crucial preprocessing step, which involves corresponding the audio instances to specific physical states or abnormal conditions. First, it is necessary to clearly define various possible bolt states or abnormal conditions, such as "loose", "broken", "normal", etc., and assign a unique category label to each state. These labels should be clear, consistent, and easy to understand for use in subsequent model training and classification tasks. For each tower monitoring audio instance, it is necessary for professional technicians or trained annotators to listen carefully to determine the specific state of the bolts in the audio. The annotator will label each audio instance as the corresponding bolt state, such as "loose" or "broken", according to the predefined category labels. To ensure the accuracy of the annotation, multiple annotation and cross-validation strategies can be implemented. That is, each audio instance can be independently annotated by multiple annotators, and then their annotation results are compared. If the annotation results are inconsistent, a higher-level technician can be arbitrated to determine the final category label. After the annotation is completed, the audio instances and their corresponding category labels need to be stored in the database for subsequent model training and classification tasks. When storing the data, it should be ensured that the association relationship between each audio instance and its category label is clear and definite. Over time and with the accumulation of experience, it may be necessary to update or adjust the definition of the category labels. For example, if a new bolt abnormal state is found, a new category label can be assigned to it. At the same time, it is also necessary to regularly review and verify the existing annotated data to ensure its accuracy and consistency. Through the above process, the accuracy and consistency of the "category labels of the initial target tower detection audio instances" can be ensured, providing reliable data support for the subsequent training of the basic voiceprint classification model.
[0069] In the embodiment of the present invention, the tower monitoring audio denoising model is obtained in the following manner. Obtain a noise-free tower monitoring audio instance, a noisy tower monitoring audio instance matching the noise-free tower monitoring audio instance, and a noise label of the noisy tower monitoring audio instance; the noise label indicates the noise data in the noisy tower monitoring audio instance;
[0070] Use a first noise cancellation model to perform a noise filtering operation on the noisy tower monitoring audio instance to determine a first tower monitoring filtered audio, and estimate the noise data of the noisy tower monitoring audio instance to determine the estimated noise data information of the noisy tower monitoring audio instance;
[0071] Obtain the bolt audio data in the noisy tower monitoring audio instance according to the predicted noise data information. From the first tower monitoring filtered audio, collect the audio data with the same time sequence node as the bolt audio data in the noisy tower monitoring audio instance, and use the bolt audio data in the noisy tower monitoring audio instance to replace the audio data in the first tower monitoring filtered audio to determine a tower monitoring filtered audio instance;
[0072] Obtain the noise filtering cost parameter of the first noise cancellation model according to the tower monitoring filtered audio instance and the noise-free tower monitoring audio instance;
[0073] Obtain the noise data prediction cost parameter of the first noise cancellation model according to the predicted noise data information and the noise label;
[0074] Perform a tuning operation on the structural parameters of the first noise cancellation model according to the noise filtering cost parameter and the noise data prediction cost parameter to determine a tower monitoring audio denoising model for performing noise filtering operations on the original tower monitoring audio.
[0075] In an embodiment of the present invention, exemplarily, the server first obtains noise-free tower monitoring audio instances, noisy tower monitoring audio instances that match these noise-free audio instances, and noise labels corresponding to these noisy audio instances from the tower monitoring system. The noise labels clearly indicate which parts of the audio are noise data. For example, the server can record the noise-free bolt tightening sound in a stable tower monitoring environment as a noise-free audio instance, and at the same time record the sound of the same operation in another tower monitoring environment affected by environmental interference (such as wind noise, traffic noise, etc.) as a noisy audio instance. The noise labels are marked by experts and clearly indicate which time periods are noise in the noisy audio. Next, the server will use an initial noise cancellation model (referred to as the first noise cancellation model) to perform noise filtering operations on the noisy tower monitoring audio instances, generating the first tower monitoring filtered audio. At the same time, this model will also estimate the noise data in the noisy audio, generating estimated noise data information. For example, the first noise cancellation model can be a deep learning-based noise reduction model that can identify and remove the noise part in the audio, and at the same time estimate the specific information of the removed noise, such as the type and intensity of the noise. The server extracts the bolt audio data in the noisy audio (i.e., the effective audio data after removing the noise) according to the estimated noise data information, and collects the audio data at the same time sequence nodes as these bolt audio data from the first tower monitoring filtered audio. Then, the server will replace the corresponding audio data in the first tower monitoring filtered audio with the bolt audio data in the noisy audio, thereby generating a tower monitoring filtered audio instance. For example, if the estimated noise data information shows that there is strong traffic noise in a certain time period, the server will extract the bolt tightening sound of the non-noise part in this time period from the noisy audio, and use this part of the sound to replace the audio data in the same time period in the first tower monitoring filtered audio. Next, the server will calculate the noise filtering cost parameter of the first noise cancellation model according to the tower monitoring filtered audio instance and the noise-free tower monitoring audio instance. This cost parameter reflects the performance of the model in noise filtering. At the same time, according to the estimated noise data information and the actual noise label, the server will also calculate the noise data estimation cost parameter, which reflects the accuracy of the model in noise estimation. For example, the noise filtering cost parameter can be calculated by comparing the similarity between the filtered audio instance and the noise-free audio instance, while the noise data estimation cost parameter can be calculated by comparing the consistency between the estimated noise data information and the actual noise label. Finally, the server will optimize the structural parameters of the first noise cancellation model according to the calculated noise filtering cost parameter and noise data estimation cost parameter to improve the performance of the model in noise filtering and noise estimation. The optimized model will be determined as the final tower monitoring audio denoising model for performing noise filtering operations on the original tower monitoring audio.For example, if the noise filtering cost parameter indicates that the model performs poorly in filtering certain types of noise, the server will adjust the weights or network structure of the model to improve the performance in this regard. Similarly, if the noise data prediction cost parameter indicates that there is a deviation in the model's prediction of noise data, the server will also adjust the model accordingly to improve the prediction accuracy.
[0076] In an embodiment of the present invention, the first noise cancellation model includes a feature enhancement module, a noise filtering module, and a noise data prediction module;
[0077] The use of the first noise cancellation model to perform a noise filtering operation on the noisy tower monitoring audio instance to determine the first tower monitoring filtered audio, and to perform noise data prediction on the noisy tower monitoring audio instance to determine the predicted noise data information of the noisy tower monitoring audio instance can be implemented through the following examples.
[0078] Using the feature enhancement module, perform feature enhancement on the noisy tower monitoring audio instance to determine the feature-enhanced audio matching the noisy tower monitoring audio instance;
[0079] Using the noise filtering module, perform a noise filtering operation on the feature-enhanced audio matching the noisy tower monitoring audio instance to determine the first tower monitoring filtered audio;
[0080] Using the noise data prediction module, perform noise data prediction on the feature-enhanced audio matching the noisy tower monitoring audio instance to determine the predicted noise data information of the noisy tower monitoring audio instance.
[0081] In an embodiment of the present invention, exemplarily, the server first uses a feature enhancement module to enhance the features of the noisy tower monitoring audio instance. The purpose of feature enhancement is to highlight the key information in the audio, especially those parts that may be masked or weakened by noise, so as to help subsequent noise filtering and prediction be more accurate. For example, the server receives a noisy tower monitoring audio instance containing wind noise and the sound of a loose bolt. The feature enhancement module will use a series of algorithms (such as spectral analysis, wavelet transform, etc.) to strengthen the features of the loose bolt sound, while suppressing or weakening the features of noise such as wind noise. In this way, in the audio after feature enhancement (i.e., the feature-enhanced audio), the features of the loose bolt sound will be more obvious. Next, the server uses a noise filtering module to perform noise filtering operations on the feature-enhanced audio. The purpose of the noise filtering module is to remove or reduce the interference of noise from the audio, so as to obtain a clearer target sound. The noise filtering module will process the feature-enhanced audio and use advanced signal processing techniques (such as adaptive filtering, spectral subtraction, etc.) to identify and remove noise components such as wind noise. After the noise filtering operation, the server will obtain the first tower monitoring filtered audio, in which the loose bolt sound is clearer and distinguishable, while noise such as wind noise is greatly suppressed or eliminated. Finally, the server uses a noise data prediction module to predict the noise data of the feature-enhanced audio. The purpose of noise data prediction is to analyze and estimate information such as the type, intensity, and distribution of noise in the audio, which is very important for evaluating the effect of noise filtering and subsequent optimization processing. In the above example, the noise data prediction module will deeply analyze the feature-enhanced audio and use machine learning algorithms or statistical models to predict the noise components in it. For example, it can identify features such as the frequency range and intensity change of wind noise and generate predicted noise data information. These information can not only help the server evaluate the performance of the noise filtering module, but also provide valuable references for subsequent model tuning and parameter adjustment.
[0082] In an embodiment of the present invention, the feature enhancement module includes a robust scaling sub-module and a feature enhancement sub-module; using the feature enhancement module to enhance the features of the noisy tower monitoring audio instance and determining the feature-enhanced audio matching the noisy tower monitoring audio instance can be implemented through the following examples.
[0083] Using the robust scaling sub-module, according to the audio attributes matching the noisy tower monitoring audio instance, perform robust scaling processing on the noisy tower monitoring audio instance to determine the robustly scaled noisy tower monitoring audio instance;
[0084] Using the feature enhancement sub-module, perform feature enhancement on the robustly scaled noisy tower monitoring audio instance to determine the feature-enhanced audio matching the noisy tower monitoring audio instance.
[0085] In an embodiment of the present invention, exemplarily, the server first performs robust scaling processing on the noisy tower monitoring audio instance using the robust scaling sub-module. The purpose of this step is to adjust the dynamic range of the audio to make it more suitable for subsequent feature enhancement processing, and at the same time reduce the impact of outliers or noise on the overall data. For example, the server receives a noisy tower monitoring audio instance with large volume fluctuations, where the volume of some parts is too high while that of other parts is relatively low. The robust scaling sub-module will analyze the audio attributes (such as volume, frequency distribution, etc.) of this audio instance, and then adopt a robust scaling method (such as median absolute deviation scaling, robust standardization, etc.) to adjust the amplitude of the audio, so as to reduce the impact of volume fluctuations on subsequent processing while maintaining the original audio features. After the robust scaling processing, the volume and dynamic range of the audio instance will be more balanced and consistent. Next, the server uses the feature enhancement sub-module to perform feature enhancement on the robustly scaled noisy tower monitoring audio instance. The goal of feature enhancement is to extract and strengthen the key features in the audio for better subsequent noise filtering and prediction. The feature enhancement sub-module will further analyze and process the audio instance that has undergone robust scaling processing. It may adopt a series of signal processing techniques (such as spectral analysis, time-frequency analysis, etc.) to extract the key features in the audio, such as the sound features of bolt loosening, the frequency characteristics of wind noise, etc. Then, by enhancing these key features (such as increasing the amplitude of the bolt loosening sound, sharpening its spectral characteristics, etc.), these features will be more easily recognized and distinguished in subsequent noise filtering and prediction. After being processed by the feature enhancement sub-module, the server will obtain a feature-enhanced audio that matches the original noisy tower monitoring audio instance, in which the key features are significantly strengthened and highlighted.
[0086] In an embodiment of the present invention, the noise filtering module includes a filtering coding sub-module, a filtering feature mapping module, and a filtering inverse robust scaling sub-module;
[0087] Using the noise filtering module, perform a noise filtering operation on the feature-enhanced audio that matches the noisy tower monitoring audio instance, and determine the first tower monitoring filtered audio, which can be implemented through the following examples.
[0088] Using the filtering coding sub-module, perform feature coding on the feature-enhanced audio that matches the noisy tower monitoring audio instance, and determine the first feature-coded audio that matches the noisy tower monitoring audio instance;
[0089] Using the filtering feature mapping module, perform feature mapping processing on the first feature-coded audio, and determine the first feature-mapped audio that matches the noisy tower monitoring audio instance;
[0090] Using the filtering inverse robust scaling sub-module, perform inverse robust scaling processing on the first feature-mapped audio to determine the first iron tower monitoring filtered audio.
[0091] In an embodiment of the present invention, exemplarily, the server first uses the filtering encoding sub-module to perform feature encoding on the feature-enhanced audio matched with the noisy iron tower monitoring audio instance. The purpose of feature encoding is to convert audio data into a format that is easier to process and analyze while retaining the key features of the audio. For example, the server has obtained a feature-enhanced iron tower monitoring audio through the feature enhancement module, which contains noises such as bolt loosening sounds and wind sounds. The filtering encoding sub-module will use a suitable encoding method (such as Mel Frequency Cepstral Coefficients MFCC, Linear Predictive Coding LPC, etc.) to encode this audio. The encoded audio (i.e., the first feature-encoded audio) converts the waveform data of the audio into a series of feature vectors, which can more accurately describe the sound characteristics in the audio, such as pitch, timbre, etc., facilitating subsequent feature mapping and noise filtering. Next, the server uses the filtering feature mapping module to perform feature mapping processing on the first feature-encoded audio. The purpose of feature mapping is to map the encoded feature vectors to a new feature space so that in this new space, the target sound (such as bolt loosening sound) and noise (such as wind sound) can be more easily distinguished. The filtering feature mapping module will use an advanced mapping algorithm (such as autoencoders, principal component analysis in deep learning, etc.) to process the first feature-encoded audio. Through training and learning, this module can identify the different distribution laws of bolt loosening sounds and wind sounds and other noises in the feature space and map them to different regions. In this way, in the audio after feature mapping processing (i.e., the first feature-mapped audio), the features of bolt loosening sounds and wind sounds and other noises will be more clearly separated. Finally, the server uses the filtering inverse robust scaling sub-module to perform inverse robust scaling processing on the first feature-mapped audio. The purpose of inverse robust scaling is to restore the audio data after feature mapping processing to the original dynamic range while maintaining the noise filtering effect. In the above example, the filtering inverse robust scaling sub-module will use the inverse operation corresponding to the robust scaling sub-module to process the first feature-mapped audio. It will restore the feature-mapped audio data to a dynamic range close to the original audio according to the dynamic range and scaling factor of the original audio. In this way, in the audio after inverse robust scaling processing (i.e., the first iron tower monitoring filtered audio), the bolt loosening sound will be more clearly distinguishable, while the wind sound and other noises are effectively filtered out or their influence is reduced. At the same time, this filtered audio also retains the dynamic range and volume change characteristics of the original audio, making subsequent audio processing and analysis more accurate and reliable.
[0092] In an embodiment of the present invention, the filtering encoding sub-module includes a plurality of cascaded encoding networks;
[0093] Using the filtering and encoding sub-module, feature encoding is performed on the feature-enhanced audio matched with the noisy tower monitoring audio instance, and the first feature-encoded audio of the noisy tower monitoring audio instance can be determined through the following examples.
[0094] Input the output current intermediate feature audio of the current encoding network into the subsequent encoding network; the input of the initial encoding network is the feature-enhanced audio matched with the noisy tower monitoring audio instance;
[0095] Using the subsequent encoding network, perform feature encoding on the output current intermediate feature audio to determine the initial feature-encoded audio matched with the subsequent encoding network;
[0096] Fuse the initial feature-encoded audio matched with the subsequent encoding network and the output current intermediate feature audio to determine the output intermediate feature audio of the subsequent encoding network;
[0097] Take the output intermediate feature audio of the last encoding network as the first feature-encoded audio of the noisy tower monitoring audio instance.
[0098] In an embodiment of the present invention, exemplarily, the server first uses the feature-enhanced audio as the input of the initial encoding network. For example, the feature-enhanced audio is an audio file containing the sound of bolt loosening and the sound of wind. This file is first fed into the first encoding network. Each encoding network performs feature encoding on the input audio and outputs an intermediate feature audio. This intermediate feature audio serves as the input for the next encoding network. For example, the first encoding network extracts key features in the feature-enhanced audio, such as the frequency and amplitude of the bolt loosening sound, and then outputs an intermediate feature audio. This intermediate feature audio is subsequently input into the second encoding network. The second encoding network further extracts and encodes these features and then outputs a second intermediate feature audio. This process continues until the last encoding network. At each encoding network stage, the server fuses the initial feature-encoded audio matched by the subsequent encoding network and the currently output intermediate feature audio. This fusion can be achieved through weighted averaging, superposition, or other complex fusion algorithms, aiming to better retain and refine the key features in the audio. For example, at the second encoding network stage, the server fuses the output of the first encoding network (i.e., the first intermediate feature audio) and the initial feature-encoded audio of the second encoding network to obtain the second intermediate feature audio. This process is repeated at each encoding network stage. When the audio data passes through all the cascaded encoding networks, the server uses the output intermediate feature audio of the last encoding network as the first feature-encoded audio of the noisy tower monitoring audio instance. This first feature-encoded audio is a highly refined and encoded audio representation that contains the key feature information in the original audio but removes a large amount of redundant and noisy data. Through this processing process of the cascaded encoding network, the server can more accurately extract and encode the key features in the noisy tower monitoring audio instance, providing a more accurate data basis for subsequent feature mapping and noise filtering.
[0099] In an embodiment of the present invention, the noise data estimation module includes an estimation encoding sub-module, an estimation feature mapping module, and an estimation inverse robust scaling sub-module;
[0100] Using the noise data estimation module to estimate the noise data of the feature-enhanced audio matched by the noisy tower monitoring audio instance and determine the estimated noise data information of the noisy tower monitoring audio instance can be implemented through the following examples.
[0101] Using the estimation encoding sub-module to perform feature encoding on the feature-enhanced audio matched by the noisy tower monitoring audio instance to determine the second feature-encoded audio matched by the noisy tower monitoring audio instance;
[0102] Using the predicted feature mapping module, perform feature mapping processing on the second feature-encoded audio to determine the second feature-mapped audio that matches the noisy iron tower monitoring audio instance;
[0103] Using the predicted inverse robust scaling sub-module, perform inverse robust scaling processing on the second feature-mapped audio to determine the second feature-mapped audio after inverse robust scaling;
[0104] Based on the second feature-mapped audio after inverse robust scaling, obtain the predicted noise data information of the noisy iron tower monitoring audio instance.
[0105] In an embodiment of the present invention, by way of example, first, the server uses the predicted encoding sub-module to perform feature encoding on the feature-enhanced audio. For example, the feature-enhanced audio contains complex noises such as bolt loosening sounds and wind sounds. The predicted encoding sub-module will adopt a specific encoding algorithm (such as the convolutional neural network CNN in deep learning, etc.) to extract the key features in the audio. After encoding, the server obtains a second feature-encoded audio, which represents the feature information in the original audio in a more compact and efficient form. Next, the server uses the predicted feature mapping module to perform feature mapping processing on the second feature-encoded audio. The purpose of this module is to map the encoded features to a new space so that the noise and the target sound can be more easily distinguished. For example, through the training and learning of a deep learning model (such as an autoencoder), the predicted feature mapping module can identify the different distribution laws of noises such as bolt loosening sounds and wind sounds in the feature space and map them to different regions. After the mapping process, the server obtains a second feature-mapped audio, in which the features of the noise and the target sound are more clearly separated. Then, the server uses the predicted inverse robust scaling sub-module to perform inverse robust scaling processing on the second feature-mapped audio. The purpose of this step is to restore the audio data after feature mapping to the original dynamic range while maintaining the prediction accuracy of the noise data. Through the inverse robust scaling process, the server obtains a second feature-mapped audio after inverse robust scaling, which not only retains the original audio dynamic range but also highlights the features of the noise data. Finally, the server analyzes and determines the predicted noise data information based on the second feature-mapped audio after inverse robust scaling. For example, by calculating parameters such as the energy and frequency distribution of the noise components in the audio, the server can obtain a detailed prediction report on the noise. This report can include key information such as the type, intensity, and duration of the noise, providing an important basis for subsequent iron tower maintenance and noise control. Through this process, the server can accurately predict the noise data information in the noisy iron tower monitoring audio instance, providing strong support for subsequent noise filtering and iron tower status assessment.
[0106] In an embodiment of the present invention, the step of obtaining the estimated noise data information of the noisy tower monitoring audio instance based on the second feature-mapped audio after inverse robust scaling may be implemented through the following example.
[0107] Perform standardization processing on the second feature-mapped audio after inverse robust scaling to determine the estimated noise label of the noisy tower monitoring audio instance; the estimated noise label represents the noise data in the noisy tower monitoring audio instance;
[0108] Use the estimated noise label of the noisy tower monitoring audio instance as the estimated noise data information of the noisy tower monitoring audio instance.
[0109] In an embodiment of the present invention, exemplarily, the server first performs standardization processing on the second feature-mapped audio after inverse robust scaling. Standardization processing is an important data preprocessing step, which can eliminate the influence of data dimensions and make different features comparable. In this scenario, the purpose of standardization processing is to convert the noise data in the second feature-mapped audio after inverse robust scaling into a unified measurement standard for subsequent analysis and comparison. For example, the server can adopt the Z-score standardization method to convert the noise data in the audio into a distribution with a mean of 0 and a standard deviation of 1. In this way, the noise data is converted into a relatively fixed scale for subsequent processing. After standardization processing, the server determines the estimated noise label of the noisy tower monitoring audio instance based on the processed noise data. This noise label is actually an identifier used to represent the noise data in the audio. It can be a simple binary classification label (such as "noisy" or "noise-free"), or a more complex label that includes information such as the type and intensity of the noise. For example, the server can set a threshold, and when the standardized noise data exceeds this threshold, it is marked as "noisy", otherwise it is marked as "noise-free". Or, the server can also use more complex machine learning algorithms to classify and label the standardized noise data. Finally, the server uses the estimated noise label of the noisy tower monitoring audio instance as the estimated noise data information of this audio instance. This information can be stored in a database or sent to other systems or users through the network for subsequent processing and analysis. For example, the server can associate the estimated noise label with the original audio file and store them in the same database. In this way, when a user needs to query the noise situation of a certain audio file, the server can directly provide the estimated noise label of this file as a reference. Through this process, the server can accurately extract and label the noise data in the noisy tower monitoring audio instance, providing an important basis for subsequent tower status evaluation and noise control.
[0110] In an embodiment of the present invention, obtaining the bolt audio data in the noisy tower monitoring audio instance based on the estimated noise data information can be implemented through the following examples.
[0111] Obtain the noise data in the noisy tower monitoring audio instance according to the estimated noise data information;
[0112] Regard the audio data other than the noise data in the noisy tower monitoring audio instance as the bolt audio data in the noisy tower monitoring audio instance.
[0113] In an embodiment of the present invention, exemplarily, the server will first identify and extract the noise data in the noisy tower monitoring audio instance according to the estimated noise data information. The estimated noise data information may include the type of noise, timestamp, frequency range, etc. These information help the server accurately locate and extract the noise part in the audio. For example, if the estimated noise data information indicates that there is wind noise between the 10th second and the 20th second of the audio, the server will extract the data of this time period from the original audio and mark it as noise data. After identifying and extracting the noise data, the server will regard the audio data other than the noise data in the noisy tower monitoring audio instance as the bolt audio data. This part of the data is considered to be the sound related to the bolt state because they are not identified as noise. For example, if the total length of the original audio is 30 seconds and the server has extracted the wind noise between the 10th second and the 20th second, then the remaining audio data from the 1st second to the 10th second and from the 20th second to the 30th second will be regarded as the bolt audio data. This process is actually a data cleaning process. The server removes the noise data and retains the audio information directly related to the bolt state, providing a more accurate data basis for the subsequent bolt state analysis. In summary, through these two steps, the server can use the estimated noise data information to extract the bolt audio data from the noisy tower monitoring audio, providing important reference information for the maintenance and repair of the tower.
[0114] In an embodiment of the present invention, performing a tuning operation on the structural parameters of the first noise cancellation model according to the noise filtering cost parameter and the noise data estimation cost parameter, and determining a tower monitoring audio denoising model for performing noise filtering operation on the original tower monitoring audio can be implemented through the following examples.
[0115] Obtain the contribution degrees respectively matching the noise filtering cost parameter and the noise data estimation cost parameter;
[0116] According to the contribution degrees respectively matched by the noise filtering cost parameter and the noise data estimation cost parameter, perform multi-objective optimization processing on the noise filtering cost parameter and the noise data estimation cost parameter to determine the overall cost parameter of the first noise cancellation model;
[0117] According to the overall cost parameter of the first noise cancellation model, perform a tuning operation on the structural parameters of the first noise cancellation model to determine a tower monitoring audio denoising model for performing noise filtering operations on the original tower monitoring audio.
[0118] In the embodiment of the present invention, by way of example, the server first obtains the contribution degrees respectively matched by the noise filtering cost parameter and the noise data estimation cost parameter. The contribution degree can be understood as the importance and influence of each cost parameter in the model optimization process. For example, the server can determine the contribution degrees of these two cost parameters through historical data analysis or expert evaluation. For example, after analysis, the contribution degree of the noise filtering cost parameter is determined to be 0.6, while the contribution degree of the noise data estimation cost parameter is determined to be 0.4. This means that in the model optimization process, the effect of noise filtering is more important. Next, the server performs multi-objective optimization processing on these two cost parameters according to their contribution degrees. Multi-objective optimization is an optimization method that simultaneously considers multiple objectives, aiming to find a balance so that all objectives can be relatively well satisfied. In this scenario, the server will adopt an optimization algorithm (such as genetic algorithm, particle swarm optimization, etc.) and simultaneously consider the noise filtering effect and the accuracy of noise data estimation. Through continuous iteration and optimization, the server will finally determine an overall cost parameter, which reflects the optimal trade-off between the two cost parameters under the given contribution degrees. Finally, the server performs a tuning operation on the structural parameters of the model according to the overall cost parameter of the first noise cancellation model. The structural parameters can include the number of layers of the model, the number of neurons, the selection of activation functions, etc. For example, if the overall cost parameter indicates that the noise filtering effect is more important, the server will increase the complexity of the model (such as increasing the number of layers or neurons) to improve the denoising ability of the model. On the contrary, if the accuracy of noise data estimation is more important, the server will adjust the structure of the model to better capture and estimate noise data. Through this series of tuning operations, the server will finally determine a tower monitoring audio denoising model for performing noise filtering operations on the original tower monitoring audio. This model achieves a better balance between noise filtering and noise data estimation and can more effectively remove the noise components in the tower monitoring audio.
[0119] In the embodiment of the present invention, the acquisition of the noise-free tower monitoring audio instance, the noisy tower monitoring audio instance matched with the noise-free tower monitoring audio instance, and the noise label of the noisy tower monitoring audio instance can be implemented through the following examples.
[0120] Collect the tower monitoring audio with audio quality greater than the audio quality threshold from the tower monitoring audio data pool as the noise-free tower monitoring audio instance;
[0121] Degrade the sampling values of the audio sample points in the noise-free tower monitoring audio instance to determine the processed noise-free tower monitoring audio instance;
[0122] Perform signal enhancement processing on the processed noise-free tower monitoring audio instance to determine the to-be-determined noisy tower monitoring audio matched by the noise-free tower monitoring audio instance;
[0123] Receive the noise label for the to-be-determined noisy tower monitoring audio;
[0124] Based on the noise label of the to-be-determined noisy tower monitoring audio, the to-be-determined noisy tower monitoring audio, and the noise-free tower monitoring audio instance, obtain the noisy tower monitoring audio instance matched by the noise-free tower monitoring audio instance, and based on the noise label of the to-be-determined noisy tower monitoring audio, obtain the noise label of the noisy tower monitoring audio instance.
[0125] In an embodiment of the present invention, exemplarily, the server screens tower monitoring audio with audio quality higher than a set threshold from the tower monitoring audio data pool. These high-quality audios are regarded as noise-free tower monitoring audio instances. For example, the server sets an audio quality threshold of 0.9 (ranging from 0 to 1), and only when the comprehensive score of quality indicators such as audio clarity and signal-to-noise ratio is higher than 0.9, will the audio be selected as a noise-free tower monitoring audio instance. To simulate the possible noise situations in the real environment, the server degrades the collected noise-free tower monitoring audio instances. This usually involves adding different types of noise or distorting the audio signal to a certain extent. For example, the server adds Gaussian white noise to the sampling values of audio sample points, or simulates the loss during signal transmission to reduce the audio quality. After the degradation process, the server performs signal enhancement on these processed audios, aiming to restore or improve certain characteristics of the audio while retaining the added noise components. This step is to simulate the process of enhancing noisy audio in reality. For example, the server uses an equalizer to adjust the audio spectrum, or uses a noise reduction algorithm to try to reduce part of the noise while ensuring that the added noise is not completely removed. After processing the to-be-determined noisy tower monitoring audios, the server receives the noise labels for these audios from users or experts. These labels usually indicate the type and intensity of the noise in the audio. For example, users will listen to the audio through an interface and score it or select predefined noise labels such as "slight wind noise" and "strong electromagnetic interference". Finally, the server determines the noisy tower monitoring audio instances and their noise labels based on the received noise labels, the to-be-determined noisy tower monitoring audios, and the original noise-free tower monitoring audio instances. For example, if a to-be-determined audio is labeled as "slight wind noise", the server will store this audio, its corresponding noise-free version, and the noise label together as a dataset for subsequent model training and testing. Through the above steps, the server can construct a dataset containing noise-free and noisy tower monitoring audio instances and their corresponding noise labels, providing a rich data basis for subsequent audio processing and analysis.
[0126] In an embodiment of the present invention, to obtain the noisy tower monitoring audio instances matching the noise-free tower monitoring audio instances according to the noise labels of the to-be-determined noisy tower monitoring audios, the to-be-determined noisy tower monitoring audios, and the noise-free tower monitoring audio instances, the following example can be executed.
[0127] According to the noise labels, obtain the bolt audio data in the to-be-determined noisy tower monitoring audio, and the audio data corresponding to the bolt audio data in the noise-free tower monitoring audio instance;
[0128] Adopt the audio data corresponding to the bolt audio data in the noise-free iron tower monitoring audio example to replace the bolt audio data in the to-be-determined noisy iron tower monitoring audio, and determine the noisy iron tower monitoring audio example matched with the noise-free iron tower monitoring audio example.
[0129] In the embodiments of the present invention, exemplarily, the server will first identify and extract the audio data related to bolts from the to-be-determined noisy iron tower monitoring audio according to the noise label. These bolt audio data are the specific sounds of the iron tower bolts recorded in a noisy environment, and may include characteristic sounds such as bolt loosening and tightening. For example, if the noise label contains "bolt loosening sound", the server will use audio processing technologies, such as filtering and spectrum analysis, to locate and extract the audio segments corresponding to these loosening sounds from the to-be-determined noisy iron tower monitoring audio. Next, the server will search for the audio data corresponding to the extracted bolt audio data in the noise-free iron tower monitoring audio example. This is because the bolt sounds recorded in a noise-free environment can more clearly reflect the true state of the bolts, which is helpful for subsequent analysis and judgment. For example, the server can find the clear audio segment corresponding to the bolt loosening sound in the noise-free iron tower monitoring audio example by comparing spectrum features, time-domain waveforms, etc. Finally, the server will adopt the clear audio data corresponding to the bolt audio data in the noise-free iron tower monitoring audio example to replace the bolt audio data in the to-be-determined noisy iron tower monitoring audio. The purpose of this is to retain the noise environment in the original audio, and at the same time replace the noise-interfered bolt sounds with clearer bolt sounds, so as to obtain a noisy iron tower monitoring audio example that contains both the actual noise environment and clear bolt sounds. For example, if the original to-be-determined audio records the noise-interfered bolt loosening sound within the time period T, the server will replace it with the clear bolt loosening sound in the same time period of the noise-free audio. Through the above steps, the server can generate a noisy iron tower monitoring audio example that matches the noise-free iron tower monitoring audio example. This new audio example not only retains the actual noise environment, but also contains clear bolt sound data, providing more accurate data support for subsequent iron tower status monitoring and fault diagnosis.
[0130] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the disclosure to the precise forms disclosed. Numerous modifications and variations are possible in light of the above teachings. The embodiments are chosen and described in order to best illustrate the principles of the disclosure and its practical application, thereby enabling those skilled in the art to best utilize the disclosure and to utilize various embodiments with different modifications to suit the particular applications contemplated.
Claims
1. A tower bolt abnormality monitoring method based on deep learning and voiceprint, characterized in that: The method comprises: Get the original tower monitoring audio; Calling a pre-trained tower monitoring audio denoising model to process the original tower monitoring audio to obtain a target tower monitoring audio; Obtaining the bolt soundprint to be analyzed of the target tower monitoring audio; Loading the bolt voiceprint to be analyzed into a pre-trained bolt voiceprint anomaly classification model for voiceprint classification processing to obtain a target voiceprint anomaly classification result of the target tower monitoring audio; The tower monitoring audio denoising model is obtained in the following manner, including: collecting tower monitoring audio with audio quality greater than an audio quality threshold from a tower monitoring audio data pool as a noise-free tower monitoring audio instance; Degrading the sampling values of the audio sample points in the noise-free iron tower monitoring audio instance to determine a processed noise-free iron tower monitoring audio instance; Performing signal enhancement processing on the processed noise-free iron tower monitoring audio instance to determine the pending noisy iron tower monitoring audio matched by the noise-free iron tower monitoring audio instance; Receiving a noise label for the monitoring audio of the pending noisy iron tower; According to the noise label, the bolt audio data in the pending noisy iron tower monitoring audio and the audio data corresponding to the bolt audio data in the noise-free iron tower monitoring audio instance are obtained; The audio data corresponding to the bolt audio data in the noise-free tower monitoring audio instance is used to replace the bolt audio data in the pending noisy tower monitoring audio, and the noisy tower monitoring audio instance that matches the noise-free tower monitoring audio instance is determined; and the noise label of the noisy tower monitoring audio instance is obtained according to the noise label of the pending noisy tower monitoring audio; the noise label indicates the noise data in the noisy tower monitoring audio instance; Using the first noise elimination model, performing a noise filtering operation on the noisy iron tower monitoring audio instance to determine the first iron tower monitoring filtered audio, performing noise data estimation on the noisy iron tower monitoring audio instance to determine the estimated noise data information of the noisy iron tower monitoring audio instance; According to the estimated noise data information, the bolt audio data in the noisy iron tower monitoring audio instance is obtained, and from the first iron tower monitoring filtered audio, the audio data having the same timing node as the bolt audio data in the noisy iron tower monitoring audio instance is collected, and the bolt audio data in the noisy iron tower monitoring audio instance is used to replace the audio data in the first iron tower monitoring filtered audio, so as to determine the iron tower monitoring filtered audio instance; Obtaining a noise filtering cost parameter of the first noise elimination model according to the tower monitoring filtered audio instance and the noise-free tower monitoring audio instance; Obtaining a noise data estimation cost parameter of the first noise elimination model according to the estimated noise data information and the noise label; Obtaining the contribution of the noise filtering cost parameter and the noise data estimation cost parameter respectively matching each other; According to the contribution degrees of the noise filtering cost parameter and the noise data estimation cost parameter respectively matched, multi-objective optimization processing is performed on the noise filtering cost parameter and the noise data estimation cost parameter to determine the overall cost parameter of the first noise elimination model; According to the overall cost parameter of the first noise elimination model, a tuning operation is performed on the structural parameters of the first noise elimination model to determine a tower monitoring audio denoising model for performing a noise filtering operation on the original tower monitoring audio.
2. The method according to claim 1, characterized in that The bolt voiceprint anomaly classification model is obtained by the following methods, including: Obtaining a current bolt voiceprint instance of a target tower monitoring audio instance; the current bolt voiceprint instance is configured with a current voiceprint abnormality target value; According to the basic voiceprint classification model, an integrated voiceprint classification model and a target voiceprint classification model are constructed; the basic voiceprint classification model is obtained by loading the basic bolt voiceprint instance of the target tower monitoring audio instance into a preset classification model and performing bolt voiceprint abnormal classification training; the basic bolt voiceprint instance is configured with an initial target tower monitoring audio instance category label; the initial target tower monitoring audio instance category label is different from the current voiceprint abnormality target value; Loading the current bolt voiceprint instance into the integrated voiceprint classification model for voiceprint classification processing to obtain a first sample voiceprint type recognition result; Integrate the first sample voiceprint type recognition result and the current voiceprint abnormality target value to obtain a sample voiceprint abnormality target value; Loading the current bolt voiceprint instance into the target voiceprint classification model for voiceprint classification processing to obtain a second sample voiceprint type recognition result; According to the deviation between the second sample voiceprint type identification result and the sample voiceprint abnormal target value, a training process is performed on the target voiceprint classification model. The trained target voiceprint classification model is used to determine the bolt voiceprint abnormal classification model. The bolt voiceprint abnormal classification model is used to determine the type of the current voiceprint abnormal target value for the bolt voiceprint instance.
3. The method according to claim 2, characterized in that After obtaining the current bolt voiceprint instance of the target tower monitoring audio instance, the method further includes: Performing data cleaning processing on the current bolt voiceprint instance to obtain a transition bolt voiceprint instance, wherein the data cleaning processing is used to change the composition of the current bolt voiceprint instance; Loading the transition bolt voiceprint instance into the integrated voiceprint classification model for voiceprint classification processing to obtain a third sample voiceprint type recognition result; Loading the transition bolt voiceprint instance into the target voiceprint classification model for voiceprint classification processing to obtain a fourth sample voiceprint type recognition result; determining a first error parameter according to a deviation between the fourth sample voiceprint type recognition result and the third sample voiceprint type recognition result; According to the deviation between the second sample voiceprint type recognition result and the sample voiceprint abnormality target value, a training process is performed on the target voiceprint classification model, including: determining a second error parameter according to a deviation between the second sample voiceprint type recognition result and the sample voiceprint abnormality target value; Determining a target error parameter according to the first error parameter and the second error parameter; According to the target error parameter, the structural parameters of the target voiceprint classification model are adjusted.
4. The method according to claim 3, characterized in that The method further comprises: Adjusting the integrated voiceprint classification model according to the structural parameters of the target voiceprint classification model to obtain a candidate integrated voiceprint classification model; When the training cycle is the first preset cycle, after training for a training cycle of the second preset cycle, obtaining current structural parameters of the target voiceprint classification model; Processing the current structural parameters according to the state space model to obtain target structural parameters; Adjusting the candidate integrated voiceprint classification model according to the target structure parameter; When the preset training termination condition is reached, the candidate integrated voiceprint classification model is determined as the bolt voiceprint abnormality classification model, and the training cycle is the cycle of training the target voiceprint classification model using the current bolt voiceprint instance.
5. The method according to claim 4, characterized in that The step of processing the current structural parameters according to the state space model to obtain target structural parameters includes: Determine, according to the training cycle corresponding to the current structural parameter, a first influencing factor of the integrated voiceprint classification model and a second influencing factor of the target voiceprint classification model; the first influencing factor has a negative feedback relationship with the training cycle, and the second influencing factor has a positive feedback relationship with the training cycle; The current structural parameters are processed according to the first influencing factor, the second influencing factor and the state space model to obtain the target structural parameters.
6. The method according to claim 2, characterized in that The training method of the basic voiceprint classification model includes: Obtain the foundation bolt voiceprint instance of the target tower monitoring audio instance; Loading the basic bolt voiceprint instance into the preset classification model for voiceprint classification processing to obtain an initial sample voiceprint type recognition result; Determining a sample error parameter according to a deviation between the initial sample voiceprint type recognition result and the initial target tower monitoring audio instance category label; The structural parameters of the preset classification model are adjusted according to the sample error parameters, and the preset classification model that has completed training is determined as the basic voiceprint classification model.
7. The method according to claim 1, characterized in that The first noise elimination model includes a feature enhancement module, a noise filtering module and a noise data estimation module; the feature enhancement module includes a robust scaling submodule and a feature enhancement submodule; The noise filtering module includes a filtering coding submodule, a filtering feature mapping module and a filtering reverse robust scaling submodule; the noise data estimation module includes an estimation coding submodule, an estimation feature mapping module and an estimation reverse robust scaling submodule; Using the first noise elimination model, performing a noise filtering operation on the noisy iron tower monitoring audio instance, determining a first iron tower monitoring filtered audio, estimating noise data on the noisy iron tower monitoring audio instance, and determining estimated noise data information of the noisy iron tower monitoring audio instance, including: Using the robust scaling submodule, the noisy tower monitoring audio instance is robustly scaled according to the audio attribute matched by the noisy tower monitoring audio instance, and a robustly scaled noisy tower monitoring audio instance is determined; Using the feature enhancement submodule, feature enhancement is performed on the robustly scaled noisy iron tower monitoring audio instance to determine feature enhanced audio that matches the noisy iron tower monitoring audio instance; The output current intermediate feature audio of the current coding network is input into the subsequent coding network; the input of the initial coding network is the feature enhanced audio matched by the noisy tower monitoring audio instance; Using the post-order coding network, feature coding the output current intermediate feature audio to determine the initial feature coded audio matched by the post-order coding network; The initial feature coded audio matched by the post-order coding network and the output current intermediate feature audio are merged to determine the output intermediate feature audio of the post-order coding network; The output intermediate feature audio of the last encoding network is used as the first feature encoded audio of the noisy tower monitoring audio instance; Using the filtering feature mapping module, the first feature-coded audio is subjected to feature mapping processing to determine the first feature-mapped audio matched by the noisy tower monitoring audio instance; Using the filtering reverse robust scaling submodule, the first feature mapping audio is subjected to reverse robust scaling processing to determine the first tower monitoring filtered audio; Using the estimation coding submodule, feature coding is performed on the feature-enhanced audio matched by the noisy iron tower monitoring audio instance, and a second feature-coded audio matched by the noisy iron tower monitoring audio instance is determined; Using the estimated feature mapping module, the second feature-encoded audio is subjected to feature mapping processing to determine the second feature-mapped audio matched by the noisy tower monitoring audio instance; Using the estimated reverse robust scaling submodule, the second feature map audio is subjected to reverse robust scaling processing to determine the reverse robustly scaled second feature map audio; The second feature map audio after the reverse robust scaling is subjected to standardization processing to determine an estimated noise label of the noisy iron tower monitoring audio instance; the estimated noise label represents the noise data in the noisy iron tower monitoring audio instance; The estimated noise label of the noisy iron tower monitoring audio instance is regarded as the estimated noise data information of the noisy iron tower monitoring audio instance.
8. The method according to claim 1, characterized in that The step of obtaining the bolt audio data in the noisy iron tower monitoring audio instance according to the estimated noise data information includes: According to the estimated noise data information, noise data in the noisy tower monitoring audio instance is obtained; The audio data other than the noise data in the noisy iron tower monitoring audio instance is regarded as the bolt audio data in the noisy iron tower monitoring audio instance.
9. A server system, characterized in that: The method comprises a server, wherein the server is used to execute the method described in any one of claims 1 to 8.
Citation Information
Patent Citations
Iron tower bolt health detection method and system based on artificial intelligence and voiceprint
CN118212938A