Iron tower bolt state maintenance method and system
By enhancing the voiceprint feature and multi-layer perception processing of tower monitoring audio data, the problem of errors in bolt status monitoring in the prior art is solved, real-time monitoring and early warning of tower bolt status is realized, and maintenance efficiency and safety are improved.
Patent Information
- Application Number
- CN202510095911.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-01-22
- Publication Date
- 2025-05-23
AI Technical Summary
In the prior art, tower bolt status monitoring is prone to errors and fails to effectively use sample data for processing.
By obtaining tower monitoring audio data, extracting original voiceprints and inputting pre-trained voiceprint feature enhancement model, performing feature extraction and multi-layer perception conversion, determining the voiceprint type classification results, and generating alarm information based on the results.
Real-time monitoring and early warning of tower bolt status is realized, and maintenance efficiency, safety and accuracy are improved.
Smart Images

Figure CN120030472A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a method and system for maintaining the status of iron tower bolts. Background Art
[0002] In industries such as electricity and communications, iron towers are important infrastructure. The stability and safety of iron towers are of vital importance, and the tightening state of bolts is a key factor affecting the safety of iron tower structures. Traditional bolt inspection methods mainly rely on manual inspections, which are inefficient and easily affected by human factors.
[0003] There are also bolt status maintenance methods based on artificial intelligence in the prior art, but there are still certain problems: for example, the invention patent with publication number CN118782086A discloses a tower bolt maintenance method and system based on artificial intelligence and voiceprint, the method: de-interference processing is performed on the tower detection audio to be analyzed, and the voiceprint to be analyzed is extracted from the de-interferenced voiceprint to be analyzed; the voiceprint to be analyzed is input into a pre-trained voiceprint analysis model to obtain the voiceprint analysis result corresponding to the voiceprint to be analyzed; when the voiceprint analysis result is characterized as an abnormal result, the target emergency strategy is obtained from the preset tower bolt maintenance strategy library according to the emergency strategy association relationship pre-configured for the abnormal result, but it only considers the classification coding of the benchmark sample and the target sample, and does not consider whether the sample data is correct, that is, the data is not processed, therefore, the monitoring accuracy of the bolt status is questionable. Summary of the invention
[0004] Purpose of the invention: In order to overcome the deficiencies of the prior art, the present invention provides a method for maintaining the status of tower bolts, which solves the problem that errors are prone to occur in monitoring the status of tower bolts. In addition, the present invention also provides a system for maintaining the status of tower bolts.
[0005] Technical solution: According to a first aspect of the present invention, a method for maintaining the state of a tower bolt is provided, the method comprising:
[0006] Obtain the tower monitoring audio data to be processed;
[0007] Acquire an original voiceprint from the tower monitoring audio data, and input the original voiceprint into a pre-trained voiceprint feature enhancement model to obtain a target enhanced voiceprint;
[0008] Performing feature extraction and multi-layer perceptual conversion on the target enhanced voiceprint to determine a target voiceprint type classification result of the target enhanced voiceprint;
[0009] When the target voiceprint type classification result is characterized as a normal result, the normal result is recorded in a preset maintenance log;
[0010] When the target voiceprint type classification result is characterized as an abnormal result, an alarm message is generated and sent to the target interface.
[0011] Further, including:
[0012] The voiceprint feature enhancement model is obtained by the following methods, including:
[0013] Obtain the original voiceprint instance and the target voiceprint instance;
[0014] The original voiceprint instance is subjected to voiceprint feature transfer through a preset basic model to obtain the original voiceprint prediction confidence, and the original voiceprint instance is subjected to feature combination processing through a feature combination unit of a target model to obtain an original transition feature;
[0015] The voiceprint feature conversion unit and the voiceprint feature enhancement unit of the target model respectively perform feature corresponding conversion operations on the original transition features to obtain a first target voiceprint prediction confidence and a first target voiceprint feature enhancement confidence;
[0016] Determine the confidence difference between the original voiceprint prediction confidence and the first target voiceprint prediction confidence as the voiceprint prediction error of the original voiceprint instance;
[0017] According to different state representations of the original target value of the original voiceprint instance, different voiceprint correction functions are correspondingly set;
[0018] Wherein, when the state of the original target value is represented as an inactive state, the voiceprint correction function is adjusted based on the larger deviation between the voiceprint prediction error and the inactive state; when the state of the original target value is represented as an active state, the voiceprint correction function is set based on the smaller deviation between the voiceprint prediction error and the inactive state;
[0019] When the original target value state is represented as an inactive state and the voiceprint prediction error is positive or zero, the original target value is corrected by the voiceprint target value, and the state of the original corrected target value is represented as the voiceprint prediction error;
[0020] When the original target value state is represented as an inactive state and the voiceprint prediction error is negative, the original target value is corrected by the voiceprint target value, and the state of the original corrected target value is represented as an inactive state;
[0021] When the original target value state is represented as an activated state and the voiceprint prediction error is positive or zero, the original target value is corrected by the voiceprint target value, and the state of the original corrected target value is represented as an activated state;
[0022] When the original target value state is represented as an activated state and the voiceprint prediction error is negative, the original target value is corrected by the voiceprint target value to obtain a state representation of the original corrected target value adjusted based on the voiceprint prediction error and the activated state;
[0023] Performing feature combination processing on the target voiceprint instance through the feature combination unit of the target model to obtain a target transition feature;
[0024] The voiceprint feature conversion unit and the voiceprint feature enhancement unit of the target model respectively perform feature corresponding conversion operations on the target transition feature to obtain a second target voiceprint prediction confidence and a second target voiceprint feature enhancement confidence;
[0025] Performing voiceprint probability fusion processing on the second target voiceprint prediction confidence and the second target voiceprint feature enhancement confidence through the voiceprint probability fusion component of the target model to obtain a fused prediction confidence;
[0026] Based on the original correction target value, the target value of the target voiceprint instance, the first target voiceprint feature enhancement confidence and the fusion prediction confidence, a training process is performed on the target model to obtain a voiceprint feature enhancement model.
[0027] Further, including:
[0028] The voiceprint probability fusion component of the target model performs voiceprint probability fusion processing on the second target voiceprint prediction confidence and the second target voiceprint feature enhancement confidence to obtain a fused prediction confidence, including:
[0029] By means of the receiving unit of the voiceprint probability fusion component, a selective ignoring operation is performed on the enhanced confidence of the second target voiceprint feature to obtain the enhanced confidence of the second target voiceprint feature after the operation is performed;
[0030] The second target voiceprint prediction confidence and the second target voiceprint feature enhancement confidence after the execution operation are subjected to voiceprint probability fusion processing through the combining unit of the voiceprint probability fusion component to obtain the fused prediction confidence.
[0031] Further, including:
[0032] The receiving unit of the voiceprint probability fusion component performs a selective ignoring operation on the second target voiceprint feature enhancement confidence to obtain the second target voiceprint feature enhancement confidence after the operation is performed, including:
[0033] Acquiring the confidence of the ignore selection through the receiving unit of the voiceprint probability fusion component;
[0034] Determining a retention selection confidence of the second target voiceprint feature enhancement confidence based on the ignore selection confidence;
[0035] The second target voiceprint feature enhancement confidence is retained with the retention selection confidence, or the second target voiceprint feature enhancement confidence is updated to an initial value with the ignore selection confidence, to obtain the second target voiceprint feature enhancement confidence after the execution operation.
[0036] Further, including:
[0037] After performing a selective ignoring operation on the second target voiceprint feature enhancement confidence through the receiving unit of the voiceprint probability fusion component to obtain the second target voiceprint feature enhancement confidence after the operation is performed, the method further includes:
[0038] Acquire, through the receiving unit of the voiceprint probability fusion component, a first influence coefficient of the prediction confidence of the second target voiceprint and a second influence coefficient of the enhanced confidence of the second target voiceprint feature after the execution of the operation;
[0039] By means of the modified linear subunit in the receiving unit, based on the first influence coefficient and the second influence coefficient, the prediction confidence of the second target voiceprint and the enhanced confidence of the second target voiceprint feature after the operation are adjusted to obtain an input ratio of the combining unit;
[0040] The combining unit of the voiceprint probability fusion component performs voiceprint probability fusion processing on the second target voiceprint prediction confidence and the second target voiceprint feature enhancement confidence after the execution of the operation to obtain the fused prediction confidence, including:
[0041] The input ratio is subjected to voiceprint probability fusion processing by the combining unit of the voiceprint probability fusion component to obtain the fusion prediction confidence.
[0042] Further, including:
[0043] The step of performing a training process on the target model based on the original correction target value, the target value of the target voiceprint instance, the first target voiceprint feature enhancement confidence and the fusion prediction confidence to obtain a voiceprint feature enhancement model includes:
[0044] Setting a correction cost function based on the original correction target value and the first target voiceprint feature enhanced confidence;
[0045] Setting a target cost function based on the target objective value and the fusion prediction confidence;
[0046] Based on the correction cost function and the target cost function, a training process is performed on the target model to obtain the voiceprint feature enhancement model.
[0047] Further, including:
[0048] The step of executing a training process on the target model based on the correction cost function and the target cost function to obtain the voiceprint feature enhancement model includes:
[0049] Performing cost calculations on the correction cost function and the target cost function respectively, and obtaining correction cost parameters and target cost parameters accordingly;
[0050] Determining a total cost parameter of the target model according to the correction cost parameter and the target cost parameter;
[0051] Based on the total cost parameter, the model variables in the target model are adjusted and optimized according to preset loop conditions to obtain the voiceprint feature enhancement model.
[0052] Further, including:
[0053] The step of performing feature extraction and multi-layer perceptual conversion on the target enhanced voiceprint to determine a target voiceprint type classification result of the target enhanced voiceprint includes:
[0054] Acquire the target enhanced voiceprint;
[0055] Extracting features of the target enhanced voiceprint based on a shared feature extraction network to obtain a shared feature vector of the target enhanced voiceprint;
[0056] Performing MLP transformation on the common feature vector based on multiple fully connected networks to obtain comparative voiceprint features corresponding to each fully connected network, wherein the comparative voiceprint features output by each fully connected network are features of different voiceprint categories in the target enhanced voiceprint;
[0057] Determine the categories of multiple comparative voiceprint features respectively, and obtain the initial voiceprint type classification result corresponding to each comparative voiceprint feature;
[0058] The target voiceprint type classification result of the target enhanced voiceprint is determined according to the initial voiceprint type classification result corresponding to each compared voiceprint feature.
[0059] Further, including:
[0060] Each fully connected network is configured with a preprocessing component, and the common feature vector is subjected to MLP conversion based on multiple fully connected networks to obtain the comparative voiceprint features corresponding to each fully connected network, including:
[0061] A preprocessing component based on a target fully connected network configuration performs feature extraction on the common feature vector to obtain a pending feature vector corresponding to the target fully connected network, wherein the target fully connected network is any one of the multiple fully connected networks; the pending feature vector includes a first feature vector and a second feature vector, and the first feature vector and the second feature vector have different data acquisition rates;
[0062] Performing a first feature processing on the first feature vector based on the target fully connected network to obtain a processed first feature vector; the first feature processing is filtering processing;
[0063] Performing parallel data extraction on the second feature vector based on multiple local perception layers to obtain multiple data extraction features, wherein each local perception layer in the multiple local perception layers has a different data collection rate;
[0064] Based on the downsampling layer, the plurality of data extraction features are reduced in feature space, and the plurality of data extraction features after the feature space reduction are fused to obtain a response result corresponding to the second feature vector;
[0065] Performing dimension conversion on the response result corresponding to the second eigenvector;
[0066] Performing upward reconstruction processing on the response result after the dimension conversion to obtain a processed second eigenvector, wherein the feature space size of the processed second eigenvector is consistent with the feature space size of the first eigenvector;
[0067] Fusing the processed first feature vector and the processed second feature vector to obtain a fused feature vector;
[0068] The fused feature vector is subjected to filtering and upward reconstruction in sequence to obtain an upwardly reconstructed feature vector, wherein the feature space size of the upwardly reconstructed feature vector corresponds to the feature space size of the target enhanced voiceprint;
[0069] The upwardly reconstructed feature vector is used as the comparative voiceprint feature corresponding to the target fully connected network.
[0070] On the other hand, the present invention further provides a server system, characterized in that it includes a server, and the server is used to execute any one of the methods described above.
[0071] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0072] The present invention first obtains the tower monitoring audio data, extracts the original soundprint, and processes it using the pre-trained soundprint feature enhancement model to obtain the enhanced soundprint. Subsequently, the enhanced soundprint is classified through feature extraction and multi-layer perception conversion to determine whether the bolt status is normal. If the classification result is abnormal, the system will generate an alarm message and send it to the target interface.
[0073] In the voiceprint enhancement process, the present invention sets different voiceprint correction functions according to the different state representations of the original target value of the original voiceprint instance, and performs classification and calibration processing on the sample data, which greatly increases the validity of the sample and provides more accurate results for the classification results below.
[0074] Therefore, the present invention realizes real-time monitoring and early warning of the status of the tower bolts, and improves maintenance efficiency, safety and accuracy. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings required for use in the embodiments are briefly introduced below. It should be understood that the following drawings only illustrate certain embodiments of the present invention and should not be regarded as limiting the scope. For those of ordinary skill in the art, other related drawings can also be obtained based on these drawings without creative work.
[0076] Figure 1 A schematic diagram of the steps of the tower bolt emergency strategy management method provided by an embodiment of the present invention;
[0077] Figure 2 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0078] In order to make the purpose, technical solutions and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are part of the embodiments of the present invention, not all of the embodiments. Generally, the components of the embodiments of the present invention described and shown in the drawings here can be arranged and designed in various different configurations.
[0079] The specific implementation modes of the present invention are described in detail below in conjunction with the accompanying drawings.
[0080] In order to solve the technical problems in the aforementioned background technology, Figure 1 A flow chart of a tower bolt maintenance method based on deep learning and voiceprint provided in an embodiment of the present disclosure is provided. The tower bolt maintenance method based on deep learning and voiceprint is introduced in detail below.
[0081] Step S201, obtaining the tower monitoring audio data to be processed;
[0082] Step S202, obtaining an original voiceprint from the tower monitoring audio data, and inputting the original voiceprint into a pre-trained voiceprint feature enhancement model to obtain a target enhanced voiceprint;
[0083] Step S203, performing feature extraction and multi-layer perceptual conversion on the target enhanced voiceprint to determine a target voiceprint type classification result of the target enhanced voiceprint;
[0084] Step S204, when the target voiceprint type classification result is characterized as a normal result, the normal result is recorded in a preset maintenance log;
[0085] Step S205: when the target voiceprint type classification result is characterized as an abnormal result, an alarm message is generated and sent to the target interface.
[0086] In an embodiment of the present invention, exemplarily, the server automatically receives and stores audio data from audio monitoring equipment installed at key parts of the tower. For example, at key node locations of a 5G communication tower, such as bolt connections, highly sensitive audio acquisition equipment is installed. These devices continuously monitor and record the sounds that may be generated by bolts due to wind, temperature changes or loose bolts. The server periodically obtains these audio data from the audio monitoring equipment for subsequent processing. The server uses audio processing software to perform a preliminary analysis on the collected tower monitoring audio data and extracts the original voiceprint information therein. Subsequently, the server inputs these original voiceprint data into a voiceprint feature enhancement model that has been trained with a large number of bolt sound samples. This model can highlight the features in the sound related to the bolt state, filter out noise and other irrelevant information, and thus obtain an enhanced voiceprint data, namely the target enhanced voiceprint. The server further processes the target enhanced voiceprint and uses deep learning technology to extract key features in the sound, such as frequency, amplitude, etc. Next, the server uses a multi-layer perceptron (a neural network model) to classify these features. The multi-layer perception machine will classify the current target enhanced voiceprint as "normal" or "abnormal" based on the previously learned normal and abnormal voiceprint feature patterns of bolts. If the target voiceprint type classification result is characterized as normal, it means that the current state of the tower bolt is good and there is no abnormal sound. The server will record this normal result in the preset maintenance log. In this way, maintenance personnel can view the historical status records of the tower bolts at any time. If the target voiceprint type classification result is characterized as abnormal, such as detecting a specific sound generated when the bolt is loose, the server will immediately generate an alarm message. This alarm message will be sent to the target interface of the maintenance personnel, such as a mobile phone application or a computer monitoring system, so that they can know and deal with the problems with the tower bolts in a timely manner. In this way, real-time monitoring and early warning of the tower bolt status can be achieved, improving maintenance efficiency and safety.
[0087] In the embodiment of the present invention, the voiceprint feature enhancement model is obtained in the following manner.
[0088] Obtain the original voiceprint instance and the target voiceprint instance;
[0089] The original voiceprint instance is subjected to voiceprint feature transfer through a preset basic model to obtain an original voiceprint prediction confidence, and the original voiceprint instance is subjected to voiceprint feature transfer through a voiceprint feature conversion unit and a voiceprint feature enhancement unit of a target model to obtain a first target voiceprint prediction confidence and a first target voiceprint feature enhancement confidence;
[0090] Based on the original voiceprint prediction confidence and the first target voiceprint prediction confidence, performing voiceprint target value correction on the original target value of the original voiceprint instance to obtain an original corrected target value of the original voiceprint instance;
[0091] The target voiceprint instance is respectively transferred with the voiceprint feature conversion unit and the voiceprint feature enhancement unit of the target model to obtain a second target voiceprint prediction confidence and a second target voiceprint feature enhancement confidence;
[0092] Performing voiceprint probability fusion processing on the second target voiceprint prediction confidence and the second target voiceprint feature enhancement confidence through the voiceprint probability fusion component of the target model to obtain a fused prediction confidence;
[0093] Based on the original correction target value, the target value of the target voiceprint instance, the first target voiceprint feature enhancement confidence and the fusion prediction confidence, a training process is performed on the target model to obtain a voiceprint feature enhancement model.
[0094] In the embodiment of the present invention, illustratively, the server selects two types of voiceprint data from the previously collected tower bolt monitoring audio database. One type is the original voiceprint instance, which is the sound record of the bolt working normally under various environmental conditions. The other type is the target voiceprint instance, which specifically marks the sound characteristics of the bolt when it is abnormal (such as loosening, precursors to fracture, etc.). The server first uses a preset basic model (which can be a preliminary deep learning model) to perform feature analysis and transmission on the original voiceprint instance. This model will output an original voiceprint prediction confidence, indicating the model's confidence in the normal working state of these voiceprint data. Then, the server uses a more complex target model, which includes a voiceprint feature conversion unit and a voiceprint feature enhancement unit. These two units perform more in-depth feature extraction and enhancement processing on the original voiceprint instance respectively. After the processing is completed, the model will give the first target voiceprint prediction confidence and the first target voiceprint feature enhancement confidence, which respectively indicate the model's confidence in voiceprint classification and voiceprint feature enhancement. The server compares the prediction confidences given by the base model and the target model. If a difference is found, the target value (i.e., label) of the original voiceprint instance will be corrected. This correction is intended to reduce the error during model training and ensure the accuracy of the label. The corrected target value is called the original corrected target value. For the target voiceprint instance marked as abnormal, the server also uses the target model for processing. The voiceprint feature conversion unit and the voiceprint feature enhancement unit will extract and enhance the features of these data respectively, and then give the second target voiceprint prediction confidence and the second target voiceprint feature enhancement confidence. The server uses a special component in the target model, the voiceprint probability fusion component, to fuse the above two confidences. This fusion can comprehensively consider the performance of the model in both classification and feature enhancement, and derive a more comprehensive and accurate fusion prediction confidence. Finally, the server uses all the above data and confidence information to train the target model. This training process is intended to enable the model to more accurately identify and classify normal and abnormal voiceprints of tower bolts, while enhancing key features to improve recognition accuracy. After training, this model is called a voiceprint feature enhancement model, which can be used for subsequent tower bolt status monitoring and early warning.
[0095] In order to more clearly describe the solution provided by the embodiment of the present invention, a more detailed description is given below.
[0096] In the embodiment of the present invention, the original voiceprint instance refers to a voiceprint data sample collected in a specific scenario (such as the environment when the tower bolts are working normally) without any processing or enhancement. These data samples are usually used to train the model to identify the voiceprint features in the "normal" state. For example, audio acquisition equipment is installed on the tower of a wind power station, and these devices will continuously record the sound of the tower bolts in normal working state. These sound data, such as the sound when the bolts are tightened, the slight sound when the wind blows through the bolts, etc., will be collected and used as the original voiceprint instance. The target voiceprint instance refers to those voiceprint data samples that are specifically marked with a specific state (such as abnormal conditions such as loose bolts and damage). These data samples are crucial for training the model to identify "abnormal" voiceprint features. If the bolts of a tower become loose due to long-term use, then when the bolts emit sounds different from the normal state under the action of wind, these abnormal sound data will be recorded and used as target voiceprint instances. These instances will contain the sound frequency, amplitude and other characteristics unique to the loose bolts.
[0097] In addition, the preset base model is a pre-built model that is usually used for preliminary processing or analysis of data. In the context of voiceprint feature transfer, this base model can be a trained deep learning model that is used to extract basic features from the input voiceprint. For example, there is an audio classification model based on a convolutional neural network (CNN). This model has been trained with a large amount of audio data and can recognize basic features in the audio. Now, this trained model is used as the preset base model for preliminary feature extraction and analysis of the original voiceprint instance. Voiceprint feature transfer is a process in which the parameters and structure of the model are used to extract features from the input voiceprint data. These features can include the frequency, rhythm, pitch, etc. of the sound, which are crucial for subsequent classification and recognition. When an original voiceprint instance is input into the preset base model, the model processes the input layer by layer and finally outputs a feature vector. This process is the transfer of voiceprint features, which converts the original audio data into a form that the model can understand and analyze. The original voiceprint prediction confidence is the confidence of the base model in its prediction results. In the context of voiceprint recognition, it indicates the degree of certainty that the model believes that the input original voiceprint instance belongs to a certain category. For example, after the basic model analyzes an input original voiceprint instance, it gives a prediction result and outputs a confidence score of 0.9. This means that the model has 90% confidence that this voiceprint instance belongs to the category it predicts (such as "normal bolt sound"). The target model is the main model to be trained or optimized, which can contain more complex structures and algorithms to handle more advanced tasks or improve the accuracy of predictions. In the scenario of voiceprint recognition, the target model can be a deep learning network that contains multiple layers and complex connections to extract more refined features from the input voiceprint data and perform accurate classification. The voiceprint feature conversion unit and the voiceprint feature enhancement unit are specific parts of the target model, which are responsible for converting and enhancing the input voiceprint features, respectively. The conversion unit can convert the original features into a form more suitable for model processing, while the enhancement unit may highlight certain key features or suppress unimportant information. For example, the target model of the embodiment of the present invention includes a convolution layer specifically used to convert voiceprint features and an attention mechanism layer used to enhance specific features. The convolutional layer can convert the input voiceprint data into a higher-level feature representation, while the attention mechanism layer can emphasize or weaken the importance of certain features according to the learning objectives of the model. The first target voiceprint prediction confidence and the first target voiceprint feature enhancement confidence respectively represent the confidence of the target model in its prediction results after voiceprint feature conversion and enhancement. The first target voiceprint prediction confidence focuses on the accuracy of classification, while the first target voiceprint feature enhancement confidence focuses on the effectiveness of feature enhancement.After being processed by the target model, you may get a classification prediction and a corresponding confidence score (such as a prediction confidence of 0.85) for the input voiceprint instance. At the same time, the model will also give a confidence score for the feature enhancement effect (such as a feature enhancement confidence of 0.9). These scores can help you understand the model's confidence level in its predictions and feature enhancement effects.
[0098] In addition, voiceprint target value correction is a process in which the original target value is adjusted or corrected by comparing and analyzing the prediction confidence given by different models. The purpose of correction is to improve the accuracy and reliability of the data, especially in the presence of noise or labeling errors. If the original voiceprint instance is labeled as "normal", but both the base model and the target model give a low prediction confidence (for example, both are below a certain threshold), then the original target value of this instance needs to be corrected to reflect the model's more accurate judgment of the true state of this instance.
[0099] The original corrected target value is the new target value obtained after the voiceprint target value is corrected. It may be the same as the original target value or different, depending on the results of the correction process. If the correction process determines that the original voiceprint instance may actually be "abnormal", then the original corrected target value will be set to "abnormal". The second target voiceprint prediction confidence and the second target voiceprint feature enhancement confidence respectively represent the confidence of the target model in its prediction results after voiceprint feature conversion and enhancement. The second target voiceprint prediction confidence focuses on the classification accuracy after feature conversion, while the second target voiceprint feature enhancement confidence focuses on the effectiveness of feature enhancement. When the target model processes a target voiceprint instance, it outputs a classification prediction (for example, bolt loose or normal) and a confidence score about this prediction (i.e., the second target voiceprint prediction confidence). At the same time, the model also gives a confidence score about the feature enhancement effect (i.e., the second target voiceprint feature enhancement confidence) to indicate the contribution of the enhanced features to the prediction.
[0100] In addition, the voiceprint probability fusion component is a special component in the target model, which is used to fuse the confidences from different sources (here, the second target voiceprint prediction confidence and the second target voiceprint feature enhancement confidence) to produce a more comprehensive and reliable confidence indicator. This fusion is usually based on certain probability or statistical methods. For example, there are two confidence scores, one is the prediction confidence based on feature conversion (0.85), and the other is the confidence based on feature enhancement (0.9). The voiceprint probability fusion component can use some weighted average or other statistical methods to fuse these two scores, such as taking the average value of the two (0.875) as the final fused prediction confidence. The fused prediction confidence refers to the comprehensive confidence obtained by fusing the second target voiceprint prediction confidence and the second target voiceprint feature enhancement confidence through the voiceprint probability fusion component. It reflects the overall prediction confidence of the model after comprehensively considering feature conversion and feature enhancement. If the prediction confidence after fusion is 0.88, it means that the model has 88% confidence that its prediction result is correct after comprehensively considering all factors. The voiceprint feature enhancement model refers to a model obtained through a specific training process, which has high performance in voiceprint feature enhancement. This model can more effectively use the enhanced voiceprint features for prediction and classification. For example, there is a deep learning network that performs well in recognizing voiceprint data with enhanced features after being trained with specific training data and processes. This trained network is the voiceprint feature enhancement model of the embodiment of the present invention.
[0101] In an embodiment of the present invention, based on the original voiceprint prediction confidence and the first target voiceprint prediction confidence, the original target value of the original voiceprint instance is corrected to obtain the original corrected target value of the original voiceprint instance, which can be implemented through the following examples.
[0102] Determine the confidence difference between the original voiceprint prediction confidence and the first target voiceprint prediction confidence as the voiceprint prediction error of the original voiceprint instance;
[0103] Based on the voiceprint prediction error, setting a voiceprint correction function;
[0104] The voiceprint target value correction is performed on the original target value of the original voiceprint instance through the voiceprint correction function to obtain the original corrected target value of the original voiceprint instance.
[0105] In the embodiment of the present invention, for example, the server first obtains the original voiceprint prediction confidence of an original voiceprint instance, such as 0.8, which means that the model has 80% confidence in the classification or recognition of this voiceprint instance. Then, the server obtains the first target voiceprint prediction confidence of the same voiceprint instance, such as 0.9. The difference between the two confidences, i.e., 0.9-0.8=0.1, is determined by the server as the voiceprint prediction error of this voiceprint instance. Based on the voiceprint prediction error of 0.1 calculated in the previous step, the server needs to set a voiceprint correction function to adjust the original target value. This function can consider a variety of factors, such as the size and direction of the error, and other possible correction parameters. For example, the server can set a linear correction function, in which the error value is used as an adjustment coefficient to linearly interpolate or extrapolate the original target value to obtain a corrected target value. The server now uses the voiceprint correction function set in the previous step to correct the original target value of the original voiceprint instance. For example, if the original target value is a label representing a sound category (such as "normal sound" or "abnormal sound"), the server can adjust the confidence of this label or directly change the label according to the voiceprint prediction error. For example, if the original target value is "normal sound", but the voiceprint prediction error shows that the model is not confident enough about this judgment, the server can adjust the label to "may be a normal sound" or change it to "abnormal sound" according to the correction function. In the embodiment of the present invention, the voiceprint correction function can be any one of a linear correction function, a nonlinear correction function, and a correction function based on machine learning, which is not limited here.
[0106] In the embodiment of the present invention, the voiceprint correction function is set based on the voiceprint prediction error, which can be implemented through the following examples.
[0107] According to different state representations of the original target value of the original voiceprint instance, different voiceprint correction functions are correspondingly set;
[0108] Among them, when the state of the original target value is represented as an inactive state, the voiceprint correction function is adjusted based on the larger deviation between the voiceprint prediction error and the inactive state; when the state of the original target value is represented as an active state, the voiceprint correction function is set based on the smaller deviation between the voiceprint prediction error and the inactive state.
[0109] In an embodiment of the present invention, the server first analyzes the original target value of the original voiceprint instance, which can represent different states. For example, in the scenario of tower bolt maintenance, the target value can represent the state of the bolt, such as "tightening" or "loosening". For these two different states, the server will set different voiceprint correction functions. For example, the server recognizes that the original target value of a certain original voiceprint instance indicates that the bolt is "tightened" (i.e., inactive), but the voiceprint prediction error shows that the model predicts that the bolt may be in a "loose" state, which has a large deviation. In order to correct this deviation, the server will set a voiceprint correction function, which will focus on this large prediction error and make a large adjustment to the original target value to ensure that the model can more accurately predict the true state of the bolt. For example, if the original target value is 0 (indicating tightening), but the voiceprint prediction error is large, such as reaching 0.3 (within a standardized range), then the voiceprint correction function can adjust the original target value upward by a certain amount to reflect the risk of loosening of the bolt. On the contrary, if the server recognizes that the original target value of a certain original voiceprint instance indicates that the bolt is "loose" (i.e., activated), but the voiceprint prediction error is relatively small, it means that the model's prediction is closer to the actual state. In this case, the server will set a relatively mild voiceprint correction function to fine-tune the original target value. For example, if the original target value is 1 (indicating looseness) and the voiceprint prediction error is only 0.1, then the voiceprint correction function will only make a small adjustment to the original target value to maintain the accuracy of the model's prediction. With this setting, the server can flexibly adjust the voiceprint correction function according to different situations, thereby improving the model's prediction accuracy for the bolt status of the tower. This helps to promptly detect and deal with potential bolt loosening problems and ensure the safe operation of the tower.
[0110] In the embodiment of the present invention, the voiceprint target value correction is performed on the original target value of the original voiceprint instance by using the voiceprint correction function to obtain the original corrected target value of the original voiceprint instance, which can be implemented through the following examples.
[0111] When the original target value state is represented as an inactive state and the voiceprint prediction error is positive or zero, the original target value is corrected by the voiceprint target value, and the state of the original corrected target value is represented as the voiceprint prediction error;
[0112] When the original target value state is represented as an inactive state and the voiceprint prediction error is negative, the original target value is corrected by the voiceprint target value, and the state of the original corrected target value is represented as an inactive state;
[0113] When the original target value state is represented as an activated state and the voiceprint prediction error is positive or zero, the original target value is corrected by the voiceprint target value, and the state of the original corrected target value is represented as an activated state;
[0114] When the original target value state is represented as an activated state and the voiceprint prediction error is negative, the original target value is subjected to voiceprint target value correction to obtain a state representation of the original corrected target value that is adjusted based on the voiceprint prediction error and the activated state.
[0115] In an embodiment of the present invention, for example, in the scenario of tower bolt maintenance, for example, the server recognizes the voiceprint data of a bolt, and its original target value state is represented as "tightened" (inactive state), that is, the bolt should be tightened. However, through model prediction, it is found that the voiceprint prediction error is positive or zero, which means that the model predicts that the bolt is not completely tightened or has a tendency to loosen. The server will use the voiceprint correction function to correct the original target value. In this case, the state of the original corrected target value will be adjusted to a state corresponding to the voiceprint prediction error. For example, if the voiceprint prediction error is positive, the corrected target value can indicate that the bolt is slightly loose. Similarly, in the scenario of tower bolt maintenance, if the server recognizes that the voiceprint data of the bolt indicates that the bolt is in a "tightened" state, but the voiceprint prediction error is negative, this means that the model predicts that the bolt is tighter than it actually is, or too tight. In this case, the server will correct the original target value through the voiceprint correction function, but the corrected state is still represented as "tightened" (inactive state), because the negative error indicates that the actual state may be tighter than predicted. For example, the server recognizes that the voiceprint data of a bolt indicates that the bolt is "loose" (activated state). If the voiceprint prediction error is positive or zero, this means that the model predicts that the bolt is indeed loose or the degree of looseness is consistent with the prediction. The server will use the voiceprint correction function to correct the original target value, and the corrected state is still expressed as "loose" (activated state), because a positive error or a zero error indicates that the model's prediction is consistent with the actual state or the bolt may be looser. In the scenario of tower bolt maintenance, if the server recognizes that the voiceprint data of the bolt indicates that the bolt is in a "loose" state, but the voiceprint prediction error is negative, this means that the model predicts that the bolt may not be so loose in fact, or it may be tighter than it actually is. In this case, the server will use the voiceprint correction function to correct the original target value. The corrected state will be adjusted based on the voiceprint prediction error and the "loose" state, and a state between "tight" and "loose" may be obtained to more accurately reflect the actual state of the bolt. Through such a correction process, the server can more accurately evaluate the actual state of the tower bolt, so as to take effective maintenance measures in time.
[0116] In the embodiment of the present invention, the voiceprint feature conversion unit and the voiceprint feature enhancement unit of the target model respectively transfer the voiceprint features of the original voiceprint instance to obtain the first target voiceprint prediction confidence and the first target voiceprint feature enhancement confidence, which can be implemented through the following examples.
[0117] Performing feature combination processing on the original voiceprint instance through the feature combination unit of the target model to obtain the original transition feature;
[0118] The voiceprint feature conversion unit and the voiceprint feature enhancement unit of the target model respectively perform feature correspondence conversion operations on the original transition features to obtain the first target voiceprint prediction confidence and the first target voiceprint feature enhancement confidence.
[0119] In an embodiment of the present invention, exemplarily, in a voiceprint recognition system, the server receives an original voiceprint instance, which can be a voice recording of a person. The server first processes this original voiceprint instance using the feature combination unit in the target model. The task of the feature combination unit is to extract and combine key features in the voiceprint, such as pitch, timbre, formants, etc., to form a more comprehensive and richer feature set, which is called the original transition feature. For example, the server can extract acoustic features such as the fundamental frequency, formants, and spectral envelope of the speaker from the recording and combine these features together to form a multi-dimensional feature vector, which is the original transition feature. After extracting the original transition features, the server passes these features to two different units in the target model: the voiceprint feature conversion unit and the voiceprint feature enhancement unit. The task of the voiceprint feature conversion unit is to convert the original transition features into a form that is easier to classify and recognize. For example, it can use a certain transformation (such as linear discriminant analysis LDA, principal component analysis PCA, etc.) to reduce the dimension of the features while retaining the most useful information. The converted features are used for voiceprint recognition and a prediction confidence level is given, that is, the first target voiceprint prediction confidence level. This confidence level represents the degree of certainty of the model that the voiceprint instance belongs to a specific speaker. Different from the voiceprint feature conversion unit, the purpose of the voiceprint feature enhancement unit is to further enhance the useful information in the original transition features to highlight the unique acoustic features of the speaker. This can be achieved by weighting, normalizing the features or using other enhancement techniques. The enhanced features are used to verify the accuracy of voiceprint recognition and a feature enhancement confidence level is given, that is, the first target voiceprint feature enhancement confidence level. This confidence level reflects the contribution degree of the enhanced features to improving the accuracy of voiceprint recognition. For example, when the server processes a voice recording, the voiceprint feature conversion unit can reduce the dimension of the features through the LDA transformation and give a prediction confidence level of 0.9 (indicating a 90% certainty level). At the same time, the voiceprint feature enhancement unit can weight certain key features to highlight the unique timbre of the speaker and give a feature enhancement confidence level of 0.85 (indicating that there is an 85% probability that the enhanced features will improve the recognition accuracy).
[0120] In an embodiment of the present invention, the voiceprint feature transfer of the target voiceprint instance is respectively performed through the voiceprint feature conversion unit and the voiceprint feature enhancement unit of the target model to obtain the second target voiceprint prediction confidence level and the second target voiceprint feature enhancement confidence level, which can be implemented through the following examples.
[0121] The target transition features are obtained by performing feature combination processing on the target voiceprint instance through the feature combination unit of the target model;
[0122] The voiceprint feature conversion unit and the voiceprint feature enhancement unit of the target model respectively perform feature correspondence conversion operations on the target transition features to obtain the second target voiceprint prediction confidence and the second target voiceprint feature enhancement confidence.
[0123] In the embodiment of the present invention, the server first performs feature combination processing on the collected target voiceprint instance through the feature combination unit of the target model. The purpose of this step is to effectively combine multiple features in the original voiceprint data, so as to extract more representative feature information, namely, the target transition feature. For example, the original voiceprint data may contain multiple features such as the vibration frequency, amplitude, and duration of the bolt. The feature combination unit will fuse these features to form a comprehensive feature vector, namely, the target transition feature. This feature vector can more comprehensively characterize the voiceprint characteristics of the bolt, and provide a basis for subsequent feature conversion and enhancement. Next, the server performs feature corresponding conversion operation on the target transition feature through the voiceprint feature conversion unit of the target model. The purpose of this step is to convert the combined features into a form that is easier to analyze and identify, and at the same time extract key information closely related to the bolt state. For example, the voiceprint feature conversion unit can use structures such as convolutional neural network (CNN) or recurrent neural network (RNN) in deep learning technology to perform deep feature extraction and conversion on the target transition feature. Through this step, the server can obtain the prediction confidence of the second target voiceprint, which reflects the degree of correlation between the converted voiceprint feature and the bolt state. Finally, the server strengthens the target transition features through the voiceprint feature strengthening unit of the target model. The purpose of this step is to further highlight the feature information most relevant to the bolt state and suppress irrelevant or redundant features, thereby improving the accuracy of subsequent bolt state recognition. For example, the voiceprint feature strengthening unit can use technical means such as attention mechanism to weighted strengthen the key information in the target transition features. Through this step, the server can obtain the second target voiceprint feature strengthening confidence, which reflects the reliability and importance of the strengthened voiceprint feature in bolt state recognition. In summary, in the tower bolt maintenance scenario, the server processes and analyzes the voiceprint features of the target voiceprint instance through the deep learning model, and performs operations such as feature combination, voiceprint feature conversion and voiceprint feature strengthening in turn, and finally obtains the key information closely related to the bolt state and the corresponding confidence index. This information provides an important basis for subsequent bolt state monitoring and early warning.
[0124] In an embodiment of the present invention, the voiceprint probability fusion component of the target model is used to perform voiceprint probability fusion processing on the second target voiceprint prediction confidence and the second target voiceprint feature enhancement confidence to obtain a fused prediction confidence, which can be implemented through the following examples.
[0125] By means of the receiving unit of the voiceprint probability fusion component, a selective ignoring operation is performed on the enhanced confidence of the second target voiceprint feature to obtain the enhanced confidence of the second target voiceprint feature after the operation is performed;
[0126] The second target voiceprint prediction confidence and the second target voiceprint feature enhancement confidence after the execution operation are subjected to voiceprint probability fusion processing through the combining unit of the voiceprint probability fusion component to obtain the fused prediction confidence.
[0127] In the embodiment of the present invention, exemplarily, the server first performs a selective ignoring operation on the enhanced confidence of the second target voiceprint feature through the receiving unit of the voiceprint probability fusion component. The purpose of this step is to remove those inaccurate or abnormal confidence values that may be caused by noise or other interference factors. For example, when the server processes the voiceprint data of a group of bolts, it is found that the enhanced confidence of the second target voiceprint feature of one of the bolts is abnormally high or abnormally low. This may be because the bolt is interfered by external noise when collecting voiceprint data. In order to avoid the adverse effects of such abnormal values on subsequent fusion processing, the server selectively ignores this abnormal confidence value, thereby obtaining a corrected and more accurate enhanced confidence of the second target voiceprint feature. Next, the server performs voiceprint probability fusion processing on the second target voiceprint prediction confidence and the enhanced confidence of the second target voiceprint feature after the operation is performed through the combining unit of the voiceprint probability fusion component. The purpose of this step is to effectively combine the two confidence values to obtain a more comprehensive and accurate fusion prediction confidence. For example, the server has obtained a second target voiceprint prediction confidence of 0.8 for a bolt, and the second target voiceprint feature enhancement confidence after the selective ignoring operation is 0.7. In the voiceprint probability fusion processing, the server can use weighted average or other fusion algorithms to fuse these two confidence values into a fused prediction confidence. This fused prediction confidence can more comprehensively reflect the actual status of the bolt and provide more accurate information for subsequent status assessment and early warning. Through such a processing flow, the server can more effectively use deep learning and voiceprint technology to monitor and evaluate the status of the tower bolts, promptly discover and deal with potential problems, and ensure the safe operation of the tower.
[0128] In the embodiment of the present invention, the receiving unit of the voiceprint probability fusion component performs a selective ignoring operation on the enhanced confidence of the second target voiceprint feature to obtain the enhanced confidence of the second target voiceprint feature after the operation is performed, which can be implemented through the following examples.
[0129] Acquiring the confidence of the ignore selection through the receiving unit of the voiceprint probability fusion component;
[0130] Determining a retention selection confidence of the second target voiceprint feature enhancement confidence based on the ignore selection confidence;
[0131] The second target voiceprint feature enhancement confidence is retained with the retention selection confidence, or the second target voiceprint feature enhancement confidence is updated to an initial value with the ignore selection confidence, to obtain the second target voiceprint feature enhancement confidence after the execution operation.
[0132] In the embodiment of the present invention, exemplarily, when the server processes the voiceprint data of the tower bolt, it first evaluates whether the current voiceprint feature enhancement confidence should be ignored through the receiving unit of the voiceprint probability fusion component. This evaluation can be based on a series of factors, such as the environmental noise level, the performance of the recording equipment, and the accuracy of the previous voiceprint recognition. For example, if the server detects that the surrounding noise is high during recording, this can affect the accuracy of the voiceprint feature enhancement. Therefore, the receiving unit calculates an ignore selection confidence, which represents the possibility that the voiceprint feature enhancement confidence is ignored under the current conditions. After determining the ignore selection confidence, the server calculates the retain selection confidence based on this value. The retain selection confidence reflects the extent to which the second target voiceprint feature enhancement confidence is still valid and should be retained under the current conditions. If the ignore selection confidence is low (meaning that the current environment has a small impact on the voiceprint feature enhancement), the retain selection confidence will be relatively high. Conversely, if the ignore selection confidence is high (meaning that the current environment has a large impact on the voiceprint feature enhancement), the retain selection confidence will be relatively low. Finally, the server will decide how to handle the enhanced confidence of the second target voiceprint feature based on the retention selection confidence and the ignore selection confidence. If the retention selection confidence is higher than a preset threshold (such as 0.6), the server will choose to retain the second target voiceprint feature enhanced confidence and use it for subsequent voiceprint probability fusion processing. If the ignore selection confidence is higher than the preset threshold (such as 0.4), the server will choose to update the second target voiceprint feature enhanced confidence to an initial value (such as set to 0 or a default value) to reduce its impact on the final fusion prediction confidence. For example, in the tower bolt maintenance scenario, if the server detects that the voiceprint data of a bolt is interfered by strong wind noise, resulting in a high ignore selection confidence, then it will choose to reset the corresponding second target voiceprint feature enhanced confidence to the initial value to ensure that the final fusion prediction confidence is not affected by such interference.
[0133] In an embodiment of the present invention, after performing a selective ignoring operation on the enhanced confidence of the second target voiceprint feature through the receiving unit of the voiceprint probability fusion component and obtaining the enhanced confidence of the second target voiceprint feature after the operation is performed, it can be implemented in the following manner.
[0134] Acquire, through the receiving unit of the voiceprint probability fusion component, a first influence coefficient of the prediction confidence of the second target voiceprint and a second influence coefficient of the enhanced confidence of the second target voiceprint feature after the execution of the operation;
[0135] By means of the modified linear subunit in the receiving unit, based on the first influence coefficient and the second influence coefficient, the prediction confidence of the second target voiceprint and the enhanced confidence of the second target voiceprint feature after the operation are adjusted to obtain an input ratio of the combining unit;
[0136] The combining unit of the voiceprint probability fusion component performs voiceprint probability fusion processing on the second target voiceprint prediction confidence and the second target voiceprint feature enhancement confidence after the execution operation to obtain the fused prediction confidence, which can be implemented through the following examples.
[0137] The input ratio is subjected to voiceprint probability fusion processing by the combining unit of the voiceprint probability fusion component to obtain the fusion prediction confidence.
[0138] In the embodiment of the present invention, exemplarily, the server first obtains the first influence coefficient of the prediction confidence of the second target voiceprint through the receiving unit of the voiceprint probability fusion component. This coefficient may be based on multiple factors, such as the historical accuracy of the prediction model, the richness of the training data, etc. For example, if the prediction model has performed well in the prediction of similar voiceprint data in the past, the first influence coefficient can be relatively high. At the same time, the receiving unit also obtains the second influence coefficient of the second target voiceprint feature enhancement confidence after the operation is performed. This coefficient may take into account factors such as the effectiveness of feature enhancement and the quality of data used in the enhancement process. If the feature enhancement process can significantly improve the accuracy of recognition, the second influence coefficient can be increased accordingly. The corrected linear subunit in the receiving unit adjusts the second target voiceprint prediction confidence and the second target voiceprint feature enhancement confidence after the operation is performed according to the first influence coefficient and the second influence coefficient. This process is similar to weighting two confidences to reflect their respective importance in the fusion process. For example, if the first influence coefficient is 0.7 and the second influence coefficient is 0.3, the corrected linear subunit can adjust the weights of the two confidences according to this ratio to obtain an input ratio. This ratio will be used as the input of the combination unit of the voiceprint probability fusion component. Finally, the combination unit of the voiceprint probability fusion component will perform voiceprint probability fusion processing according to the input ratio. This process can be a weighted average or a more complex probability fusion algorithm, which aims to comprehensively consider the prediction confidence of the second target voiceprint and the enhanced confidence of the second target voiceprint feature after the operation is performed, so as to obtain a more accurate and reliable fusion prediction confidence. For example, in the voiceprint recognition of tower bolts, if the voiceprint data of a bolt is similar to the voiceprint of the loose bolt in the training data, but has a certain difference from the voiceprint of the tightened bolt, then by fusing the prediction confidence, the server can more accurately judge the status of the bolt, so as to issue a maintenance alarm or take other necessary measures in time.
[0139] In an embodiment of the present invention, the training process is performed on the target model based on the original correction target value, the target value of the target voiceprint instance, the first target voiceprint feature enhancement confidence and the fusion prediction confidence to obtain a voiceprint feature enhancement model, which can be implemented through the following examples.
[0140] Setting a correction cost function based on the original correction target value and the first target voiceprint feature enhanced confidence;
[0141] Setting a target cost function based on the target objective value and the fusion prediction confidence;
[0142] Based on the correction cost function and the target cost function, a training process is performed on the target model to obtain the voiceprint feature enhancement model.
[0143] In the embodiment of the present invention, the server first obtains a set of original voiceprint data of the bolt state, and obtains an original correction target value after preliminary processing, which reflects the baseline situation of the bolt state. At the same time, the server also obtains a first target voiceprint feature enhancement confidence through the deep learning model, and this confidence represents the confidence of the model in the extracted bolt voiceprint feature. For example, the original correction target value represents the standard voiceprint feature of the bolt tightening state, and the first target voiceprint feature enhancement confidence is the recognition confidence of the model for this feature. The server will set a correction cost function based on the gap between the two, which is used to measure the difference between the model prediction and the actual state, and provide guidance for subsequent model training. The server also obtains the bolt state target value (i.e., the target target value) annotated by the expert, which represents the actual situation of the bolt state. At the same time, through the previous steps, the server has obtained a fusion prediction confidence, which is the prediction confidence of the model for the bolt state after considering multiple factors. For example, the target target value indicates that the bolt has loosened, and the fusion prediction confidence is the confidence of the model in this judgment. Based on the difference between the two, the server sets a target cost function, which is used to measure the accuracy of the model prediction and provide optimization direction for subsequent model training. The server now has two cost functions: the correction cost function and the target cost function. These two functions reflect the performance of the model in identifying the bolt soundprint features and predicting the bolt status, respectively. Next, the server will use these two cost functions to train the deep learning model. During the training process, the model will continuously adjust its parameters to minimize the two cost functions, thereby improving the ability to recognize the bolt soundprint features and the accuracy of predicting the bolt status. Finally, after multiple rounds of iterative training, the server will obtain an optimized soundprint feature enhancement model. This model can more accurately identify the soundprint features of the bolts and more reliably predict the status of the bolts, thereby providing strong support for the maintenance of the tower bolts.
[0144] In the embodiment of the present invention, the training process is performed on the target model based on the correction cost function and the target cost function to obtain the voiceprint feature enhancement model, which can be implemented through the following examples.
[0145] Performing cost calculations on the correction cost function and the target cost function respectively, and obtaining correction cost parameters and target cost parameters accordingly;
[0146] Determining a total cost parameter of the target model according to the correction cost parameter and the target cost parameter;
[0147] Based on the total cost parameter, the model variables in the target model are adjusted and optimized according to preset loop conditions to obtain the voiceprint feature enhancement model.
[0148] In an embodiment of the present invention, illustratively, the server first calculates the values of the correction cost function and the target cost function respectively, and these two values are called the correction cost parameter and the target cost parameter respectively. The correction cost parameter reflects the gap between the model prediction and the original correction target value, while the target cost parameter reflects the gap between the model prediction and the target target value marked by the expert. For example, the server can calculate a specific value, such as the correction cost parameter is 0.8 and the target cost parameter is 0.6, and these two values respectively represent the error degree of the model in identifying the bolt soundprint feature and predicting the bolt state. The server will then determine the total cost parameter of the target model based on the correction cost parameter and the target cost parameter. This total cost parameter comprehensively considers the performance of the model in both recognition and prediction. For example, the server can perform a weighted average of the correction cost parameter and the target cost parameter to obtain a total cost parameter, such as 0.7. This value represents the overall performance of the model. Finally, the server will adjust and optimize the model variables in the target model according to the total cost parameter and the preset loop conditions. This process may include modifying the parameters of the model, adjusting the structure of the model, or changing the training strategy of the model. For example, if the total cost parameter is high, it means that the performance of the model needs to be improved. The server can optimize the model by increasing the depth of the model, increasing the amount of training data, adjusting the learning rate, etc. After multiple rounds of iterative training and optimization, the server will obtain a voiceprint feature enhancement model with better performance. This optimized model can more accurately identify the voiceprint features of the tower bolts and more reliably predict the status of the bolts. For example, in actual applications, the server can use this model to monitor the status of the tower bolts in real time, and promptly detect and warn of potential bolt loosening or damage, thereby ensuring the safe and stable operation of the tower.
[0149] In the embodiment of the present invention, the feature extraction and multi-layer perceptual conversion of the target enhanced voiceprint to determine the target voiceprint type classification result of the target enhanced voiceprint can be implemented through the following examples.
[0150] Acquire the target enhanced voiceprint;
[0151] Extracting features of the target enhanced voiceprint based on a shared feature extraction network to obtain a shared feature vector of the target enhanced voiceprint;
[0152] Performing MLP transformation on the common feature vector based on multiple fully connected networks to obtain comparative voiceprint features corresponding to each fully connected network, wherein the comparative voiceprint features output by each fully connected network are features of different voiceprint categories in the target enhanced voiceprint;
[0153] According to the obtained comparative voiceprint features among the multiple comparative voiceprint features, a target voiceprint type classification result of the target enhanced voiceprint is determined.
[0154] In an embodiment of the present invention, illustratively, the server first obtains the target enhanced voiceprint from the voiceprint database. These voiceprint data can be bolt voiceprint signals collected at different times and in different environments, and are preprocessed and enhanced. For example, the server can obtain the voiceprint signal of a specific bolt under wind load as the target enhanced voiceprint. The server then uses a pre-trained common feature extraction network to extract features from the target enhanced voiceprint. This network can be a convolutional neural network (CNN) based on deep learning, which can effectively extract common feature vectors from voiceprint signals. For example, the server extracts common features such as frequency and energy in the target enhanced voiceprint through the CNN network to form a high-dimensional feature vector. The server then uses multiple fully connected networks (also known as multi-layer perceptrons, MLPs) to transform the extracted common feature vectors. Each fully connected network is trained for different voiceprint categories (such as bolt tightening, bolt loosening, etc.) to output comparative voiceprint features of the category. For example, the server may have three fully connected networks, corresponding to the three states of bolt tightening, bolt loosening and bolt damage. The common feature vectors will be input into these three networks respectively to obtain the comparative voiceprint features output by each network. Finally, the server will determine the target voiceprint type classification result of the target enhanced voiceprint based on the multiple comparative voiceprint features obtained. This is usually achieved by comparing the confidence of the comparative voiceprint features output by different fully connected networks. For example, if the server finds that the comparative voiceprint feature output by the fully connected network corresponding to the loose bolt has the highest confidence, then it will judge that the bolt state corresponding to the target enhanced voiceprint is loose. In this way, the server can accurately identify the current state of the tower bolts, thereby providing timely warnings and maintenance suggestions to maintenance personnel.
[0155] In an embodiment of the present invention, each fully connected network is configured with a preprocessing component, and the common feature vector is subjected to MLP conversion based on multiple fully connected networks to obtain the comparative voiceprint features corresponding to each fully connected network, which can be implemented through the following examples.
[0156] Performing feature extraction on the common feature vector based on a preprocessing component configured in a target fully connected network to obtain a pending feature vector corresponding to the target fully connected network, wherein the target fully connected network is any one of the multiple fully connected networks;
[0157] Based on the target fully connected network, the undetermined feature vector is subjected to MLP transformation to obtain a comparative voiceprint feature corresponding to the target fully connected network.
[0158] In an embodiment of the present invention, exemplarily, before processing the common feature vector, the server will further extract features from these features through the preprocessing component configured in each fully connected network. The preprocessing component may include various data normalization, dimensionality reduction or feature selection techniques. For example, the server is processing a common feature vector, which contains multiple features of the bolt soundprint, such as frequency, amplitude, etc. The preprocessing component will screen and optimize these features, extract the most representative features for the subsequent fully connected network, and generate a pending feature vector. Specifically, if a fully connected network focuses on identifying the tightening state of the bolt, its preprocessing component can emphasize features related to the tightening state of the bolt, such as the soundprint intensity in a specific frequency range, while suppressing other irrelevant or redundant features. After processing by the preprocessing component, the server will obtain a pending feature vector for a specific fully connected network. This vector will then be input into the corresponding fully connected network. The fully connected network is a multi-layer perceptron (MLP), which learns and maps the relationship between input features and output labels through nonlinear transformations of multiple layers of neurons. For example, the server inputs a preprocessed undetermined feature vector into a fully connected network that focuses on identifying the loose state of bolts. This network has learned the characteristic pattern of bolt loosening soundprints through training. After the network's forward propagation calculation, the server will obtain a comparative soundprint feature, which represents the similarity between the input soundprint and the loose state of the bolt. In this way, the server can use multiple fully connected networks specifically for different bolt states to refine the input common feature vectors, thereby improving the accuracy and reliability of soundprint recognition. This helps to detect various states of tower bolts in a timely and accurate manner, providing strong support for maintenance work.
[0159] In an embodiment of the present invention, the undetermined feature vector includes a first feature vector and a second feature vector, and the data acquisition rates corresponding to the first feature vector and the second feature vector are different; the MLP conversion of the undetermined feature vector based on the target fully connected network to obtain the comparative voiceprint features corresponding to the target fully connected network can be implemented through the following examples.
[0160] Performing a first feature processing on the first feature vector based on the target fully connected network to obtain a processed first feature vector;
[0161] Performing second feature processing on the second feature vector to obtain a processed second feature vector;
[0162] Fusing the processed first feature vector and the processed second feature vector to obtain a fused feature vector;
[0163] The fused feature vector is processed to obtain a comparative voiceprint feature corresponding to the target fully connected network.
[0164] In an embodiment of the present invention, exemplarily, the server first obtains a first feature vector, which may be obtained at a higher data acquisition rate, such as collecting data thousands of times per second, so as to capture more details of the bolt soundprint. For example, these high-frequency collected data may include the tiny vibration mode of the bolt under wind load. The server processes the first feature vector through a fully connected network (i.e., the target fully connected network), which may include data normalization, denoising or feature enhancement operations, and finally obtains the processed first feature vector. Unlike the first feature vector, the second feature vector may be obtained at a lower data acquisition rate, such as collecting data dozens of times per second. Although these data have fewer details, they can reflect the long-term trend or periodic change of the bolt state. The server performs a second feature processing on the second feature vector, which may include data smoothing, feature extraction and other operations to adapt to its different data acquisition rates, and obtain the processed second feature vector. The server then fuses the processed first feature vector and the processed second feature vector. The fusion method can be simple splicing or more complex algorithms, such as weighted fusion or feature transformation. For example, the server can align two feature vectors in a certain dimension and then concatenate them into a longer feature vector. This fused feature vector contains both high-frequency and low-frequency bolt voiceprint information. Finally, the server processes the fused feature vector to obtain the comparative voiceprint features corresponding to the target fully connected network. This processing can include nonlinear transformations through fully connected layers, application of activation functions, and possible dropout operations to prevent overfitting. Ultimately, the server outputs a comparative voiceprint feature that combines the information of high-frequency and low-frequency data, providing a richer and more accurate basis for subsequent bolt status classification or prediction. In this way, the server can comprehensively utilize voiceprint information at different data acquisition rates to improve the accuracy and reliability of bolt status detection.
[0165] In the embodiment of the present invention, the processing of the fused feature vector to obtain the comparative voiceprint feature corresponding to the target fully connected network can be implemented through the following examples.
[0166] The fused feature vector is subjected to filtering and upward reconstruction in sequence to obtain an upwardly reconstructed feature vector, wherein the feature space size of the upwardly reconstructed feature vector corresponds to the feature space size of the target enhanced voiceprint;
[0167] The upwardly reconstructed feature vector is used as the comparative voiceprint feature corresponding to the target fully connected network.
[0168] In an embodiment of the present invention, illustratively, first, the server performs filtering on the fused feature vector. The filtering process is intended to remove noise and redundant information in the fused feature vector and highlight the key features of the bolt voiceprint. For example, the server can apply a low-pass filter to remove high-frequency noise, or apply a band-pass filter to retain only signals within a specific frequency range that is closely related to the bolt state. Next, the server performs upward reconstruction on the filtered feature vector. Upward reconstruction is a technique that aims to map low-dimensional features back to a high-dimensional space so that it corresponds to the feature space size of the original data (i.e., the target enhanced voiceprint). This is usually achieved through a deep learning model (such as an autoencoder), which learns how to restore compressed features to the original high-dimensional space during training. For example, if the size of the fused feature vector is 100 dimensions and the feature space size of the target enhanced voiceprint is 1000 dimensions, the server will use a trained autoencoder model to upwardly reconstruct the 100-dimensional fused feature vector into a 1000-dimensional feature vector. After filtering and upward reconstruction, the upwardly reconstructed feature vector obtained by the server has the same feature space size as the target enhanced voiceprint. At this point, the server will use this upwardly reconstructed feature vector as the comparative voiceprint feature corresponding to the target fully connected network. This comparative voiceprint feature now contains information of the same dimension as the original voiceprint data, but has been optimized by deep learning and feature engineering, so it is more representative and discriminative. This helps the server to more accurately judge the status of the bolts in subsequent classification or recognition tasks. For example, in the maintenance scenario of tower bolts, the server can use this comparative voiceprint feature to train a classifier to distinguish whether the bolts are tightened, loose, or damaged. In this way, when new voiceprint data is input, the server can quickly and accurately identify the current status of the bolts and issue maintenance alerts in a timely manner.
[0169] In the embodiment of the present invention, the first feature processing is filtering processing, and the second feature processing is performed on the second feature vector to obtain the processed second feature vector, which can be implemented through the following examples.
[0170] Performing data extraction and fusion processing on the second feature vector based on the feature analysis component to obtain a response result corresponding to the second feature vector;
[0171] Performing dimension conversion on the response result corresponding to the second eigenvector;
[0172] The response result after the dimension conversion is upwardly reconstructed to obtain a processed second eigenvector, and the feature space size of the processed second eigenvector is consistent with the feature space size of the first eigenvector.
[0173] In an embodiment of the present invention, exemplarily, for the first eigenvector, the server will perform filtering processing. Considering that the first eigenvector may contain high-frequency data, filtering processing helps to remove noise and interference and highlight useful signals. For example, the server can use a bandpass filter to retain only signals within a specific frequency range related to bolt state monitoring, thereby improving the accuracy of subsequent analysis. For the second eigenvector, since it may contain data with lower frequencies or longer time spans, the server will use the feature parsing component to perform data extraction and fusion processing. The purpose of this step is to extract the key information in the second eigenvector and fuse it into a more representative response result. For example, the feature parsing component can identify periodic signals or long-term trends in the second eigenvector, and fuse this information into a comprehensive indicator to reflect the long-term changes in the bolt state. After obtaining the response result, the server will perform dimensional conversion to ensure that it is consistent with the feature space size of the first eigenvector. Dimension conversion can include dimensionality reduction or dimensionality increase operations, depending on the initial dimension of the response result and the dimension of the first eigenvector. For example, if the dimension of the response result is lower than the first eigenvector, the server can use methods such as principal component analysis (PCA) to increase the dimension; conversely, if the dimension of the response result is higher than the first eigenvector, it can be reduced in dimension. Finally, the server will reconstruct the response result after the dimension conversion. The purpose of this step is to ensure that the processed second eigenvector is consistent with the first eigenvector in the feature space, so as to facilitate subsequent feature fusion and classification processing. Upward reconstruction processing may be achieved through a deep learning model (such as an autoencoder), which can learn how to map low-dimensional features to high-dimensional space while keeping the intrinsic structure of the data unchanged. Through the above steps, the server can effectively process the first and second eigenvectors, extract key information, and fuse them into more representative feature vectors, providing strong support for subsequent tower bolt status monitoring and classification.
[0174] In an embodiment of the present invention, the feature analysis component includes a downsampling layer and multiple local perception layers. The feature analysis component performs data extraction and fusion processing on the two feature vectors to obtain a response result corresponding to the second feature vector, which can be implemented through the following example.
[0175] Performing parallel data extraction on the second feature vector based on the multiple local perception layers to obtain multiple data extraction features, wherein the data collection rate of each local perception layer in the multiple local perception layers is different;
[0176] Based on the downsampling layer, feature space reduction is performed on the multiple data extraction features, and the multiple data extraction features after feature space reduction are fused to obtain a response result corresponding to the second feature vector.
[0177] In an embodiment of the present invention, exemplarily, multiple local perception layers in the feature parsing component are designed to perform parallel data extraction on the second feature vector at different data acquisition rates. For example, there are three local perception layers, corresponding to three different data acquisition rates of high, medium and low. High-level data acquisition rate: This layer focuses on capturing high-frequency details in the second feature vector, which can be the immediate response or slight changes of the bolts in a specific environment. For example, it can capture the immediate vibration mode of the bolts of the tower under strong wind. Middle-level data acquisition rate: This layer focuses on medium-frequency data changes, which can reflect daily fluctuations or periodic changes in the state of the bolts. For example, it can detect slight deformations of the bolts caused by temperature changes. Low-level data acquisition rate: This layer pays more attention to long-term trends and slow changes, and can be used to identify signs of wear or aging of the bolts. For example, by analyzing long-term data changes, the service life of the bolts can be predicted. Through the parallel processing of these three local perception layers, the server can simultaneously capture the changes in the state of the bolts on different time scales, providing more comprehensive information for subsequent feature fusion and state judgment. After multiple local perception layers complete data extraction, the downsampling layer will reduce and fuse the feature space of these extracted features. The main function of the downsampling layer is to reduce the dimension of the features while retaining the most important information for subsequent classification and identification. For example, each local perception layer extracts dozens of features, and the downsampling layer will reduce these features to a few most representative features through specific algorithms (such as maximum pooling, average pooling, etc.). Then, these reduced features will be fused into a comprehensive response result, which contains both high-frequency details of the bolt status and reflects medium-term fluctuations and long-term trends. Through such a processing flow, the server can generate a comprehensive and refined second feature vector response result, which provides strong data support for the maintenance and condition monitoring of tower bolts.
[0178] In the embodiment of the present invention, the determination of the target voiceprint type classification result of the target enhanced voiceprint according to the comparative voiceprint features among the obtained multiple comparative voiceprint features can be implemented through the following examples.
[0179] Determine the categories of the multiple comparative voiceprint features respectively, and obtain an initial voiceprint type classification result corresponding to each comparative voiceprint feature;
[0180] The target voiceprint type classification result of the target enhanced voiceprint is determined according to the initial voiceprint type classification result corresponding to each compared voiceprint feature.
[0181] In the embodiment of the present invention, the server first determines the category of each comparison voiceprint feature. This step is achieved by deep learning models (such as convolutional neural networks or recurrent neural networks), which have been trained on a large amount of voiceprint data and can accurately identify the category corresponding to the voiceprint feature. For example, the server can have three comparison voiceprint features, corresponding to three different states of the bolt: tightened, loose, and damaged. For each comparison voiceprint feature, the server uses the trained deep learning model to infer and obtain the initial voiceprint type classification result corresponding to each feature. After obtaining the initial voiceprint type classification result of each comparison voiceprint feature, the server will combine these results to determine the target voiceprint type classification result of the target enhanced voiceprint. For example, if two of the three comparison voiceprint features are classified as "loose" and one is classified as "tight", then the server can determine the target voiceprint type classification result of the target enhanced voiceprint as "loose" according to the majority principle. In addition, the server can also consider the weight relationship between the comparison voiceprint features. For example, some features may be more important or more decisive in the classification. In this case, the server can adjust the final classification result according to the weight of the feature. Through this processing flow, the server can accurately determine the status of the tower bolts, so that timely maintenance and replacement can be carried out to ensure the safe operation of the tower. This bolt status monitoring method based on deep learning and voiceprint is efficient, accurate and automated, and can greatly improve the efficiency and safety of tower maintenance.
[0182] The embodiment of the present invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the above-mentioned tower bolt maintenance method based on deep learning and voiceprint. Figure 2 As shown, Figure 2 The block diagram of the computer device 100 provided in the embodiment of the present invention. The computer device 100 includes a memory 111, a processor 112 and a communication unit 113. To achieve data transmission or interaction, the memory 111, the processor 112 and the communication unit 113 are directly or indirectly electrically connected to each other. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.
[0183] For illustrative purposes, the foregoing description is made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the present disclosure to the precise form disclosed. Numerous modifications and variations are possible in accordance with the above teachings. These embodiments are selected and described in order to best illustrate the principles of the present disclosure and its practical application, so that those skilled in the art can best utilize the present disclosure and utilize various embodiments with different modifications to suit the intended specific application.
Claims
1. A method for maintaining the status of a tower bolt, characterized in that: The method includes: Obtain the tower monitoring audio data to be processed; Acquire an original voiceprint from the tower monitoring audio data, and input the original voiceprint into a pre-trained voiceprint feature enhancement model to obtain a target enhanced voiceprint; Performing feature extraction and multi-layer perceptual conversion on the target enhanced voiceprint to determine a target voiceprint type classification result of the target enhanced voiceprint; When the target voiceprint type classification result is characterized as a normal result, the normal result is recorded in a preset maintenance log; When the target voiceprint type classification result is characterized as an abnormal result, an alarm message is generated and sent to the target interface.
2. The method according to claim 1, characterized in that The voiceprint feature enhancement model is obtained by the following methods, including: Obtain the original voiceprint instance and the target voiceprint instance; The original voiceprint instance is subjected to voiceprint feature transfer through a preset basic model to obtain the original voiceprint prediction confidence, and the original voiceprint instance is subjected to feature combination processing through a feature combination unit of a target model to obtain an original transition feature; The voiceprint feature conversion unit and the voiceprint feature enhancement unit of the target model respectively perform feature corresponding conversion operations on the original transition features to obtain a first target voiceprint prediction confidence and a first target voiceprint feature enhancement confidence; Determine the confidence difference between the original voiceprint prediction confidence and the first target voiceprint prediction confidence as the voiceprint prediction error of the original voiceprint instance; According to different state representations of the original target value of the original voiceprint instance, different voiceprint correction functions are correspondingly set; Wherein, when the state of the original target value is represented as an inactive state, the voiceprint correction function is adjusted based on the larger deviation between the voiceprint prediction error and the inactive state; when the state of the original target value is represented as an active state, the voiceprint correction function is set based on the smaller deviation between the voiceprint prediction error and the inactive state; When the original target value state is represented as an inactive state and the voiceprint prediction error is positive or zero, the original target value is corrected by the voiceprint target value, and the state of the original corrected target value is represented as the voiceprint prediction error; When the original target value state is represented as an inactive state and the voiceprint prediction error is negative, the original target value is corrected by the voiceprint target value, and the state of the original corrected target value is represented as an inactive state; When the original target value state is represented as an activated state and the voiceprint prediction error is positive or zero, the original target value is corrected by the voiceprint target value, and the state of the original corrected target value is represented as an activated state; When the original target value state is represented as an activated state and the voiceprint prediction error is negative, the original target value is corrected by the voiceprint target value to obtain a state representation of the original corrected target value adjusted based on the voiceprint prediction error and the activated state; Performing feature combination processing on the target voiceprint instance through the feature combination unit of the target model to obtain a target transition feature; The voiceprint feature conversion unit and the voiceprint feature enhancement unit of the target model respectively perform feature correspondence conversion operations on the target transition features to obtain a second target voiceprint prediction confidence and a second target voiceprint feature enhancement confidence; Performing voiceprint probability fusion processing on the second target voiceprint prediction confidence and the second target voiceprint feature enhancement confidence through the voiceprint probability fusion component of the target model to obtain a fused prediction confidence; Based on the original correction target value, the target value of the target voiceprint instance, the first target voiceprint feature enhancement confidence and the fusion prediction confidence, a training process is performed on the target model to obtain a voiceprint feature enhancement model.
3. The method according to claim 2, characterized in that The voiceprint probability fusion component of the target model performs voiceprint probability fusion processing on the second target voiceprint prediction confidence and the second target voiceprint feature enhancement confidence to obtain a fused prediction confidence, including: Performing a selective ignoring operation on the enhanced confidence of the second target voiceprint feature through the receiving unit of the voiceprint probability fusion component to obtain the enhanced confidence of the second target voiceprint feature after the operation is performed; The second target voiceprint prediction confidence and the second target voiceprint feature enhancement confidence after the execution operation are subjected to voiceprint probability fusion processing through the combining unit of the voiceprint probability fusion component to obtain the fused prediction confidence.
4. The method according to claim 3, characterized in that: The receiving unit of the voiceprint probability fusion component performs a selective ignoring operation on the second target voiceprint feature enhancement confidence to obtain the second target voiceprint feature enhancement confidence after the operation is performed, including: Acquiring the confidence of the ignore selection through the receiving unit of the voiceprint probability fusion component; Determining a retention selection confidence of the second target voiceprint feature enhancement confidence based on the ignore selection confidence; The second target voiceprint feature enhancement confidence is retained with the retention selection confidence, or the second target voiceprint feature enhancement confidence is updated to an initial value with the ignore selection confidence, to obtain the second target voiceprint feature enhancement confidence after the execution operation.
5. The method according to claim 4, characterized in that After performing a selective ignoring operation on the second target voiceprint feature enhancement confidence through the receiving unit of the voiceprint probability fusion component to obtain the second target voiceprint feature enhancement confidence after the operation is performed, the method further includes: Acquire, through the receiving unit of the voiceprint probability fusion component, a first influence coefficient of the prediction confidence of the second target voiceprint and a second influence coefficient of the enhanced confidence of the second target voiceprint feature after the execution of the operation; By means of the modified linear subunit in the receiving unit, based on the first influence coefficient and the second influence coefficient, the prediction confidence of the second target voiceprint and the enhanced confidence of the second target voiceprint feature after the operation are adjusted to obtain an input ratio of the combining unit; The combining unit of the voiceprint probability fusion component performs voiceprint probability fusion processing on the second target voiceprint prediction confidence and the second target voiceprint feature enhancement confidence after the execution of the operation to obtain the fused prediction confidence, including: The input ratio is subjected to voiceprint probability fusion processing by the combining unit of the voiceprint probability fusion component to obtain the fusion prediction confidence.
6. The method according to claim 2, characterized in that The step of performing a training process on the target model based on the original correction target value, the target value of the target voiceprint instance, the first target voiceprint feature enhancement confidence and the fusion prediction confidence to obtain a voiceprint feature enhancement model includes: Setting a correction cost function based on the original correction target value and the first target voiceprint feature enhanced confidence; Setting a target cost function based on the target objective value and the fusion prediction confidence; Based on the correction cost function and the target cost function, a training process is performed on the target model to obtain the voiceprint feature enhancement model.
7. The method according to claim 6, characterized in that The step of executing a training process on the target model based on the correction cost function and the target cost function to obtain the voiceprint feature enhancement model includes: Performing cost calculations on the correction cost function and the target cost function respectively, and obtaining correction cost parameters and target cost parameters correspondingly; Determining a total cost parameter of the target model according to the correction cost parameter and the target cost parameter; Based on the total cost parameter, the model variables in the target model are adjusted and optimized according to preset loop conditions to obtain the voiceprint feature enhancement model.
8. The method according to claim 1, characterized in that The step of performing feature extraction and multi-layer perceptual conversion on the target enhanced voiceprint to determine a target voiceprint type classification result of the target enhanced voiceprint includes: Acquire the target enhanced voiceprint; Extracting features of the target enhanced voiceprint based on a shared feature extraction network to obtain a shared feature vector of the target enhanced voiceprint; Performing MLP transformation on the common feature vector based on multiple fully connected networks to obtain comparative voiceprint features corresponding to each fully connected network, wherein the comparative voiceprint features output by each fully connected network are features of different voiceprint categories in the target enhanced voiceprint; Determine the categories of multiple comparative voiceprint features respectively, and obtain the initial voiceprint type classification result corresponding to each comparative voiceprint feature; The target voiceprint type classification result of the target enhanced voiceprint is determined according to the initial voiceprint type classification result corresponding to each compared voiceprint feature.
9. The method according to claim 8, characterized in that Each fully connected network is configured with a preprocessing component, and the common feature vector is subjected to MLP conversion based on multiple fully connected networks to obtain the comparative voiceprint features corresponding to each fully connected network, including: A preprocessing component based on a target fully connected network configuration performs feature extraction on the common feature vector to obtain a pending feature vector corresponding to the target fully connected network, wherein the target fully connected network is any one of the multiple fully connected networks; the pending feature vector includes a first feature vector and a second feature vector, and the first feature vector and the second feature vector have different data acquisition rates; Performing a first feature processing on the first feature vector based on the target fully connected network to obtain a processed first feature vector; the first feature processing is filtering processing; Performing parallel data extraction on the second feature vector based on multiple local perception layers to obtain multiple data extraction features, wherein each local perception layer in the multiple local perception layers has a different data collection rate; Based on the downsampling layer, the plurality of data extraction features are reduced in feature space, and the plurality of data extraction features after the feature space reduction are fused to obtain a response result corresponding to the second feature vector; Performing dimension conversion on the response result corresponding to the second eigenvector; Performing upward reconstruction processing on the response result after the dimension conversion to obtain a processed second eigenvector, wherein the feature space size of the processed second eigenvector is consistent with the feature space size of the first eigenvector; Fusing the processed first feature vector and the processed second feature vector to obtain a fused feature vector; The fused feature vector is subjected to filtering and upward reconstruction in sequence to obtain an upwardly reconstructed feature vector, wherein the feature space size of the upwardly reconstructed feature vector corresponds to the feature space size of the target enhanced voiceprint; The upwardly reconstructed feature vector is used as the comparative voiceprint feature corresponding to the target fully connected network.
10. A server system, characterized in that: The method comprises a server, wherein the server is used to execute the method described in any one of claims 1 to 9.
Citation Information
Patent Citations
Iron tower bolt maintenance method and system based on artificial intelligence and voiceprint
CN118782086A