Tower bolt maintenance method and system based on artificial intelligence and voiceprint
Through the tower bolt maintenance method based on artificial intelligence and voiceprint, the voiceprint analysis model is used to automatically identify the bolt status and obtain emergency strategies, which solves the problem of traditional manual inspection being time-consuming and inaccurate, and realizes efficient and accurate bolt detection and maintenance.
Patent Information
- Application Number
- CN202410904274.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2024-07-08
- Publication Date
- 2025-09-30
- Estimated Expiration
- 2044-07-08
AI Technical Summary
Traditional tower bolt maintenance methods rely on manual inspections, which are time-consuming and error-prone. Existing AI-based detection methods lack accuracy and cannot effectively guarantee the safety and stability of tower bolts.
A tower bolt maintenance method based on artificial intelligence and voiceprints is adopted. By collecting tower detection audio, removing interference and then using a pre-trained voiceprint analysis model to perform voiceprint analysis, the bolt status is automatically identified and emergency strategies are obtained based on abnormal results.
It improves the accuracy and efficiency of tower bolt detection, realizes automated maintenance, and reduces the cost and risk of manual inspection.
Smart Images

Figure CN118782086B_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a tower bolt maintenance method and system based on artificial intelligence and voiceprint. Background Art
[0002] Power towers are a crucial piece of infrastructure for power transmission lines. Typically constructed of steel, they are elongated, tower-like structures with high wind resistance and stability. They are primarily used to support transmission lines. With the development of the communications and power industries, the safety and stability of these towers, as critical infrastructure, are paramount. However, the bolts on these towers can become loose or damaged due to various factors, such as weather, age, and vandalism, posing safety risks. Traditional tower bolt maintenance relies primarily on manual inspections, a time-consuming and error-prone method.
[0003] Some related existing technologies disclose methods for detecting abnormalities in towers. For example, the invention patent with publication number CN117952960A discloses an artificial intelligence-based method for detecting defects in power tower components. However, it still processes images and its accuracy cannot be guaranteed. The invention patent with publication number CN116910667B discloses a method and system for analyzing abnormal conditions of communication towers based on a decision-making algorithm. The method collects historical tower environmental data and uses it to construct relevant decision-making models, which can effectively reduce the cost of analyzing abnormal conditions of communication towers. However, due to the limited amount of historical data, the accuracy of detection cannot be guaranteed. Summary of the Invention
[0004] Purpose of the invention: In order to overcome the shortcomings of the existing technology, the present invention provides a tower bolt maintenance method based on artificial intelligence and voiceprint, and the present invention also provides a tower bolt maintenance system based on artificial intelligence and voiceprint.
[0005] Technical solution: According to a first aspect of the present invention, a tower bolt maintenance method based on artificial intelligence and voiceprint is provided, comprising:
[0006] Collect tower inspection audio to be analyzed;
[0007] Perform interference removal on the tower detection audio to be analyzed, and extract the voiceprint to be analyzed from the voiceprint to be analyzed after interference removal;
[0008] Inputting the voiceprint to be analyzed into a pre-trained voiceprint analysis model to obtain a voiceprint analysis result corresponding to the voiceprint to be analyzed;
[0009] When the voiceprint analysis result is characterized as an abnormal result, a target emergency strategy is obtained from a preset tower bolt maintenance strategy library according to the emergency strategy association relationship pre-configured for the abnormal result.
[0010] Further, including:
[0011] The step of inputting the voiceprint to be analyzed into a pre-trained voiceprint analysis model to obtain a voiceprint analysis result corresponding to the voiceprint to be analyzed includes:
[0012] Performing feature extraction and feature conversion on the voiceprint to be analyzed using the voiceprint analysis model to obtain a vector representation of the voiceprint to be analyzed;
[0013] Performing type inference based on the vector representation to be analyzed to obtain a type inference result corresponding to the voiceprint to be analyzed, and performing sparse coding based on the type inference result to obtain an analysis type vector;
[0014] The voiceprint analysis model is used to perform transform coding processing on the vector representation to be analyzed in the voiceprint classification domain corresponding to the analysis type vector to obtain the voiceprint coding feature of the voiceprint to be analyzed, which is determined as the voiceprint feature to be analyzed; the voiceprint analysis model is a network model obtained by integrated learning of voiceprint type recognition and voiceprint feature extraction operations;
[0015] Using the voiceprint analysis model, performing voiceprint type identification and voiceprint feature extraction operations on each bolt reference voiceprint in the bolt voiceprint database to obtain the corresponding voiceprint type vector and comparative voiceprint feature of each bolt reference voiceprint;
[0016] Converting the voiceprint type vector corresponding to each bolt reference voiceprint to obtain a voiceprint type representation corresponding to the bolt reference voiceprint;
[0017] Determine, based on the voiceprint type representation corresponding to each bolt reference voiceprint, a reference voiceprint of the same type of bolt corresponding to the same voiceprint type representation, and determine the comparative voiceprint feature corresponding to the reference voiceprint of the same type of bolt as the comparative voiceprint feature corresponding to the same voiceprint type representation, thereby obtaining a primary mapping relationship between each voiceprint type representation and the comparative voiceprint feature in the voiceprint type representation;
[0018] For each comparative voiceprint feature in the comparative voiceprint features, determining the bolt reference voiceprint corresponding to the same comparative voiceprint feature according to the comparative voiceprint features corresponding to each bolt reference voiceprint, thereby obtaining a high-order mapping relationship between each comparative voiceprint feature in the comparative voiceprint features and the bolt reference voiceprint;
[0019] Determining the primary-order mapping relationship and the high-order mapping relationship as a preset type association relationship;
[0020] Performing conversion processing on the analysis type vector to obtain an analysis type representation;
[0021] Calculating the similarity between the analysis type representation and each voiceprint type representation, and determining the voiceprint type representation whose similarity meets a preset deviation degree condition as the pending voiceprint type representation;
[0022] According to the preliminary mapping relationship, the pending comparative voiceprint features corresponding to the pending voiceprint type representation are determined as the pending comparative voiceprint feature dataset; the preset type association relationship includes a mapping relationship between the voiceprint type representation and the comparative voiceprint features, and a mapping relationship between the comparative voiceprint features and the bolt reference voiceprints in the bolt voiceprint database;
[0023] Calculating the feature matching degree between each pending comparative voiceprint feature in the pending comparative voiceprint feature data set and the voiceprint feature to be analyzed;
[0024] Determine the pending comparison voiceprint feature whose feature matching degree meets the preset matching degree condition as the target comparison voiceprint feature, obtain the target comparison voiceprint feature dataset, and obtain the target bolt reference voiceprint dataset corresponding to the target comparison voiceprint feature dataset;
[0025] A voiceprint analysis result corresponding to the voiceprint to be analyzed is obtained according to the target bolt reference voiceprint dataset.
[0026] Further, including:
[0027] The interference removal processing of the tower detection audio to be analyzed includes:
[0028] Acquire target monitoring audio data of the tower detection audio to be analyzed, wherein the target monitoring audio data includes a segment of the tower detection audio to be analyzed;
[0029] Determining a target audio segment in the target monitoring audio data, wherein the target audio segment is a segment containing interfering noise, and the interfering noise is invalid data;
[0030] Determining an initial time node of the interfering noise in each target audio segment;
[0031] Determining an accurate time node in each of the initial time nodes according to segment sequence information of each of the initial time nodes in all the target audio segments in the segment to which it belongs;
[0032] For each target audio segment, a sequence identifier of the interference noise in the target audio segment is determined based on the accurate time node of the target audio segment, and interference removal processing is performed based on the sequence identifier.
[0033] Further, including:
[0034] The method further comprises:
[0035] Using the voiceprint analysis model, a deep learning feature extraction operation is performed on each bolt reference voiceprint in the bolt voiceprint database to obtain a deep reference voiceprint feature corresponding to each bolt reference voiceprint;
[0036] Performing a deep learning feature extraction operation on the voiceprint to be analyzed to obtain a deep voiceprint feature to be analyzed corresponding to the voiceprint to be analyzed;
[0037] Based on the target bolt reference voiceprint dataset, the feature similarity between the deep voiceprint feature to be analyzed and the deep reference voiceprint feature corresponding to each target bolt reference voiceprint is calculated;
[0038] According to the sorting of the feature similarities based on the error values, the target bolt reference voiceprint with the smallest error value is extracted and determined as the voiceprint analysis result.
[0039] Further, including:
[0040] Before performing feature extraction and feature conversion operations on the voiceprint to be analyzed by the voiceprint analysis model to obtain a vector representation to be analyzed of the voiceprint to be analyzed, the method further includes:
[0041] Obtaining a similar voiceprint group data set; each similar voiceprint group data set includes at least one similar voiceprint group; each similar voiceprint group includes sample voiceprints with the same preset target value;
[0042] Performing feature extraction and feature conversion operations on each sample voiceprint in each similar voiceprint group data set using an initial voiceprint analysis model to obtain a sample vector representation of each sample voiceprint;
[0043] Perform type inference and sparse coding based on the sample vector representation to obtain a sample type vector corresponding to each sample voiceprint;
[0044] Based on the sample type vector and the preset target value, obtaining the category error corresponding to each similar voiceprint group data set;
[0045] By using the initial voiceprint analysis model, transform coding is performed on the sample vector representation in the voiceprint classification domain corresponding to the sample type vector to obtain the classification domain coding features corresponding to each sample voiceprint;
[0046] For each group of similar voiceprints, in the similar voiceprint groups with the same preset target value in the data sets of the similar voiceprint groups, sample training voiceprint extraction is performed according to the classification domain coding features corresponding to each sample voiceprint, and a training voiceprint array corresponding to each group of similar voiceprints is obtained, thereby obtaining a training voiceprint array set corresponding to each data set of the similar voiceprint groups;
[0047] According to the classification domain coding features corresponding to the sample voiceprints, sample classification domain feature similarity calculation and transform coding error calculation are performed on the training voiceprint arrays to obtain the classification domain coding errors corresponding to the data sets of the similar voiceprint groups;
[0048] performing training data difference calculation on each training voiceprint array in the training voiceprint array set according to the sample vector representation of each sample voiceprint, and obtaining the matching characteristic error corresponding to each similar voiceprint group data set;
[0049] Obtaining a final error based on the category error, the classification domain encoding error, and the matching feature error;
[0050] Based on the final error, the model parameters of the initial voiceprint analysis model are tuned to obtain the trained voiceprint analysis model.
[0051] Further, including:
[0052] The sample voiceprints in each group of similar voiceprints include a reference voiceprint and a target sample voiceprint; in the similar voiceprint groups with the same preset target value in the data sets of the similar voiceprint groups, sample training voiceprint extraction is performed according to the classification domain coding features corresponding to each sample voiceprint, and the training voiceprint array corresponding to each group of similar voiceprints is obtained, including:
[0053] Extracting sample voiceprints having the same target value as the preset target value of each group of similar voiceprints from the data sets of the similar voiceprint groups to obtain a set of voiceprints of the same type;
[0054] Calculating the voiceprint matching degree between each voiceprint of the same type in the voiceprint set and a reference voiceprint in each group of the same type of voiceprints based on the classification domain coding features corresponding to each sample voiceprint, and determining the non-target sample voiceprint corresponding to the reference voiceprint based on the voiceprint matching degree and a preset non-target sample determination criterion;
[0055] Each non-target sample voiceprint in the non-target sample voiceprint is integrated with the reference voiceprint and the target sample voiceprint to obtain a training voiceprint array corresponding to each group of similar voiceprints.
[0056] Further, including:
[0057] The method of performing sample classification domain feature similarity calculation and transform coding error calculation on each training voiceprint array based on the classification domain coding features corresponding to each sample voiceprint, and obtaining the classification domain coding errors corresponding to each similar voiceprint group data set, includes:
[0058] According to the classification domain coding features corresponding to the sample voiceprints, obtaining the reference sample classification domain coding features corresponding to the reference voiceprints, the target sample classification domain coding features corresponding to the target sample voiceprints, and the non-target sample classification domain coding features corresponding to the non-target sample voiceprints in the training voiceprint arrays;
[0059] Comparing the similarity between the classification domain coding feature of the reference sample and the classification domain coding feature of the target sample to obtain a first similarity coefficient;
[0060] Calculating the feature similarity between the classification domain coding feature of the reference sample and the classification domain coding feature of the non-target sample to obtain a second similarity coefficient;
[0061] Obtaining a similarity coefficient error according to an absolute difference between the first similarity coefficient and the second similarity coefficient;
[0062] Generate the target type representations corresponding to the reference sample classification domain coding feature, the target sample classification domain coding feature, and the non-target sample classification domain coding feature by using a preset symbol coding function, and calculate the root mean square error between the reference sample classification domain coding feature, the target sample classification domain coding feature, the non-target sample classification domain coding feature, and their corresponding target type representations to obtain a prediction error;
[0063] The similarity coefficient error and the prediction error are linearly superimposed to obtain the classification domain coding error.
[0064] Further, including:
[0065] The method further comprises:
[0066] Extracting sample training voiceprints from the similar voiceprint groups with different preset target values in the respective similar voiceprint group data sets to obtain the overall sample training voiceprint data sets corresponding to the respective similar voiceprint group data sets;
[0067] For each overall sample training voiceprint in the overall sample training voiceprint dataset, performing feature fusion on the sample type vector corresponding to each sample voiceprint in the overall sample training voiceprint and the classification domain coding feature to obtain a fusion feature;
[0068] Based on the fusion features of each sample voiceprint in the overall sample training voiceprints, sample classification domain feature similarity calculation is performed to obtain the corresponding fusion feature errors of each similar voiceprint group data set.
[0069] Further, including:
[0070] The obtaining of a final error based on the category error, the classification domain encoding error, and the matching feature error includes:
[0071] The category error, the classification domain encoding error, the fusion feature error and the matching feature error are linearly superimposed to obtain the final error.
[0072] In a second aspect, an embodiment of the present invention provides a readable storage medium, wherein the readable storage medium includes a computer program, and when the computer program is executed, the computer device where the readable storage medium is located is controlled to execute the method described in the first aspect.
[0073] Beneficial effects: Compared with the prior art, the present invention has the following advantages:
[0074] The disclosed method and system for tower bolt maintenance based on artificial intelligence and voiceprints collects tower detection audio and performs interference removal to extract clear voiceprints. These voiceprints are then analyzed using a pre-trained voiceprint analysis model. If the analysis results indicate an anomaly, the system automatically selects the corresponding emergency strategy from a library of tower bolt maintenance strategies based on pre-set emergency strategy associations. This design not only improves the accuracy and efficiency of tower bolt detection but also enables automated maintenance, significantly reducing the cost and risk of manual inspections. BRIEF DESCRIPTION OF THE DRAWINGS
[0075] To more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly describes the drawings required for use in the embodiments. It should be understood that the following drawings illustrate only certain embodiments of the present invention and should not be construed as limiting the scope of the present invention. Those skilled in the art can, without inventive effort, derive other relevant drawings from these drawings.
[0076] Figure 1 A schematic flow chart of the steps of a tower bolt maintenance method based on artificial intelligence and voiceprint provided by an embodiment of the present invention;
[0077] Figure 2 A schematic diagram of the structure of a computer device provided in an embodiment of the present invention. DETAILED DESCRIPTION
[0078] The following will clearly and completely describe the technical solutions in the embodiments of the present invention in conjunction with the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of the present invention and not all of the embodiments. All other embodiments obtained by ordinary technicians in this field based on the embodiments of the present invention without making any creative efforts shall fall within the scope of protection of the present invention.
[0079] The specific embodiments of the present invention are described in detail below with reference to the accompanying drawings.
[0080] In order to solve the technical problems in the above background technology, Figure 1 This is a flow chart of a tower bolt maintenance method based on artificial intelligence and voiceprint provided in an embodiment of the present disclosure. The tower bolt maintenance method based on artificial intelligence and voiceprint is introduced in detail below.
[0081] Step S201, collecting the tower detection audio to be analyzed;
[0082] Step S202, performing interference removal processing on the tower detection audio to be analyzed, and extracting the voiceprint to be analyzed from the voiceprint to be analyzed after interference removal;
[0083] Step S203: inputting the voiceprint to be analyzed into a pre-trained voiceprint analysis model to obtain a voiceprint analysis result corresponding to the voiceprint to be analyzed;
[0084] Step S204: when the voiceprint analysis result is characterized as an abnormal result, a target emergency strategy is obtained from a preset tower bolt maintenance strategy library according to the emergency strategy association relationship pre-configured for the abnormal result.
[0085] In an embodiment of the present invention, for example, a server receives a real-time audio stream from audio capture devices installed on a tower. These audio capture devices are placed at key locations on the tower, particularly at bolt joints, to capture sounds generated by loose bolts or other abnormalities. For example, highly sensitive microphones are installed on a communication tower in a certain city. These microphones continuously capture ambient sound and transmit the audio data in real time to the server for analysis. After receiving the audio data, the server first runs a denoising algorithm to eliminate non-critical sounds such as wind and ambient noise, while highlighting critical sounds such as loose bolts. For example, the server can run a denoising algorithm based on spectral subtraction, which can identify and eliminate both steady-state and transient noise in the audio. After denoising, the server further extracts key voiceprint features from the audio, such as frequency and amplitude, for subsequent analysis. The server inputs the extracted voiceprint features into a pre-trained voiceprint analysis model. This model is trained using a machine learning algorithm (such as deep learning) on a large number of labeled bolt sound samples, and can accurately identify abnormalities such as loose bolts. For example, when the server inputs a denoised voiceprint feature into the model, the model will output a probability distribution representing the status of the bolt, such as "normal", "slightly loose" or "severely loose". If the result output by the voiceprint analysis model shows that the bolt status is abnormal (such as "slightly loose" or "severely loose"), the server will obtain the corresponding emergency strategy from the preset tower bolt maintenance strategy library based on this abnormal result. For example, if the model output is "slightly loose", the server will select a preventive maintenance strategy, such as increasing the inspection frequency; if the model output is "severely loose", the server will trigger an emergency maintenance strategy, such as immediately dispatching maintenance personnel to the site for tightening operations. These strategies are pre-configured based on historical data and expert knowledge to ensure the safe operation of the tower.
[0086] In the embodiment of the present invention, the aforementioned step S203 can be implemented through the following example.
[0087] Using the voiceprint analysis model, the voiceprint to be analyzed is identified as a type of voiceprint, and an analysis type vector corresponding to the voiceprint to be analyzed is obtained;
[0088] In the voiceprint classification domain corresponding to the analysis type vector, a voiceprint feature extraction operation is performed on the voiceprint to be analyzed to obtain a voiceprint feature to be analyzed corresponding to the voiceprint to be analyzed; the voiceprint analysis model is a network model obtained by integrated learning of voiceprint type recognition and voiceprint feature extraction operation processing; based on the degree of deviation of the voiceprint type representation in the analysis type vector and the preset type association relationship, a pending voiceprint type representation is determined, and a pending comparative voiceprint feature data set corresponding to the pending voiceprint type representation is determined; the preset type association relationship includes a mapping relationship between the voiceprint type representation and the comparative voiceprint feature, and a mapping relationship between the comparative voiceprint feature and the bolt reference voiceprint in the bolt voiceprint database;
[0089] Determine, in the pending comparative voiceprint feature dataset, a target comparative voiceprint feature dataset whose matching degree meets a preset matching requirement with the voiceprint feature to be analyzed, and obtain a target bolt reference voiceprint dataset corresponding to the target comparative voiceprint feature dataset;
[0090] A voiceprint analysis result corresponding to the voiceprint to be analyzed is obtained according to the target bolt reference voiceprint dataset.
[0091] In an embodiment of the present invention, upon receiving a voiceprint to be analyzed, the server will first identify the voiceprint type using a trained voiceprint analysis model. This model can identify different voiceprint types, such as wind, machinery, and human voice, and generates an analysis type vector for each voiceprint type. For example, when the server receives audio data from a tower, the model will identify different types of soundprints, such as loose bolts, wind, and human voice, and generate a unique vector representation for each type. After determining the voiceprint type, the server will perform feature extraction in the corresponding voiceprint classification domain. This process primarily extracts key features that represent the voiceprint type, such as frequency, amplitude, and waveform. For example, for the sound of loose bolts, the server will extract the unique waveform variations and frequency characteristics of this sound to form a feature set for the voiceprint to be analyzed. Based on the association between the analysis type vector and pre-set types, the server will determine a candidate voiceprint type representation. This representation is used to further narrow the possible voiceprints and improve recognition accuracy. Simultaneously, based on this candidate voiceprint type representation, the server will retrieve a candidate comparison voiceprint feature dataset from the bolt voiceprint database. These data sets are used to compare with the voiceprint features to be analyzed.
[0092] The server compares the voiceprint features to be analyzed with the pending comparison voiceprint feature datasets one by one, identifying a target comparison voiceprint feature dataset that meets the preset matching requirements. This matching process is based on metrics such as cosine similarity and Euclidean distance. Once a matching target comparison voiceprint feature dataset is found, the server further obtains the bolt reference voiceprint dataset corresponding to these feature datasets.
[0093] Finally, the server will provide the voiceprint analysis results of the target bolt based on the reference voiceprint dataset. This result is a specific description of the bolt status (such as "bolt loose", "bolt normal", etc.), as well as a score of the bolt health index. This analysis result will be sent to the relevant maintenance personnel or system so that they can take necessary maintenance measures in a timely manner. For example, if the analysis results show that the bolt is loose, the maintenance personnel will receive an alert prompting them to go to the site for tightening operations.
[0094] In order to more clearly describe the solution provided by the embodiment of the present application, a more detailed description is given below.
[0095] In an embodiment of the present invention, the voiceprint analysis model is exemplarily a model trained using machine learning or deep learning techniques to analyze and identify specific patterns or features in audio data. In the scenario of tower bolt maintenance, this model is trained to identify sound features related to the bolt status. Assume that a deep learning network (such as a convolutional neural network (CNN) or a recurrent neural network (RNN)) is used to build this model. During the model training phase, the model is presented with a large amount of labeled audio data. These labels indicate the bolt status (e.g., normal, loose, etc.) in the audio. By learning the features of this audio data, the model gradually becomes able to independently identify the bolt status in new, unlabeled audio. The voiceprint to be analyzed refers to audio signals or features extracted from the tower inspection audio that have not yet been analyzed by the model. These signals or features may contain information about the bolt status. A microphone installed on the tower captures an audio segment that may contain various sounds such as wind, birdsong, and loose bolts.
[0096] Before voiceprint analysis begins, the raw, unprocessed audio signal is the voiceprint to be analyzed. Voiceprint type identification is the process of classifying audio data using a voiceprint analysis model. The goal is to identify different sound types present in the audio, such as loose bolts, wind, and human voices. When the voiceprint to be analyzed is input into the voiceprint analysis model, the model analyzes it and attempts to identify the different sound types. For example, the model may identify that a particular audio segment contains the sounds of loose bolts, wind, and distant human voices.
[0097] An analysis type vector is a data representation of the output of a voiceprint analysis model. It is typically a numerical vector that quantitatively describes the characteristics and attributes of different sound types identified in the audio. This vector can contain information about the sound type, intensity, duration, and other aspects. After the model completes voiceprint type identification, it generates an analysis type vector. This vector is a multidimensional array of numbers, with each dimension representing a different sound feature or attribute. For example, one dimension might represent the intensity of the sound of a loose bolt, while another dimension represents the duration of that sound. This vector allows for a more precise understanding of the specific characteristics of each sound in the audio.
[0098] Furthermore, a voiceprint classification domain refers to a specific processing space or category defined during voiceprint analysis based on soundprint type (e.g., loose bolts, wind noise, etc.). Within this domain, more refined feature extraction and classification can be performed on the soundprint being analyzed. For example, in the soundprint analysis of tower bolt maintenance, there is a dedicated voiceprint classification domain for "loose bolts." When the model identifies the soundprint type as "loose bolts," subsequent feature extraction and comparison operations will be performed within this specific voiceprint classification domain. Voiceprint feature extraction is the process of extracting key features from an audio signal that represent the audio signal. These features can include frequency, amplitude, waveform, and other features, which are used for subsequent voiceprint comparison and recognition. In the soundprint analysis of tower bolts, voiceprint feature extraction can include extracting specific frequency components, duration, and energy distribution from the sound of loose bolts. These features are used to distinguish between normal and abnormal bolt states. In voiceprint analysis, ensemble learning can be used to simultaneously identify voiceprint types and extract features.
[0099] Assume that the voiceprint analysis model in this embodiment of the present invention consists of two sub-models: one responsible for voiceprint type identification and the other for voiceprint feature extraction. Through ensemble learning, these two sub-models can work together, making the entire voiceprint analysis process more accurate and efficient. The association between the analysis type vector output by the model and a predefined type refers to the relationship between the analysis type vector output by the model and a predefined sound type. The predefined type association sets a standard or range for each sound type, which is used for comparison with the actual analysis type vector. If there is a predefined type association that defines certain conditions that the analysis type vector for "loose bolt sound" should meet (such as the presence of a specific frequency range), then when the model outputs an analysis type vector, this association can be used to determine whether it belongs to the "loose bolt sound" category. The degree of deviation of the voiceprint type representation refers to the degree of difference between the analysis type vector output by the model and the standard or range defined in the predefined type association. This degree of deviation is used to determine the accuracy of the pending voiceprint type. If the analysis type vector output by the model deviates slightly from the predefined type for "loose bolt sound," the pending voiceprint type can be considered to be "loose bolt sound." The pending voiceprint type representation is a voice type preliminarily determined based on the degree of deviation; and the pending comparative voiceprint feature dataset is a set of voice features associated with this preliminarily determined voice type and used for further comparison and confirmation.
[0100] Suppose the model initially identifies a certain audio segment as a "loose bolt sound." The associated pending comparison voiceprint feature dataset may contain multiple known, typical features of loose bolt sounds. This data will be compared with the features of the voiceprint to be analyzed to confirm the accuracy of the initial judgment. Matching is a quantitative metric used to measure the degree of similarity between the voiceprint feature to be analyzed and a feature in the pending comparison voiceprint feature dataset. Typically, matching is calculated by calculating the similarity or distance between two voiceprint features. If cosine similarity is used as the matching metric, the matching value ranges from -1 to 1, with values closer to 1 indicating greater similarity between the two voiceprint features. The preset matching requirement is a preset threshold or standard used to determine whether the voiceprint feature to be analyzed is sufficiently similar to the features in the pending comparison voiceprint feature dataset. A matching feature is considered found only when the matching degree meets or exceeds this preset matching requirement.
[0101] Assuming the preset matching requirement is a cosine similarity greater than or equal to 0.9, then the two features are considered a match only if the cosine similarity between the voiceprint feature to be analyzed and a comparison voiceprint feature is greater than or equal to 0.9. The target comparison voiceprint feature dataset is the set of voiceprint features in the pending comparison voiceprint feature dataset that meet the preset matching requirement for the voiceprint feature to be analyzed. These features will be used for further analysis and recognition. If the matching degree of three features in the pending comparison voiceprint feature dataset all meet the preset matching requirement, then these three features constitute the target comparison voiceprint feature dataset.
[0102] The target bolt reference voiceprint dataset is associated with the target comparison voiceprint feature dataset and contains detailed voiceprint data for bolts corresponding to the target comparison voiceprint features. This data is typically collected from actual bolts and serves as a reference standard to accurately identify the bolt's condition. Assuming that the target comparison voiceprint feature dataset in this embodiment of the present invention contains a voiceprint feature corresponding to a loose bolt, the target bolt reference voiceprint dataset should also contain detailed voiceprint data for this loose bolt, such as waveform and frequency distribution. This data will be used for more detailed comparison and identification with the voiceprint to be analyzed to accurately determine the bolt's condition. The voiceprint analysis result is the conclusion or result obtained after comparing and analyzing the voiceprint to be analyzed. It typically indicates the category or status of the voiceprint to be analyzed, as well as possible other relevant information. In the context of bolt condition detection, the voiceprint analysis result is a clear judgment, such as "bolt loose" or "bolt normal," but it can also include more detailed information, such as the degree of looseness and recommended maintenance measures.
[0103] In the embodiment of the present invention, the aforementioned step S202 can be implemented through the following example.
[0104] Acquire target monitoring audio data of the tower detection audio to be analyzed, wherein the target monitoring audio data includes a segment of the tower detection audio to be analyzed;
[0105] Determining a target audio segment in the target monitoring audio data, wherein the target audio segment is a segment containing interfering noise, and the interfering noise is invalid data;
[0106] Determining an initial time node of the interfering noise in each target audio segment;
[0107] Determining an accurate time node in each of the initial time nodes according to segment sequence information of each of the initial time nodes in all the target audio segments in the segment to which it belongs;
[0108] For each target audio segment, a sequence identifier of the interference noise in the target audio segment is determined based on the accurate time node of the target audio segment, and interference removal processing is performed based on the sequence identifier.
[0109] In an embodiment of the present invention, for example, a server receives a segment of tower monitoring audio from an audio acquisition device. This segment records the sound conditions of various components (including bolts) on the tower. The server first stores this segment as target monitoring audio data for subsequent analysis and processing. This segment contains multiple segments, including the sound of bolts working normally, as well as interference noise such as wind noise and current noise. The server processes the target monitoring audio data using audio analysis software. Through techniques such as spectrum analysis and waveform recognition, it identifies segments containing interference noise, namely target audio segments. For example, the server may detect an abnormal audio waveform in a segment, with a frequency distribution that clearly differs from the sound of normal bolt operation, and thus determine that segment as a target audio segment containing interference noise. For each target audio segment identified as containing interference noise, the server further analyzes the audio signal within this segment and, using a specific algorithm (such as an energy-based noise detection algorithm), determines the time when the interference noise begins to appear, namely the initial time node. For example, in a target audio segment, the server detects that the energy of the audio signal suddenly increases and the frequency distribution becomes complex starting at the 3rd second, thus determining the 3rd second as the initial time node of the interference noise in this segment.
[0110] Since the initial time node may only be based on the relative time within the segment, in order to more accurately locate the position of the interference noise in the entire audio, the server needs to convert the initial time node into an accurate time node based on the sequence information of each target audio segment in the overall audio. For example, if a target audio segment is the 5th segment of the entire audio, and each segment is 10 seconds long, then the exact time node of the interference noise that appears at the 3rd second of the segment in the entire audio is 43 seconds (that is, the total length of the first 4 segments, 40 seconds, plus the 3 seconds within the segment). The server performs interference removal processing on these segments based on the exact time node of the interference noise in each target audio segment. The processing methods may include filtering, noise reduction, audio repair, etc. For example, for interference noise with an exact time node of 43 seconds, the server will apply a bandpass filter before and after this time point to suppress the noise frequency component, or use audio repair technology to restore the original audio signal covered by the noise. After the processing is completed, the server will obtain a clean audio data with the interference noise removed, providing more accurate sound information for subsequent tower status detection.
[0111] In an embodiment of the present invention, before executing the step of determining the pending voiceprint type representation based on the degree of deviation between the analyzed type vector and the voiceprint type representation in the preset type association relationship, the embodiment of the present invention further provides the following implementation.
[0112] Using the voiceprint analysis model, performing voiceprint type identification and voiceprint feature extraction operations on each bolt reference voiceprint in the bolt voiceprint database to obtain a voiceprint type vector and comparative voiceprint features corresponding to each bolt reference voiceprint;
[0113] Converting the voiceprint type vector corresponding to each bolt reference voiceprint to obtain a voiceprint type representation corresponding to the bolt reference voiceprint;
[0114] According to the voiceprint type representations and comparative voiceprint features corresponding to the reference voiceprints of the bolts, a primary mapping relationship between each voiceprint type representation in the voiceprint type representation and the comparative voiceprint features is generated, as well as a high-order mapping relationship between each comparative voiceprint feature in the comparative voiceprint features and the bolt reference voiceprint is generated;
[0115] The primary-order mapping relationship and the high-order mapping relationship are determined as the preset type association relationship.
[0116] In an embodiment of the present invention, illustratively, the server first starts the voiceprint analysis model, which is a trained deep learning model specifically used to identify voiceprint types and extract voiceprint features. Then, the server retrieves the reference voiceprint data of each bolt from the bolt voiceprint database. These data are bolt sounds recorded under different conditions, and contain voiceprint information of various bolt states (such as normal, loose, damaged, etc.). The server inputs the bolt reference voiceprint data into the voiceprint analysis model, and the model automatically performs voiceprint type identification and voiceprint feature extraction. For example, for the voiceprint data of a loose bolt, the model will identify its type as "loose bolt voiceprint" and extract the unique features of the voiceprint, such as specific frequency distribution, amplitude changes, etc.
[0117] After processing the voiceprint analysis model, the server obtains a series of voiceprint type vectors corresponding to the bolt reference voiceprint. These vectors are the model's internal mathematical representation of the voiceprint type. The server needs to transform these voiceprint type vectors into more interpretable voiceprint type representations. This can be achieved through dimensionality reduction techniques (such as principal component analysis (PCA) or t-SNE), which map high-dimensional vector space to a lower-dimensional space while preserving the relative relationships between vectors. The transformed voiceprint type representation is easier for humans to understand and visualize. After obtaining the voiceprint type representation and comparative voiceprint features of the bolt reference voiceprint, the server begins to establish a mapping relationship between them. The primary mapping relationship describes the relationship between each voiceprint type representation and the corresponding comparative voiceprint feature. For example, the "loose bolt voiceprint" type representation may be associated with specific frequency fluctuations and amplitude reduction voiceprint features. Higher-order mapping relationships are more complex, considering the interactions and correlations between multiple comparative voiceprint features, as well as their overall relationship with the bolt reference voiceprint. This mapping relationship reveals the underlying connections and patterns between different voiceprint features.
[0118] Finally, the server integrates the primary and higher-order mappings to form pre-defined type associations. These associations provide an important reference for subsequent voiceprint recognition and analysis. When new bolt voiceprint data is input, the server can use these pre-defined type associations to quickly and accurately identify the voiceprint type and extract key features.
[0119] In an embodiment of the present invention, the aforementioned steps of generating a primary mapping relationship between each voiceprint type representation and the comparative voiceprint feature in the voiceprint type representation according to the voiceprint type representation and the comparative voiceprint feature corresponding to each bolt reference voiceprint, and a high-order mapping relationship between each comparative voiceprint feature in the comparative voiceprint feature and the bolt reference voiceprint, can be implemented through the following examples.
[0120] Determine, based on the voiceprint type representation corresponding to each bolt reference voiceprint, a reference voiceprint of the same type of bolt corresponding to the same voiceprint type representation, and determine the comparative voiceprint feature corresponding to the reference voiceprint of the same type of bolt as the comparative voiceprint feature corresponding to the same voiceprint type representation, thereby obtaining a primary mapping relationship between each voiceprint type representation and the comparative voiceprint feature in the voiceprint type representation;
[0121] For each comparative voiceprint feature in the comparative voiceprint features, the bolt reference voiceprint corresponding to the same comparative voiceprint feature is determined according to the comparative voiceprint features corresponding to each bolt reference voiceprint, thereby obtaining a high-order mapping relationship between each comparative voiceprint feature in the comparative voiceprint features and the bolt reference voiceprint.
[0122] In an embodiment of the present invention, illustratively, the server first classifies and organizes the obtained voiceprint type representations of each bolt reference voiceprint. For example, if the server finds that the voiceprint type representations of three bolt reference voiceprints are all marked as "loose type", then these three bolt reference voiceprints are determined to be similar bolt reference voiceprints. After determining the similar bolt reference voiceprints, the server further extracts the corresponding comparative voiceprint features of these similar bolt reference voiceprints. Taking the "loose type" voiceprint type representation as an example, the server will extract common comparative voiceprint features from these three similar bolt reference voiceprints, such as specific frequency fluctuations, amplitude changes, etc. These features are determined as the corresponding comparative voiceprint features of the "loose type" voiceprint type representation.
[0123] Through the above steps, the server can identify the corresponding comparative voiceprint features for each voiceprint type representation, thereby establishing a preliminary mapping relationship between the voiceprint type representation and the comparative voiceprint features. This mapping relationship helps the server quickly identify the voiceprint type of new voiceprint data with similar voiceprint features during subsequent analysis. Next, for each comparative voiceprint feature, the server finds all bolt reference voiceprints that share that feature. For example, for a specific frequency fluctuation, the server searches for all bolt reference voiceprints that contain that frequency fluctuation and groups them together. Through the classification step above, the server can identify certain comparative voiceprint features that co-occur in multiple bolt reference voiceprints, thereby establishing a higher-order mapping relationship between these comparative voiceprint features and the bolt reference voiceprints. This higher-order mapping relationship reveals the potential connections between different voiceprint features, helping the server gain a deeper understanding of the relationship between bolt status and voiceprint features. Through the above steps, the server establishes a preliminary mapping relationship from voiceprint type representation to comparative voiceprint features, as well as a higher-order mapping relationship from comparative voiceprint features to bolt reference voiceprints, providing strong support for subsequent voiceprint recognition and analysis.
[0124] In an embodiment of the present invention, the aforementioned steps of determining the pending voiceprint type representation based on the degree of deviation between the voiceprint type representation in the association relationship between the analysis type vector and the preset type, and determining the pending comparative voiceprint feature data set corresponding to the pending voiceprint type representation, can be implemented through the following examples.
[0125] Performing conversion processing on the analysis type vector to obtain an analysis type representation;
[0126] Calculating the similarity between the analysis type representation and each voiceprint type representation, and determining the voiceprint type representation whose similarity meets a preset deviation degree condition as the pending voiceprint type representation;
[0127] According to the primary-order mapping relationship, the pending comparative voiceprint features corresponding to the pending voiceprint type representation are determined to be the pending comparative voiceprint feature dataset.
[0128] In an embodiment of the present invention, for example, a server receives new bolt voiceprint data and first extracts an analysis type vector for the data using a voiceprint analysis model. This vector is the model's mathematical representation of the new voiceprint data. To facilitate subsequent processing, the server needs to transform this analysis type vector. For example, the server can use the same method used to convert the bolt reference voiceprint, such as dimensionality reduction techniques (PCA or t-SNE), to convert it into a more interpretable analysis type representation. This representation can be a low-dimensional vector or a well-understood label. After obtaining the analysis type representation, the server compares it with pre-set voiceprint type representations to determine the likely type of the new voiceprint data. To this end, the server calculates the similarity between the analysis type representation and each voiceprint type representation stored in the database. This similarity can be calculated using a variety of methods, such as cosine similarity and Euclidean distance. Taking cosine similarity as an example, the server calculates the cosine of the angle between the analysis type representation and each voiceprint type representation. The closer the value is to 1, the more similar the two are.
[0129] After calculating the similarity, the server sets a preset deviation condition to determine which preset voiceprint type the new voiceprint data most closely matches. For example, the server can set a similarity threshold; only when the similarity between the analysis type representation and a particular voiceprint type representation exceeds this threshold will the voiceprint type representation be identified as a pending voiceprint type representation. Suppose the server finds that the similarity between the analysis type representation and the "loose" voiceprint type representation reaches 0.95 (cosine similarity), while the similarity with other types is less than 0.8. If the threshold is set to 0.9, "loose" will be identified as the pending voiceprint type representation. After determining the pending voiceprint type representation, the server needs to find the corresponding comparative voiceprint features based on the primary mapping relationship. Taking "loose" as an example, the server will search for comparative voiceprint features associated with the "loose" voiceprint type representation. These features can include specific frequency fluctuations, amplitude changes, and other characteristics. These comparative voiceprint features associated with the pending voiceprint type representation constitute the pending comparative voiceprint feature dataset. This dataset will be used for further analysis and verification of new voiceprint data. For example, the server can use this dataset to train a new classifier to more accurately identify whether the newly collected bolt voiceprint data is indeed "loose".
[0130] In an embodiment of the present invention, the aforementioned step of determining a target comparative voiceprint feature dataset whose matching degree meets a preset matching requirement with the voiceprint feature to be analyzed in the pending comparative voiceprint feature dataset can be implemented through the following example.
[0131] Calculating the feature matching degree between each pending comparative voiceprint feature in the pending comparative voiceprint feature data set and the voiceprint feature to be analyzed;
[0132] The pending comparison voiceprint feature whose feature matching degree meets the preset matching degree condition is determined as the target comparison voiceprint feature, and the target comparison voiceprint feature data set is obtained.
[0133] In an embodiment of the present invention, illustratively, after determining the dataset of pending comparative voiceprint features, the server needs to calculate the degree of match between each pending comparative voiceprint feature in this dataset and the voiceprint feature to be analyzed. For example, suppose the server has determined a dataset containing multiple pending comparative voiceprint features, one of which is "frequency fluctuation range of 200-300Hz", and the voiceprint feature to be analyzed also contains a similar frequency fluctuation range of "210-290Hz". In order to quantify the degree of similarity between the two features, the server can use a specific algorithm to calculate their matching degree, such as calculating the proportion of the overlapping part of the two frequency fluctuation ranges to the total range. After calculating the matching degree between all pending comparative voiceprint features and the voiceprint feature to be analyzed, the server will set a preset matching degree condition to filter out the comparative voiceprint feature that best matches the target voiceprint feature. Taking the frequency fluctuation range mentioned above as an example, if the server sets a preset matching condition of "match greater than 0.8," then a pending comparison voiceprint feature will only be identified as a target comparison voiceprint feature if its matching degree with the voiceprint feature to be analyzed is greater than 0.8. By screening all pending comparison voiceprint features that meet the preset matching condition, the server ultimately obtains a target comparison voiceprint feature dataset. This dataset contains all comparison voiceprint features that most closely match the voiceprint feature to be analyzed and can be used for subsequent voiceprint recognition, classification, or other related tasks. For example, if the server ultimately determines three target comparison voiceprint features: "frequency fluctuation range of 200-300Hz," "amplitude variation pattern A," and "sound wave waveform feature B," these three features constitute the target comparison voiceprint feature dataset. This dataset can be used to further analyze the specific type and status of the voiceprint data to be analyzed, or to train a more accurate voiceprint recognition model.
[0134] In the embodiment of the present invention, the aforementioned steps of identifying the voiceprint type of the voiceprint to be analyzed by using the voiceprint analysis model and obtaining the analysis type vector corresponding to the voiceprint to be analyzed can be implemented through the following examples.
[0135] Performing feature extraction and feature conversion on the voiceprint to be analyzed using the voiceprint analysis model to obtain a vector representation of the voiceprint to be analyzed;
[0136] Performing type inference based on the vector representation to be analyzed to obtain a type inference result corresponding to the voiceprint to be analyzed, and performing sparse coding based on the type inference result to obtain the analysis type vector;
[0137] The step of performing a voiceprint feature extraction operation on the voiceprint to be analyzed in the voiceprint classification domain corresponding to the analysis type vector to obtain a voiceprint feature to be analyzed corresponding to the voiceprint to be analyzed includes:
[0138] The voiceprint analysis model is used to perform transform coding processing on the vector representation to be analyzed in the voiceprint classification domain corresponding to the analysis type vector to obtain the voiceprint coding feature of the voiceprint to be analyzed, which is determined as the voiceprint feature to be analyzed.
[0139] In an embodiment of the present invention, for example, a server receives voiceprint data of a bolt to be analyzed and first processes it using a pre-trained voiceprint analysis model. This model possesses feature extraction and feature conversion capabilities, extracting key information from the voiceprint and converting it into a mathematical vector. For example, the voiceprint analysis model uses a series of complex mathematical operations and signal processing techniques, such as Fast Fourier Transform (FFT) or Mel-Frequency Cepstral Coefficients (MFCC), to extract key features such as frequency and amplitude from the voiceprint to be analyzed. These features are then converted into a high-dimensional vector, which serves as the vector representation of the analyzed vector. After obtaining the vector representation of the analyzed vector, the server uses the model's type inference mechanism to determine the likely type of the voiceprint. Type inference can be based on machine learning methods, such as support vector machines (SVMs) or neural networks. For example, if the vector representation of the analyzed vector is highly similar to the feature vector of a "loose" bolt voiceprint, the type inference result is "loose." The server then performs sparse encoding on this type inference result, converting the type inference result into a sparse vector, which serves as the analysis type vector. In this vector, the values of elements corresponding to "loose" will be high, while the values of other elements will be low or zero. After determining the possible type of the voiceprint to be analyzed, the server further extracts features of the voiceprint to be analyzed within the corresponding voiceprint classification domain. During this process, the voiceprint analysis model performs more refined transform coding on the vector representation to be analyzed. For example, if the voiceprint to be analyzed is inferred to be "loose," the server will perform specific transform coding on the vector representation to be analyzed within the classification domain of the "loose" bolt voiceprint, such as principal component analysis (PCA) or linear discriminant analysis (LDA), to extract more discriminative voiceprint features. These features are called voiceprint encoding features, or the features of the voiceprint to be analyzed. After the above steps, the server ultimately obtains a voiceprint encoding feature vector containing key information about the voiceprint to be analyzed. This vector is the voiceprint feature to be analyzed. This feature vector can be used in subsequent tasks such as bolt condition identification and fault diagnosis. For example, the server can compare this voiceprint feature to known bolt fault voiceprint features to determine whether the bolt to be analyzed is faulty, as well as the type and severity of the fault.
[0140] In the embodiments of the present invention, the following implementation modes are also provided.
[0141] Using the voiceprint analysis model, a deep learning feature extraction operation is performed on each bolt reference voiceprint in the bolt voiceprint database to obtain a deep reference voiceprint feature corresponding to each bolt reference voiceprint;
[0142] Performing a deep learning feature extraction operation on the voiceprint to be analyzed to obtain a deep voiceprint feature to be analyzed corresponding to the voiceprint to be analyzed;
[0143] Based on the target bolt reference voiceprint dataset, the feature similarity between the deep voiceprint feature to be analyzed and the deep reference voiceprint feature corresponding to each target bolt reference voiceprint is calculated;
[0144] According to the sorting of the feature similarities based on the error values, the target bolt reference voiceprint with the smallest error value is extracted and determined as the voiceprint analysis result.
[0145] In an embodiment of the present invention, illustratively, the server first uses a pre-trained voiceprint analysis model to perform deep learning feature extraction on each bolt reference voiceprint in the bolt voiceprint database. During this process, the model will automatically learn and extract deep features in the bolt voiceprint, which are more discriminative and robust than traditional manual features. For example, for a specific "loose" bolt voiceprint, the deep learning model will extract features related to the specific frequency and amplitude changes generated when the bolt is loose. These features are encoded into a high-dimensional vector, namely the deep reference voiceprint feature. Similarly, the server also uses the voiceprint analysis model to perform deep learning feature extraction on the bolt voiceprint to be analyzed. This process is similar to step one, but the object of processing is the bolt voiceprint to be analyzed. For example, for a newly collected bolt voiceprint data, the server will use the model to extract its deep voiceprint feature to be analyzed, which is also a high-dimensional vector for subsequent similarity calculation.
[0146] After obtaining the target bolt reference voiceprint dataset (i.e., the set of bolt reference voiceprints most relevant to the voiceprint to be analyzed), the server calculates the similarity between the deep voiceprint features to be analyzed and the deep reference voiceprint features of each target bolt reference voiceprint. Similarity can be calculated using various methods, such as cosine similarity and Euclidean distance. Taking cosine similarity as an example, the server calculates the cosine value of the angle between the deep voiceprint features to be analyzed and each deep reference voiceprint feature. The closer the value is to 1, the more similar the two are. The server then sorts the calculated similarity values by error value (1 minus the similarity value). The smaller the error value, the more similar the voiceprint to be analyzed is to the corresponding bolt reference voiceprint. For example, if the deep reference voiceprint features of a "loose" bolt reference voiceprint have the highest similarity (i.e., the smallest error value) with the deep reference voiceprint features of the voiceprint to be analyzed, the server will determine this "loose" bolt reference voiceprint as the voiceprint analysis result. In this way, through deep learning feature extraction and similarity calculation, the server can accurately identify the type and status of the bolt voiceprint to be analyzed, providing strong support for subsequent bolt status monitoring and fault diagnosis.
[0147] In an embodiment of the present invention, before the aforementioned step of performing feature extraction and feature conversion operations on the voiceprint to be analyzed through the voiceprint analysis model to obtain the vector representation to be analyzed of the voiceprint to be analyzed, the embodiment of the present invention also provides the following implementation method.
[0148] Obtaining a similar voiceprint group data set; each similar voiceprint group data set includes at least one similar voiceprint group; each similar voiceprint group includes sample voiceprints with the same preset target value;
[0149] Performing feature extraction and feature conversion operations on each sample voiceprint in each similar voiceprint group data set using an initial voiceprint analysis model to obtain a sample vector representation of each sample voiceprint;
[0150] Perform type inference and sparse coding based on the sample vector representation to obtain a sample type vector corresponding to each sample voiceprint;
[0151] Based on the sample type vector and the preset target value, obtaining the category error corresponding to each similar voiceprint group data set;
[0152] By using the initial voiceprint analysis model, transform coding is performed on the sample vector representation in the voiceprint classification domain corresponding to the sample type vector to obtain the classification domain coding features corresponding to each sample voiceprint;
[0153] For each group of similar voiceprints, in the similar voiceprint groups with the same preset target value in the data sets of the similar voiceprint groups, sample training voiceprint extraction is performed according to the classification domain coding features corresponding to each sample voiceprint, and a training voiceprint array corresponding to each group of similar voiceprints is obtained, thereby obtaining a training voiceprint array set corresponding to each data set of the similar voiceprint groups;
[0154] According to the classification domain coding features corresponding to the sample voiceprints, sample classification domain feature similarity calculation and transform coding error calculation are performed on the training voiceprint arrays to obtain the classification domain coding errors corresponding to the data sets of the similar voiceprint groups;
[0155] performing training data difference calculation on each training voiceprint array in the training voiceprint array set according to the sample vector representation of each sample voiceprint, and obtaining the matching characteristic error corresponding to each similar voiceprint group data set;
[0156] Obtaining a final error based on the category error, the classification domain encoding error, and the matching feature error;
[0157] Based on the final error, the model parameters of the initial voiceprint analysis model are tuned to obtain the trained voiceprint analysis model.
[0158] In an embodiment of the present invention, the server illustratively obtains multiple similar voiceprint group data sets from a voiceprint database. For example, one of the similar voiceprint group data sets contains multiple groups of "normal" bolt voiceprints, each group containing several sample voiceprints with the same preset target value (i.e., all "normal"). The server uses the initial voiceprint analysis model to perform feature extraction and feature conversion on the sample voiceprints in each similar voiceprint group data set. For example, for the "normal" bolt voiceprint, the model extracts key features such as its frequency and amplitude and converts these features into vector form, i.e., a sample vector representation. Based on the sample vector representation, the server performs type inference and generates a sample type vector for each sample voiceprint. For example, in the sample type vector of the "normal" bolt voiceprint, the element value corresponding to "normal" will be higher. The server compares the sample type vector with the preset target value and calculates the error for each category. For example, if the sample type vector of a "normal" sample voiceprint differs significantly from the "normal" target value, then its category error will be higher.
[0159] Within the voiceprint classification domain corresponding to the sample type vector, the server uses the initial voiceprint analysis model to perform transform coding on the sample vector representation, extracting the classification domain encoding features for each sample voiceprint. These features are more discriminating and facilitate subsequent voiceprint classification and recognition. For each group of similar voiceprints, the server extracts sample training voiceprints based on the classification domain encoding features, obtaining a training voiceprint array for each group of similar voiceprints. For example, for the "normal" bolt voiceprint group, the server extracts a training voiceprint array that represents the voiceprint characteristics of that group.
[0160] Based on the classification domain encoding features, the server calculates classification domain feature similarity and transform encoding error for the training voiceprint arrays, obtaining the classification domain encoding error for each similar voiceprint group dataset. This error reflects the difference in classification domain encoding features between the training voiceprint arrays and the original sample voiceprints. The server also calculates the training data difference for each training voiceprint array in the training voiceprint array set based on the sample vector representation, obtaining the matching feature error for each similar voiceprint group dataset. This error reflects the degree of overall feature match between the training voiceprint arrays and the original sample voiceprints. The server combines the category error, classification domain encoding error, and matching feature error to obtain a final error. Based on this final error, the model parameters of the initial voiceprint analysis model are then fine-tuned to obtain a trained voiceprint analysis model. This fine-tuning process may include adjusting model parameters such as weights and biases to improve the model's accuracy and generalization capabilities.
[0161] In an embodiment of the present invention, the sample voiceprints in each group of similar voiceprints include a baseline voiceprint and a target sample voiceprint; the aforementioned steps of extracting sample training voiceprints according to the classification domain coding features corresponding to each sample voiceprint in the similar voiceprint group data set and obtaining the training voiceprint array corresponding to each group of similar voiceprints can be implemented through the following examples.
[0162] Extracting sample voiceprints having the same target value as the preset target value of each group of similar voiceprints from the data sets of the similar voiceprint groups to obtain a set of voiceprints of the same type;
[0163] Calculating the voiceprint matching degree between each voiceprint of the same type in the voiceprint set and a reference voiceprint in each group of the same type of voiceprints based on the classification domain coding features corresponding to each sample voiceprint, and determining the non-target sample voiceprint corresponding to the reference voiceprint based on the voiceprint matching degree and a preset non-target sample determination criterion;
[0164] Each non-target sample voiceprint in the non-target sample voiceprint is integrated with the reference voiceprint and the target sample voiceprint to obtain a training voiceprint array corresponding to each group of similar voiceprints.
[0165] In an embodiment of the present invention, illustratively, the server first extracts sample voiceprints that are the same as the preset target value of each group of similar voiceprints from the voiceprint database for each group of similar voiceprints. For example, for the similar voiceprint group of "loose type" bolt voiceprints, the server will filter out all sample voiceprints marked as "loose type" to form a "loose type" voiceprint set. Next, the server uses the classification domain coding features of each sample voiceprint that has been extracted to calculate the matching degree between each voiceprint in the same type of voiceprint set and the benchmark voiceprint in each group of similar voiceprints. The matching degree here can be obtained by calculating the similarity between the two voiceprint feature vectors, such as using methods such as cosine similarity. Taking the "loose type" bolt voiceprint as an example, the server will select a typical "loose type" voiceprint as the benchmark voiceprint, and then calculate the matching degree between each "loose type" voiceprint in the set and this benchmark voiceprint.
[0166] The server determines which voiceprints are non-target samples that do not match the baseline voiceprint based on the calculated voiceprint matching degree and the preset non-target sample judgment standard. This judgment standard can be a threshold. When the matching degree is lower than this threshold, the voiceprint is considered a non-target sample voiceprint. For example, if the matching degree of a "loose" voiceprint with the baseline voiceprint is lower than 0.8 (the assumed threshold), the server will judge it as a non-target sample voiceprint.
[0167] Finally, the server integrates each non-target sample voiceprint with the baseline voiceprint and the target sample voiceprint, forming a training voiceprint array corresponding to each similar voiceprint group. This integration process can be a simple merging process or a filtering and combination process based on specific rules. Taking the "loose" bolt voiceprint as an example, the server will ultimately obtain a training voiceprint array containing the baseline voiceprint, the target sample voiceprint, and multiple non-target sample voiceprints. This array will be used in the subsequent voiceprint analysis model training to improve the model's recognition accuracy for "loose" bolt voiceprints.
[0168] In an embodiment of the present invention, the aforementioned step of performing sample classification domain feature similarity calculation and transform coding error calculation on each training voiceprint array based on the classification domain coding features corresponding to each sample voiceprint, and obtaining the classification domain coding errors corresponding to each similar voiceprint group data set, can be implemented through the following examples.
[0169] According to the classification domain coding features corresponding to the sample voiceprints, obtaining the reference sample classification domain coding features corresponding to the reference voiceprints, the target sample classification domain coding features corresponding to the target sample voiceprints, and the non-target sample classification domain coding features corresponding to the non-target sample voiceprints in the training voiceprint arrays;
[0170] Comparing the similarity between the classification domain coding feature of the reference sample and the classification domain coding feature of the target sample to obtain a first similarity coefficient;
[0171] Calculating the feature similarity between the classification domain coding feature of the reference sample and the classification domain coding feature of the non-target sample to obtain a second similarity coefficient;
[0172] Obtaining a similarity coefficient error according to an absolute difference between the first similarity coefficient and the second similarity coefficient;
[0173] Generate the target type representations corresponding to the reference sample classification domain coding feature, the target sample classification domain coding feature, and the non-target sample classification domain coding feature by using a preset symbol coding function, and calculate the root mean square error between the reference sample classification domain coding feature, the target sample classification domain coding feature, the non-target sample classification domain coding feature, and their corresponding target type representations to obtain a prediction error;
[0174] The similarity coefficient error and the prediction error are linearly superimposed to obtain the classification domain coding error.
[0175] In an embodiment of the present invention, illustratively, the server first extracts the classification domain coding features of the baseline voiceprint, target sample voiceprint, and non-target sample voiceprint for each training voiceprint array. For example, in a training voiceprint array containing "normal", "loose" and "broken" bolt voiceprints, the server will extract the classification domain coding features of these three types of voiceprints respectively. The server will then compare the similarity between the classification domain coding features of the baseline sample voiceprint (such as the "normal" bolt voiceprint) and the classification domain coding features of the target sample voiceprint (such as the "loose" bolt voiceprint). This similarity is called the first similarity coefficient. This coefficient can be obtained by calculating the cosine similarity between the two feature vectors, etc.
[0176] Then, the server calculates the similarity between the classification domain coding features of the baseline sample voiceprint and the classification domain coding features of the non-target sample voiceprint (such as the "broken type" bolt voiceprint). This similarity is called the second similarity coefficient. The purpose of this step is to quantify the degree of similarity between the baseline voiceprint and other types of voiceprints. The server calculates their absolute difference based on the first similarity coefficient and the second similarity coefficient. This difference is called the similarity coefficient error. This error reflects the degree of difference in the similarity between the baseline voiceprint and the target voiceprint and the non-target voiceprint. Next, the server generates respective target type representations for the classification domain coding features of the baseline sample voiceprint, the target sample voiceprint, and the non-target sample voiceprint through a preset symbol encoding function.
[0177] The server then calculates the root mean square error between these classification domain encoding features and their corresponding target type representations. This error is called the prediction error. The prediction error reflects the accuracy of the model in predicting various voiceprint types.
[0178] Finally, the server linearly superimposes the similarity coefficient error and the prediction error to obtain the classification domain encoding error. This error comprehensively considers the similarity between voiceprints and the accuracy of the model's predictions, and is a key metric for evaluating the performance of voiceprint analysis models. By optimizing this error, the model's classification and recognition capabilities can be further improved.
[0179] In the embodiments of the present invention, the following implementation modes are also provided.
[0180] Extracting sample training voiceprints from the similar voiceprint groups with different preset target values in the respective similar voiceprint group data sets to obtain the overall sample training voiceprint data sets corresponding to the respective similar voiceprint group data sets;
[0181] For each overall sample training voiceprint in the overall sample training voiceprint dataset, performing feature fusion on the sample type vector corresponding to each sample voiceprint in the overall sample training voiceprint and the classification domain coding feature to obtain a fusion feature;
[0182] Based on the fusion features of each sample voiceprint in the overall sample training voiceprints, sample classification domain feature similarity calculation is performed to obtain the corresponding fusion feature errors of each similar voiceprint group data set.
[0183] In an embodiment of the present invention, illustratively, the server extracts not only sample voiceprints with the same preset target value as each group of similar voiceprints from each similar voiceprint group data set, but also sample voiceprints from similar voiceprint groups with different preset target values. For example, in addition to extracting "normal" bolt voiceprints, other types of bolt voiceprints such as "loose" and "broken" are also extracted. In this way, the server can obtain an overall sample training voiceprint data set containing multiple types of voiceprints. For each overall sample training voiceprint in the overall sample training voiceprint data set (for example, a training set containing "normal", "loose" and "broken" bolt voiceprints), the server performs feature fusion on the sample type vector of each sample voiceprint with the classification domain encoding feature.
[0184] Specifically, the server uses a feature fusion technique (such as concatenation or weighted summation) to combine the two features, resulting in a more comprehensive and rich fused feature. The server then uses the fused feature to calculate feature similarity in the sample classification domain. For example, the server can compare the similarity between the fused feature of a "normal" bolt voiceprint and the fused feature of a "loose" or "broken" bolt voiceprint. Through this comparison, the server can quantify the degree of difference between different voiceprint types. To evaluate the effectiveness of the fused feature, the server calculates the fused feature error. This error is calculated by comparing the consistency between the fused feature and the true class label. For example, if a fused feature is closer to the "loose" bolt voiceprint than its true "normal" label, it will have a larger fused feature error. Through this step, the server can evaluate the accuracy of its feature fusion and classification methods and make further optimizations and adjustments accordingly. A reduction in the fused feature error indicates that the server's classification and recognition capabilities are improving, and it can more accurately distinguish different voiceprint types.
[0185] In the embodiment of the present invention, the aforementioned step of obtaining the final error based on the category error, the classification domain encoding error and the matching feature error can be implemented through the following example.
[0186] The category error, the classification domain encoding error, the fusion feature error and the matching feature error are linearly superimposed to obtain the final error.
[0187] In an embodiment of the present invention, exemplarily, after the server completes the calculation of various errors, namely, the category error, the classification domain coding error, the fusion feature error, and the matching feature error, it will perform a linear superposition process to obtain the final error. This process is like adding the various error values together, but each error will be assigned a different weight according to its importance. For example, suppose the category error is 0.1, the classification domain coding error is 0.05, the fusion feature error is 0.08, and the matching feature error is 0.07. The server will assign different weights to these errors, such as a category error weight of 0.4, a classification domain coding error weight of 0.2, a fusion feature error weight of 0.2, and a matching feature error weight of 0.2. Then, the server will add these weighted errors to obtain the final error. The specific calculation is as follows:
[0188] Final error = 0.1*0.4+0.05*0.2+0.08*0.2+0.07*0.2=0.04+0.01+0.016+0.014=0.08. This final error is a key metric used by the server to evaluate the performance of its voiceprint recognition system. By continuously optimizing model parameters and algorithms, the server can reduce this error and improve voiceprint recognition accuracy. In practice, the server will regularly recalculate this final error to ensure that the system's performance remains optimal.
[0189] An embodiment of the present invention provides a tower bolt maintenance system based on artificial intelligence and voiceprint, comprising:
[0190] The acquisition module is used to collect the tower detection audio to be analyzed;
[0191] A module is proposed for performing interference removal on the tower detection audio to be analyzed and extracting the voiceprint to be analyzed from the voiceprint to be analyzed after interference removal;
[0192] An analysis module, configured to input the voiceprint to be analyzed into a pre-trained voiceprint analysis model to obtain a voiceprint analysis result corresponding to the voiceprint to be analyzed;
[0193] The processing module is used to obtain a target emergency strategy from a preset tower bolt maintenance strategy library according to the emergency strategy association relationship pre-configured for the abnormal result when the voiceprint analysis result is characterized as an abnormal result.
[0194] The embodiment of the present invention provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned tower bolt maintenance method based on artificial intelligence and voiceprint. Figure 2 As shown, Figure 2This is a block diagram of the structure of a computer device 100 provided in an embodiment of the present invention. Computer device 100 includes a memory 111, a processor 112, and a communication unit 113. To enable data transmission or exchange, memory 111, processor 112, and communication unit 113 are electrically connected to each other, directly or indirectly. For example, these components can be electrically connected via one or more communication buses or signal lines.
[0195] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the above illustrative discussion is not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Numerous modifications and variations are possible in light of the above teachings. These embodiments have been selected and described in order to best illustrate the principles of the present disclosure and its practical application, thereby enabling those skilled in the art to best utilize the present disclosure and to utilize various embodiments with various modifications as appropriate for the specific application contemplated.
[0196] It will be understood by those skilled in the art that embodiments of the present invention may be provided as methods, systems, or computer program products. Thus, the present invention may take the form of an entirely hardware embodiment, an entirely software embodiment, or an embodiment combining software and hardware. Furthermore, the present invention may take the form of a computer program product implemented on one or more computer-usable storage media (including but not limited to magnetic disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.
[0197] Although the preferred embodiments of the present invention have been described, those skilled in the art may make additional changes and modifications to these embodiments once they have learned the basic creative concept. Therefore, the appended claims are intended to be interpreted as including the preferred embodiments and all changes and modifications that fall within the scope of the present invention.
[0198] Obviously, those skilled in the art may make various changes and modifications to the embodiments of the present invention without departing from the spirit and scope of the embodiments of the present invention. Thus, if such changes and modifications of the embodiments of the present invention fall within the scope of the claims and their equivalents, the present invention is intended to include such changes and modifications.
Claims
1. A tower bolt maintenance method based on artificial intelligence and voiceprint, characterized in that: include: Collect tower inspection audio to be analyzed; Perform interference removal on the tower detection audio to be analyzed, and extract the voiceprint to be analyzed from the voiceprint to be analyzed after interference removal; Inputting the voiceprint to be analyzed into a pre-trained voiceprint analysis model to obtain a voiceprint analysis result corresponding to the voiceprint to be analyzed; When the voiceprint analysis result is characterized as an abnormal result, a target emergency strategy is obtained from a preset tower bolt maintenance strategy library according to the emergency strategy association relationship pre-configured for the abnormal result; The step of inputting the voiceprint to be analyzed into a pre-trained voiceprint analysis model to obtain a voiceprint analysis result corresponding to the voiceprint to be analyzed includes: Performing feature extraction and feature conversion on the voiceprint to be analyzed using the voiceprint analysis model to obtain a vector representation of the voiceprint to be analyzed; Performing type inference based on the vector representation to be analyzed to obtain a type inference result corresponding to the voiceprint to be analyzed, and performing sparse coding based on the type inference result to obtain an analysis type vector; The voiceprint analysis model is used to perform transform coding processing on the vector representation to be analyzed in the voiceprint classification domain corresponding to the analysis type vector to obtain the voiceprint coding feature of the voiceprint to be analyzed, which is determined as the voiceprint feature to be analyzed; the voiceprint analysis model is a network model obtained by integrated learning of voiceprint type recognition and voiceprint feature extraction operations; Using the voiceprint analysis model, performing voiceprint type identification and voiceprint feature extraction operations on each bolt reference voiceprint in the bolt voiceprint database to obtain the corresponding voiceprint type vector and comparative voiceprint feature of each bolt reference voiceprint; Converting the voiceprint type vector corresponding to each bolt reference voiceprint to obtain a voiceprint type representation corresponding to the bolt reference voiceprint; Determine, based on the voiceprint type representation corresponding to each bolt reference voiceprint, a reference voiceprint of the same type of bolt corresponding to the same voiceprint type representation, and determine the comparative voiceprint feature corresponding to the reference voiceprint of the same type of bolt as the comparative voiceprint feature corresponding to the same voiceprint type representation, thereby obtaining a primary mapping relationship between each voiceprint type representation and the comparative voiceprint feature in the voiceprint type representation; For each comparative voiceprint feature in the comparative voiceprint features, determining the bolt reference voiceprint corresponding to the same comparative voiceprint feature according to the comparative voiceprint features corresponding to each bolt reference voiceprint, thereby obtaining a high-order mapping relationship between each comparative voiceprint feature in the comparative voiceprint features and the bolt reference voiceprint; Determining the primary-order mapping relationship and the high-order mapping relationship as a preset type association relationship; Performing conversion processing on the analysis type vector to obtain an analysis type representation; Calculating the similarity between the analysis type representation and each voiceprint type representation, and determining the voiceprint type representation whose similarity meets a preset deviation degree condition as the pending voiceprint type representation; According to the preliminary mapping relationship, the pending comparative voiceprint features corresponding to the pending voiceprint type representation are determined to be a pending comparative voiceprint feature dataset; the preset type association relationship includes a mapping relationship between the voiceprint type representation and the comparative voiceprint features, and a mapping relationship between the comparative voiceprint features and the bolt reference voiceprints in the bolt voiceprint database; Calculating the feature matching degree between each pending comparative voiceprint feature in the pending comparative voiceprint feature data set and the voiceprint feature to be analyzed; Determine the pending comparison voiceprint feature whose feature matching degree meets the preset matching degree condition as the target comparison voiceprint feature, obtain a target comparison voiceprint feature dataset, and obtain a target bolt reference voiceprint dataset corresponding to the target comparison voiceprint feature dataset; A voiceprint analysis result corresponding to the voiceprint to be analyzed is obtained according to the target bolt reference voiceprint dataset.
2. The method according to claim 1, characterized in that The interference removal processing of the tower detection audio to be analyzed includes: Acquire target monitoring audio data of the tower detection audio to be analyzed, wherein the target monitoring audio data includes a segment of the tower detection audio to be analyzed; Determining a target audio segment in the target monitoring audio data, wherein the target audio segment is a segment containing interfering noise, and the interfering noise is invalid data; Determining an initial time node of the interfering noise in each target audio segment; Determining an accurate time node in each of the initial time nodes according to segment sequence information of each of the initial time nodes in all the target audio segments in the segment to which it belongs; For each target audio segment, a sequence identifier of the interference noise in the target audio segment is determined based on the accurate time node of the target audio segment, and interference removal processing is performed based on the sequence identifier.
3. The method according to claim 1, characterized in that The method further comprises: Using the voiceprint analysis model, a deep learning feature extraction operation is performed on each bolt reference voiceprint in the bolt voiceprint database to obtain a deep reference voiceprint feature corresponding to each bolt reference voiceprint; Performing a deep learning feature extraction operation on the voiceprint to be analyzed to obtain a deep voiceprint feature to be analyzed corresponding to the voiceprint to be analyzed; Based on the target bolt reference voiceprint dataset, the feature similarity between the deep voiceprint feature to be analyzed and the deep reference voiceprint feature corresponding to each target bolt reference voiceprint is calculated; According to the sorting of the feature similarities based on the error values, the target bolt reference voiceprint with the smallest error value is extracted and determined as the voiceprint analysis result.
4. The method according to claim 1, wherein Before performing feature extraction and feature conversion operations on the voiceprint to be analyzed by the voiceprint analysis model to obtain a vector representation to be analyzed of the voiceprint to be analyzed, the method further includes: Obtaining a similar voiceprint group data set; each similar voiceprint group data set includes at least one similar voiceprint group; each similar voiceprint group includes sample voiceprints with the same preset target value; Performing feature extraction and feature conversion operations on each sample voiceprint in each similar voiceprint group data set using an initial voiceprint analysis model to obtain a sample vector representation of each sample voiceprint; Performing type inference and sparse coding based on the sample vector representation to obtain a sample type vector corresponding to each sample voiceprint; Based on the sample type vector and the preset target value, obtaining the category error corresponding to each similar voiceprint group data set; By using the initial voiceprint analysis model, transform coding is performed on the sample vector representation in the voiceprint classification domain corresponding to the sample type vector to obtain the classification domain coding features corresponding to each sample voiceprint; For each group of similar voiceprints, in the similar voiceprint groups with the same preset target value in the data sets of the similar voiceprint groups, sample training voiceprint extraction is performed according to the classification domain coding features corresponding to each sample voiceprint, and a training voiceprint array corresponding to each group of similar voiceprints is obtained, thereby obtaining a training voiceprint array set corresponding to each data set of the similar voiceprint groups; According to the classification domain coding features corresponding to the sample voiceprints, sample classification domain feature similarity calculation and transform coding error calculation are performed on the training voiceprint arrays to obtain the classification domain coding errors corresponding to the data sets of the similar voiceprint groups; performing training data difference calculation on each training voiceprint array in the training voiceprint array set according to the sample vector representation of each sample voiceprint, and obtaining the matching characteristic error corresponding to each similar voiceprint group data set; Obtaining a final error based on the category error, the classification domain encoding error, and the matching feature error; Based on the final error, the model parameters of the initial voiceprint analysis model are tuned to obtain the trained voiceprint analysis model.
5. The method according to claim 4, characterized in that The sample voiceprints in each group of similar voiceprints include a reference voiceprint and a target sample voiceprint; in the similar voiceprint groups with the same preset target value in the data sets of the similar voiceprint groups, sample training voiceprint extraction is performed according to the classification domain coding features corresponding to each sample voiceprint, and the training voiceprint array corresponding to each group of similar voiceprints is obtained, including: Extracting sample voiceprints having the same target value as the preset target value of each group of similar voiceprints from the data sets of the similar voiceprint groups to obtain a set of voiceprints of the same type; Calculating the voiceprint matching degree between each voiceprint of the same type in the voiceprint set and a reference voiceprint in each group of the same type of voiceprints based on the classification domain coding features corresponding to each sample voiceprint, and determining the non-target sample voiceprint corresponding to the reference voiceprint based on the voiceprint matching degree and a preset non-target sample determination criterion; Each non-target sample voiceprint in the non-target sample voiceprint is integrated with the reference voiceprint and the target sample voiceprint to obtain a training voiceprint array corresponding to each group of similar voiceprints.
6. The method according to claim 5, characterized in that The method of performing sample classification domain feature similarity calculation and transform coding error calculation on each training voiceprint array based on the classification domain coding features corresponding to each sample voiceprint, and obtaining the classification domain coding errors corresponding to each similar voiceprint group data set, includes: According to the classification domain coding features corresponding to the sample voiceprints, obtaining the reference sample classification domain coding features corresponding to the reference voiceprints, the target sample classification domain coding features corresponding to the target sample voiceprints, and the non-target sample classification domain coding features corresponding to the non-target sample voiceprints in the training voiceprint arrays; Comparing the similarity between the classification domain coding feature of the reference sample and the classification domain coding feature of the target sample to obtain a first similarity coefficient; Calculating the feature similarity between the classification domain coding feature of the reference sample and the classification domain coding feature of the non-target sample to obtain a second similarity coefficient; Obtaining a similarity coefficient error according to an absolute difference between the first similarity coefficient and the second similarity coefficient; Generate the target type representations corresponding to the reference sample classification domain coding feature, the target sample classification domain coding feature, and the non-target sample classification domain coding feature by using a preset symbol coding function, and calculate the root mean square error between the reference sample classification domain coding feature, the target sample classification domain coding feature, the non-target sample classification domain coding feature, and their corresponding target type representations to obtain a prediction error; The similarity coefficient error and the prediction error are linearly superimposed to obtain the classification domain coding error.
7. The method according to claim 5, characterized in that The method further comprises: Extracting sample training voiceprints from the similar voiceprint groups with different preset target values in the respective similar voiceprint group data sets to obtain the overall sample training voiceprint data sets corresponding to the respective similar voiceprint group data sets; For each overall sample training voiceprint in the overall sample training voiceprint dataset, performing feature fusion on the sample type vector corresponding to each sample voiceprint in the overall sample training voiceprint and the classification domain coding feature to obtain a fusion feature; Based on the fusion features of each sample voiceprint in the overall sample training voiceprints, sample classification domain feature similarity calculation is performed to obtain the corresponding fusion feature errors of each similar voiceprint group data set.
8. The method according to claim 7, characterized in that The obtaining of a final error based on the category error, the classification domain encoding error, and the matching feature error includes: The category error, the classification domain encoding error, the fusion feature error and the matching feature error are linearly superimposed to obtain the final error.
9. A readable storage medium, characterized in that: The readable storage medium includes a computer program, and when the computer program is executed, the computer device where the readable storage medium is located is controlled to execute the method according to any one of claims 1 to 8.
Citation Information
Patent Citations
A Method and System for Analyzing Abnormal States of Communication Towers Based on Decision Algorithms
CN116910667B
Electric iron tower part defect detection method based on artificial intelligence
CN117952960A
Abnormal equipment monitoring method and system based on equipment operation audio
CN117854245A
Iron tower bolt health detection method and system based on artificial intelligence and voiceprint
CN118212938A