A method and system for processing iron tower bolts based on a voiceprint classification model

By constructing a voiceprint knowledge graph to optimize the bolt voiceprint instance dataset, training a target voiceprint classification model, identifying abnormal voiceprints, and constructing emergency handling strategies, the problem of low accuracy in bolt monitoring in existing technologies is solved. This enables real-time monitoring and automatic emergency handling of tower bolt status, improving monitoring efficiency and safety.

CN120108422BActive Publication Date: 2025-12-02JIANGSU ELECTRIC POWER CO RUDONG COUNTY POWER SUPPLY CO +2
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510140401.5
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-02-08
Publication Date
2025-12-02
Estimated Expiration
2045-02-08

AI Technical Summary

Technical Problem

In existing technologies, AI-based methods for monitoring the health of tower bolts fail to effectively consider the similarity and correlation between bolt acoustic signatures, resulting in low monitoring accuracy, inefficient traditional inspections, and difficulty in timely detection of potential safety hazards.

Method used

By constructing a voiceprint knowledge graph, optimizing the bolt voiceprint instance dataset, determining sample similarity and correlation influencing factors, training a target voiceprint classification model, identifying abnormal voiceprints, and constructing emergency response strategies, real-time monitoring and automatic emergency response can be achieved.

Benefits of technology

This significantly improves the accuracy and efficiency of monitoring the condition of tower bolts, ensuring the stability and safety of the towers and enabling the timely detection and handling of potential safety hazards.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120108422B_ABST
    Figure CN120108422B_ABST
Patent Text Reader

Abstract

This invention relates to a method and system for handling tower bolts based on a voiceprint classification model, belonging to the field of network communication technology. The method includes: first, extracting the voiceprint to be analyzed from tower monitoring audio, and classifying it using a pre-trained target voiceprint classification model. When an abnormal voiceprint is detected, the system receives and processes the relevant emergency request. Subsequently, based on the abnormal voiceprint and the emergency request, a current emergency handling instruction is constructed, and an emergency handling strategy matching basis is obtained through an emergency scenario recognition model. Finally, a suitable target emergency strategy is selected from a preset emergency handling strategy database and executed. This invention achieves real-time monitoring and automatic emergency handling of tower bolt status, significantly improving the efficiency and safety of tower maintenance.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of network communication technology, and more specifically, to a method and system for processing tower bolts based on a voiceprint classification model. Background Technology

[0002] With rapid societal development, the stability and safety of communication towers, as crucial infrastructure, are paramount. However, issues such as loose bolts on the towers can affect their structural stability, potentially leading to safety accidents. Traditional inspection methods are inefficient and struggle to detect potential safety hazards in a timely manner.

[0003] Existing technologies also include artificial intelligence-based methods for monitoring bolt loosening, but certain problems remain. For example, invention patent CN118212938A discloses a method and system for detecting the health of tower bolts based on artificial intelligence and voiceprints. This method acquires pending feedback voiceprint instance data containing real values, which represent the corresponding health status of the tower bolts. Then, this data is balanced, and voiceprint features are extracted. Next, a classification model for the health status of tower bolts is trained based on the extracted voiceprint features. Finally, the current set of feedback voiceprints is collected, its voiceprint features are extracted, and the trained model is used to identify the health status of the tower bolts. By predicting the real values, the health status of the tower bolts corresponding to the current feedback voiceprint can be understood. However, this patent does not consider the similarity and correlation between bolt voiceprint instances, thus the accuracy of monitoring the health status of tower bolts needs improvement. Summary of the Invention

[0004] To overcome the shortcomings of existing technologies, this invention provides a method for handling tower bolts based on a voiceprint classification model, which solves the problems of poor monitoring effect and low safety of tower bolt loosening. In addition, this invention also provides a tower bolt handling system based on a voiceprint classification model.

[0005] According to a first aspect of the present invention, a method for processing tower bolts based on a voiceprint classification model is provided, the method comprising:

[0006] The voiceprint to be analyzed is obtained from the audio monitored by the tower;

[0007] The voiceprint to be analyzed is input into a pre-trained target voiceprint classification model to obtain the category classification result of the voiceprint to be analyzed; the process of obtaining the target voiceprint classification model includes: constructing an initial voiceprint classification model and a bolt voiceprint instance dataset, optimizing the bolt voiceprint instance dataset, and training the initial voiceprint classification model based on the optimized bolt voiceprint instance dataset. The optimization process of the bolt voiceprint instance dataset includes:

[0008] The sample similarity between bolt soundprint instances is determined. Based on the sample similarity, the target sample similarity between the bolt soundprint instances and the corresponding bolt soundprint instances in the bolt soundprint instance dataset of the soundprint knowledge graph is extracted, thereby determining the association influence factor of the bolt soundprint instances. The soundprint knowledge graph is obtained based on the association factor. The degree of difference between the bolt soundprint instances is determined based on the soundprint knowledge graph. The diffusion target value information of the bolt soundprint instances is obtained based on the weight corresponding to the degree of difference. The optimized bolt soundprint instance dataset is obtained based on the diffusion target value information of the bolt soundprint instances.

[0009] When the classification result is characterized as an abnormal voiceprint result, an emergency processing request for the voiceprint to be analyzed is received;

[0010] Construct current emergency response instructions based on the voiceprint to be analyzed and the emergency response requirements;

[0011] The current emergency response instructions are input into a pre-trained emergency scenario recognition model to obtain an emergency response strategy matching set.

[0012] The target emergency strategy is selected from the preset emergency strategy database based on the emergency response strategy matching set.

[0013] Furthermore, including:

[0014] The process of obtaining the target voiceprint classification model specifically includes:

[0015] Obtain a bolt acoustic signature instance dataset, which includes at least one bolt acoustic signature instance configured with an initial target value;

[0016] Based on the initial voiceprint classification model, feature extraction is performed on the bolt voiceprint instances in the bolt voiceprint instance dataset to obtain the voiceprint feature dataset.

[0017] The voiceprint features corresponding to each bolt voiceprint instance are extracted from the voiceprint feature dataset, and the sample similarity between the bolt voiceprint instances is determined based on the voiceprint features of the bolt voiceprint instances.

[0018] Based on the sample similarity, the bolt voiceprint knowledge graph bolt voiceprint instance is extracted from the bolt voiceprint instance dataset to obtain the bolt voiceprint instance dataset.

[0019] The similarity between the bolt voiceprint instance and the bolt voiceprint instance in the corresponding voiceprint knowledge graph bolt voiceprint instance dataset is extracted from the sample similarity.

[0020] Clustering is performed on the similarity of the target samples to obtain the association between the bolt soundprint instances and the bolt soundprint instances in the soundprint knowledge graph bolt soundprint instance dataset;

[0021] Based on the aforementioned correlation, determine the correlation influence factor of the bolt sound pattern instance;

[0022] Based on the aforementioned correlation influencing factors, the bolt soundprint instance is determined as an element to generate an initial soundprint knowledge graph, and the initial soundprint knowledge graph is balanced to obtain the soundprint knowledge graph.

[0023] Based on the initial target value of the bolt acoustic print instance, generate initial target value information corresponding to the bolt acoustic print instance dataset; the initial target value information includes an initial target value vector corresponding to each bolt acoustic print instance;

[0024] The degree of difference between the bolt acoustic print instances is determined based on the acoustic print knowledge graph.

[0025] Obtain the voiceprint correlation weight corresponding to the degree of difference, and assign weights to the initial target value vector of the bolt voiceprint instance according to the voiceprint correlation weight.

[0026] Cluster the initial target value vector after weight assignment to obtain the diffusion target value information of the bolt soundprint instance; the diffusion target value information is obtained by spreading the initial target value of each bolt soundprint instance in the previously constructed soundprint knowledge graph, that is, each bolt soundprint instance node will pass its target value to other nodes connected to it, so as to obtain the target value information after considering the influence of its neighboring nodes.

[0027] Obtain the diffusion target value vector of the bolt acoustic pattern instance from the diffusion target value information;

[0028] Extract the target value component with the largest target value from the diffusion target value vector;

[0029] The index information of the target value component is determined in the diffusion target value vector;

[0030] Obtain the candidate target value corresponding to the index information, and determine the candidate target value as the diffusion target value of the bolt sound pattern instance;

[0031] Perform a comparison operation between the diffusion target value and the initial target value configured for the corresponding bolt acoustic signature instance;

[0032] When the diffusion target value is inconsistent with the initial target value, the bolt sound pattern instance is determined to be the target bolt sound pattern instance to be optimized.

[0033] The initial target value of the target bolt acoustic signature instance is changed to the corresponding diffusion target value to obtain the optimized bolt acoustic signature instance dataset;

[0034] The initial voiceprint classification model is trained based on the optimized bolt voiceprint instance dataset to obtain the trained target voiceprint classification model.

[0035] Furthermore, including:

[0036] The training process for the initial soundprint classification model based on the optimized bolt soundprint instance dataset includes:

[0037] Based on the acoustic features and target values ​​of bolt acoustic instances in the optimized bolt acoustic instance dataset, the model parameters of the initial acoustic classification model are adjusted.

[0038] Based on the initial voiceprint classification model, feature extraction is performed on the bolt voiceprint instances in the optimized bolt voiceprint instance dataset to obtain the target voiceprint feature dataset.

[0039] Based on the target acoustic signature feature dataset, the target value of the bolt acoustic signature instance is optimized;

[0040] Repeat the steps of adjusting the model parameters of the initial voiceprint classification model based on the voiceprint features and target values ​​of bolt voiceprint instances in the optimized bolt voiceprint instance dataset, until the initial voiceprint classification model meets the preset training termination condition, and obtain the trained target voiceprint classification model.

[0041] Furthermore, including:

[0042] The step of adjusting the model parameters of the initial soundprint classification model based on the soundprint features and target values ​​of bolt soundprint instances in the optimized bolt soundprint instance dataset includes:

[0043] Based on the target value of the bolt sound pattern instance in the optimized bolt sound pattern instance dataset, determine the target value error parameter of the bolt sound pattern instance;

[0044] Based on the acoustic features of bolt acoustic instances in the optimized bolt acoustic instance dataset, the acoustic error parameter of the bolt acoustic instance is determined.

[0045] The target value error parameter and the voiceprint error parameter are integrated to obtain the integrated error parameter, and the model parameters of the initial voiceprint classification model are adjusted according to the integrated error parameter.

[0046] Furthermore, including:

[0047] The step of determining the acoustic error parameter of the bolt acoustic instance based on the acoustic features of the bolt acoustic instances in the optimized bolt acoustic instance dataset includes:

[0048] Based on the target value of the bolt sound pattern instance in the optimized bolt sound pattern instance dataset, perform a type recognition operation on the bolt sound pattern instance to obtain a subset of bolt sound pattern instance data corresponding to each target value.

[0049] Based on the acoustic features of bolt acoustic features in the bolt acoustic feature instance subset, the target acoustic features corresponding to the bolt acoustic feature instance subset are determined.

[0050] The acoustic features of the bolt acoustic feature instance and the target acoustic features corresponding to the bolt acoustic feature instance data subset are integrated to obtain the acoustic error parameter of the bolt acoustic feature instance.

[0051] Furthermore, including:

[0052] The step of integrating the acoustic features of the bolt acoustic signature instance and the target acoustic features corresponding to the bolt acoustic signature instance data subset to obtain the acoustic error parameter of the bolt acoustic signature instance includes:

[0053] Based on the acoustic features of the bolt acoustic features, determine the feature difference between bolt acoustic features in the bolt acoustic feature data subset and obtain the first feature difference.

[0054] Based on the target feature difference amount corresponding to the bolt soundprint instance data subset, determine the feature difference amount between the bolt soundprint instance data subsets, and obtain the second feature difference amount;

[0055] Determine the feature difference between the first feature difference and the second feature difference, obtain the third feature difference, and integrate the third feature difference with the critical difference value to obtain the integration difference degree;

[0056] If the integration difference is greater than a preset difference threshold, the average feature difference of the integration difference is determined to correct the target acoustic feature, so as to obtain the acoustic error parameter of the bolt acoustic instance.

[0057] Furthermore, including:

[0058] The step of inputting the current emergency response instructions into a pre-trained emergency scenario recognition model to obtain an emergency response strategy matching set includes:

[0059] For the voiceprint to be analyzed in the current emergency response instructions, the voiceprint matching model is used to search for reference bolt voiceprints that meet the matching criteria to obtain the matching bolt voiceprints. The voiceprint matching model is obtained through prior training.

[0060] Search for historical emergency response requirements related to the matching bolt acoustic signature;

[0061] For the emergency handling needs in the current emergency handling instructions, reference emergency handling needs with a matching degree exceeding the first semantic matching degree threshold are searched in the historical emergency handling needs through a semantic analysis model to obtain matching emergency handling instructions. The semantic analysis model is obtained through prior training.

[0062] The emergency scenario recognition model is used to classify the current emergency handling instruction and the matching emergency handling instruction into scenario types to obtain the scenario type of the current emergency handling instruction and the scenario type of each matching emergency handling instruction. The emergency scenario recognition model is trained using the processing requirements with scenario labels configured in past data as samples.

[0063] The matching emergency response instructions that are consistent with the scenario type of the current emergency response instruction are used as the emergency response strategy matching set.

[0064] Furthermore, including:

[0065] The process of obtaining the voiceprint matching model includes:

[0066] Acquire voiceprint samples for analysis;

[0067] Multiple voiceprint conversion strategies are used to perform voiceprint conversion operations on the sample voiceprint to be analyzed, resulting in multiple converted voiceprints.

[0068] Multiple target bolt voiceprint datasets were obtained by using various voiceprint conversion strategies on the same sample voiceprint to be analyzed.

[0069] Based on the voiceprints of different samples to be analyzed, multiple converted voiceprints are generated into non-target bolt voiceprint datasets by applying arbitrary voiceprint conversion strategies.

[0070] Based on the target bolt acoustic print dataset and the non-target bolt acoustic print dataset, the initial acoustic print matching model is pre-trained using a self-supervised strategy to obtain the pre-trained acoustic print matching model.

[0071] Obtain the target voiceprint to be analyzed corresponding to the business scenario of the current emergency response instruction;

[0072] Based on the target voiceprint to be analyzed, generate target ensemble learning samples and non-target ensemble learning samples;

[0073] Based on the target ensemble learning samples and non-target ensemble learning samples, the pre-trained voiceprint matching model is ensemble-learned using a self-supervised strategy to obtain the voiceprint matching model.

[0074] Furthermore, including:

[0075] The process of obtaining the semantic analysis model includes:

[0076] Obtain emergency sample processing requirements;

[0077] The sample emergency processing requirements are obtained by performing semantic transformation operations on the sample emergency processing requirements through multiple semantic transformation strategies.

[0078] Based on the emergency processing needs of the same sample, the semantic transformation processing needs obtained by applying multiple semantic transformation strategies are used to generate the target sample processing needs.

[0079] Based on the emergency processing needs of different samples, the semantic transformation processing needs obtained by applying arbitrary semantic transformation strategies generate non-target sample processing needs.

[0080] Based on the target sample processing requirements and the non-target sample processing requirements, a training process is performed on the initial semantic analysis model using a self-supervised strategy to obtain the semantic analysis model.

[0081] On the other hand, the present invention also provides a tower bolt processing system based on a voiceprint classification model, the system comprising:

[0082] The acquisition module is used to obtain the voiceprint to be analyzed from the audio of the tower monitoring.

[0083] The model training module is used to input the voiceprint to be analyzed into a pre-trained target voiceprint classification model to obtain the category classification result of the voiceprint to be analyzed. The process of obtaining the target voiceprint classification model includes: constructing an initial voiceprint classification model and a bolt voiceprint instance dataset, optimizing the bolt voiceprint instance dataset, and training the initial voiceprint classification model based on the optimized bolt voiceprint instance dataset. The optimization process of the bolt voiceprint instance dataset includes:

[0084] The similarity calculation module is used to determine the sample similarity between bolt soundprint instances. Based on the sample similarity, the module extracts the target sample similarity between the bolt soundprint instances and the corresponding bolt soundprint instances in the bolt soundprint instance dataset of the soundprint knowledge graph, thereby determining the association influence factor of the bolt soundprint instances. Based on the association factor, the module obtains the soundprint knowledge graph. Based on the soundprint knowledge graph, the module determines the degree of difference between the bolt soundprint instances. Based on the weight corresponding to the degree of difference, the module obtains the diffusion target value information of the bolt soundprint instances. Based on the diffusion target value information of the bolt soundprint instances, the module obtains the optimized bolt soundprint instance dataset.

[0085] The result processing module receives an emergency processing request for the voiceprint to be analyzed when the classification result is characterized as an abnormal voiceprint result.

[0086] Construct current emergency response instructions based on the voiceprint to be analyzed and the emergency response requirements;

[0087] The current emergency response instructions are input into a pre-trained emergency scenario recognition model to obtain an emergency response strategy matching set;

[0088] The target emergency strategy is selected from the preset emergency strategy database based on the emergency response strategy matching set.

[0089] Beneficial effects: Compared with the prior art, the present invention has the following advantages:

[0090] This invention first extracts the voiceprints to be analyzed from the audio monitored by the tower and then classifies them using a pre-trained target voiceprint classification model. When an abnormal voiceprint is detected, the system receives and processes the relevant emergency request. Subsequently, based on the abnormal voiceprint and the emergency request, a current emergency handling instruction is constructed, and an emergency handling strategy matching set is obtained through an emergency scenario recognition model. Finally, a suitable target emergency strategy is selected from a preset emergency handling strategy database for execution.

[0091] This invention also optimizes the bolt acoustic signature instance dataset. The optimization process specifically includes: determining the sample similarity between bolt acoustic signature instances; based on the sample similarity, extracting the target sample similarity between the bolt acoustic signature instances and the corresponding acoustic signature knowledge graph bolt acoustic signature instance dataset, thereby determining the association influence factor of the bolt acoustic signature instances; obtaining the acoustic signature knowledge graph based on the association factor; determining the degree of difference between the bolt acoustic signature instances based on the acoustic signature knowledge graph; obtaining the diffusion target value information of the bolt acoustic signature instances based on the weight corresponding to the degree of difference; and obtaining the optimized bolt acoustic signature instance dataset based on the diffusion target value information of the bolt acoustic signature instances.

[0092] Therefore, this invention enables real-time monitoring and automatic emergency handling of the status of tower bolts, significantly improving the efficiency, safety, and accuracy of tower maintenance. Attached Figure Description

[0093] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the embodiments will be briefly described below. It should be understood that the following drawings only show some embodiments of the present invention and should not be considered as limiting the scope. For those skilled in the art, other related drawings can be obtained based on these drawings without creative effort.

[0094] Figure 1 A flowchart illustrating the steps of a method for processing iron tower bolts based on a voiceprint classification model, provided in an embodiment of the present invention;

[0095] Figure 2 A schematic block diagram of the structure of a computer device provided in an embodiment of the present invention. Detailed Implementation

[0096] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. The components of the embodiments of the present invention described and shown in the accompanying drawings can generally be arranged and designed in various different configurations.

[0097] The specific embodiments of the present invention will now be described in detail with reference to the accompanying drawings.

[0098] In order to solve the technical problems in the aforementioned background art, such as Figure 1 This is a flowchart illustrating a method for processing tower bolts based on a voiceprint classification model, as provided in this embodiment of the present disclosure. The following is a detailed description of this method for processing tower bolts based on a voiceprint classification model.

[0099] Step S201: Obtain the voiceprint to be analyzed from the tower monitoring audio;

[0100] Step S202: Input the voiceprint to be analyzed into a pre-trained target voiceprint classification model to obtain the category classification result of the voiceprint to be analyzed;

[0101] Step S203: When the classification result is characterized as an abnormal voiceprint result, receive an emergency processing request for the voiceprint to be analyzed.

[0102] Step S204: Construct current emergency response instructions based on the voiceprint to be analyzed and the emergency response requirements;

[0103] Step S205: Input the current emergency response instruction into the pre-trained emergency scenario recognition model to obtain an emergency response strategy matching set;

[0104] Step S206: Select a target emergency strategy from the preset emergency strategy database according to the emergency response strategy matching set.

[0105] In this embodiment of the invention, for example, the server receives an audio file from an audio monitoring device installed on the tower. This audio file contains sounds emitted from the tower bolts. The server uses professional voiceprint extraction technology to successfully extract voiceprint data to be analyzed from this audio file. The server inputs the extracted voiceprint data into a pre-trained voiceprint classification model. This model is trained using a large amount of normal and abnormal bolt voiceprint data and can accurately identify whether a voiceprint is normal. After automatic analysis by the model, a classification result is output, indicating that this voiceprint belongs to an abnormal voiceprint. After determining that the voiceprint is abnormal, the server automatically triggers an emergency handling mechanism. Simultaneously, the server receives an emergency handling request from the tower maintenance team, requesting further analysis and processing of the abnormal voiceprint. Based on the specific characteristics of the abnormal voiceprint and the received emergency handling request, the server constructs an emergency handling instruction. This instruction includes a detailed description of the abnormal voiceprint, possible causes of the problem, and suggested preliminary handling measures. The server inputs the constructed emergency handling instruction into a pre-trained emergency scenario recognition model. This model can provide the server with a matching basis for emergency handling strategies based on different emergency scenarios. Through automatic analysis by the model, the server obtained the basis for an emergency response strategy matching the current abnormal soundprint. Based on the matching criteria provided by the emergency scenario recognition model, the server selected the most suitable target emergency strategy from the preset emergency response strategy database. This strategy includes detailed emergency response steps and the required resource allocation plan, providing clear operational guidelines for the tower maintenance team.

[0106] In this embodiment of the invention, the target voiceprint classification model is obtained in the following manner.

[0107] Obtain a bolt acoustic signature instance dataset, which includes at least one bolt acoustic signature instance configured with an initial target value;

[0108] Based on the initial voiceprint classification model, feature extraction is performed on the bolt voiceprint instances in the bolt voiceprint instance dataset to obtain the voiceprint feature dataset.

[0109] Based on the aforementioned voiceprint feature dataset, the bolt voiceprint instance is identified as an element to generate a voiceprint knowledge graph;

[0110] The initial target value of the bolt voiceprint instance is optimized based on the voiceprint knowledge graph to obtain the optimized bolt voiceprint instance dataset.

[0111] The initial voiceprint classification model is trained based on the optimized bolt voiceprint instance dataset to obtain the trained target voiceprint classification model.

[0112] In this embodiment of the invention, for example, the server retrieves a bolt acoustic signature instance dataset from a large database. This dataset contains hundreds or thousands of bolt acoustic signature instances, each configured with an initial target value. This initial target value represents a preliminary judgment of whether the acoustic signature instance is normal or abnormal. The server uses an initial acoustic signature classification model to extract features from each bolt acoustic signature instance in the dataset. This process mainly involves extracting key information from the acoustic signature, such as frequency and amplitude, thereby forming an acoustic signature feature dataset containing these key features. Based on the extracted acoustic signature feature dataset, the server begins to construct an acoustic signature knowledge graph. In this graph, each bolt acoustic signature instance is represented as a node, and the connections between nodes represent the similarity or correlation between these acoustic signature instances. In this way, the server can more intuitively understand the relationships between different acoustic signature instances. By referring to the relationships between various bolt acoustic signature instances in the acoustic signature knowledge graph, the server begins to optimize the initial target value. For example, if some initially labeled normal voiceprint instances are found to be closely linked to multiple abnormal voiceprint instances in the graph, the initial target values ​​of these voiceprint instances will be adjusted to be abnormal. Through this optimization process, the server obtains a more accurate and reliable dataset of bolt voiceprint instances. Finally, the server uses the optimized bolt voiceprint instance dataset to train the initial voiceprint classification model. During training, the model continuously learns and adjusts its internal parameters to better adapt to and optimize the dataset. After multiple rounds of training and adjustment, the server finally obtains a fully trained target voiceprint classification model that can more accurately identify and classify the voiceprints of tower bolts.

[0113] In this embodiment of the invention, the step of optimizing the initial target value of the bolt voiceprint instance based on the voiceprint knowledge graph to obtain the optimized bolt voiceprint instance dataset can be implemented through the following example.

[0114] The initial target value of the bolt acoustic print instance is diffused among the elements of the acoustic print knowledge graph to obtain the diffusion target value information of the bolt acoustic print instance.

[0115] Based on the diffusion target value information, the initial target value of the bolt sound pattern instance is optimized to obtain the optimized bolt sound pattern instance dataset.

[0116] In this embodiment of the invention, for example, the server first diffuses the initial target value (e.g., normal or abnormal) of each bolt acoustic signature instance within the previously constructed acoustic signature knowledge graph. This process is similar to information propagation, where each bolt acoustic signature instance node transmits its target value to other nodes connected to it. Through this diffusion mechanism, the server can obtain the diffusion target value information of each bolt acoustic signature instance after considering the influence of its neighboring nodes. For example, if a bolt acoustic signature instance originally marked as normal is closely connected to multiple abnormal nodes in the graph, then the diffusion target value information of this normal node will be affected by these abnormal nodes. After obtaining the diffusion target value information, the server uses this information to optimize the initial target value of the bolt acoustic signature instances. The optimization process mainly involves adjusting the initial target value based on the diffusion target value information to better align with the overall trends and correlations reflected in the acoustic signature knowledge graph. For example, for normal nodes strongly influenced by abnormal nodes, the server will adjust their initial target value from normal to abnormal to reflect their true state in the graph. Through this optimization process, the server can obtain a more accurate and realistic optimized bolt acoustic signature instance dataset. This dataset will serve as the basis for subsequent training and application of the voiceprint classification model, helping to improve the model's accuracy and reliability.

[0117] In this embodiment of the invention, the step of diffusing the initial target value of the bolt soundprint instance among the elements of the soundprint knowledge graph to obtain the diffusion target value information of the bolt soundprint instance can be implemented through the following example.

[0118] Based on the initial target value of the bolt sound pattern instance, generate the initial target value information corresponding to the bolt sound pattern instance dataset;

[0119] Based on preset diffusion rules, the initial target value information is diffused among the elements of the voiceprint knowledge graph to obtain the diffusion target value information of the bolt voiceprint instance.

[0120] In this embodiment of the invention, for example, the server first traverses each instance in the bolt acoustic signature instance dataset and examines the initial target value assigned to each instance. These initial target values ​​may be initially labeled by experts based on acoustic signature features, such as normal or abnormal. The server organizes these initial target values ​​into an initial target value information table, which corresponds to the bolt acoustic signature instance dataset and provides basic data for the subsequent diffusion process. The server then diffuses the target values ​​on the acoustic signature knowledge graph according to preset diffusion rules. This diffusion process can be compared to information dissemination in a social network, where each bolt acoustic signature instance, as a node in the graph, transmits its initial target value to adjacent nodes according to certain rules. For example, the rule may stipulate that if a bolt acoustic signature instance is marked as abnormal, other instances connected to it will be affected, and their diffusion target values ​​will shift towards the abnormal direction. This process continues until a stable diffusion state is reached. The server obtains the diffusion target value information of the bolt acoustic signature instances by calculating the final diffusion target value of each node. This information not only considers the characteristics of each instance itself, but also incorporates the influence of other instances in the graph, thus providing a more comprehensive reflection of the instance's true state. Through this diffusion process, the server can obtain a diffusion target value dataset that integrates global information from the voiceprint knowledge graph. This dataset will provide more accurate and richer information for subsequent target value optimization and model training.

[0121] In this embodiment of the invention, the initial target value information includes an initial target value vector corresponding to each bolt acoustic signature instance. The step of diffusing the initial target value information among the elements of the acoustic signature knowledge graph based on a preset diffusion rule to obtain the diffusion target value information of the bolt acoustic signature instance can be implemented through the following example.

[0122] The degree of difference between the bolt acoustic print instances is determined based on the acoustic print knowledge graph.

[0123] Obtain the voiceprint correlation weight corresponding to the degree of difference, and assign weights to the initial target value vector of the bolt voiceprint instance according to the voiceprint correlation weight.

[0124] Cluster the initial target value vector after weight assignment to obtain the diffusion target value information of the bolt sound pattern instance.

[0125] In this embodiment of the invention, for example, the server first analyzes the relationships between various bolt voiceprint instances in the voiceprint knowledge graph. These relationships can be based on factors such as voiceprint similarity, frequency characteristics, and amplitude characteristics. By analyzing these factors, the server can calculate the degree of difference between each pair of bolt voiceprint instances. For example, if two voiceprint instances have similar frequency distributions and consistent amplitude changes, then the degree of difference between them is small; conversely, if the frequency and amplitude characteristics differ significantly, then the degree of difference between them is large. The server determines the voiceprint correlation weight between each pair of bolt voiceprint instances based on the degree of difference calculated in the previous step. Instance pairs with smaller degrees of difference are assigned larger weights because they are more correlated. Next, the server uses these weights to assign weights to the initial target value vectors of the bolt voiceprint instances. For example, if two bolt voiceprint instances are very similar (small degree of difference), then their target value vectors will influence each other more during the diffusion process, that is, changes in the target value of one instance will have a greater impact on the target value of another similar instance. After completing the weight assignment, the server performs cluster analysis on these weighted initial target value vectors. The purpose of clustering is to group bolt acoustic fingerprint instances with similar target value vectors and high correlation together to form a diffusion target value cluster. During this process, the server considers the weight relationships between instances to ensure that the clustering results accurately reflect the correlations within the acoustic fingerprint knowledge graph. After clustering, the server obtains the diffusion target value information of the bolt acoustic fingerprint instances, which serves as an important basis for subsequent optimization and model training. Through this processing flow, the server can more accurately capture the inherent connections and patterns between bolt acoustic fingerprint instances, thereby improving the performance of the acoustic fingerprint classification model.

[0126] In this embodiment of the invention, the step of optimizing the initial target value of the bolt sound pattern instance based on the diffusion target value information to obtain an optimized bolt sound pattern instance dataset can be implemented through the following example.

[0127] Obtain the diffusion target value vector of the bolt acoustic pattern instance from the diffusion target value information;

[0128] Based on the diffusion target value vector, determine the diffusion target value of the bolt acoustic pattern instance;

[0129] Based on the diffusion target value, the initial target value of the bolt sound pattern instance is optimized to obtain the optimized bolt sound pattern instance dataset.

[0130] In this embodiment of the invention, for example, the server extracts the corresponding diffusion target value vector for each bolt acoustic signature instance from the calculated diffusion target value information. This vector consists of multiple elements, each representing the diffusion target value of the instance in a specific dimension, such as the score or classification result of frequency features, amplitude features, etc. The server performs a comprehensive analysis on each bolt acoustic signature instance based on the extracted diffusion target value vector. This can be achieved by calculating the mean, median, or other statistics of the vector to determine a specific diffusion target value. For example, if most elements in the diffusion target value vector point to the "abnormal" category, and the mean or median of the vector exceeds a certain preset threshold, then the diffusion target value of the bolt acoustic signature instance may be determined as "abnormal". After determining the diffusion target value for each bolt acoustic signature instance, the server compares these diffusion target values ​​with the instance's initial target value. If the diffusion target value is inconsistent with the initial target value, or if the diffusion target value provides additional information (such as higher confidence), the server optimizes the initial target value according to preset rules or algorithms. For example, if an initial target value for a bolt acoustic signature instance is labeled "normal," but the diffusion target value strongly points to "abnormal," and this instance is closely linked to multiple known anomalous instances in the acoustic signature knowledge graph, then the server will optimize the initial target value of this instance from "normal" to "abnormal." After optimizing the initial target values ​​of all bolt acoustic signature instances, the server obtains an optimized bolt acoustic signature instance dataset. Through this optimization process, the server can leverage global information and relationships within the acoustic signature knowledge graph to improve the original bolt acoustic signature instance dataset, thereby enhancing the training effect and performance of subsequent acoustic signature classification models.

[0131] In this embodiment of the invention, the step of determining the diffusion target value of the bolt sound pattern instance based on the diffusion target value vector can be implemented through the following example.

[0132] Extract the target value component with the largest target value from the diffusion target value vector;

[0133] The index information of the target value component is determined in the diffusion target value vector;

[0134] Obtain the candidate target value corresponding to the index information, and determine the candidate target value as the diffusion target value of the bolt sound pattern instance.

[0135] In this embodiment of the invention, exemplarily, assume the server has calculated a diffusion target value vector for a bolt acoustic signature instance. This vector may contain multiple components, each representing a different target value (such as the confidence level of a classification label). The server iterates through this vector, comparing the values ​​of each component to find the component with the largest value. For example, if the vector is represented as [0.2, 0.5, 0.8, 0.3], then the component with the largest value is 0.8. After finding the target value component with the largest value, the server records the index information of this component. The index information indicates the position of the component in the vector. Continuing the example above, the component with the largest value, 0.8, in the vector [0.2, 0.5, 0.8, 0.3] is located in the third position, so its index information is 2 (indices usually start counting from 0). With the index information, the server retrieves the corresponding candidate target value from a predefined list or mapping of target values ​​based on this index. This candidate target value is the diffusion target value of the bolt acoustic signature instance. For example, if the target value corresponding to index 2 is "abnormal", then the server will determine the diffusion target value of this bolt soundprint instance as "abnormal".

[0136] In this embodiment of the invention, the step of optimizing the initial target value of the bolt sound pattern instance based on the diffusion target value to obtain the optimized bolt sound pattern instance dataset can be implemented through the following example.

[0137] Perform a comparison operation between the diffusion target value and the initial target value configured for the corresponding bolt acoustic signature instance;

[0138] When the diffusion target value is inconsistent with the initial target value, the bolt sound pattern instance is determined to be the target bolt sound pattern instance to be optimized.

[0139] The initial target value of the target bolt acoustic signature instance is changed to the corresponding diffusion target value to obtain the optimized bolt acoustic signature instance dataset.

[0140] In this embodiment of the invention, for example, the server first obtains the diffusion target value for each bolt acoustic signature instance. These diffusion target values ​​are calculated through the preceding steps and represent the target classification or state of the bolt acoustic signature instance after diffusion in the acoustic signature knowledge graph. Next, the server compares the diffusion target value of each bolt acoustic signature instance with its configured initial target value. For example, the initial target value of a bolt acoustic signature instance might be "normal," while the diffusion target value calculated through diffusion might be "abnormal." During the comparison, if the server finds that the diffusion target value of a bolt acoustic signature instance is inconsistent with the initial target value, it marks this bolt acoustic signature instance as a target bolt acoustic signature instance to be optimized. Continuing the example above, since the initial target value of this bolt acoustic signature instance is "normal," while the diffusion target value is "abnormal," the two are inconsistent, therefore this instance will be determined as a target bolt acoustic signature instance to be optimized. Once the target bolt acoustic signature instances to be optimized are determined, the server changes the initial target value of these instances to the corresponding diffusion target value. In the example above, the server changes the initial target value of the bolt acoustic signature instance from "normal" to "abnormal." This modification process is applied to all target bolt acoustic print instances marked as requiring optimization. The server ultimately obtains an optimized bolt acoustic print instance dataset, which includes bolt acoustic print instances optimized by the diffusion target value. Through this optimization process, the server can leverage global information and relationships within the acoustic print knowledge graph to improve the initial target value of the bolt acoustic print instances, thereby enhancing the accuracy and reliability of the dataset and providing a better data foundation for subsequent acoustic print classification model training.

[0141] In this embodiment of the invention, the step of determining the bolt voiceprint instance as an element to generate a voiceprint knowledge graph based on the voiceprint feature dataset can be implemented through the following example.

[0142] The voiceprint features corresponding to each bolt voiceprint instance are extracted from the voiceprint feature dataset, and the sample similarity between the bolt voiceprint instances is determined based on the voiceprint features of the bolt voiceprint instances.

[0143] Based on the sample similarity, the bolt voiceprint knowledge graph bolt voiceprint instance is extracted from the bolt voiceprint instance dataset to obtain the bolt voiceprint instance dataset.

[0144] Based on the bolt voiceprint instance dataset of the voiceprint knowledge graph, the bolt voiceprint instances are identified as elements to generate the voiceprint knowledge graph.

[0145] In this embodiment of the invention, exemplarily, the server first processes the voiceprint feature dataset, which contains voiceprint data of multiple bolt voiceprint instances. For each bolt voiceprint instance, the server extracts its unique voiceprint features, such as frequency distribution and amplitude variation. Next, the server uses a similarity algorithm (such as cosine similarity, Euclidean distance, etc.) to calculate the sample similarity between each pair of bolt voiceprint instances. For example, if there are two bolt voiceprint instances A and B, the server calculates the similarity between their voiceprint features, obtaining a specific numerical value representing the degree of similarity between A and B. After calculating the sample similarity between all bolt voiceprint instances, the server extracts bolt voiceprint instances from the voiceprint knowledge graph based on these similarity values. Specifically, the server sets a similarity threshold; only when the similarity between two bolt voiceprint instances exceeds this threshold are they considered to have a direct association in the voiceprint knowledge graph. In this way, the server can extract a new dataset from the original bolt voiceprint instance dataset, which only contains those bolt voiceprint instances that are directly associated in the voiceprint knowledge graph. Finally, the server uses the extracted bolt voiceprint instance dataset from the voiceprint knowledge graph to generate a voiceprint knowledge graph. In this graph, each bolt voiceprint instance is represented as a node (element), and the connections between nodes represent their similarity relationships. This voiceprint knowledge graph can be used for subsequent tasks such as voiceprint analysis and diffusion target value calculation. For example, by analyzing the connections and node attributes in the graph, the server can more accurately determine the classification or state of a bolt voiceprint instance. Through these steps, the server can construct a voiceprint knowledge graph based on the similarity of bolt voiceprint instances, which provides an important data structure and information foundation for subsequent target value diffusion and optimization.

[0146] In this embodiment of the invention, the step of determining the bolt voiceprint instance as an element to generate a voiceprint knowledge graph based on the bolt voiceprint instance dataset of the voiceprint knowledge graph can be implemented through the following example.

[0147] Obtain the association between the bolt soundprint instance and the bolt soundprint instance in the corresponding soundprint knowledge graph bolt soundprint instance dataset, and obtain the association influence factor of the bolt soundprint instance.

[0148] Based on the aforementioned correlation influencing factors, the bolt soundprint instance is identified as an element to generate an initial soundprint knowledge graph, and the initial soundprint knowledge graph is then balanced to obtain the final soundprint knowledge graph.

[0149] In this embodiment of the invention, for example, when processing bolt voiceprint instances, the server first analyzes the association between these instances and instances in the bolt voiceprint instance dataset of the voiceprint knowledge graph. This association can be based on various factors, such as the similarity of voiceprint features, temporal proximity, or other features that indicate a connection between two instances. For example, if a bolt voiceprint instance is highly similar in frequency features to a known anomalous instance in the graph, then there is a strong association between them. To quantify this association, the server calculates the association influence factor for each bolt voiceprint instance. This factor can be a value between 0 and 1, representing the strength of the association between the instance and instances in the graph. For example, if a bolt voiceprint instance is very similar to an instance in the graph, its association influence factor may be close to 1; if the similarity is low or there is no direct association, the factor value may be close to 0. After determining the association influence factor for each bolt voiceprint instance, the server uses this information to construct an initial voiceprint knowledge graph. In this knowledge graph, each bolt voiceprint instance is represented as a node, and the lines connecting nodes represent their relationships. The thickness or color intensity of the lines indicates the magnitude of the association's influence factor. Next, the server balances this initial voiceprint knowledge graph. The purpose of balancing is to optimize the graph's structure, ensuring the uniformity and accuracy of information distribution. This can include adjusting the connection weights between nodes, deleting weak connections, or adding necessary connections. For example, if the association influence factor between two nodes is very low, the server will choose to disconnect them; if a new strong association is found, a new connection will be added. After balancing, the server finally obtains an optimized voiceprint knowledge graph. This graph more accurately reflects the relationships between bolt voiceprint instances, providing a solid foundation for subsequent target value diffusion, optimization, and model training.

[0150] In this embodiment of the invention, obtaining the association between the bolt soundprint instance and the bolt soundprint instance in the soundprint knowledge graph bolt soundprint instance dataset, and obtaining the association influence factor of the bolt soundprint instance, can be implemented through the following example.

[0151] The similarity between the bolt voiceprint instance and the bolt voiceprint instance in the corresponding voiceprint knowledge graph bolt voiceprint instance dataset is extracted from the sample similarity.

[0152] Clustering is performed on the similarity of the target samples to obtain the association between the bolt soundprint instances and the bolt soundprint instances in the soundprint knowledge graph bolt soundprint instance dataset;

[0153] Based on the aforementioned correlation, the correlation influence factor of the bolt sound pattern instance is determined.

[0154] In this embodiment of the invention, for example, the server has already calculated the sample similarity between bolt soundprint instances and stored it in a similarity matrix. Now, the server needs to extract the target sample similarity between each bolt soundprint instance and existing bolt soundprint instances in the soundprint knowledge graph from this matrix. For example, suppose the server has a new bolt soundprint instance A, and bolt soundprint instances B, C, and D already exist in the soundprint knowledge graph. The server will extract the similarity values ​​between A and B, A and C, and A and D from the similarity matrix; these values ​​are the target sample similarities. After extracting the target sample similarities, the server will use a clustering algorithm (such as K-means, hierarchical clustering, etc.) to cluster these similarity values. The purpose of clustering is to group bolt soundprint instances with high similarity into one category, thereby revealing the association between them. Taking K-means clustering as an example, the server will set K cluster centers, and then assign each target sample similarity to the nearest cluster center. After multiple iterations of optimization, a stable clustering result is formed. In this way, the server can determine which bolt soundprint instances have a high degree of association based on the clustering results. After clustering, the server analyzes the correlation between each bolt soundprint instance and other instances in its cluster. The correlation influence factor is a metric used to quantify this correlation. For example, if bolt soundprint instance A has a high similarity to other instances (such as B and C) in its cluster, then A's correlation influence factor will be large. This factor can be a weighted average that considers the similarity between A and all other instances in the cluster. The server calculates the correlation influence factor for each bolt soundprint instance and stores it in the knowledge graph's metadata for later use. Through this process, the server can accurately quantify the correlation between bolt soundprint instances and existing instances in the soundprint knowledge graph, providing crucial information for subsequent graph construction, target value diffusion, and optimization tasks.

[0155] In this embodiment of the invention, the process of training the initial soundprint classification model based on the optimized bolt soundprint instance dataset can be implemented through the following example.

[0156] Based on the acoustic features and target values ​​of bolt acoustic instances in the optimized bolt acoustic instance dataset, the model parameters of the initial acoustic classification model are adjusted.

[0157] Based on the initial voiceprint classification model, feature extraction is performed on the bolt voiceprint instances in the optimized bolt voiceprint instance dataset to obtain the target voiceprint feature dataset.

[0158] This embodiment does not impose restrictions on the initial voiceprint classification model. Deep learning-based methods, such as d-vector, x-vector, ResNet, etc., can be used, or mainstream voiceprint recognition models, such as ECAPA-TDNN, can be used.

[0159] Based on the target acoustic signature feature dataset, the target value of the bolt acoustic signature instance is optimized;

[0160] Repeat the steps of adjusting the model parameters of the initial voiceprint classification model based on the voiceprint features and target values ​​of bolt voiceprint instances in the optimized bolt voiceprint instance dataset, until the initial voiceprint classification model meets the preset training termination condition, and obtain the trained target voiceprint classification model.

[0161] In this embodiment of the invention, exemplarily, the server first loads an optimized bolt acoustic signature instance dataset. Each instance in this dataset contains acoustic signature features and a target value (such as a classification label like "normal" or "abnormal"). Next, the server uses this data to train an initial acoustic signature classification model. During training, the model adjusts its internal parameters based on the acoustic signature features and target value of each instance to make the model's predictions closer to the true target value of the instance. For example, if the model incorrectly predicts an abnormal bolt acoustic signature instance as normal, the model will increase its sensitivity to abnormal features during parameter adjustment to better identify abnormal instances. After the initial adjustment of the model parameters, the server uses the adjusted initial acoustic signature classification model to extract features from the bolt acoustic signature instances in the dataset. During this process, the model extracts the most representative acoustic signature features from each instance, forming a new target acoustic signature feature dataset. This data can be used for further model optimization and target value adjustment. After obtaining the target acoustic signature feature dataset, the server compares and analyzes the model's predictions with the actual target values. If the server detects significant discrepancies between the predicted values ​​and the actual target values ​​for certain instances, it will optimize and adjust the target values ​​for these instances. For example, for instances misclassified by the model, the server will adjust their target labels to more accurately reflect the true state of the instances. The server will continuously repeat the above process of adjusting model parameters and extracting features until the performance of the initial voiceprint classification model reaches the preset training termination conditions. These conditions may include the model's accuracy, recall, F1 score reaching a certain threshold, or the maximum number of training epochs being reached. Once the model meets these conditions, the server will stop training and save the currently trained target voiceprint classification model for later use.

[0162] In this embodiment of the invention, the step of adjusting the model parameters of the initial soundprint classification model based on the soundprint features and target values ​​of bolt soundprint instances in the optimized bolt soundprint instance dataset can be implemented through the following example.

[0163] Based on the target value of the bolt sound pattern instance in the optimized bolt sound pattern instance dataset, determine the target value error parameter of the bolt sound pattern instance;

[0164] Based on the acoustic features of bolt acoustic instances in the optimized bolt acoustic instance dataset, the acoustic error parameter of the bolt acoustic instance is determined.

[0165] The target value error parameter and the voiceprint error parameter are integrated to obtain the integrated error parameter, and the model parameters of the initial voiceprint classification model are adjusted according to the integrated error parameter.

[0166] In this embodiment of the invention, for example, the server first traverses the optimized bolt acoustic signature instance dataset. For each bolt acoustic signature instance, the server compares its actual target value with the target value predicted by the initial acoustic signature classification model. For example, if the actual target value of a bolt acoustic signature instance is "abnormal," but the model predicts it as "normal," then the target value error parameter for this instance will be relatively large. The server quantifies this error by calculating the difference between the predicted and actual values, such as using mean squared error or cross-entropy loss, thereby obtaining the target value error parameter for each bolt acoustic signature instance. Next, the server analyzes the acoustic signature features of each bolt acoustic signature instance. It compares the differences between the features extracted by the model and the acoustic signature features of the instance itself. This difference may be due to inaccuracies or omissions in the feature extraction process. To quantify this difference, the server uses metrics such as Euclidean distance and cosine similarity between features to calculate the acoustic signature error parameter. For example, if a specific frequency component of a bolt acoustic signature instance is not accurately reflected in the features extracted by the model, then the acoustic signature error parameter for that instance will increase accordingly. After obtaining the target error parameter and the soundprint error parameter for each bolt soundprint instance, the server integrates these two parameters. The integration can be a simple weighted average or a more complex error fusion strategy designed based on the specific situation. The integrated error parameter, or integrated error parameter, more comprehensively reflects the model's overall performance in both predicting the target and extracting features. Finally, the server adjusts the parameters of the initial soundprint classification model based on this integrated error parameter. The purpose of this adjustment is to reduce the integrated error parameter, thereby improving the model's prediction accuracy. For example, if the model has a large error when predicting bolt soundprint instances of a specific category, the server will adjust the model weights or biases related to that category so that the model can better learn the features of that category in subsequent training iterations. Through this adjustment process, the server aims to obtain a soundprint classification model with better performance.

[0167] In this embodiment of the invention, the step of determining the acoustic error parameter of the bolt acoustic instance based on the acoustic features of the bolt acoustic instances in the optimized bolt acoustic instance dataset can be implemented through the following example.

[0168] Based on the target value of the bolt sound pattern instance in the optimized bolt sound pattern instance dataset, perform a type recognition operation on the bolt sound pattern instance to obtain a subset of bolt sound pattern instance data corresponding to each target value.

[0169] Based on the acoustic features of bolt acoustic features in the bolt acoustic feature instance subset, the target acoustic features corresponding to the bolt acoustic feature instance subset are determined.

[0170] The acoustic features of the bolt acoustic feature instance and the target acoustic features corresponding to the bolt acoustic feature instance data subset are integrated to obtain the acoustic error parameter of the bolt acoustic feature instance.

[0171] In this embodiment of the invention, exemplarily, the server first traverses the optimized bolt acoustic signature instance dataset, identifying the target value for each bolt acoustic signature instance. These target values ​​can represent different states of the bolt, such as "normal," "loose," or "damaged." Next, the server classifies the bolt acoustic signature instances according to these target values, forming different subsets of bolt acoustic signature instance data. For example, all bolt acoustic signature instances with a target value of "normal" are grouped into one subset, while all those with a target value of "loose" are grouped into another subset. After obtaining each subset of bolt acoustic signature instance data, the server analyzes the acoustic signature features within each subset. For each subset, the server calculates the mean, median, or other statistics of all instance acoustic signature features, using these as the target acoustic signature features for that subset. These target acoustic signature features represent the typical characteristics of the bolt acoustic signature instances in that subset. For example, in the subset of bolt acoustic signature instances in the "normal" state, the server might find a higher acoustic signature energy distribution in a certain frequency band, which would become one of the target acoustic signature features for that subset. Finally, the server compares the acoustic features of each bolt acoustic signature instance with the target acoustic features of its corresponding subset of data. This difference can be measured by calculating the similarity or distance between the two, such as using Euclidean distance or Pearson correlation coefficient. The server quantifies this difference as an acoustic error parameter, which is used as a basis for subsequent model parameter adjustments. For example, if the acoustic features of a bolt acoustic signature instance differ significantly from the target acoustic features of its "loose" state subset, the acoustic error parameter for that instance will increase accordingly, indicating that the model may have a large error in extracting the features of that instance, requiring appropriate adjustments.

[0172] In this embodiment of the invention, the step of integrating the acoustic features of the bolt acoustic pattern instance and the target acoustic features corresponding to the bolt acoustic pattern instance data subset to obtain the acoustic error parameter of the bolt acoustic pattern instance can be implemented through the following example.

[0173] Based on the acoustic features of the bolt acoustic features, determine the feature difference between bolt acoustic features in the bolt acoustic feature data subset and obtain the first feature difference.

[0174] Based on the target feature difference amount corresponding to the bolt soundprint instance data subset, determine the feature difference amount between the bolt soundprint instance data subsets, and obtain the second feature difference amount;

[0175] Determine the feature difference between the first feature difference and the second feature difference, obtain the third feature difference, and integrate the third feature difference with the critical difference value to obtain the integration difference degree;

[0176] If the integration difference is greater than a preset difference threshold, the average feature difference of the integration difference is determined to correct the target acoustic feature, so as to obtain the acoustic error parameter of the bolt acoustic instance.

[0177] In this embodiment of the invention, for example, the server first calculates the feature difference between each bolt acoustic signature instance in each subset of bolt acoustic signature data. This can be done by comparing the acoustic signature feature vectors between instances, for example, using methods such as cosine similarity or Euclidean distance to measure the difference between features. In this way, the server can obtain the average difference between instances within the subset, which is the first feature difference. This difference reflects the consistency of acoustic signature features between instances within the subset. Next, the server compares the target acoustic signature feature differences between different subsets of bolt acoustic signature data. This involves calculating the difference in target acoustic signature features between different subsets, which can be calculated using a method similar to the first step. In this way, the server can quantify the feature differences between different state subsets, which is the second feature difference. This difference reveals the acoustic signature feature discriminability between different state subsets. The server then calculates the difference between the first and second feature difference, which can be understood as a comparison between differences within and outside the subset, i.e., the third feature difference. This difference reflects the relative relationship between the consistency of acoustic signature features within a subset and the discriminability of acoustic signature features between subsets. The server then compares and integrates this third feature difference with a preset critical difference value to obtain an integrated difference score. This integrated difference score comprehensively considers the differences in voiceprint features within and outside the subset. If the calculated integrated difference score exceeds a preset difference score threshold, it means that the differences in voiceprint features within the subset are not significant enough compared to the differences between subsets, leading to inaccurate classification. In this case, the server uses the average feature difference of the integrated difference score to correct the target voiceprint features to improve the discriminative power of the voiceprint features. This correction helps the model better identify bolt voiceprint instances in different states. The difference between the corrected target voiceprint features and the original voiceprint features is the voiceprint error parameter of the bolt voiceprint instance, which will be used for subsequent model parameter adjustment and optimization.

[0178] In this embodiment of the invention, the step of inputting the current emergency response instruction into a pre-trained emergency scenario recognition model to obtain an emergency response strategy matching set can be implemented through the following example.

[0179] For the voiceprint to be analyzed in the current emergency response instructions, the voiceprint matching model is used to search for reference bolt voiceprints that meet the matching criteria to obtain the matching bolt voiceprints. The voiceprint matching model is obtained through prior training.

[0180] Based on the matched bolt soundprint, and targeting the emergency handling needs in the current emergency handling instructions, a semantic analysis model is used to search for reference emergency handling needs that meet the matching criteria to obtain the matched emergency handling instructions. The semantic analysis model is obtained through prior training.

[0181] The emergency scenario recognition model is used to classify the current emergency handling instruction and the matching emergency handling instruction into scenario types to obtain the scenario type of the current emergency handling instruction and the scenario type of each matching emergency handling instruction. The emergency scenario recognition model is trained using the processing requirements with scenario labels configured in past data as samples.

[0182] The matching emergency response instructions that are consistent with the scenario type of the current emergency response instruction are used as the emergency response strategy matching set.

[0183] In this embodiment of the invention, for example, the server receives a current emergency handling instruction, which includes a bolt soundprint to be analyzed. The server first uses a pre-trained soundprint matching model to match the soundprint to be analyzed. The soundprint matching model searches its built-in bolt soundprint database for reference bolt soundprints that match the soundprint to be analyzed according to preset matching criteria. For example, if the characteristics of the soundprint to be analyzed are highly similar to the soundprint characteristics of a loose bolt in the database, then this loose bolt soundprint will be selected as the matching bolt soundprint. After determining the matching bolt soundprint, the server, based on the bolt state reflected by the soundprint (e.g., loose, broken, etc.), uses a pre-trained semantic analysis model to search for emergency handling requirements corresponding to this state in the current emergency handling instruction. The semantic analysis model analyzes the text content in the handling instruction to find emergency handling requirements related to the state of the matching bolt soundprint. For example, if the matching bolt soundprint indicates a loose bolt, the semantic analysis model will search the handling instructions for requirements on how to handle the loose bolt and find the corresponding matching emergency handling instruction. The server will then use an emergency scenario recognition model trained on past data to classify the current emergency handling instruction and matching emergency handling instructions into different scenario types. This model categorizes the instructions based on their content, such as bolt condition and processing requirements, into different scenario types, such as "emergency repair" and "preventative maintenance." For example, if the current emergency handling instruction mentions a bolt breakage problem requiring immediate attention, it will be classified as an "emergency repair" scenario. After classifying the scenarios, the server will filter out matching emergency handling instructions that match the current emergency handling instruction's scenario type. These instructions will provide the server with a basis for matching emergency handling strategies. For example, if the current emergency handling instruction is classified as an "emergency repair" scenario, the server will find all matching emergency handling instructions also classified as this scenario type and refer to the processing strategies in these instructions to formulate an emergency handling plan for the current problem.

[0184] In this embodiment of the invention, the step of searching for reference emergency handling requirements that meet the matching criteria based on the matching bolt soundprint and the emergency handling requirements in the current emergency handling instructions can be implemented through the following example.

[0185] Search for historical emergency response requirements related to the matching bolt acoustic signature;

[0186] For the emergency handling needs in the current emergency handling instructions, reference emergency handling needs with a matching degree exceeding the first semantic matching degree threshold are searched in the historical emergency handling needs through a semantic analysis model to obtain the matching emergency handling instructions.

[0187] In this embodiment of the invention, exemplarily, after identifying a bolt soundprint that matches the bolt soundprint in the current emergency handling instruction (i.e., a matching bolt soundprint), the server begins searching for historical emergency handling requests associated with this matching bolt soundprint. Specifically, the server accesses its internal database or external data storage system to find past emergency handling measures taken for the same or similar bolt soundprints (i.e., bolt condition or problem). This historical data may include previous maintenance records, fault handling reports, or user feedback. For example, if the matching bolt soundprint indicates a loose bolt problem, the server searches for all past emergency handling requests for the loose bolt problem, such as tightening the bolt or replacing the bolt. After obtaining the specific emergency handling request in the current emergency handling instruction (e.g., tightening the loose bolt), the server uses a pre-trained semantic analysis model to search for reference emergency handling requests in historical emergency handling requests that are semantically highly matched to the current request. This search process takes into account factors such as the textual description of the emergency handling request and the similarity of the handling actions. To find highly relevant reference requests, the server sets a first semantic matching threshold (e.g., 80% matching degree). Only when the semantic match between a historical emergency handling requirement and the current emergency handling requirement exceeds this threshold will it be considered a valid reference. This ensures that the found reference requirements are closely related to the current problem, thereby improving the accuracy and efficiency of emergency handling. For example, if the requirement in the current emergency handling instruction is "tighten loose bolts," the server will search for similar handling requirements in the history, such as "tighten loose bolts" or "resolve bolt looseness issues," and calculate their semantic match with the current requirement. Only requirements with a match exceeding a preset threshold will be selected as references. Through the above search and matching process, the server will eventually obtain a set of reference emergency handling requirements that highly match the current emergency handling instruction. These reference requirements constitute the matched emergency handling instructions, providing the server with specific handling methods and strategies for the current bolt problem. The server can formulate or optimize the current emergency handling plan based on these matched emergency handling instructions. For example, after finding multiple historical emergency handling requirements that highly match the requirement of "tightening loose bolts," the server will derive a comprehensive matched emergency handling instruction, including steps such as using specific tools and methods to tighten the bolts and checking the bolt's tightness. This instruction will directly guide on-site operators in carrying out effective emergency response.

[0188] In this embodiment of the invention, the step of searching for reference emergency handling requirements that meet the matching criteria based on the matching bolt soundprint and the emergency handling requirements in the current emergency handling instructions can be implemented through the following example.

[0189] In response to the emergency response needs in the current emergency response instructions, reference emergency response needs with a matching degree exceeding the second semantic matching degree threshold are searched in the preset archived logs using a semantic analysis model to obtain the matching response needs;

[0190] Based on the matching bolt sound pattern and the matching processing requirements, an emergency matching processing instruction is obtained.

[0191] In an embodiment of the invention, exemplarily, the server receives a current emergency handling instruction, which includes an emergency handling requirement, such as "urgently repair loose bolts." To find reference emergency handling requirements that match this requirement, the server uses its built-in semantic analysis model to search through preset archived logs. These archived logs may contain records and solutions for handling similar problems in the past. The semantic analysis model analyzes the text data in the logs to find handling requirements that are semantically similar to the requirement of "urgently repair loose bolts." To ensure the accuracy of the found reference requirements, the server sets a second semantic matching threshold, such as 75%. Only when a log entry's semantic matching degree with the current emergency handling requirement exceeds this threshold is it considered a matched reference emergency handling requirement. For example, if an entry in the archived log records the operation of "quickly tightening loose bolts to prevent equipment failure," then this entry may have a high semantic matching degree with "urgently repair loose bolts" and is thus selected as a matched reference requirement. Through the search and matching process in the previous step, the server obtains a set of reference emergency handling requirements that highly match the current emergency handling requirement. These matched handling requirements provide the server with possible methods and strategies for solving the current problem. In this example, "quickly tightening loose bolts to prevent equipment failure" is a matched processing requirement. After obtaining the matched bolt sound pattern and the matched processing requirement, the server further integrates this information to generate a specific matched emergency handling instruction. This instruction combines the current state of the bolt (reflected by the matched bolt sound pattern) and known effective handling methods (provided by the matched processing requirement) to provide operators with a clear and feasible emergency handling solution. For example, the server might generate a matched emergency handling instruction like this: "Based on bolt sound pattern analysis, a bolt loosening problem has been identified. Please refer to historical successful cases and adopt the 'quickly tightening loose bolts to prevent equipment failure' handling method to ensure safe equipment operation." Such instructions directly guide operators on how to perform emergency handling, improving processing efficiency and accuracy.

[0192] In this embodiment of the invention, the following implementation methods are also provided.

[0193] Acquire voiceprint samples for analysis;

[0194] Based on the sample acoustic prints to be analyzed, a target bolt acoustic print dataset generated from reference bolt acoustic prints and a non-target bolt acoustic print dataset generated from non-reference bolt acoustic prints are generated.

[0195] Based on the target bolt acoustic print dataset and the non-target bolt acoustic print dataset, the initial acoustic print matching model is pre-trained using a self-supervised strategy to obtain the pre-trained acoustic print matching model.

[0196] Based on the current emergency response instructions and business scenario, the previously trained voiceprint matching model is integrated and learned to obtain a voiceprint matching model.

[0197] In this embodiment of the invention, for example, to train and optimize the voiceprint matching model, the server first needs to acquire a certain number of sample voiceprints to be analyzed. These samples can be extracted from historical databases or obtained through field collection. For example, the server can extract audio files containing bolt voiceprints as samples from past mechanical equipment maintenance records. These samples should contain various types of bolt voiceprints to reflect the sound characteristics of bolts in different states. After acquiring the sample voiceprints to be analyzed, the server will create a dataset. First, the server will select voiceprints with obvious bolt characteristics (such as loose bolts, broken bolts, etc.) as the target bolt voiceprint dataset. These data will serve as positive samples during model training. At the same time, the server will also generate a non-target bolt voiceprint dataset, which contains voiceprints without bolt characteristics or other types of voiceprints, as negative samples. These two datasets will be used together for model training. After the datasets are prepared, the server will use a self-supervised strategy to pre-train the initial voiceprint matching model. Self-supervised learning is a method of learning using the inherent structure of the data itself, which does not require a large amount of manually labeled data. In this process, the server will teach the model how to distinguish between target bolt voiceprints and non-target bolt voiceprints. Through continuous iteration and optimization, the model gradually learns to extract bolt features from complex voiceprint signals and accurately identify the target bolt's voiceprint. After completing this training phase, the server obtains a pre-trained voiceprint matching model. Following this model, the server performs further ensemble learning based on the current emergency response instructions' business context. For example, if the current business context involves the rapid identification and handling of a specific type of bolt fault, the server collects more bolt voiceprint data related to this fault type and adds it to the training set. Then, the server uses ensemble learning to combine the pre-trained voiceprint matching model with this new data to improve the model's recognition performance in specific business scenarios. Through this process, the server ultimately obtains a more accurate and efficient voiceprint matching model to meet the needs of the current emergency response instructions.

[0198] In this embodiment of the invention, the generation of a target bolt acoustic signature dataset generated from a reference bolt acoustic signature and a non-target bolt acoustic signature dataset generated from a non-reference bolt acoustic signature based on the sample acoustic signature to be analyzed can be performed through the following example.

[0199] Multiple voiceprint conversion strategies are used to perform voiceprint conversion operations on the sample voiceprint to be analyzed, resulting in multiple converted voiceprints.

[0200] Multiple target bolt voiceprint datasets were obtained by using various voiceprint conversion strategies on the same sample voiceprint to be analyzed.

[0201] Based on the voiceprints of different samples to be analyzed, multiple non-target bolt voiceprint datasets are generated by applying arbitrary voiceprint conversion strategies.

[0202] In this embodiment of the invention, for example, after the server acquires a batch of sample voiceprints to be analyzed, it applies multiple voiceprint conversion strategies to these voiceprints in order to enhance the generalization ability of the model. For example, the server may use conversion strategies such as changing the playback speed of the audio, increasing background noise, and adjusting the volume or pitch of the audio. Through these strategies, the same sample voiceprint to be analyzed can be converted into multiple versions, each representing the bolt voiceprint under different environments or conditions. In this way, the model can better adapt to various real-world situations during learning. For the same sample voiceprint to be analyzed, the server will obtain a series of converted voiceprints after performing multiple voiceprint conversions. Although these converted voiceprints differ in audio characteristics, they all originate from the same original voiceprint, and therefore they all retain the specific characteristics of bolts. The server will classify these converted voiceprints into the target bolt voiceprint dataset. This dataset will be used to train the model to recognize the specific sound features of bolts. In addition to converting the same sample, the server will also perform voiceprint conversion on different sample voiceprints to be analyzed. These different samples may contain other types of voiceprints, such as the sounds of other mechanical parts, environmental noise, etc. The converted voiceprints obtained after these voiceprints are transformed will not be classified as bolt voiceprints, but will be added to the non-target bolt voiceprint dataset. This dataset will be used to train the model to distinguish bolt voiceprints from other types of voiceprints, improving the model's discrimination ability. For example, the server has a sample library containing various sounds of mechanical equipment. It selects voiceprints that do not contain bolt features for transformation, such as bearing rotation sounds and motor running sounds. These transformed voiceprints will be added to the non-target bolt voiceprint dataset as negative samples during model training. In this way, the model can learn how to accurately distinguish bolt voiceprints from other similar sounds, thereby improving the recognition accuracy in practical applications.

[0203] In this embodiment of the invention, the process of integrating and learning the previously trained voiceprint matching model based on the current emergency response instruction business scenario to obtain the voiceprint matching model can be implemented through the following example.

[0204] Obtain the target voiceprint to be analyzed corresponding to the business scenario of the current emergency response instruction;

[0205] Based on the target voiceprint to be analyzed, generate target ensemble learning samples and non-target ensemble learning samples;

[0206] Based on the target ensemble learning samples and non-target ensemble learning samples, the pre-trained voiceprint matching model is ensemble-learned using a self-supervised strategy to obtain the voiceprint matching model.

[0207] In an embodiment of the invention, for example, the server receives a current emergency handling instruction concerning a specific type of bolt loosening problem that requires rapid location and handling. To address this business scenario, the server acquires target voiceprints to be analyzed from relevant mechanical equipment monitoring systems or on-site recording equipment. These voiceprints are collected under this business scenario, and therefore are likely to contain characteristic sounds of bolt loosening. After acquiring the target voiceprints, the server processes them to generate two types of samples: target ensemble learning samples and non-target ensemble learning samples. Target ensemble learning samples are those voiceprints that explicitly contain bolt loosening features; these will be used as positive samples to train the model to recognize the sound of bolt loosening. Non-target ensemble learning samples are those voiceprints that do not contain bolt loosening features or contain other mechanical sounds; these will be used as negative samples to help the model distinguish sounds other than bolt loosening. With the target and non-target ensemble learning samples, the server employs a self-supervised strategy to further ensemble learn the previously trained voiceprint matching model. In this process, the model continuously learns the sound characteristics of loose bolts from the target ensemble learning samples and learns how to distinguish other sounds from non-target ensemble learning samples. Through repeated training and adjustments, the model's recognition ability gradually improves. Ultimately, the server obtains an ensemble learning-optimized voiceprint matching model that can more accurately identify the sound of loose bolts in the current emergency response instruction scenario, providing strong support for quickly locating and handling problems.

[0208] In this embodiment of the invention, the following implementation methods are also provided.

[0209] Obtain emergency sample processing requirements;

[0210] Based on the aforementioned emergency processing requirements for samples, target sample processing requirements generated from reference emergency processing requirements and non-target sample processing requirements generated from non-reference emergency processing requirements are generated.

[0211] Based on the target sample processing requirements and the non-target sample processing requirements, a training process is performed on the initial semantic analysis model using a self-supervised strategy to obtain the semantic analysis model.

[0212] In this embodiment of the invention, for example, to train and optimize the semantic analysis model, the server first needs to collect a series of sample emergency handling requirements. These requirements can be obtained from historical records, work logs, or an emergency handling database. For example, the server can collect relevant records of various emergency situations handled by the company over the past few years, which typically contain detailed emergency handling requirements and solutions. After acquiring the sample emergency handling requirements, the server categorizes these requirements. First, the server selects those clear and specific emergency handling requirements as target sample handling requirements, which will serve as positive samples during model training. For example, clear task requirements such as "urgently replace a faulty hard drive" or "quickly repair network failures" can be classified as target sample handling requirements. At the same time, the server also generates non-target sample handling requirements, which can be vague, irrelevant, or erroneous emergency handling requirements. For example, broad or vague requirements such as "check server status" or "optimize system performance" can be used as non-target sample handling requirements. These non-target sample handling requirements will help the model learn to distinguish and filter invalid or irrelevant information. After preparing the target and non-target sample handling requirements, the server will use these samples to train the initial semantic analysis model. During training, the model learns to accurately understand and identify key information in the processing requests of target samples, such as task type, urgency, and required resources. Simultaneously, the model learns to distinguish and ignore invalid or irrelevant information in the processing requests of non-target samples. Through repeated iterations and optimization of the model's parameters and structure, the server eventually obtains a trained semantic analysis model. This model can automatically analyze and understand key information in emergency processing requests, thereby providing accurate and efficient decision support for subsequent emergency responses. For example, when the server receives a new emergency processing request, the trained semantic analysis model can quickly identify the type and urgency of the request and recommend appropriate processing solutions and required resources.

[0213] In this embodiment of the invention, the generation of target sample processing requirements generated from reference emergency processing requirements and non-target sample processing requirements generated from non-reference emergency processing requirements based on the sample emergency processing requirements can be implemented through the following examples.

[0214] The sample emergency processing requirements are obtained by performing semantic transformation operations on the sample emergency processing requirements through multiple semantic transformation strategies.

[0215] Based on the emergency processing needs of the same sample, the semantic transformation processing needs obtained by applying multiple semantic transformation strategies are used to generate the target sample processing needs.

[0216] Based on the emergency processing needs of different samples, the semantic transformation processing needs obtained by applying arbitrary semantic transformation strategies generate non-target sample processing needs.

[0217] In this embodiment of the invention, for example, after acquiring a series of sample emergency handling requests, the server employs various semantic transformation strategies to transform these requests. For instance, for an original emergency handling request, "Server crash, urgent recovery required," the server uses strategies such as synonym substitution, sentence transformation, and adding or deleting detailed information for semantic transformation. Through these transformations, semantically transformed processing requests such as "Urgently repair the faulty server" and "Server malfunction, please quickly restart and recover" can be obtained. For the same original emergency handling request, after the server performs multiple semantic transformations, it will obtain multiple processing requests with different expressions but similar meanings. These requests all point to the same goal, namely, restoring the normal operation of the server. Therefore, the server will classify these requests obtained through different semantic transformation strategies into target sample processing requests. For example, for the original request "Server crash, urgent recovery required," the transformed "Urgently repair the faulty server" and "Server malfunction, please quickly restart and recover" will all be considered target sample processing requests. The server will also perform semantic transformations on different original emergency handling requests. Since these original requests themselves point to different goals, even after semantic transformation, the core intent they express is still different. These transformed requirements are categorized as non-target sample processing requirements. For example, the original requirement "Network connection error, needs to be checked and repaired" can be semantically transformed into "Please check and resolve network connection problems." This transformed requirement differs in core intent from the target sample processing requirement related to "server crash," and therefore is considered a non-target sample processing requirement. Through these steps, the server can generate a rich set of target and non-target sample processing requirements, providing a sufficient data foundation for subsequent semantic analysis model training.

[0218] This invention also provides a tower bolt processing system based on a voiceprint classification model, the system comprising:

[0219] The acquisition module is used to obtain the voiceprint to be analyzed from the audio of the tower monitoring.

[0220] The model training module is used to input the voiceprint to be analyzed into a pre-trained target voiceprint classification model to obtain the category classification result of the voiceprint to be analyzed. The process of obtaining the target voiceprint classification model includes: constructing an initial voiceprint classification model and a bolt voiceprint instance dataset, optimizing the bolt voiceprint instance dataset, and training the initial voiceprint classification model based on the optimized bolt voiceprint instance dataset. The optimization process of the bolt voiceprint instance dataset includes:

[0221] The similarity calculation module is used to determine the sample similarity between bolt soundprint instances. Based on the sample similarity, the module extracts the target sample similarity between the bolt soundprint instances and the corresponding bolt soundprint instances in the bolt soundprint instance dataset of the soundprint knowledge graph, thereby determining the association influence factor of the bolt soundprint instances. Based on the association factor, the module obtains the soundprint knowledge graph. Based on the soundprint knowledge graph, the module determines the degree of difference between the bolt soundprint instances. Based on the weight corresponding to the degree of difference, the module obtains the diffusion target value information of the bolt soundprint instances. Based on the diffusion target value information of the bolt soundprint instances, the module obtains the optimized bolt soundprint instance dataset.

[0222] The result processing module receives an emergency processing request for the voiceprint to be analyzed when the classification result is characterized as an abnormal voiceprint result.

[0223] Construct current emergency response instructions based on the voiceprint to be analyzed and the emergency response requirements;

[0224] The current emergency response instructions are input into a pre-trained emergency scenario recognition model to obtain an emergency response strategy matching set.

[0225] The target emergency strategy is selected from the preset emergency strategy database based on the emergency response strategy matching set.

[0226] This invention also provides a computer device 100, which includes a processor and a non-volatile memory storing computer instructions. When the computer instructions are executed by the processor, the computer device 100 executes the aforementioned method for processing iron tower bolts based on artificial intelligence and voiceprint. Figure 2 As shown, Figure 2 This is a structural block diagram of a computer device 100 provided in an embodiment of the present invention. The computer device 100 includes a memory 111, a processor 112, and a communication unit 113. To enable data transmission or interaction, the memory 111, processor 112, and communication unit 113 are electrically connected to each other directly or indirectly. For example, these components can be electrically connected to each other through one or more communication buses or signal lines.

[0227] For illustrative purposes, the foregoing description has been made with reference to specific embodiments. However, the foregoing illustrative discussions are not intended to be exhaustive or to limit the present disclosure to the precise forms disclosed. Numerous modifications and variations are possible in accordance with the foregoing teachings. These embodiments were chosen and described in order to best illustrate the principles of the present disclosure and its practical application, thereby enabling those skilled in the art to best utilize the disclosure and to employ various embodiments with different modifications to suit a particular intended application.

Claims

1. A method for processing iron tower bolts based on a voiceprint classification model, characterized in that, The method includes: The voiceprint to be analyzed is obtained from the audio monitored by the tower; The voiceprint to be analyzed is input into a pre-trained target voiceprint classification model to obtain the category classification result of the voiceprint to be analyzed; the process of obtaining the target voiceprint classification model includes: constructing an initial voiceprint classification model and a bolt voiceprint instance dataset, optimizing the bolt voiceprint instance dataset, and training the initial voiceprint classification model based on the optimized bolt voiceprint instance dataset. The optimization process of the bolt voiceprint instance dataset includes: The sample similarity between bolt soundprint instances is determined. Based on the sample similarity, the target sample similarity between the bolt soundprint instances and the corresponding bolt soundprint instances in the bolt soundprint instance dataset of the soundprint knowledge graph is extracted, thereby determining the association influence factor of the bolt soundprint instances. The soundprint knowledge graph is obtained based on the association factor. The degree of difference between the bolt soundprint instances is determined based on the soundprint knowledge graph. The diffusion target value information of the bolt soundprint instances is obtained based on the weight corresponding to the degree of difference. The optimized bolt soundprint instance dataset is obtained based on the diffusion target value information of the bolt soundprint instances. When the classification result is characterized as an abnormal voiceprint result, an emergency processing request for the voiceprint to be analyzed is received; Construct current emergency response instructions based on the voiceprint to be analyzed and the emergency response requirements; The current emergency response instructions are input into a pre-trained emergency scenario recognition model to obtain an emergency response strategy matching set. The target emergency strategy is selected from the preset emergency strategy database based on the emergency response strategy matching set.

2. The method according to claim 1, characterized in that, The process of obtaining the target voiceprint classification model further includes: Obtain a bolt acoustic signature instance dataset, which includes at least one bolt acoustic signature instance configured with an initial target value; Based on the initial voiceprint classification model, feature extraction is performed on the bolt voiceprint instances in the bolt voiceprint instance dataset to obtain the voiceprint feature dataset. The voiceprint features corresponding to each bolt voiceprint instance are extracted from the voiceprint feature dataset, and the sample similarity between the bolt voiceprint instances is determined based on the voiceprint features of the bolt voiceprint instances. Based on the sample similarity, the bolt voiceprint knowledge graph bolt voiceprint instance is extracted from the bolt voiceprint instance dataset to obtain the bolt voiceprint instance dataset. The similarity between the bolt voiceprint instance and the bolt voiceprint instance in the corresponding voiceprint knowledge graph bolt voiceprint instance dataset is extracted from the sample similarity. Clustering is performed on the similarity of the target samples to obtain the association between the bolt soundprint instances and the bolt soundprint instances in the soundprint knowledge graph bolt soundprint instance dataset; Based on the aforementioned correlation, determine the correlation influence factor of the bolt sound pattern instance; Based on the aforementioned correlation influencing factors, the bolt soundprint instance is determined as an element to generate an initial soundprint knowledge graph, and the initial soundprint knowledge graph is balanced to obtain the soundprint knowledge graph. Based on the initial target value of the bolt acoustic print instance, generate initial target value information corresponding to the bolt acoustic print instance dataset; the initial target value information includes an initial target value vector corresponding to each bolt acoustic print instance; The degree of difference between the bolt acoustic print instances is determined based on the acoustic print knowledge graph. Obtain the voiceprint correlation weight corresponding to the degree of difference, and assign weights to the initial target value vector of the bolt voiceprint instance according to the voiceprint correlation weight. Cluster the initial target value vector after weight assignment to obtain the diffusion target value information of the bolt soundprint instance; the diffusion target value information is obtained by spreading the initial target value of each bolt soundprint instance in the previously constructed soundprint knowledge graph, that is, each bolt soundprint instance node will pass its target value to other nodes connected to it, so as to obtain the target value information after considering the influence of its neighboring nodes. Obtain the diffusion target value vector of the bolt acoustic pattern instance from the diffusion target value information; Extract the target value component with the largest target value from the diffusion target value vector; The index information of the target value component is determined in the diffusion target value vector; Obtain the candidate target value corresponding to the index information, and determine the candidate target value as the diffusion target value of the bolt sound pattern instance; Perform a comparison operation between the diffusion target value and the initial target value configured for the corresponding bolt acoustic signature instance; When the diffusion target value is inconsistent with the initial target value, the bolt sound pattern instance is determined to be the target bolt sound pattern instance to be optimized. The initial target value of the target bolt acoustic signature instance is changed to the corresponding diffusion target value to obtain the optimized bolt acoustic signature instance dataset; The initial voiceprint classification model is trained based on the optimized bolt voiceprint instance dataset to obtain the trained target voiceprint classification model.

3. The method according to claim 2, characterized in that, The training process for the initial soundprint classification model based on the optimized bolt soundprint instance dataset includes: Based on the acoustic features and target values ​​of bolt acoustic instances in the optimized bolt acoustic instance dataset, the model parameters of the initial acoustic classification model are adjusted. Based on the initial voiceprint classification model, feature extraction is performed on the bolt voiceprint instances in the optimized bolt voiceprint instance dataset to obtain the target voiceprint feature dataset. Based on the target acoustic signature feature dataset, the target value of the bolt acoustic signature instance is optimized; Repeat the steps of adjusting the model parameters of the initial voiceprint classification model based on the voiceprint features and target values ​​of bolt voiceprint instances in the optimized bolt voiceprint instance dataset, until the initial voiceprint classification model meets the preset training termination condition, and obtain the trained target voiceprint classification model.

4. The method according to claim 3, characterized in that, The step of adjusting the model parameters of the initial soundprint classification model based on the soundprint features and target values ​​of bolt soundprint instances in the optimized bolt soundprint instance dataset includes: Based on the target value of the bolt sound pattern instance in the optimized bolt sound pattern instance dataset, determine the target value error parameter of the bolt sound pattern instance; Based on the acoustic features of bolt acoustic instances in the optimized bolt acoustic instance dataset, the acoustic error parameter of the bolt acoustic instance is determined. The target value error parameter and the voiceprint error parameter are integrated to obtain the integrated error parameter, and the model parameters of the initial voiceprint classification model are adjusted according to the integrated error parameter.

5. The method according to claim 4, characterized in that, The step of determining the acoustic error parameter of the bolt acoustic instance based on the acoustic features of the bolt acoustic instances in the optimized bolt acoustic instance dataset includes: Based on the target value of the bolt sound pattern instance in the optimized bolt sound pattern instance dataset, perform a type recognition operation on the bolt sound pattern instance to obtain a subset of bolt sound pattern instance data corresponding to each target value. Based on the acoustic features of bolt acoustic features in the bolt acoustic feature instance subset, the target acoustic features corresponding to the bolt acoustic feature instance subset are determined. The acoustic features of the bolt acoustic feature instance and the target acoustic features corresponding to the bolt acoustic feature instance data subset are integrated to obtain the acoustic error parameter of the bolt acoustic feature instance.

6. The method according to claim 5, characterized in that, The step of integrating the acoustic features of the bolt acoustic signature instance and the target acoustic features corresponding to the bolt acoustic signature instance data subset to obtain the acoustic error parameter of the bolt acoustic signature instance includes: Based on the acoustic features of the bolt acoustic features, determine the feature difference between bolt acoustic features in the bolt acoustic feature data subset and obtain the first feature difference. Based on the target feature difference amount corresponding to the bolt soundprint instance data subset, determine the feature difference amount between the bolt soundprint instance data subsets, and obtain the second feature difference amount; Determine the feature difference between the first feature difference and the second feature difference, obtain the third feature difference, and integrate the third feature difference with the critical difference value to obtain the integration difference degree; If the integration difference is greater than a preset difference threshold, the average feature difference of the integration difference is determined to correct the target acoustic feature, so as to obtain the acoustic error parameter of the bolt acoustic instance.

7. The method according to claim 1, characterized in that, The step of inputting the current emergency response instructions into a pre-trained emergency scenario recognition model to obtain an emergency response strategy matching set includes: For the voiceprint to be analyzed in the current emergency response instructions, the voiceprint matching model is used to search for reference bolt voiceprints that meet the matching criteria to obtain the matching bolt voiceprints. The voiceprint matching model is obtained through prior training. Search for historical emergency response requirements related to the matching bolt acoustic signature; For the emergency handling needs in the current emergency handling instructions, reference emergency handling needs with a matching degree exceeding the first semantic matching degree threshold are searched in the historical emergency handling needs through a semantic analysis model to obtain matching emergency handling instructions. The semantic analysis model is obtained through prior training. The emergency scenario recognition model is used to classify the current emergency handling instruction and the matching emergency handling instruction into scenario types to obtain the scenario type of the current emergency handling instruction and the scenario type of each matching emergency handling instruction. The emergency scenario recognition model is trained using the processing requirements with scenario labels configured in past data as samples. The matching emergency response instructions that are consistent with the scenario type of the current emergency response instruction are used as the emergency response strategy matching set.

8. The method according to claim 7, characterized in that, The process of obtaining the voiceprint matching model includes: Acquire voiceprint samples for analysis; Multiple voiceprint conversion strategies are used to perform voiceprint conversion operations on the sample voiceprint to be analyzed, resulting in multiple converted voiceprints. Multiple target bolt voiceprint datasets were obtained by using various voiceprint conversion strategies on the same sample voiceprint to be analyzed. Based on the voiceprints of different samples to be analyzed, multiple converted voiceprints are generated into non-target bolt voiceprint datasets by applying arbitrary voiceprint conversion strategies. Based on the target bolt acoustic print dataset and the non-target bolt acoustic print dataset, the initial acoustic print matching model is pre-trained using a self-supervised strategy to obtain the pre-trained acoustic print matching model. Obtain the target voiceprint to be analyzed corresponding to the business scenario of the current emergency response instruction; Based on the target voiceprint to be analyzed, generate target ensemble learning samples and non-target ensemble learning samples; Based on the target ensemble learning samples and non-target ensemble learning samples, the pre-trained voiceprint matching model is ensemble-learned using a self-supervised strategy to obtain the voiceprint matching model.

9. The method according to claim 7, characterized in that, The process of obtaining the semantic analysis model includes: Obtain emergency sample processing requirements; The sample emergency processing requirements are obtained by performing semantic transformation operations on the sample emergency processing requirements through multiple semantic transformation strategies. Based on the emergency processing needs of the same sample, the semantic transformation processing needs obtained by applying multiple semantic transformation strategies are used to generate the target sample processing needs. Based on the emergency processing needs of different samples, the semantic transformation processing needs obtained by applying arbitrary semantic transformation strategies generate non-target sample processing needs. Based on the target sample processing requirements and the non-target sample processing requirements, a training process is performed on the initial semantic analysis model using a self-supervised strategy to obtain the semantic analysis model.

10. A tower bolt processing system based on a voiceprint classification model, characterized in that, The system includes: The acquisition module is used to obtain the voiceprint to be analyzed from the audio of the tower monitoring. The model training module is used to input the voiceprint to be analyzed into a pre-trained target voiceprint classification model to obtain the category classification result of the voiceprint to be analyzed. The process of obtaining the target voiceprint classification model includes: constructing an initial voiceprint classification model and a bolt voiceprint instance dataset, optimizing the bolt voiceprint instance dataset, and training the initial voiceprint classification model based on the optimized bolt voiceprint instance dataset. The optimization process of the bolt voiceprint instance dataset includes: The similarity calculation module is used to determine the sample similarity between bolt soundprint instances. Based on the sample similarity, the module extracts the target sample similarity between the bolt soundprint instances and the corresponding bolt soundprint instances in the bolt soundprint instance dataset of the soundprint knowledge graph, thereby determining the association influence factor of the bolt soundprint instances. Based on the association factor, the module obtains the soundprint knowledge graph. Based on the soundprint knowledge graph, the module determines the degree of difference between the bolt soundprint instances. Based on the weight corresponding to the degree of difference, the module obtains the diffusion target value information of the bolt soundprint instances. Based on the diffusion target value information of the bolt soundprint instances, the module obtains the optimized bolt soundprint instance dataset. The result processing module receives an emergency processing request for the voiceprint to be analyzed when the classification result is characterized as an abnormal voiceprint result. Construct current emergency response instructions based on the voiceprint to be analyzed and the emergency response requirements; The current emergency response instructions are input into a pre-trained emergency scenario recognition model to obtain an emergency response strategy matching set. The target emergency strategy is selected from the preset emergency strategy database based on the emergency response strategy matching set.

Citation Information

Patent Citations

  • Iron tower bolt health detection method and system based on artificial intelligence and voiceprint

    CN118212938A

  • Bolt detection system and implementation method thereof

    CN108269249A

  • Power system operation and maintenance management system and method based on intelligent voice recognition

    CN117577134A