Equipment detection operation and maintenance method and device, related equipment and computer program product

By using multimodal fusion of voiceprint signals and fault description information, and utilizing a target network model for equipment fault diagnosis, the problems of light dependence and disassembly inspection in existing technologies are solved, achieving efficient and accurate equipment fault detection.

CN121580137APending Publication Date: 2026-02-27CHINA TELECOM CORP LTD TECHNOLOGY INNOVATION CENTER +1
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202511852145.9
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-12-09
Publication Date
2026-02-27

AI Technical Summary

Technical Problem

Existing equipment fault diagnosis methods are highly dependent on lighting conditions, require disassembly for inspection, lack real-time performance and accuracy, and are difficult to detect early and subtle signs, increasing the risk of network outages.

Method used

By acquiring voiceprint signals and fault description information when equipment malfunctions, and using a target network model for feature extraction and causal reasoning, a voice-language shared semantic space is constructed to achieve real-time intelligent diagnosis of equipment malfunctions.

Benefits of technology

It enables real-time intelligent diagnosis without disassembly or light dependence, significantly improving fault location efficiency and diagnostic accuracy, and reducing the need for manual intervention and the risk of network downtime.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121580137A_ABST
    Figure CN121580137A_ABST
Patent Text Reader

Abstract

The invention provides an equipment detection operation and maintenance method and device, related equipment and a computer program product, and relates to the technical field of computers and Internet. The method comprises the following steps: acquiring a first voiceprint signal acquired when a first device has a first fault, and first fault description information describing the first fault; acquiring actual fault information of the first fault; inputting the first voiceprint signal and the first fault description information into a target network model to obtain predicted fault information output by the target network model for the first fault; determining a first reward function based on the actual fault information and the predicted fault information; and carrying out training optimization on the target network model by utilizing the first reward function so as to carry out equipment fault detection and analysis according to the target network model. According to the method, more accurate, efficient and self-adaptive intelligent fault diagnosis and operation and maintenance can be realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This disclosure relates to the field of computer and Internet technology, and in particular to a method, apparatus, electronic device, computer-readable storage medium, and computer program product for equipment testing and maintenance. Background Technology

[0002] This section is intended to provide background or context for the embodiments of this disclosure as set forth in the claims. The description herein is not intended to be a prior art simply because it is included in this section.

[0003] In the field of equipment operation and maintenance, existing technologies mainly rely on traditional image processing for equipment fault diagnosis and prediction. These methods generally have several limitations when dealing with network equipment faults, such as strong dependence on lighting conditions and the need for disassembly for inspection, which limits their application in practical operation and maintenance scenarios.

[0004] In addition, existing technologies are insufficient in terms of real-time performance and accuracy, making it difficult to effectively capture early signs of minor mechanical wear or electrical breakdown, thereby increasing the risk of network outages. Summary of the Invention

[0005] The purpose of this disclosure is to provide a method, apparatus, electronic device, computer-readable storage medium, and computer program product for equipment testing and maintenance, which can improve the accuracy and efficiency of equipment fault detection.

[0006] Other features and advantages of this disclosure will become apparent from the following detailed description, or may be learned in part from practice of this disclosure.

[0007] This disclosure provides a device detection and maintenance method, comprising: acquiring a first voiceprint signal collected when a first device experiences a first fault, and first fault description information describing the first fault; acquiring actual fault information of the first fault, the actual fault information including actual fault type, actual fault root cause description, actual verification steps, and actual repair instructions; inputting the first voiceprint signal and the first fault description information into a target network model to obtain predicted fault information output by the target network model for the first fault, the predicted fault information including predicted fault type, predicted fault root cause, predicted verification steps, and predicted repair instructions; determining a first reward function based on the actual fault information and the predicted fault information; and training and optimizing the target network model using the first reward function to perform device fault detection and analysis based on the target network model.

[0008] In some embodiments, the method further includes: acquiring a second voiceprint signal collected when the second device experiences a second fault, and the actual fault type, actual fault root cause description, actual verification steps, and actual repair instructions of the second fault; acquiring second fault description information, which is used to describe faults other than the second fault; inputting the second voiceprint signal and the second fault description information into the target network model to obtain the predicted fault type, predicted fault root cause, predicted verification steps, and predicted repair instructions output by the target network for the second fault; determining a second reward function based on the actual fault type, actual fault root cause description, actual verification steps, and actual repair instructions of the second device, and the predicted fault type, predicted fault root cause, predicted verification steps, and predicted repair instructions of the second device; and training and optimizing the target network model using the second reward function.

[0009] In some embodiments, the target network model includes a feature encoder and a causal inferencer; wherein, inputting the first voiceprint signal and the first fault description information into the target network model to obtain the predicted fault information output by the target network model for the first fault includes: extracting frequency domain features and time domain features from the first voiceprint signal using the feature encoder to obtain voiceprint time domain features and voiceprint frequency domain features; extracting features from the first fault description information using the feature encoder to obtain natural language features; fusing the voiceprint time domain features, the voiceprint frequency domain features, and the natural language features using the feature encoder to obtain fused fault features; and inputting the fused fault features into the target network model to obtain the predicted fault information.

[0010] In some embodiments, the first voiceprint signal is extracted using the feature encoder to obtain frequency domain features and time domain features, respectively, to obtain voiceprint time domain features and voiceprint frequency domain features. This includes: performing windowing and framing processing on the first voiceprint signal to obtain multiple frames of voiceprint sub-signals; extracting the frequency domain features and time domain features of each frame of voiceprint sub-signals; determining the voiceprint frequency domain features based on the frequency domain features of each frame of voiceprint sub-signals; and determining the voiceprint time domain features based on the time domain features of each frame of voiceprint sub-signals.

[0011] In some embodiments, the feature encoder fuses the voiceprint temporal features, the voiceprint frequency features, and the natural language features to obtain fused fault features, including: fusing the voiceprint temporal features and the voiceprint frequency features to obtain voiceprint fused features; constructing an overcomplete sparse dictionary; solving for the sparse coefficients of the voiceprint fused features using the overcomplete sparse dictionary; determining a sparse fault spectrum fingerprint based on the sparse coefficients and the overcomplete sparse dictionary; and fusing the sparse fault spectrum fingerprint and the natural language features using the feature encoder to obtain the fused fault features.

[0012] In some embodiments, the method further includes: acquiring a vibration signal collected when the first device experiences the first fault; wherein, inputting the first acoustic signature signal and the first fault description information into a target network model to obtain predicted fault information output by the target network model for the first fault includes: inputting the first acoustic signature signal, the first fault description information, and the vibration signal into the target network model to obtain predicted fault information output by the target network model for the first fault.

[0013] In some embodiments, the method further includes: acquiring third voiceprint information collected from a third device; performing feature matching between the third voiceprint signal and multiple known fault voiceprint signals; if the third voiceprint signal successfully matches at least one fault voiceprint signal, then inputting the third voiceprint signal into the target network model so that the target network model outputs predicted fault information for the third device.

[0014] In some embodiments, the method further includes: acquiring fourth voiceprint information collected from a fourth device; performing feature matching between the fourth voiceprint signal and multiple known fault voiceprint signals to determine the similarity between the fourth voiceprint signal and each known fault voiceprint signal; if the confidence of the predicted fault root cause in the predicted fault information corresponding to the fourth device is less than a first preset threshold or the similarity between the fourth voiceprint signal and all known fault voiceprint signals is less than a second preset threshold, then determining the fourth voiceprint information as a candidate fault sample; and updating the known fault voiceprint signal based on the candidate fault sample.

[0015] In some embodiments, the device detection and maintenance is performed by the target edge device; the target network model includes a feature encoder and a causal inferencer; the method further includes: the target edge device locally stores the parameters corresponding to the feature encoder, uploads the parameters of the causal inferencer to a central server, so that the central server updates the global parameters corresponding to the causal inferencer based on the parameters uploaded by multiple edge devices; receives the global parameters from the central server; and updates the parameters of the target network model based on the global parameters.

[0016] This disclosure provides an equipment detection and maintenance device, including: a first voiceprint signal acquisition module, an actual fault information acquisition module, a prediction module, a reward function determination module, and an optimization module.

[0017] The first voiceprint signal acquisition module is used to acquire a first voiceprint signal collected when the first device experiences a first fault, and first fault description information describing the first fault; the actual fault information acquisition module can be used to acquire actual fault information of the first fault, which includes actual fault type, actual fault root cause description, actual verification steps, and actual repair instructions; the prediction module can be used to input the first voiceprint signal and the first fault description information into a target network model to obtain predicted fault information output by the target network model for the first fault, which includes predicted fault type, predicted fault root cause, predicted verification steps, and predicted repair instructions; the reward function determination module can be used to determine a first reward function based on the actual fault information and the predicted fault information; the optimization module can be used to train and optimize the target network model using the first reward function, so as to perform device fault detection and analysis based on the target network model.

[0018] This disclosure provides an electronic device comprising: a memory and a processor; the memory is used to store computer program instructions; the processor invokes the computer program instructions stored in the memory to implement the device detection and maintenance method described above.

[0019] This disclosure provides a computer-readable storage medium storing computer program instructions to implement the device testing and maintenance method as described in any of the preceding embodiments.

[0020] This disclosure provides a computer program product or computer program that includes computer program instructions stored in a computer-readable storage medium. The computer program instructions are read from the computer-readable storage medium, and the processor executes the computer program instructions to implement the aforementioned device detection and maintenance method.

[0021] The equipment detection and maintenance method, apparatus, electronic device, computer-readable storage medium, and computer program product provided in this disclosure achieve more accurate, efficient, and adaptive intelligent operation and maintenance by fusing equipment acoustic signals with manual fault description information and continuously optimizing the model using a reward function constructed based on the difference between actual and predicted fault information. This method effectively improves the accuracy and reliability of fault diagnosis, promotes the automation and intelligence of the operation and maintenance process, and significantly reduces the need for manual intervention.

[0022] It should be understood that the above general description and the following detailed description are merely exemplary and do not limit this disclosure. Attached Figure Description

[0023] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this disclosure and, together with the description, serve to explain the principles of this disclosure. It is obvious that the drawings described below are merely some embodiments of this disclosure, and those skilled in the art can obtain other drawings based on these drawings without any inventive effort.

[0024] Figure 1 A schematic diagram of a scenario that can be applied to the equipment detection and maintenance method or equipment detection and maintenance device in the embodiments of this disclosure is shown.

[0025] Figure 2 This is a flowchart illustrating a device testing and maintenance method according to an exemplary embodiment.

[0026] Figure 3 This is a flowchart illustrating a training method for a target network model according to an exemplary embodiment.

[0027] Figure 4 This is a flowchart illustrating a device testing and maintenance method according to an exemplary embodiment.

[0028] Figure 5 This is a flowchart illustrating a device testing and maintenance method according to an exemplary embodiment.

[0029] Figure 6 This is a fault information prediction method illustrated according to an exemplary embodiment.

[0030] Figure 7 This is a flowchart illustrating a method for updating a known voiceprint signal according to an exemplary embodiment.

[0031] Figure 8 This is a deployment architecture diagram of a distributed acoustic sensor network according to an exemplary embodiment.

[0032] Figure 9This is a flowchart illustrating a time-frequency joint feature extraction method according to an exemplary embodiment.

[0033] Figure 10 This is an architecture diagram illustrating a shared semantic space construction method according to an exemplary embodiment.

[0034] Figure 11 This is a flowchart illustrating a model training method according to an exemplary embodiment.

[0035] Figure 12 This is a schematic diagram illustrating an online prediction and adaptive enhancement update method according to an exemplary embodiment.

[0036] Figure 13 This is a block diagram illustrating a device detection and maintenance apparatus according to an exemplary embodiment.

[0037] Figure 14 A schematic diagram of the structure of an electronic device suitable for implementing embodiments of the present disclosure is shown. Detailed Implementation

[0038] Exemplary embodiments will now be described more fully with reference to the accompanying drawings. However, these exemplary embodiments can be implemented in many forms and should not be construed as limited to the embodiments set forth herein; rather, they are provided so that this disclosure will be thorough and complete, and will fully convey the concept of the exemplary embodiments to those skilled in the art. The same reference numerals in the drawings denote the same or similar parts, and therefore repeated descriptions of them will be omitted.

[0039] Those skilled in the art will recognize that embodiments of this disclosure can be a system, apparatus, device, method, or computer program product. Therefore, this disclosure can be implemented in the following forms: entirely hardware, entirely software (including firmware, resident software, microcode, etc.), or a combination of hardware and software.

[0040] The features, structures, or characteristics described in this disclosure can be combined in any suitable manner in one or more embodiments. Numerous specific details are provided in the following description to give a thorough understanding of embodiments of this disclosure. However, those skilled in the art will recognize that the technical solutions of this disclosure can be practiced with one or more specific details omitted, or other methods, components, apparatuses, steps, etc., can be employed. In other instances, well-known methods, apparatuses, implementations, or operations are not shown or described in detail to avoid obscuring various aspects of this disclosure.

[0041] In this disclosure, the terms "module" or "unit" refer to a computer program or part of a computer program that has a predetermined function and works with other related parts to achieve a predetermined goal, and can be implemented wholly or partially using software, hardware (such as processing circuitry or memory), or a combination thereof. Similarly, a processor (or multiple processors or memory) can be used to implement one or more modules or units. Furthermore, each module or unit can be part of an overall module or unit that includes the functionality of that module or unit.

[0042] The accompanying drawings are merely illustrative of this disclosure, and the same reference numerals in the drawings denote the same or similar parts, thus omitting repeated descriptions of them. Some block diagrams shown in the drawings do not necessarily correspond to physically or logically independent entities. These functional entities may be implemented in software, in one or more hardware modules or integrated circuits, or in different network and / or processor devices and / or microcontroller devices.

[0043] The flowchart shown in the accompanying drawings is merely illustrative and does not necessarily include all content and steps, nor does it require execution in the described order. For example, some steps may be broken down, while others may be combined or partially combined; therefore, the actual execution order may change depending on the specific circumstances.

[0044] In the description of this disclosure, unless otherwise stated, " / " means "or," for example, A / B can mean A or B. "And / or" in this document is merely a description of the relationship between related objects, indicating that three relationships can exist. For example, A and / or B can represent: A alone, A and B simultaneously, and B alone. Furthermore, "at least one" means one or more, and "multiple" means two or more. The terms "first," "second," etc., do not limit the quantity or order of execution, and "first," "second," etc., do not necessarily imply differences; the terms "contains," "includes," and "has" are used to indicate an open-ended meaning of inclusion and refer to the existence of additional elements / components / etc. besides those listed.

[0045] To better understand the above-mentioned objectives, features and advantages of the present invention, the present invention will be further described in detail below with reference to the accompanying drawings and specific embodiments. It should be noted that, unless otherwise specified, the embodiments and features in the embodiments of the present disclosure can be combined with each other.

[0046] The following section will first explain some of the terms used in the embodiments of this disclosure so that those skilled in the art can understand them.

[0047] Voice-speech fusion self-diagnostic network operation and maintenance big model platform: With the fine alignment of device acoustic characteristics and operation and maintenance verbal descriptions as the core, it is a platform that realizes intelligent operation and maintenance diagnosis through multi-stage collaboration.

[0048] Equipment acoustic characteristics refer to the sound signals and physical information contained therein collected by deployed acoustic sensors (such as MEMS microphones) that reflect the operating status of network devices (such as server fans, relays, power supplies, etc.). These acoustic characteristics (such as acoustic signals) can refer to the sound generated by physical processes such as mechanical vibration and electrical activity during equipment operation. These acoustic characteristics (such as acoustic signals) can be the acoustic characteristics of mechanical components, the acoustic characteristics of electrical components, etc. For example, mechanical components may produce friction noise of a specific frequency due to fan bearing wear; electrical components may produce high-frequency pulse noise or resonant sound, such as that generated by relay switches, capacitor aging, or partial discharge.

[0049] Distributed acoustic sensor network: A network consisting of multiple acoustic sensor nodes distributed in key locations such as computer rooms and equipment compartments. The nodes are interconnected through wireless communication technology and transmit the collected data to a central server.

[0050] Frequency-time joint encoder: The original voiceprint data is processed by first dividing the time-domain signal into frames and windowing it, then extracting spectral features by combining short-time Fourier transform, and analyzing time-frequency localization information through wavelet transform, and finally fusing them to form an encoder with multi-dimensional voiceprint representation.

[0051] Sparse fault spectrum fingerprint: After sparsification, it can effectively distinguish between normal and fault states, retain key features in the voiceprint that are strongly correlated with faults, and suppress the feature fingerprint of interference factors such as environmental noise.

[0052] Voice-Shared Semantic Space: A space constructed to achieve semantic alignment between voiceprints and operational spoken descriptions. Through contrastive learning, voiceprint features with similar semantics are more closely embedded in the space with the text.

[0053] Chain-based causal reinforcement fine-tuning: This training method uses multi-dimensional samples of "voiceprint-text-action" with root cause-disposal labels. With voiceprint features as input, the causal reasoning module sequentially generates root cause inference, verification steps, and repair instructions. The matching degree between the output of each step and the label is used as a reward signal to adjust the model parameters.

[0054] Online active learning and adaptive incremental update mechanism: The platform supports a mechanism that automatically marks suspicious samples and pushes them to experts for confirmation when new types of acoustic anomalies occur. It also updates the fault spectrum library based on feedback and dynamically adjusts the model inference strategy based on real-time data.

[0055] BERT (Bidirectional Encoder Representations from Transformers): A pre-trained language model used to embed spoken descriptions from operations and maintenance personnel to extract natural language features.

[0056] Proximal Policy Optimization algorithm: A reinforcement learning algorithm used to adjust model parameters in chained causal reinforcement fine-tuning during model training.

[0057] The preceding text introduced some terms and concepts involved in the embodiments of this disclosure. The following text introduces the technical features involved in the embodiments of this disclosure.

[0058] The exemplary embodiments of this disclosure will now be described in detail with reference to the accompanying drawings.

[0059] Figure 1 A schematic diagram of a scenario that can be applied to the equipment detection and maintenance method or equipment detection and maintenance device in the embodiments of this disclosure is shown.

[0060] Please refer to Figure 1 The diagram illustrates an implementation environment provided by an exemplary embodiment of this disclosure.

[0061] like Figure 1 As shown, system architecture 100 may include terminal devices 101, 102, and 103, a network 104, and a server 105. Network 104 serves as the medium for providing communication links between terminal devices 101, 102, and 103 and server 105. Network 104 may include various connection types, such as wired or wireless communication links, or fiber optic cables, etc.

[0062] Users can use terminal devices 101, 102, and 103 to interact with server 105 via network 104 to receive or send messages, etc. Terminal devices 101, 102, and 103 can be various electronic devices with displays and web browsing capabilities, including but not limited to smartphones, tablets, laptops, desktop computers, wearable devices, virtual reality devices, smart home devices, etc.

[0063] Server 105 can be a server that provides various services, such as a backend management server that supports the devices operated by users using terminal devices 101, 102, and 103. The backend management server can analyze and process received requests and other data, and feed the processing results back to the terminal devices.

[0064] A server can be a standalone physical server, a server cluster or a distributed system consisting of multiple physical servers, or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communication, middleware services, domain name services, security services, CDN (Content Delivery Network), and big data and artificial intelligence platforms. This disclosure does not impose any restrictions on this.

[0065] Server 105 may, for example, acquire a first voiceprint signal collected when the first device experiences a first fault, and first fault description information describing the first fault; server 105 may, for example, acquire actual fault information of the first fault, including actual fault type, actual fault root cause description, actual verification steps, and actual repair instructions; server 105 may, for example, input the first voiceprint signal and the first fault description information into a target network model to obtain predicted fault information output by the target network model for the first fault, including predicted fault type, predicted fault root cause, predicted verification steps, and predicted repair instructions; server 105 may, for example, determine a first reward function based on the actual fault information and the predicted fault information; server 105 may, for example, use the first reward function to train and optimize the target network model so as to perform device fault detection and analysis based on the target network model.

[0066] It should be understood that Figure 1 The number of terminal devices, networks, and servers shown is merely illustrative. Server 105 can be a single physical server or a combination of multiple servers. Depending on actual needs, it can have any number of terminal devices, networks, and servers.

[0067] Figure 2 This is a flowchart illustrating a device testing and maintenance method according to an exemplary embodiment.

[0068] The method provided in this disclosure can be executed by any electronic device with computing power, for example, the method can be executed by the above-described... Figure 1 The execution can be performed by a server or terminal device in the embodiments, or it can be performed by both a server and a terminal device. In the following embodiments, the server is used as the execution subject for illustration, but this disclosure is not limited to this.

[0069] In some embodiments, a distributed acoustic sensor network can be deployed first. This network consists of multiple acoustic sensor nodes distributed in key locations such as the equipment room and equipment compartment. The sensor nodes can be selected as microphones, and the deployment location is determined based on the equipment structure and the acoustic propagation path of the fault. For example, for vulnerable components such as fans and relays, sensors are deployed in unobstructed areas 0.5-1.5 meters away from the target equipment; for large equipment compartments, a grid-like deployment (3-5 meters apart) can be used to cover the acoustic characteristics of various operating scenarios such as mechanical vibration and electrical discharge.

[0070] In some embodiments, nodes can be interconnected via low-power long-distance LoRa (a low-power wide-area network technology) or high-bandwidth Wi-Fi (wireless local area network access) technology. Data is transmitted to the central server via Transmission Control Protocol (TCP) / Internet Protocol (IP). The server uses NTP (Network Time Protocol) to achieve time synchronization of data among multiple nodes, ensuring the spatiotemporal consistency of the voiceprint signal.

[0071] In some embodiments, the sensor can selectively capture low-frequency mechanical vibration noises below 200Hz (such as fan bearing wear) and high-frequency electrical discharge noises above 2kHz (such as relay pulses and power supply resonance), covering typical sound signature signals such as buzzing and fan whistling.

[0072] "Fault acoustic propagation path" is a core concept in acoustic fault diagnosis. It refers to the entire process by which the sound signal propagates through multiple physical paths to the sensor (receiving point) after a fault occurs inside the equipment and an abnormal sound (sound source) is generated.

[0073] Spatiotemporal consistency refers to the ability of voiceprint signals collected from different locations (spaces) and at different times to be accurately aligned and correlated, thereby forming a unified and meaningful overall picture.

[0074] Reference Figure 2 The equipment testing and maintenance method provided in this embodiment may include the following steps.

[0075] Step S202: Obtain the first voiceprint signal collected when the first device experiences a first fault, and the first fault description information describing the first fault.

[0076] The first acoustic signature signal can refer to a raw, unprocessed acoustic waveform data collected by an acoustic sensor deployed near a first device (such as a specific server, router, or switch) when the first device (such as a specific, identified fault type, such as fan bearing wear, relay arcing, or capacitor aging) experiences a first fault.

[0077] In some embodiments, "first fault description information" may refer to a textual description of a "first fault" that has occurred in the "first device" provided by maintenance personnel in natural language. The textual description may include a textual description of the phenomenon, type, root cause, historical experience, verification method or handling measures corresponding to the first fault of the first device.

[0078] The aforementioned first fault description information can be textual information converted from verbal descriptions provided by maintenance personnel.

[0079] Step S204: Obtain the actual fault information of the first fault. The actual fault information includes the actual fault type, the actual fault root cause description, the actual verification steps, and the actual repair instructions.

[0080] A fault type can be a standardized and categorized label for a fault phenomenon. It is usually an item in a predefined list of faults. For example, a fault type could be: fan bearing wear, power capacitor bulging, relay contact burning, or abnormal hard drive read / write noises, etc.

[0081] Root Cause Description: A detailed textual description of the root cause of the fault. It is richer than "Fault Type" and provides a deeper causal explanation. For example, "Due to long-term operation and lubrication failure, the fan bearing has experienced physical wear, resulting in clearance and abnormal friction noise."

[0082] Verification steps: A textual description of a series of checks or tests that need to be performed to confirm the fault. For example, "1. Use monitoring software to check if the fan speed fluctuates periodically or exceeds the threshold; 2. Under safe conditions, listen closely to the source of the abnormal noise to confirm that the sound source is the fan module; 3. Check the equipment log to see if there are any alarm records related to the fan."

[0083] Repair instructions: A textual description of the specific repair operations that need to be performed to resolve the fault.

[0084] Step S206: Input the first voiceprint signal and the first fault description information into the target network model to obtain the predicted fault information output by the target network model for the first fault. The predicted fault information includes the predicted fault type, the predicted fault root cause, the predicted verification steps, and the predicted repair instructions.

[0085] In some embodiments, to align the semantics of the first voiceprint signal with the operational verbal description (such as the first fault description information), a shared semantic space for voice and language can be constructed. For example, on the text side, a pre-trained language model BERT-base-uncased (a bidirectional encoder representation based on Transformer) is used to embed the operational verbal description: the input text (such as the first fault description information) is segmented and truncated / padded to a length of 512, and [CLS] and [SEP] tags are added. A 768-dimensional vector at the [CLS] position is extracted as a natural language feature through a 12-layer Transformer encoder. This feature contains semantic information such as fault phenomena (e.g., "the equipment emits a continuous hum") and historical experience (e.g., "similar problems were caused by capacitor aging"). On the voiceprint side, the features corresponding to the voiceprint signal can be projected into the same 768-dimensional space as the text embedding through two fully connected layers (256→768 dimensions, ReLU activation function).

[0086] In some embodiments, multi-dimensional samples can be used to implement chain-like causal reinforcement fine-tuning during the model training phase. The sample set may include acoustic signals (corresponding to voiceprint features) recorded in historical operations and maintenance, descriptions of root causes of faults (such as the classification label "fan bearing wear"), verification steps (such as "check fan speed fluctuations", text sequence), and repair instructions (such as "replace bearing", text sequence). The target network model architecture may include a voiceprint feature encoder (outputting 768 dimensions) and a causal inference module (based on a Transformer decoder, 6 layers, 768-dimensional hidden layers, 12 multi-head attention heads). During training, the model can take voiceprint fusion features as input and sequentially perform root cause inference (classification task, Softmax outputs fault category probability), verification step generation (autoregressive text generation, vocabulary size 30522), and repair instruction generation (same as above) through the causal inference module.

[0087] Step S208: Determine the first reward function based on the actual fault information and the predicted fault information.

[0088] In some embodiments, the matching degree between the output (actual fault information) and the label (predicted fault information) of each step of the target network model can be used as a reward signal: root cause inference reward ( (Correctly categorized) or ( (Error); Verification steps and repair instructions reward ( ).

[0089] Step S210: The target network model is trained and optimized using the first reward function so as to perform equipment fault detection and analysis based on the target network model.

[0090] In some embodiments, the Proximal Policy Optimization (PPO) algorithm can be used to adjust the parameters of the target network model: the policy network is a causal inference module, and the value network evaluates the state value (mean squared error loss); when collecting trajectories, 32 samples are collected in each batch, and output is generated and cumulative reward is calculated at each step. , Represents the reward at time t. This represents the cumulative reward at time t+1, where t is a number greater than 0, and the discount factor. The advantage function A(s, a) is calculated using the Generalized Advantage Estimation (GAE). During the update phase, the policy network is updated in four rounds (batch size 8) to ensure the logical coherence of the generation process (e.g., the probability of generating the verification step "check speed fluctuation" after the root cause "bearing wear" is increased to over 0.8) and interpretability (key frequency components in the voiceprint features are located through attention visualization).

[0091] This disclosure discloses an embodiment that collects device voiceprint signals by deploying a distributed acoustic sensor network, constructs a shared semantic space for voice and speech to achieve cross-modal semantic alignment between acoustic features and maintenance text, and fine-tunes the target network model based on chain-like causal reinforcement training. This enables the model to sequentially generate accurate fault types, root cause descriptions, verification steps, and repair instructions based on voiceprint signals. Finally, the model parameters are continuously optimized through a reinforcement learning reward mechanism, achieving real-time intelligent diagnosis and maintenance of early device faults without disassembly or light dependence. This significantly improves fault location efficiency and diagnostic accuracy, reduces the need for manual intervention, and lowers the risk of network downtime. This method can capture early signs of weak mechanical wear and electrical breakdown without disassembly or light dependence, significantly improving the real-time performance of maintenance, reducing manual intervention, and lowering the risk of network downtime due to hardware aging or latent faults.

[0092] Figure 3 This is a flowchart illustrating a training method for a target network model according to an exemplary embodiment.

[0093] refer to Figure 3 The training method for the aforementioned target network model may include the following steps.

[0094] Step S302: Obtain the second voiceprint signal collected when the second device experiences a second fault, the actual fault type of the second fault, the actual fault root cause description, the actual verification steps, and the actual repair instructions.

[0095] Step S304: Obtain second fault description information. The second fault description information is used to describe faults other than the second fault.

[0096] Step S306: Input the second voiceprint signal and the second fault description information into the target network model to obtain the predicted fault type, predicted fault root cause, predicted verification steps, and predicted repair instructions output by the target network for the second fault.

[0097] Step S308: Determine the second reward function based on the actual fault type, actual fault root cause description, actual verification steps, and actual repair instructions of the second device, and the predicted fault type, predicted fault root cause, predicted verification steps, and predicted repair instructions of the second device.

[0098] Step S310: Train and optimize the target network model using the second reward function.

[0099] In some embodiments, the parameters of the target network model can be optimized through contrastive learning: positive sample pairs are defined as the acoustic fingerprint projection and text embedding of the same fault (e.g., the acoustic fingerprint corresponding to "fan bearing wear" and the text of "bearing friction noise"), and negative sample pairs are combinations of different faults (e.g., the acoustic fingerprint of "bearing wear" and the text of "power supply resonance"); the InfoNCE (Information Noise-Contrastive Estimation) loss function can be used, with temperature parameters... The optimizer is AdamW (Adaptive Moment Estimation with Weight Decay), which is trained until the loss converges, so that the cosine distance between the voiceprint features with similar semantics and the text embedding in space is less than 0.3, forming a cross-modal semantic association.

[0100] This embodiment of the disclosure introduces interfering fault descriptions (second fault description information) that are not related to the current fault as negative sample inputs, which forces the target network model to extract features that are truly related to the current fault more accurately from the voiceprint signal during the training process and suppresses the response to irrelevant semantic information. This enhances the model's feature discrimination ability and robustness in complex multi-fault backgrounds and improves the accuracy and generalization ability of fault diagnosis.

[0101] Figure 4 This is a flowchart illustrating a device testing and maintenance method according to an exemplary embodiment.

[0102] In some embodiments, the target network model may include a feature encoder and a causal inferencer.

[0103] refer to Figure 4 The above-mentioned equipment testing and maintenance methods may include the following steps.

[0104] Step S402: Extract frequency domain features and time domain features from the first voiceprint signal using a feature encoder to obtain voiceprint time domain features and voiceprint frequency domain features.

[0105] In some embodiments, the first voiceprint signal can be windowed and framed to obtain multiple frames of voiceprint sub-signals; then, the frequency domain features (e.g., frequency domain features extracted by Fourier transform) and time domain features (e.g., time domain features extracted by wavelet transform) of each frame of voiceprint sub-signals are extracted; then, the voiceprint frequency domain features are determined according to the frequency domain features of each frame of voiceprint sub-signals (e.g., the voiceprint frequency domain features corresponding to each voiceprint sub-signal are added or spliced); and the voiceprint time domain features are determined according to the time domain features of each frame of voiceprint sub-signals (e.g., the voiceprint time domain features corresponding to each voiceprint sub-signal are added or spliced).

[0106] In some embodiments, during the feature extraction stage, a frequency-time domain joint encoder can be used to process the raw voiceprint data (such as the first voiceprint signal). The encoder can first perform frame-by-frame windowing processing on the time-domain signal (such as the first voiceprint signal): using a Hamming window (window function expression) The frame length is 512 points (corresponding to a sampling rate of 10ms@48kHz), and the frame shift is 256 points (overlap rate of 50%). N represents the window length, and n represents the sample point index, where n is an integer variable. The framed and windowed signal is subjected to a Short-Time Fourier Transform (STFT) to extract spectral features. Specifically, it is implemented as a 1024-point Fast Fourier Transform (FFT) to calculate the power spectral density (PSD) of each frame and extract frequency domain statistics such as center frequency, bandwidth, and energy concentration. At the same time, a Continuous Wavelet Transform (CWT) can be performed using the Morlet mother wavelet, with a scale range of 1-64 (corresponding to a frequency resolution of 0.75-48kHz) to extract time-frequency localization features such as time-frequency energy distribution, instantaneous frequency, and time-frequency entropy. Frequency domain and wavelet features are spliced ​​together and then input into a fully connected layer (256→128 dimensions) to achieve feature dimensionality reduction and fusion, forming a multi-dimensional voiceprint representation that includes frequency distribution, energy change, and instantaneous phase.

[0107] Step S404: Extract features from the first fault description information using a feature encoder to obtain natural language features.

[0108] Step S406: The speaker's time-domain features, frequency-domain features, and natural language features are fused using a feature encoder to obtain fused fault features.

[0109] In some embodiments, the time-domain features and frequency-domain features of the voiceprint can be fused to obtain the voiceprint fusion features; then an overcomplete sparse dictionary is constructed; next, the sparse coefficients of the voiceprint fusion features are solved by combining the overcomplete sparse dictionary; further, based on the sparse coefficients and the overcomplete sparse dictionary, the sparse fault spectrum fingerprint is determined; finally, the sparse fault spectrum fingerprint and natural language features are fused by a feature encoder to obtain the fused fault features.

[0110] In some embodiments, to suppress environmental noise interference, a K-SVD dictionary (K-Singular Value Decomposition dictionary) learning algorithm can be used to sparsify the fused features: an overcomplete dictionary is constructed, and the sparsity of the sparse coefficients (x) (s=10) is solved using the Orthogonal Matching Pursuit (OMP) algorithm, ultimately generating a sparse fault spectrum fingerprint. This fingerprint retains key features strongly correlated with the fault (such as the 120Hz harmonic component corresponding to bearing wear) while suppressing irrelevant information such as background white noise.

[0111] Step S408: Input the fused fault features into the target network model to obtain the predicted fault information.

[0112] The above technical solution, based on a multimodal fusion method of voiceprint signals and natural language descriptions, effectively extracts key features of equipment faults through frequency-time domain joint feature extraction and sparse dictionary learning techniques. It achieves high-precision fault diagnosis and status prediction while suppressing environmental noise interference, significantly improving the accuracy and robustness of equipment detection and maintenance.

[0113] Figure 5 This is a flowchart illustrating a device testing and maintenance method according to an exemplary embodiment.

[0114] refer to Figure 5 The above-mentioned equipment testing and maintenance methods may include the following steps.

[0115] Step S502: Obtain the first voiceprint signal collected when the first device experiences a first fault, and the first fault description information describing the first fault.

[0116] Step S504: Obtain the actual fault information of the first fault. The actual fault information includes the actual fault type, the actual fault root cause description, the actual verification steps, and the actual repair instructions.

[0117] Step S506: Obtain the vibration signal collected when the first device experiences the first fault.

[0118] Vibration signals can be a record of how a physical quantity changes over time, used to describe the patterns and characteristics of an object or mechanical system reciprocating (oscillating) around its equilibrium position. In the field of engineering and equipment monitoring, it can refer to time-series data that quantifies the dynamic behavior of a mechanical system, collected by sensors (such as accelerometers, velocity sensors, or displacement sensors).

[0119] Step S508: Input the first acoustic signature signal, the first fault description information and the vibration signal into the target network model to obtain the predicted fault information output by the target network model for the first fault.

[0120] Step S510: Determine the first reward function based on the actual fault information and the predicted fault information.

[0121] Step S512: The target network model is trained and optimized using the first reward function so as to perform equipment fault detection and analysis based on the target network model.

[0122] This technical solution integrates multi-source information such as acoustic signatures, vibration signals, and natural language descriptions to construct a multimodal fault feature representation system. It also combines a reward mechanism based on the comparison of actual and predicted fault information to optimize the model through reinforcement learning, which significantly improves the accuracy, robustness, and generalization ability of equipment fault diagnosis. At the same time, it realizes accurate judgment and autonomous evolution of fault type, root cause, and repair strategy, providing a comprehensive and reliable solution for intelligent operation and maintenance in complex industrial scenarios.

[0123] Figure 6 This is a fault information prediction method illustrated according to an exemplary embodiment.

[0124] refer to Figure 6 The above-mentioned fault information prediction method may include the following steps.

[0125] Step S602: Obtain the third voiceprint information collected by the third device.

[0126] Step S604: Perform feature matching between the third voiceprint signal and multiple known fault voiceprint signals.

[0127] Step S606: If the third voiceprint signal successfully matches at least one fault voiceprint signal, the third voiceprint signal is input into the target network model so that the target network model outputs predicted fault information for the third device.

[0128] This technical solution achieves preliminary screening and precise triggering of faults by rapidly matching the real-time voiceprint signal of the device under test with a known fault voiceprint database. When a voiceprint feature that highly matches a known fault pattern is detected, the system automatically activates the subsequent deep analysis model. This "matching-triggering" mechanism not only significantly improves the real-time performance and efficiency of fault detection and avoids the computational resource consumption caused by continuously running complex models, but also ensures that the input deep model consists of highly suspected fault signals through pre-matching and screening, thereby improving the accuracy and reliability of the overall diagnostic process and providing an efficient and energy-saving intelligent solution for industrial equipment condition monitoring.

[0129] Figure 7 This is a flowchart illustrating a method for updating a known voiceprint signal according to an exemplary embodiment.

[0130] refer to Figure 7 The method for updating the known voiceprint signal can include the following steps.

[0131] Step S702: Obtain the fourth voiceprint information collected from the fourth device.

[0132] Step S704: Perform feature matching between the fourth voiceprint signal and multiple known fault voiceprint signals to determine the similarity between the fourth voiceprint signal and each known fault voiceprint signal.

[0133] Step S706: If the confidence level of the predicted fault root cause in the predicted fault information corresponding to the fourth device is less than the first preset threshold or the similarity between the fourth voiceprint signal and all known fault voiceprint signals is less than the second preset threshold, then the fourth voiceprint information is determined as a candidate fault sample.

[0134] Step S708: Update the known fault voiceprint signal based on the candidate fault samples.

[0135] In some embodiments, when a new type of acoustic anomaly occurs, suspicious samples can be automatically labeled using two conditions: the confidence level of the model root cause inference (a quantitative indicator of the degree of confidence of the model in determining the "root cause of the fault" after analyzing a voiceprint feature) < 0.7, or the cosine similarity between the voiceprint feature and all samples in the fault spectrum library < 0.6.

[0136] In some embodiments, suspicious samples can be pushed to operations and maintenance experts for confirmation (which may include a time spectrum of the voiceprint and historical similar cases). After the feedback data is processed to generate a sparse fault spectrum fingerprint through feature extraction, it is clustered by DBSCAN (Density-Based Spatial Clustering of Applications with Noise) (neighborhood radius 0.4, minimum number of samples 5) to determine whether it is a new fault class (if the average distance within the cluster is >0.5, a new class is added), and the sparse dictionary of the fault spectrum library is updated (K-SVD is retrained to include new samples, and the dictionary size is expanded to (K=256+N), where N is the number of new samples).

[0137] This technical solution automatically filters out abnormal voiceprint samples with low confidence prediction and low matching degree as candidate fault samples by setting a dual threshold mechanism of confidence and similarity. After expert confirmation and cluster analysis, the known fault voiceprint database is dynamically updated, realizing online evolution and adaptive optimization of the fault voiceprint database. This effectively solves the bottleneck of the traditional method's insufficient ability to identify new fault modes, and significantly improves the fault diagnosis system's ability to discover unknown fault types and its generalization performance for long-term application.

[0138] In some embodiments, the device detection and maintenance method provided in this application can be executed by the target edge device, and the target network model can include a feature encoder and a causal inferencer.

[0139] In some embodiments, the target edge device can locally store the parameters corresponding to the feature encoder and upload the parameters of the causal inference to the central server, so that the central server can update the global parameters corresponding to the causal inference based on the parameters uploaded by multiple edge devices; then it can receive the global parameters from the central server; and finally update the parameters of the target network model based on the global parameters.

[0140] This technical solution utilizes a distributed architecture that stores feature encoder parameters locally on the edge device and uploads causal inference parameters to the central server for collaborative updates. This architecture ensures the real-time performance and data privacy of device voiceprint feature extraction through edge computing, while also achieving global optimization of model capabilities by aggregating fault reasoning experience from various edge nodes in the cloud. This effectively balances detection efficiency and model generalization performance, constructing a distributed intelligent operation and maintenance system that balances real-time response and continuous evolution.

[0141] Below, this application will explain and illustrate the equipment operation and maintenance method with reference to specific embodiments.

[0142] Figure 8 This is a deployment architecture diagram of a distributed acoustic sensor network according to an exemplary embodiment.

[0143] like Figure 8 As shown, the aforementioned distributed acoustic sensor network can be composed of multiple acoustic sensor nodes distributed in key locations such as computer rooms and equipment compartments. The sensor nodes are selected as MEMS microphones (sampling rate 48kHz, 16-bit quantization, frequency range 20Hz-20kHz). Deployment locations are determined based on equipment structure and fault acoustic propagation paths: for vulnerable components such as fans and relays, sensors are deployed in unobstructed areas 0.5-1.5 meters away from the target equipment; for large equipment compartments, a grid deployment (spacing 3-5 meters) is used to cover the acoustic characteristics of various operating scenarios such as mechanical vibration and electrical discharge. Each node is interconnected via low-power long-distance LoRa (a low-power wide-area network technology) (transmission rate 0.3-50kbps, coverage radius 3-15km) or high-bandwidth Wi-Fi (802.11n / ac, transmission rate 150-866Mbps) technology. Data is transmitted to the central server via TCP / IP protocol. The server uses NTP protocol to achieve multi-node data time synchronization, ensuring the spatiotemporal consistency of the acoustic signature signal. The sensor can specifically capture low-frequency abnormal noises of mechanical vibration below 200Hz (such as fan bearing wear) and high-frequency noises of electrical discharge above 2kHz (such as relay pulses and power supply resonance), covering typical sound signature signals such as buzzing and fan whistling.

[0144] Figure 9 This is a flowchart illustrating a time-frequency joint feature extraction method according to an exemplary embodiment.

[0145] like Figure 9 As shown, in the feature extraction stage, a frequency-time domain co-encoder can be used to process the raw speaker data. The encoder first performs frame-by-frame windowing on the time-domain signal: using a Hamming window, a frame length of 512 points, and a frame shift of 256 points. The frame-by-frame windowed signal is then processed using Short-Time Fourier Transform (STFT) to extract spectral features, specifically a 1024-point FFT, to calculate the power spectral density (PSD) of each frame and extract frequency domain statistics such as center frequency, bandwidth, and energy concentration. Simultaneously, a Morlet mother wavelet is used for Continuous Wavelet Transform (CWT), with a scale range of 1-64 (corresponding to a frequency resolution of 0.75-48kHz), to extract time-frequency localization features such as time-frequency energy distribution, instantaneous frequency, and time-frequency entropy. The frequency domain and wavelet features are concatenated and input into a fully connected layer (256→128 dimensions) to achieve feature dimensionality reduction and fusion, forming a multi-dimensional speaker representation that includes frequency distribution, energy change, and instantaneous phase. To suppress environmental noise interference, the K-SVD dictionary learning algorithm is used to sparsify the fused features: an overcomplete dictionary is constructed, and the sparsity of the sparse coefficients (x) (s=10) is solved using the Orthogonal Matching Pursuit (OMP) algorithm, ultimately generating a sparse fault spectrum fingerprint. This fingerprint retains key features strongly correlated with the fault (such as the 120Hz harmonic component corresponding to bearing wear) while suppressing irrelevant information such as background white noise.

[0146] Figure 10 This is an architecture diagram illustrating a shared semantic space construction method according to an exemplary embodiment.

[0147] refer to Figure 10 To align the semantics of the first voiceprint signal with the operational verbal description (such as the first fault description information), a shared semantic space for voice and language can be constructed. For example, on the text side, a pre-trained language model BERT-base-uncased (based on a Transformer-based bidirectional encoder representation) is used to embed the operational verbal description: the input text (such as the first fault description information) is segmented and truncated / padded to a length of 512, and [CLS] and [SEP] tags are added. A 768-dimensional vector at the [CLS] position is extracted as a natural language feature through a 12-layer Transformer encoder. This feature contains semantic information such as fault phenomena (e.g., "the equipment emits a continuous hum") and historical experience (e.g., "similar problems were caused by capacitor aging"). On the voiceprint side, the features corresponding to the voiceprint signal can be projected into the same 768-dimensional space as the text embedding through two fully connected layers (256→768 dimensions, ReLU activation function).

[0148] Optimizing projection parameters through contrastive learning: The parameters of the target network model are optimized through contrastive learning. Positive sample pairs are defined as the acoustic fingerprint projection and text embedding of the same fault (e.g., the acoustic fingerprint corresponding to "fan bearing wear" and the text for "bearing friction noise"). Negative sample pairs are combinations of different faults (e.g., the acoustic fingerprint for "bearing wear" and the text for "power supply resonance"). The InfoNCE loss function and temperature parameter can be used. The optimizer is AdamW (Adaptive Moment Estimation with Weight Decay), which is trained until the loss converges, so that the cosine distance between the voiceprint features with similar semantics and the text embedding in space is less than 0.3, forming a cross-modal semantic association.

[0149] Figure 11 This is a flowchart illustrating a model training method according to an exemplary embodiment.

[0150] refer to Figure 11 The above model training method may include the following steps.

[0151] During model training, multi-dimensional samples can be used to implement chain-like causal reinforcement fine-tuning. The sample set can include acoustic signals (corresponding to voiceprint features) recorded in historical operations, fault root cause descriptions (such as the "fan bearing wear" classification label), verification steps (such as "check fan speed fluctuations", text sequence), and repair instructions (such as "replace bearing", text sequence). The target network model architecture can include a voiceprint feature encoder (output 768 dimensions) and a causal inference module (based on a Transformer decoder, 6 layers, 768-dimensional hidden layers, 12 multi-head attention heads). During training, the model can take voiceprint fusion features as input and sequentially execute root cause inference (classification task, Softmax output fault category probability), verification step generation (autoregressive text generation, vocabulary size 30522), and repair instruction generation (same as above) through the causal inference module. The matching degree between the output and the label at each step is used as a reward signal: root cause inference reward ( (Correctly categorized) or ( (Error); Verification steps and repair instructions reward\ (Generate, label) (range 0-1, weight 0.5), and compare it with the label text in the training data (i.e., the manually labeled, standard verification step "check if the fan speed fluctuates too much").

[0152] In some embodiments, the proximal policy optimization (PPO) algorithm can be used to adjust the model parameters: the policy network is a causal inference module, and the value network evaluates the state value (mean squared error loss); when collecting trajectories, 32 samples are collected in each batch, and output is generated and cumulative reward is calculated at each step. Discount factor The advantage function A(s, a) is calculated through generalized advantage estimation (GAE); during the update phase, four rounds of small-batch updates (batch size 8) are performed on the policy network to ensure the logical coherence of the generation process (e.g., the probability of generating the verification step "check speed fluctuation" after the root cause "bearing wear" is increased to above 0.8) and interpretability (key frequency components in the voiceprint features are located through attention visualization).

[0153] Figure 12 This is a schematic diagram illustrating an online prediction and adaptive enhancement update method according to an exemplary embodiment.

[0154] refer to Figure 12 The online prediction and adaptive enhancement update method described above may include the following process.

[0155] In some embodiments, when a new type of acoustic anomaly occurs, suspicious samples can be automatically labeled using two conditions: the model root cause inference confidence is <0.7, or the cosine similarity between the voiceprint features and all samples in the fault spectrum library is <0.6. Suspicious samples are pushed to maintenance experts for confirmation (with a voiceprint spectrum and historical similar cases). After the feedback data is processed to generate a sparse fault spectrum fingerprint through feature extraction, it is clustered by DBSCAN (neighborhood radius 0.4, minimum number of samples 5) to determine whether it is a new fault class (if the average distance within the cluster is >0.5, a new class is added), and the sparse dictionary of the fault spectrum library is updated (K-SVD is retrained to include new samples, and the dictionary size is expanded to (K=256+N), where N is the number of new samples). Meanwhile, based on real-time collected voiceprint data and the effects of maintenance actions (such as voiceprint returning to normal after repair), an online stochastic gradient descent dynamic adjustment strategy is adopted to adjust the model inference strategy: the fully connected layer parameters of the feature extraction encoder are updated every 100 new samples (learning rate 0.001), and the fusion weights of spectrum and wavelet features are optimized (such as increasing the weight of low-frequency components related to mechanical wear to 0.7); the causal inference module adjusts the attention weights of the Transformer through backpropagation (such as enhancing the association weight between the root cause "bearing wear" and the verification step "check rotation speed"), and updates the causal graph (adding association rules between the root cause of the fault and the verification and repair steps, such as "electrical discharge → verification 'detect partial discharge amount' → repair 'replace insulation component'"). This mechanism continuously improves the model's ability to capture early weak anomalies: for example, the recognition accuracy of low-frequency abnormal noises of 50-150Hz in the early stage of mechanical wear (energy is only 1.2 times that of the normal state) increases from the initial 65% to 85%, and the false negative rate of 3-5kHz partial discharge noise (duration <200ms) before electrical breakdown decreases from 22% to 5%.

[0156] Compared to traditional multimodal solutions that rely on image or fixed-parameter sensors, the method provided in this application avoids the complexity of disassembly and testing and the limitations of lighting conditions by leveraging the advantages of auditory perception. It can monitor in real time while the device is running (latency <200ms), effectively reducing the risk of network downtime caused by hardware aging or hidden faults, significantly improving the real-time performance of operation and maintenance (fault location time reduced from 30 minutes to 5 minutes) and reducing the need for manual intervention (manual intervention rate reduced by 40%).

[0157] This application's embodiments can solve the technical problems of traditional multimodal solutions for image or fixed-parameter sensors, such as complex disassembly inspection, susceptibility to lighting conditions, and difficulty in timely detection of early weak anomalies. It achieves the following: avoids the complexity of disassembly inspection and lighting limitations; monitors equipment operating status in real time; effectively reduces the risk of network downtime due to hardware aging or hidden faults; significantly improves the real-time performance of maintenance; and reduces the need for manual intervention.

[0158] The above method focuses on the precise alignment of equipment acoustic features with verbal descriptions used in operations and maintenance (O&M), achieving intelligent O&M diagnosis through multi-stage collaboration. The technical solution is detailed below from five key aspects: distributed acoustic sensor network deployment, multimodal feature extraction, semantic alignment of voice and speech, causal reinforcement training, and online adaptive updating. Below, this application will provide a detailed explanation and description of the above-mentioned equipment operation and maintenance methods in conjunction with specific application scenarios.

[0159] In a network operation and maintenance scenario of a 5G core equipment room, this voice-enabled self-diagnostic network operation and maintenance large model platform was deployed. First, 32 MEMS microphones (sampling rate 48kHz, 16-bit quantization, frequency range 20Hz-20kHz) were deployed near vulnerable components such as fans (1.2 meters away) and relays (0.8 meters away) and in the equipment compartment (grid spacing 4 meters). A hybrid network of LoRa (transmission rate 10kbps, coverage radius 8km) and Wi-Fi (802.11ac, transmission rate 433Mbps) was established, and the NTP protocol was used to achieve ±5ms time synchronization. The original voiceprint data is framed and windowed (Hamming window, frame length 512 points, frame shift 256 points), and then 1024-point FFT is used to extract PSD features (center frequency, bandwidth, etc.). Combined with Morlet mother wave CWT (scale 1-64) to extract time-frequency energy distribution and other features, after being fused by a fully connected layer (256→128 dimensions), a sparse fault spectrum fingerprint is generated by using a K-SVD dictionary (K=256) and the OMP algorithm (sparseness 10).

[0160] On the text side, BERT-base-uncased is used to extract 768-dimensional embeddings of operational speech (e.g., "the equipment keeps humming"). On the voiceprint side, a fully connected layer (256→768 dimensions) is projected onto the same space. InfoNCE loss (τ=0.07) is used to optimize contrastive learning, ensuring that the cosine distance between similar semantic voiceprint-text pairs is ≤0.28. The causal inference module takes voiceprint features as input and uses a Transformer decoder to sequentially generate root causes (e.g., "fan bearing wear," classification accuracy 92%), verification steps (e.g., "check speed fluctuations," BLEU-4=0.85), and repair instructions (e.g., "replace the bearing," BLEU-4=0.82). The PPO algorithm (γ=0.95, ε=0.2) is used for optimization to ensure logical coherence in the generation. During the online phase, when a suspicious sample with a confidence level of <0.7 is detected (such as a low-frequency abnormal noise in the 50-150Hz range), it is pushed to an expert for confirmation. The fault spectrum library is then updated through DBSCAN clustering (neighborhood radius of 0.4, minimum number of samples of 5). For every 100 new samples, the feature fusion weights (low-frequency component weight is increased to 0.7) and causal inference attention weights are dynamically adjusted, thereby increasing the accuracy of early weak anomaly identification from 65% to 85%.

[0161] In some embodiments, to further enhance multimodal sensing capabilities, vibration sensors (accelerometers, sampling rate 1kHz, range ±5g) can be deployed in conjunction with acoustic sensors. Vibration signals are processed using Empirical Mode Decomposition (EMD) to extract Intrinsic Mode Functions (IMFs). Combined with the STFT features of the acoustic signature, a cross-modal self-attention network (6 layers, 12 heads) is used to achieve deep fusion of acoustic and vibration features, generating a composite fault characterization including vibration amplitude, frequency, and acoustic signature energy distribution. On the text side, a network model can be used to extract context-aware embeddings of operational speech. These embeddings are then aligned with the fused acoustic and vibration features through a symmetric projection layer (768→768 dimensions, BatchNorm), and the semantic space is optimized to enhance cross-modal matching robustness. The causal reasoning module introduces a graph neural network (GNN) to construct a knowledge graph of "fault root cause - verification steps - repair instructions" (nodes include 200+ fault types, 500+ verification items, and 300+ repair actions). The interpretability of root cause inference is enhanced through a message passing mechanism (aggregating features of neighboring nodes) (e.g., the edge weight between the bearing wear node and the "speed fluctuation" verification node is increased to 0.85).

[0162] In some embodiments, the online update phase can adopt a federated learning framework, where local models in each data center (preserving acoustic and vibration feature encoder parameters) periodically upload gradients to the central server, and the global model is updated by weighted averaging (distributing weights according to sample size). While protecting local data privacy, this accelerates cross-data center knowledge sharing for new fault types (such as new power module resonance), enabling the accuracy of identifying similar faults to be improved by 15%-20% simultaneously in the three data centers.

[0163] A vibration sensor is a device that converts the mechanical vibration of an object into a measurable electrical signal.

[0164] The above embodiments, on the one hand, embed the operational verbal descriptions into the text using a pre-trained language model, and on the other hand, project the sparse fault spectrum fingerprint onto the same 768-dimensional space as the text embedding. Through contrastive learning to optimize the projection parameters, cross-modal semantic associations are formed, achieving semantic alignment between voiceprint and operational verbal descriptions, effectively combining acoustic information with operational knowledge. Traditional multimodal solutions struggle to achieve fine-grained alignment of voice and language, while this application, by constructing a shared semantic space, enables the system to understand the corresponding fault semantics based on voiceprint features, improving the accuracy and intelligence of fault diagnosis. On the other hand, chain-like causal reinforcement fine-tuning is implemented using multi-dimensional samples with root cause-action labels ("voiceprint-text-action"). The model takes voiceprint features as input and sequentially executes root cause inference, verification step generation, and repair instruction generation through the causal inference module. Each step outputs the matching degree with the label as a reward signal. The proximal policy optimization (PPO) algorithm is used to adjust the model parameters, enabling the model to perform causal inference based on voiceprint features and generate reasonable fault root causes, verification steps, and repair instructions. Traditional solutions lack this chain-like causal reasoning and reinforcement learning mechanism. The embodiments of this application ensure the logical coherence and interpretability of the generation process through this training method, thereby improving the scientific nature of fault diagnosis and handling. In addition, by automatically marking suspicious samples under two conditions and pushing them to operation and maintenance experts for confirmation, new fault classes are determined using DBSCAN clustering, and the fault spectrum library and sparse dictionary are updated. By adopting online stochastic gradient descent to dynamically adjust the model inference strategy and update the causal graph, the model's ability to capture early weak anomalies can be continuously improved, and it can continuously adapt to new fault types and scenarios. Traditional models struggle to dynamically adapt to new situations. This application's online active learning and adaptive update mechanism allows the model to continuously learn and evolve during operation, improving its generalization ability and practicality. Furthermore, this application employs a frequency-time domain joint encoder to process the raw voiceprint data. It extracts spectral features and time-frequency localization features through frame-by-frame windowing, Short-Time Fourier Transform (STFT), and Continuous Wavelet Transform (CWT), concatenating these features before inputting them into a fully connected layer for feature dimensionality reduction and fusion. A K-SVD dictionary learning algorithm is used for sparsity processing. This enables the extraction of multi-dimensional voiceprint representations including frequency distribution, energy changes, and instantaneous phase, suppressing environmental noise interference and retaining key features strongly correlated with faults. Traditional feature extraction methods may fail to comprehensively and accurately extract voiceprint features; this application's joint feature extraction method is more comprehensive and targeted, improving feature quality. Finally, this application deploys a distributed acoustic sensor network consisting of multiple acoustic sensor nodes. The nodes are selected as MEMS microphones, and the deployment location is determined based on the device structure and the acoustic propagation path of the fault. Each node is interconnected via LoRa or Wi-Fi technology, and the data is transmitted to the central server via TCP / IP protocol. The NTP protocol is used to realize the time synchronization of data of multiple nodes, which can capture different types of acoustic features in a targeted manner and ensure the spatiotemporal consistency of the voiceprint signal.Traditional sensor deployments may not be able to fully cover the acoustic characteristics of a device. The distributed deployment method of this invention can more accurately collect acoustic information during device operation, providing a more reliable data foundation for subsequent fault diagnosis.

[0165] This application's technical solution constructs a large-scale intelligent operation and maintenance system for equipment, integrating perception, fusion, reasoning, and evolution. At the perception layer, through the collaborative deployment of acoustic and vibration sensors and a cross-modal self-attention network, deep fusion of acoustic and vibration features is achieved, generating a composite fault representation containing multi-dimensional physical information, significantly improving the comprehensiveness of state perception. At the semantic understanding layer, the RoBERTa-large model and symmetric projection technology are employed to effectively optimize the semantic alignment between natural language descriptions and physical features, enhancing the ability to understand the context of operational language. At the reasoning layer, a graph neural network is innovatively introduced to construct a causal knowledge graph. Through a message passing mechanism, the logical connections between fault root causes, verification steps, and repair instructions are clarified, significantly improving the interpretability of the diagnostic process and the reliability of decision-making. At the system evolution layer, a federated learning framework is used to achieve collaborative optimization between edge nodes and the central server, realizing cross-data center fault knowledge sharing and continuous model evolution while strictly protecting local data privacy. This solution, through multi-level technological innovation, comprehensively addresses the core challenges of fault detection accuracy, interpretability, adaptability, and data security in complex industrial scenarios, providing a complete technological paradigm for intelligent operation and maintenance in modern industry.

[0166] It should be particularly noted that the steps in each embodiment of the above-described equipment testing and maintenance method can be overlapped, substituted, added, or deleted. Therefore, these reasonable permutations and combinations of the equipment testing and maintenance method should also fall within the protection scope of this disclosure, and the protection scope of this disclosure should not be limited to the described embodiments.

[0167] Based on the same inventive concept, this disclosure also provides a device for equipment detection and maintenance, as described in the following embodiments. Since the principle by which this device solves the problem is similar to that of the method embodiments described above, the implementation of this device embodiment can refer to the implementation of the method embodiments described above, and repeated details will not be repeated.

[0168] Figure 13 This is a block diagram illustrating a device for equipment testing and maintenance according to an exemplary embodiment. (Refer to...) Figure 13 The equipment detection and maintenance device 1300 provided in this embodiment may include: a first voiceprint signal acquisition module 1301, an actual fault information acquisition module 1302, a prediction module 1303, a reward function determination module 1304, and an optimization module 1305.

[0169] The first voiceprint signal acquisition module 1301 can be used to acquire the first voiceprint signal acquired when the first device experiences a first fault, and the first fault description information describing the first fault; the actual fault information acquisition module 1302 can be used to acquire the actual fault information of the first fault, the actual fault information including the actual fault type, the actual fault root cause description, the actual verification steps, and the actual repair instructions; the prediction module 1303 can be used to input the first voiceprint signal and the first fault description information into the target network model to obtain the predicted fault information output by the target network model for the first fault, the predicted fault information including the predicted fault type, the predicted fault root cause, the predicted verification steps, and the predicted repair instructions; the reward function determination module 1304 can be used to determine the first reward function based on the actual fault information and the predicted fault information; the optimization module 1305 can be used to train and optimize the target network model using the first reward function, so as to perform device fault detection and analysis based on the target network model.

[0170] It should be noted that the aforementioned first voiceprint signal acquisition module 1301, actual fault information acquisition module 1302, prediction module 1303, reward function determination module 1304, and optimization module 1305 correspond to S202 to S210 in the method embodiment. The examples and application scenarios implemented by these modules and their corresponding steps are the same, but they are not limited to the content disclosed in the above method embodiment. It should also be noted that these modules, as part of the device, can be executed in a computer system such as a set of computer-executable instructions.

[0171] In some embodiments, the device detection and maintenance apparatus 1300 may further include: a second voiceprint signal acquisition module, a second fault description acquisition module, a second information input module, a second reward function determination module, and a second training module.

[0172] The second voiceprint signal acquisition module can be used to acquire the second voiceprint signal collected when the second device experiences a second fault, as well as the actual fault type, actual fault root cause description, actual verification steps, and actual repair instructions of the second fault; the second fault description acquisition module can be used to acquire second fault description information, which is used to describe faults other than the second fault; the second information input module can be used to input the second voiceprint signal and the second fault description information into the target network model to obtain the predicted fault type, predicted fault root cause, predicted verification steps, and predicted repair instructions output by the target network for the second fault; the second reward function determination module can be used to determine the second reward function based on the actual fault type, actual fault root cause description, actual verification steps, and actual repair instructions of the second device, and the predicted fault type, predicted fault root cause, predicted verification steps, and predicted repair instructions of the second device; the second training module can be used to train and optimize the target network model using the second reward function.

[0173] In some embodiments, the target network model includes a feature encoder and a causal inferencer; wherein, the prediction module 1303 may include: a voiceprint temporal feature determination submodule, a natural language feature determination submodule, a fused fault feature determination submodule, and a fused fault feature input submodule.

[0174] The submodule for determining the voiceprint temporal features can be used to extract frequency domain features and temporal features from the first voiceprint signal using the feature encoder to obtain voiceprint temporal features and voiceprint frequency domain features; the submodule for determining the natural language features can be used to extract features from the first fault description information using the feature encoder to obtain natural language features; the submodule for determining the fused fault features can be used to fuse the voiceprint temporal features, the voiceprint frequency domain features, and the natural language features using the feature encoder to obtain fused fault features; and the submodule for inputting the fused fault features can be used to input the fused fault features into the target network model to obtain the predicted fault information.

[0175] In some embodiments, the voiceprint time-domain feature determination submodule may include: a multi-frame voiceprint sub-signal determination unit, a time-frequency feature extraction unit, a frequency domain feature determination unit, and a time-domain feature determination unit.

[0176] The multi-frame voiceprint sub-signal determination unit can be used to perform windowing and framing processing on the first voiceprint signal to obtain multi-frame voiceprint sub-signals; the time-frequency feature extraction unit can be used to extract the frequency domain features and time domain features of each frame voiceprint sub-signal; the frequency domain feature determination unit can be used to determine the voiceprint frequency domain features based on the frequency domain features of each frame voiceprint sub-signal; and the time domain feature determination unit can be used to determine the voiceprint time domain features based on the time domain features of each frame voiceprint sub-signal.

[0177] In some embodiments, the fusion fault feature determination submodule may include: a voiceprint fusion feature determination unit, a sparse dictionary determination unit, a sparse coefficient determination unit, a sparse fault spectrum fingerprint determination unit, and a feature fusion determination unit.

[0178] The voiceprint fusion feature determination unit can be used to fuse the voiceprint time-domain features and the voiceprint frequency-domain features to obtain voiceprint fusion features; the sparse dictionary determination unit can be used to construct an overcomplete sparse dictionary; the sparse coefficient determination unit can be used to solve the sparse coefficients of the voiceprint fusion features in combination with the overcomplete sparse dictionary; the sparse fault spectrum fingerprint determination unit can be used to determine the sparse fault spectrum fingerprint based on the sparse coefficients and the overcomplete sparse dictionary; and the feature fusion determination unit can be used to fuse the sparse fault spectrum fingerprint and the natural language features through the feature encoder to obtain the fused fault features.

[0179] In some embodiments, the equipment detection and maintenance device 1300 may further include a vibration signal determination module.

[0180] The vibration signal determination module can be used to acquire the vibration signal collected when the first device experiences the first fault.

[0181] The prediction module 1303 may include a prediction submodule.

[0182] The prediction submodule can be used to input the first acoustic signature signal, the first fault description information and the vibration signal into the target network model to obtain the predicted fault information output by the target network model for the first fault.

[0183] In some embodiments, the device detection and maintenance device 1300 may further include: a third voiceprint information determination module, a feature matching module, and a matching success module.

[0184] The third voiceprint information determination module can be used to acquire the third voiceprint information collected from the third device; the feature matching module can be used to perform feature matching between the third voiceprint signal and multiple known fault voiceprint signals; the matching success module can be used to input the third voiceprint signal into the target network model if the third voiceprint signal is successfully matched with at least one fault voiceprint signal, so that the target network model outputs predicted fault information for the third device.

[0185] In some embodiments, the equipment detection and maintenance device 1300 may further include: a fourth voiceprint information determination module, a similarity determination module, a candidate fault sample determination module, and a signal update module.

[0186] The fourth voiceprint information determination module can be used to acquire the fourth voiceprint information collected from the fourth device; the similarity determination module can be used to perform feature matching between the fourth voiceprint signal and multiple known fault voiceprint signals to determine the similarity between the fourth voiceprint signal and each known fault voiceprint signal; the candidate fault sample determination module can be used to determine the fourth voiceprint information as a candidate fault sample if the confidence of the predicted fault root cause in the predicted fault information corresponding to the fourth device is less than a first preset threshold or the similarity between the fourth voiceprint signal and all known fault voiceprint signals is less than a second preset threshold; and the signal update module can be used to update the known fault voiceprint signal based on the candidate fault sample.

[0187] In some embodiments, the device detection and maintenance is performed by the target edge device; the target network model includes a feature encoder and a causal inferencer; the device detection and maintenance device 1300 may also include a local storage module, a global parameter receiving module, and a parameter update module.

[0188] The local storage module can be used by the target edge device to locally store the parameters corresponding to the feature encoder and upload the parameters of the causal inference to the central server, so that the central server can update the global parameters corresponding to the causal inference based on the parameters uploaded by multiple edge devices; the global parameter receiving module can be used to receive the global parameters from the central server; and the parameter updating module can be used to update the parameters of the target network model based on the global parameters.

[0189] Since the functions of the device 1300 have been described in detail in their respective method embodiments, they will not be repeated here.

[0190] The modules and / or sub-modules and / or units described in the embodiments of this disclosure can be implemented in software or hardware. The described modules and / or sub-modules and / or units can also be located in a processor. The names of these modules and / or sub-modules and / or units do not, in some cases, constitute a limitation on the module and / or sub-module and / or unit itself.

[0191] The flowcharts and block diagrams in the accompanying drawings illustrate the architecture, functionality, and operation of possible implementations of systems, methods, and computer program products according to various embodiments of the present disclosure. In this regard, each block in a flowchart or block diagram may represent a portion of a module or program segment containing one or more executable instructions for implementing a specified logical function. It should also be noted that in some alternative implementations, the functions indicated in the blocks may occur in a different order than those indicated in the drawings. For example, two consecutively indicated blocks may actually be executed substantially in parallel, and they may sometimes be executed in reverse order, depending on the functions involved. It should also be noted that each block in a block diagram or flowchart, and combinations of blocks in a block diagram or flowchart, may be implemented using a dedicated hardware-based system that performs the specified function or operation, or using a combination of dedicated hardware and computer program instructions.

[0192] Furthermore, the above figures are merely illustrative of the processes included in the method according to exemplary embodiments of this disclosure and are not intended to be limiting. It is readily understood that the processes shown in the above figures do not indicate or limit the temporal order of these processes. Additionally, it is readily understood that these processes may be executed synchronously or asynchronously, for example, in multiple modules.

[0193] Figure 14 A schematic diagram of an electronic device suitable for implementing embodiments of the present disclosure is shown. It should be noted that... Figure 14 The illustrated electronic device 1400 is merely an example and should not be construed as limiting the functionality and scope of use of the embodiments disclosed herein.

[0194] like Figure 14 As shown, the electronic device 1400 includes a central processing unit (CPU) 1401, which can perform various appropriate actions and processes according to a program stored in a read-only memory (ROM) 1402 or a program loaded from a storage section 1408 into a random access memory (RAM) 1403. The RAM 1403 also stores various programs and data required for the operation of the electronic device 1400. The CPU 1401, ROM 1402, and RAM 1403 are interconnected via a bus 1404. An input / output (I / O) interface 1405 is also connected to the bus 1404.

[0195] The following components are connected to I / O interface 1405: an input section 1406 including a keyboard, mouse, etc.; an output section 1407 including a cathode ray tube (CRT), liquid crystal display (LCD), etc., and speakers, etc.; a storage section 1408 including a hard disk, etc.; and a communication section 1409 including a network interface card such as a LAN card, modem, etc. The communication section 1409 performs communication processing via a network such as the Internet. Drive 1410 is also connected to I / O interface 1405 as needed. Removable media 1411, such as a disk, optical disk, magneto-optical disk, semiconductor memory, etc., are installed on drive 1410 as needed so that computer programs read from it can be installed into storage section 1408 as needed.

[0196] In particular, according to embodiments of this disclosure, the processes described above with reference to the flowcharts can be implemented as computer software programs. For example, embodiments of this disclosure include a computer program product comprising a computer program carried on a computer-readable storage medium, the computer program containing computer program instructions for performing the methods shown in the flowcharts. In such embodiments, the computer program can be downloaded and installed from a network via communication section 1409, and / or installed from removable medium 1411. When the computer program is executed by central processing unit (CPU) 1401, it performs the functions defined above in the system of this disclosure.

[0197] It should be noted that the computer-readable storage medium disclosed herein may be a computer-readable signal medium or a computer-readable storage medium, or any combination thereof. A computer-readable storage medium may be, for example,—but not limited to—an electrical, magnetic, optical, electromagnetic, infrared, or semiconductor system, apparatus, or device, or any combination thereof. More specific examples of a computer-readable storage medium may include, but are not limited to: an electrical connection having one or more wires, a portable computer disk, a hard disk, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), optical fiber, portable compact disk read-only memory (CD-ROM), optical storage device, magnetic storage device, or any suitable combination thereof. In this disclosure, a computer-readable storage medium may be any tangible medium containing or storing a program that can be used by or in conjunction with an instruction execution system, apparatus, or device. In this disclosure, a computer-readable signal medium may include a data signal propagated in baseband or as part of a carrier wave, carrying computer-readable computer program instructions. Such propagated data signals may take various forms, including but not limited to electromagnetic signals, optical signals, or any suitable combination thereof. Computer-readable signal media can also be any computer-readable storage medium other than a computer-readable storage medium, which can send, propagate, or transmit a program for use by or in connection with an instruction execution system, apparatus, or device. Computer program instructions contained on a computer-readable storage medium can be transmitted using any suitable medium, including but not limited to: wireless, wire, optical fiber, RF, etc., or any suitable combination thereof.

[0198] In another aspect, this disclosure also provides a computer-readable storage medium, which may be included in the device described in the above embodiments; or it may exist independently and not assembled into the device. The computer-readable storage medium carries one or more programs, which, when executed by the device, enable the device to perform the following functions: acquiring a first voiceprint signal collected when a first fault occurs in the first device, and first fault description information describing the first fault; acquiring actual fault information of the first fault, including actual fault type, actual fault root cause description, actual verification steps, and actual repair instructions; inputting the first voiceprint signal and the first fault description information into a target network model to obtain predicted fault information output by the target network model for the first fault, including predicted fault type, predicted fault root cause, predicted verification steps, and predicted repair instructions; determining a first reward function based on the actual fault information and the predicted fault information; and training and optimizing the target network model using the first reward function to perform device fault detection and analysis based on the target network model.

[0199] According to one aspect of this disclosure, a computer program product or computer program is provided, comprising computer program instructions stored in a computer-readable storage medium. The computer program instructions are read from the computer-readable storage medium, and a processor executes the computer program instructions to implement the methods provided in various optional implementations of the above embodiments.

[0200] From the above description of the embodiments, those skilled in the art will readily understand that the exemplary embodiments described herein can be implemented by software or by combining software with necessary hardware. Therefore, the technical solutions of the embodiments of this disclosure can be embodied in the form of a software product, which can be stored in a non-volatile storage medium (such as a CD-ROM, USB flash drive, or portable hard drive) and includes several computer program instructions to cause an electronic device (such as a server or terminal device) to execute the method according to the embodiments of this disclosure.

[0201] Other embodiments of this disclosure will readily occur to those skilled in the art upon consideration of the specification and practice disclosed herein. This disclosure is intended to cover any variations, uses, or adaptations of this disclosure that follow the general principles of this disclosure and include common knowledge or customary techniques in the art not disclosed herein. The specification and examples are to be considered exemplary only, and the true scope and spirit of this disclosure are indicated by the claims.

[0202] It should be understood that this disclosure is not limited to the detailed structures, drawing arrangements or implementations shown herein; rather, this disclosure is intended to cover various modifications and equivalent arrangements contained within the spirit and scope of the appended claims.

Claims

1. A device detection operation and maintenance method, characterized in that, The method comprises: acquiring a first voiceprint signal collected when a first device has a first fault, and first fault description information describing the first fault; acquiring actual fault information of the first fault, the actual fault information comprising an actual fault type, an actual fault root cause description, an actual verification step, and an actual repair instruction; inputting the first voiceprint signal and the first fault description information into a target network model to obtain predicted fault information output by the target network model for the first fault, the predicted fault information comprising a predicted fault type, a predicted fault root cause, a predicted verification step, and a predicted repair instruction; determining a first reward function based on the actual fault information and the predicted fault information; training and optimizing the target network model using the first reward function, so as to perform device fault detection and analysis according to the target network model.

2. The method of claim 1, wherein, The method further comprises: acquiring a second voiceprint signal collected when a second device has a second fault, and an actual fault type, an actual fault root cause description, an actual verification step, and an actual repair instruction of the second fault; acquiring second fault description information for describing faults other than the second fault; inputting the second voiceprint signal and the second fault description information into the target network model to obtain a predicted fault type, a predicted fault root cause, a predicted verification step, and a predicted repair instruction output by the target network for the second fault; determining a second reward function according to the actual fault type, the actual fault root cause description, the actual verification step, and the actual repair instruction of the second device, and the predicted fault type, the predicted fault root cause, the predicted verification step, and the predicted repair instruction of the second device; training and optimizing the target network model using the second reward function.

3. The method of claim 1, wherein, The target network model comprises a feature encoder; wherein inputting the first voiceprint signal and the first fault description information into the target network model to obtain the predicted fault information output by the target network model for the first fault comprises: extracting frequency domain features and time domain features of the first voiceprint signal through the feature encoder to obtain voiceprint time domain features and voiceprint frequency domain features; extracting features of the first fault description information through the feature encoder to obtain natural language features; performing feature fusion on the voiceprint time domain features, the voiceprint frequency domain features, and the natural language features through the feature encoder to obtain fused fault features; inputting the fused fault features into the target network model to obtain the predicted fault information.

4. The method of claim 3, wherein, Extracting frequency domain features and time domain features of the first voiceprint signal through the feature encoder to obtain voiceprint time domain features and voiceprint frequency domain features comprises: performing windowing and framing processing on the first voiceprint signal to obtain multiple voiceprint sub-signals; extracting frequency domain features and time domain features of each voiceprint sub-signal; determining the voiceprint frequency domain features according to the frequency domain features of each voiceprint sub-signal; determining the voiceprint time domain features according to the time domain features of each voiceprint sub-signal.

5. The method of claim 4, wherein, The feature encoder is used for fusing the voiceprint time domain feature, the voiceprint frequency domain feature and the natural language feature to obtain a fusion fault feature, including: The voiceprint time domain feature and the voiceprint frequency domain feature are fused to obtain a voiceprint fusion feature; An over-complete sparse dictionary is constructed; The over-complete sparse dictionary is combined to solve sparse coefficients of the voiceprint fusion feature; Based on the sparse coefficients and the over-complete sparse dictionary, a sparse fault spectrum fingerprint is determined; The feature encoder is used for fusing the sparse fault spectrum fingerprint and the natural language feature to obtain the fusion fault feature.

6. The method of claim 1, wherein, The method further includes: An vibration signal collected when the first device occurs the first fault is acquired; The first voiceprint signal and the first fault description information are input into a target network model to obtain predicted fault information output by the target network model for the first fault, including: The first voiceprint signal, the first fault description information and the vibration signal are input into the target network model to obtain predicted fault information output by the target network model for the first fault.

7. The method of claim 1, wherein, The method further includes: Third voiceprint information collected for a third device is acquired; The third voiceprint signal is matched with a plurality of known fault voiceprint signals; If the third voiceprint signal is successfully matched with at least one fault voiceprint signal, the third voiceprint signal is input into the target network model so that the target network model outputs predicted fault information for the third device.

8. The method of claim 1, wherein, The method further includes: Fourth voiceprint information collected for a fourth device is acquired; The fourth voiceprint signal is matched with a plurality of known fault voiceprint signals to determine a similarity of the fourth voiceprint signal with each known fault voiceprint signal; If a confidence of a predicted fault root cause in predicted fault information corresponding to the fourth device is less than a first preset threshold or the similarity of the fourth voiceprint signal with all known fault voiceprint signals is less than a second preset threshold, the fourth voiceprint information is determined as a candidate fault sample; The known fault voiceprint signals are updated based on the candidate fault sample.

9. The method of claim 8, wherein, The device detection operation and maintenance are performed by a target edge device; The target network model includes a feature encoder and a causal reasoner; the method further includes: The target edge device locally stores parameters corresponding to the feature encoder and uploads parameters of the causal reasoner to a center server, so that the center server updates global parameters corresponding to the causal reasoner based on parameters uploaded by a plurality of edge devices; The global parameters are received from the center server; Parameters of the target network model are updated based on the global parameters.

10. An equipment detection operation and maintenance device, characterized in that, including: A first voiceprint signal acquisition module is configured to acquire a first voiceprint signal collected when a first device occurs a first fault, and first fault description information describing the first fault; An actual fault information acquisition module is configured to acquire actual fault information of the first fault, the actual fault information including an actual fault type, an actual fault root cause description, an actual verification step and an actual repair instruction; A prediction module is configured to input the first voiceprint signal and the first fault description information into a target network model to obtain predicted fault information output by the target network model for the first fault, the predicted fault information including a predicted fault type, a predicted fault root cause, a predicted verification step, and a predicted repair instruction. A reward function determination module is configured to determine a first reward function based on the actual fault information and the predicted fault information. An optimization module is configured to train and optimize the target network model using the first reward function, so as to perform device fault detection and analysis according to the target network model.

11. An electronic device, comprising: Comprise: a memory and a processor; The memory is used to store computer program instructions; the processor invokes the computer program instructions stored in the memory, and is used to implement the device detection operation and maintenance method according to any one of claims 1-9.

12. A computer-readable storage medium having stored thereon computer program instructions, wherein, The computer program instructions are executed by the processor to implement the device detection operation and maintenance method according to any one of claims 1-9.

13. A computer program product comprising computer program instructions stored in a computer readable storage medium, characterized in that, The computer program instructions are executed by the processor to implement the method according to any one of claims 1-9.