Sound equipment fault diagnosis system and method based on artificial intelligence

By building an AI-based audio fault diagnosis system and using graph convolution and attention mechanisms to extract audio system node features, the accuracy and adaptation efficiency issues of intelligent diagnosis of audio systems in existing technologies are solved, and rapid model adaptation and intelligent closed-loop control of audio systems are achieved.

CN120705734AInactive Publication Date: 2025-09-26CAF
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510805862.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-17
Publication Date
2025-09-26
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

Existing technologies make it difficult to achieve accurate, real-time, and transferable intelligent fault diagnosis of audio systems. This is especially true in multi-channel line array systems, where slight changes in a single node may affect the overall sound field through structural links. Traditional diagnostic methods lack the ability to deeply model the evolutionary trends of equipment operation and structural dependencies. In addition, the hardware structure differences between different audio equipment models lead to high model deployment and maintenance costs and low adaptation efficiency.

Method used

An artificial intelligence-based audio fault diagnosis system is adopted. By obtaining training data to train the equipment diagnosis model, a sound system structure diagram is constructed, and node features are extracted using graph convolution and attention mechanisms. Combined with dynamic behavior analysis, intelligent diagnosis and closed-loop control of the sound system are achieved.

Benefits of technology

It achieves rapid model adaptation of the audio system, improves the model migration efficiency in cross-model scenarios, enhances the robustness to local drift and differences between devices, can automatically trigger actions such as alarms, adjustments and isolation, and realizes intelligent closed-loop control from fault identification to system response.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120705734A_ABST
    Figure CN120705734A_ABST
Patent Text Reader

Abstract

The invention discloses a sound equipment fault diagnosis system and method based on artificial intelligence, and belongs to the technical field of fault diagnosis, and the method comprises the steps: inputting training data into an equipment diagnosis model for training, and outputting the parameters of the trained equipment diagnosis model; collecting original data of the sound equipment unit; calculating a standard residual tensor of the preset window; calculating a trend abnormal strength index of each unit; collecting equipment structure topology; constructing a sound system structure diagram; extracting a context feature of each node in a structural neighborhood through graph convolution, fusing a node state and a neighbor change degree through an attention mechanism to obtain a joint embedding vector, and calculating based on an embedding distance and a structural edge weight to obtain an attention matrix; and calculating a response score, and outputting a response action set and a response record. The embedded driving response module is diagnosed, actions such as alarm, adjustment and isolation are automatically triggered based on the node state and structure propagation relation, and intelligent closed-loop control from fault recognition to system response is truly achieved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of fault diagnosis, and in particular relates to an audio fault diagnosis system and method based on artificial intelligence. Background Art

[0002] With the widespread application of audio systems in diverse scenarios such as cultural performances, conference projects, and government and enterprise venues, the complexity of their equipment structures and the requirements for operational stability have increased simultaneously. Modern audio systems are often composed of multiple speaker units, amplifier channels, and DSP control modules, forming a large-scale, strongly coupled, time-varying device network. In this context, any sub-module of the system, such as unit mismatch, voltage fluctuation, structural installation offset, and abnormal control parameters, may cause problems such as degraded sound field quality, nonlinear distortion, and amplifier overload. Traditional operation and maintenance methods mostly rely on manual experience judgment, single-point measurement, or static threshold alarms. They lack the ability to deeply model the evolution of equipment operation trends and structural dependencies, making it difficult to achieve accurate, real-time, and portable intelligent diagnosis. Especially in multi-channel line array systems, slight changes in a single node may affect the overall sound field through structural links, making it difficult to fully identify system-level faults based solely on local electrical or acoustic characteristics.

[0003] In addition, there are many types of audio equipment and different deployment methods. There are significant differences in hardware structure, signal path, and control interface between different products. The existing solutions based on fixed models are difficult to use universally. The model deployment and maintenance costs are high and the adaptation efficiency is low, which has become an important bottleneck restricting the improvement of the industry's intelligent level.

[0004] To this end, we propose an artificial intelligence-based audio fault diagnosis system and method to solve the above problems. Summary of the Invention

[0005] The purpose of the present invention is to solve the problem in the prior art that it is difficult to achieve accurate, real-time and transferable intelligent diagnosis, and to propose an audio fault diagnosis system and method based on artificial intelligence.

[0006] In order to achieve the above object, the present invention adopts the following technical solutions:

[0007] The audio fault diagnosis method based on artificial intelligence includes:

[0008] S1: Acquire training data; input the training data into the equipment diagnostic model for training, and output the trained equipment diagnostic model parameters;

[0009] S2: Collecting raw data from the audio unit, forming the raw data into raw tensors and performing normalization processing to obtain a structured input tensor. The structured input tensor is a feature sequence within a time window, wherein a data validity mask is introduced. The data validity mask is used to record the completeness status of each channel data;

[0010] S3: Calculate a standard residual tensor of a preset window, wherein the standard residual tensor is obtained by combining a residual result and the data validity mask by element-by-element multiplication, and the residual result is calculated using the structured input tensor;

[0011] Extracting dynamic behavior from the standard residual tensor to obtain an output tensor;

[0012] The output tensors extracted from different dynamic behaviors are concatenated to output the embedding vector;

[0013] Collect the normal central trend of the same model of equipment, average the standard residual tensor in the time dimension, and obtain the average residual of each feature channel;

[0014] Calculating the difference between the average residual and the normal central trend, and taking the norm to obtain a trend anomaly strength index for each unit, where the trend anomaly strength index represents the intensity of deviation of the overall operating trend of the device from the normal mode within the time window;

[0015] S4: collecting device structure topology; constructing an audio system structure diagram, wherein the node set of the audio system structure diagram is composed of audio units, and the node features are the embedding vectors; constructing an edge set based on the device structure topology, introducing a basic edge weight for each edge, and the basic edge weight is used to represent the strength of the connection between nodes; and using the edge set to form a weighted adjacency matrix;

[0016] S5: Graph convolution is used to extract the contextual features of each node in the structural neighborhood. The node state and the degree of change of the neighbors are fused through the attention mechanism to obtain a joint embedding vector. The attention matrix is ​​calculated based on the embedding distance and the structural edge weight.

[0017] S6: Input the joint embedding vector and attention matrix to calculate the response score;

[0018] Execute a corresponding preset strategy based on the response score range; the corresponding strategy includes a response action set;

[0019] Output response action set and response record.

[0020] Preferably, the training data includes operating sample data of the device and corresponding labels, wherein the operating sample data includes acoustic characteristics, power amplifier status and DSP control parameters.

[0021] Preferably, the dynamic behavior in step S3 includes short-term mutations, structural disturbances and slow-changing trends.

[0022] Preferably, the dynamic behavior extraction of the standard residual tensor in step S3 is completed by a multi-convolution module, which includes three one-dimensional convolution layers with convolution kernel sizes of 3, 5, and 7, respectively, and channel outputs of 32; each layer is followed by a ReLU activation function and a maximum pooling operation to extract three types of dynamic behaviors: short-term mutations, structural perturbations, and slow-changing trends.

[0023] Preferably, the device structure topology in step S4 indicates whether there is a physical connection between units, and the format is an adjacency list or an adjacency matrix.

[0024] Preferably, a trend perception mechanism is also introduced into the edge weight in step S4, and the trend perception mechanism assigns a higher connection weight to a node combination with a drastic state change, and the degree of state change is measured by a trend anomaly intensity index.

[0025] Preferably, in step S6, the following corresponding preset strategies are executed according to the interval in which the response score falls:

[0026] When the response score is less than the first threshold, no action is taken and only the status is recorded;

[0027] When the response score is greater than or equal to the first threshold and less than the second threshold, a local alarm is triggered, calling the central alarm system to generate an alert. The fields include node number, anomaly type, and adjacency impact information;

[0028] When the response score is greater than or equal to the second threshold and less than the third threshold, the parameter adjustment module interface is called to adjust the audio gain, limiter, and EQ coefficient, and marked as being processed;

[0029] When the response score is greater than or equal to a third threshold and the node state difference is greater than a mismatch threshold, the hardware isolation interface is immediately called to shut down the node power amplifier channel to prevent system cascading failures; wherein, the node state difference is calculated by a joint embedding vector and is used to represent the diagnostic state difference of each node.

[0030] The audio fault diagnosis system based on artificial intelligence includes:

[0031] A model training module is used to obtain training data; input the training data into the device diagnostic model for training, and output the trained device diagnostic model parameters;

[0032] A data acquisition module is used to collect raw data from the audio unit, organize the raw data into raw tensors, and perform normalization processing to obtain a structured input tensor. The structured input tensor is a feature sequence within a time window, wherein a data validity mask is introduced to record the completeness status of each channel data;

[0033] A feature extraction module, wherein the feature extraction module is used to calculate a standard residual tensor of a preset window, wherein the standard residual tensor is obtained by combining the residual result and the data validity mask through element-by-element multiplication, and the residual result is calculated through the structured input tensor; dynamic behavior is extracted from the standard residual tensor to obtain an output tensor; the output tensors corresponding to different dynamic behaviors are spliced ​​and then output as an embedding vector; the normal central trend of devices of the same model is collected, and the standard residual tensors are averaged in the time dimension to obtain the average residual of each feature channel; the difference between the average residual and the normal central trend is calculated, and the norm is taken to obtain a trend anomaly strength index for each unit, wherein the trend anomaly strength index represents the deviation intensity of the overall operating trend of the device from the normal mode within the time window;

[0034] A structure graph construction module is used to collect device structure topology; construct an audio system structure graph, wherein the node set of the audio system structure graph is composed of audio units, and the node features are the embedding vectors; an edge set is constructed based on the device structure topology, and a basic edge weight is introduced for each edge, wherein the basic edge weight is used to represent the strength of the connection between nodes; and a weighted adjacency matrix is ​​formed using the edge set;

[0035] The diagnostic reasoning module is used to extract the contextual features of each node in the structural neighborhood through graph convolution, fuse the node state with the degree of change of the neighborhood through the attention mechanism to obtain a joint embedding vector, and calculate the attention matrix based on the embedding distance and the structural edge weight;

[0036] A strategy generation module is used to input a joint embedding vector and an attention matrix to calculate a response score; execute a corresponding preset strategy based on the interval in which the response score is located; the corresponding strategy includes a response action set; and output a response action set and a response record.

[0037] To sum up, the technical effects and advantages of the present invention are as follows: the present invention completes rapid model adaptation for new equipment through small sample embedding alignment and prototype fine-tuning mechanism, significantly improving the migration efficiency of the model in cross-model scenarios; then constructs a multi-source data access and structure-sensitive standardization mechanism for the operation period, introduces dynamic correction terms to enhance the robustness of the model to local drift and differences between devices, and embeds the diagnostic drive response module to automatically trigger alarms, adjustments, isolation and other actions based on the node status and structural propagation relationship, truly realizing intelligent closed-loop control from fault identification to system response. BRIEF DESCRIPTION OF THE DRAWINGS

[0038] Figure 1 is a flow chart of the method of the present invention;

[0039] Figure 2 Schematic diagram of the system structure of the present invention. DETAILED DESCRIPTION

[0040] The technical solutions in the embodiments of the present invention will be clearly and completely described below in conjunction with the drawings in the embodiments of the present invention. Obviously, the described embodiments are only part of the embodiments of the present invention, rather than all the embodiments.

[0041] like Figure 1 As shown, the audio fault diagnosis method based on artificial intelligence includes:

[0042] S1: Acquire training data; input the training data into the equipment diagnostic model for training, and output the trained equipment diagnostic model parameters;

[0043] S2: Collecting raw data from the audio unit, forming the raw data into raw tensors and performing normalization processing to obtain a structured input tensor. The structured input tensor is a feature sequence within a time window, wherein a data validity mask is introduced. The data validity mask is used to record the completeness status of each channel data;

[0044] S3: Calculate a standard residual tensor of a preset window, wherein the standard residual tensor is obtained by combining a residual result and the data validity mask by element-by-element multiplication, and the residual result is calculated using the structured input tensor;

[0045] Extracting dynamic behavior from the standard residual tensor to obtain an output tensor;

[0046] The output tensors extracted from different dynamic behaviors are concatenated to output the embedding vector;

[0047] Collect the normal central trend of the same model of equipment, average the standard residual tensor in the time dimension, and obtain the average residual of each feature channel;

[0048] Calculating the difference between the average residual and the normal central trend, and taking the norm to obtain a trend anomaly strength index for each unit, where the trend anomaly strength index represents the intensity of deviation of the overall operating trend of the device from the normal mode within the time window;

[0049] S4: collecting device structure topology; constructing an audio system structure diagram, wherein the node set of the audio system structure diagram is composed of audio units, and the node features are the embedding vectors; constructing an edge set based on the device structure topology, introducing a basic edge weight for each edge, and the basic edge weight is used to represent the strength of the connection between nodes; and using the edge set to form a weighted adjacency matrix;

[0050] S5: Graph convolution is used to extract the contextual features of each node in the structural neighborhood. The node state and the degree of change of the neighbors are fused through the attention mechanism to obtain a joint embedding vector. The attention matrix is ​​calculated based on the embedding distance and the structural edge weight.

[0051] S6: Input the joint embedding vector and attention matrix to calculate the response score;

[0052] Execute a corresponding preset strategy based on the response score range; the corresponding strategy includes a response action set;

[0053] Output response action set and response record.

[0054] The specific steps are as follows:

[0055] Step 1: Small Sample Adaptation Mechanism

[0056] The purpose of this step is to enable rapid adaptation of audio equipment diagnostic models in the absence of large-scale labeled data. This is especially true in the context of multi-model deployments of audio systems. Differences in structure, parameters, and signal behavior between different amplifier modules, line array speakers, or portable speakers can lead to performance degradation in unified models during initial deployment. Therefore, when initially deploying equipment, the model must be initialized or fine-tuned based on a small number of labeled samples to ensure fault detection in the field.

[0057] The input data consists of a small number of target device operation samples X′(i, t) and their corresponding manually confirmed labels y′(i). The structure of X′(i, t) is consistent with X(i, t) in step 2, including acoustic characteristics, amplifier status, DSP control parameters, etc. Data collection is carried out through the standard debugging process:

[0058] During the testing phase after equipment installation, a high-frequency, wideband standard test tone (such as pink noise) is used to drive the speaker unit. A microphone array collects the output of each speaker unit at a sampling rate of 44.1kHz. Simultaneously, the power amplifier module collects voltage, current, and temperature data at a sampling rate of 100Hz. The DSP controller reports current gain, EQ settings, and delay configuration parameters via the CAN bus or serial port. This data is synchronized and cached locally by an edge processing device (such as an industrial tablet or control host), generating a feature tensor X′(i,t) for each unit. During debugging, maintenance personnel play a preset audio sequence, observe the sound response curve and amplifier status, and label any abnormal units (e.g., "unit 1 voltage is low," "unit 3 frequency response is distorted"), generating a label y′(i) with a value range of {0,1}.

[0059] The model initialization adopts the embedding alignment + prototype fine-tuning strategy. First, the deployment end introduces a set of embedding models that have been trained by the headquarters. The model structure is a two-layer GRU (Gated Recurrent Unit) with 64 hidden units in each layer, which is dedicated to extracting dynamic representations of time series features. Given a set of collected data X′(i, t), the embedding network is input to obtain the embedding vector h′ of each sample. i :

[0060] h′ i =GRU2(GRU1(X′(i,t)))

[0061] Among them, GRU1 and GRU2 are the first and second layer gated recurrent units, respectively, and the hidden dimension of each layer is 64. X′(i,t) represents the feature tensor of the i-th device at time t, which contains acoustic, electrical, and DSP parameters and is in the form of a T×d matrix, where T is the number of time frames and d is the feature dimension (e.g., 54). h′ i Represents the final embedded state vector extracted by the unit, with a dimension of 64.

[0062] In order to achieve rapid adaptation, instead of retraining the complete classifier on the current site, the “prototype representation” method is used. The headquarters model maintains an embedded prototype c for each type of fault. k , represents the average embedding feature of the kth type of fault. These prototypes are obtained by clustering a large number of existing training samples and stored in the deployment package. The target device is embedded in the vector h′ i With each prototype c k Perform Euclidean distance calculation to find the fault type k corresponding to the minimum distance * , and use the following method to fine-tune the model parameters:

[0063]

[0064] Among them, f′ θ represents the fine-tuned model parameters in the current task; c y′(i) represents the prototype vector of the fault category corresponding to the label y′(i); h′ i is the state vector output by the embedding model; ||·|| represents the Euclidean distance; and θ is the adjustable weight parameter of the classification module (e.g., the output layer weight matrix). This process is optimized through gradient descent for approximately 5–10 rounds to achieve rapid adaptation.

[0065] The output is the adapted diagnostic model parameter f′ θ , used in the dynamic modeling process starting in step 3. This output is generated entirely based on actual collected data. The model structure has a clear source, consistent variable definitions, and a controllable processing process. It can be directly embedded in subsequent steps without repeated collection or conversion. This entire process, through a "light modeling + prototype constraints" approach, injects intelligent capabilities into the deployment phase, a key technical step in resolving engineering deployment obstacles.

[0066] Step 2: Multi-source data collection and standardization

[0067] The goal of this step is to complete the model adaptation of step 1 (i.e. f′ θ After having the initial diagnostic capability for the current equipment, the data access process for the operation period is officially started to construct a standardized input tensor X(i,t) for subsequent time series modeling. Unlike the small sample X′(i,t) in step one, this step processes continuous, real-time, multimodal operation data, and needs to be synchronized, screened, and dynamically corrected according to the model structure and normalization template generated in step one. This step not only completes the conventional data organization, but also specifically introduces structural correction factors and dynamic perception regularization terms designed for the operating characteristics of the audio system (such as non-constant noise interference, electrical temperature drift, and structural imbalance coupling), providing a more robust input structure for subsequent time embedding.

[0068] The input of this step is the diagnostic model parameter f′ output in the previous step θ , which contains the input structure definition accepted by the model (i.e., feature dimension, order, and normalization range) and the feature distribution learned from small samples. Real-time data on acoustic, electrical, and DSP control parameters collected after the equipment enters full operation also serves as an input source. All data is uniformly cached and synchronously aligned by edge nodes before participating in this processing step.

[0069] When the system is running, the acoustic signal from each audio unit i is collected and processed by the microphone array. The sampling rate is 44.1kHz, each frame is 25ms, and the frame shift is 10ms. The Mel filter bank is used to extract the 40-dimensional energy band and perform logarithmic compression to form x s(t,i). The power amplifier module collects voltage, current, and temperature data 100 times per second through a power monitoring chip (such as INA219); the DSP module regularly uploads control states such as gain, delay, and EQ at a frequency of approximately 1Hz. These signals are synchronized to a unified time axis through edge devices and form the original tensor X raw (i,t).

[0070] According to f′ θ Each type of characteristic normalization parameters recorded in the (such as the maximum and minimum values ​​of the acoustic channel, voltage range, EQ coefficient interval, etc.), each channel is normalized to form To enhance the system’s sensitivity to abnormal changes during operation, especially to handle local drift of speakers under high load or asymmetrical hanging, we designed and introduced a dynamic drift sensitivity term δ(i,t), which is used to adjust the impact of different channel characteristics on the overall embedding.

[0071] The tensors that can be used for modeling are defined as follows:

[0072]

[0073] Among them, α is the model adaptive adjustment coefficient, which is initially set to 0.85, and is determined by f′ θ Input or preset; δ(i,t) is the dynamic structure correction term, calculated as follows:

[0074]

[0075] in, represents the structural adjacency units connected to node i (generated from the graph in the subsequent step 4, but provided directly in this step based on the device configuration file); γ is a regularization factor, indicating the allowable tolerance for local structural drift, with a default value of 0.15. This design reflects a key feature of the patent: relative behavioral differences between different loudspeakers (such as acoustic attenuation and amplifier thermal drift) are often more meaningful in diagnosis than absolute values. Therefore, this term enhances the model's sensitivity to "relative anomalies."

[0076] Special Note: Because actual audio system deployments may experience missing nodes or sensor failures, we introduced a data validity mask m(i, t) during the tensor generation phase to improve the stability of subsequent modeling. This mask records the completeness of each channel's data in each frame. This mask value is passed as metadata to step 3 and used as an attention adjustment factor in subsequent embedding calculations.

[0077] Output:

[0078] Output 1: Structured input tensor X(i,t) for embedding modeling, with dimension T×d, where all channels are normalized and structure-sensitively adjusted;

[0079] Output 2: Data validity mask tensor m(i, t), which is used for node weight correction in subsequent model processing.

[0080] This step is to call the model structure f′ generated in step 1 θ , implementing a data stream standardization scheme customized for each device model and establishing a structured tensor X(i,t) that is fully consistent with the diagnostic model input. Unlike conventional "data preprocessing," this step incorporates a structure-sensitive regularization term δ(i,t) based on the unique physical characteristics of the audio system during operation. This not only improves the ability to perceive local drift but also ensures that the features of units at different physical locations maintain structural consistency during the modeling phase. This mechanism provides a more robust and engineering-adaptable input representation for graph structure modeling and temporal feature extraction.

[0081] It is important to note that although the name of step 2 contains “multi-source data collection”, this step is completely different from the small sample collection mechanism in step 1 in terms of function positioning and data source. Step 1 serves the initialization stage before model deployment. Its collection object is a very small number of manually labeled samples X′(i, t). The data comes from sampling in a debugging environment (such as playing test sounds, manual diagnosis confirmation), and the purpose is to use these representative samples to analyze the initial model structure f′. θ Make quick fine-tuning to adapt it to the characteristics of the current device model.

[0082] The "collection" in step 2 refers to the online data stream during the formal operation of the equipment, which is characterized by high frequency, no label, and continuous generation. The goal is to convert these raw signals into θ The standardized input tensor X(i,t) with the same defined structure is used for subsequent time series embedding and spatial modeling. Therefore, while both steps 1 and 2 involve data collection, the former is a small sample for initialization, while the latter is the main input pipeline for the ongoing inference system. The two do not overlap in terms of process stages, data targets, or processing strategies.

[0083] Step 3: Time Series Embedding Modeling

[0084] The goal of this step is to construct a state embedding vector h for each audio unit i’s running behavior within the time window T based on the standardized time series tensor X(i,t) output from step 2. i, providing highly expressive and robust node feature representations for subsequent structural graph modeling and diagnostic reasoning. Since the actual deployment of audio systems involves suspended array structures, high-power electrical loads, and complex and diverse audio environments, their operating status is often accompanied by the following significant characteristics: resonant interference caused by structural coupling, temperature drift effects caused by power supply fluctuations, and short-term distortion caused by spectral anomalies. These anomalies may not only appear in the "current frame", but are also likely to gradually accumulate in the form of slight changes in the "time evolution process." Therefore, traditional time series modeling solutions find it difficult to stably capture this "multimodal anomaly evolution trajectory" under limited sample conditions.

[0085] The key innovation of this step is to propose a time series residual-aware embedding structure. By modeling the residual trajectory of the standardized sequence, it introduces structural consistency regularization and variation trend suppression mechanisms, effectively enhancing the model's sensitivity to "mutation-type" and "accumulation-type" faults. Compared with the general GRU or Transformer solution, this method can combine the validity mask m(i,t) output from step 2 with the early adaptation model f′ θ The structural template realizes dynamic perception, robust compression and sparse highlighting of fault features.

[0086] enter:

[0087] The feature sequence within the time window output in step 2 has been normalized and structurally unified;

[0088] m(i,t)∈{0,1} T×d : The validity mask output from step 2 indicates whether the data of each channel in each frame is valid;

[0089] f′ θ : The adapted model parameters output in step 1 are used to load the pre-trained feature extraction layer weights and standard residual template.

[0090] To model the evolutionary behavior and weak mutation trends in the sequence, we first define a "residual reference trajectory" It is defined as the step 1 model f′ θ The average characteristic trajectory of the same type of equipment under standard working conditions is used as a reference benchmark to determine whether the current behavior deviates from the normal operating mode.

[0091] Then calculate the standard residual tensor R(i,t) of the current window:

[0092]

[0093] Here, ⊙ represents an element-by-element multiplication, which combines the residual result with the validity mask to retain only valid data channels. R(i,t) represents the frame-by-frame deviation trend of device i’s current state relative to normal operation and is an important input for diagnostic modeling.

[0094] Next, we construct a multi-scale convolutional module to perform sequence abstraction on R(i,t). The module structure consists of three one-dimensional convolutional layers (Conv1D) with kernel sizes of 3, 5, and 7, respectively, and a channel output of 32. Each layer is followed by a ReLU activation function and a max pooling operation to extract three types of dynamic behaviors: short-term mutations, structural perturbations, and slowly changing trends. The output tensors of these three paths are concatenated into a unified temporal representation Z(i), which is then connected to a fully connected layer to obtain the final embedding:

[0095] h i =σ(W z ·Concat[Conv3,Conv5,Conv7]+b)

[0096] Among them, W z is the output mapping matrix, σ is the activation function (using tanh), and the output dimension is 64. This structure uses multi-scale filters to jointly perceive different types of fluctuation behaviors in the residual trajectory. It is a key module designed for mixed abnormal conditions in audio operation (such as frequency drift, temperature rise, and gain jitter).

[0097] To further suppress non-diagnostic related periodic disturbances, we introduce an abnormal trend regularization term based on the statistical characteristics of the residuals The definition is as follows:

[0098]

[0099] Among them, μ R f′ θ The normal trend center of the same model device is represented by λ, which is a regulation weight (recommended setting of 0.1 to 0.3). This regularization term encourages the embedded model to automatically suppress behaviors with small mean deviations but significant distribution shifts, helping to improve the ability to identify small but long-term accumulated anomalies (such as slow current rise and frequency response drift).

[0100] Output:

[0101] Output 1: Embedding vector of each audio unit Contains its multi-scale dynamic behavior characteristics within the time window;

[0102] Output 2: Trend anomaly strength indicator for each unit Used to initialize and update edge connection weights in subsequent graph structures.

[0103] This step starts with the actual operating characteristics of the audio system, focusing on three types of diagnostic difficulties: "micro-changes," "slow drift," and "structural disturbances." It introduces a multi-scale residual modeling mechanism based on traditional sequence modeling, effectively enhancing the abnormal perception capability of state expression. At the same time, based on the model structure provided in step one and the data mask in step two, a time series processing mechanism of "alignment benchmark—differential modeling—dynamic screening" is constructed, making state embedding not only "encoding," but also "interpretative," "device adaptable," and "trend filtering capable." This embedding structure provides a robust input foundation for subsequent structural diagram construction and multi-channel cross-inference, and is the key interface for the entire patented system to achieve accurate fault perception.

[0104] Step 4: Structural diagram construction

[0105] This step aims to embed the node vector h based on the output of the previous stage i and trend anomaly indicators i , constructing a sound system structure diagram G = (V, E, H). This diagram will serve as the foundation for subsequent graph neural network reasoning, representing the physical relationships and logical dependencies between devices.

[0106] The input for this step is:

[0107] Node embedding vector The output from step 3 reflects the dynamic behavior of unit i;

[0108] Trend Anomaly Indicator The output from step 3 represents the state fluctuation intensity of the unit;

[0109] Topology Configuration x g (i, j), provided by the configuration file during deployment, indicates whether there is a physical connection between units i and j (such as cable connection or amplifier distribution order). The format is an adjacency list or adjacency matrix, and the unit is Boolean.

[0110] First, initialize the node set V = {v1,...,v n}, each node v i Corresponding to an audio unit. The feature of each node is directly taken from step 3 h i , forming a node feature matrix H = [h1,h2,...,h n ]. The matrix dimension is n×64, where n is the total number of devices and 64 is the embedding dimension.

[0111] Then, according to the structure topology x collected during deployment g (i,j) constructs the edge set E. If x g (i,j)=1, indicating that there is a structural connection between i and j, then an undirected edge e is added to the graph. ij .

[0112] Considering that the structural layout of audio equipment usually has spatial regularity (such as front-to-back adjacency and horizontal symmetry in a line array), we ij Introducing the basic edge weight w ij , used to indicate the strength of the connection between nodes. This edge weight not only relies on the physical structure but also introduces a trend-aware mechanism: for node pairs with drastic state changes, a higher connection weight is assigned to strengthen their propagation path in subsequent reasoning. This design is defined by the following formula:

[0113]

[0114] Among them, λ is the adjustment factor, which controls the degree of enhancement of the trend indicator to the edge weight, and the recommended value is 0.5; s i With s j Indicates the trend strength of the corresponding node, the unit is a dimensionless normalized indicator; x g (i,j) is the connection information read from the deployment file, and its value is 0 or 1. g (i,j)=0, then w ij =0, indicating that there is no structural connection between the two nodes.

[0115] The edge set E constructed in this way = {e ij |x g (i,j)=1} to form a sparse adjacency matrix Among them A ij =w ij The adjacency matrix will be used as one of the inputs of the graph neural network to model the structural dependencies between nodes.

[0116] In actual operation, the structural connection information x g (i, j) is typically entered into the system as a configuration file (e.g., JSON or XML format) containing a table of physical connection sequences. The deployment tool automatically imports this table through the configuration interface during device initialization and converts it into an adjacency matrix format in the graph construction module.

[0117] Output:

[0118] Output 1: structure graph G = (V, E, H), including node set, edge set and node feature matrix;

[0119] Output 2: Weighted adjacency matrix A, dimension n×n, used for propagation calculation of graph neural network.

[0120] This step is to fully use the output h of the previous stage i With s i Under the premise of combining the physical structure relationship provided during installation x g(i, j), a trend-aware graph structure G is constructed. This graph not only expresses the physical connections of the devices but also embeds node operating trends through edge weights, providing subsequent graph neural networks with structural input that better reflects the dynamic behavior characteristics of the audio system.

[0121] Step 5: Joint spatial-temporal diagnostic reasoning

[0122] The goal of this step is to achieve space-time joint feature modeling based on the structure graph G = (V, E, H) and adjacency matrix A constructed in step 4, that is, to embed h into the physical structure relationship between audio equipment and the time evolution state of each node. i , building a neural network model that can comprehensively analyze the dependencies between nodes and temporal dynamics. The model outputs the spatial-temporal diagnostic embedding z for each node i i , used by subsequent response modules to perform fault identification and system control.

[0123] In order to achieve spatial dependency modeling and temporal dynamic feature preservation of each node in the sound system, this step designs a graph-aware attentive embedding layer. This structure consists of two stages:

[0124] Graph adjacency-enhanced embedding extraction:

[0125] The contextual features of each node in the structural neighborhood are extracted through graph convolution operations, which are defined as follows:

[0126]

[0127] in, is the normalized adjacency matrix after adding self-connection; is the input feature matrix; W (1) is the graph convolution weight matrix (64×64); σ is the ReLU activation function. This operation enables each node i to obtain the weighted representation of the features of its structural neighbors.

[0128] Attention fusion layer:

[0129] Considering that faults between units in an audio system may propagate along the "physical coupling link", node state differences need to be paid special attention to. Therefore, an attention mechanism is designed to fuse the node state and the degree of change of the neighboring state to form a joint embedding:

[0130]

[0131] in, represents the set of nodes adjacent to i, α ij is the attention coefficient based on the embedding distance and structural edge weight, which is calculated as follows:

[0132]

[0133] Attention value α ij Higher weights are assigned to neighbors with strong structural connections and high state similarity, enabling the model to have the capability of "structure-guided fault attention enhancement", which is particularly suitable for handling local collaborative anomalies caused by power amplifier resonance, linear array offset, etc.

[0134] After combining the above two steps, we can get the space-time joint embedding It not only contains the temporal embedding information of node i itself, but also integrates the weighted dynamic states of its structural neighbors, providing accurate input for subsequent decision-making.

[0135] Output:

[0136] Output 1: Joint spatial-temporal embedding z i , represents the state representation of each node i under the comprehensive structural relationship and temporal dynamics;

[0137] Output 2: Attention matrix α ij , visualize the spread of attention between nodes and provide a basis for system interpretability analysis.

[0138] Step 6: Diagnosis-driven system response

[0139] This step is the end point of this patent solution. The goal is to complete the spatial-temporal joint diagnosis modeling of each device unit of the sound system in the previous stage, based on its output results z i With the structural attention weight α ij , automatically executing a set of clearly defined and controllable system response actions. These responses directly trigger the control interface to complete operations such as alarm push, operating parameter adjustment, and unit isolation control, ensuring that the system can react immediately when potential anomalies are discovered, and forming a closed loop of response records.

[0140] enter:

[0141] Joint spatial-temporal embedding vector Indicates the diagnostic status of each device node;

[0142] Node attention weight matrix α ij , represents the state propagation strength under the structural dependence between nodes;

[0143] The graph structure meta-information (such as channel number, device model, and controllable interface table) is imported by the device configuration system.

[0144] The system response logic first passes the input embedding vector z i With neighbor attention α ij Calculate the response score r i, used to determine whether node i needs to respond and the response level, defined as follows:

[0145]

[0146] in:

[0147] w r is the response weight vector (predefined or pretrained);

[0148] ||z i -z j || 2 is the node status difference;

[0149] β is the structural enhancement factor (default 0.5);

[0150] σ is the Sigmoid function, so that r i ∈[0,1].

[0151] According to r i The system performs the following operations in the range, all of which are issued through the device control interface and do not require manual intervention:

[0152] r i <0.4: no action, only recording status;

[0153] 0.4≤r i <0.7: triggers a local alarm and calls the central alarm system (via CAN bus or TCP API) to generate an alert. The fields include node number, anomaly type, and adjacency impact information.

[0154] 0.7≤r i <0.85: Call the parameter adjustment module interface (such as DSP control API) to adjust audio gain, limiter, EQ coefficient, etc., and mark the "processing" status;

[0155] r i ≥0.85 and (Structural mismatch is obvious, γ is the default value of 0.6): Immediately call the hardware isolation interface to shut down the power amplifier channel of the node (through power amplifier control instructions or power module relay control) to prevent system cascading failures.

[0156] The system performs actions in the following specific ways:

[0157] DSP parameter adjustment: Send predefined control packets (JSON format) through the serial port or TCP to update the target DSP control parameters;

[0158] Alarm output: Generate alarm notifications through the platform's central control system API, with the level matching the response level;

[0159] Power isolation: Calling the control instructions of the subsequent power amplifier (such as setting the channel Mute state or physical circuit breaking);

[0160] All control actions are recorded as "response record entries", including execution time, interface call log, instruction content and action results.

[0161] Output:

[0162] Output 1: Response action set Each element is the action type (alarm, adjustment, isolation) and execution interface of node i;

[0163] Output 2: Response record packet L = {l i}, each l i Contains response timestamp, response indicator r i , instruction log, action status code and other fields are used for operation and maintenance backtracking and security auditing.

[0164] The technical solutions in the above-mentioned embodiments of the present application have at least the following technical effects or advantages: the present invention completes rapid model adaptation for new equipment through a small sample embedding alignment and prototype fine-tuning mechanism, significantly improving the migration efficiency of the model in cross-model scenarios; then constructs a multi-source data access and structure-sensitive standardization mechanism for the operation period, introduces dynamic correction terms to enhance the robustness of the model to local drift and differences between devices, and embeds a diagnostic drive response module to automatically trigger alarms, adjustments, isolation and other actions based on the node status and structural propagation relationship, truly realizing intelligent closed-loop control from fault identification to system response.

[0165] The present application also provides an artificial intelligence-based audio fault diagnosis system, such as Figure 2 As shown, including:

[0166] A model training module is used to obtain training data; input the training data into the device diagnostic model for training, and output the trained device diagnostic model parameters;

[0167] A data acquisition module is used to collect raw data from the audio unit, organize the raw data into raw tensors, and perform normalization processing to obtain a structured input tensor. The structured input tensor is a feature sequence within a time window, wherein a data validity mask is introduced to record the completeness status of each channel data;

[0168] A feature extraction module, wherein the feature extraction module is used to calculate a standard residual tensor of a preset window, wherein the standard residual tensor is obtained by combining the residual result and the data validity mask through element-by-element multiplication, and the residual result is calculated through the structured input tensor; dynamic behavior is extracted from the standard residual tensor to obtain an output tensor; the output tensors corresponding to different dynamic behaviors are spliced ​​and then output as an embedding vector; the normal central trend of devices of the same model is collected, and the standard residual tensors are averaged in the time dimension to obtain the average residual of each feature channel; the difference between the average residual and the normal central trend is calculated, and the norm is taken to obtain a trend anomaly strength index for each unit, wherein the trend anomaly strength index represents the deviation intensity of the overall operating trend of the device from the normal mode within the time window;

[0169] A structure graph construction module is used to collect device structure topology; construct an audio system structure graph, wherein the node set of the audio system structure graph is composed of audio units, and the node features are the embedding vectors; an edge set is constructed based on the device structure topology, and a basic edge weight is introduced for each edge, wherein the basic edge weight is used to represent the strength of the connection between nodes; and a weighted adjacency matrix is ​​formed using the edge set;

[0170] The diagnostic reasoning module is used to extract the contextual features of each node in the structural neighborhood through graph convolution, fuse the node state with the degree of change of the neighborhood through the attention mechanism to obtain a joint embedding vector, and calculate the attention matrix based on the embedding distance and the structural edge weight;

[0171] A strategy generation module is used to input a joint embedding vector and an attention matrix to calculate a response score; execute a corresponding preset strategy based on the interval in which the response score is located; the corresponding strategy includes a response action set; and output a response action set and a response record.

[0172] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. An artificial intelligence-based audio fault diagnosis method, characterized in that: include: S1: Acquire training data; input the training data into the equipment diagnostic model for training, and output the trained equipment diagnostic model parameters; S2: Collecting raw data from the audio unit, forming the raw data into raw tensors and performing normalization processing to obtain a structured input tensor. The structured input tensor is a feature sequence within a time window, wherein a data validity mask is introduced. The data validity mask is used to record the completeness status of each channel data; S3: Calculate a standard residual tensor of a preset window, wherein the standard residual tensor is obtained by combining a residual result and the data validity mask by element-by-element multiplication, and the residual result is calculated using the structured input tensor; Extracting dynamic behavior from the standard residual tensor to obtain an output tensor; The output tensors extracted from different dynamic behaviors are concatenated to output the embedding vector; Collect the normal central trend of the same model of equipment, average the standard residual tensor in the time dimension, and obtain the average residual of each feature channel; Calculating the difference between the average residual and the normal central trend, and taking the norm to obtain a trend anomaly strength index for each unit, where the trend anomaly strength index represents the intensity of deviation of the overall operating trend of the device from the normal mode within the time window; S4: collecting device structure topology; constructing an audio system structure diagram, wherein the node set of the audio system structure diagram is composed of audio units, and the node features are the embedding vectors; constructing an edge set based on the device structure topology, introducing a basic edge weight for each edge, and the basic edge weight is used to represent the strength of the connection between nodes; and using the edge set to form a weighted adjacency matrix; S5: Graph convolution is used to extract the contextual features of each node in the structural neighborhood. The node state and the degree of change of the neighbors are fused through the attention mechanism to obtain a joint embedding vector. The attention matrix is ​​calculated based on the embedding distance and the structural edge weight. S6: Input the joint embedding vector and attention matrix to calculate the response score; Execute a corresponding preset strategy based on the response score range; the corresponding strategy includes a response action set; Output response action set and response record.

2. The method for diagnosing audio faults based on artificial intelligence according to claim 1, characterized in that: The training data includes operating sample data of the device and corresponding labels, and the operating sample data includes acoustic characteristics, power amplifier status and DSP control parameters.

3. The method for diagnosing audio faults based on artificial intelligence according to claim 1, characterized in that: The dynamic behaviors described in step S3 include short-term mutations, structural disturbances, and slow-changing trends.

4. The method for diagnosing audio faults based on artificial intelligence according to claim 1, characterized in that: The dynamic behavior extraction of the standard residual tensor described in step S3 is completed by a multi-convolution module, which contains three one-dimensional convolution layers with convolution kernel sizes of 3, 5, and 7, respectively, and channel outputs of 32; each layer is followed by a ReLU activation function and a maximum pooling operation to extract three types of dynamic behaviors: short-term mutations, structural perturbations, and slow-changing trends.

5. The artificial intelligence-based audio fault diagnosis method according to claim 1, characterized in that: The device structure topology described in step S4 indicates whether there is a physical connection between units, and the format is an adjacency list or an adjacency matrix.

6. The method for diagnosing audio faults based on artificial intelligence according to claim 1, characterized in that: A trend perception mechanism is also introduced into the edge weights described in step S4. The trend perception mechanism assigns higher connection weights to node combinations with drastic state changes. The degree of state change is measured by a trend anomaly intensity index.

7. The method for diagnosing audio faults based on artificial intelligence according to claim 1, characterized in that: In step S6, the following corresponding preset strategies are executed according to the interval in which the response score falls: When the response score is less than the first threshold, no action is taken and only the status is recorded; When the response score is greater than or equal to the first threshold and less than the second threshold, a local alarm is triggered, calling the central alarm system to generate an alert. The fields include node number, anomaly type, and adjacency impact information; When the response score is greater than or equal to the second threshold and less than the third threshold, the parameter adjustment module interface is called to adjust the audio gain, limiter, and EQ coefficient, and marked as being processed; When the response score is greater than or equal to a third threshold and the node state difference is greater than a mismatch threshold, the hardware isolation interface is immediately called to shut down the node power amplifier channel to prevent system cascading failures; wherein, the node state difference is calculated by a joint embedding vector and is used to represent the diagnostic state difference of each node.

8. The audio fault diagnosis system based on artificial intelligence is characterized by: include: A model training module, wherein the model training module is used to obtain training data; Inputting the training data into a device diagnostic model for training, and outputting trained device diagnostic model parameters; A data acquisition module is used to collect raw data from the audio unit, organize the raw data into raw tensors, and perform normalization processing to obtain a structured input tensor. The structured input tensor is a feature sequence within a time window, wherein a data validity mask is introduced to record the completeness status of each channel data; A feature extraction module is used to calculate a standard residual tensor of a preset window, wherein the standard residual tensor is obtained by combining a residual result and the data validity mask by element-by-element multiplication, and the residual result is calculated using the structured input tensor; Extracting dynamic behavior from the standard residual tensor to obtain an output tensor; splicing the output tensors corresponding to different dynamic behaviors to output an embedded vector; collecting the normal central trend of devices of the same model, averaging the standard residual tensors in the time dimension to obtain the average residual of each feature channel; calculating the difference between the average residual and the normal central trend, and taking the norm to obtain the trend anomaly strength index of each unit, the trend anomaly strength index indicating the intensity of deviation of the overall operating trend of the device from the normal mode within the time window; A structure graph construction module is used to collect device structure topology; construct an audio system structure graph, wherein the node set of the audio system structure graph is composed of audio units, and the node features are the embedding vectors; an edge set is constructed based on the device structure topology, and a basic edge weight is introduced for each edge, wherein the basic edge weight is used to represent the strength of the connection between nodes; and a weighted adjacency matrix is ​​formed using the edge set; The diagnostic reasoning module is used to extract the contextual features of each node in the structural neighborhood through graph convolution, fuse the node state with the degree of change of the neighborhood through the attention mechanism to obtain a joint embedding vector, and calculate the attention matrix based on the embedding distance and the structural edge weight; A strategy generation module is used to input a joint embedding vector and an attention matrix to calculate a response score; execute a corresponding preset strategy based on the interval in which the response score is located; the corresponding strategy includes a response action set; and output a response action set and a response record.

Citation Information

Cited By

  • Electrical fire monitoring identification method and system based on electrical load fluctuation

    CN121075048A