Fault diagnosis method and system based on differential attention double-branch network and large language model interpretation

By combining differential attention dual-branch networks and large language models, the robustness and interpretability issues of multi-sensor fault diagnosis under complex working conditions are solved, achieving high-accuracy fault diagnosis and transparent user explanation, thus improving engineering efficiency.

CN121980480APending Publication Date: 2026-05-05SHANGHAI UNIV
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
SHANGHAI UNIV
Filing Date
2026-04-08
Publication Date
2026-05-05

AI Technical Summary

Technical Problem

Existing data-driven multi-sensor fault diagnosis methods lack robustness under complex operating conditions, and lack stable and verifiable sensor-level evidence chains and user-friendly interpretations, which affects the efficiency of engineering implementation.

Method used

We employ a differential attention dual-branch network and a large language model to extract statistical deviation features of sensor channels through paired window differential modeling, generate channel importance scores, and use channel selection probability to guide cross-attention calculation to output structured diagnostic evidence. We then combine the large language model to generate a natural language report.

Benefits of technology

It improves the robustness and reliability of diagnosis, forms a stable sensor-level evidence chain, outputs transparent and interpretable diagnostic reports, and enhances engineering usability.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121980480A_ABST
    Figure CN121980480A_ABST
Patent Text Reader

Abstract

The invention discloses a fault diagnosis method and system based on a differential attention double-branch network and large language model interpretation, and relates to the technical field of industrial system intelligent operation and maintenance and multi-sensor fault diagnosis. The method comprises the following steps: constructing paired input of a to-be-diagnosed window and a reference window and carrying out differential modeling; in a differential branch, a channel cross attention mechanism modulated by channel selection probability prior is adopted, and adaptive focusing and contribution redistribution are carried out on a key sensor channel; meanwhile, a sharing representation branch is introduced to describe a working condition generality and a cross-sample stability mode, and fusion with a differential branch is carried out at a decision-making layer; diagnosis bases such as key sensors, statistical characteristics of the key sensors, channel attention weights and the like are normalized into structured evidences, and the structured evidences are input into a large language model to generate a natural language diagnosis report. According to the method, high diagnosis accuracy can be kept under complex working conditions, and stable and reviewable fault diagnosis results can be output.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of intelligent operation and maintenance and multi-sensor fault diagnosis technology for industrial systems, and in particular to a fault diagnosis method and system based on differential attention dual-branch network and large language model interpretation. Background Technology

[0002] Industrial systems (such as refrigeration units, HVAC systems, and industrial process units) are typically equipped with numerous sensors for monitoring and control. Due to load variations, environmental disturbances, and control compensation, multi-sensor time series data exhibit strong coupling, high noise levels, and significant operational drift. While existing data-driven diagnostic methods can output fault categories, they generally suffer from the following shortcomings: 1) Changes in operating conditions cause fault characteristics to overlap with normal changes, resulting in insufficient robustness; 2) The model output is mainly based on categories and lacks a transparent chain of evidence of "key sensors - key time segments - evidence indicators", making it difficult to verify; 3) Importance or attention fluctuates significantly across samples, making it difficult to form stable and consistent sensor attribution results; 4) The lack of natural language explanations and handling suggestions for operation and maintenance personnel affects the efficiency of project implementation.

[0003] Therefore, there is an urgent need for a fault diagnosis solution that can maintain high diagnostic accuracy under complex operating conditions and output stable, verifiable sensor-level evidence and user-friendly interpretation. Summary of the Invention

[0004] The technical problem to be solved by this invention is how to provide a method that can maintain a high diagnostic accuracy under complex working conditions and output stable and verifiable fault diagnosis results.

[0005] To solve the above-mentioned technical problems, the technical solution adopted by the present invention is: a fault diagnosis method based on differential attention dual-branch network and large language model interpretation, comprising the following steps: Acquire time series data from multiple sensors, construct a window to be diagnosed and a reference window to form a pair of inputs, and calculate the difference information between the window to be diagnosed and the reference window. Based on the differential information, the statistical deviation features of each sensor channel are extracted, a channel importance score is generated, and the channel selection probability is obtained through an end-to-end trainable channel selection module. The channel selection probability is used to guide the channel cross-attention calculation to enhance key channels and obtain differential discrimination features; common information of working conditions is extracted to obtain shared discrimination features, and the classification outputs corresponding to the differential discrimination features and the shared discrimination features are fused to obtain the fault category and confidence level; Output the set of key channels and their corresponding statistical features, and output the channel attention weights to form structured diagnostic evidence; The structured diagnostic evidence, along with the fault category and confidence level, is input into a large language model to generate and output a natural language diagnostic report.

[0006] This invention also discloses a fault diagnosis system based on differential attention dual-branch network and large language model interpretation, comprising: The paired window differential modeling module is used to construct a window to be diagnosed and a reference window and form a paired input, and calculate differential information to characterize the deviation; wherein, the reference window is retrieved from the normal data pool according to the similarity of the operating conditions, and the similarity of the operating conditions is based on at least one or more of environmental variables, load variables, and control variables; The statistically guided channel selection module is used to extract statistical deviation features of each channel based on the difference sequence to form a channel importance score, and outputs the channel selection probability by the end-to-end trainable channel selection module. Differential Attention Module: Used to guide channel cross-attention calculation using channel selection probability, so that key channels are enhanced in attention matching and aggregation, thereby forming differential discriminative features; Shared representation and fusion classification module: used to extract common information about operating conditions in parallel and fuse it with differential branches to output fault categories and confidence levels; The structured evidence and large language model interpretation module is used to output key channel sets and their statistical features and attention weights to form structured diagnostic evidence, which is then input into the large language model to generate a natural language diagnostic report. It also incorporates device knowledge base retrieval enhancements and performs consistency checks and filtering / corrections on the output.

[0007] The beneficial effects of adopting the above technical solution are as follows: 1) By highlighting deviation features through paired window differential modeling, the interference of operating condition drift on diagnosis is reduced, and robustness is improved; 2) The linkage between channel selection and differential attention enables key sensors to not only be identified, but also play an actual role in feature calculation, thereby improving the credibility of sensor-level attribution; 3) Shared representations supplement common operating condition information and are fused with differential output, thereby improving class separability and generalization ability; 4) Outputting structured evidence (key sensors + statistical comparison + attention) to form a verifiable evidence chain; 5) The large language model interpretation and generation driven by structured evidence can output transparent and interpretable diagnostic reports and treatment suggestions for users, thereby improving engineering usability. Attached Figure Description

[0008] The present invention will now be described in further detail with reference to the accompanying drawings and specific embodiments.

[0009] Figure 1 This is a flowchart of the detection method described in Embodiment 1 of the present invention; Figure 2 This is a schematic diagram illustrating the selection of statistical guidance channels and their guidance on channel cross-attention in the detection method described in Embodiment 1 of the present invention; Figure 3 This is a schematic diagram of the structured diagnostic evidence-driven large language model interpretation generation and consistency verification process in the detection method described in Embodiment 1 of the present invention; Figure 4 This is a schematic diagram of the system described in Embodiment 2 of the present invention; Figure 5 This is a schematic diagram of the device described in Embodiment 3 of the present invention. Detailed Implementation

[0010] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of the embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of the present invention.

[0011] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.

[0012] Example 1 like Figure 1 As shown in the figure, this invention discloses a method for detecting defects in multiple categories of industrial products based on multi-source semantic anchor point constraint reconstruction. The method specifically includes the following steps: Acquire time series data from multiple sensors, construct a window to be diagnosed and a reference window to form a pair of inputs, and calculate the difference information between the window to be diagnosed and the reference window. Based on the differential information, the statistical deviation features of each sensor channel are extracted, a channel importance score is generated, and the channel selection probability is obtained through an end-to-end trainable channel selection module. The channel selection probability is used to guide the calculation of channel cross-attention to enhance key channels and obtain differential discrimination features; common information of working conditions is extracted to obtain shared discrimination features, and the classification outputs corresponding to the differential discrimination features and the shared discrimination features are fused to obtain the fault category and confidence level; Output the set of key channels and their corresponding statistical features, and output the channel attention weights to form structured diagnostic evidence; The structured diagnostic evidence, along with the fault category and confidence level, is input into a large language model to generate and output a natural language diagnostic report.

[0013] The above steps will be explained in detail below with reference to specific content: Furthermore, the reference window is retrieved from the normal data pool based on operating condition similarity, which is based on at least one or more of environmental variables, load variables, and control variables.

[0014] Furthermore, the paired window representation and residual construction method includes the following steps: The multi-sensor time series is sliced ​​according to the window length T to obtain the diagnostic window X and the reference window. Where C represents the number of sensor categories and T represents the window length; The reference window is preferably retrieved from the normal data pool based on operating condition similarity, which is based on at least one or more of environmental variables, load variables, and control variables (such as temperature, flow rate, valve position, frequency command, etc.).

[0015] After constructing paired inputs, the difference information is calculated, preferably in the form of residuals: The reference window is used to provide a "normal baseline under the same operating conditions", so that the model focuses on deviations rather than absolute value changes.

[0016] The statistical deviation features include at least the mean and standard deviation of the difference sequence, and the channel importance score is obtained by combining the statistical deviation features with a learnable correction term output by a lightweight network.

[0017] Furthermore, the difference branch first constructs a channel anomaly prior. In this embodiment, the statistical deviation features only use the mean and variance (standard deviation) information of the difference sequences. For each channel i, the difference sequence... Calculate the deviation amplitude statistic: in This represents the standard deviation calculation. The aggregation norm with respect to the time dimension is preferably implemented as the average absolute aggregation, so as to reflect both the average deviation and the deviation fluctuation without introducing additional complex statistical terms. This represents the data of the i-th channel of the window to be diagnosed. This represents the data of the i-th channel of the reference window.

[0018] To enhance the reliability and task relevance of channel anomaly priors, a learnable channel-guided score is introduced: in , , , For learnable parameters, Let represent the learnable guided score of the i-th channel at time step t. This represents the learnable guided score of the i-th channel within the window, used to supplement statistical deviation features and guide channel priors; Finally, the prior score of the channel anomaly was obtained. : .

[0019] To obtain the key channel set, this embodiment uses a channel selection module that can be trained end-to-end to output the channel selection probability vector m. The channel selection module can be implemented using continuous relaxation, differentiable sorting, gating networks, etc., and this invention does not limit its specific form. As a preferred implementation example, a continuously relaxed Top-K selection, such as LapSum, can be used: Solve for the threshold b such that: Where F For continuous gating in the form of cumulative distribution function, For smoothing coefficients, Let m be the desired number of critical channels to select, m be the importance probability of the critical sensors, and b be the adaptive threshold. This represents the soft selection value for channel i. This represents the normalized channel selection probability.

[0020] In differential channel cross-attention, channels are treated as tokens, and the target window is used as... Use the reference window as , Furthermore, a prior selection is introduced to highlight key channels, and the differential attention logit is represented as: , Where Q represents the set of query vectors obtained by linear projection of the target window X, and K and V represent the set of query vectors obtained by linear projection of the reference window X. The set of key vectors and the set of value vectors obtained by linear projection; m represents the normalized channel selection probability; Map the channel selection prior to the logit space; For temperature coefficient, This is a numerical stability constant. In addition, the key values ​​of the reference window can be weighted with a lower bound to avoid the complete inactivation of non-critical channels and improve numerical stability.

[0021] The cross-attention guidance channels include at least one or more of the following methods: a) Construct an attention bias term based on the channel selection probability and apply it to the attention similarity calculation (basic idea); b) Weight the key features of the reference window based on the channel selection probability and set a non-zero lower bound to avoid the complete deactivation of non-critical channels (engineering design). c) Mask the attention weights based on the channel selection probability and renormalize them (engineering design).

[0022] like Figure 2 The diagram illustrates the statistical guidance channel selection and its guidance on channel cross-attention. Shared branches are used to extract common operating condition information, employing channel cross-attention. in, This represents the shared branch channel attention weight matrix, where d represents the feature dimension of the query or key vector. Indicates shared branch attention output; Residual-protected fusion is performed in the differential branch to enhance the bias component and suppress common interference: in The coefficients are fusion coefficients. Temporal features are then extracted using multi-branch one-dimensional convolution, and temporal attention pooling is applied to obtain the fusion coefficients. The classifier can be in the form of cosine classification: in For class prototype vectors, This is a learnable scaling factor. The difference branch and the shared branch each output logits, which are then weighted and fused to obtain the fault category and confidence score. The confidence score is the maximum category probability obtained by softmaxing the fused logits.

[0023] The structured diagnostic evidence is in tabular, key-value pair, or JSON format, and includes at least the fault category, confidence level, key channel identifier, and key channel statistical features.

[0024] Furthermore, for operational-oriented explanation, the channel contribution is obtained by aggregating from the differential attention tensor: Where H represents the number of attention heads. This represents the attention weight assigned to key channel i by query channel q under attention head h; the Top-K channels with the largest contribution are selected as the set of key sensors. The mean and standard deviation of the key sensors are calculated in the target window and reference window respectively, and together with the fault category and confidence level, they constitute structured diagnostic evidence.

[0025] Before generating the natural language diagnostic report, text fragments related to fault categories or critical channels are retrieved from the device knowledge base and input into the large language model along with structured diagnostic evidence to achieve retrieval-enhanced large language model interpretation generation.

[0026] Furthermore, structured diagnostic evidence is input into a large language model to generate a natural language diagnostic report. The diagnostic report preferably includes: explanation of the fault category, evidence of key channel anomalies, possible causes, and remedial suggestions. The structured diagnostic evidence-driven large language model interpretation generation and consistency verification process is as follows: Figure 3 As shown.

[0027] Before generating a diagnostic report, text fragments (such as maintenance manuals, fault databases, and control logic descriptions) related to the fault category or key channels can be retrieved from the equipment knowledge base and input into the large language model along with structured diagnostic evidence to achieve enhanced interpretation generation.

[0028] To improve reliability, consistency checks are performed on the output of the large language model. The consistency check includes at least: whether the sensor name in the output belongs to the key channel set of the structured diagnostic evidence, whether the key statistical values ​​in the output are consistent with the structured diagnostic evidence or fall within its allowable range, and whether the fault category in the output matches the structured diagnostic evidence.

[0029] When consistency checks fail, the optional procedures for filtering or correction are as follows: Filtering: Delete or block inconsistent content, including illegal sensor names, statistical values ​​that do not match the evidence, and fault category descriptions that do not match the evidence; Correction: Replace the above content with the corresponding fields in the structured diagnostic evidence, or rewrite the marked paragraphs by calling the large language model under evidence constraints; Review: Perform a consistency check on the revised report again; if it still fails, output a simplified report consisting only of verifiable evidence fields (including at least the fault category, confidence level, key channels and their statistical comparison and time location weights).

[0030] Example 2 Corresponding to the method described in Example 1, such as Figure 4 As shown, Embodiment 2 of the present invention discloses a multi-category industrial product defect detection system based on multi-source semantic anchor point constraint reconstruction. The system includes: The paired window differential modeling module is used to construct a window to be diagnosed and a reference window to form a paired input, and calculate differential information to characterize the deviation. The reference window is retrieved from the normal data pool based on operating condition similarity, which is based on at least one or more of environmental variables, load variables, and control variables.

[0031] The statistically guided channel selection module is used to extract statistical deviation features of each channel based on the difference sequence to form a channel importance score, and outputs the channel selection probability by the end-to-end trainable channel selection module. Differential Attention Module: Used to guide channel cross-attention calculation using channel selection probability, so that key channels are enhanced in attention matching and aggregation, thereby forming differential discriminative features; Shared representation and fusion classification module: used to extract common information about operating conditions in parallel and fuse it with differential branches to output fault categories and confidence levels; The structured evidence and large language model interpretation module is used to output key channel sets and their statistical features and attention weights to form structured diagnostic evidence, which is then input into the large language model to generate a natural language diagnostic report. It also incorporates device knowledge base retrieval enhancements and performs consistency checks and filtering / corrections on the output.

[0032] It should be noted that the specific implementation methods of each module in the system described in this embodiment two can refer to the methods described in embodiment one, and will not be repeated here.

[0033] Example 3 In one exemplary embodiment, the present invention also provides a computer device, which may be a server or a terminal, and its internal structure diagram may be as follows. Figure 5 As shown, the computer device includes a processor, memory, input / output interfaces (I / O), and a communication interface. The processor, memory, and I / O interfaces are connected via a system bus, and the communication interface is also connected to the system bus via the I / O interfaces. The processor provides computational and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and a database. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The I / O interfaces are used for exchanging information between the processor and external devices. The communication interface is used for communicating with external terminals via a network connection. When the computer program is executed by the processor, it implements the multi-category industrial product defect detection method based on multi-source semantic anchor constraint reconstruction described in Embodiment 1.

[0034] Those skilled in the art will understand that Figure 5The structures shown are merely block diagrams of some structures related to the present application and do not constitute a limitation on the computer device to which the present application is applied. Specific computer devices may include more or fewer components than shown in the figures, or combine certain components, or have different component arrangements. In an exemplary embodiment, a computer device is provided, including a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the steps in the above-described method embodiments.

[0035] In one exemplary embodiment, the present invention also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0036] In one exemplary embodiment, the present invention also provides a computer program product, including a computer program that, when executed by a processor, implements the steps in the above-described method embodiments.

[0037] It should be noted that the user information (including but not limited to user device information, user personal information, etc.) and data (including but not limited to data used for analysis, data stored, data displayed, etc.) involved in this application are all information and data authorized by the user or fully authorized by all parties, and the collection, use and processing of the relevant data must comply with relevant regulations.

[0038] Those skilled in the art will understand that all or part of the processes in the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments described above. Any references to memory, databases, or other media used in the embodiments provided in this application can include at least one of non-volatile and volatile memory. Non-volatile memory can include read-only memory (ROM), magnetic tape, floppy disk, flash memory, optical memory, high-density embedded non-volatile memory, resistive random access memory (ReRAM), magnetic random access memory (MRAM), ferroelectric random access memory (FRAM), phase change memory (PCM), graphene memory, etc. Volatile memory can include random access memory (RAM) or external cache memory, etc. By way of illustration and not limitation, RAM can take many forms, such as Static Random Access Memory (SRAM) or Dynamic Random Access Memory (DRAM).

[0039] The databases involved in the embodiments provided in this application may include at least one type of relational database and non-relational database. Non-relational databases may include, but are not limited to, blockchain-based distributed databases. The processors involved in the embodiments provided in this application may be general-purpose processors, central processing units, graphics processing units, digital signal processors, programmable logic units, data processing logic units, etc., and are not limited to these.

[0040] The technical features of the above embodiments can be combined in any way. For the sake of brevity, not all possible combinations of the technical features in the above embodiments are described. However, as long as there is no contradiction in the combination of these technical features, they should be considered to be within the scope of this specification.

[0041] This document uses specific examples to illustrate the principles and implementation methods of this application. The descriptions of the above embodiments are only for the purpose of helping to understand the methods and core ideas of this application. Furthermore, those skilled in the art will recognize that, based on the ideas of this application, there will be changes in the specific implementation methods and application scope. Therefore, the content of this specification should not be construed as a limitation of this application.

Claims

1. A fault diagnosis method based on differential attention dual-branch network and large language model interpretation, characterized in that, Includes the following steps: Acquire time series data from multiple sensors, construct a window to be diagnosed and a reference window to form a pair of inputs, and calculate the difference information between the window to be diagnosed and the reference window. Based on the differential information, the statistical deviation features of each sensor channel are extracted, a channel importance score is generated, and the channel selection probability is obtained through an end-to-end trainable channel selection module. The channel selection probability is used to guide the channel cross-attention calculation to enhance key channels and obtain differential discrimination features. Common information of working conditions is extracted to obtain shared discrimination features. The classification outputs corresponding to the differential discrimination features and the shared discrimination features are fused to obtain the fault category and confidence level. Output the set of key channels and their corresponding statistical features, and output the channel attention weights to form structured diagnostic evidence; The structured diagnostic evidence, along with the fault category and confidence level, is input into a large language model to generate and output a natural language diagnostic report.

2. The fault diagnosis method based on differential attention dual-branch network and large language model interpretation as described in claim 1, characterized in that, The method for generating window representations and residual constructions includes the following steps: The multi-sensor time series is sliced ​​according to the window length T to obtain the diagnostic window X and the reference window. : Where C represents the number of categories and T represents the window length; The reference window is retrieved from the data pool based on operating condition similarity, which is based on at least one or more of environmental variables, load variables, and control variables. After constructing paired inputs, the difference information is calculated and expressed in residual form: The reference window is used to provide a normal baseline under the same operating conditions, so that the model focuses on the deviation rather than the absolute value change.

3. The fault diagnosis method based on differential attention dual-branch network and large language model interpretation as described in claim 2, characterized in that, The statistically guided prior of the difference branch includes the following steps: The differential branch first constructs a channel anomaly prior, then uses the mean and variance information of the difference sequences to calculate statistical deviation characteristics for each channel i's difference sequence. Calculate the deviation amplitude statistic: in, This represents the standard deviation calculation. Denotes the aggregation norm with respect to the time dimension. This represents the data of the i-th channel of the window to be diagnosed. This represents the data of the i-th channel of the reference window; Introducing learnable pathways to guide scores: in , , , For learnable parameters, Let represent the learnable guided score of the i-th channel at time step t. This represents the learnable guided score of the i-th channel within the window, used to supplement statistical deviation features and guide channel priors; Finally, the prior score of the channel anomaly was obtained. : 。 4. The fault diagnosis method based on differential attention dual-branch network and large language model interpretation as described in claim 3, characterized in that, The channel selection probability is obtained using the following method: Using a continuously relaxed Top-K selection, the threshold b is solved such that: Where F For continuous gating in the form of cumulative distribution function, The smoothing coefficient controls the sharpness of continuous relaxation. Let m be the desired number of critical channels to select, m be the importance probability of the critical sensors, and b be the adaptive threshold. This represents the soft selection value for channel i. This represents the normalized channel selection probability.

5. The fault diagnosis method based on differential attention dual-branch network and large language model interpretation as described in claim 4, characterized in that: In differential channel cross-attention, channels are treated as tokens, and the target window is used as... Use the reference window as , Furthermore, a prior selection is introduced to highlight key channels, and the differential attention logit is represented as: ; Where Q represents the set of query vectors obtained by linear projection of the target window X, and K and V represent the set of query vectors obtained by linear projection of the reference window. The set of key vectors and the set of value vectors obtained by linear projection; m represents the normalized channel selection probability; Map the channel selection prior to the logit space; For temperature coefficient, is the numerical stability constant.

6. The fault diagnosis method based on differential attention dual-branch network and large language model interpretation as described in claim 5, characterized in that, The method for obtaining the fault category and confidence level includes the following steps: Shared branches are used to extract common information about operating conditions, employing channel cross-attention: in, This represents the shared branch channel attention weight matrix, where d represents the feature dimension of the query or key vector. Indicates shared branch attention output; Residual-protected fusion is performed in the differential branch to enhance the bias component and suppress common interference: in The fusion coefficient; Subsequently, temporal features are extracted through multi-branch one-dimensional convolution, and then obtained through temporal attention pooling. The classifier uses a cosine classification method. in For class prototype vectors, A learnable scaling factor; The differential branch and the shared branch output logits respectively, and then perform weighted fusion to obtain the fault category and confidence level. The confidence level is the maximum category probability obtained by softmax of the fused logits.

7. The fault diagnosis method based on differential attention dual-branch network and large language model interpretation as described in claim 6, characterized in that, The methods for forming structured diagnostic evidence include the following steps: Channel contributions are obtained by aggregating from the differential attention tensor: Where H represents the number of attention heads. This indicates the attention weight assigned to key channel i by query channel q under attention head h; the Top-K channels with the largest contribution are selected as the set of key sensors; the mean and standard deviation of the key sensors are calculated in the target window and reference window respectively, and together with the fault category and confidence level, they constitute structured diagnostic evidence; the structured diagnostic evidence is represented by tables, key-value pairs or JSON, and includes at least: fault category, confidence level, key channel identifier and key channel statistical characteristics.

8. The fault diagnosis method based on differential attention dual-branch network and large language model interpretation as described in claim 1, characterized in that, The natural language diagnostic report includes explanations of fault categories, evidence of critical channel anomalies, possible causes, and recommendations for handling them.

9. The fault diagnosis method based on differential attention dual-branch network and large language model interpretation as described in claim 1, characterized in that, When consistency verification fails, the following process is executed: Filtering: Delete or block inconsistent content, including illegal sensor names, statistical values ​​that do not match the evidence, and fault category descriptions that do not match the evidence; Correction: Replace the above content with the corresponding fields in the structured diagnostic evidence, or rewrite the marked paragraphs by calling the large language model under evidence constraints; Review: Perform a consistency check on the revised report again; if it still fails, output a simplified report consisting only of verifiable evidence fields.

10. A fault diagnosis system based on differential attention dual-branch network and large language model interpretation, characterized in that... include: The paired window differential modeling module is used to construct a window to be diagnosed and a reference window and form a paired input, and calculate differential information to characterize the deviation; wherein, the reference window is retrieved from the normal data pool according to the similarity of the operating conditions, and the similarity of the operating conditions is based on at least one or more of environmental variables, load variables, and control variables; The statistically guided channel selection module is used to extract statistical deviation features of each channel based on the difference sequence to form a channel importance score, and outputs the channel selection probability by the end-to-end trainable channel selection module. Differential Attention Module: Used to guide channel cross-attention calculation using channel selection probability, so that key channels are enhanced in attention matching and aggregation, thereby forming differential discriminative features; Shared representation and fusion classification module: used to extract common information about operating conditions in parallel and fuse it with differential branches to output fault categories and confidence levels; The structured evidence and large language model interpretation module is used to output key channel sets and their statistical features and attention weights to form structured diagnostic evidence, which is then input into the large language model to generate a natural language diagnostic report. It also incorporates device knowledge base retrieval enhancements and performs consistency checks and filtering / corrections on the output.

Citation Information

Patent Citations

  • Data attribution analysis task processing method, system and device based on large language model and storage medium

    CN120524150A

  • Industrial Internet of Things anomaly detection method based on time sequence and text joint modeling

    CN121093216A

  • Method and system for domain knowledge augmented multi-head attention based robust universal lesion detection

    US20230177678A1