Coal preparation plant full-category equipment anomaly identification method and system based on industrial large model
Patent Information
- Application Number
- CN202610902316.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-23
- Publication Date
- 2026-09-04
- Estimated Expiration
- 2046-06-23
AI Technical Summary
上述设备长期处于高粉尘、强振动、高湿度、重载荷的恶劣工况下,机械磨损与介质腐蚀耦合作用显著,极易发生各类故障
在本发明中,构建选煤厂全品类设备的深度学习知识库,并采用微调后的工业大模型将设备的融合特征向量映射至深度学习知识库的向量空间,实现设备运行特征与故障知识的空间对齐,为后续精准检索奠定了空间一致性基础;将初筛异常结果与映射对齐后的融合特征向量作为检索词,初筛异常结果作为检索的硬性过滤条件,可将检索范围限定在当前设备对应的知识子空间内,消除跨设备语义混淆,同时确保召回的知识条目与当前设备匹配,显著提升检索准确率;再由工业大模型综合这些知识与实时多模态运行数据进行深度推理,确保输出的故障类型、故障位置定位、根因分析、故障严重程度等均有知识库中的依据支撑,提高了设备异常识别的准确性,同时满足工业场景对高可靠性、实时性的要求。
Smart Images

Figure CN122471304B_ABST
Abstract
Description
Technical Field
[0001] This invention belongs to the technical field of equipment anomaly identification in coal preparation plants, and particularly relates to a method and system for anomaly identification of all types of equipment in coal preparation plants based on a large industrial model. Background Technology
[0002] The statements in this section are merely background information related to the present invention and do not necessarily constitute prior art.
[0003] Coal preparation plants are the core link in the clean and efficient utilization of coal. Their production process includes multiple stages such as conveying, crushing, sorting, and storage. Core equipment includes belt conveyors, crushers, jigs, chutes, valves, and coal bunkers. These devices operate under harsh conditions of high dust, strong vibration, high humidity, and heavy loads for extended periods, resulting in significant coupling effects of mechanical wear and media corrosion, making them highly susceptible to various malfunctions.
[0004] Currently, most anomaly identification in coal preparation plant equipment is based on individual devices. Therefore, dozens or even more independent anomaly identification models need to be maintained simultaneously, resulting in extremely high deployment and maintenance costs. Furthermore, upgrades and optimizations of any one model cannot be migrated to other devices, leading to poor generalization ability. In addition, due to the large number of devices in a coal preparation plant, if the feature vectors of multiple devices are stored and retrieved in a unified manner in a multi-device coexistence scenario, cross-device semantic confusion can easily occur due to the overlap or similarity of fault features in the numerical or semantic spaces of different devices, leading to inaccurate anomaly identification. Moreover, searching the entire knowledge base of all devices involves a huge search scope and high computational overhead, resulting in reduced real-time performance and search accuracy, which cannot meet the industrial scenario requirements of coal preparation plants with multiple devices, high real-time performance, and high reliability. Summary of the Invention
[0005] To overcome the shortcomings of the prior art, this invention provides a method and system for anomaly identification of all types of equipment in a coal preparation plant based on a large industrial model. In the scenario of all types of equipment in a coal preparation plant, it improves the accuracy of equipment anomaly identification while meeting the requirements of high reliability and real-time performance in industrial scenarios.
[0006] To achieve the above objectives, the present invention adopts the following technical solution: In a first aspect, the present invention provides a method for anomaly identification of all types of equipment in a coal preparation plant based on a large industrial model, including: Single-modal feature extraction and cross-modal fusion are performed on the multimodal operation data of each piece of equipment in the coal preparation plant to obtain the fused feature vector of each piece of equipment; By fine-tuning the industrial big model, the fused feature vectors are mapped to the vector space of the deep learning knowledge base, thereby achieving spatial alignment between operational features and fault knowledge. The fused feature vectors are input into a pre-trained anomaly classification network to identify abnormal device states, thereby obtaining preliminary anomaly screening results; wherein, the preliminary anomaly screening results include device type and suspected fault type; The initial screening results and the fused feature vector after mapping alignment are used as search terms to search in the vector knowledge base of the deep learning knowledge base to obtain knowledge entries that are highly related to the abnormal state of the equipment; wherein, the deep learning knowledge base is constructed based on all types of equipment in the coal preparation plant and corresponding multi-dimensional knowledge. Based on the retrieved knowledge entries, initial screening results of anomalies, and multimodal operational data, a fine-tuned industrial large model is used to achieve one or more of the following: accurate identification of fault types, fault location, fault severity assessment, and root cause analysis.
[0007] Secondly, this invention provides an anomaly identification system for all types of equipment in a coal preparation plant based on a large industrial model, including: The modal fusion module is configured to: extract single-modal features and fuse cross-modal features from the multimodal operating data of each piece of equipment in the coal preparation plant to obtain the fused feature vector of each piece of equipment; The spatial alignment module is configured to map the fused feature vectors to the vector space of the deep learning knowledge base through the fine-tuned industrial big model, thereby achieving spatial alignment between the running features and fault knowledge. The anomaly screening module is configured to: input the fused feature vector into a pre-trained anomaly classification network to identify the abnormal state of the device and obtain the initial screening anomaly result; wherein, the initial screening anomaly result includes the device type and the suspected fault type; The retrieval module is configured to: use the initial screening anomaly results and the fused feature vector after mapping and alignment as search terms to search in the vector knowledge base of the deep learning knowledge base to obtain knowledge entries that are highly related to the abnormal state of the equipment; wherein, the deep learning knowledge base is constructed based on all types of equipment in the coal preparation plant and corresponding multi-dimensional knowledge; The anomaly identification and diagnosis module is configured to: based on the retrieved knowledge entries, initial anomaly results, and multimodal operational data, utilize a fine-tuned industrial large model to achieve one or more of the following: accurate identification of fault types, fault location, fault severity assessment, and root cause analysis.
[0008] Thirdly, the present invention provides an electronic device including a memory and a processor, and computer instructions stored in the memory and running on the processor, wherein the computer instructions, when executed by the processor, perform the method described in the first aspect.
[0009] Fourthly, the present invention provides a computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in the first aspect.
[0010] Fifthly, the present invention provides a computer program product, including a computer program that, when executed by a processor, implements the method described in the first aspect.
[0011] The above one or more technical solutions have the following beneficial effects: In this invention, a deep learning knowledge base for all types of equipment in a coal preparation plant is constructed. A fine-tuned industrial big data model is used to map the fused feature vectors of the equipment to the vector space of the deep learning knowledge base, achieving spatial alignment between equipment operating features and fault knowledge, laying a spatial consistency foundation for subsequent accurate retrieval. The initial screening anomaly results and the mapped fused feature vectors are used as search terms, and the initial screening anomaly results are used as hard filtering conditions for retrieval. This limits the search scope to the knowledge subspace corresponding to the current equipment, eliminating cross-equipment semantic confusion, while ensuring that the recalled knowledge items match the current equipment, significantly improving retrieval accuracy. Then, the industrial big data model integrates this knowledge with real-time multimodal operating data for deep reasoning, ensuring that the output fault type, fault location, root cause analysis, fault severity, etc., are all supported by the knowledge base, improving the accuracy of equipment anomaly identification, while meeting the requirements of high reliability and real-time performance in industrial scenarios.
[0012] Advantages of additional aspects of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description
[0013] The accompanying drawings, which form part of this invention, are used to provide a further understanding of the invention. The illustrative embodiments of the invention and their descriptions are used to explain the invention and do not constitute an improper limitation of the invention.
[0014] Figure 1 This is an overall flowchart of the anomaly identification method for all types of equipment in a coal preparation plant based on a large industrial model, as described in Embodiment 1 of the present invention. Detailed Implementation
[0015] It should be noted that the following detailed descriptions are exemplary and intended to provide further illustration of the invention. Unless otherwise specified, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this invention pertains.
[0016] It should be noted that the terminology used herein is for the purpose of describing particular implementations only and is not intended to limit the exemplary implementations of the present invention.
[0017] Where there is no conflict, the embodiments and features in the embodiments of the present invention can be combined with each other.
[0018] Example 1 like Figure 1 As shown, this embodiment discloses a method for anomaly identification of all types of equipment in a coal preparation plant based on a large industrial model, including: Single-modal feature extraction and cross-modal fusion are performed on the multimodal operation data of each piece of equipment in the coal preparation plant to obtain the fused feature vector of each piece of equipment; By fine-tuning the industrial big model, the fused feature vectors are mapped to the vector space of the deep learning knowledge base, thereby achieving spatial alignment between operational features and fault knowledge. The fused feature vectors are input into a pre-trained anomaly classification network to identify abnormal device states, resulting in preliminary anomaly screening results. These results include device type and suspected fault type. The initial screening results and the fused feature vector after mapping and alignment are used as search terms to search in the vector knowledge base of the deep learning knowledge base to obtain knowledge entries that are highly related to the abnormal state of the equipment; the deep learning knowledge base is constructed based on the full range of equipment in the coal preparation plant and the corresponding multi-dimensional knowledge. Based on the retrieved knowledge entries, initial screening results of anomalies, and multimodal operational data, a fine-tuned industrial large model is used to achieve one or more of the following: accurate identification of fault types, fault location, fault severity assessment, and root cause analysis.
[0019] In this embodiment, a deep learning knowledge base for all types of equipment in a coal preparation plant is constructed. A fine-tuned industrial big data model is used to map the fused feature vectors of the equipment to the vector space of the deep learning knowledge base, achieving spatial alignment between equipment operating features and fault knowledge, laying a spatial consistency foundation for subsequent accurate retrieval. The initial screening anomaly results and the mapped fused feature vectors are used as search terms, and the initial screening anomaly results are used as hard filtering conditions for retrieval. This limits the search scope to the knowledge subspace corresponding to the current equipment, eliminates cross-equipment semantic confusion, reduces the computational load of retrieval, and ensures that the recalled knowledge items match the current equipment, significantly improving retrieval accuracy. Then, the industrial big data model integrates this knowledge with real-time multimodal operating data for deep reasoning, ensuring that the output fault type, fault location, root cause analysis, fault severity, etc., are all supported by the knowledge base, meeting the requirements of high reliability and real-time performance in industrial scenarios.
[0020] The following is a detailed description of the anomaly identification method for all types of equipment in a coal preparation plant based on a large industrial model, which is involved in this embodiment: S1: Construct a deep learning knowledge base based on all types of equipment in a coal preparation plant and their corresponding multi-dimensional knowledge.
[0021] This embodiment constructs a five-dimensional knowledge system covering all categories of equipment in coal preparation plants, including equipment body parameter library, fault mechanism library, sensor feature library, fault case library, and expert handling experience library.
[0022] The equipment body parameter library includes the design parameters, structural drawings, operating procedures, and maintenance standard data for each piece of equipment.
[0023] The fault mechanism database includes the occurrence mechanism, evolution law, and characteristic data of faults corresponding to various equipment; for example, for belt conveyors, it includes fault mechanisms such as belt misalignment, tearing, idler roller overheating, and abnormal noise; for crushers, it includes fault mechanisms such as bearing overheating and jamming; for jigs, it includes fault mechanisms such as air valve failure and bed abnormality; for chutes, it includes fault mechanisms such as blockage and coal leakage; for valves, it includes fault mechanisms such as abnormal status and internal leakage; and for coal bunkers, it includes fault mechanisms such as abnormal material level and bunker wall detachment.
[0024] The sensor feature library includes feature thresholds, temporal patterns, and spectral characteristics of multi-dimensional data such as vibration, temperature, sound, vision, and electrical quantities corresponding to various fault types, as well as the feature threshold range under normal operating conditions.
[0025] The fault case database contains full-process data of historical fault events in coal preparation plants, including pre-fault operating data, fault phenomena, handling process, and maintenance results.
[0026] The expert experience database includes fault handling procedures, emergency measures, maintenance standards, and prevention plans from equipment operation and maintenance experts in the coal preparation industry.
[0027] This embodiment is designed for the specific operation and maintenance scenarios of all types of equipment in a coal preparation plant. It defines exclusive entities and relationships, which is different from the general industrial broad entity relationships. It is precisely adapted to the fault knowledge system of six core equipment: belt conveyors, crushers, jigs, chutes, valves, and coal bunkers.
[0028] In this embodiment, knowledge from the five-dimensional knowledge system is extracted into entities, relationships, and semantics through an industrial big data model to generate standardized knowledge units. Then, all standardized knowledge units are converted into knowledge vectors through a word embedding model to form a unified vector set. The knowledge base type to which each piece of knowledge belongs is identified by metadata tags.
[0029] As a specific implementation method, based on the coal preparation production process and equipment failure characteristics, five major categories of exclusive entities are defined: equipment entities, failure entities, failure characteristic entities, operating condition case entities, and operation and maintenance handling entities. This fully covers all the content of the five-dimensional knowledge system in this embodiment, with no general redundant entities.
[0030] The specific equipment entities include belt conveyors, crushers, jigs, feed chutes, pneumatic valves, raw coal bunkers, and product bunkers. Each entity is bound to unique equipment IDs, installation locations, process levels, rated operating parameters, structural attributes, and other inherent parameters, corresponding to the "equipment entity parameter library" of the five-dimensional knowledge system.
[0031] The fault entities are unique fault tags for each piece of equipment, specifically including belt conveyor misalignment, belt longitudinal tear, idler roller overheating, idler roller abnormal noise, crusher bearing overheating, crusher rotor jamming, jig air valve failure, jig bed abnormality, chute blockage, chute coal leakage, valve internal leakage, valve abnormality, coal bunker material level abnormality, and coal bunker wall detachment, corresponding to the "fault mechanism library" of the five-dimensional knowledge system.
[0032] The fault feature entities include vibration spectrum features, temperature change trends, equipment acoustic features, equipment image texture features, motor current fluctuation features, pressure difference features, material level change features, and stress anomaly features, corresponding to the "sensor feature library" of the five-dimensional knowledge system.
[0033] The case studies include normal operating conditions, early-stage deterioration conditions, moderate failure conditions, severe failure conditions, and historical failure events in coal preparation plants, along with their timing, operating conditions, environment, and preconditions for handling, forming a "failure case library" corresponding to a five-dimensional knowledge system.
[0034] The operation and maintenance entity includes troubleshooting steps, emergency response plans, maintenance standards, spare parts compatibility types, inspection cycles, and preventive maintenance strategies, corresponding to an "expert handling experience base" of five-dimensional knowledge system.
[0035] For the aforementioned coal preparation-specific entities, fixed and computable scenario-based relationships are constructed to achieve logical connection of fault knowledge, distinguishing it from the undifferentiated relationships in general industry. Specifically, this includes: (1) Equipment-fault relationship: Equipment [occurs / can trigger] fault. For example: belt conveyor - can occur - idler roller overheating fault, jig - can occur - bed abnormality fault.
[0036] (2) Fault-feature relationship: Fault [corresponds to / triggers] exclusive abnormal features. For example: Roller overheating fault - corresponds to - sudden rise in bearing temperature, slight vibration abnormal features; Chute blockage fault - corresponds to - material accumulation, high frequency vibration features of the chute.
[0037] (3) Fault-mechanism relationship: Faults follow a specific evolution mechanism. For example, bearing wear failure follows the mechanism of long-term frictional heating and gradual deterioration failure.
[0038] (4) Fault-Operating Condition Relationship: The fault occurs under specific production conditions. Example: Abnormal coal bunker level occurs under conditions of fluctuating raw coal feed and poor unloading. (5) Fault-Resolution Relationship: Fault [Matching] Dedicated Operation and Maintenance Resolution Plan. Example: Valve internal leakage fault - Matching - Valve body seal replacement, pressure adjustment and handling process; (6) Fault-Level Relationship: Fault [Corresponds to] Preset Warning Level. Example: Longitudinal tear of belt - corresponds to Level 1 Emergency Warning; Slight overheating of idler roller - corresponds to Level 3 General Warning.
[0039] This embodiment constructs a deep learning knowledge base covering all categories of core equipment in coal preparation plants, including belt conveyors, crushers, jigs, chutes, valves, and coal bunkers. Through a large industrial model, it standardizes and vectorizes fault mechanisms, expert experience, and case data, establishing a unified knowledge base for fault diagnosis of coal preparation plant equipment. This solves the problems of model fragmentation, poor generalization, and difficulty in cross-equipment migration in existing technologies, and significantly reduces the cost of model deployment and maintenance.
[0040] The original five-dimensional knowledge system contains a large amount of unstructured text, scattered parameters, experience descriptions, case records, and other raw data, which cannot be directly used for large-scale model retrieval and reasoning. This embodiment addresses the five-dimensional knowledge system, which includes equipment ontology parameter library, fault mechanism library, sensor feature library, fault case library, and expert handling experience library. Through entity and relation extraction and semantic normalization, all the original five-dimensional knowledge systems are uniformly decomposed, cleaned, and reconstructed into standardized knowledge units. Each knowledge unit is formed by the matching logical relationship between five entities: equipment, fault, feature, operating condition, and handling. Each knowledge unit contains only a single entity relationship, i.e., a simple logical relationship between two types of entities. It is the smallest basic unit of the knowledge base that can be quantified, vectorized, and retrieved, realizing the transformation of messy raw knowledge into structured and computable knowledge.
[0041] This embodiment uses a pre-trained word embedding model to convert standardized knowledge units into knowledge vectors, constructs a vector knowledge base, and establishes index mapping relationships.
[0042] In this embodiment, the word embedding model is designed based on the Transformer encoder architecture, and is fully adaptable to fault diagnosis scenarios for all types of equipment in coal preparation plants. Specifically, the word embedding model uses a lightweight Transformer encoder as its core architecture and is specifically adapted to the semantic characteristics of the coal preparation field.
[0043] The word embedding model employs a 6-layer Transformer encoder as its core architecture. Each layer contains a multi-head self-attention standard force module (MHSA) with fault entity constraints and a feedforward neural network module (FFN). The embedding dimension is... Setting it to 768 ensures complete alignment with the vector space dimension of the subsequent large industrial model, achieving seamless adaptation of the semantic space.
[0044] The word embedding model processes the following steps: (1) Input layer.
[0045] The input layer takes as its input a text sequence of standardized knowledge units for the final product. , This represents the nth word. The sequence length is set to a maximum of 512; this is achieved using a specialized thesaurus for coal preparation. Map each token to an initial word embedding vector. Simultaneously, location embedding for capturing timing information is added. Segment embeddings that distinguish knowledge types The input embedding vector is obtained. for:
[0046] in, ; For embedded dimensions, The sequence length is given.
[0047] Initial word embedding vector The acquisition specifically involves: based on the constructed 35,000-word thesaurus specifically for the coal preparation field. The final standardized knowledge units after corpus supplementation and improvement are tokenized, dividing the continuous text into individual tokens to obtain a discrete token sequence. The word embedding model has a built-in embedding lookup table, in which each word token corresponds to a set of fixed-dimensional floating-point vectors. The embedding layer weights are initially loaded with general pre-trained word embedding parameters. Each token after word segmentation is traversed, and the corresponding vector is read out by index in the embedding lookup table. The vectors corresponding to all tokens are concatenated to obtain the initial word embedding vector sequence of the entire text. .
[0048] Position embedding The learnable positional encoding parameters are built into the word embedding model. Based on the maximum sequence length of 512 dimensions of the word embedding model, 512 sets of 768-dimensional positional vectors are predefined. The training process is updated iteratively with the backpropagation of the word embedding model. No manual formula design is required. It is used to represent the character order and contextual positional relationship of knowledge text, making up for the lack of temporal awareness of the Transformer encoder. This embodiment only performs dimensional adaptation and alignment, without structural changes.
[0049] Segment embedding vector Unlike the binary segment embedding design of the general model, this embodiment defines four categories of coal preparation knowledge segment labels: equipment parameters, fault mechanisms, operation and maintenance cases, and expert experience. It also predefines four sets of exclusive 768-dimensional learnable segment embedding parameters. When a standardized knowledge unit is input, the corresponding segment embedding vector is automatically matched based on its knowledge type. Furthermore, through continuous optimization during dual-task training, differentiated semantic encoding of coal preparation fault knowledge across different dimensions is achieved, addressing the shortcomings of general model knowledge semantic mixing and low discriminative power.
[0050] (2) Transformer encoder core layer.
[0051] To address the strong entity association characteristics of coal preparation fault knowledge, a fault entity attention bias is added to the self-attention mechanism. For core associated entities such as equipment entities, fault types, and characteristic parameters within the same standardized knowledge unit, the attention weight is proactively increased. The modified self-attention calculation process is as follows: For input embedding vectors Generate the query matrix through linear transformation Key matrix Value matrix :
[0052] in, , , In this embodiment, the number of attention heads is set to 12; For the embedded dimension.
[0053] Introducing the attention bias matrix for faulty entities : for Matrix, when the first The token and the first When a token belongs to a fault-associated entity pair (such as "belt conveyor - misalignment", "idler - overheating", "crusher - bearing jamming") , are trainable parameters, otherwise The revised formula for calculating attention score is:
[0054] The output of multi-head self-attention is a concatenation of the attention results from each head: ,in, , To output the linear transformation matrix, For the embedded dimension, the superscript T indicates transpose.
[0055] After multi-head self-attention output, the data is processed through layer normalization (LN) and a feedforward neural network (FFN). FFN employs... GELU The activation function is calculated using the following formula:
[0056] in, For the first linear layer, For the second linear layer, From the dimension Mapped to 3072, Map back ; For activation functions; This represents the output characteristics of multi-head self-attention; This indicates a normalization layer, and FFN stands for feedforward neural network.
[0057] (3) Output layer.
[0058] After encoding by a 6-layer Transformer encoder, the vector corresponding to the label in the output sequence is taken as the sentence embedding vector of the entire input standardized knowledge unit, which is the final knowledge vector. This enables end-to-end conversion of standardized knowledge units into fixed-dimensional vectors.
[0059] As one specific implementation method, a thesaurus specific to the coal preparation field The construction specifically includes: To address the issue of out-of-vocabulary (OOV) terms in coal preparation terminology within a general Chinese vocabulary, this paper expands the vocabulary by mining coal preparation-specific terms using a combination of mutual information (MI) and left / right entropy (LE) algorithms, building upon the BERT-based general Chinese vocabulary. (1) Mutual Information (MI): Measures the degree of binding between two characters, and is used to filter candidate strings of technical terms. The calculation formula is as follows:
[0060] in, For characters With characters The probability of consecutive co-occurrence in the corpus , characters respectively , The probability of a single occurrence in the corpus.
[0061] (2) Right entropy (LE / RE): measures the boundary degrees of freedom of candidate strings and verifies the independence of terms. The calculation formula is:
[0062] in, Candidate words left entropy, Candidate words The right entropy; Candidate words Characters appear on the left The conditional probability, Candidate words Characters appear on the right The conditional probability.
[0063] Filtering mutual information , The strings were used as specialized terms specific to coal preparation and added to the basic vocabulary, ultimately constructing a specialized vocabulary of 35,000 words for the coal preparation field. This ensures the semantic accuracy of embedded technical terms from the source.
[0064] As a specific implementation method, the word embedding model adopts a two-stage training strategy of general pre-training initialization and coal preparation-specific fine-tuning, which not only retains the general semantic understanding ability, but also deeply adapts to the specific semantic scenario of coal preparation equipment fault diagnosis.
[0065] Using the generated standardized knowledge units and their corresponding coal preparation-specific entities and relationships as the core anchors and structured benchmarks, authoritative industry data is supplemented in a targeted manner. All supplementary data is strictly aligned with the five major entities of equipment, faults, characteristics, operating conditions, and maintenance, as well as various relationships, to prevent the inclusion of general and invalid data. The data is supplemented with coal preparation industry standards, coal sorting equipment design manuals, coal preparation plant operation and maintenance specifications, core academic literature in the field, and fault case studies from the same industry, forming a coal preparation equipment-specific corpus. The corpus covers all categories of equipment, including belt conveyors, crushers, jigs, chutes, valves, and coal bunkers, providing comprehensive knowledge on parameters, faults, characteristics, cases, and handling experience.
[0066] Standardized preprocessing was performed on the corpus: deduplication, noise reduction, special symbol filtering, and sentence and word segmentation were completed. Based on the entity extraction results, forced word segmentation was performed on coal preparation-specific entities (such as belt conveyor misalignment, jig air valve failure, idler roller overheating, and bed abnormality) to ensure the semantic integrity of professional terms, and finally, a standardized training corpus was generated. .
[0067] The pre-training initialization phase specifically involves: using the weights of a word embedding model pre-trained on a large-scale general Chinese corpus as the initial weights for the word embedding model, thus completing the basic construction of the general semantic understanding capability of the word embedding model and significantly reducing the computational cost and overfitting risk of training from scratch; based on the constructed coal preparation-specific corpus... A dual-task joint training strategy is adopted, with masked language model (MLM) as the main task and coal preparation-specific fault entity matching (FEM) as the auxiliary task, to enhance the semantic representation ability of word embedding model for coal preparation fault knowledge.
[0068] The training dataset is constructed by using a standardized corpus. The text in the dataset is segmented according to a maximum sequence length of 512 to generate a training sample set. For each sample, a random mask token is created with a 15% probability. 80% of the tokens are replaced with the [MASK] tag, 10% are replaced with random words, and 10% retain the original words to generate MLM task training samples. At the same time, based on the entity extraction results, each sample is labeled with associated entity pairs of device-fault, fault-feature, and feature-treatment to generate FEM task training samples.
[0069] For the coal preparation-specific fault entity matching as an auxiliary task, this embodiment adopts a scenario-adaptive differentiated annotation strategy, which does not require mandatory annotation of entity pairs for all samples. This approach aligns with the real distribution of corpora in the coal preparation field, balancing training effectiveness and engineering feasibility. The specific annotation rules are as follows: Fully labeled samples: literature on fault mechanisms, real-world fault cases in the industry, literature on equipment fault feature analysis, and fault repair and maintenance records. These samples naturally contain strong correlations between equipment-fault, fault-feature, and fault-handling entities. Each sample is accurately labeled with positive and negative entity pairs, and all samples participate in coal preparation-specific fault entity matching and comparison learning training. Some labeled samples: coal preparation plant operation and maintenance specifications, inspection operation standards, and equipment maintenance procedures. These samples contain weak correlations between faults and operation and maintenance, and between equipment and maintenance strategies. Valid paragraph labeled entity pairs containing fault scenarios are selected and used to participate in coal preparation-specific fault entity matching auxiliary training. Unlabeled samples: Basic equipment parameters, industry standard general clauses, equipment structure descriptions, and pure process parameter descriptions. There are no fault-related entity relationships. They do not label coal preparation-specific fault entity matching entity pairs, do not participate in coal preparation-specific fault entity matching task training, and only participate in mask language model mask pre-training to learn the word order of professional terms and general semantic rules in the coal preparation field.
[0070] Based on the aforementioned differentiated annotation rules and relying on the entity extraction results, multi-dimensional associated entity pair annotation of annotable samples is completed, and a training sample set for a coal preparation-specific fault entity matching task is constructed. This not only avoids training noise caused by invalid annotations, but also accurately strengthens the word embedding model's ability to associate and represent the core knowledge of coal preparation faults.
[0071] Loss function for training word embedding models Main task loss Loss of auxiliary tasks A weighted sum that balances general semantic learning with domain knowledge adaptation:
[0072] in, The weighting factor is set to 0.4 in this embodiment.
[0073] Main task loss The cross-entropy loss function is used to calculate the error between the predicted and true values of the masked token, thereby optimizing the general semantic understanding capability of the word embedding model.
[0074] in, The number of mask tokens in the sample. The actual value of the mask token. The input sequence after masking. This represents the predicted probability of the masked token by the word embedding model.
[0075] Auxiliary task loss A contrastive learning loss function is employed to enhance the semantic similarity of fault-related entity pairs and reduce the semantic similarity of non-related entities, thereby optimizing the model's ability to represent the association of coal preparation fault knowledge.
[0076] in, This represents the number of positive sample pairs. These are two entities in a positive sample pair; Pair the fault-related entities with the sample set (e.g., "jigging machine - air valve failure" "chute - blockage"). These are two entities in a negative sample pair; For the negative sample set of non-related entities; entity embedding vector With entity embedding vector Cosine similarity; entity embedding vector With entity embedding vector Cosine similarity; The temperature coefficient is set to 0.07 in this embodiment.
[0077] The AdamW optimizer was used in the word embedding model training. The initial learning rate was set to 2e-5, the weight decay coefficient was set to 1e-2, the batch size was set to 32, and the training epochs were set to 10. A linear learning rate decay strategy was adopted, and the warm-up steps were set to 10% of the total steps. During the training process, when the validation set loss no longer decreased for two consecutive epochs, the training was terminated early to save the optimal model weights.
[0078] A validation set containing 1,000 pairs of coal preparation professional terms was constructed. Semantic similarity labels were manually annotated by coal preparation industry experts. The trained word embedding model had a Pearson correlation coefficient of ≥0.92 on the validation set, ensuring the semantic accuracy and domain adaptability of the word embedding.
[0079] The trained word embedding model is primarily used to achieve the vectorization transformation of standardized knowledge units, the construction and index mapping of vector knowledge bases, and to provide a unified vector space for the entire process of feature semantic enhancement and fault retrieval matching. The specific process is as follows: (1) Vectorization of standardized knowledge units.
[0080] For each standardized knowledge unit generated (equipment parameter entry, fault mechanism entry, fault case entry, expert experience entry, etc.), end-to-end vectorization processing is performed, including: Text preprocessing: Standardized knowledge unit texts are converted into model input format, with [CLS] tags added at the beginning and [SEP] tags added at the end; for long texts exceeding 512 characters, a sliding window is used to divide them into multiple subsequences to ensure semantic integrity.
[0081] Vector generation: The preprocessed text sequence is input into the trained word embedding model, and the vector corresponding to the [CLS] label of the output sequence is used as the core knowledge vector of the standardized knowledge unit; for long texts segmented by multiple windows, the average of the vectors of each sub-sequence is taken as the final knowledge vector. .
[0082] Vector normalization: L2 normalization is performed on the generated knowledge vectors to eliminate the interference of vector magnitude on similarity calculation. The normalization formula is:
[0083] in, Describing the L2 norm, This is the result of normalizing the knowledge vector.
[0084] (2) Vector knowledge base construction and index mapping.
[0085] Vector knowledge base construction: The normalized knowledge vectors corresponding to all standardized knowledge units are bound one by one with the metadata such as the original text of the standardized knowledge unit, knowledge type, associated equipment, and fault type, to build a deep learning vector knowledge base for all types of equipment in the coal preparation plant. Each data item is in the format of: <knowledge ID, normalized knowledge vector, original knowledge text, metadata tag>.
[0086] Efficient retrieval index construction: Using the FAISS (Facebook AI Similarity Search) library, an IVF_FLAT inverted index is built for the vector knowledge base. The number of cluster centers is set to 1024 for 768-dimensional vectors, enabling millisecond-level similarity retrieval for millions of vectors.
[0087] Bidirectional mapping relationship establishment: Establish a bidirectional mapping relationship between knowledge ID, vector index, and original knowledge unit to ensure that the retrieved vector can be quickly traced back to the corresponding original knowledge content, providing traceable knowledge support for subsequent reasoning of large industrial models.
[0088] The trained word embedding model provides a unified semantic vector space for the entire process, supports the multimodal feature semantic enhancement of S3, maps the fused feature vectors of device operation to a unified semantic space, and achieves spatial alignment between operational features and fault knowledge; it also provides core support for anomaly identification and knowledge base retrieval, converting the initial screening results of device anomalies into retrieval vectors, and performing matching retrieval in the vector knowledge base using cosine similarity. The similarity calculation formula is as follows:
[0089] in, This is an abnormal retrieval vector. For knowledge vectors in the knowledge base, The closer the value is to 1, the higher the semantic similarity. This represents the L2 norm.
[0090] All knowledge content in the five-dimensional knowledge system is 100% converted into standardized knowledge units, with no omissions or redundancies. Subsequent word embedding vectorization, knowledge base retrieval, and large model reasoning all use "standardized knowledge units" as the core carrier to achieve deep adaptation between the five-dimensional knowledge base and model reasoning, solving the defects of existing technologies such as messy industry knowledge bases and inability to accurately link with AI models.
[0091] The five-dimensional knowledge system's top-level classification and storage system is the macro-level knowledge container, while the standardized knowledge unit is the smallest standardized atomic particle after the knowledge base is broken down, i.e., the micro-level basic data unit. The two have a hierarchical relationship of "whole-part, container-content".
[0092] S2: Extract single-modal features and fuse cross-modal features from the multimodal operating data of each piece of equipment in the coal preparation plant to obtain the fused feature vector of each piece of equipment.
[0093] In this embodiment, for six major categories of equipment, explosion-proof multi-source heterogeneous sensors and mobile inspection robots are deployed to collect multimodal data. The mobile inspection robot is equipped with a dual-spectrum gimbal camera and an acoustic fingerprint sensor, which can collect visible light and infrared images in real time and has autonomous navigation, obstacle avoidance and fixed-point inspection capabilities. It can monitor the temperature data of roller bearings, frequency converters, motors, etc. in real time, and the acoustic fingerprint sensor monitors noise-sensitive areas such as jigs and flotation machines. All sensor data are filtered, normalized and feature extracted by edge computing nodes and then uniformly connected to the industrial Internet of Things platform.
[0094] For belt conveyors: deploy vibration sensors, infrared temperature sensors, sound sensors, visual inspection cameras, misalignment switches, and tear sensors to collect roller vibration signals, bearing temperature, belt running sound patterns, belt surface images, and misalignment / tear switch data.
[0095] For example, one set of vibration sensors and infrared temperature sensors are deployed every 100 meters, and one sound sensor is deployed every 50 meters. Visual inspection cameras, belt misalignment switches, and longitudinal tear sensors are deployed at the head and tail of the machine to collect data on roller vibration, bearing temperature, belt running sound, belt surface image, and belt misalignment / tear switch quantity.
[0096] For crushers: deploy vibration sensors, temperature sensors, current sensors, and sound sensors to collect bearing vibration signals, bearing temperature, motor operating current, and equipment operating sound data.
[0097] For example, one set of vibration sensor and one set of temperature sensor are deployed in each of the front and rear bearing housings, and a current sensor and a sound sensor are deployed on the motor side to collect bearing vibration, temperature, motor current and equipment operation sound data.
[0098] For jigs: deploy pressure sensors, displacement sensors, bed density sensors, valve actuator feedback sensors, and current sensors to collect data on valve inlet and outlet pressures, valve stroke, bed thickness and density, valve actuator feedback signals, and valve motor operating current.
[0099] For example, pressure sensors are deployed at the inlet and outlet of each air valve, displacement sensors and feedback sensors are deployed on the air valve actuators, bed density sensors and radar level gauges are deployed in the jigging chamber, and current sensors are deployed on the air valve motors to collect data on air valve pressure, stroke, actuator feedback, bed density and thickness, and motor current.
[0100] For chutes: Deploy radar level sensors, vibration sensors, visual inspection cameras, and pressure sensors to collect data on material level, chute vibration signals, internal images of the chute, and material impact pressure.
[0101] For example, radar level sensors are deployed at the inlet and outlet, vibration sensors and visual inspection cameras are deployed in the tank, and pressure sensors are deployed at the bottom to collect data on material level, tank vibration, internal images, and material impact pressure.
[0102] For valves: Deploy position sensors, inlet and outlet differential pressure sensors, sound sensors, and torque sensors to collect valve opening data, pressure difference between the front and rear ends of the valve, fluid acoustic signature, and actuator torque data.
[0103] For example, position sensors and torque sensors are deployed on the valve body, pressure sensors are deployed at the inlet and outlet, and sound sensors are deployed next to the valve body to collect data on valve opening, actuator torque, differential pressure between the front and rear ends, and fluid acoustic signatures.
[0104] For coal bunkers: Deploy radar level sensors, bunker wall stress sensors, vibration sensors, and visual inspection cameras to collect data on material level, bunker wall stress, bunker wall vibration signals, and internal image data of the coal bunker.
[0105] For example, a radar level sensor and a visual inspection camera are deployed on the top, and stress sensors and vibration sensors are deployed on the silo walls to collect data on coal silo level, silo wall stress, vibration, and images inside the silo.
[0106] The collected multi-source heterogeneous data is subjected to time-series alignment, noise filtering, missing value completion, and normalization. At the same time, the data is screened for outliers and assessed for data quality through a fine-tuned industrial large model to remove invalid and interfering data, and generate standardized time-series datasets and image datasets.
[0107] As a specific implementation method, the core requirement of time-series data is to capture the time evolution pattern and correlation characteristics of equipment operating parameters. It is suitable for scenarios with strong time-series faults in coal preparation plants (such as the gradual process of bearing wear from slight to severe) and long sequence data processing. It uses an improved Transformer network or an improved bidirectional LSTM network for processing, and can adaptively switch according to the data length. For example, the improved bidirectional LSTM network is selected for short sequences ≤1000 frames, and the improved Transformer network is selected for long sequences >1000 frames.
[0108] First, the preprocessed time series data is standardized to map the data to the [0,1] interval to eliminate dimensional differences. Then, the standardized time series data is divided into frame sequences of fixed length (frame length is set to 64, step size is set to 32) and input into the corresponding improved Transformer network or improved bidirectional LSTM network. The network captures the short-term correlation and long-term dependency of the time series data through feature encoding and outputs a high-dimensional time series correlation feature vector (dimension is set to 256) for subsequent cross-modal fusion.
[0109] To address the shortcomings of traditional bidirectional LSTM networks, such as gradient vanishing and high computational cost during long sequence training, the improvements to bidirectional LSTM networks are as follows: a. Add a Layer Normalization module to normalize the input of each layer, using the following formula:
[0110] in, For the input sequence The mean, For the input sequence variance To prevent tiny values with a denominator of 0 (take 1e-5). , These are learnable parameters, which effectively solve the problems of gradient vanishing and training instability in long sequence training. Representation layer normalization.
[0111] b. Introduce a residual connection to superimpose the residuals of the bidirectional LSTM network input and output, as shown in the following formula:
[0112] in, For bidirectional LSTM network input features, The output features of the bidirectional LSTM network are reduced by residual connections to reduce information loss during the transmission of long sequence features, thereby enhancing the ability to capture fault temporal evolution features. Representation layer normalization.
[0113] c. Simplify the number of neurons in the hidden layer of the network, reducing the number of neurons in the hidden layer from 256 to 128, while removing redundant fully connected layers. Without losing core features, the number of network parameters is reduced by more than 40%, adapting to the low computing power constraints of edge computing nodes.
[0114] To address the shortcomings of traditional Transformer networks, such as high computational complexity of global attention and unsuitability for edge computing, the following improvements to the Transformer network are made: a. Replace global attention with local window attention. Divide the time series into several local windows (window size set to 16), and calculate the attention weight only within each local window. The calculation formula is as follows:
[0115] in, 、 、 These are query, key, and value matrices, respectively. For the key matrix dimension, local window attention reduces computational complexity from... Down to This significantly reduces computing power consumption, among which The sequence length of the local window. The window size is [size].
[0116] b. Adaptive temporal location coding is introduced, combining the periodic characteristics of coal preparation plant time-series data (such as equipment operating cycles and fault evolution cycles). The coding formula is as follows:
[0117]
[0118] in, For timing position; For dimension indexing; The embedding dimension is set to 256 in this embodiment. It is a periodic adaptive coefficient that is dynamically adjusted according to the equipment operating cycle, with a value range of 0.01 to 0.1, which enhances the network's accuracy in extracting the temporal evolution characteristics of equipment faults.
[0119] The core requirement for image-based data is to suppress irrelevant interference such as dust and light and shadow changes in coal preparation plants, and to accurately capture the visual features of equipment failure areas (such as belt tearing, valve wear, and bin wall detachment). This embodiment adopts an improved CNN network based on a lightweight MobileNet architecture to extract the visual features of image-based data.
[0120] Specifically, the preprocessed image data is first normalized in size (uniformly adjusted to 224×224 pixels) and its RGB channels are standardized. Then, it is input into an improved CNN network, passing through a feature extraction layer, a feature enhancement layer, and a feature pooling layer in sequence. Visual features such as texture, contour, and fault areas of the image are extracted through convolution operations, activation functions, and pooling operations. Finally, a 256-dimensional visual feature vector is output through a global average pooling layer to ensure the globality and representativeness of the features.
[0121] To address the shortcomings of traditional CNNs, such as weak anti-interference capability, large parameter count, and inaccurate feature extraction of fault regions, this embodiment improves the CNN network as follows: a. The two core computational steps of MobileNet's native depthwise separable convolution, namely group convolution and dilated convolution, are internally fused. The characteristics of group convolution and dilated convolution are directly embedded into the original process of depthwise separable convolution. Group convolution reduces the amount of computation, while the dilated convolution expands the feature receptive field without increasing the number of convolution kernel parameters by utilizing the dilated convolution's dilated interval sampling characteristic. The complete calculation logic of the convolution kernel is as follows: Let the original convolution kernel size be... The convolution dilation rate is (This plan is adopted) The number of input feature channels is Number of groups .
[0122] First, dilated grouped depthwise convolutions are performed, and dilated sampling convolutions are applied to each group of features. This abandons the traditional dense sampling method and expands the feature perception range without adding new parameters through dilated convolutions with interval sampling. The calculation formula is as follows:
[0123] in: The input feature map is a single-path feature map after grouping; These are the depthwise convolutional kernel weights; The convolution dilation rate is used to control the sampling interval of the convolution kernel. The receptive field is expanded by sampling at intervals, and the number of parameters of the convolution kernel is the same as that of the ordinary convolution kernel, with no additional parameter overhead. This is the original kernel size; Represents spatial location coordinates, This represents the offset within the convolution kernel.
[0124] Subsequently, grouped point convolution is performed on the output of the grouped depthwise convolution to break the isolation of single-channel features and complete cross-channel feature fusion. At the same time, the grouping mechanism reduces the computational cost of the channel fusion stage. The calculation formula is as follows:
[0125] In the formula: for The convolution kernel is set to 4, and the number of groups is fixed. By constraining the channel operation dimension through grouped convolution, the computational cost is further reduced and the number of model parameters is decreased. Number of groups; Output the grouped depthwise convolution; This represents point convolution.
[0126] To quantify the receptive field expansion effect of dilated convolution, an equivalent receptive field calculation formula for dilated convolution is introduced to accurately characterize the feature perception range of dilated convolution:
[0127] In the formula: The equivalent receptive field size for dilated convolution. This is the original kernel size. This refers to the dilation rate. Compared to ordinary convolution, dilated convolution, by adjusting the sampling interval, can effectively expand the receptive field of features and capture a wider range of image feature information while keeping the kernel size constant.
[0128] This composite convolutional structure integrates the core advantages of grouped convolution, depthwise separable convolution, and dilated convolution. It simplifies convolution operations and compresses model parameters by relying on grouped convolution and depthwise separable convolution, further reducing the number of network parameters by 30%. At the same time, it effectively expands the feature receptive field by leveraging the receptive field expansion characteristic of dilated convolution without parameter increment, fully extracting global and local correlation features of the image, and significantly improving the model's ability to extract and identify detailed features of fault areas.
[0129] b. Introduce a channel attention module into the feature enhancement layer to adaptively adjust the feature weights of different channels, strengthen fault-related channel features, and suppress irrelevant interference channel features. The calculation formula is as follows: Step 1: Global Average Pooling
[0130] Step 2: Compression and activation of the fully connected layer:
[0131] Third step: Feature weighting:
[0132] Where Z represents the result after global average pooling. The two-dimensional feature map output by the MobileNet improved convolutional structure is in coordinates The characteristic values at that location, This is a two-dimensional feature map output by the MobileNet improved convolutional structure; 、 The height and width of the feature map, It is the Sigmoid activation function. It is the ReLU activation function. 、 As the weight of the fully connected layer, the channel attention mechanism enables the network to focus on the features of fault areas (such as the texture of a torn belt or the edge of a worn valve) and suppress irrelevant interference such as dust and light. The result is after compression and activation of the fully connected layer; The result is after feature weighting.
[0133] c. Remove redundant convolutional and pooling layers from the original MobileNet architecture, retain the core feature extraction layers, and replace fixed pooling with adaptive pooling in the pooling layers. Adaptively adjust the pooling window according to the size of the faulty region in the image to improve the flexibility and accuracy of feature extraction.
[0134] The core requirement for sound data is to match the frequency range of sound patterns from coal preparation plant equipment operation (such as bearing noise, fluid noise, and motor vibration). This embodiment uses Mel spectrum transform and an improved CNN network to accurately extract abnormal sound pattern features.
[0135] First, the preprocessed audio signal is sampled at a uniform rate (fixed at 10kHz) and pre-emphasized (using a first-order high-pass filter to suppress low-frequency noise). Then, the time-domain audio signal is converted into a Mel spectrogram, i.e., a two-dimensional feature map, through an optimized Mel spectrum transform. The Mel spectrogram is then input into an improved CNN network, which performs operations such as convolution, activation, and pooling to extract the frequency features and temporal evolution features of the voiceprint. Finally, a 256-dimensional voiceprint feature vector is output for subsequent cross-modal fusion.
[0136] To address the shortcomings of traditional Mel-frequency transform, such as insufficient frequency resolution and small receptive field of CNN network speaker features leading to overfitting, this embodiment optimizes the Mel-frequency transform. Specifically, for the speaker frequency range (100Hz~5kHz) of coal preparation plant equipment, the frequency resolution and frame length settings are optimized. The Mel-frequency transformation formula is as follows:
[0137] in, For frequency.
[0138] The number of Mel filter banks was increased from 40 to 64, the frame length was set to 256ms, and the frame shift was set to 128ms to improve the frequency resolution of the low-frequency band (100Hz~1kHz, corresponding to bearing noise and fluid noise) and ensure that abnormal voiceprint features are not missed. At the same time, spectral smoothing was introduced to reduce the interference of environmental noise on voiceprint features.
[0139] The improvement of the CNN network is to introduce dilated convolution in the feature extraction layer. The dilated convolution is directly embedded in the backbone feature extraction layer of the CNN network, which expands the receptive field of voiceprint features and can capture long-term voiceprint evolution features (such as the change from slight to severe bearing noise) without increasing the number of parameters.
[0140] The calculation logic for dilated convolution is as follows:
[0141] in, The kernel size; The void ratio is set to 2 in this embodiment; convolution kernel The weights are increased by more than 2 times through dilated convolution, thereby improving the ability to capture abnormal audio signature features of devices; Indicates spatial coordinates.
[0142] Meanwhile, a dropout layer is introduced before the fully connected layer of the CNN network, with a dropout probability set to 0.1. L2 regularization (weight decay coefficient 1e-3) is also introduced, and the regularization formula is as follows:
[0143] in, For cross-entropy loss, The regularization coefficient is . The network weight parameters effectively suppress model overfitting and improve the ability to identify and extract abnormal voiceprint features in complex environments of coal preparation plants.
[0144] This embodiment employs a cross-modal attention fusion module to perform feature-level fusion on the extracted single-modal features, mine the correlation features and coupling relationships between different modal data, and generate a fused feature vector.
[0145] This embodiment addresses the fault characteristics of different equipment in a coal preparation plant by deploying multi-source heterogeneous sensors in a differentiated manner. Through single-modal feature extraction and cross-modal attention fusion, it fully explores the coupled correlation characteristics of multi-modal data such as vibration, temperature, sound, vision, and electrical quantities. This effectively solves the problems of single sensors being susceptible to environmental interference and having low recognition accuracy, significantly improves the ability to identify early weak anomalies in equipment, and reduces the false alarm rate and missed alarm rate.
[0146] S3: By fine-tuning the industrial big model, the fused feature vectors are mapped to the vector space of the deep learning knowledge base, realizing the spatial alignment of operating features and fault knowledge.
[0147] In this embodiment, a low-rank adaptation (LoRA) and multi-task joint fine-tuning strategy is used to fine-tune the large industrial model, taking into account fine-tuning efficiency, model generalization ability and adaptability to coal preparation scenarios.
[0148] The construction of the industrial large-scale model fine-tuning dataset involves the following steps: Using historical fault annotation data from coal preparation plants over the past 5-10 years and industry expert experience datasets as the core, supplemented by fault case annotation data from similar equipment in the same industry, a dedicated fine-tuning dataset is constructed. This dataset covers all categories of equipment, including belt conveyors, crushers, jigs, chutes, valves, and coal bunkers, as well as all corresponding fault types (such as belt conveyor misalignment / tears / idler overheating / abnormal noise, crusher bearing overheating / jamming, etc.), ensuring the comprehensiveness and representativeness of the dataset. The historical fault annotation data includes comprehensive annotation information such as equipment operating parameters, fault phenomena, fault types, fault locations, root cause analysis, and handling processes. The expert experience dataset contains structured text annotations from experts on diagnostic logic, handling procedures, and emergency measures for various faults.
[0149] Standardized preprocessing was performed on the dedicated fine-tuning dataset, including: (1) Data deduplication and noise reduction, removing duplicate data, invalid data (such as abnormal data caused by sensor failure) and interference information (such as production parameters unrelated to the failure); (2) Unified labeling, normalizing fault labels with different formats and expressions (such as "overheating of idler rollers" and "overheating of idler rollers"), unifying the labeling standards for fault types, fault levels, and root cause classifications, and generating a standardized labeling system; (3) Data partitioning, dividing the dedicated fine-tuning dataset into training set, validation set, and test set in a ratio of 7:2:1, where the training set is used for updating model parameters, the validation set is used for adjusting hyperparameters and monitoring overfitting, and the test set is used for verifying the final fine-tuning effect; (4) Data format adaptation, converting the preprocessed dedicated fine-tuning dataset into an input format that can be recognized by the industrial large model, with each sample containing "input text (equipment operation description + fault phenomenon) + labeling information (fault type, fault location, root cause label)", and combining it with the constructed coal preparation field dedicated vocabulary. We uniformly encode the technical terms in the samples to ensure semantic consistency.
[0150] In this embodiment, the industrial large model sequentially includes a multi-layer embedding layer, a layer normalization built into the encoder, a feedforward network, a residual connection, a Transformer decoder, and a multi-task output head. The LoRA low-rank adaptation fine-tuning method is adopted to fine-tune only the attention weight matrix of the Transformer encoder layer of the industrial large model, freeze other parameters of the industrial large model, significantly reduce the computational power consumption and memory usage of the fine-tuning process, and avoid catastrophic forgetting of the industrial large model, while retaining the original general semantic understanding and logical reasoning capabilities of the industrial large model.
[0151] The specific design is as follows: In each multi-head self-attention module of the Transformer encoder for large industrial models, the query matrix... Key matrix weight matrix , Insert low-rank matrix pairs respectively , .in For dimension The lower projection matrix, For dimension The upper projection matrix; As a low-rank dimension, this embodiment is set to... Much smaller than the model embedding dimension During fine-tuning, only the low-rank matrix is updated. , The parameters, the original weight matrix , Keep frozen. Specifically, the compressed low-rank space is remapped back to the target high-dimensional space, completing dimension restoration, with the dimension values increasing from small to large; this is defined as the upper projection matrix. The high-dimensional space is mapped to the low-rank space, completing dimension compression, with the dimension values decreasing from large to small; this is defined as the lower projection matrix.
[0152] The fine-tuned attention weight calculation logic is corrected as follows:
[0153]
[0154] in, 、 The fine-tuned query and key matrix weights ensure that the fine-tuned industrial big data model can accurately capture the specific semantic associations of coal preparation equipment faults; , Query matrix Key matrix The weight matrix.
[0155] This embodiment designs three core fine-tuning tasks: fault type classification, fault root cause reasoning, and fault knowledge matching, to collaboratively enhance the industrial big model's semantic understanding, feature matching, and logical reasoning capabilities for coal preparation equipment faults.
[0156] Among them, the fault type classification task takes equipment operation description and fault phenomenon as input, trains industrial big data model to output accurate fault type labels (such as "belt conveyor - idler roller overheating" and "jigging machine - air valve failure"). It is essentially a multi-classification task used to enhance the industrial big data model's ability to identify fault features.
[0157] The root cause reasoning task takes equipment operating parameters and fault types as input to train the industrial large model to output the core root causes of the fault (such as "roller overheating - bearing wear" or "air valve failure - actuator jamming"). It is essentially a sequence generation task used to enhance the logical reasoning ability of the industrial large model.
[0158] The fault knowledge matching task takes fault phenomena as input and trains an industrial big data model to recall matching fault mechanisms and expert experience from a constructed vector knowledge base. Essentially, it is a semantic retrieval matching task used to enhance the collaborative ability between the industrial big data model and the deep learning knowledge base, providing support for subsequent fault diagnosis.
[0159] In the fine-tuning training of the large industrial model, the AdamW optimizer was selected, with an initial learning rate of 1e-5 (the learning rate specific to LoRA parameters) and a weight decay coefficient of 1e-2 to avoid overfitting. The batch size was set to 16, combined with a gradient accumulation strategy (gradient accumulation steps of 4) to balance memory usage and training efficiency. The training epochs were set to 8, using a linear learning rate decay strategy, with a warm-up step of 10% of the total training steps to ensure stable convergence during training.
[0160] In this embodiment, the total loss function for fine-tuning the industrial large model is a weighted sum of the losses from the three fine-tuning tasks, balancing the training priorities of each task. The total loss function is as follows:
[0161] in, Cross-entropy loss for fault type classification tasks is used to measure the error between the model's fault type prediction results and the true labels. To calculate the sequence generation error for the cross-entropy loss of the root cause reasoning task, a greedy decoding strategy is used. The contrastive loss for fault knowledge matching tasks is used to enhance the accuracy of model retrieval matching. , , The weighting coefficients are set to 0.4, 0.3, and 0.3 respectively in this embodiment to prioritize the accuracy of fault type identification.
[0162] During fine-tuning training, the loss value of the validation set, the accuracy of fault type classification, the accuracy of root cause reasoning, and the accuracy of knowledge matching are monitored in real time. An early stopping strategy is adopted, in which training is terminated immediately when the loss of the validation set no longer decreases for two consecutive rounds, and the optimal model weights are saved to avoid model overfitting. At the same time, a random dropout strategy is added during training, with the dropout probability set to 0.1, to further improve the generalization ability of the model.
[0163] The fine-tuned industrial model was validated using a test set, and three core evaluation indicators were set to ensure that the model is suitable for fault diagnosis scenarios in coal preparation equipment. Fault type classification accuracy: The calculation formula is as follows: The verification accuracy rate is required to be ≥98%, and the test accuracy rate is required to be ≥97%. Accuracy of root cause reasoning: The calculation formula is as follows: The verification accuracy rate is required to be ≥95%, and the test accuracy rate is required to be ≥94%. Fault knowledge matching accuracy: The calculation formula is as follows: The verification accuracy rate is required to be ≥96%, and the test accuracy rate is required to be ≥95%.
[0164] If the validation metrics do not meet the requirements, adjust the hyperparameters (such as increasing the low-rank dimension r, adjusting the learning rate, and optimizing the weights of the loss function) and retrain fine-tune until the metrics are met. If the metrics are met, associate and bind the fine-tuned model weights with the constructed vector knowledge base for subsequent feature semantic enhancement, anomaly detection, and root cause analysis.
[0165] In this embodiment, during the fine-tuning process, the industrial big model learns from a large amount of feature-fault knowledge pairing data, forming a nonlinear mapping function internally from the feature space to the semantic space of the knowledge base. The fine-tuned industrial big model then semantically enhances the fused feature vectors, mapping them to the vector space of the deep learning knowledge base, thus completing the spatial matching of features and fault knowledge. Specifically, the fused feature vectors serve as the input to the fine-tuned industrial big model. After initial adaptation by the model embedding layer, a multi-layer Transformer encoder completes the semantic enhancement and vector space mapping. The top-level output of the multi-layer Transformer encoder is the fused feature vector aligned with the knowledge base.
[0166] S4: Input the fused feature vector into the pre-trained anomaly classification network to identify the abnormal state of the device and obtain the initial screening anomaly results; among which, the initial screening anomaly results include the device type and the suspected fault type.
[0167] In this embodiment, taking into account the core deployment requirements of low computing power and low latency of edge computing nodes, an anomaly classification network adapted to the initial screening of anomalies of all types of equipment in coal preparation plants is constructed based on lightweight deep neural networks. The core objective is to quickly complete the determination of "normal / abnormal", equipment type location and preliminary identification of suspected faults, while taking into account both inference efficiency and identification accuracy.
[0168] In this embodiment, the anomaly classification network uses MobileNetV3-Small as its basic architecture. This architecture has the advantages of fewer parameters and faster inference speed, making it suitable for scenarios with limited computing power at edge computing nodes. Redundant convolutional layers, pooling layers, and fully connected layers in the original architecture were removed. For the 512-dimensional multimodal fusion feature vector output by S2, the network input dimension and the structure of each layer were redesigned to strictly control the total number of network parameters to within 1 million, ensuring that the single-sample inference time meets the real-time requirements of the edge.
[0169] Specifically, the anomaly classification network comprises a four-layer core structure: input layer, feature compression layer, feature enhancement layer, and classification output layer. The functions and parameter configurations of each layer are as follows: The input layer receives the 512-dimensional fused feature vector output by S2 and simultaneously completes feature normalization processing to eliminate the dimensional differences between features of different modalities; the feature compression layer uses a 1×1 convolution kernel to compress the 512-dimensional feature vector to 256 dimensions, significantly reducing the computational load without losing core features; the feature enhancement layer uses depthwise separable convolution combined with batch normalization (BN) and ReLU activation function to enhance the ability to distinguish fault features, while introducing a dropout layer (with a deactivation probability set to 0.1) to suppress overfitting; the classification output layer adopts a dual-branch design. The first branch is a binary classifier that outputs the "normal / abnormal" judgment result and confidence level of the device, and the second branch is a multi-label classifier that outputs the device type and the corresponding suspected fault type (such as belt conveyor misalignment, crusher bearing overheating, etc.) and confidence level, accurately matching the initial screening output requirements.
[0170] The parameters of each layer of the anomaly classification network are scientifically initialized. The convolutional layers are initialized using He normality to ensure the rationality of the initialized parameters and accelerate network convergence. The fully connected layers are initialized using Xavier uniformity to avoid gradient vanishing or gradient exploding problems. At the same time, L2 regularization (with the weight decay coefficient set to 1e-3) is introduced to further suppress overfitting of the anomaly classification network and improve its generalization ability to complex working conditions in coal preparation plants. The classification output layer uses the Softmax activation function to map the output results to probability values, enhancing the interpretability of the results and facilitating subsequent confidence determination.
[0171] This embodiment employs a two-stage training strategy of transfer learning pre-training and coal preparation scenario-specific fine-tuning to train the anomaly classification network. This strategy reduces training computational power consumption and shortens the training cycle, while also enabling the network to deeply adapt to the fault characteristics of coal preparation plant equipment, ensuring the accuracy and real-time performance of the initial anomaly screening.
[0172] The transfer learning pre-training process involves: firstly, using a general dataset for industrial equipment anomaly detection that covers the normal and abnormal operating characteristics of various industrial equipment, a lightweight network is pre-trained to establish the foundation for the network's general anomaly recognition capabilities. During pre-training, the Adam optimizer is used, with an initial learning rate of 1e-4, a weight decay coefficient of 1e-3, a batch size of 32, and a cross-entropy loss function. Five training epochs are performed. After pre-training, the basic network weights are saved to lay the foundation for subsequent scenario-based fine-tuning, significantly reducing the computational cost and overfitting risk of training from scratch.
[0173] Using standardized time-series and image datasets preprocessed with S2 as the core, and combining historical equipment operation data from coal preparation plants over the past 3-5 years, a coal preparation plant-specific training dataset for anomaly classification networks is constructed. This training dataset strictly covers normal and abnormal operating conditions for six core equipment categories. Abnormal operating conditions include all typical fault conditions of the six core equipment categories. The ratio of normal to abnormal operating condition data is controlled at 1:1.2 to avoid data imbalance leading to model bias. Each sample is fully labeled, including a 512-dimensional fused feature vector, a "normal / abnormal" label, an equipment type label, and a suspected fault type label. These are randomly divided into training, validation, and test sets in a 7:2:1 ratio for network parameter updates, hyperparameter adjustments, and training effect verification, respectively.
[0174] Based on a training dataset specific to coal preparation plants, the pre-trained network was fine-tuned, focusing on optimizing the parameters of the feature enhancement layer and the classification output layer, while freezing the parameters of other layers to ensure the network quickly adapts to the fault characteristics of coal preparation equipment. During the fine-tuning process, the optimizer was changed to the SGD optimizer (initial learning rate set to 5e-5, momentum set to 0.9), the batch size was adjusted to 16, and an early stopping strategy was adopted to avoid model overfitting. Training was terminated immediately when the validation set loss no longer decreased for two consecutive rounds. The total loss function was set as a weighted sum of binary classification loss and multi-label classification loss, with the binary classification loss (normal / abnormal determination) accounting for 40% of the weight and the multi-label classification loss (equipment type + suspected fault type determination) accounting for 60% of the weight, balancing the training priority of the two determination tasks.
[0175] The trained anomaly classification network was fully validated using a test set. Core evaluation metrics were set as follows: anomaly identification accuracy ≥ 95%, device type determination accuracy ≥ 98%, suspected fault type determination accuracy ≥ 90%, and single-sample inference time ≤ 50ms, ensuring that the real-time initial screening and accuracy requirements at the edge were met. If any metric failed to meet the standard, the network structure (e.g., adjusting feature compression dimensionality, dropout probability) or hyperparameters (e.g., learning rate, batch size) was adjusted, and fine-tuning was performed again until all metrics met the requirements. The optimal model weights were then saved for deployment on edge computing nodes.
[0176] The trained anomaly classification network is deployed on edge computing nodes. The edge computing nodes call the pre-trained anomaly classification network, start the fast inference mode, and turn off the dropout layer in the training phase to improve the inference speed. The inference time of a single sample is strictly controlled within 50ms to ensure that the initial screening response delay of the entire process does not exceed 200ms, which is suitable for the real-time monitoring needs of coal preparation plant equipment.
[0177] The trained anomaly classification network receives the 512-dimensional multimodal fusion feature vector output by S2. It performs L2 normalization on the input fusion feature vector to eliminate the interference of feature modulus differences on the classification results, ensuring the consistency and stability of the input data and providing standardized input for subsequent real-time inference. The trained anomaly classification network outputs initial screening results through two branches: first, a "normal / abnormal" judgment result and corresponding confidence level (a confidence level ≥ 0.8 is considered a valid judgment, while a confidence level below 0.8 is marked as a suspected anomaly requiring further verification); second, the device type, suspected fault type, and corresponding confidence level, clarifying the possible fault type to which the abnormal device belongs. Integrating these results, a standardized initial screening report is generated, including core information such as the name of the abnormal device, the suspected fault type, the judgment confidence level, and the inference time.
[0178] In addition, the standardized initial screening report is simultaneously pushed to the knowledge base retrieval section as a core component of the search terms for subsequent retrieval of relevant content such as fault mechanisms and cases in the knowledge base. At the same time, the initial screening results are fed back to the monitoring terminal of the edge computing node in real time for on-site inspection personnel to view in real time, enabling rapid response to anomalies. If the sample is determined to be in normal working condition, it is marked as normal operating data and simultaneously fed back to the data preprocessing section for subsequent incremental fine-tuning of the network, continuously improving the network's recognition accuracy.
[0179] For belt conveyor idler roller overheating and abnormal noise faults, the anomaly classification network can accurately identify the location of the faulty idler roller and analyze the root cause of the fault as bearing wear, with an identification confidence level of 99.2%; for jig machine air valve faults, the model can accurately distinguish between air valve leakage and actuator jamming faults, with an identification accuracy rate of 98.5%.
[0180] S5: Use the initial screening results and the fused feature vector after mapping and alignment as search terms to search the vector knowledge base of the deep learning knowledge base and obtain knowledge entries that are highly related to the abnormal state of the device.
[0181] In this embodiment, based on the retrieval enhancement generation framework, the initial screening of abnormal results and the fused feature vectors after mapping and alignment are used as search terms. Similarity retrieval is performed in the deep learning knowledge base to recall matching fault mechanisms, cases, feature data and expert experience. S6: Based on the knowledge entries retrieved, the initial screening results of anomalies, and the multimodal operational data, the fine-tuned industrial large model is used to achieve one or more of the following: accurate identification of fault types, fault location, fault severity assessment, and root cause analysis.
[0182] In this embodiment, the recalled knowledge base content, real-time multimodal operation data, and initial screening results are input into the fine-tuned industrial big model. The industrial big model then completes the accurate identification of fault types, fault location, fault severity assessment, and analyzes the root cause of the fault, outputting a standardized fault diagnosis report.
[0183] This embodiment also includes automatically parsing newly added fault cases, handling experience, and equipment parameters through the industrial big data model, and then synchronously updating them to the vector knowledge base to achieve self-iteration and self-optimization of the knowledge base.
[0184] The specific mechanism and update process are as follows: (1) Iterative update triggering mechanism: A triple triggering method of timed triggering + manual triggering + automatic triggering is adopted to ensure that no new knowledge is missed and that updates are timely. Automatic triggering: Once the closed-loop process is completed, if new fault cases, handling experience, maintenance records are generated, or new equipment operating parameters or fault characteristic data are collected, the knowledge base update process will be triggered without manual intervention. Scheduled Trigger: Set a fixed monthly update window (at the end of each month) to centrally summarize and update the newly added knowledge accumulated in that month (including the latest industry standards, failure cases in the same industry, and new experience from experts) to ensure the timeliness of the knowledge base; Manual triggering: When a coal mine adds new equipment, updates equipment parameters, or when the industry releases new operation and maintenance standards or experts provide new handling experience, the operation and maintenance administrator can manually trigger the update process to flexibly adapt to the update needs of special scenarios.
[0185] (2) Expand the scope of knowledge collection: Clarify the core content of iterative updates to ensure that the knowledge base is comprehensive and up-to-date, specifically including: New knowledge related to faults: new fault cases generated by closed-loop handling (including pre-fault operating data, fault phenomena, root cause analysis, handling process, and maintenance results), new fault handling experience from experts, and mechanism and characteristic data of new faults; New equipment-related knowledge: main body parameters, structural drawings, operating procedures, and maintenance standards of newly added equipment in the coal preparation plant; parameter updates and modification records of existing equipment. New industry and standards related knowledge: the latest operation and maintenance standards and safety specifications released by the coal preparation industry, typical failure cases and advanced handling technologies of similar equipment in the same industry; New data-related knowledge includes: newly collected sensor data, feature thresholds of newly added faults, temporal patterns and spectral characteristics, and optimized feature extraction rules during data preprocessing.
[0186] (3) Intelligent parsing and standardization: Following the processing logic of the industrial large model, new knowledge is automatically parsed and standardized to ensure consistency with the original knowledge base in terms of format and semantics. Unstructured knowledge parsing: After fine-tuning the industrial big data model, entity extraction (extracting core entities such as equipment, faults, features, and handling measures) and relationship extraction (mining the relationships between entities, such as "new fault type - corresponding features" and "handling measures - applicable scenarios") are performed on newly added unstructured text such as fault cases, expert experience, and maintenance records. Semantic normalization: According to the established standardization rules, the expression of new knowledge is unified (such as unified fault type naming and handling process specifications), eliminating expression differences, generating standardized knowledge units, and ensuring that the format is consistent with the original knowledge units; Structured knowledge adaptation: For newly added structured data such as equipment parameters and industry standards, classify and organize them according to the classification rules of the original knowledge base (equipment body parameter library, fault mechanism library, etc.), and supplement the corresponding metadata tags (such as equipment type, knowledge category, update time).
[0187] (4) Vectorization and synchronous update of knowledge base: Based on the constructed word embedding model, the vectorization transformation of newly added standardized knowledge units is completed, and the vector knowledge base and the original knowledge base are updated synchronously. Knowledge vectorization: The newly added standardized knowledge units are input into the trained word embedding model to generate 768-dimensional knowledge vectors, and L2 normalization is performed to ensure alignment with the original knowledge vector space; On the one hand, newly added knowledge units (original text + metadata tags) are synchronized to the original knowledge base to supplement the corresponding category directory; on the other hand, normalized knowledge vectors are synchronized to the vector knowledge base, the vector index is updated, the FAISSIVF_FLAT inverted index is used, and the cluster centers are updated synchronously to ensure retrieval accuracy. Mapping relationship update: Synchronously update the bidirectional mapping relationship between knowledge ID, vector index, and original knowledge unit to ensure that newly added knowledge can be quickly retrieved and traced, and seamlessly connected with the knowledge base retrieval process.
[0188] (5) Update verification and anomaly correction: Establish a dual verification mechanism to ensure the accuracy and adaptability of newly added knowledge and avoid the entry of incorrect knowledge. Automatic verification: The semantic consistency and logical rationality of the newly added knowledge are verified through the fine-tuned industrial big model (such as verifying the matching of fault types and characteristics and the feasibility of handling measures). If the verification fails, the abnormality is automatically marked and fed back to the operation and maintenance administrator. Manual review: For newly added knowledge that has passed automatic verification, coal preparation industry operation and maintenance experts will conduct a sample review (the sampling ratio shall not be less than 10%), focusing on reviewing key knowledge such as new fault cases and core equipment parameters to ensure that there are no omissions and no errors. Anomaly Correction: For erroneous knowledge or semantically inconsistent knowledge discovered during verification, the administrator corrects them and re-executes the parsing, vectorization, and update process to ensure the accuracy of the knowledge base.
[0189] (6) Self-iterative optimization mechanism: Combining knowledge base usage scenarios and feedback, the knowledge base is dynamically optimized to improve knowledge usability. Based on the knowledge base retrieval records, the retrieval frequency and matching accuracy of each knowledge unit are statistically analyzed. Knowledge units with high frequency of retrieval and high matching degree are given higher weights, and the retrieval ranking logic is optimized to improve retrieval efficiency. Knowledge units that have not been retrieved for a long time and have low matching accuracy (such as outdated disposal experience and parameters of obsolete equipment) are marked as needing optimization, and experts will decide whether to retain or delete them after evaluation. The newly added knowledge is used as supplementary training data. The word embedding model and the fine-tuned industrial model are incrementally fine-tuned regularly (every quarter) to enhance the model's semantic understanding and matching ability for new faults and new equipment, and realize a closed loop of "knowledge base update - model optimization - recognition accuracy improvement". Record the content, time, triggering method, and operator of each update to form a complete update log, which facilitates traceability, auditing, and troubleshooting, and ensures that the knowledge base update process is controllable.
[0190] This embodiment also includes: based on the fault identification results, and combined with the fault's impact on the coal preparation plant's production, its urgency, and the risk of shutdown, classifying the fault into four warning levels: Level 1 warning is for emergency faults: faults that directly cause equipment shutdowns and safety accidents; such as longitudinal tearing of conveyor belts, complete blockage of chutes, and large-scale detachment of coal bunker walls. These faults directly trigger the highest priority response. Level 2 warning indicates a critical fault: a fault that will cause equipment failure and affect core production processes in the short term; such as severe overheating of crusher bearings, core fault of jig air valve, or severe internal leakage of valves. Such faults must be responded to and dealt with within 4 hours. Level 3 warning is for general faults: the equipment shows an abnormal deterioration trend, which does not affect current production, but will cause faults if it is operated for a long time, such as slight overheating of the idler roller, slight belt misalignment, and slight deviation of valve opening. Such faults need to be confirmed by inspection within 24 hours. Level 4 warning is a reminder: the equipment operating parameters fluctuate slightly and require continuous monitoring; such as slight abnormal noises during equipment operation or slight fluctuations in material level, these faults should be included in the daily inspection plan.
[0191] Tiered early warning push: Based on the early warning level, corresponding push methods and handling procedures are adopted. Level 1 early warnings are directly pushed to the plant-level dispatch and maintenance manager, supporting the triggering of emergency shutdown linkage; Level 2 early warnings are pushed to the workshop director and maintenance team; Level 3 and 4 early warnings are pushed to equipment inspection personnel and included in the inspection plan; through the industrial big data model combined with the expert handling experience in the deep learning knowledge base, corresponding standardized handling plans, emergency measures, spare parts requirements and maintenance schedule suggestions are generated for different levels of early warnings. At the same time, the handling progress is tracked and the handling results are fed back to the deep learning knowledge base to complete the closed-loop management of operation and maintenance.
[0192] This embodiment deeply integrates the retrieval enhancement generation, logical reasoning, and few-shot learning capabilities of industrial large models with the operation and maintenance scenarios of coal preparation plant equipment. Through the RAG framework, it achieves accurate fault diagnosis and root cause analysis, effectively solving the industry pain point of few rare fault samples and high diagnostic difficulty in coal preparation plants. At the same time, it establishes a four-level hierarchical early warning mechanism adapted to the coal preparation production scenario, realizing the intelligentization of the entire process of anomaly identification, hierarchical early warning, and closed-loop handling, and promoting the transformation of coal preparation plant operation and maintenance mode from passive emergency repair to proactive prevention.
[0193] This embodiment adopts an edge-cloud collaborative architecture, which completes real-time data collection and initial anomaly screening at the edge, and completes deep reasoning of large models and knowledge base management in the cloud. This not only ensures the real-time nature of anomaly warnings, but also fully leverages the deep reasoning capabilities of large models. It can be directly adapted to the existing automated system architecture of coal preparation plants, is easy to implement in engineering, and has strong practicality and promotional value.
[0194] Example 2 The purpose of this embodiment is to provide an anomaly identification system for all types of equipment in a coal preparation plant based on a large industrial model, including: The modal fusion module is configured to: extract single-modal features and fuse cross-modal features from the multimodal operating data of each piece of equipment in the coal preparation plant to obtain the fused feature vector of each piece of equipment; The spatial alignment module is configured to map the fused feature vectors to the vector space of the deep learning knowledge base through the fine-tuned industrial big model, thereby achieving spatial alignment between the running features and fault knowledge. The anomaly screening module is configured to: input the fused feature vector into a pre-trained anomaly classification network to identify the abnormal state of the device and obtain the initial screening anomaly result; wherein, the initial screening anomaly result includes the device type and the suspected fault type; The retrieval module is configured to: use the initial screening anomaly results and the fused feature vector after mapping and alignment as search terms to search in the vector knowledge base of the deep learning knowledge base to obtain knowledge entries that are highly related to the abnormal state of the equipment; wherein, the deep learning knowledge base is constructed based on all types of equipment in the coal preparation plant and corresponding multi-dimensional knowledge; The anomaly identification and diagnosis module is configured to: based on retrieved knowledge entries, initial anomaly screening results, and multimodal operational data, utilize a fine-tuned industrial large-scale model to achieve one or more of the following: accurate identification of fault types, fault location, fault severity assessment, and root cause analysis. In further embodiments, the following is also provided: An electronic device includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor. When executed by the processor, the computer instructions perform the method described in Embodiment 1. For brevity, further details are omitted here.
[0195] It should be understood that in this embodiment, the processor can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor or any conventional processor, etc.
[0196] Memory may include read-only memory and random access memory, and provides instructions and data to the processor. A portion of memory may also include non-volatile random access memory. For example, memory may also store information about the device type.
[0197] A computer-readable storage medium for storing computer instructions, which, when executed by a processor, perform the method described in Embodiment 1.
[0198] The method in Embodiment 1 can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor. The software modules can reside in readily available storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in memory, and the processor reads information from the memory and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, a detailed description is not provided here.
[0199] A computer program product includes a computer program that, when executed by a processor, implements the method described in Embodiment 1.
[0200] The present invention also provides at least one computer program product tangibly stored on a non-transitory computer-readable storage medium. The computer program product includes computer-executable instructions, such as instructions included in program modules, which execute in a device on a target real or virtual processor to perform the processes / methods described above. Typically, program modules include routines, programs, libraries, objects, classes, components, data structures, etc., that perform specific tasks or implement specific abstract data types. In various embodiments, the functionality of program modules can be combined or divided among program modules as needed. The machine-executable instructions for the program modules can execute within a local or distributed device. In a distributed device, the program modules can reside in both local and remote storage media.
[0201] The computer program code used to implement the methods of the present invention may be written in one or more programming languages. This computer program code may be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing device, such that when executed by the computer or other programmable data processing device, the program code causes the functions / operations specified in the flowcharts and / or block diagrams to be implemented. The program code may be executed entirely on a computer, partially on a computer, as a stand-alone software package, partially on a computer and partially on a remote computer, or entirely on a remote computer or server.
[0202] In the context of this invention, computer program code or related data may be carried by any suitable carrier to enable a device, apparatus, or processor to perform the various processes and operations described above. Examples of carriers include signals, computer-readable media, and the like. Examples of signals may include electrical, optical, radio, sound, or other forms of propagation signals, such as carrier waves, infrared signals, etc.
[0203] Those skilled in the art will recognize that the units and algorithm steps described in conjunction with the embodiments herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0204] While the specific embodiments of the present invention have been described above in conjunction with the accompanying drawings, this is not intended to limit the scope of protection of the present invention. Those skilled in the art should understand that various modifications or variations that can be made by those skilled in the art without creative effort based on the technical solutions of the present invention are still within the scope of protection of the present invention.
Claims
1. A method for anomaly identification of all types of equipment in a coal preparation plant based on a large industrial model, characterized in that, include: Single-modal feature extraction and cross-modal fusion are performed on the multimodal operation data of each piece of equipment in the coal preparation plant to obtain the fused feature vector of each piece of equipment; By fine-tuning the industrial big model, the fused feature vectors are mapped to the vector space of the deep learning knowledge base, thereby achieving spatial alignment between operational features and fault knowledge. The fused feature vectors are input into a pre-trained anomaly classification network to identify abnormal device states, thereby obtaining preliminary anomaly screening results; wherein, the preliminary anomaly screening results include device type and suspected fault type; The initial screening results and the fused feature vector after mapping alignment are used as search terms to search in the vector knowledge base of the deep learning knowledge base to obtain knowledge entries that are highly related to the abnormal state of the equipment; wherein, the deep learning knowledge base is constructed based on all types of equipment in the coal preparation plant and corresponding multi-dimensional knowledge. Based on the retrieved knowledge entries, initial screening results of anomalies, and multimodal operational data, a fine-tuned industrial large model is used to achieve one or more of the following: accurate identification of fault types, fault location, fault severity assessment, and root cause analysis.
2. The method for anomaly identification of all types of equipment in a coal preparation plant based on a large industrial model as described in claim 1, characterized in that, The construction of the vector knowledge base of the deep learning knowledge base specifically involves: constructing a five-dimensional knowledge system covering all types of equipment in a coal preparation plant; extracting entities, extracting relations, and normalizing the semantics of the knowledge in the five-dimensional knowledge system through an industrial big data model to generate standardized knowledge units; then converting all standardized knowledge units into knowledge vectors through a word embedding model to form a unified vector set; and identifying the knowledge base type to which each piece of knowledge belongs through metadata tags; the five-dimensional knowledge system includes an equipment ontology parameter library, a fault mechanism library, a sensor feature library, a fault case library, and an expert handling experience library.
3. The method for anomaly identification of all types of equipment in a coal preparation plant based on a large industrial model as described in claim 1, characterized in that, For each type of equipment in the coal preparation plant, single-modal features are extracted from the multimodal operation data of each piece of equipment. Then, cross-modal feature-level fusion is performed on the single-modal features to obtain the fused feature vector for each piece of equipment. Specifically: Temporal correlation features are extracted from time-series data using an improved bidirectional LSTM network or an improved Transformer network to obtain temporal correlation feature vectors. The improvement of the bidirectional LSTM network is as follows: a normalization layer is added before each input layer, and residual connections are introduced between the input and output of the bidirectional LSTM network. The improvement of the Transformer network is as follows: local window attention is used to replace global attention, and adaptive temporal position coding is introduced to enhance the extraction of temporal evolution features of equipment faults. For image data, an improved CNN network is used to extract visual features to obtain visual feature vectors. The improvement of the CNN network is that a channel attention mechanism is introduced in the feature enhancement layer to strengthen the features of fault regions. For audio data, a Mel spectrogram is generated through Mel spectrum transform, and then a CNN network is used to extract voiceprint features to obtain a voiceprint feature vector. The cross-modal attention fusion module performs cross-modal feature-level fusion of temporal correlation feature vectors, visual feature vectors, and voiceprint feature vectors to obtain the fused feature vector for each device.
4. The method for anomaly identification of all types of equipment in a coal preparation plant based on a large industrial model as described in claim 1, characterized in that, By mapping the fused feature vector to the vector space of the deep learning knowledge base through the fine-tuned industrial big model, the spatial alignment of the running features and fault knowledge is achieved. Specifically, the fused feature vector is semantically enhanced and dimensionally transformed using the fine-tuned industrial big model, so that the aligned feature vector is in the same semantic space as the knowledge vector in the vector knowledge base, and the aligned feature vector is obtained with the same dimension as the knowledge vector in the deep learning knowledge base. Specifically, the fine-tuning of the industrial big model involves: using low-rank adaptation and multi-task joint fine-tuning strategies based on historical fault labeling data of coal preparation plant equipment and expert experience datasets; the multi-tasks include fault type classification task, fault root cause reasoning task, and fault knowledge matching task.
5. The method for anomaly identification of all types of equipment in a coal preparation plant based on a large industrial model as described in claim 1, characterized in that, The training process of the anomaly classification network includes: Based on historical operating data of coal preparation plant equipment under normal and abnormal operating conditions, a general dataset for industrial equipment anomaly detection is constructed. The anomaly classification network is pre-trained using the general dataset for industrial equipment anomaly detection through transfer learning. The pre-trained anomaly classification network is fine-tuned based on a coal preparation plant-specific training dataset. During the fine-tuning process, network parameters other than the feature enhancement layer and the classification output layer are frozen to adapt to the fault characteristics of the coal preparation equipment. The construction of the coal preparation plant-specific training dataset includes: constructing normal operating condition datasets and abnormal operating condition datasets for the core equipment of the coal preparation plant based on preprocessed standardized time-series datasets and image datasets, combined with historical equipment operation data of the coal preparation plant. The abnormal operating condition dataset includes a training dataset of all typical fault conditions of the core equipment.
6. The method for anomaly identification of all types of equipment in a coal preparation plant based on a large industrial model as described in claim 2, characterized in that, The training of the word embedding model is specifically as follows: Based on the coal preparation-specific entities and relationships corresponding to standardized knowledge units, a corpus dedicated to the coal preparation field is constructed. Based on mutual information and left-right entropy algorithms, coal preparation-specific professional terms are mined from the coal preparation-specific corpus to construct a coal preparation-specific lexicon. Based on the constructed coal preparation-specific vocabulary and entity extraction results, the word embedding model is jointly trained using the masked language model task as the main training task and the fault entity matching task as the auxiliary training task. Among them, an improved Transformer encoder is used as the word embedding model. The improvement of the Transformer encoder is as follows: a fault entity attention bias matrix is introduced into the self-attention mechanism. When the i-th token and the j-th token in the input sequence belong to a predefined fault-related entity pair, the corresponding element of the fault entity attention bias matrix is set to a trainable non-zero bias value.
7. A coal preparation plant full-category equipment anomaly identification system based on an industrial large-scale model, characterized in that, include: The modal fusion module is configured to: extract single-modal features and fuse cross-modal features from the multimodal operating data of each piece of equipment in the coal preparation plant to obtain the fused feature vector of each piece of equipment; The spatial alignment module is configured to map the fused feature vectors to the vector space of the deep learning knowledge base through the fine-tuned industrial large model, thereby achieving spatial alignment between the running features and fault knowledge. The anomaly screening module is configured to: input the fused feature vector into a pre-trained anomaly classification network to identify the abnormal state of the device and obtain the initial screening anomaly result; wherein, the initial screening anomaly result includes the device type and the suspected fault type; The retrieval module is configured to: use the initial screening anomaly results and the fused feature vector after mapping and alignment as search terms to search in the vector knowledge base of the deep learning knowledge base to obtain knowledge entries that are highly related to the abnormal state of the equipment; wherein, the deep learning knowledge base is constructed based on all types of equipment in the coal preparation plant and corresponding multi-dimensional knowledge; The anomaly identification and diagnosis module is configured to: based on the retrieved knowledge entries, initial anomaly results, and multimodal operational data, utilize a fine-tuned industrial large model to achieve one or more of the following: accurate identification of fault types, fault location, fault severity assessment, and root cause analysis.
8. An electronic device, characterized in that, It includes a memory and a processor, as well as computer instructions stored in the memory and running on the processor, which, when executed by the processor, perform the method according to any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, Used to store computer instructions, which, when executed by a processor, perform the method described in any one of claims 1-6.
10. A computer program product, characterized in that, Includes a computer program, which, when executed by a processor, implements the method described in any one of claims 1-6.
Citation Information
Patent Citations
Fault diagnosis question-answering system based on multi-modal knowledge graph and large language model
CN119988638A
Intelligent domain decision-making method and system based on knowledge base
CN120611795A