Abnormality detection method based on hybrid expert field adaptive industrial large model

By introducing hybrid expert adaptive technology in industrial big models, combining explicit industrial knowledge and multimodal data, the problems of insufficient generalization ability and low computing efficiency of industrial big models in abnormal detection are solved, and more efficient and flexible abnormal detection capabilities are achieved.

CN120011862AActive Publication Date: 2025-05-16ZHEJIANG UNIV

Patent Information

Application Number
CN202510469021.6
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-04-15
Publication Date
2025-05-16
Estimated Expiration
2045-04-15

AI Technical Summary

Technical Problem

The existing industrial large models have problems such as insufficient generalization capabilities, low computational efficiency and difficulty in making full use of explicit industrial knowledge in the anomaly detection task.

Method used

An anomaly detection method based on adaptive industrial large models in the hybrid expert field is proposed. By collecting multimodal data, preprocessing and feature alignment, combining explicit industrial knowledge, a hybrid expert model is constructed, and trained and optimized to achieve efficient anomaly detection.

Benefits of technology

Through the context learning architecture, the integration of industrial knowledge and multimodal data can be reduced to overfitting; the hybrid expert model is used to improve computing efficiency and task adaptability; efficient training and reasoning are achieved on the domestic intelligent computing platform to improve the generalization ability and adaptability of the model.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120011862A_ABST
    Figure CN120011862A_ABST
Patent Text Reader

Abstract

The invention discloses an anomaly detection method based on a hybrid expert field adaptive industrial large model. According to the method, a hybrid expert model and a field adaptive technology are introduced, and an efficient industrial anomaly detection scheme is provided for the problems of training efficiency and adaptability of an industrial large model in multi-field tasks. Comprising the steps of data preprocessing, field division and self-adaption, industrial knowledge and data fusion, hybrid expert model construction, model training and reasoning adaptation and the like. The hybrid expert model can automatically select a proper expert for reasoning according to input data, and the domain adaptive module effectively reduces the distribution difference of data in different domains, so that the migration ability and generalization performance of the model are improved. The method is based on data driving, can adapt to multi-modal data of different industrial scenes, and has high universality. Compared with the prior art, the method has higher training efficiency and better model adaptability in industrial application, and both theoretical property and practicability are enhanced.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of anomaly detection of industrial large models, and in particular to an anomaly detection method based on hybrid expert domain adaptive industrial large models. Background Art

[0002] With the rapid development of industrial automation and intelligent manufacturing, the demand for artificial intelligence (AI) and large-scale models in the industrial field continues to grow, especially in industrial data analysis, fault prediction, quality control, and optimization decision-making. However, traditional industrial models face many challenges in practical applications, including data heterogeneity, knowledge differences between fields, and the cost of large-scale computing. In particular, industrial data usually has the characteristics of high knowledge ratio and low total data volume, which is significantly different from the massive data training method that general large models rely on. As a result, industrial large models are prone to insufficient generalization ability and low computational efficiency in tasks such as anomaly detection.

[0003] Existing industrial large models mainly rely on large-scale historical data for training. However, due to the scarcity of industrial data and the difficulty in obtaining high-quality annotations, simple data-driven methods are difficult to achieve ideal results in complex industrial environments. At the same time, industrial knowledge is often explicit and highly specialized. Traditional data-driven models find it difficult to fully mine and utilize this knowledge, which can easily lead to model overfitting or lack of adaptability. In addition, when processing multimodal data (such as images, text, sensor data, etc.), industrial large models also face the challenges of multi-source data fusion and cross-modal feature alignment. Summary of the invention

[0004] The present invention aims to address the deficiencies of the prior art and propose an anomaly detection method based on a hybrid expert domain adaptive industrial large model.

[0005] The object of the present invention is achieved through the following technical solution: an anomaly detection method based on a hybrid expert domain adaptive industrial large model, comprising the following steps:

[0006] S1. Collect multimodal data from different fields in industrial environments and preprocess them according to the data type;

[0007] S2, project the preprocessed data from different fields into a unified feature space, use the pre-trained domain discriminator for feature alignment, and use the pre-trained domain autoencoder for feature conversion to make the distribution of data from different fields closer;

[0008] S3, integrating industrial knowledge with processed multimodal industrial data through explicit industrial knowledge;

[0009] S4, constructing a hybrid expert model and training it, wherein the hybrid expert model is composed of a plurality of hybrid expert modules, each of which includes a linear projection layer, a self-attention layer, a gating network, a hybrid expert layer and an output layer, and the hybrid expert layer is an expert layer for different data processing methods;

[0010] S5. Optimize the intelligent computing platform and ensure hardware compatibility, and use the optimized hybrid expert model to perform anomaly detection and reasoning on the data after multi-modal industrial data fusion.

[0011] Furthermore, the multimodal data includes: one or more of text records, industrial images, video streams, and sensor data.

[0012] Furthermore, the preprocessing according to the data type is specifically as follows:

[0013] For time series data, the sliding window method is used for segmentation, and Fourier transform or wavelet transform is used to extract time-frequency features;

[0014] For text data, perform word segmentation and part-of-speech tagging, and use the word vector model for vectorization;

[0015] For image data, convolutional neural networks are used to extract visual features, and data enhancement technology is used to improve data diversity.

[0016] Furthermore, the pre-training process of the domain discriminator specifically includes: making the domain discriminator unable to distinguish the data distribution of the source domain and the target domain through adversarial training, and using a gradient reversal layer to optimize the loss function.

[0017] Furthermore, the domain autoencoder performs joint training of multiple domains during the pre-training process.

[0018] Furthermore, the fusion of industrial knowledge with processed multimodal industrial data through explicit industrial knowledge is specifically as follows:

[0019] S3.1. Extract explicit industrial knowledge from the knowledge base of the industrial field and transform it into structured data, rule sets or construct industrial knowledge graphs;

[0020] S3.2, use the multimodal encoder to extract features from different data types respectively, and map the data of multiple modes into a unified feature space.

[0021] S3.3. Contextual learning is performed by building a contextual learning architecture based on the domain knowledge graph and combining it with the domain knowledge graph.

[0022] Furthermore, in the hybrid expert model, the expert layers for different data processing methods specifically include the following experts:

[0023] Fully connected experts: contain one or more fully connected layers

[0024] Convolution Expert: Contains one or more convolutional layers

[0025] Transformer Expert: Contains the Transformer architecture with attention mechanism

[0026] Zero Expert: directly discard the input;

[0027] Copy Expert: Skip the current hybrid expert layer;

[0028] Constant Expert: performs constant correction and substitution on input;

[0029] The expert layers in each hybrid expert module include four of them.

[0030] Furthermore, the gated neural network scores the input features, calculates the activation probability of each expert, and selects the activated top-2 experts. The activation probability is specifically:

[0031] .

[0033] On the other hand, the specification also provides an anomaly detection device based on a hybrid expert domain adaptive industrial large model, including a memory and one or more processors, wherein the memory stores executable code, and when the processor executes the executable code, it implements the anomaly detection method based on a hybrid expert domain adaptive industrial large model.

[0034] On the other hand, the specification also provides a computer-readable storage medium on which a program is stored. When the program is executed by a processor, the anomaly detection method based on a hybrid expert domain adaptive industrial large model is implemented.

[0035] Beneficial effects of the present invention:

[0036] In the first aspect, the present invention first integrates explicit industrial knowledge and multimodal industrial data through a contextual learning architecture. The architecture is based on the continuous training phase of pre-training and contextual learning reasoning (i.e., model warm-up), which further optimizes the knowledge expression of the model on the basis of improving the contextual learning ability of the large model. By combining explicit industrial knowledge (such as industry specifications, operating manuals, etc.) with multimodal data in the industrial field (such as images, texts, sensor data, etc.), the overfitting phenomenon in the data-driven model is significantly reduced.

[0037] On the second aspect, the present invention proposes an innovative hybrid expert (MoE) model. By introducing multiple expert networks and a dynamic selection mechanism, the model can dynamically activate the appropriate expert network for reasoning when facing a variety of industrial tasks, thereby improving computing efficiency and task adaptability. In the hybrid expert architecture, by designing multiple types of neural network experts, the operating efficiency is improved, the computing overhead is reduced, and the reasoning efficiency is improved. On the third aspect, in order to achieve efficient training and reasoning on different intelligent computing platforms, the present invention optimizes the large model training and reasoning process on the domestic intelligent computing platform. On the one hand, the present invention can be adapted to domestic chips (such as Ascend, Kunlun) and deep learning frameworks (such as MindSpore, PaddlePaddle), and optimize compilation options and computing libraries. For example, on the Ascend chip, the CANN operator library is used to optimize matrix calculations to reduce computing delays. On the other hand, the present invention adjusts the underlying operators and adopts methods such as operator fusion (BatchNorm is fused with Conv to reduce memory access, which can accelerate convolution calculations in image anomaly detection tasks), precision optimization (using BF16 / INT8 to accelerate reasoning, and using low-precision calculations to improve reasoning speed in time series anomaly detection tasks), and computational parallelization (using data parallelism and model parallelism strategies to improve computational efficiency in a multi-card training environment). The industrial large model of the present invention can perform efficient multi-task parallel training and reasoning on a domestic platform with a high-performance computing cluster and large-scale data storage.

[0038] Fourthly, the model of the present invention can be effectively migrated from one industrial field to another through domain adaptation technology. The method improves the adaptability and generalization ability of the model in unknown fields by jointly training data from multiple fields, using shared knowledge and cross-domain learning. Specifically, when processing different industrial scenarios, the model can dynamically adjust the activation mode of its expert network, thereby improving the processing ability of different tasks and data. Fifthly, the method of the present invention adopts a strategy combining large-scale pre-training with fine-tuning. First, the large model is pre-trained through self-supervised learning or generative models to learn the common features of the data. Then, using domain adaptation technology, the pre-trained model is applied to specific industrial tasks, such as equipment fault diagnosis, production line optimization, etc., to further improve the performance of the model. BRIEF DESCRIPTION OF THE DRAWINGS

[0039] Figure 1 It is a flow chart of the hybrid expert-based domain adaptive industrial large model anomaly detection method implemented by the present invention;

[0040] Figure 2 It is the functional architecture specifically included in each module of the present invention;

[0041] Figure 3It is a structural diagram of the hybrid expert model for industrial large models implemented by the present invention;

[0042] Figure 4 It is a structural diagram of a hybrid expert layer (MoE Layer) in the hybrid expert model of the present invention;

[0043] Figure 5 It is a flow chart of data processing by the hybrid expert model in the present invention;

[0044] Figure 6 are various experts involved in the hybrid expert model of the present invention;

[0045] Figure 7 It is a schematic diagram of an anomaly detection device based on a hybrid expert domain adaptive industrial large model provided by an embodiment of the present invention. DETAILED DESCRIPTION

[0046] The specific implementation modes of the present invention are further described in detail below with reference to the accompanying drawings.

[0047] like Figure 1 As shown, the present invention provides an anomaly detection method based on a hybrid expert domain adaptive industrial large model, and the specific implementation method is as follows:

[0048] Step 1: Collect multimodal data in an industrial environment, including but not limited to text records (production logs, maintenance reports, fault descriptions, etc.), industrial images (image data such as equipment operating status and product appearance inspection), video streams (production line monitoring videos, equipment operation process, etc.), and sensor data (time series data such as temperature, pressure, and vibration).

[0049] Step 2: Clean and standardize the collected data, including denoising, filling missing values, anomaly detection, and removing redundant data. This method has different processing methods for data of different modalities.

[0050] For time series data, the present invention adopts a sliding window method for slicing and uses Fourier transform or wavelet transform to extract time-frequency features;

[0051] For text data, the present invention performs word segmentation and part-of-speech tagging, and uses a word vector model (such as Word2Vec or BERT) for vectorization;

[0052] For image data, the present invention adopts convolutional neural network (CNN) to extract visual features and uses data enhancement technology to improve data diversity.

[0053] Step 3: Project the preprocessed data into a unified feature space and use domain adaptation technology to reduce the differences in data distribution between different domains. The domain discriminator network structure can use a fully connected neural network (FCNN) to perform adversarial training through input features and domain labels; the domain autoencoder can use a convolutional neural network (CNN) or a multi-layer perceptron (MLP) to ensure that the domain features of the input data are effectively extracted by minimizing the reconstruction error. Taking the semiconductor manufacturing and petrochemical industries as examples, it can specifically include the following sub-steps:

[0054] (1) Map industrial data from different fields to a unified feature space to make it easier to align in high-dimensional space. Use self-supervised learning methods, such as contrastive learning or variational autoencoders (VAE), to enhance feature expression capabilities.

[0055] For example, during training, the data from a semiconductor factory (source domain) and the data from a petrochemical factory (target domain) are paired, and the feature distances between normal samples and abnormal samples are compared to ensure that normal samples under normal process conditions have similar feature vectors.

[0056] (2) Introduce a domain discriminator and use adversarial training to make it impossible for the domain discriminator to distinguish the data distribution of the source domain and the target domain, thereby achieving feature alignment. Specifically, a gradient reversal layer (GRL) is used to optimize the loss function so that the domain discriminator can simultaneously minimize the task loss and maximize the domain discrimination loss.

[0057] Building a domain discriminator , after inputting the feature, we determine its source (semiconductor factory or petrochemical factory). Through gradient reversal (GRL), the feature distribution is made as close as possible during training to weaken the domain differences. The loss function formula can be as follows:

[0058]

[0059] in, are the characteristics of the input data, is the predicted probability of the domain classifier.

[0060] (3) Using Domain Autoencoder for feature conversion makes the distribution of data in different domains closer. In the pre-training stage, multi-source domain data is used for joint training to enhance the cross-domain generalization ability of the Domain Autoencoder.

[0061] For example, a domain autoencoder is trained to reconstruct the data of semiconductor plants and petrochemical plants in the same feature space. Joint training is performed in the pre-training phase to ensure cross-domain generalization capabilities. In the inference phase, the model can automatically adjust the features of different plants to make them closer to the same distribution.

[0062] In simple terms, this step achieves alignment of feature distribution by constructing adversarial loss between the source domain and the target domain. At the same time, the domain transfer learning method is adopted to improve the generalization ability of large models for data in different domains through cross-domain joint training and knowledge sharing mechanisms while retaining domain characteristics and enhancing feature expression capabilities.

[0063] Step 4: Extract explicit industrial knowledge (such as rules, specifications, failure modes, etc.) from the knowledge base of the industrial field, and integrate the industrial knowledge with multimodal industrial data (text, images, etc.) to enhance the model's understanding and application capabilities for industrial scenarios. This process includes multiple sub-steps to ensure the seamless integration of knowledge and data, which are implemented as follows:

[0064] (1) First, explicit industrial knowledge (such as rules, specifications, failure modes, etc.) needs to be extracted from the knowledge base in the industrial field. Common sources of industrial knowledge include industrial standards and specifications (such as equipment operation standards, fault diagnosis specifications, and maintenance manuals), failure mode and effect analysis (this method can help identify and analyze potential failure modes of equipment and their possible impacts), expert experience (through expert knowledge and intuition to conduct qualitative analysis of equipment operation and identify early warning signs of equipment failure), knowledge graphs (structuring industrial knowledge and representing it in the form of graphs so that the system can acquire knowledge through query, reasoning, etc.), etc.

[0065] When this knowledge is converted into a form that can be processed by the model, rules such as fault diagnosis and maintenance processes can be converted into structured data or rule sets, so that the system can be directly applied to anomaly detection or prediction tasks. It can also be achieved by constructing an industrial knowledge graph, presenting elements such as equipment, faults, maintenance, and operations as graph nodes, and performing reasoning and queries through the connection relationships of the graph.

[0066] (2) In this stage, the goal is to map data from multiple modalities, such as text, images, and sensor data, into a unified feature space so that data from different modalities can interact with each other and extract cross-modal information. This process involves using multimodal encoders (such as variants such as Transformer and BERT) to extract features from different data types separately. For example, CNN is used to extract features from image data, RNN or Transformer is used to process time series data, and BERT is used to process text data. These encoders convert their respective features into a unified vector representation. In order to enhance the spatial and temporal association of multimodal data, relative position encoding is used to adjust the relationship between different modalities. For example, in the fusion of image and text data, one-dimensional or two-dimensional position encoding can be used to optimize the representation of feature sequences and ensure effective spatial alignment of images and text.

[0067] (3) In the process of multimodal data fusion, it is crucial to combine domain knowledge graphs for contextual learning of data. In the process of data fusion, by building a contextual learning architecture based on domain knowledge graphs, a Transformer with cross-attention mechanism is used to interpret and reason about abnormal patterns in the data based on known industrial standards and knowledge background.

[0068] For example, in a smart manufacturing environment, the model predicts and diagnoses faults in a factory's CNC (numerical control machine tool) equipment. By learning explicit knowledge such as industrial standards (GB / T 20759-2020 "Reliability Evaluation of CNC Machine Tools") and expert experience, the model can infer the type of anomaly. For example, when the spindle vibration frequency exceeds 5kHz and the temperature rises by more than 10°C, it may be spindle bearing wear. If the motor power consumption increases abnormally (exceeding the normal value by 20%), it may be tool jamming or spindle overload.

[0069] Step 5: Design a hybrid expert model and complete the core capability training of the large model through an efficient hybrid expert architecture.

[0070] In this step, refer to Figure 3 The overall architecture diagram of the hybrid expert model and Figure 4 The MoE layer structure diagram in the figure shows that the hybrid expert model consists of N MoE Blocks. The training and reasoning process of the model can be referred to Figure 5 The sub-steps are described in detail as follows:

[0071] (1) Input data ,in is the input sequence length, is the batch size, is the feature dimension. Query, )、Key(Key, ) and Value, ) to generate the basic input of the attention mechanism: , , ,in , , is the projection weight matrix.

[0072] (2) The features are passed in parallel to the attention mechanism, where the attention mechanism is used to calculate the global attention feature, calculated as ,in is the dimension of the attention space. This step uses multiple attention heads to calculate in parallel and output feature dimensions Keeping it unchanged, global features are extracted to support shared knowledge across domains.

[0073] (3) The output of the attention mechanism is passed to the routing mechanism, which dynamically selects the top-2 experts to be activated. That is, the gating network is used to score the input features and calculate the activation probability of each expert. .in is the full-text feature representation obtained by the attention mechanism in (2), is the learnable parameter matrix of the gating network, is the bias term of the gating network.

[0074] (4) The expert head processes domain-specific features. Activate the most relevant experts to avoid all experts from participating in the calculation to improve efficiency. The activated experts receive Input and perform feature extraction respectively.

[0075] (5) Weighted fusion of the outputs of the shared head and the expert head, and the output of the attention mechanism and expert header output Fusion by weight ,in , is the fusion weight parameter.

[0076] (6) Generate the final output through a feedforward neural network.

[0077] The expert basic neural network structure included in this step can be referred to Figure 6 , as follows:

[0078] Fully Connected Expert: Contains one or more fully connected layers, suitable for processing low-dimensional, unstructured feature inputs, such as sequence embedding or flattened image features;

[0079] Convolutional Experts: Contains one or more convolutional layers and is suitable for processing data with local correlation, such as images, time series, speech signals, etc.

[0080] Transformer Experts: Contains the Transformer architecture with an attention mechanism, emphasizing the modeling of long-range dependencies and global semantic information;

[0081] Zero Expert: directly discard the input;

[0082] Copy Expert: Skip the current mixed expert layer;

[0083] Constant Expert: Performs constant modification and substitution on inputs.

[0084] (7) Loss function: The cross-entropy loss of the anomaly detection task is used as the main task loss. In order to prevent some experts from being overused, the load balancing loss is used as the expert regularization loss to encourage the balanced use of different experts:

[0085]

[0086] For the routing mechanism, in order to ensure the sparse activation of the expert, an L1 regularization loss is usually added as a gating loss:

[0087]

[0088] in is the weight of the router.

[0089] (8) Backpropagation and optimization: Calculate the gradient of the loss function and use an optimization algorithm (such as Adam or SGD) to update the parameters.

[0090] The various Experts designed in the present invention expand the combination mode of model processing input from the original single mode to multiple modes, thereby improving the model's expressiveness and task processing flexibility. All inputs will be processed by two experts together, so different data processing methods can be applied to different types of data, thereby improving the model's expressiveness and task processing flexibility. In addition, the three types of experts that do not require computing costs, namely Constant Expert, Zero Experts and Copy Experts, can significantly reduce computing overhead and achieve the purpose of reducing computing resource consumption.

[0091] Step 6: For the intelligent computing platform, optimize the training architecture and underlying operators to achieve efficient model training and reasoning, and then use the operator-tuned model to perform industrial anomaly detection to obtain the final anomaly detection results.

[0092] This step mainly optimizes the training and reasoning process of large models, including operator tuning, distributed training, and hardware compatibility design, to improve the performance of the model on the local intelligent computing platform. In the training phase, the underlying operators are optimized to achieve efficient processing of multimodal data, support multi-task parallel training, and effectively improve model performance. In the reasoning phase, model optimization technologies such as pruning and quantization are combined to reduce computational complexity and improve reasoning efficiency.

[0093] According to the above method embodiment, the present invention specification also provides an embodiment of an anomaly detection system based on a hybrid expert domain adaptive industrial large model, such as Figure 2 As shown, the system includes a data preprocessing module, a domain adaptation module, an industrial knowledge and data fusion module, and a hybrid expert model;

[0094] The data preprocessing module preprocesses the collected multimodal data in different fields according to the data type;

[0095] The domain adaptation module projects the pre-processed data of different domains into a unified feature space, uses a pre-trained domain discriminator for feature alignment, and uses a pre-trained domain autoencoder for feature conversion, so that the distribution of data in different domains is closer;

[0096] The industrial knowledge and data fusion module fuses the industrial knowledge with the processed multimodal industrial data through explicit industrial knowledge;

[0097] The hybrid expert model is composed of several hybrid expert modules, each of which includes a linear projection layer, a self-attention layer, a gating network, a hybrid expert layer and an output layer. The hybrid expert layer is an expert layer for different data processing methods.

[0098] The specific implementation process of the present invention is as follows: build and deploy the above system. After the data preprocessing module completes the formatting and feature extraction of the input data, the domain adaptation module calibrates the feature distribution, the fusion module performs joint modeling of data and knowledge, and finally completes the training of the large model in the hybrid attention expert module. In the model reasoning stage, the optimized intelligent computing platform is used to realize the real-time and efficient reasoning process to meet the application requirements of industrial field anomaly detection. The remaining detailed steps are detailed in the implementation process of the corresponding steps in the above method, which will not be repeated here.

[0099] Corresponding to the aforementioned embodiment of an anomaly detection method based on a hybrid expert domain adaptive industrial large model, the present invention also provides an embodiment of an anomaly detection device based on a hybrid expert domain adaptive industrial large model.

[0100] See also Figure 7 An embodiment of the present invention provides an anomaly detection device based on a hybrid expert domain adaptive industrial large model, including a memory and one or more processors, wherein the memory stores executable code, and when the processor executes the executable code, it is used to implement an anomaly detection method based on a hybrid expert domain adaptive industrial large model in the above embodiment.

[0101] An embodiment of an anomaly detection device based on a hybrid expert domain adaptive industrial large model provided by the present invention can be applied to any device with data processing capabilities, and the any device with data processing capabilities can be a device or apparatus such as a computer. The device embodiment can be implemented through software, or through hardware or a combination of software and hardware. Taking software implementation as an example, as a device in a logical sense, it is formed by the processor of any device with data processing capabilities in which it is located reading the corresponding computer program instructions in the non-volatile memory into the internal memory for execution. From a hardware perspective, if Figure 7 As shown, it is a hardware structure diagram of any device with data processing capability where an abnormality detection device based on a hybrid expert domain adaptive industrial large model provided by the present invention is located, except Figure 7 In addition to the processor, memory, network interface, and non-volatile memory shown, any device with data processing capabilities in which the apparatus in the embodiments is located may also include other hardware, generally based on the actual functions of the device with data processing capabilities, which will not be described in detail.

[0102] The implementation process of the functions and effects of each unit in the above-mentioned device is specifically described in the implementation process of the corresponding steps in the above-mentioned method, and will not be repeated here.

[0103] For the device embodiment, since it basically corresponds to the method embodiment, the relevant parts can refer to the partial description of the method embodiment. The device embodiment described above is only schematic, wherein the units described as separate components may or may not be physically separated, and the components displayed as units may or may not be physical units, that is, they may be located in one place, or they may be distributed on multiple network units. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of the present invention. Ordinary technicians in this field can understand and implement it without paying creative work.

[0104] An embodiment of the present invention further provides a computer-readable storage medium having a program stored thereon. When the program is executed by a processor, an anomaly detection method based on a hybrid expert domain adaptive industrial large model in the above embodiment is implemented.

[0105] The computer-readable storage medium may be an internal storage unit of any device with data processing capability described in any of the aforementioned embodiments, such as a hard disk or a memory. The computer-readable storage medium may also be an external storage device of any device with data processing capability, such as a plug-in hard disk, a smart media card (SMC), an SD card, a flash card, etc. equipped on the device. Furthermore, the computer-readable storage medium may also include both an internal storage unit and an external storage device of any device with data processing capability. The computer-readable storage medium is used to store the computer program and other programs and data required by any device with data processing capability, and may also be used to temporarily store data that has been output or is to be output.

[0106] The present invention also provides a computer program product, including a computer program, and when the computer program is executed by a processor, the anomaly detection method based on a hybrid expert domain adaptive industrial large model is implemented.

[0107] Those skilled in the art will readily appreciate other embodiments of the present application after considering the description and practicing the contents disclosed herein. The present application is intended to cover any modification, use or adaptation of the present application, which follows the general principles of the present application and includes common knowledge or customary techniques in the art that are not disclosed in the present application. The description and examples are intended to be exemplary only, and the true scope and spirit of the present application are indicated by the claims.

[0108] It should be understood that the above general description and the detailed description below are only exemplary and explanatory and cannot limit the present application. The present application is not limited to the precise structure described above and shown in the drawings, and various modifications and changes can be made without departing from the scope thereof. The scope of the present application is limited only by the attached claims.

Claims

1. An anomaly detection method based on a hybrid expert domain adaptive industrial large model, characterized in that: The steps include: S1. Collect multimodal data from different fields in industrial environments and preprocess them according to the data type; S2, project the preprocessed data from different fields into a unified feature space, use the pre-trained domain discriminator for feature alignment, and use the pre-trained domain autoencoder for feature conversion to make the distribution of data from different fields closer; S3, integrating industrial knowledge with processed multimodal industrial data through explicit industrial knowledge; S4, constructing a hybrid expert model and training it, wherein the hybrid expert model is composed of a plurality of hybrid expert modules, each of which includes a linear projection layer, a self-attention layer, a gating network, a hybrid expert layer and an output layer, and the hybrid expert layer is an expert layer for different data processing methods; S5. Optimize the intelligent computing platform and ensure hardware compatibility, and use the optimized hybrid expert model to perform anomaly detection and reasoning on the data after multi-modal industrial data fusion.

2. The anomaly detection method based on hybrid expert domain adaptive industrial big model according to claim 1 is characterized in that: The multimodal data includes: one or more of text records, industrial images, video streams, and sensor data.

3. The anomaly detection method based on hybrid expert domain adaptive industrial big model according to claim 1 is characterized in that: The preprocessing according to the data type is specifically as follows: For time series data, the sliding window method is used for segmentation, and Fourier transform or wavelet transform is used to extract time-frequency features; For text data, perform word segmentation and part-of-speech tagging, and use the word vector model for vectorization; For image data, convolutional neural networks are used to extract visual features, and data enhancement technology is used to improve data diversity.

4. The anomaly detection method based on hybrid expert domain adaptive industrial big model according to claim 1 is characterized in that: The domain discriminator is a fully connected neural network, and the pre-training process specifically includes: using adversarial training to make the domain discriminator unable to distinguish the data distribution of the source domain and the target domain, and using a gradient reversal layer to optimize the loss function.

5. The anomaly detection method based on hybrid expert domain adaptive industrial big model according to claim 1 is characterized in that: The domain autoencoder uses a convolutional neural network or a multi-layer perceptron to perform joint training of multiple domains during a pre-training process, and ensures that the domain features of the input data are effectively extracted by minimizing the reconstruction error.

6. The anomaly detection method based on hybrid expert domain adaptive industrial big model according to claim 1 is characterized in that: The fusion of industrial knowledge with processed multimodal industrial data through explicit industrial knowledge is specifically as follows: S3.

1. Extract explicit industrial knowledge from the knowledge base of the industrial field and transform it into structured data, rule sets or construct industrial knowledge graphs; S3.2, use the multimodal encoder to extract features from different data types respectively, and map the data of multiple modes into a unified feature space. S3.

3. Contextual learning is performed by building a contextual learning architecture based on the domain knowledge graph and combining it with the domain knowledge graph.

7. The anomaly detection method based on hybrid expert domain adaptive industrial big model according to claim 1 is characterized in that: In the hybrid expert model, the expert layer for different data processing methods specifically includes the following experts: Fully connected experts: contain one or more fully connected layers; Convolution Expert: contains one or more convolutional layers; Transformer Expert: Contains the Transformer architecture with attention mechanism; Zero Expert: directly discard the input; Copy Expert: Skip the current hybrid expert layer; Constant Expert: performs constant correction and substitution on input; The expert layers in each hybrid expert module include four of them.

8. The anomaly detection method based on hybrid expert domain adaptive industrial big model according to claim 7 is characterized in that: The loss function in the training of the hybrid expert model is specifically as follows: using the cross entropy loss of the anomaly detection task as the main task loss, using the negative entropy loss as the expert regularization loss to encourage the balanced use of different experts, and for the routing mechanism, adding the L1 regularization loss as the gating loss.

9. An anomaly detection device based on a hybrid expert domain adaptive industrial large model, comprising a memory and one or more processors, wherein the memory stores executable code, characterized in that: When the processor executes the executable code, an anomaly detection method based on a hybrid expert domain adaptive industrial large model as described in any one of claims 1 to 8 is implemented.

10. A computer-readable storage medium having a program stored thereon, characterized in that: When the program is executed by a processor, an anomaly detection method based on a hybrid expert domain adaptive industrial large model as described in any one of claims 1 to 8 is implemented.

Citation Information

Patent Citations

  • Unsupervised depth field adaptation method based on distributed confrontation

    CN113011523A

  • GNN encoder and abnormal point detection method based on graph context learning

    CN113076738A

  • Intelligent building fault diagnosis method and device based on domain self-adaption

    CN116756634A

  • Fine adjustment method, system and equipment based on large language model and medium

    CN117290480A

  • Image generation method adopting generative adversarial network based on hybrid expert model

    CN118968258A

Cited By

  • Feature extraction and fusion method based on three-way hybrid coding and MOE architecture

    CN120180057A

  • A Feature Extraction and Fusion Method Based on Three-Way Hybrid Coding and MOE Architecture

    CN120180057B

  • Three-dimensional scene understanding and instruction analysis method based on multi-modal large model

    CN120849867A