Metering laboratory fault diagnosis method and device based on machine learning
By preprocessing and using machine learning to diagnose multi-source heterogeneous data from metrology laboratories, combined with resource scheduling models and manual re-inspection, the problems of low accuracy and efficiency in fault diagnosis of metrology laboratories have been solved, achieving intelligent and efficient management.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- 国网河北省电力有限公司营销服务中心
- Filing Date
- 2025-12-30
- Publication Date
- 2026-05-01
AI Technical Summary
Existing metrology laboratories rely on human experience for fault diagnosis, which leads to misjudgments and omissions. Machine learning methods have limited generalization ability, and resource scheduling lacks dynamic adjustment, resulting in low operational efficiency.
By preprocessing the initial multi-source heterogeneous data of the power metering equipment, using a pre-trained fault diagnosis model to diagnose machine faults, combining a resource scheduling model to generate a detection equipment allocation scheme, and performing manual re-inspection to generate a fault diagnosis report.
It significantly improves the accuracy and response efficiency of fault diagnosis, optimizes resource utilization, and realizes intelligent and efficient management of the metrology laboratory.
Smart Images

Figure CN121958828A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the interdisciplinary fields of metrology laboratory operation and maintenance and artificial intelligence, and in particular to a method and apparatus for fault diagnosis in metrology laboratories based on machine learning. Background Technology
[0002] Metrology laboratories are a core component of power system operation, undertaking key functions such as testing, calibration, and fault diagnosis of power metering equipment. With the continuous expansion of power system scale, increasingly complex electricity consumption environments, and the constant introduction of new power metering equipment, the industry is placing higher demands on the operational efficiency, diagnostic accuracy, and refined resource management of metrology laboratories.
[0003] Currently, fault diagnosis in metrology laboratories mainly relies on manual experience combined with historical fault records. Some laboratories are attempting to introduce electronic fault case databases for reference, while other studies are using machine learning algorithms such as support vector machines and random forests for equipment status classification or anomaly detection. Resource allocation is generally based on static rules and experience-based scheduling, lacking dynamic adjustment capabilities.
[0004] Existing technologies have significant limitations: human diagnosis is highly subjective and inconsistent, prone to misjudgment and omission, and has a delayed response; electronic systems only focus on data storage and retrieval, lacking in-depth analysis and intelligent mining capabilities; existing machine learning methods are mostly targeted at specific devices or isolated scenarios, with limited generalization capabilities; resource scheduling is difficult to adapt to real-time task loads, often resulting in a situation where idle equipment and scarce resources coexist, and overall operating efficiency needs to be improved. Summary of the Invention
[0005] This invention provides a method and apparatus for fault diagnosis in metrology laboratories based on machine learning, in order to solve the problems of low accuracy and efficiency in fault diagnosis of metrology laboratories in the prior art.
[0006] In a first aspect, embodiments of the present invention provide a method for fault diagnosis in metrology laboratories based on machine learning, comprising: The initial multi-source heterogeneous data collected from the power metering equipment is preprocessed, and the obtained target multi-source heterogeneous data is input into the pre-trained fault diagnosis model to obtain the machine fault diagnosis results. The real-time task load is obtained, and the real-time task load and the machine fault diagnosis results are input into a pre-trained resource scheduling model to obtain a detection equipment allocation scheme. Based on the machine fault diagnosis results, the power metering equipment is manually re-inspected by the detection equipment allocated by the detection equipment allocation scheme to obtain the manual fault diagnosis results. A fault diagnosis report is generated based on the machine fault diagnosis results and the manual fault diagnosis results.
[0007] Secondly, embodiments of the present invention provide an apparatus for fault diagnosis in metrology laboratories based on machine learning, comprising: The data processing module is used to preprocess the initial multi-source heterogeneous data collected from the power metering equipment; The fault diagnosis module is used to input the obtained target multi-source heterogeneous data into the pre-trained fault diagnosis model to obtain machine fault diagnosis results. The resource allocation module is used to obtain the real-time task load, input the real-time task load and the machine fault diagnosis results into the pre-trained resource scheduling model, and obtain the detection equipment allocation scheme. The acquisition module is used to perform manual re-inspection of the power metering equipment based on the machine fault diagnosis results and through the detection equipment allocation scheme of the detection equipment allocation scheme, and to acquire the manual fault diagnosis results. The report generation module is used to generate a fault diagnosis report based on the machine fault diagnosis results and the manual fault diagnosis results.
[0008] Thirdly, embodiments of the present invention provide a terminal including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the machine learning-based metrology laboratory fault diagnosis method as described in the first aspect or any possible implementation thereof.
[0009] Fourthly, embodiments of the present invention provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the steps of the machine learning-based metrology laboratory fault diagnosis method as described in the first aspect or any possible implementation thereof.
[0010] This invention provides a method and apparatus for fault diagnosis in a metrology laboratory based on machine learning. The method involves preprocessing initial multi-source heterogeneous data collected from power metering equipment and inputting the resulting target multi-source heterogeneous data into a pre-trained fault diagnosis model to obtain machine fault diagnosis results. Real-time task load is acquired, and the real-time task load and machine fault diagnosis results are input into a pre-trained resource scheduling model to obtain a detection equipment allocation scheme. Based on the machine fault diagnosis results, the power metering equipment is manually re-inspected using the detection equipment allocated according to the detection equipment allocation scheme to obtain manual fault diagnosis results. A fault diagnosis report is generated based on the machine fault diagnosis results and the manual fault diagnosis results. This invention significantly improves the accuracy and response efficiency of fault diagnosis by real-time acquisition of initial multi-source heterogeneous data and automated fault diagnosis using a machine learning model, combined with intelligent resource scheduling and human-machine collaboration mechanisms, while also optimizing the resource utilization rate of the metrology laboratory. Attached Figure Description
[0011] To more clearly illustrate the technical solutions in the embodiments of the present invention, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0012] Figure 1 This is a flowchart illustrating the implementation of a machine learning-based method for diagnosing faults in a metrology laboratory, as provided in this embodiment of the invention. Figure 2 This is a schematic diagram of the structure of the machine learning-based metrology laboratory fault diagnosis device provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the terminal provided in an embodiment of the present invention. Detailed Implementation
[0013] In the following description, specific details such as particular system architectures and techniques are set forth for illustrative purposes and not for limitation, in order to provide a thorough understanding of the embodiments of the invention. However, those skilled in the art will understand that the invention can be implemented in other embodiments without these specific details. In other instances, detailed descriptions of well-known systems, apparatuses, circuits, and methods are omitted so as not to obscure the description of the invention with unnecessary detail.
[0014] To make the objectives, technical solutions, and advantages of the present invention clearer, specific embodiments will be described below in conjunction with the accompanying drawings.
[0015] Figure 1 A flowchart illustrating the implementation of a machine learning-based method for fault diagnosis in metrology laboratories, as provided in this embodiment of the invention, is detailed below: Step 101: Preprocess the initial multi-source heterogeneous data collected from the power metering equipment, and input the obtained target multi-source heterogeneous data into the pre-trained fault diagnosis model to obtain the machine fault diagnosis result.
[0016] By deploying laboratory management terminals in the metrology laboratory, initial multi-source heterogeneous data from power metering equipment is collected in real time using communication interfaces, IoT gateways, and application programming interfaces (APIs). This initial multi-source heterogeneous data can include equipment operating status data, environmental monitoring data, and testing task data.
[0017] The communication interface is used to directly connect to power metering equipment and collect equipment operating status data. It can be a serial port, Ethernet interface, or General-Purpose Interface Bus (GPIB) interface. The IoT gateway connects to a sensor group containing temperature sensors, humidity sensors, air pressure sensors, laser scattering dust sensors, and electromagnetic interference monitors. It is used to collect environmental monitoring data through the sensor group. The API is used to access the database to read detection task data.
[0018] Equipment operating status data should include at least electrical parameter data, internal equipment status data, and process monitoring data; Electrical parameter data includes at least one of the following: voltage, current, power (active, reactive, apparent), power factor, frequency, harmonic content (total harmonic distortion (THD), each harmonic), and phase angle; internal equipment status data includes at least one of the following: temperature (e.g., chip, winding, ambient temperature and humidity), operating status signals (e.g., switch quantity, relay status, gear information), power supply status (e.g., supply voltage, battery charge), error codes, and self-test logs; process monitoring data is time sequence data collected during the execution of specified tasks (e.g., startup, creep, error testing); Environmental monitoring data should include at least climate and electromagnetic data (ambient background noise level, presence of strong interference sources nearby (such as switch operation, radio transmission)); climate data should include at least ambient temperature, ambient humidity, atmospheric pressure, and dust concentration. The testing task data should include at least the equipment number, model, specifications, manufacturer, warehousing time, testing items, standard and procedural requirements, accuracy indicators, task priority, and planned completion time of the power metering equipment to be tested.
[0019] In this embodiment, initial multi-source heterogeneous data such as equipment operating status, environmental monitoring, and testing tasks are collected in real time through communication interfaces, IoT gateways, and API interfaces, realizing comprehensive and automated management of the metrology laboratory. Its advantages include improving the real-time performance of data acquisition and system interoperability, supporting fine environmental monitoring to enhance testing reliability, optimizing resource allocation and improving work efficiency through integrated data management, and ensuring that the testing process complies with standard procedures, thus guaranteeing the comprehensiveness and compliance of the data.
[0020] In one embodiment, preprocessing of initial multi-source heterogeneous data may include data cleaning, data parsing and format standardization, and data fusion and correlation to improve data quality, consistency and usability. This effectively solves the problems of noise, format differences and information silos in the integration of initial multi-source heterogeneous data, thereby enhancing the reliability, interoperability and intelligent analysis capabilities of laboratory management terminals, and laying a solid foundation for efficient and adaptive laboratory decision-making and management.
[0021] Data cleaning includes at least missing value handling and outlier handling. Missing value handling employs forward imputation, backward imputation, or linear interpolation methods, while outlier handling employs 3D interpolation. Principles, IQR method, or filtering algorithm.
[0022] For time-series data, such as process monitoring data, forward filling, backward filling, or linear interpolation methods can be used. For non-critical data points with severe missing data, the entire record can be removed directly. For categorical data, such as certain device status signals, a separate "unknown" or "missing" category can be set.
[0023] Outlier handling refers to identifying and removing obviously unreasonable data caused by momentary sensor malfunctions, electromagnetic interference, or transmission errors. Generally, 3... The principle (Raida criterion) or IQR (interquartile range) method is used to identify and remove values that exceed the reasonable range. Thresholds are set according to the physical limits of the equipment (such as voltage cannot be negative or extremely high). Business rules are used to treat data that exceeds the threshold as invalid. Filtering algorithms are used for high-frequency collected electrical parameter time-series data (such as voltage and current). Filtering algorithms can include moving average filtering, Kalman filtering, etc., to smooth noise and preserve the true signal trend.
[0024] Data parsing and format standardization include at least protocol parsing, timestamp alignment and synchronization, and data format unification.
[0025] Protocol parsing involves analyzing the raw byte stream data collected from interfaces such as serial ports and GPIBs, according to the communication protocol documents (such as Modbus, IEC-61850, etc.) provided by the equipment manufacturer, and extracting meaningful physical quantities (such as voltage values and status codes). Timestamp alignment and synchronization require all data to have a unified and accurate timestamp. Since different data sources may have different acquisition frequencies and clocks, they need to be synchronized to the same timeline. For low-frequency data (such as environmental data), it may be necessary to align and resample with high-frequency electrical data through timestamps.
[0026] The data format is unified by converting all parsed data into a standardized data format. For example, all numeric data is converted to Float or Int; all status and text data (such as error codes and device models) are converted to String or Category; and timestamps are unified to ISO 8601 format or Unix timestamps.
[0027] Data fusion and correlation include at least spatial correlation and task correlation.
[0028] Spatial association involves linking environmental monitoring data (from a sensor group at a specific location) with the operational status data of nearby power metering equipment; for example, linking the temperature and humidity sensor data inside cabinet A with the internal temperature data of the power meter inside the cabinet. Task association involves using detection task data (such as equipment number and detection items) as key indexes to associate with equipment operating status data and environmental monitoring data; in this way, the model can know "which equipment", "which test" and "under what environmental conditions" a certain voltage time series data was generated.
[0029] This step, by collecting initial multi-source heterogeneous data from power metering equipment in real time, can comprehensively cover potential factors for fault diagnosis and avoid the limitations of a single data source. This data fusion method allows the fault diagnosis model to reason based on richer contextual information, significantly improving the accuracy and reliability of fault type, level, and location identification, and reducing the risk of misjudgment or omission.
[0030] In one embodiment, the preprocessed target multi-source heterogeneous data is input into a pre-trained fault diagnosis model for machine fault diagnosis. Before performing machine fault diagnosis, it is necessary to first construct a fault diagnosis model and then pre-train the fault diagnosis model.
[0031] Optionally, a fault diagnosis model is first constructed, comprising a first feature extraction layer, a modality fusion layer, a feature enhancement layer, and a diagnostic prediction layer. The first feature extraction layer extracts temporal feature vectors, state feature vectors, environmental feature vectors, and task feature vectors from the target multi-source heterogeneous data. The modality fusion layer aligns and fuses all feature vectors extracted by the first feature extraction layer to capture cross-modal correlation information, resulting in a first fused feature vector. The feature enhancement layer enhances the first fused feature vector through self-supervised learning and knowledge graphs, improving the robustness and semantic richness of the fused features, resulting in a second enhanced feature vector. The diagnostic prediction layer performs multi-task prediction based on the second enhanced feature vector, outputs the machine fault diagnosis result, and estimates the confidence level.
[0032] By integrating multiple heterogeneous data sources such as temporal, state, environmental, and task data through the first feature extraction layer, and using a modality fusion layer for feature alignment and fusion, this design can comprehensively capture the diverse influencing factors of equipment failure, avoiding diagnostic biases caused by single data modalities in traditional methods. For example, the temporal feature extraction module combines 1D CNN, LSTM, and self-attention mechanisms to simultaneously capture local temporal patterns and long-term dependencies. Meanwhile, the multimodal fusion module dynamically weights the importance of different modalities through cross-attention mechanisms and gated fusion, thereby improving the accuracy and robustness of diagnosis. This multimodal fusion strategy solves the problem of insufficient data utilization in existing technologies.
[0033] A hierarchical feature processing flow is adopted, including a feature extraction layer and a feature enhancement layer, in which a self-supervised contrastive learning module and a knowledge graph enhancement module work together. Self-supervised contrastive learning enhances feature representations and improves the model's generalization ability in scenarios with few samples by generating positive and negative sample pairs and using the InfoNCE loss function, without the need for a large amount of labeled data. The knowledge graph enhancement module integrates external domain knowledge and fuses semantic information through graph neural networks, making the diagnostic results more interpretable and reliable. This combination not only improves feature quality but also reduces the risk of overfitting.
[0034] The first feature extraction layer is constructed based on the temporal feature extraction module, the state feature extraction module, the environment feature extraction module, and the task feature extraction module.
[0035] The temporal feature extraction module is used to extract local temporal features from target multi-source heterogeneous data through a 1D convolutional neural network, extract sequence-dependent features from each local temporal feature through a long short-term memory network, and apply attention weights to each sequence-dependent feature through a self-attention mechanism to obtain a temporal feature vector.
[0036] The 1D convolutional neural network (CNN) takes electrical parameter data (such as voltage and current time series) and process monitoring data (time series during testing) as input. It extracts local time-series features through convolutional kernels, capturing short-term patterns (such as fluctuations and abnormal peaks). The Long Short-Term Memory (LSTM) network receives the output of the 1D CNN, processes sequence dependencies, and learns long-term time dynamics, such as equipment operating trends or precursors to failures. A self-attention mechanism applies attention weights to the LTM output, highlighting important features at key time points and enhancing sensitivity to abnormal events. The time-series feature extraction module extracts robust feature representations (time-series feature vectors) from the time-series data, suitable for high-frequency data acquisition. The 1D CNN efficiently captures local features, the LTM network handles long sequence dependencies, and the self-attention mechanism enhances the model's focus on key events; combined, they effectively handle nonlinear time-series data.
[0037] The state feature extraction module is used to convert categorical data from the device's internal state data in the target multi-source heterogeneous data into a first embedding vector through a first embedding layer. Then, a first multilayer perceptron (constructed from multiple fully connected layers (FCN)) learns the feature interactions and nonlinear relationships between the numerical data and the embedding vectors in the device's internal state data to obtain a state feature vector. The state feature extraction module is used to extract abstract features from the device's internal state data to encode the device's health status. The first embedding layer processes categorical variables to reduce the curse of dimensionality; the first multilayer perceptron captures complex feature interactions to improve representational capabilities.
[0038] The first embedding layer takes the device's internal state data (such as temperature values, status codes, and error logs) as input and converts the categorical data (such as status signals) in the device's internal state data into a dense vector representation, namely the first embedding vector. The first multilayer perceptron receives the first embedding vector and numerical data (such as temperature) and learns feature interactions and nonlinear relationships.
[0039] The environmental feature extraction module is used to smooth the environmental monitoring data in the target multi-source heterogeneous data through the moving average filtering unit, and learn the high-level representation of environmental features from the smoothed environmental monitoring data through the first fully connected network to obtain the environmental feature vector.
[0040] The moving average filtering unit takes environmental monitoring data as input, including climate data such as temperature and humidity time series, and electromagnetic data such as noise levels. It smooths the time series data to reduce noise. The first fully connected network receives the filtered data and learns a high-level representation of the environmental characteristics. The environmental feature extraction module extracts features from the environmental monitoring data to reflect the impact of external conditions on the equipment. It reduces environmental noise interference through preprocessing by the moving average filtering unit and extracts robust features through the first fully connected network, suitable for point data or low-frequency time series.
[0041] The task feature extraction module is used to convert the classification attributes of the detection task data in the target multi-source heterogeneous data into a second embedding vector through the second embedding layer, and learn the combination and relative importance of different task features (such as priority and device model) from the second embedding vector through the second fully connected network to obtain the task feature vector; The second embedding layer takes detection task data as input (e.g., structured data such as device number, model, and priority), and converts the classification attributes into second embedding vectors. The second fully connected network learns the combination and weights of task features, such as the impact of task priority on fault risk. The task feature extraction module extracts semantic features from the detection task data and associates them with the detection context. The second embedding layer processes high-cardinality classification data; the second fully connected network integrates task information, enhancing the model's context awareness.
[0042] The modality fusion layer is constructed based on the feature alignment module and the multimodal fusion module.
[0043] The feature alignment module is used to align the temporal feature vector, state feature vector, environment feature vector, and task feature vector through Deformable Convolution (DC) (e.g., features with different sampling rates), and then perform temporal alignment through Dynamic Time Warping (DTW) network to complete feature alignment (synchronizing data from different timelines to a unified timestamp).
[0044] The feature alignment module ensures that multimodal features are aligned in time and space, eliminating biases caused by differences in acquisition frequency. It uses deformable convolution to process feature dimensions and DTW to handle time series alignment, ensuring consistency in both feature and time dimensions across different modalities, thus addressing the fusion challenges posed by the heterogeneity of multi-source data. This refined alignment reduces information loss, making subsequent fusion more efficient. Furthermore, the multimodal fusion module dynamically adjusts weights based on a cross-attention mechanism, further optimizing feature interactions and improving the model's stability and adaptability in complex industrial environments.
[0045] The multimodal fusion module is used to calculate the first relevance weights of the temporal feature vector, state feature vector, environment feature vector and task feature vector after feature alignment through the cross-attention mechanism. The first gated fusion unit performs feature fusion on the temporal feature vector, state feature vector, environment feature vector and task feature vector based on each first relevance weight to obtain the first fused feature vector. The cross-attention mechanism is used to calculate the first correlation weights between aligned features, such as how environmental features affect temporal features. The first gated fusion unit dynamically adjusts the contributions of each modality feature based on the attention weights, fusing them into a unified feature representation (the first fused feature vector). The multimodal fusion module is used to fuse multimodal features, capture cross-modal interactions, and enhance information complementarity. The cross-attention mechanism focuses on relevant modalities; gated fusion uses adaptive weighting to avoid information redundancy and effectively improve the fusion effect.
[0046] The feature enhancement layer is constructed based on a self-supervised contrastive learning module and a knowledge graph enhancement module.
[0047] The self-supervised contrastive learning module constructs units using positive and negative sample pairs. It then augments the first fused feature vector by adding noise or temporal shifts to generate positive sample pairs and randomly samples negative sample pairs. A contrastive loss unit uses the InfoNCE loss function to bring positive sample pairs closer together and distance negative sample pairs further apart, resulting in a first-level augmented feature vector. This module enhances the robustness and generalization ability of features through self-supervised learning, reducing overfitting. It requires no additional labeled data and utilizes unsupervised signals to improve feature quality.
[0048] The knowledge graph enhancement module retrieves relevant entities and relationships (such as equipment failure history and environmental impact rules) from a pre-created laboratory knowledge base via the knowledge graph query unit. It then integrates semantic information by performing graph convolution operations on the first-level enhanced feature vectors, entities, and relationships using a first-level graph neural network (GNN) to obtain the second-level enhanced feature vectors. This module injects domain knowledge to enhance the semantic richness and interpretability of the features.
[0049] The diagnostic prediction layer is built upon a multi-task prediction module and an uncertainty estimation module.
[0050] The multi-task prediction module is constructed based on fault type prediction head, fault level prediction head, fault location prediction head, and handling suggestion prediction head. The fault type prediction head (Softmax output layer, predicts fault category (e.g., short circuit, overload)), the fault level prediction head (regression output layer, predicts fault severity (e.g., continuous 0-1 values or discrete levels)), the fault location prediction head (classification output layer, predicts the location of the fault (e.g., equipment component)), and the handling suggestion prediction head (sequence generation network (e.g., Transformer decoder, generates text-based handling suggestions) share the input secondary enhanced feature vector, and output the fault type, fault level, fault location, and fault handling suggestions respectively. The multi-task prediction module learns the shared secondary enhanced feature vector through multiple tasks, reducing the number of parameters, improving training efficiency, and outputting comprehensive diagnostic information.
[0051] The uncertainty estimation module is used to evaluate the confidence level of fault type, fault level, fault location, and fault handling suggestions through a Bayesian neural network, and outputs machine fault diagnosis results with the fault type, fault level, fault location, and fault handling suggestions carrying the confidence level.
[0052] The diagnostic prediction layer employs a multi-task prediction module that shares feature inputs and simultaneously outputs fault type, level, location, and recommendations. This optimizes computational efficiency and avoids resource waste caused by repetitive modeling. In addition, the uncertainty estimation module provides the confidence level of the prediction results through a Bayesian neural network, enabling users to assess the reliability of the diagnosis and thus support the decision-making process.
[0053] Optionally, after the fault diagnosis model is constructed, it is trained. When the similarity between the fault classification output by the fault diagnosis model and the real fault classification reaches the threshold, it indicates that the fault diagnosis accuracy of the fault diagnosis model has met the standard, and the training is terminated to obtain the pre-trained fault diagnosis model. Otherwise, the parameters in the fault diagnosis model are modified and retrained.
[0054] In this embodiment, by employing real-time data acquisition and preprocessing, combined with a pre-trained fault diagnosis model, the system achieves instant fault detection and automated output results, greatly shortening the response time from fault occurrence to diagnosis, avoiding the delays of traditional manual inspections, and improving the operational efficiency of the metrology laboratory. It is especially suitable for high-load or emergency scenarios, ensuring that faults can be quickly identified and handled.
[0055] In this embodiment, a hierarchical modular construction is adopted. For example, the first feature extraction layer is decomposed into multiple dedicated modules, each of which is optimized for a specific data type (such as the time series feature extraction module using LSTM to process sequence data). This design not only improves code maintainability and reusability, but also allows for flexible expansion to new modalities or tasks, reducing system upgrade costs.
[0056] In this embodiment, fault diagnosis utilizes the fusion of initial multi-source heterogeneous data such as time series, state, environment, and task, combined with feature alignment and cross-attention mechanisms, to improve the comprehensiveness and robustness of the diagnosis. Through self-supervised contrastive learning and knowledge graph enhancement, feature representation is strengthened, improving generalization ability and interpretability. Multi-task prediction and uncertainty estimation optimize efficiency and output confidence scores to support the decision-making process. The modular architecture enhances the system's scalability and applicability. Overall, it significantly improves the accuracy, efficiency, and reliability of diagnosis in metrology laboratory applications, demonstrating outstanding technological progress and practical value.
[0057] Step 102: Obtain the real-time task load, input the real-time task load and machine fault diagnosis results into the pre-trained resource scheduling model, and obtain the detection equipment allocation scheme.
[0058] Real-time task load includes at least task queue details, resource usage status, and system performance metrics.
[0059] The task queue details should include at least the task list (a collection of all assigned but not yet started and currently executing inspection tasks), the number of tasks (the total number of tasks currently queued for execution), and task attributes (details for each task). Task attributes should include at least the task priority (the importance and urgency of different tasks, the core basis for scheduling decisions), the planned completion time (the deadline for the task, used to calculate time urgency), the inspection items (the specific test type required for the task (e.g., startup test, creep test, error test, etc.), which determines the type of inspection equipment needed), the accuracy index (the required accuracy level for the task, which affects the accuracy and status requirements of the inspection equipment), and the equipment model (the model of the equipment to be inspected, which may be related to dedicated inspection fixtures or adapters). Resource occupancy status includes at least the status of the testing equipment (idle / busy: whether the equipment is currently performing a task; estimated idle time: for busy equipment, the estimated completion time of its current task; equipment health status: whether the equipment itself is faulty or needs calibration; equipment capability attributes: the testing items, measurement range, accuracy level, etc. supported by the equipment) and the status of personnel (idle / busy: whether personnel are operating the equipment; skills and qualifications: which testing equipment the personnel are qualified and capable of operating). System performance metrics should include at least average waiting time (the average waiting time of a task in the queue), system throughput (the number of tasks completed by the laboratory per unit time), and resource utilization (the percentage of time that testing equipment and personnel are busy).
[0060] In other words, real-time task load is a multi-dimensional data body that integrates task requirements, resource supply, and system performance; it tells the resource scheduling model: "How many tasks are there now, what are the requirements of these tasks, what forces and equipment we have available, and how busy the system is at present."
[0061] By collecting multi-dimensional data such as task queue details, resource occupancy status, and system performance indicators in real time, scheduling decisions are ensured to be based on a comprehensive and dynamic information foundation. Task attributes include details such as priority and planned completion time, resource status covers equipment and personnel, and performance indicators such as average waiting time and throughput are integrated. This comprehensive data integration avoids decision-making biases caused by information isolation in traditional scheduling systems, and can significantly improve scheduling accuracy and system adaptability.
[0062] In one embodiment, the resource scheduling model is constructed based on a second feature extraction layer, a feature fusion layer, a scheduling decision layer, and a scheme generation layer.
[0063] The second feature extraction layer is used to extract a second feature vector in parallel from the input real-time task load and machine fault diagnosis results. The second feature vector includes task feature vector, resource feature vector, performance feature vector, and fault feature vector. The feature fusion layer is used to dynamically fuse the extracted second feature vectors to capture the interaction relationship between tasks, resources, performance, and faults, and generate a unified second fused feature vector. The scheduling decision layer is used to generate an original scheduling decision using a policy network based on the second fused feature vector. The scheme generation layer is used to convert the original scheduling decision into a structured detection equipment allocation scheme and add scheme metadata.
[0064] The second feature extraction layer is constructed based on the task feature extraction module, resource feature extraction module, performance feature extraction module, and fault feature extraction module; The task feature extraction module is used to convert the task queue details in the real-time task load into a third embedding vector (high-dimensional vector) through the third embedding layer, capture long-term dependencies from the third embedding vector through the Transformer encoder, and reduce the dimensionality of the long-term dependencies through the feature projection unit to obtain the task feature vector. The task feature extraction module is used to extract contextual features of the task sequence from the task queue details, including task priority, planned completion time, etc. The self-attention mechanism of the Transformer encoder can capture long-term dependencies between tasks and improve the accuracy of feature representation.
[0065] The resource feature extraction module is used to convert the resource occupancy status in the real-time task load into a graph structure (nodes represent equipment or personnel, and edges represent relationships) through the graph construction unit. The module learns the node features in the graph structure through the second graph neural network, and aggregates the node features through the attention pooling unit to obtain the resource feature vector (aggregating global resource features). The resource feature extraction module is used to extract the status features of equipment and personnel from the resource occupancy status, including equipment idle status and personnel skills; GNN can model complex relationships between resources (such as equipment-personnel association) and enhance feature representation.
[0066] The performance feature extraction module is used to extract performance feature vectors from system performance indicators in real-time task loads through a second multilayer perceptron (constructed by a fully connected layer, activation function units (such as ReLU), and batch normalization units). The performance feature extraction module performs a linear transformation on the input data through a fully connected layer, then introduces nonlinearity through activation function units, and performs stable training with batch normalization units, finally outputting a performance feature vector. The performance feature extraction module is used to extract numerical features from system performance indicators, such as average latency and system throughput. MLP has a simple structure and high computational efficiency, making it suitable for processing numerical data.
[0067] The fault feature extraction module (attention-enhanced MLP) is used to extract fault feature vectors from machine fault diagnosis results through the fourth embedding layer, the third fully connected network, and the attention mechanism. The input data is first processed by the fourth embedding layer to process categorical fault data (such as fault type), and then features are extracted through the third fully connected network. The attention mechanism unit focuses on key fault information to obtain fault feature vectors. The fault feature extraction module is used to extract fault-related features from the machine fault diagnosis results, including fault level, handling suggestions, etc. The attention mechanism can highlight important fault attributes and improve feature quality.
[0068] The feature fusion layer is used to calculate the second correlation weights of the task feature vector, resource feature vector, performance feature vector and fault feature vector through the cross attention mechanism. The second gated fusion unit fuses the task feature vector, resource feature vector, performance feature vector and fault feature vector based on each second correlation weight (complementary features are enhanced by the modality contrast learning unit during the fusion process) to obtain the second fused feature vector. The feature fusion layer is used to fuse multi-source features to generate a unified second fused feature vector for scheduling decisions; the cross-attention mechanism can adaptively adjust the importance of each feature, gating fusion ensures robustness, and contrastive learning improves feature discriminativeness.
[0069] Through the cross-attention mechanism and gated fusion unit of the feature fusion layer, the correlation weights of task, resource, performance and fault feature vectors can be dynamically calculated and weighted fusion is performed. This mechanism ensures the complementarity of information from different feature sources, avoids overfitting or information redundancy, and enables the scheduling model to adaptively balance various factors, improve the robustness and flexibility of decision-making, especially in changing environments.
[0070] The scheduling decision layer (based on the Actor-Critic framework) is used to infer the action probability distribution from the second fused feature vector through the policy network unit (Actor network), evaluate the state value of the second fused feature vector through the value network unit (Critic network), and output the original scheduling decision based on the action probability distribution and state value through the task output head, carrying equipment assignment data, personnel assignment data and task scheduling data. The Actor-Critic framework can handle continuous decision problems, and multi-task output improves efficiency and adapts to real-time scheduling requirements.
[0071] The scheduling decision layer integrates policy networks and value networks, simulating a reinforcement learning framework, and outputs action probability distributions and state value assessments. This enables the decision-making process to consider not only the current state but also long-term benefits, and can automatically optimize resource allocation strategies, such as equipment assignment and task sequencing, thereby improving system throughput and resource utilization while reducing the need for human intervention.
[0072] The scheme generation layer is used to structure the original scheduling decisions and add scheme metadata through the Softmax unit (for classification output), regression unit (for numerical output), and metadata addition unit. The output carries the detection equipment allocation scheme carrying equipment assignment data (clearly specifying which specific detection equipment (or which group of equipment) is assigned to re-inspect the faulty power metering equipment, the selection criteria including: equipment capability matching: the selected equipment must be able to complete the detection items that need to be re-inspected as indicated in the fault diagnosis results; equipment accuracy matching: the accuracy level of the equipment must meet the accuracy index requirements of the original task of the equipment to be inspected; equipment status optimization: priority is given to idle or soon-to-be-idle equipment with good health status), personnel assignment data, task scheduling data, and scheme metadata. The scheduling decision layer outputs categorized equipment and personnel assignment data through the Softmax unit, generates task scheduling numerical outputs through the regression unit, and adds scheme IDs, generation timestamps, and expected targets through the metadata addition unit. The scheme generation layer transforms the raw scheduling decisions into structured detection equipment allocation schemes, ensuring standardized output formats for easy execution and tracking.
[0073] The scheme generation layer uses Softmax units, regression units, and other methods to structure the original scheduling decisions and add scheme metadata such as scheme ID and generation timestamp, making the output scheme standardized and traceable. This design enhances the readability and usability of the scheme, facilitates subsequent monitoring, auditing and optimization, and meets the compliance and maintainability requirements in industrial applications.
[0074] Task scheduling data should include at least the execution order (determining the insertion position of the re-inspection task in the existing task queue; high-priority (e.g., high fault level) re-inspection tasks may be scheduled first and executed in the queue), start time (the specific start time of the planned re-inspection task), and estimated time (estimated time required to complete the re-inspection task based on historical data or models); scheme metadata should include at least the scheme ID (a unique identifier for allocating the scheme for subsequent tracking), generation timestamp (the time when the scheme was generated), and expected goals (the goals to be achieved, such as "complete as quickly as possible", "minimize the impact on the original task", "highest accuracy re-inspection", etc.)).
[0075] In this embodiment, by integrating multi-dimensional real-time data such as task queues, resource status, performance indicators, and fault information, and using technologies such as Transformer and graph neural networks for deep feature extraction and intelligent fusion, a dynamic and adaptive resource scheduling model is constructed. This model employs a reinforcement learning framework for optimization decision-making, efficiently generating structured and traceable scheduling schemes, significantly improving system throughput, resource utilization, and response speed. It also possesses excellent robustness and scalability, making it suitable for efficient task management in metrology laboratories. By introducing the resource scheduling model and combining real-time task load with machine fault diagnosis results, a dynamic allocation scheme for testing equipment is generated. This optimized scheduling not only considers fault priority but also takes into account the overall workload of the metrology laboratory, thereby rationally allocating human and equipment resources, reducing idle or overused resources, and improving the productivity and cost-effectiveness of the metrology laboratory.
[0076] Step 103: Based on the machine fault diagnosis results, the power metering equipment is manually re-inspected by the detection equipment allocated by the detection equipment allocation scheme to obtain the manual fault diagnosis results.
[0077] In this step, a manual review process is introduced on the basis of machine diagnosis. By comparing the machine and manual fault diagnosis results, a comprehensive report is generated. This human-machine collaborative design not only leverages the efficiency of machine learning but also retains the judgment of human experts, reducing the false alarm rate of the automated system and enhancing the credibility and practicality of the diagnostic results.
[0078] The results of manual fault diagnosis include the following: verification and confirmation information of machine diagnosis results, in-depth diagnostic information not covered by the machine, detailed records and verification of the handling process, additional information on the re-inspection process and environment, and an index of evidentiary materials.
[0079] Verification and confirmation information for machine diagnostic results includes: consistency determination. This involves clearly recording whether the manual re-inspection conclusion is consistent with the machine fault diagnosis result. If inconsistent, the final and accurate fault type, fault level, and fault location confirmed by the human investigator must be used to correct the original fault type, fault level, and fault location.
[0080] In-depth diagnostic information not covered by the machine includes: root cause analysis, fault mechanism description, and discovery of latent faults.
[0081] Machines may diagnose phenomena, while humans can analyze the root cause; for example, a machine might diagnose "error exceeding tolerance," while human analysis might conclude: "A sudden voltage surge in the laboratory yesterday damaged the equipment's reference source, leading to the error exceeding tolerance." Fault mechanism descriptions are textual descriptions and analyses of how a fault occurs and develops. Latent fault discovery records other potential faults or performance degradation points newly discovered during re-inspection that were not detected by the machine.
[0082] Detailed records and verification of the handling process include: adopted handling recommendations, actual handling measures taken, and verification of handling effects. Adopted handling recommendations: Record which part (if any) of the machine-generated handling recommendations was ultimately executed. Actual handling measures taken: Record in detail the specific operational steps actually taken manually, such as repair, calibration, and component replacement (e.g., "Replaced the internal fuse F1 of the FLUKE 741B calibrator"). Verification of handling effects: Record the results of retesting after taking handling measures to prove that the fault has been eliminated (e.g., "After repair, error testing was performed again; the error at each point was within ±0.1%, which is acceptable").
[0083] Additional information regarding the re-inspection process and environment includes: the re-inspection equipment and tools used, environmental conditions during the re-inspection, and personnel information. Re-inspection equipment and tools used: Record the specific testing equipment, calibrators, software tools, and their serial numbers used in the re-inspection (e.g., "Voltage measurement was performed using an Agilent 3458A digital multimeter (serial number: LAB-008)"). Environmental conditions during the re-inspection: Record any transient environmental conditions observed during the re-inspection that may be related to the fault (e.g., "During the re-inspection, it was found that the air conditioning vent in the computer room was directly facing this equipment, which may have caused excessively rapid local temperature rise"). Personnel information: The ID or name of the technician performing the re-inspection, and the timestamp of the re-inspection, which increases the traceability and accountability of the results.
[0084] The evidentiary materials index includes: data snapshots and screenshots, and measurement data records. Data snapshots and screenshots are the storage paths or unique identifiers for visual evidence such as waveforms, spectrum diagrams, error code screenshots, and test software interface screenshots saved during the re-examination process. Measurement data records are the sequences of key measurement values recorded during the re-examination process.
[0085] Step 104: Generate a fault diagnosis report based on the machine fault diagnosis results and the manual fault diagnosis results.
[0086] The laboratory management terminal generates fault diagnosis reports based on machine fault diagnosis results and manual fault diagnosis results, and can display the fault diagnosis reports in real time on the screen.
[0087] In one embodiment, after generating a fault diagnosis report based on machine fault diagnosis results and manual fault diagnosis results, the process may further include: Record the execution effect data of the testing equipment allocation plan; The target multi-source heterogeneous data, real-time task load, detection equipment allocation scheme, fault diagnosis report and execution effect data are stored in the laboratory knowledge base, and the laboratory knowledge base is called to continuously optimize the fault diagnosis model and resource scheduling model.
[0088] The execution effect data of the allocation plan should include at least the task execution efficiency data, task execution quality data, resource consumption data, and cost data.
[0089] Task execution efficiency data can include actual completion time, total task processing time, and resource idle rate. Task execution quality data can include re-inspection accuracy and first-pass yield; re-inspection accuracy is a comparison of the consistency between manual re-inspection results and machine diagnostic results. If there is a discrepancy, the finally confirmed correct result is recorded. The first-pass yield indicates whether the allocated equipment can efficiently complete the re-inspection in one go without needing to change equipment midway. Resource consumption data includes actual working hours and spare parts consumption. Cost data, i.e., comprehensive cost estimation, includes the labor costs, equipment depreciation costs, and consumable costs involved in the execution of this scheduling plan.
[0090] In one embodiment, the laboratory management terminal stores target multi-source heterogeneous data, real-time task load, testing equipment allocation scheme, fault diagnosis report, and allocation scheme execution effect data into a pre-created laboratory knowledge base, and continuously optimizes the fault diagnosis model and resource scheduling model by calling the laboratory knowledge base through a federated learning mechanism.
[0091] In one embodiment, continuously optimizing the fault diagnosis model and resource scheduling model by invoking a laboratory knowledge base may include: Based on a preset optimization cycle, incremental knowledge is read from the laboratory knowledge base to construct an incremental dataset. This incremental dataset is then used to train the current fault diagnosis model and the current resource scheduling model. The first local model parameters of the current fault diagnosis model and the second local model parameters of the current resource scheduling model are obtained. Differential privacy noise is added to both the first and second local model parameters to obtain the corresponding first and second desensitized model parameters. These parameters are then uploaded to a server, where they are aggregated to obtain the first and second global model parameters. Based on these parameters, the current fault diagnosis model and the current resource scheduling model are updated, and the updated models are used for subsequent machine fault diagnosis.
[0092] Optionally, each metrology laboratory's laboratory management terminal can read incremental knowledge from the local laboratory knowledge base to construct an incremental dataset based on a preset optimization cycle. The fault diagnosis model and resource scheduling model can be trained using the incremental dataset. During the training process, the hyperparameters of the fault diagnosis model and resource scheduling model are continuously adjusted until the performance indicators of the fault diagnosis model and resource scheduling model meet the preset indicator thresholds.
[0093] The server aggregates the first de-identified model parameters of each metrology laboratory to obtain the first global model parameters, and aggregates the second de-identified model parameters of each metrology laboratory to obtain the second global model parameters. The first global model parameters and the second global model parameters are then sent to the management terminals of each laboratory. Based on the received first global model parameters and second global model parameters, the laboratory management terminals iteratively update the fault diagnosis model and the resource scheduling model, respectively.
[0094] The laboratory knowledge base also stores historical data on models and optimizations, full lifecycle data of equipment and assets, domain knowledge and rule bases (standards and procedures base: storing relevant national / industry verification procedures, technical standards, safety specifications, operating procedures (SOPs) and other structured or unstructured documents. Models can use this to verify the compliance of their diagnoses and recommendations; expert experience rules: transforming successful and reusable experiences from past manual diagnoses into structured rules (such as "IF condition THEN conclusion"), which can be used for model reference or direct use, enhancing the interpretability of the model; typical failure case base: storing details of typical and complex failure cases in history, including phenomena, data, analysis process, handling methods and final conclusions, serving as a valuable resource for model training and manual reference), system logs and audit trail data, historical performance indicators and KPI data, external knowledge and environmental context data, and is dynamically updated based on a preset update mechanism.
[0095] By leveraging federated learning mechanisms and laboratory knowledge bases, historical data and execution results can be stored, and fault diagnosis models and resource scheduling models can be continuously optimized. This self-evolutionary function enables the system to adapt to equipment aging, environmental changes, or new fault modes over time, maintaining long-term effectiveness and advancement, and reducing the need for manual model updates.
[0096] In this embodiment of the invention, a complete closed-loop process is provided from data acquisition, processing, diagnosis to report generation and storage, realizing comprehensive digital management of fault diagnosis. This integrated design simplifies the operation and maintenance of metrology laboratories, displays reports in real time on the screen for easy monitoring and decision-making, and records execution effect data to support subsequent analysis, thereby enhancing overall maintainability and transparency.
[0097] By combining local incremental learning in each metrology laboratory with global aggregation on the server side, continuous optimization of the fault diagnosis model and resource scheduling model is achieved. Its advantages include significantly improving the accuracy and adaptability of the model, ensuring that the system can dynamically respond to environmental changes through regular incremental training and hyperparameter tuning; at the same time, the introduction of differential privacy noise enhances data security, effectively protecting sensitive information without affecting collaborative learning; the distributed architecture also improves learning efficiency and scalability, reduces communication burden, and enhances model robustness through adaptive mechanisms, ultimately improving the maintainability and cross-laboratory collaboration capabilities of the overall system, making it suitable for efficient and secure metrology network management.
[0098] This invention provides a machine learning-based method for fault diagnosis in a metering laboratory. The method involves preprocessing initial multi-source heterogeneous data collected from power metering equipment and inputting the resulting target multi-source heterogeneous data into a pre-trained fault diagnosis model to obtain machine fault diagnosis results. Real-time task load is acquired, and the real-time task load and machine fault diagnosis results are input into a pre-trained resource scheduling model to obtain a detection equipment allocation scheme. Based on the machine fault diagnosis results, the power metering equipment is manually re-inspected using the detection equipment allocated according to the detection equipment allocation scheme to obtain manual fault diagnosis results. A fault diagnosis report is generated based on the machine fault diagnosis results and the manual fault diagnosis results. This invention significantly improves the accuracy and response efficiency of fault diagnosis by real-time acquisition of initial multi-source heterogeneous data and automated fault diagnosis using a machine learning model, combined with intelligent resource scheduling and human-machine collaboration mechanisms, while optimizing the utilization rate of metering laboratory resources. Furthermore, the human-machine interaction design ensures the reliability of the results, while the continuous optimization function based on federated learning and a laboratory knowledge base enables the system to adapt to dynamic environments, ultimately forming an end-to-end integrated solution that enhances the convenience and long-term effectiveness of operation and maintenance.
[0099] In this embodiment of the invention, a manual re-inspection step is introduced after automatic diagnosis, and a fault diagnosis report is generated by comparing the machine and manual results, forming a closed loop of "diagnosis-re-inspection-feedback". This design not only ensures the reliability of the diagnosis results, but also provides real feedback for model optimization by recording the execution effect data of the allocation scheme. In addition, the model is continuously optimized by calling the laboratory knowledge base using the federated learning mechanism, avoiding the risk of centralized data storage. At the same time, data security is protected by differential privacy technology, enabling the system to adapt to environmental changes and maintain high performance in the long term, demonstrating strong self-learning ability and practicality.
[0100] By acquiring and processing multi-source heterogeneous data in real time and combining it with advanced machine learning models (such as multimodal fusion diagnosis and dynamic resource scheduling), the system achieves accurate and automated fault diagnosis. Its innovation lies in the closed-loop verification mechanism of human-machine collaboration and the continuous optimization capability based on federated learning, which not only ensures the reliability of diagnostic results but also achieves secure knowledge sharing through privacy protection technology. The overall design is highly modular and scalable, which can significantly improve the operational efficiency, resource utilization, and scientific decision-making of metrology laboratories.
[0101] It should be understood that the sequence number of each step in the above embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.
[0102] The following are device embodiments of the present invention. For details not described in detail, please refer to the corresponding method embodiments described above.
[0103] Figure 2 The diagram shows a structural schematic of a machine learning-based fault diagnosis device for metrology laboratories according to an embodiment of the present invention. For ease of explanation, only the parts relevant to the embodiment of the present invention are shown, and are described in detail below: like Figure 2 As shown, the device for fault diagnosis in a metrology laboratory based on machine learning includes: a data processing module 21, a fault diagnosis module 22, a resource allocation module 23, an acquisition module 24, and a report generation module 25.
[0104] Data processing module 21 is used to preprocess the initial multi-source heterogeneous data collected from the power metering equipment; The fault diagnosis module 22 is used to input the obtained target multi-source heterogeneous data into the pre-trained fault diagnosis model to obtain machine fault diagnosis results. The resource allocation module 23 is used to obtain the real-time task load, input the real-time task load and machine fault diagnosis results into the pre-trained resource scheduling model, and obtain the detection equipment allocation scheme. The acquisition module 24 is used to perform manual re-inspection of the power metering equipment based on the machine fault diagnosis results and the detection equipment allocated by the detection equipment allocation scheme, and to obtain the manual fault diagnosis results. The report generation module 25 is used to generate a fault diagnosis report based on the machine fault diagnosis results and the manual fault diagnosis results.
[0105] In one possible implementation, after generating a fault diagnosis report based on the machine fault diagnosis results and the manual fault diagnosis results, the data processing module 21 is further configured to: Record the execution effect data of the testing equipment allocation plan; The target multi-source heterogeneous data, real-time task load, detection equipment allocation scheme, fault diagnosis report and execution effect data are stored in the laboratory knowledge base, and the laboratory knowledge base is called to continuously optimize the fault diagnosis model and resource scheduling model.
[0106] In one possible implementation, when the data processing module 21 calls the laboratory knowledge base to continuously optimize the fault diagnosis model and resource scheduling model, it is used for: Based on a preset optimization cycle, incremental knowledge is read from the laboratory knowledge base to construct an incremental dataset, and the current fault diagnosis model and the current resource scheduling model are trained using the incremental dataset. Obtain the first local model parameters of the current fault diagnosis model and the second local model parameters of the current resource scheduling model. Add differential privacy noise to the first local model parameters and the second local model parameters respectively to obtain the first desensitized model parameters and the second desensitized model parameters. The first desensitization model parameters and the second desensitization model parameters are uploaded to the server, and the server aggregates them to obtain the first global model parameters and the second global model parameters. Based on the first global model parameters and the second global model parameters, the current fault diagnosis model and the current resource scheduling model are updated respectively, and the updated fault diagnosis model and resource scheduling model are used for subsequent machine fault diagnosis.
[0107] In one possible implementation, the fault diagnosis model includes a first feature extraction layer, a modality fusion layer, a feature enhancement layer, and a diagnostic prediction layer. The first feature extraction layer is used to extract a first feature vector from the target multi-source heterogeneous data. The first feature vector includes a time-series feature vector, a state feature vector, an environmental feature vector, and a task feature vector. The modality fusion layer is used to align and fuse the first feature vector to obtain the first fused feature vector; The feature enhancement layer is used to enhance the first fused feature vector through self-supervised learning and knowledge graph to obtain the second enhanced feature vector. The diagnostic prediction layer is used for multi-task prediction based on the second enhanced feature vector and outputs machine fault diagnosis results.
[0108] In one possible implementation, the modal fusion layer includes: a feature alignment module and a multimodal fusion module; The feature alignment module is used to align the feature dimensions through deformable convolutions and to align the temporal dimension through a temporally dynamic alignment network. The multimodal fusion module is used to calculate the first correlation weights of the temporal feature vector, state feature vector, environment feature vector and task feature vector through the cross-attention mechanism, and perform feature fusion through the first gating fusion unit based on the first correlation weights to obtain the first fused feature vector.
[0109] In one possible implementation, the feature enhancement layer includes: a self-supervised contrastive learning module and a knowledge graph enhancement module; The self-supervised contrastive learning module is used to augment the first fused feature vector, generate positive sample pairs, and randomly sample to generate negative sample pairs. The InfoNCE loss function is used to bring the positive sample pairs closer and push the negative sample pairs further apart to obtain the first-level augmented feature vector. The knowledge graph enhancement module is used to retrieve relevant entities and relationships from the laboratory knowledge base. It performs graph convolution operations on the first-level enhanced feature vector, entities, and relationships through the first graph neural network to obtain the second enhanced feature vector.
[0110] In one possible implementation, the resource scheduling model includes: a second feature extraction layer, a feature fusion layer, a scheduling decision layer, and a scheme generation layer; The second feature extraction layer is used to extract a second feature vector in parallel from the input real-time task load and machine fault diagnosis results. The second feature vector includes task feature vector, resource feature vector, performance feature vector and fault feature vector. The feature fusion layer is used to dynamically fuse the second feature vector to generate a second fused feature vector; The scheduling decision layer is used to generate the original scheduling decision based on the second fusion feature vector; The scheme generation layer is used to convert the original scheduling decisions into structured detection equipment allocation schemes and add scheme metadata.
[0111] The above embodiments provide a machine learning-based fault diagnosis device for a metrology laboratory. A data processing module preprocesses the initial multi-source heterogeneous data collected from power metering equipment. A fault diagnosis module inputs the obtained target multi-source heterogeneous data into a pre-trained fault diagnosis model to obtain machine fault diagnosis results. A resource allocation module acquires the real-time task load and inputs the real-time task load and machine fault diagnosis results into a pre-trained resource scheduling model to obtain a detection equipment allocation scheme. Based on the machine fault diagnosis results, the power metering equipment is manually re-inspected using the detection equipment allocated according to the detection equipment allocation scheme, and an acquisition module obtains the manual fault diagnosis results. A report generation module generates a fault diagnosis report based on the machine fault diagnosis results and the manual fault diagnosis results. This invention significantly improves the accuracy and response efficiency of fault diagnosis by real-time acquisition of initial multi-source heterogeneous data and automated fault diagnosis using a machine learning model, combined with intelligent resource scheduling and human-machine collaboration mechanisms, while optimizing the utilization rate of metrology laboratory resources. Furthermore, the human-machine interaction design ensures the reliability of the results, while the continuous optimization function based on federated learning and a laboratory knowledge base enables the system to adapt to dynamic environments, ultimately forming an end-to-end integrated solution that enhances the convenience of operation and maintenance and its long-term effectiveness.
[0112] Figure 3 This is a schematic diagram of a terminal provided in an embodiment of the present invention. Figure 3 As shown, the terminal 3 in this embodiment includes a processor 30, a memory 31, and a computer program 32 stored in the memory 31 and executable on the processor 30. When the processor 30 executes the computer program 32, it implements the steps in the various machine learning-based metrology laboratory fault diagnosis method embodiments described above, for example... Figure 1 Steps 101 to 104 are shown. Alternatively, when processor 30 executes computer program 32, it implements the functions of each module / unit in the above-described device embodiments, for example... Figure 2 The functions of each module / unit are shown.
[0113] For example, computer program 32 can be divided into one or more modules / units, one or more of which are stored in memory 31 and executed by processor 30 to complete the present invention. One or more modules / units can be a series of computer program instruction segments capable of performing a specific function, which describe the execution process of computer program 32 in terminal 3. For example, computer program 32 can be divided into... Figure 2 The modules / units shown are shown.
[0114] Terminal 3 may include, but is not limited to, processor 30 and memory 31. Those skilled in the art will understand that... Figure 3This is merely an example of terminal 3 and does not constitute a limitation on terminal 3. It may include more or fewer components than shown, or combine certain components, or different components. For example, the terminal may also include input / output devices, network access devices, buses, etc.
[0115] The processor 30 may be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor may be a microprocessor or any conventional processor.
[0116] The memory 31 can be an internal storage unit of the terminal 3, such as a hard disk or RAM of the terminal 3. The memory 31 can also be an external storage device of the terminal 3, such as a plug-in hard disk, Smart MediaCard (SMC), Secure Digital (SD) card, or Flash Card equipped on the terminal 3. Furthermore, the memory 31 can include both internal and external storage units of the terminal 3. The memory 31 is used to store computer programs and other programs and data required by the terminal. The memory 31 can also be used to temporarily store data that has been output or will be output.
[0117] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this application. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.
[0118] In the above embodiments, the descriptions of each embodiment have different focuses. For parts that are not described in detail or recorded in a certain embodiment, please refer to the relevant descriptions of other embodiments.
[0119] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementations should not be considered beyond the scope of this invention.
[0120] In the embodiments provided by this invention, it should be understood that the disclosed devices / terminals and methods can be implemented in other ways. For example, the device / terminal embodiments described above are merely illustrative. For instance, the division of modules or units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between devices or units may be electrical, mechanical, or other forms.
[0121] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0122] Furthermore, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit.
[0123] If integrated modules / units are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the methods of the above embodiments of the present invention can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various machine learning-based metrology laboratory fault diagnosis method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying computer program code, recording media, USB flash drives, portable hard drives, magnetic disks, optical disks, computer memory, read-only memory (ROM), random access memory (RAM), electrical carrier signals, telecommunication signals, and software distribution media, etc.
[0124] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for fault diagnosis in metrology laboratories based on machine learning, characterized in that, include: The initial multi-source heterogeneous data collected from the power metering equipment is preprocessed, and the obtained target multi-source heterogeneous data is input into the pre-trained fault diagnosis model to obtain the machine fault diagnosis results. The real-time task load is obtained, and the real-time task load and the machine fault diagnosis results are input into a pre-trained resource scheduling model to obtain a detection equipment allocation scheme. Based on the machine fault diagnosis results, the power metering equipment is manually re-inspected by the detection equipment allocated by the detection equipment allocation scheme to obtain the manual fault diagnosis results. A fault diagnosis report is generated based on the machine fault diagnosis results and the manual fault diagnosis results.
2. The method for fault diagnosis in metrology laboratories based on machine learning according to claim 1, characterized in that, After generating a fault diagnosis report based on the machine fault diagnosis results and the manual fault diagnosis results, the process also includes: Record the execution effect data of the aforementioned detection equipment allocation scheme; The target multi-source heterogeneous data, the real-time task load, the detection equipment allocation scheme, the fault diagnosis report, and the execution effect data are stored in the laboratory knowledge base, and the laboratory knowledge base is called to continuously optimize the fault diagnosis model and the resource scheduling model.
3. The method for fault diagnosis in metrology laboratories based on machine learning according to claim 2, characterized in that, The step of continuously optimizing the fault diagnosis model and the resource scheduling model by calling the laboratory knowledge base includes: Based on a preset optimization cycle, incremental knowledge is read from the laboratory knowledge base to construct an incremental dataset, and the current fault diagnosis model and the current resource scheduling model are trained using the incremental dataset. Obtain the first local model parameters of the current fault diagnosis model and the second local model parameters of the current resource scheduling model, and add differential privacy noise to the first local model parameters and the second local model parameters respectively to obtain the first desensitized model parameters and the second desensitized model parameters. The first desensitization model parameters and the second desensitization model parameters are uploaded to the server, and the server aggregates them to obtain the first global model parameters and the second global model parameters. Based on the first global model parameters and the second global model parameters, the current fault diagnosis model and the current resource scheduling model are updated respectively, and the updated fault diagnosis model and resource scheduling model are used for subsequent machine fault diagnosis.
4. The method for fault diagnosis in metrology laboratories based on machine learning according to any one of claims 1-3, characterized in that, The fault diagnosis model includes a first feature extraction layer, a modality fusion layer, a feature enhancement layer, and a diagnosis prediction layer; The first feature extraction layer is used to extract a first feature vector from the target multi-source heterogeneous data. The first feature vector includes a time-series feature vector, a state feature vector, an environmental feature vector, and a task feature vector. The modality fusion layer is used to align and fuse the first feature vector to obtain a first fused feature vector; The feature enhancement layer is used to enhance the first fused feature vector through self-supervised learning and knowledge graph to obtain a second enhanced feature vector. The diagnostic prediction layer is used to perform multi-task prediction based on the second enhanced feature vector and output machine fault diagnosis results.
5. The method for fault diagnosis in metrology laboratories based on machine learning according to claim 4, characterized in that, The modality fusion layer includes: a feature alignment module and a multimodal fusion module; The feature alignment module is used to perform feature dimension alignment through deformable convolution and time dimension alignment through a time dynamic alignment network. The multimodal fusion module is used to calculate the first correlation weights of the temporal feature vector, the state feature vector, the environment feature vector, and the task feature vector through a cross-attention mechanism, and perform feature fusion through a first gating fusion unit based on the first correlation weights to obtain the first fused feature vector.
6. The method for fault diagnosis in metrology laboratories based on machine learning according to claim 5, characterized in that, The feature enhancement layer includes: a self-supervised contrastive learning module and a knowledge graph enhancement module; The self-supervised contrastive learning module is used to perform data augmentation on the first fused feature vector, generate positive sample pairs, and randomly sample to generate negative sample pairs. The InfoNCE loss function is used to bring the positive sample pairs closer and push the negative sample pairs further apart to obtain a first-level augmented feature vector. The knowledge graph enhancement module is used to retrieve relevant entities and relationships from the laboratory knowledge base, and to perform graph convolution operations on the first-level enhanced feature vector, the entities, and the relationships through a first graph neural network to obtain a second enhanced feature vector.
7. The method for fault diagnosis in metrology laboratories based on machine learning according to claim 6, characterized in that, The resource scheduling model includes: a second feature extraction layer, a feature fusion layer, a scheduling decision layer, and a scheme generation layer; The second feature extraction layer is used to extract a second feature vector in parallel from the input real-time task load and the machine fault diagnosis results. The second feature vector includes a task feature vector, a resource feature vector, a performance feature vector, and a fault feature vector. The feature fusion layer is used to dynamically fuse the second feature vector to generate a second fused feature vector. The scheduling decision layer is used to generate an original scheduling decision based on the second fused feature vector; The scheme generation layer is used to convert the original scheduling decision into a structured detection equipment allocation scheme and add scheme metadata.
8. A device for fault diagnosis in a metrology laboratory based on machine learning, characterized in that, include: The data processing module is used to preprocess the initial multi-source heterogeneous data collected from the power metering equipment; The fault diagnosis module is used to input the obtained target multi-source heterogeneous data into the pre-trained fault diagnosis model to obtain machine fault diagnosis results. The resource allocation module is used to obtain the real-time task load, input the real-time task load and the machine fault diagnosis results into the pre-trained resource scheduling model, and obtain the detection equipment allocation scheme. The acquisition module is used to perform manual re-inspection of the power metering equipment based on the machine fault diagnosis results and through the detection equipment allocation scheme of the detection equipment allocation scheme, and to acquire the manual fault diagnosis results. The report generation module is used to generate a fault diagnosis report based on the machine fault diagnosis results and the manual fault diagnosis results.
9. A terminal, comprising a memory and a processor, the memory for storing a computer program, the processor for calling and running the computer program stored in the memory, characterized in that, When the processor executes the computer program, it implements the steps of the machine learning-based method for diagnosing faults in a metrology laboratory as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by a processor, it implements the steps of the machine learning-based method for diagnosing faults in a metrology laboratory as described in any one of claims 1 to 7.
Citation Information
Patent Citations
A packaging system
IE61850B1