Power operation and maintenance risk prediction method based on multi-modal fusion large model
By collecting multi-source heterogeneous data to form a standardized multimodal dataset, generating spatiotemporal semantic fusion feature vectors and inputting them into a large multimodal fusion model, the problem of insufficient utilization of multi-source data in existing technologies is solved, efficient and automated prediction of power operation and maintenance risks is achieved, the timeliness and accuracy of fault warnings are improved, and dependence on manual experience is reduced.
Patent Information
- Application Number
- CN202510882655.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-27
- Publication Date
- 2025-10-10
AI Technical Summary
Existing technologies lack the ability to comprehensively utilize multi-source heterogeneous data in power operation and maintenance risk prediction, and it is difficult to efficiently capture the spatiotemporal correlations and semantic characteristics behind the data. As a result, the timeliness and accuracy of risk predictions are unable to meet the operation and maintenance needs of modern power grids. In addition, reliance on manual experience analysis is inefficient, has a high missed detection rate, and is difficult to detect early hidden faults.
Collect multi-source heterogeneous data to form a normalized multimodal dataset, generate spatiotemporal semantic fusion feature vectors, and input them into the multimodal fusion model for forward transmission. Utilize the learning ability of the large model and the implicit domain knowledge graph to achieve cross-modal feature fusion and risk reasoning, automatically associate the spatial fault propagation paths of equipment, and improve the timeliness and accuracy of fault warnings.
It realizes real-time automated analysis of massive data, reduces dependence on the experience of senior operation and maintenance personnel, significantly improves operation and maintenance efficiency, reduces labor costs, can automatically associate equipment topology with time series, improves the prediction and generalization capabilities of new faults, and generates risk transmission path explanations that meet power industry standards.
Smart Images

Figure CN120767799A_ABST
Abstract
Description
Technical Field
[0001] The present invention belongs to the field of smart grid processing, and specifically relates to a power operation and maintenance risk prediction method based on a multimodal fusion large model. Background Art
[0002] Against the backdrop of large-scale interconnection and complex operation of smart grids, operating status monitoring of power equipment and operation and maintenance risk prediction have become core links in ensuring the safe and stable operation of power grids. With the popularization of sensor technology and the Internet of Things, the multi-source heterogeneous data generated in power systems (such as real-time equipment operating values, inspection images, operation and maintenance log texts, etc.) has grown exponentially. These data contain key information about the health status of equipment. Accurate power operation and maintenance risk prediction can identify potential equipment failures (such as insulation aging, joint overheating, etc.) in advance, avoid sudden power outages, reduce operation and maintenance costs, and improve power supply reliability. However, existing technologies lack the ability to comprehensively utilize multi-source data, making it difficult to efficiently capture the spatiotemporal correlations and semantic features behind the data, resulting in the timeliness and accuracy of risk predictions being unable to meet the operation and maintenance needs of modern power grids.
[0003] Current power operation and maintenance risk prediction schemes primarily rely on a combination of manual experience and traditional algorithms. This model has significant flaws. First, manual analysis relies on the accumulated experience of experienced operation and maintenance personnel. Faced with massive amounts of multimodal data, it is subject to strong subjectivity, low analytical efficiency, and a high rate of missed detections. This makes it particularly difficult to detect early hidden faults (such as weak discharges and progressive insulation degradation). Second, traditional algorithms (such as those based on statistical models) have limited ability to integrate multi-source heterogeneous data and cannot effectively address the semantic alignment and temporal and spatial consistency of cross-modal data such as numerical, image, and text. For example, traditional methods struggle to correlate fault transmission paths within the device topology space (such as the spatial correlation between main transformer overload and adjacent busbar current anomalies). They also fail to embed domain knowledge rules related to the "overload → temperature rise → insulation aging" domain. This results in weak prediction capabilities for new or rare fault modes, and the prediction results lack interpretability, making them difficult to directly apply to standardized power operation and maintenance decision-making processes. Summary of the Invention
[0004] Based on the above problems, an embodiment of the present application provides a power operation and maintenance risk prediction method based on a multimodal fusion large model to solve the problems existing in the above-mentioned prior art.
[0005] The present application provides a method for predicting electric power operation and maintenance risks based on a multimodal fusion large model. The method includes:
[0006] Step 1: Collect multi-source heterogeneous data related to power operation and maintenance to form a normalized multimodal dataset;
[0007] Step 2: Generate spatiotemporal semantic fusion feature vectors based on the normalized multimodal dataset;
[0008] Step 3: Input the spatiotemporal semantic fusion feature vector into the multimodal fusion model for forward transmission to predict the power operation and maintenance risk and obtain the power operation and maintenance risk prediction result.
[0009] This application provides a power operation and maintenance risk prediction method based on a multimodal fusion large model, which has the following technical advantages:
[0010] The power operation and maintenance risk prediction method based on a multimodal fusion large model provided in this application has technical benefits that directly address the pain points of background technologies, as follows:
[0011] This application uses models to replace manual labor in the collection, normalization, feature fusion, and risk prediction of multimodal data, reducing reliance on the experience of experienced operations and maintenance personnel. This approach enables real-time automated analysis of massive amounts of data, avoids human under-detection, significantly improves operations and maintenance efficiency, and reduces labor costs, resolving the issues of high labor input and low efficiency in traditional operations and maintenance models. BRIEF DESCRIPTION OF THE DRAWINGS
[0012] Figure 1 This is a structural diagram of a method for predicting power operation and maintenance risks based on a multimodal fusion large model in an embodiment of the present application. DETAILED DESCRIPTION
[0013] Figure 1 This is a schematic diagram of the structure of a power operation and maintenance risk prediction method based on a multi-modal fusion model according to an embodiment of the present application. Figure 1 As shown, it includes:
[0014] Step 1: Collect multi-source heterogeneous data related to power operation and maintenance to form a normalized multimodal dataset;
[0015] Step 2: Generate spatiotemporal semantic fusion feature vectors based on the normalized multimodal dataset;
[0016] Step 3: Input the spatiotemporal semantic fusion feature vector into the multimodal fusion model for forward transmission to predict the power operation and maintenance risk and obtain the power operation and maintenance risk prediction result.
[0017] In summary, this application collects multi-source heterogeneous data related to power operation and maintenance to form a normalized multimodal data set, and uniformly normalizes the scattered multi-source data, breaks down the barriers of different modal data, and realizes systematic integration. This operation makes full use of the equipment health status information contained in the data, avoids the information loss caused by single modal data, provides standardized input for subsequent cross-modal feature fusion, and significantly improves the ability to mine data value. Secondly, by generating a "spatiotemporal semantic fusion feature vector" based on the normalized multimodal data set, the spatiotemporal dimensions such as equipment topological location and time series are deeply integrated with cross-modal semantics such as image defects and text alarms. This processing method constructs a unified feature representation that includes the spatiotemporal association and semantic logic of the equipment, which can capture the multimodal feature combination of early hidden faults such as weak discharge and progressive insulation degradation, thereby improving the timeliness and accuracy of fault warning and solving the problem of difficulty in detecting early faults. In addition, by inputting the spatiotemporal semantic fusion feature vector into the multimodal fusion large model for forward transmission, the learning ability of the large model and the implicit domain knowledge graph are used to realize risk reasoning based on knowledge rules. This process It can automatically associate the spatial fault propagation paths of equipment, and perform semantic reasoning on new or rare fault modes through large models, reducing dependence on manual experience and effectively improving the model's predictive generalization ability for new faults. Finally, the multimodal fusion large model in this application uses the aforementioned spatiotemporal semantic fusion feature vector as input, so that the prediction process can associate specific equipment topology, time series and cross-modal evidence (such as the conduction path of "main transformer overload → adjacent busbar current abnormality → connector color change image"). The prediction results thus generated can be traced back to specific spatiotemporal scenarios and multimodal features, forming a risk conduction path explanation that meets the standards of the power industry, which can directly support operation and maintenance decisions and solve the industry pain points where the results are unexplainable.
[0018] Optionally, the step 1 of collecting multi-source heterogeneous data related to power operation and maintenance to form a normalized multimodal data set specifically includes the following steps:
[0019] Step 11: Based on the configured multi-source heterogeneous data interface, collect multi-source heterogeneous data related to power operation and maintenance from multi-source heterogeneous data sources to form an original multi-source data set;
[0020] Step 12: Process the original multi-source dataset to generate a cleaned multi-source dataset;
[0021] Step 13: Processing the structured data in the cleaned multi-source dataset to generate a standardized structured dataset;
[0022] Step 14: Process the unstructured data in the cleaned multi-source dataset to generate a labeled unstructured dataset;
[0023] Step 15: Process the standardized structured dataset and the annotated unstructured dataset to generate a normalized multimodal dataset.
[0024] Optionally, step 11, based on the configured multi-source heterogeneous data interface, collecting multi-source heterogeneous data related to power operation and maintenance from multi-source heterogeneous data sources to form an original multi-source data set, specifically includes the following steps:
[0025] Step 111: Based on the configured multi-source heterogeneous data interface, real-time data from power equipment sensors, operation and maintenance log text, inspection images / videos, equipment operation audio, historical fault reports, and industry standard documents are collected from multiple heterogeneous data sources to obtain multi-source heterogeneous data;
[0026] Step 112: Based on the configured data integration framework, perform protocol type parsing on the multi-source heterogeneous data, convert the multi-source heterogeneous data into structured units in a unified format, and annotate the data with protocol type tags to generate protocol-parsed data units.
[0027] Step 113: Based on the data integration rules, the data units after protocol parsing are integrated and indexed to generate an original multi-source data set associated by device ID and time dimension and carrying a temporary index.
[0028] Preferably, for step 111, a plug-in architecture is used for the multi-source heterogeneous data interface, and independent data acquisition plug-ins are configured for different data source types (such as sensors using the Modbus protocol, log servers using the RESTful API, and video surveillance systems using the RTSP protocol). Each plug-in has a built-in protocol parsing module. For example, the Modbus plug-in can parse binary numerical data sent by the sensor, and the RESTful plug-in can process operation and maintenance log text in JSON format, ensuring that the interface has cross-protocol compatibility. The above-mentioned multi-source heterogeneous data interface can adapt to different types of data: numerical type: real-time collection of power equipment sensor data (such as current, voltage, and temperature) through the OPC UA protocol; text type: pulling operation and maintenance log text from the Elasticsearch log cluster and parsing information such as fault time and operation records; image / video type: accessing inspection cameras through the RTSP protocol to periodically capture images of the equipment appearance; audio type: using the built-in microphone array of the edge computing node to collect partial discharge and mechanical vibration audio during equipment operation; document type: synchronizing historical fault reports and industry standard documents (such as DL / T 596) from a document management system (such as SharePoint). During data collection, a breakpoint-resume mechanism is adopted. If data is lost due to network interruption during the collection process, the interface automatically records the timestamp of the last successful collection and continues collection from the breakpoint after the network is restored to ensure data continuity and integrity.
[0029] Preferably, for the above step 112, the protocol identification engine is configured to automatically determine the data protocol type using the following strategy:
[0030] Port feature matching: For example, port 502 corresponds to the Modbus protocol, and port 8080 corresponds to the RESTful API;
[0031] Data header parsing: parsing the protocol identification field in the data message (such as the "HTTP / 1.1" prefix of the HTTP protocol);
[0032] Signature code matching: Use regular expressions to match data features of specific protocols (such as the curly bracket structure of JSON data).
[0033] After completing protocol identification, call the corresponding parser (such as JSON parser or Modbus parser) to convert the corresponding raw data into a key-value pair format. For example, parse the sensor binary data into `{"Device ID":"T1","Temperature":65,"Acquisition time":"2024-10-01 10:00:00"}`.
[0034] Furthermore, in this application, a data mapping template is configured to uniformly convert the parsed data into a standard structured unit, that is, a structured unit in a unified format. The template predefines fields (such as device ID, timestamp, data type, data value) and supports custom extensions. For example, the inspection image data is mapped to `{"device ID":"T1","timestamp":"2024-10-0110:00:00","data type":"image","data value":"image file path"}` to ensure that all data is consistent in format.
[0035] Furthermore, protocol type tags are added to each structural unit to prioritize and trace data during subsequent processing. For example, real-time sensor data is labeled with "Protocol Tag": "Modbus" and log text is labeled with "Protocol Tag": "RESTful", facilitating targeted processing during cleaning and integration.
[0036] For step 113, the data units after protocol parsing are grouped and aggregated according to the device ID (such as "main transformer T1") and the acquisition timestamp. For example, the sensor data, inspection images, and log records of the same device at different time points are merged into a device time series unit to form a `{"device ID":"T1","time series":[{"time stamp":"2024-10-0110:00:00","data":{...}},{"time stamp":"2024-10-01 10:10:00","data":{...}}]}` structure to achieve spatiotemporal alignment of cross-modal data. Furthermore, a composite index is constructed for the classified and aggregated data units, which includes three dimensions: device ID, timestamp, and data type:
[0037] Device ID index: accelerates batch queries of data on the same device and supports device-level health status analysis;
[0038] Timestamp index: supports retrieving data by time range (e.g. querying the operating data of a device in the past 24 hours);
[0039] Data type indexing: Quickly locate specific modal data (such as extracting only image data for defect detection).
[0040] The index is stored in a B+ tree structure to ensure the efficiency and stability of data query.
[0041] Based on this composite index, a temporary index ID is assigned to each data unit, which is used to uniquely identify the data in subsequent cleaning, labeling, and fusion processes. The temporary index, along with the device ID and timestamp, forms a data traceability chain, ensuring traceability throughout the data lifecycle and providing a reliable basis for model training and result interpretation.
[0042] Optionally, step 12, processing the original multi-source dataset to generate a cleaned multi-source dataset, specifically includes the following steps:
[0043] Step 121: Perform sub-modal mixed anomaly detection processing on the original multi-source dataset to generate a sub-modal preliminary cleaned dataset;
[0044] Step 122: Perform cross-modal consistency graph verification on the sub-modal preliminary cleansing dataset to generate a consistency-verified dataset;
[0045] Step 123: Perform dynamic threshold reinforcement learning optimization processing on the consistency-checked dataset to generate a cleaned multi-source dataset.
[0046] Optionally, step 121, performing sub-modal mixed anomaly detection processing on the original multi-source dataset to generate a sub-modal preliminary cleaned dataset, specifically includes the following steps:
[0047] Step 1211, for structured numerical data in the original multi-source data set, based on the Transformer-isolation forest hybrid model, the time-dependent features in the structured numerical data are captured, and the isolation forest algorithm is used to score the "isolation degree" of the data according to the time-dependent features, to identify noise data deviating from the normal distribution and filter, to obtain a structured numerical cleaning subset;
[0048] Step 1212, for text data in the original multi-source data set, based on domain large model semantic verification, the semantic similarity between text data is calculated to determine noise data with semantic contradiction or format anomaly and filter, to obtain a text data cleaning subset;
[0049] Step 1213, for image / audio data in the original multi-source data set, the CNN extracted visual features and MFCC acoustic features are spliced and input into the GAN network using a multi-modal feature fusion detector, to identify media data with abnormal clarity or discontinuous features and filter, to obtain a multimedia data subset;
[0050] Step 1214, align and merge the structured numerical cleaning subset, the text data cleaning subset, and the multimedia data subset according to the space-time dimension, to generate a preliminary cleaning data set.
[0051] Preferably, in step 1211, for time series numerical data collected by power equipment sensors (such as current, temperature) (structured numerical data), first capture the long and short term time-dependent relationship in it through the Transformer encoder to obtain time-dependent features. The model uses a self-attention mechanism to dynamically calculate the correlation weight between data at different time points, such as identifying the time series causal relationship of "voltage fluctuation -> temperature rise". To enhance the expression of time series features, position encoding is added to the input layer of the Transformer encoder, which converts timestamp information into a vector representation that the model can learn, so that the model can perceive the time sequence of the data.
[0052] Further, the extracted time-dependent features are input into the isolation forest algorithm to calculate the "isolation degree" score of each data point. The higher the score, the greater the degree of deviation of the data point from the normal mode (such as a sudden current spike). Based on the set dynamic threshold (such as a score exceeding 95% quantile), data points higher than the threshold are determined as abnormal and filtered. For example, if a transformer temperature data jumps from 60-80℃ to 150℃, it is determined as an outlier to be filtered.
[0053] Preferably, in step 1212, a pre-trained language model in the power field (such as ERNIE-Power) is used for text data such as operation and maintenance logs and fault reports. The model is pre-trained on millions of power texts (including standard documents and fault cases) and has the ability to understand power professional terms (such as "differential protection action" and "insulation aging"). Entity recognition and relationship extraction are performed on the text data. For example, entities such as "equipment ID", "fault type" and "processing time" are extracted from the logs, and logical relationships between entities are constructed (such as "main transformer T1 overtemperature → trigger alarm"). The logical relationships between entities are further represented as semantic vectors, and the cosine similarity between the semantic vector representations is calculated to evaluate semantic consistency. For example, by comparing the fault descriptions of the same device at different times, if the frequency of "excessive temperature" and "cooling system abnormality" appearing at the same time is significantly higher than other combinations, then the two descriptions are considered to be semantically related. Based on a set similarity threshold (such as 0.3), the cosine similarity between semantic vector representations is compared to filter out semantically contradictory texts (such as two logs at the same time point recording "device normal" and "severe overload" respectively) or texts with abnormal formats (such as log entries containing garbled characters).
[0054] Preferably, in step 1213, multimodal feature fusion technology is used to perform anomaly detection on image / audio data. The implementation principle is as follows: for images: ResNet50 is used to extract device appearance features (such as cracks in transformer bushings and heating of joints), and the attention mechanism is used to focus on key areas (such as heat sinks and terminal blocks) to obtain visual features; for audio: MFCC (Mel-Frequency Cepstral Coefficients) is used to extract time-frequency features of device operation audio, such as high-frequency pulse signals of partial discharge.
[0055] Visual and acoustic features are concatenated at the feature vector level and fed into the GAN network for anomaly detection. The generator attempts to reconstruct normal data features, while the discriminator identifies abnormal features that cannot be reconstructed (such as the audio features corresponding to abnormal sounds).
[0056] Anomaly identification is achieved by calculating the reconstruction error. When image clarity falls below a threshold (e.g., PSNR < 20dB) or audio features are incoherent (e.g., spectral mutation exceeding 30%), the data is identified as abnormal. Multi-scale analysis is also introduced to simultaneously detect macroscopic anomalies (e.g., equipment tilt) and microscopic anomalies (e.g., loose bolts) in images, as well as persistent anomalies (e.g., unusual noise from motor bearings) and transient anomalies (e.g., discharge pulses) in audio, to identify abnormal data and generate multimedia data subsets.
[0057] Preferably, in step 1214, cross-modal data associations are established for the cleaned structured numerical, textual, and image / audio datasets, using the device ID and timestamp as core index keys. For example, the temperature data (numerical), inspection log (text), and appearance image (image) for device T1 at "2024-10-01 10:00:00" are associated with the same index node. For data with inconsistent temporal precision (e.g., sensor sampling at the second level vs. image acquisition at the minute level), linear interpolation or nearest neighbor matching is used for temporal alignment. The aligned datasets of different modalities are then mapped into a unified multidimensional tensor space. For example, numerical data is converted into a time series tensor, textual data into a semantic embedding tensor, and image / audio data into a feature tensor. A modality identification dimension is then added to the tensor dimensions. For example, a fourth dimension [modality type] is added to the three-dimensional tensor [time, device, feature] to clearly distinguish between data sources such as numerical, textual, and image data. This results in a merged dataset, and further logical consistency verification is performed on the different modal data at the same spatiotemporal point. For example, if the image shows that the device is smoking (abnormal) at a certain moment, but the temperature sensor data at the corresponding time is normal, the data conflict check process is triggered, and the data reliability is further verified manually or by models.
[0058] Optionally, step 122, performing cross-modal consistency graph verification on the sub-modal preliminary cleaned dataset to generate a consistency-verified dataset, specifically includes the following steps:
[0059] Step 1221: Perform dynamic multimodal knowledge graph construction processing on the sub-modal preliminary cleansing data set to generate a real-time associated knowledge graph;
[0060] Step 1222: Perform graph neural network consistency reasoning on the real-time associated knowledge graph to generate a consistency score graph;
[0061] Step 1223: Perform dynamic threshold filtering and knowledge completion processing on the consistency score map to generate a consistency-verified data set.
[0062] Optionally, step 1221 performs dynamic multimodal knowledge graph construction processing on the sub-modal preliminary cleaned data set to generate a real-time associated knowledge graph, specifically: performing multimodal entity extraction processing on the sub-modal preliminary cleaned data set to generate an initial graph node set; performing time-series dynamic edge generation processing on the initial graph node set to generate a real-time dynamic graph; performing unified semantic space mapping processing on the image features and numerical features in the real-time dynamic graph to generate a cross-modal semantic vector; performing graph edge enhancement processing on the cross-modal semantic vector to generate a real-time associated knowledge graph.
[0063] Optionally, the step 1222 performs graph neural network consistency reasoning on the real-time associated knowledge graph to generate a consistency score graph, specifically: performing graph attention feature aggregation processing on the real-time associated knowledge graph to generate a preliminary attention weight graph; performing reinforcement learning parameter optimization processing on the preliminary attention weight graph to generate an optimized attention weight graph; performing time perception graph modeling processing on the optimized attention weight graph to generate a temporal dynamic graph model; performing logical consistency verification processing on the temporal dynamic graph model to generate a consistency score graph.
[0064] Optionally, step 1223 performs dynamic threshold filtering and knowledge completion processing on the consistency score map to generate a consistency-verified data set, specifically: performing Gaussian mixture modeling processing on the historical scoring data of the consistency score map to generate a dynamic filtering threshold set; performing threshold dynamic filtering processing on the consistency score map to generate a preliminary filtered data set; performing knowledge embedding reasoning processing on the missing associations in the preliminary filtered data set to generate a completed association candidate set; performing fusion verification processing on the completed association candidate set and the preliminary filtered data set to generate a consistency-verified data set.
[0065] Preferably, in the specific implementation, in step 1221, for the sub-modal preliminary cleaning data set (including the structured numerical cleaning subset, the text data cleaning subset, and the multimedia data subset): for the structured numerical data, extract the device ID, parameter value, and timestamp (such as `<Device ID = T1, parameter = temperature, value = 85 ° C, time = 2024-10-01 10:00:00>`; For text data, entity recognition is performed through the ERNIE-Power model to extract the device ID, fault type, and processing status (such as `<Device ID = T1, Fault Type = Overtemperature, Status = Alarm>`); For multimedia data, device components (such as sleeves, heat sinks) and abnormal features (such as cracks, heat generation) are extracted from the image, and abnormal sound features are extracted from the audio, and finally an initial graph node set (including numerical nodes, text nodes, image nodes, and audio nodes) is generated. The following steps are further performed on the initial graph node set: establish a temporal relationship edge based on the timestamp (such as `Temperature rise @ 10:00 → Alarm trigger @ 10:01`); establish a cross-modal association edge based on the device ID (such as `Temperature node of device T1 → Inspection log node of device T1`); The time distance between nodes (e.g., Δt = 60 seconds) is calculated as the edge weight to generate a real-time dynamic graph (including nodes and time series / association edges with weights). Furthermore, for the real-time dynamic graph, the feature vectors of different modal features are extracted through ResNet50 and mapped to a unified semantic space, thereby generating a cross-modal semantic vector (e.g., the distance between `temperature 85°C` and `fever image` in the same vector space is <0.2). Furthermore, for the cross-modal semantic vector, the semantic similarity is calculated: if the cosine similarity is >0.8, the edge weight is enhanced; domain knowledge edges are added: such as "temperature > 80°C → overtemperature alarm" (weight 0.9), a time decay factor is introduced: the edge weight decays with time (e.g., Δt > 24 hours, weight × 0.5), and finally a real-time associative knowledge graph (including enhanced cross-modal semantic edges) is generated.
[0066] Preferably, in step 1222, for the real-time associated knowledge graph, GAT (graph attention network) is used to calculate the attention weights between nodes, and for each node, different attention weights are assigned when aggregating neighbor node features (such as giving higher weights to adjacent nodes in time series), thereby generating a preliminary attention weight map (node features contain weighted aggregation information of neighbor nodes). Further, for the preliminary attention weight map, rewards are calculated based on historical consistency cases annotated by domain experts, and the PPO algorithm is further used to optimize the GAT parameters to maximize the cumulative rewards, thereby generating an optimized attention weight map (parameters optimized by reinforcement learning). Next, for the optimized attention weight map, a time dimension (such as a time embedding vector) is added to the graph structure, and T-GCN (temporal graph convolutional network) is used to capture the dynamic changes of the map to generate a temporal dynamic graph model (capable of expressing the evolution of the map over time). Finally, for the time-series dynamic graph model, we define logical rules in the power field (such as "overheating is inevitably accompanied by a temperature rise trend") and perform logical verification on the associated paths in the graph, calculate the consistency score, and thus generate a consistency score graph (each associated path is accompanied by a consistency score, such as 0.9 indicates high consistency).
[0067] Preferably, in step 1223, for the historical scoring data in the consistency scoring map, the EM algorithm is used to fit the Gaussian mixture model to calculate the probability density of each scoring interval, thereby generating a dynamic filtering threshold set (such as the threshold corresponding to the 95% confidence interval). In addition, for each association path in the consistency scoring map, the corresponding threshold is selected according to the current time and device status to filter out associations below the threshold (such as associations with a score <0.7 are marked as suspicious), thereby generating a preliminary filtered data set (including associations that pass the threshold verification). Further, for the missing associations in the preliminary filtered data set, the TransE model is used to convert the map into a knowledge embedding vector to infer possible missing associations based on the embedding vector (such as "equipment T1 temperature is high" → "cooling system abnormality") to generate a completed association candidate set (including possible missing associations and confidence levels). Finally, for the completed association candidate set and the preliminary filtered data set, the completed association candidate set is added to the preliminary filtered data set, and a logical consistency check is performed again (such as checking whether the newly added association conflicts with existing knowledge) to generate a consistency-verified data set (including completed and verified associations).
[0068] Optionally, step 123, performing dynamic threshold reinforcement learning optimization processing on the consistency-checked data set to generate a cleaned multi-source data set, specifically includes the following steps:
[0069] Step 1231: Perform historical threshold pattern retrieval and scene matching processing on the consistency-checked data set to generate a dynamic threshold configuration set;
[0070] Step 1232: Perform real-time data feedback fine-tuning and abnormal circuit breaking processing on the dynamic threshold configuration set to generate a cleaned multi-source data set.
[0071] Optionally, step 1231 performs historical threshold pattern retrieval and scene matching processing on the consistency-verified data set to generate a dynamic threshold configuration set, specifically: performing incremental feature extraction processing on the historical threshold knowledge base to generate an updated historical threshold knowledge base; performing multi-dimensional scene feature encoding processing on the consistency-verified data set to generate a current scene feature vector; performing hierarchical similarity matching processing on the current scene feature vector and the updated historical threshold knowledge base to generate a matching historical scene threshold set; performing device health correction and dynamic fusion processing on the matching historical scene threshold set to generate a dynamic threshold configuration set.
[0072] Optionally, step 1232 performs real-time data feedback fine-tuning and abnormal fuse processing on the dynamic threshold configuration set to generate a cleaned multi-source data set, specifically: threshold application and preliminary filtering processing are performed on the dynamic threshold configuration set and the consistency-checked data set to generate a preliminary filtered data set; false alarm and missed alarm statistics and regularized threshold fine-tuning processing are performed on the preliminary filtered data set to generate a fine-tuned threshold configuration set; secondary filtering and abnormal fuse processing are performed on the fine-tuned threshold configuration set and the preliminary filtered data set to generate a fuse mark data set; fusion verification and knowledge base update processing are performed on the fuse mark data set and the secondary filtered data set to generate a cleaned multi-source data set.
[0073] Preferably, in step 1231, a sliding window mechanism (e.g., data from the last 30 days) is used to extract time series features from the historical threshold knowledge base (containing historical equipment operation data, abnormal cases, and threshold configurations) to obtain an abnormal pattern library, including: periodic features: such as daily load peaks and valleys, and seasonal temperature changes; trend features: such as baseline drift caused by equipment aging (e.g., a temperature increase of 0.5°C / year); and sudden changes: such as parameter changes before and after equipment maintenance. In addition, an incremental learning algorithm (e.g., iForest) is used to continuously update the historical threshold knowledge base, and the weight of the newly added pattern decays over time (e.g., a weight of 0.8 on day t-1 and a weight of 0.3 on day t-7) to generate an updated historical threshold knowledge base (containing time-varying features and dynamic abnormal patterns). For the data set after consistency verification, a three-dimensional scene feature vector is constructed: equipment dimension: equipment type (such as main transformer, switch), operating years, and health index (HI); environmental dimension: regional climate (such as tropical, temperate), season, and real-time weather (temperature, humidity); operation dimension: load level (light load, full load, overload), and operation mode (single ring network, dual power supply). The three-dimensional scene feature vector is fused and encoded using the Transformer encoder to generate a fixed-length scene feature vector as the current scene feature vector (such as `[device type = main transformer, HI = 0.75, season = summer, load = full load]`). Next, a three-layer similarity matching mechanism is constructed based on the current scenario feature vector and the updated historical threshold knowledge base: device layer: prioritizes matching of devices of the same type (e.g., main transformer T1 matches historical main transformer cases); environment layer: further matches similar environments in the device layer matching results (e.g., summer → matches historical summer cases); operation layer: based on the first two layers, matches the closest operating state (e.g., full load → matches historical full load cases), and uses a mixed metric of cosine similarity + Euclidean distance to calculate the similarity scores of scenarios at different layers (e.g., similarity > 0.8 is considered a strong match) to generate a matching historical scenario threshold set (e.g., a threshold configuration containing 10 historical cases). Finally, for the matched historical scene threshold set, the threshold is adjusted according to the real-time health index (HI) of the device. The formula is: Corrected threshold = historical threshold × (1 + α × (1-HI)) (α is the correction coefficient, such as the temperature threshold α = 0.2, HI = 0.7, the threshold is reduced by 6%), and the corrected threshold set is weighted fused, with the weight based on the scene similarity and time proximity: Final threshold = ∑ (similarity × time weight × corrected threshold), to generate a dynamic threshold configuration set based on the final threshold (such as the temperature threshold is dynamically adjusted to 75-85 ° C, changing with the health of the device).
[0074] Preferably, in step 1232, for the dynamic threshold configuration set and the data set after consistency verification, each data point in the data set is filtered by applying the corresponding threshold. For example, for numerical data, if the temperature is >85°C or <50°C, it is marked as abnormal; for text data, if it contains contradictory keywords (such as "normal" and "alarm" appearing at the same time), it is marked as abnormal; for image / audio data, if the PSNR is <20dB or the spectrum mutation is >30%, it is marked as abnormal. A sliding window is further used to count the abnormality ratio. When the abnormality ratio in the window is >30%, an early warning is triggered, thereby generating a preliminary filtered data set (containing normal data and marked abnormal data). For the preliminary filtered data set, a false positive and false negative evaluation matrix is constructed: false positive: the ratio of manually confirmed normal data that is marked as abnormal; false negative: the ratio of manually confirmed abnormal data that is not marked. The threshold is further fine-tuned based on the evaluation results: if the false alarm rate is high (>15%), the threshold is relaxed (for example, the temperature threshold is changed from 85°C to 90°C); if the false alarm rate is high (>10%), the threshold is tightened (for example, the temperature threshold is changed from 85°C to 80°C), and the step size is adaptively fine-tuned: dynamically adjusted according to the degree of false alarm / false alarm (for example, the step size is 5°C when the false alarm rate is 20%, and the step size is 1°C when the false alarm rate is 5%), thereby generating a fine-tuned threshold configuration set (for example, the temperature threshold is changed from 85°C to 88°C). Next, for the fine-tuned threshold configuration set and the preliminary filtered data set, the fine-tuned threshold is applied to re-screen the preliminary filtered data and perform continuous anomaly detection: if the same device has an anomaly at five consecutive time points, a fuse is triggered (forcibly marked as an anomaly), and cross-modal consistency fuse: when the image shows smoke (anomaly) but the temperature is normal, if such a conflict occurs more than 3 times / h, a fuse is triggered (the temperature data is marked as suspicious), and device status fuse: when the device is in maintenance status, related anomalies are automatically ignored (such as the high temperature of the main transformer under maintenance does not alarm), thereby generating a fuse mark data set (including normal data after secondary filtering, marked abnormal data and fuse mark) as the secondary filtered data set. Finally, the fuse mark dataset and the secondary filtered dataset are fused and verified: cross-modal consistency check: for example, when a temperature anomaly occurs, check whether there is corroborating image / audio evidence at the corresponding time; temporal coherence check: for example, check whether the anomaly conforms to the development pattern of equipment failure (such as a slow temperature rise → a sharp rise → failure); knowledge base update: add confirmed new anomaly patterns to the historical threshold knowledge base; update the equipment health index (HI): adjust the HI value based on the severity of the anomaly (for example, a severe anomaly causes the HI to drop by 0.1). This ultimately generates a cleaned multi-source dataset (including the final filtered normal data and traceable anomaly data).
[0075] Optionally, step 13, processing the structured data in the cleaned multi-source dataset to generate a standardized structured dataset, specifically includes the following steps:
[0076] Step 131: Perform feature dimensional distribution modeling on the structured data in the cleaned multi-source data set to generate a feature dimensional distribution matrix;
[0077] Step 132: performing a standardized strategy tensor generation process on the characteristic dimension distribution matrix to generate a standardized strategy tensor;
[0078] Step 133: Perform tensor conversion and semantic verification on the standardized policy tensor and structured data to generate a standardized structured data set.
[0079] Optionally, step 131 performs characteristic dimensional distribution modeling on the structured data in the cleaned multi-source data set to generate a characteristic dimensional distribution matrix, specifically: performing dynamic dimensional analysis and distribution estimation on the cleaned structured data to generate a dimension-distribution feature vector; performing multidimensional matrix construction and clustering on the dimension-distribution feature vector to generate a characteristic dimensional clustering matrix; and performing domain knowledge injection and matrix enhancement on the characteristic dimensional clustering matrix to generate a characteristic dimensional distribution matrix.
[0080] Optionally, the step 132 performs standardized policy tensor generation processing on the feature dimensional distribution matrix to generate a standardized policy tensor, specifically: automatically generates a policy candidate set on the feature dimensional distribution matrix to generate a policy candidate tensor; performs knowledge graph enhanced reasoning processing on the policy candidate tensor to generate an optimized policy tensor; and performs dynamic device health adaptation processing on the optimized policy tensor to generate a standardized policy tensor.
[0081] Optionally, step 133 performs tensor conversion and semantic verification on the standardized policy tensor and structured data to generate a standardized structured data set, specifically: performs parallel tensor operation conversion processing on the structured data and the standardized policy tensor to generate a preliminary standardized feature tensor; performs cross-modal semantic association verification processing on the preliminary standardized feature tensor to generate a semantic verification score matrix; performs feedback strategy fine-tuning processing on the semantic verification score matrix and the preliminary standardized feature tensor to generate a standardized structured data set.
[0082] Preferably, in step 131, a systematic analysis is performed on the structured data (such as sensor values of power equipment current, voltage, temperature, etc.) in the cleaned multi-source data set to construct a matrix reflecting its dimensionality and distribution characteristics, which is specifically implemented as follows:
[0083] First, dynamic dimensionality analysis and distribution estimation are performed. Data dimensions are analyzed by identifying the physical units of the data (e.g., amperes for current, volts for voltage) and the data source (e.g., transformer oil temperature sensor, line current transformer). Kernel density estimation is then used to model the probability distribution of the data, calculating statistics such as mean, variance, and quantiles. This generates a dimension-distribution feature vector containing both dimensional information and distribution parameters. For example, for transformer oil temperature data, the feature vector [dimension = °C, mean = 65, standard deviation = 5, 95% quantile = 75] is generated. Next, a multidimensional matrix is constructed and clustered. Using these feature vectors as matrix elements, a multidimensional feature dimension matrix is constructed, with each row corresponding to a data feature (e.g., oil temperature, current) and each column corresponding to a dimension or distribution parameter. Subsequently, a hierarchical clustering algorithm is used to group data features based on dimensional type and distribution similarity. Features with similar properties are clustered into the same category, forming a feature dimension clustering matrix, achieving a preliminary classification and integration of the data features. Finally, domain knowledge injection and matrix enhancement are implemented. Power domain knowledge (such as equipment rated parameters, safe operating ranges, and industry standard thresholds) is incorporated into the clustering matrix in the form of rules or weights. For example, the transformer's rated oil temperature limit of 85°C is used as a constraint, the weight of the corresponding data features is increased, and missing domain knowledge attributes (such as safety margins and warning thresholds) are supplemented. Ultimately, a feature dimension distribution matrix is generated that comprehensively reflects the data's dimensional and distribution characteristics, providing a basis for the formulation of standardization strategies.
[0084] Preferably, in step 132, based on the characteristic dimension distribution matrix obtained in step 131, a standardized strategy tensor adapted to the characteristics of the power data is generated through multi-stage processing, and the specific implementation process is as follows:
[0085] First, a set of strategy candidates is automatically generated. For different data features in the matrix (e.g., highly volatile current data, low-volatility voltage data), combined with Z-score normalization, a variety of standardized strategy combinations are generated. For example, for current data, candidate strategies such as "Z-score normalization" and "normalization by historical extreme values" are generated. The parameters of each strategy (e.g., mean, standard deviation, scaling factor) are encoded as tensor elements to form a strategy candidate tensor.
[0086] Secondly, knowledge graph-enhanced reasoning is performed. Based on the power sector knowledge graph (including equipment topology, fault transmission paths, and operating rules), the impact of each candidate strategy on data characteristics is analyzed. Through knowledge graph reasoning, the rationality and applicability of the strategy are evaluated. For example, for voltage data close to the rated value, a normalization strategy that retains the original dimension is preferred to avoid loss of physical meaning due to over-normalization. Based on the reasoning results, the strategy combination is screened and optimized, generating an optimized strategy tensor to ensure that the strategy meets the actual needs of the power sector.
[0087] Finally, dynamic device health adaptation is performed: The parameters in the optimization strategy tensor are dynamically adjusted based on the device health assessment results (such as aging level and historical fault records). For example, for severely aged devices, data fluctuations may be caused by decreased device performance. In this case, the threshold range is appropriately relaxed during normalization to enhance the strategy's robustness to abnormal data, ultimately generating a standardized strategy tensor that matches the device status.
[0088] Preferably, in step 133, the normalization strategy tensor is applied to the structured data, and the rationality and consistency of the normalization result are ensured through semantic verification. The specific implementation process is as follows:
[0089] First, parallel tensor conversion processing is performed. Structured data is synchronously converted according to the policy parameters specified in the normalization policy tensor. For example, if the policy is "Z-score normalize current data," the original structured data is converted into a preliminary normalized feature tensor, achieving uniform data format and value range.
[0090] Next, cross-modal semantic association verification is performed: the initially standardized structured data features are associated with the semantic information of other modal data (such as operation and maintenance log text and inspection images). Natural language processing technology is used to extract fault descriptions (such as "overtemperature alarm") in the text, and computer vision algorithms are used to identify abnormal equipment features in the image (such as discoloration of the connector) and perform semantic matching with structured data (such as temperature values). Cosine similarity calculation, knowledge graph entity alignment and other technologies are used to evaluate the consistency of the semantic expression of data features in different modalities, generate a semantic verification score matrix, and quantify the degree of semantic association of each data feature.
[0091] Finally, feedback-based strategy fine-tuning is performed: Based on the feedback from the semantic verification score matrix, the strategy is optimized for the initial standardized feature tensor. For features with low semantic relevance (such as inconsistencies between the normalized temperature value and the image hot spot information), the normalization strategy parameters are readjusted or additional constraints are added (such as setting an upper temperature threshold), and the optimization process is iteratively performed. After multiple rounds of feedback adjustment, a standardized structured dataset is ultimately generated that meets the standards of the power industry at the dimensional, numerical, and semantic levels, providing high-quality input for subsequent risk prediction models.
[0092] Optionally, step 14, processing the unstructured data in the cleaned multi-source dataset to generate a labeled unstructured dataset, specifically includes the following steps:
[0093] Step 141: performing sub-modal feature extraction and dynamic annotation processing on the unstructured data in the cleaned multi-source dataset to generate a preliminary annotated dataset;
[0094] Step 142: Perform cross-modal graph association and logic verification on the preliminary annotated dataset to generate a consistent annotated dataset;
[0095] Step 143: Perform expert knowledge fusion and dynamic update processing on the consistent annotated dataset to generate an annotated unstructured dataset.
[0096] Optionally, step 141 performs sub-modal feature extraction and dynamic annotation processing on the unstructured data in the cleaned multi-source data set to generate a preliminary annotated data set, specifically: performing domain large model enhanced semantic extraction processing on the text logs in the cleaned unstructured data to generate a text semantic annotation set; performing multi-modal feature fusion annotation processing on the inspection images in the cleaned unstructured data to generate an image visual annotation set; performing time-frequency feature fusion annotation processing on the device audio in the cleaned unstructured data to generate an audio acoustic annotation set; performing cross-modal graph association verification processing on the text semantic annotation set, the image visual annotation set, and the audio acoustic annotation set to generate a preliminary annotated data set.
[0097] Optionally, step 142, performs cross-modal graph association and logic verification processing on the preliminary annotated dataset to generate a consistent annotated dataset, specifically: performs cross-modal knowledge graph construction processing on the preliminary annotated dataset to generate a dynamic association graph; performs graph neural network logic reasoning processing on the dynamic association graph to generate a consistency scoring matrix; performs dynamic threshold filtering and rule updating processing on the consistency scoring matrix to generate a consistent annotated dataset.
[0098] Optionally, step 143 performs expert knowledge fusion and dynamic update processing on the consistent annotated dataset to generate an annotated unstructured dataset, specifically: performs expert feedback reinforcement learning processing on the consistent annotated dataset to generate an expert optimized annotation set; performs industry standard dynamic mapping processing on the expert optimized annotation set to generate a standard adapted annotation set; performs closed-loop model iterative update processing on the standard adapted annotation set to generate an annotated unstructured dataset.
[0099] Preferably, in step 14, for the unstructured data (text logs, inspection images, equipment audio) in the cleaned multi-source dataset, a preliminary annotated dataset is generated through sub-modal processing and cross-modal verification. The specific implementation process is as follows:
[0100] First, domain-wide model-enhanced semantic extraction is performed on text logs. Using pre-trained language models for the power sector (such as ERNIE-Power), combined with named entity recognition (NER) and relation extraction techniques, key semantic information, such as device ID, fault type, and treatment measures, is extracted from operation and maintenance logs. For example, the log "Main transformer T1 experienced an overtemperature alarm at 10:00 AM on 2024-10-01, and the cooling system has been activated" is parsed into the structured annotations `{"Device ID":"T1","Fault Type":"Overtemperature","Time":"2024-10-01 10:00","Treatment":"Start cooling system"}`, forming a text semantic annotation set.
[0101] Secondly, the inspection images are annotated using multimodal feature fusion. The ResNet50 convolutional neural network is used to extract visual features (such as cracks on transformer bushings and discoloration on joints), and the attention mechanism is used to focus on key areas. At the same time, OCR technology is used to identify textual information in the image (such as equipment nameplates and parameter labels). The visual features are then fused with the textual information to generate an image visual annotation set containing the defect location, type, and related parameters. For example, `{"Equipment ID":"T1","Defect Type":"Bushing Crack","Location":[x1,y1,x2,y2],"Severity Level":"Intermediate"}`.
[0102] Next, the device audio is annotated using time-frequency feature fusion. MFCC (Mel-Frequency Cepstral Coefficients) are used to extract audio time-frequency features and identify abnormal sound patterns such as partial discharge and mechanical vibration. A short-time Fourier transform (STFT) is combined to generate a spectrogram. Deep learning models (such as CNN-LSTM) are then used to classify audio events, annotating the anomaly type, occurrence time, and duration. This creates an audio acoustic annotation set, such as `{"Device ID":"T1","Anomaly Type":"Partial Discharge","Time":"2024-10-01 10:05:00","Duration":5s}`.
[0103] Finally, a cross-modal graph correlation check is performed on the three types of annotation sets: using device ID and timestamp as indexes, a preliminary correlation graph is constructed to verify the consistency of annotation information across different modalities. For example, if a text log records an "overtemperature alarm," the corresponding image is checked for hot spots and the audio is checked for abnormal sounds. If any inconsistencies are found, the annotations are corrected, ultimately generating a preliminary annotated dataset.
[0104] Preferably, in step 142, a consistent annotation dataset is generated through graph construction, logical reasoning, and threshold filtering. The specific technical implementation process is as follows:
[0105] First, a cross-modal knowledge graph is constructed: the entities (equipment, fault, time) and relationships (occurrence, processing, association) in the preliminary annotated data set are mapped into knowledge graph nodes and edges. For example, a relationship chain of "main transformer T1 → occurrence → overtemperature alarm → trigger → cooling system start" is constructed; at the same time, prior knowledge of the knowledge graph in the power field (such as equipment topology and fault conduction path) is introduced to supplement missing associations and form a dynamic association graph.
[0106] Next, graph neural network logical reasoning is performed: Graph Attention Networks (GAT) or Graph Convolutional Networks (GCN) are used to aggregate graph features and calculate the confidence level of logical associations between nodes. For example, by reasoning about the contradictory relationship between "overtemperature alarm" and "cooling system not started," a consistency score is generated for each edge. Combined with time series information (such as requiring action within 5 minutes of a fault), the temporal logic rationality of the annotations is evaluated to generate a consistency score matrix.
[0107] Finally, dynamic threshold filtering and rule updating are performed: based on the distribution of historical annotation data, a Gaussian mixture model is used to calculate the dynamic filtering threshold, and annotations with scores below the threshold (such as contradictory or ambiguous associations) are marked as pending correction; at the same time, new logical rules discovered during the verification process (such as "abnormal pressure of hydrogen energy equipment → accompanied by abnormal noise") are updated to the rule base, and after filtering abnormal annotations, a consistent annotation dataset is generated.
[0108] Preferably, in step 143, by integrating expert experience, industry standards and model iteration, a final annotated unstructured dataset is generated. The specific technical implementation process is as follows:
[0109] First, expert feedback reinforcement learning processing is performed: the consistent labeled dataset is submitted to power experts for manual review, and the experts provide feedback on the accuracy and completeness of the annotations (such as correcting incorrect annotations and supplementing missing information). Then, a reinforcement learning algorithm (such as PPO) is used to optimize the annotation model parameters with expert feedback as the reward signal, so that the model learns the expert annotation logic and generates an expert-optimized annotation set.
[0110] Secondly, dynamic mapping processing of industry standards is carried out: the expert-optimized annotation set is aligned with the power industry standards (such as DL / T596 and GB / T 38591), custom annotation terms are converted into standard terms (such as mapping "overtemperature" to "overtemperature operation"), and necessary fields (such as safety level and processing time limit) are supplemented according to standard requirements to generate a standard-adapted annotation set that meets the specifications.
[0111] Finally, the closed-loop model is iteratively updated: the standard adaptation annotation set is used as training data to retrain the annotation model; the model's prediction results on the new data are again verified by experts, forming a closed loop of "annotation-verification-optimization"; at the same time, the newly added annotation cases and rules are incorporated into the knowledge base, and the model is continuously updated, and finally a labeled unstructured data set is generated to ensure the accuracy, standardization and timeliness of the annotation results.
[0112] Optionally, step 15, processing the standardized structured dataset and the annotated unstructured dataset to generate a normalized multimodal dataset, specifically includes the following steps:
[0113] Step 151: Perform spatiotemporal alignment and tensor fusion processing on the standardized structured dataset and the annotated unstructured dataset to generate a multimodal fusion feature tensor:.
[0114] Step 152: Perform domain knowledge graph constraint verification processing on the multimodal fusion feature tensor to generate a consistency verification tensor;
[0115] Step 153: Perform dynamic threshold filtering and knowledge completion processing on the consistency check tensor to generate a normalized multimodal dataset.
[0116] Optionally, step 151 performs spatiotemporal alignment and tensor fusion processing on the standardized structured dataset and the annotated unstructured dataset to generate a multimodal fusion feature tensor, specifically: performing spatiotemporal index construction processing on the standardized structured dataset and the annotated unstructured dataset to generate a cross-modal spatiotemporal index table; performing feature tensor conversion processing on the data associated with the cross-modal spatiotemporal index table to generate a heterogeneous modal feature tensor; and performing attention weighted fusion processing on the heterogeneous modal feature tensors to generate a multimodal fusion feature tensor.
[0117] Optionally, step 152 performs domain knowledge graph constraint verification processing on the multimodal fusion feature tensor to generate a consistency verification tensor, specifically: performing domain knowledge graph access and rule mapping processing on the multimodal fusion feature tensor to generate a constraint rule tensor; performing graph neural network inference processing on the constraint rule tensor and the multimodal fusion feature tensor to generate a consistency score tensor; performing dynamic rule update and tensor fusion processing on the consistency score tensor to generate a consistency verification tensor.
[0118] Optionally, step 153 performs dynamic threshold filtering and knowledge completion processing on the consistency check tensor to generate a normalized multimodal dataset, specifically: performing historical distribution modeling and dynamic threshold generation processing on the consistency check tensor to generate an adaptive filtering threshold set; performing threshold filtering and rare pattern retention processing on the consistency check tensor to generate a preliminary filtered dataset; performing knowledge graph reasoning and completion processing on the missing associations in the preliminary filtered dataset to generate a completed association candidate set; performing fusion verification and standard normalization processing on the completed association candidate set and the preliminary filtered dataset to generate a normalized multimodal dataset.
[0119] Preferably, step 151 takes a standardized structured data set (such as equipment operation values) and annotated unstructured data set (such as inspection image annotations and log text) as objects, and generates a multimodal fusion feature tensor through three steps. The specific technical implementation process is as follows:
[0120] First, a spatiotemporal index is constructed. Using the device ID and timestamp as the core index, the numerical acquisition time (such as the sampling time of current data) is extracted from the standardized structured data, and information such as the image capture time and log recording time is parsed from the annotated unstructured data. For data with inconsistent time accuracy (such as sensor data at the second level and inspection images at the minute level), linear interpolation or nearest neighbor matching algorithms are used to align and construct a cross-modal spatiotemporal index table. For example, the temperature data of device T1 at "2024-10-01 10:00:00", the inspection image annotation at the corresponding time, and the operation and maintenance log records are associated with the same index entry to clarify the spatiotemporal correspondence of multimodal data.
[0121] Secondly, feature tensor conversion processing is performed. For data associated with the cross-modal spatiotemporal index table, structured data (such as temperature and current values) are converted into time series tensors and encoded into vectors containing time series features through Transformer. Text semantic annotations in unstructured data are converted into semantic embedding tensors through the BERT-Power model. Image visual annotations and audio acoustic annotations are processed into feature tensors using ResNet50 and MFCC+CNN, respectively. Finally, heterogeneous modal feature tensors are generated, achieving a unified mathematical representation of multimodal data.
[0122] Finally, attention-weighted fusion processing is performed. A multimodal attention mechanism is designed to dynamically assign weights based on the operating characteristics of power equipment. For example, when the equipment is overloaded, numerical data such as current and temperature are given higher weights; when an image anomaly is detected, the weight of visual features is increased. Through weighted summation or tensor concatenation, heterogeneous modal feature tensors are fused into a multimodal fusion feature tensor, preserving the key information of each modal data and strengthening its relevance.
[0123] Preferably, in step 152, based on the multimodal fusion feature tensor, this step generates a consistency check tensor through knowledge graph constraints and reasoning. The specific technical implementation process is as follows:
[0124] First, domain knowledge graph access and rule mapping are performed. The power domain knowledge graph is introduced, covering information such as equipment topology (such as the connection between transformers and busbars), fault conduction rules (such as "overtemperature → insulation aging"), and industry standards and specifications (such as GB / T power equipment safety thresholds). Rules in the knowledge graph (such as "main transformer oil temperature exceeding 85°C triggers an alarm") are mapped into computable constraints and converted into a constraint rule tensor, where each tensor element corresponds to the parameters and logical expression of a rule.
[0125] Next, graph neural network inference processing is performed. The multimodal fusion feature tensor and the constraint rule tensor are input into a graph neural network (such as a GNN). A computational graph is constructed with the tensor elements as nodes and the data associations as edges. Through graph convolution operations, the multimodal data is inferred to determine whether it conforms to the domain knowledge constraints. For example, if the oil temperature data exceeds a threshold but there is no alarm record in the log, the contradiction score of this association is calculated. Finally, a consistency score tensor is generated, which quantifies the compliance level of each data association.
[0126] Finally, dynamic rule updates and tensor fusion processing are performed. Based on the consistency score results, if a new data association pattern is discovered (such as "abnormal pressure of hydrogen energy equipment accompanied by a specific audio signal"), it is updated to the domain knowledge graph and constraint rule tensor. At the same time, abnormal data with low scores are marked, and the consistency score tensor is fused with the original multimodal fusion feature tensor to generate a consistency verification tensor containing the verification results.
[0127] Preferably, in step 153, a final normalized multimodal dataset is generated through threshold filtering, knowledge reasoning, and normalization processing. The specific technical implementation process is as follows:
[0128] First, historical distribution modeling and dynamic threshold generation are performed. Gaussian mixture modeling is performed on the historical data in the consistency check tensor to analyze data distribution characteristics (such as the range of numerical fluctuations during normal operation and the confidence distribution of each modal association). Combining the health status of the equipment (such as allowing a wider fluctuation range for aging equipment) and the operating scenario (such as relaxing the temperature threshold during high temperatures in summer), an adaptive filtering threshold set is generated to achieve dynamic adjustment of the threshold.
[0129] Next, threshold filtering and rare pattern retention are performed. The consistency check tensor is filtered according to the threshold set, filtering out contradictory or abnormal data with scores below the threshold (such as records with temperatures exceeding the threshold but no hot spots in the image). At the same time, data that does not meet the current rules but may represent new failure modes (such as rare equipment combination anomalies) is marked and retained to form a preliminary filtered data set.
[0130] Next, the knowledge graph is used for reasoning and completion. For missing associations in the initial filtered dataset (e.g., situations where the logs don't record them but the images show anomalies), the power knowledge graph is used for reasoning. By using the device association rules in the graph (e.g., "switch anomaly → affects downstream voltage") and the case library, possible associations are predicted and candidate sets of completed associations are generated. For example, the logical association "discoloration of the connector in the image → corresponding current data anomaly" is supplemented.
[0131] Finally, fusion verification and standardization are performed. The completed association candidate set is merged with the preliminary filtered data set, and the data consistency is verified again. The data set is standardized according to power industry standards (such as data format specifications and term definitions), and the data units and label formats are unified. Finally, a standardized multimodal data set that meets the standards is generated, providing high-quality input for subsequent data analysis and model training.
[0132] Optionally, step 2, generating a spatiotemporal semantic fusion feature vector based on the normalized multimodal dataset, specifically includes the following steps:
[0133] Step 21: performing spatiotemporal feature extraction and tensor encoding processing on the normalized multimodal dataset to generate a spatiotemporal feature tensor;
[0134] Step 22: Perform cross-modal semantic fusion processing on the spatiotemporal feature tensor to generate a semantic fusion tensor;
[0135] Step 23: Perform knowledge graph constraint and feature dimensionality reduction processing on the semantic fusion tensor to generate a spatiotemporal semantic fusion feature vector.
[0136] Optionally, step 21, performing spatiotemporal feature extraction and tensor encoding processing on the normalized multimodal dataset to generate a spatiotemporal feature tensor, specifically includes the following steps:
[0137] Step 211: performing dynamic time window feature extraction processing on the normalized multimodal dataset to generate a time series feature matrix;
[0138] Step 212: performing device topology-aware spatial feature extraction processing on the temporal feature matrix to generate a spatial correlation tensor;
[0139] Step 213: performing cross-modal semantic embedding processing on the unstructured features in the normalized multimodal dataset to generate a modality semantic tensor;
[0140] Step 214: Perform spatiotemporal dimension fusion encoding processing on the temporal feature matrix, the spatial correlation tensor, and the modal semantic tensor to generate a spatiotemporal feature tensor.
[0141] Optionally, step 211: performing dynamic time window feature extraction processing on the normalized multimodal dataset to generate a time series feature matrix, specifically: performing sampling frequency identification and window division processing on the numerical features in the normalized multimodal dataset to generate a dynamic time window set; performing time series pattern extraction processing on the numerical features associated with the dynamic time window set to generate a time series feature vector set; performing event sequence encoding processing on the unstructured data timestamps in the normalized multimodal dataset to generate a time series trigger signal matrix; performing time axis alignment and feature splicing processing on the time series feature vector set and the time series trigger signal matrix to generate a time series feature matrix.
[0142] Optionally, step 212: performs device topology-aware spatial feature extraction processing on the time series feature matrix to generate a spatial association tensor, specifically: performs association mapping processing on the time series feature matrix and the power equipment topology knowledge graph to generate a device topology association graph; performs graph attention feature aggregation processing on the device topology association graph to generate a spatial feature aggregation tensor; performs device health weighting and dimension encoding processing on the spatial feature aggregation tensor to generate a spatial association tensor.
[0143] Optionally, step 213: performing cross-modal semantic embedding processing on the unstructured features in the normalized multimodal dataset to generate a modal semantic tensor, specifically: performing CLIP model semantic embedding processing on the image features in the normalized multimodal dataset to generate a visual semantic tensor; performing BERT-ERNIE semantic encoding processing on the text features in the normalized multimodal dataset to generate a text semantic tensor; performing MFCC-Transformer semantic mapping processing on the audio features in the normalized multimodal dataset to generate an audio semantic tensor; performing modal splicing and unified dimension processing on the visual semantic tensor, the text semantic tensor and the audio semantic tensor to generate a modal semantic tensor.
[0144] Optionally, step 214: performs spatiotemporal dimension fusion encoding processing on the temporal feature matrix, spatial correlation tensor and modal semantic tensor to generate a spatiotemporal feature tensor, specifically: performs spatiotemporal dimension alignment processing on the temporal feature matrix and the spatial correlation tensor to generate a spatiotemporal correlation intermediate tensor; performs cross-modal attention fusion processing on the spatiotemporal correlation intermediate tensor and the modal semantic tensor to generate a preliminary spatiotemporal feature tensor; performs device health gating and dimension compression processing on the preliminary spatiotemporal feature tensor to generate a spatiotemporal feature tensor.
[0145] Preferably, in step 211, the focus is on normalizing the multimodal dataset to extract the temporal features therein and form a matrix. The specific technical implementation process is as follows:
[0146] First, the sampling frequency and window division process are performed for the numerical features in the dataset (such as equipment current, voltage, and temperature). The system automatically detects the sampling frequency of each numerical feature, for example, identifying that the oil temperature data of a certain transformer is sampled once per second, and the current data of a certain line is sampled once every 5 seconds. Based on the sampling frequency and the operating patterns of the equipment, the time window is dynamically divided, such as using a 1-minute sliding window for the high-frequency sampled oil temperature data and a 5-minute sliding window for the low-frequency sampled current data, ultimately generating a dynamic time window set.
[0147] Next, the numerical features associated with the dynamic time window set are processed for temporal pattern extraction. Signal processing methods such as Fourier transform and wavelet transform are used to analyze the trends, periodicity, and mutation characteristics of the data within each window. LSTM and Transformer are also used to extract long-term dependencies, such as identifying the temporal causal chain of "voltage fluctuation → frequency change → current anomaly," thereby generating a set of temporal feature vectors.
[0148] Subsequently, the timestamps of unstructured data in the dataset (such as inspection logs and image annotations) are encoded into event sequences. These timestamps are converted into event trigger signals, such as high-level signals for equipment failure alarms and low-level signals for normal operation. This forms a timing trigger signal matrix that records the order in which each event occurs and its state changes along the timeline.
[0149] Finally, the time series feature vector set and the time series trigger signal matrix are aligned on the time axis, and through feature splicing operations, the time series pattern of numerical features is integrated with the event sequence of unstructured data to generate a time series feature matrix, thereby realizing the unified expression of multimodal data time series information.
[0150] Preferably, in step 212, based on the time series feature matrix generated in step 211 and combined with the power equipment topology knowledge, spatial correlation features are extracted. The specific process is as follows:
[0151] First, an association mapping process is performed to associate the time series feature matrix with the power equipment topology knowledge graph. Using the device ID as a bridge, the operating data of each device (such as the numerical characteristics and event information of transformers and circuit breakers) is mapped to the corresponding node in the topology graph. This constructs a device topology association graph, which intuitively displays the connection relationships between devices (such as the connection between transformers and busbars and the on / off control of circuit breakers) and the data flow path.
[0152] Next, graph attention feature aggregation is performed on the device topology graph. Using a graph attention network (GAT), attention coefficients are dynamically allocated based on the physical connection weights and operational relevance between devices, focusing on key nodes and edges. For example, when a current anomaly occurs on a line, the feature weights of the device nodes at both ends of the line and the associated devices are enhanced. This is then aggregated to form a spatial feature aggregation tensor, highlighting the range of devices potentially affected by the anomaly and the degree of relevance.
[0153] Finally, based on the device health assessment results (calculated from historical fault records and real-time monitoring data), the spatial feature aggregation tensor is weighted and dimensionally encoded. Devices with lower health scores are given higher weights to emphasize their impact on the overall system. At the same time, the device's spatial location information (such as substation coordinates and device installation location) and connection level are encoded into the tensor dimensions to generate a spatial correlation tensor that comprehensively reflects the coupling relationship between the device's spatial topology and operating status.
[0154] Preferably, in step 213: semantic embedding and fusion are achieved through different models for the unstructured features in the normalized multimodal dataset, and the specific steps are as follows:
[0155] The CLIP model semantic embedding process is used to process image features. Through contrastive learning, the CLIP model maps equipment images (such as transformer appearance and line joint status) and text descriptions (such as "bushing crack" and "joint overheating") into the same semantic space. This extracts visual semantic information from the image, such as equipment defects and operating status, and generates a visual semantic tensor.
[0156] BERT-ERNIE semantic encoding is used to process text features. Combining BERT's bidirectional Transformer architecture with ERNIE's pre-training advantages in the power sector, this method performs word segmentation and encoding on text such as operation and maintenance logs and fault reports, extracting semantic information such as device ID, fault type, and treatment measures, and converting it into a text semantic tensor.
[0157] For audio features, we implement MFCC-Transformer semantic mapping. MFCCs are first used to extract the time-frequency characteristics of the equipment's operating audio. Then, through the Transformer model, these time-frequency characteristics are semantically associated with abnormal sound patterns of power equipment (such as partial discharge and mechanical vibration) to generate an audio semantic tensor.
[0158] Finally, the visual, text, and audio semantic tensors are modally concatenated and dimensionalized. This tensor concatenation integrates semantic information from different modalities and adjusts the dimensions to make them consistent, generating a modal semantic tensor. This enables the fusion of cross-modal semantics for unstructured data.
[0159] Preferably, in step 214, the temporal feature matrix, the spatial correlation tensor and the modal semantic tensor are deeply fused to generate a spatiotemporal feature tensor. The specific technical implementation is as follows:
[0160] First, the time series feature matrix and the spatial correlation tensor are aligned in the time and space dimensions. Using timestamps and device IDs as a benchmark, the data points in the time series are mapped one-to-one with the nodes in the device spatial topology. The tensor dimensions are adjusted to align the two in time and space, generating a temporal and spatial correlation intermediate tensor and establishing a preliminary connection between the temporal evolution and spatial distribution of the data.
[0161] Next, the spatiotemporal correlation intermediate tensor is fused with the modal semantic tensor using cross-modal attention. A multi-head attention mechanism dynamically assigns weights to different modal data based on their contribution to device status assessment. For example, when an abnormal device image is detected, the association weight between the visual semantic tensor and the spatiotemporal data is enhanced. A weighted fusion operation is then used to generate a preliminary spatiotemporal feature tensor, achieving deep fusion of multimodal data in both the spatiotemporal and temporal dimensions.
[0162] Finally, based on the device health assessment results, the preliminary spatiotemporal feature tensor is subjected to device health gating and dimensionality compression. A health gating unit is set up to adjust the importance of each feature in the tensor based on the device health status, with features of devices with low health being given higher weights. Simultaneously, a dimensionality reduction algorithm, principal component analysis (PCA), is used to remove redundant information and compress the tensor dimensions, ultimately generating a spatiotemporal feature tensor that expresses the spatiotemporal semantic characteristics of multimodal data in a compact and efficient manner.
[0163] Optionally, step 22, performing cross-modal semantic fusion processing on the spatiotemporal feature tensor to generate a semantic fusion tensor, specifically includes the following steps:
[0164] Step 221: Perform intra-modal attention enhancement processing on the spatiotemporal feature tensor to generate a modality-enhanced feature tensor;
[0165] Step 222: Perform cross-modal attention fusion processing on the modality-enhanced feature tensor to generate a cross-modal correlation tensor;
[0166] Step 223: Perform domain knowledge graph constraint and semantic completion processing on the cross-modal association tensor to generate a semantic fusion tensor.
[0167] Optionally, step 221 performs intra-modal attention enhancement processing on the spatiotemporal feature tensor to generate a modal enhanced feature tensor, specifically: performing time-space self-attention enhancement processing on the numerical features in the spatiotemporal feature tensor to generate a numerical enhanced feature tensor; performing defect area channel attention enhancement processing on the image features in the spatiotemporal feature tensor to generate an image enhanced feature tensor; performing fault entity self-attention enhancement processing on the text features in the spatiotemporal feature tensor to generate a text enhanced feature tensor; performing time-frequency domain attention enhancement processing on the audio features in the spatiotemporal feature tensor to generate an audio enhanced feature tensor.
[0168] Optionally, step 222 performs cross-modal cross-attention fusion processing on the modal enhancement feature tensor to generate a cross-modal association tensor, specifically: performing cross-modal attention network construction processing on the modal enhancement feature tensor to generate an attention weight matrix; performing dynamic weighted fusion processing on the cross-modal attention weight matrix and the modal enhancement feature tensor to generate a preliminary cross-modal association tensor; performing semantic consistency verification and feature optimization processing on the preliminary cross-modal association tensor to generate a cross-modal association tensor.
[0169] Optionally, step 223 performs domain knowledge graph constraint and semantic completion processing on the cross-modal association tensor to generate a semantic fusion tensor, specifically: performing domain knowledge graph access and rule mapping processing on the cross-modal association tensor to generate a constraint rule set; performing graph neural network reasoning processing on the cross-modal association tensor and the constraint rule set to generate a semantic consistency score tensor; performing knowledge reasoning completion and health weighting processing on the semantic consistency score tensor to generate a semantic completion tensor; performing rule iteration and tensor fusion processing on the semantic completion tensor to generate a semantic fusion tensor.
[0170] Preferably, in step 221, for different modal data in the spatiotemporal feature tensor, key features are enhanced through a customized attention mechanism. The specific implementation is as follows: for the numerical features in the spatiotemporal feature tensor (such as current, voltage, temperature, etc.), a two-layer attention architecture is adopted. The self-attention module of the Transformer encoder captures the long-term and short-term dependencies in the time series (such as the temporal correlation of "load increase → oil temperature lag increase"). At the same time, a spatial self-attention matrix is constructed in combination with the device topology knowledge to calculate the correlation between the values of adjacent devices at the same time (such as the impact weight of bus current anomalies on adjacent transformers). Dynamic threshold gating is introduced. When the value fluctuation exceeds the historical mean ±2σ, the attention weight of the corresponding time point is automatically increased, thereby generating a numerically enhanced feature tensor that focuses on abnormal changes.
[0171] Preferably, for image features in the spatiotemporal feature tensor (such as equipment inspection images), YOLOv8 is first used to locate key equipment areas in the image (such as transformer bushings and cable connectors), and then the SENet channel attention mechanism is used to calculate the importance of each feature channel. For example, in infrared thermal images, the thermal feature channel in the temperature anomaly area is assigned a weight of more than 0.8, and the background environment channel weight is suppressed to less than 0.2. At the same time, the spatial attention mask is used to further focus on the defect location, thereby generating an image enhancement feature tensor.
[0172] Preferably, the text features in the spatiotemporal feature tensor (such as operation and maintenance logs and fault report text) are extracted based on the BERT-ERNIE pre-trained model, and entities such as device ID and fault type are constructed in the text. A multi-head self-attention mechanism is used to construct an entity association graph (such as the correlation strength between "main transformer T1" and "overtemperature alarm"). A knowledge bias in the power domain is introduced, and a weight offset of 0.3 is applied to key fault terms specified in the "Guidelines for the Status Assessment of Power Equipment" (such as "inter-turn short circuit") to generate a text-enhanced feature tensor that highlights the fault semantics.
[0173] Preferably, for the audio features in the spatiotemporal feature tensor (such as the device operating audio), the audio signal is converted into a time-frequency diagram through MFCC, and the CBAM convolutional attention module is used to capture the timing pattern of local discharge pulses in the time dimension (such as a high-frequency pulse once every 10 seconds) and lock the discharge feature frequency band (30-300kHz) in the frequency dimension. By comparing the spectral distribution of normal device audio, the attention weight of the frequency band that deviates from the mean is automatically increased, thereby generating an audio enhancement feature tensor.
[0174] Finally, the above four types of enhanced tensors are concatenated according to the modal dimension to obtain the modal enhanced feature tensor containing the key information of each modality.
[0175] Preferably, in step 222, a cross-modal attention network is constructed to achieve deep correlation of multimodal features of the modality-enhanced feature tensor. The specific implementation steps are as follows:
[0176] For modality-enhanced feature tensors (including numerical, image, text, and audio), a four-layer cross-modal attention network is designed. Each layer contains an eight-head attention mechanism, each focusing on a specific modality pair (such as "numerical value-image" and "text-audio"). The network is trained through contrastive learning. For example, the attention weight of "abnormal oil temperature value" and "infrared hot spot image" is forced to be higher than the random combination. This generates a dynamic attention weight matrix, in which the element value represents the strength of the association between modalities (such as the numerical value-image association weight of 0.72).
[0177] The attention weight matrix and modality-enhanced feature tensor are weighted element-by-element, dynamically adjusting the contribution of each modality based on real-time data. When an "urgent defect" occurs in a text log, the image feature weight is automatically increased to 0.6, and the numerical feature weight is adjusted to 0.3. Tensor concatenation and residual connections are used to generate a preliminary cross-modal correlation tensor.
[0178] A triplet loss function is introduced to the initial cross-modal association tensor to verify the semantic logic between the modalities. If the image display device is normal but the value exceeds the threshold, the semantic contradiction loss between the two is calculated and backpropagated to optimize the attention parameter. Combined with a reinforcement learning strategy, iterative optimization is performed using expert-annotated cross-modal association examples (such as "oil temperature 90°C + hot spot image + overtemperature warning text") as a reward signal, resulting in a logically consistent cross-modal association tensor.
[0179] Preferably, in step 223, the cross-modal association tensor is verified and enhanced in combination with the knowledge graph in the electric power field, specifically including the following steps:
[0180] First, for the cross-modal association tensor, a knowledge graph containing equipment topology and fault rules is imported, and the production rules (such as "the cooling system should be started when the main transformer oil temperature is ≥85°C") are mapped into computable constraints to generate a constraint rule set. For example, the logical rule "oil temperature exceeds the threshold → cooling system status" is converted into a Boolean expression tensor.
[0181] Furthermore, the cross-modal association tensor and constraint rule set are fed into a GNN network, where tensor elements serve as nodes and modal associations serve as edges to construct a computational graph. Graph convolutions are used to infer data compliance, such as determining whether the association "oil temperature 92°C and cooling system not started" violates the rules. This generates a semantic consistency score tensor, where each element represents the probability of compliance (e.g., 0.92 indicates high compliance).
[0182] Furthermore, for associations with a semantic consistency score tensor score below 0.7 (e.g., "casing crack image" has no corresponding text record), knowledge graph reasoning and completion are performed to retrieve the transmission path from "casing crack → partial discharge → audio anomaly" to generate potential association candidates. The completed tensor is weighted based on the device health index (HI), with the association weight for devices with an HI < 0.5 increased by 0.2. This ultimately generates a semantic completion tensor.
[0183] For the semantic completion tensor and the original rule set, if a new association is found during the completion process (such as "abnormal hydrogen tank pressure → specific audio signal"), the knowledge graph will be updated and the constraint rule set will be reconstructed. Finally, the semantic completion tensor and the original rule set will be fused through the tensor dot product operation to generate a semantic fusion tensor that meets industry standards, providing structured semantic features for subsequent risk prediction.
[0184] Optionally, step 23: performing knowledge graph constraint and feature dimensionality reduction processing on the semantic fusion tensor to generate a spatiotemporal semantic fusion feature vector specifically includes the following steps:
[0185] Step 231: Perform domain knowledge graph constraint verification processing on the semantic fusion tensor to generate a constraint verification feature tensor;
[0186] Step 232: Perform attention pooling and dimensionality reduction processing on the constraint check feature tensor to generate a low-dimensional feature vector set;
[0187] Step 233: Perform health weighting and semantic enhancement processing on the low-dimensional feature vector set to generate a spatiotemporal semantic fusion feature vector.
[0188] Optionally, the step 231 performs domain knowledge graph constraint verification processing on the semantic fusion tensor to generate a constraint verification feature tensor, specifically: performing domain knowledge graph rule extraction and mapping processing on the semantic fusion tensor to generate a constraint rule tensor; performing graph neural network inference processing on the semantic fusion tensor and the constraint rule tensor to generate a consistency scoring matrix; performing weighted fusion processing on the consistency scoring matrix and the semantic fusion tensor to generate a constraint verification feature tensor.
[0189] Optionally, step 232 performs attention pooling and dimensionality reduction processing on the constraint check feature tensor to generate a low-dimensional feature vector set, specifically: performing multi-head attention pooling processing on the constraint check feature tensor to generate spatiotemporal semantic aggregation features; performing knowledge distillation dimensionality reduction processing on the spatiotemporal semantic aggregation features to generate semantic compression feature vectors; performing consistency check and health weighting processing on the semantic compression feature vectors to generate a low-dimensional feature vector set.
[0190] Optionally, the step 233, the health degree weighting and semantic enhancement processing are performed on the low-dimensional feature vector set to generate a spatiotemporal semantic fusion feature vector, specifically: performing device health degree dynamic weighting processing on the low-dimensional feature vector set to generate a health degree enhanced feature vector; performing spatiotemporal semantic embedding processing on the health degree enhanced feature vector to generate a spatiotemporal semantic feature vector; performing domain knowledge graph association and confidence score processing on the spatiotemporal semantic feature vector to generate a spatiotemporal semantic fusion feature vector.
[0191] Preferably, in step 231, the semantic fusion tensor is taken as the processing object, and by incorporating the rules and logic of the power domain knowledge graph, a constraint verification feature tensor is generated, and the specific implementation process is as follows:
[0192] Firstly, the knowledge contained in the power domain knowledge graph, such as device topology structure, fault conduction rules, industry standard specifications (such as threshold limits in the “Power Equipment Operation Regulations”), etc., is parsed, and the computable rules (such as “the temperature of the transformer oil exceeds 85℃ and lasts for 10 minutes to trigger a level one alarm”) are extracted, and these rules are encoded as constraint rule tensors. In the encoding process, the conditions, parameters, and logical relationships in the rules are converted into elements and operator symbols in the tensor, for example, the temperature threshold, time threshold, etc. parameters are taken as specific dimension values of the tensor, and logical judgments (such as “and” “or” relationships) are converted into tensor operation instructions.
[0193] Secondly, the semantic fusion tensor and the constraint rule tensor are jointly input into a graph neural network (GNN). The data features in the semantic fusion tensor are taken as nodes, and the association relationships between the modal data are taken as edges to construct a computation graph; the constraint rule tensor is taken as the basis and constraint condition for reasoning. Through graph convolution operation, the compliance of each data association in the semantic fusion tensor is analyzed, for example, whether the data association “the temperature of the transformer oil reaches 90℃ and no alarm record” conforms to the rules. After reasoning, a consistency score matrix is generated for each data association, and each element in the matrix corresponds to the compliance probability (value range 0-1) of a data association, such as 0.9 indicating high compliance with the rules and 0.2 indicating obvious contradictions.
[0194] Finally, according to the score results in the consistency score matrix, the semantic fusion tensor is weighted and adjusted. For data associations with high scores (close to 1), the original feature weights are retained; for data associations with low scores (close to 0), the feature weights are reduced, and the score matrix and the semantic fusion tensor are fused through tensor element multiplication. Finally, a constraint verification feature tensor is generated, which not only contains the feature information of the original semantic fusion tensor, but also embodies the constraint verification results of the domain knowledge graph, highlighting the data features that conform to the rules and suppressing the data features that are contradictory or abnormal.
[0195] Preferably, in step 232, for the constraint check feature tensor, a low-dimensional feature vector set is generated by reducing the feature dimension and retaining key information, specifically including the following steps:
[0196] First, the constraint-check feature tensor is fed into a multi-head attention network, which aggregates the spatiotemporal semantic features within the tensor from multiple perspectives. Each attention head focuses on a different feature subspace, for example, some focusing on time series features, others on spatial topological associations, and still others on cross-modal semantic relationships. Through the multi-head attention mechanism, the importance weights of different feature regions are calculated and the tensor is weightedly aggregated. This compresses the high-dimensional constraint-check feature tensor into spatiotemporal semantic aggregate features, reducing feature dimensionality and redundant information while retaining key information.
[0197] Secondly, knowledge distillation techniques are used to transfer the knowledge of spatiotemporal semantic aggregation features to a low-dimensional space. A pre-trained high-dimensional feature model is used as the teacher model, and a simpler low-dimensional model is constructed as the student model. By minimizing the difference between the outputs of the teacher and student models (such as mean squared error and KL divergence), the student model learns the core semantic information of the high-dimensional features and further compresses the spatiotemporal semantic aggregation features into a semantically compressed feature vector, significantly reducing the feature dimensionality while maintaining the semantic expressiveness of the data.
[0198] Finally, the semantically compressed feature vector is checked for consistency, examining whether the features within the vector contain logical contradictions or conflicts with domain knowledge. For example, if the vector contains contradictory semantic features indicating both normal device operation and fault alarms, these features are corrected or reweighted. Furthermore, the semantically compressed feature vector is weighted based on the device health assessment results, assigning higher weights to features indicating low device health to highlight potential risk information. This ultimately generates a low-dimensional feature vector set with streamlined dimensions that contains key spatiotemporal semantic information and features associated with device health status.
[0199] Optionally, step 3, inputting the spatiotemporal semantic fusion feature vector into the multimodal fusion model for forward transmission to predict the power operation and maintenance risk and obtain the power operation and maintenance risk prediction result, specifically includes the following steps:
[0200] Step 31: Perform cross-modal feature fusion and temporal dynamic reasoning processing on the spatiotemporal semantic fusion feature vector to generate a risk feature representation;
[0201] Step 32: Perform domain knowledge graph enhancement and graph neural network reasoning on the risk feature representation to generate a risk association graph;
[0202] Step 33: Perform hierarchical risk scoring and multi-label classification processing on the risk association map to generate power operation and maintenance risk prediction results.
[0203] Optionally, the multimodal fusion model includes a constructed multi-layer Transformer encoder, a temporal dynamic reasoning layer, and a feature fusion layer; step 31, performing cross-modal feature fusion and temporal dynamic reasoning processing on the spatiotemporal semantic fusion feature vector to generate a risk feature representation, specifically includes the following steps:
[0204] Step 311: Based on the constructed multi-layer Transformer encoder, cross-modal feature fusion is performed on the spatiotemporal semantic fusion feature vector to generate a cross-modal fusion feature;
[0205] Step 312: Based on the constructed temporal dynamic reasoning layer, temporal dynamic reasoning is performed on the spatiotemporal semantic fusion feature vector to generate temporal dynamic features;
[0206] Step 313: Based on the feature fusion layer, the cross-modal fusion features and the temporal dynamic features are fused to generate a risk feature representation.
[0207] Preferably, in step 311, the spatiotemporal semantic fusion feature vector is used as the processing object, and a multi-layer Transformer encoder is used to implement deep fusion of multimodal features, which specifically includes the following steps:
[0208] First, the spatiotemporal semantic fusion feature vector (containing multimodal semantic information such as numerical values, images, text, and audio) is input into a multi-layer Transformer encoder. This encoder uses a multi-head attention mechanism, with each attention head independently calculating the association weights between different modal features. For example, when processing power equipment data, one attention head focuses on the association between "abnormal oil temperature values" and "hot spots in infrared thermal images," while another head focuses on the correspondence between "fault text descriptions" and "abnormal audio signals."
[0209] Second, through a self-attention mechanism, the encoder dynamically calculates the importance of each modal feature at each time step. For features containing historical equipment operating data, the model adjusts attention allocation based on dependencies within the time series (e.g., "continuously increasing load → gradually increasing oil temperature"), strengthening the weights of modal features with causal relationships. Furthermore, the position encoding layer incorporates temporal order information into the feature vector, ensuring that the model captures the temporal dynamics of multimodal data.
[0210] Finally, through iterative calculations in multiple layers of Transformers, feature information from different modalities is fused in the semantic space, forming cross-modal fusion features that contain cross-modal semantic associations. For example, the abnormal temperature value of the equipment, the corresponding thermal imaging image features, and the semantics of the fault description text are encoded into a unified feature vector, achieving deep integration and complementary expression of multimodal information.
[0211] Preferably, in step 312, the time series information in the spatiotemporal semantic fusion feature vector is used to mine the potential risk evolution law through the time series dynamic reasoning layer. The specific implementation steps are as follows:
[0212] First, the time series dynamic inference layer uses a hybrid architecture combining LSTM and Transformer. For short-term time series dependencies (such as second-level current fluctuations), the LSTM gating mechanism can effectively memorize and transmit key state information, capturing rapidly changing device parameter trends. For long-term dependencies (such as monthly device aging trends), the Transformer's self-attention mechanism can directly model the relationship between any time steps, avoiding the vanishing gradient problem in long sequences.
[0213] Secondly, the model incorporates a dynamic time window mechanism to automatically adjust the analysis period based on the device's operating status. For example, when a device is in a fault warning state, the time window is shortened to minutes to monitor parameter changes in real time; during normal operation, hourly or daily windows are used to analyze trend characteristics. Furthermore, by calculating statistics such as first-order differences and moving averages of time series characteristics, the model's sensitivity to abnormal fluctuations is enhanced.
[0214] Finally, the inference layer outputs dynamic time series features. These features not only capture historical changes in equipment parameters but also predict potential trends in future time steps. For example, based on historical oil temperature data, the probability and magnitude of a temperature rise within the next 30 minutes can be predicted, providing a dynamic time series basis for risk assessment.
[0215] Preferably, in step 313, the cross-modal fusion features are deeply integrated with the temporal dynamic features to generate a risk feature representation for risk assessment. The specific implementation steps are as follows:
[0216] First, the feature fusion layer uses a gated fusion mechanism, dynamically adjusting the fusion ratio of cross-modal features and time series features through learnable gating units. For example, when a sudden abnormality occurs in the equipment (such as a sudden current drop), the weight of the time series dynamic features is increased to highlight short-term risk signals. If there are persistent risks (such as insulation aging trends), the influence of cross-modal fusion features is enhanced, and multimodal evidence is integrated to assess risks.
[0217] Secondly, the two types of features are fused using tensor concatenation and a multi-layer perceptron (MLP). The cross-modal fusion features are first concatenated with the temporal dynamic features in a dimensionally oriented manner to form a joint feature vector containing spatial, semantic, and temporal information. The MLP then performs a nonlinear transformation on this joint vector to extract high-order feature interaction information. For example, the "historical oil temperature rise trend" (a temporal feature) is fused with the "hot spot distribution in infrared images" (a cross-modal feature) to uncover potential correlations between temperature anomalies and physical damage to equipment.
[0218] Finally, the risk feature representation is output, which integrates the spatial semantic information and temporal evolution laws of multimodal data, characterizes the key characteristics of equipment operation risks in a compact vector form, and provides a unified input basis for subsequent risk association map construction and risk scoring.
[0219] Optionally, the multimodal fusion model includes a knowledge graph embedding layer, a graph neural network layer, and a semantic verification layer; step 32, performing domain knowledge graph enhancement and graph neural network reasoning processing on the risk feature representation to generate a risk association graph, specifically includes the following steps:
[0220] Step 321: Based on the knowledge graph embedding layer, perform rule extraction and embedding fusion on the risk feature representation to generate knowledge-enhanced features;
[0221] Step 322: Based on the graph neural network layer, perform graph modeling and reasoning on the knowledge enhancement features to generate a risk association graph structure;
[0222] Step 323: Based on the semantic verification layer, perform semantic verification and weight optimization on the risk association graph structure to generate a risk association graph.
[0223] Preferably, in step 321, first, structured rules are extracted from the power domain knowledge graph. For device topology relationships, connection rules such as "main transformer T1 is connected to circuit breaker CB1 via busbar B1" are parsed and converted into a triplet of `<main transformer T1, connection, busbar B1>`. For fault conduction logic, "oil temperature continuously exceeding the threshold for 2 hours will cause insulation aging" is converted into a conditional expression. Industry standards and specifications, such as "casing surface cracks exceeding 3mm in length require immediate repair" in DL / T 596, are mapped into numerical constraints. Simultaneously, using models such as TransE-Power, entities (such as "main transformer T1" and "insulation aging") and relationships (such as "lead to" and "connection") in the knowledge graph are mapped into low-dimensional embedding vectors to prepare for subsequent fusion. Secondly, the risk feature representation is fused with the knowledge graph embedding. Using a multi-head attention mechanism, the association weights between the risk feature vector and the embedding vectors of each knowledge entity are calculated. For example, when the risk feature contains information about "oil temperature 95°C", the model will calculate its similarity with the embeddings of knowledge entities such as "overtemperature fault" and "insulation aging". If the similarity score with "overtemperature fault" reaches 0.85, a higher attention weight will be given to this entity. Subsequently, with the help of the self-attention mechanism, the importance of each knowledge entity is dynamically adjusted at different time steps, and the position encoding information is combined to ensure the accuracy of feature fusion in the time dimension. Finally, after multiple layers of calculation and iteration, the risk feature representation is fused with the high-weight knowledge entity embedding to generate a knowledge-enhanced feature. This feature not only retains the multimodal information in the original risk feature, but also incorporates professional knowledge in the power field. For example, the numerical feature of "abnormal oil temperature" is combined with the knowledge embedding of "overtemperature fault" to form a unified vector representation containing semantic information such as fault type and severity.
[0224] Preferably, in step 322, a risk association graph is first constructed based on the knowledge-enhanced features. The equipment, status, and event information in the knowledge-enhanced features are defined as graph nodes. For example, "main transformer T1," "oil temperature 95°C," and "alarm occurred at 10:00 AM on 2024-10-01" are each considered independent nodes. Based on device topology, fault transmission rules, and temporal relationships, graph edges are constructed, such as the "connection" edge between "main transformer T1" and "busbar B1" and the "cause" edge between "oil temperature 95°C" and "insulation aging." Initial weights are assigned to represent the closeness of these connections. Next, a graph neural network (such as GATv2) is used to reason about the risk association graph. Using a multi-head attention mechanism, each node aggregates feature information from its neighboring nodes. For example, the "main transformer T1" node will integrate the features of its connected neighboring nodes, such as "busbar B1" and "oil temperature 95°C," and combine these features with attention weights to calculate an updated node representation. Furthermore, the model dynamically adjusts edge weights based on factors such as device health. For devices with lower health, the weights of their associated risk transmission edges are increased accordingly. Furthermore, through graph convolution, risk features are propagated within the graph structure, inferring potential risk nodes and transmission paths. For example, the risk of "winding burnout" can be inferred from the "oil temperature 95°C" node and added to the graph structure. Finally, after multiple rounds of graph neural network calculations and updates, a risk association graph structure containing rich risk association information is generated. This structure clearly displays the topological relationships, causal connections, and temporal evolution between devices, states, and events, providing an intuitive graphical model foundation for subsequent risk assessment.
[0225] Preferably, in step 323, first, a semantic check is performed on the risk association graph structure. The node and edge information in the graph is compared with knowledge rules in the power sector. Using rules defined in the OWL ontology language (e.g., "Overtemperature and a normal cooling system should trigger a sensor fault alarm"), the graph associations are checked for contradictions or irrationalities. Cross-modal semantic alignment is also performed. For example, the semantic similarity between the numerical feature of "oil temperature 95°C" and the visual features of the hot spot area in the infrared thermal image is calculated. If the similarity falls below a threshold, a semantic conflict is flagged. Second, weight optimization is performed based on the semantic check results. Edges with semantic contradictions have their weights reduced; while associations that conform to domain knowledge and have been verified to be correct are appropriately weighted. A reinforcement learning algorithm is employed, using historical risk cases annotated by experts as reward signals. For example, if the model correctly identifies the risk association of "casing cracks leading to partial discharge," a positive reward is given; otherwise, negative feedback is given. Edge weights are adjusted through continuous iterative optimization of the graph neural network parameters. Furthermore, for newly discovered risk associations not documented in the knowledge graph, initial weights are generated through few-shot learning and dynamically updated as data accumulates. Finally, after semantic validation and weight optimization, the risk association graph structure is normalized, low-confidence edges are removed, and the core risk transmission paths are retained to generate the final risk association graph. Each edge in this graph is accompanied by a semantic confidence score and rule traceability information, and each node contains a complete description of the risk characteristics, providing a reliable basis for the accurate prediction of power operation and maintenance risks.
[0226] Optionally, the multimodal fusion model includes a hierarchical scoring layer, a multi-label classification layer, and an uncertainty quantification layer; step 33, performing hierarchical risk scoring and multi-label classification processing on the risk association map to generate a power operation and maintenance risk prediction result, includes the following steps:
[0227] Step 331: Based on the multi-label classification layer, perform fault type probability prediction on the risk association map to generate a primary classification result;
[0228] Step 332: Based on the hierarchical scoring layer, perform a risk level assessment on the primary classification results to generate a high-level risk score.
[0229] Step 333: Based on the uncertainty quantification layer, confidence assessment and report generation are performed on the risk scoring results to generate power operation and maintenance risk prediction results.
[0230] Preferably, in step 331, the node features and edge weight information in the risk association graph are first input into a multi-label classification layer. This layer utilizes a multi-label classification model based on a graph neural network. This layer uses the graph structural features extracted by the GNN as input, combined with node attribute information (such as equipment type, operating parameters, and fault symptom descriptions). A multi-layer perceptron (MLP) performs feature transformation, mapping the graph structural information into a semantic space of fault types. Next, a Softmax function is used to calculate the probability score for each fault type. The model predefines a set of labels for common fault types in power equipment, such as "insulation aging," "partial discharge," and "overload." For example, for a risk association graph containing nodes with "persistently high oil temperature" and "abnormal infrared hot spots," the model calculates the similarity between the features and each label and concludes that the probability of "insulation aging" is 0.75, the probability of "overload" is 0.6, and the probability of "partial discharge" is 0.3. Furthermore, an attention mechanism is introduced to assign higher weights to nodes and edges in the graph that are closely associated with specific fault types, thereby strengthening the influence of key risk features on the classification results. Finally, the primary classification result is output. This result contains the probability vector of each fault type, such as `[insulation aging: 0.75, overload: 0.6, partial discharge: 0.3]`, which represents the possibility of different faults occurring in the equipment reflected by the spectrum and provides basic data for subsequent risk level assessment.
[0231] Preferably, in step 332, first, a hierarchical risk scoring system is constructed. Power operation and maintenance risks are divided into multiple levels, such as "Emergency," "Critical," "General," and "Caution," with each level corresponding to different risk thresholds and handling strategies. Based on industry standards (such as DL / T 596) and historical fault data, a basic risk score is set for each fault type and weighted based on the probability of the fault. For example, the basic score for an "insulation aging" fault is 80 points. If its probability is 0.75, the adjusted score is 80 * 0.75 = 60 points. Next, the analytic hierarchy process (AHP) is used to comprehensively assess the risk levels of multiple fault types. This considers the correlation and cumulative effects between fault types. For example, when "insulation aging" and "partial discharge" occur simultaneously, the risk level needs to be nonlinearly increased based on the individual fault scores. The final risk level is determined by calculating the weighted sum of each fault score and comparing it with the hierarchical threshold. If the total score reaches 70 points and exceeds the "Critical" level threshold (65 points), the risk level is assessed as "Critical." Finally, an advanced risk score is output, which includes a risk level label (such as "serious") and a quantitative score (such as 70 points), clearly reflecting the severity of the current operating risk of the equipment and providing a quantitative basis for operation and maintenance decisions.
[0232] Preferably, in step 333, the uncertainty of the risk score is first assessed. Methods such as Monte Carlo dropout and Bayesian neural networks are used to quantify parameter and data uncertainty in the model prediction process. For example, neurons in the neural network are randomly discarded multiple times to simulate risk scoring results under different model parameters, and the variance and confidence interval of the score are calculated. If the mean of the "critical" risk score in 100 simulations is 70 points and the standard deviation is 5 points, the confidence interval is [65, 75], indicating that the risk assessment result has a certain range of fluctuation. Next, a power operation and maintenance risk prediction report is generated. The report includes the risk level, quantitative score, confidence interval, and risk description. For example, "The current risk level of main transformer T1 is 'critical', with a score of 70 points (confidence interval [65, 75]). There is a risk of insulation aging (probability 0.75). Immediate infrared retesting and oil chromatography analysis are recommended." Furthermore, a historical case library is used to provide experience and reference solutions for similar risk scenarios, enhancing the practicality of the report. Finally, the power operation and maintenance risk prediction results are output in the form of a structured report, covering the quantitative conclusions of risk assessment, uncertainty analysis and operation and maintenance recommendations, helping operation and maintenance personnel to fully understand the equipment risk status and formulate targeted strategies.
[0233] The above is merely one specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of the present application. Therefore, the scope of protection of the present application should be based on the scope of protection of the claims.
Claims
1. A method for predicting power operation and maintenance risks based on a multimodal fusion large model, characterized in that: include: Step 1: Collect multi-source heterogeneous data related to power operation and maintenance to form a normalized multimodal dataset; Step 2: Generate spatiotemporal semantic fusion feature vectors based on the normalized multimodal dataset; Step 3: Input the spatiotemporal semantic fusion feature vector into the multimodal fusion model for forward transmission to predict the power operation and maintenance risk and obtain the power operation and maintenance risk prediction result.
2. The method according to claim 1, characterized in that The step 1 comprises: Step 11: Based on the configured multi-source heterogeneous data interface, collect multi-source heterogeneous data related to power operation and maintenance from multi-source heterogeneous data sources to form an original multi-source data set; Step 12: Process the original multi-source dataset to generate a cleaned multi-source dataset; Step 13: Processing the structured data in the cleaned multi-source dataset to generate a standardized structured dataset; Step 14: Process the unstructured data in the cleaned multi-source dataset to generate a labeled unstructured dataset; Step 15: Process the standardized structured dataset and the annotated unstructured dataset to generate a normalized multimodal dataset.
3. The method according to claim 2, characterized in that The step 11 comprises: Step 111: Based on the configured multi-source heterogeneous data interface, real-time data from power equipment sensors, operation and maintenance log text, inspection images / videos, equipment operation audio, historical fault reports, and industry standard documents are collected from multiple heterogeneous data sources to obtain multi-source heterogeneous data; Step 112: Based on the configured data integration framework, perform protocol type parsing on the multi-source heterogeneous data, convert the multi-source heterogeneous data into structured units in a unified format, and annotate the data with protocol type tags to generate protocol-parsed data units. Step 113: Based on the data integration rules, the data units after protocol parsing are integrated and indexed to generate an original multi-source data set associated by device ID and time dimension and carrying a temporary index.
4. The method according to claim 2, characterized in that The step 12 includes: Step 121: Perform sub-modal mixed anomaly detection processing on the original multi-source dataset to generate a sub-modal preliminary cleaned dataset; Step 122: Perform cross-modal consistency graph verification on the sub-modal preliminary cleansing dataset to generate a consistency-verified dataset; Step 123: Perform dynamic threshold reinforcement learning optimization processing on the consistency-checked dataset to generate a cleaned multi-source dataset.
5. The method according to claim 4, characterized in that The step 12 comprises: Step 1211: For the structured numerical data in the original multi-source dataset, the temporal dependency features in the structured numerical data are captured based on the Transformer-Isolation Forest hybrid model. The "degree of isolation" of the data is scored using the Isolation Forest algorithm based on the temporal dependency features to identify and filter out noise data that deviates from the normal distribution, thereby obtaining a structured numerical cleansing subset. Step 1212: For the text data in the original multi-source dataset, based on the domain model semantic verification, the semantic similarity between the text data is calculated to determine the noise data with semantic contradictions or abnormal formats, and filter them to obtain the text data cleaning subset; Step 1213: For the image / audio data in the original multi-source dataset, a multimodal feature fusion detector is used to combine the visual features extracted by the CNN with the MFCC acoustic features, and then input them into the GAN network to identify and filter media data with abnormal clarity or incoherent features to obtain a multimedia data subset. Step 1214: align and merge the structured numerical cleaning subset, the text data cleaning subset, and the multimedia data subset according to the spatiotemporal dimensions to generate a sub-modal preliminary cleaning data set.
6. The method according to claim 4, characterized in that The step 122 includes: Step 1221: Perform dynamic multimodal knowledge graph construction processing on the sub-modal preliminary cleansing data set to generate a real-time associated knowledge graph; Step 1222: Perform graph neural network consistency reasoning on the real-time associated knowledge graph to generate a consistency score graph; Step 1223: Perform dynamic threshold filtering and knowledge completion processing on the consistency score map to generate a consistency-verified data set.
7. The method according to claim 4, characterized in that Step 123, performing dynamic threshold reinforcement learning optimization processing on the consistency-checked data set to generate a cleaned multi-source data set, specifically includes the following steps: Step 1231: Perform historical threshold pattern retrieval and scene matching processing on the consistency-checked data set to generate a dynamic threshold configuration set; Step 1232: Perform real-time data feedback fine-tuning and abnormal circuit breaking processing on the dynamic threshold configuration set to generate a cleaned multi-source data set.
8. The method according to claim 2, characterized in that The step 13 comprises: Step 131: Perform feature dimensional distribution modeling on the structured data in the cleaned multi-source dataset to generate a feature dimensional distribution matrix; Step 132: performing a standardized strategy tensor generation process on the characteristic dimension distribution matrix to generate a standardized strategy tensor; Step 133: Perform tensor conversion and semantic verification on the standardized policy tensor and structured data to generate a standardized structured data set.
9. The method according to claim 2, characterized in that Step 2 includes: Step 21: performing spatiotemporal feature extraction and tensor encoding processing on the normalized multimodal dataset to generate a spatiotemporal feature tensor; Step 22: Perform cross-modal semantic fusion processing on the spatiotemporal feature tensor to generate a semantic fusion tensor; Step 23: Perform knowledge graph constraint and feature dimensionality reduction processing on the semantic fusion tensor to generate a spatiotemporal semantic fusion feature vector.
10. The method according to claim 2, characterized in that Step 3 includes: Step 31: Perform cross-modal feature fusion and temporal dynamic reasoning processing on the spatiotemporal semantic fusion feature vector to generate a risk feature representation; Step 32: Perform domain knowledge graph enhancement and graph neural network reasoning on the risk feature representation to generate a risk association graph; Step 33: Perform hierarchical risk scoring and multi-label classification processing on the risk association map to generate power operation and maintenance risk prediction results.
Citation Information
Cited By
Multi-mode sensing fusion electric power operation and maintenance safety early warning method and system
CN120952556A
Ocean three-dimensional temperature field reconstruction method and system
CN121350580A
A method and system for reconstructing a three-dimensional temperature field of the ocean
CN121350580B
Acoustic detection method and system for bolt axial force of wind driven generator
CN121453263A
Power edge monitoring method and system based on multi-modal collaboration and environmental knowledge enhancement
CN122313405A