Fusion and management system for multi-source heterogeneous science and technology information resources

By constructing a knowledge graph and multimodal semantic fusion module and combining it with graph neural networks, the problems of detection accuracy and reasoning adaptability of multi-source heterogeneous data in the power grid are solved, efficient fault diagnosis and preventive maintenance are achieved, and the robustness and operation and maintenance efficiency of the power grid system are improved.

CN120654190APending Publication Date: 2025-09-16SUN YAT SEN UNIV
View PDF 0 Cites 12 Cited by

Patent Information

Application Number
CN202510775899.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-06-11
Publication Date
2025-09-16

AI Technical Summary

Technical Problem

When existing technologies use graph neural networks in power grids to detect conflicts in the feature space of multi-source heterogeneous data resources, there are problems with data quality and poor adaptability of inference rules, resulting in decreased detection accuracy and inaccurate tracing of the root causes of data conflicts.

Method used

A knowledge graph module is constructed to extract fault chains and parameter constraints from historical operation and maintenance logs and equipment manuals, and a hypergraph structure is generated by combining real-time sensor data. Data standardization and semantic mapping are performed through a multimodal semantic fusion module, and graph neural networks are used to trace feature contradictions and locate the root cause equipment.

Benefits of technology

It improves the robustness and reasoning ability of the power grid system for sparse or abnormal data, enhances the depth and breadth of equipment status perception, realizes efficient fault diagnosis and preventive maintenance, and improves system reliability and operation and maintenance efficiency.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120654190A_ABST
    Figure CN120654190A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of information resource fusion management, and particularly discloses a fusion and management system for multi-source heterogeneous science and technology information resources. Analyzing an equipment fault chain from an unstructured text of a historical operation and maintenance log, extracting a rated parameter constraint from a structured table of an equipment manual, and collecting an operation feature vector from a real-time sensing data stream to generate a knowledge graph containing N entity relationships; based on an entity attribute constraint rule of the knowledge graph, designing a bidirectional attention mapping network to calculate semantic similarity weights of multi-source data and knowledge nodes, and generating a graph embedding vector set with weight marks through Hadamard product operation; according to the method, the embedded vector set is input into the pre-trained graph neural network model, and the root cause equipment set causing feature offset is positioned, so that efficient fault diagnosis and positioning are realized, decision support is provided for a subsequent preventive maintenance strategy, and the reliability and the operation and maintenance efficiency of the system are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention belongs to the technical field of information resource fusion management, and relates to a fusion and management system for multi-source heterogeneous scientific and technological information resources. Background Art

[0002] In modern power systems, power grids face the challenge of dealing with multi-source, heterogeneous data resources. This data, sourced from diverse sensors, monitoring equipment, and information systems, is diverse and complex. Therefore, employing graph neural networks to detect conflicts in feature spaces and tracing the root causes of data conflicts through path reasoning is crucial and necessary. First, power grid operation and management rely on accurate data analysis. Any data conflict can lead to erroneous decisions, impacting grid security and stability. By analyzing the feature spaces of multi-source data using graph neural networks, we can effectively identify conflicts and inconsistencies, and then trace the root causes of data conflicts through path reasoning. This process not only improves data quality but also provides a reliable basis for optimized grid scheduling and fault diagnosis, ensuring efficient power system operation.

[0003] However, current technologies still suffer from significant flaws and drawbacks when using graph neural networks to detect feature space conflicts in multi-source, heterogeneous data resources within power grids. First, the performance of graph neural networks often relies on the quality and quantity of training data. However, in practical power grid applications, data often suffers from missing, noisy, and inconsistent data. These data quality issues directly impact training effectiveness, resulting in reduced model accuracy in detecting feature space conflicts. Second, existing technologies often rely on fixed inference rules and graph structures during path reasoning, which makes them less adaptable to the dynamically changing power grid environment. During power grid operation, system states and data flows may change frequently, and fixed inference rules may not be able to reflect these changes in a timely manner, resulting in inaccurate data conflict tracing. Furthermore, the lack of sufficient contextual information during path reasoning can lead to an incomplete understanding of the root causes of data conflicts, making it impossible to effectively resolve practical problems. Summary of the Invention

[0004] In view of the above problems existing in the prior art, the present invention provides a fusion and management system for multi-source heterogeneous scientific and technological information resources, which is used to solve the above technical problems.

[0005] In order to achieve the above-mentioned and other purposes, the technical solutions adopted by the present invention are as follows: The present invention provides a fusion and management system for multi-source heterogeneous scientific and technological information resources, the system comprising: The knowledge graph construction module parses equipment failure chains from the unstructured text of historical operation and maintenance logs, extracts rated parameter constraints from structured tables in equipment manuals, and collects operational feature vectors from real-time sensor data streams. The hypergraph storage engine is then used to fuse device topological connections with failure mode causal relationships to generate a knowledge graph containing N entity relationships. Multimodal semantic fusion module: Based on the entity attribute constraints of the knowledge graph, it implements sliding window normalization on temperature sensor time series data, performs wavelet time-frequency transform on vibration signals to generate spectrograms, and extracts current harmonic feature vectors. It also designs a bidirectional attention mapping network to calculate the semantic similarity weights between multi-source data and knowledge nodes, and generates a weighted graph embedding vector set through Hadamard product operations. Feature contradiction tracing module: The embedded vector set is input into the pre-trained graph neural network model to construct a multi-dimensional feature space distribution map of the device status; the cosine similarity matrix between nodes of similar devices is calculated to identify abnormal node pairs below the set threshold; the conduction path is inferred based on the topological connection rules of the knowledge graph to locate the root cause device set that causes feature deviation.

[0006] Exemplarily, the operation logic of the knowledge graph construction module includes: Step S11: Perform unstructured text parsing on historical operation and maintenance logs, extract entity relationship triples {equipment ID, fault type, maintenance measures} through the fault event causal relationship chain, and generate a fault chain entity relationship set; Step S12: parsing the structured table in the equipment manual, matching the rated voltage and maximum load rate parameter constraint rules based on regular expressions, and generating an equipment parameter constraint rule set; Step S13: collecting sensor data streams in real time, extracting transformer temperature mean, vibration spectrum main frequency, and current harmonic distortion rate features, and constructing a device operation status feature vector set; Step S14: establishing a dynamic conflict resolution mechanism. When it is detected that the deviation between the manual parameters and the statistical mean of the equipment operating status feature vector set exceeds a set threshold, the constraint rules are dynamically updated based on the sensor data in the sliding window. Step S15: The fault chain set, updated constraint rule set, and state feature set are integrated through the hypergraph storage engine to associate the topological connection relationship of the transformer nodes with the historical fault mode, and generate a knowledge graph containing N entity relationships, where each transformer entity node contains an attribute group including a load rate threshold, an ambient temperature and humidity coupling coefficient, and the number of maintenance times in the past three years.

[0007] Exemplarily, the operation logic of step S14 is: Step S141: Calculate the absolute deviation between the equipment manual parameters and the statistical mean of the equipment operation status feature vector set to generate a parameter deviation vector D; Step S142: When it is detected that the deviation of any dimension in D exceeds the set threshold, a real-time data collection instruction with a sliding time window T = 30 minutes is triggered; Step S143: Perform a time series stationarity test on the sensor data stream within the window T, and select a set Q of continuous sampling points that meets the normal distribution requirements; Step S144: Based on the statistical distribution characteristics of the set Q, a second-order polynomial fitting algorithm is used to generate a parameter correction coefficient matrix C; Step S145: injecting the correction coefficient C into the device parameter constraint rule set to dynamically update the rated voltage allowable range and maximum load rate threshold of the corresponding device; Step S146: Generate an updated constraint rule set P' and perform version binding marking with the knowledge graph node.

[0008] For example, the operation logic of the multimodal semantic fusion module is: Step S21: normalize the temperature sensor time series data according to a 5-minute sliding window, calculate the z-score value of the data in each window, and generate a standardized temperature matrix; Step S22: performing wavelet packet decomposition on the vibration sensor signal, extracting energy proportion characteristics of the 0 to 500 Hz frequency band, and generating a vibration time-frequency spectrum; Step S23: Collect current sensor readings, extract the amplitudes of the 3rd, 5th, and 7th harmonic components through fast Fourier transform, and construct a current harmonic feature vector; Step S24: Encode the device rated parameter node attributes in the knowledge graph into knowledge node vectors, and design a bidirectional attention network to calculate the three sets of semantic similarity weights between the temperature matrix, vibration spectrum, and harmonic vector and the knowledge nodes respectively; Step S25: Perform a Hadamard product operation on the weight matrix and the multimodal feature dataset to generate a weighted graph embedding vector set.

[0009] Exemplarily, the operation logic of step S24 includes: Step S241: extracting equipment rated parameter node attributes from the knowledge graph, including rated voltage value, maximum load rate, and allowable temperature rise threshold, and encoding them into numerical vectors of the same dimension; Step S242: normalizing the standardized temperature matrix generated by the temperature sensor, and calculating the root mean square value of the temperature change rate in each time window; Step S243: converting the vibration spectrum into a grayscale pixel matrix, and extracting image texture features as vibration representation vectors; Step S244: construct a bidirectional attention network model, set the temperature change rate, vibration texture characteristics, and current harmonic vector as the query end, and the knowledge node vector as the key value end; Step S245: Calculate the cosine similarity between the query end and the key-value end through the multi-head attention mechanism, and generate three sets of semantic similarity weights of the temperature matrix, vibration spectrum, harmonic vector and knowledge node respectively.

[0010] Exemplarily, step S25 further includes: Step S251: performing softmax normalization processing on the three sets of semantic similarity weights to output a semantic similarity weight matrix with unified dimension; loading the semantic similarity weight matrix generated by the bidirectional attention network and a multimodal feature dataset, wherein the multimodal feature dataset includes a normalized temperature matrix, a reference vibration time-frequency spectrum, and a reference current harmonic vector; Step S252: Verify the consistency of the weight matrix column dimension and the multimodal feature row dimension. If the dimensions do not match, trigger the feature vector interpolation and completion algorithm. Step S253: performing Hadamard product operations in the order of feature channels, the calculation method is to multiply the weight matrix elements by the corresponding position eigenvalues ​​bit by bit; Step S254: adding a weight source identifier to each product result, including three types of tags: {temperature weight, vibration weight, harmonic weight}; Step S255: Aggregate all labeled product results, reorganize the dimensions according to the device number and timestamp, and generate a graph embedding vector set.

[0011] Exemplarily, the feature conflict tracing module includes: Step S31: Loading a graph embedding vector set and a pre-trained graph neural network model, wherein the embedding vector set includes device temperature, vibration, harmonic characteristics and weight labels; Step S32: Input the embedded vector into the graph neural network for spatial mapping to generate a multi-dimensional feature space distribution graph, where each node represents the state vector of a single device; Step S33: Calculate the cosine similarity between the nodes of the same type of devices, and mark them as abnormal node pairs when the similarity is lower than the set threshold; Step S34: Draw the electrical connection conduction path of the abnormal node based on the device topology connection relationship in the knowledge graph; Step S35: performing reverse gradient propagation calculation along the conduction path to locate the initial abnormal device node that causes the feature shift; Step S36: Aggregate the initial abnormal nodes and their associated devices to generate a root cause device set and mark the deviation type.

[0012] Exemplarily, the sub-steps of step S34 include: Step S341: Loading the device topology connection relationship data and abnormal node set in the knowledge graph, obtaining the electrical connection direction information between the devices, where the electrical connection direction information is divided into transmission / distribution direction; Step S342: Starting from the abnormal node, perform a bidirectional breadth-first search along the electrical connection direction, and traverse a path with no more than 6 cascaded devices; Step S343: Filter the effective conduction paths according to the device type and remove the reactive compensation device path branches; the device types include transformers / circuit breakers / capacitors; Step S344: Perform physical connection verification on the screened conductive paths, detect impedance matching of adjacent device ports, and remove mismatched path segments; Step S345: Mark the valid conduction path as a topological link diagram with arrows, and generate conduction path data including {path length, device type sequence, impedance value}.

[0013] Exemplarily, the sub-steps of step S35 include: Step S351: Loading conduction path data and device characteristic offset values, and establishing a path node gradient calculation index table; Step S352: Initialize the back propagation gradient using the feature offset of the device node at the end of the path as the starting point; Step S353: Backtrack the calculation hop by hop along the conduction path, and update the formula as Δi = Δj × (impedance matching × 0.8 + device health × 0.2); where Δj represents the gradient value of the downstream node of the current node in the conduction path, and Δi represents the gradient value of the upstream node currently being calculated; Step S354: When it is detected that the gradient change rate between adjacent nodes exceeds a set threshold, the node is marked as a candidate abnormal source; Step S355: Select the candidate node with the largest cumulative gradient value and confirm it as the initial abnormal device node.

[0014] As described above, the present invention provides a fusion and management system for multi-source heterogeneous scientific and technological information resources, which has at least the following beneficial effects: This method analyzes equipment failure chains from historical operation and maintenance logs, extracts rated parameter constraints from equipment manuals, and combines real-time sensor data streams to collect operational feature vectors. This method constructs a hypergraph-structured knowledge graph that integrates equipment topological connectivity with the causal relationships between failure modes. This not only enables semantic modeling of complex industrial systems but also provides a highly interpretable and logically clear framework for subsequent analysis. Compared to traditional purely data-driven approaches, this knowledge-driven modeling approach significantly improves the system's robustness and reasoning capabilities when faced with sparse or anomalous data. Secondly, in the multimodal semantic fusion module, heterogeneous sensor data such as temperature, vibration, and current are normalized, wavelet transformed, and feature extracted. This data is then semantically mapped to entity nodes in the knowledge graph using a bidirectional attention mechanism, further enhancing the depth and breadth of the system's perception of device status. This module utilizes the Hadamard product operation to generate a weighted graph embedding vector set, enabling data from different sources to be integrated into a unified semantic space. This effectively addresses the semantic gap between heterogeneous data sources and improves the overall perception accuracy and generalization capabilities of the system. Finally, the feature conflict tracing module uses a graph neural network to construct a multidimensional feature space distribution map of device status. It then calculates cosine similarity to identify abnormal device nodes. Combined with the topological structure of the knowledge graph, it performs transmission path inference to accurately locate the root cause of feature deviations. This process not only enables efficient fault diagnosis and location but also provides decision support for subsequent preventive maintenance strategies, thereby improving system reliability and operational efficiency. BRIEF DESCRIPTION OF THE DRAWINGS

[0015] In order to more clearly illustrate the technical solutions of the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For ordinary technicians in this field, other drawings can be obtained based on these drawings without creative work.

[0016] Figure 1 Schematic diagram of the connection of various modules of the system of the present invention. DETAILED DESCRIPTION

[0017] The above contents described below in conjunction with the implementation of the present invention are merely examples and explanations of the concept of the present invention. Those skilled in the art may make various modifications or additions to the described specific embodiments or replace them in a similar manner. As long as they do not deviate from the concept of the invention or exceed the scope defined by the claims, they shall fall within the scope of protection of the present invention.

[0018] Example 1 See also Figure 1 As shown, a fusion and management system for multi-source heterogeneous scientific and technological information resources includes a knowledge graph construction module, a multimodal semantic fusion module, and a feature contradiction tracing module. The above modules are connected by wired and / or wireless connections to achieve data transmission between the modules; Knowledge graph construction module: parses equipment failure chains from the unstructured text of historical operation and maintenance logs, extracts rated parameter constraints from structured tables in equipment manuals, and collects operation feature vectors from real-time sensor data streams. Simultaneously, the hypergraph storage engine is used to fuse device topology connection relationships with failure mode causal relationships to generate a knowledge graph containing N entity relationships.

[0019] The operational logic of the knowledge graph construction module includes: Step S11: Perform unstructured text parsing on historical operation and maintenance logs, extract entity relationship triples {equipment ID, fault type, maintenance measures} through the fault event causal relationship chain, and generate a fault chain entity relationship set; Step S12: parsing the structured table in the equipment manual, matching the rated voltage and maximum load rate parameter constraint rules based on regular expressions, and generating an equipment parameter constraint rule set; Step S13: collecting sensor data streams in real time, extracting transformer temperature mean, vibration spectrum main frequency, and current harmonic distortion rate features, and constructing a device operation status feature vector set; Step S14: establishing a dynamic conflict resolution mechanism. When it is detected that the deviation between the manual parameters and the statistical mean of the equipment operating status feature vector set exceeds a set threshold, the constraint rules are dynamically updated based on the sensor data in the sliding window. Step S15: The fault chain set, updated constraint rule set, and state feature set are integrated through the hypergraph storage engine to associate the topological connection relationship of the transformer nodes with the historical fault mode, and generate a knowledge graph containing N entity relationships, where each transformer entity node contains an attribute group including a load rate threshold, an ambient temperature and humidity coupling coefficient, and the number of maintenance times in the past three years.

[0020] This embodiment of the present invention first performs deep text mining on historical operation and maintenance logs, employing natural language processing techniques to identify equipment identifiers, anomalies, and remediation measures in fault events. It then constructs a fault propagation chain through event-time correlation analysis, automatically generating a set of triplets consisting of equipment, fault, and maintenance action, and assigning weights to each relationship based on the fault frequency and repair time. Optical character recognition and table parsing techniques are used to extract rated parameters from structured data in equipment manuals. Regular matching rules are designed to capture voltage ranges and load factor thresholds. An aging compensation factor is introduced based on the equipment's operational history to dynamically adjust the threshold, forming a parameter constraint rule library with version identification. A streaming computing framework is deployed for real-time data processing, applying sliding window statistics to raw sensor data. Temperature features are weighted averaged within the window (recent data is given a higher weight). The main vibration frequency is determined using spectral energy peak detection. The harmonic distortion rate is calculated as the root sum of squares of the fundamental and third harmonic amplitudes, constructing a standard feature vector containing time and space stamps. Dynamic conflict resolution is achieved by setting a dual-threshold mechanism. When the monitored feature mean continuously exceeds the manual parameter setting range and the fluctuation variance is higher than the warning line, the parameter self-correction program is activated. Based on the recent equipment health status data, the data quantile method is used to recalculate the allowable fluctuation range, and the constraint rule version in the knowledge graph is simultaneously updated. The knowledge graph storage uses a hypergraph relational model to perform a three-dimensional mapping of the rated parameters, real-time characteristics, and historical fault records of the equipment entity nodes. The load rate threshold is calculated by adding the average annual aging rate to the initial parameters of the equipment. The ambient temperature and humidity coupling coefficient is derived from the regression analysis of the equipment's full life cycle data. The number of maintenance times in the past three years is based on the actual fault maintenance records statistically counted by the work order system and weighted by the time decay factor, ultimately forming a knowledge graph that integrates the physical characteristics, operating status, and historical knowledge of the equipment.

[0021] The operation logic of step S14 is: Step S141: Calculate the absolute deviation between the equipment manual parameters and the statistical mean of the equipment operation status feature vector set to generate a parameter deviation vector D; Step S142: When it is detected that the deviation of any dimension in D exceeds the set threshold, a real-time data collection instruction with a sliding time window T = 30 minutes is triggered; Step S143: Perform a time series stationarity test on the sensor data stream within the window T, and select a set Q of continuous sampling points that meets the normal distribution requirements; Step S144: Based on the statistical distribution characteristics of the set Q, a second-order polynomial fitting algorithm is used to generate a parameter correction coefficient matrix C; Step S145: injecting the correction coefficient C into the device parameter constraint rule set to dynamically update the rated voltage allowable range and maximum load rate threshold of the corresponding device; Step S146: Generate an updated constraint rule set P' and perform version binding marking with the knowledge graph node.

[0022] The examples in this specification first calculate the deviation by comparing the rated parameters in the equipment manual with the statistical mean of the real-time operating characteristic vector, using a sliding window mechanism to ensure data timeliness. Specifically, for key parameters such as rated voltage and maximum load factor, the sensor data within the last 5-minute window is averaged in real time and the absolute difference is calculated against the manual parameters. If the deviation of any parameter exceeds a set threshold for three consecutive cycles (e.g., voltage deviation exceeds 15%, load factor deviation exceeds 20%), a 30-minute high-density data collection window is triggered. The raw data collected during this window period is verified for time series stationarity, outliers caused by sudden interference are removed, and a set of continuous sampling points that conform to normal distribution characteristics is retained. Based on the filtered valid data, a parameter correction model is established using second-order polynomial regression analysis. By calculating the gradient relationship between the actual operating parameters and the theoretical values, a correction coefficient matrix is ​​dynamically generated. The voltage correction value is the weighted average of the 95th percentile of the window data, and the load factor threshold is dynamically adjusted based on the equipment aging factor and 10% of the recent peak load. When the updated constraint rule set is version-associated with the knowledge graph node, the load rate threshold is compensated in stages based on the equipment's age. The ambient temperature and humidity coupling coefficient is determined through regression analysis of historical failure data. When counting the number of repairs over the past three years, planned maintenance is distinguished from fault repairs (only the latter are counted) and weighted by time decay. The updated parameter set is incrementally injected into the knowledge graph.

[0023] Multimodal semantic fusion module: Based on the entity attribute constraint rules of the knowledge graph, sliding window normalization is implemented on the temperature sensor time series data, wavelet time-frequency transform is performed on the vibration signal to generate a spectrum diagram, and the current harmonic feature vector is extracted; a bidirectional attention mapping network is designed to calculate the semantic similarity weights between multi-source data and knowledge nodes, and a graph embedding vector set with weight labels is generated through Hadamard product operation.

[0024] The operating logic of the multimodal semantic fusion module is: Step S21: normalize the temperature sensor time series data according to a 5-minute sliding window, calculate the z-score value of the data in each window, and generate a standardized temperature matrix; Step S22: performing wavelet packet decomposition on the vibration sensor signal, extracting energy proportion characteristics of the 0 to 500 Hz frequency band, and generating a vibration time-frequency spectrum; Step S23: Collect current sensor readings, extract the amplitudes of the 3rd, 5th, and 7th harmonic components through fast Fourier transform, and construct a current harmonic feature vector; Step S24: Encode the device rated parameter node attributes in the knowledge graph into knowledge node vectors, and design a bidirectional attention network to calculate the three sets of semantic similarity weights between the temperature matrix, vibration spectrum, and harmonic vector and the knowledge nodes respectively; Step S25: Perform a Hadamard product operation on the weight matrix and the multimodal feature dataset to generate a weighted graph embedding vector set.

[0025] In the multimodal data fusion process, this embodiment of the present invention first performs dynamic normalization on the temperature sensor time series data. The mean and standard deviation of the temperature values ​​within each 5-minute time window are calculated in a rolling manner. The original temperature values ​​are converted to z-scores in a standard normal distribution to eliminate baseline drift caused by environmental changes, forming a standardized temperature matrix arranged in time series. A six-layer wavelet packet decomposition technique is used to process vibration signals. After decomposition, the signal is reconstructed into the effective frequency band of 0 to 500 Hz. The energy percentage of each sub-band relative to the total energy is calculated and converted into a 128×128 pixel grayscale image to represent the spatiotemporal distribution of vibration energy. For current sensor data, a fast Fourier transform with a Hanning window is used to accurately separate the third, fifth, and seventh harmonic components based on a 50 Hz fundamental frequency. The ratio of each harmonic amplitude to the fundamental amplitude is used to construct a three-dimensional harmonic feature vector. Device parameter nodes in the knowledge graph are converted into a numerical matrix through vectorized encoding. Attributes such as rated voltage, load factor threshold, and environmental coefficient are mapped to different dimensions of the vector, forming a semantically annotated knowledge node matrix. The bidirectional attention network employs a three-channel parallel structure to calculate the semantic similarity weights of the temperature matrix, vibration spectrogram, and harmonic vectors with knowledge nodes. The temperature channel calculates association strength through sliding window mean comparison, the vibration channel uses the intersection-over-union ratio of the spectral texture feature histogram to assess matching, and the harmonic channel determines similarity based on the correlation coefficient of the frequency band energy distribution. Finally, a Hadamard product operation is used to achieve spatial alignment of multimodal features and weights. The weight matrix is ​​element-wise multiplied with the normalized feature matrix of the corresponding modality to generate a weighted three-dimensional embedding vector for each time point. The temperature weights are marked as the red channel, the vibration weights as the green channel, and the harmonics as the blue channel, forming a set of weighted spectrum embedding vectors.

[0026] The operation logic of step S24 includes: Step S241: extracting equipment rated parameter node attributes from the knowledge graph, including rated voltage value, maximum load rate, and allowable temperature rise threshold, and encoding them into numerical vectors of the same dimension; Step S242: normalizing the standardized temperature matrix generated by the temperature sensor, and calculating the root mean square value of the temperature change rate in each time window; Step S243: converting the vibration spectrum into a grayscale pixel matrix, and extracting image texture features as vibration representation vectors; Step S244: construct a bidirectional attention network model, set the temperature change rate, vibration texture characteristics, and current harmonic vector as the query end, and the knowledge node vector as the key value end; Step S245: Calculate the cosine similarity between the query end and the key-value end through the multi-head attention mechanism, and generate three sets of semantic similarity weights of the temperature matrix, vibration spectrum, harmonic vector and knowledge node respectively.

[0027] During the semantic weight calculation process, this embodiment of the present invention first extracts the device's rated parameter nodes from the knowledge graph, including key attributes such as rated voltage, load rate threshold, and allowable temperature rise range. Each parameter is scaled to the [0, 1] range through linear normalization and then concatenated into a fixed-dimensional numeric vector. For temperature sensor data, based on a pre-normalized temperature matrix, the temperature gradient between adjacent sampling points is calculated within a 5-minute time window. The severity of the temperature change is quantified by taking the square root of the squared average of the gradient values ​​within the sliding window, forming a timestamped temperature dynamic feature vector. After grayscale conversion, the vibration spectrum graph is extracted using the local binary pattern algorithm to extract image texture features. The edge direction histogram within each 3×3 pixel region is calculated to generate a vibration feature vector representing the device's mechanical state. A bidirectional attention network employs a three-way parallel structure, treating the temperature change rate, vibration texture features, and harmonic vector as independent query input streams. The knowledge node vector serves as a shared key-value comparison endpoint. Each attention head calculates the cosine similarity between the query vector and the key-value vector in different projection spaces to obtain association weights for the three dimensions of temperature-knowledge, vibration-knowledge, and harmonic-knowledge. In the specific implementation, the temperature channel focuses on comparing the degree of matching between the current dynamic changes and the rated temperature rise threshold, the vibration channel focuses on analyzing the similarity between the spectrum texture and historical failure modes, and the harmonic channel evaluates the correlation between the distortion characteristics and the load rate constraints. The final output is three sets of semantic similarity weight values.

[0028] Step S25 further includes: Step S251: performing softmax normalization processing on the three sets of semantic similarity weights to output a semantic similarity weight matrix with unified dimension; loading the semantic similarity weight matrix generated by the bidirectional attention network and a multimodal feature dataset, wherein the multimodal feature dataset includes a normalized temperature matrix, a reference vibration time-frequency spectrum, and a reference current harmonic vector; Step S252: Verify the consistency of the weight matrix column dimension and the multimodal feature row dimension. If the dimensions do not match, trigger the feature vector interpolation and completion algorithm. Step S253: performing Hadamard product operations in the order of feature channels, the calculation method is to multiply the weight matrix elements by the corresponding position eigenvalues ​​bit by bit; Step S254: adding a weight source identifier to each product result, including three types of tags: {temperature weight, vibration weight, harmonic weight}; Step S255: Aggregate all labeled product results, reorganize the dimensions according to the device number and timestamp, and generate a graph embedding vector set.

[0029] During the feature fusion stage, this embodiment of the present invention first applies softmax normalization to the three sets of semantic similarity weights output by the bidirectional attention network. The weight values ​​for temperature, vibration, and harmonics are compressed to the range [0, 1] and their sum is maintained at 1, forming a semantic weight probability matrix. Simultaneously, the pre-processed standardized temperature matrix (obtained through windowed z-score calculation), the reference vibration time-frequency spectrum (a grayscale image of the effective frequency band energy reconstructed via wavelet packet decomposition), and the reference current harmonic vector (a three-dimensional vector constructed from the amplitude ratio of the harmonic components) are loaded. When performing spatial alignment of the weight matrix and multimodal features, a dynamic interpolation strategy is used to address dimensional inconsistencies. When a mismatch is detected between the number of temperature matrix rows (time series length) and the number of weight columns, linear interpolation is used to generate supplementary feature points at equal intervals along the time axis. When the resolution of the vibration spectrum does not match the weight dimension, bicubic interpolation is used to adjust the image size. During the Hadamard product calculation, the normalized value of each temperature data point is multiplied by its corresponding weight. The vibration spectrum pixels are applied point by point to the weights, and the three components of the harmonic vector are multiplied by the three weight channels, respectively. This ensures that features of different physical dimensions produce comparable fusion effects under the action of the weights. The product results are categorized by source type: temperature features are marked with red spatial channel weights, vibration features are marked with green spectral channel weights, and harmonics are marked with blue frequency domain channel weights. Finally, through spatiotemporal reorganization, the three-channel weight product results for the same device number and the same timestamp are superimposed and combined to form a set of graph embedding vectors with multidimensional semantic labels.

[0030] Feature contradiction tracing module: The embedded vector set is input into the pre-trained graph neural network model to construct a multi-dimensional feature space distribution map of the device status; the cosine similarity matrix between nodes of similar devices is calculated to identify abnormal node pairs below the set threshold; the conduction path is inferred based on the topological connection rules of the knowledge graph to locate the root cause device set that causes feature deviation.

[0031] The feature conflict tracing module includes: Step S31: Loading a graph embedding vector set and a pre-trained graph neural network model, wherein the embedding vector set includes device temperature, vibration, harmonic characteristics and weight labels; Step S32: Input the embedded vector into the graph neural network for spatial mapping to generate a multi-dimensional feature space distribution graph, where each node represents the state vector of a single device; Step S33: Calculate the cosine similarity between the nodes of the same type of devices, and mark them as abnormal node pairs when the similarity is lower than the set threshold; Step S34: Draw the electrical connection conduction path of the abnormal node based on the device topology connection relationship in the knowledge graph; Step S35: performing reverse gradient propagation calculation along the conduction path to locate the initial abnormal device node that causes the feature shift; Step S36: Aggregate the initial abnormal nodes and their associated devices to generate a root cause device set and mark the deviation type.

[0032] During the feature conflict tracing process, this embodiment of the present invention first loads a graph embedding vector set containing device temperature, vibration, and harmonic features and their weighted labels, and invokes a pre-trained graph neural network model. Using the graph neural network's hierarchical spatial mapping, the multidimensional feature vectors are projected into a low-dimensional space, generating a feature distribution map representing the device's operating status. Each node corresponds to a unique device, and edges between nodes represent the device's electrical connections. When evaluating the state similarity of similar device nodes, the cosine similarity between node vectors is calculated (a measure of similarity based on the cosine of the angle between the vectors). Node pairs with similarities below a set threshold are marked as abnormal, and their spatial coordinates and device attributes are extracted. Based on the topological connectivity information in the knowledge graph, a bidirectional breadth search is performed starting from the abnormal node, mapping the complete conduction path including associated devices such as switchgear, lines, and transformers. The path search depth is limited to six cascaded devices. Backward gradient propagation analysis is performed along the conduction path, starting from the terminal abnormal node, and gradient attenuation values ​​are weighted based on impedance matching and device health status, tracing back to the initial abnormal node. Finally, the initial abnormal node and its associated equipment under the same power supply bus are aggregated, and the fault modes (such as overload, insulation aging, and harmonic excess) in the historical maintenance records are combined to mark the deviation type of the root cause equipment set according to the frequency of occurrence of the fault type.

[0033] The sub-steps of step S34 include: Step S341: Loading the device topology connection relationship data and abnormal node set in the knowledge graph, obtaining the electrical connection direction information between the devices, where the electrical connection direction information is divided into transmission / distribution direction; Step S342: Starting from the abnormal node, perform a bidirectional breadth-first search along the electrical connection direction, and traverse a path with no more than 6 cascaded devices; Step S343: Filter the effective conduction paths according to the device type and remove the reactive compensation device path branches; the device types include transformers / circuit breakers / capacitors; Step S344: Perform physical connection verification on the screened conductive paths, detect impedance matching of adjacent device ports, and remove mismatched path segments; Step S345: Mark the valid conduction path as a topological link diagram with arrows, and generate conduction path data including {path length, device type sequence, impedance value}.

[0034] During the conduction path mapping process, this embodiment of the present invention first loads the device topology connection relationships and a set of marked abnormal nodes from the knowledge graph. The electrical connection directions between devices are analyzed (transmission direction refers to the upstream substation to the current busbar, and distribution direction refers to the current busbar to the user end), and an adjacency matrix with directional attributes is established. Using the abnormal node as the search starting point, a bidirectional breadth-first search is performed simultaneously along both the transmission and distribution directions, limiting the path traversal depth to six cascaded devices, such as an abnormal transformer → switchgear → line → downstream transformer. During this process, the device type sequence of the path nodes is recorded in real time. Based on the power system's conduction characteristics, path branches containing reactive compensation devices such as capacitors and reactors are eliminated, while valid conduction paths for key equipment such as transformers, circuit breakers, and lines are retained. Physical connectivity verification is performed on the initially screened paths by calculating the impedance ratio of adjacent device ports. The calculation logic is based on the percentage matching between the output impedance of the upstream device and the input impedance of the downstream device. If the impedance mismatch between three consecutive nodes exceeds a set threshold, the path segment is deemed invalid and removed. The final effective conduction path is presented as a topological diagram with directional arrows. The total number of hops in the path is marked as a length indicator. The equipment types are arranged in the conduction order (such as transformer-circuit breaker-line). The total impedance of the path is taken as the geometric mean of the impedance of each node port to generate a standardized conduction path dataset.

[0035] The sub-steps of step S35 include: Step S351: Loading conduction path data and device characteristic offset values, and establishing a path node gradient calculation index table; Step S352: Initialize the back propagation gradient using the feature offset of the device node at the end of the path as the starting point; Step S353: Backtrack the calculation hop by hop along the conduction path, and update the formula as Δi = Δj × (impedance matching × 0.8 + device health × 0.2); where Δj represents the gradient value of the downstream node of the current node in the conduction path, and Δi represents the gradient value of the upstream node currently being calculated; The calculation process example is: in the conduction path A→B→C→D: Calculate backward from the end D (the initial gradient of D is Δ=1.0); Calculate the gradient of node C: ΔC = ΔD × (impedance matching from C to D × 0.8 + health of C × 0.2); Calculate the gradient of node B: ΔB = ΔC × (impedance matching from B to C × 0.8 + health of B × 0.2); At this time: when calculating node B, Δj=ΔC (the gradient of downstream node C); when calculating node B, Δi=ΔB (the gradient of node B currently being calculated).

[0036] Step S354: When it is detected that the gradient change rate between adjacent nodes exceeds a set threshold, the node is marked as a candidate abnormal source; Step S355: Select the candidate node with the largest cumulative gradient value and confirm it as the initial abnormal device node.

[0037] During the anomaly source location process, the present embodiment first loads conduction path data containing device electrical connection relationships and characteristic offset values. A gradient calculation table is constructed, indexed by device nodes and attributed by impedance matching and health. Starting from the device node at the end of the conduction path, its characteristic offset value is normalized to an initial gradient value of 1.0, and the calculation is performed backtracking step by step in the opposite direction of the electrical connection. The gradient value of each upstream node is determined by multiplying the gradient value of its immediately downstream node by a comprehensive weight coefficient. The weight coefficient is composed of the impedance matching between adjacent devices and the device's own health score. Impedance matching is calculated as the ratio of the upstream device's output impedance to the downstream device's input impedance. The health score is derived from a comprehensive assessment of the device's operating age, maintenance history, and real-time monitoring data. If an increase in the gradient value between adjacent nodes exceeds a set threshold, the node is marked as a candidate anomaly source. After traversing the entire conduction path, the cumulative gradient values ​​of each candidate node are ranked. The cumulative value is the weighted sum of the gradient values ​​of all downstream paths from that node. The node with the largest cumulative gradient value is selected as the initial anomaly device node. Breakpoints exceeding the threshold are also recorded as auxiliary location information.

[0038] It should be noted that the intervals and thresholds are set for ease of comparison. The threshold size depends on the amount of sample data and the cardinality set by those skilled in the art for each set of sample data, as long as it does not affect the proportional relationship between the parameter and the quantized value. Furthermore, the above formulas are all dimensionless numerical calculations. These formulas are derived from software simulations of the most recent real-world conditions using large amounts of data. The preset parameters in these formulas are set by those skilled in the art based on actual conditions.

[0039] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0040] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

[0041] Finally: The above description is only a preferred embodiment of the present invention and is not intended to limit the present invention. Any modifications, equivalent substitutions, improvements, etc. made within the spirit and principles of the present invention should be included in the scope of protection of the present invention.

Claims

1. A fusion and management system for multi-source heterogeneous scientific and technological information resources, characterized by: include: The knowledge graph construction module parses equipment failure chains from the unstructured text of historical operation and maintenance logs, extracts rated parameter constraints from structured tables in equipment manuals, and collects operational feature vectors from real-time sensor data streams. The hypergraph storage engine is then used to fuse device topological connections with failure mode causal relationships to generate a knowledge graph containing N entity relationships. Multimodal semantic fusion module: Based on the entity attribute constraints of the knowledge graph, it implements sliding window normalization on temperature sensor time series data, performs wavelet time-frequency transform on vibration signals to generate spectrograms, and extracts current harmonic feature vectors. It also designs a bidirectional attention mapping network to calculate the semantic similarity weights between multi-source data and knowledge nodes, and generates a weighted graph embedding vector set through Hadamard product operations. Feature Conflict Tracing Module: This module inputs the embedded vector set into a pre-trained graph neural network model to construct a multi-dimensional feature space distribution map of the device status. Calculate the cosine similarity matrix between nodes of the same type of devices to identify abnormal node pairs below the set threshold; perform conduction path reasoning based on the topological connection rules of the knowledge graph to locate the root cause device set that causes feature deviation.

2. The system for integrating and managing multi-source heterogeneous scientific and technological information resources according to claim 1, characterized in that: The operational logic of the knowledge graph construction module includes: Step S11: Perform unstructured text parsing on historical operation and maintenance logs, extract entity relationship triples {equipment ID, fault type, maintenance measures} through the fault event causal relationship chain, and generate a fault chain entity relationship set; Step S12: parsing the structured table in the equipment manual, matching the rated voltage and maximum load rate parameter constraint rules based on regular expressions, and generating an equipment parameter constraint rule set; Step S13: collecting sensor data streams in real time, extracting transformer temperature mean, vibration spectrum main frequency, and current harmonic distortion rate features, and constructing a device operation status feature vector set; Step S14: establishing a dynamic conflict resolution mechanism. When it is detected that the deviation between the manual parameters and the statistical mean of the equipment operating status feature vector set exceeds a set threshold, the constraint rules are dynamically updated based on the sensor data in the sliding window. Step S15: The fault chain set, updated constraint rule set, and state feature set are integrated through the hypergraph storage engine to associate the topological connection relationship of the transformer nodes with the historical fault mode, and generate a knowledge graph containing N entity relationships, where each transformer entity node contains an attribute group including a load rate threshold, an ambient temperature and humidity coupling coefficient, and the number of maintenance times in the past three years.

3. The system for integrating and managing multi-source heterogeneous scientific and technological information resources according to claim 2, characterized in that: The operation logic of step S14 is: Step S141: Calculate the absolute deviation between the equipment manual parameters and the statistical mean of the equipment operation status feature vector set to generate a parameter deviation vector D; Step S142: When it is detected that the deviation of any dimension in D exceeds the set threshold, a real-time data collection instruction with a sliding time window T = 30 minutes is triggered; Step S143: Perform a time series stationarity test on the sensor data stream within the window T, and select a set Q of continuous sampling points that meets the normal distribution requirements; Step S144: Based on the statistical distribution characteristics of the set Q, a second-order polynomial fitting algorithm is used to generate a parameter correction coefficient matrix C; Step S145: injecting the correction coefficient C into the device parameter constraint rule set to dynamically update the rated voltage allowable range and maximum load rate threshold of the corresponding device; Step S146: Generate an updated constraint rule set P' and perform version binding marking with the knowledge graph node.

4. The system for integrating and managing multi-source heterogeneous scientific and technological information resources according to claim 1, characterized in that: The operating logic of the multimodal semantic fusion module is: Step S21: normalize the temperature sensor time series data according to a 5-minute sliding window, calculate the z-score value of the data in each window, and generate a standardized temperature matrix; Step S22: performing wavelet packet decomposition on the vibration sensor signal, extracting energy proportion characteristics of the 0 to 500 Hz frequency band, and generating a vibration time-frequency spectrum; Step S23: Collect current sensor readings, extract the amplitudes of the 3rd, 5th, and 7th harmonic components through fast Fourier transform, and construct a current harmonic feature vector; Step S24: Encode the device rated parameter node attributes in the knowledge graph into knowledge node vectors, and design a bidirectional attention network to calculate the three sets of semantic similarity weights between the temperature matrix, vibration spectrum, and harmonic vector and the knowledge nodes respectively; Step S25: Perform a Hadamard product operation on the weight matrix and the multimodal feature dataset to generate a weighted graph embedding vector set.

5. The system for integrating and managing multi-source heterogeneous scientific and technological information resources according to claim 4, characterized in that: The operation logic of step S24 includes: Step S241: extracting equipment rated parameter node attributes from the knowledge graph, including rated voltage value, maximum load rate, and allowable temperature rise threshold, and encoding them into numerical vectors of the same dimension; Step S242: normalizing the standardized temperature matrix generated by the temperature sensor, and calculating the root mean square value of the temperature change rate in each time window; Step S243: converting the vibration spectrum into a grayscale pixel matrix, and extracting image texture features as vibration representation vectors; Step S244: construct a bidirectional attention network model, set the temperature change rate, vibration texture characteristics, and current harmonic vector as the query end, and the knowledge node vector as the key value end; Step S245: Calculate the cosine similarity between the query end and the key-value end through the multi-head attention mechanism, and generate three sets of semantic similarity weights of the temperature matrix, vibration spectrum, harmonic vector and knowledge node respectively.

6. The system for integrating and managing multi-source heterogeneous scientific and technological information resources according to claim 4, characterized in that: Step S25 further includes: Step S251: performing softmax normalization processing on the three sets of semantic similarity weights to output a semantic similarity weight matrix with unified dimension; loading the semantic similarity weight matrix generated by the bidirectional attention network and a multimodal feature dataset, wherein the multimodal feature dataset includes a normalized temperature matrix, a reference vibration time-frequency spectrum, and a reference current harmonic vector; Step S252: Verify the consistency of the weight matrix column dimension and the multimodal feature row dimension. If the dimensions do not match, trigger the feature vector interpolation and completion algorithm. Step S253: performing Hadamard product operations in the order of feature channels, the calculation method is to multiply the weight matrix elements by the corresponding position eigenvalues ​​bit by bit; Step S254: adding a weight source identifier to each product result, including three types of tags: {temperature weight, vibration weight, harmonic weight}; Step S255: Aggregate all labeled product results, reorganize the dimensions according to the device number and timestamp, and generate a graph embedding vector set.

7. The system for integrating and managing multi-source heterogeneous scientific and technological information resources according to claim 1, characterized in that: The feature conflict tracing module includes: Step S31: Loading a graph embedding vector set and a pre-trained graph neural network model, wherein the embedding vector set includes device temperature, vibration, harmonic characteristics and weight labels; Step S32: Input the embedded vector into the graph neural network for spatial mapping to generate a multi-dimensional feature space distribution graph, where each node represents the state vector of a single device; Step S33: Calculate the cosine similarity between the nodes of the same type of devices, and mark them as abnormal node pairs when the similarity is lower than the set threshold; Step S34: Draw the electrical connection conduction path of the abnormal node based on the device topology connection relationship in the knowledge graph; Step S35: performing reverse gradient propagation calculation along the conduction path to locate the initial abnormal device node that causes the feature shift; Step S36: Aggregate the initial abnormal nodes and their associated devices to generate a root cause device set and mark the deviation type.

8. The system for integrating and managing multi-source heterogeneous scientific and technological information resources according to claim 7, characterized in that: The sub-steps of step S34 include: Step S341: Loading the device topology connection relationship data and abnormal node set in the knowledge graph, obtaining the electrical connection direction information between the devices, where the electrical connection direction information is divided into transmission / distribution direction; Step S342: Starting from the abnormal node, perform a bidirectional breadth-first search along the electrical connection direction, and traverse a path with no more than 6 cascaded devices; Step S343: Filter the effective conduction paths according to the device type and remove the reactive compensation device path branches; the device types include transformers / circuit breakers / capacitors; Step S344: Perform physical connection verification on the screened conductive paths, detect impedance matching of adjacent device ports, and remove mismatched path segments; Step S345: Mark the valid conduction path as a topological link diagram with arrows, and generate conduction path data including {path length, device type sequence, impedance value}.

9. The system for integrating and managing multi-source heterogeneous scientific and technological information resources according to claim 8, characterized in that: The sub-steps of step S35 include: Step S351: Loading conduction path data and device characteristic offset values, and establishing a path node gradient calculation index table; Step S352: Initialize the back propagation gradient using the feature offset of the device node at the end of the path as the starting point; Step S353: Backtrack the calculation hop by hop along the conduction path, and update the formula as Δi = Δj × (impedance matching × 0.8 + device health × 0.2); where Δj represents the gradient value of the downstream node of the current node in the conduction path, and Δi represents the gradient value of the upstream node currently being calculated; Step S354: When it is detected that the gradient change rate between adjacent nodes exceeds a set threshold, the node is marked as a candidate abnormal source; Step S355: Select the candidate node with the largest cumulative gradient value and confirm it as the initial abnormal device node.

Citation Information

Cited By

  • Software improvement suggestion generation method based on fault feature matching and role guidance

    CN121233477A

  • Software improvement suggestion generation method based on fault feature matching and role orientation

    CN121233477B

  • Operation and maintenance knowledge graph dynamic construction method and system based on monitoring and disposal process

    CN121327155A

  • Enterprise operation index prediction system based on data mining

    CN121352599A

  • Knowledge graph-based refrigerating machine room operation and maintenance robot and operation and maintenance method

    CN121353792A