A method and system for identifying and classifying supply chain breakpoints based on multi-source data fusion

CN122571330APending Publication Date: 2026-08-14BEIJING SGITG ACCENTURE INFORMATION TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-03-31
Publication Date
2026-08-14

AI Technical Summary

Technical Problem

[0003]现有技术虽然在各自的应用场景下取得了一定成效,但在面对复杂、动态且涉及多环节的电力系统产业链断点识别时,存在多源异构数据利用不充分、量化分析与语义理解相割裂、难以处理未知断点类型,以及难以预测断点在复杂时空依赖结构中的传播路径问题

Benefits of technology

本发明提供了一种基于多源数据融合的产业链断点识别与分类方法,包括:采集多源数据,并对多源数据进行融合处理,构建多源异构数据集合;基于多源异构数据集合,抽取构成产业链结构的实体及实体间的关联关系,构建表征电力系统产业链结构的知识图谱;基于多源异构数据集合中的电网拓扑结构,结合知识图谱,融合形成统一图模型;基于统一图模型中的各节点,从多源异构数据集合中提取出各节点的多维度特征;将统一图模型与多维度特征输入到时空图神经网络模型,计算各节点的断点风险及断点沿依赖路径的传播趋势,获得断点风险识别结果;基于多源异构数据集合中的文本数据,利用大语言模型对文本数据进行语义理解与推理,得到文本数据中的断点事件信息,将断点事件信息作为语义断点识别结果;基于断点风险识别结果与语义断点识别结果的融合结果进行断点分类,生成断点识别结果并进行预警。本发明通过构建多源异构数据集合与产业链知识图谱,融合电网拓扑结构与知识图谱形成统一图模型,结合时空图神经网络与大语言模型分别从量化数据与语义文本两个维度识别断点风险,并进一步将两种不同的识别结果进行融合分类与可视化预警,实现了对电力系统产业链断点的全域感知、时空传播推演、未知事件语义识别及精准分级预警,显著提升了断点识别的全面性、准确性与实时性,降低了对外部标注数据的依赖,为电力系统产业链安全保障提供了数据驱动的智能化决策支持。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122571330A_ABST
    Figure CN122571330A_ABST
Patent Text Reader

Abstract

This invention provides a method and system for identifying and classifying supply chain breakpoints based on multi-source data fusion, relating to the field of electronic system supply chain security technology. The method includes: collecting multi-source data to construct a multi-source heterogeneous data set; constructing a knowledge graph based on the multi-source heterogeneous data set; fusing the knowledge graph with the power grid topology to form a unified graph model; extracting multi-dimensional features based on the unified graph model; inputting the unified graph model and multi-dimensional features into a spatiotemporal graph neural network model to obtain breakpoint risk identification results; obtaining breakpoint event information from text data, using the breakpoint event information as a semantic breakpoint identification result for breakpoint classification, generating breakpoint identification results, and issuing early warnings. This invention achieves full-domain perception, spatiotemporal propagation simulation, semantic identification of unknown events, and precise hierarchical early warning for power system supply chain breakpoints, providing intelligent decision support for power system supply chain security.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of electronic system supply chain security technology, specifically to a method and system for identifying and classifying supply chain breakpoints based on multi-source data fusion. Background Technology

[0002] The power system industry chain covers multiple links such as power generation, transmission, distribution, power consumption and equipment manufacturing. The identification and classification of its breakpoints (such as equipment supply interruption, shortage of key components, logistics stagnation, etc.) is the key to ensuring the safety of power supply. At present, the industry mainly adopts the following identification methods: (1) Rule-based fault diagnosis method: expert experience is transformed into fixed logical rules or decision trees. When the monitoring data triggers the preset threshold, it is judged as a breakpoint. This method responds quickly to scenarios with clear rules, but its reasoning ability is limited when facing unknown or complex breakpoint types. (2) Method based on a single artificial intelligence model: support vector machine, convolutional neural network and other single models are used to train and classify power grid operation data or equipment status data. This method has a high accuracy, but it is difficult to integrate multi-source heterogeneous data and has a strong dependence on labeled data. (3) Topology analysis method based on graph neural network: graph neural network is used to analyze the relationship between power grid nodes to locate fault nodes. This method is good at processing structured topology data, but it is difficult to process time series dynamic features and unstructured text information. (4) Time-series data analysis method: This method uses models such as recurrent neural networks and long short-term memory networks to process time-series data and identifies abnormal patterns by capturing time-dependent features. This method performs well in scenarios with strong time-series dependencies, but it ignores the power grid topology and supply chain dependencies.

[0003] While existing technologies have achieved certain results in their respective application scenarios, when faced with the identification of breakpoints in the complex, dynamic, and multi-stage power system supply chain, there are still problems such as insufficient utilization of multi-source heterogeneous data, separation between quantitative analysis and semantic understanding, difficulty in handling unknown breakpoint types, and difficulty in predicting the propagation path of breakpoints in complex spatiotemporal dependency structures. Summary of the Invention

[0004] To overcome the shortcomings of the existing technology, this invention provides a method for identifying and classifying supply chain breakpoints based on multi-source data fusion, comprising: Collect data from multiple sources and fuse the data to construct a multi-source heterogeneous data set; Based on a multi-source heterogeneous data set, the entities that constitute the industrial chain structure and the relationships between entities are extracted to construct a knowledge graph that represents the industrial chain structure of the power system. Based on the power grid topology in a multi-source heterogeneous data set, a unified graph model is formed by combining knowledge graphs; based on each node in the unified graph model, multi-dimensional features of each node are extracted from the multi-source heterogeneous data set. By inputting a unified graph model and multi-dimensional features into a spatiotemporal graph neural network model, the breakpoint risk of each node and the propagation trend of the breakpoint along the dependent path are calculated to obtain the breakpoint risk identification result. Based on text data in a multi-source heterogeneous dataset, a large language model is used to perform semantic understanding and reasoning on the text data to obtain breakpoint event information in the text data, and the breakpoint event information is used as the semantic breakpoint recognition result. Breakpoints are classified based on the fusion of breakpoint risk identification results and semantic breakpoint identification results, and breakpoint identification results are generated and warnings are issued.

[0005] Preferably, based on a multi-source heterogeneous data set, entities constituting the industrial chain structure and the relationships between entities are extracted to construct a knowledge graph representing the power system industrial chain structure, including: For text data in multi-source heterogeneous datasets, a sequence labeling recognition model is used to identify multiple entities that constitute the industrial chain structure. These entities include at least one of equipment entities, enterprise entities, raw material entities, and project entities. A remote supervision-based relationship classification model is adopted, combined with a predefined set of relationship types, to determine the relationships between entities. These relationships include at least one of supply relationships, dependency relationships, and transportation relationships. Each identified entity is treated as a node, and each determined relationship is treated as a relationship edge. A knowledge graph is constructed and stored in a graph database using an attribute graph model.

[0006] Preferably, a unified graph model is formed by fusing the power grid topology from a multi-source heterogeneous dataset with a knowledge graph, including: Map the power grid nodes in the power grid topology to the first node set of the unified graph model. The power grid nodes include substation nodes and line nodes. The entities that constitute the industrial chain structure in the knowledge graph are mapped to the second node set of the unified graph model, where the second node set and the first node set have at least one shared node. The shared node is a device entity that exists in both the power grid topology and the knowledge graph. The physical connection edges in the power grid topology are mapped to the first edge set of the unified graph model, and the association relationship edges in the knowledge graph are mapped to the second edge set of the unified graph model. The edge type is labeled for the physical connection edges and association relationship edges associated with shared nodes. The unified graph model is constructed by taking the union of the first node set and the second node set as the node set and taking the union of the first edge set and the second edge set as the edge set.

[0007] Preferably, based on each node in the unified graph model, multi-dimensional features of each node are extracted from a multi-source heterogeneous data set, including: Based on the nodes in the unified graph model, static attribute feature vectors of each node are extracted from a multi-source heterogeneous data set; the static attribute feature vectors include at least one or more of the following: equipment capacity, annual production capacity of the enterprise, geographical location coordinates, and node type code; Time series data associated with each node are extracted from a multi-source heterogeneous dataset and normalized to form a dynamic time series feature vector. The dynamic time series feature vector includes one or more of the following: equipment operating parameter sequence, inventory change sequence, raw material price sequence, and logistics status sequence. Static attribute feature vectors are concatenated with dynamic temporal feature vectors to form multi-dimensional feature vectors for each node.

[0008] Preferably, a unified graph model and multi-dimensional features are input into a spatiotemporal graph neural network model to calculate the breakpoint risk of nodes and the propagation trend of breakpoints along dependent paths, thereby obtaining breakpoint risk identification results, including: The unified graph model and multi-dimensional features are input into the spatiotemporal graph convolutional network. Through the forward propagation of the spatiotemporal graph convolutional network, the spatiotemporal graph convolutional network is used to extract temporal features in the time dimension using a one-dimensional convolutional layer, and in the spatial dimension, the graph convolutional layer is used to aggregate the feature information of adjacent nodes according to the edge set of the unified graph model. Based on temporal features and feature information, the breakpoint risk probability value of each node at the current time and the propagation probability value of each edge at the current time are output. Nodes with breakpoint risk probability values ​​exceeding a preset threshold are used as source nodes, and propagation probability values ​​are used as edge weights. Multiple breakpoint propagation paths are generated using a depth-first search algorithm. Each breakpoint propagation path includes propagation direction and propagation time estimation. The probability value of breakpoint risk and the breakpoint propagation path are determined as the breakpoint risk identification results.

[0009] Preferably, based on text data from a multi-source heterogeneous dataset, a large language model is used to perform semantic understanding and reasoning on the text data to obtain breakpoint event information in the text data. This breakpoint event information is then used as the semantic breakpoint identification result, including: New text data is selected from a multi-source heterogeneous dataset. The new text data includes at least one of the following: news information, corporate announcements, policy documents, and industry reports. Obtain a preset number of high-risk nodes with the highest risk probability values ​​from the breakpoint risk identification results, and extract subgraph information associated with each high-risk node from the knowledge graph; Based on the subgraph information of each high-risk node, a prompt word is constructed. The prompt word includes one or more of the following fields: role setting field, text content field to be analyzed field, context reference field, and output format constraint field. Natural language output is obtained using a large language model based on prompt words; By combining regular expression matching with a rule-based parser, one or more of the following are extracted from the natural language output: event type, affected entity identifier, event occurrence time, event duration, and confidence score, to form breakpoint event information. This breakpoint event information is then used as the semantic breakpoint recognition result.

[0010] Preferably, breakpoints are classified based on the breakpoint risk identification results and semantic breakpoint identification results, and breakpoint identification results are generated and warnings are issued, including: The breakpoint risk probability value of each node in the breakpoint risk identification result is used as the first evidence, and the confidence score of each event record in the semantic breakpoint identification result is used as the second evidence. The DS evidence theory is used to calculate the integrated breakpoint confidence after fusion. When the overall breakpoint confidence exceeds the first threshold, the industrial chain link corresponding to the node or event is identified as a breakpoint. Based on the identified breakpoints, and according to the node type or event type that triggered the breakpoint, they are mapped to a preset breakpoint classification system to determine the category and nature of the breakpoint. The urgency level of each breakpoint is determined based on the comprehensive breakpoint confidence level and the estimated propagation time in the breakpoint propagation path. A breakpoint identification report is generated, which includes one or more of the following: breakpoint identifier, link category, nature category, scope of impact, expected impact time window, and suggested countermeasures. The breakpoint identification report and the breakpoint propagation path diagram are displayed through a visual interface for graded early warning.

[0011] Based on the same inventive concept, this invention also provides a data fusion-based supply chain breakpoint identification and classification system, which includes: The multi-source data set construction module is used to collect multi-source data, fuse the multi-source data, and construct a multi-source heterogeneous data set. The knowledge graph construction module is used to extract entities that constitute the industrial chain structure and the relationships between entities based on multi-source heterogeneous data sets, and to construct a knowledge graph that represents the industrial chain structure of the power system. The model feature determination module is used to integrate the power grid topology in the multi-source heterogeneous data set with knowledge graph to form a unified graph model; based on each node in the unified graph model, it extracts multi-dimensional features of each node from the multi-source heterogeneous data set. The breakpoint risk identification module is used to input the unified graph model and multi-dimensional features into the spatiotemporal graph neural network model, calculate the breakpoint risk of each node and the propagation trend of the breakpoint along the dependent path, and obtain the breakpoint risk identification result. The semantic breakpoint identification module is used to perform semantic understanding and reasoning on text data based on multi-source heterogeneous data sets, and obtain breakpoint event information in the text data, and use the breakpoint event information as the semantic breakpoint identification result. The breakpoint classification module is used to classify breakpoints based on the fusion result of breakpoint risk identification and semantic breakpoint identification, generate breakpoint identification results, and issue warnings.

[0012] Preferably, the knowledge graph construction module is specifically used for: For text data in multi-source heterogeneous datasets, a sequence labeling recognition model is used to identify multiple entities that constitute the industrial chain structure. These entities include at least one of equipment entities, enterprise entities, raw material entities, and project entities. A remote supervision-based relationship classification model is adopted, combined with a predefined set of relationship types, to determine the relationships between entities. These relationships include at least one of supply relationships, dependency relationships, and transportation relationships. Each identified entity is treated as a node, and each determined relationship is treated as a relationship edge. A knowledge graph is constructed and stored in a graph database using an attribute graph model.

[0013] Preferably, the model feature determination module is specifically used for: Map the power grid nodes in the power grid topology to the first node set of the unified graph model. The power grid nodes include substation nodes and line nodes. The entities that constitute the industrial chain structure in the knowledge graph are mapped to the second node set of the unified graph model, where the second node set and the first node set have at least one shared node. The shared node is a device entity that exists in both the power grid topology and the knowledge graph. The physical connection edges in the power grid topology are mapped to the first edge set of the unified graph model, and the association relationship edges in the knowledge graph are mapped to the second edge set of the unified graph model. The edge type is labeled for the physical connection edges and association relationship edges associated with shared nodes. The unified graph model is constructed by taking the union of the first node set and the second node set as the node set and taking the union of the first edge set and the second edge set as the edge set.

[0014] Preferably, the model feature determination module is also specifically used for: Based on the nodes in the unified graph model, static attribute feature vectors of each node are extracted from a multi-source heterogeneous data set; the static attribute feature vectors include at least one or more of the following: equipment capacity, annual production capacity of the enterprise, geographical location coordinates, and node type code; Time series data associated with each node are extracted from a multi-source heterogeneous dataset and normalized to form a dynamic time series feature vector. The dynamic time series feature vector includes one or more of the following: equipment operating parameter sequence, inventory change sequence, raw material price sequence, and logistics status sequence. Static attribute feature vectors are concatenated with dynamic temporal feature vectors to form multi-dimensional feature vectors for each node.

[0015] Preferably, the breakpoint risk identification and acquisition module is specifically used for: The unified graph model and multi-dimensional features are input into the spatiotemporal graph convolutional network. Through the forward propagation of the spatiotemporal graph convolutional network, the spatiotemporal graph convolutional network is used to extract temporal features in the time dimension using a one-dimensional convolutional layer, and in the spatial dimension, the graph convolutional layer is used to aggregate the feature information of adjacent nodes according to the edge set of the unified graph model. Based on temporal features and feature information, the breakpoint risk probability value of each node at the current time and the propagation probability value of each edge at the current time are output. Nodes with breakpoint risk probability values ​​exceeding a preset threshold are used as source nodes, and propagation probability values ​​are used as edge weights. Multiple breakpoint propagation paths are generated using a depth-first search algorithm. Each breakpoint propagation path includes propagation direction and propagation time estimation. The probability value of breakpoint risk and the breakpoint propagation path are determined as the breakpoint risk identification results.

[0016] Preferably, the semantic breakpoint recognition module is specifically used for: New text data is selected from a multi-source heterogeneous dataset. The new text data includes at least one of the following: news information, corporate announcements, policy documents, and industry reports. Obtain a preset number of high-risk nodes with the highest risk probability values ​​from the breakpoint risk identification results, and extract subgraph information associated with each high-risk node from the knowledge graph; Based on the subgraph information of each high-risk node, a prompt word is constructed. The prompt word includes one or more of the following fields: role setting field, text content field to be analyzed field, context reference field, and output format constraint field. Natural language output is obtained using a large language model based on prompt words; By combining regular expression matching with a rule-based parser, one or more of the following are extracted from the natural language output: event type, affected entity identifier, event occurrence time, event duration, and confidence score, to form breakpoint event information. This breakpoint event information is then used as the semantic breakpoint recognition result.

[0017] Preferably, the breakpoint classification module is specifically used for: The breakpoint risk probability value of each node in the breakpoint risk identification result is used as the first evidence, and the confidence score of each event record in the semantic breakpoint identification result is used as the second evidence. The DS evidence theory is used to calculate the integrated breakpoint confidence after fusion. When the overall breakpoint confidence exceeds the first threshold, the industrial chain link corresponding to the node or event is identified as a breakpoint. Based on the identified breakpoints, and according to the node type or event type that triggered the breakpoint, they are mapped to a preset breakpoint classification system to determine the category and nature of the breakpoint. The urgency level of each breakpoint is determined based on the comprehensive breakpoint confidence level and the estimated propagation time in the breakpoint propagation path. A breakpoint identification report is generated, which includes one or more of the following: breakpoint identifier, link category, nature category, scope of impact, expected impact time window, and suggested countermeasures. The breakpoint identification report and the breakpoint propagation path diagram are displayed through a visual interface for graded early warning.

[0018] Based on the same inventive concept, the present invention also provides an electronic device, comprising: at least one processor and a memory; wherein the memory and the processor are connected via a bus; The memory is used to store one or more programs; When the one or more programs are executed by the at least one processor, a data fusion-based method for identifying and classifying supply chain breakpoints is implemented as described above.

[0019] Based on the same inventive concept, the present invention also provides a readable storage medium having an executable program stored thereon, wherein when the executable program is executed, it implements the aforementioned method for identifying and classifying supply chain breakpoints based on data fusion.

[0020] Compared with the closest existing technology, the present invention has the following beneficial effects: This invention provides a method for identifying and classifying supply chain breakpoints based on multi-source data fusion, comprising: collecting multi-source data and fusing the multi-source data to construct a multi-source heterogeneous data set; based on the multi-source heterogeneous data set, extracting entities constituting the supply chain structure and the relationships between entities to construct a knowledge graph representing the supply chain structure of a power system; based on the power grid topology in the multi-source heterogeneous data set, combining the knowledge graph to form a unified graph model; based on each node in the unified graph model, extracting multi-dimensional features of each node from the multi-source heterogeneous data set; inputting the unified graph model and multi-dimensional features into a spatiotemporal graph neural network model to calculate the breakpoint risk of each node and the propagation trend of the breakpoint along the dependent path, obtaining the breakpoint risk identification result; based on the text data in the multi-source heterogeneous data set, using a large language model to perform semantic understanding and reasoning on the text data to obtain breakpoint event information in the text data, and using the breakpoint event information as the semantic breakpoint identification result; classifying breakpoints based on the fusion result of the breakpoint risk identification result and the semantic breakpoint identification result, generating breakpoint identification results and issuing warnings. This invention constructs a multi-source heterogeneous data set and an industry chain knowledge graph, integrates the power grid topology and knowledge graph to form a unified graph model, and combines spatiotemporal graph neural networks and large language models to identify breakpoint risks from both quantitative data and semantic text dimensions. Furthermore, it fuses, classifies, and visualizes the two different identification results for early warning, realizing full-domain perception, spatiotemporal propagation simulation, semantic identification of unknown events, and accurate hierarchical early warning of breakpoints in the power system industry chain. This significantly improves the comprehensiveness, accuracy, and real-time performance of breakpoint identification, reduces dependence on external labeled data, and provides data-driven intelligent decision support for the security of the power system industry chain. Attached Figure Description

[0021] Figure 1 This invention provides a flowchart illustrating a method for identifying and classifying supply chain breakpoints based on multi-source data fusion. Figure 2 The present invention provides a structural diagram of a supply chain breakpoint identification and classification system based on multi-source data fusion; Figure 3 A schematic diagram of the electronic device provided by the present invention. Detailed Implementation

[0022] Embodiments of the present invention are described in detail below, examples of which are illustrated in the accompanying drawings, wherein the same or similar reference numerals denote the same or similar elements or elements having the same or similar functions throughout. The embodiments described below with reference to the accompanying drawings are exemplary and intended to explain the present invention, and should not be construed as limiting the present invention.

[0023] Example 1: This invention provides a method for identifying and classifying supply chain breakpoints based on multi-source data fusion. Specifically, Figure 1 The flowchart of the supply chain breakpoint identification and classification method based on multi-source data fusion provided in this embodiment of the invention is shown in the figure, and includes the following steps: S101: Collect multi-source data, fuse the multi-source data, and construct a multi-source heterogeneous data set; S102: Based on a multi-source heterogeneous data set, extract the entities that constitute the industrial chain structure and the relationships between entities, and construct a knowledge graph that represents the industrial chain structure of the power system. S103: Based on the power grid topology in a multi-source heterogeneous data set, and combined with a knowledge graph, a unified graph model is formed; based on each node in the unified graph model, multi-dimensional features of each node are extracted from the multi-source heterogeneous data set. S104: Input the unified graph model and multi-dimensional features into the spatiotemporal graph neural network model, calculate the breakpoint risk of each node and the propagation trend of the breakpoint along the dependent path, and obtain the breakpoint risk identification result. S105: Based on text data in a multi-source heterogeneous dataset, use a large language model to perform semantic understanding and reasoning on the text data, obtain breakpoint event information in the text data, and use the breakpoint event information as the semantic breakpoint recognition result. S106: Based on the fusion result of the breakpoint risk identification result and the semantic breakpoint identification result, breakpoint classification is performed, breakpoint identification results are generated, and early warning is issued.

[0024] This invention constructs a multi-source heterogeneous data set and an industry chain knowledge graph, integrates the power grid topology and knowledge graph to form a unified graph model, and combines spatiotemporal graph neural networks and large language models to identify breakpoint risks from both quantitative data and semantic text dimensions. Furthermore, it fuses, classifies, and visualizes the two different identification results for early warning, realizing full-domain perception, spatiotemporal propagation simulation, semantic identification of unknown events, and accurate hierarchical early warning of breakpoints in the power system industry chain. This significantly improves the comprehensiveness, accuracy, and real-time performance of breakpoint identification, reduces dependence on external labeled data, and provides data-driven intelligent decision support for the security of the power system industry chain.

[0025] In this invention, multi-source data includes, but is not limited to: Internal power grid operation data includes: telemetry and telemetry data output from the SCADA (Supervisory Control and Data Acquisition) system; equipment ledgers and maintenance records from the PMS (Production Management System); and the power grid topology from the GIS (Geographic Information System). The SCADA system collects real-time power grid operation data (voltage, current, frequency, switch status, etc.) and is the source of core internal power grid operation data. The PMS system manages equipment ledgers, maintenance records, and equipment status assessments, providing static attribute data for equipment. The GIS system manages the power grid topology (geographical location and connection relationships of substations and lines), providing spatial structure data of the power grid.

[0026] Supply chain data includes: purchase orders and supplier information from ERP (Enterprise Resource Planning) systems; raw material price indices and key component inventory levels from third-party supply chain platforms; and real-time location data of logistics vehicles collected by IoT devices. ERP systems manage purchase orders, supplier information, contract fulfillment, and inventory management, providing structured data across the supply chain.

[0027] External environment data: publicly available policy documents, industry media news, corporate announcements, and social media texts.

[0028] After obtaining multi-source data, the fusion process includes: data cleaning (filling missing values ​​and removing outliers) of each data source, timestamp alignment (unifying all data to the same time coordinate system), and entity alignment (mapping different names describing the same device or enterprise in different data sources to a unified identifier), forming a multi-source heterogeneous data set indexed by a unified entity identifier and timestamp.

[0029] In some optional implementations, based on multi-source heterogeneous data sets, entities constituting the industrial chain structure and the relationships between entities are extracted to construct a knowledge graph representing the power system industrial chain structure, including: For text data in multi-source heterogeneous datasets, a sequence labeling recognition model is used to identify multiple entities that constitute the industrial chain structure. These entities include at least one of equipment entities, enterprise entities, raw material entities, and project entities. A remote supervision-based relationship classification model is adopted, combined with a predefined set of relationship types, to determine the relationships between entities. These relationships include at least one of supply relationships, dependency relationships, and transportation relationships. Each identified entity is treated as a node, and each determined relationship is treated as a relationship edge. A knowledge graph is constructed and stored in a graph database using an attribute graph model.

[0030] Specifically, for text data in multi-source heterogeneous datasets (mainly historical texts used to construct knowledge graphs), sequence labeling recognition models (such as BiLSTM-CRF (Bidirectional Long Short-Term Memory - Conditional Random Field)) are used to identify multiple entities constituting the industrial chain structure, including equipment entities, enterprise entities, raw material entities, and project entities. A remotely supervised relationship classification model is adopted, combined with a predefined set of relationship types (including supply relationships, dependency relationships, and transportation relationships), to determine the relationships between entities; Each identified entity is treated as a node, and each determined relationship is treated as an edge. These are stored in a graph database (such as Neo4j) using an attribute graph model. Each node is assigned static attributes (such as enterprise capacity and geographical location), and each edge is assigned attributes of relationship strength and time validity.

[0031] Before using BiLSTM-CRF, the BiLSTM-CRF model needs to be pre-trained using labeled corpora in the power system industry chain.

[0032] The training data is constructed as follows: documents related to the power industry chain are selected from historical texts of the power industry (such as power grid bidding announcements, equipment manufacturer news, and industry research reports); equipment entities, enterprise entities, raw material entities, and project entities in the documents are labeled manually to form a labeled corpus; the labeled corpus has no less than 5,000 sentences and covers typical entity types in each link of the power system industry chain.

[0033] During training, the model learns the contextual features of each character or word using both forward LSTM (Long Short-Term Memory) and backward LSTM. These feature vectors are then input into the CRF layer. The CRF layer learns the transition probabilities between labels (e.g., the probability of "B-ORG" being followed by "I-ORG" is much higher than that of "B-PER"), and outputs the globally optimal label sequence. After training, the model can accurately identify various industry chain entities in power industry texts.

[0034] Using the BiLSTM-CRF model has the following advantages: Context sensitivity: The bidirectional structure of BiLSTM can simultaneously capture the contextual information of characters or words, demonstrating good recognition capabilities for common entity features in power industry texts (such as "Limited Company," "Power Grid Company," and "Substation"). For example, in the text "Baobian Electric won the bid for the UHV transformer project," the forward LSTM captures the "electric" suffix feature, while the backward LSTM captures the "won the bid" action feature, jointly identifying "Baobian Electric" as a corporate entity.

[0035] Label sequence validity: The CRF layer ensures the validity of the output label sequence through the label transition probability matrix. For example, the CRF layer avoids unreasonable sequences such as "I-ORG" directly followed by "B-PER", thereby ensuring the accuracy of entity boundaries.

[0036] Domain adaptability: Through training with labeled corpora in the power industry, the model can learn the unique entity naming patterns in the power industry chain. For example, equipment entities often end with "transformer", "circuit breaker", "GIS", etc., enterprise entities often end with "group", "company", "factory", etc., and raw material entities often end with "steel", "copper", "silicon", etc.

[0037] In some optional implementations, a unified graph model is formed by fusing the power grid topology from a multi-source heterogeneous dataset with a knowledge graph. This includes: mapping power grid nodes in the power grid topology to a first node set in the unified graph model, where power grid nodes include substation nodes and line nodes; mapping entities constituting the industry chain structure in the knowledge graph to a second node set in the unified graph model, where the second node set and the first node set share at least one node, and the shared node is a device entity that exists in both the power grid topology and the knowledge graph; mapping physical connection edges in the power grid topology to a first edge set in the unified graph model, mapping relational edges in the knowledge graph to a second edge set in the unified graph model, and labeling the edge types of the physical connection edges and relational edges associated with the shared nodes; and constructing the unified graph model by using the union of the first node set and the second node set as the node set of the unified graph model and the union of the first edge set and the second edge set as the edge set of the unified graph model.

[0038] Specifically, the power grid nodes (including substation nodes and line nodes) in the power grid topology are mapped to the first node set of the unified graph model; entities constituting the industry chain structure in the knowledge graph are mapped to the second node set of the unified graph model. The first and second node sets share at least one node, which is a device entity (such as a transformer) existing in both the power grid topology and the knowledge graph, serving as an anchor point for the fusion of the two graph structures. Physical connection edges in the power grid topology are mapped to the first edge set of the unified graph model, and relational edges in the knowledge graph are mapped to the second edge set of the unified graph model. Edge type labels are applied to the physical connection edges and relational edges associated with the shared node. The union of the first and second node sets is used as the node set of the unified graph model, and the union of the first and second edge sets is used as the edge set of the unified graph model, thus constructing the unified graph model.

[0039] Preferably, based on each node in the unified graph model, multi-dimensional features of each node are extracted from a multi-source heterogeneous data set, including: extracting static attribute feature vectors of each node from the multi-source heterogeneous data set based on each node in the unified graph model; the static attribute feature vectors include at least one or more of equipment capacity, enterprise annual production capacity, geographical location coordinates, and node type codes; extracting time series data associated with each node from the multi-source heterogeneous data set, and forming dynamic time series feature vectors after normalization; the dynamic time series feature vectors include one or more of equipment operating parameter sequences, inventory change sequences, raw material price sequences, and logistics status sequences; and concatenating the static attribute feature vectors and the dynamic time series feature vectors to form multi-dimensional feature vectors of each node.

[0040] Specifically, for each node in the unified graph model, the static attribute feature vector of the node is extracted from the multi-source heterogeneous data set, including equipment capacity, annual production capacity of the enterprise, geographical location coordinates and node type encoding; For each node, time series data associated with that node are extracted from a multi-source heterogeneous data set, and after normalization, a dynamic time series feature vector is formed, including equipment operating parameter sequence, inventory change sequence, raw material price sequence, and logistics status sequence. Static attribute feature vectors are concatenated with dynamic temporal feature vectors to form multi-dimensional feature vectors for each node.

[0041] Preferably, a unified graph model and multi-dimensional features are input into a spatiotemporal graph neural network model to calculate the breakpoint risk of nodes and the propagation trend of breakpoints along dependent paths, thereby obtaining breakpoint risk identification results, including: The unified graph model and multi-dimensional features are input into the spatiotemporal graph convolutional network. Through the forward propagation of the spatiotemporal graph convolutional network, the spatiotemporal graph convolutional network is used to extract temporal features in the time dimension using a one-dimensional convolutional layer, and in the spatial dimension, the graph convolutional layer is used to aggregate the feature information of adjacent nodes according to the edge set of the unified graph model. Based on temporal features and feature information, the breakpoint risk probability value of each node at the current time and the propagation probability value of each edge at the current time are output. Nodes with breakpoint risk probability values ​​exceeding a preset threshold are used as source nodes, and propagation probability values ​​are used as edge weights. Multiple breakpoint propagation paths are generated using a depth-first search algorithm. Each breakpoint propagation path includes propagation direction and propagation time estimation. The probability value of breakpoint risk and the breakpoint propagation path are determined as the breakpoint risk identification results.

[0042] Specifically, a Spatial Temporal Graph Convolutional Network (ST-GCN) is used as the spatiotemporal graph neural network model. A unified graph model and multi-dimensional features are input into the network. Through forward propagation, one-dimensional convolutional layers are used to extract temporal features in the time dimension, and graph convolutional layers are used in the spatial dimension to aggregate the feature information of adjacent nodes according to the edge set of the unified graph model.

[0043] Based on the extracted temporal features and aggregated spatial features, the network output layer outputs the breakpoint risk probability value of each node at the current time and the propagation probability value of each edge at the current time.

[0044] Nodes with a breakpoint risk probability value exceeding a preset threshold (e.g., 0.7) are used as source nodes, and the propagation probability value is used as the edge weight. Multiple breakpoint propagation paths are generated using a depth-first search algorithm, and each path includes the propagation direction and propagation time estimate.

[0045] The probability value of breakpoint risk and the breakpoint propagation path are determined as the breakpoint risk identification results.

[0046] In this invention, the Spatiotemporal Graph Convolutional Network (ST-GCN) model is used to model the spatiotemporal propagation of breakpoint risk in the supply chain dependency network by alternately performing graph convolution in the spatial dimension and one-dimensional convolution in the temporal dimension.

[0047] (1) The structure of the model is as follows: ① Input layer The model input consists of two parts: Graph structure input: An adjacency matrix of a unified graph model, recording the connections between nodes, including physical connection edges in the power grid topology and logical dependency edges (supply relationships, dependencies, and transportation relationships) in the knowledge graph. Each edge is assigned a different weight according to its type to distinguish the difference in the impact of physical connections and logical dependencies on breakpoint propagation.

[0048] Node feature input: A multi-dimensional feature matrix for each node, with a time window length of 30 days and 128 features per time step. This feature matrix integrates the node's static attribute features (equipment capacity, enterprise production capacity, geographical location) and dynamic time-series features (operating parameters, inventory changes, price fluctuations).

[0049] ② Spatial graph convolutional layer The spatial dimension employs a graph convolutional network for feature aggregation, consisting of two graph convolutional layers. The operation of each graph convolutional layer is as follows: for each node, its own features are weighted and aggregated with the features of all its neighboring nodes. The weighting coefficients are determined by the edge weights between nodes and the node degree, essentially allowing each node to "hear" information from its neighbors. The first graph convolutional layer maps the input feature dimension from 128 dimensions to 256 dimensions, and the second layer maps the 256 dimensions to 128 dimensions. Each graph convolutional layer is followed by a ReLU activation function to introduce non-linear expressive power.

[0050] ③ Temporal convolutional layer Temporal features are extracted using a one-dimensional convolutional network, consisting of three layers. Each convolutional kernel has a size of 3, meaning the convolution operation is performed every three time steps along the time axis. The convolutional layers use "same" padding to ensure the input sequence length remains constant. The first convolutional layer reduces the number of channels from 128 to 64, the second from 64 to 32, and the third from 32 to 16. Each convolutional layer is followed by a ReLU activation function and a Dropout layer (randomly discarding 20% ​​of neurons) to prevent overfitting.

[0051] ④ Spatiotemporal fusion method To simultaneously model spatiotemporal dependencies, the model adopts an alternating "space-time" stacking structure: first, a graph convolutional layer aggregates spatial features, then a temporal convolutional layer extracts temporal features, and this stacking is repeated twice. This alternating structure enables breakpoint risk information to propagate along the industrial chain path in the spatial dimension, while evolving along historical trends in the temporal dimension, achieving comprehensive modeling of the breakpoint propagation pattern.

[0052] ⑤ Output layer After spatiotemporal feature extraction, the model calculates the results through two independent output branches: Node breakpoint risk probability: The final feature vector of each node is input into a fully connected layer, and after passing through the Sigmoid activation function, a probability value between 0 and 1 is output, representing the risk level of the node becoming a breakpoint. The Sigmoid function maps any real number to the interval (0,1), making it easy to interpret as a probability.

[0053] Edge propagation probability: For each edge in the unified graph model, the propagation probability value between 0 and 1 is calculated through a fully connected layer, representing the probability that the breakpoint will propagate from the source node to the target node.

[0054] (2) The model training process is as follows: Before applying the model, the ST-GCN model needs to be pre-trained using historical breakpoint event data.

[0055] ① Training data construction Known breakpoint events (such as equipment supply interruptions, raw material shortages, and logistics delays) are selected from historical power system operation data. Using the breakpoint occurrence time as a baseline, multi-source data from 30 time steps (i.e., 30 days) prior to the breakpoint occurrence are extracted as positive samples for training. Simultaneously, nodes and time periods without breakpoints are randomly selected as negative samples to construct a balanced training dataset.

[0056] ② Loss Function Design The model employs a multi-task learning framework, simultaneously optimizing both node risk prediction and edge propagation prediction tasks. Node prediction loss: A binary cross-entropy loss is used to measure the difference between the node risk probability predicted by the model and the true label. The loss value is small when the predicted probability is close to the true label, and large when the predicted probability is far from the true label.

[0057] Edge prediction loss: Also using binary cross-entropy loss, it measures the difference between the edge propagation probability predicted by the model and the actual propagation.

[0058] Total loss: The node prediction loss and the edge prediction loss are weighted and summed. The weight coefficient of the edge prediction loss is set to 0.5 to balance the optimization objectives of the two tasks.

[0059] ③ Optimizer and Training Strategy The Adam optimizer is used for parameter updates. The initial learning rate is set to 0.001, and the learning rate is decayed to 50% of the original rate every 50 training epochs.

[0060] During each training session, 32 samples are randomly selected from the dataset as a batch for parameter updates.

[0061] The training is planned for 200 epochs, with 20% of the training data used as a validation set. Training will be terminated early if the validation set loss stops decreasing for 20 consecutive epochs to prevent overfitting.

[0062] (3) The process of model reasoning and application is as follows: After training, the unified graph model and multi-dimensional features are input into the trained ST-GCN model, and the following processing is performed through forward propagation: In the time dimension, the one-dimensional convolutional layer extracts the temporal variation features of each node; In the spatial dimension, graph convolutional layers aggregate the feature information of adjacent nodes according to the edge set of a unified graph model; Based on the extracted temporal features and aggregated spatial features, the output layer outputs the breakpoint risk probability value of each node at the current time and the propagation probability value of each edge at the current time.

[0063] Preferably, based on text data from a multi-source heterogeneous dataset, a large language model is used to perform semantic understanding and reasoning on the text data to obtain breakpoint event information in the text data. This breakpoint event information is then used as the semantic breakpoint identification result, including: New text data is selected from a multi-source heterogeneous dataset. The new text data includes at least one of the following: news information, corporate announcements, policy documents, and industry reports. Obtain a preset number of high-risk nodes with the highest risk probability values ​​from the breakpoint risk identification results, and extract subgraph information associated with each high-risk node from the knowledge graph; Based on the subgraph information of each high-risk node, a prompt word is constructed. The prompt word includes one or more of the following fields: role setting field, text content field to be analyzed field, context reference field, and output format constraint field. Natural language output is obtained using a large language model based on prompt words; By combining regular expression matching with a rule-based parser, one or more of the following are extracted from the natural language output: event type, affected entity identifier, event occurrence time, event duration, and confidence score, to form breakpoint event information. This breakpoint event information is then used as the semantic breakpoint recognition result.

[0064] The text data processed in this step is mainly newly acquired text data in real time, used to identify currently occurring breakpoint events. Specifically, (1) Model selection and domain adaptation This embodiment uses a general-purpose large language model such as GPT-4o or Llama 3. To ensure the general-purpose model accurately identifies power system supply chain disruption events, the following domain adaptations are required: Search Enhancement Generation (RAG) Mechanism Since the general-purpose large model has not been specifically trained in the power system industry chain, directly inputting text may lead to inaccurate output or produce illusions. Therefore, this solution adopts a retrieval-enhanced generation mechanism, injecting subgraph information related to high-risk nodes from the knowledge graph into the prompt words as the contextual background for model inference.

[0065] The subgraph information is constructed as follows: Obtain nodes whose risk probability value exceeds a preset threshold from the breakpoint risk identification results; Extract the neighbors and relationships within two hops of these nodes from the knowledge graph; Convert the nodes and relationships in the subgraph into text descriptions, such as: "Oriented silicon steel is supplied by company D and downstream to supplier A for the production of ultra-high voltage transformers."

[0066] By injecting domain knowledge, large language models can reason based on an understanding of the industry chain structure, reducing illusions and improving the accuracy of breakpoint event identification.

[0067] Furthermore, for application scenarios requiring higher recognition accuracy, fine-tuning can be used to adapt the basic large model to the specific domain.

[0068] The method for constructing fine-tuned training data is as follows: Filter documents related to breakpoint events from historical texts of the power industry; The event type (materials disruption / logistics disruption / policy disruption, etc.), affected entities, start time of impact, and duration of impact are labeled for each document using manual annotation. Convert the labeled data into an "input text-output JSON" format to form a fine-tuning training set with a size of no less than 500 records; By employing efficient parameter fine-tuning methods such as low-rank adaptation (LoRA), the model's ability to identify breakpoint events in the power industry chain is enhanced while maintaining its basic capabilities.

[0069] (2) Logic of constructing prompt words The quality of the prompt words directly determines the accuracy of the output of a large language model. This embodiment uses a structured prompt word template, which includes the following fields: ① Prompt word template design The prompts use a multi-segment structure, employing explicit delimiters and identifiers to distinguish different fields: [System Role] You are a power system supply chain security analysis expert, skilled at identifying supply chain disruption risks from news and announcements.

[0070] [Task Background] The following are the supply chain dependencies relevant to this analysis: {Subgraph Information Text Description} [Text to be analyzed] {Add text content} [Analysis Steps] Please think about this step by step: First, determine whether the event may have an impact on the power system industry chain; If there is an impact, identify the specific entities affected (businesses, equipment, raw materials, or projects). Determine the event type (material disruption / logistics disruption / funding disruption / technology disruption / policy disruption); Estimate the start time and duration of the impact.

[0071] [Output Format] Please strictly adhere to the following JSON format when outputting; do not add any extra text: { "Does it have an impact?": "Yes / No" "Event Type": "Material Disruption / Logistics Disruption / Funding Disruption / Technology Disruption / Policy Disruption" “Affected Entities”: [“Entity 1”, “Entity 2”], "Affects start time": "YYYY-MM-DD" Duration of impact: "Number of days or 'unknown'" "Confidence level": A number between 0 and 1 } ② Field descriptions are shown in Table 1.

[0072] Table 1 ③ The principle of mind chain design Adding an "Analysis Steps" field to the prompt words forces the model to reason step by step, rather than directly guessing the result. This is mainly based on the following considerations: Reduce illusions: Require the model to first determine "whether there is an impact". If the determination is "no", subsequent fields can output null values ​​to avoid the model forcibly fabricating event types when there is no impact. To improve accuracy, complex tasks are broken down into four relatively simple sub-steps, making it easier for the model to complete them correctly. Enhanced interpretability: The reasoning process of the model can be traced through the output of the thought chain, which facilitates the subsequent optimization of prompt words.

[0073] (3) Model calling and output parsing ① Model call Send the completed prompt words to the Large Language Model API (such as the OpenAI API or a locally deployed Llama 3 model) and set the model parameters: The temperature is set to 0.2 to reduce the randomness of the output and ensure the stability of the results; Set the maximum output length (Max Tokens) to 500 to ensure there is enough space to output the complete JSON; Set the stop sequences to "\n\n" to avoid output exceeding the expected range.

[0074] ② Output Analysis The natural language output returned by the model may contain additional explanatory text, therefore the JSON (JavaScript Object Notation) portion needs to be extracted using a parser. The parsing process is as follows: First, use regular expressions to match the JSON objects in the output, with the matching pattern "{[^}". }”; If multiple JSON objects are matched, the last one is taken (the model may output multiple JSON objects during inference, and the final result is usually at the end). If a complete JSON cannot be matched, try to extract key information field by field (such as "event type: material breakpoint") and use a rule-based parser for fallback processing; The extracted fields are converted into structured data to form semantic breakpoint event records.

[0075] Preferably, breakpoints are classified based on the breakpoint risk identification results and semantic breakpoint identification results, and breakpoint identification results are generated and warnings are issued, including: The breakpoint risk probability value of each node in the breakpoint risk identification result is used as the first evidence, and the confidence score of each event record in the semantic breakpoint identification result is used as the second evidence. The DS evidence theory is used to calculate the integrated breakpoint confidence after fusion. When the overall breakpoint confidence exceeds the first threshold, the industrial chain link corresponding to the node or event is identified as a breakpoint. Based on the identified breakpoints, and according to the node type or event type that triggered the breakpoint, they are mapped to a preset breakpoint classification system to determine the category and nature of the breakpoint. The urgency level of each breakpoint is determined based on the comprehensive breakpoint confidence level and the estimated propagation time in the breakpoint propagation path. A breakpoint identification report is generated, which includes one or more of the following: breakpoint identifier, link category, nature category, scope of impact, expected impact time window, and suggested countermeasures. The breakpoint identification report and the breakpoint propagation path diagram are displayed through a visual interface for graded early warning.

[0076] Specifically, the DS evidence theory requires mapping the numerical values ​​of different pieces of evidence to the same frame of discernment, that is, the set of all possible propositions. In this embodiment, the frame of discernment is defined as: Θ = {Breakpoint occurred, breakpoint did not occur} The numerical sources and physical meanings of the two pieces of evidence are shown in Table 2:

[0077] Table 2 The two values ​​differ in a physical sense: the quantified risk probability is based on historical data statistical patterns, while the semantic confidence is based on text semantic understanding. To integrate them within a unified framework, the two values ​​need to be mapped to the basic trust allocation (mass function) in DS evidence theory.

[0078] The basic trust assignment function is constructed as follows: ① Construction of the mass function for quantifying risk evidence Let p be the probability value of the node breakpoint risk output by the spatiotemporal graph neural network (0≤p≤1). This value represents the likelihood that the breakpoint will occur, as the model believes it is possible. However, the quantization model itself has uncertainties and cannot be 100% certain. Therefore, a confidence coefficient α (α∈[0,1]) is introduced to represent the degree of confidence in the output of the quantization model, and the remaining part is allocated as uncertainty.

[0079] The mapping rules are as follows: Trust assignment for the proposition "breakpoint occurrence": m1(breakpoint occurrence) = α × p Assigning confidence to the proposition "the breakpoint did not occur": m1(the breakpoint did not occur) = α × (1-p) Trust allocation (uncertainty) for the overall identification framework: m1(Θ)=1-α ②Construction of the mass function for semantic evidence Let s be the semantic confidence score output by the large language model (0 ≤ s ≤ 1). This value represents the model's confidence level in judging text breakpoint events. Semantic models also exhibit uncertainty; therefore, a confidence coefficient β is introduced to represent the degree of trust in the semantic model's output.

[0080] The mapping rules are as follows: Assignment of confidence to the proposition "breakpoint occurs": m²(breakpoint occurs) = β × s Trust assignment for the proposition "breakpoint did not occur": m2(breakpoint did not occur) = 0 (the semantic model only outputs breakpoint event information and does not actively output the judgment of "did not occur") Trust allocation for the entire recognition framework: m2(Θ)=1-β×s ③ Explanation of the rationality of the mapping The core idea of ​​the above mapping method is to treat the original numerical values ​​as the "tendency" of evidence towards a specific proposition, while preserving the "uncertainty" about the source of evidence itself through the confidence coefficient. This mapping has the following advantages: Preserve the relative relationship of the original values: the larger the p value, the larger m1 (the breakpoint occurs), which is intuitive; Reflecting the confidence level of the evidence source: The values ​​of α and β are based on the model's performance on the validation set and have an engineering basis; Uncertainty is transitive: The existence of m(Θ) allows the fusion result to retain uncertainty, avoiding forced decision-making.

[0081] In practical applications, the confidence coefficient can be adjusted based on the performance of the specific model. When the model accuracy improves, the confidence coefficient can be increased accordingly; when the model accuracy is low, the confidence coefficient should be decreased and the uncertainty allocation increased.

[0082] When the overall breakpoint confidence exceeds the first threshold (e.g., 0.6), the industrial chain link corresponding to the node or event is identified as a breakpoint. For each identified breakpoint, based on the node type or event type that triggered the breakpoint, it is mapped to a preset breakpoint classification system to determine the breakpoint's category (power generation side, power grid side, power consumption side, equipment manufacturing side) and nature category (material breakpoint, logistics breakpoint, funding breakpoint, technology breakpoint, policy breakpoint). Based on the overall confidence level of the breakpoints and the estimated propagation time in the breakpoint propagation path, the urgency level of each breakpoint is determined (already occurred, about to occur, potential risk). Generate a breakpoint identification report, including breakpoint identifier, link category, nature category, urgency level, impact scope, expected impact time window, and suggested countermeasures. Display the breakpoint identification report and breakpoint propagation path diagram through a visual interface, and provide tiered early warnings (e.g., red, orange, and yellow alerts).

[0083] The beneficial effects of this invention are: (1) Global perception ability By integrating internal power grid operation data, upstream and downstream industry chain data, and external environmental data, a multi-source heterogeneous data set is constructed. This overcomes the limitations of traditional methods that rely solely on internal operation data and can identify hidden disruptions in the upstream of the industry chain in advance. For example, by using external data such as raw material price fluctuations, changes in supplier capacity, and abnormal logistics status, early warnings can be issued regarding supply risks of key components.

[0084] (2) Spatiotemporal propagation deduction By integrating the power grid topology with knowledge graphs into a unified graph model, and combining it with a spatiotemporal graph neural network, it is possible to simulate the propagation path of breakpoints in the physical power grid and the supply chain dependency network, enabling advanced prediction of the impact of breakpoints. For example, it can predict when raw material shortages will spread to equipment manufacturers and then to projects under construction, providing a precise time window for emergency dispatch.

[0085] (3) Intelligent semantic understanding By leveraging large language models to process unstructured text data and combining them with contextual information provided by knowledge graphs, the system's ability to identify unknown and sudden breakpoint events is enhanced. For example, for events that are difficult to perceive through quantitative data, such as "a country implements export controls" or "a factory experiences a sudden fire," large language models can directly identify and extract breakpoint information from news reports, compensating for the insufficient generalization ability of rule-driven models.

[0086] (4) Reduce data dependence By using knowledge graphs to enhance the generation of large language models, the quantitative analysis results (high-risk nodes) are used as the focus of semantic analysis, and the industrial chain dependencies in the knowledge graph are used as the context. This reduces the reliance on massive amounts of labeled data and enables relatively accurate breakpoint classification even with a small number of samples.

[0087] (5) Tiered early warning and decision support By integrating quantitative analysis and semantic recognition results using DS evidence theory, a breakpoint identification report is generated, which includes breakpoint identifiers, link categories, nature categories, urgency levels, impact scope, expected impact time windows, and suggested countermeasures. The report is then used to provide tiered early warnings through a visual interface, providing data-driven intelligent decision support for the security of the power system industry chain.

[0088] The following section uses the identification of supply chain breakpoints in an ultra-high voltage transformer in a power transmission and transformation project of a regional power grid company as an example to illustrate the data fusion-based industrial chain breakpoint identification and classification method provided by this invention.

[0089] Step 1: Data Acquisition and Fusion The system collects the following data in advance: Internal data: power grid planning information (the estimated start date of Project C is April 1, 2024, and the equipment requirement list includes one UHV transformer), and historical procurement records (supplier A has a historical fulfillment rate of 95%).

[0090] Supply chain data: Supplier A's production capacity (annual capacity of 50 units), price trend of core raw material oriented silicon steel (price series over the past 30 days, up 20% in the last week), and GPS trajectory of logistics company B's vehicle (currently located at supplier A's factory area, expected to arrive at project C in 3 days).

[0091] External data: Weather forecast for Supplier A's location (no abnormalities expected in the coming week), industry news ("Company D's grain-oriented silicon steel production base has suspended production for rectification due to environmental inspection, with the suspension period expected to be one month").

[0092] The above data is cleaned (missing value filling, outlier removal), timestamp aligned (unified to 0:00 every day), and entity aligned (mapping "Baobian Electric" and "Supplier A" as the same entity), forming a multi-source heterogeneous data set indexed by unified entity identifiers and timestamps.

[0093] Step 2: Construct a knowledge graph The BiLSTM-CRF model was used to extract the entities "Company D", "Oriented Silicon Steel", and "Suspension of Production for Rectification" from the industry news article "Company D's Oriented Silicon Steel Production Base Suspends Production for Rectification Due to Environmental Inspection". A remote supervision-based relation classification model, combined with a predefined set of relation types, was used to determine the following relations: Company D has a "supply relationship" with grain-oriented silicon steel. There is a "dependency relationship" between grain-oriented silicon steel and ultra-high voltage transformers. Supplier A has a "supply relationship" with the ultra-high voltage transformer. There is a dependency relationship between Project C and the ultra-high voltage transformer. There is a "transportation relationship" between logistics company B and supplier A. The identified entities are treated as nodes, and the relationships as edges, stored in the Neo4j graph database using an attribute graph model. Nodes are assigned static attributes: Supplier A node is assigned "annual production capacity of 50 units," and Logistics Company B node is assigned "transportation time of 3 days." Edges are assigned attributes: Enterprise D → oriented silicon steel edge is assigned "supply share of 60%," and Supplier A → UHV transformer edge is assigned "historical fulfillment rate of 95%." This forms a knowledge graph of the industry chain.

[0094] Step 3: Construct a unified graph model Map substation nodes (such as "Project C Substation") in the power grid topology to the first node set. Map entities "UHV Transformer", "Supplier A", "Company D", "Oriented Silicon Steel", "Project C", and "Logistics Company B" in the knowledge graph to the second node set. Among them, "Project C Substation" and "Project C" are shared nodes (they are the same physical object, i.e., the substation connected to Project C).

[0095] Map the physical connection edges in the power grid topology (such as the connection line between "Project C Substation" and the upstream substation) to the first edge set. Map the supply, dependency, and transportation relationship edges in the knowledge graph to the second edge set. Label the physical connection edges and knowledge graph edges associated with the shared node "Project C" with edge type as "physical-dependency composite edge".

[0096] We construct a unified graph model by taking the union of the first and second node sets as the node set and the union of the first and second edge sets as the edge set. This unified graph model simultaneously includes the physical topology of the power grid and the logical dependency structure of the industry chain.

[0097] Step 4: Extract multi-dimensional features For each node in the unified graph model, extract static attribute features from a multi-source heterogeneous dataset: UHV transformer node: Equipment capacity = 1000MVA, Node type code = Equipment (01) Supplier A Node: Annual production capacity = 50 units, Geographic coordinates = (116.40°E, 39.90°N), Node type code = Enterprise (02) Oriented silicon steel joint: Raw material type code = Raw material (03) Project C node: Project type code = Project (04) For each node, extract dynamic temporal features: Grain-oriented silicon steel node: The price series over the past 30 days [3200, 3250, ..., 3800] yuan / ton, after normalization, forms a vector [0.2, 0.3, ..., 0.8]. Supplier A node: Inventory change sequence over the past 30 days [100, 98, ..., 70] tons, normalized to form a vector [0.9, 0.8, ..., 0.3]. Logistics company node B: The transportation trajectory sequence of the past 7 days, which is normalized to form a vector.

[0098] Static and dynamic features are concatenated to form multi-dimensional feature vectors for each node, with a unified dimension of 128.

[0099] Step 5: Quantitative Breakpoint Risk Identification A unified graph model and multi-dimensional feature vectors are input into a spatiotemporal graph convolutional network. The network structure consists of 3 one-dimensional convolutional layers (kernel size 3, stride 1), 2 graph convolutional layers (aggregating 2-hop neighbor node information), and a fully connected output layer (activation function sigmoid).

[0100] Through forward propagation, the network outputs: Risk probability value for grain-oriented silicon steel nodes: 0.85; Supplier A node risk probability value: 0.78; The probability value of risk at the UHV transformer node is 0.72. Risk probability value for node C of the project: 0.65.

[0101] Side propagation probability value: Probability of propagation from grain-oriented silicon steel to supplier A: 0.82; Probability of propagation from supplier A to the UHV transformer: 0.75; Probability of propagation from UHV transformer to side C of the project: 0.68.

[0102] Using nodes with a risk probability value > 0.7 as source nodes (oriented silicon steel, supplier A, UHV transformer), and weighted by propagation probability values, a depth-first search algorithm is used to generate breakpoint propagation paths: Among them, Path 1: Grain-oriented silicon steel (risk 0.85) → Supplier A (propagation time 7 days) → UHV transformer (propagation time 14 days) → Project C (propagation time 21 days); Path 2: Supplier A (risk 0.78) → UHV transformer (propagation time 7 days) → Project C (propagation time 14 days); The probability value of breakpoint risk and the breakpoint propagation path are determined as the breakpoint risk identification results.

[0103] Step 6: Semantic Breakpoint Identification New text was selected from the multi-source heterogeneous dataset: "The company's D-oriented silicon steel production base has suspended production for rectification due to environmental inspection, and the suspension period is expected to be 1 month."

[0104] Obtain nodes with a risk probability value > 0.7 (oriented silicon steel, supplier A, UHV transformer) from the breakpoint risk identification results, and extract subgraph information for these nodes from the knowledge graph: Grain-oriented silicon steel node: Upstream dependent company D (supply relationship), downstream supplier A; Supplier A node: upstream relies on grain-oriented silicon steel, downstream supplies ultra-high voltage transformers; Ultra-high voltage transformer node: upstream supplier A, downstream supplier project C.

[0105] The prompt words are constructed as follows: Character setting: You are a power system supply chain security analysis expert; Contextual Reference: Grain-oriented silicon steel is a key raw material supplied by company D, which is then supplied downstream to supplier A for the production of ultra-high voltage transformers. Supplier A has an annual production capacity of 50 units and a historical contract fulfillment rate of 95%. Ultra-high voltage transformers are core equipment in project C. Text to be analyzed: The company's D-oriented silicon steel production base has suspended production for rectification due to environmental inspection, and the suspension period is expected to be one month. Output format: Please output in JSON format, including the following fields: event type (material breakpoint / logistics breakpoint / policy breakpoint), affected entities, start time of impact, duration of impact (days), and confidence level (0-1).

[0106] Send the prompt words to the GPT-4o model to obtain natural language output: { "Event Type": "Supply Breakdown" "Affected Entity": "Company D, Grain-Oriented Silicon Steel" "Affect start time": "2024-03-15" "Duration of impact": "30" Confidence level: 0.92 } By combining regular expression matching with a rule-based parser, the above fields are extracted from the output to form structured breakpoint event information, which serves as the semantic breakpoint recognition result.

[0107] Step 7: Integrate Classification and Early Warning The risk probability value of 0.85 for the oriented silicon steel node in the breakpoint risk identification results was used as the first piece of evidence, and the confidence level of 0.92 for the semantic breakpoint event record was used as the second piece of evidence.

[0108] The combined breakpoint confidence level after fusion is calculated using DS evidence theory: Let the recognition frame Θ = {breakpoint occurred, breakpoint did not occur}. mass function and These represent the basic trust allocation of the two pieces of evidence for each proposition in the identification framework. Among them, This indicates that a trust breakpoint has occurred. This indicates that the trust breakpoint did not occur. This indicates an uncertain trust that cannot distinguish between "breakpoint occurrence" and "breakpoint non-occurrence".

[0109] Assigning a breakpoint risk probability value of 0.85 as the confidence level for the first piece of evidence:

[0110] Using a semantic breakpoint confidence score of 0.92 as the trust assignment for the second piece of evidence:

[0111] Fusion formula:

[0112] Results

[0113] The overall breakpoint confidence level of 0.95 exceeds the first threshold of 0.6, thus identifying the corresponding industrial chain link of oriented silicon steel as a breakpoint.

[0114] Based on the node type (raw material node) and event type (material breakpoint) that triggered the breakpoint, it is mapped to a preset breakpoint classification system: Segment Category: Equipment Manufacturing Side (Oriented Silicon Steel is an upstream raw material for UHV transformers) Category: Material Disruption (Interruption of Key Raw Material Supply) Based on a comprehensive breakpoint confidence level of 0.95 and the estimated propagation time in the propagation path (affecting supplier A within 7 days), the urgency level is determined as follows: The overall confidence level is >0.8 and the propagation time is estimated to be within 7 days, which will affect downstream nodes (supplier A). It was determined that "a breakpoint is about to occur".

[0115] The generated breakpoint identification report is shown in Table 3.

[0116]

[0117] Table 3 The tiered early warning system and its visualization are shown below: The breakpoint identification report and breakpoint propagation path diagram are displayed through a visual interface: The breakpoint propagation path graph is presented in the form of a directed graph, with the risk probability value of each node and the propagation probability value of each edge labeled. High-risk nodes (grain-oriented silicon steel, supplier A, UHV transformer) are highlighted in red. The propagation path is marked with an orange arrow, the direction of the arrow indicates the direction of propagation, and the width of the arrow indicates the probability of propagation. Warning Level: Orange Warning (Imminent Breakdown), sent to Project Manager C and Supply Chain Management Department. The warning information includes: SMS notification: "[Orange Alert] Risk of supply disruption for grain-oriented silicon steel. It is expected to affect supplier A within 7 days and project C within 21 days. Please initiate the emergency procurement process immediately." A system pop-up window displays a summary of the breakpoint identification report and provides a "View Details" link to jump to the full report page. Example 2: Based on the same inventive concept, this invention also provides a data fusion-based supply chain breakpoint identification and classification system, the structure of which is as follows: Figure 2 As shown, the system includes: The multi-source data set construction module 201 is used to collect multi-source data, perform fusion processing on the multi-source data, and construct a multi-source heterogeneous data set. The knowledge graph construction module 202 is used to extract entities that constitute the industrial chain structure and the relationships between entities based on a multi-source heterogeneous data set, and to construct a knowledge graph that represents the industrial chain structure of the power system. The model feature determination module 203 is used to integrate the power grid topology in the multi-source heterogeneous data set with the knowledge graph to form a unified graph model; and to extract the multi-dimensional features of each node from the multi-source heterogeneous data set based on each node in the unified graph model. The breakpoint risk identification module 204 is used to input the unified graph model and multi-dimensional features into the spatiotemporal graph neural network model, calculate the breakpoint risk of each node and the propagation trend of the breakpoint along the dependent path, and obtain the breakpoint risk identification result. The semantic breakpoint recognition module 205 is used to perform semantic understanding and reasoning on text data based on text data in a multi-source heterogeneous data set, and obtain breakpoint event information in the text data, and use the breakpoint event information as the semantic breakpoint recognition result. The breakpoint classification module 206 is used to classify breakpoints based on the fusion result of breakpoint risk identification and semantic breakpoint identification, generate breakpoint identification results, and issue warnings.

[0118] Preferably, the knowledge graph construction module 202 is specifically used for: For text data in multi-source heterogeneous datasets, a sequence labeling recognition model is used to identify multiple entities that constitute the industrial chain structure. These entities include at least one of equipment entities, enterprise entities, raw material entities, and project entities. A remote supervision-based relationship classification model is adopted, combined with a predefined set of relationship types, to determine the relationships between entities. These relationships include at least one of supply relationships, dependency relationships, and transportation relationships. Each identified entity is treated as a node, and each determined relationship is treated as a relationship edge. A knowledge graph is constructed and stored in a graph database using an attribute graph model.

[0119] Preferably, the model feature determination module 203 is specifically used for: Map the power grid nodes in the power grid topology to the first node set of the unified graph model. The power grid nodes include substation nodes and line nodes. The entities that constitute the industrial chain structure in the knowledge graph are mapped to the second node set of the unified graph model, where the second node set and the first node set have at least one shared node. The shared node is a device entity that exists in both the power grid topology and the knowledge graph. The physical connection edges in the power grid topology are mapped to the first edge set of the unified graph model, and the association relationship edges in the knowledge graph are mapped to the second edge set of the unified graph model. The edge type is labeled for the physical connection edges and association relationship edges associated with shared nodes. The unified graph model is constructed by taking the union of the first node set and the second node set as the node set and taking the union of the first edge set and the second edge set as the edge set.

[0120] Preferably, the model feature determination module 203 is also specifically used for: Based on the nodes in the unified graph model, static attribute feature vectors of each node are extracted from a multi-source heterogeneous data set; the static attribute feature vectors include at least one or more of the following: equipment capacity, annual production capacity of the enterprise, geographical location coordinates, and node type code; Time series data associated with each node are extracted from a multi-source heterogeneous dataset and normalized to form a dynamic time series feature vector. The dynamic time series feature vector includes one or more of the following: equipment operating parameter sequence, inventory change sequence, raw material price sequence, and logistics status sequence. Static attribute feature vectors are concatenated with dynamic temporal feature vectors to form multi-dimensional feature vectors for each node.

[0121] Preferably, the breakpoint risk identification and acquisition module 204 is specifically used for: The unified graph model and multi-dimensional features are input into the spatiotemporal graph convolutional network. Through the forward propagation of the spatiotemporal graph convolutional network, the spatiotemporal graph convolutional network is used to extract temporal features in the time dimension using a one-dimensional convolutional layer, and in the spatial dimension, the graph convolutional layer is used to aggregate the feature information of adjacent nodes according to the edge set of the unified graph model. Based on temporal features and feature information, the breakpoint risk probability value of each node at the current time and the propagation probability value of each edge at the current time are output. Nodes with breakpoint risk probability values ​​exceeding a preset threshold are used as source nodes, and propagation probability values ​​are used as edge weights. Multiple breakpoint propagation paths are generated using a depth-first search algorithm. Each breakpoint propagation path includes propagation direction and propagation time estimation. The probability value of breakpoint risk and the breakpoint propagation path are determined as the breakpoint risk identification results.

[0122] Preferably, the semantic breakpoint recognition module 205 is specifically used for: New text data is selected from a multi-source heterogeneous dataset. The new text data includes at least one of the following: news information, corporate announcements, policy documents, and industry reports. Obtain a preset number of high-risk nodes with the highest risk probability values ​​from the breakpoint risk identification results, and extract subgraph information associated with each high-risk node from the knowledge graph; Based on the subgraph information of each high-risk node, a prompt word is constructed. The prompt word includes one or more of the following fields: role setting field, text content field to be analyzed field, context reference field, and output format constraint field. Natural language output is obtained using a large language model based on prompt words; By combining regular expression matching with a rule-based parser, one or more of the following are extracted from the natural language output: event type, affected entity identifier, event occurrence time, event duration, and confidence score, to form breakpoint event information. This breakpoint event information is then used as the semantic breakpoint recognition result.

[0123] Preferably, the breakpoint classification module 206 is specifically used for: The breakpoint risk probability value of each node in the breakpoint risk identification result is used as the first evidence, and the confidence score of each event record in the semantic breakpoint identification result is used as the second evidence. The DS evidence theory is used to calculate the integrated breakpoint confidence after fusion. When the overall breakpoint confidence exceeds the first threshold, the industrial chain link corresponding to the node or event is identified as a breakpoint. Based on the identified breakpoints, and according to the node type or event type that triggered the breakpoint, they are mapped to a preset breakpoint classification system to determine the category and nature of the breakpoint. The urgency level of each breakpoint is determined based on the comprehensive breakpoint confidence level and the estimated propagation time in the breakpoint propagation path. A breakpoint identification report is generated, which includes one or more of the following: breakpoint identifier, link category, nature category, scope of impact, expected impact time window, and suggested countermeasures. The breakpoint identification report and the breakpoint propagation path diagram are displayed through a visual interface for graded early warning.

[0124] Example 3: Based on the same inventive concept, such as Figure 3As shown, the present invention also provides an electronic device, which may be a computer device, a microcontroller device, a smart mobile device, etc. The electronic device in this embodiment may include a processor, a memory, a transceiver component, etc. The memory, processor, and transceiver component are connected via a bus; the memory can be used to store executable programs, and an exemplary executable program may include instructions; the processor is used to execute the instructions stored in the memory. The memory can also be used to store data, which can be accessed and / or modified when instructions are executed.

[0125] The processor may be a Central Processing Unit (CPU), or it may be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. It is the computing core and control core of the terminal, and it is suitable for implementing one or more instructions. Specifically, it is suitable for loading and executing one or more instructions in a readable storage medium to realize the corresponding method flow or corresponding function, so as to realize the steps of the data fusion-based industrial chain breakpoint identification and classification method in the above embodiments.

[0126] Example 4: Based on the same inventive concept, this invention also provides a readable storage medium, specifically an electronic device readable storage medium (Memory). This readable storage medium is a memory device within an electronic device used to store programs and data. It is understood that the readable storage medium here can include both the built-in storage medium within the electronic device and extended storage media supported by the electronic device. The storage medium provides storage space, which stores the terminal's operating system. Furthermore, this storage space also stores one or more instructions suitable for loading and execution by a processor. These instructions can be one or more executable programs (including program code). It should be noted that the storage medium here can be high-speed RAM or non-volatile memory, such as at least one disk storage device. The processor can load and execute one or more instructions stored in the storage medium to implement the steps of the data fusion-based supply chain breakpoint identification and classification method described in the above embodiments.

[0127] Those skilled in the art will understand that embodiments of the present invention can be provided as methods, systems, or computer program products. Therefore, the present invention can take the form of a completely hardware embodiment, a completely software embodiment, or an embodiment combining software and hardware aspects. Furthermore, the present invention can take the form of a computer program product embodied on one or more computer-usable storage media (including, but not limited to, disk storage, CD-ROM, optical storage, etc.) containing computer-usable program code.

[0128] This invention is described with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of the invention. It will be understood that each block of the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, special-purpose computer, embedded processor, or other programmable data processing apparatus to produce a machine, such that the instructions, which execute via the processor of the computer or other programmable data processing apparatus, generate instructions for implementing the flowchart illustrations and / or block diagrams. Figure 1 One or more processes and / or boxes Figure 1 A device that provides the functions specified in one or more boxes.

[0129] These computer program instructions may also be stored in a computer-readable storage medium that can direct a computer or other programmable data processing device to function in a particular manner, such that the instructions stored in the computer-readable storage medium produce an article of manufacture including instruction means, which are implemented in a process Figure 1 One or more processes and / or boxes Figure 1 The function specified in one or more boxes.

[0130] These computer program instructions may also be loaded onto a computer or other programmable data processing equipment to cause a series of operational steps to be performed on the computer or other programmable equipment to produce a computer-implemented process, thereby providing instructions that execute on the computer or other programmable equipment for implementing the process. Figure 1 One or more processes and / or boxes Figure 1 The steps of the function specified in one or more boxes.

[0131] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and not to limit its scope of protection. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that after reading the present invention, they can still make various changes, modifications or equivalent substitutions to the specific implementation of the application, but these changes, modifications or equivalent substitutions are all within the scope of protection of the claims pending approval.

Claims

1. A method for identifying and classifying supply chain breakpoints based on multi-source data fusion, characterized in that, include: Collect multi-source data and perform fusion processing on the multi-source data to construct a multi-source heterogeneous data set; Based on the aforementioned multi-source heterogeneous data set, entities constituting the industrial chain structure and the relationships between entities are extracted to construct a knowledge graph representing the power system industrial chain structure. Based on the power grid topology in the multi-source heterogeneous data set, and combined with the knowledge graph, a unified graph model is formed. Based on each node in the unified graph model, multi-dimensional features of each node are extracted from the multi-source heterogeneous data set; The unified graph model and the multi-dimensional features are input into the spatiotemporal graph neural network model to calculate the breakpoint risk of each node and the propagation trend of the breakpoint along the dependent path, thereby obtaining the breakpoint risk identification result. Based on the text data in the multi-source heterogeneous data set, a large language model is used to perform semantic understanding and reasoning on the text data to obtain breakpoint event information in the text data, and the breakpoint event information is used as the semantic breakpoint identification result. Based on the fusion result of the breakpoint risk identification result and the semantic breakpoint identification result, breakpoint classification is performed, breakpoint identification results are generated, and warnings are issued.

2. The method according to claim 1, characterized in that, Based on the multi-source heterogeneous data set, the entities constituting the industrial chain structure and the relationships between entities are extracted to construct a knowledge graph representing the power system industrial chain structure, including: For the text data in the multi-source heterogeneous data set, a sequence labeling recognition model is used to identify multiple entities that constitute the industrial chain structure. The entities include at least one of equipment entities, enterprise entities, raw material entities, and project entities. A remotely supervised relationship classification model is adopted, combined with a predefined set of relationship types, to determine the relationships between entities. The relationships include at least one of supply relationships, dependency relationships, and transportation relationships. Each identified entity is treated as a node, and each determined relationship is treated as a relationship edge to construct the knowledge graph. The knowledge graph is then stored in a graph database using an attribute graph model.

3. The method according to claim 1, characterized in that, The unified graph model, formed by fusing the power grid topology from the multi-source heterogeneous data set with the knowledge graph, includes: Map the power grid nodes in the power grid topology to the first node set of the unified graph model, wherein the power grid nodes include substation nodes and line nodes; The entities constituting the industrial chain structure in the knowledge graph are mapped to the second node set of the unified graph model, wherein the second node set and the first node set have at least one shared node, and the shared node is a device entity that exists in both the power grid topology and the knowledge graph. The physical connection edges in the power grid topology are mapped to the first edge set of the unified graph model, the association relationship edges in the knowledge graph are mapped to the second edge set of the unified graph model, and the physical connection edges and association relationship edges associated with shared nodes are labeled with edge type. The unified graph model is constructed by taking the union of the first node set and the second node set as the node set of the unified graph model, and taking the union of the first edge set and the second edge set as the edge set of the unified graph model.

4. The method according to claim 3, characterized in that, The step of extracting multi-dimensional features of each node from the multi-source heterogeneous data set based on each node in the unified graph model includes: Based on each node in the unified graph model, static attribute feature vectors of each node are extracted from the multi-source heterogeneous data set; the static attribute feature vectors include at least one or more of the following: equipment capacity, annual enterprise production capacity, geographical location coordinates, and node type code; Time series data associated with each node are extracted from the multi-source heterogeneous data set and normalized to form a dynamic time series feature vector; the dynamic time series feature vector includes one or more of the following: equipment operating parameter sequence, inventory change sequence, raw material price sequence, and logistics status sequence. The static attribute feature vector and the dynamic temporal feature vector are concatenated to form the multi-dimensional feature vector of each node.

5. The method according to claim 4, characterized in that, The step of inputting the unified graph model and the multi-dimensional features into the spatiotemporal graph neural network model to calculate the breakpoint risk of nodes and the propagation trend of breakpoints along dependent paths, and obtaining the breakpoint risk identification result, includes: The unified graph model and the multi-dimensional features are input into the spatiotemporal graph convolutional network. Through the forward propagation of the spatiotemporal graph convolutional network, the spatiotemporal graph convolutional network is used to extract temporal features in the time dimension using a one-dimensional convolutional layer, and in the spatial dimension, a graph convolutional layer is used to aggregate the feature information of adjacent nodes according to the edge set of the unified graph model. Based on the temporal features and the feature information, output the breakpoint risk probability value of each node at the current time and the propagation probability value of each edge at the current time; Using nodes whose breakpoint risk probability value exceeds a preset threshold as source nodes and propagation probability value as edge weights, a depth-first search algorithm is used to generate multiple breakpoint propagation paths. Each breakpoint propagation path includes propagation direction and propagation time estimation. The breakpoint risk probability value and the breakpoint propagation path are determined as the breakpoint risk identification result.

6. The method according to claim 1, characterized in that, The text data based on the multi-source heterogeneous dataset is used to perform semantic understanding and reasoning on the text data using a large language model to obtain breakpoint event information in the text data. This breakpoint event information is then used as the semantic breakpoint identification result, including: New text data is selected from the multi-source heterogeneous data set, and the new text data includes at least one of news information, corporate announcements, policy documents and industry reports; Obtain a preset number of high-risk nodes with the highest risk probability values ​​from the breakpoint risk identification results, and extract subgraph information associated with each of the high-risk nodes from the knowledge graph; Based on the subgraph information of each high-risk node, a prompt word is constructed. The prompt word includes one or more of the following: role setting field, text content field to be analyzed field, context reference field, and output format constraint field. Based on the prompt words, natural language output is obtained using a large language model; By combining regular expression matching with a rule-based parser, one or more of the following are extracted from the natural language output: event type, affected entity identifier, event occurrence time, event duration, and confidence score, to form breakpoint event information. This breakpoint event information is then used as the semantic breakpoint identification result.

7. The method according to claim 1, characterized in that, The step of classifying breakpoints based on the breakpoint risk identification result and the semantic breakpoint identification result, generating breakpoint identification results, and issuing warnings includes: The breakpoint risk probability value of each node in the breakpoint risk identification result is used as the first evidence, and the confidence score of each event record in the semantic breakpoint identification result is used as the second evidence. The DS evidence theory is used to calculate the integrated breakpoint confidence after fusion. When the overall breakpoint confidence exceeds the first threshold, the industrial chain link corresponding to the node or event is identified as a breakpoint. Based on the identified breakpoints, and according to the node type or event type that triggered the breakpoint, they are mapped to a preset breakpoint classification system to determine the category and nature of the breakpoint. Based on the comprehensive breakpoint confidence level and the estimated propagation time in the breakpoint propagation path, the urgency level of each breakpoint is determined. A breakpoint identification report is generated, which includes one or more of the following: breakpoint identifier, link category, nature category, scope of impact, expected impact time window, and suggested countermeasures. The breakpoint identification report and the breakpoint propagation path diagram are displayed through a visual interface for graded early warning.

8. A supply chain breakpoint identification and classification system based on data fusion, characterized in that, The system includes: A multi-source data set construction module is used to collect multi-source data and perform fusion processing on the multi-source data to construct a multi-source heterogeneous data set; The knowledge graph construction module is used to extract entities that constitute the industrial chain structure and the relationships between entities based on the multi-source heterogeneous data set, and to construct a knowledge graph that represents the industrial chain structure of the power system. The model feature determination module is used to fuse the power grid topology in the multi-source heterogeneous data set with the knowledge graph to form a unified graph model; and to extract multi-dimensional features of each node from the multi-source heterogeneous data set based on each node in the unified graph model. The breakpoint risk identification module is used to input the unified graph model and the multi-dimensional features into the spatiotemporal graph neural network model, calculate the breakpoint risk of each node and the propagation trend of the breakpoint along the dependent path, and obtain the breakpoint risk identification result. The semantic breakpoint identification module is used to perform semantic understanding and reasoning on the text data in the multi-source heterogeneous data set using a large language model, to obtain breakpoint event information in the text data, and to use the breakpoint event information as the semantic breakpoint identification result. The breakpoint classification module is used to classify breakpoints based on the fusion result of the breakpoint risk identification result and the semantic breakpoint identification result, generate breakpoint identification results, and issue warnings.

9. An electronic device, characterized in that, include: At least one processor and memory; The memory and processor are connected via a bus; The memory is used to store one or more programs; When the one or more programs are executed by the at least one processor, the method for identifying and classifying supply chain breakpoints based on multi-source data fusion as described in any one of claims 1 to 7 is implemented.

10. A readable storage medium, characterized in that, It contains an execution program, which, when executed, implements the supply chain breakpoint identification and classification method based on multi-source data fusion as described in any one of claims 1 to 7.