A government affair data management intelligent operation and maintenance method based on big data
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- SHAANXI HAIZHI ZHITONG ELECTRONIC TECHNOLOGY CO LTD
- Filing Date
- 2026-05-15
- Publication Date
- 2026-08-07
AI Technical Summary
然而,现有技术对于政务数据治理过程中的风险传播分析、运行状态映射与智能运维调度缺少统一协同机制,难以根据实时运行状态动态调整治理策略与运维指令,容易出现运维响应滞后与治理效率下降的问题
首先,本发明通过GraphSAGE网络对政务关系图执行关系特征传播处理,并结合数字孪生映射技术构建数字孪生治理体,实现了政务数据治理过程中运行状态与数据关联结构的动态映射能力,从而提高了政务数据治理过程中的状态感知能力与风险识别能力。
Smart Images

Figure CN122529464A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent operation and maintenance technology, and in particular to an intelligent operation and maintenance method for government data governance based on big data. Background Technology
[0002] With the development of e-government, the scale of data and business relationships within e-government data governance platforms are constantly increasing, and a large amount of e-government data forms complex data relationship structures during sharing, exchange, and collaboration. When the platform's operational status changes, problems such as abnormal data propagation, resource imbalance, and abnormal service status can easily occur, thereby affecting the efficiency of e-government data governance and the stability of platform operation.
[0003] Most existing government data governance solutions focus on data cleaning, data quality testing, and anomaly handling, with some solutions incorporating knowledge graphs and artificial intelligence to analyze data relationships. However, existing technologies lack a unified and collaborative mechanism for risk propagation analysis, operational status mapping, and intelligent operation and maintenance scheduling in the government data governance process. This makes it difficult to dynamically adjust governance strategies and operation and maintenance instructions based on real-time operational status, easily leading to problems such as delayed operation and maintenance response and decreased governance efficiency.
[0004] Meanwhile, existing technologies lack a digital twin mapping mechanism for the governance process of government data, making it impossible to synchronously extrapolate and analyze the operational status, data association structure, and governance strategies during the governance process. This results in discrepancies between the governance strategies and the actual operational status, making it difficult to meet the intelligent operation and maintenance needs in complex government data governance scenarios.
[0005] Therefore, how to provide a smart operation and maintenance method for government data governance based on big data is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention
[0006] One objective of this invention is to propose an intelligent operation and maintenance method for government data governance based on big data. This invention fully utilizes GraphSAGE network, personalized PageRank algorithm, multi-dimensional assignment algorithm and digital twin mapping technology to dynamically process risk propagation analysis, governance strategy deduction and intelligent operation and maintenance scheduling in the process of government data governance. It has the advantages of high accuracy in risk identification, high efficiency in operation and maintenance scheduling, strong real-time state mapping and high governance stability.
[0007] According to an embodiment of the present invention, a smart operation and maintenance method for government data governance based on big data includes the following steps: S1. Collect operational data and government data from the government data governance platform, and preprocess them to generate operational status sequences and government data sequences; S2. Perform entity recognition and relationship extraction operations on the government data sequence to construct a government relationship graph. Then, perform neighborhood aggregation and relationship feature propagation operations on the government relationship graph through the GraphSAGE network. Combine the running state sequence to perform digital twin mapping processing on the feature propagation results to generate a digital twin governance entity. S3. Through the digital twin governance system, perform data quality inspection, correlation anomaly identification and risk propagation analysis on the government data sequence, extract the risk propagation path, and use the personalized PageRank algorithm to calculate the propagation impact value of each node in the risk propagation path to construct a governance risk sequence; S4. Based on the governance risk sequence, the governance strategy simulation and deduction operation is performed in the digital twin governance body according to the preset governance strategy library to construct the strategy risk matrix. The multi-dimensional assignment algorithm is used to perform joint allocation calculation on the strategy risk matrix to select the target strategy sequence. S5. Perform policy adaptation analysis in the digital twin governance body according to the target policy sequence, generate operation and maintenance scheduling sequence, and map the operation and maintenance scheduling sequence into operation and maintenance instruction sequence according to a unified instruction template. S6. Perform operation and maintenance control operations on the government data governance platform according to the operation and maintenance instruction sequence, and update the state mapping of the digital twin governance entity according to the operation and maintenance results.
[0008] Optionally, the operational data refers to the resource usage information, task execution information, and service status information generated by the government data governance platform during operation; the government data refers to the data set originating from government business activities; the preprocessing includes format unification, anomaly removal, field standardization, and time alignment processing; the governance strategy library refers to the data set that pre-stores multiple governance strategies; and the unified instruction template refers to the preset standardized operation and maintenance instruction generation rules.
[0009] Optionally, S2 specifically includes: S21. Perform field parsing, semantic segmentation and entity tagging on the government data sequence, extract business entities and behavioral entities, and generate entity node sequence; S22. Extract the data call relationship, business association relationship and access association relationship between each entity node based on the entity node sequence, construct the entity association edge set, and construct the government affairs relationship graph based on the entity node sequence and the entity association edge set; S23. Input the government relationship graph into the GraphSAGE network, perform neighborhood feature aggregation operation on the neighboring nodes corresponding to each entity node, and perform relationship feature propagation processing based on the aggregation results to generate a node feature sequence. S24. Extract the running state features of each time step based on the running state sequence, and perform association and fusion processing between the running state features and the corresponding node features in the node feature sequence to generate a state association sequence. S25. Perform virtual-real state association mapping operation on the state association sequence, construct digital twin mapping nodes corresponding to each entity node, and construct digital twin association structure based on the entity association edge set to form a digital twin governance body corresponding to the government data governance platform.
[0010] Optionally, S23 specifically includes: S231. Map each entity node in the government affairs relationship diagram to a corresponding node vector according to the preset vector mapping rules, and construct the set of adjacent nodes of each entity node based on the set of entity association edges. S232. Perform neighborhood feature aggregation operation on the node vectors of each adjacent node, and arrange the aggregation results in layers according to the connection level of entity nodes in the entity association edge set to generate an aggregated feature sequence. S233. Perform multi-level relation feature propagation operation on the aggregated feature sequence, pass the aggregated features corresponding to each entity node along the corresponding entity association edge in a hierarchical manner, and perform iterative update operation on the corresponding aggregated features after the propagation of each layer of relation features is completed, to generate the propagation feature sequence. S234. Extract the running state features corresponding to each time step based on the running state sequence, and adjust the propagation order of each entity node in the propagation feature sequence according to the running state features to generate a dynamic propagation sequence. S235. Perform node association strength calculation on each entity node in the dynamic propagation sequence, and perform feature fusion processing on the dynamic propagation features corresponding to each entity node based on the node association strength to generate the node feature sequence corresponding to each entity node.
[0011] Optionally, S25 specifically includes: S251. Perform state mapping identification processing on each entity node in the state association sequence, and construct the corresponding digital twin mapping node based on the node characteristics of each entity node. S252. Determine the node connection relationships between each digital twin mapping node based on the entity association edge set, and construct the digital twin association structure based on the node connection relationships; S253. Extract the running state features corresponding to each time step from the running state sequence, map the running state features to the corresponding digital twin mapping nodes, perform state change analysis on the node features in the digital twin mapping nodes, and generate a state change sequence. S254. Adjust the connection order of the node connection relationship according to the state change sequence, and perform node association update operation on the digital twin association structure according to the adjusted node connection relationship to generate a dynamic association structure sequence. S255. Based on the dynamic association structure sequence, the virtual and real state correspondence of each digital twin mapping node is verified, and the node features in each digital twin mapping node are updated according to the verification results to form a digital twin governance body corresponding to the government data governance platform.
[0012] Optionally, S3 specifically includes: S31. Based on the digital twin governance system, perform data integrity detection, data consistency detection, and data association validity detection on each data node in the government data sequence, screen out abnormal data nodes, and generate an abnormal node sequence; S32. Extract the corresponding node connection relationship in the digital twin governance body based on the abnormal node sequence, and perform node association traversal operation along the digital twin association structure to identify the associated abnormal nodes corresponding to the abnormal nodes and generate the associated abnormal sequence. S33. Extract the node propagation relationship between each associated abnormal node in the associated abnormal sequence, and construct the corresponding risk propagation path based on the node propagation relationship; S34. Using a personalized PageRank algorithm, the propagation impact value is calculated for each node in the risk propagation path. The propagation impact value is iteratively updated based on the node propagation relationship and the digital twin association structure to generate a node impact sequence. S35. Based on the node impact sequence, perform risk aggregation analysis on each risk propagation path, extract the risk diffusion range and risk propagation intensity corresponding to each risk propagation path, and generate a risk assessment sequence. S36. Based on the risk assessment sequence, classify the risk levels of each associated abnormal node, and prioritize each risk propagation path according to the risk level to construct a governance risk sequence.
[0013] Optionally, S32 specifically includes: S321. Based on the abnormal node sequence, extract the digital twin mapping nodes corresponding to each abnormal node in the digital twin governance body, and determine the node connection relationship between each digital twin mapping node based on the digital twin association structure to generate an abnormal association node set. S322. Perform node association traversal operation on each digital twin mapping node in the abnormal association node set, extract the associated node path corresponding to each digital twin mapping node along the node connection relationship, and construct the node traversal sequence. S323. Extract the features of each running state in the running state sequence, and perform state association analysis on each associated node in the node traversal sequence based on the running state features to generate a state association node sequence. S324. Perform node propagation association analysis on each associated node in the state association node sequence, and extract the abnormal propagation association results between each associated node based on the node connection relationship to generate an abnormal propagation sequence. S325. Perform an associated anomaly determination operation on each associated node according to the anomaly propagation sequence, determine the associated nodes that meet the preset anomaly propagation conditions as associated anomaly nodes, and generate an associated anomaly sequence.
[0014] Optionally, S34 specifically includes: S341. Extract the node connection relationships and associated abnormal nodes in each risk propagation path, and construct the corresponding node propagation structure based on the node connection relationships; S342. Based on the node propagation structure, count the number of node accesses, the number of abnormal associations, and the number of node association paths corresponding to each associated abnormal node, and generate the propagation weight sequence corresponding to each associated abnormal node based on the statistical results. S343. Based on the propagation weight sequence, a personalized PageRank algorithm is used to calculate the propagation impact value of each associated abnormal node in the node propagation structure, and the propagation impact value between each associated abnormal node is transferred according to the node connection relationship to generate an intermediate impact sequence. S344. Extract the operational status features corresponding to each digital twin mapping node in the digital twin governance body, and perform dynamic adjustment operations on the propagation impact values in the intermediate impact sequence based on the operational status features to generate a dynamic impact sequence. S345. Perform iterative update operations on each propagation influence value in the dynamic influence sequence until the change results of each propagation influence value in two consecutive iterations reach the preset convergence threshold to obtain a stable influence sequence. S346. Based on the propagation impact values in the stable impact sequence, determine the node propagation priority in each risk propagation path, and generate the node impact sequence corresponding to each risk propagation path based on the node propagation priority.
[0015] Optionally, S4 specifically includes: S41. Based on the governance risk sequence, extract the risk level, risk propagation intensity and node propagation priority corresponding to each risk propagation path, construct the corresponding risk features, and perform strategy matching operation in the preset governance strategy library according to each risk feature to generate candidate strategy sequence. S42. Map the candidate strategy sequence to the digital twin governance body, perform governance process simulation and deduction operations on the digital twin mapping nodes corresponding to each candidate strategy, and extract the state change results corresponding to each candidate strategy based on the digital twin association structure to generate a twin simulation sequence. S43. Based on the twin simulation sequence, statistically analyze the risk diffusion range, node state change results, and resource consumption results of each candidate strategy, and construct a strategy risk matrix based on the statistical results; S44. Perform risk association mapping operation on each strategy risk data in the strategy risk matrix, and sort the strategy risk data according to the node propagation priority corresponding to each risk propagation path to generate a risk sorting sequence. S45. A multidimensional assignment algorithm is used to perform joint allocation calculations on the risk data of each strategy in the risk ranking sequence, and a strategy allocation relationship is established based on the risk propagation path in the governance risk sequence and the digital twin mapping node in the digital twin governance body to generate a strategy allocation sequence. S46. Based on the strategy allocation sequence, perform correlation matching analysis on the resource occupancy results, node status change results and risk diffusion range corresponding to each candidate strategy, screen candidate strategies that meet the preset governance conditions, and obtain the target strategy sequence.
[0016] Optionally, S5 specifically includes: S51. Extract the node propagation priority and resource consumption results corresponding to each target strategy based on the target strategy sequence, and map them to the corresponding digital twin mapping nodes in the digital twin governance body to generate a strategy mapping sequence. S52. Based on the strategy mapping sequence, perform strategy adaptation analysis on each digital twin mapping node in the digital twin governance body, and extract the resource occupancy information, task execution information and service status information corresponding to each digital twin mapping node by combining the running status sequence, and generate a status adaptation sequence. S53. Perform association analysis on each digital twin mapping node based on the state adaptation sequence, and generate a scheduling association sequence in combination with the digital twin association structure; S54. Arrange the resource usage results corresponding to each target strategy in order according to the scheduling association sequence, and adjust the priority of the association sorting results in combination with the node propagation priority to generate the operation and maintenance scheduling sequence. S55. Perform a unified instruction template mapping operation on the operation and maintenance scheduling sequence, and generate the corresponding operation and maintenance instruction sequence based on the node connection relationship, resource occupation results and status adaptation sequence.
[0017] The beneficial effects of this invention are: First, this invention performs relation feature propagation processing on the government affairs relationship graph through the GraphSAGE network, and constructs a digital twin governance body by combining digital twin mapping technology. This realizes the dynamic mapping capability between the operating status and data association structure in the process of government affairs data governance, thereby improving the status perception capability and risk identification capability in the process of government affairs data governance.
[0018] Secondly, this invention dynamically calculates the propagation impact value in the risk propagation path using a personalized PageRank algorithm, and combines governance strategy simulation and multi-dimensional assignment algorithm to perform joint allocation processing, thereby realizing the collaborative processing capability between risk propagation analysis and governance strategy scheduling, thus improving the accuracy of governance strategy matching and the efficiency of operation and maintenance scheduling.
[0019] Finally, based on the digital twin governance system, this invention performs dynamic analysis and state mapping update processing on the operation and maintenance scheduling sequence, thereby realizing intelligent operation and maintenance control capabilities in the process of government data governance, thus improving the operational stability and governance reliability of the government data governance platform. Attached Figure Description
[0020] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Fig. 1 This is a flowchart of an intelligent operation and maintenance method for government data governance based on big data, proposed in this invention. Fig. 2 This is a flowchart illustrating the government data association process of a big data-based intelligent operation and maintenance method for government data governance proposed in this invention. Fig. 3 This is a flowchart illustrating the intelligent operation and maintenance scheduling control process of an intelligent operation and maintenance method for government data governance based on big data proposed in this invention. Detailed Implementation
[0021] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.
[0022] refer to Figs. 1-3 A smart operation and maintenance method for government data governance based on big data includes the following steps: S1. Collect operational data and government data from the government data governance platform, and preprocess them to generate operational status sequences and government data sequences; S2. Perform entity recognition and relationship extraction operations on the government data sequence to construct a government relationship graph. Then, perform neighborhood aggregation and relationship feature propagation operations on the government relationship graph through the GraphSAGE network. Combine the running state sequence to perform digital twin mapping processing on the feature propagation results to generate a digital twin governance entity. S3. Through the digital twin governance system, perform data quality inspection, correlation anomaly identification and risk propagation analysis on the government data sequence, extract the risk propagation path, and use the personalized PageRank algorithm to calculate the propagation impact value of each node in the risk propagation path to construct a governance risk sequence; S4. Based on the governance risk sequence, the governance strategy simulation and deduction operation is performed in the digital twin governance body according to the preset governance strategy library to construct the strategy risk matrix. The multi-dimensional assignment algorithm is used to perform joint allocation calculation on the strategy risk matrix to select the target strategy sequence. S5. Perform policy adaptation analysis in the digital twin governance body according to the target policy sequence, generate operation and maintenance scheduling sequence, and map the operation and maintenance scheduling sequence into operation and maintenance instruction sequence according to a unified instruction template. S6. Perform operation and maintenance control operations on the government data governance platform according to the operation and maintenance instruction sequence, and update the state mapping of the digital twin governance entity according to the operation and maintenance results.
[0023] In this embodiment, the operational data refers to the resource usage information, task execution information, and service status information generated by the government data governance platform during operation. The government data refers to the data set originating from government business activities. Preprocessing includes format unification, anomaly removal, field standardization, and time alignment. The governance strategy library refers to the data set that pre-stores various governance strategies. The unified instruction template refers to the preset standardized operation and maintenance instruction generation rules.
[0024] In this embodiment, S2 specifically includes: S21. Perform field parsing processing on the user identifier field, business approval field, data exchange field and access record field in the government data sequence, and perform semantic segmentation operation on the business description content in the government data sequence using word segmentation rules; Entity tagging is performed on the approval object, data call object, access subject and business processing action in the semantic segmentation result. The approval department, data resource node and business behavior node are respectively identified as business entity and behavior entity, and finally the corresponding entity node sequence is formed. S22. Perform correlation analysis on the data exchange records between business entities in the entity node sequence, extract the data call relationship between different business entities, and at the same time perform correlation extraction processing on the upper and lower level approval paths in the business approval process to form business relationship; Perform access association analysis on access records and resource call records in behavioral entities to generate access association relationships; The call relationships, business relationships, and access relationships between different entity nodes are respectively constructed as corresponding entity association edges, and a government affairs relationship graph is constructed based on the entity node sequence and the entity association edge set; S23. Input the government relationship graph into the GraphSAGE network, perform neighborhood feature aggregation operation on the neighboring nodes corresponding to each entity node, and perform relationship feature propagation processing based on the aggregation results to generate a node feature sequence. S24. Extract the running state features of each time step based on the running state sequence, and perform association and fusion processing between the running state features and the corresponding node features in the node feature sequence to generate a state association sequence. S25. Perform virtual-real state association mapping operation on the state association sequence, construct digital twin mapping nodes corresponding to each entity node, and construct digital twin association structure based on the entity association edge set to form a digital twin governance body corresponding to the government data governance platform.
[0025] In this embodiment, S23 specifically includes: S231. According to the 128-dimensional vector mapping rule, the business entity nodes and behavior entity nodes in the government affairs relationship graph are vectorized and encoded, and the corresponding set of adjacent nodes is constructed based on the data call relationship, business relationship and access relationship in the entity association edge set. For entity nodes with more than 15 connections, the corresponding entity nodes are identified as high-frequency associated nodes, and the set of adjacent nodes corresponding to the high-frequency associated nodes is extracted first. S232. Perform neighborhood feature aggregation processing on the node vectors corresponding to each adjacent node, and arrange the aggregation results in layers according to the node connection level in the entity association edge set; the aggregation features corresponding to directly connected entity nodes are determined as the first layer aggregation features, and the aggregation features corresponding to indirectly connected entity nodes are determined as the second layer aggregation features. At the same time, the aggregation features corresponding to entity association edges with a data access frequency of more than 300 times are assigned an aggregation weight of 0.8, and the aggregation features corresponding to entity association edges with an access frequency of less than 50 times are assigned an aggregation weight of 0.3, and finally form an aggregation feature sequence. S233. Perform 4-layer relation feature propagation processing on the aggregated feature sequence, so that the first-layer aggregated feature is propagated to the adjacent entity nodes along the corresponding entity association edge, and the second-layer aggregated feature is passed hierarchically to the indirect entity nodes in the association path. After the propagation of relational features at each level is completed, iterative update processing is performed on the aggregated features of the corresponding entity nodes, and stable state marking processing is performed on the aggregated features with feature differences less than 0.002, finally generating a propagation feature sequence; S234. Extract resource occupancy information, task execution information and service status information corresponding to each time step from the running status sequence, and perform propagation order adjustment processing on the entity nodes in the propagation feature sequence according to the resource occupancy ratio corresponding to each entity node. For entity nodes with a resource utilization rate greater than 75%, the corresponding entity node will be adjusted to a priority propagation position. For entity nodes with a task blocking count greater than 10 times, the propagation priority of the corresponding entity node will be increased, and a dynamic propagation sequence will be generated. S235. Count the number of entity association edges, node access times and data call times corresponding to each entity node, and calculate the node association strength between each entity node based on the statistical results. For entity nodes with a node association strength greater than 0.85, high-intensity feature fusion processing is performed on the corresponding dynamic propagation features. For entity nodes with a node association strength less than 0.4, low-weight fusion processing is performed on the corresponding dynamic propagation features, ultimately generating node feature sequences for each entity node.
[0026] In this embodiment, S25 specifically includes: S251. Perform state mapping identification processing on the business entity nodes and behavior entity nodes in the state association sequence, and construct the corresponding digital twin mapping node based on the 128-dimensional node features of each entity node. For entity nodes that have been accessed more than 500 times, increase the state refresh frequency of the corresponding digital twin mapping node so that the corresponding digital twin mapping node performs state synchronization processing once every 2 seconds. S252. Determine the node connection relationship between each digital twin mapping node based on the data call relationship, business relationship and access relationship in the entity association edge set, and perform hierarchical division processing on each node connection relationship according to the node connection frequency; For node connections with more than 300 connections, they are identified as high-strength connections, and corresponding digital twin connection structures are built first, ultimately forming a digital twin connection structure containing 230 node connections. S253. Extract the resource usage information, task execution information and service status information corresponding to each time step from the running status sequence, and map them to the corresponding digital twin mapping nodes in the order of time steps; For digital twin mapping nodes with a resource occupancy rate greater than 80%, high-frequency state change analysis and processing are performed. For digital twin mapping nodes whose task execution delay exceeds a preset threshold for three consecutive times, the state change weight of the corresponding node features is increased, and a state change sequence is generated. S254. Based on the magnitude of state changes in the state change sequence, perform connection order adjustment processing on the connection relationship of each node; For digital twin mapping nodes with a state change magnitude greater than 0.65, the corresponding node connection relationship is adjusted to the priority connection position. For digital twin mapping nodes with more than 5 service state anomalies, the association update frequency of the corresponding node connection relationship is increased, and node association update processing is performed on the digital twin association structure according to the adjusted node connection relationship to generate a dynamic association structure sequence. S255. Perform virtual-real state correspondence verification processing on each digital twin mapping node in the dynamic association structure sequence, and count the state deviation value between the state characteristics in the digital twin mapping node and the actual operating state of the government data governance platform. For digital twin mapping nodes with a state deviation value greater than 0.03, state correction processing is performed on the corresponding node features. For digital twin mapping nodes with a state deviation value less than 0.01, the current node feature state remains unchanged, ultimately forming a digital twin governance entity that is synchronized with the operating state of the government data governance platform.
[0027] In this embodiment, S3 specifically includes: S31. Perform field integrity verification on each data node in the government data sequence. When the field missing rate is greater than 12%, the corresponding data node will be identified as a data integrity abnormal node. Perform consistency comparison processing on business fields in different data nodes. When the data difference value corresponding to the same business identifier is greater than 0.08, the corresponding data node is identified as a data consistency abnormal node. Perform association validity checks on the data call relationships and business association relationships between each data node. When the association failure rate is greater than 15%, the corresponding data node is identified as an abnormal data association node, and finally a sequence of abnormal nodes is formed. S32. Extract the corresponding node connection relationship in the digital twin governance body based on the abnormal node sequence, and perform node association traversal operation along the digital twin association structure to identify the associated abnormal nodes corresponding to the abnormal nodes and generate the associated abnormal sequence. S33. Perform node propagation relationship extraction processing on each associated abnormal node in the associated abnormal sequence, and construct the corresponding node propagation relationship based on the data call order, business association direction and access connection order; For node propagation relationships with more than 3 propagation levels, the corresponding propagation path is identified as a high-risk propagation path, and the corresponding risk propagation path is constructed according to the node connection order. S34. Using a personalized PageRank algorithm, the propagation impact value is calculated for each node in the risk propagation path. The propagation impact value is iteratively updated based on the node propagation relationship and the digital twin association structure to generate a node impact sequence. S35. Perform risk aggregation analysis on each propagation impact value in the node impact sequence, and statistically analyze the number of associated abnormal nodes, the number of propagation levels, and the distribution of node impact values in each risk propagation path. When the number of associated abnormal nodes in a risk propagation path is greater than 10, the corresponding risk propagation path will be identified as a large-scale diffusion path. When the average node impact value is greater than 0.75, the corresponding risk propagation path is identified as a high-intensity propagation path, and a risk assessment sequence is finally generated. S36. Based on the risk diffusion range and risk propagation intensity in the risk assessment sequence, perform risk level classification processing on each associated abnormal node; For associated abnormal nodes with a risk propagation intensity greater than 0.8, the corresponding associated abnormal nodes will be identified as high-risk nodes; For risk propagation paths with a risk spread to more than 15 related nodes, the ranking priority of the corresponding risk propagation paths is increased, and a governance risk sequence is constructed based on the risk level and ranking results.
[0028] In this embodiment, S32 specifically includes: S321. Extract the corresponding digital twin mapping node in the digital twin governance body based on the abnormal node identifier in the abnormal node sequence, and determine the node connection relationship between each digital twin mapping node based on the data call relationship, business relationship and access relationship in the digital twin association structure. For digital twin mapping nodes with more than 18 node connections, the corresponding digital twin mapping nodes are identified as key associated nodes and are given priority to be included in the abnormal associated node set, ultimately forming an abnormal associated node set containing multiple node connection relationships. S322. Perform node association traversal processing on each digital twin mapping node in the abnormal association node set according to the node connection order, and extract the corresponding association node path along the node connection relationship. For association node paths with a node propagation level of 4 or more, increase the traversal priority of the corresponding association node path. At the same time, perform high-frequency traversal processing on association node paths with a node access frequency of more than 600 times, and finally construct a node traversal sequence. S323. Extract resource usage information, task execution information and service status information corresponding to each time step from the running status sequence, and perform correlation analysis and processing on the execution status of each associated node in the node traversal sequence based on the characteristics of each running status. For associated nodes with a resource utilization rate greater than 78%, increase the state association weight of the corresponding associated nodes; For associated nodes whose task execution delay exceeds a preset threshold three times consecutively, the abnormal state association level of the corresponding associated node is increased, and a state association node sequence is finally generated. S324. Perform node propagation association analysis on each associated node in the state-associated node sequence, and extract the data propagation direction, access propagation order and business propagation level between each associated node based on the node connection relationship. For associated nodes with a propagation level greater than 3 levels and a propagation frequency greater than 200 times, the corresponding node propagation relationship is determined as an abnormal propagation association result, and an abnormal propagation sequence is generated based on the abnormal propagation association result. S325. Perform association anomaly judgment processing on each associated node based on the propagation frequency, propagation level and node influence value in the abnormal propagation sequence. Nodes with a propagation frequency greater than 300 times and a node influence value greater than 0.8 are identified as high-risk anomalous nodes. Increase the priority of anomaly detection for associated nodes with a propagation level of 5 or more, and finally generate an associated anomaly sequence.
[0029] In this embodiment, S34 specifically includes: S341. Extract the node connection relationship of the digital twin mapping nodes in each risk propagation path, and construct the corresponding node propagation structure according to the data call direction, business propagation direction and access propagation order. Risk propagation paths with more than 25 node connections are identified as high-density propagation paths. At the same time, associated abnormal nodes with a propagation level of 4 or more are marked as key propagation structures. S342. Perform statistical analysis on each associated abnormal node in the node propagation structure, and count the number of node accesses, the number of abnormal associations, and the number of node association paths for the corresponding associated abnormal node. Increase the access weight of associated abnormal nodes with more than 1000 visits, increase the abnormal propagation weight of associated abnormal nodes with more than 15 abnormal associations, increase the path propagation weight of associated abnormal nodes with more than 12 associated paths, and generate a propagation weight sequence. S343. Based on the propagation weight sequence, perform personalized PageRank propagation impact value calculation for each associated abnormal node in the node propagation structure, and use node access weight, abnormal propagation weight and path propagation weight as propagation impact value calculation parameters. For associated abnormal nodes with a propagation level of more than 3 levels, the association transmission coefficient of the corresponding propagation impact value is increased. At the same time, hierarchical association transmission processing is performed on the propagation impact value between each associated abnormal node according to the node connection relationship to generate an intermediate impact sequence. S344. Extract resource occupancy information, task execution information and service status information corresponding to each digital twin mapping node from the digital twin governance body, and perform dynamic adjustment processing on the propagation impact value in the intermediate impact sequence based on the characteristics of each operating status. For digital twin mapping nodes with a resource occupancy rate greater than 82%, increase the dynamic adjustment weight of the corresponding propagation impact value. For digital twin mapping nodes with a task execution blocking time exceeding the preset threshold for four consecutive times, increase the abnormal propagation level of the corresponding propagation impact value and generate a dynamic impact sequence. S345. Perform iterative update processing on each propagation impact value in the dynamic impact sequence, and synchronously correct each propagation impact value according to the node connection relationship; When the difference in the propagation influence value between two consecutive iterations is less than 0.001, it is determined that the current propagation influence value has reached the preset convergence threshold, and a stable influence sequence is obtained. S346. Based on the propagation impact values in the stable impact sequence, determine the node propagation priority in each risk propagation path, and generate the node impact sequence corresponding to each risk propagation path based on the node propagation priority.
[0030] In this embodiment, S4 specifically includes: S41. Based on the governance risk sequence, extract the risk level, risk propagation intensity and node propagation priority corresponding to each risk propagation path, construct the corresponding risk features, and perform strategy matching operation in the preset governance strategy library according to each risk feature to generate candidate strategy sequence. S42. Map the candidate strategy sequence to the digital twin governance body, perform governance process simulation and deduction operations on the digital twin mapping nodes corresponding to each candidate strategy, and extract the state change results corresponding to each candidate strategy based on the digital twin association structure to generate a twin simulation sequence. S43. Based on the twin simulation sequence, statistically analyze the risk diffusion range, node state change results, and resource consumption results of each candidate strategy, and construct a strategy risk matrix based on the statistical results; S44. Perform risk association mapping operation on each strategy risk data in the strategy risk matrix, and sort the strategy risk data according to the node propagation priority corresponding to each risk propagation path to generate a risk sorting sequence. S45. A multidimensional assignment algorithm is used to perform joint allocation calculations on the risk data of each strategy in the risk ranking sequence, and a strategy allocation relationship is established based on the risk propagation path in the governance risk sequence and the digital twin mapping node in the digital twin governance body to generate a strategy allocation sequence. S46. Based on the strategy allocation sequence, perform correlation matching analysis on the resource occupancy results, node status change results and risk diffusion range corresponding to each candidate strategy, screen candidate strategies that meet the preset governance conditions, and obtain the target strategy sequence.
[0031] In this embodiment, S5 specifically includes: S51. Extract the node propagation priority and resource consumption results corresponding to each target strategy based on the target strategy sequence, and map them to the corresponding digital twin mapping nodes in the digital twin governance body to generate a strategy mapping sequence. S52. Based on the strategy mapping sequence, perform strategy adaptation analysis on each digital twin mapping node in the digital twin governance body, and extract the resource occupancy information, task execution information and service status information corresponding to each digital twin mapping node by combining the running status sequence, and generate a status adaptation sequence. S53. Perform association analysis on each digital twin mapping node based on the state adaptation sequence, and generate a scheduling association sequence in combination with the digital twin association structure; S54. Arrange the resource usage results corresponding to each target strategy in order according to the scheduling association sequence, and adjust the priority of the association sorting results in combination with the node propagation priority to generate the operation and maintenance scheduling sequence. S55. Perform a unified instruction template mapping operation on the operation and maintenance scheduling sequence, and generate the corresponding operation and maintenance instruction sequence based on the node connection relationship, resource occupation results and status adaptation sequence.
[0032] In this embodiment, S6 specifically includes: S61. Perform operation and maintenance control operations on the corresponding operation status data in the government data governance platform according to the operation and maintenance instruction sequence, collect the operation status change results corresponding to each operation and maintenance control operation, and generate a status feedback sequence. S62. Map the state feedback sequence to the corresponding digital twin mapping node in the digital twin governance body, and perform state synchronization update operation on the node characteristics in each digital twin mapping node according to the state feedback sequence to generate a state update sequence. S63. Perform association change analysis on the node connection relationship in the digital twin association structure based on the state update sequence, and perform dynamic association update operation on the digital twin association structure according to the association change results.
[0033] Example 1: To verify the feasibility of this invention in practice, it was applied to the intelligent operation and maintenance scenario of a government data governance platform. This platform simultaneously connects to multiple business data sources, processing approximately 4.2 million pieces of government data daily. Internally, the platform contains 86 business data nodes and 230 data association paths. During long-term operation, due to the complex data call and access relationships between different business data, when some business nodes experience abnormal states, the abnormal data can easily spread along the node connections, leading to service status anomalies, resource consumption fluctuations, and governance task blockages in multiple associated nodes. Traditional operation and maintenance methods primarily use fixed rules to screen abnormal data, failing to dynamically analyze risk propagation paths or dynamically adjust governance strategies based on real-time operational status. Therefore, they are prone to delayed operation and maintenance responses and decreased governance efficiency.
[0034] When applying this invention, the operational data and government data in the government data governance platform are first uniformly collected. The collected results are then processed to unify the format, standardize the fields, and align the time, forming an operational status sequence and a government data sequence. In actual operation, operational status features are collected once for each time step, with a single round of operational status data volume of approximately 12GB. Subsequently, business entities and behavioral entities in the government data sequence are identified and processed, and a government relationship graph is constructed based on data call relationships, business association relationships, and access association relationships. The entity nodes in the government relationship graph are processed through neighborhood aggregation and relationship feature propagation using the GraphSAGE network. The number of iterations for single-round relationship feature propagation is set to 6 rounds, and the node feature dimension is set to 128 dimensions, forming a dynamic association structure between different entity nodes. Simultaneously, the operational status features in the operational status sequence are mapped to the corresponding entity nodes, and a corresponding digital twin governance body is constructed using digital twin mapping technology, enabling real-time synchronization of operational status changes in the platform to the digital twin governance body.
[0035] When abnormal data occurs during platform operation, this invention uses a digital twin governance system to perform a correlational analysis of the node connections corresponding to the abnormal nodes and constructs risk propagation paths based on node propagation relationships. Subsequently, a personalized PageRank algorithm is used to dynamically calculate the propagation impact value in each risk propagation path. The propagation impact value iteration count is set to 20 rounds, and the convergence threshold is set to 0.001. Simultaneously, the propagation impact value is dynamically adjusted based on the operational status characteristics within the digital twin governance system, enabling the system to quickly identify high-risk associated nodes and high-risk propagation paths. After completing the risk propagation analysis, this invention further performs governance strategy simulation and deduction processing within the digital twin governance system based on the governance risk sequence. The results of state changes, risk diffusion range, and resource consumption corresponding to different governance strategies are dynamically analyzed. A multi-dimensional assignment algorithm is then used to perform joint allocation processing on each governance strategy, thereby selecting target strategy sequences with a high degree of suitability.
[0036] Subsequently, this invention generates an operation and maintenance scheduling sequence based on the target strategy sequence, and generates a corresponding operation and maintenance instruction sequence in conjunction with a unified instruction template. After the operation and maintenance instruction sequence is sent to the government data governance platform, the resource occupancy status, task execution status, and service status in the platform can be dynamically adjusted according to the real-time operating status. Simultaneously, this invention continuously updates the execution status mapping of the digital twin governance entity, ensuring that the status changes in the digital twin governance entity are synchronized with the actual operating status of the platform, thereby achieving dynamic and intelligent operation and maintenance control in the government data governance process.
[0037] Table 1 Comparison of the Effects of Intelligent Operation and Maintenance in Government Data Governance
[0038] As shown in Table 1, after adopting the method of the present invention, the accuracy rate of data anomaly identification increased from 82.1% to 96.7%, and the accuracy rate of risk propagation identification increased from 75.4% to 95.3%. This indicates that the present invention, through the combination of GraphSAGE network and digital twin governance, enhances the correlation analysis capability and anomaly propagation identification capability between government data, enabling the platform to more accurately identify abnormal nodes and risk propagation paths.
[0039] Meanwhile, the risk propagation path location time was reduced from 19.2s to 5.6s, the operation and maintenance instruction generation time was reduced from 13.1s to 3.9s, and the operation and maintenance response completion time was reduced from 28.4s to 8.7s. This indicates that the present invention can quickly complete the deduction of governance strategies and operation and maintenance scheduling based on the changes in the operating status in the digital twin governance system, thereby improving the dynamic response efficiency in the process of government data governance.
[0040] Furthermore, the resource usage volatility decreased from 27.8% to 11.6%, the number of service status anomalies decreased from 35 to 8, and the consistency of status mapping increased from 81.3% to 96.8%. This demonstrates that the present invention can achieve dynamic synchronization analysis between the operational status and the governance status through the digital twin mapping mechanism, and adjust the operation and maintenance instructions according to the real-time status, thereby improving the operational stability and governance reliability of the government data governance platform.
[0041] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.
Claims
1. A smart operation and maintenance method for government data governance based on big data, characterized in that, Includes the following steps: S1. Collect operational data and government data from the government data governance platform, and preprocess them to generate operational status sequences and government data sequences; S2. Perform entity recognition and relationship extraction operations on the government data sequence to construct a government relationship graph. Then, perform neighborhood aggregation and relationship feature propagation operations on the government relationship graph through the GraphSAGE network. Combine the running state sequence to perform digital twin mapping processing on the feature propagation results to generate a digital twin governance entity. S3. Through the digital twin governance system, perform data quality inspection, correlation anomaly identification and risk propagation analysis on the government data sequence, extract the risk propagation path, and use the personalized PageRank algorithm to calculate the propagation impact value of each node in the risk propagation path to construct a governance risk sequence; S4. Based on the governance risk sequence, the governance strategy simulation and deduction operation is performed in the digital twin governance body according to the preset governance strategy library to construct the strategy risk matrix. The multi-dimensional assignment algorithm is used to perform joint allocation calculation on the strategy risk matrix to select the target strategy sequence. S5. Perform policy adaptation analysis in the digital twin governance body according to the target policy sequence, generate operation and maintenance scheduling sequence, and map the operation and maintenance scheduling sequence into operation and maintenance instruction sequence according to a unified instruction template. S6. Perform operation and maintenance control operations on the government data governance platform according to the operation and maintenance instruction sequence, and update the state mapping of the digital twin governance entity according to the operation and maintenance results.
2. The intelligent operation and maintenance method for government data governance based on big data according to claim 1, characterized in that, The operational data refers to the resource usage information, task execution information, and service status information generated by the government data governance platform during operation. The government data refers to the data set originating from government business activities. The preprocessing includes format unification, anomaly removal, field standardization, and time alignment. The governance strategy library refers to the data set that pre-stores multiple governance strategies. The unified instruction template refers to the preset standardized operation and maintenance instruction generation rules.
3. The intelligent operation and maintenance method for government data governance based on big data according to claim 1, characterized in that, S2 specifically includes: S21. Perform field parsing, semantic segmentation and entity tagging on the government data sequence, extract business entities and behavioral entities, and generate entity node sequence; S22. Extract the data call relationship, business association relationship and access association relationship between each entity node based on the entity node sequence, construct the entity association edge set, and construct the government affairs relationship graph based on the entity node sequence and the entity association edge set; S23. Input the government relationship graph into the GraphSAGE network, perform neighborhood feature aggregation operation on the neighboring nodes corresponding to each entity node, and perform relationship feature propagation processing based on the aggregation results to generate a node feature sequence. S24. Extract the running state features of each time step based on the running state sequence, and perform association and fusion processing between the running state features and the corresponding node features in the node feature sequence to generate a state association sequence. S25. Perform virtual-real state association mapping operation on the state association sequence, construct digital twin mapping nodes corresponding to each entity node, and construct digital twin association structure based on the entity association edge set to form a digital twin governance body corresponding to the government data governance platform.
4. The intelligent operation and maintenance method for government data governance based on big data according to claim 3, characterized in that, S23 specifically includes: S231. Map each entity node in the government affairs relationship diagram to a corresponding node vector according to the preset vector mapping rules, and construct the set of adjacent nodes of each entity node based on the set of entity association edges. S232. Perform neighborhood feature aggregation operation on the node vectors of each adjacent node, and arrange the aggregation results in layers according to the connection level of entity nodes in the entity association edge set to generate an aggregated feature sequence. S233. Perform multi-level relation feature propagation operation on the aggregated feature sequence, pass the aggregated features corresponding to each entity node along the corresponding entity association edge in a hierarchical manner, and perform iterative update operation on the corresponding aggregated features after the propagation of each layer of relation features is completed, to generate the propagation feature sequence. S234. Extract the running state features corresponding to each time step based on the running state sequence, and adjust the propagation order of each entity node in the propagation feature sequence according to the running state features to generate a dynamic propagation sequence. S235. Perform node association strength calculation on each entity node in the dynamic propagation sequence, and perform feature fusion processing on the dynamic propagation features corresponding to each entity node based on the node association strength to generate the node feature sequence corresponding to each entity node.
5. The intelligent operation and maintenance method for government data governance based on big data according to claim 3, characterized in that, Specifically, S25 includes: S251. Perform state mapping identification processing on each entity node in the state association sequence, and construct the corresponding digital twin mapping node based on the node characteristics of each entity node. S252. Determine the node connection relationships between each digital twin mapping node based on the entity association edge set, and construct the digital twin association structure based on the node connection relationships; S253. Extract the running state features corresponding to each time step from the running state sequence, map the running state features to the corresponding digital twin mapping nodes, perform state change analysis on the node features in the digital twin mapping nodes, and generate a state change sequence. S254. Adjust the connection order of the node connection relationship according to the state change sequence, and perform node association update operation on the digital twin association structure according to the adjusted node connection relationship to generate a dynamic association structure sequence. S255. Based on the dynamic association structure sequence, the virtual and real state correspondence of each digital twin mapping node is verified, and the node features in each digital twin mapping node are updated according to the verification results to form a digital twin governance body corresponding to the government data governance platform.
6. The intelligent operation and maintenance method for government data governance based on big data according to claim 1, characterized in that, S3 specifically includes: S31. Based on the digital twin governance system, perform data integrity detection, data consistency detection, and data association validity detection on each data node in the government data sequence, screen out abnormal data nodes, and generate an abnormal node sequence; S32. Extract the corresponding node connection relationship in the digital twin governance body based on the abnormal node sequence, and perform node association traversal operation along the digital twin association structure to identify the associated abnormal nodes corresponding to the abnormal nodes and generate the associated abnormal sequence. S33. Extract the node propagation relationship between each associated abnormal node in the associated abnormal sequence, and construct the corresponding risk propagation path based on the node propagation relationship; S34. Using a personalized PageRank algorithm, the propagation impact value is calculated for each node in the risk propagation path. The propagation impact value is iteratively updated based on the node propagation relationship and the digital twin association structure to generate a node impact sequence. S35. Based on the node impact sequence, perform risk aggregation analysis on each risk propagation path, extract the risk diffusion range and risk propagation intensity corresponding to each risk propagation path, and generate a risk assessment sequence. S36. Based on the risk assessment sequence, classify the risk levels of each associated abnormal node, and prioritize each risk propagation path according to the risk level to construct a governance risk sequence.
7. The intelligent operation and maintenance method for government data governance based on big data according to claim 6, characterized in that, Specifically, S32 includes: S321. Based on the abnormal node sequence, extract the digital twin mapping nodes corresponding to each abnormal node in the digital twin governance body, and determine the node connection relationship between each digital twin mapping node based on the digital twin association structure to generate an abnormal association node set. S322. Perform node association traversal operation on each digital twin mapping node in the abnormal association node set, extract the associated node path corresponding to each digital twin mapping node along the node connection relationship, and construct the node traversal sequence. S323. Extract the features of each running state in the running state sequence, and perform state association analysis on each associated node in the node traversal sequence based on the running state features to generate a state association node sequence. S324. Perform node propagation association analysis on each associated node in the state association node sequence, and extract the abnormal propagation association results between each associated node based on the node connection relationship to generate an abnormal propagation sequence. S325. Perform an associated anomaly determination operation on each associated node according to the anomaly propagation sequence, determine the associated nodes that meet the preset anomaly propagation conditions as associated anomaly nodes, and generate an associated anomaly sequence.
8. The intelligent operation and maintenance method for government data governance based on big data according to claim 6, characterized in that, S34 specifically includes: S341. Extract the node connection relationships and associated abnormal nodes in each risk propagation path, and construct the corresponding node propagation structure based on the node connection relationships; S342. Based on the node propagation structure, count the number of node accesses, the number of abnormal associations, and the number of node association paths corresponding to each associated abnormal node, and generate the propagation weight sequence corresponding to each associated abnormal node based on the statistical results. S343. Based on the propagation weight sequence, a personalized PageRank algorithm is used to calculate the propagation impact value of each associated abnormal node in the node propagation structure, and the propagation impact value between each associated abnormal node is transferred according to the node connection relationship to generate an intermediate impact sequence. S344. Extract the operational status features corresponding to each digital twin mapping node in the digital twin governance body, and perform dynamic adjustment operations on the propagation impact values in the intermediate impact sequence based on the operational status features to generate a dynamic impact sequence. S345. Perform iterative update operations on each propagation influence value in the dynamic influence sequence until the change results of each propagation influence value in two consecutive iterations reach the preset convergence threshold to obtain a stable influence sequence. S346. Based on the propagation impact values in the stable impact sequence, determine the node propagation priority in each risk propagation path, and generate the node impact sequence corresponding to each risk propagation path based on the node propagation priority.
9. The intelligent operation and maintenance method for government data governance based on big data according to claim 1, characterized in that, S4 specifically includes: S41. Based on the governance risk sequence, extract the risk level, risk propagation intensity and node propagation priority corresponding to each risk propagation path, construct the corresponding risk features, and perform strategy matching operation in the preset governance strategy library according to each risk feature to generate candidate strategy sequence. S42. Map the candidate strategy sequence to the digital twin governance body, perform governance process simulation and deduction operations on the digital twin mapping nodes corresponding to each candidate strategy, and extract the state change results corresponding to each candidate strategy based on the digital twin association structure to generate a twin simulation sequence. S43. Based on the twin simulation sequence, statistically analyze the risk diffusion range, node state change results, and resource consumption results of each candidate strategy, and construct a strategy risk matrix based on the statistical results; S44. Perform risk association mapping operation on each strategy risk data in the strategy risk matrix, and sort the strategy risk data according to the node propagation priority corresponding to each risk propagation path to generate a risk sorting sequence. S45. A multidimensional assignment algorithm is used to perform joint allocation calculations on the risk data of each strategy in the risk ranking sequence, and a strategy allocation relationship is established based on the risk propagation path in the governance risk sequence and the digital twin mapping node in the digital twin governance body to generate a strategy allocation sequence. S46. Based on the strategy allocation sequence, perform correlation matching analysis on the resource occupancy results, node status change results and risk diffusion range corresponding to each candidate strategy, screen candidate strategies that meet the preset governance conditions, and obtain the target strategy sequence.
10. The intelligent operation and maintenance method for government data governance based on big data according to claim 1, characterized in that, S5 specifically includes: S51. Extract the node propagation priority and resource consumption results corresponding to each target strategy based on the target strategy sequence, and map them to the corresponding digital twin mapping nodes in the digital twin governance body to generate a strategy mapping sequence. S52. Based on the strategy mapping sequence, perform strategy adaptation analysis on each digital twin mapping node in the digital twin governance body, and extract the resource occupancy information, task execution information and service status information corresponding to each digital twin mapping node by combining the running status sequence, and generate a status adaptation sequence. S53. Perform association analysis on each digital twin mapping node based on the state adaptation sequence, and generate a scheduling association sequence in combination with the digital twin association structure; S54. Arrange the resource usage results corresponding to each target strategy in order according to the scheduling association sequence, and adjust the priority of the association sorting results in combination with the node propagation priority to generate the operation and maintenance scheduling sequence. S55. Perform a unified instruction template mapping operation on the operation and maintenance scheduling sequence, and generate the corresponding operation and maintenance instruction sequence based on the node connection relationship, resource occupation results and status adaptation sequence.