A big data aggregation analysis method and system based on typical business scenarios

By combining time series analysis, causal inference, graph neural network, attention mechanism and reinforcement learning technology in big data aggregation analysis methods and systems, the problem of difficulty in mining time characteristics and causal relationships in the existing technology when processing multi-source and multi-dimensional business data is solved, and more accurate and credible data analysis is achieved to support scientific decision-making in enterprises.

CN119415893BActive Publication Date: 2025-05-06BEIJING SHUYANG SMART TECH CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510012667.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-01-06
Publication Date
2025-05-06
Estimated Expiration
2045-01-06

AI Technical Summary

Technical Problem

When the prior art processes large-scale, multi-source, and multi-dimensional business data, it is difficult to effectively mine the time characteristics and causal relationships of the data, and it is impossible to handle the complexity of the cross-domain data-related network well, resulting in insufficient credibility and practicality of the analysis results.

Method used

It provides a big data aggregation analysis method and system based on typical business scenarios. By receiving multi-dimensional business data streams from multiple heterogeneous data sources, combining time series analysis and causal inference technology, it deeply explores the time characteristics and causal relationships of data, and builds a cross-domain data association network. Then, using graph neural network algorithm and attention mechanism, deep learning processes nodes and edge features in the network and filters out key data nodes and their combination modes. Using reinforcement learning mechanism and Q-learning algorithm, we optimize the data aggregation and analysis process, dynamically adjust the selection strategies of key data nodes, and generate comprehensive business insight reports.

Benefits of technology

Through this method and system, multi-source data can be effectively integrated, deeply revealed the hidden laws and causal relationships behind the data, improve the accuracy and credibility of the analysis, enhance the interpretability and robustness of the model, generate scientific business insight reports, and support enterprises to make more scientific decisions.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119415893B_ABST
    Figure CN119415893B_ABST
Patent Text Reader

Abstract

The present application provides a method and system for big data aggregation analysis based on a typical business scenario. Among them, a multi-dimensional business data stream is received from multiple heterogeneous data sources; the multi-dimensional business data stream is used, combined with time series analysis and causal inference technology, to deeply mine and process the time characteristics and causal relationships of the data to obtain a cross-domain data association network; based on the cross-domain data association network, a graph neural network algorithm is used to perform deep learning processing on the characteristics of nodes and edges in the network to obtain key data nodes and their combination patterns; according to the key data nodes and their combination patterns, the selection strategy of key data nodes is dynamically adjusted through the Q-learning algorithm to obtain an optimized data aggregation analysis model; based on the optimized data aggregation analysis model, the business data is comprehensively analyzed and processed to generate a comprehensive business insight report. The technical solution provided by this application can significantly improve the efficiency and accuracy of big data aggregation analysis.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The embodiments of the present application relate to the field of data processing technology, and in particular, to a method and system for big data aggregation and analysis based on typical business scenarios. Background Art

[0002] In modern enterprises, business data usually comes from multiple heterogeneous data sources, such as sales systems, customer relationship management systems, supply chain systems, etc. These data have multi-dimensional characteristics, including time series data, transaction records, user behavior data, etc. Enterprises need to conduct comprehensive analysis of these multi-dimensional business data flows to explore the time characteristics and causal relationships behind the data and build a cross-domain data association network. Through such analysis, enterprises can better understand business processes, discover potential business opportunities and risks, and make scientific decisions.

[0003] At present, many companies use traditional data analysis methods, such as statistical analysis and machine learning, to process multi-dimensional business data. These methods usually include steps such as data cleaning, feature extraction, and model training. Although these methods perform well in some aspects, they have certain limitations when processing large-scale, multi-source, and multi-dimensional data. For example, traditional methods are difficult to effectively mine the temporal characteristics and causal relationships of data, and cannot handle the complexity of cross-domain data association networks well.

[0004] When dealing with multi-source heterogeneous data, traditional methods have difficulty in data integration and are prone to losing important information, resulting in inaccurate analysis results; existing data analysis methods often lack in-depth exploration of the temporal characteristics and causal relationships of the data, and are unable to fully reveal the hidden laws behind the data; traditional models have poor interpretability and robustness when dealing with complex network structures, making it difficult to highlight the influence of key data nodes, resulting in insufficient credibility and practicality of the analysis results. Summary of the invention

[0005] The embodiments of the present application provide a big data aggregation analysis method and system based on a typical business scenario to solve the problem of insufficient credibility and practicality of analysis results in the prior art.

[0006] In a first aspect, the present application provides a big data aggregation analysis method based on a typical business scenario, including:

[0007] Receive multi-dimensional business data streams from multiple heterogeneous data sources; use the multi-dimensional business data streams, combined with time series analysis and causal inference technology, to deeply mine and process the time characteristics and causal relationships of the data to obtain a cross-domain data association network; based on the cross-domain data association network, use a graph neural network algorithm to perform deep learning processing on the features of nodes and edges in the network, introduce an attention mechanism to highlight the influence of important nodes, and obtain key data nodes and their combination patterns; based on the key data nodes and their combination patterns, use a reinforcement learning mechanism to optimize the data aggregation analysis process, dynamically adjust the selection strategy of key data nodes through a Q-learning algorithm, and obtain an optimized data aggregation analysis model; based on the optimized data aggregation analysis model, perform comprehensive analysis and processing on the business data to generate a comprehensive business insight report.

[0008] In a second aspect, the embodiment of the present application provides a big data aggregation analysis system based on a typical business scenario, including:

[0009] A receiving module, used to receive multi-dimensional business data streams from multiple heterogeneous data sources;

[0010] A processing module is used to utilize the multi-dimensional business data stream, combine time series analysis and causal inference technology, conduct in-depth mining and processing on the time characteristics and causal relationships of the data, and obtain a cross-domain data association network;

[0011] A learning module is used to perform deep learning processing on the features of nodes and edges in the network based on the cross-domain data association network using a graph neural network algorithm, introduce an attention mechanism to highlight the influence of important nodes, and obtain key data nodes and their combination patterns;

[0012] An optimization module is used to optimize the data aggregation analysis process according to the key data nodes and their combination patterns by using a reinforcement learning mechanism, dynamically adjust the selection strategy of the key data nodes by using a Q-learning algorithm, and obtain an optimized data aggregation analysis model;

[0013] The generation module is used to perform comprehensive analysis and processing on the business data based on the optimized data aggregation analysis model to generate a comprehensive business insight report.

[0014] In an embodiment of the present application, a multi-dimensional business data stream is received from multiple heterogeneous data sources; the multi-dimensional business data stream is used, combined with time series analysis and causal inference technology, to deeply mine and process the time characteristics and causal relationships of the data to obtain a cross-domain data association network; based on the cross-domain data association network, a graph neural network algorithm is used to perform deep learning processing on the features of the nodes and edges in the network, and an attention mechanism is introduced to highlight the influence of important nodes, so as to obtain key data nodes and their combination patterns; based on the key data nodes and their combination patterns, a reinforcement learning mechanism is used to optimize the data aggregation analysis process, and the selection strategy of the key data nodes is dynamically adjusted through the Q-learning algorithm to obtain an optimized data aggregation analysis model; based on the optimized data aggregation analysis model, the business data is comprehensively analyzed and processed to generate a comprehensive business insight report.

[0015] By receiving multi-dimensional business data streams from multiple heterogeneous data sources, this method can effectively integrate data from different sources, improve the comprehensiveness and accuracy of the data, and provide a solid foundation for subsequent in-depth analysis; using time series analysis and causal inference technology, the temporal characteristics and causal relationships of the data are deeply mined and processed, which can reveal the hidden laws and potential causal relationships behind the data, and provide strong support for the construction of cross-domain data association networks; based on the cross-domain data association network, the graph neural network algorithm is used to perform deep learning processing on the features of nodes and edges in the network, and the attention mechanism is introduced to highlight the influence of important nodes, which can more accurately identify key data nodes and their combination patterns, and improve the model's explanatory power and prediction accuracy; by optimizing the data aggregation analysis process using the reinforcement learning mechanism, especially by dynamically adjusting the selection strategy of key data nodes through the Q-learning algorithm, the data aggregation analysis model can be gradually optimized, so that it can show better performance and robustness in practical applications; based on the optimized data aggregation analysis model, the business data is comprehensively analyzed and processed to generate a comprehensive business insight report. The report not only provides a comprehensive description of the current business situation, but also contains predictions and suggestions for future business development, providing scientific decision-making support for enterprise managers, helping them better understand the business status, grasp market trends, and formulate effective strategic plans. Through the above method, the present invention can significantly improve the efficiency and accuracy of big data aggregation analysis, and provide strong support for the business decision-making of enterprises.

[0016] Based on the cross-domain data association network, the graph neural network algorithm is used to perform deep learning processing on the features of nodes and edges. The deep feature representation of the nodes is obtained through multi-layer nonlinear transformation, and the self-attention mechanism is combined to evaluate the relationship strength between nodes, dynamically adjust the attention weight, and use the node importance scoring algorithm to screen out key data nodes, analyze their interaction patterns, and generate combination patterns that reflect business logic; the deep feature representation of the nodes is extracted through multi-layer nonlinear transformation, while retaining the topological structure information, thereby improving the richness and accuracy of the features; the self-attention mechanism is combined to dynamically adjust the attention weight, highlighting the influence of important nodes, and enhancing the interpretability and robustness of the model; the node importance scoring algorithm is used to screen out key data nodes, which helps to focus on core information and improve the efficiency and accuracy of analysis.

[0017] The method optimizes the data aggregation analysis process according to the key data nodes and their combination patterns by using a reinforcement learning mechanism, dynamically adjusts the selection strategy of the key data nodes by using a Q-learning algorithm, and obtains an optimized data aggregation analysis model, including: constructing a reinforcement learning framework, defining state space and actions, using a Q-learning algorithm to update the Q value through multiple iterations of learning, selecting the action with the highest Q value as the selection strategy of the key data node, observing environmental feedback after execution to obtain an immediate reward, dynamically adjusting the selection strategy based on the immediate reward until converging to the optimal strategy, and finally obtaining the optimized data aggregation analysis model; dynamically adjusting the selection strategy of the key data nodes by using a Q-learning algorithm, and being able to continuously optimize the selection strategy according to environmental feedback to improve the adaptability and robustness of the model; optimizing the data aggregation analysis process by using a reinforcement learning mechanism, so that the model can gradually learn more effective strategies and improve the efficiency and accuracy of data aggregation analysis; through a continuous iterative learning process, the model can gradually converge to the optimal strategy, ensuring the best performance in practical applications, and providing more scientific decision support for enterprises; analyzing the interaction pattern of key data nodes, generating a combination pattern reflecting business logic, and providing enterprises with deeper business insights and decision support.

[0018] These and other aspects of the present application will become more clearly understood in the description of the following embodiments. BRIEF DESCRIPTION OF THE DRAWINGS

[0019] In order to more clearly illustrate the embodiments of the present application or the technical solutions in the prior art, a brief introduction will be given below to the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are some embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without paying any creative work.

[0020] Figure 1A flowchart of a big data aggregation analysis method based on a typical business scenario provided in an embodiment of the present application;

[0021] Figure 2 A schematic diagram of the structure of a big data aggregation and analysis system based on a typical business scenario provided in an embodiment of the present application. DETAILED DESCRIPTION

[0022] In order to enable those skilled in the art to better understand the solution of the present application, the technical solution in the embodiments of the present application will be clearly and completely described below in conjunction with the drawings in the embodiments of the present application.

[0023] In some of the processes described in the specification and claims of this application and the above-mentioned figures, multiple operations that appear in a specific order are included, but it should be clearly understood that these operations may not be executed in the order in which they appear in this article or executed in parallel. The serial numbers of the operations, such as 101, 102, etc., are only used to distinguish between different operations, and the serial numbers themselves do not represent any execution order. In addition, these processes may include more or fewer operations, and these operations may be executed in sequence or in parallel. It should be noted that the descriptions of "first", "second", etc. in this article are used to distinguish different messages, devices, modules, etc., do not represent the order of precedence, and do not limit the "first" and "second" to be different types.

[0024] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the present application, all other embodiments obtained by those skilled in the art without creative work are within the scope of protection of this application.

[0025] Figure 1 A flowchart of a big data aggregation analysis method based on a typical business scenario is provided for the embodiment of the present application. Figure 1 As shown, the method includes:

[0026] 101. Receive multi-dimensional business data streams from multiple heterogeneous data sources;

[0027] Refers to business data of various types and attributes obtained from multiple different sources, such as sales data, customer behavior data, supply chain data, etc.; different data sources may include databases, file systems, real-time streaming data, etc., and the formats and structures of these data sources may be different; data is obtained from multiple heterogeneous data sources through data collection tools or APIs, such as extracting sales records from SQL databases and reading user behavior data from log files; the collected data is cleaned, formatted, and standardized to ensure data consistency and availability.

[0028] For example, in financial transaction data, customer transaction records are extracted from the bank's transaction system, including transaction amount, transaction time, transaction type (such as deposit, withdrawal, transfer), etc.; in social media data, user posts, comments and likes data are obtained from the API of the social media platform, including user ID, post content, release time, etc. In IoT data, sensor data collected from smart home devices includes temperature, humidity, light intensity, etc.

[0029] 102. Utilizing the multi-dimensional business data stream, combined with time series analysis and causal inference technology, the temporal characteristics and causal relationships of the data are deeply mined and processed to obtain a cross-domain data association network;

[0030] By analyzing time series data, we can discover the patterns and trends of data changes over time; through statistical and machine learning methods, we can infer the causal relationship between data and identify which factors lead to specific results; we can associate data from different sources to form a network containing nodes and edges, where nodes represent data entities and edges represent the relationships between entities; we can analyze the time series part of multi-dimensional business data and extract temporal patterns and trends; we can use causal inference technology to identify the causal relationship between data, such as whether sales growth is affected by advertising; we can integrate the analysis results and build a cross-domain data association network, where nodes represent data entities and edges represent the relationships between entities.

[0031] For example, in financial transaction data, we analyze the transaction records of customers to find the trend of daily transaction volume and identify the peak and trough periods of transaction; in social media data, we analyze the frequency of users’ posts to identify active time periods, such as the difference between weekdays and weekends; in networked data, we use the sensor data of smart home devices to find the periodic changes in temperature and humidity. Causal inference: In financial transaction data, we use causal inference technology to analyze the impact of holidays and promotions on transaction volume and identify which factors lead to an increase in transaction volume; in social media data, we analyze the content of users’ posts and interactive data to identify the impact of hot topics on user activity, such as a large number of discussions triggered by an event; in networked data, we analyze the impact of temperature and humidity on the frequency of use of smart home devices, such as the increase in the frequency of air conditioning use due to hot weather. Build an association network: In financial transaction data, customers, transaction types, and transaction times are associated to form a network. Nodes include customers, transaction types, etc., and edges represent transaction relationships. In social media data, users, post content, and interactive data are associated to form a network. Nodes include users, posts, etc., and edges represent interactive relationships. In networked data, devices, sensor data, and time are associated to form a network. Nodes include devices, sensors, etc., and edges represent data transmission relationships.

[0032] 103. Based on the cross-domain data association network, a graph neural network algorithm is used to perform deep learning processing on the features of nodes and edges in the network, and an attention mechanism is introduced to highlight the influence of important nodes, so as to obtain key data nodes and their combination patterns;

[0033] A deep learning algorithm for processing graph data, which can extract the features of nodes and edges; dynamically adjust the attention weights by calculating the similarity scores between nodes to highlight the influence of important nodes; nodes that play a core role in the network have an important impact on business decisions; the interaction pattern between key data nodes reflects the business logic; use the graph neural network algorithm to perform deep learning processing on the features of nodes and edges in the cross-domain data association network to extract deep-level feature representations; combine the self-attention mechanism to evaluate the relationship strength between nodes and dynamically adjust the attention weights; use the node importance scoring algorithm to screen out key data nodes; analyze the interaction pattern between key data nodes to generate a combination pattern that reflects the business logic.

[0034] Optionally, in step 103, based on the cross-domain data association network, a graph neural network algorithm is used to perform deep learning processing on the features of nodes and edges in the network, and an attention mechanism is introduced to highlight the influence of important nodes, so as to obtain key data nodes and their combination patterns, including:

[0035] Based on the cross-domain data association network, deep learning processing is performed on the features of the nodes and edges in the network, and a deep feature representation of each node is obtained through multi-layer nonlinear transformations, while retaining the topological structure information between the nodes to obtain a deep feature representation of the nodes; using the deep feature representation of the nodes, combined with the self-attention mechanism, the strength of the relationship between each node and its neighboring nodes is evaluated, and the attention weight is dynamically adjusted by calculating the similarity score of the features between the nodes to obtain the attention weight; according to the attention weight, a node importance scoring algorithm is used to evaluate the importance of the nodes in the network, and key data nodes that play a core role in the network are screened out to obtain key data nodes; based on the key data nodes, the interaction pattern between the nodes is analyzed to generate a combination pattern of key data nodes that reflect the business logic.

[0036] In the field of energy and power, the State Grid needs to conduct comprehensive analysis on a large amount of power equipment and user data to optimize power grid operation and improve service quality. These data include equipment status data, user power consumption data, meteorological data, etc., with diverse sources and complex structures; collect power equipment status data (such as transformers, transmission line status), user power consumption data (such as power consumption, power consumption time), meteorological data (such as temperature, humidity), etc. from different data sources; build a cross-domain data association network to associate equipment, users and meteorological data to form a network containing nodes and edges. The graph neural network algorithm is used to perform deep learning processing on the features of nodes and edges in the network. The deep feature representation of each node is obtained through multi-layer nonlinear transformation, while retaining the topological structure information between nodes. The deep feature representation of nodes is used in combination with the self-attention mechanism to evaluate the strength of the relationship between each node and its neighboring nodes, and the attention weight is dynamically adjusted by calculating the similarity score of the features between nodes. According to the similarity score, the attention weight is dynamically adjusted to highlight the influence of important nodes, such as key equipment and high-power users. According to the attention weight, the node importance scoring algorithm is used to evaluate the importance of nodes in the network, and the key data nodes that play a core role in the network are screened out. For example, the equipment and users that have the greatest impact on the operation of the power grid are identified; the key data nodes screened out include key equipment (such as main transformers, transmission lines) and high-power users (such as large industrial users); based on the key data nodes, the interaction patterns between nodes are analyzed to generate a combination pattern of key data nodes that reflect the business logic. For example, the relationship between the operating status of key equipment and the power consumption of users is analyzed to identify which equipment failures may cause power outages for users; the generated combination pattern can help the State Grid optimize equipment maintenance plans, prevent equipment failures in advance, and improve the reliability and service quality of the power grid.

[0037] The deep feature representation of nodes is extracted through the graph neural network algorithm, which improves the richness and accuracy of the features; combined with the self-attention mechanism, the attention weight is dynamically adjusted to highlight the influence of important nodes, enhancing the interpretability and robustness of the model; key data nodes are screened out through the node importance scoring algorithm to help focus on core information and improve the efficiency and accuracy of analysis; the interaction patterns of key data nodes are analyzed to generate combination patterns that reflect business logic, providing the State Grid with deeper business insights and decision-making support.

[0038] Optionally, the node importance scoring algorithm is used to evaluate the importance of nodes in the network according to the attention weight, and key data nodes that play a core role in the network are screened out to obtain key data nodes, including:

[0039] By using the attention weight, each node in the network is evaluated for importance using a node importance scoring algorithm to obtain an importance scoring value; according to the importance scoring value, the nodes in the network are sorted to generate an importance ranking list of the nodes; based on the importance ranking list of the nodes, a threshold is set or the top N nodes are selected to screen the nodes in the network to obtain key data nodes that play a core role in the network; by further analyzing the key data nodes that play a core role in the network, the actual role and influence of these nodes in the network are confirmed to generate key data nodes.

[0040] In the field of energy and power, the State Grid needs to conduct a comprehensive analysis of a large amount of power equipment and user data to optimize power grid operation and improve service quality. These data include equipment status data, user power consumption data, meteorological data, etc., with diverse sources and complex structures; collect power equipment status data (such as transformers, transmission line status), user power consumption data (such as power consumption, power consumption time), meteorological data (such as temperature, humidity), etc. from different data sources; build a cross-domain data association network to associate equipment, users and meteorological data to form a network containing nodes and edges. Use the graph neural network algorithm to perform deep learning processing on the features of nodes and edges in the network, and obtain the deep feature representation of each node through multi-layer nonlinear transformation, while retaining the topological structure information between nodes; use the deep feature representation of nodes, combined with the self-attention mechanism, to evaluate the strength of the relationship between each node and its neighboring nodes, and dynamically adjust the attention weight by calculating the similarity score of the features between nodes; use the attention weight to evaluate the importance of each node in the network using the node importance scoring algorithm to obtain the importance score value. For example, the importance score of each node is calculated using the PageRank algorithm or the node centrality algorithm; the nodes in the network are sorted according to the importance score value to generate an importance ranking list of the nodes. For example, all nodes are sorted from high to low according to the importance score value; based on the importance ranking list of the nodes, a threshold is set or the top N nodes are selected, and the nodes in the network are screened to obtain key data nodes that play a core role in the network. For example, the nodes in the top 10% of the importance score value are selected as key data nodes; the key data nodes are screened out according to the set threshold value or the top N nodes selected. For example, the 100 nodes with the highest importance score value are screened out; by further analyzing the key data nodes that play a core role in the network, the actual role and influence of these nodes in the network are confirmed. For example, the relationship between the operating status of key equipment and the power consumption of users is analyzed to identify which equipment failures may cause power outages for users; based on the results of further analysis, a list of key data nodes is generated. For example, the actual role and influence of key equipment (such as main transformers, transmission lines) and high power consumption users (such as large industrial users) in the network are confirmed.

[0041] The importance of each node in the network is evaluated through the node importance scoring algorithm to ensure that the selected nodes do have a significant impact on the business; the nodes are sorted to quickly identify the most important nodes to improve analysis efficiency; key data nodes are screened out by setting thresholds or selecting the top N nodes to reduce the complexity of subsequent analysis and improve efficiency; the actual role and influence of key data nodes in the network are confirmed to provide State Grid with deeper business insights and decision-making support.

[0042] Optionally, the deep feature representation of the node is used in combination with a self-attention mechanism to evaluate the strength of the relationship between each node and its neighboring nodes, and the attention weight is dynamically adjusted by calculating the similarity score of the features between the nodes to obtain the attention weight, including:

[0043] Using the deep feature representation of the node and combining it with the self-attention mechanism, the feature similarity between each node and all its neighboring nodes is calculated to obtain a similarity score; based on the similarity score, normalization is performed through a softmax function to generate an attention weight; based on the attention weight, factors such as the distance between nodes or the type of relationship are introduced as additional inputs to refine the evaluation of the strength of the relationship between nodes and obtain a more accurate attention weight; using the more accurate attention weight, the feature representation of the node is weighted and summed to generate the final attention weight.

[0044] In the field of energy and power, the State Grid needs to conduct a comprehensive analysis of a large amount of power equipment and user data to optimize power grid operation and improve service quality. These data include equipment status data, user power consumption data, meteorological data, etc., with diverse sources and complex structures; collect power equipment status data (such as transformers, transmission line status), user power consumption data (such as power consumption, power consumption time), meteorological data (such as temperature, humidity), etc. from different data sources; build a cross-domain data association network to associate equipment, users and meteorological data to form a network containing nodes and edges. Use the graph neural network algorithm to perform deep learning processing on the features of nodes and edges in the network, and obtain the deep feature representation of each node through multi-layer nonlinear transformation; use the deep feature representation of the node, combined with the self-attention mechanism, calculate the feature similarity between each node and all its neighboring nodes, and obtain the similarity score. For example, calculate the feature similarity between the transformer and the adjacent transmission line; according to the similarity score, normalize it through the softmax function to generate the attention weight. For example, a softmax function is used to convert similarity scores into normalized attention weights; based on the attention weights, factors such as the distance between nodes or the type of relationship are introduced as additional inputs to refine the evaluation of the strength of the relationship between nodes and obtain more accurate attention weights. For example, the physical distance between the transformer and the transmission line, as well as the type of connection between them (direct connection or indirect connection) are considered; using more accurate attention weights, the feature representations of the nodes are weighted and summed to generate the final attention weights. For example, the feature representations of the transformer are weighted and summed to generate the final attention weights; using the final attention weights, each node in the network is evaluated for importance using a node importance scoring algorithm to obtain an importance score value. For example, the importance score of each node is calculated using a PageRank algorithm or a node centrality algorithm; according to the importance score value, the nodes in the network are sorted to generate an importance ranking list of the nodes. For example, all nodes are sorted from high to low according to the importance score value; based on the importance ranking list of the nodes, a threshold is set or the top N nodes are selected to screen the nodes in the network to obtain key data nodes that play a core role in the network. For example, select the top 10% of the nodes in importance score as key data nodes; filter out key data nodes according to the set threshold or the top N nodes selected. For example, filter out the 100 nodes with the highest importance score; further analyze the key data nodes that play a core role in the network to confirm the actual role and influence of these nodes in the network. For example, analyze the relationship between the operating status of key equipment and the user's power consumption, and identify which equipment failures may cause power outages for users; generate a list of key data nodes based on the results of further analysis.For example, confirm the actual role and influence of key equipment (such as major transformers, transmission lines) and high electricity users (such as large industrial users) in the network.

[0045] Calculate the feature similarity between nodes, evaluate the relationship strength between nodes, and provide a basis for the subsequent generation of attention weights; normalize the similarity scores through the softmax function to generate attention weights, and ensure that the sum of the weights is 1 for subsequent processing; introduce factors such as the distance between nodes or the type of relationship to refine the evaluation of the relationship strength between nodes and improve the accuracy of attention weights; use more accurate attention weights to perform weighted summation on the feature representation of the nodes to generate the final attention weights to provide support for the screening of important nodes; evaluate the importance of each node in the network through the node importance scoring algorithm to ensure that the screened nodes do have an important impact on the business; sort the nodes to quickly identify the most important nodes and improve analysis efficiency; screen out key data nodes by setting thresholds or selecting the top N nodes to reduce the complexity of subsequent analysis and improve efficiency; confirm the actual role and influence of key data nodes in the network to provide State Grid with deeper business insights and decision support.

[0046] This application takes into account that in the field of energy and power, the State Grid needs to conduct comprehensive analysis of a large amount of power equipment and user data to optimize power grid operation and improve service quality. These data include equipment status data, user power consumption data, meteorological data, etc., with diverse sources and complex structures.

[0047] Optionally, the generating attention weight by normalizing the similarity score through a softmax function includes:

[0048] In order to more accurately reflect the relationship strength between nodes, the feature similarity score, relationship type weight and distance between nodes are comprehensively considered, and after nonlinear adjustment and normalization, the attention weight of each node to its neighbor node is generated. :

[0049] ;

[0050] in, Representation Node Its neighbor nodes The feature similarity score between them; Representation Node Its neighbor nodes The feature similarity score between ; Representation Node With Node The distance between Representation Node With Node The distance between Representation Node With Node The weight of the relationship type between them; and is an adjustment factor used to control the influence of relationship type weight and distance on the similarity score; and It is a nonlinear adjustment factor, which is used to introduce nonlinear terms and enhance the flexibility of the model; is a threshold used to control the activation condition of the nonlinear term; Representation Node The set of all neighbor nodes of ; is a natural exponential function, which is used to convert the adjusted similarity score into a positive value for normalization; is a nonlinear term that is used to reduce the influence of the similarity score when it is low and increase its influence when it is high.

[0051] Through this formula, factors such as the feature similarity score, relationship type weight and distance between nodes are comprehensively considered, and after nonlinear adjustment and normalization, the attention weight of each node to its neighbor node is generated.

[0052] This method aims to generate the attention weight of each node to its neighboring nodes by comprehensively considering factors such as feature similarity scores, relationship type weights and distances between nodes, using nonlinear adjustment and normalization processing, so as to more accurately reflect the relationship strength between nodes and improve the flexibility and accuracy of the model.

[0053] In the attention weight In the similarity score Reflects the similarity of features between nodes and is the basis for evaluating the strength of relationships between nodes; distance Consider the physical or logical distance between nodes, which affects the evaluation of relationship strength. The closer the distance, the stronger the relationship may be. Relationship type weight Consider the type of relationship between nodes, such as direct connection or indirect connection, which affects the evaluation of relationship strength. Different types of relationships may have different importance. Adjustment factor and By adjusting and , which can balance the influence of feature similarity, distance and relationship type weight, making the model more flexible; nonlinear adjustment factor and Through nonlinear terms , reducing its influence when the similarity score is low and increasing its influence when the similarity score is high, making the model more robust; threshold By setting the appropriate , the activation conditions of nonlinear terms can be controlled to make the model better adapt to different data distributions.

[0054] Among them, the feature similarity score Through the computing node and nodes The similarity between the feature vectors (such as cosine similarity, Euclidean distance, etc.) is obtained; distance According to the node and nodes The physical or logical distance between them is calculated. For example, for power equipment, the geographical distance or network topology distance can be used. The relationship type weight According to the node and nodes The relationship type (such as direct connection, indirect connection, etc.) is pre-defined or determined through expert knowledge; the adjustment factor and Find the optimal parameter value through experiments or cross-validation adjustment; nonlinear adjustment factor and Find the optimal parameter value through experiments or cross-validation adjustment; threshold Find the optimal parameter value through experiments or cross-validation adjustment.

[0055] Collect power equipment status data (such as transformers, transmission line status), user power consumption data (such as power consumption, power consumption time), meteorological data (such as temperature, humidity), etc. from different data sources; associate equipment, users and meteorological data to form a network containing nodes and edges. Nodes include equipment, users and meteorological stations, and edges represent the relationship between them; use graph neural network algorithms to perform deep learning processing on the features of nodes and edges in the network, and obtain deep feature representation of each node through multi-layer nonlinear transformation; calculate nodes and nodes The similarity score between the feature vectors For example, the cosine similarity is used to calculate the feature similarity between the transformer and the adjacent transmission line; and nodes The physical distance between For example, calculate the geographical distance between the transformer and the transmission line; and nodes Predefine relationship type weights between relationship types For example, the weight of the direct connection is set to 1, and the weight of the indirect connection is set to 0.5; find the optimal parameter value through experiments or cross-validation adjustment ; Use the above formula to calculate the attention weight of each node to its neighboring nodes ; Using the final attention weight, each node in the network is evaluated for importance using a node importance scoring algorithm to obtain an importance score. For example, the PageRank algorithm is used to calculate the importance score of each node; according to the importance score, the nodes in the network are sorted to generate an importance ranking list of the nodes. For example, all nodes are sorted from high to low according to the importance score; based on the importance ranking list of the nodes, a threshold is set or the top N nodes are selected, and the nodes in the network are screened to obtain key data nodes that play a core role in the network. For example, the top 10% of the nodes in importance score are selected as key data nodes; according to the set threshold or the top N nodes selected, the key data nodes are screened out. For example, the 100 nodes with the highest importance score are screened out; through further analysis of the key data nodes that play a core role in the network, the actual role and influence of these nodes in the network are confirmed. For example, the relationship between the operating status of key equipment and the power consumption of users is analyzed to identify which equipment failures may cause power outages for users; based on the results of further analysis, a list of key data nodes is generated. For example, confirm the actual role and influence of key equipment (such as major transformers, transmission lines) and high electricity users (such as large industrial users) in the network.

[0056] By comprehensively considering feature similarity, distance and relationship type weights, the relationship strength between nodes can be more accurately reflected; nonlinear terms are introduced to enhance the flexibility of the model so that it can better adapt to different data distributions and scenarios; normalization is performed to ensure that the sum of attention weights is 1, which is convenient for subsequent weighted summation processing; the importance of each node in the network is evaluated through a node importance scoring algorithm to ensure that the selected nodes do have an important impact on the business; the nodes are sorted to quickly identify the most important nodes and improve analysis efficiency; key data nodes are selected by setting thresholds or selecting the top N nodes to reduce the complexity of subsequent analysis and improve efficiency; the actual role and influence of key data nodes in the network are confirmed to provide State Grid with deeper business insights and decision support.

[0057] 104. According to the key data nodes and their combination patterns, the data aggregation analysis process is optimized using a reinforcement learning mechanism, and the selection strategy of the key data nodes is dynamically adjusted through a Q-learning algorithm to obtain an optimized data aggregation analysis model;

[0058] Through interaction with the environment, the strategy is continuously optimized to maximize the cumulative reward; a reinforcement learning algorithm that updates the Q value of each state-action pair through iterative learning to reflect the expected return after adopting a certain selection strategy; a model for comprehensive analysis of business data, which improves the analysis effect by optimizing the selection strategy; the state space is defined as a collection of key data nodes and their combination patterns, and the action is the selection strategy of the key data nodes in the space; through multiple iterative learning, the Q value of each state-action pair is updated to reflect the expected return after adopting a certain selection strategy; in each iteration, the action with the highest Q value is selected as the selection strategy for the key data node according to the current state and the updated Q value table, and the environmental feedback is observed after execution to obtain an immediate reward; through a continuous iterative learning process, the selection strategy is dynamically adjusted based on the immediate reward until it converges to the optimal strategy, and finally an optimized data aggregation analysis model is obtained.

[0059] Optionally, in step 104, according to the key data nodes and their combination patterns, the data aggregation analysis process is optimized by using a reinforcement learning mechanism, and the selection strategy of the key data nodes is dynamically adjusted by a Q-learning algorithm to obtain an optimized data aggregation analysis model, including:

[0060] The key data nodes and their combination patterns are used to construct a reinforcement learning framework, and the state space is defined as a set of key data nodes and their combination patterns, and the action is a selection strategy for the key data nodes in the space, so as to obtain a reinforcement learning framework; based on the reinforcement learning framework, the data aggregation analysis process is optimized using a Q-learning algorithm, and the Q value of each state-action pair is updated through multiple iterative learning, and the Q value reflects the expected return after adopting a certain selection strategy, so as to obtain an updated Q value table; in each iteration, according to the current state and the updated Q value table, the action with the highest Q value is selected as the selection strategy for the key data node, and the environmental feedback is observed after execution to obtain an immediate reward; through a continuous iterative learning process, the selection strategy of the key data node is dynamically adjusted based on the immediate reward, so that the model can gradually learn more effective strategies until it converges to the optimal strategy, and finally obtain an optimized data aggregation analysis model.

[0061] In the field of energy and power, the State Grid needs to conduct comprehensive analysis on a large amount of power equipment and user data to optimize power grid operation and improve service quality. These data include equipment status data, user power consumption data, meteorological data, etc., with diverse sources and complex structures; collect power equipment status data (such as transformers, transmission line status), user power consumption data (such as power consumption, power consumption time), meteorological data (such as temperature, humidity), etc. from different data sources; through graph neural networks and attention mechanisms, screen out key data nodes and their combination patterns; define the set of key data nodes and their combination patterns as state space; define the strategy for selecting key data nodes as actions; initialize the Q value table and set the Q value of all state-action pairs to 0; based on the reinforcement learning framework, use the Q-learning algorithm to optimize the data aggregation analysis process, and update the Q value of each state-action pair through multiple iterative learning. For example, select a key data node in the initial state, observe environmental feedback after execution, get immediate rewards, and update the Q value table; in each iteration, select the action with the highest Q value as the selection strategy for the key data node based on the current state and the updated Q value table. For example, based on the current state, select the key data node with the highest Q value; execute the selected key data node, observe the environmental feedback, and get an immediate reward. For example, select a key device for monitoring, observe the changes in its operating status, and get an immediate reward; update the Q value table based on the immediate reward. For example, if the selected key data node brings a significant optimization effect, increase its Q value; through a continuous iterative learning process, dynamically adjust the selection strategy of the key data node based on the immediate reward, so that the model can gradually learn more effective strategies until it converges to the optimal strategy. For example, through multiple iterations, the model learns to select the most appropriate key data nodes in different situations, and finally obtains an optimized data aggregation analysis model.

[0062] By defining the state space and actions, a basic framework is provided for the Q-learning algorithm to ensure the systematic and standardized nature of the optimization process; through multiple iterative learning, the Q value of each state-action pair is updated to reflect the expected return after adopting a certain selection strategy, and the selection strategy of key data nodes is gradually optimized; according to the current state and the updated Q value table, the action with the highest Q value is selected as the selection strategy for the key data node to ensure that each selection is currently optimal; through a continuous iterative learning process, the selection strategy is dynamically adjusted based on immediate rewards, so that the model can gradually learn more effective strategies until it converges to the optimal strategy, and finally obtains the optimized data aggregation analysis model; by dynamically adjusting the selection strategy of key data nodes, the efficiency and accuracy of data aggregation analysis are improved, providing more scientific decision-making support for the State Grid.

[0063] Optionally, based on the reinforcement learning framework, the data aggregation analysis process is optimized using a Q-learning algorithm, and the Q value of each state-action pair is updated through multiple iterative learning. The Q value reflects the expected return after taking a certain selection strategy, and an updated Q value table is obtained, including:

[0064] Using the reinforcement learning framework, a Q-value table is initialized, wherein the initial Q-value of each state-action pair is set to zero or a random value, thereby obtaining an initial Q-value table; based on the initial Q-value table, a Q-learning algorithm is used to start from the initial state, select an action, and observe environmental feedback after executing the action to obtain an immediate reward; according to the immediate reward and the Q-value of the next state, the Q-value of the current state-action pair is updated so that the Q-value gradually approaches the long-term expected return after adopting the selection strategy, thereby obtaining an updated Q-value; through multiple iterative learning, the process of selecting actions, executing actions, observing feedback, and updating Q-values ​​is repeated, and the Q-value of each state-action pair is continuously adjusted until the Q-value table converges, that is, the change of the Q-value tends to be stable, thereby generating an updated Q-value table.

[0065] In the field of energy and power, the State Grid needs to conduct a comprehensive analysis of a large amount of power equipment and user data to optimize power grid operation and improve service quality. These data include equipment status data, user power consumption data, meteorological data, etc., with diverse sources and complex structures; data preparation: collect power equipment status data (such as transformers, transmission line status), user power consumption data (such as power consumption, power consumption time), meteorological data (such as temperature, humidity), etc. from different data sources; through graph neural networks and attention mechanisms, screen out key data nodes and their combination patterns, for example, key equipment (such as main transformers, transmission lines) and high power consumption users (such as large industrial users); define the set of key data nodes and their combination patterns as state space; define the strategy for selecting key data nodes as actions; use the reinforcement learning framework to initialize a Q value table, in which the initial Q value of each state-action pair is set to zero or a random value. For example, the Q value of each state-action pair in the initial Q value table is set to 0; based on the initial Q value table, use the Q-learning algorithm to start from the initial state, select an action, observe the environmental feedback after executing the action, and get an immediate reward. For example, a key data node is selected in the initial state, and the changes in the power grid operation state are observed after execution to obtain an immediate reward; the Q value of the current state-action pair is updated according to the immediate reward and the Q value of the next state, so that the Q value gradually approaches the long-term expected return after adopting the selection strategy; through multiple iterations of learning, the process of selecting actions, executing actions, observing feedback, and updating Q values ​​is repeated, and the Q value of each state-action pair is continuously adjusted until the Q value table converges, that is, the change of Q value tends to be stable. For example, through multiple iterations, the model learns to select the most appropriate key data nodes in different situations, and finally obtains the optimized data aggregation analysis model; through the Q-learning algorithm, the optimized data aggregation analysis model is finally generated, and the model can select the optimal key data nodes according to the current state, improving the efficiency and accuracy of data aggregation analysis; the optimized model is applied to the actual power grid operation, for example, the most appropriate transformer is selected for monitoring according to the current power grid state, optimizing the power grid operation and improving the service quality.

[0066] Through the Q-learning algorithm, the selection strategy of key data nodes is dynamically adjusted to improve the efficiency and accuracy of data aggregation analysis; through continuous iterative learning, the model gradually learns more effective strategies until it converges to the optimal strategy, and finally obtains the optimized data aggregation analysis model; the optimized model provides more scientific decision-making support for the State Grid, helps optimize grid operation, and improves service quality.

[0067] This application considers that in reinforcement learning, the Q-learning algorithm updates the Q value of each state-action pair through iterative learning, reflecting the expected return after taking a certain selection strategy. In order to more accurately calculate the immediate reward and update the Q value, this method enhances the flexibility and robustness of the model by introducing nonlinear adjustment terms.

[0068] Optionally, based on the initial Q value table, using a Q-learning algorithm to start from an initial state, selecting an action, observing environmental feedback after executing the action, and obtaining an immediate reward, includes:

[0069] To calculate the immediate reward obtained after performing an action, the following formula is used:

[0070] ;

[0071] in, Indicates that at time step Instant rewards when Indicates in status Next action Rewards immediately received after is the learning rate, which controls the speed of Q value update; Indicates in status Next action The current Q value of Indicates execution of an action The next state reached after Indicates that in the next state The maximum Q value of all possible actions; and It is a nonlinear adjustment factor, which is used to introduce nonlinear terms and enhance the flexibility of the model; is a threshold used to control the activation condition of the nonlinear term;

[0072] According to the calculated instant reward, the Q value of the current state-action pair is updated using the following formula: The value gradually approaches the long-term expected return after adopting this selection strategy:

[0073] ;

[0074] in: Indicates in status Next action Current value; Indicates that at time step Instant rewards when is a discount factor that controls the weight of future rewards; Indicates that in the next state The maximum Q value of all possible actions; and It is a nonlinear adjustment factor, which is used to introduce nonlinear terms and enhance the flexibility of the model; is a threshold used to control the activation condition of the nonlinear term;

[0075] Through the above formula, first calculate the instant reward , and then use the immediate reward to update the Q value of the current state-action pair, thereby gradually optimizing the key data node selection strategy in the data aggregation analysis process.

[0076] This method aims to comprehensively consider immediate rewards and future rewards, use nonlinear adjustment terms to enhance the flexibility of the model, gradually optimize the key data node selection strategy in the data aggregation analysis process, and improve the performance and stability of the model.

[0077] Instant Rewards middle, Reflected in the status Next action The reward immediately after is the basis for evaluating the effect of the current selection strategy; the learning rate Controls the speed of Q value update; maximum Q value Reflects the optimal choice in the next state; nonlinear adjustment factor and Introducing nonlinear terms to enhance the flexibility of the model; threshold Controls the activation conditions of the nonlinear terms.

[0078] Update the Q value middle, Reflected in the status Next action The current Q value is the key to guiding the optimization of the selection strategy; the learning rate Control the speed of Q value update; discount factor Controls the weight of future rewards; maximum Q value Reflects the optimal choice in the next state; nonlinear adjustment factor and Introducing nonlinear terms to enhance the flexibility of the model; threshold Controls the activation conditions of the nonlinear terms.

[0079] in, is in state Next action The rewards obtained immediately after the action can be directly obtained through environmental feedback; is the learning rate, adjusted through experiments or cross-validation; is in the next state The maximum Q value of all possible actions is found through the Q value table; and is a nonlinear adjustment factor, adjusted through experiments or cross-validation; is the threshold, adjusted through experiments or cross-validation; is in state Next action The current Q value of is found through the Q value table; is at the time step The instant reward at the time is calculated by the above formula; is the learning rate, adjusted through experiments or cross-validation; is the discount factor, adjusted through experiments or cross-validation; is in the next state The maximum Q value of all possible actions is found through the Q value table; and is a nonlinear adjustment factor, adjusted through experiments or cross-validation; is the threshold, adjusted through experiments or cross-validation.

[0080] In the field of energy and power, the State Grid needs to conduct comprehensive analysis of a large amount of power equipment and user data to optimize power grid operation and improve service quality. These data include equipment status data, user power consumption data, meteorological data, etc., with diverse sources and complex structures; collect power equipment status data (such as transformers, transmission line status), user power consumption data (such as power consumption, power consumption time), meteorological data (such as temperature, humidity), etc. from different data sources; associate equipment, users and meteorological data to form a network containing nodes and edges. Nodes include equipment, users and weather stations, and edges represent the relationship between them; define the state space as a set of key data nodes and their combination patterns, and actions are selection strategies for key data nodes in the space; initialize a Q value table, in which the initial Q value of each state-action pair is set to zero or a random value; based on the initial Q value table, starting from the initial state, select an action, observe environmental feedback after executing the action, and get an immediate reward. ; Use the immediate reward formula to calculate the immediate reward obtained after performing an action For example, select a key data node in the initial state, observe the changes in the power grid operation state after execution, and get an immediate reward ; Instant rewards calculated based on , use the Q value update formula to update the Q value of the current state-action pair, so that the Q value gradually approaches the long-term expected return after adopting the selection strategy. For example, using the Q value update formula:

[0081] ;

[0082] Through multiple iterative learning, the process of repeatedly selecting actions, executing actions, observing feedback, and updating Q values ​​is repeated, and the Q value of each state-action pair is continuously adjusted until the Q value table converges, that is, the change of Q value tends to be stable; through the Q-learning algorithm, an optimized data aggregation analysis model is finally generated. The model can select the optimal key data nodes according to the current state, thereby improving the efficiency and accuracy of data aggregation analysis; the optimized model is applied to actual power grid operation, for example, the most suitable transformer is selected for monitoring according to the current power grid state, so as to optimize power grid operation and improve service quality.

[0083] Through the Q-learning algorithm, the selection strategy of key data nodes is dynamically adjusted to improve the efficiency and accuracy of data aggregation analysis; through continuous iterative learning, the model gradually learns more effective strategies until it converges to the optimal strategy, and finally obtains the optimized data aggregation analysis model; the optimized model provides more scientific decision-making support for the State Grid, helps optimize grid operation, and improves service quality.

[0084] 105. Based on the optimized data aggregation analysis model, the business data is comprehensively analyzed and processed to generate a comprehensive business insight report.

[0085] Utilize the optimized data aggregation analysis model to comprehensively analyze business data and extract valuable information and patterns; prepare a report that includes a comprehensive description of the current business status, as well as forecasts and suggestions for future business development; Utilize the optimized data aggregation analysis model to comprehensively analyze business data and extract valuable information and patterns; prepare a comprehensive business insight report that includes a comprehensive description of the current business status, business trend analysis, identification of potential risks and opportunities, as well as future forecasts and suggestions; provide scientific decision-making support for business managers through comprehensive business insight reports to help them better understand the current business status, grasp market dynamics, and formulate effective strategic plans.

[0086] Optionally, the step 105 of performing comprehensive analysis and processing on the business data based on the optimized data aggregation analysis model to generate a comprehensive business insight report includes:

[0087] Using the optimized data aggregation analysis model, a comprehensive analysis is performed on the multi-dimensional business data streams received from multiple heterogeneous data sources to obtain valuable information and patterns;

[0088] Based on the valuable information and patterns, the potential rules and trends in the business process are deeply mined and processed to generate business insight content;

[0089] Generate a comprehensive business insight report based on the business insight content, combined with a comprehensive description of the current business situation and predictions and recommendations for future business development.

[0090] In the field of energy and power, the State Grid needs to conduct comprehensive analysis on a large amount of power equipment and user data to optimize the operation of the power grid and improve the service quality. These data include equipment status data, user power consumption data, meteorological data, etc., with diverse sources and complex structures; collect power equipment status data (such as transformers, transmission line status), user power consumption data (such as power consumption, power consumption time), meteorological data (such as temperature, humidity), etc. from different data sources; use the optimized data aggregation analysis model to conduct comprehensive analysis and processing on these multi-dimensional business data streams. For example, analyze the operating status of transformers, the changing trend of user power consumption, the impact of meteorological conditions on power grid operation, etc.; extract valuable information and patterns from the analysis results, such as the operating status of key equipment, the behavior patterns of high-power users, the impact of meteorological conditions on power grid load, etc.; based on the extracted valuable information and patterns, conduct in-depth mining and processing of potential laws and trends in business processes. For example, discover the failure modes of certain key equipment, the seasonal changes in user power consumption, the impact of extreme weather on power grid operation, etc.; analyze the time characteristics and causal relationships of business data through time series analysis and causal inference technology to generate business insight content. For example, predicting electricity demand in the future, identifying potential risks of equipment failure, etc.; based on the business insights and a comprehensive description of the current business situation, write the first part of the report. For example, describe the current operating status of the power grid, the health of key equipment, user electricity behavior, etc.; write the second part of the report to analyze the trends and patterns of business data in detail. For example, analyze the seasonal changes in electricity consumption, the cyclical characteristics of equipment failures, etc.; write the third part of the report to identify potential risks and opportunities in business processes. For example, identify high-risk equipment, potential energy-saving opportunities, new market growth points, etc.; write the fourth part of the report to propose forecasts and suggestions for future business development. For example, predict the growth trend of future electricity demand, recommend measures to optimize power grid operation, and propose new business development directions.

[0091] Through the optimized data aggregation analysis model, the multi-dimensional business data flow is comprehensively analyzed and processed to extract valuable information and patterns, providing a basis for subsequent business insights; through in-depth mining of business data, the potential rules and trends in the business process are discovered, and detailed business insights are provided to help companies better understand the current business status; a comprehensive business insight report is compiled, the report content includes a comprehensive description of the current business status, business trend analysis, identification of potential risks and opportunities, and future forecasts and suggestions, providing scientific decision-making support for corporate managers; through comprehensive business insight reports, corporate managers are helped to better understand the current business status, grasp market dynamics, formulate effective strategic plans, and improve the competitiveness and service quality of the company; the operating status data of equipment such as transformers and transmission lines are collected; the data such as users' electricity consumption and electricity consumption time are collected; meteorological data such as temperature and humidity are collected; the above data are comprehensively analyzed and processed using the optimized data aggregation analysis model. For example, analyze the operating status of transformers, the changing trend of user power consumption, the impact of meteorological conditions on power grid operation, etc.; extract valuable information and patterns from the analysis results, such as the operating status of key equipment, the behavior patterns of high-power users, the impact of meteorological conditions on power grid load, etc.; based on the extracted valuable information and patterns, conduct in-depth mining and processing of potential laws and trends in business processes. For example, discover the failure mode of certain key equipment, the seasonal change law of user power consumption, the impact of extreme weather on power grid operation, etc.; analyze the time characteristics and causal relationships of business data through time series analysis and causal inference technology to generate business insight content. For example, predict power demand in the future, identify potential equipment failure risks, etc.; write the first part of the report to describe the current operating status of the power grid, the health status of key equipment, and user power consumption behavior. For example: the current power grid is operating smoothly, and key equipment (such as transformer T1 and transmission line L1) is in good operating condition; power users are mainly concentrated in industrial areas, and power consumption increases significantly during peak hours in summer; write the second part of the report to analyze the trends and laws of business data in detail. For example: user electricity consumption shows obvious peaks in summer and winter, and is relatively stable in spring and autumn; the failure rate of key equipment increases significantly under high temperature and high humidity conditions; write the third part of the report to identify potential risks and opportunities in business processes. For example: warm weather may cause key equipment to overheat, so it is recommended to strengthen equipment maintenance and monitoring; emerging new energy projects (such as photovoltaic power stations) provide new market growth points, so it is recommended to increase investment; write the fourth part of the report to put forward forecasts and suggestions for future business development. For example: it is estimated that in the next five years, electricity demand will maintain an average annual growth rate of 5%, so it is recommended to optimize the layout of the power grid and improve power supply capacity; it is recommended to introduce advanced data analysis technology to further improve the intelligence level of power grid operation and improve service quality.

[0092] Through the optimized data aggregation analysis model, we can comprehensively analyze multi-dimensional business data and extract valuable information and patterns; through in-depth mining of business data, we can discover potential rules and trends in business processes and provide detailed business insights; and compile comprehensive business insight reports to provide scientific decision-making support for enterprise managers. The report content includes a comprehensive description of the current business situation, business trend analysis, identification of potential risks and opportunities, and future forecasts and recommendations.

[0093] Figure 2 A schematic diagram of the structure of a big data aggregation analysis system based on a typical business scenario is provided for the embodiment of the present application. Figure 2 As shown, the device comprises:

[0094] A receiving module 21 is used to receive multi-dimensional business data streams from multiple heterogeneous data sources;

[0095] Processing module 22, used to utilize the multi-dimensional business data stream, combined with time series analysis and causal inference technology, to conduct in-depth mining and processing of the time characteristics and causal relationships of the data, and obtain a cross-domain data association network;

[0096] A learning module 23 is used to perform deep learning processing on the features of nodes and edges in the network based on the cross-domain data association network using a graph neural network algorithm, introduce an attention mechanism to highlight the influence of important nodes, and obtain key data nodes and their combination patterns;

[0097] The optimization module 24 is used to optimize the data aggregation analysis process according to the key data nodes and their combination patterns by using a reinforcement learning mechanism, and dynamically adjust the selection strategy of the key data nodes by using a Q-learning algorithm to obtain an optimized data aggregation analysis model;

[0098] The generation module 25 is used to perform comprehensive analysis and processing on the business data based on the optimized data aggregation analysis model to generate a comprehensive business insight report.

[0099] Figure 2 The big data aggregation analysis system based on typical business scenarios can be executed Figure 1 The implementation principle and technical effect of the big data aggregation analysis method based on a typical business scenario described in the embodiment shown will not be repeated. The specific way in which each module and unit performs operations in the big data aggregation analysis system based on a typical business scenario in the above embodiment has been described in detail in the embodiment of the method, and will not be elaborated here.

[0100] Through the description of the above implementation methods, those skilled in the art can clearly understand that each implementation method can be implemented by means of software plus a necessary general hardware platform, and of course, can also be implemented by hardware. Based on this understanding, the above technical solution is essentially or the part that contributes to the prior art can be embodied in the form of a software product, and the computer software product can be stored in a computer-readable storage medium, such as ROM / RAM, a disk, an optical disk, etc., including a number of instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to execute the methods described in each embodiment or some parts of the embodiments.

[0101] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present application, rather than to limit it. Although the present application has been described in detail with reference to the aforementioned embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the aforementioned embodiments, or make equivalent replacements for some of the technical features therein. However, these modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present application.

Claims

1. A big data aggregation analysis method based on a typical business scenario, characterized in that: include: Receiving a multi-dimensional business data stream from multiple heterogeneous data sources, wherein the multi-dimensional business data stream includes power equipment status data, user power consumption data, and meteorological data; By using the multi-dimensional business data stream, combined with time series analysis and causal inference technology, the time characteristics and causal relationships of the data are deeply mined and processed to obtain a cross-domain data association network, wherein the cross-domain data association network refers to associating the power equipment status data, the user power consumption data, and the meteorological data to form a network including nodes and edges; Based on the cross-domain data association network, a graph neural network algorithm is used to perform deep learning processing on the features of nodes and edges in the network, and a deep feature representation of each node is obtained through multi-layer nonlinear transformation. An attention mechanism is introduced to highlight the influence of important nodes, and key data nodes and their combination patterns are obtained. The process of introducing the attention mechanism to highlight the influence of important nodes includes: using the deep feature representation of each node, combining the self-attention mechanism, evaluating the strength of the relationship between each node and its neighboring nodes, and dynamically adjusting the attention weight by calculating the similarity score of the features between the nodes to obtain the attention weight; The process of generating attention weights includes: In order to more accurately reflect the relationship strength between nodes, the feature similarity score, relationship type weight and distance between nodes are comprehensively considered, and after nonlinear adjustment and normalization, the attention weight of each node to its neighbor node is generated. : ; in: Representation Node Its neighbor nodes The feature similarity score between them; Representation Node Its neighbor nodes The feature similarity score between ; Representation Node With Node The distance between Representation Node and The distance between Representation Node With Node The weight of the relationship type between them; and is an adjustment factor used to control the influence of relationship type weight and distance on the similarity score; and It is a nonlinear adjustment factor, which is used to introduce nonlinear terms and enhance the flexibility of the model; is a threshold used to control the activation condition of the nonlinear term; Representation Node The set of all neighbor nodes of is a natural exponential function, which is used to convert the adjusted similarity score into a positive value for normalization; is a nonlinear term used to balance the influence of similarity scores; according to the key data nodes and their combination patterns, the data aggregation analysis process is optimized using the reinforcement learning mechanism, and the selection strategy of the key data nodes is dynamically adjusted through the Q-learning algorithm to obtain the optimized data aggregation analysis model; Based on the optimized data aggregation analysis model, the business data is comprehensively analyzed and processed to generate a comprehensive business insight report.

2. The method according to claim 1, characterized in that Based on the cross-domain data association network, the graph neural network algorithm is used to perform deep learning processing on the features of nodes and edges in the network, and the attention mechanism is introduced to highlight the influence of important nodes, so as to obtain key data nodes and their combination patterns, including: Based on the cross-domain data association network, deep learning processing is performed on the features of nodes and edges in the network, and a deep feature representation of each node is obtained through multi-layer nonlinear transformation, while retaining the topological structure information between nodes to obtain a deep feature representation of the node; By using the deep feature representation of the node and combining it with the self-attention mechanism, the strength of the relationship between each node and its neighboring nodes is evaluated, and the attention weight is dynamically adjusted by calculating the similarity score of the features between the nodes to obtain the attention weight; According to the attention weights, a node importance scoring algorithm is used to evaluate the importance of nodes in the network, and key data nodes that play a core role in the network are screened out to obtain key data nodes; Based on the key data nodes, the interaction patterns between the nodes are analyzed to generate a combination pattern of the key data nodes that reflects the business logic.

3. The method according to claim 2, characterized in that The deep feature representation of the node is used in combination with the self-attention mechanism to evaluate the strength of the relationship between each node and its neighboring nodes, and the attention weight is dynamically adjusted by calculating the similarity score of the features between the nodes to obtain the attention weight, including: Using the deep feature representation of the node and combining it with the self-attention mechanism, the feature similarity between each node and all its neighboring nodes is calculated to obtain a similarity score; According to the similarity score, normalization is performed through a softmax function to generate an attention weight; Based on the attention weight, factors of distance between nodes or relationship type are introduced as additional inputs to refine the evaluation of the strength of the relationship between nodes and obtain more accurate attention weights; Using more accurate attention weights, the feature representations of the nodes are weighted and summed to generate the final attention weights.

4. The method according to claim 2, characterized in that: According to the attention weight, a node importance scoring algorithm is used to evaluate the importance of nodes in the network, and key data nodes that play a core role in the network are screened out to obtain key data nodes, including: Using the attention weight, each node in the network is evaluated for importance using a node importance scoring algorithm to obtain an importance score value; According to the importance score value, the nodes in the network are sorted to generate an importance sorted list of the nodes; Based on the importance ranking list of the nodes, a threshold is set or the top N nodes are selected to screen the nodes in the network to obtain the key data nodes that play a core role in the network; By analyzing the key data nodes that play a core role in the network, the actual role and influence of these nodes in the network are confirmed, and key data nodes are generated.

5. The method according to claim 1, characterized in that According to the key data nodes and their combination patterns, the data aggregation analysis process is optimized by using a reinforcement learning mechanism, and the selection strategy of the key data nodes is dynamically adjusted by a Q-learning algorithm to obtain an optimized data aggregation analysis model, including: Using the key data nodes and their combination patterns, a reinforcement learning framework is constructed, the state space is defined as a set of key data nodes and their combination patterns, and the action is a selection strategy of the key data nodes in the space, thereby obtaining the reinforcement learning framework; Based on the reinforcement learning framework, the data aggregation analysis process is optimized using the Q-learning algorithm. Through multiple iterative learning, the Q value of each state-action pair is updated. The Q value reflects the expected return after taking a certain selection strategy, and an updated Q value table is obtained. In each iteration, according to the current state and the updated Q-value table, the action with the highest Q-value is selected as the selection strategy of the key data node, and the environmental feedback is observed after execution to obtain an immediate reward; Through a continuous iterative learning process, the selection strategy of key data nodes is dynamically adjusted based on the instant rewards, so that the model can gradually learn more effective strategies until it converges to the optimal strategy, and finally obtains an optimized data aggregation analysis model.

6. The method according to claim 5, characterized in that Based on the reinforcement learning framework, the data aggregation analysis process is optimized using the Q-learning algorithm. Through multiple iterative learning, the Q value of each state-action pair is updated. The Q value reflects the expected return after taking a certain selection strategy, and an updated Q value table is obtained, including: Using the reinforcement learning framework, a Q-value table is initialized, wherein the initial Q-value of each state-action pair is set to zero or a random value, to obtain an initial Q-value table; Based on the initial Q value table, using the Q-learning algorithm, starting from the initial state, selecting an action, observing environmental feedback after executing the action, and obtaining an immediate reward; According to the instant reward and the Q value of the next state, the Q value of the current state-action pair is updated so that the Q value gradually approaches the long-term expected return after adopting the selection strategy, thereby obtaining an updated Q value; Through multiple iterative learning, the process of repeatedly selecting actions, executing actions, observing feedback, and updating Q values ​​is repeated, and the Q value of each state-action pair is continuously adjusted until the Q value table converges, that is, the change of Q value tends to be stable, and an updated Q value table is generated.

7. The method according to claim 6, characterized in that Based on the initial Q value table, using the Q-learning algorithm to start from the initial state, selecting an action, observing environmental feedback after executing the action, and obtaining an immediate reward, includes: To calculate the immediate reward obtained after performing an action, the following formula is used: ; in, Indicates that at time step Instant rewards when Indicates in status Next action Rewards immediately received after is the learning rate, which controls the speed of Q value update; Indicates in status Next action The current Q value of Indicates execution of an action The next state reached after Indicates that in the next state The maximum Q value of all possible actions; and It is a nonlinear adjustment factor, which is used to introduce nonlinear terms and enhance the flexibility of the model; is a threshold used to control the activation condition of the nonlinear term; According to the calculated instant reward, the Q value of the current state-action pair is updated using the following formula: The value gradually approaches the long-term expected return after adopting this selection strategy: ; in, Indicates in status Next action Current value; Indicates that at time step Instant rewards when is a discount factor that controls the weight of future rewards; Indicates that in the next state The maximum Q value of all possible actions; and It is a nonlinear adjustment factor, which is used to introduce nonlinear terms and enhance the flexibility of the model; is a threshold used to control the activation condition of the nonlinear term; Take advantage of instant rewards Update the Q-value of the current state-action pair to gradually optimize the key data node selection strategy in the data aggregation analysis process.

8. The method according to claim 1, characterized in that The process of comprehensively analyzing and processing the business data based on the optimized data aggregation analysis model to generate a comprehensive business insight report includes: Using the optimized data aggregation analysis model, a comprehensive analysis is performed on the multi-dimensional business data streams received from multiple heterogeneous data sources to obtain valuable information and patterns; Based on the valuable information and patterns, the potential rules and trends in the business process are deeply mined and processed to generate business insight content; Generate a comprehensive business insight report based on the business insight content, combined with a comprehensive description of the current business situation and predictions and recommendations for future business development.

9. A big data aggregation analysis system based on typical business scenarios, characterized in that: include: A receiving module, used to receive multi-dimensional business data streams from multiple heterogeneous data sources, wherein the multi-dimensional business data streams include power equipment status data, user power consumption data, and meteorological data; A processing module, used to utilize the multi-dimensional business data stream, in combination with time series analysis and causal inference technology, to perform in-depth mining and processing on the time characteristics and causal relationships of the data, and obtain a cross-domain data association network, wherein the cross-domain data association network refers to associating the power equipment status data, the user power consumption data, and the meteorological data to form a network including nodes and edges; A learning module is used to perform deep learning processing on the features of nodes and edges in the network based on the cross-domain data association network using a graph neural network algorithm, obtain a deep feature representation of each node through multi-layer nonlinear transformation, introduce an attention mechanism to highlight the influence of important nodes, and obtain key data nodes and their combination patterns; The process of introducing the attention mechanism to highlight the influence of important nodes includes: using the deep feature representation of each node, combining the self-attention mechanism, evaluating the strength of the relationship between each node and its neighboring nodes, and dynamically adjusting the attention weight by calculating the similarity score of the features between the nodes to obtain the attention weight; The process of generating attention weights includes: In order to more accurately reflect the relationship strength between nodes, the feature similarity score, relationship type weight and distance between nodes are comprehensively considered, and after nonlinear adjustment and normalization, the attention weight of each node to its neighbor node is generated. : ; in: Representation Node Its neighbor nodes The feature similarity score between them; Representation Node Its neighbor nodes The feature similarity score between ; Representation Node With Node The distance between Representation Node and The distance between Representation Node With Node The weight of the relationship type between them; and is an adjustment factor used to control the influence of relationship type weight and distance on the similarity score; and It is a nonlinear adjustment factor, which is used to introduce nonlinear terms and enhance the flexibility of the model; is a threshold used to control the activation condition of the nonlinear term; Representation Node The set of all neighbor nodes of is a natural exponential function, which is used to convert the adjusted similarity score into a positive value for normalization; is a nonlinear term used to balance the impact of similarity scores; An optimization module is used to optimize the data aggregation analysis process according to the key data nodes and their combination patterns by using a reinforcement learning mechanism, dynamically adjust the selection strategy of the key data nodes by using a Q-learning algorithm, and obtain an optimized data aggregation analysis model; The generation module is used to perform comprehensive analysis and processing on the business data based on the optimized data aggregation analysis model to generate a comprehensive business insight report.

Citation Information

Patent Citations

  • Energy storage power station evaluation method and system oriented to source network load multivariate application

    CN118484666A