Execution decision and optimization method and device based on graph structure, equipment and medium

By converting the original state information into graph structure data and using graph neural network to process inter-entity relationships, generating enhanced state representations and updating graph structures, the problem of insufficient modeling of inter-entity relationships in the prior art is solved, and the accuracy of state representation and the decision-making ability of reinforcement learning models are improved.

CN120509437APending Publication Date: 2025-08-19PING AN TECH (SHENZHEN) CO LTD
View PDF 0 Cites 1 Cited by

Patent Information

Application Number
CN202510643853.5
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-19
Publication Date
2025-08-19

AI Technical Summary

Technical Problem

Existing reinforcement learning methods fail to effectively capture and model complex relationships and multi-level interactive information between entities, resulting in redundant state representations or missing critical information, especially in the fields of fintech and health care, affecting the accuracy and adaptability of decisions.

Method used

The original state information is converted into graph structure data, the relationship between entities is processed through the graph neural network, the enhanced state representation is generated, and the nodes and edges in the graph structure data are updated based on the feedback information, and the graph neural network is used to aggregate multi-hop neighborhood information for decision-making.

Benefits of technology

It improves the accuracy of state representation and intelligence of processing, optimizes the performance and generalization capabilities of reinforcement learning models, especially decision-making effects in complex and dynamic environments.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120509437A_ABST
    Figure CN120509437A_ABST
Patent Text Reader

Abstract

The invention relates to the technical field of artificial intelligence, can be applied to business scenes of financial science and technology, medical health and the like, and discloses an execution decision and optimization method and device based on a graph structure, equipment and a medium. The graph structure data comprises a plurality of nodes representing entities and one or more edges representing relationships among the entities, processing the graph structure data through a graph neural network, aggregating multi-hop neighborhood information of the nodes based on the relationships among the entities defined by the edges, generating enhanced state representation, determining an action to be executed based on the enhanced state representation, and executing the action to be executed based on the action to be executed. And updating nodes and edges in the graph structure data according to feedback information triggered by the execution action. According to the method, state representation is modeled through the graph neural network, the multilevel relation between entities in the environment is effectively captured, the problem of state representation redundancy or key information missing is solved, and therefore the accuracy of state representation and the intelligence of processing are improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a graph-based execution decision and optimization method, device, equipment and storage medium. Background Art

[0002] In the existing field of reinforcement learning (RL), traditional neural network methods, such as Deep Q-Network (DQN) and Asynchronous Advantage Actor-Critic (A3C), typically treat states as sets of independent features, such as pixel matrices or numerical vectors. These methods extract features through fully connected or convolutional neural networks, but this approach ignores the structured relationships between entities in the environment, such as spatial location and the interaction logic between entities. This simplified approach leads to redundant state representations or fails to capture key environmental information. This is particularly true when dealing with complex tasks, as it lacks sufficient modeling of the relationships between entities in the environment.

[0003] In the fintech sector, existing reinforcement learning methods typically focus on treating market data, trading behavior, or user information as independent features. However, this approach fails to fully consider the complex relationships between various market entities, such as the interactions between market participants, trends in capital flows, and potential connections between macroeconomic indicators. Traditional methods are unable to effectively capture these multi-layered relationships, resulting in redundant state representations or omissions of key factors. Therefore, when faced with the rapidly changing and dynamically interactive financial markets, existing technologies have significant deficiencies in the completeness and accuracy of state representation. The limitations of traditional methods are particularly prominent in scenarios such as high-frequency trading and risk management, which require real-time response to market fluctuations.

[0004] In the healthcare business, traditional reinforcement learning methods typically model patient health data, medical records, and other features as independent features. However, complex relationships exist between patients, between doctors and patients, and between medical resources, and these relationships have a significant impact on treatment outcomes and health management. Existing technologies fail to fully utilize these complex relationships, resulting in a lack of necessary contextual information and interactive relationship modeling in state representation in many cases. This is particularly true in scenarios such as personalized treatment and cross-departmental collaboration, where the limitations of existing methods become more pronounced. Although some methods have attempted to apply graph neural networks to multi-agent system modeling in the medical field, they often focus on optimizing diagnosis or treatment recommendations, and fail to effectively integrate these complex relationships into the state representation process. Therefore, existing technologies fail to fully explore the complex relationships behind the data when processing healthcare data, resulting in limited effectiveness in practical applications.

[0005] Existing methods often rely on manually designed rules to supplement relational information in the environment, especially in dynamic and complex environments. Although these methods can work in some scenarios, they fail to effectively address the impact of multi-level relationships and dynamic changes on reinforcement learning state representation, resulting in a lack of sufficient adaptability of the model in the face of ever-changing environments. Summary of the Invention

[0006] The main purpose of the present invention is to provide a graph-based execution decision and optimization method, device, equipment and storage medium, aiming to solve the technical problem that the existing technology fails to effectively capture and model the complex relationships and multi-level interactive information between entities, resulting in redundancy in state representation or missing key information.

[0007] To achieve the above objectives, the present invention provides a graph-based execution decision and optimization method, comprising:

[0008] Acquire original state information and convert the original state information into graph structure data, wherein the graph structure data includes a plurality of nodes representing entities and one or more edges representing relationships between the entities;

[0009] Processing the graph structure data through a graph neural network, aggregating multi-hop neighborhood information of the node based on the inter-entity relationships defined by the edges, and generating an enhanced state representation;

[0010] determining an action to be performed based on the enhanced state representation;

[0011] The nodes and edges in the graph structure data are updated according to the feedback information triggered when the execution action is executed.

[0012] Furthermore, to achieve the above-mentioned objectives, the present invention provides an execution decision and optimization device based on a graph structure, comprising:

[0013] A graph structure building module, configured to obtain original state information and convert the original state information into graph structure data, wherein the graph structure data includes a plurality of nodes representing entities and one or more edges representing relationships between the entities;

[0014] A graph neural network processing module, configured to process the graph structure data through a graph neural network, aggregate multi-hop neighborhood information of the nodes based on the inter-entity relationships defined by the edges, and generate an enhanced state representation;

[0015] an action determination module, configured to determine an action to be performed based on the enhanced state representation;

[0016] A graph updating module is used to update the nodes and edges in the graph structure data according to feedback information triggered when the execution action is executed.

[0017] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer device, which includes a memory, a processor, and a graph-based execution decision and optimization program stored in the memory and run on the processor. When the graph-based execution decision and optimization program is executed by the processor, the steps of the graph-based execution decision and optimization method described above are implemented.

[0018] Furthermore, to achieve the above-mentioned purpose, the present invention also provides a computer-readable storage medium, on which a graph-structure-based execution decision and optimization program is stored. When the graph-structure-based execution decision and optimization program is executed by a processor, the steps of the graph-structure-based execution decision and optimization method as described above are implemented.

[0019] Beneficial effects: The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as financial technology and medical health. It discloses a graph-based execution decision and optimization method, device, equipment and medium, including: obtaining original state information, converting the original state information into graph structure data, wherein the graph structure data includes multiple nodes representing entities and one or more edges representing relationships between entities, processing the graph structure data through a graph neural network, aggregating multi-hop neighborhood information of nodes based on the relationships between entities defined by the edges, generating an enhanced state representation, determining the action to be executed based on the enhanced state representation, and updating the nodes and edges in the graph structure data according to the feedback information triggered by the execution of the action. The present invention models the state representation through a graph neural network, effectively capturing the multi-level relationships between entities in the environment, overcoming the problems of redundant state representation or missing key information, thereby improving the accuracy of the state representation and the intelligence of the processing, and optimizing the performance and generalization ability of the reinforcement learning model. BRIEF DESCRIPTION OF THE DRAWINGS

[0020] The present invention will be further described below with reference to the accompanying drawings and embodiments, in which:

[0021] Figure 1 A schematic diagram of an application environment of an execution decision and optimization method based on a graph structure in one embodiment of the present invention;

[0022] Figure 2 This is a flow chart of an embodiment of a graph-based execution decision and optimization method of the present invention;

[0023] Figure 3 A schematic diagram of functional modules of a preferred embodiment of the graph-based execution decision and optimization device of the present invention;

[0024] Figure 4 A schematic diagram of the structure of a computer device according to an embodiment of the present invention;

[0025] Figure 5FIG. 2 is another structural diagram of a computer device according to an embodiment of the present invention. DETAILED DESCRIPTION

[0026] It should be understood that the specific embodiments described herein are only used to explain the present invention and are not intended to limit the present invention.

[0027] The graph-based execution decision and optimization method provided by the embodiment of the present invention can be applied in Figure 1 In an application environment, the user terminal communicates with the server terminal through a network. The server terminal can obtain the original state information through the user terminal, convert the original state information into graph structure data, and the graph structure data includes multiple nodes representing entities and one or more edges representing the relationship between entities. The graph structure data is processed by a graph neural network, and the multi-hop neighborhood information of the nodes is aggregated based on the relationship between entities defined by the edges to generate an enhanced state representation. The action to be executed is determined based on the enhanced state representation, and the nodes and edges in the graph structure data are updated according to the feedback information triggered by the execution of the action. The present invention models the state representation through a graph neural network, effectively captures the multi-level relationship between entities in the environment, overcomes the problem of redundant state representation or missing key information, thereby improving the accuracy of state representation and the intelligence of processing, and optimizing the performance and generalization ability of the reinforcement learning model. Among them, the user terminal can be but is not limited to various personal computers, laptops, smart phones, tablet computers and portable wearable devices. The server terminal can be implemented with an independent server or a server cluster consisting of multiple servers. The present invention is described in detail below through specific embodiments.

[0028] See also Figure 2 , Figure 2 This is a flow chart of an embodiment of the graph-based execution decision and optimization method provided by the present invention. It should be noted that although the flow chart shows a logical order, in some cases, the steps shown or described may be performed in a different order than here.

[0029] like Figure 2 As shown, the graph-based execution decision and optimization method proposed in the present invention includes the following steps:

[0030] S10, acquiring original state information, and converting the original state information into graph structure data, wherein the graph structure data includes a plurality of nodes representing entities and one or more edges representing relationships between the entities;

[0031] In this embodiment, the purpose of obtaining raw state information is to extract the most basic, unprocessed state data from the environment. These raw state information are usually feature representations of entities in the environment and have not undergone any preprocessing or transformation so that they can be further analyzed and processed in subsequent steps. The way to obtain raw state information depends on the specific application field and environment settings, and usually includes sensor data, user input, transaction records, etc. These data contain preliminary information about the environment that the system needs to understand. In the financial field, this may include the user's account balance, transaction records, credit score, etc., while in the medical and health field, the raw state information may include the patient's physiological data, medical records, disease history, etc. Whether in the field of financial technology or medical and health, this raw state information provides the basis for subsequent model training and decision making.

[0032] This can be achieved by directly acquiring data from the environment, typically stored in sensors, databases, or other external systems. To extract raw state information from these sources, the system typically integrates and collects data through APIs, database queries, or real-time data streaming technologies. This step does not involve complex data processing or feature extraction; it simply acquires the data, ensuring that all important environmental variables are recorded for subsequent processing and analysis.

[0033] In practical applications, the technical implementation methods for obtaining raw state information may vary. In the financial sector, a system may need to obtain customer transaction data, account balances, and other information from a bank database. This data is typically stored in a structured database, and the system accesses it in real time through SQL queries or API interfaces. For example, a financial institution could use a RESTful API interface to obtain transaction records and account balances from a user's account system and use this information as raw state data to input into a reinforcement learning system.

[0034] In the healthcare sector, raw status information is typically obtained from the hospital's electronic health record (EHR) system. This data includes patient diagnosis and treatment information, historical medical records, and test results. By integrating with the hospital information system (HIS) or other medical data platforms, the system can automatically obtain real-time patient vital signs such as blood pressure, heart rate, and body temperature as raw status information input. To ensure the accuracy and timeliness of this data, the system often needs to connect with the hospital's real-time monitoring equipment to ensure that this raw status information can be efficiently obtained.

[0035] Example: In the financial sector, suppose a system needs to analyze a customer's behavior and predict their credit risk. When acquiring raw state information, the system will obtain the customer's account information, transaction history, credit score, and other data in real time. This raw state information provides the basis for the system's subsequent decision-making process. In subsequent steps, the system will convert this data into graph-structured data and process it using a graph neural network to obtain a credit risk prediction for the customer.

[0036] In the healthcare sector, a system may need to predict changes in a patient's condition by monitoring their vital signs in real time (such as blood sugar levels, heart rate, and body temperature). When acquiring raw status information, the system extracts real-time patient data from monitoring devices or electronic health record systems, providing a foundation for subsequent graph structure conversion and decision-making. This data is further processed to help the system identify patients' health risks and provide optimized medical plans.

[0037] Obtaining raw state information provides fundamental data support for the subsequent reinforcement learning process. By ensuring real-time and accurate acquisition of raw data from the environment, the system can provide a reliable data source for subsequent steps such as graph structure conversion, state representation generation, and action decision-making. This process not only provides the reinforcement learning model with a comprehensive perspective of the environment but also provides the necessary input for subsequent complex operations (such as graph neural network processing, feature extraction, and reward signal analysis), ensuring the model's accurate environmental perception during training and application.

[0038] The raw state information is converted into graph-structured data. Graph-structured data consists of multiple nodes and edges, where nodes represent entities and edges represent relationships between entities. This step transforms the raw data into a graph structure, facilitating subsequent deep learning and processing by graph neural networks.

[0039] First, the system identifies entity objects from the raw state information. Entity objects can be any data unit with independent existence and meaning. For example, in the financial field, they may be users, accounts, transactions, etc., and in the medical and health field, they may be patients, doctors, medical records, etc. These entity objects will be mapped as nodes in the graph. Each node represents an independent entity, and the node attributes are defined by the specific characteristics of the entity. The attributes of the node may include but are not limited to the basic information of the entity (such as account balance, patient's medical history, drug treatment plan, etc.). These attributes will be used as input for calculation and processing in the graph neural network.

[0040] Next, based on the relationships between entities in the original data, edges are generated to represent the interactions between entities. The construction of edges is key to the graph structure, as edges represent interactions, connections, or dependencies between entities. In finance, edges might represent fund transfers or transactions between accounts; in healthcare, edges might represent therapeutic relationships between patients and doctors, or between patients and medications. The generation of these edges depends on the type of relationship between entities, typically based on spatial location, time series relationships, or interaction logic. For example, when a funds transfer occurs between two users, the two accounts in the graph will be connected by an edge, which can be accompanied by additional information such as the transfer amount and transaction time.

[0041] When generating graph-structured data, edge definitions are not limited to simple relationships but may also include edge attributes. These attributes may be dynamically adjusted based on the needs of the application scenario. For example, in finance, edge attributes may include transaction amount and frequency; in healthcare, edge attributes may include treatment intensity and medication duration. These edge attributes will help graph neural networks better capture subtle relationships between nodes in subsequent processing stages.

[0042] This conversion process not only transforms data from raw information into a graph structure but also provides structured input for subsequent graph neural network operations. The introduction of a graph structure enables the model to effectively process and analyze complex relationships between entities. Unlike traditional flat data processing methods, graph structures preserve the relative relationships and structural information between entities. This advantage is particularly evident when dealing with tasks involving multiple entities and complex interactions.

[0043] In the financial sector, graph-structured data can be constructed from customer account information, transaction history, and fund flows within banking systems. For example, each customer account is a node, and fund transfers between accounts are represented by edges. Node attributes can include account balances and financial products held, while edge attributes can represent transaction amounts and types. Once the graph structure is constructed, graph neural networks can analyze fund flow patterns and identify potential risks or investment opportunities.

[0044] In the healthcare field, graph-structured data can be constructed based on entities such as patients, doctors, and medications in electronic health records (EHRs). Each patient, doctor, medication, or treatment plan can be represented as a node, and the relationships between them (for example, the patient-doctor relationship, the patient-drug treatment relationship) are represented by edges. Node attributes can include basic patient information, disease type, and doctor's professional background, while edge attributes can include treatment duration and medication dosage. By constructing this information into a graph structure, the system can conduct a more in-depth analysis of the patient's treatment process and optimize resource allocation and treatment plans.

[0045] When generating graph data, the selection of data sources is crucial. The system can collect the required raw data from various databases. Data sources include internal databases (such as financial transaction databases and medical records) and external data sources (such as public financial data and public medical databases). This data is cleaned, preprocessed, and standardized according to specific requirements to ensure data consistency and integrity. Then, graph construction algorithms transform the data into a graph structure, forming a structured input that meets the processing requirements of neural networks.

[0046] This embodiment effectively provides structured input data for subsequent graph neural network learning by converting raw state information into graph-structured data. This conversion not only simplifies information processing but also enhances the system's ability to model complex relationships. Graph-structured data provides a more efficient information flow for neural networks, enabling the system to capture richer contextual information when processing relationships between entities, improving the model's learning effectiveness and generalization capabilities.

[0047] S20, processing the graph structure data through a graph neural network, aggregating multi-hop neighborhood information of the node based on the inter-entity relationships defined by the edges, and generating an enhanced state representation;

[0048] In this embodiment, the purpose of processing graph-structured data with a graph neural network (GNN) is to aggregate multi-hop neighborhood information of nodes based on the relationships between entities defined by edges, thereby generating an enhanced state representation. By aggregating nodes and their neighborhoods, GNNs can propagate information within the network, allowing each node to obtain information from its adjacent nodes or multi-hop adjacent nodes, resulting in a richer feature representation.

[0049] First, a graph neural network takes graph structure data as input. The characteristics of each node in the graph (such as node attributes and edge attributes) are embedded in the graph neural network and participate in subsequent calculations. The relationships between nodes are defined by edges in the graph. Edge attributes can be distance, type, time, weight, etc., which are used to characterize the interactions between nodes. The graph neural network gradually updates the state of each node by iteratively aggregating information from these neighboring nodes.

[0050] Specifically, during multi-hop neighborhood aggregation, information from the first-hop neighborhood nodes is passed to the target node through an aggregation operation. The target node then continues to aggregate information based on the second-hop and higher-order neighborhood nodes until the preset number of hops is reached. In this way, graph neural networks can capture the relationships between nodes at greater distances, thereby generating a more comprehensive and accurate enhanced state representation.

[0051] The goal of enhanced state representation generation is to extract more contextual information through multi-hop relationships between nodes, providing richer feature input for subsequent tasks such as decision-making and prediction. Aggregating multi-hop neighborhood information ensures that each node not only relies on the information of its immediate neighbors but also obtains valuable information from more distant nodes, improving the model's expressive power.

[0052] Multi-hop neighborhood aggregation in graph neural networks can be implemented in several ways, and the specific implementation method usually depends on the graph neural network architecture used. For example, the hierarchical graph convolutional network (HGCN) can implement multi-hop neighborhood aggregation through a hierarchical structure. In the first layer, the graph neural network updates the representation of the node through the node's direct neighborhood (first hop); in the second layer, the node representation is further updated based on the aggregated information of the first-hop node and the features of the second-hop neighbors. Through this hierarchical aggregation mechanism, the network can capture deeper inter-node relationships layer by layer.

[0053] In the financial sector, if we need to analyze the correlations between different financial products, graph neural networks can use multi-hop neighborhood aggregation to obtain relevant market information from multiple levels of nodes. For example, the first-hop neighborhood nodes may be other financial products in the same product category, while the second-hop nodes may be other products in adjacent markets or the same industry. This allows for a comprehensive market analysis.

[0054] In healthcare, graph neural networks can be used to analyze the relationship between patients and medical resources. The first-hop neighborhood might be the doctor, medication, or treatment plan a patient has directly contacted, while the second-hop neighborhood might be other hospitals, doctors, or other patients with the same disease whom the patient has previously treated. By aggregating these multi-hop neighborhoods, the model can integrate medical information from multiple dimensions to help predict treatment outcomes or disease trends.

[0055] Example: In the financial sector, graph neural networks can be used to analyze multi-level capital flows. Each account is represented as a node, and the fund transfer relationships between accounts are represented as edges. In this graph structure, graph neural networks can capture indirect capital flows between accounts by aggregating information from multi-hop neighborhoods, such as related financial products and market behavior. This approach allows the system to identify potential patterns and risk points in capital flows, thereby optimizing investment strategies and risk management.

[0056] In healthcare, graph neural networks can be used to analyze the relationships between patients, doctors, drugs, and treatment plans. Each patient, doctor, or drug serves as a node, while the patient-doctor relationship and the treatment relationship serve as edges. Through multi-hop neighborhood aggregation, the system can capture indirect relationships within a patient's treatment process, such as the patient's indirect treatment history with different doctors, thereby providing doctors with more comprehensive treatment recommendations or resource optimization solutions.

[0057] This embodiment processes graph-structured data and aggregates multi-hop neighborhood information through a graph neural network, which can significantly improve the model's ability to capture complex relationships. Multi-hop neighborhood aggregation not only relies on the node's direct neighbor information, but also enriches the node's state representation with information from distant nodes. This approach can more comprehensively understand and model the dependencies between nodes, making the generated enhanced state representation more accurate and expressive, especially when dealing with dynamic and complex environments.

[0058] S30, determining an action to be performed based on the enhanced state representation;

[0059] In this embodiment, the process of determining the action to be performed based on the generated enhanced state representation mainly involves the decision generation module in reinforcement learning. The enhanced state representation aggregates multi-hop neighborhood information through a graph neural network, reflecting the comprehensive information of the current environment, including the relationship between entities, the characteristics of each node and their interaction methods. In reinforcement learning, the determination of the action to be performed usually depends on the current state representation and the assessment of the state by the reinforcement learning policy network. Specifically, the enhanced state representation is used as input, and the corresponding action output is generated through the policy network.

[0060] The enhanced state representation provides multidimensional contextual information for each node, including information about the node's immediate neighborhood and its connections to more distant nodes. This information helps make more reasonable decisions through dynamic environmental feedback. By inputting the enhanced state representation into a decision module (such as a policy network), the module outputs an optimal action based on the policy evaluation values obtained during training. Reinforcement learning algorithms (such as Q-learning and Actor-Critic) can continuously adjust their strategies based on the current state and historical feedback information to achieve the goal of optimizing long-term rewards.

[0061] In reinforcement learning, the actions to be taken are typically derived from the output of a policy network. For example, in Q-learning, a state-based Q-value function calculates the expected reward for each possible action, and the policy network uses these values to select an optimal action. In the actor-critic algorithm, the actor generates actions, while the critic evaluates the value of those actions. Based on the generated enhanced state representation, the policy network can make reasonable choices to maximize long-term reward.

[0062] In specific implementations, various policy network architectures can be used to determine the action to be executed. In a Deep Q-Network (DQN), the network inputs the augmented state representation into a deep neural network. The network outputs the Q value for each possible action, and the action with the highest Q value is selected as the action to be executed. In A3C (Asynchronous Advantage Actor-Critic), the actor generates an action based on the current state, while the critic evaluates the value of the action to ultimately determine the optimal action.

[0063] In finance, reinforcement learning uses state representation to inform trading decisions. For example, an enhanced state representation can be generated based on historical stock market data, correlations between stocks, and market dynamics. In this case, the reinforcement learning model selects the optimal buy, sell, or hold action based on the market state. Multi-hop neighborhood aggregation helps capture complex relationships in the stock market that are not easily observed directly, such as indirect influences between stocks and market volatility.

[0064] In healthcare, enhanced state representation can integrate information such as a patient's medical history, doctor's recommendations, and treatment history to help generate treatment actions. Policy networks can predict treatment outcomes based on a patient's health status and select the optimal treatment plan or medication. This enhanced state representation allows the system to provide personalized treatment recommendations based on the patient's overall condition and historical connections with other patients, improving treatment effectiveness.

[0065] This embodiment improves decision-making accuracy and adaptability by determining pending actions based on an enhanced state representation. The enhanced state representation aggregates information from multiple nodes and their multi-hop neighborhoods, providing a more comprehensive picture of the environment. Processing this rich state information through a reinforcement learning policy network allows for better response to complex dynamic environments, enabling more informed decisions and ultimately optimizing long-term benefits. This approach improves decision-making accuracy, particularly in multi-dimensional and complex tasks.

[0066] S40: Update the nodes and edges in the graph structure data according to the feedback information triggered when the execution action is executed.

[0067] In this embodiment, the nodes and edges in the graph structure data are updated using feedback information generated by executing actions. Feedback information in reinforcement learning typically refers to environmental feedback or reward signals associated with an action, which is crucial for the model's next decision. Updating nodes and edges is typically used to dynamically adjust the graph structure to accurately reflect changes in the current environment.

[0068] Nodes in graph-structured data represent entities, while edges represent relationships between entities. In many application scenarios, entities and their relationships change over time. This is particularly true in multi-agent systems, where interactions between agents influence their states and behaviors. When an action is performed, the system receives feedback, typically in the form of a reward signal. This reward signal can be positive or negative, depending on whether the action leads to the desired outcome.

[0069] Using this feedback, the system updates the nodes and edges in the graph structure data. Specifically, it adjusts the states of the corresponding nodes and edges in the graph based on the magnitude of the reward signal. First, by parsing the state change data in the feedback information, the system identifies which entities have been added, invalidated, or changed. Then, based on this change data, the system dynamically updates the nodes and edges in the graph. For example, invalid nodes and their associated edges are removed, while new nodes and their associated edges are added to the graph structure.

[0070] Edge updates involve adjusting edge weights or changing edge connections. In reinforcement learning, edge weights are often related to the frequency, importance, or other metrics of interactions between nodes. Based on feedback, edge weights may be increased or decreased to better represent the strength of the connection between nodes. For example, if a node frequently interacts with other nodes, the edge weight between them can be increased; conversely, if interactions are less frequent, the edge weight can be decreased.

[0071] In practical applications, implementing feedback updates may rely on a variety of models and strategies. In multi-agent systems, each agent's behavior is dynamically adjusted based on feedback from other agents. For example, in reinforcement learning, methods such as Q-learning or A3C use real-time feedback to update the states of nodes in a graph. These updates are typically based on reward signals associated with the current action, which are dynamically adjusted as the model learns to reflect the interactions between entities.

[0072] In graph neural networks (GNNs), node and edge state updates are typically implemented through a "message passing" mechanism. Each node updates its state based on information from neighboring nodes, and edge weights are adjusted based on this interaction information. In financial applications, the system can adjust edge weights between stocks based on market feedback to reflect market dynamics. For example, when the prices of two stocks are highly correlated, the edge weight between them will be increased, and vice versa.

[0073] In healthcare, patient feedback can influence adjustments to treatment plans. For example, patient feedback after a treatment (such as drug response or changes in physical signs) can influence the system's assessment of the patient's condition and, in turn, update the treatment path. Node updates can involve changes to the patient's health data, while edge updates can involve adjustments between different treatment plans.

[0074] This embodiment dynamically adjusts the graph structure by updating the nodes and edges in the graph structure data based on feedback information triggered by execution actions, enabling the graph structure to more accurately reflect the current environmental state. This dynamic update mechanism enhances the adaptability of the graph structure, allowing the system to better respond to environmental changes and make more reasonable decisions. In reinforcement learning, timely updating of nodes and edges in the graph structure makes state representation more accurate, helping to improve decision quality and optimize the long-term performance of the model.

[0075] The present invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as financial technology and medical health. It discloses a graph-based execution decision and optimization method, device, equipment and medium, including: obtaining original state information, converting the original state information into graph structure data, wherein the graph structure data includes multiple nodes representing entities and one or more edges representing relationships between entities, processing the graph structure data through a graph neural network, aggregating multi-hop neighborhood information of nodes based on the relationships between entities defined by the edges, generating an enhanced state representation, determining the action to be executed based on the enhanced state representation, and updating the nodes and edges in the graph structure data according to feedback information triggered by the execution of the action. The present invention models the state representation through a graph neural network, effectively capturing the multi-level relationships between entities in the environment, overcoming the problems of redundant state representation or missing key information, thereby improving the accuracy of the state representation and the intelligence of the processing, and optimizing the performance and generalization ability of the reinforcement learning model.

[0076] In one embodiment, in step S10, the original state information is converted into graph structure data, where the graph structure data includes a plurality of nodes representing entities and one or more edges representing relationships between the entities, including:

[0077] S101, identifying entity objects with independent attributes as candidate nodes from the original state information;

[0078] S102, analyzing the Euclidean distances between candidate nodes according to the spatial distribution of the candidate nodes;

[0079] S103, generating edges representing spatial relationships between entities based on the Euclidean distance;

[0080] S104, extracting the movement speed and category label of the candidate node as node attributes;

[0081] S105: Combining the node attributes and the edges representing the spatial relationship between entities into graph structure data.

[0082] In this embodiment, the system extracts entity objects with independent attributes from the original state information. Entity objects refer to objects that have independence in a specific situation or environment. For example, in an intelligent transportation system, vehicles and pedestrians are independent entities. In an intelligent medical system, patients and doctors may be independent entities. Each entity has its own independent characteristics, such as location, speed, health status, etc. The key to this step is to identify which objects should be used as nodes of the graph in a given scenario, and these nodes will contain independent attribute information. Assume that the original state information contains a set of objects (such as vehicles, pedestrians, patients, etc.), each object can be identified as a node based on its attributes such as type, location and state. In implementation, the data processing algorithm can be combined with the identifier and attributes of the object to extract each independent information as a candidate node.

[0083] Once candidate nodes are identified, the system analyzes their spatial distribution. Spatial distribution analysis is typically based on the location coordinates of objects, which can be two-dimensional or three-dimensional. In this step, by calculating the Euclidean distance between nodes, the system can determine the relative position of each node in space. In intelligent transportation, the spatial distribution of vehicles and pedestrians can be calculated using coordinates acquired by GPS or sensors. In healthcare, the location data of patients or doctors can be obtained through positioning systems, allowing for similar distance calculations.

[0084] After calculating the Euclidean distances between nodes, the system generates edges based on these distances to represent the spatial relationships between entities. The existence of an edge indicates that there is a certain relationship between the nodes. For example, in a traffic system, the distance between vehicles can be used to represent the connection between them through an edge. A shorter distance may indicate a stronger connection or a higher priority. If the Euclidean distance between two nodes is less than a preset threshold, the two nodes are connected by an edge. This edge indicates that there is some kind of interaction or association between the two. For example, in a traffic management system, an edge may represent the influence area between vehicles, while in a medical system, an edge may represent the connection between a patient and a medical device or doctor.

[0085] Each node not only represents a static entity, but also should contain some dynamic attributes, such as speed, category, etc. These attributes help to further analyze the behavior or characteristics of the node. The speed of movement indicates the movement rate of the node in a specific time, and the category label can indicate the category to which the node belongs, such as "vehicle", "pedestrian", "patient", etc. The attribute extraction of each node can be obtained through sensors or data sources. For example, the speed of a vehicle can be obtained through radar or GPS equipment, while the speed of a pedestrian can be estimated through video analysis technology. Category labels are usually determined based on object recognition technology or predefined classification rules.

[0086] The attributes of nodes and edges are combined to form the final graph structure data. Node attributes include speed, category labels, etc., while edges represent the spatial relationships between nodes. By combining this information, the final graph structure can accurately represent the complex relationships between entities and provide structured data for subsequent graph neural network processing. Each node will contain a set of attributes (such as speed, category label), and each edge will include a weight representing the spatial relationship. Through the processing of graph neural networks, this structured data can be used to generate richer node representations, which can then be used in applications such as prediction, analysis, and decision-making.

[0087] In practical applications, this process can be accomplished through the collaboration of multiple sensors, data sources, and processing modules. For example, in intelligent transportation systems, vehicles and pedestrians obtain location data through sensors or cameras. The system calculates Euclidean distances to generate edges connecting vehicles and pedestrians. These edges not only represent physical proximity but also provide deeper information through node attributes (such as speed and category). The system further processes this data to generate a graph structure, ultimately performing complex state analysis using graph neural networks.

[0088] In healthcare, nodes can represent patients, doctors, or medical devices, while edges represent the physical or logical relationships between them (e.g., the treatment relationship between a patient and a doctor, or the usage relationship between a patient and a device). Node attributes (e.g., the patient's health status, the doctor's specialty) and edge attributes (e.g., the frequency of interaction during treatment) are combined into graph-structured data for subsequent analysis and prediction.

[0089] This embodiment converts raw state information into graph-structured data and constructs a graph based on node attributes and spatial relationships, enabling the system to more accurately represent and process complex relationships between entities. Graph-structured data effectively integrates multi-dimensional information between entities, facilitating in-depth analysis and processing by graph neural networks, thereby improving the system's predictive capabilities and decision-making accuracy.

[0090] In one embodiment, in step S10, the original state information is converted into graph structure data, where the graph structure data includes a plurality of nodes representing entities and one or more edges representing relationships between the entities, including:

[0091] S106, when the original state information includes table data, defining each data record in the table data as a node;

[0092] S107, determining the Pearson correlation coefficient between the eigenvectors of each node;

[0093] S108, selecting a node pair whose absolute value of the Pearson correlation coefficient is greater than a preset coefficient threshold, and generating an edge based on the node pair;

[0094] S109, encoding the discrete features in the tabular data into node category labels;

[0095] S110, generating node attributes according to the normalized values of the continuous features in the table data;

[0096] S111 , combining the nodes, node category labels, node attributes, and edges into graph structure data.

[0097] In this embodiment, the system first identifies the tabular data in the original state information. Tabular data usually consists of multiple records, each record representing an entity or an object. Each data record will be defined as a node in the graph structure, and the node represents the position of these entities in the graph. The node is the basic building block of the graph and represents each specific entity in the original state information. Assume that the original state information contains a user information table, in which each record includes the user's name, age, health status and other data. Each record will be converted into a node, and the node contains all the data of this record as its attributes. For example, if the tabular data record is the personal information of the user "Zhang San", the node generated by the record may contain "Zhang San's" name, age, gender and other attributes.

[0098] After generating the nodes, the next step is to calculate the feature correlation between the nodes. Each node has a feature vector, which can be the attributes of the node (for example, the user's age, health status, etc.). The Pearson correlation coefficient is a statistic that measures the linear correlation between two variables. By calculating the Pearson correlation coefficient of the feature vectors between nodes, the system can determine which nodes are more similar in characteristics and thus decide whether a connection should be established between them. Assuming that the feature vector of each node contains information such as the user's age, gender, health status, etc., the system evaluates the similarity between them by calculating the Pearson correlation coefficient between the feature vectors of each two nodes.

[0099] After calculating the Pearson correlation coefficient between nodes, the next step is to select node pairs whose absolute value of the correlation coefficient is greater than the preset threshold and generate edges based on these node pairs. Edges indicate that there is a certain relationship between nodes. In this step, the setting of the threshold is key. Only when the feature correlation between nodes is strong enough will a connection be established between them. Assuming the preset threshold is 0.8, the system will filter out all node pairs with an absolute value of the Pearson correlation coefficient greater than 0.8 and generate edges between these nodes. The generated edges will be marked as "high similarity" and connect the relevant nodes in the graph. For example, if the health status and age of two users are highly correlated, an edge will be established between them, and the weight of the edge may indicate the strength of the similarity.

[0100] Some features in tabular data are discrete, such as gender, occupation type, etc. These features are usually categorical variables. In order to convert these discrete features into graph-structured data, they can be encoded as category labels of nodes. Category labels are another attribute of nodes that can help the system group or classify nodes. Assuming that the gender field in the user table is "male" or "female", these discrete features will be encoded as numerical values (for example, male = 0, female = 1) or represented as category labels using strings (such as "male" and "female"). Each node will have a category label attribute that indicates the category to which the node belongs.

[0101] For continuous features in tabular data, such as age, income, weight, etc., these features need to be normalized. Normalization can ensure that all features are on the same scale, thereby avoiding the impact of a feature's scale being too large or too small on the graph structure. The normalized feature values will be used to generate node attributes to help the system better process these numerical data. Suppose the user's age feature is 30 years old and his income is 5,000 yuan. Through normalization, the age feature can be converted into a value between 0 and 1 through the formula, for example: Normalized value = (current value - minimum value) / (maximum value - minimum value). In this way, the values of continuous features such as age and income will be scaled to a uniform range, which is convenient for subsequent graph neural network processing.

[0102] After extracting node features and generating edges, the system combines the nodes, their attributes, category labels, and edges into the final graph structure data. This graph structure data serves as the input to the graph neural network and contains information about the relationships between nodes as well as the node characteristics themselves. Generating this graph structure data is the ultimate goal of this step, effectively supporting subsequent processing by the graph neural network for deeper analysis. Assume that each node has an attribute vector, including continuous and discrete features, as well as a category label. By combining this information with the edge definitions, a complete graph structure data is generated, which can be directly fed into the graph neural network for further processing.

[0103] By converting tabular data into graph-structured data, this embodiment can effectively capture the complex relationships and interactions between entities, making subsequent graph neural network processing more accurate and comprehensive. Graph-structured data provides a more flexible and structured representation for modeling relationships between entities, avoiding the structured relationships that may be overlooked by simple feature extraction methods in traditional methods. This can improve the model's expressiveness and predictive accuracy in a variety of application scenarios, particularly in areas such as traffic management, social networks, and healthcare.

[0104] In one embodiment, in step S10, the original state information is converted into graph structure data, where the graph structure data includes a plurality of nodes representing entities and one or more edges representing relationships between the entities, including:

[0105] S112, when the original state information includes three-dimensional point cloud data, reducing the density of discrete point clouds in the three-dimensional point cloud data by voxel rasterization processing;

[0106] S113, using a region growing module to cluster discrete point clouds in the three-dimensional point cloud data whose spatial distance is less than a preset distance threshold to form nodes;

[0107] S114, determining the spatial coordinates of the center point of the bounding box of the node, and using the spatial coordinates as the node position attribute;

[0108] S115, generating edges according to the visibility detection results between the nodes;

[0109] S116, generating edge attributes based on the reflection intensity of the discrete point cloud in the three-dimensional point cloud data;

[0110] S117: Combine the nodes, node position attributes, edges, and edge attributes into graph structure data.

[0111] In this embodiment, the system converts raw 3D point cloud data into graph-structured data for processing by a graph neural network. 3D point cloud data typically consists of a set of points collected in space, with each point representing a spatial location and often used to represent the shape of an object or environment. During this conversion process, nodes represent points in 3D space, while edges represent the spatial relationships between these points.

[0112] Point cloud data usually has high density and noise, so voxel rasterization is used to reduce the density of point clouds and convert them into discretized grid data. Voxel rasterization divides the three-dimensional space into small cubic voxels and maps all points to these voxels, thereby reducing the number of points in the original data and making them more evenly distributed. This process helps to simplify the data and reduce the computational burden in subsequent processing while retaining the overall structural characteristics of the point cloud. Assuming that in the point cloud data, the coordinates of each point in the space are (x, y, z), voxel rasterization sets a voxel size (such as 1cm 3 ), assigning each point to a corresponding voxel, thereby reducing the number of points to be processed. For example, if multiple points within a region fall within the same voxel, they are considered the same representative point. Voxel rasterization reduces the density of the point cloud, improving the efficiency of subsequent processing.

[0113] Region growing is a distance-based clustering algorithm that starts from an initial point (seed point) and gradually adds adjacent points to the cluster until a preset distance threshold is reached. For each point, the spatial distance from other points is determined to determine whether to add it to the current cluster. This process can effectively group discrete point cloud data in three-dimensional space into several regions based on spatial position relationships. In practical applications, the region growing module will start from a certain initial point and calculate the Euclidean distance between this point and other points. If the distance of a point is less than a preset threshold, the point is classified into the same category as the initial point. Continue to perform the same process on the surrounding points until the region growing is completed. Ultimately, each point group (cluster) formed corresponds to a node, and each node represents a set of points with spatial adjacency in three-dimensional space.

[0114] For each generated node, the system calculates the center point of its bounding box and uses the spatial coordinates of the center point as the node's attributes. The bounding box is a minimal rectangular box that completely contains the point group represented by the node, and the coordinates of the center point represent the position of the node in three-dimensional space. In three-dimensional space, assuming that the point set of the node contains several discrete points, the center point of the bounding box is calculated as follows:

[0115]

[0116] Among them, x i ,y i ,z i is the coordinate of each point in the node set, and n is the number of points contained in the node. The center point, as a node position attribute, provides the representative position of the node in three-dimensional space.

[0117] Visibility detection refers to determining whether objects or point clouds can directly "see" each other in three-dimensional space based on the occlusion relationship between them. If two nodes are not blocked by other objects or obstacles and the distance between them is close, it is considered that there is a direct visibility relationship between them, and an edge is generated between these nodes. This process helps to determine the spatial relationship between nodes. Visibility detection can be accomplished through technologies such as ray tracing. The system sends a ray from one node to another and checks whether the ray is blocked by any obstacles. If the ray is not blocked, an edge is generated between the two nodes. For example, in an autonomous driving system, the visibility relationship between the vehicle and the obstacles ahead is crucial, and the system uses visibility detection to determine which obstacles are potential risks.

[0118] Reflection intensity refers to the intensity of the light signal reflected from each point collected from the point cloud data, which is usually related to the properties of the object's surface (such as material, surface smoothness, etc.). Reflection intensity can be used as an attribute of the edge to describe the strength of the association between nodes. The higher the reflection intensity, the stronger the connection between the nodes. In a laser radar (LiDAR) system, the reflection intensity usually varies with the degree of reflection of the laser beam. The reflection intensity value of each point cloud sampling point can be used as additional information of the point as the weight of the edge. For each pair of connected nodes, the attribute of the edge can be set to the weighted average of the reflection intensities of the two. For example, if the reflection intensities between node A and node B are 0.8 and 0.6 respectively, the attribute of the edge connecting them may be 0.7.

[0119] Ultimately, all the attributes of the nodes (such as node features, position attributes) and the edges between them (including edge attributes) will be combined into complete graph structure data for subsequent graph neural network processing. Graph structure data includes not only basic entity information but also relationship information between nodes, which is the basis for graph neural network training. The system organizes the characteristics of each node (such as position, movement speed, category label) and edge information (such as visibility, reflection intensity, etc.) into a graph data structure. Each node and edge has its own attributes, and the graph structure data will be used as input for subsequent processing of the graph neural network for feature aggregation and graph structure analysis.

[0120] This embodiment converts 3D point cloud data into graph-structured data, enabling a more precise description of the relationships between different entities (such as objects, obstacles, and organs), avoiding the simplification of spatial relationships in traditional methods. The graph structure not only preserves spatial information but also flexibly represents complex interactions and structures, providing a more efficient and accurate analysis method suitable for a variety of application scenarios, including autonomous driving and medical image analysis.

[0121] In one embodiment, the above step S20 includes:

[0122] S201, in a first level of a hierarchical graph convolutional network, determining an inter-node attention weight based on the inter-entity relationship defined by the edge;

[0123] S202, performing weighted aggregation on the feature vectors of the node's immediate neighboring nodes according to the inter-node attention weights to generate a first-hop neighborhood aggregate feature;

[0124] S203, in the second level of the hierarchical graph convolutional network, taking the first-hop neighborhood aggregation feature as input, aggregating feature vectors of indirect neighboring nodes of the node based on the inter-entity relationship defined by the edge, to generate a second-hop neighborhood aggregation feature;

[0125] S204, performing node compression on the second-hop neighborhood aggregation features through a graph pooling operation, retaining the key subgraph structure and generating high-order relationship features;

[0126] S205, performing a residual connection on the first-hop neighborhood aggregation feature and the high-order relationship feature to generate a multi-level neighborhood aggregation result;

[0127] S206: Map the multi-level neighborhood aggregation results to the enhanced state representation through a fully connected layer.

[0128] In this embodiment, a graph neural network (GNN) processes graph-structured data through multi-level neighborhood aggregation to generate an enhanced state representation. This enhanced state representation can more accurately capture the deep relationships and complex interactions between nodes for subsequent decision-making. Specifically, the graph neural network aggregates multi-hop neighborhood information based on the relationships between nodes (defined by edges) through a layered convolutional network.

[0129] In the first level of the graph neural network, the importance of each pair of connected nodes in information transmission is determined by calculating the attention weight between them. The relationship between nodes is defined by the attributes of the edge (such as spatial position, interaction relationship, etc.). In this way, the graph neural network can assign different weights to each pair of connected nodes. Nodes with stronger relationships are given higher weights, thus occupying a more important position in the subsequent aggregation process. The relationship between nodes is weighted by the attention mechanism within the graph neural network. Specifically, the network calculates the strength of the relationship between each pair of nodes (for example, the similarity between the two nodes) and determines their contribution to the subsequent information aggregation based on the calculated value.

[0130] First-hop neighborhood aggregate features are generated by weighted aggregation of the feature vectors of a node's immediate neighborhood. Attention weights determine the importance of each neighboring node in the aggregation process. Nodes are connected by edges, and the weight of the edge is directly related to the strength of the relationship between them. Within a node's neighborhood, the features of its immediate neighbors are weighted averaged according to the attention weights to generate an aggregated feature representation. This step improves the network's representational capabilities by calculating each node's neighborhood features and aggregating them to generate a new node representation.

[0131] At the second level, the graph neural network further processes the first-hop aggregated features, aggregating the feature vectors of indirect neighboring nodes. This step helps capture longer-range relationships between nodes and further enriches the node's feature information. Indirect neighbors refer to nodes that are indirectly connected through one or more intermediary nodes. The node features obtained after the first-level aggregation serve as input to further aggregate the node's indirect neighborhood information. This considers not only the features of directly neighboring nodes but also those of nodes indirectly connected through other nodes. This step further enhances the node's representation capabilities, enabling it to reflect relationships at more levels.

[0132] Graph pooling is used to compress node information, reduce redundancy, and retain the most critical parts of the graph structure. This process helps improve the computational efficiency of the model while retaining the most representative features of the graph. Graph pooling reduces the number of nodes and retains the nodes that best represent the graph structure by selecting nodes with the highest eigenvalues or filtering them through other pooling methods (such as maximum pooling). This enables graph neural networks to effectively process larger-scale graph data while still retaining key information.

[0133] Residual connections help prevent information loss that can occur in deep networks. By performing a residual connection on the first-hop neighborhood aggregated features and the second-hop aggregated features, low-order features are preserved while enhancing the influence of higher-order features. Residual connections sum the first-hop and second-hop aggregated features to ensure smooth information transfer and preserve more underlying information. This connection method improves network training efficiency and helps prevent vanishing and exploding gradients.

[0134] Finally, a fully connected layer maps the aggregated multi-level features into the final enhanced state representation. The function of a fully connected layer is to map the features after graph convolution and aggregation into the target state representation space. This representation is used for subsequent decision-making or other tasks. A fully connected layer uses a linear transformation to map the multi-level aggregated features into the final state representation space, generating an enhanced state representation. A fully connected layer can be viewed as a standard neural network layer, transforming the input features through a weight matrix and bias to produce the final output.

[0135] This embodiment effectively captures complex dependencies and high-level interactions between nodes through multi-level neighborhood aggregation in a layered graph convolutional network. This approach enables the system to maintain strong adaptability and accuracy in complex and dynamic environments, significantly improving the model's representation capabilities and decision-making quality.

[0136] In one embodiment, the above step S40 includes:

[0137] S401, parsing the state change data in the feedback information to identify newly added entity objects and invalid entity objects;

[0138] S402, removing the node corresponding to the invalid entity object and its associated edges from the node set of the original graph structure data, and retaining the original nodes that have not failed as the remaining nodes;

[0139] S403, extracting the spatial coordinates and node attributes of the newly added entity object, and generating a newly added node based on the spatial coordinates and node attributes;

[0140] S404, generating a new edge based on the spatial relationship between the new node and the remaining nodes;

[0141] S405, adjusting the edge attribute according to the reward signal in the feedback information;

[0142] S406: Generate updated graph structure data based on the newly added nodes, newly added edges, adjusted edge attributes, and remaining nodes.

[0143] In this embodiment, based on the feedback information triggered by the execution action, the nodes and edges in the graph structure data are updated.

[0144] The generation and acquisition of feedback information is a critical component of graph updates. Feedback typically comes from the actual results or environmental changes following system actions (such as agent operations or system decisions). This information is used to adjust nodes and edges in the graph, thereby improving the accuracy of subsequent state representations and the system's decision-making capabilities. In reinforcement learning or adaptive systems, the system takes an action based on its current state (for example, steering or accelerating in autonomous driving; implementing a treatment plan in healthcare). These actions are selected based on the current graph data and are typically determined by graph neural networks or other machine learning models based on node and edge state information. Following the execution of an action, the environment undergoes corresponding changes, which constitute the system's feedback. For example, in autonomous driving, the system may receive data such as collision detection results, changes in vehicle position, or changes in speed. In healthcare, feedback may include evaluations of treatment effectiveness, changes in patient health status, or a doctor's assessment of a treatment plan. To make feedback information processable and computable, the system needs to convert these environmental responses into quantified feedback signals. These signals may include: Reward signals: In reinforcement learning, reward signals typically represent the quality of an action and can be positive (such as improved efficiency or accuracy) or negative (such as errors or failures). Reward signals are used to guide the model to learn the optimal strategy; State change data: Feedback information also includes state change data between entities, such as the addition of new nodes (entities), the removal of failed nodes, or changes in node attributes (such as location, category, health status, etc.).

[0145] Based on the aforementioned environmental responses and quantified feedback signals, the system generates feedback information through a specific algorithm. This feedback information may include: the identification and attributes of newly added entities: for example, newly identified vehicles, patients, or other business entities; the identification of failed entities: for example, failed sensors, invalid operation histories, or no longer relevant business nodes; and reward signals: scores, rewards, or penalties generated based on the system's execution results.

[0146] The resulting feedback is then passed back to the graph processing module in the system for dynamic graph structure updates. This feedback influences the connection strength between nodes (edge properties), node updates, and adjustments, thereby optimizing subsequent state representation and decision-making processes.

[0147] Extract data about state changes from the feedback information and determine which entities are new and which entities have become invalid. By analyzing historical data and real-time feedback, the system can identify entities whose states have changed and make corresponding adjustments to the graph structure based on these changes. The system parses the state data in the feedback information, compares the current state with the previous state, and then identifies new entity objects (such as new vehicles, patients, users, etc.) and invalid entity objects (such as disappeared objects or invalid operations). This process is usually implemented using data mining technology, differential analysis and other methods.

[0148] For entity objects that have been identified as invalid, the system removes their corresponding nodes and edges from the graph structure, while retaining those nodes that are still valid. This is a core step in graph structure updates, ensuring that only active nodes and connection relationships are retained in the graph. The system extracts a node set from the original graph structure data and determines which nodes and edges are invalid based on the state change data. Once an invalid entity object is identified, the relevant nodes and edges are deleted from the graph, and the remaining valid nodes and their connected edges are retained to form a new node set.

[0149] The spatial coordinates and attributes of the newly added entity objects will serve as the basis for generating new nodes. Node attributes include entity characteristics (such as location, category, speed, etc.), which will help the new node establish connections with other nodes. The system extracts its spatial coordinates (such as GPS coordinates or other spatial location information) and its attributes (such as type, speed, etc.) from the newly added entity objects. These data are processed and used to generate new nodes, which are then added to the graph structure. This step injects new entities into the graph structure, increasing the complexity and diversity of the graph.

[0150] New edges are generated between the newly added nodes and the remaining nodes based on their spatial relationships. These edges represent the associations between the newly added nodes and other nodes in the graph, such as location proximity or other interactive characteristics. The system calculates the spatial relationships (e.g., distance, relative position, interaction relationships, etc.) between the newly added nodes and the remaining nodes and generates new edges based on these relationships. Edge generation can be based on preset distance thresholds, similarity metrics, and other rules to ensure that there is actual spatial or relational dependency between the connected nodes.

[0151] The system adjusts edge weights based on the reward signal in the feedback information. The reward signal reflects the frequency or importance of interactions between nodes. Therefore, for node pairs that interact frequently, the system increases the edge weights between them, while for node pairs that interact less frequently, the system decreases the edge weights. The system receives feedback signals and adjusts edge weights based on preset threshold rules. For example, when the reward signal is above a preset threshold, the system considers the relationship between the node pair to be more important and increases the edge weight between them. Conversely, when the reward signal is above a certain threshold, the system decreases the edge weight of node pairs with less frequent interactions. Adjusting edge weights helps strengthen or weaken interactions between nodes, thereby optimizing the subsequent state representation generation and decision-making process.

[0152] Finally, the system merges all updated nodes, edges, and their attributes to generate new graph data. This process integrates the newly added nodes and edges, as well as the adjusted edge weights, to form a complete graph structure. The newly added nodes, generated edges, updated edge attributes, and remaining node data are combined into a new graph structure. This new graph structure not only contains all new and updated entity information, but also includes the impact of reinforcement learning signals (such as reward signals) reflecting the current state on node and edge relationships. This updated graph structure will be used for subsequent processing and decision-making tasks.

[0153] Example: In the financial sector, graph neural networks are used to process graph-structured data based on historical trading data, market behavior, and decision feedback for prediction and decision optimization. Training datasets can include past market data, trading strategies, and the resulting feedback (such as market changes and trading results). Each historical record in this data can be considered an independent interaction scenario, where each market state and corresponding strategy decision data are represented using a graph structure. In this application, graph neural network processing focuses on the relationships between various market factors (such as different stocks, asset classes, and trading timing). Each asset, stock, or investment product is represented as a node in the graph, and the relationships between assets (such as correlations and mutual influence) are represented as edges. By aggregating multi-hop neighborhood information, the model automatically captures the underlying interaction patterns and causal relationships between these nodes. Data augmentation methods generate multiple enhanced views, enabling the model to cope with various market scenarios, including varying market volatility and different investment portfolios. These enhanced views are processed by the graph neural network to form corresponding enhanced state representations, which are used to train and optimize the decision model. Positive sample pairs are generated by combining effective market behavior and decision feedback, while negative sample pairs are generated by randomly perturbing historical market data. Through comparative learning, the model can distinguish effective from ineffective decisions, thereby improving decision accuracy and risk control capabilities. Ultimately, it can help financial institutions automatically optimize trading strategies, assess risks, and improve investment returns based on market data and historical feedback. Through continuous iterative training, the model can adaptively adjust in an ever-changing market, optimize strategies in real time, and enhance the risk management and return prediction capabilities of financial products.

[0154] This embodiment dynamically updates the nodes and edges in the graph structure based on feedback information, enabling the system to respond to environmental changes in real time and adjust state representation promptly, thereby providing more accurate decision support. This significantly improves the adaptability and representation capabilities of the graph structure, enabling the system to more flexibly handle dynamic and complex environments.

[0155] In one embodiment, before step S20, the method further includes:

[0156] S2001, collecting original state information of multiple historical interaction scenarios, converting the original state information of each historical interaction scenario into corresponding historical graph structure data, and associating the corresponding historical execution actions and historical feedback information to form a training data set;

[0157] S2002: Select target historical graph structure data from the training dataset, generate an enhanced view of the target historical graph structure data using a data augmentation method, process the enhanced view using the graph neural network to obtain a corresponding enhanced state representation, and combine multiple enhanced state representations of the same target historical graph structure data into a positive sample pair;

[0158] S2003: selecting other historical graph structure data other than the target historical graph structure data from the training dataset, generating enhanced state representations of the other historical graph structure data through the graph neural network, and combining the enhanced state representations of the other historical graph structure data with the enhanced state representations of the positive sample pairs to form negative sample pairs;

[0159] S2004, determining the cosine similarity between the enhanced state representation of the target history graph structure data and the enhanced state representation of the enhanced view in the positive sample pair;

[0160] S2005, determining the cosine similarity between the enhanced state representation of the positive sample pair and the enhanced state representation of the negative sample pair;

[0161] S2006, optimizing the parameters of the graph neural network by maximizing the cosine similarity between the target historical graph structure data in the positive sample pair and the enhanced state representation of the enhanced view, and minimizing the cosine similarity between the enhanced state representation of the positive sample pair and the enhanced state representation of the negative sample pair through a contrast loss function.

[0162] In this embodiment, a training data set is constructed using historical interaction data. Collecting raw state information means extracting the system's past state data from historical records. This state data includes multidimensional information related to system operations or events, such as environmental conditions, system inputs, execution actions, and feedback. Each historical interaction scenario represents the state of the system at a specific moment, containing specific entity information, contextual data, and corresponding operations. In actual implementation, historical data can be obtained from the system database through a log management system, or automatically captured through a monitoring and recording module. All data needs to be uniformly formatted and constructed into a training data set in combination with information such as timestamps and interaction content.

[0163] The original state information is converted into graph-structured data, which can effectively express the relationships and interaction information between entities. During the conversion process, the state information of each interaction scenario will be represented as a node in the graph, and the relationships between nodes are connected through edges. The edges in the graph represent the interaction or dependency relationships between different entities. By mapping each entity and its attributes (such as position, status, etc.) in the original state information into nodes, the relationships between entities (such as physical proximity, logical relationships, etc.) are converted into edges. In addition, historical execution actions and feedback information will serve as additional information, forming graph data together with nodes and edges. These graph-structured data provide the input required for training graph neural networks.

[0164] The goal of data augmentation is to artificially perturb the original data so that it can be trained in a wider range of contexts. Augmenting views involves perturbing or transforming the original graph structure to a certain extent without changing the underlying structure of the data. Various data augmentation techniques can be used, such as random addition of noise, subtle changes to the graph topology, and slight perturbations of nodes or edges. The resulting augmented views retain the same data attributes as the original graph, but may alter topological relationships or other graph features. Processing augmented views through graph neural networks can improve the robustness and generalization of the model.

[0165] Graph neural networks (GNNs) are used to process graph data and extract enhanced state representations. GNNs can effectively aggregate node information and learn the complex relationships between nodes in graph-structured data, thereby generating a global representation of the graph. Processing the enhanced view through GNNs yields richer node and edge representations, which are crucial for subsequent decision-making and prediction tasks. GNNs use a message passing mechanism to aggregate information about neighboring nodes. The information of each node is passed through the network and updated, ultimately generating an enhanced state representation. This enhanced state representation incorporates multi-level semantic information about nodes and edges.

[0166] The multiple enhanced state representations generated are used to construct positive pairs. These pairs contain different enhanced views of the same target historical graph data and are used to optimize the model's feature representation capabilities during training. In practice, positive pairs consist of multiple enhanced state representations generated from the same target graph data, representing the state changes of the same entity under different environments or perturbations. These enhanced state representations are paired during contrastive learning and serve as an important foundation for training the model.

[0167] Negative data is selected from the training dataset and the corresponding augmented state representations are generated using a graph neural network. Negative data is typically unrelated to the target data and helps the model learn to distinguish between correct and incorrect state-action pairs. By randomly selecting historical graph data with different structures from the target data, the corresponding augmented state representations are generated using a graph neural network. This negative data can be compared with the positive data to train the model to identify valid and invalid state associations.

[0168] The enhanced state representations of negative samples are paired with the enhanced state representations of positive sample pairs to form negative sample pairs, which are used to optimize the model in contrastive learning. The enhanced state representation of each negative sample is paired with the enhanced state representation of the corresponding positive sample pair for further similarity calculation and loss function optimization. The generation of negative samples helps the model learn more refined feature distinctions, thereby improving the model's recognition ability.

[0169] Calculate the similarity between the augmented state representation of the target historical graph data and the augmented state representation of the augmented view in the positive example pair. Similarity calculation typically uses cosine similarity to measure the angle between two vectors. A smaller angle indicates a higher similarity. Calculate the cosine similarity between the augmented state representation of the target historical graph data and the augmented state representation of the augmented view in the positive example pair. A value closer to 1 indicates a closer relationship between the two. Calculating cosine similarity helps guide the model to learn similarities between different graph structures.

[0170] Computing the similarity between the enhanced state representations of positive and negative sample pairs helps the model learn to distinguish valid from invalid state associations. Using cosine similarity to calculate the difference between positive and negative samples, positive samples should have a high similarity, while negative samples should have a low similarity. This similarity calculation serves as the basis for model optimization, promoting the model's enhanced learning of positive samples and suppressing the influence of negative samples.

[0171] Use the contrastive loss function for model optimization. The contrastive loss function aims to improve the model's discriminative ability by maximizing the similarity between positive samples and minimizing the similarity between positive and negative samples. Based on the calculated similarity values, the contrastive loss function is used to optimize the parameters of the graph neural network. Based on the similarity difference between positive and negative sample pairs, the model parameters are adjusted to improve the fit of positive samples and reduce misjudgments of negative samples.

[0172] In practical applications, the entire training process can be adapted across multiple fields, such as autonomous driving, financial forecasting, and healthcare. For autonomous driving, training data can include vehicle sensor data, current environmental conditions, and control actions. In finance, training data might include historical transaction data, market behavior data, and corresponding strategic decisions. In healthcare, training data might include patient health records, treatment decisions, and clinical feedback.

[0173] This embodiment uses a graph neural network to process graph-structured data and optimizes the model through comparative learning, effectively extracting complex relationships and interaction patterns between entities. Data augmentation and the construction of positive and negative sample pairs enhance the model's generalization and adaptability to diverse data, improving the accuracy and reliability of the decision-making model. In practical applications, it can effectively address the issues of redundant and missing state representations in traditional methods, particularly in complex dynamic environments, enabling adaptive learning and updating of optimal strategies.

[0174] In one embodiment, a graph-based execution decision and optimization device is provided, which corresponds one-to-one to the graph-based execution decision and optimization method in the above embodiment. Figure 3 , Figure 3This is a functional module diagram of a preferred embodiment of the graph-based execution decision and optimization device of the present invention. It includes a data acquisition module 10, a graph structure construction module 20, a graph neural network processing module 30, an action determination module 40, and a graph update module 50. Each functional module is described in detail below:

[0175] A graph structure building module 10 is configured to obtain original state information and convert the original state information into graph structure data, wherein the graph structure data includes a plurality of nodes representing entities and one or more edges representing relationships between the entities;

[0176] A graph neural network processing module 20 is configured to process the graph structure data through a graph neural network, aggregate multi-hop neighborhood information of the nodes based on the inter-entity relationships defined by the edges, and generate an enhanced state representation;

[0177] an action determination module 30, configured to determine an action to be performed based on the enhanced state representation;

[0178] The graph updating module 40 is configured to update the nodes and edges in the graph structure data according to feedback information triggered when the execution action is executed.

[0179] In one embodiment, the graph structure building module 10 is specifically configured to:

[0180] Identifying entity objects with independent attributes as candidate nodes from the original state information;

[0181] Analyzing the Euclidean distances between candidate nodes according to the spatial distribution of the candidate nodes;

[0182] generating edges representing spatial relationships between entities based on the Euclidean distance;

[0183] Extracting the movement speed and category label of the candidate node as node attributes;

[0184] The node attributes and the edges representing the spatial relationship between entities are combined into graph structure data.

[0185] In one embodiment, the graph structure building module 10 is specifically configured to:

[0186] When the original state information includes table data, each data record in the table data is defined as a node;

[0187] Determine the Pearson correlation coefficient between the eigenvectors of each node;

[0188] Selecting a node pair whose absolute value of the Pearson correlation coefficient is greater than a preset coefficient threshold, and generating an edge based on the node pair;

[0189] Encoding discrete features in the tabular data into node category labels;

[0190] generating node attributes according to normalized values of continuous features in the tabular data;

[0191] The nodes, node category labels, node attributes and edges are combined into graph structure data.

[0192] In one embodiment, the graph structure building module 10 is specifically configured to:

[0193] When the original state information includes three-dimensional point cloud data, reducing the density of discrete point clouds in the three-dimensional point cloud data by voxel rasterization processing;

[0194] Clustering discrete point clouds whose spatial distance in the three-dimensional point cloud data is less than a preset distance threshold into nodes using a region growing module;

[0195] Determine the spatial coordinates of the center point of the bounding box of the node, and use the spatial coordinates as the node position attribute;

[0196] generating edges according to the visibility detection results between the nodes;

[0197] generating edge attributes based on reflection intensities of discrete point clouds in the three-dimensional point cloud data;

[0198] The nodes, node position attributes, edges and edge attributes are combined into graph structure data.

[0199] In one embodiment, the graph neural network processing module 20 is specifically configured to:

[0200] In a first level of the hierarchical graph convolutional network, attention weights between nodes are determined based on the relationships between entities defined by the edges;

[0201] Perform weighted aggregation on the feature vectors of the node's immediate neighboring nodes according to the inter-node attention weights to generate first-hop neighborhood aggregate features;

[0202] In a second level of the hierarchical graph convolutional network, taking the first-hop neighborhood aggregation feature as input, aggregating feature vectors of indirect neighboring nodes of the node based on the inter-entity relationship defined by the edge to generate a second-hop neighborhood aggregation feature;

[0203] The second-hop neighborhood aggregation features are node-compressed through graph pooling operations to retain key subgraph structures and generate high-order relationship features;

[0204] Performing a residual connection between the first-hop neighborhood aggregation feature and the high-order relationship feature to generate a multi-level neighborhood aggregation result;

[0205] The multi-level neighborhood aggregation results are mapped to the enhanced state representation through a fully connected layer.

[0206] In one embodiment, the map updating module 40 is specifically configured to:

[0207] Parsing the state change data in the feedback information to identify newly added entity objects and invalid entity objects;

[0208] Remove the node and its associated edges corresponding to the invalid entity object from the node set of the original graph structure data, and retain the original nodes that have not failed as the remaining nodes;

[0209] Extracting the spatial coordinates and node attributes of the newly added entity object, and generating a new node based on the spatial coordinates and node attributes;

[0210] Generate a new edge based on the spatial relationship between the new node and the remaining nodes;

[0211] Adjusting edge attributes according to the reward signal in the feedback information;

[0212] Updated graph structure data is generated based on the newly added nodes, the newly added edges, the adjusted edge attributes, and the remaining nodes.

[0213] In one embodiment, the graph neural network processing module 20 is specifically configured to:

[0214] Collect the original state information of multiple historical interaction scenarios, convert the original state information of each historical interaction scenario into corresponding historical graph structure data, and associate the corresponding historical execution actions and historical feedback information to form a training data set;

[0215] Selecting target historical graph structure data from the training dataset, generating an enhanced view of the target historical graph structure data using a data augmentation method, processing the enhanced view using the graph neural network to obtain a corresponding enhanced state representation, and combining multiple enhanced state representations of the same target historical graph structure data into a positive sample pair;

[0216] Selecting other historical graph structure data other than the target historical graph structure data from the training data set, generating enhanced state representations of the other historical graph structure data through the graph neural network, and combining the enhanced state representations of the other historical graph structure data with the enhanced state representations of the positive sample pairs to form a negative sample pair;

[0217] determining a cosine similarity between an augmented state representation of the target history graph structure data and an augmented state representation of the augmented view in the positive sample pair;

[0218] Determining a cosine similarity between the augmented state representation of the positive sample pair and the augmented state representation of the negative sample pair;

[0219] The parameters of the graph neural network are optimized by maximizing the cosine similarity between the target historical graph structure data and the enhanced state representation of the enhanced view in the positive sample pair and minimizing the cosine similarity between the enhanced state representation of the positive sample pair and the enhanced state representation of the negative sample pair through the contrast loss function.

[0220] In one embodiment, a computer device is provided. The computer device may be a server, and its internal structure diagram may be as follows: Figure 4 As shown. The computer device includes a processor, memory, network interface and database connected via a system bus. The processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage medium stores an operating system, a computer program and a database. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external user terminal via a network connection. When the computer program is executed by the processor, it realizes the functions or steps of a server-side execution decision and optimization method based on a graph structure.

[0221] In one embodiment, a computer device is provided. The computer device may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown. The computer device includes a processor, memory, network interface, display screen and input device connected via a system bus. Among them, the processor of the computer device is used to provide computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and an internal memory. The non-volatile storage medium stores an operating system and a computer program. The internal memory provides an environment for the operation of the operating system and computer program in the non-volatile storage medium. The network interface of the computer device is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it realizes the functions or steps of a user-side execution decision and optimization method based on a graph structure.

[0222] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor. When the processor executes the computer program, the following steps are performed:

[0223] Acquire original state information and convert the original state information into graph structure data, wherein the graph structure data includes a plurality of nodes representing entities and one or more edges representing relationships between the entities;

[0224] Processing the graph structure data through a graph neural network, aggregating multi-hop neighborhood information of the node based on the inter-entity relationships defined by the edges, and generating an enhanced state representation;

[0225] determining an action to be performed based on the enhanced state representation;

[0226] The nodes and edges in the graph structure data are updated according to the feedback information triggered when the execution action is executed.

[0227] In one embodiment, a computer-readable storage medium is provided, on which a computer program is stored. When the computer program is executed by a processor, the following steps are implemented:

[0228] Acquire original state information and convert the original state information into graph structure data, wherein the graph structure data includes a plurality of nodes representing entities and one or more edges representing relationships between the entities;

[0229] Processing the graph structure data through a graph neural network, aggregating multi-hop neighborhood information of the node based on the inter-entity relationships defined by the edges, and generating an enhanced state representation;

[0230] determining an action to be performed based on the enhanced state representation;

[0231] The nodes and edges in the graph structure data are updated according to the feedback information triggered when the execution action is executed.

[0232] It should be noted that the above functions or steps that can be implemented by the computer-readable storage medium or computer device can be found in the relevant descriptions of the server side and the user side in the aforementioned method embodiment. To avoid repetition, they will not be described one by one here.

[0233] Those skilled in the art will appreciate that all or part of the processes in the above-mentioned embodiments can be implemented by instructing the relevant hardware through a computer program. The computer program can be stored in a non-volatile computer-readable storage medium. When the computer program is executed, it can include the processes of the embodiments of the above-mentioned methods. Among them, any reference to memory, storage, database or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM) or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), double data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link (Synchlink) DRAM (SLDRAM), memory bus (Rambus) direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM).

[0234] Those skilled in the art will clearly understand that for the sake of convenience and brevity of description, only the division of the above-mentioned functional units and modules is used as an example. In actual applications, the above-mentioned functions can be distributed and completed by different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

[0235] It should be noted that if any software tools or components other than those of the Company appear in the embodiments of this application, they are merely for illustration and do not represent actual use. The above embodiments are intended only to illustrate the technical solutions of the present invention, not to limit them. Although the present invention has been described in detail with reference to the above embodiments, those skilled in the art should understand that they can still modify the technical solutions described in the above embodiments, or replace some of the technical features therein with equivalents. These modifications or replacements do not deviate the essence of the corresponding technical solutions from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included in the scope of protection of the present invention.

Claims

1. A graph-based execution decision and optimization method, characterized in that: The following steps are involved: Acquire original state information and convert the original state information into graph structure data, wherein the graph structure data includes a plurality of nodes representing entities and one or more edges representing relationships between the entities; Processing the graph structure data through a graph neural network, aggregating multi-hop neighborhood information of the node based on the inter-entity relationships defined by the edges, and generating an enhanced state representation; determining an action to be performed based on the enhanced state representation; The nodes and edges in the graph structure data are updated according to the feedback information triggered when the execution action is executed.

2. The graph-based execution decision and optimization method according to claim 1, wherein: Converting the original state information into graph structure data, wherein the graph structure data includes a plurality of nodes representing entities and one or more edges representing relationships between the entities, including: Identifying entity objects with independent attributes as candidate nodes from the original state information; Analyzing the Euclidean distances between candidate nodes according to the spatial distribution of the candidate nodes; generating edges representing spatial relationships between entities based on the Euclidean distance; Extracting the movement speed and category label of the candidate node as node attributes; The node attributes and the edges representing the spatial relationship between entities are combined into graph structure data.

3. The graph-based execution decision and optimization method according to claim 1, wherein: Converting the original state information into graph structure data, wherein the graph structure data includes a plurality of nodes representing entities and one or more edges representing relationships between the entities, including: When the original state information includes table data, each data record in the table data is defined as a node; Determine the Pearson correlation coefficient between the eigenvectors of each node; Selecting a node pair whose absolute value of the Pearson correlation coefficient is greater than a preset coefficient threshold, and generating an edge based on the node pair; Encoding discrete features in the tabular data into node category labels; generating node attributes according to normalized values of continuous features in the tabular data; The nodes, node category labels, node attributes and edges are combined into graph structure data.

4. The graph-based execution decision and optimization method according to claim 1, wherein: Converting the original state information into graph structure data, wherein the graph structure data includes a plurality of nodes representing entities and one or more edges representing relationships between the entities, including: When the original state information includes three-dimensional point cloud data, reducing the density of discrete point clouds in the three-dimensional point cloud data by voxel rasterization processing; Clustering discrete point clouds whose spatial distance in the three-dimensional point cloud data is less than a preset distance threshold into nodes using a region growing module; Determine the spatial coordinates of the center point of the bounding box of the node, and use the spatial coordinates as the node position attribute; generating edges according to the visibility detection results between the nodes; generating edge attributes based on reflection intensities of discrete point clouds in the three-dimensional point cloud data; The nodes, node position attributes, edges and edge attributes are combined into graph structure data.

5. The graph-based execution decision and optimization method according to claim 1, wherein: Processing the graph structure data through a graph neural network, aggregating multi-hop neighborhood information of the node based on the inter-entity relationships defined by the edges, and generating an enhanced state representation, including: In a first level of the hierarchical graph convolutional network, attention weights between nodes are determined based on the relationships between entities defined by the edges; Perform weighted aggregation on the feature vectors of the node's immediate neighboring nodes according to the inter-node attention weights to generate first-hop neighborhood aggregate features; In a second level of the hierarchical graph convolutional network, taking the first-hop neighborhood aggregation feature as input, aggregating feature vectors of indirect neighboring nodes of the node based on the inter-entity relationship defined by the edge to generate a second-hop neighborhood aggregation feature; The second-hop neighborhood aggregation features are node-compressed through graph pooling operations to retain key subgraph structures and generate high-order relationship features; Performing a residual connection between the first-hop neighborhood aggregation feature and the high-order relationship feature to generate a multi-level neighborhood aggregation result; The multi-level neighborhood aggregation results are mapped to the enhanced state representation through a fully connected layer.

6. The graph-based execution decision and optimization method according to claim 1, wherein: Updating the nodes and edges in the graph structure data according to feedback information triggered when the execution action is executed, including: Parsing the state change data in the feedback information to identify newly added entity objects and invalid entity objects; Remove the node and its associated edges corresponding to the invalid entity object from the node set of the original graph structure data, and retain the original nodes that have not failed as the remaining nodes; Extracting the spatial coordinates and node attributes of the newly added entity object, and generating a new node based on the spatial coordinates and node attributes; Generate a new edge based on the spatial relationship between the new node and the remaining nodes; Adjusting edge attributes according to the reward signal in the feedback information; Updated graph structure data is generated based on the newly added nodes, the newly added edges, the adjusted edge attributes, and the remaining nodes.

7. The graph-based execution decision and optimization method according to claim 1, wherein: Before processing the graph structure data by a graph neural network and aggregating multi-hop neighborhood information of the node based on the inter-entity relationships defined by the edges to generate an enhanced state representation, the method further includes: Collect the original state information of multiple historical interaction scenarios, convert the original state information of each historical interaction scenario into corresponding historical graph structure data, and associate the corresponding historical execution actions and historical feedback information to form a training data set; Selecting target historical graph structure data from the training dataset, generating an enhanced view of the target historical graph structure data using a data augmentation method, processing the enhanced view using the graph neural network to obtain a corresponding enhanced state representation, and combining multiple enhanced state representations of the same target historical graph structure data into a positive sample pair; Selecting other historical graph structure data other than the target historical graph structure data from the training data set, generating enhanced state representations of the other historical graph structure data through the graph neural network, and combining the enhanced state representations of the other historical graph structure data with the enhanced state representations of the positive sample pairs to form a negative sample pair; determining a cosine similarity between an augmented state representation of the target history graph structure data and an augmented state representation of the augmented view in the positive sample pair; Determining a cosine similarity between the augmented state representation of the positive sample pair and the augmented state representation of the negative sample pair; The parameters of the graph neural network are optimized by maximizing the cosine similarity between the target historical graph structure data and the enhanced state representation of the enhanced view in the positive sample pair and minimizing the cosine similarity between the enhanced state representation of the positive sample pair and the enhanced state representation of the negative sample pair through the contrast loss function.

8. An execution decision and optimization device based on a graph structure, characterized in that: The graph-based execution decision and optimization device includes: A graph structure building module, configured to obtain original state information and convert the original state information into graph structure data, wherein the graph structure data includes a plurality of nodes representing entities and one or more edges representing relationships between the entities; A graph neural network processing module, configured to process the graph structure data through a graph neural network, aggregate multi-hop neighborhood information of the nodes based on the inter-entity relationships defined by the edges, and generate an enhanced state representation; an action determination module, configured to determine an action to be performed based on the enhanced state representation; A graph updating module is used to update the nodes and edges in the graph structure data according to feedback information triggered when the execution action is executed.

9. A computer device, characterized in that: The computer device includes a memory, a processor, and a graph-based execution decision and optimization program stored in the memory and capable of running on the processor. When the graph-based execution decision and optimization program is executed by the processor, the steps of the graph-based execution decision and optimization method as described in any one of claims 1 to 7 are implemented.

10. A computer-readable storage medium, characterized in that The storage medium stores a graph-structure-based execution decision and optimization program, which, when executed by a processor, implements the steps of the graph-structure-based execution decision and optimization method as described in any one of claims 1 to 7.

Citation Information

Cited By

  • Large-scale GPS point cloud data tobacco field contour extraction method based on graph neural network

    CN120807962A