Decision-making method and device based on environment semantic map, equipment and medium
By acquiring multimodal input data, constructing and updating the environmental semantic graph, and optimizing the decision path using a flow matching network, the problem of insufficient combination of multimodal fusion features in existing technologies is solved, and efficient and accurate action sequence decision-making in complex environments is achieved.
Patent Information
- Application Number
- CN202511060290.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-11-14
AI Technical Summary
Existing VLA models struggle to effectively incorporate multimodal fusion features in fintech, healthcare, and intelligent robotics fields. They lack globally optimal action decision-making and dynamic adjustment capabilities, resulting in low task execution efficiency, insufficient accuracy, and an inability to adapt to complex environmental changes.
By acquiring multimodal input data, extracting and fusing features, constructing and dynamically updating the environmental semantic graph, using a flow matching network to learn transport mapping to generate action sequences, adjusting the action sequences based on environmental interaction feedback, and updating network parameters in conjunction with a preset reward function.
It enables efficient and accurate action sequence decision-making in complex and dynamic environments, improves multimodal information processing capabilities, and enhances the understanding of semantic relationships in the environment and decision-making performance.
Smart Images

Figure CN120952169A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a decision-making method, apparatus, device, and storage medium based on environmental semantic graphs. Background Technology
[0002] In the fintech business field, existing Visual-Language-Action (VLA) models face numerous technical bottlenecks in applications such as automated risk control and intelligent counter services. Traditional models typically process visual data (such as document images and customer behavior monitoring), linguistic data (such as voice commands and text input), and business action data (such as interaction process records) independently. This isolated approach makes it difficult to accurately understand the complex context of financial business. For example, in automated counter service scenarios, when faced with multi-step business instructions from customers, traditional models lack a comprehensive understanding of the relevant objects in the counter environment, business rules, and customer intentions, resulting in low task execution efficiency and insufficient accuracy. Furthermore, current models generally adopt static multimodal fusion strategies, which are difficult to cope with complex situations such as dynamic changes in customer behavior and updates to service rules in the business environment, resulting in untimely decision-making responses and easily leading to service anomalies and potential risks. Especially in the decision-making process, traditional methods lack the ability to effectively utilize environmental feedback during business execution to adjust their action sequences, failing to achieve the optimal customer service path under the global task objective.
[0003] In the healthcare field, existing VLA models also have significant shortcomings in scenarios such as multimodal assisted diagnosis and rehabilitation training guidance. Current models often employ action decision-making methods based on local rules, making it difficult to efficiently integrate complex visual inputs (such as patient posture and rehabilitation equipment location), linguistic information (such as medical orders and training instructions), and action execution requirements (such as precise training movements) in the medical environment. This results in a lack of adaptability to dynamic medical environments. For example, when rehabilitation robots guide patient training, traditional models cannot dynamically adjust based on real-time changes in the patient's state and the specific distribution of equipment in the environment, leading to inaccurate guidance and a lack of individualized optimization of training programs. Furthermore, the lack of full utilization of real-time interactive feedback with patients limits action decisions to predefined paths, hindering continuous optimization and improvement of training effectiveness, and impacting the safety and accuracy of medical services.
[0004] In the fields of intelligent robotics and other technologies, existing VLA models still primarily rely on local information-driven action selection when facing complex environmental tasks, making it difficult to globally optimize decision sequences. Traditional multimodal fusion methods have limitations in handling dynamic object changes, complex semantic instructions, and spatial constraints, failing to fully understand the relationships between entities in the environment, leading to a disconnect between perception results and action planning. Especially during task execution, the model fails to fully utilize environmental interaction feedback to dynamically adjust decisions, often relying on static strategies and lacking optimization for the global task objective, easily resulting in inefficiency or decision errors. Existing methods have limited ability to understand the overall semantics of the environment, causing robots to struggle to coordinate the complex relationships between visual perception, language understanding, and action execution when performing multi-step tasks such as "organizing objects and classifying them by position," resulting in poor overall task performance. Summary of the Invention
[0005] The main objective of this invention is to provide a decision-making method, apparatus, device, and storage medium based on environmental semantic graphs, aiming to solve the technical problem that existing technologies lack the ability to combine environmental semantic graphs with multimodal fusion features and combine them with flow matching networks to achieve globally optimal action decision-making and dynamic adjustment.
[0006] To achieve the above objectives, the present invention provides a decision-making method based on environmental semantic graphs, comprising:
[0007] Acquire multimodal input data, and perform feature extraction and fusion on the multimodal input data to generate initial multimodal fusion features;
[0008] Based on the entity and relationship information contained in the initial multimodal fusion features, an environmental semantic graph is constructed and dynamically updated.
[0009] The environmental semantic map is fused with the initial multimodal fusion features to generate deep fusion features;
[0010] The current action distribution is determined based on the deep fusion features, and the transmission mapping from the current action distribution to the preset target action distribution is learned through the flow matching network to generate the mapped action distribution.
[0011] An action sequence is generated based on the mapped action distribution, and the action sequence is adjusted based on real-time environmental interaction feedback.
[0012] The network parameters of the flow matching network are updated based on the environmental interaction feedback and the preset reward function.
[0013] Furthermore, to achieve the above objectives, the present invention provides a decision-making device based on an environmental semantic graph, comprising:
[0014] A multimodal feature processing module is used to acquire multimodal input data, and to extract and fuse features from the multimodal input data to generate initial multimodal fusion features;
[0015] The environmental semantic graph module is used to construct and dynamically update the environmental semantic graph based on the entity and relationship information contained in the initial multimodal fusion features.
[0016] The feature fusion module is used to fuse the environmental semantic map with the initial multimodal fusion features to generate deep fusion features;
[0017] The flow matching decision module is used to determine the current action distribution based on the deep fusion features, and learn the transmission mapping from the current action distribution to the preset target action distribution through the flow matching network to generate the mapped action distribution;
[0018] An action adjustment module is used to generate an action sequence based on the mapped action distribution, and to adjust the action sequence based on real-time environmental interaction feedback.
[0019] The strategy optimization module is used to update the network parameters of the flow matching network based on the environmental interaction feedback and the preset reward function.
[0020] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a decision-making program based on an environmental semantic graph stored in the memory and executable on the processor, wherein when the decision-making program based on the environmental semantic graph is executed by the processor, it implements the steps of the decision-making method based on the environmental semantic graph as described above.
[0021] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a decision-making program based on an environmental semantic graph, wherein when the decision-making program based on the environmental semantic graph is executed by a processor, it implements the steps of the decision-making method based on the environmental semantic graph as described above.
[0022] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech, healthcare, and intelligent robots. It discloses a decision-making method, apparatus, device, and medium based on an environmental semantic graph, comprising: acquiring multimodal input data and extracting and fusing it to generate initial multimodal fusion features; constructing and dynamically updating an environmental semantic graph based on entity and relationship information contained in the initial multimodal fusion features; fusing the environmental semantic graph and the initial multimodal fusion features to generate deep fusion features; determining the current action distribution based on the deep fusion features; using a flow matching network to learn the transmission mapping from the current action distribution to a preset target action distribution to generate a mapped action distribution; generating an action sequence and adjusting the action sequence according to real-time environmental interaction feedback; and updating the network parameters of the flow matching network by combining environmental interaction feedback and a preset reward function. This invention, by combining an environmental semantic graph with multimodal fusion features and utilizing a flow matching network to optimize the decision path and dynamically adjust the action sequence at the action distribution level, can accurately understand environmental semantic relationships, improve multimodal information processing and decision-making performance in complex dynamic environments, and achieve better action sequence decision results. Attached Figure Description
[0023] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:
[0024] Figure 1 This is a schematic diagram of an application environment for a decision-making method based on environmental semantic graphs according to an embodiment of the present invention;
[0025] Figure 2 This is a flowchart illustrating an embodiment of the decision-making method based on environmental semantic graphs of the present invention.
[0026] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the decision-making device based on environmental semantic graph of the present invention;
[0027] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;
[0028] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0029] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0030] The decision-making method based on environmental semantic graphs provided in this invention can be applied to, for example... Figure 1In this application environment, the user terminal communicates with the server via a network. The server can obtain multimodal input data from the user terminal, extract and fuse it to generate initial multimodal fusion features, construct and dynamically update an environmental semantic graph based on the entity and relationship information contained in the initial multimodal fusion features, fuse the environmental semantic graph with the initial multimodal fusion features to generate deep fusion features, determine the current action distribution based on the deep fusion features, use a flow matching network to learn the transmission mapping from the current action distribution to a preset target action distribution to generate a mapped action distribution, generate an action sequence, and adjust the action sequence according to real-time environmental interaction feedback. The network parameters of the flow matching network are updated in combination with environmental interaction feedback and a preset reward function. This invention, by combining environmental semantic graphs and multimodal fusion features, and using a flow matching network to optimize decision paths and dynamically adjust action sequences at the action distribution level, can accurately understand environmental semantic relationships, improve multimodal information processing and decision-making performance in complex dynamic environments, and achieve better action sequence decision results. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The present invention will now be described in detail through specific embodiments.
[0031] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the decision-making method based on environmental semantic graphs provided by the present invention. It should be noted that although a logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0032] like Figure 2 As shown, the decision-making method based on environmental semantic graphs proposed in this invention includes the following steps:
[0033] S10, acquire multimodal input data, and perform feature extraction and fusion on the multimodal input data to generate initial multimodal fusion features;
[0034] In this embodiment, acquiring multimodal input data involves real-time collection and structured input of heterogeneous data from different sources. Visual data can be acquired through image sensors such as high-resolution cameras. The pixel matrix in the image sequence needs to be standardized to ensure compatibility with subsequent processing modules. Language data comes from human-computer interaction command text. This type of data is represented as a character sequence and requires a unified encoding format such as UTF-8 and preprocessing to remove redundant characters and noise symbols, ensuring that the text content can be processed by the semantic model. Action data comes from historical execution records. The data format includes time-seriesd action identifiers, location coordinates, and action parameter value sets, ensuring the consistency and integrity of timestamps and spatial locations. Environmental structure data can be obtained through LiDAR scanning. Its data format is a dense or sparse three-dimensional point cloud set containing spatial geometric structure information. The integrity and temporal consistency of multimodal input data are achieved through a data synchronization mechanism, using a global clock or sensor timestamp alignment algorithm to ensure that data from different sources correspond correctly on the same reference time axis. Feature extraction and fusion of multimodal input data requires the use of dedicated feature extraction networks for different data types. Image data can be captured using convolutional neural networks to capture local spatial features, generating feature tensors that represent object shape, edge, and texture information. Text data is converted into dense vector representations through embedding models such as semantic vector models to capture semantic associations and contextual dependencies between words. Action data can be learned using temporal convolutional networks or recurrent neural networks to learn temporal dependencies and historical behavior patterns, while point cloud data is mapped into a dense representation of spatial location distribution through point cloud feature extraction models. Multimodal feature fusion can be achieved through feature concatenation operations, combining visual, linguistic, action, and spatial features according to a unified dimensional standard into a fused feature tensor, or through interactive attention mechanisms to calculate the association weights between different modalities, enabling complementary use of different information. The final generated initial multimodal fusion feature is a multidimensional tensor with the ability to jointly represent the environment, instructions, and historical behavior from multiple perspectives, serving as a unified input for subsequent processing modules. Each operation is tightly integrated, with the result of the previous processing serving as the input for the next, ensuring complete data transmission and contextual consistency.
[0035] A distributed data acquisition module can be used to independently complete data format conversion and timestamp synchronization at the visual sensor, speech recognition device, and motion control unit before sending the data to the central processing node. The layer depth and kernel size of the convolutional neural network can be adjusted to adapt to input image data of different resolutions, ensuring the level of detail in feature representation. During multimodal feature fusion, splicing, weighted averaging, or multimodal interactive attention can be selected as the fusion method according to task requirements to optimize the contribution of each modality in different task scenarios. In healthcare applications, the text preprocessing module can be adjusted to handle the special format of electronic health record text, and the point cloud processing module can be adjusted to adapt to the coordinate system of 3D medical imaging data. Furthermore, in fintech applications, the processing method for historical action data can be adjusted to support the time window statistical feature representation of multidimensional transaction sequences and customer behavior logs.
[0036] Example description: In healthcare operations, multimodal input data may include patient activity images captured by ward cameras, nursing instruction text input by medical staff, patient historical behavior records, and LiDAR point cloud scan data of the ward structure. After feature extraction and fusion, a multidimensional feature representation that can be used for nursing robot-assisted decision-making is generated.
[0037] In fintech operations, multimodal input data can include on-site banking video, customer voice input of business instructions, customer's past business transaction history, and point cloud data of the counter environment. This data is extracted and fused through a multi-channel network to form input for comprehensive customer behavior analysis.
[0038] In intelligent robot scenarios, intelligent robots can acquire indoor panoramic camera images of the home environment, user-inputted task commands such as "clean the dining table", historical task trajectory data of the robot, and 3D point cloud data of the current indoor layout. Through feature extraction and fusion, global initial multimodal fusion features are formed, which serve as the perception basis for the intelligent robot's subsequent path planning, target detection, and action sequence generation, thereby improving its task adaptation and execution efficiency in dynamic and complex home environments.
[0039] This embodiment achieves efficient integration of multi-source, multi-structure, and multi-temporal data by unifying the input and processing of multi-modal data such as visual, linguistic, action, and spatial data, combined with a dedicated multi-modal feature extraction network and fusion mechanism. This enhances data representation and environmental understanding capabilities, ensuring that subsequent systems can obtain accurate and consistent contextual input in complex environments.
[0040] S20, Based on the entity and relationship information contained in the initial multimodal fusion features, construct and dynamically update the environmental semantic graph;
[0041] In this embodiment, the entity and relation information contained in the initial multimodal fusion features need to be obtained through decoding of multidimensional perception results and contextual semantics. Entity information includes object categories, physical boundaries, geometric shapes, and spatial coordinates presented in visual, linguistic, and spatial data. Relationship information includes spatial proximity, relative direction, semantic association, and action dependencies in instructions between objects. When identifying entities, a category label and spatial coordinate representation method embedded in space is used. Relationship parsing can be completed based on geometric proximity analysis and instruction context analysis. When constructing the environmental semantic graph, a graph structure is used to organize the above entities as nodes and relationships as edges. Nodes contain attributes such as category identifiers, position vectors, shape descriptions, and dynamic states, while edges record spatial distance, semantic dependency strength, and historical task relationships. During dynamic updates, changes in the environment are automatically parsed from continuously collected multimodal input data to detect changes in node or edge attributes, triggering add, delete, and modify operations to ensure that the graph is consistent with the current environmental perception state. Node addition and deletion involve monitoring the addition or disappearance of objects, while edge adjustment involves changes in the relative positions between objects or changes in semantic instruction associations. The graph is updated using graph algorithms. The goal of the update is to maintain structural consistency and real-time content. Attribute alignment rules are used in the update to ensure the correct mapping between historical nodes and new nodes.
[0042] A combination of multi-channel convolutional networks and region detection networks can be employed, using each object detection box in the visual data as the source of nodes, and action target words and spatial labeling phrases in the language data as supplementary sources. Spatial proximity calculation methods can be used to analyze point cloud coordinate distribution and detection box overlap to determine spatial relationships as edges. Furthermore, a command parsing module can identify dependencies between actions and objects, establishing semantic relationship edges. During dynamic updates, a threshold-based monitoring mechanism can be designed, triggering updates only when the spatial coordinate changes of a node exceed a preset threshold or the semantic weight changes of an edge reach a significant level. The update algorithm can utilize graph comparison and differential update procedures, modifying only nodes and edges inconsistent with the previous graph state, reducing computational load while maintaining real-time performance.
[0043] Example: In healthcare scenarios, nursing robots automatically represent the spatial and task relationships between beds, patients, medical staff, and medical equipment in wards using environmental semantic graphs, maintain the correspondence between patients and beds and the safe distance around beds, and support the robot to dynamically adjust its path and actions.
[0044] In fintech business scenarios, counter service robots use environmental semantic graphs to describe the entities and business relationships of customers, counter windows, and business documents, ensuring the understanding and maintenance of queuing order and service processes, and improving the on-site service experience.
[0045] In smart robot home scenarios, service robots use environmental semantic graphs to fully describe the relationship between home items such as tables, chairs, coffee tables, and cabinets and the user-specified operating area. As user behavior changes, the relationship between furniture positions and instructions in the graph is updated in real time, ensuring that the robot can efficiently and accurately perform complex household chores such as cleaning, putting away items, and organizing.
[0046] This embodiment constructs and dynamically updates an environmental semantic graph based on entity and relation information extracted from initial multimodal fusion features. This enables the physical structure, semantic context, and task instruction content of the environment to be expressed in a unified graph structure and continuously maintained. It can reflect changes in the location, state, and relationships of objects in the environment in real time, providing stable and accurate environmental semantic support for subsequent complex task processing and dynamic decision-making, and improving the adaptability to dynamic changes in the environment.
[0047] S30, the environmental semantic map is fused with the initial multimodal fusion feature to generate a deep fusion feature;
[0048] In this embodiment, the environmental semantic graph is a graph-structured data where nodes represent entity features such as location, category, and state, and edges represent spatial, semantic, and constraint relationships between entities, providing a structured representation of the current environment. The initial multimodal fusion features are the result of deep joint encoding of visual, linguistic, spatial, and historical action data, containing a comprehensive perceptual representation of the environmental state.
[0049] When fusing the environmental semantic graph with the initial multimodal fusion features, the graph data first needs to be encoded into a feature space format consistent with the initial multimodal fusion features. This process can employ graph convolutional networks, which aggregate the feature information of each node in the environmental semantic graph from its neighboring nodes and perform nonlinear mapping to generate a graph embedding representation containing structural context. Each layer of the graph convolution achieves the aggregation of information on multi-order relationships by weighted summation of the features of neighboring nodes and mapping them to the latent space.
[0050] The initial multimodal fusion features can be further transformed using a fully connected neural network to align them with the graph embedding representation in terms of distribution, scale, and semantic granularity. The correlation between the aligned features is determined through an attention mechanism, using the graph embedding as the query and the initial multimodal fusion features as the key and value, respectively. The attention weight between the two is calculated to reflect the degree of coupling between them in different semantic dimensions.
[0051] Finally, these attention weights are used to sum the initial multimodal fusion features in a weighted manner, and aggregated features containing global structural context and multimodal details are obtained as deep fusion features to fully express the multimodal state and semantic structure of the current environment.
[0052] Node features can be encoded using multi-layer graph convolutional networks. The input to the graph convolutional network includes node attribute vectors and a graph adjacency matrix. Contextual information is extracted through layer-by-layer feature aggregation. The initial multimodal fusion features can be passed through a feature standardization network to ensure their mean and variance are consistent with the graph embedding, improving the fusion effect. A self-attention mechanism can be used to match the graph embedding as a query vector with the initial multimodal fusion features, calculating attention weights. The weighted aggregation operation is efficiently implemented using matrix multiplication, and the output aggregated feature dimension is consistent with the initial multimodal fusion features, facilitating subsequent computation.
[0053] Furthermore, the cross-attention module can be used to alternately update the graph embedding and the initial multimodal fusion features, enabling deep information fusion after multiple rounds of interaction. Residual connections can be used to preserve the original initial multimodal fusion feature components in the aggregated output, thereby enhancing feature stability and richness.
[0054] Example: In healthcare scenarios, operating room service robots can accurately grasp the relationship between the operating table, medical instruments, and doctors' actions by fusing environmental semantic graphs with initial multimodal fusion features, thus supporting precise instrument delivery.
[0055] In fintech business scenarios, intelligent service robots can effectively understand customer needs and plan service processes by integrating graphical information about customers, business documents, and counter space layout, as well as customer language interaction characteristics.
[0056] In smart robot home scenarios, cleaning robots dynamically perceive changes in the home environment and precisely adjust their cleaning paths by integrating the positional relationships of furniture such as coffee tables, sofas, and carpets with the map information and multimodal features of voice commands.
[0057] This embodiment fuses the environmental semantic map with the initial multimodal fusion features to generate deep fusion features that not only contain detailed information from multimodal perception but also incorporate global structural context and semantic relationships. This enables subsequent decision-making modules to perceive physical spatial layout, object semantic relationships, and task dependencies, thereby improving the understanding of complex environments and enhancing adaptability to dynamic changes. This provides a comprehensive and consistent input representation for subsequent action decisions.
[0058] S40, determine the current action distribution based on the deep fusion features, and learn the transmission mapping from the current action distribution to the preset target action distribution through the stream matching network to generate the mapped action distribution;
[0059] In this embodiment, deep fusion features are a comprehensive representation of environmental multimodal information and environmental semantic graph, including visual perception, language commands, spatial layout, and historical action information. The process of determining the current action distribution based on deep fusion features can be implemented through a policy model. The policy model is a probabilistic prediction structure based on a deep neural network. After inputting the deep fusion features, it outputs a set of probability values for each possible action under the current environmental state, representing the probability tendency to choose each action under the current conditions.
[0060] The mapping between the current action distribution and the preset target action distribution is learned through a flow matching network. A flow matching network is a model with reversible transformation properties, such as a normalized flow, which can represent the optimal transmission path between two distributions in a continuous space. The core of learning the transmission mapping is to calculate the distribution distance between the current action distribution and the target action distribution; common distances include the Wasserstein distance or the KL divergence. By minimizing this distribution distance, the reversible transformation parameters of the flow matching network are adjusted so that the current action distribution, after being mapped by the flow matching network, is as close as possible to the target action distribution.
[0061] Finally, the reversible transformation learned in the flow matching network is applied to map the current action distribution to a new distribution, generating a mapped action distribution. This mapped action distribution not only reflects the current environmental perception state but also embeds the prior requirements of the target task, serving as the output of the decision module for subsequent action sampling and planning.
[0062] A deep neural network can be used as the policy model. This network takes deeply fused features as input and outputs a probability distribution of the action space through a multilayer perceptron structure. This network can be pre-trained using historical environmental interaction data to improve the quality of initial decisions. The flow matching network can employ a normalized flow architecture with affine coupling layers to ensure the invertibility and differentiability of the mapping. Distribution distance calculation can use an approximate estimation method based on batch samples to reduce computational overhead. During training, the flow matching network parameters are updated by minimizing the distribution distance loss function, gradually approximating the target action distribution.
[0063] Furthermore, adaptive adjustments can be achieved for different tasks by embedding task context information into the parameters of the target action distribution. Flow matching networks can use multi-scale coupled layer structures to enhance their ability to fit complex distributions in high-dimensional action spaces.
[0064] Example: In healthcare scenarios, when robots assist in nursing decision-making, they can map the deep integration features of ward layout, patient status and medical instructions into the optimal distribution of nursing actions, and accurately generate nursing sequences.
[0065] In fintech business scenarios, intelligent counter robots can map the deep integration features of customer interactions, document images and business semantics into a priority sequence of service actions, enabling them to quickly and efficiently complete services.
[0066] In intelligent robot home services, the robot maps the deep integration features of the multimodal state of the home and user commands into action distribution, ensuring that the robot's path selection, action sequence and operation details meet user needs and spatial constraints, thereby improving the user experience.
[0067] This embodiment determines the current action distribution based on deep fusion features and learns the transmission mapping through a flow matching network. This enables the action distribution under the current environment perception to be quickly aligned with the target action distribution corresponding to the task requirements, ensuring that subsequent action sampling can reflect task orientation and global optimality, effectively improving the accuracy and robustness of decision-making, and enhancing the adaptability to dynamic environments and complex task requirements.
[0068] S50, generate an action sequence based on the mapped action distribution, and adjust the action sequence based on the real-time acquired environmental interaction feedback;
[0069] In this embodiment, the mapped action distribution is a probability representation of actions adjusted by combining deep fusion features and task requirements in the preceding processing steps, reflecting a comprehensive conditional probability distribution of environmental state, task semantics, and target needs. The process of generating an action sequence based on the mapped action distribution involves sequentially selecting actions through a sampling mechanism. In specific implementations, an adaptive importance sampling strategy can be adopted to ensure dense sampling in high-probability regions, thereby enhancing the feasibility and robustness of action decisions. The length and content of the action sequence can be associated with task context information; for example, the length of the action sequence can be dynamically adjusted based on user instructions or environmental state.
[0070] While executing the action sequence, real-time environmental feedback is acquired through continuous monitoring of environmental changes by sensor perception modules, collecting information such as changes in spatial layout, the state of interactive objects, and dynamic obstacles. As external observation information during real-time operation, environmental feedback can promptly reflect the effectiveness and deviations of the action execution.
[0071] Adjusting the action sequence involves comparing and evaluating the acquired environmental interaction feedback with the previously generated action sequence to determine whether the current action sequence is consistent with the environmental state. In practice, a real-time correction mechanism can be employed, updating the action distribution weights, adjusting subsequent actions to be executed, or replanning the entire sequence to ensure the feasibility of the action sequence in the current environment and the task completion rate.
[0072] An online sampling method based on Bayesian updates can be used to sample action sequences from the mapped action distribution, while setting a dynamic length upper limit to adapt to changes in environmental complexity. Real-time environmental interaction feedback can be acquired through a multimodal perception module, such as a visual camera sensing spatial changes, a force sensor detecting object interaction states, and a voice module detecting new user commands. Specific methods for adjusting action sequences can include: adjusting the weights and resampling actions that have not yet been executed, or triggering a complete sequence replanning when the environmental state changes significantly.
[0073] A priority buffering mechanism can also be used to record the correspondence between generated action sequences and environmental states, and to quickly filter candidate actions after new environmental interaction feedback occurs, thereby reducing latency.
[0074] Example: In healthcare scenarios, rehabilitation training robots generate rehabilitation training action sequences based on the mapped action distribution, and adjust the training rhythm and intensity according to the patient's real-time limb reactions and training status to ensure personalized and safe rehabilitation process.
[0075] In fintech business scenarios, intelligent interactive terminals generate a sequence of user service steps based on the mapped action distribution, and dynamically adjust subsequent guidance steps in combination with customer behavior feedback to improve service efficiency and user satisfaction.
[0076] In smart robot home services, the robot generates a sequence of cleaning and tidying actions based on the mapped action distribution, and adjusts the cleaning route and operation order according to real-time changes in home layout and user instructions to improve execution effect and user experience.
[0077] This embodiment generates action sequences from the mapped action distribution and makes real-time adjustments based on environmental interaction feedback. This ensures the dynamic feasibility and high task completion rate of the action sequences, enables rapid response to environmental changes, effectively improves the adaptability and security of system decision-making, and enhances the adaptability to complex tasks and changing environments.
[0078] S60, update the network parameters of the flow matching network based on the environmental interaction feedback and the preset reward function.
[0079] In this embodiment, environmental interaction feedback refers to real-time state data perceived and collected from the environment during the execution of the action sequence, including but not limited to the physical state, positional changes, and interaction response results of objects in the environment. The preset reward function is a multi-dimensional evaluation standard predefined according to task requirements, used to measure the effectiveness of the action sequence execution, typically including indicators such as task completion degree, efficiency performance, and the degree to which safety constraints are met. Environmental interaction feedback, as input to the reward function, is combined with the task objective to generate reward data, quantitatively evaluating the contribution of each action execution to the task objective.
[0080] Updating the network parameters of a flow matching network involves adjusting the weights of the trained network using optimization algorithms to improve its performance on newly collected environmental feedback and reward data. Specifically, the collected environmental interaction feedback is first structured and converted into a computationally usable state representation. Then, combined with the reward values at corresponding time points, an advantage function estimate is calculated to measure the superiority of the current policy compared to the average policy. Based on the advantage function estimate and the policy loss function, a loss expression is constructed as the optimization objective.
[0081] The gradient of the policy loss with respect to the network parameters is calculated layer by layer using the backpropagation algorithm. Gradient descent or its variants (such as the Adam optimizer) are then used to iteratively update the parameters of each layer in the flow matching network. In this way, the network can continuously improve based on new environmental feedback and reward evaluations, thereby gradually learning a better action distribution mapping in dynamic environments.
[0082] A comprehensive reward function can be formed by combining multi-dimensional reward functions, such as setting rewards for task completion metrics, penalties for task execution time, and rewards for compliance with safety rules. After environmental interaction feedback and reward data collection, the correspondence between environmental state, action sequence, and reward data can be established through time series alignment. The dominance function estimation can be implemented based on the temporal difference (TD) method, and the policy loss function can adopt the proximal policy optimization (PPO) loss function to ensure update stability.
[0083] When updating network parameters, a distributed training mechanism can be used to use data collected from different tasks or environments in parallel for training, thereby improving training efficiency. The learning rate can also be dynamically adjusted to adapt the convergence speed based on the changing trend of the reward function.
[0084] Example: In healthcare scenarios, rehabilitation robots update the parameters of the flow matching network based on patient action feedback and the doctor's preset rehabilitation progress reward function, making the subsequent training process more aligned with the patient's recovery progress and needs.
[0085] In fintech business scenarios, intelligent financial assistants update the parameters of the flow matching network based on user interaction feedback and multi-dimensional user satisfaction reward functions, thereby improving the personalization level of subsequent recommendation services and user stickiness.
[0086] In intelligent robot home services, service robots adjust the parameters of the flow matching network based on changes in the home environment and the reward function set by the user for task completion, thereby optimizing the accuracy and efficiency of action distribution mapping in the next round of cleaning tasks.
[0087] This embodiment combines environmental interaction feedback with a preset reward function for parameter updates, enabling the flow matching network to adapt to constantly changing environments, improving the network's generalization ability and long-term learning performance on complex tasks, achieving better action distribution mapping, and continuously improving decision quality and environmental adaptability.
[0088] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as fintech, healthcare, and intelligent robots. It discloses a decision-making method, apparatus, device, and medium based on an environmental semantic graph, comprising: acquiring multimodal input data and extracting and fusing it to generate initial multimodal fusion features; constructing and dynamically updating an environmental semantic graph based on entity and relationship information contained in the initial multimodal fusion features; fusing the environmental semantic graph and the initial multimodal fusion features to generate deep fusion features; determining the current action distribution based on the deep fusion features; using a flow matching network to learn the transmission mapping from the current action distribution to a preset target action distribution to generate a mapped action distribution; generating an action sequence and adjusting the action sequence according to real-time environmental interaction feedback; and updating the network parameters of the flow matching network by combining environmental interaction feedback and a preset reward function. This invention, by combining an environmental semantic graph with multimodal fusion features and utilizing a flow matching network to optimize the decision path and dynamically adjust the action sequence at the action distribution level, can accurately understand environmental semantic relationships, improve multimodal information processing and decision-making performance in complex dynamic environments, and achieve better action sequence decision results.
[0089] In one embodiment, step S10 above includes:
[0090] S101 acquires visual data through a camera, acquires environmental structure data through a lidar, receives user command text as language command data, and acquires historical data of action execution.
[0091] S102, The visual data is processed by the object detection model and the instance segmentation model to generate object category and location information and object outline information, respectively;
[0092] S103, The object category and location information and the object contour information are processed by a convolutional neural network to generate a visual feature representation;
[0093] S104, The language instruction data is processed through a semantic encoding model and a dependency parsing model to generate word vector representations and syntactic structure information, respectively;
[0094] S105, uses a temporal convolutional network to process historical action execution data and generate local action features;
[0095] S106, The local action features are processed by a recurrent neural network to generate an action history feature representation;
[0096] S107, The environmental structure data is processed by a point cloud processing model to generate spatial location features;
[0097] S108, The spatial location features are processed through a spatial coding model to generate a three-dimensional structural representation;
[0098] S109, the visual feature representation, the word vector representation, the grammatical structure information, the action history feature representation, and the three-dimensional structure representation are fused to generate an initial multimodal fusion feature.
[0099] In this embodiment, the operation of acquiring multimodal input data includes the simultaneous acquisition of data from multiple sources. Visual data comes from a high-resolution camera installed on the system or device. The camera acquires a sequence of two-dimensional images of target objects in the environment. This process supports different lighting conditions and viewing angles to ensure the diversity and completeness of the captured information. Environmental structure data is acquired through LiDAR. LiDAR uses laser scanning technology to collect depth information of objects and space in the environment in real time, reflecting the three-dimensional spatial structure in the form of point cloud data. User command text, as language command data, is collected through natural language input devices such as microphones or terminal input interfaces to ensure that the commands can cover the rich semantics of natural language expression. Action execution history data is collected through built-in sensors or operation recording modules, reflecting the temporal trajectory and execution characteristics of historical operation behaviors.
[0100] The processing of visual data involves the collaborative application of object detection models and instance segmentation models. Object detection models identify object categories and their locations in images, outputting category labels and coordinate bounding boxes. Instance segmentation models further segment the fine contours of individual objects, providing pixel-level boundary information. Convolutional neural networks (CNNs) jointly process the object category and location information from object detection and the object contour information from instance segmentation. CNNs efficiently extract spatial features through local receptive fields and shared weight mechanisms, generating a compact representation of the semantic and spatial understanding of the visual scene, thus forming a visual feature representation.
[0101] For the processing of language instruction data, the semantic coding model is used to convert the original text sequence into the corresponding word vector representation. The word vector captures the semantic relevance between words through the embedding space. The dependency analysis model parses the syntactic structure in the text to form grammatical structure information, revealing the grammatical dependency paths such as subject-verb-object relations and modification relations in the instructions, ensuring accurate parsing of user instructions.
[0102] The processing of historical action data adopts a temporal convolutional network. The temporal convolutional network extracts local features from continuous action segments in the time series through one-dimensional convolution operations to generate local action features. Then, a recurrent neural network processes the local action features. The recurrent neural network has the ability to process sequence data. Through its recurrent structure, it models the long-term dependencies in the action sequence and outputs a representation of historical action features, providing an overall dynamic context description of the historical behavior sequence.
[0103] Point cloud processing models are used to process environmental structure data acquired by LiDAR. These models extract spatial location features from 3D point clouds using methods such as spatial sampling and feature aggregation, capturing the geometric relationships of objects and structures within the scene. Based on these spatial location features, spatial coding models encode the spatial relationships and local geometric structures between points in the point cloud, generating a 3D structural representation that provides accurate spatial context information for subsequent fusion.
[0104] Finally, the processed visual feature representations, word vector representations, syntactic structure information, action history feature representations, and 3D structural representations are jointly represented through a fusion module for learning. This fusion module can employ methods such as multilayer perceptrons, attention mechanisms, or tensor fusion to map features from different modalities to a shared semantic space, integrating multimodal information in a unified feature space to generate initial multimodal fused features. These fused features possess a unified ability to represent vision, language, action history, and spatial environment, and can serve as input for subsequent modules, supporting multidimensional perception and understanding in the decision-making process.
[0105] This embodiment acquires visual data, environmental structure data, language command data, and historical action execution data in parallel through multiple channels, including cameras, LiDAR, user input interfaces, and historical data acquisition modules. It utilizes various advanced models and algorithms, such as object detection, instance segmentation, convolutional neural networks, semantic coding, dependency parsing, temporal convolution, recurrent neural networks, point cloud processing, and spatial coding, combined with a fusion module for unified representation learning, forming an initial multimodal fusion feature with rich semantic expressive capabilities. In this way, visual scene, language commands, spatial geometry, and historical behavior information can be simultaneously captured in multimodal perception and fused within a unified feature space. This significantly improves the ability of subsequent processing to understand and utilize multi-source heterogeneous data in complex environments, solving the problems of information fragmentation, missing context, and inconsistent expression caused by isolated processing of multimodal data in existing technologies. It achieves close semantic and spatial associations across modalities, providing globally consistent and detailed input data representation for environmental perception, task execution, and intelligent decision-making.
[0106] In one embodiment, step S20 above includes:
[0107] S201, extract object entity information, semantic entity information, spatial region entity information, spatial position relationship information, semantic logical relationship information, and operational constraint relationship information from the initial multimodal fusion features;
[0108] S202, generate environmental semantic graph nodes based on the object entity information, the semantic entity information and the spatial region entity information;
[0109] S203, generate environmental semantic graph edges based on the spatial location relationship information, the semantic logical relationship information, and the operational constraint relationship information;
[0110] S204, combine the environmental semantic graph nodes and the environmental semantic graph edges to generate an environmental semantic graph;
[0111] S205, monitor the state changes of the object entity information, the semantic entity information and the spatial region entity information, and add new environmental semantic graph nodes in response to the monitored state changes;
[0112] S206, monitor the changes in the spatial location relationship information, the semantic logic relationship information, and the operational constraint relationship information, and update the environmental semantic graph edges in the environmental semantic graph;
[0113] S207, Update the feature representation of the environmental semantic graph nodes in the environmental semantic graph using a graph neural network.
[0114] In this embodiment, when extracting object entity information from the initial multimodal fusion features, the object entity information includes the identified specific object category and its unique identifier. Combined with the previous visual feature representation results, the object category is associated with attributes such as spatial location and size to represent an independent physically existing object. Semantic entity information comes from keywords and semantic fragments in language instruction data and historical action descriptions, representing abstract concepts or functional attributes with task-related meaning, such as "cleaning object" or "placement area". Spatial region entity information refers to the results of regional division in the environment identified from environmental structure data, usually obtained through spatial clustering, plane fitting, etc., such as a room, desktop, or designated work area.
[0115] Extracting spatial location relationship information involves structuring the relative spatial positions of entities, such as proximity, containment, and intersection. Semantic logical relationship information is determined by analyzing the grammatical structure of language instructions in multimodal input and the task context, identifying semantic instruction associations such as "the object to be cleaned and the target area." Operational constraint relationship information is represented by cross-analysis of action history data and environmental spatial encoding, indicating the physical or logical preconditions for certain actions, such as state constraints like objects being immovable or areas currently occupied.
[0116] Environmental semantic graph nodes are generated based on object entity information, semantic entity information, and spatial region entity information. Each node is represented as a vertex in the graph structure by a node encoder. Each node carries corresponding attributes, such as category, location, size, and semantic label. The uniqueness of a node is determined by its physical location and semantic context.
[0117] The environment semantic graph edges are generated based on spatial location relationship information, semantic logical relationship information, and operational constraint relationship information. The edges are represented by an edge encoder as directed or undirected relationships connecting nodes. The attributes of the edges include relationship type, weight, and timeliness marker. The environment semantic graph is generated by combining nodes and edges. The entire graph structure is stored in the form of an adjacency list or sparse matrix to efficiently represent entities and their relationships.
[0118] It monitors state changes in object entity information, semantic entity information, and spatial region entity information, involving real-time detection of newly appearing, disappearing, or significantly relocated objects or regions. It triggers new node operations through state comparison mechanisms and threshold judgment rules, such as introducing new object categories as new graph nodes. It also monitors changes in spatial location relationship information, semantic logical relationship information, and operational constraint relationship information, involving updates to relationship edges, such as changes in spatial proximity relationships caused by object movement or adjustments to logical relationships caused by task state updates.
[0119] The feature representations of nodes in the environmental semantic graph are updated by using graph neural networks. The message passing mechanism of graph neural networks calculates the contextual feature representations of nodes in multi-hop neighbor aggregation. The node embedding is calculated by combining the attributes of neighbor nodes and the edge weights, so that the updated node features can dynamically reflect the semantic context and local structure of the current global environment.
[0120] This embodiment extracts and structures information such as objects, semantics, and space from the initial multimodal fusion features in a unified manner. It constructs a graph structure with entities as nodes and relationships as edges, dynamically monitors and updates the state of nodes and edges, and uses graph neural networks to propagate and aggregate contextual information in the structure. Ultimately, it achieves unified modeling and efficient updating of multidimensional dynamic elements in the environment, solving the problem of existing technologies lacking global correlation and real-time change perception between multimodal information in complex environments. This enables the system to accurately characterize the semantic and spatial structure of the environment, providing a real-time consistent and highly expressive input representation for subsequent decision-making processes.
[0121] In one embodiment, step S30 above includes:
[0122] S301, the environmental semantic graph is encoded through a graph convolutional network to generate a graph embedding representation;
[0123] S302, The initial multimodal fusion features are processed by a feature transformation network to generate transformed features;
[0124] S303, determine the attention weights between the graph embedding representation and the transformation features;
[0125] S304, the transformation features are weighted and aggregated according to the attention weights to generate aggregated features, and the aggregated features are used as deep fusion features.
[0126] In this embodiment, when the system needs to handle multimodal scenarios with highly complex environmental semantics, simple feature splicing and static fusion cannot fully utilize the potential contextual associations and dynamic dependencies between different modalities. Therefore, the operation first uses a graph convolutional network to encode the environmental semantic graph. The environmental semantic graph consists of nodes and edges. Nodes include object entities, semantic entities, spatial region entities, etc., while edges express the spatial positional relationships, semantic logical relationships, and operational constraints between nodes. The graph convolutional network iteratively aggregates the information of node neighbors, combines the topological structure and node feature attributes, and propagates and updates the node representation layer by layer. The computation form of each graph convolution layer is to extract local and global semantic context information by multiplying the adjacency matrix and the trainable weight matrix, superimposing activation functions and normalization processing, and finally outputting a graph embedding representation matrix, where each row of the matrix represents a multidimensional embedding vector of a node.
[0127] After obtaining the graph embedding representation, the system performs feature transformation operations on the initial multimodal fusion features in parallel. The feature transformation network uses a multilayer perceptron or convolutional units to map the initial multimodal fusion features into the same vector space as the graph embedding representation, enabling the two types of features to be aligned in both numerical and semantic dimensions. This process, through weight matrix training, allows visual features, linguistic features, and spatial structure features to be compared and interact with each other in the same embedding space, thereby eliminating statistical distribution differences between modalities in subsequent interaction mechanisms.
[0128] The system further employs an attention weight calculation module to calculate the matching degree between the graph embedding representation and the transformed features, addressing the cross-spatial and cross-semantic coupling. Specifically, this involves using a scaled dot product attention mechanism, treating the graph embedding representation as the query vector and the transformed features as keys and values. Matching scores are calculated through matrix multiplication and scaling operations, and then normalized using Softmax to obtain the attention weights between each pair of features. These attention weights not only measure the association strength between a single node and a single multimodal feature but also express the conditional dependence of global graph information on the multimodal landscape, possessing the ability to dynamically adjust to the local spatial context.
[0129] After obtaining the attention weights, the system performs a weighted aggregation operation on the transformed features based on these weights. The aggregation operation maps each transformed feature to a corresponding attention weight, calculates a weighted sum, and forms the feature representation for each channel after fusion. This allows the aggregated features to fully express the global adjustment and local correction of the original multimodal perception data by the environmental information in the current semantic graph. Residual connections can be further superimposed during the weighted aggregation process, enabling the output to enhance contextual coupling while maintaining the integrity and stability of the original multimodal features, preventing gradient vanishing or overfitting.
[0130] Finally, the aggregated features are output as deep fusion features. These deep fusion features are represented in high-dimensional tensor form, containing rich spatial, semantic, and action context dependencies, providing a unified and fully expressive input for subsequent action distribution determination and dynamic adjustment. The output of deep fusion features not only ensures the integrated processing of multimodal data and graph structure information but also enhances the model's global modeling ability for dynamic environmental changes, complex object semantic relationships, and task intent.
[0131] Throughout the entire operation process, graph convolutional network encoding, feature transformation network processing, attention weight calculation, and weighted aggregation operations can be efficiently implemented on a parallel hardware architecture, ensuring real-time requirements. At the same time, it is easy to adjust and adapt to different embedding dimensions and network depths, supporting deployment in multiple scenarios such as financial risk control systems, medical image analysis, and autonomous decision-making for intelligent robots, and possessing high scalability and applicability.
[0132] This embodiment introduces graph convolutional network encoding, cross-modal feature alignment, and attention-based dynamic weighted aggregation operations between multimodal perception and environmental semantic understanding. This significantly enhances the deep fusion features across four dimensions: spatial, semantic, contextual, and multimodal representation. It effectively overcomes the shortcomings of traditional multimodal fusion, which relies solely on static feature concatenation, leading to contextual information loss and insufficient representation. This allows for a more comprehensive understanding of objects, relationships, and tasks within dynamically changing environments. Figure 3 The coupling modeling capability between the system and the user is improved. Through graph convolution and attention weighting mechanisms, the system can dynamically adjust the importance and contextual adaptability of perceptual features in different scenarios, thereby improving the accuracy and safety of subsequent decision-making processes in response to complex environments and task semantics.
[0133] In one embodiment, step S40 above includes:
[0134] S401, Construct a policy network that includes a deep neural network;
[0135] S402, The deep fusion features are processed by the policy network to generate a set of action probability values;
[0136] S403, establish an action probability distribution based on the set of action probability values, and define the action probability distribution as the current action distribution;
[0137] S404, Set the target action distribution according to the mission objective and environmental constraints;
[0138] S405, construct a flow matching network that includes a reversible transformation layer;
[0139] S406, determine the distribution distance between the current action distribution and the target action distribution;
[0140] S407, Minimize the distribution distance through the flow matching network to learn the transmission mapping;
[0141] S408, apply the transfer mapping to the current action distribution to generate a mapped action distribution.
[0142] In this embodiment, starting with deep fusion features, a policy network is first established. This policy network employs a deep neural network architecture, utilizing deep perceptual units and nonlinear activation functions to perform high-dimensional abstraction of the input deep fusion features. These deep fusion features include globally coupled information from visual semantics, spatial structure, historical actions, and the environmental graph. The weights of the policy network are trained through supervised learning and reinforcement learning, enabling it to adapt to multi-dimensional input patterns. The output of the policy network is a set of action probability values, encompassing discrete or continuous action units that the system can execute. Each value represents the confidence level at which the action is selected. Through standardization, these probability values are normalized to a set that meets the probability distribution requirements, further establishing the action probability distribution, which is defined as the current action distribution. The current action distribution characterizes the system's global tendency to select different actions under the current environmental state and input context.
[0143] To optimize action decision-making performance under given task objectives and environmental constraints, a target action distribution needs to be introduced. This distribution is modeled using a preset task strategy or historical best strategy experience, reflecting the probabilistic patterns of the global action configuration required for overall task completion. The target action distribution is set through a target design module, taking into account the optimal strategy prior under multi-dimensional constraints such as task completion time, operational costs, and environmental safety.
[0144] To optimize the mapping between the current action distribution and the target action distribution, a flow matching network is constructed in the process. The flow matching network is an invertible mapping neural network based on probability density modeling, belonging to the Flow-Matching framework. Its core objective is to learn a smooth, invertible mapping function between the current and target action distributions to capture the optimal transmission path between the two distributions. Internally, the flow matching network employs an invertible transformation layer structure, including an affine coupling transformation module, a permutation operation module, and an invertible normalization unit module. These modules collectively construct a highly expressive invertible neural network. The affine coupling transformation module performs conditional affine transformations on a portion of the input distribution's dimensions; the permutation operation module shuffles the variable order to improve model complexity; and the invertible normalization unit module ensures that the input and output data are standardized while preserving invertibility. Through this combined structure, the flow matching network can achieve continuous and smooth transformations of the high-dimensional distribution space without losing information, gradually adjusting the current action distribution to approach the preset target action distribution. During the training process of the flow matching network, the transmission mapping is optimized by minimizing the distribution distance between the current action distribution and the target action distribution, such as the Wasserstein distance, so that the final generated action distribution has global optimality and dynamic adaptability.
[0145] To achieve optimal mapping, the distance calculation module first measures the distribution distance between the current action distribution and the target action distribution. This distribution distance can be achieved using Wasserstein distance, Kullback-Leibler divergence, or other rigorously defined probabilistic distribution metrics. The resulting distance quantifies the degree of deviation of the current action decision. By using the distribution distance as the loss function, the flow matching network optimizes the parameters through backpropagation, minimizing the distance to gradually approximate the target action distribution, thus completing the learning of the transport mapping.
[0146] Finally, the learned transport mapping is applied to the current action distribution to obtain the mapped action distribution. This mapped action distribution not only closely approximates the target action distribution at the probabilistic level but also retains the dynamic characteristics of the current environment, achieving a balance between local adaptation and global optimization. Throughout the process, the structural design and parameter tuning of the deep neural network and the flow matching network support hardware acceleration and multi-task parallel training, ensuring efficient processing of complex input features and dynamic environment states in real-time environments.
[0147] This embodiment uses deeply fused features as input to build a policy network and generate the current action distribution. This allows for the full utilization of multimodal environment perception information. Simultaneously, a flow matching network is used to perform precise optimization learning on the current action distribution based on distribution distance, making the output mapped action distribution closer to the globally optimal decision configuration under the conditions of task objectives and environmental constraints. The overall processing flow not only achieves a closed-loop process from local state perception to globally optimal decision-making but also adapts to dynamic environmental changes, improving the decision robustness of action sequences and task completion efficiency.
[0148] In one embodiment, step S50 above includes:
[0149] S501, an initial action sequence is generated by sampling from the mapped action distribution;
[0150] S502, execute the initial action sequence and obtain environmental interaction feedback in real time;
[0151] S503, integrate the environmental interaction feedback with the deep fusion features to generate updated fusion features;
[0152] S504, The updated fusion features are input into the flow matching network to generate the updated action distribution;
[0153] S505, adjust the action sequence based on the updated action distribution.
[0154] In this embodiment, an initial action sequence is generated by sampling from the mapped action distribution. The mapped action distribution provides the probability weights of each candidate action in the current context. These weights reflect the joint modeling results of the multimodal environment and task constraints previously performed using a flow matching network. During this sampling process, importance sampling or random sampling methods can be used to ensure that the probabilistic characteristics of the action sequence are consistent with those of the mapped action distribution, avoiding overfitting or overexploration. Each element of the initial action sequence corresponds to an executable atomic operation, and the operation sequence is strictly arranged according to the sampling order, ensuring traceability and consistency in subsequent decisions.
[0155] During the execution of the initial action sequence, the motion control module maps discrete or continuous action units into specific control signals required by the execution platform in real time. During this process, environmental feedback is acquired through a multimodal sensor array. This feedback includes, but is not limited to, physical data such as visual, depth, and force data collected by external sensors, as well as system status and action execution results monitored by internal sensors. This feedback not only reflects the results of action execution but also implies changes in the environmental state and the sources of uncertainty.
[0156] After collecting environmental interaction feedback, information fusion processing is immediately performed, combining the environmental interaction feedback with deep fusion features to generate updated fusion features. This step, through feature splicing and multi-channel normalization, preserves the multimodal contextual relationships in the original deep fusion features while seamlessly introducing real-time environmental change information into the updated representation space, enabling data representation to take into account both global environmental modeling and the dynamism of current context awareness.
[0157] The updated fusion features are fed as input into the pre-trained flow matching network, which recalculates the current action distribution. During this process, the network's internal parameters remain fixed, and the output distribution is updated only under the drive of the input features. The updated action distribution instantly reflects the impact of changes in the current environment on the optimal action strategy. This step enables action decision-making to have rapid convergence and environmental adaptability.
[0158] The action sequence is adjusted based on the updated action distribution. Specifically, this adjustment involves mechanisms such as probabilistic resampling, dynamic adjustment of action sequence length, and local subsequence optimization to replace or rearrange some or all action units in the sampled action sequence. The adjusted action sequence not only conforms to the probabilistic characteristics of the updated action distribution but also adapts to changes in environmental conditions with minimal action perturbation, reducing unnecessary action overhead and execution risks.
[0159] This embodiment achieves closed-loop adaptive action decision-making and environmental change response based on multimodal perception by sampling based on the mapped action distribution, acquiring real-time environmental interaction feedback, updating fused features through feedback information, and regenerating the action distribution through a flow matching network. Finally, it adjusts the action sequence. This process not only improves the global optimality of the action sequence but also possesses dynamic robustness. Under conditions of uncertain environmental states, complex task semantics, or execution process disturbances, it can promptly adjust decision-making strategies, thereby improving task success rate and execution efficiency.
[0160] In one embodiment, step S60 above includes:
[0161] S601, design a multi-dimensional reward function that includes task completion indicators, efficiency indicators and safety indicators, and determine the reward data in the environmental interaction process based on the multi-dimensional reward function;
[0162] S602, collect state data, action sequence data and reward data during the environmental interaction process;
[0163] S603, determine the advantage function estimate based on the state data, action sequence data and reward data;
[0164] S604, Calculate the strategy loss based on the advantage function estimation and the preset strategy loss function;
[0165] S605, the network parameters of the flow matching network are updated using the backpropagation algorithm to minimize the policy loss, and an updated flow matching network is generated.
[0166] In this embodiment, a multi-dimensional reward function, comprising task completion metrics, efficiency metrics, and safety metrics, is first designed to clarify the quantitative standards for action decisions and task goal achievement. The task completion metric measures whether the task is fully completed after the action sequence is executed, such as the success rate of a material handling task. The efficiency metric measures execution speed or energy consumption, such as the minimization of total task time or resource consumption. The safety metric reflects the degree of impact of the behavior on the environment and the individual, such as the avoidance of collision risks or overload states. Each metric is mapped to a reward value in functional form. The reward function is weighted by weight parameters, which can be set according to domain requirements to adapt to different task scenarios. Based on the defined multi-dimensional reward function, reward values are calculated from continuously sampled state sequences, action sequences, and environmental response data during environmental interaction. The reward data serves as feedback on the effectiveness of the current strategy execution in the environment.
[0167] The collected data includes state data, action sequence data, and reward data. State data reflects the dynamic characteristics of the current environment and interactive objects, such as spatial position, velocity, and posture. Action sequence data consists of discrete or continuous action streams previously output by the decision module and actually executed, including action type, execution parameters, and execution timestamps. Reward data is calculated in real time by combining the current state and action execution effect with a preset reward function. Data acquisition uses a streaming storage mechanism to ensure data consistency and support subsequent gradient calculations.
[0168] The advantage function estimation is based on state data, action sequence data, and reward data. The advantage function quantifies the relative value of action selection by comparing the expected return of taking a specific action at a given moment with the average return under the current policy. The calculation employs time difference methods and discounted cumulative returns. The advantage function estimate is expressed as the sample mean or sliding window mean. High-frequency noise can be reduced and the stability and representativeness of the estimate improved through techniques such as Gaussian smoothing or exponentially weighted averaging.
[0169] The strategy loss calculation is based on the advantage function estimation and the preset strategy loss function. The strategy loss function usually adopts the form of proximal policy optimization loss to ensure that the parameter changes are smooth during the policy update process and avoid policy degradation caused by excessive update amplitude. The calculation process involves the ratio between the action probability output by the current policy and the output probability of the old policy, the upper and lower bounds of the pruning interval, and the product of the advantage function. The entire process strictly maintains the consistency between the action distribution probability structure and the policy feasible region constraints.
[0170] Finally, the backpropagation algorithm is employed to update the network parameters of the flow matching network, aiming to minimize the policy loss. During the update process, the chain rule is used to propagate gradients to the parameters of each layer of the network, and an adaptive learning rate adjustment algorithm, such as Adam or RMSProp, is used to control the update magnitude, ensuring convergence speed and training stability. The updated network parameters are stored in real time, forming the updated flow matching network for the current cycle, ensuring that the latest policy model is used in subsequent action distribution generation processes.
[0171] Example: In the field of healthcare, an intelligent medical robot system is used in surgical assistance scenarios. Its task is to deliver instruments, locate surgical sites, and make dynamic adjustments based on the doctor's natural language instructions, intraoperative visual and structural data, and historical surgical records.
[0172] The medical robot first acquires multimodal input data, including visual data captured by a high-resolution intraoperative camera, environmental structural data of the operating table and patient's anatomical area captured by a 3D laser scanner, natural language text data transcribed from the doctor's verbal instructions, and the patient's corresponding historical surgical operation records. The visual data is processed by an object detection model to detect surgical instruments, patient anatomical features, and surgical area boundaries, and an instance segmentation model to extract instrument and tissue contours. Object categories, positions, and contour features are encoded using a convolutional neural network to generate visual feature representations. Language instruction data is converted into word vectors by a semantic encoding model, and a dependency parsing model parses the structural information in the instructions. Historical surgical records are processed by a temporal convolutional network to form local action pattern features, which are then refined into action history features by a recurrent neural network. Environmental structural data has spatial location features extracted by a point cloud processing model and generated into a three-dimensional spatial representation by a spatial encoding network. All these features are fused to form initial multimodal fusion features, used to describe the comprehensive information state of the patient's surgical area, instruction semantics, historical behavior, and environmental space.
[0173] The medical robot constructs and dynamically updates an environmental semantic graph based on anatomical entities (such as organs and blood vessels) and semantic relationships (such as "delivering scissors to the edge of the right lobe of the liver") identified from initial multimodal fusion features. By extracting physical entities, spatial regions, positional relationships, and semantic logical relationships, semantic graph nodes and edges are generated and combined in real time to form the semantic graph. Monitoring instrument status, changes in anatomical features, and intraoperative environmental adjustments, the robot dynamically adds nodes or updates edges to describe changes in relationships, such as newly labeled vascular branches due to bleeding or the current position of instruments. A graph neural network continuously updates the feature representations of the graph nodes, ensuring the graph closely follows environmental changes and reflects the state of the surgical scene in real time.
[0174] The medical robot encodes the semantic graph using a graph convolutional network to generate a graph embedding representation. Simultaneously, it maps the initial multimodal fusion features through a feature transformation network to generate transformed features. The system calculates the attention weights between the graph embedding and the transformed features to measure the relevance of different semantic nodes to relevant entities in the instructions. The transformed features are then weighted and aggregated to ultimately generate deep fusion features, which serve as input for downstream decision-making.
[0175] The robot outputs the current action distribution, such as a probability set of candidate actions like "delivering forceps" and "adjusting the camera angle," based on deep fusion features through a policy network. A target action distribution is set according to the patient's current anatomical state, the surgeon's operational goals, and safety constraints. A flow matching network is used to learn the optimal transmission mapping from the current action distribution to the target action distribution. The mapping is optimized by minimizing the distribution distance (e.g., Wasserstein distance) to make the mapped action distribution more aligned with the surgical goals, ultimately generating the mapped action distribution.
[0176] The robot samples from the mapped motion distribution to generate an initial motion sequence, such as "moving the robotic arm to the right lobe of the liver → adjusting the viewing angle → delivering the scissors." The robot begins executing the motion sequence and acquires real-time environmental feedback through visual and force sensors, such as instrument contact position deviation, viewing angle shift, or slight patient position movement. The system combines the environmental feedback with deep fusion features to form updated fusion features. The updated fusion features are re-inputted into the flow matching network to generate an updated motion distribution, thereby adjusting the motion sequence, such as "correcting the robotic arm position → compensating for the viewing angle," to achieve high-precision motion correction.
[0177] Throughout the surgical procedure, the medical robot combines environmental feedback with multi-dimensional reward metrics such as task completion rate (e.g., accurate delivery success rate), efficiency (e.g., action duration), and safety (e.g., instrument-tissue contact pressure) to calculate reward data. The robot collects complete state data, action sequence data, and reward data to determine an estimated advantage function, measuring the relative advantage of the current action selection. Based on the advantage function and the proximal policy optimization loss function, the strategy loss is calculated. A backpropagation algorithm is used to update the flow matching network parameters to minimize the strategy loss, thereby generating an updated flow matching network. The updated network is used to generate the next action distribution, enabling the medical robot to continuously adaptively optimize its decision-making strategy throughout the surgical process, improving operational accuracy, efficiency, and safety, and ensuring stable and reliable intraoperative operations.
[0178] In the fintech field, an intelligent investment advisory system is used for client asset allocation and risk management. Its task is to make multi-asset portfolio decisions, dynamically control risks, and adjust portfolios based on the client's natural language investment needs, client's historical investment behavior data, real-time market conditions, and market structure data.
[0179] The system first acquires multimodal input data, including multi-dimensional financial market data obtained in real-time through a market data interface, market structure data (such as market participant behavior distribution and trading book liquidity data) obtained through an API, user-submitted textual investment requests, and historical investment transaction records in the client's account. The system extracts market price sequence features and volatility characteristics from the market data using a time-series modeling module; the market structure data is processed by a graph structure modeling module to form a structured representation of the market state. User textual requests are parsed using a semantic understanding module to parse investment intent, obtaining word vector representations and syntactic structure information. The client's historical investment records are modeled using time series to generate historical investment behavior features. These features are then fused to form initial multimodal fusion features, comprehensively describing the client's goals, historical preferences, current market environment, and structural state.
[0180] The system constructs and dynamically updates an environmental semantic graph based on investment targets (such as stocks, bonds, and funds), market relationships (such as sector correlation and asset liquidity constraints), and user intent identified from initial multimodal fusion features. The system extracts asset entity information, industry sector entity information, liquidity region classification information, price correlations, semantic logic relationships of user needs, and regulatory compliance constraints to form nodes and edges of the environmental semantic graph, which are then combined into a complete graph. The system monitors market price fluctuations, sector rotation, and changes in customer intent, dynamically adding new nodes (e.g., adding popular assets or industries) and updating edges (e.g., enhancing asset correlation). A graph neural network continuously updates the feature representations of the graph nodes, ensuring that the semantic graph reflects real-time changes in market and customer needs.
[0181] The system encodes the environmental semantic graph using a graph convolutional network to generate a graph embedding representation. It then processes the initial multimodal fusion features through a feature transformation network to generate transformed features. Attention weights are calculated between the graph embedding and the transformed features to identify the market and asset relationship nodes that have the greatest impact on the customer's current needs. By weighted aggregation of the transformed features, a deep fusion feature is generated for decision input.
[0182] The system determines the current action distribution through a policy network based on deep fusion features, such as weight allocation suggestions for various assets in a multi-asset portfolio. Combining current market conditions, client investment objectives, and regulatory restrictions, the system sets a target action distribution and learns the transmission mapping from the current action distribution to the target action distribution through a flow matching network. By minimizing the distribution distance and optimizing the mapping, the final mapped action distribution better aligns with the client's overall investment goals.
[0183] The system samples the mapped action distribution to generate an initial portfolio recommendation sequence, such as "increase the weight of Equity A by 5% → decrease the weight of US Treasury bonds by 3% → increase the weight of gold ETFs by 2%". During the execution of portfolio recommendations, the system obtains real-time feedback on market changes, client account trading activity, and compliance checks. After fusing environmental feedback with deep fusion features, updated fusion features are generated and input into the stream matching network to regenerate the updated action distribution. This allows for adjustments to portfolio recommendations, such as reducing the weight of assets affected by sharp market fluctuations and increasing defensive asset allocations, thereby rebalancing portfolio risk.
[0184] In the decision-making loop, the system combines task completion metrics (such as reduction in portfolio risk exposure), efficiency metrics (such as decision and execution delays), and safety metrics (such as compliance risks) to design a multi-dimensional reward function and calculate reward data. The system collects complete market state data, adjusted recommendation sequence data, and reward data to determine the advantage function estimate. Based on the advantage function and the near-end strategy optimization loss function, it calculates the strategy loss and updates the flow matching network parameters through a backpropagation algorithm, generating an updated flow matching network. This network is used for future portfolio decisions, enabling the system to continuously adaptively optimize asset allocation recommendations through market and client interaction, improving the scientific rigor, robustness, and compliance of investment recommendations, and helping clients dynamically manage investment risks and returns.
[0185] In the field of intelligent robots, a service robot is used for task collaboration in complex home environments, such as responding to user voice commands to complete tasks like "clearing the dining table, tidying up the tableware and putting it away".
[0186] The robot first acquires multimodal input data, including visual data collected by an onboard camera, environmental structure data obtained by LiDAR, task commands issued by the user in natural language, and the robot's own historical action data within the environment. The visual data is processed by an object detection model and an instance segmentation model to generate object category and location information, as well as object contour information. A convolutional neural network further integrates this information to form visual feature representations. User commands are processed by a semantic encoding model and a dependency parsing model to obtain word vector representations and syntactic structure information. The robot's past action sequences are processed by a temporal convolutional network to form local action features, which are then refined into action history features by a recurrent neural network. A point cloud processing model analyzes the spatial relationships of the LiDAR data, and a spatial encoding module generates a three-dimensional structural representation. These multimodal data are fused to generate initial multimodal fusion features, comprehensively describing the robot's current physical environment, user intent, and historical interaction behavior.
[0187] Based on initial multimodal fusion features, the robot identifies entity information of objects in the environment, such as tables, tableware, and storage cabinets; semantic entity information of sorting in user commands; spatial area entity information (e.g., table area, storage area); spatial positional relationships (e.g., "bowls on the table"); semantic logical relationships (e.g., "put them back in their place after sorting"); and operational constraint relationships (e.g., "avoid collisions"). These entities and relationships are used to generate nodes and edges of the environmental semantic graph, which are combined to form the current environmental semantic graph. The robot monitors changes in object state (e.g., tableware being moved) and changes in spatial relationships (e.g., the user temporarily placing new objects) in real time, dynamically updating the nodes and edges of the environmental semantic graph. Simultaneously, it updates the feature representation of each node through a graph neural network to ensure that the semantic graph is consistent with the state of the physical world.
[0188] The robot encodes the semantic graph of the environment using a graph convolutional network to obtain a graph embedding representation, and generates transformed features from the initial multimodal fusion features using a feature transformation network. The robot calculates the attention weights between the two, identifies key objects and relationship nodes in the current task, uses them as the focus of weighted aggregation, obtains aggregated features, and uses them as deep fusion features.
[0189] Based on deep fusion features, the robot generates a current action distribution through a policy network, representing the probability of various actions (e.g., grasping, carrying, placing) in the current state. A target action distribution (e.g., efficiently and safely completing a tidying task) is set according to the user's goals and home environment constraints. A flow matching network minimizes the distribution distance between the current action distribution and the target action distribution, learns a transport mapping, and applies it to the current action distribution to generate a mapped action distribution, which guides the robot's action decisions.
[0190] The robot samples from the mapped action distribution to generate an initial action sequence, such as "move to the dining table → grab a bowl → place it in the cupboard → move to the dining table → grab a plate → place it in the dish cabinet." During the execution of the action sequence, the robot collects real-time environmental feedback, including visually perceived object states, obstacle detection information, and feedback from robot joint movements. The robot combines this feedback with deep fusion features to generate updated fusion features. The input stream matching network updates the action distribution and dynamically adjusts the action sequence based on the updated distribution. For example, it avoids clutter temporarily placed in the path by the user or changes the grasping strategy based on the risk of objects slipping, continuously improving the safety and efficiency of task execution.
[0191] The entire decision-making process combines task completion metrics (e.g., tableware return integrity rate), efficiency metrics (e.g., task completion time), and safety metrics (e.g., collision event rate) to design a multi-dimensional reward function and calculate reward data in real time. The robot collects state data, action sequence data, and reward data during the task process, determines the advantage function estimate, calculates the policy loss, and updates the flow matching network parameters through the backpropagation algorithm, generating an updated flow matching network. This updated network enables the robot to better understand complex home environments, multimodal task instructions, and dynamic environmental changes in subsequent tasks, gradually improving task completion quality and user experience, achieving an evolutionary process from single execution to adaptive optimization.
[0192] This embodiment introduces a multi-dimensional reward function and calculates reward data within the closed loop of action execution and environmental interaction feedback. It then combines state, action, and reward data to calculate an advantage function and updates network parameters based on the policy loss function, achieving adaptive online learning capabilities for the flow matching network. This process enables action decisions to be progressively optimized during continuous interaction. Under dynamic environments and complex multi-objective tasks, through reward-driven mechanisms, advantage function evaluation, and loss minimization, it improves the overall robustness of the action policy, task achievement rate, and resource utilization efficiency.
[0193] In one embodiment, a decision-making device based on an environmental semantic graph is provided, which corresponds one-to-one with the decision-making method based on an environmental semantic graph in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the decision-making device based on environmental semantic graph of the present invention. The modules include a multimodal feature processing module 10, an environmental semantic graph module 20, a feature fusion module 30, a flow matching decision module 40, an action adjustment module 50, and a strategy optimization module 60. Detailed descriptions of each functional module are as follows:
[0194] The multimodal feature processing module 10 is used to acquire multimodal input data and perform feature extraction and fusion on the multimodal input data to generate initial multimodal fusion features;
[0195] The environmental semantic graph module 20 is used to construct and dynamically update the environmental semantic graph based on the entity and relationship information contained in the initial multimodal fusion features.
[0196] Feature fusion module 30 is used to fuse the environmental semantic map with the initial multimodal fusion features to generate deep fusion features;
[0197] The flow matching decision module 40 is used to determine the current action distribution based on the deep fusion features, and learn the transmission mapping from the current action distribution to the preset target action distribution through the flow matching network to generate the mapped action distribution;
[0198] The action adjustment module 50 is used to generate an action sequence based on the mapped action distribution, and to adjust the action sequence based on real-time environmental interaction feedback.
[0199] The strategy optimization module 60 is used to update the network parameters of the flow matching network based on the environmental interaction feedback and the preset reward function.
[0200] In one embodiment, the multimodal feature processing module 10 is specifically used for:
[0201] The system acquires visual data through cameras, environmental structure data through LiDAR, receives user command text as language command data, and acquires historical data of action execution.
[0202] The visual data is processed by an object detection model and an instance segmentation model to generate object category and location information and object outline information, respectively.
[0203] The object category and location information and the object outline information are processed by a convolutional neural network to generate a visual feature representation.
[0204] The language instruction data is processed by a semantic encoding model and a dependency parsing model to generate word vector representations and syntactic structure information, respectively.
[0205] Historical action execution data is processed through a temporal convolutional network to generate local action features;
[0206] The local action features are processed by a recurrent neural network to generate an action history feature representation.
[0207] The environmental structure data is processed using a point cloud processing model to generate spatial location features;
[0208] The spatial location features are processed using a spatial coding model to generate a three-dimensional structural representation.
[0209] The visual feature representation, word vector representation, grammatical structure information, action history feature representation, and three-dimensional structure representation are fused to generate an initial multimodal fusion feature.
[0210] In one embodiment, the environmental semantic graph module 20 is specifically used for:
[0211] Extract object entity information, semantic entity information, spatial region entity information, spatial positional relationship information, semantic logical relationship information, and operational constraint relationship information from the initial multimodal fusion features;
[0212] Based on the object entity information, the semantic entity information, and the spatial region entity information, an environmental semantic graph node is generated;
[0213] An environmental semantic graph edge is generated based on the spatial location relationship information, the semantic logical relationship information, and the operational constraint relationship information.
[0214] Combine the environmental semantic graph nodes and the environmental semantic graph edges to generate an environmental semantic graph;
[0215] Monitor the state changes of the object entity information, the semantic entity information, and the spatial region entity information, and add new environmental semantic graph nodes in response to the detected state changes;
[0216] Monitor changes in the spatial location relationship information, the semantic logical relationship information, and the operational constraint relationship information, and update the environmental semantic graph edges in the environmental semantic graph.
[0217] The feature representations of the environmental semantic graph nodes in the environmental semantic graph are updated using a graph neural network.
[0218] In one embodiment, the feature fusion module 30 is specifically used for:
[0219] The environmental semantic graph is encoded using a graph convolutional network to generate a graph embedding representation;
[0220] The initial multimodal fusion features are processed by a feature transformation network to generate transformed features;
[0221] Determine the attention weights between the graph embedding representation and the transformed features;
[0222] The transformation features are weighted and aggregated according to the attention weights to generate aggregated features, and the aggregated features are used as deep fusion features.
[0223] In one embodiment, the flow matching decision module 40 is specifically used for:
[0224] Construct a policy network that includes a deep neural network;
[0225] The deep fusion features are processed through the policy network to generate a set of action probability values;
[0226] An action probability distribution is established based on the set of action probability values, and the action probability distribution is defined as the current action distribution;
[0227] The distribution of target actions is set according to the mission objectives and environmental constraints;
[0228] Construct a flow matching network that includes a reversible transformation layer;
[0229] Determine the distribution distance between the current action distribution and the target action distribution;
[0230] The distribution distance is minimized through the flow matching network to learn the transport mapping;
[0231] The transport mapping is applied to the current action distribution to generate the mapped action distribution.
[0232] In one embodiment, the motion adjustment module 50 is specifically used for:
[0233] An initial action sequence is generated by sampling from the mapped action distribution;
[0234] Execute the initial action sequence and obtain environmental interaction feedback in real time;
[0235] By integrating the environmental interaction feedback with the deep fusion features, an updated fusion feature is generated;
[0236] The updated fusion features are input into the flow matching network to generate the updated action distribution;
[0237] The action sequence is adjusted based on the updated action distribution.
[0238] In one embodiment, the strategy optimization module 60 is specifically used for:
[0239] Design a multi-dimensional reward function that includes task completion metrics, efficiency metrics, and safety metrics, and determine reward data during the environmental interaction process based on the multi-dimensional reward function;
[0240] Collect state data, action sequence data, and reward data during the environmental interaction process;
[0241] The advantage function estimate is determined based on the state data, action sequence data, and reward data.
[0242] The strategy loss is calculated based on the advantage function estimation and the preset strategy loss function;
[0243] The network parameters of the flow matching network are updated using the backpropagation algorithm to minimize the policy loss, thereby generating an updated flow matching network.
[0244] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements server-side functions or steps of a decision-making method based on an environmental semantic graph.
[0245] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5 As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides decision-making and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements user-side functions or steps of a decision-making method based on an environmental semantic graph.
[0246] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:
[0247] Acquire multimodal input data, and perform feature extraction and fusion on the multimodal input data to generate initial multimodal fusion features;
[0248] Based on the entity and relationship information contained in the initial multimodal fusion features, an environmental semantic graph is constructed and dynamically updated.
[0249] The environmental semantic map is fused with the initial multimodal fusion features to generate deep fusion features;
[0250] The current action distribution is determined based on the deep fusion features, and the transmission mapping from the current action distribution to the preset target action distribution is learned through the flow matching network to generate the mapped action distribution.
[0251] An action sequence is generated based on the mapped action distribution, and the action sequence is adjusted based on real-time environmental interaction feedback.
[0252] The network parameters of the flow matching network are updated based on the environmental interaction feedback and the preset reward function.
[0253] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0254] Acquire multimodal input data, and perform feature extraction and fusion on the multimodal input data to generate initial multimodal fusion features;
[0255] Based on the entity and relationship information contained in the initial multimodal fusion features, an environmental semantic graph is constructed and dynamically updated.
[0256] The environmental semantic map is fused with the initial multimodal fusion features to generate deep fusion features;
[0257] The current action distribution is determined based on the deep fusion features, and the transmission mapping from the current action distribution to the preset target action distribution is learned through the flow matching network to generate the mapped action distribution.
[0258] An action sequence is generated based on the mapped action distribution, and the action sequence is adjusted based on real-time environmental interaction feedback.
[0259] The network parameters of the flow matching network are updated based on the environmental interaction feedback and the preset reward function.
[0260] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0261] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0262] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0263] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A decision-making method based on environmental semantic graphs, characterized in that, Includes the following steps: Acquire multimodal input data, and perform feature extraction and fusion on the multimodal input data to generate initial multimodal fusion features; Based on the entity and relationship information contained in the initial multimodal fusion features, an environmental semantic graph is constructed and dynamically updated. The environmental semantic map is fused with the initial multimodal fusion features to generate deep fusion features; The current action distribution is determined based on the deep fusion features, and the transmission mapping from the current action distribution to the preset target action distribution is learned through the flow matching network to generate the mapped action distribution. An action sequence is generated based on the mapped action distribution, and the action sequence is adjusted based on real-time environmental interaction feedback. The network parameters of the flow matching network are updated based on the environmental interaction feedback and the preset reward function.
2. The decision-making method based on environmental semantic graph as described in claim 1, characterized in that, Acquire multimodal input data, and perform feature extraction and fusion on the multimodal input data to generate initial multimodal fusion features, including: The system acquires visual data through cameras, environmental structure data through LiDAR, receives user command text as language command data, and acquires historical data of action execution. The visual data is processed by an object detection model and an instance segmentation model to generate object category and location information and object outline information, respectively. The object category and location information and the object outline information are processed by a convolutional neural network to generate a visual feature representation. The language instruction data is processed by a semantic encoding model and a dependency parsing model to generate word vector representations and syntactic structure information, respectively. Historical action execution data is processed through a temporal convolutional network to generate local action features; The local action features are processed by a recurrent neural network to generate an action history feature representation. The environmental structure data is processed using a point cloud processing model to generate spatial location features; The spatial location features are processed using a spatial coding model to generate a three-dimensional structural representation. The visual feature representation, word vector representation, grammatical structure information, action history feature representation, and three-dimensional structure representation are fused to generate an initial multimodal fusion feature.
3. The decision-making method based on environmental semantic graphs as described in claim 1, characterized in that, Based on the entity and relation information contained in the initial multimodal fusion features, an environmental semantic graph is constructed and dynamically updated, including: Extract object entity information, semantic entity information, spatial region entity information, spatial positional relationship information, semantic logical relationship information, and operational constraint relationship information from the initial multimodal fusion features; Based on the object entity information, the semantic entity information, and the spatial region entity information, an environmental semantic graph node is generated; An environmental semantic graph edge is generated based on the spatial location relationship information, the semantic logical relationship information, and the operational constraint relationship information. Combine the environmental semantic graph nodes and the environmental semantic graph edges to generate an environmental semantic graph; Monitor the state changes of the object entity information, the semantic entity information, and the spatial region entity information, and add new environmental semantic graph nodes in response to the detected state changes; Monitor changes in the spatial location relationship information, the semantic logical relationship information, and the operational constraint relationship information, and update the environmental semantic graph edges in the environmental semantic graph. The feature representations of the environmental semantic graph nodes in the environmental semantic graph are updated using a graph neural network.
4. The decision-making method based on environmental semantic graph as described in claim 1, characterized in that, The environmental semantic map is fused with the initial multimodal fusion features to generate deep fusion features, including: The environmental semantic graph is encoded using a graph convolutional network to generate a graph embedding representation; The initial multimodal fusion features are processed by a feature transformation network to generate transformed features; Determine the attention weights between the graph embedding representation and the transformed features; The transformed features are weighted and aggregated according to the attention weights to generate aggregated features, and the aggregated features are used as deep fusion features.
5. The decision-making method based on environmental semantic graph as described in claim 1, characterized in that, Based on the deep fusion features, the current action distribution is determined, and a transmission mapping from the current action distribution to a preset target action distribution is learned through a stream matching network to generate the mapped action distribution, including: Construct a policy network that includes a deep neural network; The deep fusion features are processed through the policy network to generate a set of action probability values; An action probability distribution is established based on the set of action probability values, and the action probability distribution is defined as the current action distribution; The distribution of target actions is set according to the mission objectives and environmental constraints; Construct a flow matching network that includes a reversible transformation layer; Determine the distribution distance between the current action distribution and the target action distribution; The distribution distance is minimized through the flow matching network to learn the transport mapping; The transport mapping is applied to the current action distribution to generate the mapped action distribution.
6. The decision-making method based on environmental semantic graph as described in claim 1, characterized in that, An action sequence is generated based on the mapped action distribution, and the action sequence is adjusted based on real-time environmental interaction feedback, including: An initial action sequence is generated by sampling from the mapped action distribution; Execute the initial action sequence and obtain environmental interaction feedback in real time; By integrating the environmental interaction feedback with the deep fusion features, an updated fusion feature is generated; The updated fusion features are input into the flow matching network to generate the updated action distribution; The action sequence is adjusted based on the updated action distribution.
7. The decision-making method based on environmental semantic graph as described in claim 1, characterized in that, Based on the environmental interaction feedback and the preset reward function, the network parameters of the flow matching network are updated, including: Design a multi-dimensional reward function that includes task completion metrics, efficiency metrics, and safety metrics, and determine reward data during the environmental interaction process based on the multi-dimensional reward function; Collect state data, action sequence data, and reward data during the environmental interaction process; The advantage function estimate is determined based on the state data, action sequence data, and reward data. The strategy loss is calculated based on the advantage function estimation and the preset strategy loss function; The network parameters of the flow matching network are updated using the backpropagation algorithm to minimize the policy loss, thereby generating an updated flow matching network.
8. A decision-making device based on environmental semantic graphs, characterized in that, The decision-making device based on environmental semantic graphs includes: A multimodal feature processing module is used to acquire multimodal input data, and to extract and fuse features from the multimodal input data to generate initial multimodal fusion features; The environmental semantic graph module is used to construct and dynamically update the environmental semantic graph based on the entity and relationship information contained in the initial multimodal fusion features. The feature fusion module is used to fuse the environmental semantic map with the initial multimodal fusion features to generate deep fusion features; The flow matching decision module is used to determine the current action distribution based on the deep fusion features, and learn the transmission mapping from the current action distribution to the preset target action distribution through the flow matching network to generate the mapped action distribution; An action adjustment module is used to generate an action sequence based on the mapped action distribution, and to adjust the action sequence based on real-time environmental interaction feedback. The strategy optimization module is used to update the network parameters of the flow matching network based on the environmental interaction feedback and the preset reward function.
9. A computer device, characterized in that, The computer device includes a memory, a processor, and an environment semantic graph-based decision-making program stored in the memory and executable on the processor, wherein the environment semantic graph-based decision-making program, when executed by the processor, implements the steps of the environment semantic graph-based decision-making method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a decision-making program based on an environmental semantic graph, which, when executed by a processor, implements the steps of the decision-making method based on an environmental semantic graph as described in any one of claims 1-7.
Citation Information
Cited By
Smart home control method and system based on large language model and graph global optimization
CN121165521A