Robot action decision-making method and device based on spatial relationship, equipment and medium
By collecting and processing multimodal data to construct a three-dimensional spatial topology map and perform cross-modal alignment, the problem of lack of correlation in spatial relationships in robot motion decision-making is solved, and efficient decision-making and precise action execution in complex scenarios are achieved.
Patent Information
- Application Number
- CN202511063671.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-30
- Publication Date
- 2025-09-12
- Estimated Expiration
- 2045-07-30
AI Technical Summary
Existing models lack effective association when processing spatial relationships in multimodal data, resulting in deviations in robot action decisions. In particular, when complex spatial scenes change, the understanding of spatial relationships cannot be quickly updated, affecting the application of intelligent logistics, robotic navigation, and operating room robots.
Collect multimodal data of the target space, construct a three-dimensional spatial topology map through spatial feature extraction, and use cross-modal projection functions to perform cross-modal spatial projection alignment, perform spatial relationship reasoning and action planning, and generate target action decision data.
It improves the robot's decision-making ability and the feasibility and accuracy of action execution in complex spatial scenarios, and can quickly adapt to environmental changes and ensure that the action path meets the task objectives and spatial constraints.
Smart Images

Figure CN120620218A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of artificial intelligence technology, and in particular to a robot action decision-making method, device, equipment and medium based on spatial relationship. Background Art
[0002] In the field of Vision-Language-Action (VLA), existing models have significant deficiencies in handling spatial relationships in multimodal data.
[0003] Specifically, traditional approaches often process data from each modality independently, lacking an effective connection between the spatial layout of objects in the visual scene, the spatial orientation vocabulary used in verbal descriptions, and the spatial constraints of action execution. For example, in a robotic task in a banking lobby, when faced with the instruction to "place the water cup on the coffee table to the right of the reception desk," traditional approaches struggle to accurately match the visually recognized positions of objects like the coffee table, water cup, and reception desk with the spatial relationships in the verbal instruction, leading to discrepancies in action planning.
[0004] Furthermore, existing multimodal fusion methods are weak in spatial reasoning. When faced with complex spatial scene changes (such as object occlusion and spatial layout changes), they are unable to quickly update their understanding of spatial relationships and struggle to make reasonable action decisions based on spatial information.
[0005] Moreover, most models lack a unified expression and processing mechanism for spatial information, which makes the fusion of multimodal data in the spatial dimension inefficient, seriously restricting its application in scenarios with high requirements for spatial perception, such as intelligent logistics, robot navigation, and operating room robots. Summary of the Invention
[0006] In view of the above, it is necessary to provide a robot motion decision-making method, device, equipment and medium based on spatial relationships, aiming to solve the problem of robot motion decision-making deviation caused by the lack of association and processing of spatial relationships.
[0007] A robot motion decision method based on spatial relations, the robot motion decision method based on spatial relations comprising:
[0008] In response to an action decision instruction for a target robot in a target space, collecting multimodal data of the target space;
[0009] Performing spatial feature extraction on the multimodal data to obtain multimodal spatial features;
[0010] Constructing a three-dimensional spatial topology map using the multimodal spatial features;
[0011] Performing cross-modal spatial projection alignment on the multimodal spatial features using a cross-modal projection function to obtain a multimodal spatial alignment feature;
[0012] Performing spatial relationship reasoning based on the three-dimensional spatial topology graph and the multimodal spatial alignment features to obtain a target spatial relationship;
[0013] Action planning is performed according to the target spatial relationship and the three-dimensional spatial topology graph to obtain target action decision data.
[0014] A robot motion decision-making device based on spatial relations, comprising:
[0015] an acquisition unit, configured to acquire multimodal data of the target space in response to an action decision instruction for a target robot in the target space;
[0016] an extraction unit, configured to extract spatial features from the multimodal data to obtain multimodal spatial features;
[0017] A construction unit, configured to construct a three-dimensional spatial topology map using the multimodal spatial features;
[0018] an alignment unit, configured to perform cross-modal spatial projection alignment on the multimodal spatial features using a cross-modal projection function to obtain a multimodal spatial alignment feature;
[0019] An inference unit, configured to perform spatial relationship inference based on the three-dimensional spatial topology graph and the multimodal spatial alignment features to obtain a target spatial relationship;
[0020] A planning unit is used to perform action planning based on the target spatial relationship and the three-dimensional space topology map to obtain target action decision data.
[0021] A computer device, comprising:
[0022] a memory storing at least one instruction; and
[0023] A processor executes instructions stored in the memory to implement the robot action decision method based on spatial relations.
[0024] A computer-readable storage medium stores at least one instruction, and the at least one instruction is executed by a processor in a computer device to implement the robot action decision method based on spatial relations.
[0025] It can be seen from the above technical solutions that the present invention can collect multimodal data of the target space, and perform spatial feature extraction on the multimodal data to obtain multimodal spatial features, so as to accurately capture the spatial information in the multimodal data; use the multimodal spatial features to construct a three-dimensional spatial topology map to clarify the three-dimensional spatial structure and the relationship between the elements in the target space; use the cross-modal projection function to perform cross-modal spatial projection alignment on the multimodal spatial features, so that different modal features can be fused and compared in the same dimension; according to the three-dimensional spatial topology map and the multimodal spatial alignment features, spatial relationship reasoning is performed to obtain the target spatial relationship, and according to the target spatial relationship and the three-dimensional spatial topology map, action planning is performed to obtain target action decision data, thereby planning an action path that meets the task objectives and spatial constraints, improving the decision-making ability in complex spatial scenes and the feasibility and accuracy of robot action execution. BRIEF DESCRIPTION OF THE DRAWINGS
[0026] Figure 1 It is a flow chart of a preferred embodiment of the robot action decision-making method based on spatial relations of the present invention.
[0027] Figure 2 It is a functional module diagram of a preferred embodiment of the robot motion decision-making device based on spatial relations of the present invention.
[0028] Figure 3 It is a structural diagram of a computer device of a preferred embodiment of the present invention for implementing a robot action decision-making method based on spatial relations. DETAILED DESCRIPTION
[0029] In order to make the objectives, technical solutions and advantages of the present invention more clear, the present invention is described in detail below with reference to the accompanying drawings and specific embodiments.
[0030] like Figure 1 FIG. 1 is a flow chart of a preferred embodiment of the robot action decision method based on spatial relationship of the present invention. According to different requirements, the order of the steps in the flow chart can be changed, and some steps can be omitted.
[0031] The spatial relationship-based robot motion decision method is applied to one or more computer devices, which are devices that can automatically perform numerical calculations and / or information processing according to pre-set or stored instructions. Their hardware includes but is not limited to microprocessors, application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), digital signal processors (DSPs), embedded devices, etc.
[0032] The computer device can be any electronic product that can interact with a user, such as a personal computer, a tablet computer, a smart phone, a personal digital assistant (PDA), a game console, an interactive network television (IPTV), a smart wearable device, etc.
[0033] The computer device may also include a network device and / or a user device, wherein the network device includes, but is not limited to, a single network server, a server group consisting of multiple network servers, or a cloud based on cloud computing consisting of a large number of hosts or network servers.
[0034] The server can be an independent server or a cloud server that provides basic cloud computing services such as cloud services, cloud databases, cloud computing, cloud functions, cloud storage, network services, cloud communications, middleware services, domain name services, security services, content delivery networks (CDNs), and big data and artificial intelligence platforms.
[0035] Among them, artificial intelligence (AI) is the theory, method, technology and application system that uses digital computers or machines controlled by digital computers to simulate, extend and expand human intelligence, perceive the environment, acquire knowledge and use knowledge to obtain the best results.
[0036] Fundamental AI technologies generally include sensors, dedicated AI chips, cloud computing, distributed storage, big data processing, operating / interaction systems, and mechatronics. AI software technologies primarily encompass computer vision, robotics, biometrics, speech processing, natural language processing, and machine learning / deep learning.
[0037] The network where the computer device is located includes but is not limited to the Internet, wide area network, metropolitan area network, local area network, virtual private network (VPN), etc.
[0038] S10 , in response to an action decision instruction for a target robot in a target space, collecting multimodal data of the target space.
[0039] In this embodiment, the target space may be the robot's moving space, such as a bank lobby in a financial scenario, an operating room in a medical scenario, etc.
[0040] Accordingly, the target robot may be a service robot in a bank lobby, or a robot that assists in delivering surgical instruments in an operating room.
[0041] In this embodiment, the action decision instruction can be automatically triggered when it is detected that the target robot is started, so as to achieve full control of the target robot during the execution of the action.
[0042] In this embodiment, the multimodal data may be collected using an image acquisition device, a voice acquisition device, various sensor devices, etc. deployed in the target space.
[0043] The multimodal data may include visual data, voice data, motion data, etc.
[0044] S11, performing spatial feature extraction on the multimodal data to obtain multimodal spatial features.
[0045] In this embodiment, extracting spatial features from the multimodal data to obtain multimodal spatial features includes:
[0046] A Transformer-based 3D object detection network is used to identify target objects in the multimodal data to obtain three-dimensional bounding box information of the target objects; surface geometric features and spatial distribution features of the target objects are extracted using point cloud technology; and visual spatial features are generated based on the three-dimensional bounding box information of the target objects, the surface geometric features and spatial distribution features of the target objects;
[0047] Performing word segmentation and part-of-speech tagging on the language text in the multimodal data to obtain spatial location vocabulary, spatial relationship vocabulary, and entity vocabulary related to space as spatial text; encoding the spatial text using a pre-trained language model to obtain spatial semantic information; and converting the spatial semantic information into language space features through spatial relationship analysis;
[0048] Acquire motion data from the multimodal data and positioning data collected by an inertial sensor and a laser radar; determine the starting position, ending position, and motion path of the motion data in three-dimensional space based on the positioning data; and convert the starting position, ending position, and motion path into motion space features using a spatial trajectory modeling algorithm;
[0049] The visual space feature, the language space feature and the action space feature are integrated to obtain the multimodal space feature.
[0050] The Transformer-based 3D object detection network may be a Swin Transformer3D (Shifted Window Transformer 3D) network.
[0051] The three-dimensional bounding box information may include position coordinate information and size information.
[0052] The point cloud technology is used to process image depth information.
[0053] The language model may include a RoBERTa (Robustly Optimized BERT Pretraining Approach) model.
[0054] The motion space features may include information such as spatial displacement and direction change.
[0055] In financial security monitoring, visual spatial feature extraction can identify the location and posture of people within the monitoring area. Language spatial feature extraction can parse spatial descriptions in alarm commands, such as "There's an unusual gathering of people in the left corner of the hall." Motion spatial feature extraction can capture the trajectory of unusual movements of people, providing multimodal spatial information for security decision-making. In robotic-assisted medical surgeries, visual spatial feature extraction can obtain information such as the three-dimensional position and size of organs within the patient's body. Language spatial feature extraction can parse spatial relationships in doctor's instructions, such as "Move the instrument above the liver." Motion spatial feature extraction can capture the robot's trajectory, ensuring the accuracy of the surgical operation.
[0056] Through the above embodiments, effective extraction of spatial features in three modal data, namely vision, language and action, is achieved, providing basic data support for subsequent three-dimensional spatial topology modeling and cross-modal projection alignment, so that spatial information in multimodal data can be accurately captured and represented.
[0057] S12, constructing a three-dimensional spatial topology map using the multimodal spatial features.
[0058] In this embodiment, constructing a three-dimensional spatial topology map using the multimodal spatial features includes:
[0059] Determine the target object, the entity vocabulary, the starting position, and the ending position as nodes;
[0060] determining the actual spatial relationship between the target objects according to the spatial distribution characteristics;
[0061] Determining the spatial logical relationship between the entity words according to the spatial text;
[0062] determining the spatial constraint relationship during the action execution process according to the action space characteristics;
[0063] Determine an edge according to the actual spatial relationship, the spatial logical relationship, and the spatial constraint relationship;
[0064] The nodes and edges are processed using a message passing mechanism of a graph neural network to obtain the three-dimensional space topology graph.
[0065] The actual spatial relationship may include adjacent, occluded, etc.
[0066] For example, in a financial scenario, the three-dimensional spatial topology map can be used to identify monitoring equipment and personnel in a bank hall as nodes, and the positional relationships between monitoring equipment and personnel as edges. This topology map can clearly identify the spatial layout and relationships between elements in the bank hall, facilitating robot deployment.
[0067] For example, in a medical scenario, the three-dimensional spatial topology map can be used to represent hospital beds, medical equipment, and channels as nodes, and the spatial relationships between nodes as edges, thereby optimizing the configuration of patient transfer routes and medical resources such as robots.
[0068] In the above embodiment, the constructed three-dimensional spatial topology map can accurately reflect the three-dimensional spatial structure of the scene and the relationship between the elements, and uniformly organize and represent the spatial information in the multimodal data in the form of a graph, providing structured spatial knowledge for subsequent spatial reasoning and decision-making.
[0069] S13, performing cross-modal spatial projection alignment on the multimodal spatial features using a cross-modal projection function to obtain a multimodal spatial alignment feature.
[0070] In this embodiment, in order to achieve better results when processing data of various modalities in a unified manner, it is also necessary to spatially align the multimodal data.
[0071] Specifically, the cross-modal spatial projection alignment of the multimodal spatial features using a cross-modal projection function to obtain a multimodal spatial alignment feature includes:
[0072] Construct a three-dimensional space coordinate system;
[0073] The visual space features are converted to the three-dimensional space coordinate system using a projection matrix, the language space features are mapped to the three-dimensional space coordinate system using a semantic-space mapping model, and the action space features are subjected to coordinate transformation according to the three-dimensional space coordinate system to obtain an initial coordinate system;
[0074] An attention-based alignment algorithm is used to calculate the alignment scores of different modalities in the three-dimensional space coordinate system;
[0075] performing alignment processing on the initial coordinate system according to the alignment score to obtain the modal space alignment feature;
[0076] Among them, in the three-dimensional space coordinate system, for the visual space feature, the upper left corner of the image is used as the origin, and the scale of the coordinate system is determined according to the image size and the actual scene ratio of the target space; for the language space feature, the spatial orientation vocabulary is mapped to the three-dimensional space coordinate system; for the action space feature, the starting position is determined as the origin of the coordinate system, and the direction and scale of the coordinate system are determined according to the action direction and displacement reflected by the motion path.
[0077] The three-dimensional space coordinate system can be used as a mapping reference for multimodal spatial information.
[0078] Among them, by adopting the attention mechanism, accurate alignment of each modal feature can be achieved by continuously adjusting parameters.
[0079] Through the above embodiments, precise alignment of multimodal spatial information in a unified coordinate system is achieved, breaking the spatial information barriers between different modalities, enabling the spatial features of visual modalities, language modalities, and motion modalities to be fused and compared in the same dimension, thereby improving the fusion efficiency and accuracy of multimodal spatial information.
[0080] S14: Perform spatial relationship reasoning based on the three-dimensional spatial topology graph and the multimodal spatial alignment features to obtain a target spatial relationship.
[0081] In this embodiment, performing spatial relationship reasoning based on the three-dimensional spatial topology graph and the multimodal spatial alignment features to obtain a target spatial relationship includes:
[0082] fusing the three-dimensional spatial topology map with the multimodal spatial alignment feature to obtain a fusion feature;
[0083] Using a graph neural network to perform message transmission and node feature update on the three-dimensional spatial topology graph based on the fusion features to obtain the target spatial relationship;
[0084] The target spatial relationship includes the hidden spatial relationship between the target objects and the influence of language and action execution on the spatial relationship.
[0085] Among them, the graph neural network is used to perform message transmission and node feature update on the three-dimensional space topology map based on the fusion features, so as to infer the changes in spatial relationships after the objects are moved.
[0086] For example, in the path planning of a financial escort robot, the spatial relationships of obstacles, target locations, etc. on the escort route can be inferred based on the three-dimensional spatial topology map.
[0087] For another example, in the motion planning of a medical surgical robot, the spatial relationship between surgical instruments and patient organs can be inferred based on the three-dimensional spatial topology map.
[0088] S15, performing action planning according to the target spatial relationship and the three-dimensional spatial topology graph to obtain target action decision data.
[0089] In this embodiment, performing action planning based on the target spatial relationship and the three-dimensional spatial topology graph to obtain target action decision data includes:
[0090] According to the action decision instruction and the target spatial relationship, a reinforcement learning algorithm is used to search for an optimal path on the three-dimensional spatial topology graph to obtain a search result;
[0091] The target action decision data is generated according to the search results.
[0092] Specifically, based on the target spatial relationship, the spatial layout of objects in the current scene, the relationship between objects (such as "the box is on the left side of the shelf"), and language instructions or task requirements (such as "put the box on the shelf") can be clarified.
[0093] Specifically, according to the action decision instruction, the core goal of the task can be clarified, including the starting state (such as the current position of the robot, the initial position of the target object) and the ending state (such as the placement position of the target object).
[0094] Furthermore, key spatial points such as the robot's feasible position, the position of the target object, and the position of obstacles can be extracted from the three-dimensional space topology map as new topology map nodes based on the action decision instructions and the target spatial relationship.
[0095] Furthermore, based on physical feasibility (such as collision-free paths and action execution range), effective movement paths or operation steps between nodes can be defined as edges to form a sub-topology graph based on the task scenario.
[0096] Furthermore, reinforcement learning algorithms (such as DQN (Deep Q-Network) and PPO (Proximal Policy Optimization)) are used to transform the path search in the sub-topology graph into a "state-action" decision-making problem. Specifically,
[0097] State: the spatial relationship between the current node (such as the robot position) and the target node (such as the placement position), obstacle distribution, etc.
[0098] Action: Movement or operation from the current node to the adjacent node (such as "move in front of an object", "grab an object").
[0099] Reward function: designed based on path length (e.g., shorter paths have higher rewards), safety (e.g., avoiding obstacles has higher rewards), and task completion (e.g., reaching the target node has the highest reward).
[0100] Through iterative training, the model learns the optimal path from the starting node to the target node and the action sequence corresponding to the optimal path, such as: move to point A → grab an object → move to point B → place the object.
[0101] The action sequence corresponding to the optimal path is further used as the search result and converted into specific control instructions, such as robot joint angles, movement speed, operation force and other parameters, to obtain the target action decision data.
[0102] For example, in a financial scenario, in the path planning of a financial escort robot, after inferring the spatial relationships of obstacles, target locations, etc. on the escort route based on the three-dimensional spatial topology map, the optimal escort path is further planned through reinforcement learning, thereby ensuring the safe and efficient transportation of funds.
[0103] For example, in a medical scenario, during the motion planning of a surgical robot, after inferring the spatial relationship between the surgical instrument and the patient's organs based on the three-dimensional spatial topology map, a precise instrument motion path is further planned in combination with the surgical goal, thereby avoiding damage to surrounding tissues.
[0104] Through the above embodiments, effective spatial relationship reasoning can be performed based on the topological map and aligned multimodal spatial features, and action paths that meet task objectives and spatial constraints can be accurately planned, thereby improving the model's decision-making ability and the feasibility and accuracy of action execution in complex spatial scenarios.
[0105] In this embodiment, after obtaining the target action decision data, the method further includes:
[0106] Controlling the target robot to respond to the action decision instruction according to the target action decision data;
[0107] During the process of the target robot performing an action, continuously monitoring the real-time changes of each modal data in the target space;
[0108] When a change in modal data is detected, the three-dimensional space topology map and the multimodal space alignment feature are updated according to the detected change data.
[0109] Specifically, it can continuously monitor changes in visual, language, and motion modal data to perceive spatial changes in the environment, such as object movement, language command updates, and changes in spatial state feedback from sensors. When environmental changes are detected, the 3D spatial topology map and the multimodal spatial alignment features are updated in a timely manner.
[0110] Among them, new data and decision results can also be used as training samples to train various network models involved in the robot's action decision-making process. For example, the backpropagation algorithm can be used to optimize the parameters of each model.
[0111] Through the above embodiments, the model can quickly adapt to environmental changes, update spatial relationship understanding and action decision-making strategies in real time, ensure the model's processing capabilities and decision-making accuracy in a dynamic environment, and enhance the model's robustness and practicality.
[0112] It can be seen from the above technical solutions that the present invention can collect multimodal data of the target space, and perform spatial feature extraction on the multimodal data to obtain multimodal spatial features, so as to accurately capture the spatial information in the multimodal data; use the multimodal spatial features to construct a three-dimensional spatial topology map to clarify the three-dimensional spatial structure and the relationship between the elements in the target space; use the cross-modal projection function to perform cross-modal spatial projection alignment on the multimodal spatial features, so that different modal features can be fused and compared in the same dimension; according to the three-dimensional spatial topology map and the multimodal spatial alignment features, spatial relationship reasoning is performed to obtain the target spatial relationship, and according to the target spatial relationship and the three-dimensional spatial topology map, action planning is performed to obtain target action decision data, thereby planning an action path that meets the task objectives and spatial constraints, improving the decision-making ability in complex spatial scenes and the feasibility and accuracy of robot action execution.
[0113] like Figure 2 Figure 1 shows a functional block diagram of a preferred embodiment of a spatial relationship-based robot motion decision-making device according to the present invention. The spatial relationship-based robot motion decision-making device 11 comprises an acquisition unit 110, an extraction unit 111, a construction unit 112, an alignment unit 113, an inference unit 114, and a planning unit 115. As used herein, a module / unit refers to a series of computer program segments that can be executed by a processor and perform a fixed function, and are stored in a memory. The functions of each module / unit in this embodiment will be described in detail in subsequent embodiments.
[0114] The acquisition unit 110 is configured to acquire multimodal data of the target space in response to an action decision instruction for the target robot in the target space.
[0115] In this embodiment, the target space may be the robot's moving space, such as a bank lobby in a financial scenario, an operating room in a medical scenario, etc.
[0116] Accordingly, the target robot may be a service robot in a bank lobby, or a robot that assists in delivering surgical instruments in an operating room.
[0117] In this embodiment, the action decision instruction can be automatically triggered when it is detected that the target robot is started, so as to achieve full control of the target robot during the execution of the action.
[0118] In this embodiment, the multimodal data may be collected using an image acquisition device, a voice acquisition device, various sensor devices, etc. deployed in the target space.
[0119] The multimodal data may include visual data, voice data, motion data, etc.
[0120] The extraction unit 111 is configured to extract spatial features from the multimodal data to obtain multimodal spatial features.
[0121] In this embodiment, the extraction unit 111 extracts spatial features from the multimodal data, and the obtained multimodal spatial features include:
[0122] A Transformer-based 3D object detection network is used to identify target objects in the multimodal data to obtain three-dimensional bounding box information of the target objects; surface geometric features and spatial distribution features of the target objects are extracted using point cloud technology; and visual spatial features are generated based on the three-dimensional bounding box information of the target objects, the surface geometric features and spatial distribution features of the target objects;
[0123] Performing word segmentation and part-of-speech tagging on the language text in the multimodal data to obtain spatial location vocabulary, spatial relationship vocabulary, and entity vocabulary related to space as spatial text; encoding the spatial text using a pre-trained language model to obtain spatial semantic information; and converting the spatial semantic information into language space features through spatial relationship analysis;
[0124] Acquire motion data from the multimodal data and positioning data collected by an inertial sensor and a laser radar; determine the starting position, ending position, and motion path of the motion data in three-dimensional space based on the positioning data; and convert the starting position, ending position, and motion path into motion space features using a spatial trajectory modeling algorithm;
[0125] The visual space feature, the language space feature and the action space feature are integrated to obtain the multimodal space feature.
[0126] The Transformer-based 3D object detection network may be a Swin Transformer3D (Shifted Window Transformer 3D) network.
[0127] The three-dimensional bounding box information may include position coordinate information and size information.
[0128] The point cloud technology is used to process image depth information.
[0129] The language model may include a RoBERTa (Robustly Optimized BERT Pretraining Approach) model.
[0130] The motion space features may include information such as spatial displacement and direction change.
[0131] In financial security monitoring, visual spatial feature extraction can identify the location and posture of people within the monitoring area. Language spatial feature extraction can parse spatial descriptions in alarm commands, such as "There's an unusual gathering of people in the left corner of the hall." Motion spatial feature extraction can capture the trajectory of unusual movements of people, providing multimodal spatial information for security decision-making. In robotic-assisted medical surgeries, visual spatial feature extraction can obtain information such as the three-dimensional position and size of organs within the patient's body. Language spatial feature extraction can parse spatial relationships in doctor's instructions, such as "Move the instrument above the liver." Motion spatial feature extraction can capture the robot's trajectory, ensuring the accuracy of the surgical operation.
[0132] Through the above embodiments, effective extraction of spatial features in three modal data, namely vision, language and action, is achieved, providing basic data support for subsequent three-dimensional spatial topology modeling and cross-modal projection alignment, so that spatial information in multimodal data can be accurately captured and represented.
[0133] The construction unit 112 is configured to construct a three-dimensional spatial topology map using the multimodal spatial features.
[0134] In this embodiment, the construction unit 112 constructs a three-dimensional space topology map using the multimodal space features, including:
[0135] Determine the target object, the entity vocabulary, the starting position, and the ending position as nodes;
[0136] determining the actual spatial relationship between the target objects according to the spatial distribution characteristics;
[0137] Determining the spatial logical relationship between the entity words according to the spatial text;
[0138] determining the spatial constraint relationship during the action execution process according to the action space characteristics;
[0139] Determine an edge according to the actual spatial relationship, the spatial logical relationship, and the spatial constraint relationship;
[0140] The nodes and edges are processed using a message passing mechanism of a graph neural network to obtain the three-dimensional space topology graph.
[0141] The actual spatial relationship may include adjacent, occluded, etc.
[0142] For example, in a financial scenario, the three-dimensional spatial topology map can be used to identify monitoring equipment and personnel in a bank hall as nodes, and the positional relationships between monitoring equipment and personnel as edges. This topology map can clearly identify the spatial layout and relationships between elements in the bank hall, facilitating robot deployment.
[0143] For example, in a medical scenario, the three-dimensional spatial topology map can be used to represent hospital beds, medical equipment, and channels as nodes, and the spatial relationships between nodes as edges, thereby optimizing the configuration of patient transfer routes and medical resources such as robots.
[0144] In the above embodiment, the constructed three-dimensional spatial topology map can accurately reflect the three-dimensional spatial structure of the scene and the relationship between the elements, and uniformly organize and represent the spatial information in the multimodal data in the form of a graph, providing structured spatial knowledge for subsequent spatial reasoning and decision-making.
[0145] The alignment unit 113 is configured to perform cross-modal spatial projection alignment on the multimodal spatial features using a cross-modal projection function to obtain a multimodal spatial alignment feature.
[0146] In this embodiment, in order to achieve better results when processing data of various modalities in a unified manner, it is also necessary to spatially align the multimodal data.
[0147] Specifically, the alignment unit 113 performs cross-modal spatial projection alignment on the multimodal spatial features using a cross-modal projection function, and obtains multimodal spatial alignment features including:
[0148] Construct a three-dimensional space coordinate system;
[0149] The visual space features are converted to the three-dimensional space coordinate system using a projection matrix, the language space features are mapped to the three-dimensional space coordinate system using a semantic-space mapping model, and the action space features are subjected to coordinate transformation according to the three-dimensional space coordinate system to obtain an initial coordinate system;
[0150] An attention-based alignment algorithm is used to calculate the alignment scores of different modalities in the three-dimensional space coordinate system;
[0151] performing alignment processing on the initial coordinate system according to the alignment score to obtain the modal space alignment feature;
[0152] Among them, in the three-dimensional space coordinate system, for the visual space feature, the upper left corner of the image is used as the origin, and the scale of the coordinate system is determined according to the image size and the actual scene ratio of the target space; for the language space feature, the spatial orientation vocabulary is mapped to the three-dimensional space coordinate system; for the action space feature, the starting position is determined as the origin of the coordinate system, and the direction and scale of the coordinate system are determined according to the action direction and displacement reflected by the motion path.
[0153] The three-dimensional space coordinate system can be used as a mapping reference for multimodal spatial information.
[0154] Among them, by adopting the attention mechanism, accurate alignment of each modal feature can be achieved by continuously adjusting parameters.
[0155] Through the above embodiments, precise alignment of multimodal spatial information in a unified coordinate system is achieved, breaking the spatial information barriers between different modalities, enabling the spatial features of visual modalities, language modalities, and motion modalities to be fused and compared in the same dimension, thereby improving the fusion efficiency and accuracy of multimodal spatial information.
[0156] The reasoning unit 114 is configured to perform spatial relationship reasoning based on the three-dimensional spatial topology graph and the multimodal spatial alignment features to obtain a target spatial relationship.
[0157] In this embodiment, the inference unit 114 performs spatial relationship inference based on the three-dimensional spatial topology graph and the multimodal spatial alignment features, and obtains the target spatial relationship including:
[0158] fusing the three-dimensional spatial topology map with the multimodal spatial alignment feature to obtain a fusion feature;
[0159] Using a graph neural network to perform message transmission and node feature update on the three-dimensional spatial topology graph based on the fusion features to obtain the target spatial relationship;
[0160] The target spatial relationship includes the hidden spatial relationship between the target objects and the influence of language and action execution on the spatial relationship.
[0161] Among them, the graph neural network is used to perform message transmission and node feature update on the three-dimensional space topology map based on the fusion features, so as to infer the changes in spatial relationships after the objects are moved.
[0162] For example, in the path planning of a financial escort robot, the spatial relationships of obstacles, target locations, etc. on the escort route can be inferred based on the three-dimensional spatial topology map.
[0163] For another example, in the motion planning of a medical surgical robot, the spatial relationship between surgical instruments and patient organs can be inferred based on the three-dimensional spatial topology map.
[0164] The planning unit 115 is configured to perform action planning based on the target spatial relationship and the three-dimensional spatial topology graph to obtain target action decision data.
[0165] In this embodiment, the planning unit 115 performs action planning based on the target spatial relationship and the three-dimensional spatial topology graph, and obtains target action decision data including:
[0166] According to the action decision instruction and the target spatial relationship, a reinforcement learning algorithm is used to search for an optimal path on the three-dimensional spatial topology graph to obtain a search result;
[0167] The target action decision data is generated according to the search results.
[0168] Specifically, based on the target spatial relationship, the spatial layout of objects in the current scene, the relationship between objects (such as "the box is on the left side of the shelf"), and language instructions or task requirements (such as "put the box on the shelf") can be clarified.
[0169] Specifically, according to the action decision instruction, the core goal of the task can be clarified, including the starting state (such as the current position of the robot, the initial position of the target object) and the ending state (such as the placement position of the target object).
[0170] Furthermore, key spatial points such as the robot's feasible position, the position of the target object, and the position of obstacles can be extracted from the three-dimensional space topology map as new topology map nodes based on the action decision instructions and the target spatial relationship.
[0171] Furthermore, based on physical feasibility (such as collision-free paths and action execution range), effective movement paths or operation steps between nodes can be defined as edges to form a sub-topology graph based on the task scenario.
[0172] Furthermore, reinforcement learning algorithms (such as DQN (Deep Q-Network) and PPO (Proximal Policy Optimization)) are used to transform the path search in the sub-topology graph into a "state-action" decision-making problem. Specifically,
[0173] State: the spatial relationship between the current node (such as the robot position) and the target node (such as the placement position), obstacle distribution, etc.
[0174] Action: Movement or operation from the current node to the adjacent node (such as "move in front of an object", "grab an object").
[0175] Reward function: designed based on path length (e.g., shorter paths have higher rewards), safety (e.g., avoiding obstacles has higher rewards), and task completion (e.g., reaching the target node has the highest reward).
[0176] Through iterative training, the model learns the optimal path from the starting node to the target node and the action sequence corresponding to the optimal path, such as: move to point A → grab an object → move to point B → place the object.
[0177] The action sequence corresponding to the optimal path is further used as the search result and converted into specific control instructions, such as robot joint angles, movement speed, operation force and other parameters, to obtain the target action decision data.
[0178] For example, in a financial scenario, in the path planning of a financial escort robot, after inferring the spatial relationships of obstacles, target locations, etc. on the escort route based on the three-dimensional spatial topology map, the optimal escort path is further planned through reinforcement learning, thereby ensuring the safe and efficient transportation of funds.
[0179] For example, in a medical scenario, during the motion planning of a surgical robot, after inferring the spatial relationship between the surgical instrument and the patient's organs based on the three-dimensional spatial topology map, a precise instrument motion path is further planned in combination with the surgical goal, thereby avoiding damage to surrounding tissues.
[0180] Through the above embodiments, effective spatial relationship reasoning can be performed based on the topological map and aligned multimodal spatial features, and action paths that meet task objectives and spatial constraints can be accurately planned, thereby improving the model's decision-making ability and the feasibility and accuracy of action execution in complex spatial scenarios.
[0181] In this embodiment, after obtaining the target action decision data, the target robot is controlled to respond to the action decision instruction according to the target action decision data;
[0182] During the process of the target robot performing an action, continuously monitoring the real-time changes of each modal data in the target space;
[0183] When a change in modal data is detected, the three-dimensional space topology map and the multimodal space alignment feature are updated according to the detected change data.
[0184] Specifically, it can continuously monitor changes in visual, language, and motion modal data to perceive spatial changes in the environment, such as object movement, language command updates, and changes in spatial state feedback from sensors. When environmental changes are detected, the 3D spatial topology map and the multimodal spatial alignment features are updated in a timely manner.
[0185] Among them, new data and decision results can also be used as training samples to train various network models involved in the robot's action decision-making process. For example, the backpropagation algorithm can be used to optimize the parameters of each model.
[0186] Through the above embodiments, the model can quickly adapt to environmental changes, update spatial relationship understanding and action decision-making strategies in real time, ensure the model's processing capabilities and decision-making accuracy in a dynamic environment, and enhance the model's robustness and practicality.
[0187] It can be seen from the above technical solutions that the present invention can collect multimodal data of the target space, and perform spatial feature extraction on the multimodal data to obtain multimodal spatial features, so as to accurately capture the spatial information in the multimodal data; use the multimodal spatial features to construct a three-dimensional spatial topology map to clarify the three-dimensional spatial structure and the relationship between the elements in the target space; use the cross-modal projection function to perform cross-modal spatial projection alignment on the multimodal spatial features, so that different modal features can be fused and compared in the same dimension; according to the three-dimensional spatial topology map and the multimodal spatial alignment features, spatial relationship reasoning is performed to obtain the target spatial relationship, and according to the target spatial relationship and the three-dimensional spatial topology map, action planning is performed to obtain target action decision data, thereby planning an action path that meets the task objectives and spatial constraints, improving the decision-making ability in complex spatial scenes and the feasibility and accuracy of robot action execution.
[0188] like Figure 3 , which is a structural diagram of a computer device of a preferred embodiment of the present invention for implementing a robot action decision-making method based on spatial relations.
[0189] The computer device 1 may include a memory 12, a processor 13 and a bus (the arrow in the figure represents the bus), and may also include a computer program stored in the memory 12 and executable on the processor 13, such as a robot motion decision program based on spatial relationships.
[0190] Those skilled in the art will understand that the schematic diagram is merely an example of the computer device 1 and does not constitute a limitation on the computer device 1. The computer device 1 may have either a bus structure or a star structure. The computer device 1 may also include more or less other hardware or software than shown in the figure, or a different arrangement of components. For example, the computer device 1 may also include input and output devices, network access devices, etc.
[0191] It should be noted that the computer device 1 is only an example. Other existing or future electronic products that are suitable for the present invention should also be included in the scope of protection of the present invention and included here by reference.
[0192] The memory 12 includes at least one type of readable storage medium, including a flash memory, a mobile hard disk, a multimedia card, a card-type memory (e.g., SD or DX memory), a magnetic memory, a magnetic disk, an optical disk, etc. In some embodiments, the memory 12 may be an internal storage unit of the computer device 1, such as a mobile hard disk of the computer device 1. In other embodiments, the memory 12 may also be an external storage device of the computer device 1, such as a plug-in mobile hard disk, a smart memory card (SMC), a secure digital (SD) card, a flash card, etc. equipped on the computer device 1. Furthermore, the memory 12 may include both an internal storage unit of the computer device 1 and an external storage device. The memory 12 can be used not only to store application software installed in the computer device 1 and various types of data, such as the code of a robot motion decision program based on spatial relationships, but also to temporarily store data that has been output or is about to be output.
[0193] In some embodiments, the processor 13 may be comprised of an integrated circuit, such as a single packaged integrated circuit or a plurality of packaged integrated circuits with the same or different functions, including one or more central processing units (CPUs), microprocessors, digital processing chips, graphics processors, and combinations of various control chips. The processor 13 is the control core (Control Unit) of the computer device 1, connecting the various components of the entire computer device 1 using various interfaces and circuits. It executes or runs programs or modules stored in the memory 12 (e.g., executing a robot action decision program based on spatial relationships) and calls data stored in the memory 12 to perform various functions and process data.
[0194] The processor 13 executes the operating system of the computer device 1 and various installed applications. The processor 13 executes the applications to implement the steps in the above-mentioned embodiments of the robot action decision method based on spatial relationship, for example Figure 1 Steps shown.
[0195] Exemplarily, the computer program may be divided into one or more modules / units, which are stored in the memory 12 and executed by the processor 13 to implement the present invention. The one or more modules / units may be a series of computer-readable instruction segments capable of performing specific functions, and the instruction segments are used to describe the execution process of the computer program in the computer device 1. For example, the computer program may be divided into an acquisition unit 110, an extraction unit 111, a construction unit 112, an alignment unit 113, an inference unit 114, and a planning unit 115.
[0196] The above-mentioned integrated unit implemented in the form of a software functional module can be stored in a computer-readable storage medium. The above-mentioned software functional module stored in a storage medium includes a number of instructions for causing a computer device (which can be a personal computer, computer device, or network device, etc.) or a processor to execute the portion of the spatial relationship-based robot motion decision-making method described in various embodiments of the present invention.
[0197] If the modules / units integrated in the computer device 1 are implemented as software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the present invention can also implement all or part of the processes in the above-mentioned method embodiments by instructing relevant hardware devices through a computer program. The computer program can be stored in a computer-readable storage medium. When executed by a processor, the computer program can implement the steps of each of the above-mentioned method embodiments.
[0198] The computer program includes computer program code, which may be in source code form, object code form, executable file, or some intermediate form. The computer-readable medium may include any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a mobile hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory (ROM), a random access memory, etc.
[0199] Furthermore, the computer-readable storage medium may mainly include a program storage area and a data storage area, wherein the program storage area may store an operating system, an application required for at least one function, etc.; the data storage area may store data created according to the use of the blockchain node, etc.
[0200] Blockchain, as used in this article, refers to a novel application model for computer technologies such as distributed data storage, peer-to-peer transmission, consensus mechanisms, and encryption algorithms. Blockchain is essentially a decentralized database, a series of data blocks generated using cryptographic methods. Each block contains information about a batch of online transactions, used to verify the validity of this information (to prevent counterfeiting) and generate the next block. Blockchain can include the underlying blockchain platform, the platform product service layer, and the application service layer.
[0201] The bus can be a peripheral component interconnect (PCI) bus or an extended industry standard architecture (EISA) bus. The bus can be divided into an address bus, a data bus, a control bus, etc. For ease of representation, Figure 3 The figure shows that only one straight line is used, but it does not mean that there is only one bus or one type of bus. The bus is configured to realize the connection and communication between the memory 12 and at least one processor 13.
[0202] Although not shown, the computer device 1 may also include a power supply (such as a battery) to power various components. Preferably, the power supply can be logically connected to the at least one processor 13 via a power management device, thereby implementing functions such as charging management, discharging management, and power consumption management through the power management device. The power supply may also include one or more DC or AC power supplies, a recharging device, a power failure detection circuit, a power converter or inverter, a power status indicator, and other arbitrary components. The computer device 1 may also include various sensors, Bluetooth modules, Wi-Fi modules, etc., which will not be detailed here.
[0203] Furthermore, the computer device 1 may also include a network interface. Optionally, the network interface may include a wired interface and / or a wireless interface (such as a WI-FI interface, a Bluetooth interface, etc.), which is usually used to establish a communication connection between the computer device 1 and other computer devices.
[0204] Optionally, the computer device 1 may further include a user interface, which may be a display or an input unit (such as a keyboard). Optionally, the user interface may also be a standard wired interface or a wireless interface. Optionally, in some embodiments, the display may be an LED display, a liquid crystal display, a touch-sensitive liquid crystal display, or an OLED (Organic Light-Emitting Diode) touchscreen. The display may also be appropriately referred to as a display screen or a display unit, and is used to display information processed in the computer device 1 and to display a visual user interface.
[0205] It should be understood that the embodiment is for illustration only and the scope of the patent application is not limited to this structure.
[0206] It will be understood by those skilled in the art that Figure 3 The structure shown does not constitute a limitation on the computer device 1 , and may include fewer or more components than shown in the figure, or combine certain components, or arrange the components differently.
[0207] Combine Figure 1 The memory 12 in the computer device 1 stores a plurality of instructions to implement a robot action decision method based on spatial relationships, and the processor 13 can execute the plurality of instructions to implement:
[0208] In response to an action decision instruction for a target robot in a target space, collecting multimodal data of the target space;
[0209] Performing spatial feature extraction on the multimodal data to obtain multimodal spatial features;
[0210] Constructing a three-dimensional spatial topology map using the multimodal spatial features;
[0211] Performing cross-modal spatial projection alignment on the multimodal spatial features using a cross-modal projection function to obtain a multimodal spatial alignment feature;
[0212] Performing spatial relationship reasoning based on the three-dimensional spatial topology graph and the multimodal spatial alignment features to obtain a target spatial relationship;
[0213] Action planning is performed according to the target spatial relationship and the three-dimensional spatial topology graph to obtain target action decision data.
[0214] Specifically, the specific implementation method of the processor 13 for the above instructions can refer to Figure 1 The description of the relevant steps in the corresponding embodiments will not be repeated here.
[0215] It should be noted that the data involved in this case were all obtained legally. The software tools or components not produced by our company that appear in the embodiments of this application are merely examples and do not represent actual use.
[0216] In the several embodiments provided herein, it should be understood that the disclosed systems, devices, and methods may be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the module division is merely a logical functional division, and actual implementation may employ other division methods.
[0217] The present invention can be used in a wide variety of general or special computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, and the like. The present invention can be described in the general context of computer-executable instructions executed by a computer, such as program modules. Generally, program modules include routines, programs, objects, components, data structures, etc. that perform specific tasks or implement specific abstract data types. The present invention can also be practiced in distributed computing environments in which tasks are performed by remote processing devices connected via a communications network. In a distributed computing environment, program modules can be located in local and remote computer storage media, including storage devices.
[0218] The modules described as separate components may or may not be physically separate, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed across multiple network elements. Some or all of the modules may be selected to achieve the purpose of the solution of this embodiment according to actual needs.
[0219] In addition, the functional modules in various embodiments of the present invention may be integrated into a single processing unit, each unit may exist physically separately, or two or more units may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or hardware plus software functional modules.
[0220] It will be apparent to those skilled in the art that the present invention is not limited to the details of the exemplary embodiments described above, and that the present invention can be implemented in other specific forms without departing from the spirit or essential characteristics of the present invention.
[0221] Therefore, the embodiments should be considered in all respects as illustrative and non-restrictive, and the scope of the invention is defined by the appended claims rather than the foregoing description, and all changes that come within the meaning and range of equivalents of the claims are intended to be embraced therein. Any reference to a figure in a claim should not be construed as limiting the claim to which it relates.
[0222] Furthermore, it is clear that the word "comprising" does not exclude other units or steps, and the singular does not exclude the plural. Multiple units or devices described in the present invention may also be implemented by a single unit or device through software or hardware. Terms such as first and second are used to indicate names and do not imply any particular order.
[0223] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not limiting. Although the present invention has been described in detail with reference to the preferred embodiments, those skilled in the art should understand that the technical solutions of the present invention may be modified or replaced by equivalents without departing from the spirit and scope of the technical solutions of the present invention.
Claims
1. A robot action decision method based on spatial relations, characterized in that: The robot action decision method based on spatial relationship includes: In response to an action decision instruction for a target robot in a target space, collecting multimodal data of the target space; Performing spatial feature extraction on the multimodal data to obtain multimodal spatial features; Constructing a three-dimensional spatial topology map using the multimodal spatial features; Performing cross-modal spatial projection alignment on the multimodal spatial features using a cross-modal projection function to obtain a multimodal spatial alignment feature; Performing spatial relationship reasoning based on the three-dimensional spatial topology graph and the multimodal spatial alignment features to obtain a target spatial relationship; Action planning is performed according to the target spatial relationship and the three-dimensional spatial topology graph to obtain target action decision data.
2. The robot action decision method based on spatial relationship according to claim 1, characterized in that: The extracting spatial features from the multimodal data to obtain multimodal spatial features includes: A Transformer-based 3D object detection network is used to identify target objects in the multimodal data to obtain three-dimensional bounding box information of the target objects; surface geometric features and spatial distribution features of the target objects are extracted using point cloud technology; and visual spatial features are generated based on the three-dimensional bounding box information of the target objects, the surface geometric features and spatial distribution features of the target objects; Performing word segmentation and part-of-speech tagging on the language text in the multimodal data to obtain spatial location vocabulary, spatial relationship vocabulary, and entity vocabulary related to space as spatial text; encoding the spatial text using a pre-trained language model to obtain spatial semantic information; and converting the spatial semantic information into language space features through spatial relationship analysis; Acquire motion data from the multimodal data and positioning data collected by an inertial sensor and a laser radar; determine the starting position, ending position, and motion path of the motion data in three-dimensional space based on the positioning data; and convert the starting position, ending position, and motion path into motion space features using a spatial trajectory modeling algorithm; The visual space feature, the language space feature and the action space feature are integrated to obtain the multimodal space feature.
3. The robot action decision method based on spatial relationship according to claim 2, characterized in that: The constructing of a three-dimensional spatial topology map using the multimodal spatial features includes: Determine the target object, the entity vocabulary, the starting position, and the ending position as nodes; determining the actual spatial relationship between the target objects according to the spatial distribution characteristics; Determining the spatial logical relationship between the entity words according to the spatial text; determining the spatial constraint relationship during the action execution process according to the action space characteristics; Determine an edge according to the actual spatial relationship, the spatial logical relationship, and the spatial constraint relationship; The nodes and edges are processed using a message passing mechanism of a graph neural network to obtain the three-dimensional space topology graph.
4. The robot action decision method based on spatial relationship according to claim 2, characterized in that: The step of performing cross-modal spatial projection alignment on the multimodal spatial features by using a cross-modal projection function to obtain a multimodal spatial alignment feature includes: Construct a three-dimensional space coordinate system; The visual space features are converted to the three-dimensional space coordinate system using a projection matrix, the language space features are mapped to the three-dimensional space coordinate system using a semantic-space mapping model, and the action space features are subjected to coordinate transformation according to the three-dimensional space coordinate system to obtain an initial coordinate system; An attention-based alignment algorithm is used to calculate the alignment scores of different modalities in the three-dimensional space coordinate system; performing alignment processing on the initial coordinate system according to the alignment score to obtain the modal space alignment feature; Among them, in the three-dimensional space coordinate system, for the visual space feature, the upper left corner of the image is used as the origin, and the scale of the coordinate system is determined according to the image size and the actual scene ratio of the target space; for the language space feature, the spatial orientation vocabulary is mapped to the three-dimensional space coordinate system; for the action space feature, the starting position is determined as the origin of the coordinate system, and the direction and scale of the coordinate system are determined according to the action direction and displacement reflected by the motion path.
5. The robot action decision method based on spatial relationship according to claim 2, characterized in that: The performing spatial relationship reasoning based on the three-dimensional spatial topology graph and the multimodal spatial alignment features to obtain the target spatial relationship includes: fusing the three-dimensional spatial topology map with the multimodal spatial alignment feature to obtain a fusion feature; Using a graph neural network to perform message transmission and node feature update on the three-dimensional spatial topology graph based on the fusion features to obtain the target spatial relationship; The target spatial relationship includes the hidden spatial relationship between the target objects and the influence of language and action execution on the spatial relationship.
6. The robot action decision method based on spatial relationship according to claim 1, characterized in that: The performing of action planning according to the target spatial relationship and the three-dimensional spatial topology graph to obtain target action decision data includes: According to the action decision instruction and the target spatial relationship, a reinforcement learning algorithm is used to search for an optimal path on the three-dimensional spatial topology graph to obtain a search result; The target action decision data is generated according to the search results.
7. The robot motion decision method based on spatial relationship according to claim 1, characterized in that: After obtaining the target action decision data, the method further includes: Controlling the target robot to respond to the action decision instruction according to the target action decision data; During the process of the target robot performing an action, continuously monitoring the real-time changes of each modal data in the target space; When a change in modal data is detected, the three-dimensional space topology map and the multimodal space alignment feature are updated according to the detected change data.
8. A robot action decision-making device based on spatial relations, characterized in that: The robot action decision-making device based on spatial relationship includes: an acquisition unit, configured to acquire multimodal data of the target space in response to an action decision instruction for a target robot in the target space; an extraction unit, configured to extract spatial features from the multimodal data to obtain multimodal spatial features; A construction unit, configured to construct a three-dimensional spatial topology map using the multimodal spatial features; an alignment unit, configured to perform cross-modal spatial projection alignment on the multimodal spatial features using a cross-modal projection function to obtain a multimodal spatial alignment feature; An inference unit, configured to perform spatial relationship inference based on the three-dimensional spatial topology graph and the multimodal spatial alignment features to obtain a target spatial relationship; A planning unit is used to perform action planning based on the target spatial relationship and the three-dimensional space topology map to obtain target action decision data.
9. A computer device, characterized in that: The computer device comprises: a memory storing at least one instruction; and A processor executes instructions stored in the memory to implement the robot action decision method based on spatial relations as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores at least one instruction, and the at least one instruction is executed by a processor in a computer device to implement the robot action decision method based on spatial relations as described in any one of claims 1 to 7.
Citation Information
Patent Citations
Man-machine cooperation method and system based on multi-modal behavior online prediction
CN113524175A
Cross-domain multi-robot cooperation brain-like mapping method and device based on multi-modal perception
CN117685952A
Autonomous robot decision-making system based on multi-modal perception fusion and method thereof
CN119295883A
Mechanical arm trajectory tracking control double-Q reinforcement learning method with disturbance
CN119910656A