Mechanical arm task planning and control method, device, equipment and medium
By acquiring visual and linguistic information to generate an environment model, and combining dynamic programming and reinforcement learning to optimize the path, the problem of low path planning efficiency and weak generalization ability of robotic arms in complex environments is solved, and efficient and precise operation control is achieved.
Patent Information
- Application Number
- CN202511051743.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-07-29
- Publication Date
- 2025-11-11
AI Technical Summary
Existing technologies suffer from low efficiency in path planning for robotic arms in complex dynamic environments, weak task understanding and generalization capabilities, and poor motion execution accuracy. In particular, in the fields of embodied intelligence, fintech, and healthcare, there are problems such as incomplete acquisition of environmental models, insufficient fusion of multi-source information, and poor adaptability of path planning.
By acquiring visual information and task description language information, an environment model is generated and encoded. The language model is used for task planning, and dynamic programming is combined to generate a basic path. The path is adjusted through reinforcement learning to generate the final path, and the robotic arm movement is optimized based on joint control commands.
It improves the robotic arm's path adaptability and task generalization ability in complex dynamic environments, achieves efficient and precise operation control, and enhances the system's stability and the reliability of task execution.
Smart Images

Figure CN120921371A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of artificial intelligence technology, and in particular to a method, apparatus, device and storage medium for robot arm task planning and control. Background Technology
[0002] In the field of embodied intelligence, especially for tasks where fixed robotic arms achieve precise grasping through dual-arm collaboration and possess a certain degree of generalization ability, existing technologies still face significant technical bottlenecks in perception, decision-making, path planning, and cross-task generalization. Traditional embodied intelligence systems, in complex and ever-changing environments, often rely on single-modal environmental perception information, resulting in incomplete environmental model acquisition and a lack of effective mechanisms for fusing multi-source information. This limitation leads to lower accuracy in perception results when robots face dynamically changing or unstructured environments, thus affecting subsequent task decision-making and path planning performance.
[0003] In the fintech sector, embodied intelligent devices are gradually being applied to systems such as smart teller machines, smart safes, and document transfer robots. However, due to insufficient multimodal information fusion capabilities, these devices suffer from low accuracy in handling environmental interference, autonomous recognition, and path planning in complex financial office scenarios. This results in a high task execution failure rate and an inability to flexibly adapt to environmental changes. Furthermore, existing technologies lack deep understanding and planning based on task semantics, making it difficult for devices to autonomously parse task logic and efficiently generate action sequences based on financial business instructions, thus limiting the level of operational intelligence in dynamic financial transactions.
[0004] In the healthcare sector, intelligent equipment such as fixed robotic arms are widely used for high-precision tasks such as drug sorting, medical supply handling, and sample collection. However, existing embodied intelligence methods are easily affected by factors such as personnel movement and equipment interference in complex medical environments. Environmental modeling relies on single visual perception information, and dynamic obstacle recognition and environmental model updates are not timely, resulting in poor path planning adaptability, low accuracy of robotic arm grasping paths, and problems such as grasping failures, path conflicts, and even affecting overall operational safety. In addition, current technologies lack cross-task generalization capabilities to meet different medical operation needs, making it difficult for robotic arms to quickly adapt to different types of operation tasks, thus hindering the improvement of the intelligence level of medical systems. Summary of the Invention
[0005] The main objective of this invention is to provide a method, apparatus, device, and storage medium for robot arm task planning and control, aiming to solve the technical problems of low efficiency in robot arm path planning, weak task understanding and generalization ability, and poor motion execution accuracy in complex dynamic environments.
[0006] To achieve the above objectives, the present invention provides a robotic arm task planning and control method, comprising:
[0007] Acquire visual information and task description language information, and generate an environment model based on the visual information;
[0008] The visual information and the task description language information are encoded to obtain visual features and language features, respectively.
[0009] Based on the visual features and the language features, a language model is used for task planning to generate a target action sequence.
[0010] Under the aforementioned environmental model, a basic path for the robotic arm's movement is generated through dynamic programming;
[0011] The basic path is adjusted according to changes in the environment through reinforcement learning to generate the final path;
[0012] The target motion sequence is decoded into joint control commands, and motion control signals are generated based on the final path;
[0013] During the execution of the joint control commands, the joint control commands are optimized through reinforcement learning based on feedback rewards obtained from the environment.
[0014] Furthermore, to achieve the above objectives, the present invention provides a robotic arm task planning and control device, comprising:
[0015] An environment perception module is used to acquire visual information and task description language information, and generate an environment model based on the visual information.
[0016] The information encoding module is used to encode the visual information and the task description language information to obtain visual features and language features, respectively;
[0017] The task planning module is used to perform task planning based on the visual features and the language features, and generate a target action sequence using a language model.
[0018] The path generation module is used to generate a basic path for the movement of the robotic arm through dynamic programming under the environment model.
[0019] The path optimization module is used to adjust the basic path according to environmental changes through reinforcement learning to generate the final path;
[0020] A control command generation module is used to decode the target motion sequence into joint control commands and generate motion control signals based on the final path;
[0021] The instruction optimization module is used to optimize the joint control instructions through reinforcement learning based on feedback rewards obtained from the environment during the execution of the joint control instructions.
[0022] Furthermore, to achieve the above objectives, the present invention also provides a computer device, the computer device including a memory, a processor, and a robotic arm task planning and control program stored in the memory and executable on the processor, wherein when the robotic arm task planning and control program is executed by the processor, it implements the steps of the robotic arm task planning and control method as described above.
[0023] Furthermore, to achieve the above objectives, the present invention also provides a computer-readable storage medium storing a robotic arm task planning and control program, wherein when the robotic arm task planning and control program is executed by a processor, it implements the steps of the robotic arm task planning and control method described above.
[0024] Beneficial Effects: This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as embodied intelligence, fintech, and healthcare. It discloses a method, device, equipment, and medium for robotic arm task planning and control, including: acquiring visual information and task description language information; generating an environment model based on the visual information; encoding the visual information and task description language information respectively to obtain visual features and language features; using the language model to perform task planning on the visual features and language features to generate a target action sequence; generating a basic path for robotic arm movement through dynamic programming under the environment model; adjusting the basic path according to environmental changes using reinforcement learning to generate a final path; decoding the target action sequence into joint control commands; generating motion control signals based on the final path; and optimizing the joint control commands through reinforcement learning in conjunction with environmental feedback rewards during the execution of the joint control commands. This invention achieves task planning under multimodal information through the fusion perception of visual information and task description language information, combined with a language model. It adopts a hybrid path planning mechanism combining dynamic programming and reinforcement learning, and further optimizes joint control commands through environmental feedback, improving the path adaptability and task generalization ability of the robotic arm in complex dynamic environments, and achieving efficient and precise robotic arm operation control. Attached Figure Description
[0025] The present invention will be further described below with reference to the accompanying drawings and embodiments. In the accompanying drawings:
[0026] Figure 1 This is a schematic diagram of an application environment for the robotic arm task planning and control method in one embodiment of the present invention;
[0027] Figure 2 This is a flowchart illustrating an embodiment of the robotic arm task planning and control method of the present invention;
[0028] Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the robotic arm task planning and control device of the present invention;
[0029] Figure 4 This is a schematic diagram of the structure of a computer device according to an embodiment of the present invention;
[0030] Figure 5 This is another structural schematic diagram of a computer device according to one embodiment of the present invention. Detailed Implementation
[0031] It should be understood that the specific embodiments described herein are for illustrative purposes only and are not intended to limit the scope of the invention.
[0032] The robotic arm task planning and control method provided in this embodiment of the invention can be applied to, for example... Figure 1 In this application environment, the user terminal communicates with the server via a network. The server can obtain visual information and task description language information from the user terminal, generate an environment model based on the visual information, encode the visual information and task description language information respectively to obtain visual features and language features, use the language model to perform task planning on the visual features and language features, generate a target action sequence, generate a basic path for the robotic arm movement through dynamic programming under the environment model, and combine reinforcement learning to adjust the basic path according to environmental changes to generate a final path. The target action sequence is decoded into joint control commands, and motion control signals are generated based on the final path. During the execution of joint control commands, reinforcement learning is used to optimize the joint control commands in combination with environmental feedback rewards. This invention achieves task planning under multimodal information through the fusion perception of visual information and task description language information, combined with a language model, and adopts a hybrid path planning mechanism that combines dynamic programming and reinforcement learning. It further optimizes joint control commands through environmental feedback, improves the path adaptability and task generalization ability of the robotic arm in complex dynamic environments, and achieves efficient and precise robotic arm operation control. The user terminal can be, but is not limited to, various personal computers, laptops, smartphones, tablets, and portable wearable devices. The server can be implemented using a standalone server or a server cluster consisting of multiple servers. The invention will be described in detail below through specific embodiments.
[0033] Please see Figure 2 , Figure 2 This is a flowchart illustrating an embodiment of the robotic arm task planning and control method provided by the present invention. It should be noted that although the logical order is shown in the flowchart, in some cases, the steps shown or described may be performed in a different order than that shown here.
[0034] like Figure 2 As shown, the robotic arm task planning and control method proposed in this invention includes the following steps:
[0035] S10, acquire visual information and task description language information, and generate an environment model based on the visual information;
[0036] In this embodiment, the process of acquiring visual information is typically accomplished through optical acquisition devices mounted on the robotic arm or in the workspace. These devices include, but are not limited to, RGB cameras, depth cameras, LiDAR, structured light sensors, and binocular vision systems. RGB cameras acquire two-dimensional image information, depth cameras and LiDAR acquire three-dimensional spatial information, and structured light sensors and binocular vision systems acquire higher-precision spatial structure information. The acquired visual information refers to image data, point cloud data, or depth information collected by the aforementioned devices. Specific content includes, but is not limited to, the geometric shape, spatial position, surface color, texture features, edge contours, shadow distribution, and spatial depth data of objects in the scene. Visual information, as a data source for environmental understanding and spatial modeling, is a crucial foundation for subsequent environmental modeling, path planning, and motion control.
[0037] Task description language information refers to the language data provided by users through voice input, text input, or other information interaction methods to specify the target of the robotic arm operation. The sources of language information include, but are not limited to, natural language voice commands, written text input, symbolic encoding information, and console interface commands. The language information includes various information such as the category of the object being operated on, its spatial location, operation method, action sequence requirements, and execution order restrictions. The flexibility of language information in its expression reduces the operational threshold between users and the robotic arm system, facilitating the rapid deployment and adjustment of human-machine collaborative operation tasks.
[0038] After acquiring visual information, an environment model is generated using this information. The environment model is a spatial data representation that combines visual information to reflect the spatial structure of the work environment, obstacle locations, accessible areas, and information about the objects being manipulated. The environment model can be expressed as a 3D point cloud map, a spatial grid diagram, a topology diagram, a spatial voxel representation, a continuous surface reconstruction result, or a data integration model that combines multiple representations. In the specific generation process, the acquired visual information is first preprocessed, including image denoising, color correction, geometric distortion correction, coordinate transformation, and scale normalization, to improve the accuracy and reliability of the data. Subsequently, through image segmentation, object detection, boundary extraction, deep fusion, and spatial registration, the image information is combined with 3D spatial information to extract obstacles, objects being manipulated, and other spatial structures in the scene. Based on this information, an environment model containing both static and dynamic elements is gradually constructed through point cloud stitching, deep fusion, or spatial mapping. This environment model not only describes spatial location and structural information but also retains the boundary features of obstacles, the movement trends of dynamic obstacles, and the semantic information of the objects being manipulated, providing spatial data support for subsequent path planning and action generation.
[0039] The acquisition of visual information can be achieved by selecting different sensing devices depending on the specific work scenario. In industrial production lines, fixed RGB-D cameras can be used to collect continuous scene images and depth data. In the healthcare sector, miniature 3D camera systems mounted on the end effector of robotic arms can be used to acquire data on the patient's body surface structure or surgical environment, improving the safety and accuracy of spatial operations. In the fintech sector, for automated physical asset inspection within data centers, LiDAR combined with high-precision cameras can be used to acquire information on equipment layout and spatial obstacles.
[0040] Task description language information can be flexibly adjusted based on different human-computer interaction systems. In industrial scenarios, operation commands and task parameters can be input through a configured touch screen input interface, industrial control panel, or mobile application, improving operational efficiency. In the healthcare field, doctors can issue operation goals through voice interaction systems, such as specifying the location of medicines and operation procedures, improving the convenience and intelligence of medical operations assisted by robotic arms. In the fintech field, maintenance personnel can issue inspection tasks, asset inventory instructions, and spatial positioning parameters through natural language text systems, enabling efficient remote deployment.
[0041] The generation process of the environmental model can be adjusted according to the selected visual information expression form. If a monocular camera is used to acquire two-dimensional image data, an environmental structure model can be constructed by combining image recognition, planar reconstruction, and geometric reasoning. If a depth camera or LiDAR is used, spatial registration, mesh reconstruction, and obstacle recognition can be performed directly using the acquired three-dimensional point cloud data to form a complete spatial environment model. In the field of healthcare, the environmental model can construct three-dimensional spatial information of the surgical environment by fusing image data and spatial scanning data, assisting robot path planning and precise operation. In the field of fintech, the environmental model generates a high-precision three-dimensional spatial map by scanning the spatial structure and asset distribution information of the computer room, improving the efficiency and safety of inspection and operation path planning.
[0042] Example Description: In the field of embodied intelligence, for fixed robotic arms performing complex tasks, the system deploys high-resolution RGB-D cameras and LiDAR to jointly acquire visual information of the work area. This captures in real-time the geometric shape and spatial position of objects in the environment, the trajectories of dynamic obstacles, and changes in background structure. Combined with the task description information given by the operator via a voice system, specifying the target item category, operational requirements, and action sequence, the system fuses and processes this information to generate a high-precision 3D environment model containing information on environmental structure, obstacle distribution, dynamic trends, and the manipulated object. This environment model supports dynamic updates and multi-scenario adaptation, ensuring that the robotic arm can still perform path planning and operational control based on complete and accurate spatial information even when the environment changes or the task is adjusted. This improves the robotic arm's adaptability, robustness, and operational efficiency in complex environments and diverse tasks, meeting the requirements for flexible deployment and generalization of embodied intelligence systems in typical scenarios such as manufacturing, logistics, and collaboration.
[0043] In the field of healthcare, a high-precision vision system installed at the end of a surgical robotic arm is used to acquire environmental information of the surgical area. Combined with the operation target instructions given by the doctor through the voice system, an environmental model is generated in real time. The model accurately describes the surgical space structure, the location of obstacles, and the operation path, effectively assisting the robotic arm in completing precise operations in complex spaces and improving surgical safety and efficiency.
[0044] In the fintech business, for data center environments, multi-sensor systems are used to acquire information on equipment spatial layout, obstacle locations, and environmental changes. Operation and maintenance personnel issue inspection and maintenance task instructions through a natural language system. The system generates an environmental model based on comprehensive environmental information and task requirements, ensuring that the robotic arm can safely and efficiently complete inspection, retrieval, or emergency intervention tasks in high-density spaces, reducing the risk of manual operation and improving the intelligent management level of fintech infrastructure.
[0045] This embodiment combines visual information with task description language information, and utilizes multi-source data to jointly participate in environment modeling. This effectively improves the completeness of the environment model, the accuracy of spatial structure expression, and the semantic relevance of operation task instructions. It solves the problems in the prior art where robotic arms rely on a single data source, resulting in incomplete modeling information, inaccurate understanding of spatial structure, and missing task semantic information, leading to high failure rates in path planning and operation execution and poor generalization ability. This improves the operational stability and task adaptability of the robotic arm in complex environments.
[0046] S20, the visual information and the task description language information are encoded to obtain visual features and language features respectively;
[0047] In this embodiment, visual information refers to multi-dimensional data acquired through image, video, or 3D sensing devices that reflects information such as the structure of the external environment, object shape, spatial location, surface features, and color texture. Sources include data sets output by devices such as monocular or binocular cameras, depth cameras, LiDAR, and structured light sensors. The spatial location, geometric contours, edge features, color distribution, and material information contained in the visual information constitute the basic perception data of the external environment that the robotic arm needs to acquire before performing operations. Task description language information refers to language data expressed through speech recognition, natural language input, or symbol sequences that can describe the operation target, action requirements, task logic, or operation instruction sequences. Sources include text input interfaces, voice interaction systems, and host computer instruction systems. Language information provides a complete expression of the operation object, task intent, action type, and execution conditions. Encoding visual information refers to converting multi-dimensional visual information into low-dimensional, structured numerical representations that express the spatial characteristics of the environment through deep learning models or other information compression and abstraction mechanisms. Specifically, convolutional neural network structures can be used to extract spatial features. Local texture and boundary features are extracted through multi-layer convolution operations, and global receptive fields are superimposed to capture spatial structural information. Finally, the visual features are output after dimensionality compression. Visual features are a set of vectors that reflect the environmental structure, object distribution, and spatial relationships.
[0048] Encoding task description language information involves converting natural language task information into numerical, computable semantic feature representations through word embedding mechanisms, sequence models, or language understanding structures. Specifically, language information can be mapped to a sequence of word vectors using a word embedding matrix. Further, recurrent neural networks, long short-term memory networks, or transformer structures are used to extract the temporal dependencies and semantic information of the language sequence, ultimately generating a fixed-dimensional language feature vector. These language features are used to express task objectives, operational requirements, and logical dependencies. Visual and language features work together in subsequent information fusion processes to guide task planning and action generation, ensuring that the system possesses both environmental awareness and the ability to understand task objectives and operational requirements.
[0049] In the specific implementation process, the visual information acquisition part can connect to various types of sensing devices. For indoor environments, a combination of RGB-D cameras and stereo vision is preferred to acquire high-precision scene images and depth information. For outdoor or large-space scenes, LiDAR can be deployed to acquire 3D structural information of the environment. The encoding part is built based on a deep convolutional network structure. The convolutional layers extract edge, color, and texture features of local regions through sliding windows, and the pooling layers compress spatial dimensions to enhance feature representation capabilities. Finally, a fixed-length visual feature vector is output through a multi-layer structure. The language information input part can be configured with a speech recognition module or a text command interface according to actual task requirements. The text information is mapped into a word vector sequence after word segmentation. The sequence feature extraction structure can adopt a bidirectional long short-term memory network, a gated recurrent unit, or a multi-layer transformer structure to extract temporal dependencies, logical structures, and key semantic information in the language sequence. Finally, a fixed-length language feature representation is output through feature aggregation operations. The system supports multi-task concurrency and multimodal input fusion. Visual features and language features are jointly input into the downstream task planning module to achieve efficient generation of operation strategies.
[0050] Example Description: In the field of healthcare, surgical assistive robotic arm systems use cameras to collect real-time visual information of the operating table area. Combined with the surgical task description information input by the doctor through voice or text, the system encodes and obtains high-dimensional environmental structure and surgical target expression. The system uses visual features to identify the location of organs, surgical instruments and operating space, and uses linguistic features to express the operation steps and key parts information. These are used together for subsequent surgical path planning and precise operation, improving the safety and accuracy of the surgical process.
[0051] In the fintech business, robotic automated teller equipment uses cameras to collect image information of the operating area, combines it with the user's input of transaction operation language information, encodes and obtains the counter environment structure and business operation instructions, uses visual features to locate the operation interface and physical buttons, and uses language features to guide transaction logic and instruction parsing. The system efficiently completes automated transaction operations based on these two types of information, improving the intelligence level of the financial service system and the user experience.
[0052] In the field of embodied intelligence, mobile collaborative robotic arm systems acquire visual information of the workspace through multimodal sensors, combine it with task description language information received during human-computer interaction, and encode and generate environmental structural features and task intent expressions. This enables the system to accurately understand the operation objects and action goals in complex dynamic environments, and improves multi-task collaboration capabilities and generalization adaptability.
[0053] This embodiment maps environmental perception data and task description information into structured visual and linguistic features, respectively, compressing high-dimensional complex information into low-dimensional expressions, preserving key spatial structures, environmental attributes, and task intentions, improving information fusion and decision-making efficiency, reducing the system's computational dependence on raw data, and enhancing the robotic arm's task understanding and operational accuracy in complex dynamic environments.
[0054] S30, Based on the visual features and the language features, use a language model to perform task planning and generate a target action sequence;
[0055] In this embodiment, visual features refer to fixed-length numerical representations that reflect the spatial structure of the environment, the distribution of object positions, and spatial relationships, obtained by performing spatial structure extraction, local detail compression, and global relationship modeling on visual information. Visual features can be extracted through multi-layer convolutional neural networks or transformer structures, and have the ability to express the appearance attributes of objects, spatial layout, and environmental boundaries. Linguistic features are obtained by mapping task description language information into a low-dimensional, structured set of semantic vectors that express task goals and operational requirements through word embedding, sequence models, and semantic aggregation. Linguistic features reflect operational logic, task goals, and intent information.
[0056] Based on visual and linguistic features, the language model performs cross-modal information fusion, which means jointly expressing the visual features that express the structure of the environment and the linguistic features that express the intention of the task in a high-dimensional feature space. The fusion process can use multilayer perceptron, cross-modal transformer structure or attention mechanism to build a dynamic correlation mapping between visual and linguistic information, capture the intrinsic relationship between the environmental state and the task logic, and form a joint feature expression.
[0057] A language model is a deep neural network structure that has the ability to understand task logic and generate sequences. Specifically, it includes sequence-to-sequence models based on transformer structures, semantic reasoning networks with conditional encoding mechanisms, or multi-layer recurrent neural networks. Based on the fused joint features, a language model can perform task logic analysis, operation intent parsing, and action sequence generation.
[0058] Task planning refers to the process by which a language model, based on joint feature input, breaks down the global task intent into ordered operational steps that satisfy environmental constraints through semantic reasoning and task logic decomposition. During task planning, the language model performs multi-level feature transformations and semantic abstraction on the joint features, extracting key task objectives, operational constraints, and action dependencies. It then constructs a task logic dependency graph, clarifying the execution order, spatial location, and relationships between the operational objects of each action step.
[0059] Based on the task logic dependency graph, the language model further transforms the high-level task description into a low-level, atomic sequence of actions oriented towards robotic arm operation. The action sequence includes specific instructions such as spatial movement, posture adjustment, and end effector operation. Each target action is an executable operation unit of the system in physical space. The final action sequence reflects all the operation steps required for the robotic arm to complete the overall task, and has spatial continuity, task dependency integrity, and environmental adaptability.
[0060] Visual features can be extracted using deep convolutional neural network structures. Inputting environmental images or point cloud information, multi-layer convolution and pooling operations are used to extract spatial structure, texture boundaries, and object location distribution information. Combined with a global receptive field structure, the overall environmental layout is captured, and a fixed-length visual feature vector is output, supporting multi-dimensional environmental perception representation. Linguistic features can be extracted using text embedding and sequence modeling structures. Task description language information is processed through word segmentation and mapped to a sequence of word vectors. Inputting this sequence into a long short-term memory network, gated recurrent units, or transformer structures, task logic, temporal dependencies, and semantic information are extracted, and a fixed-length linguistic feature representation is output.
[0061] In the process of cross-modal information fusion, a multimodal transformer structure based on an attention mechanism can be adopted to dynamically calculate the correlation weights between visual and linguistic features, jointly express environmental structure and task intent information, and generate fused features. Based on the fused feature input, the language model uses a multi-layer semantic reasoning structure, graph neural network, or sequence generation structure to analyze task logical dependencies, construct a task logic graph, and output ordered task operation steps.
[0062] In the specific task planning process, the language model combines the input of the environmental model and the kinematic constraints of the robotic arm. Based on the spatial layout, obstacle information and task requirements, it dynamically generates a sequence of target actions that meets the requirements of safety, accessibility and efficiency. The action sequence includes the target position, posture parameters, end effector type and time constraints, ensuring the spatial continuity and task adaptability of the robotic arm operation.
[0063] Example Description: In the field of healthcare, surgical robot systems acquire images of surgical scenes through visual sensors, generate environmental spatial structure features, and combine these with surgical operation descriptions input by doctors via voice to generate language features. Based on the joint feature input, the system uses a language model for task planning, generating target action sequences that include operation steps such as cutting, suturing, and object movement. This guides the robotic arm to perform operation tasks with high precision in complex surgical environments, improving surgical safety and operational efficiency.
[0064] In the fintech business sector, intelligent service robot systems acquire images of the operating space through environmental visual sensing devices, generate counter environment layout information, and combine them with the user's input of business processing needs in language information to generate language features. Based on the joint features, the system uses a language model to generate the sequence of operational actions required in the processing flow, including task steps such as identity verification, information entry, and document transfer, thereby improving the operational accuracy and service efficiency in the financial business processing process.
[0065] In the field of embodied intelligence, the multi-task collaborative robotic arm platform acquires visual information in complex dynamic environments through a multimodal sensing system. Combined with task language information in human-computer interaction, it generates a joint feature input language model. Based on semantic reasoning and task logic modeling, the system generates target action sequences that are adaptable across tasks and scenarios. This supports the robotic arm in achieving flexible and efficient operation execution and task collaboration in unstructured environments, improving the system's environmental adaptability and multi-task processing capabilities.
[0066] This embodiment utilizes joint information expression based on visual and linguistic features, combined with multi-layer semantic reasoning and task logic modeling capabilities of language models. This enables the system to efficiently and accurately parse environmental structures and task requirements, generating target action sequences with spatial executability, task logic integrity, and environmental adaptability. This improves the robotic arm's task understanding and operational planning efficiency in complex environments, and enhances the system's generalization ability and task execution reliability.
[0067] S40, Under the aforementioned environmental model, a basic path for the movement of the robotic arm is generated through dynamic programming;
[0068] In this embodiment, the environment model refers to a three-dimensional spatial representation structure that reflects all known information within the robotic arm's operating space, formed through a structured representation of environmental data such as surrounding spatial information, object distribution, obstacle locations, and boundary structures. The environment model can be represented using a three-dimensional raster map, point cloud dataset, or mesh structure, and includes detailed descriptions of spatial occupancy information, obstacle shapes, and passable areas, providing a basic source of spatial perception information.
[0069] Generating a basic path within the environmental model refers to constructing a mathematical space representation suitable for path search and optimization based on the aforementioned environmental model, combined with the robotic arm's own kinematic constraints, joint reachability, and spatial obstacle distribution. Then, dynamic programming is used to plan a path sequence from the initial position to the target position for the robotic arm's motion. The basic path, under the environmental structure and robotic arm motion constraints, is a robotic arm motion sequence that satisfies feasibility, safety, and path coherence, obtained recursively by traversing the state space and action space using dynamic programming.
[0070] Dynamic programming is a path optimization method based on the theory of optimal substructure and overlapping subproblems. It finds the global optimal solution step-by-step through recursive decomposition and state transitions. In its implementation, the environment state space is first defined based on an environment model. This state space includes a combination of the robot arm's position, posture, joint angles, and spatial obstacle information, covering all possible robot arm state configurations. Simultaneously, the robot arm motion space is defined, reflecting the set of controllable movements of each joint, including position changes, posture adjustments, and opening / closing movements of the manipulator units.
[0071] Using dynamic programming, for each environmental state, the optimal value function for reaching the target state from that state is recursively calculated. The value function can be set based on path length, operational energy consumption, obstacle avoidance efficiency, or time cost, reflecting the quality of the path under different states. During the dynamic programming recursion process, based on the target state and initial state calibrated in the environmental model, the value iteration process is executed in reverse, gradually backtracking to find the state sequence that connects each state node and meets the path optimality requirement.
[0072] Finally, based on the above state sequence, a basic path is generated. The basic path is a continuous motion path that the robotic arm can directly execute in the spatial environment. It ensures that the path meets the requirements of spatial accessibility, obstacle avoidance safety and operational stability, thus forming a preliminary path scheme for the robotic arm's movement.
[0073] The environment model can be constructed by fusing visual sensing data, depth sensing information, and prior map data. The generated environment model can be represented using a multi-resolution 3D raster, including the positions of static obstacles, the states of dynamic environmental objects, and information on passable areas. The environmental state space is discretized, dividing the robotic arm's operating space into a finite set of state nodes. Each state node describes the robotic arm's spatial position, posture, and relationship with the environment. The motion space is set according to the robotic arm's structure and joint mobility, including linear translation, posture adjustment, or end effector commands.
[0074] In dynamic programming methods, value iteration, policy iteration, or heuristic-guided dynamic programming algorithms can be employed. Based on the environmental state space and action space, the optimal value function is recursively calculated, and combined with environmental model information, the optimal path connecting the initial state and the target state is searched step by step. In the specific path generation process, spatial obstacle avoidance constraints, robotic arm joint kinematic constraints, and operational safety indicators can be introduced. During dynamic programming, a cost function is incorporated to balance path length, operational stability, and obstacle avoidance efficiency, generating a basic path.
[0075] Example Description: In the field of healthcare, surgical robotic arms construct an environmental model of the operating room space, calibrating the positions of the operating table, medical equipment, and operating area. Based on this environmental model, the system uses dynamic programming to plan the basic path, ensuring that the robotic arm can move smoothly and avoid obstacles safely during tasks such as suturing and retrieving objects, effectively improving the accuracy and safety of surgical operations.
[0076] In the fintech business, intelligent handling robotic arms construct spatial environment models in bank vaults or self-service equipment maintenance environments. Combining the on-site environment layout and equipment location, they use dynamic programming methods to generate basic paths, enabling the robotic arms to efficiently and safely move valuables or perform equipment maintenance tasks within the vault, thereby improving the security and efficiency of operations in financial venues.
[0077] In the field of embodied intelligence, the dual-arm collaborative robot system constructs a dynamic environment model by acquiring visual information and depth perception data of the work area in real time. Based on this environment model, the system uses dynamic programming to generate basic paths for the dual-arm collaborative tasks of the robotic arms, ensuring that the path planning takes into account spatial obstacle avoidance, collaborative operation and motion stability. This supports the robotic arms to achieve high-precision operation and flexible path adjustment in complex dynamic environments, and enhances the autonomous operation capability and complex task processing capability of the embodied intelligence system.
[0078] This embodiment utilizes dynamic programming path generation based on an environment model to fully leverage environmental structural information and the kinematic characteristics of the robotic arm. The dynamic programming algorithm globally searches for paths within the state space, ensuring that the generated basic paths possess optimality, feasibility, and environmental adaptability. This effectively avoids local optima traps in path planning, improves the efficiency and quality of path generation for the robotic arm in complex environments, provides a reliable initial path foundation for subsequent path optimization and dynamic adjustment, and enhances the stability of the overall operating system and the reliability of task execution.
[0079] S50, through reinforcement learning, adjusts the basic path according to changes in the environment to generate the final path;
[0080] In this embodiment, environmental change refers to changes in the position of obstacles, objects, dynamic interference factors, or other environmental states within the spatial environment during the operation of the robotic arm. Environmental changes may include the movement of obstacles, changes in the posture of objects, adjustments to the structure of the work area, or interference from external factors. The basic path is a preliminary path generated by the robotic arm based on the environmental model using a dynamic programming method. This path has theoretical optimality under static environmental assumptions. However, due to the dynamic changes in the actual operating environment, the basic path is difficult to fully adapt to the real-time environmental state, and may suffer from path deviation, obstacle avoidance failure, or reduced operating efficiency.
[0081] Reinforcement learning is a machine learning method that uses an agent to interact with its environment, continuously collect feedback information, optimize policy models based on reward mechanisms, and gradually improve the agent's decision-making ability and adaptability. In the process of adjusting the path of a robotic arm, reinforcement learning dynamically corrects the basic path based on changes in the environment, generating a final path that adapts to the current environmental state, has higher operational efficiency, and stronger environmental adaptability.
[0082] In practice, the system first monitors changes in the environmental state within the robotic arm's operating area in real time. This is achieved through visual sensing, depth perception, or position tracking devices to acquire environmental change information, including the new positions of obstacles, the movement trajectories of dynamic objects, and the detection results of unforeseen environmental events. Based on this environmental change information and the basic path, a reinforcement learning state representation is constructed. This representation encompasses the robotic arm's current position, environmental structure, obstacle distribution, and basic path information, forming the input state space for reinforcement learning decisions.
[0083] The reinforcement learning state representation is input into the reinforcement learning policy network. The policy network outputs path adjustment actions based on the current environment state. The path adjustment actions include offset adjustment of basic path nodes, replanning of path segments, or local optimization of motion trajectory. The path adjustment actions can correct path segments in the basic path that are affected by environmental changes in real time, thereby improving the feasibility and safety of the path.
[0084] The robotic arm performs path adjustment actions, moving according to the updated path segment. The system generates reward signals based on environmental feedback, with the reward signals set according to path execution effectiveness, obstacle avoidance success rate, path smoothness, or operational efficiency, quantifying the degree of optimization in the path adjustment. Through a reinforcement learning strategy, the policy network parameters are updated based on the reward signals. After the policy network parameters are optimized, the agent's path adjustment capability under similar environmental changes is improved.
[0085] The system determines whether the adjusted path meets the optimization criteria, which include shortening the path length, improving obstacle avoidance success rate, or improving operational efficiency. If the optimization criteria are met, the adjusted path is used as the final path for the robotic arm to execute. If the optimization criteria are not met, the system continues to collect environmental change information and repeats the path adjustment and strategy optimization process until the final path meets the optimization criteria, ensuring that the path planning results dynamically adapt to environmental changes.
[0086] Environmental change information can be acquired based on high-frequency visual image acquisition and depth perception, capturing obstacle movement, object deformation, or changes in environmental structure through real-time 3D reconstruction technology. Reinforcement learning state representation is constructed by fusing information from the robotic arm's position, joint angles, environmental grid map, and basic path nodes. The policy network can employ deep reinforcement learning structures, such as deep Q-networks, proximal policy optimization networks, or graph-based policy generation models.
[0087] During the generation of path adjustment actions, the policy network outputs specific path node offsets or local path segment replacement schemes by combining basic path information and environmental change status. The system then integrates the local path adjustment results with the unaffected basic path segments through the path replanning module to form a continuous and executable adjustment path.
[0088] The reward signal can be set based on the smoothness of the path, the success rate of obstacle avoidance, and the overall operational efficiency of the robotic arm operation after path adjustment. The system monitors the path execution effect in real time, improves the path optimization effect through positive incentives, and suppresses suboptimal path selection through negative feedback. The policy network updates parameters through gradient descent to improve the adaptive capability of the path adjustment strategy.
[0089] Example Description: In the field of healthcare, when the surgical environment changes during surgery, such as changes in patient positioning, surgical instrument position, or changes in real-time image data, the system dynamically adjusts the basic path based on reinforcement learning to generate a final path that adapts to the changes in the surgical environment. This ensures the accuracy of the robotic arm operation and the safety of the surgery, reduces the operational risks caused by environmental interference, and improves the efficiency of medical operations and surgical outcomes.
[0090] In the fintech business, intelligent logistics robotic arms face challenges such as personnel movement, changes in the stacking of goods, or adjustments to the position of equipment in bank vaults, financial data centers, or self-service equipment maintenance sites. The system monitors the environmental status in real time and dynamically optimizes the basic path through reinforcement learning to generate the final path that adapts to the real-time environment. This improves the flexibility and safety of the robotic arm's handling, inspection, and maintenance operations, ensuring the operational continuity and equipment stability in financial business scenarios.
[0091] In the field of embodied intelligence, dual-arm collaborative robot systems face challenges such as moving obstacles, changes in collaborating objects, or external environmental interference in complex industrial production lines or dynamic warehousing environments. The system dynamically adjusts the basic path through reinforcement learning strategies to generate a real-time adapted final path, supporting the robot system to maintain the feasibility of the path and operational stability in highly dynamic environments, and improving the adaptive capability and multi-task collaboration level of embodied intelligence systems.
[0092] This embodiment utilizes a reinforcement learning-based path dynamic adjustment method to achieve dynamic adaptive optimization of the robotic arm's basic path in actual operating environments. By interacting with the environment in real time, reinforcement learning captures information about environmental changes and dynamically adjusts the basic path, effectively improving the flexibility and adaptability of path planning. This avoids path failure or operational risks caused by environmental changes, enhances the stability and operational efficiency of the robotic arm in complex and dynamic environments, and improves the overall intelligent decision-making level and task completion quality of the operating system.
[0093] S60, decode the target motion sequence into joint control commands, and generate motion control signals based on the final path;
[0094] In this embodiment, the target action sequence refers to a sequence of action instructions generated by a language model based on visual and linguistic features, expressing the actions required by the robotic arm during task execution. The target action sequence is typically expressed in abstract action codes or symbolic form, containing information such as the operation target, spatial position, action category, or action parameters. It lacks a direct mapping relationship with the specific control system of the robotic arm and cannot directly drive the robotic arm to perform actual operations. To realize the conversion of the target action sequence into the underlying control instructions of the robotic arm, the target action sequence needs to be decoded into specific joint control instructions.
[0095] Joint control commands refer to the specific control parameters for each joint execution unit of a robotic arm. These typically include target position, velocity, acceleration, or torque commands for each joint. Joint control commands possess clear physical meaning and real-time capability, directly driving the movement of the robotic arm joints to complete spatial manipulation tasks. The process of decoding a target motion sequence into joint control commands actually involves using a motion decoder to perform feature analysis and parameter mapping on the abstract motion sequence, generating specific control parameters adapted to the robotic arm execution system.
[0096] During the target motion sequence decoding process, the current joint state information of the robotic arm is first acquired. This information includes the position, angle, velocity, torque, or sensor feedback data of each joint, reflecting the current physical state and motion capability of the robotic arm. The current joint state and the target motion sequence are then input into the motion decoder. The motion decoder analyzes the target motion sequence based on the robotic arm's kinematic model and control constraints, and calculates and generates joint control parameters corresponding to the robotic arm's physical structure.
[0097] Joint control parameters are converted into specific joint control commands through instruction encoding or signal encapsulation. The commands cover joint angle setting, target trajectory following, motion speed control and torque limitation, ensuring that the robotic arm has high-precision and high-stability control capabilities during the execution of the target action sequence.
[0098] Based on the generated joint control commands, motion control signals are generated by combining the final path. The final path is an operation path formed by dynamically adjusting the base path through reinforcement learning, possessing real-time environmental adaptability and optimal execution efficiency. The motion control signal refers to the spatial trajectory information that guides the robotic arm to perform the operation task according to the final path. The motion control signal typically includes the spatial position of the end effector, attitude parameters, path node sequence, or trajectory interpolation information, reflecting the robotic arm's motion trajectory and path following requirements in space.
[0099] In practical implementation, based on the final path information, the system analyzes the sequence of key points in the path and generates continuous and smooth spatial trajectory instructions through algorithms such as path interpolation, trajectory smoothing, or inverse kinematics, forming motion control signals. These motion control signals, along with joint control instructions, are input into the robotic arm control system to guide the robotic arm to precisely execute spatial operation tasks according to the predetermined final path and action sequence, achieving collaborative control and efficient task execution through multi-source information fusion.
[0100] The motion decoder can be designed based on a neural network structure, such as a recurrent neural network, a converter network, or a multimodal fusion network. Combined with the kinematic constraints and dynamic model of the robotic arm, it analyzes the target motion sequence and joint state information to generate joint control parameters that meet the structural and control requirements of the robotic arm. After the joint control parameters are generated, they are encapsulated into joint control commands using standard control command protocols, such as CAN bus protocol, EtherCAT protocol, or industrial Ethernet protocol, ensuring the real-time performance and reliability of command transmission.
[0101] The final path information is obtained through path reconstruction, 3D environment modeling, and dynamic path adjustment algorithms. Based on the sequence of key points of the final path parsing trajectory, the system uses spline interpolation, B-spline curves, or fifth-order polynomial trajectory planning methods to generate spatial trajectory commands, forming motion control signals. The motion control signals are converted into joint spatial position commands through inverse kinematics algorithms. These commands work in conjunction with the joint control commands to generate underlying drive signals through the robotic arm controller, driving each execution unit of the robotic arm in real time to complete complex spatial operation tasks.
[0102] Example: In the healthcare field, surgical robotic arms decode the sequence of surgical tasks into specific joint control commands, combine them with the final path dynamically adjusted during the operation, and generate high-precision motion control signals to drive the robotic arm to complete minimally invasive surgical operations, thereby improving the accuracy and safety of surgical operations and reducing the risk of trauma to patients.
[0103] In the fintech business sector, logistics and delivery robotic arms are jointly controlled by target motion sequence decoding and final path control signals to adapt to the complex obstacle distribution and dynamic changes in financial data center or vault environments. This improves the flexibility and efficiency of goods handling, equipment maintenance and inspection operations, and ensures the continuity of operations and equipment stability in financial business scenarios.
[0104] In the field of embodied intelligence, dual-arm collaborative robot systems decode the target motion sequence of complex task planning into joint control instructions, combine them with the final path information after dynamic adjustment by multiple robot systems, and generate multi-channel motion control signals to achieve multi-arm collaboration, dynamic obstacle avoidance and adaptive operation in complex environments, thereby improving the task execution capability and multi-task collaboration level of embodied intelligence systems in dynamic environments.
[0105] This embodiment achieves the effective transformation of abstract task planning results into the underlying control system of the robotic arm through target action sequence decoding and final path motion control signal generation. This ensures the accuracy and operational flexibility of the robotic arm in complex environments, and improves its spatial manipulation capabilities and environmental adaptability. By jointly considering the target action sequence and final path information, the system implements a dynamic control strategy based on multimodal information fusion. This enhances the robotic arm's autonomous operation capabilities and task completion efficiency in highly dynamic and unstructured environments, reduces errors and safety risks during operation, and improves the overall system reliability and intelligence.
[0106] S70, during the execution of the joint control command, the joint control command is optimized through reinforcement learning based on the feedback reward obtained from the environment.
[0107] In this embodiment, joint control commands refer to a set of specific parameters generated by combining the target motion sequence, the robotic arm's kinematic model, and the final path information. These commands are used to directly control the execution units of each joint of the robotic arm, typically including joint angle, position, velocity, acceleration, or torque commands. This ensures that the robotic arm accurately executes the operation task according to the planned motion sequence and path trajectory. During the execution of joint control commands, the robotic arm completes posture adjustment, spatial movement, or operation tasks within the physical space according to the set instructions. This process is susceptible to environmental uncertainties, system errors, or external interference, leading to deviations between the joint control commands and the actual operation results.
[0108] To dynamically compensate for environmental disturbances and system uncertainties, and to achieve adaptive optimization of joint control commands, the system dynamically updates and optimizes joint control commands based on feedback rewards obtained from the environment, combined with a reinforcement learning mechanism. Feedback rewards in the environment refer to performance indicators reflecting task performance and environmental adaptability obtained through an environmental sensing system or state observation module during the robotic arm's task execution. These rewards can be quantified numerically to characterize key performance aspects such as task completion quality, path following error, energy efficiency, or operational safety.
[0109] Reinforcement learning is a learning method based on interaction with the environment, autonomous learning, and strategy optimization. The system dynamically adjusts control strategies or command parameters by analyzing feedback rewards to improve operational performance and environmental adaptability. In practice, the system continuously monitors the execution effect of joint control commands, acquiring real-time information on the robotic arm's end-effector pose, task completion status, or external environmental changes through environmental sensors, vision systems, or force feedback modules. Combined with the expected task outcome, it calculates the feedback reward value.
[0110] Based on feedback rewards, the system utilizes reinforcement learning strategies to adjust or optimize joint control commands. Specifically, this includes using a policy network or value function to calculate the direction and amplitude of command optimization based on reward signals, updating control parameters, dynamically compensating for system errors and environmental interference, and improving the accuracy of joint control command execution and task completion efficiency. This process can employ proximal policy optimization (PPO), deep deterministic policy gradient (DDPG), or other adaptive reinforcement learning algorithms to ensure the system possesses continuous learning and autonomous optimization capabilities in complex dynamic environments.
[0111] Environmental feedback rewards can be obtained based on high-precision vision sensors, LiDAR, torque sensors, or inertial measurement units to perceive in real time the robotic arm's end-effector position deviation, path tracking error, grasping stability, or obstacle avoidance performance, and generate quantified reward signals by combining these with pre-set task indicators. The reinforcement learning strategy can be implemented through a policy network or value network constructed using deep neural networks, combining environmental states and historical rewards to output optimized adjustment parameters for joint control commands.
[0112] In practical systems, the reward function design can be dynamically adjusted according to specific needs for different operational tasks and environmental conditions. A multi-dimensional reward structure can be set to balance operational efficiency, energy consumption optimization, and system safety, thereby improving the effectiveness and stability of the reinforcement learning process. During the joint control command optimization process, the system uses soft update mechanisms, parameter constraints, or safety boundary settings to ensure that optimization adjustments do not cause system instability or operational loss of control, thus guaranteeing the high reliability and operational safety of the robotic arm in complex environments.
[0113] Example: In the healthcare field, robotic arms assisting in minimally invasive surgery dynamically acquire feedback rewards based on surgical field feedback, tissue contact information, and force feedback signals. Through reinforcement learning, they optimize joint control commands, compensate for operational deviations caused by changes in patient position, differences in tissue flexibility, or minor errors, and improve the precision and safety of surgical procedures.
[0114] In the fintech business, intelligent warehousing robotic arms, based on environmental monitoring information and operational feedback, optimize joint control commands in real time within financial data center environments. This allows them to adapt to dynamic obstacles, equipment layout changes, or high-density environments within the storage area, improving the flexibility and efficiency of financial equipment maintenance, goods handling, and data center operations, and ensuring the stable operation of financial systems.
[0115] In the field of embodied intelligence, collaborative multi-arm systems jointly acquire multi-source environmental feedback rewards and dynamically optimize the joint control commands of each arm based on reinforcement learning mechanisms. This enables multi-arm collaboration, dynamic obstacle avoidance, and adaptive control in complex operation tasks, thereby enhancing the task execution capability and multi-task parallel collaboration level of embodied intelligence systems in unstructured environments.
[0116] This embodiment addresses the execution deviation issues caused by system errors, environmental changes, or external interference during robotic arm operation in complex dynamic environments by optimizing joint control commands based on environmental feedback rewards and reinforcement learning. This improves the accuracy of robotic arm joint control and the system's environmental adaptability. Through real-time acquisition of environmental feedback and autonomous learning based on reward signals, the system possesses continuous optimization and adaptive adjustment capabilities, enhancing the robotic arm's robustness and task completion efficiency in complex tasks and dynamic environments, and improving the overall system's intelligence level and operational stability.
[0117] This invention relates to the field of artificial intelligence technology and can be applied to business scenarios such as embodied intelligence, fintech, and healthcare. It discloses a method, device, equipment, and medium for robotic arm task planning and control, comprising: acquiring visual information and task description language information; generating an environment model based on the visual information; encoding the visual information and task description language information respectively to obtain visual features and language features; using the language model to perform task planning on the visual features and language features to generate a target action sequence; generating a basic path for robotic arm movement through dynamic programming under the environment model; adjusting the basic path according to environmental changes using reinforcement learning to generate a final path; decoding the target action sequence into joint control commands; generating motion control signals based on the final path; and optimizing the joint control commands through reinforcement learning in conjunction with environmental feedback rewards during the execution of the joint control commands. This invention achieves task planning under multimodal information through the fusion perception of visual information and task description language information, combined with a language model, and employs a hybrid path planning mechanism combining dynamic programming and reinforcement learning. Furthermore, it optimizes joint control commands through environmental feedback, improving the path adaptability and task generalization ability of the robotic arm in complex dynamic environments, and achieving efficient and precise robotic arm operation control.
[0118] In one embodiment, step S10 includes:
[0119] S101 acquires scene image sequences as visual information using an optical sensor;
[0120] S102 receives task instructions as task description language information through a speech recognition module or text input interface;
[0121] S103, Perform object detection on the scene image sequence to obtain object spatial coordinates and object category labels;
[0122] S104, Identify the outlines of dynamic obstacles and the boundaries of static obstacles in the scene image sequence;
[0123] S105, construct a three-dimensional environment grid map based on the object's spatial coordinates, dynamic obstacle outlines, and static obstacle boundaries;
[0124] S106, the three-dimensional environment grid map is labeled as an environment model.
[0125] In this embodiment, visual information refers to a set of data acquired through optical imaging devices that reflects the spatial structure, object features, and dynamic changes of the environment. It is typically represented as a continuous sequence of two-dimensional images or a group of multiple frames. Optical sensors are used to acquire visual information, and their implementation includes vision systems based on RGB cameras, depth cameras, structured light sensors, or multimodal fusion. The system acquires visual information within the operating space by real-time acquisition of scene image sequences. The image sequences must have sufficient spatial resolution and temporal frame rate to ensure comprehensive capture of object distribution, dynamic changes, and spatial structure information in the environment.
[0126] Task description language information refers to the task requirements issued by the operator based on natural language input, voice commands, or text instructions, expressing the operational goals, task content, or behavioral norms. Through the speech recognition module, the system converts speech signals into text-formatted task instructions; through the text input interface, the system directly receives text data based on keyboard, touch, or other input devices. As a crucial source for the system to understand user intent and operational goals, the accuracy and completeness of task description language information directly impact the subsequent task planning and execution effectiveness.
[0127] Scene image sequences undergo object detection processing to extract object spatial coordinates and category labels. Object spatial coordinates refer to the position and orientation of an object in three-dimensional space, while category labels identify object attributes, functions, or categories. Object detection employs an image recognition algorithm based on convolutional neural networks, accurately identifying various objects within the scene through multi-scale feature extraction, bounding box regression, and category classification. This generates corresponding spatial coordinate data and category label sets, ensuring detailed and structured object information input for the environment modeling process.
[0128] The system further identifies dynamic obstacle outlines and static obstacle boundaries in the scene image sequence. Dynamic obstacle outlines refer to the external boundary information of moving objects whose position or shape changes over time, commonly found in mobile devices, people, or non-fixed structures. Static obstacle boundaries refer to the spatial boundaries of relatively fixed, immovable objects in the environment, such as walls, shelves, or equipment infrastructure. The distinction between dynamic and static obstacles is achieved through inter-frame image differencing, optical flow estimation, or multi-frame temporal feature analysis. The system extracts the spatial distribution and boundary outlines of obstacles, ensuring the accuracy and real-time performance of environmental modeling.
[0129] A 3D environment grid map is a spatial structure representation built upon object spatial coordinates, dynamic obstacle outlines, and static obstacle boundary information. This map describes the 3D layout relationships within the operational space using a regular grid or hierarchical data structure, reflecting the positions of objects, obstacle distribution, and drivable area divisions. The generation of the environment grid map employs methods such as point cloud reconstruction, voxel partitioning, or occupancy grid modeling to ensure high accuracy, real-time performance, and dynamic updating capabilities in spatial modeling.
[0130] The 3D environment grid map is labeled as the environment model. The environment model serves as the basic data input for subsequent task planning, path generation, and control execution. Based on the spatial structure information and dynamic change data provided by the environment model, the system supports multimodal information fusion and operation strategy generation, ensuring that robotic arms and other operating devices can safely and accurately perform tasks in complex dynamic environments.
[0131] This embodiment constructs an environment model based on visual information acquired by optical sensors and task description language information, achieving comprehensive fusion and structured expression of multi-source information within the operational space, thus improving the completeness and accuracy of environmental perception. The system combines object detection and obstacle recognition technologies to dynamically distinguish between static structures and dynamic interference factors in the environment, optimizing the expression accuracy and update efficiency of the 3D environment grid map. This enhances the system's real-time modeling capability for complex operational environments, solving the problems of incomplete information, delayed response to dynamic changes, or inaccurate spatial structure representation in traditional environmental modeling. Ultimately, it improves the environmental understanding level of the multimodal fusion system and the stability and reliability of subsequent task execution.
[0132] In one embodiment, step S20 above includes:
[0133] S201, Perform multi-layer convolution operation on the visual information to obtain a spatial feature map;
[0134] S202, perform a pooling operation on the spatial feature map to compress the feature dimension and obtain visual features;
[0135] S203, perform word embedding processing on the task description language information to generate a word vector sequence;
[0136] S204, Input the word vector sequence into a recurrent neural network to extract the temporal features of the word vector sequence;
[0137] S205, aggregate the temporal features to generate a language feature vector.
[0138] In this embodiment, the visual information is scene image sequence data acquired by the aforementioned optical sensor, including environmental structure, object information, and dynamic changes within the operating space, possessing multi-dimensional spatial attributes and rich visual elements such as texture, edge, and color. The visual information undergoes multi-layer convolution operations, specifically utilizing multiple consecutive convolutional layers in a convolutional neural network to perform local feature extraction and spatial pattern capture. Each convolutional operation slides a convolutional kernel across the two-dimensional or three-dimensional image data, extracting structural features such as edges, textures, and shapes at different levels. As the layers deepen, the extracted spatial structural information becomes increasingly rich, forming a spatial feature map reflecting both global and local features of the scene.
[0139] Spatial feature maps undergo pooling operations to compress feature dimensionality. Pooling involves dividing the spatial feature map into fixed-size regions and extracting statistical feature values within these regions using max pooling, average pooling, or other dimensionality reduction methods. This reduces data dimensionality, minimizes redundant information, and retains key information, improving the compactness and robustness of feature representation. Pooling generates high-dimensional visual features that include the position, shape, and distribution relationships of objects in the operating environment, serving as the foundational input for subsequent multimodal fusion and task planning.
[0140] Task description language information consists of natural language instructions, speech recognition text, or manually input text, and includes task objectives, operational behaviors, or environmental constraints. Word embedding processing of task description language information refers to transforming words, phrases, or symbols in the language information into dense vector representations in a low-dimensional, continuous space using embedding mapping methods. Word embedding processing captures the semantic relationships, contextual associations, and latent features of language units through a training corpus, generating word vector sequences that preserve the semantic structure and sequential features of the language information, reduce data dimensionality, and facilitate subsequent processing by deep learning models.
[0141] Word vector sequences are input into a recurrent neural network (RNN) for temporal feature extraction. The RNN uses a recursive structure to progressively pass state information over time, capturing contextual and long-term dependencies in the language sequence. The RNN can employ a standard RNN structure or incorporate structures such as Long Short-Term Memory (LSTM) networks and Gated Recurrent Units (GRUs) to enhance its ability to model complex language structures and long-distance dependencies, thereby improving the accuracy and stability of semantic information extraction.
[0142] Finally, the system aggregates the temporal features, integrates the sequence state information output by the recurrent neural network, and generates language feature vectors. The language feature vectors express the overall semantic structure, task intent, and logical relationship of the language information with fixed-dimensional numerical values. As the basic input for multimodal information fusion and task planning, the system ensures that visual and linguistic information have a unified and fusionable expression in the feature space, thereby improving the system's multimodal understanding and task reasoning capabilities.
[0143] This embodiment encodes visual information and task description language information into compact, high-dimensional, and structured visual and linguistic features, respectively. This enables the system to effectively represent and compress information across different modalities in the feature space, enhancing the structural integrity and semantic clarity of the information. Visual information undergoes multi-layer convolution and pooling operations, preserving key visual features such as scene structure, object positions, and spatial relationships, reducing data dimensionality, and optimizing subsequent computational efficiency. Linguistic information is extracted through word embedding and recurrent neural networks, capturing contextual dependencies and semantic relationships within the language sequence. This forms a semantically complete linguistic feature expression, overcoming the limitations of traditional language processing methods in maintaining semantic consistency and contextual integrity. It improves the system's ability to understand task requirements and operational objectives, ensuring the accuracy and stability of the multimodal information fusion process.
[0144] In one embodiment, step S30 above includes:
[0145] S301, Perform cross-modal fusion processing on the visual features and the language features through a language model to generate fused features;
[0146] S302, Based on the fused features, semantic reasoning is performed through the language model to generate a task logic dependency graph;
[0147] S303, the task logic dependency graph is decomposed into a sequence of atomic operation steps using the language model;
[0148] S304, the atomic operation step sequence is transformed into a target action sequence through the language model based on the kinematic constraints of the robotic arm.
[0149] In this embodiment, visual features are spatial structure information extracted through multi-layer convolution and pooling operations in the preceding process, capable of reflecting the position, shape, distribution, and overall spatial layout of objects in the environment. Linguistic features are semantic expression vectors formed through word embedding and recurrent neural network processing, embodying task instructions, operational goals, and logical relationships. These two types of features originate from visual information and task description language information, respectively, and differ in data type, dimensional structure, and semantic expression. To achieve information integration and joint understanding, the system utilizes a language model to perform cross-modal fusion processing on visual and linguistic features.
[0150] Cross-modal fusion processing refers to mapping data from different modalities to a unified feature space through the fusion mechanism within a language model. This preserves the informational advantages of each modality while capturing the semantic connections and logical mappings between visual and linguistic data. Fusion methods can employ multilayer perceptron structures, interactive attention mechanisms, or Transformer structures. By calculating the joint representation of visual and linguistic features, fused features are generated. These fused features possess both spatial perception and semantic understanding capabilities, reflecting the overall degree of matching between environmental information and task requirements.
[0151] Based on the fused features, the language model further performs semantic reasoning. This semantic reasoning process utilizes multi-layered nonlinear transformations, attention mechanisms, and relational graph learning structures within deep neural networks to analyze the logical relationships and operational requirements within the fused features, generating a task logic dependency graph. This graph expresses the sequence, dependencies, and parallel structures between operational steps in a graph structure. Nodes represent specific operational units, and edges represent the sequential relationships and constraints between operations. The task logic dependency graph reveals the overall structural logic of the task, helping the system understand the multi-step collaborative relationships in complex tasks.
[0152] Based on the generated task logic dependency graph and relying on the language model structure, the system decomposes the overall task into a more basic and executable sequence of atomic operation steps. Atomic operation steps refer to basic operation units that cannot be further decomposed, typically including standard operation actions such as grasping, moving, placing, rotating, and adjusting posture. The sequence of steps clearly defines the order and dependency logic of the operations, ensuring that the operation plan is executable and coherent.
[0153] Based on this, the language model transforms the sequence of atomic operation steps into a target action sequence according to the kinematic constraints of the robotic arm. These kinematic constraints encompass the physical conditions of the robotic arm's joint structure, workspace, degree-of-freedom restrictions, maximum speed, and load limits. The transformation process utilizes inverse kinematics calculations, path planning algorithms, and motion constraint parsing to ensure that the target action sequence conforms to the robotic arm's actual execution capabilities and environmental constraints. The target action sequence is output as a set of explicit action instructions, including the predetermined motion trajectories of each joint of the robotic arm, the processing method for the manipulated object, and the order of operations. This serves as the input for subsequent path planning and action execution, ensuring the accuracy and feasibility of the task planning results.
[0154] This embodiment achieves joint expression and deep understanding of different information sources by fusing visual and linguistic features across modalities. This overcomes the problems of incomplete expression and weak semantic connections inherent in single-modal information, enhancing the system's adaptability to complex environments and diverse task requirements. The fused features are used to generate a task logic dependency graph through semantic reasoning. The system obtains a structured logical expression between operation steps, overcoming the limitations of traditional task planning methods that lack hierarchical structure and dependency descriptions for complex tasks. By combining the decomposition of the task logic dependency graph with the kinematic constraints of the robotic arm, the system generates a target action sequence that possesses physical feasibility and operational coherence. This ensures that the robotic arm can accurately and efficiently execute predetermined tasks in complex environments, improving the accuracy, environmental adaptability, and execution stability of task planning.
[0155] In one embodiment, step S40 above includes:
[0156] S401, Define the environmental state space and the robotic arm motion space based on the environmental model;
[0157] S402, Based on the environmental state space and the robotic arm motion space, the optimal value function for each environmental state is determined through a dynamic programming module;
[0158] S403, extract the target state and initial state from the environment model;
[0159] S404, Perform reverse value iteration from the target state to the initial state;
[0160] S405, In each iteration, select the predecessor state corresponding to the action that maximizes the optimal value function, and add the predecessor state to the path node sequence;
[0161] S406, Stop iteration when the predecessor state is the initial state;
[0162] S407, Generate a basic path based on the path node sequence.
[0163] In this embodiment, the environment model is a three-dimensional environment grid map constructed using visual information in the previous steps. It includes the spatial coordinates of objects, the outlines of dynamic obstacles, and the boundaries of static obstacles, possessing complete spatial structure information and obstacle distribution information. Based on the environment model, the system first defines the environment state space, which is the set of all possible states that may exist in the environment model. A state typically represents the specific values of parameters such as the position, posture, and joint angle combination of the robotic arm or the manipulated object in three-dimensional space, reflecting the overall state information of the robotic arm in different spatial positions.
[0164] Meanwhile, the system defines the robotic arm's motion space, which is the set of all basic actions that the robotic arm can perform in the environmental state space. These actions include joint angle changes, end effector position adjustments, path node switching, and other operations. The motion space is limited by the robotic arm's structural parameters, degrees of freedom, load capacity, and spatial obstacles to ensure that the generated path is physically feasible.
[0165] Based on the aforementioned state space and action space, the system calculates the optimal value function through a dynamic programming module. The dynamic programming module iteratively updates the value function using a recursive structure and based on the Bellman optimality principle to determine the cumulative value obtainable in each environmental state when taking the optimal action strategy. The value function measures the comprehensive cost of reaching the target state from the current state, typically considering factors such as path length, energy consumption, obstacle avoidance cost, and task time. A larger value function indicates a better path. The dynamic programming module obtains the globally optimal path information by performing value function calculations on all environmental states.
[0166] The system extracts the target state and initial state from the environment model. The target state is the final position or posture that the robotic arm is expected to reach, and the initial state is the current spatial position and posture of the robotic arm. Both are derived from the specific state parameters in the environment model, ensuring that the path planning process is based on real spatial information.
[0167] During path planning, the system starts from the target state and executes a backward value iteration process. This iteration utilizes a dynamic programming-based backward reasoning strategy to gradually backtrack from the target state to the initial state, avoiding getting trapped in local optima and ensuring global path optimality. In each iteration, the system selects the predecessor state corresponding to the action that maximizes the optimal value function. The predecessor state represents the previous environmental state that can be reached from the current state by performing a specific action. Selecting the optimal predecessor state ensures that the path backtracks along the optimal value trajectory.
[0168] The system adds the selected predecessor state to the path node sequence in sequence. The path node sequence is a set of environmental states arranged in reverse order, representing the complete path structure from the target state to the initial state. When the predecessor state backtracks to the initial state, the iteration process terminates, indicating that the path backtracking is complete. The system generates a basic path based on the path node sequence. The basic path is a set of complete action paths from the initial state along the optimal trajectory to the target state. The basic path takes into account the environmental structure, obstacle distribution, and robotic arm action constraints, and has high efficiency, executability, and obstacle avoidance capabilities.
[0169] This embodiment defines the environmental state space and the robotic arm's motion space based on an environmental model. The system comprehensively describes the robotic arm's motion capabilities and environmental constraints in space, avoiding state omissions and unreasonable actions during path planning. Combined with the optimal value function calculated by the dynamic programming module, the system optimizes the path structure globally, overcoming the shortcomings of traditional path planning methods that rely on local information, leading to low path quality. Through backward value iteration and predecessor state selection, the path generation process possesses global optimality, avoiding paths getting trapped in suboptimal solutions or local loops. The final generated basic path strictly adheres to environmental constraints and the robotic arm's motion capabilities, ensuring the path's efficiency, consistency, and obstacle avoidance capabilities, thus improving the robotic arm's task execution efficiency and safety in complex environments.
[0170] In one embodiment, step S50 above includes:
[0171] S501 monitors changes in environmental conditions and generates environmental change information;
[0172] S501, Based on the environmental change information and the basic path, construct a reinforcement learning state representation;
[0173] S501, input the reinforcement learning state representation into the policy network and output a path adjustment action;
[0174] S501, Execute the path adjustment action to obtain the adjusted path;
[0175] S501, Generate a reward signal based on the environmental interaction result, and update the parameters of the policy network based on the reward signal;
[0176] S501, determine whether the adjusted path simultaneously satisfies the conditions for shortening the path length and improving the obstacle avoidance success rate.
[0177] S501, when the path length reduction condition and obstacle avoidance success rate improvement condition are met, the adjusted path is taken as the final path.
[0178] In this embodiment, the system first monitors changes in the environmental state. These changes refer to the real-time alterations in environmental structure, obstacle distribution, and target object position caused by dynamic external environmental factors or task execution feedback during the robotic arm's operation. Environmental state changes are acquired through a sensor system, including image information captured by a vision sensor, spatial point cloud data acquired by a depth sensor, and spatial coordinate information fed back by a position sensor. Through the fusion and real-time analysis of this multi-source data, the system generates environmental change information. This information is a structured set of environmental state data, specifically including parameters such as obstacle position updates, target position offsets, and changes in spatial accessibility, accurately reflecting the dynamic characteristics of the current environment.
[0179] Based on environmental change information and the basic path, the system constructs a reinforcement learning state representation. The state representation comprehensively includes environmental change information, basic path structure, and the current position and posture information of the robotic arm. Through the state encoding module, the high-dimensional environmental information and path information are transformed into low-dimensional numerical feature vectors suitable for reinforcement learning models. The state representation ensures that the reinforcement learning model can perceive the dynamic features of the environment and changes in the path structure.
[0180] The system inputs the reinforcement learning state representation into the policy network. The policy network is a parameterized deep neural network structure with the ability to map the environment state to path adjustment actions end-to-end. The policy network structure may include multilayer perceptrons, convolutional networks, or recurrent networks. The policy network outputs path adjustment actions through forward inference. The path adjustment actions are fine-tuning operations on the current basic path structure and environment state, specifically including path node position adjustment, path smoothing, obstacle avoidance action insertion, etc. The path adjustment actions improve path adaptability and safety.
[0181] The system performs path adjustment actions to obtain the adjusted path. While maintaining the global structural characteristics of the original basic path, the adjusted path is dynamically modified in response to changes in the local environment, thereby improving the obstacle avoidance capability and execution efficiency of the path.
[0182] The system generates reward signals based on the environmental interaction results. The environmental interaction results refer to the path performance indicators collected by the system in real time during the execution of the robotic arm along the adjusted path, including parameters such as path execution time, path length, energy consumption, and obstacle avoidance attempts. The reward signals are generated based on a preset optimization objective function. Positive rewards encourage shorter path lengths and improved obstacle avoidance success rates, while negative rewards penalize longer path lengths or obstacle avoidance failures. The reward signals serve as the basis for the reinforcement learning model to optimize the network parameters of the strategy.
[0183] The system updates the policy network parameters based on reward signals. The parameter update process uses gradient descent or proximal policy optimization methods to improve the policy network's path adjustment capability and convergence speed in dynamic environments.
[0184] The system determines whether the adjusted path meets the conditions for shortening the path length and improving the obstacle avoidance success rate. The condition for shortening the path length means that the total length of the adjusted path is reduced by more than a preset threshold compared to the basic path. The condition for improving the obstacle avoidance success rate means that the success rate of the adjusted path in the obstacle avoidance test is higher than that of the basic path. If both conditions are met, it means that the path optimization effect has reached the expected level.
[0185] When both conditions are met, the system adopts the adjusted path as the final path, which possesses global optimality, local adaptability, and high obstacle avoidance efficiency. If either condition is not met, the system returns to the step of monitoring environmental state changes, continues iterative optimization, and forms a closed-loop path adjustment mechanism to ensure that the path always adapts to the dynamic environment, thereby improving the efficiency and safety of the robotic arm's task execution.
[0186] This embodiment utilizes reinforcement learning to dynamically adjust the basic path, overcoming the poor adaptability of traditional path planning methods in changing environments. It constructs a state representation based on environmental state changes and the basic path structure, ensuring the policy network perceives environmental dynamics and path structure, thus improving the targeted nature of path adjustments. By outputting path adjustment actions through the policy network, local adaptive correction of the path is achieved, preventing the path from getting trapped in local optima or suboptimal solutions. The introduction of environmental interaction reward signals and a policy network parameter update mechanism forms a closed loop of path optimization based on performance feedback, ensuring shorter path lengths and improved obstacle avoidance success rates. The system possesses real-time dynamic adjustment and self-optimization capabilities, enhancing the path planning efficiency and task completion quality of the robotic arm in complex dynamic environments.
[0187] In one embodiment, step S60 above includes:
[0188] S601, Obtain the current joint status of the robotic arm;
[0189] S602, The target motion sequence and the current robotic arm joint state are processed by the policy network in the motion decoder to generate joint control parameters;
[0190] S603, convert the joint control parameters into joint control commands;
[0191] S604, parse the sequence of key points of the trajectory in the final path;
[0192] S605, perform spline interpolation processing on the trajectory key point sequence to generate a Cartesian space trajectory instruction, and use the Cartesian space trajectory instruction as a motion control signal;
[0193] S606, convert the motion control signal into a joint spatial position command through inverse kinematics solution;
[0194] S607, The joint spatial position command and the joint control command are input to the motor controller, and the motor controller generates a pulse width modulation motor drive signal;
[0195] S608, the motor driving the robotic arm executes the pulse width modulation motor drive signal.
[0196] In this embodiment, the system first acquires the current joint state of the robotic arm. The joint state information includes the angular position, angular velocity, joint load, and actual posture data fed back by sensors during execution for each joint. The joint state is acquired in real time through a high-precision angle encoder, inertial measurement unit, and other posture sensors, ensuring that the subsequent motion decoding process can be combined with the actual motion state of the robotic arm, thereby improving the accuracy and responsiveness of control commands.
[0197] The system uses a policy network in the motion decoder to jointly process the target motion sequence and the current joint state of the robotic arm. The target motion sequence is a structured set of motion instructions generated from visual and linguistic information in the preceding steps, containing the parameters and execution order of each sub-motion after task decomposition. The policy network is a deep neural network structure with nonlinear mapping capabilities, employing a multilayer perceptron or temporal recurrent network. Combining the target motion sequence and the features of the robotic arm joint state, it outputs joint control parameters. These parameters include the desired pose change, velocity planning, and torque distribution information for each joint, ensuring that the generated control commands conform to the current physical state of the robotic arm and the task requirements.
[0198] The system converts joint control parameters into joint control commands. The joint control commands are standardized joint drive signal input formats, specifically including joint angle increments, target position commands, or target speed commands. They conform to the control system communication protocol, ensuring that the commands can be directly parsed and executed by the lower-level control system.
[0199] The system analyzes the sequence of key points in the final path. These key points are a pre-set set of spatial coordinates generated during path generation, describing the spatial positions and attitudes that the robotic arm's end effector needs to traverse sequentially. The system performs spline interpolation on the sequence of key points, using cubic B-splines or quintic polynomials to generate a smooth and continuous spatial trajectory between the key points, thus avoiding abrupt changes or discontinuities in the robotic arm's trajectory during path execution.
[0200] The Cartesian space trajectory command generated by the interpolation process describes the spatial position change trajectory and time sequence information of the robotic arm's end effector. As a motion control signal, the trajectory command guides the control system to perform spatial path tracking tasks, ensuring the continuity and accuracy of path execution.
[0201] The system converts motion control signals into joint spatial position commands through inverse kinematics. Based on the robotic arm's structural parameters, joint topology, and kinematic equations, inverse kinematics calculates the changes in joint position corresponding to a given Cartesian trajectory. The joint spatial position commands include the target position coordinates or angle information of each joint, ensuring that the robotic arm's end effector completes the spatial path tracking task according to the Cartesian trajectory.
[0202] The system inputs joint spatial position commands and joint control commands into the motor controller. The motor controller is the underlying drive module at the execution level. Based on the input commands, it generates pulse width modulation motor drive signals. The pulse width modulation signals control the motor's on / off duty cycle to achieve precise control of the motor speed and output torque, ensuring that each joint moves synchronously and in a coordinated manner according to the planned position, speed and dynamic response requirements.
[0203] The system ultimately drives the robotic arm's motor to execute pulse-width modulated motor drive signals. The robotic arm completes precise operations according to the planned path and task action sequence, ensuring that the end effector can efficiently and smoothly execute the predetermined task objectives in complex environments.
[0204] Example Description: In an embodied intelligent dual-arm collaborative operation scenario, the robotic arm system is deployed in a complex and dynamic industrial production environment. The objective is to precisely grasp and transport irregular items on shelves using the coordinated action of both arms. The system first acquires visual information and task description language information through multimodal sensing devices deployed around the robotic arm or its working area. The visual information comes from real-time scene image sequences captured by high-resolution stereo vision cameras, depth cameras, or multi-angle optical sensors. The scene includes static shelves, dynamic handling equipment, and various types of items to be grasped. The task description language information is obtained through voice input devices used by on-site personnel or through text task instructions issued by the production scheduling system. An example of a task instruction is "Transport the blue boxes on the third shelf to the designated workstation."
[0205] After acquiring data, the system generates an environment model based on visual information. The specific process includes: using a deep learning object detection network to identify various objects in the scene image sequence, obtaining the spatial coordinates and category label of each object; distinguishing between dynamic and static obstacles through dynamic contour detection and background modeling methods, extracting the contours of dynamic obstacles and the boundary information of static obstacles; and based on the above information, constructing a high-precision 3D environment grid map. The environment model comprehensively describes the distribution of objects, obstacle structures, and dynamic changes within the workspace, ensuring the system possesses complete environmental awareness capabilities.
[0206] Subsequently, the system encodes the acquired visual information and task description language information separately. The visual information is processed through a multi-layer convolutional neural network to extract multi-scale spatial features, which are then compressed using feature pooling to form a stable and compact visual feature representation. The task description language information is first processed through word embedding to generate a word vector sequence. This sequence is then further input into a bidirectional recurrent neural network to extract temporal language features, ultimately aggregating to generate a language feature vector. These visual and language features provide the foundation for cross-modal information fusion in subsequent task planning.
[0207] Based on the aforementioned multimodal features, the system utilizes an integrated language model for task planning. The language model first fuses visual and linguistic features through a cross-modal feature fusion mechanism to form a unified high-level abstract feature expression. Building upon the fused features, the system performs semantic reasoning, constructing a task logic dependency graph specific to the current operation task. This graph clearly defines the temporal and causal relationships of each operation step. The language model further decomposes the task logic dependency graph into a refined sequence of atomic operation steps, ensuring that complex tasks are broken down into clear and controllable basic action units. Combining the robotic arm's structure, kinematic constraints, and environmental model information, the system ultimately transforms the atomic operation step sequence into a target action sequence. This target action sequence fully describes the spatial position adjustments, dual-arm collaborative movements, and interaction logic involved in this grasping and handling task.
[0208] Based on the established environment model, the system generates a basic path for the robotic arm's motion using dynamic programming. First, the system defines the environment state space and the robotic arm motion space according to the environment model. The state space comprehensively covers all possible positions and postures of the robotic arm's end effector within the working area, while the motion space encompasses the set of movements of each joint of the robotic arm within the physically and structurally permissible range. The system calculates the optimal value function corresponding to each state in the environment state space using a dynamic programming module. The value function evaluates the optimal path cost from each state to the target state. The system extracts the current initial state and the task target state of the robotic arm from the environment model. Starting from the target state, the system performs backward value iteration. During the iteration, the optimal predecessor state is selected according to the value function and added to the path node sequence until the predecessor state backtracks to the initial state. The system generates a basic path based on the path node sequence, which serves as the initially planned motion route for the robotic arm's end effector.
[0209] During actual execution, the environmental state may change due to dynamic obstacles, object displacement, or operational errors. The system dynamically adjusts the basic path using reinforcement learning to generate the final path. The system continuously monitors environmental state changes, generates environmental change information, and combines this with the basic path information to construct a reinforcement learning state representation that comprehensively reflects environmental changes and path structure characteristics. The system inputs this state representation into the reinforcement learning policy network and outputs path adjustment actions. These actions optimize path nodes or local trajectories in real time in response to environmental changes. The system executes the path adjustment actions, obtains the adjusted path, and generates a reward signal in real time based on the environmental interaction results. The reward signal is dynamically calculated based on indicators such as path length reduction and obstacle avoidance success rate. The system updates the policy network parameters using the reward signal, improving the policy network's path adjustment capability in dynamic environments. The system iterates through this process until the adjusted path simultaneously meets preset threshold standards in terms of path length reduction and obstacle avoidance success rate improvement. Finally, the optimized path is output as the final path.
[0210] After acquiring the target motion sequence and final path, the system performs motion decoding and motion control signal generation. The system obtains the current joint state parameters of the robotic arm, inputs them into the policy network of the motion decoder along with the target motion sequence, and outputs joint control parameters for the current joint state. These joint control parameters are standardized and converted into joint control commands. The system further analyzes the trajectory key point sequence of the final path, smooths the trajectory using spline interpolation, and generates continuous Cartesian space trajectory commands. These trajectory commands are input as motion control signals to the control system. Based on the inverse kinematics model of the robotic arm, the system converts the motion control signals into spatial position commands for each joint. These joint control commands are then input into the motor controller to generate pulse-width modulated motor drive signals. Ultimately, these signals drive the motors of each joint of the robotic arm to coordinate and execute the task path, ensuring that the end effector accurately tracks the planned trajectory and completes the specified operation.
[0211] During the execution of joint control commands, the system acquires environmental feedback information in real time. Based on the feedback rewards, the system continuously optimizes the joint control commands using reinforcement learning methods. Specifically, this process includes collecting the actual pose data of the end effector, calculating the error between the actual and expected poses, and dynamically generating feedback reward signals based on the error values. The system uses the feedback rewards to update the policy network parameters, optimizes the joint control compensation commands, and superimposes the compensation commands onto the original joint control commands. This effectively improves the dynamic adaptability and error correction capability of the joint control commands, ensuring path tracking accuracy and overall motion stability during robotic arm operation, and enhancing the system's operational reliability and task completion efficiency in complex dynamic environments.
[0212] In the healthcare sector, embodied intelligent robotic arms are used to assist medical staff in performing tasks such as handling and sorting medical consumables. The system first utilizes vision sensors installed in operating rooms or drug storage areas to acquire visual information and task description language information. Visual information is obtained by collecting real-time image sequences of medical workbenches, drug cabinets, and consumable storage areas. Task description language information is simultaneously transmitted through voice input commands from medical staff or text scheduling commands from the intelligent management system. Task commands include scenario requirements such as "moving a specified batch of drugs to a designated ward" and "grabbing disposable medical devices and placing them on a delivery vehicle."
[0213] The system utilizes collected visual information and a medical supply identification algorithm to accurately detect the location, category, and expiration date labels of medical items. A dynamic obstacle detection function monitors the movement trajectories of medical personnel and transport equipment in real time, while also identifying the structural information of static obstacles such as hospital beds, fixed equipment, and medicine cabinet boundaries. Based on these identification results, the system generates a 3D environmental grid map. The map annotations include the distribution of medical supply categories, path obstacles, and dynamic safety zones, providing complete environmental information for the robotic arm's motion planning.
[0214] Visual information is used to extract medical scene features through a multi-layer convolutional neural network. Task description language information is processed through word embedding and recurrent neural network temporal processing to extract key operational vocabulary and operational logic from medical instructions. The system inputs visual and linguistic features into a fusion model to associate information and construct a task logic dependency graph based on medical task requirements. The task logic dependency graph is further decomposed into atomic operation steps such as grasping, carrying, and placing, and combined with the robotic arm motion constraints to generate target action sequences adapted to medical scenarios.
[0215] The system uses dynamic programming based on an environmental model to generate basic paths in medical scenarios. Path planning fully considers the layout of medicine cabinets, the width of passageways, and the dynamic range of personnel movement. During actual handling, the basic paths are adjusted in real time according to environmental changes. Reinforcement learning dynamically optimizes path nodes to cope with complex situations such as temporary path occupancy by medical staff and movement of mobile equipment. Feedback and rewards are used to continuously adjust the path structure. The system executes the adjusted paths to ensure that medicines or medical devices are accurately transported to their target locations within a limited time, while simultaneously optimizing path length and obstacle avoidance success rate.
[0216] During the generation and execution of joint control commands, the system analyzes the current joint state of the robotic arm in real time. The target motion sequence is dynamically adapted to the current posture of the robotic arm by the motion decoder. The system analyzes the key points of the path and generates continuous motion control signals. The system adjusts the joint spatial position commands in real time based on inverse kinematics. The joint control commands and motion control signals are synchronously input into the motor controller to generate pulse width modulation motor drive signals, which drive the robotic arm to accurately perform drug grasping, handling and placement operations.
[0217] During execution, the system optimizes joint control commands in real time through feedback rewards. By collecting the actual pose of the robotic arm end effector, calculating the deviation from the expected pose, and dynamically generating compensation commands to superimpose and correct the current control commands, the system improves the operational stability and response sensitivity of the robotic arm in the medical environment. This ensures that there is no misplacement or deviation during the handling of medicines or medical devices, and achieves efficient, accurate, and stable automatic handling assistance for medical supplies.
[0218] In the fintech sector, embodied intelligent robotic arms are widely used in scenarios such as self-service equipment maintenance, intelligent document management, and high-value document handling. The system acquires visual information through optical sensors deployed within intelligent vaults, document processing centers, or high-end self-service equipment. This visual information is combined with task descriptions input via voice recognition modules or dedicated maintenance terminals to form the data foundation for operation execution. Visual information includes image data such as vault shelf structure, document placement status, and document box barcode labels. Task descriptions include text or voice commands issued by maintenance personnel, such as "Move Class A documents to the designated storage area" or "Grab document B and deliver it to the risk control review area."
[0219] Based on the acquired visual information, the system performs high-precision detection of tickets, vouchers, and archives, extracts the spatial coordinate information and category labels of the items, and the dynamic obstacle detection module identifies the movement trajectories of operators, maintenance robots, and equipment pallets in real time. Static obstacle boundary information accurately marks vault shelves, aisles, and protective structures. Combining the above information, a three-dimensional environmental grid map is constructed to reflect the spatial layout and dynamic changes within the financial operation area, ensuring that the robotic arm has complete environmental perception capabilities.
[0220] Image data is processed through a multi-layer convolutional network to extract spatial features relevant to financial scenarios. Linguistic information is extracted using word embeddings and recurrent neural networks to extract task logic keywords and temporal features. The system inputs visual and linguistic features into a multimodal fusion module to generate a fused feature representation that reflects the needs of financial operations. Based on these fused features, the system uses a deep language model to infer the task's logical dependency graph, automatically breaking it down into basic operational steps such as picking, moving, and placing. Combined with the kinematic constraints of the robotic arm, it generates a target action sequence suitable for financial business processes.
[0221] For the environmental model, the system uses dynamic programming algorithms to formulate the basic path of the robotic arm. The path design fully considers the arrangement of vault shelves, personnel operation areas, and equipment safety boundaries. During actual execution, the system dynamically adjusts the path structure according to environmental changes through reinforcement learning. The reinforcement learning model optimizes path nodes based on a feedback reward mechanism to adapt to sudden personnel intervention, passage congestion, or temporary adjustments to high-value documents, ensuring that the path has optimal length and maximum obstacle avoidance efficiency.
[0222] In the robotic arm control phase, the system collects joint state information of the robotic arm in real time. Combined with the target motion sequence, it adaptively generates joint control parameters that match the current state through a motion decoder, and transforms them into specific joint control commands. The system further analyzes the trajectory key points in the final path, performs spline interpolation to generate smooth and continuous Cartesian space trajectory commands, and calculates joint spatial position commands in real time using inverse kinematics. The joint control commands and motion control signals are input to the motor controller, ultimately generating pulse width modulation motor drive signals to drive the robotic arm to complete the voucher grasping, handling, and delivery operations with high precision.
[0223] During the execution of joint control commands, the system monitors the actual pose data of the robotic arm's end effector in real time, calculates the error between the actual pose and the expected position, generates feedback rewards based on the error, and the reinforcement learning model generates optimized joint control compensation commands by updating the policy network parameters. The compensation commands are superimposed on the original control commands, effectively improving the stability and response speed of the robotic arm in financial business operations, ensuring the safe, stable, and accurate handling of high-value documents, important files, and other items, and meeting the high standards of intelligent operation and maintenance and automated management in the financial technology field.
[0224] This embodiment inputs the target motion sequence and the current joint state of the robotic arm into the strategy network. The system can dynamically adjust control parameters according to the real-time state of the robotic arm, improving the matching degree between control commands and the actual physical state. By parsing the key points of the final path trajectory and performing spline interpolation, the system generates a continuous and smooth spatial path, avoiding trajectory abrupt changes and instability during the robotic arm's movement. Combined with the inverse kinematics process, the system effectively transforms the high-level trajectory planning results into the low-level control parameters of each joint, ensuring the accuracy of path execution and high precision of end-effector position control. Based on pulse width modulation technology for motor control signal generation and execution, the system improves the response speed and torque control accuracy during motor drive, ensuring good motion stability and operational reliability of the robotic arm in dynamic environments. The overall structure effectively connects task planning, path tracking, and dynamic control, improving the accuracy, stability, and adaptability of the robotic arm operation to complex tasks.
[0225] In one embodiment, a robotic arm task planning and control device is provided, which corresponds one-to-one with the robotic arm task planning and control method described in the above embodiments. (Refer to...) Figure 3 , Figure 3 This is a schematic diagram of the functional modules of a preferred embodiment of the robotic arm task planning and control device of the present invention. The modules include: environmental perception module 10, information encoding module 20, task planning module 30, path generation module 40, path optimization module 50, control command generation module 60, and command optimization module 70. Detailed descriptions of each functional module are as follows:
[0226] The environment perception module 10 is used to acquire visual information and task description language information, and generate an environment model based on the visual information.
[0227] Information encoding module 20 is used to encode the visual information and the task description language information to obtain visual features and language features respectively;
[0228] Task planning module 30 is used to perform task planning based on the visual features and the language features using a language model, and generate a target action sequence.
[0229] The path generation module 40 is used to generate a basic path for the movement of the robotic arm through dynamic programming under the environment model.
[0230] The path optimization module 50 is used to adjust the basic path according to environmental changes through reinforcement learning to generate the final path;
[0231] The control command generation module 60 is used to decode the target motion sequence into joint control commands and generate motion control signals based on the final path;
[0232] The instruction optimization module 70 is used to optimize the joint control instruction through reinforcement learning based on the feedback reward obtained from the environment during the execution of the joint control instruction.
[0233] In one embodiment, the environment sensing module 10 is specifically used for:
[0234] Scene image sequences are acquired using optical sensors as visual information;
[0235] The task instructions are received as task description language information through a speech recognition module or text input interface.
[0236] Object detection is performed on the scene image sequence to obtain the object spatial coordinates and object category labels;
[0237] Identify the outlines of dynamic obstacles and the boundaries of static obstacles in the scene image sequence;
[0238] A three-dimensional environment grid map is constructed based on the object's spatial coordinates, dynamic obstacle outlines, and static obstacle boundaries.
[0239] The three-dimensional environment grid map is labeled as an environment model.
[0240] In one embodiment, the information encoding module 20 is specifically used for:
[0241] The visual information is subjected to multi-layer convolution operations to obtain a spatial feature map;
[0242] The spatial feature map is pooled to compress the feature dimensions, thus obtaining visual features;
[0243] The task description language information is subjected to word embedding processing to generate a word vector sequence;
[0244] The word vector sequence is input into a recurrent neural network to extract the temporal features of the word vector sequence;
[0245] The temporal features are aggregated to generate a language feature vector.
[0246] In one embodiment, the task planning module 30 is specifically used for:
[0247] The visual features and the linguistic features are fused across modally using a language model to generate fused features.
[0248] Based on the fused features, semantic reasoning is performed through the language model to generate a task logic dependency graph;
[0249] The language model is used to decompose the task logic dependency graph into a sequence of atomic operation steps.
[0250] The language model is used to transform the sequence of atomic operation steps into a sequence of target actions based on the kinematic constraints of the robotic arm.
[0251] In one embodiment, the path generation module 40 is specifically used for:
[0252] The environmental state space and the robotic arm motion space are defined based on the aforementioned environmental model;
[0253] Based on the environmental state space and the robotic arm motion space, the optimal value function for each environmental state is determined through a dynamic programming module.
[0254] Extract the target state and initial state from the environment model;
[0255] Starting from the target state, perform reverse value iterations towards the initial state;
[0256] In each iteration, the predecessor state corresponding to the action that maximizes the optimal value function is selected, and the predecessor state is added to the path node sequence.
[0257] The iteration stops when the predecessor state is the initial state.
[0258] A basic path is generated based on the path node sequence.
[0259] In one embodiment, the path optimization module 50 is specifically used for:
[0260] Monitor changes in environmental conditions and generate environmental change information;
[0261] Based on the environmental change information and the basic path, a reinforcement learning state representation is constructed;
[0262] The reinforcement learning state representation is input into the policy network, and the output path adjustment action is generated.
[0263] Perform the path adjustment action to obtain the adjusted path;
[0264] A reward signal is generated based on the environmental interaction results, and the parameters of the policy network are updated based on the reward signal.
[0265] Determine whether the adjusted path simultaneously satisfies the conditions for shortening the path length and improving the obstacle avoidance success rate.
[0266] When the conditions for shortening the path length and improving the obstacle avoidance success rate are met, the adjusted path will be used as the final path.
[0267] In one embodiment, the control instruction generation module 60 is specifically used for:
[0268] Get the current state of the robotic arm joints;
[0269] The target motion sequence and the current robotic arm joint state are processed by the policy network in the motion decoder to generate joint control parameters.
[0270] The joint control parameters are converted into joint control commands;
[0271] Analyze the sequence of key points in the final path;
[0272] The sequence of key points of the trajectory is subjected to spline interpolation to generate a Cartesian space trajectory instruction, and the Cartesian space trajectory instruction is used as a motion control signal.
[0273] The motion control signal is converted into joint spatial position commands through inverse kinematics;
[0274] The joint spatial position command and the joint control command are input into the motor controller, and the motor controller generates a pulse width modulation motor drive signal.
[0275] The motor driving the robotic arm executes the pulse-width modulated motor drive signal.
[0276] In one embodiment, a computer device is provided, which may be a server, and its internal structure diagram may be as follows: Figure 4 As shown, the computer device includes a processor, memory, network interface, and database connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile and / or volatile storage media and internal memory. The non-volatile storage media stores the operating system, computer programs, and database. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used for communication with external user terminals via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a robotic arm task planning and control method on the server side.
[0277] In one embodiment, a computer device is provided, which may be a user terminal, and its internal structure diagram may be as follows: Figure 5As shown, the computer device includes a processor, memory, network interface, display screen, and input devices connected via a system bus. The processor provides determination and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage media. The network interface is used to communicate with an external server via a network connection. When the computer program is executed by the processor, it implements the functions or steps of a robotic arm task planning and control method on the user side.
[0278] In one embodiment, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to perform the following steps:
[0279] Acquire visual information and task description language information, and generate an environment model based on the visual information;
[0280] The visual information and the task description language information are encoded to obtain visual features and language features, respectively.
[0281] Based on the visual features and the language features, a language model is used for task planning to generate a target action sequence.
[0282] Under the aforementioned environmental model, a basic path for the robotic arm's movement is generated through dynamic programming;
[0283] The basic path is adjusted according to changes in the environment through reinforcement learning to generate the final path;
[0284] The target motion sequence is decoded into joint control commands, and motion control signals are generated based on the final path;
[0285] During the execution of the joint control commands, the joint control commands are optimized through reinforcement learning based on feedback rewards obtained from the environment.
[0286] In one embodiment, a computer-readable storage medium is provided having a computer program stored thereon, the computer program performing the following steps when executed by a processor:
[0287] Acquire visual information and task description language information, and generate an environment model based on the visual information;
[0288] The visual information and the task description language information are encoded to obtain visual features and language features, respectively.
[0289] Based on the visual features and the language features, a language model is used for task planning to generate a target action sequence.
[0290] Under the aforementioned environmental model, a basic path for the robotic arm's movement is generated through dynamic programming;
[0291] The basic path is adjusted according to changes in the environment through reinforcement learning to generate the final path;
[0292] The target motion sequence is decoded into joint control commands, and motion control signals are generated based on the final path;
[0293] During the execution of the joint control commands, the joint control commands are optimized through reinforcement learning based on feedback rewards obtained from the environment.
[0294] It should be noted that the functions or steps that can be implemented by the computer-readable storage medium or computer device described above can be referred to the relevant descriptions on the server side and user side in the foregoing method embodiments. To avoid repetition, they will not be described one by one here.
[0295] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.
[0296] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.
[0297] It should be noted that if any software tools or components not belonging to this company appear in the embodiments of this application, they are merely illustrative examples and do not represent actual use. The embodiments described above are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.
Claims
1. A method for task planning and control of a robotic arm, characterized in that, Includes the following steps: Acquire visual information and task description language information, and generate an environment model based on the visual information; The visual information and the task description language information are encoded to obtain visual features and language features, respectively. Based on the visual features and the language features, a language model is used for task planning to generate a target action sequence. Under the aforementioned environmental model, a basic path for the robotic arm's movement is generated through dynamic programming; The basic path is adjusted according to changes in the environment through reinforcement learning to generate the final path; The target motion sequence is decoded into joint control commands, and motion control signals are generated based on the final path; During the execution of the joint control commands, the joint control commands are optimized through reinforcement learning based on feedback rewards obtained from the environment.
2. The robotic arm task planning and control method as described in claim 1, characterized in that, Acquiring visual information and task description language information, and generating an environment model based on the visual information, including: Scene image sequences are acquired using optical sensors as visual information; The task instructions are received as task description language information through a speech recognition module or text input interface. Object detection is performed on the scene image sequence to obtain the object spatial coordinates and object category labels; Identify the outlines of dynamic obstacles and the boundaries of static obstacles in the scene image sequence; A three-dimensional environment grid map is constructed based on the object's spatial coordinates, dynamic obstacle outlines, and static obstacle boundaries. The three-dimensional environment grid map is labeled as an environment model.
3. The robotic arm task planning and control method as described in claim 1, characterized in that, The visual information and the task description language information are encoded to obtain visual features and language features, respectively, including: The visual information is subjected to multi-layer convolution operations to obtain a spatial feature map; The spatial feature map is pooled to compress the feature dimensions, thus obtaining visual features; The task description language information is subjected to word embedding processing to generate a word vector sequence; The word vector sequence is input into a recurrent neural network to extract the temporal features of the word vector sequence; The temporal features are aggregated to generate a language feature vector.
4. The robotic arm task planning and control method as described in claim 1, characterized in that, Based on the visual features and the language features, a language model is used for task planning to generate a target action sequence, including: The visual features and the linguistic features are fused across modally using a language model to generate fused features. Based on the fused features, semantic reasoning is performed through the language model to generate a task logic dependency graph; The language model is used to decompose the task logic dependency graph into a sequence of atomic operation steps. The language model is used to transform the sequence of atomic operation steps into a sequence of target actions based on the kinematic constraints of the robotic arm.
5. The robotic arm task planning and control method as described in claim 1, characterized in that, Under the aforementioned environmental model, a basic path for the robotic arm's movement is generated through dynamic programming, including: The environmental state space and the robotic arm motion space are defined based on the aforementioned environmental model; Based on the environmental state space and the robotic arm motion space, the optimal value function for each environmental state is determined through a dynamic programming module. Extract the target state and initial state from the environment model; Starting from the target state, perform reverse value iterations towards the initial state; In each iteration, the predecessor state corresponding to the action that maximizes the optimal value function is selected, and the predecessor state is added to the path node sequence. The iteration stops when the predecessor state is the initial state. A basic path is generated based on the path node sequence.
6. The robotic arm task planning and control method as described in claim 1, characterized in that, By using reinforcement learning to adjust the basic path according to changes in the environment, a final path is generated, including: Monitor changes in environmental conditions and generate environmental change information; Based on the environmental change information and the basic path, a reinforcement learning state representation is constructed; The reinforcement learning state representation is input into the policy network, and the output path adjustment action is generated. Perform the path adjustment action to obtain the adjusted path; A reward signal is generated based on the environmental interaction results, and the parameters of the policy network are updated based on the reward signal. Determine whether the adjusted path simultaneously satisfies the conditions for shortening the path length and improving the obstacle avoidance success rate. When the conditions for shortening the path length and improving the obstacle avoidance success rate are met, the adjusted path will be used as the final path.
7. The robotic arm task planning and control method as described in claim 1, characterized in that, Decoding the target motion sequence into joint control commands and generating motion control signals based on the final path includes: Get the current state of the robotic arm joints; The target motion sequence and the current robotic arm joint state are processed by the policy network in the motion decoder to generate joint control parameters. The joint control parameters are converted into joint control commands; Analyze the sequence of key points in the final path; The sequence of key points of the trajectory is subjected to spline interpolation to generate a Cartesian space trajectory instruction, and the Cartesian space trajectory instruction is used as a motion control signal. The motion control signal is converted into joint spatial position commands through inverse kinematics; The joint spatial position command and the joint control command are input into the motor controller, and the motor controller generates a pulse width modulation motor drive signal. The motor driving the robotic arm executes the pulse-width modulated motor drive signal.
8. A robotic arm task planning and control device, characterized in that, The robotic arm task planning and control device includes: An environment perception module is used to acquire visual information and task description language information, and generate an environment model based on the visual information. The information encoding module is used to encode the visual information and the task description language information to obtain visual features and language features, respectively; The task planning module is used to perform task planning based on the visual features and the language features, and generate a target action sequence using a language model. The path generation module is used to generate a basic path for the movement of the robotic arm through dynamic programming under the environment model. The path optimization module is used to adjust the basic path according to environmental changes through reinforcement learning to generate the final path; A control command generation module is used to decode the target motion sequence into joint control commands and generate motion control signals based on the final path; The instruction optimization module is used to optimize the joint control instructions through reinforcement learning based on feedback rewards obtained from the environment during the execution of the joint control instructions.
9. A computer device, characterized in that, The computer device includes a memory, a processor, and a robotic arm task planning and control program stored in the memory and executable on the processor. When the robotic arm task planning and control program is executed by the processor, it implements the steps of the robotic arm task planning and control method as described in any one of claims 1-7.
10. A computer-readable storage medium, characterized in that, The storage medium stores a robotic arm task planning and control program, which, when executed by a processor, implements the steps of the robotic arm task planning and control method as described in any one of claims 1-7.
Citation Information
Cited By
Intelligent robot with body, material transfer method, storage medium and program product
CN121374644A
Production robot control system based on semantic recognition
CN121403393A
Motion control method and device based on vision and motion generation model
CN121973248A
Fruit maturity detection and grading picking method and system
CN121999481A