Humanoid robot control method and device based on visual large model, and robot

CN121468602BActive Publication Date: 2026-09-15SHANGHAI SPIDER-MAN ROBOT CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202610024426.3
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2026-01-09
Publication Date
2026-09-15
Estimated Expiration
2046-01-09

AI Technical Summary

Technical Problem

1、基于位置伺服的开环控制;该方式的可靠性高度依赖固定、标准化的生产环境,虽能适配传统工业人形机械臂的应用场景,但因灵活性不足、适配成本较高,而且缺乏实时反馈调整机制,无法补偿机械误差或复杂工业场景下的扰动,严重依赖人形机械臂的重复定位精度

Benefits of technology

[0021] Compared with existing technologies, the advantages of this invention are as follows: It perceives the target scene using a large visual model, constructs a local world model, and extracts the pose data of the target workpiece from the local world model; it extracts feature data from the end effector of the humanoid robotic arm based on the large visual model, performs mapping processing, and calculates the actual posture data of the target workpiece after it is grasped by the humanoid robotic arm; it calculates control parameters based on the pose data and the actual posture data; and it controls the humanoid robotic arm to operate on the target workpiece according to the control parameters. By constructing a local world model around the target workpiece and associating the constructed local world model with the humanoid robotic arm, it utilizes low-cost visual perception technology to provide visual guidance for the grasping and placement of the humanoid robotic arm, achieving high-precision control with low computational load, thus meeting the usage requirements of humanoid robotic arms and having a wide range of applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN121468602B_ABST
    Figure CN121468602B_ABST
Patent Text Reader

Abstract

The application relates to the field of robot control and provides a humanoid robot arm control method and device based on a visual large model and a robot. The visual large model is used to perceive a target scene, construct a local world model, and extract pose data of a target workpiece from the local world model. The visual large model is used to extract feature data of the end of the humanoid robot arm, and the feature data is subjected to mapping processing to calculate actual attitude data of the target workpiece after the target workpiece is grabbed by the humanoid robot arm. Control parameters are obtained by calculating the pose data and the actual attitude data. The humanoid robot arm is controlled according to the control parameters to operate the target workpiece. The local world model is constructed around the target workpiece, the constructed local world model is associated with the humanoid robot arm, the low-cost visual perception technology is used to visually guide the grabbing and placing of the humanoid robot arm, high-precision control is realized, the calculation amount is small, the use demand of the humanoid robot arm is met, and the application range is wide.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot control, and in particular to a method, device and robot for controlling a humanoid robotic arm based on a large visual model. Background Technology

[0002] Traditional industrial humanoid robotic arms, through long-term technological evolution, possess core characteristics of high precision and strong rigidity. They are widely used in automated production for workpiece gripping and placement, involving tasks such as assembly, sorting, and loading / unloading. Their core objective is to achieve stable gripping and precise placement of workpieces within material frames and fixtures while maintaining high precision, high speed, and high reliability.

[0003] Currently, the grasping and placement operations of robots in industrial scenarios are mainly achieved through three methods: 1. Open-loop control based on position servo: The reliability of this method is highly dependent on a fixed and standardized production environment. Although it can be adapted to the application scenarios of traditional industrial humanoid robotic arms, it is not flexible enough, has high adaptation costs, and lacks a real-time feedback adjustment mechanism. It cannot compensate for mechanical errors or disturbances in complex industrial scenarios and is heavily dependent on the repeatability of the humanoid robotic arm.

[0004] 2. Visual servo-based closed-loop control: This method requires the humanoid robotic arm to have microsecond-level real-time control response capability. When there are obstacles around the workpiece, its trajectory planning and control logic are difficult to adapt to complex environments, and the application scenarios are obviously limited. Moreover, the large amount of real-time calculation is not suitable for narrow and restricted environments.

[0005] 3. Impedance / admittance control based on force feedback: Traditional force control solutions are only applicable to feedback control after the humanoid robotic arm comes into contact with the target workpiece. In non-contact scenarios, they can only achieve basic collision detection functions. At the same time, the high cost of force sensors also limits their large-scale promotion. Even if the production environment is modified to adapt to industrial robots, or a six-dimensional force sensor is installed to provide force feedback, the application scenarios are limited, the cost is high, and it is difficult to universally adapt force feedback technology to the environment.

[0006] With technological advancements, the penetration of humanoid robots into industrial settings and their replacement of manual labor has become an inevitable trend in industry development. While humanoid robotic arms offer advantages such as high flexibility, low cost, and a large operating space, their control precision lags significantly behind traditional industrial humanoid robotic arms. Furthermore, existing control methods are insufficient to meet the precision control requirements of humanoid robotic arms in industrial settings. Therefore, how to apply humanoid robotic arms to high-precision loading and unloading scenarios with stringent industrial constraints has become a key technological challenge in promoting the practical application of humanoid robots in the industrial field.

[0007] Therefore, there is an urgent need for humanoid robotic arm control methods, devices, and robots based on large visual models to improve the above problems. Summary of the Invention

[0008] This invention provides a humanoid robotic arm control method, device, and robot based on a large visual model. This invention is used to solve the control accuracy problem in strongly constrained scenarios and can achieve high-precision control of humanoid robotic arm operations.

[0009] According to a first aspect of the present invention, a humanoid robotic arm control method based on a large visual model is provided, comprising: perceiving a target scene through a large visual model, constructing a local world model, and extracting pose data of a target workpiece from the local world model; extracting feature data of the end effector of the humanoid robotic arm based on the large visual model, performing mapping processing, and calculating the actual posture data of the target workpiece after it is grasped by the humanoid robotic arm; calculating control parameters based on the pose data and the actual posture data; and controlling the humanoid robotic arm to operate on the target workpiece according to the control parameters.

[0010] In one implementation, perceiving the target scene through a large visual model and constructing a local world model includes: acquiring multi-view images of the target scene through a rotatable camera; processing the multi-view images based on the large visual model to construct a local world model, and associating the local world model with a humanoid robotic arm.

[0011] In one implementation, the process of extracting feature data from the end effector of the humanoid robotic arm based on a large visual model and performing mapping processing to calculate the actual posture data of the target workpiece after it is grasped by the humanoid robotic arm includes: extracting feature data from the end effector of the humanoid robotic arm based on a large visual model; mapping the feature data to a tool coordinate system; and calculating the actual posture data of the target workpiece after it is grasped by the humanoid robotic arm based on a fixed transformation relationship between the tool coordinate system and the end effector of the humanoid robotic arm.

[0012] In one implementation, the control parameters are calculated based on the pose data and actual posture data, including: controlling the humanoid robotic arm to grasp the target workpiece and perform movement and / or rotation operations, acquiring multiple actual posture data during the process, calculating the deviation value based on the pose data and multiple actual posture data, and compensating for the pose deviation of the target workpiece based on the deviation value to obtain the control parameters.

[0013] In one implementation, the local world model includes the target workpiece, obstacles, and object distribution relationships.

[0014] In one embodiment, the method further includes: assigning different safety distances and weights to each obstacle according to its type, wherein the safety distance is used to provide a relative distance limit between the humanoid robotic arm and the corresponding obstacle, and the weight is used to distinguish the risk level of each obstacle so that the humanoid robotic arm maintains a greater relative distance from obstacles with higher risk levels; planning a collision-free path for the target workpiece based on the distribution relationship of obstacles and objects, safety distances, and weights in the local world model; and controlling the humanoid robotic arm to move along the collision-free path based on the control parameters.

[0015] In one implementation, the method further includes: updating the perception data of the target scene in real time through a large visual model to dynamically adjust the object information, obstacle information, and object distribution relationships within the local world model.

[0016] According to a second aspect of the present invention, a humanoid robotic arm control device based on a large visual model is provided, for the method of any one of the first aspects, comprising: a data acquisition unit, configured to perceive a target scene through a large visual model, construct a local world model, and extract pose data of a target workpiece from the local world model; a mapping unit, configured to extract feature data of the end effector of the humanoid robotic arm based on the large visual model, perform mapping processing, and calculate the actual posture data of the target workpiece after it is grasped by the humanoid robotic arm; a calculation unit, configured to calculate control parameters based on the pose data and the actual posture data; and a control unit, configured to control the humanoid robotic arm to operate on the target workpiece according to the control parameters.

[0017] According to a third aspect of the present invention, an electronic device is provided, including a memory and a processor, wherein the memory is used to store a computer program executable by the processor; and the processor is used to execute the computer program in the memory to implement the method described above.

[0018] According to a fourth aspect of the present invention, a computer-readable storage medium is provided having a computer program stored thereon, which, when executed by a processor, enables the implementation of the above-described method.

[0019] According to a fifth aspect of the present invention, a computer program product is provided, comprising a computer program that, when executed by a processor, implements the method described above.

[0020] According to a sixth aspect of the present invention, a robot is provided, including a rotatable camera, a humanoid robotic arm, and a controller; A rotatable camera is mounted on the robot's head and can rotate relative to the robot, used to perceive the target scene through a large visual model; The controller is used to execute the methods described above.

[0021] Compared with existing technologies, the advantages of this invention are as follows: It perceives the target scene using a large visual model, constructs a local world model, and extracts the pose data of the target workpiece from the local world model; it extracts feature data from the end effector of the humanoid robotic arm based on the large visual model, performs mapping processing, and calculates the actual posture data of the target workpiece after it is grasped by the humanoid robotic arm; it calculates control parameters based on the pose data and the actual posture data; and it controls the humanoid robotic arm to operate on the target workpiece according to the control parameters. By constructing a local world model around the target workpiece and associating the constructed local world model with the humanoid robotic arm, it utilizes low-cost visual perception technology to provide visual guidance for the grasping and placement of the humanoid robotic arm, achieving high-precision control with low computational load, thus meeting the usage requirements of humanoid robotic arms and having a wide range of applications. Attached Figure Description

[0022] Figure 1 A flowchart illustrating a humanoid robotic arm control method based on a large visual model, provided for an embodiment of the present invention; Figure 2 A schematic diagram of the structure of a humanoid robotic arm control device based on a large visual model provided in an embodiment of the present invention; Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of the present invention. Detailed Implementation

[0023] Unless otherwise defined, the technical or scientific terms used in this specification should have the ordinary meaning understood by one of ordinary skill in the art to which this invention pertains. Specific embodiments of the invention will be described below with reference to the accompanying drawings. It should be noted that, in order to provide a concise description, this specification cannot provide a detailed description of all features of the actual embodiments. Without departing from the spirit and scope of the invention, those skilled in the art can make modifications and substitutions to the embodiments of the invention, and the resulting embodiments are also within the protection scope of the invention.

[0024] like Figure 1 As shown, the first embodiment of the present invention provides a humanoid robotic arm control method based on a large visual model, including: S101 uses a large visual model to perceive the target scene, constructs a local world model, and extracts the pose data of the target workpiece from the local world model. The pose data includes the position and orientation information of the target workpiece.

[0025] In some embodiments, perceiving the target scene through a large visual model and constructing a local world model includes: acquiring multi-view images of the target scene using a rotatable camera (such as a rotatable camera mounted on the head of a humanoid robot); processing the multi-view images based on the large visual model to construct a local world model; associating the local world model with the humanoid robotic arm; using a large visual model with zero-sample recognition and segmentation capabilities; and acquiring multi-view information through the active rotation of the robot's head to achieve real-time, intelligent perception and understanding of the working environment, thereby improving the generalized recognition capability of unknown workpieces. The humanoid robot can understand the environment without pre-entering a precise 3D model, thus eliminating the need for cumbersome training and avoiding the cost of purchasing expensive 3D scanning equipment, thereby reducing the hardware cost of the entire perception system.

[0026] S102: Based on the large visual model, feature data of the humanoid robotic arm's end effector is extracted and mapped to calculate the actual posture data of the target workpiece after it is grasped by the humanoid robotic arm. The feature data includes geometric parameters of the robotic arm's end effector and visually recognizable marker information.

[0027] In other embodiments, feature data of the humanoid robotic arm's end effector is extracted based on a large visual model and mapped to calculate the actual posture data of the target workpiece after it is grasped by the humanoid robotic arm. This includes: extracting feature data of the humanoid robotic arm's end effector based on a large visual model; mapping the feature data to a tool coordinate system (i.e., the TCP coordinate system); and calculating the actual posture data of the target workpiece after it is grasped by the humanoid robotic arm based on a fixed transformation relationship between the tool coordinate system and the humanoid robotic arm's end effector. For example, a humanoid robot continuously identifies feature data of the robotic arm's end effector (such as AprilTag codes, markers of specific colors, or geometric features of the robotic arm's end effector identified through a large model) using a rotatable camera on its head. The identified end effector position is then associated with the workpiece held on the end effector through a fixed transformation, thereby estimating the precise pose of the workpiece in the rotatable camera coordinate system in real time. In this way, the control system directly uses the "workpiece" as the control object, rather than the "robotic arm joint," forming a workpiece-level closed-loop control with visual feedback as the core. This achieves absolute accuracy in prompting workpiece gripping and placement, while compensating for repetitive positioning errors caused by insufficient rigidity of the robotic arm and gear backlash, thus realizing high-precision operation that is truly decoupled from the precision of the robot body.

[0028] By extracting semantic information and precise pixel coordinates of target workpieces, obstacles, and robotic arm end effectors from 2D images, and then combining the multi-view 2D semantic information with the robot's own pose (position and angle of the head-rotating camera), an advanced end-to-end full-scene 3D understanding model (Omni-Scene) is adopted to achieve direct and efficient reconstruction from multi-view 2D images to 3D semantic scenes.

[0029] S103 calculates the control parameters based on the pose data and the actual attitude data.

[0030] In some embodiments, the control parameters are calculated based on pose data and actual posture data, including: controlling the humanoid robotic arm to grasp the target workpiece and perform movement and / or rotation operations, acquiring multiple actual posture data during this process, calculating the deviation value based on the pose data and multiple actual posture data; compensating for the pose deviation of the target workpiece based on the deviation value to obtain the control parameters. Since the body positioning feedback of the humanoid robotic arm with low precision is unreliable, this invention abandons traditional joint space control and instead uses a rotatable camera to continuously identify the end-effector features of the robotic arm to directly calculate the real-time pose of the target workpiece. This makes the feedback source of the control system the actual relative relationship between the workpiece and the target, thereby bypassing the error of the robotic arm body and directly improving the end-effector execution accuracy.

[0031] S104 controls the humanoid robotic arm to operate on the target workpiece (such as gripping, placing, moving and rotating) according to the control parameters.

[0032] In some other embodiments, the local world model includes the target workpiece, obstacles, and object distribution relationships. Obstacles include airtight equipment, material frames, and other workpieces (such as other workpieces identical to the target workpiece).

[0033] In some specific embodiments, the method further includes assigning different safety distances and weights to each type of obstacle based on its type. For example, a first safety distance is set for obstacles labeled "airtight equipment," and a second safety distance is set for obstacles labeled "material frame," wherein the first safety distance is greater than the second safety distance.

[0034] In some specific embodiments, the method further includes: planning a collision-free path for the target workpiece in Cartesian space based on the distribution relationship of obstacles and objects, safety distance, and weights in the local world model. The safety distance is used to provide a relative distance limit between the humanoid robotic arm and the corresponding obstacle, and the weights are used to distinguish the risk level of each obstacle, so that the humanoid robotic arm maintains a greater relative distance from obstacles with higher risk levels. This ensures that the robot can not only avoid collisions but also actively avoid high-risk areas, greatly improving reliability and safety in complex industrial environments. It enables safe and reliable obstacle avoidance in narrow spaces or areas with dense obstacles, adapts to dynamically changing environments, has higher operational robustness, and allows for different safety strategies to be set for obstacles of different natures, protecting precision equipment. For example, coarse guidance: Based on the constructed local world model, a collision-free path is planned from the starting point to a preparatory position near the target workpiece. This stage utilizes the semantic information and rough three-dimensional positions of objects and obstacles in the world model to guide the robotic arm to move quickly and safely to the vicinity of the target. Fine closed-loop: After the robot reaches the preparatory position, it switches to high-precision visual servo closed-loop control. At this point, the system continuously compares the target workpiece pose calculated in real-time by vision with the target workpiece pose obtained from the world model, generates control commands, and drives the robotic arm to make fine adjustments until the target workpiece is perfectly aligned with the target position and pose. This strategy effectively divides the work: coarse positioning solves large-scale movement and obstacle avoidance, while fine positioning solves the final high-precision operation, avoiding the limitations of a single control loop when handling complex tasks. While ensuring final accuracy, it significantly improves overall work efficiency, meets industrial cycle time, avoids the limitations of a single control mode in complex tasks, and makes the system more stable. Given that directly using high-gain vision servoing throughout the entire process can easily cause oscillations or instability during large-scale movements, this invention is divided into two stages: first, using the semantic information of the world model for rapid and safe coarse guidance to the preparatory position; then, initiating a close-range vision servoing fine closed loop to eliminate all accumulated errors. This ensures both rapid approach and final accuracy, optimizing the total task time. This strategy effectively divides the work: coarse positioning solves large-scale movement and obstacle avoidance, while fine positioning solves the final high-precision operation, avoiding the limitations of a single control loop when handling complex tasks.

[0035] In some specific embodiments, the method further includes: controlling the humanoid robotic arm to move along a collision-free path in Cartesian space based on control parameters (such as moving to the vicinity of the target workpiece or moving to the destination after grasping the target workpiece).

[0036] In other embodiments, the method further includes: updating the perception data of the target scene in real time through a large visual model to dynamically adjust object information, obstacle information, and object distribution relationships within the local world model.

[0037] The advantages of the embodiments of the present invention are as follows: 1. To address the issue of low control precision in highly constrained industrial scenarios for humanoid robotic arms, the system first uses a large-scale vision model to perceive the scene, constructing a local world model and extracting the pose data of the target workpiece. Simultaneously, the system uses the same vision system to identify feature data at the robotic arm's end effector in real time, mapping this data to the tool coordinate system (TCP coordinate system). Based on the fixed transformation relationship between the tool coordinate system and the robotic arm's end effector, the system calculates the actual pose data of the target workpiece after it is grasped by the robotic arm. Finally, by continuously comparing the real-time pose of the workpiece with the target workpiece's pose, a closed-loop control loop is formed. This loop directly compensates for workpiece pose deviations, obtaining the deviation value. Based on this deviation value, the system compensates for the target workpiece's pose deviation, obtaining control parameters. These control parameters are then used to control the robotic arm to grasp, place, move, and rotate the target workpiece, thus bypassing robotic arm body errors and achieving high-precision operation.

[0038] 2. To enable robots to operate safely and efficiently in confined spaces, fine-grained collision-free path planning is achieved using semantic information provided by a large visual model. In the constructed local world model, not only is the geometric information of obstacles identified, but they are also assigned semantic labels, including "airtight equipment" and "material box." Based on this, the path planning algorithm can adopt differentiated safety strategies according to different semantic meanings. Guided by both the global map and semantic information, a safe and efficient path is planned from the starting point to a pre-positioned location near the target point. This strategy ensures the reliability of the robotic arm's movement in complex environments, while significantly improving operational efficiency by optimizing the movement trajectory, meeting the cycle time requirements of industrial scenarios.

[0039] 3. The entire process, from environmental perception (eyes) to decision-making and planning (brain) to action execution (hands), is built and driven by unified visual information, thus forming a highly collaborative intelligent agent. Furthermore, because core perception does not rely on pre-defined fixed models, when tasks or environments change, only high-level instructions (such as the description of the target object) need to be adjusted. The system can then automatically adapt through rescanning and modeling, achieving a highly flexible and automated deployment.

[0040] 4. This invention utilizes only visual information and introduces low-cost distance sensors, such as single-point ToF or force sensors. The semantic world model constructed from the large visual model is fused with the physical information provided by these sensors. Vision + Distance: When the robotic arm approaches the workpiece for the final gripping, the depth estimation from monocular vision may have millimeter-level errors. At this time, a single-point ToF sensor installed at the end of the arm can provide precise absolute distance information to the workpiece surface, which is fused with the visually estimated depth to further improve depth estimation accuracy and ensure a high gripping success rate. Vision + Torque: When the workpiece is placed into the fixture, pure visual guidance may jam due to minute errors. At this time, a force sensor detects sudden changes in force / torque during assembly contact, triggering a compliant assembly strategy to achieve high-precision hole-shaft assembly and solve the "last millimeter" problem.

[0041] like Figure 2 As shown, based on the above control method, the second embodiment of the present invention provides a humanoid robotic arm control device based on a large visual model, including: a data acquisition unit 201, used to perceive the target scene through the large visual model, construct a local world model, and extract the pose data of the target workpiece from the local world model; a mapping unit 202, used to extract the feature data of the end effector of the humanoid robotic arm based on the large visual model, perform mapping processing, and calculate the actual posture data of the target workpiece after it is grasped by the humanoid robotic arm; a calculation unit 203, used to calculate control parameters based on the pose data and the actual posture data; and a control unit 204, used to control the humanoid robotic arm to operate on the target workpiece according to the control parameters.

[0042] It should be understood that all relevant content of each step involved in the above method embodiments can be referenced to the functional description of the corresponding functional module, and will not be repeated here. In addition, the use of suffixes such as "module," "part," or "unit" to indicate elements is only for the convenience of the description of the present invention, and has no specific meaning in itself. Therefore, "module," "part," or "unit" can be used in combination.

[0043] A third embodiment of the present invention provides an electronic device, including a memory and a processor, wherein the memory is used to store a computer program executable by the processor; and the processor is used to execute the computer program in the memory to implement the method of any of the above embodiments.

[0044] Figure 3 This is a block diagram illustrating an electronic device according to an exemplary embodiment. For example, electronic device 900 may be provided as a server. (Refer to...) Figure 3The electronic device 900 includes a processing component 922, which further includes one or more processors, and memory resources represented by memory 932 for storing instructions, such as application programs, that can be executed by the processing component 922. The application programs stored in memory 932 may include one or more modules, each corresponding to a set of instructions. Furthermore, the processing component 922 is configured to execute instructions to perform the methods described above.

[0045] The processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. A general-purpose processor can be a microprocessor or any conventional processor.

[0046] The memory 932 can be an internal storage unit of the electronic device 900, such as a hard disk or RAM of the electronic device 900. The memory 932 can also be an external storage device of the electronic device 900, such as a plug-in hard disk, SmartMediaCard (SMC), SecureDigital (SD) card, or FlashCard. Furthermore, the memory 932 can include both internal and external storage units of the electronic device 900. The memory 932 is used to store computer programs and other programs and data required by the electronic device. The memory 932 can also be used to temporarily store data that has been output or will be output.

[0047] Electronic device 900 may also include a power supply component 926 configured to perform power management of electronic device 900, a wired or wireless network interface 950 configured to connect electronic device 900 to a network, and an input / output (I / O) interface 958. Electronic device 900 may operate on an operating system stored in memory 932, such as Windows Server™, MacOS X™, Unix™, Linux™, FreeBSD™, or similar.

[0048] In an exemplary embodiment, a non-transitory computer-readable storage medium including instructions is also provided, such as a memory 932 including instructions, which can be executed by a processing component 922 of an electronic device 900 to perform the above-described method. For example, the non-transitory computer-readable storage medium may be a ROM, random access memory (RAM), CD-ROM, magnetic tape, floppy disk, and optical data storage device.

[0049] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is merely an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above. The functional units and modules in the embodiments can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit. The integrated unit can be implemented in hardware or as a software functional unit. Furthermore, the specific names of the functional units and modules are only for easy differentiation and are not intended to limit the scope of protection of this invention. The specific working process of the units and modules in the above system can be referred to the corresponding process in the foregoing method embodiments, and will not be repeated here.

[0050] A fourth embodiment of the present invention provides a readable storage medium storing a program, which, when executed, implements the method of any of the above embodiments.

[0051] The fifth embodiment of the present invention provides a computer program product, including a computer program, which, when executed, implements the method of any of the above embodiments.

[0052] The sixth embodiment of the present invention provides a robot, including a rotatable camera, a humanoid robotic arm, and a controller; the rotatable camera is disposed on the head of the robot and is rotatable relative to the robot, and is used to perceive a target scene through a large visual model; the controller is used to execute the method of any of the above embodiments.

[0053] In summary, the humanoid robotic arm control method, device, and robot based on a large visual model disclosed in this invention construct a local world model around the target workpiece and associate the constructed local world model with the humanoid robotic arm. It utilizes low-cost visual perception technology to provide visual guidance for the humanoid robotic arm's grasping and placement, achieving high-precision control with minimal computational load. Furthermore, it leverages the semantic information provided by the large visual model for refined collision-free path planning. The constructed local world model not only identifies the geometric information of obstacles but also assigns them semantic labels such as "airtight equipment" and "material frame," enabling the robot to operate safely and efficiently in confined spaces, with a wide range of applications.

[0054] In this invention, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance. The term "multiple" refers to two or more objects unless otherwise explicitly defined. The term "and / or" describes the relationship between related objects, indicating that three relationships can exist; for example, A and / or B can represent: A alone, A and B simultaneously, or B alone, where A and B can be singular or plural. The character " / " generally indicates that the preceding and following related objects are in an "or" relationship. References to "one embodiment" or "some embodiments" described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," and "in still other embodiments" appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized. The term "connection" includes direct connections and indirect connections, unless otherwise stated.

[0055] In embodiments of the present invention, "exemplarily" or "for example" are used to indicate that they are examples, illustrations, or descriptions. Any embodiment or design described as "exemplarily" or "for example" in embodiments of the present invention should not be construed as being more preferred or advantageous than other embodiments or design solutions. Rather, the use of "exemplarily" or "for example" is intended to present the relevant concepts in a specific manner.

[0056] The above description of the embodiments is intended to enable those skilled in the art to understand and apply the present invention. It will be apparent to those skilled in the art that various modifications can be made to these embodiments, and the general principles described herein can be applied to other embodiments without creative effort. Therefore, the present invention is not limited to the embodiments described herein, and any improvements and modifications made by those skilled in the art based on the disclosure of the present invention without departing from the scope and spirit of the invention are within the scope of the present invention.

Claims

1. A control method for a humanoid robotic arm based on a large visual model, characterized in that, include: The target scene is perceived through a large visual model, a local world model is constructed, and multi-view images of the target scene are acquired through a rotatable camera. The multi-view images are processed based on a large visual model to construct a local world model, and the local world model is associated with a humanoid robotic arm. The pose data of the target workpiece is then extracted from the local world model. Based on the aforementioned large visual model, feature data of the end effector of the humanoid robotic arm is extracted and mapped to calculate the actual posture data of the target workpiece after it is grasped by the humanoid robotic arm. The control parameters are calculated based on the pose data and the actual pose data. The humanoid robotic arm is controlled to operate on the target workpiece according to the control parameters. The visual big model updates the perception data of the target scene in real time to dynamically adjust the object information, obstacle information and object distribution relationship within the local world model; the types of obstacles include airtight equipment, material frames and other workpieces that are the same as the target workpiece; Based on the type of obstacle, each obstacle is assigned a different safety distance and weight. The safety distance is used to provide a relative distance limit between the humanoid robotic arm and the corresponding obstacle, and the weight is used to distinguish the risk level of each obstacle so that the humanoid robotic arm maintains a greater relative distance from obstacles with higher risk levels. Based on the distribution relationship of obstacles and objects in the local world model, the safety distance, and the weight, a collision-free path for the target workpiece is planned. Based on the control parameters, the humanoid robotic arm is controlled to move along the collision-free path.

2. The method according to claim 1, characterized in that, Based on the aforementioned large visual model, feature data of the humanoid robotic arm's end effector are extracted and mapped to calculate the actual posture data of the target workpiece after it is grasped by the humanoid robotic arm, including: Based on the aforementioned large visual model, feature data of the humanoid robotic arm's end effector were extracted. The feature data is mapped to the tool coordinate system, and the actual posture data of the target workpiece after being grasped by the humanoid robotic arm is calculated based on the fixed transformation relationship between the tool coordinate system and the end of the humanoid robotic arm.

3. The method according to claim 1, characterized in that, Based on the pose data and actual pose data, control parameters are calculated, including: The humanoid robotic arm is controlled to grasp the target workpiece and perform movement and / or rotation operations. During this process, multiple actual posture data are acquired, and a deviation value is calculated based on the posture data and the multiple actual posture data. The positional deviation of the target workpiece is compensated based on the deviation value to obtain control parameters.

4. The method according to claim 1, characterized in that, The local world model includes the target workpiece, obstacles, and the distribution relationships of objects.

5. A humanoid robotic arm control device based on a large visual model, used in the method according to any one of claims 1 to 4, characterized in that, include: The acquisition unit is used to acquire multi-view images of the target scene through a rotatable camera; process the multi-view images based on a large visual model to construct a local world model, associate the local world model with the humanoid robotic arm, and extract the pose data of the target workpiece from the local world model; The mapping unit is used to extract feature data of the end of the humanoid robotic arm based on the large visual model, perform mapping processing, and calculate the actual posture data of the target workpiece after it is grasped by the humanoid robotic arm. The calculation unit is used to calculate the control parameters based on the pose data and the actual posture data; The control unit is used to control the humanoid robotic arm to operate on the target workpiece according to the control parameters.

6. An electronic device, characterized in that, The system includes a memory and a processor, wherein the memory stores a computer program executable by the processor; and the processor executes the computer program in the memory to implement the method as described in any one of claims 1 to 4.

7. A computer-readable storage medium having a computer program stored thereon, characterized in that, When the executable computer program in the storage medium is executed by a processor, it can implement the method as described in any one of claims 1 to 4.

8. A computer program product, comprising a computer program, characterized in that, When the computer program is executed by a processor, it implements the method as described in any one of claims 1 to 4.

9. A robot, characterized in that, Includes a rotating camera, a humanoid robotic arm, and a controller; The rotatable camera is mounted on the head of the robot and is rotatable relative to the robot, and is used to perceive the target scene through a large visual model; The controller is used to perform the method as described in any one of claims 1 to 4.

Citation Information

Patent Citations

  • Humanoid robot grabbing method based on three-dimensional vision

    CN119458364A

  • Manipulator grabbing planning system and method based on visual identification

    CN120382479A

  • Industrial robot walking control system based on obstacle recognition

    CN120595816A