A robot material handling method, robot, and storage medium

By using simple line drawings to extract the shape features of the work target and align them with environmental data in industrial robots, the problem of misjudgment of voice commands in noisy environments is solved, enabling efficient material handling under pure vision commands and improving operational reliability and automation efficiency.

CN120680517BActive Publication Date: 2026-04-1458 INTELLIGENT TECH (HANGZHOU) CO LTD
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2026-04-14

AI Technical Summary

Technical Problem

Existing industrial robot material handling technologies are prone to misinterpretation of voice commands due to noise interference in noisy industrial environments, and lack support from pure vision commands, which affects operational reliability and automation efficiency.

Method used

By extracting the shape features of the work target from the simple sketches drawn by the operator, combining them with environmental data for visual feature alignment, using a contrast loss function to filter background interference, calculating the target pose and material type, and generating material handling instructions to control the movement of the actuator.

Benefits of technology

Material handling with pure vision commands was achieved in noisy environments, improving the robot's environmental adaptability and operational reliability in complex industrial environments, reducing reliance on voice interaction, and increasing the automation efficiency of material handling.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120680517B_ABST
    Figure CN120680517B_ABST
Patent Text Reader

Abstract

The machine material carrying method, the robot and the storage medium disclosed by the application obtain a sketch depicting an appearance of a work target and carrying destination position information, extract a shape feature of the work target from the sketch, collect environment data in a current task scene, extract a visual feature from the environment data, align the visual feature with the shape feature by using a contrast loss function, extract a shape feature of the work target from a target region, calculate a target pose and confirm a material type from the environment data based on the target region, finally query a preset control information library based on the target object type to obtain corresponding action constraint information, combine the action constraint information, the target pose and the destination position information, generate a material carrying instruction, and control each actuator to move to complete a material carrying task. The material carrying can be realized in a noisy industrial scene by only using a pure visual instruction without inputting a complex text instruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of robot control technology, and in particular to a robot material handling method, a robot, and a storage medium. Background Technology

[0002] In industrial robot applications, material handling technology plays a crucial role in production efficiency and automation levels in manufacturing, warehousing, and logistics. However, existing industrial robot material handling technologies suffer from significant technical deficiencies in environmental interaction. For instance, in noisy industrial environments such as assembly lines and heavy machinery operating areas, voice commands are easily misinterpreted due to noise interference from machinery and equipment operation, leading to deviations in robot actions or even task failure. Furthermore, existing multimodal interaction methods rely excessively on voice input and lack support for purely visual commands. When dust or oil in industrial settings cause voice pickup devices to malfunction, or when operators are unable to communicate via voice due to safety concerns, robots struggle to receive accurate operational instructions. This limitation in interaction significantly reduces the operational reliability of industrial robots in noisy industrial environments, especially in scenarios requiring high-frequency human-machine collaboration or rapid command responses, severely hindering the automation efficiency of production processes. Summary of the Invention

[0003] This invention addresses the shortcomings of existing technologies by providing a robotic material handling method, comprising:

[0004] Obtain a simple line drawing depicting the shape of the task target drawn by the operator and the location information of the transport destination, and extract the shape features of the task target from the line drawing;

[0005] Collect environmental data within the current task scenario, extract visual features from the environmental data, align the visual features in the environmental data with the shape features in the sketch using a contrast loss function, and retain the target area that matches the target shape in the sketch; extract the shape features of the task target from the target area, and calculate the target pose and confirm the material type based on the target area from the environmental data;

[0006] Based on the target object type, query the preset control information database to obtain the corresponding motion constraint information, combine the motion constraint information, target pose and destination location information to generate material handling instructions; control each actuator to move according to the material handling instructions to complete the material handling operation.

[0007] Preferably, acquiring a simplified sketch of the target object drawn by the operator and the location information of the transport destination includes:

[0008] Collect simple line drawings of the target object drawn by the operator and text information describing the destination location. Extract the target object's shape features from the line drawings and extract the destination location information of the current transport task from the text information. The destination location information includes platform attribute information and coordinate information of the destination where the target object is stored; or

[0009] The system collects a simple sketch of the target object drawn by the operator and a voice message describing the location of the transport destination. The shape features of the target object are extracted from the sketch, and the destination location information of the current transport task is extracted from the voice message. The destination location information includes platform attribute information and coordinate information of the destination where the target object is transported and stored.

[0010] Preferably, the operator draws a simplified sketch of the target object and the location information of the transport destination. The target object's shape features are then extracted from the sketch, including:

[0011] The operator collects a sketch, which includes a first pattern depicting the shape of the target and a second pattern depicting the shape of the storage point where the target is located.

[0012] Extract the target shape features of the work target from the first pattern, and extract the platform shape features of the storage point where the transport destination is located from the second pattern.

[0013] Preferably, extracting the shape features of the task target from the simplified drawing specifically includes:

[0014] The simplified drawing is subjected to pattern recognition and differentiation. Based on the relative position or connection mark of each segmented independent pattern in the simplified drawing, the first pattern and the second pattern are distinguished according to the preset simplified drawing drawing rules.

[0015] Preferably, extracting the shape features of the work target from the target area, and calculating the target pose and confirming the material type from the environmental data based on the target area, further includes:

[0016] The system acquires images of the target object using its onboard camera, extracts the object's edges using a contour detection algorithm, and calculates the object's three-dimensional dimensions by combining the camera's calibration parameters.

[0017] The target density value is obtained by querying a preset density database according to the target material type, and the weight of the target object is calculated based on the three-dimensional dimensions of the object and the target density value.

[0018] Based on the weight and three-dimensional dimensions of the target object, determine whether it exceeds the current equipment load range; if it exceeds the current equipment load range and the target object can be split, then execute the task splitting process to generate a sub-task sequence for moving multiple small objects; otherwise, abandon the current moving task and issue a prompt.

[0019] Preferably, the target density value is obtained by querying a preset density database based on the target material type, specifically including:

[0020] If the target material type cannot be confirmed or a unique target density value cannot be matched, the image texture features on the target object image are identified, and the image texture features are matched with materials in a preset density database. The image texture features include roughness information, and the matched material density value is used as the current target density value.

[0021] Preferably, the robot material handling method further includes: if an obstacle is detected on the preset path during the movement to the material handling destination, the movement path is replanned based on the shape characteristics of the work target to avoid the obstacle.

[0022] The present invention also discloses a robot, including a body on which a controller, a memory, and an optical depth sensing device are mounted, the memory being used to store a computer program executable by the processor, wherein the processor is configured to execute the computer program in the memory to implement the method as described in any of the preceding claims.

[0023] Preferably, the optical depth sensing device includes a lidar, a depth camera, or multiple two-dimensional cameras capable of simultaneously capturing the same target.

[0024] The present invention also discloses a computer-readable storage medium that, when an executable computer program in the storage medium is executed by a processor, enables the implementation of the method described in any of the preceding claims.

[0025] The present invention discloses a robot material handling method, robot, and storage medium. The method involves acquiring a simplified sketch of the target object drawn by an operator, along with the destination location information. The target object's shape features are extracted from the sketch. Environmental data within the current task scenario is collected, and visual features are extracted from this data. A contrastive loss function is used to align the visual features in the environmental data with the shape features in the sketch, retaining the target area that matches the target shape in the sketch. The target object's shape features are extracted from this target area, and the target pose and material type are calculated from the environmental data based on this target area. Finally, based on the target object type, a preset control information database is queried to obtain corresponding motion constraint information. The motion constraint information, target pose, and destination location information are combined to generate a material handling command, which then controls each actuator to move and complete the material handling operation. This effectively solves the technical problems in existing industrial robot material handling technology, such as the easy misinterpretation of voice commands due to the strong noise environment in the industry, and the cumbersome description of complex work objects or processes by inputting commands through pure text. It enables material handling in noisy industrial scenarios by using only visual commands, without the need to input complex coordinates or commands, thereby improving the robot's environmental adaptability and operational reliability in complex industrial environments, reducing the reliance on voice interaction, and improving the automation efficiency of material handling.

[0026] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0027] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0028] Figure 1 This is a schematic diagram illustrating the specific process of a robot material handling method disclosed in an embodiment of the present invention.

[0029] Figure 2 This is a schematic diagram illustrating the specific process of load capacity verification steps disclosed in an embodiment of the present invention.

[0030] Figure 3 This is a simplified line drawing diagram of a transportation task disclosed in an embodiment of the present invention.

[0031] Figure 4 This is a schematic diagram of the structure of a dual-branch visual encoder model disclosed in an embodiment of the present invention. Detailed Implementation

[0032] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0033] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains. The terms “first,” “second,” and similar terms used in the specification and claims of this patent application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an” or “a” and similar terms do not indicate a limitation of quantity, but rather indicate the presence of at least one.

[0034] In this embodiment, as shown in the appendix Figure 1 As shown, a robot material handling method is disclosed, which may specifically include the following steps.

[0035] Step S1: Obtain a simple sketch of the target shape drawn by the operator and the location information of the transport destination, and extract the target shape features from the sketch.

[0036] Specifically, the robot can capture images of simple sketches manually drawn by the operator on physical or electronic media using a camera. For example, the operator can use paper and pen to draw a simple sketch of the target object to be moved, and then show it to the camera mounted on the industrial robot. The industrial robot can then recognize the sketch and input it for processing. Alternatively, the operator can also manually draw a simple sketch on a control screen connected to the robot, i.e., draw the object to be moved via a touchscreen connected to the industrial robot, thus inputting the sketch into the robot for processing.

[0037] In this embodiment, the data can be collected from a simple sketch drawn by the operator depicting the shape of the work target and text information describing the location of the transport destination. The shape features of the work target are extracted from the sketch, and the destination location information of the current transport task is extracted from the text information. The destination location information includes the platform attribute information and coordinate information of the destination where the work target is stored. For example, the operator draws a rectangular sketch representing a square cardboard box and labels it with the text information "shelf coordinates (10,20,3)". From this labeled text information, the information that the platform for storage is a shelf and the specific storage location is coordinates (10,20,3) can be obtained.

[0038] In another embodiment, the system can collect a simple sketch of the target object drawn by the operator and voice information describing the destination location. The shape features of the target object are extracted from the sketch, and the destination location information for the current transport task is extracted from the voice information. The destination location information includes platform attribute information and coordinate information of the destination where the target object is stored. By describing the shape features of the transported object—which are difficult to express clearly with voice—in the form of a simple sketch, the shape feature information of the transported object can be input more accurately for subsequent recognition and matching.

[0039] In another embodiment, the storage point where the target transport destination is located is also drawn in the same line drawing as the line drawing depicting the shape of the work target. By collecting the line drawings drawn by the operator, the line drawing includes a first pattern for depicting the shape of the work target and a second pattern for depicting the shape of the storage point where the target transport destination is located; the target shape features of the work target are extracted from the first pattern, and the platform shape features of the storage point where the transport destination is located are extracted from the second pattern.

[0040] For example, when drawing simple sketches to indicate a handling task, the sketches of materials such as cardboard boxes, gears, and cylindrical workpieces serve as the first sketches depicting the shape of the task target, while the sketches of objects where materials are placed, such as workpiece platforms, assembly lines, and shelves, serve as the second sketches depicting the shape of the objects where materials are placed, i.e., the target handling destination.

[0041] In this embodiment, a third pattern for constraining the execution action can be added to the simplified drawing. Scene constraint information can be generated based on the second pattern. The third pattern includes a first type of identifier for identifying the scene constraint information of grasping the work target, and a second type of identifier for identifying the operation scene constraint information of the work target to be transported, or indicating the transport, transfer, or storage location of the work target, wherein the robot is configured to perform the operation action after completing the grasping action of the work target.

[0042] For example, special marking symbols, such as the arrow symbol "→", can be designed to indicate the two, with the starting end of the arrow symbol indicating the object to be moved and the end of the arrow symbol indicating the object to be placed; such special marking symbols are the third pattern used to constrain the execution of actions.

[0043] In this embodiment, a library of identifiers can also be pre-configured as a third pattern to configure and store specific information of each identifier for subsequent simplified diagram recognition. These identifiers may include the following:

[0044] For situations where there are different types of materials to be grabbed or moved, you can label each type of material with a simple line drawing, such as "①", "②", or "③", to indicate the order of handling.

[0045] For situations where there are multiple similar materials to be moved, Arabic numerals such as "1", "2", and "3" can be marked next to the simple line drawing of the materials to indicate the number of materials to be moved; and special marking symbols can be designed to indicate that all materials to be moved need to be moved, such as "@".

[0046] In situations where there are multiple similar materials to be moved or objects to be placed, to identify the actual materials to be moved or objects to be placed, all similar materials to be moved or objects to be placed can be drawn, and the identified materials to be moved or objects to be placed can be specially marked, such as by marking them with a special symbol. ".

[0047] For scenarios where multiple materials to be transported need to be stacked sequentially, a stacking task symbol can be defined, such as the symbol "‡".

[0048] For scenarios requiring the sequential laying of multiple materials to be moved, a laying task symbol can be defined, such as the symbol "". ".

[0049] In this embodiment, the simplified drawing can also be pattern recognized and distinguished. Based on the relative position or connection mark of each segmented independent pattern in the simplified drawing, the first pattern and the second pattern can be distinguished according to the preset simplified drawing drawing rules.

[0050] Specifically, pattern recognition is performed on the simplified drawings. Based on the relative positions of each third pattern on the simplified drawings to the first and second patterns, it is distinguished between a first-type identifier used to identify constraints in the grasping scenario of the work target, and a second-type identifier used to identify constraints in the operational scenario of the work target. For example, it can be defined that simplified drawings of the materials to be moved are uniformly drawn on the left side of the drawing or panel, and the objects to be placed are uniformly drawn on the right side. Then, based on the positions of the different independent patterns on the drawing or panel, the first and second patterns can be determined. Third patterns marked on or near the first pattern are identified as first-type identifiers used to identify constraints in the grasping scenario of the work target. Third patterns marked on or near the second pattern are identified as second-type identifiers used to identify constraints in the operational scenario of the work target.

[0051] Furthermore, first scenario constraint information for constraining and controlling the grasping action of the work target can be generated based on the first type of identifier, and second scenario constraint information for constraining and controlling the handling and transfer action or handling and storage location of the work target can be generated based on the second type of identifier. After identifying the first type of identifier, the constraint information represented by the first type of identifier is combined with the work target to generate the first scenario constraint information for constraining and controlling the grasping action of the work target. The constraint information represented by the identified second type of identifier is combined with the handling and transfer action or handling and storage location to form the second scenario constraint information for constraining and controlling the handling and transfer action or handling and storage location of the work target.

[0052] Step S2: Collect environmental data within the current task scenario, extract visual features from the environmental data, align the visual features in the environmental data with the shape features in the sketch using a contrast loss function, and retain the target area that matches the target shape in the sketch; extract the target shape features from the target area, and calculate the target pose and confirm the material type based on the target area from the environmental data.

[0053] This can be achieved by using multiple cameras mounted on the robot to acquire 2D images of the task scene in real time, detecting key points of objects in the 2D images, and calculating the 3D pose of each object by combining camera intrinsic parameters and prior dimensions. The 3D pose is then matched with target shape features to obtain target location information and target object type. Alternatively, a depth camera or LiDAR mounted on the robot can acquire point cloud data of the task scene in real time, identify and acquire the 3D pose of each object from the point cloud data, and match the 3D pose with target shape features to obtain target location information and target object type.

[0054] Specifically, the shape purity of simple line drawings can be used as a prior constraint. Through cross-modal contrastive learning, the target features of complex background images can be aligned with the shape features of simple line drawings, thereby achieving accurate recognition under background interference.

[0055] The first step involves cross-modal encoding of line drawings and real images. For line drawings, such as those depicting rectangular objects like cardboard boxes to be moved, and real images against complex backgrounds, convolutional neural networks (CNNs) are used to extract the target shape features from the line drawings, respectively. Visual features in real images By comparing losses Forced Towards Alignment is used to filter out background interference. The design incorporates contrast loss. The calculation formula is as follows:

[0056]

[0057] In this formula, The shape feature vector representing the i-th simple drawing, such as the outline feature of a rectangle, has a dimension of d and is extracted and normalized to a unit vector by CNN. The visual feature vector representing the i-th complex background image, and... Same dimension and normalized; N represents the batch size, i.e., the number of sample pairs processed simultaneously; τ is a temperature parameter used to adjust the sharpness of the feature distribution, and its value range is usually (0.1, 0.5). The core logic of this contrastive loss function is that positive sample pairs... Representing simplified line drawings and real images of the same type of material, such as a rectangular simplified line drawing and an image of a cardboard box with a complex background, requires that the features of the two be as similar as possible in the embedding space; negative sample pairs To represent different types of materials, their characteristics are forced to be far apart. The optimization objective is to minimize... Make the feature dot product of similar samples It is significantly larger than negative samples, thus achieving target feature focusing under background interference.

[0058] Then, dynamic mask generation is performed. Based on the aligned features, a U-Net network is used to generate a target region mask, retaining only pixels within areas overlapping with the sketch shape, such as the rectangular outline, while masking background clutter. Thus, the target is forced to focus through the sketch shape anchor points, and even with complex backgrounds such as stacking, reflections, and occlusions, the material can still be accurately located based on shape priors.

[0059] For example, an industrial robot can be equipped with two 2D cameras to acquire image data in real time. Key points of objects, such as corners of shelves or midpoints of material boxes, can be detected in the 2D images. Using a camera pose estimation algorithm (PnP) combined with camera intrinsics and prior dimensions, the 3D pose of the object can be calculated. For instance, given the pixel coordinates of the four corners of a shelf, combined with the actual dimensions of the shelf (e.g., width 150cm, height 200cm), the pose matrix of the shelf in the camera coordinate system can be solved, and the material type can be confirmed by matching it with simple line drawing features. If the industrial robot is equipped with a depth camera, it can also acquire point cloud data in real time, estimate the object pose using a pose estimation algorithm (DenseFusion), and confirm the material type by matching it with simple line drawing features.

[0060] In this embodiment, as shown in the appendix Figure 2 As shown, step S2 may also include verifying the load capacity of the operation instructions to improve the success rate of robot task completion, which may specifically include the following.

[0061] Step S101: Acquire images of the target object using the mounted camera, extract the edges of the target object using a contour detection algorithm, and calculate the three-dimensional dimensions of the object by combining the camera calibration parameters.

[0062] The robot acquires images of materials using its onboard camera, extracts the object's edges using OpenCV's contour detection algorithm, and calculates the 3D dimensions by combining these with camera calibration parameters (intrinsic matrix K, extrinsic matrix R,t). For example, for a rectangular object, the pixel coordinates of the top, bottom, left, and right edges are detected and converted into length, width, and height in the world coordinate system using triangulation.

[0063] Step S102: Query the preset density database according to the target material type to obtain the corresponding target density value, and calculate the weight of the target object according to the three-dimensional dimensions of the object and the target density value.

[0064] If the target material type cannot be confirmed or a target density value cannot be uniquely matched, the image texture features on the target object image are identified, and the image texture features are matched with materials in a preset density database. The image texture features include roughness information, and the matched material density value is used as the current target density value.

[0065] Specifically, a material type-density mapping table can be established based on a preset density database, such as cardboard boxes at 0.5 kg / m³. 3 Metal parts 7800kg / m 3 It supports dynamic updates. If the material type is unknown, it automatically assigns a default density by matching typical materials in the database with image texture features such as roughness.

[0066] Reuse weight calculation formula:

[0067] Calculate the weight of the object to be moved. For example, a cardboard box with dimensions of 0.5m × 0.4m × 0.3m and a density of 500kg / m³. 3 Therefore, the weight is 0.5 × 0.4 × 0.3 × 500 = 30 kg.

[0068] Step S103: Determine whether the target object's weight and three-dimensional dimensions exceed the current equipment's load range. If the target object exceeds the current equipment's load range and can be separated, execute the task separation process to generate a sub-task sequence for moving multiple small objects. Otherwise, abandon the current moving task and issue a prompt.

[0069] Specifically, if the original instruction is to move a large box to a designated material placement location, and this large box contains multiple smaller boxes; after obtaining the material dimensions and estimating the weight, the large box exceeds the robot's load capacity, while the smaller boxes are within the load capacity range, then in this case, the robot will break down the task into a sequence of sub-tasks to move multiple smaller boxes. The robot will then present the solution to the user through a human-machine interface for confirmation, and execute the tasks according to a predefined priority, such as moving the leftmost box first. Alternatively, the robot can choose to abandon the task and provide a prompt to change equipment: in scenarios where the task cannot be broken down, such as with fragile items, the excessive material will be highlighted on the interface, prompting the user to select a robot with a higher load capacity, such as a 10kg-class robotic arm. Priority can also be marked, and load risk labels such as high load or exceeding limits requiring splitting can be added to the instructions, allowing the scheduling system to optimize task allocation and avoid frequent operation of overloaded equipment.

[0070] Step S3: Based on the target object type, query the preset control information database to obtain the corresponding motion constraint information, combine the motion constraint information, target pose and destination location information to generate a material handling instruction; control each actuator to move according to the material handling instruction to complete the material handling operation.

[0071] The preset control information database stores grasping action templates corresponding to different object types. These grasping action templates contain action constraint information for the current robot to perform grasping actions, such as joint angle range, maximum gripper opening, and load limits. An initial action template is generated for each material type in the sketch using a rule base. Objects of the same shape share the same action template parameters by adjusting only the variable ones.

[0072] In addition, if an obstacle is detected on the preset path during the movement to the destination, the movement path is replanned based on the shape characteristics of the target to avoid the obstacle.

[0073] The following example illustrates a robotic cardboard box handling task in a noisy industrial environment, such as a steel mill's building materials warehouse. Workers need to direct robots to move building materials of different specifications, such as steel pipes, planks, and bricks, from the stacking area to the production line buffer area. In this scenario, the workshop contains dynamic equipment such as overhead cranes and conveyor belts; floating dust causes visual blurring; alternating bright welding light and shadow creates significant background interference; and workers cannot interact via voice due to the noise, relying instead on simple line drawings to communicate with the robot. The specific handling method is as follows:

[0074] Workers draw the following content using a touch tablet, as shown in the attached image. Figure 3 As shown.

[0075] Target object: Draw a simple line drawing representing a steel pipe, marked with "①" to indicate the highest priority; draw a simple rectangle drawing representing a wooden board, marked with "②" to indicate the second highest priority;

[0076] Target transport location: Draw a simple sketch of a storage table with the text "Temporary Storage Area" pointing towards the production line entrance. The meaning of this instruction is: "Prioritize transporting steel pipe No. 1 to the production line temporary storage area, then transport wooden plank No. 2 to the production line temporary storage area."

[0077] The robot controller performs dual-modal feature extraction and interference filtering using a simple line drawing encoder: a simple line drawing of a steel pipe (straight line) is used to extract shape features via the ViT-B / 16 algorithm. (Length, straightness); Extraction of a simple line drawing of a wooden board (rectangle) (Aspect ratio, right angle features). Scene image encoder used: Images of the stacking area are captured by an industrial camera, and visual features are extracted using the ViT-B / 16 algorithm. However, due to the influence of dust and strong light, the original features contain a large number of noises such as dust particles and high-brightness clumps of welding spots.

[0078] Perform cross-modal contrastive learning: using contrastive loss Forced Align with simple line drawing features to filter out interference:

[0079]

[0080] in Features including steel pipes, wooden boards, dust, and light spots are considered, with τ=0.3 balancing feature discrimination and noise tolerance.

[0081] A Visual Graph Neural Network (VGNN) is used for scene graph construction and priority reasoning. The nodes are defined as follows:

[0082] Physical nodes: steel pipe (position (3,1.5,2), length 3m), wooden board (position (2,0.8,1), size 2m×1m), stacking area obstacles (conveyor belt, dynamic coordinate (x,0,0));

[0083] Logical node: Target buffer (location (5,0,1.2), capacity limit: only one steel pipe can be placed).

[0084] Modeling edge relationships: The edge weights between nodes are determined by spatial distance (e.g., the distance from the steel pipe to the temporary storage area = 2m), dynamic conflict probability (e.g., the overlap between the conveyor belt movement path and the steel pipe handling path = 40%), and priority markers (① > ②).

[0085] Generate and dynamically adjust action sequences:

[0086] The Transformer module takes a scene graph and a marker vector T=[①,②, temporary area] as input and outputs the predicted action sequence: [

[0088] {"type":"grasp","target":"steel pipe①","gripper":"V-shaped gripper","force":80N"},#Adapts to cylindrical steel pipe surfaces

[0089] {"type":"move","path":[(3,1.5,2)→(4,1.5,2)→(5,0,1.2)],"speed":0.2m / s"},# Avoid dynamic path of conveyor belt

[0090] {"type":"release","check":"Temporary storage space ≤ 1 root"},#Verify capacity limit

[0091] {"type":"grasp","target":"Wooden Board ②","gripper":"Flat Board Gripper","force":50N"},#Adapts to flat wooden boards

[0092] {"type":"move","path":[(2,0.8,1)→(5,0,1.2)],"speed":0.15m / s"}#Low-speed anti-shake ]

[0094] During robot operation, adjustments are made in real time to address dynamic disturbances. For example, for dust filtering: the U-Net model generates a mask based on a simple line drawing shape (such as a straight line mask for a steel pipe) to block dust particles and correct the actual position of the steel pipe to (3.1, 1.6, 2). For strong light suppression: the contrast of the steel pipe edge is enhanced by histogram equalization to avoid misjudgment of the contour caused by light spots.

[0095] Additionally, it can include multi-object collaboration and conflict resolution steps: For example, when transporting steel pipes, if it detects that the conveyor belt is about to enter the path, the path is dynamically adjusted to (3,1.5,2)→(3.5,2,2)→(5,0,1.2), avoiding it 2 seconds in advance. If, when transporting wooden planks, it is found that the buffer area has been filled by another robot with one steel pipe, triggering "secondary confirmation": the robot highlights the buffer area on the screen, and the worker remotely marks "temporarily expand the right side of the buffer area," updating the target position to (5.2,0,1.2).

[0096] In this embodiment, workers only need to draw simple sketches and markings, without needing to input complex coordinates or commands, meeting the needs of efficient operation in noisy environments and with low interaction costs. Furthermore, through marking and VGNN relationship reasoning, a prioritized material handling order is achieved, reducing production line waiting time by more than 30%. Real-time updates to the scene map and path planning also address dynamic obstacles such as conveyor belts and overhead cranes, increasing task completion rate from 60% to over 90% compared to other methods. This embodiment, through simple sketch abstraction, cross-modal learning, and dynamic reasoning, solves the problems of difficult identification, chaotic planning, and slow interaction in building material handling in noisy industrial environments, providing a highly robust solution for intelligent manufacturing scenarios.

[0097] The robot material handling method disclosed in this embodiment obtains a simplified sketch of the target shape drawn by the operator and the location information of the destination. The shape features of the target are extracted from the sketch. Environmental data within the current task scene is collected, and visual features are extracted from the environmental data. A contrastive loss function is used to align the visual features in the environmental data with the shape features in the sketch, retaining the target area that matches the target shape in the sketch. The shape features of the target are extracted from the target area, and the target pose and material type are calculated from the environmental data based on the target area. Finally, based on the target object type, a preset control information database is queried to obtain corresponding motion constraint information. The motion constraint information, target pose, and destination location information are combined to generate a material handling command, which then controls each actuator to move and complete the material handling operation. This effectively solves the technical problems in existing industrial robot material handling technology, such as the easy misinterpretation of voice commands due to the strong noise environment in the industry, and the cumbersome description of complex work objects or processes by inputting commands through pure text. It enables material handling in noisy industrial scenarios by using only visual commands, without the need to input complex coordinates or commands, thereby improving the robot's environmental adaptability and operational reliability in complex industrial environments, reducing the reliance on voice interaction, and improving the automation efficiency of material handling.

[0098] In another embodiment, the processing of the operator's sketch in the above steps can also be performed directly by the trained model to output a sequence of action instructions that the robot can execute, as follows.

[0099] The operator's sketches and the collected environmental data (real-world images) of the current task scenario are input into the trained dual-branch visual encoder model to generate basic motion sequence instructions for controlling the robot to complete material handling operations.

[0100] As shown in the appendix Figure 4As shown, the dual-branch visual encoder model includes a line drawing encoder, a real image encoder, a feature embedding layer, a cross-modal comparison module, and an action generation network module. The line drawing encoder extracts features from the input line drawing to obtain a first feature vector containing the target shape features. The real image encoder extracts features from the input real scene image to obtain a second feature vector containing texture and color information. The first and second feature vectors have the same dimension. The feature embedding layer normalizes the first and second feature vectors to eliminate feature vector length differences and maps them to a cross-modal semantic space through nonlinear transformation. The cross-modal comparison module performs comparative learning on the embedded features, forcing line drawing features of similar objects to be close to real image features in the embedding space, while keeping features of dissimilar objects far apart, generating cross-modal aligned fused features. The action generation network module decodes the fused features through multiple layers to generate basic action sequence instructions.

[0101] Specifically, the real image encoder can use image coding models such as ViT-B / 16, SigLIP, or DINO, taking a real image as input and outputting a feature vector I containing texture, color, etc. The line drawing encoder can use the same image encoder as the real image encoder, such as ViT-B / 16, SigLIP, or DINO, taking a line drawing image as input and outputting a shape feature vector S, which focuses on extracting abstract features such as contours and geometric centers. Since the same encoder is used, S and I have the same dimensions, facilitating subsequent cross-modal comparative learning.

[0102] Feature Embedding Layer: The feature embedding layer consists of a normalization layer and a nonlinear neural network mapping layer. First, L2 normalization is performed on the real image feature vector I and the sketch feature vector S output by the dual-branch encoder to eliminate the influence of feature vector length differences. The normalization calculation method is as follows:

[0103]

[0104] Subsequently, a shared fully connected layer, such as a single-layer MLP, performs a nonlinear transformation, mapping the data to a cross-modal semantic space while maintaining the same output dimension as the input. Here, the feature embedding layer unifies the feature space by normalizing and performing nonlinear transformations, ensuring that features from different modalities, including textured real-image features and shape-based line drawings, reside in the same metric space, facilitating subsequent comparison learning and similarity calculation. Furthermore, it enhances semantic alignment, highlighting cross-modal "shape-instance" associations, such as the contour semantics of a rectangular line drawing and a cardboard box image, while suppressing background noise, such as complex textures in real-image data. The input to the feature embedding layer is the feature vectors I∈R^768 (real image) and S∈R^768 (line drawing) output from the dual-branch encoder; the output is the embedded feature vectors I'∈R^768 and S'∈R^768, used by the subsequent cross-modal comparison module to calculate the contrast loss. .

[0105] Cross-modal contrast module: Employs a contrastive learning method to learn the target, forcing the simplified line drawing features of similar objects to be close to the features of real images in the embedding space, while keeping the features of dissimilar objects away. The loss function is calculated as follows:

[0106]

[0107] Where τ is a temperature parameter used to control the exponential scaling of feature vector similarity, and its value range is usually (0, 1], I i The features of a single negative sample are the true image features.

[0108] Action generation network module: Input fused features S=F+I, decode through multiple Transformer layers to generate action sequences such as ("Move to (x,y,z)→Adjust gripper angle θ→Grab"), output as action sequences.

[0109] In this embodiment, semantic alignment can be achieved by constructing a marker-object association graph GNN through cross-modal graph structure modeling, as detailed below:

[0110] Define nodes: Object nodes: O i ={shape, size, position}, such as the aspect ratio of a rectangular object, or the number of shelves; Action node: A j ={type,params}, such as the gripper type and movement path speed for the "grab" action; Attached node: C k ={constraint,priority}, such as the "handle with care" constraint and the "①" priority marker.

[0111] Edge relationship construction:

[0112] Object-Action Edge: Connections are established based on shape matching degree (e.g., the cosine similarity between a cylindrical object and the "ring gripper" action is ≥0.8), with the weight being the action suitability score;

[0113] Object-Attached Edges: Numerical priority is directly mapped to node attributes, and color constraints are associated with physical parameters through table lookup (e.g., red → gripping force ≤ 5N).

[0114] Run the structured instruction generation algorithm again:

[0115] Based on the GNN output, operation instructions are generated through a Transformer encoder-decoder architecture: First, graph feature encoding is performed, encoding node attributes and edge weights into feature vectors. A multi-head attention mechanism is used to capture cross-node dependencies (e.g., the spatial relationship between "second shelf" and "move to z=150cm"). Then, instruction template matching is performed. A predefined instruction template library (e.g., "move [object] to [location] [action parameters] [constraints]") is used, and a pointer network selects templates and fills in the parameters.

[0116] Example: Move rectangular object #1 to the second shelf, priority 1, handle with care.

[0117] {

[0118] "object": "rectangle①",

[0119] "target": "Second shelf level (z=150cm)",

[0120] "action": ["parallel gripper grasp", "movement (speed=0.1m / s)"],

[0121] "constraint": ["force≤5N", "priority=1"]

[0122] }

[0123] Perform physical feasibility verification: Verify whether the joint angles and paths in the instructions are within the robot's workspace using an inverse kinematics model. If they exceed the limits, trigger template adjustments such as changing the gripper type or splitting the task.

[0124] Adaptive adjustment of identifiers in dynamic scenes:

[0125] Multi-identifier conflict resolution: When the same object is associated with multiple action identifiers (such as "→" translation and "‡" stacking), they are sorted by priority and executed sequentially in the output screen according to priority from left to right, from top to bottom, or number identifiers are greater than symbol identifiers, etc.

[0126] Fuzzy symbol completion: For incomplete symbols such as unclosed arrows, the CLIP model is used to generate the most likely semantic completion, such as "→" being completed as "move to". The semantic consistency before and after completion is verified through comparative learning, such as setting a similarity threshold greater than 0.9 before execution.

[0127] Through the above steps, end-to-end parsing from simple line drawing symbols to operation instructions is achieved, ensuring a symbol recognition accuracy of 98% and an instruction generation time of <200ms, significantly improving the reliability of industrial robots in executing complex instructions.

[0128] In this embodiment, training the dual-branch visual encoder model may include the following:

[0129] Specifically, in the pre-training phase, the aforementioned dual-branch encoder can be pre-trained using an open-source image-text pair dataset to establish a basic association between shape and semantics. Then, model fine-tuning is performed, preferably using both supervised learning and reinforcement learning simultaneously. In supervised learning, the input is a real image-drawing pair, and the output is a standardized action label, optimized using cross-entropy loss.

[0130] Where C is the number of action categories, Tag it as a real action. To predict probabilities.

[0131] In reinforcement learning, the Proximal Policy Optimization (PPO) algorithm is used to optimize action sequences, and the reward function can be designed as follows:

[0132] Additionally, a constraint violation penalty term can be added to the reward function.

[0133] In this embodiment, the acquisition of the cross-modal dataset for model training can be performed through the following steps.

[0134] Step S201: The real material image and the line drawing are encoded by a dual-channel input network respectively. The real image channel is processed by a neural network and connected to a dilated convolutional layer to extract multi-scale geometric features, while the line drawing channel is processed by a simple convolutional neural network to extract shape features.

[0135] Step S202: Perform cross-modal fusion on real material images and line drawing images, align multi-scale geometric features and shape features through contrastive learning and construct an association matrix, and assign the most similar line drawing category to each image;

[0136] Step S203: Construct a graph structure, define the image and sketch pair as object nodes, the action template as action nodes, and connect the object nodes and action nodes by executing actions. The edge weights are determined by the reinforcement learning reward value.

[0137] Step S204: Automatic generation of motion templates. For each material type of the sketch, an initial motion template is generated through the rule library. Objects of the same shape share the same motion template parameters by adjusting only the variable. The kinematic constraints of the industrial robot are integrated through the motion generator.

[0138] Step S205: Input the real image into the graph neural network to automatically match the material type of the line drawing, extract the corresponding optimized action template from the graph action node, generate a triplet annotation consisting of image, line drawing and action, and store it in the cross-modal dataset for model training.

[0139] Specifically, the automatic generation process of the action template in step S204 may include the following:

[0140] First, initial motion generation is performed, starting with zero-sample initialization. For each sketch category, an initial motion template is generated based on rules. For example, for rectangles: the grippers close parallel (opening = object width × 0.8), and the gripping center coordinates are the geometric center of the object; for cylinders: the grippers close in a ring (radius = object radius + 1cm), and the gripping height is at half the object height. For objects of the same shape, such as rectangles of different sizes, the motion template parameters are shared through the VGNN model, and only variable parameters such as gripper opening and gripping height are adjusted.

[0141] The motion template generator can be a multi-layer transformer-based neural network with embedded physical constraints. Industrial robot kinematic constraints, such as the maximum gripper opening, are integrated into the motion generator, and the Sigmoid function is used at the end of the network.

[0142]

[0143] Here, x is the input value, and the function output value ranges from (0,1). This function maps the probability of illegal actions of industrial robot movements to the interval (0,1), thereby implementing soft constraints on physical constraints, such as setting the probability of actions exceeding the joint angle range to a value close to 0. This filters illegal actions and ensures that the generated action template is physically feasible.

[0144] The methods for integrating industrial robot kinematic constraints into the motion generator are as follows:

[0145] First, the constraint parameters are encoded, such as the robot's kinematic constraints like the joint angle range: \left [{{θ}_{min},{θ}_{max}} \right ] Maximum opening of the gripper Load limit Encoded as a constraint vector C = \left [ {{c}_{1}{,c}_{2},...,{c}_{n}} \right ] Each element corresponds to a one-dimensional constraint parameter, such as the normalized upper bound of the joint angle. C is mapped to a constraint feature vector through a learnable neural network embedding layer. This is concatenated with the action features input from the aforementioned Transformer.

[0146] A conditional attention mechanism is introduced into the Transformer decoder, that is, when calculating self-attention, the constraint features are included. With current action characteristics Perform dot product interaction: This mechanism forces the alignment of motion features with constraint features, ensuring that the generated motion parameters, such as joint angles and gripper openings, contain implicit constraint information.

[0147] Add a constraint projection layer after the Transformer output layer to map the motion parameters to the physical feasible space through linear transformation:

[0148] ;

[0149] The Clip function truncates its output based on the upper and lower limits of the constraint vector C, such as the gripper opening. Ensure that g∈[0,G max ].

[0150] Through the above steps, kinematic constraints are integrated into the entire process of input encoding, attention calculation, and output mapping of Transformer, achieving end-to-end constraint awareness for action generation.

[0151] In another embodiment, a robot is also disclosed, including a body on which a controller, a memory, and an optical depth sensing device are mounted. The memory stores a computer program executable by the processor, wherein the processor is configured to execute the computer program in the memory to implement the steps of the robot material handling method as described in any of the preceding embodiments. The optical depth sensing device mounted on the robot can be a lidar, a depth camera, or multiple two-dimensional cameras capable of simultaneously capturing images of the same target.

[0152] In another embodiment, if the above-described robot material handling method is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, all or part of the processes in the above-described embodiments can also be implemented by a computer program instructing related hardware. The computer program can be stored in a computer-readable storage medium, and when executed by a processor, it can implement the steps of the various robot material handling method embodiments described above. The computer program includes computer program code, which can be in the form of source code, object code, executable files, or certain intermediate forms. The computer-readable medium can include: any entity or device capable of carrying the computer program code, a recording medium, a USB flash drive, a portable hard drive, a magnetic disk, an optical disk, a computer memory, a read-only memory, a random access memory, an electrical carrier signal, a telecommunication signal, and a software distribution medium, etc.

[0153] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

[0154] In summary, the above description is only a preferred embodiment of the present invention. All equivalent changes and modifications made within the scope of the claims of the present invention should be covered by the present invention.

Claims

1. A robot material handling method, characterized in that, include: Obtain a simple line drawing depicting the shape of the task target drawn by the operator and the location information of the transport destination, and extract the shape features of the task target from the line drawing; Environmental data within the current task scenario is collected, and visual features are extracted from the environmental data. A contrastive loss function is used to align the visual features in the environmental data with shape features in a simple line drawing, retaining the target region that matches the target shape in the line drawing. The shape features of the task target are extracted from the target region, and the target pose and material type are calculated from the environmental data based on the target region. Convolutional Neural Networks (CNNs) are then used to extract the target shape features from the line drawing. Visual features in real images By comparing losses Forced Towards Alignment, wherein the contrastive loss function for: , The shape feature vector of the i-th simple drawing has dimension d and is extracted and normalized to a unit vector by CNN. Represents the visual feature vector in the i-th environmental data, and Same dimension and normalized; N represents the batch size, i.e. the number of sample pairs processed at the same time; This is a temperature parameter used to adjust the sharpness of the characteristic distribution, with a value range of (0.1, 0.5). Based on the target object type, query the preset control information database to obtain the corresponding motion constraint information, combine the motion constraint information, target pose and destination location information to generate a material handling instruction; control each actuator to move according to the material handling instruction to complete the material handling operation.

2. The robot material handling method according to claim 1, characterized in that, Obtain a simplified sketch of the target object drawn by the operator and the location information of the transport destination, including: Collect simple line drawings of the target object drawn by the operator and text information describing the destination location. Extract the target object's shape features from the line drawings and extract the destination location information of the current transport task from the text information. The destination location information includes platform attribute information and coordinate information of the destination where the target object is stored; or The system collects a simple sketch of the target object drawn by the operator and a voice message describing the location of the transport destination. The shape features of the target object are extracted from the sketch, and the destination location information of the current transport task is extracted from the voice message. The destination location information includes platform attribute information and coordinate information of the destination where the target object is transported and stored.

3. The robot material handling method according to claim 1, characterized in that, Obtain a simplified sketch of the target object drawn by the operator and the location information of the transport destination. Extract the shape features of the target object from the simplified sketch, including: The operator collects a sketch, which includes a first pattern depicting the shape of the target and a second pattern depicting the shape of the storage point where the target is located. Extract the target shape features of the work target from the first pattern, and extract the platform shape features of the storage point where the transport destination is located from the second pattern.

4. The robot material handling method according to claim 3, characterized in that, Extracting the shape features of the task target from the simplified drawing specifically includes: The simplified drawing is subjected to pattern recognition and differentiation. Based on the relative position or connection mark of each segmented independent pattern in the simplified drawing, the first pattern and the second pattern are distinguished according to the preset simplified drawing drawing rules.

5. The robot material handling method according to any one of claims 1-4, characterized in that, Extracting the shape features of the target from the target area, and calculating the target pose and confirming the material type based on the target area from the environmental data, further includes: The system acquires images of the target object using its onboard camera, extracts the object's edges using a contour detection algorithm, and calculates the object's three-dimensional dimensions by combining the camera calibration parameters. The target density value is obtained by querying a preset density database based on the target material type, and the weight of the target object is calculated based on the three-dimensional dimensions of the object and the target density value. Based on the weight and three-dimensional dimensions of the target object, determine whether it exceeds the current equipment load range; if it exceeds the current equipment load range and the target object can be split, then execute the task splitting process to generate a sub-task sequence for moving multiple small objects; otherwise, abandon the current moving task and issue a prompt.

6. The robot material handling method according to claim 5, characterized in that, Based on the target material type, the corresponding target density value is obtained by querying the preset density database. Specifically, this includes: If the target material type cannot be confirmed or a unique target density value cannot be matched, the image texture features on the target object image are identified, and the image texture features are matched with materials in a preset density database. The image texture features include roughness information, and the matched material density value is used as the current target density value.

7. The robot material handling method according to claim 6, characterized in that, Also includes: If an obstacle is detected on the preset path during the movement to the destination, the movement path is replanned based on the shape characteristics of the target to avoid the obstacle.

8. A robot, characterized in that, The device includes a body on which a controller, a memory, and an optical depth sensing device are mounted, the memory being used to store a computer program executable by the processor, wherein the processor is configured to execute the computer program in the memory to implement the method as described in any one of claims 1-7.

9. The robot according to claim 8, characterized in that: The optical depth sensing device includes a lidar, a depth camera, or multiple two-dimensional cameras capable of simultaneously capturing the same target.

10. A computer-readable storage medium, characterized in that, When the executable computer program in the storage medium is executed by a processor, it can implement the method as described in any one of claims 1-7.

Citation Information

Patent Citations

  • Material carrying and moving composite robot

    CN109202885A

  • Visual mechanical arm grabbing method and device applied to parameterized parts

    CN111251295A