A control instruction generation method, an interaction device and an industrial robot

By combining operator-drawn sketches with optical depth sensing devices to generate industrial robot control instructions, the problems of cumbersome text input and difficult voice recognition in complex tasks are solved, thus improving the reliability of task execution.

CN120680515BActive Publication Date: 2026-02-2458 INTELLIGENT TECH (HANGZHOU) CO LTD
View PDF 3 Cites 0 Cited by

Patent Information

Application Number
CN202510971246.1
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-15
Publication Date
2026-02-24
Estimated Expiration
2045-07-15

AI Technical Summary

Technical Problem

Existing industrial robot control instructions require complex text descriptions for complex work objects or processes, operators need to have a certain level of training, and speech recognition is difficult in noisy industrial environments, resulting in low reliability of task execution.

Method used

The operator draws a simple sketch containing the shape and motion constraints of the target, and then uses pattern recognition to extract the target shape and motion constraint information. Combined with the collection of environmental data by an optical depth sensing device, control commands that can be executed by the industrial robot are generated.

Benefits of technology

It lowers the barrier to entry for operators, is suitable for noisy environments, and improves the reliability of industrial robots in complex tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120680515B_ABST
    Figure CN120680515B_ABST
Patent Text Reader

Abstract

The control instruction generation method, the interactive device and the industrial robot disclosed by the application extract a work target shape feature from a first pattern and generate scene constraint information according to a second pattern by collecting a sketch drawn by an operator, the sketch containing the first pattern for depicting the shape of the work target and the second pattern for constraining an execution action; the target position information and the target object type are obtained by matching the visual feature with the target shape feature through collecting environmental data in a current task scene, extracting a visual feature from the environmental data and matching the visual feature with the target shape feature; finally, the corresponding action constraint information is obtained based on the target object type, and the action constraint information, the target position information and the scene constraint information are combined and converted into an executable control instruction of the industrial robot. Therefore, the input of the control instruction can not depend on complex text input or voice recognition, and is suitable for the input of the control instruction in a complex environment such as strong noise or sound obstruction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of data processing, and more particularly to a method for generating control commands, an interactive device, and an industrial robot. Background Technology

[0002] Existing industrial robot control commands typically generate understandable instructions by recognizing and analyzing voice or text commands issued by the operator. However, pure text commands are complex to describe complex work objects or processes, cumbersome to input, and require operators to have a certain level of training. While using deep learning models to process user voice commands to generate industrial robot control commands has significant limitations in its application environment, it struggles to recognize voice commands in noisy industrial environments, easily leading to misjudgments and severely impacting the reliability of industrial robot task execution. Summary of the Invention

[0003] This invention addresses the shortcomings of existing technologies by providing a method for generating control commands, comprising:

[0004] Collect simple sketches drawn by the operator, the sketches containing a first pattern for depicting the shape of the work target and a second pattern for constraining the execution actions;

[0005] The simplified drawing is subjected to pattern recognition, the shape features of the task target are extracted from the first pattern, and scene constraint information is generated based on the second pattern;

[0006] The system collects environmental data within the current task scene using an optical depth sensing device, extracts visual features from the environmental data, and matches the visual features with target shape features to obtain target location information and target object type.

[0007] Based on the target object type, obtain the corresponding motion constraint information, and combine the motion constraint information, target position information, and scene constraint information to convert them into executable control commands for the industrial robot.

[0008] Preferably, the simple sketches drawn by the operator are collected, specifically including: simple sketches manually drawn by the operator on a physical or electronic medium and captured by a camera, or simple sketches manually drawn by the operator on a control screen connected to the current industrial robot.

[0009] Preferably, environmental data within the current task scenario is collected using an optical depth sensing device, specifically including:

[0010] Multiple cameras are used to acquire 2D images of the task scene in real time. Key points of objects are detected in the 2D images, and the 3D pose of each object is calculated by combining camera intrinsic parameters and prior dimensions. The 3D pose is then matched with target shape features to obtain target location information and target object type; or

[0011] Point cloud data within the task scene is acquired in real time using a depth camera or LiDAR. The three-dimensional poses of each object are identified and obtained from the point cloud data. The three-dimensional poses are then matched with the target shape features to obtain the target location information and the target object type.

[0012] Preferably, the control command generation method further includes: the sketch also includes a third pattern for depicting the shape of the target transport destination, the sketch is subjected to pattern recognition and differentiation, and the platform shape features of the destination for transporting and storing the operation target are extracted from the third pattern.

[0013] Preferably, the method involves acquiring environmental data within the current task scene using an optical depth sensing device, extracting visual features from the environmental data, and matching the visual features with target shape features to obtain target location information and target object type. The method further includes:

[0014] The environmental data of the current task scene is collected by an optical depth sensing device, and various visual features are extracted from the environmental data.

[0015] The visual features are matched and identified with the target shape features, and a target mask is generated based on the matching result. The target location information and target object type are obtained based on the target mask. The visual features are matched and identified with the platform shape features, and a platform mask is generated based on the matching result. The platform location information is obtained based on the platform mask.

[0016] Preferably, the second pattern includes a first type of identifier for identifying the grasping scenario constraint information of the work target and a second type of identifier for identifying the operation scenario constraint information of the work target, wherein the industrial robot is configured to perform an operation action after completing the grasping action of the work target.

[0017] Preferably, the second type of identifier also includes constraint information for identifying the transport, transfer, or storage location of the work target.

[0018] Preferably, pattern recognition is performed on the simplified drawing to extract the shape features of the task target from the first pattern, and scene constraint information is generated based on the second pattern, specifically including:

[0019] Pattern recognition is performed on the simplified drawing, and the first and second types of identifiers are distinguished based on the relative positions of each second pattern on the simplified drawing with respect to the first and third patterns.

[0020] First scenario constraint information is generated based on the first type of identifier to constrain and control the grasping action of the work target, and second scenario constraint information is generated based on the second type of identifier to constrain and control the handling and transfer action or handling and storage location of the work target.

[0021] The present invention also discloses an interactive device including a touch screen capable of drawing simplified diagrams, a controller, and a memory for storing a computer program executable by the processor, wherein the processor is configured to execute the computer program in the memory to implement a control instruction generation method as described above.

[0022] The present invention also discloses an industrial robot, including a body on which a controller, a memory, and an optical depth sensing device are mounted. The memory is used to store a computer program executable by the processor, wherein the processor is configured to execute the computer program in the memory to implement a control instruction generation method as described above.

[0023] The control command generation method, interactive device, and industrial robot provided by this invention collect a simple line drawing drawn by the operator, which includes a first pattern of the target shape and a second pattern of motion constraints. The target shape features and scene constraint information are extracted through pattern recognition, and visual feature matching is performed using environmental data collected by an optical depth sensing device to generate control commands executable by the industrial robot. This eliminates the need for complex text input or reliance on voice recognition when inputting control commands to the industrial robot, lowering the barrier to entry for operators, and making it particularly suitable for harsh environments such as high noise levels. It effectively solves the problems of existing technologies where pure text input commands are cumbersome for describing complex work objects or processes and require operators to have a certain level of training; and voice input commands are difficult to recognize in complex environments such as industrial high noise or sound obstruction, easily leading to misjudgments and affecting the reliability of the industrial robot's task execution.

[0024] Additional aspects and advantages of the invention will be set forth in part in the description which follows, and in part will be obvious from the description, or may be learned by practice of the invention. Attached Figure Description

[0025] The accompanying drawings, which are included to provide a further understanding of the invention and form part of this application, illustrate exemplary embodiments of the invention and, together with their description, serve to explain the invention and do not constitute an undue limitation thereof. In the drawings:

[0026] Figure 1This is a schematic diagram illustrating the specific process of a control instruction generation method disclosed in an embodiment of the present invention.

[0027] Figure 2 This is a simplified line drawing illustration of an embodiment of the present invention.

[0028] Figure 3 This is a schematic diagram of the structure of a dual-branch visual encoder model disclosed in an embodiment of the present invention.

[0029] Figure 4 This is a schematic diagram of the specific structure of an interactive device disclosed in an embodiment of the present invention. Detailed Implementation

[0030] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the described embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0031] Unless otherwise defined, the technical or scientific terms used herein shall have the ordinary meaning as understood by one of ordinary skill in the art to which this invention pertains. The terms “first,” “second,” and similar terms used in the specification and claims of this patent application do not indicate any order, quantity, or importance, but are merely used to distinguish different components. Similarly, the terms “an” or “a” and similar terms do not indicate a limitation of quantity, but rather indicate the presence of at least one.

[0032] In this embodiment, a control command generation method is also disclosed, which can be used to generate control commands for various industrial robots to execute, as detailed in the appendix. Figure 1 As shown, it can include the following content.

[0033] Step S1: Collect the sketch drawn by the operator. The sketch includes a first pattern for depicting the shape of the work target and a second pattern for constraining the execution action.

[0034] The industrial robot can capture images of simple sketches drawn manually by the operator on physical or electronic media via a camera. For example, the operator can use paper and pen to draw a simple sketch of the object to be moved or the material to be placed, and then show it to the camera mounted on the industrial robot. The industrial robot can then recognize the sketch and input it into the robot for processing. Alternatively, the operator can also manually draw a simple sketch directly on the control screen connected to the industrial robot via wired or wireless means, or on the touch screen built into the industrial robot. That is, the operator can draw a simple sketch of the object to be moved or the material to be placed via the touch screen, which is connected to the industrial robot, and thus input the sketch into the robot for processing.

[0035] Step S2: Perform pattern recognition on the sketch, extract the target shape features from the first pattern, and generate scene constraint information based on the second pattern.

[0036] In this embodiment, the second pattern includes a first type of identifier for identifying the constraint information of the grasping scenario of the work target, and a second type of identifier for identifying the constraint information of the operation scenario of the work target, wherein the industrial robot is configured to perform an operation action after completing the grasping action of the work target.

[0037] In this embodiment, the sketch may further include a third pattern depicting the shape of the target transport destination. Pattern recognition and differentiation are performed on the sketch, and the platform shape features serving as the destination for transporting and storing the target are extracted from the third pattern. In this case, the second type of identifier may also include constraint information for identifying the transport, transfer, or storage location of the target.

[0038] For example, when drawing simple sketches to indicate handling tasks, the sketches of materials such as cardboard boxes, gears, and cylindrical workpieces serve as the first sketch depicting the shape of the target. The sketches of objects where materials are placed, such as work platforms, assembly lines, and shelves, serve as the third sketch depicting the shape of the target handling destination. Additionally, special symbols, such as the arrow symbol "→", can be designed to indicate both. This symbol represents a translation or movement handling operation, with the starting point of the arrow symbol being the object to be handled and the ending point being the object to be placed. These special symbols serve as the second sketch used to constrain the execution of the action.

[0039] In this embodiment, a second pattern identifier library can also be pre-configured as needed to configure and store the specific information of each second pattern identifier for subsequent simplified diagram recognition. These second pattern identifiers may include the following:

[0040] For situations where there are different types of materials to be grabbed or moved, you can label each type of material with a simple line drawing, such as "①", "②", or "③", to indicate the order of handling.

[0041] For situations where there are multiple similar materials to be moved, Arabic numerals such as "1", "2", and "3" can be marked next to the simple line drawing of the materials to indicate the number of materials to be moved; and special marking symbols can be designed to indicate that all materials to be moved need to be moved, such as "@".

[0042] In situations where there are multiple similar materials to be moved or objects to be placed, to identify the actual materials to be moved or objects to be placed, all similar materials to be moved or objects to be placed can be drawn, and the identified materials to be moved or objects to be placed can be specially marked, such as by marking them with a special symbol. ".

[0043] For scenarios where multiple materials to be transported need to be stacked sequentially, a stacking task symbol can be defined, such as the symbol "‡".

[0044] For scenarios requiring the sequential laying of multiple materials to be moved, a laying task symbol can be defined, such as the symbol " ".

[0045] In one specific embodiment, as shown in the appendix Figure 2 Example of a simple line drawing:

[0046] Target object:

[0047] A simple line drawing represents a steel pipe, marked with "①" for highest priority and "@" for all. A simple rectangular line drawing represents a wooden board, marked with "②" for second priority and "@" for all.

[0048] Location for moving and storing:

[0049] Draw a simple sketch of a storage table with the words "Temporary Storage Area" on it, pointing towards the entrance of the production line. Draw a "‡" to indicate that the target objects to be transported are stacked in sequence.

[0050] The instruction means: "Prioritize moving all steel pipes No. 1 to the production line temporary storage area for stacking, and then move all wooden boards No. 2 to the production line temporary storage area for stacking."

[0051] In this embodiment, pattern recognition is performed on the simplified drawings. The first type of identifier (for identifying the constraint information of the grasping scenario of the work target) and the second type of identifier (for identifying the constraint information of the operation scenario of the work target) can be distinguished based on the relative positions of each second pattern on the simplified drawings with the first and third patterns. For example, it can be defined that the simplified drawings of the materials to be moved are uniformly drawn on the left side of the drawing or panel, and the objects to be placed are uniformly drawn on the right side of the drawing or panel. Then, based on the positions of the different independent patterns on the drawing or panel, the first and third patterns can be determined. The second patterns identified on or near the first pattern are identified as the first type of identifier for identifying the constraint information of the grasping scenario of the work target. The second patterns identified on or near the third pattern are identified as the second type of identifier for identifying the constraint information of the operation scenario of the work target.

[0052] Furthermore, based on the first type of identifier, first scenario constraint information is generated for constraining and controlling the grasping action of the work target; based on the second type of identifier, second scenario constraint information is generated for constraining and controlling the handling and transfer action or handling and storage location of the work target. After identifying the first type of identifier, the constraint information represented by the first type of identifier is combined with the work target to generate the first scenario constraint information for constraining and controlling the grasping action of the work target. The constraint information represented by the identified second type of identifier is combined with the handling and transfer action or handling and storage location to form the second scenario constraint information for constraining and controlling the handling and transfer action or handling and storage location of the work target.

[0053] Step S3: Collect environmental data within the current task scene using an optical depth sensing device, extract visual features from the environmental data, and match the visual features with target shape features to obtain target location information and target object type.

[0054] In this embodiment, multiple cameras mounted on an industrial robot can be used to acquire two-dimensional images of the task scene in real time. Key points of objects are detected in the two-dimensional images, and the three-dimensional poses of each object are calculated by combining camera intrinsic parameters and prior dimensions. The three-dimensional poses are then matched with target shape features to obtain target location information and target object type. Alternatively, a depth camera or LiDAR mounted on the industrial robot can be used to acquire point cloud data of the task scene in real time. The three-dimensional poses of each object are identified and obtained from the point cloud data, and the three-dimensional poses are then matched with target shape features to obtain target location information and target object type.

[0055] If the first type of identifier includes a filtering identifier, the relative position information of the object pattern attached to it as the target object is obtained within the first pattern composed of multiple objects. Based on the target position information of the identified and matched first pattern, the true position information of the target object is calculated. For example, in the case of multiple similar materials to be transported or material placement objects, in order to mark the actual materials to be transported or material placement objects, all similar materials to be transported or material placement objects can be drawn, and the determined materials to be transported or material placement objects can be specially marked, such as by marking them with a special symbol. ".

[0056] In this embodiment, step S3 may further include: acquiring environmental data within the current task scene using an optical depth sensing device, and extracting various visual features from the environmental data; performing target matching and recognition between the visual features and target shape features, and generating a target mask based on the matching result; obtaining target location information and target object type based on the target mask; performing platform matching and recognition between the visual features and platform shape features, and generating a platform mask based on the matching result; obtaining platform location information based on the platform mask.

[0057] Step S4: Obtain the corresponding motion constraint information based on the target object type, and combine the motion constraint information, target position information, and scene constraint information into executable control instructions for the industrial robot.

[0058] The control command generation method provided in this embodiment collects a simple sketch drawn by the operator, containing a first pattern of the target's shape and a second pattern of motion constraints. The target's shape features and scene constraint information are extracted through pattern recognition. Visual feature matching is then performed using environmental data collected by an optical depth sensing device to generate control commands executable by the industrial robot. This eliminates the need for complex text input or reliance on voice recognition when inputting control commands to the industrial robot, lowering the barrier to entry for operators, and making it particularly suitable for harsh environments such as high noise levels. It effectively solves the problems in existing technologies where pure text input commands are cumbersome for describing complex work objects or processes and require operators to have a certain level of training; and where voice input commands are difficult to recognize in complex environments such as industrial high noise or sound obstruction, easily leading to misjudgments and affecting the reliability of the industrial robot's task execution.

[0059] In another embodiment, the processing of the operator's sketch in the above steps can also be performed directly by the trained model to output a sequence of action instructions that can be executed by the industrial robot, as follows.

[0060] The operator's sketches and the collected environmental data (real-world images) of the current task scenario are input into the trained dual-branch visual encoder model to generate basic motion sequence instructions for controlling the industrial robot to complete material handling operations.

[0061] As shown in the appendix Figure 3 As shown, the dual-branch visual encoder model includes a line drawing encoder, a real image encoder, a feature embedding layer, a cross-modal comparison module, and an action generation network module. The line drawing encoder extracts features from the input line drawing to obtain a first feature vector containing the target shape features. The real image encoder extracts features from the input real scene image to obtain a second feature vector containing texture and color information. The first and second feature vectors have the same dimension. The feature embedding layer normalizes the first and second feature vectors to eliminate feature vector length differences and maps them to a cross-modal semantic space through nonlinear transformation. The cross-modal comparison module performs comparative learning on the embedded features, forcing line drawing features of similar objects to be close to real image features in the embedding space, while keeping features of dissimilar objects far apart, generating cross-modal aligned fused features. The action generation network module decodes the fused features through multiple layers to generate basic action sequence instructions.

[0062] Specifically, the real image encoder can use image coding models such as ViT-B / 16, SigLIP, or DINO, taking a real image as input and outputting a feature vector I containing texture, color, etc. The line drawing encoder can use the same image encoder as the real image encoder, such as ViT-B / 16, SigLIP, or DINO, taking a line drawing image as input and outputting a shape feature vector S, which focuses on extracting abstract features such as contours and geometric centers. Since the same encoder is used, S and I have the same dimensions, facilitating subsequent cross-modal comparative learning.

[0063] Feature Embedding Layer: The feature embedding layer consists of a normalization layer and a nonlinear neural network mapping layer. First, L2 normalization is performed on the real image feature vector I and the sketch feature vector S output by the dual-branch encoder to eliminate the influence of feature vector length differences. The normalization calculation method is as follows:

[0064] .

[0065] Subsequently, a shared fully connected layer, such as a single-layer MLP, performs a nonlinear transformation, mapping the data to a cross-modal semantic space while maintaining the same output dimension as the input. Here, the feature embedding layer unifies the feature space by normalizing and performing nonlinear transformations, ensuring that features from different modalities, including textured real-image features and shape-based line drawings, reside in the same metric space, facilitating subsequent comparison learning and similarity calculation. Furthermore, it enhances semantic alignment, highlighting cross-modal "shape-instance" associations, such as the contour semantics of a rectangular line drawing and a cardboard box image, while suppressing background noise, such as complex textures in real-image data. The input to the feature embedding layer is the feature vectors I∈R^768 (real image) and S∈R^768 (line drawing) output from the dual-branch encoder; the output is the embedded feature vectors I'∈R^768 and S'∈R^768, used by the subsequent cross-modal comparison module to calculate the contrast loss. .

[0066] Cross-modal contrast module: Employs a contrastive learning method to learn the target, forcing the simplified line drawing features of similar objects to be close to the features of real images in the embedding space, while keeping the features of dissimilar objects away. The loss function is calculated as follows:

[0067] ,

[0068] in, This is a temperature parameter used to control the exponential scaling of feature vector similarity; its value typically ranges from (0, 1). The features of a single negative sample are the true image features.

[0069] Action generation network module: Input fused features The action sequence is generated by decoding through multiple layers of Transformer, such as "move to (x,y,z) → adjust gripper angle θ → grab", and the output is an action sequence.

[0070] In this embodiment, semantic alignment can be achieved by constructing a marker-object association graph GNN through cross-modal graph structure modeling, as detailed below:

[0071] Define nodes: Object nodes: Examples include the aspect ratio of a rectangular object and the number of shelves; action nodes: Such as the gripper type and movement path speed for the "grab" action; additional nodes: Such as the "handle with care" constraint and the "①" priority marker.

[0072] Edge relationship construction:

[0073] Object-Action Edge: Connections are established based on shape matching degree (e.g., the cosine similarity between a cylindrical object and the "ring gripper" action is ≥0.8), with the weight being the action suitability score;

[0074] Object-Attached Edges: Numerical priority is directly mapped to node attributes, and color constraints are associated with physical parameters through table lookup (e.g., red → gripping force ≤ 5N).

[0075] Run the structured instruction generation algorithm again:

[0076] Based on the GNN output, operation instructions are generated through a Transformer encoder-decoder architecture: First, graph feature encoding is performed, encoding node attributes and edge weights into feature vectors. A multi-head attention mechanism is used to capture cross-node dependencies (e.g., the spatial association between "second shelf" and "move to z=150cm"). Then, instruction template matching is performed. A predefined instruction template library (e.g., "move [object] to [location] [action parameters] [constraints]") is used, and a pointer network selects the template and fills in the parameters.

[0077] Example: Move rectangular object #1 to the second shelf, priority 1, handle with care.

[0078] {

[0079] "object": "rectangle①",

[0080] "target": "Second shelf level (z=150cm)",

[0081] "action": ["parallel gripper grasp", "movement (speed=0.1m / s)"],

[0082] "constraint": ["force≤5N", "priority=1"]

[0083] }

[0084] Perform physical feasibility verification: Verify whether the joint angles and paths in the instructions are within the workspace of the industrial robot using an inverse kinematics model. If they exceed the limits, trigger template adjustment (such as changing the gripper type or splitting the task).

[0085] Adaptive adjustment of identifiers in dynamic scenes:

[0086] Multi-identifier conflict resolution: When the same object is associated with multiple action identifiers (such as "→" translation and "‡" stacking), they are sorted by priority and executed sequentially in the output screen according to priority from left to right, from top to bottom, or number identifiers are greater than symbol identifiers, etc.

[0087] Fuzzy symbol completion: For incomplete symbols such as unclosed arrows, the CLIP model is used to generate the most likely semantic completion, such as "→" being completed as "move to". The semantic consistency before and after completion is verified through comparative learning, such as setting a similarity threshold greater than 0.9 before execution.

[0088] Through the above steps, end-to-end parsing from simple line drawing symbols to operation instructions is achieved, ensuring a symbol recognition accuracy of 98% and an instruction generation time of <200ms, significantly improving the reliability of industrial robots in executing complex instructions.

[0089] In this embodiment, for cases where the background of the material to be transported or the object to which the material is placed is particularly complex, the background decoupling recognition method based on simple line drawings can be used to improve the recognition success rate of the material to be transported or the object to which the material is placed in such cases.

[0090] Specifically, by using the shape purity of simple line drawings as a prior constraint, cross-modal contrastive learning is used to align the target features of complex background images with the shape features of simple line drawings, thereby achieving accurate recognition under background interference.

[0091] First, cross-modal encoding is performed between the simplified line drawing and the real image. For the simplified line drawing (e.g., a rectangular line drawing) and the real image against a complex background, a convolutional neural network (CNN) is used to extract shape features, respectively. With visual features By comparing losses Forced Towards Alignment is used to filter out background interference. The design incorporates contrast loss. The calculation formula is as follows:

[0092] .

[0093] In this formula, The shape feature vector representing the i-th simple drawing, such as the outline feature of a rectangle, has a dimension of d and is extracted and normalized to a unit vector by CNN. The visual feature vector representing the i-th complex background image, and... Same dimension and normalized; N represents the batch size, i.e. the number of sample pairs processed at the same time; The temperature parameter is used to adjust the sharpness of the feature distribution, and its value range is typically (0.1, 0.5). The core logic of designing this contrastive loss function is that positive samples... Representing simplified line drawings and real images of the same type of material, such as a rectangular simplified line drawing and an image of a cardboard box with a complex background, requires that the features of the two be as similar as possible in the embedding space; negative sample pairs To represent different types of materials, their characteristics are forced to be far apart. The optimization objective is to minimize... Make the feature dot product of similar samples It is significantly larger than negative samples, thus achieving target feature focusing under background interference.

[0094] Then, dynamic mask generation is performed. Based on the aligned features, a U-Net network is used to generate a mask for the target region, retaining only pixels in areas that overlap with the shape of the sketch, such as within the outline of a rectangle, while masking background clutter.

[0095] Finally, self-supervised feature enhancement is performed by randomly adding more complex disturbances, such as lighting changes and occlusions, to the complex background image. The training model recovers the target features using simple line drawing priors, improving robustness. This step forces focus on the target by using simple line drawing shape anchors, enabling accurate material location based on shape priors even in complex backgrounds such as stacking, reflections, and occlusions.

[0096] In this embodiment, training the dual-branch visual encoder model may include the following:

[0097] Specifically, in the pre-training phase, the aforementioned dual-branch encoder can be pre-trained using an open-source image-text pair dataset to establish a basic association between shape and semantics. Then, model fine-tuning is performed, preferably using both supervised learning and reinforcement learning simultaneously. In supervised learning, the input is a real image-drawing pair, and the output is a standardized action label, optimized using cross-entropy loss. Where C is the number of action categories, Tag it as real action. To predict probabilities.

[0098] In reinforcement learning, the Proximal Policy Optimization (PPO) algorithm is used to optimize action sequences, and the reward function can be designed as follows: Additionally, a constraint violation penalty term can be added to the reward function.

[0099] In this embodiment, the acquisition of the cross-modal dataset for model training can be performed through the following steps.

[0100] Step S101: The real material image and the line drawing are encoded by a dual-channel input network respectively. The real image channel is processed by a neural network and connected to a dilated convolutional layer to extract multi-scale geometric features, while the line drawing channel is processed by a simple convolutional neural network to extract shape features.

[0101] Step S102: Perform cross-modal fusion on real material images and line drawing images, align multi-scale geometric features and shape features through contrastive learning and construct an association matrix, and assign the most similar line drawing category to each image;

[0102] Step S103: Construct a graph structure, define the image and sketch pair as object nodes, the action template as action nodes, and connect the object nodes and action nodes by executing actions. The edge weights are determined by the reinforcement learning reward value.

[0103] Step S104: Automatic generation of motion templates. For each material type of the sketch, an initial motion template is generated through the rule library. Objects of the same shape share the same motion template parameters by adjusting only the variable. The kinematic constraints of the industrial robot are integrated through the motion generator.

[0104] Step S105: Input the real image into the graph neural network to automatically match the material type of the line drawing, extract the corresponding optimized action template from the graph action node, generate a triplet annotation consisting of image, line drawing and action, and store it in the cross-modal dataset for model training.

[0105] Specifically, the automatic generation process of the action template in step S104 may include the following:

[0106] First, initial motion generation is performed, starting with zero-sample initialization. For each sketch category, an initial motion template is generated based on rules. For example, for rectangles: the grippers close parallel (opening = object width × 0.8), and the gripping center coordinates are the geometric center of the object; for cylinders: the grippers close in a ring (radius = object radius + 1cm), and the gripping height is at half the object height. For objects of the same shape, such as rectangles of different sizes, the motion template parameters are shared through the VGNN model, and only variable parameters such as gripper opening and gripping height are adjusted.

[0107] The motion template generator can be a multi-layer transformer-based neural network with embedded physical constraints. Industrial robot kinematic constraints, such as the maximum gripper opening, are integrated into the motion generator, and the Sigmoid function is used at the end of the network.

[0108] ,

[0109] Here, x is the input value, and the function output value ranges from (0,1). This function maps the probability of illegal actions of industrial robot movements to the interval (0,1), thereby implementing soft constraints on physical constraints, such as setting the probability of actions exceeding the joint angle range to a value close to 0. This filters illegal actions and ensures that the generated action template is physically feasible.

[0110] The methods for integrating industrial robot kinematic constraints into the motion generator are as follows:

[0111] First, constrain the parameter encoding, and then encode the kinematic constraints of the industrial robot, such as the range of joint angles. Maximum opening of the gripper Load limit Encoded as constraint vectors Each element corresponds to a one-dimensional constraint parameter, such as the normalized upper bound of the joint angle. This is achieved through a learnable neural network embedding layer. Mapped to constrained feature vectors This is concatenated with the action features input from the aforementioned Transformer.

[0112] A conditional attention mechanism is introduced into the Transformer decoder, that is, when calculating self-attention, the constraint features are included. With current action characteristics Perform dot product interaction: This mechanism forces the alignment of motion features with constraint features, ensuring that the generated motion parameters, such as joint angles and gripper openings, contain implicit constraint information.

[0113] Add a constraint projection layer after the Transformer output layer to map the motion parameters to the physical feasible space through linear transformation: ;in, The function is based on the constraint vector Upper and lower limit truncation output, such as gripper opening. ,make sure .

[0114] Through the above steps, kinematic constraints are integrated into the entire process of input encoding, attention calculation, and output mapping of Transformer, achieving end-to-end constraint awareness for action generation.

[0115] In another embodiment, as shown in the appendix Figure 4 As shown, an interactive device is also disclosed, including a touch screen 1 capable of drawing simplified diagrams, a controller 2, and a memory 3, wherein the memory is used to store a computer program executable by the processor, and wherein the processor is configured to execute the computer program in the memory to implement the control instruction generation method disclosed in the foregoing embodiments.

[0116] In another embodiment, an industrial robot is also disclosed, comprising a body on which a controller, a memory, and an optical depth sensing device are mounted. The memory is used to store a computer program executable by the processor, wherein the processor is configured to execute the computer program in the memory to implement the control instruction generation method disclosed in the foregoing embodiments.

[0117] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention, and not to limit them; although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features; and these modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of the present invention.

[0118] In summary, the above description is only a preferred embodiment of the present invention. All equivalent changes and modifications made within the scope of the claims of the present invention should be covered by the present invention.

Claims

1. A method for generating control commands, characterized in that, include: Collect simple sketches drawn by the operator, the sketches containing a first pattern for depicting the shape of the work target and a second pattern for constraining the execution actions; The simplified drawing is subjected to pattern recognition, the shape features of the task target are extracted from the first pattern, and scene constraint information is generated based on the second pattern; The system collects environmental data within the current task scene using an optical depth sensing device, extracts visual features from the environmental data, and matches the visual features with target shape features to obtain target location information and target object type. Based on the target object type, obtain the corresponding motion constraint information, and combine the motion constraint information, target position information and scene constraint information to convert them into executable control instructions for the industrial robot. The simplified drawing and the collected environmental data within the current task scene are input into the trained dual-branch visual encoder model to generate basic motion sequence instructions for controlling the industrial robot to complete material handling operations. The acquisition of the cross-modal dataset used for training the dual-branch visual encoder model is performed through the following steps: Step S101: The real material image and the line drawing are encoded by a dual-channel input network respectively. The real image channel is processed by a neural network and connected to a dilated convolutional layer to extract multi-scale geometric features, while the line drawing channel is processed by a simple convolutional neural network to extract shape features. Step S102: Perform cross-modal fusion on real material images and line drawing images, align multi-scale geometric features and shape features through contrastive learning and construct an association matrix, and assign the most similar line drawing category to each image; Step S103: Construct a graph structure, define the image and sketch pair as object nodes, the action template as action nodes, and connect the object nodes and action nodes by executing actions. The edge weights are determined by the reinforcement learning reward value. Step S104: Automatic generation of motion templates. For each material type of the sketch, an initial motion template is generated through the rule library. Objects of the same shape share the same motion template parameters by adjusting only the variable. The kinematic constraints of the industrial robot are integrated through the motion generator. Step S105: Input the real image into the graph neural network to automatically match the material type of the line drawing, extract the corresponding optimized action template from the graph action node, generate a triplet annotation consisting of image, line drawing and action, and store it in the cross-modal dataset for model training.

2. The control command generation method according to claim 1, characterized in that, The simple sketches drawn by the data acquisition operator include: The data is captured by a camera, showing simple sketches drawn manually by the operator on physical or electronic media. The image shows a simple sketch manually drawn by the operator on the control screen connected to the current industrial robot.

3. The control command generation method according to claim 2, characterized in that, Environmental data within the current task scenario is collected using an optical depth sensing device, specifically including: Multiple cameras are used to acquire 2D images of the task scene in real time. Key points of objects are detected in the 2D images, and the 3D pose of each object is calculated by combining camera intrinsic parameters and prior dimensions. The 3D pose is then matched with target shape features to obtain target location information and target object type; or Point cloud data within the task scene is acquired in real time using a depth camera or LiDAR. The three-dimensional poses of each object are identified and obtained from the point cloud data. The three-dimensional poses are then matched with the target shape features to obtain the target location information and the target object type.

4. The control command generation method according to claim 3, characterized in that, Also includes: The sketch also includes a third pattern for depicting the shape of the target transport destination. The sketch is then subjected to pattern recognition and differentiation, and the platform shape features that serve as the destination for transporting and storing the target are extracted from the third pattern.

5. The control command generation method according to claim 4, characterized in that, The system acquires environmental data within the current task scene using an optical depth sensing device, extracts visual features from the environmental data, and matches these visual features with target shape features to obtain target location information and target object type. The system also includes: The environmental data of the current task scene is collected by an optical depth sensing device, and various visual features are extracted from the environmental data. The visual features are matched and identified with the target shape features, and a target mask is generated based on the matching result. The target location information and target object type are obtained based on the target mask. The visual features are matched and identified with the platform shape features, and a platform mask is generated based on the matching result. The platform location information is obtained based on the platform mask.

6. The control command generation method according to claim 5, characterized in that: The second pattern includes a first type of identifier for identifying the constraint information of the grasping scenario of the work target, and a second type of identifier for identifying the constraint information of the operation scenario of the work target, wherein the industrial robot is configured to perform an operation action after completing the grasping action of the work target.

7. The control command generation method according to claim 6, characterized in that: The second type of identifier also includes constraint information for identifying the location of the transport or storage of the work target.

8. The control command generation method according to claim 7, characterized in that, The simplified drawing is subjected to pattern recognition. The shape features of the task target are extracted from the first pattern, and scene constraint information is generated based on the second pattern. Specifically, this includes: Pattern recognition is performed on the simplified drawing, and the first and second types of identifiers are distinguished based on the relative positions of each second pattern on the simplified drawing with respect to the first and third patterns. First scenario constraint information is generated based on the first type of identifier to constrain and control the grasping action of the work target, and second scenario constraint information is generated based on the second type of identifier to constrain and control the handling and transfer action or handling and storage location of the work target.

9. An interactive device, characterized in that: The device includes a touch screen capable of drawing diagrams, a controller, and a memory for storing a computer program executable by the controller, wherein the controller is configured to execute the computer program in the memory to implement the method as described in any one of claims 1-8.

10. An industrial robot, characterized in that: The device includes a body on which a controller, a memory, and an optical depth sensing device are mounted. The memory stores a computer program executable by the controller, wherein the controller is configured to execute the computer program in the memory to implement the method as described in any one of claims 1-8.

Citation Information

Patent Citations

  • Control method and device for network access equipment

    CN111190351A

  • Universal intelligent agent and control method thereof

    CN118261192A

  • Interactive robot grabbing method based on large language model

    CN120206528A