Multi-modal interaction method and interaction system applied to intelligent robot
By employing a collaborative approach between the Gemini Robotics-ER model and the vision-language-motion model, deep correlation and real-time decision-making of multimodal information were achieved, solving the accuracy and real-time issues in multimodal interaction of industrial robots and improving robot interaction efficiency.
Patent Information
- Application Number
- CN202511341194.6
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-09-19
- Publication Date
- 2025-11-04
- Estimated Expiration
- Not applicable · inactive patent
AI Technical Summary
Existing industrial robots lack multimodal interaction capabilities, exhibiting issues such as insufficient accuracy in multimodal feature fusion and poor real-time interaction response, making it difficult to meet the demands for high precision and high adaptability.
By employing a collaborative approach combining the Gemini Robotics-ER model and the vision-language-action model, and through multispectral visual acquisition, feature extraction, multimodal attention mechanisms, and real-time feedback control, we achieve deep association of multimodal information and rapid decision-making.
It improves the accuracy of multimodal information fusion and the real-time performance of interactive response, reduces the deviation in object recognition and the misjudgment of action intent, and ensures that the robot can adjust its actions in a timely manner to adapt to environmental changes, meeting the needs of high-precision and high-flexibility industrial interaction.
Smart Images

Figure CN120886265A_ABST
Abstract
Description
TECHNICAL FIELD
[0001] The present application relates to the field of industrial robot interaction, and in particular to a multi-modal interaction method and system applied to an intelligent robot. BACKGROUND
[0002] With the rapid evolution of industrial automation towards intelligence and flexibility, intelligent robots need to process multi-source heterogeneous information such as vision, language, and action in production and manufacturing scenarios to achieve natural interaction with the environment and humans. The current interaction capability of industrial robots is limited to a single mode or simple mode combination, which is difficult to meet the high precision and high adaptability requirements in complex assembly and flexible handling scenarios. The fusion application of Gemini Robotics-ER model and vision-language-action model provides a technical possibility to break through the multi-modal information barrier, and through the construction of a perception, decision, and execution closed loop, it promotes the upgrade of industrial robots from preset program execution to autonomous interaction decision-making, and becomes a key technology direction to improve production efficiency and flexibility.
[0003] The prior art has two significant shortcomings in multi-modal interaction. First, the accuracy of multi-modal feature fusion is insufficient, and the geometric features of visual information and the semantic features of language instructions lack a deep correlation mechanism, often resulting in operation object recognition deviation or action intention misjudgment, leading to robot execution inconsistent with instructions. Second, the real-time performance of interaction response is poor. In dynamic scenarios, the serial processing mode of visual information processing, language instruction analysis, and action planning will produce cumulative delay, and when the environment or instructions change suddenly, the robot is difficult to quickly adjust the action trajectory, which easily leads to collision risk or operation failure, and cannot meet the real-time interaction requirements of high-paced industrial production. SUMMARY
[0004] In order to overcome the shortcomings and deficiencies of the prior art, the present application provides a multi-modal interaction method and system applied to an intelligent robot.
[0005] The technical solution adopted by the present application is a multi-modal interaction method applied to an intelligent robot, comprising the following steps:
[0006] Step S1, through the multi-spectral vision acquisition device carried by the industrial robot, a dynamic image sequence of the interaction scene is captured, the robot joint angle, end effector pose, and environmental light intensity parameters at the image acquisition time are recorded synchronously, and the dynamic image sequence is decomposed into three-dimensional point cloud data containing depth information and two-dimensional texture feature matrix;
[0007] Step S2, calling the perception layer interface of the Gemini Robotics-ER model, feature extraction is performed on the three-dimensional point cloud data to generate a structured description vector containing geometric topological relations, and multi-scale convolution operations are performed on the two-dimensional texture feature matrix to output a high-dimensional feature tensor containing color distribution, texture gradient and edge contour;
[0008] Step S3, receiving an externally input natural language instruction, performing word segmentation processing through the language analysis module of the vision-language-action model to generate a semantic dependency tree composed of verbs, nouns and spatial relationship words, and performing entity mapping on the semantic dependency tree in combination with a preset industry term library to convert it into a machine recognizable instruction sequence containing action types, operation objects and spatial coordinates;
[0009] Step S4, inputting the structured description vector, high-dimensional feature tensor and machine recognizable instruction sequence into the decision layer of the Gemini Robotics-ER model, performing feature fusion on the three types of data through a multi-modal attention mechanism to generate an interaction decision matrix containing time sequence information, each element in the interaction decision matrix corresponding to joint motion parameters and end effector operation parameters of the industrial robot at a calibration time;
[0010] Step S5, based on the interaction decision matrix, performing inverse kinematics solving through the action planning module of the vision-language-action model to generate trajectory planning data containing joint angular displacement sequence, motion velocity curve and acceleration threshold, and synchronously calculating robot dynamics parameters at each time in the trajectory execution process;
[0011] Step S6, transmitting the trajectory planning data and dynamics parameters to the motion control unit of the industrial robot to drive each joint of the robot to perform corresponding actions according to the trajectory planning data, and simultaneously capturing scene change images in the action execution process through a vision acquisition device and feeding back to the perception layer of the Gemini Robotics-ER model to trigger the next cycle of interaction loop.
[0012] Further, in step S2, when the Gemini Robotics-ER model performs feature extraction on the three-dimensional point cloud data, a feature descriptor generation method based on spatial neighborhood search is adopted, a local geometric feature matrix is constructed by calculating the normal vector angle, Euclidean distance and curvature change rate of each sampling point and its k nearest neighbors in the point cloud, and the local geometric feature matrix is enhanced through the following formula:
[0013]
[0014] wherein F geowhere n is the total number of sampling points in the point cloud, k is the number of neighboring points of each sampling point, θ ij is the normal vector angle between the i th sampling point and the j th neighboring point, d ij is the Euclidean distance between the i th sampling point and the j th neighboring point, d max is the maximum Euclidean distance between all pairs of points in the point cloud, κ ij is the curvature change rate between the i th sampling point and the j th neighboring point, α, β, and γ are weight coefficients of the normal vector angle, the Euclidean distance, and the curvature change rate, respectively, V ij is the unit vector from the i th sampling point to the j th neighboring point.
[0015] In step S2, when performing multi-scale convolution operation on the two-dimensional texture feature matrix, three different size convolution kernels are set to extract features, and the feature maps output by each convolution kernel are fused by the following formula:
[0016]
[0017] where T tex is the fused texture feature tensor, m is the convolution kernel number, λ m is the weight coefficient of the feature output by the m th convolution kernel, T in is the input two-dimensional texture feature matrix, K m is the parameter matrix of the m th convolution kernel, s m is the step parameter of the m th convolution kernel, and Conv(·) represents convolution operation.
[0018] Further, in step S3, when the language analysis module of the visual-language-action model performs word segmentation processing on the natural language instruction, a word segmentation algorithm based on bidirectional long short-term memory network is adopted to determine the word segmentation boundary by calculating the probability distribution of each character in the context, and generate a word sequence containing part-of-speech tagging; in the construction process of the semantic dependency tree, the subject-predicate, verb-object, and bias-positive relationships are determined by calculating the dependency probability between words, and the dependency probability is calculated by the following formula:
[0019]
[0020] where P dep (w i , w j , r) is the probability of the existence of the dependency relationship r between the words w i and w j , W r is the weight matrix corresponding to the dependency relationship r, h i and h j are the words w i and w jword vector representation of the word, Hadamard product of the representation vectors, R is the set of all possible dependency relations;
[0021] The entity mapping process in step S3 is achieved by calculating the semantic similarity between the terms in the industry term library and the semantics of the nodes in the semantic dependency tree, and the semantic similarity is calculated by the following formula:
[0022]
[0023] Where S(w, t) is the semantic similarity between the word w and the term t, w k , t k are the projection values of the word w and the term t in the k-dimensional semantic space, d is the dimension of the semantic space, δ(w, t) is the field correlation coefficient, which is 1 when the word w and the term t belong to the same industry field, otherwise it is 0.3.
[0024] Further, in step S4, the multi-modal attention mechanism determines the attention weight by calculating the mutual information entropy between the structured description vector, the high-dimensional feature tensor and the machine recognizable instruction sequence, and the mutual information entropy is calculated by the following formula:
[0025]
[0026] Where I(X, Y) is the mutual information entropy of modal X and modal Y, P(x, y) is the joint probability distribution of x and y, P(x), P(y) are the marginal probability distribution of x, y respectively;
[0027] In the generation process of the interaction decision matrix, the fusion features are calculated by the following formula:
[0028] D t,j =σ(W d ·[F geo,t ;T tex,t ;L cmd,t ]+b d )
[0029] Where D t,j is the output value of the interaction decision matrix at time t and the jth decision dimension, σ is the sigmoid activation function, W d is the decision layer weight matrix, F geo,t is the structured description vector at time t, T tex,t is the high-dimensional feature tensor at time t, L cmd,t is the machine recognizable instruction sequence feature at time t, b d is the decision layer bias term.
[0030] Further, in step S5, the inverse kinematics solving process is achieved by iterative calculation, and in each iteration process, the joint angle is corrected by the following formula:
[0031] Δθ = (J T J+λI) -1 J T ΔX
[0032] Wherein, Δθ is the joint angle correction vector, J is the robot Jacobian matrix, λ is the damping coefficient, I is the unit matrix, and ΔX is the end effector pose error vector;
[0033] The dynamics parameter calculation includes the calculation of joint driving torque, which is achieved by the following formula:
[0034]
[0035] Wherein, τ is the joint driving torque vector, M(θ) is the inertia matrix, is the joint angular acceleration vector, is the Coriolis force and centrifugal force matrix, is the joint angular velocity vector, and G(θ) is the gravity term vector.
[0036] Further, in step S6, the driving process of the robot joint by the motion control unit is achieved by using the model predictive control method, and the control amount is calculated by the following formula:
[0037]
[0038] Wherein, u(t) is the control amount at time t, N is the prediction time domain, x(t+k|t) is the system state predicted based on the state at time t at time t+k, x r (t+k) is the reference state at time t+k, Q and R are the weight matrices of state error and control amount, respectively;
[0039] In the feedback process of the scene change image, the difference degree of the current image and the last period image is calculated by the following formula:
[0040]
[0041] Wherein, D img is the image difference degree, H and W are the height and width of the image, respectively, I c (i,j) is the gray value of the current image at pixel point (i,j), and I p (i,j) is the gray value of the last period image at pixel point (i,j). When D img is greater than the preset threshold, the next period of interaction loop is triggered.
[0042] Further, step S3 comprises the following sub-steps: S31, after receiving the externally input natural language instruction, converting the instruction text into a UTF-8 encoded character stream, performing sentence segmentation according to a preset punctuation symbol set to obtain a plurality of clauses, and marking the positions of numbers, letters and special symbols in each clause through character-level traversal; S32, calling a word segmentation interface of the visual-language-action model to perform word boundary detection on the segmented clauses, determining word segmentation points based on a bidirectional maximum matching algorithm, and generating a word segmentation result list containing word sequence numbers, word lengths and part-of-speech tags, wherein the part-of-speech tags include verbs, nouns, adjectives, adverbs and prepositions; S33, constructing a semantic dependency tree with a verb as a root node, determining the connection relationship between parent nodes and child nodes by calculating the co-occurrence frequency between words, taking nouns as action object nodes, adverbs as action modification nodes, and prepositions as spatial relationship nodes, and forming a tree diagram with a hierarchical structure; S34, performing string matching between each node of the semantic dependency tree and a term in a preset industry term library, calculating the machine code corresponding to the term with a matching degree exceeding a set threshold, replacing the text content of the corresponding node in the semantic dependency tree, and generating an identifiable instruction sequence containing machine code, node type and connection relationship.
[0043] Further, step S4 comprises the following sub-steps: S41, aligning the dimensions of the structured description vector output in step S2 and the high-dimensional feature tensor, keeping the time sequence lengths of the two consistent through zero padding, converting the machine recognizable instruction sequence into a feature matrix with the same time sequence length as the instruction sequence, and each time step corresponding to an operation instruction; S42, initializing the weight parameter matrix of the multi-modal attention mechanism, setting the number of attention heads and the feature dimension, and inputting the structured description vector, the high-dimensional feature tensor and the machine recognizable instruction feature matrix into different attention sub-layers; S43, each attention sub-layer performs linear transformation on the input features to generate query vectors, key vectors and value vectors, determines attention scores by calculating the dot product of the query vectors and the key vectors, and after normalization by the softmax function, multiplies the value vectors to obtain single-modal attention features; S44, splicing the single-modal attention features output by each attention sub-layer, performing feature fusion through a fully connected layer to generate an interaction decision matrix containing a time dimension, and each row of the matrix corresponding to a decision parameter at a time step and each column corresponding to a decision dimension.
[0044] Further, the step S5 comprises the following sub-steps: S51, extracting target poses of the end effector at each time step from the interaction decision matrix, including position coordinates and attitude angles, arranging the target poses in time sequence to form a pose sequence; S52, performing smoothing processing on the pose sequence, calculating transition points between adjacent target poses through a cubic spline interpolation algorithm to keep the pose change rate continuous, and generating a dense pose trajectory containing more intermediate points; S53, based on the dense pose trajectory, calculating angle values of each joint at each time step through inverse kinematics to form a joint angle sequence, performing velocity and acceleration constraint checking on the joint angle sequence, and adjusting the joint angle values of the corresponding time steps when the preset upper limit of velocity or acceleration is exceeded; S54, calculating angular velocity and angular acceleration of each joint at each time step according to the adjusted joint angle sequence, and combining inertia parameters, mass distribution and friction coefficient of the robot to determine required driving torque and power consumption of each joint, and generating trajectory planning data containing joint angle, velocity, acceleration and torque.
[0045] A multi-modal interaction system applied to an intelligent robot, comprising:
[0046] A multi-spectral vision information acquisition and preprocessing unit, an input end of which is connected with a vision sensor group of the industrial robot, and an output end of which is connected with a perception layer unit of the Gemini Robotics-ER model through a data bus, for converting collected image data into three-dimensional point cloud and two-dimensional texture features;
[0047] A Gemini Robotics-ER model processing unit, comprising a perception layer interface, a feature extraction module and a decision layer module, the perception layer interface being connected with the output end of the multi-spectral vision information acquisition and preprocessing unit, and the output end of the decision layer module being connected with an action planning unit of the vision-language-action model, for performing extraction and fusion decision of multi-modal features;
[0048] A natural language instruction analysis unit, an input end of which receives external language instructions, and an output end of which is connected with a semantic mapping unit of the vision-language-action model, for converting natural language into a semantic dependency tree;
[0049] A vision-language-action collaborative processing unit, connected with the Gemini Robotics-ER model processing unit, the natural language instruction analysis unit and the trajectory planning unit respectively, for performing associated mapping of language instructions and vision features;
[0050] A trajectory planning and dynamics calculation unit, an input end of which is connected with the decision layer module of the Gemini Robotics-ER model processing unit, and an output end of which is connected with a motion control unit of the industrial robot, for generating joint motion trajectory and dynamics parameters;
[0051] A motion control and feedback unit is connected with the trajectory planning and dynamics calculation unit at the input end and drives each joint actuator of the industrial robot at the output end, and is connected with the multispectral vision information acquisition and preprocessing unit through an internal feedback interface to perform closed-loop control of the interactive process.
[0052] Beneficial effects: The application provides a multi-modal interaction method and system applied to an intelligent robot, which effectively improves the interaction efficiency of the industrial robot through the cooperation of the Gemini Robotics-ER model and the vision-language-action model. In the multi-modal information fusion, the structured processing of the visual data by the perception layer of the Gemini Robotics-ER model is combined with the semantic analysis of the language instruction by the vision-language-action model, and the feature depth correlation is realized through the multi-modal attention mechanism, so that the visual features and the language instruction are accurately associated, and the problem of insufficient multi-modal fusion accuracy in the prior art is solved, and the operation object recognition deviation and the action intention misjudgment are reduced. In terms of interaction real-time performance, a perception, decision and execution closed loop is constructed, the decision layer of the Gemini Robotics-ER model quickly generates an interaction decision matrix, the action planning module of the vision-language-action model efficiently completes trajectory planning, and the motion control unit drives the execution in real time, while the visual feedback triggers the next cycle in real time, solves the problem of dynamic scene response lag, ensures that the robot adjusts the action in time to adapt to the environmental changes, improves the action execution precision and stability, and meets the high-precision and high-flexibility industrial interaction requirements. BRIEF DESCRIPTION OF DRAWINGS
[0053] Figure 1 The method step flowchart of the application is shown in the figure.
[0054] Figure 2 The system unit composition diagram of the application is shown in the figure. DETAILED DESCRIPTION
[0055] It should be noted that the embodiments in the present application and the features in the embodiments can be combined with each other without conflict, and the present application will be further described in detail in combination with the drawings and specific embodiments.
[0056] As shown in the figure, a multi-modal interaction method applied to an intelligent robot comprises the following steps: Figure 1
[0057] Step S1: A multispectral vision acquisition device carried by an industrial robot captures a dynamic image sequence of an interactive scene, synchronously records the robot joint angle, end effector pose and environmental light intensity parameters at the image acquisition time, and decomposes the dynamic image sequence into three-dimensional point cloud data containing depth information and a two-dimensional texture feature matrix.
[0058] Specifically, this step is the basic perception link of multi-modal interaction. A multispectral vision acquisition device carried by an industrial robot captures a dynamic image sequence, and the robot joint angle, end effector pose, and ambient light intensity parameters are recorded synchronously, providing multi-dimensional raw data support for subsequent processing. The dynamic image sequence is decomposed into three-dimensional point cloud data and two-dimensional texture feature matrix, which can respectively retain the spatial geometric information and surface texture information of the scene, providing structured input for feature extraction of the Gemini Robotics-ER model, ensuring the comprehensiveness and accuracy of the perception data, and being the premise of realizing accurate interaction.
[0059] In specific implementation, a multispectral camera with a resolution of 1920x1080 is used as the vision acquisition device, and the frame rate is set to 30 frames / second, which can cover a spectral range of 400-1000 nm. While capturing the dynamic image sequence, the encoder of the robot control system is used to acquire the joint angles in real time, with an accuracy of ±0.01°. The end effector pose is measured by a laser tracker, with a positional accuracy of ±0.05 mm and an attitude accuracy of ±0.02°. The ambient light intensity is acquired by an integrated illumination sensor, with a range of 0-10000 lux and a sampling frequency consistent with the image acquisition frame rate. Subsequently, a point cloud generation algorithm is used to convert each frame of image into three-dimensional point cloud data containing 50000-100000 points, with a point cloud density of 1000 points per cubic meter. A texture extraction algorithm is used to separate a two-dimensional texture feature matrix from the image, with a matrix dimension of 1024x1024 and each element representing the texture feature value of the corresponding pixel.
[0060] In step S2, the perception layer interface of the Gemini Robotics-ER model is called to extract features from the three-dimensional point cloud data, generating a structured description vector containing geometric topological relationships, and to perform multi-scale convolution operations on the two-dimensional texture feature matrix, outputting a high-dimensional feature tensor containing color distribution, texture gradient, and edge contour.
[0061] Specifically, this step is the core link of multi-modal feature extraction. The perception layer interface of the Gemini Robotics-ER model is called to process the three-dimensional point cloud data and the two-dimensional texture feature matrix, and the generated structured description vector and high-dimensional feature tensor can accurately represent the geometric and texture information of the scene. The geometric topological relationships contained in the structured description vector help the robot understand the spatial layout of objects in the scene, and the color distribution, texture gradient, and edge contour information in the high-dimensional feature tensor provide a basis for object recognition and state judgment. The output of these two features lays the foundation for subsequent multi-modal fusion and directly affects the accuracy of interaction decisions.
[0062] In the implementation process, the perception layer interface of the Gemini Robotics-ER model uses a 512-dimensional feature vector as the output dimension of the structured description vector. When processing three-dimensional point cloud data, the neighborhood search radius is set to 0.05 m, and the geometric topological relationship is generated by calculating the spatial relationship of each point in the point cloud. For the two-dimensional texture feature matrix, multi-scale convolution operation is performed, and the convolution kernel size is 3x3, 5x5 and 7x7, respectively, with a step size of 1. After 3 layers of convolution operation, the output high-dimensional feature tensor has a dimension of 256x256x64. During feature extraction, the model's operation frame rate remains consistent with the image acquisition frame rate, i.e., 30 frames / second, ensuring real-time performance. At the same time, through GPU acceleration processing, the single-frame feature extraction time is controlled within 30 ms, meeting the demand of real-time interaction of industrial robots.
[0063] In step S3, the natural language instruction input from the outside is received, and the language parsing module of the visual-language-action model is used for word segmentation processing to generate a semantic dependency tree composed of verbs, nouns and spatial relationship words. The semantic dependency tree is mapped to entities in combination with a preset industry term library, and is converted into a machine recognizable instruction sequence containing action types, operation objects and spatial coordinates.
[0064] Specifically, this step realizes the conversion of natural language instructions into machine recognizable information, and is a key bridge connecting human intentions and robot actions. The word segmentation processing of natural language instructions generates a semantic dependency tree, which can analyze the grammatical structure and semantic relationship in the instructions. In combination with the industry term library, the abstract language instructions are converted into specific action types, operation objects and spatial coordinates, etc. machine recognizable information, ensuring accurate understanding of human instructions by the robot and avoiding interaction errors caused by semantic ambiguity, providing clear target guidance for subsequent decision planning.
[0065] In specific implementation, the received natural language instruction is first converted into a unified format through character encoding, and the word segmentation processing uses a statistical-based word segmentation algorithm with an accuracy rate of more than 98%. The generated semantic dependency tree contains nodes such as verbs, nouns and spatial relationship words, and the number of nodes is determined by the length of the instruction, generally between 5-20. The industry term library covers common equipment names, action terms and spatial description vocabularies in the industrial field, totaling more than 10,000. In the entity mapping process, the term matching threshold is set to 0.8, and entity replacement is performed when the semantic similarity exceeds the threshold. The converted machine recognizable instruction sequence is stored in JSON format, containing action type fields (such as grabbing, carrying, assembling, etc.), operation object fields (such as part number, equipment name, etc.), and spatial coordinate fields (accurate to millimeter level). The generation time of each instruction sequence is controlled within 50 ms.
[0066] Step S4, input the structured description vector, high-dimensional feature tensor and machine recognizable instruction sequence into the decision layer of the Gemini Robotics-ER model, perform feature fusion on the three types of data through a multi-modal attention mechanism, and generate an interaction decision matrix containing time sequence information, each element in the interaction decision matrix corresponding to joint motion parameters and end effector operation parameters of the industrial robot at a calibration time;
[0067] Specifically, this step is the core link of multi-modal information fusion and decision making. The structured description vector, high-dimensional feature tensor and machine recognizable instruction sequence are input into the decision layer of the Gemini Robotics-ER model, and the multi-modal attention mechanism is used to realize effective fusion of different types of data. The complementarity of each modality information can be fully utilized to improve the comprehensiveness and accuracy of decision making. The generated interaction decision matrix contains joint motion parameters and end effector operation parameters of the robot at different times, which provides specific decision basis for subsequent action planning and ensures the consistency of robot action with scene demand and instruction requirement.
[0068] In the implementation process, the multi-modal attention mechanism sets 8 attention heads, each with a feature dimension of 64. The structured description vector, high-dimensional feature tensor and machine recognizable instruction sequence are weighted and fused. The attention weights in the fusion process are dynamically adjusted according to the correlation of each modality data. The correlation is calculated based on cosine similarity, and the range is between 0 and 1. The generated interaction decision matrix has a dimension of 100x20, where 100 represents the number of time steps, corresponding to an interaction period of 10 seconds (time step length is 0.1 second), and 20 decision dimensions correspond to 6 joint motion parameters (angle, velocity, acceleration) of the industrial robot and 4 operation parameters (position, attitude, jaw opening, force) of the end effector. The model operation of the decision layer adopts batch processing, 10 frames of data are processed per batch, and the processing time is controlled within 20 ms to ensure the real-time performance of the decision output.
[0069] Step S5, based on the interaction decision matrix, inverse kinematics is solved through the action planning module of the vision-language-action model to generate trajectory planning data containing joint angular displacement sequence, motion velocity curve and acceleration threshold, and the robot dynamics parameters at each time in the trajectory execution process are calculated synchronously;
[0070] Specifically, this step is responsible for converting the decision matrix into specific robot action trajectory, which is the key link between decision and execution. Through the inverse kinematics solving of the visual-language-action model's action planning module, the motion trajectory of each joint of the robot can be determined according to the interactive decision matrix, and the trajectory planning data such as joint angle displacement sequence, motion speed curve and acceleration threshold generated ensure the smoothness and safety of the robot motion. The dynamically calculated dynamics parameters provide a mechanical basis for the drive control of the robot, avoiding equipment damage caused by excessive load or improper motion, and ensuring the stability and reliability of the interaction process.
[0071] In specific implementation, the inverse kinematics solving adopts an iterative method, the number of iterations is set to 50, and the accuracy error of each iteration is controlled within 0.01°. The joint angle displacement sequence obtained by solving has a sampling frequency of 100 Hz, i.e. each time step is 0.01 seconds. The motion speed curve adopts a trapezoidal speed curve, the maximum speed is set according to the joint performance of the robot, generally between 30° / s-60° / s, and the acceleration threshold is set to 50° / s 2 , ensuring smooth motion. The dynamics parameter calculation includes joint driving torque and power consumption, the driving torque calculation considers the inertia, friction and load of the joint, and the accuracy reaches ±0.5 N·m, and the power consumption calculation error is controlled within 5%. The trajectory planning data is stored in XML format, containing timestamp, joint number, angular displacement, speed, acceleration and other information, and the generation time of a single trajectory planning data is controlled within 40 ms.
[0072] Step S6, transmitting the trajectory planning data and dynamics parameters to the motion control unit of the industrial robot to drive the joints of the robot to perform corresponding actions according to the trajectory planning data, and simultaneously capturing the scene change images in the action execution process through the visual acquisition device and feeding back to the perception layer of the Gemini Robotics-ER model to trigger the next cycle of interactive loop.
[0073] Specifically, this step is the execution and feedback link of the robot action, which transmits the trajectory planning data and dynamics parameters to the motion control unit to drive the robot to perform corresponding actions, realizing the conversion of decision intention to physical action. At the same time, the visual acquisition device captures the scene changes in real time and feeds back to the perception layer, forming a complete closed-loop control, which can timely discover the deviation in the action execution process and trigger the next cycle of interactive adjustment, ensuring the dynamic adaptability and accuracy of the interaction process, avoiding the failure of interaction caused by environmental changes or execution errors, and is an important guarantee for ensuring the continuous and effective multi-modal interaction.
[0074] In a specific implementation, the motion control unit adopts a servo system based on PID control, the control cycle is 1 ms, the received trajectory planning data is converted into an analog signal after digital-to-analog conversion, the output control signal ranges from 0 to 10 V, and the robot joints are driven to move according to the set trajectory. The position error of joint movement is controlled within ±0.1°, and the speed error is controlled within ±1° / s. The visual acquisition device continuously captures scene images during the action execution process, the frame rate is maintained at 30 frames / s, the image resolution is 1920x1080, the scene changes are detected in real time through an image difference algorithm, and when the pixel proportion of the scene changes exceeds 5%, a feedback mechanism is triggered. The feedback data is transmitted to the perception layer of the Gemini Robotics-ER model through Ethernet, the transmission delay is controlled within 20 ms, the next cycle of interaction is ensured to start in time, and the total delay of the whole action execution and feedback process is controlled within 100 ms.
[0075] Preferably, in step S2, when the Gemini Robotics-ER model extracts features from the three-dimensional point cloud data, a feature descriptor generation method based on spatial neighborhood search is adopted, a local geometric feature matrix is constructed by calculating the normal vector angle, Euclidean distance and curvature change rate of each sampling point and its k-neighbor points in the point cloud, and the local geometric feature matrix is enhanced by the following formula:
[0076]
[0077] wherein F geo is the enhanced local geometric feature vector, n is the total number of sampling points in the point cloud, k is the number of neighbor points of each sampling point, θ ij is the normal vector angle between the i th sampling point and the j th neighbor point, d ij is the Euclidean distance between the i th sampling point and the j th neighbor point, d max is the maximum Euclidean distance of all point pairs in the point cloud, κ ij is the curvature change rate between the i th sampling point and the j th neighbor point, and α, β and γ are weight coefficients of the normal vector angle, Euclidean distance and curvature change rate, respectively, V ij is the unit vector of the i th sampling point pointing to the j th neighbor point.
[0078] When the two-dimensional texture feature matrix is subjected to a multi-scale convolution operation in step S2, three convolution kernels of different sizes are set to extract features, and the feature maps output by the convolution kernels are fused by the following formula:
[0079]
[0080] wherein T tex is the fused texture feature tensor, m is the convolution kernel serial number, and λ mis the weight coefficient of the output feature of the mth convolution kernel, T in is the input two-dimensional texture feature matrix, K m is the parameter matrix of the mth convolution kernel, s m is the step parameter of the mth convolution kernel, and Conv(·) represents the convolution operation.
[0081] Specifically, the accuracy and richness of feature extraction are improved through specific feature enhancement formulas and multi-scale convolution fusion formulas. For feature extraction of three-dimensional point cloud data, a feature descriptor generation method based on spatial neighborhood search is adopted. The local geometric feature matrix is constructed by calculating the normal vector angle, Euclidean distance and curvature rate of change between each sampling point and its k nearest neighbors in the point cloud. The value of k is dynamically adjusted according to the point cloud density, generally set to 10-30, the normal vector angle calculation accuracy is controlled within ±0.5°, the Euclidean distance measurement error is not more than 0.01mm, and the curvature rate of change is calculated by using the second derivative method, with an accuracy of 0.001 / mm. In the feature enhancement formula, α, β, γ are weight coefficients, respectively taking values of 0.4, 0.3, 0.3, and the enhancement of local geometric features is realized by weighted summation, so that the generated structured descriptor vector can better reflect the spatial topological relationship of the point cloud. For the two-dimensional texture feature matrix, when performing multi-scale convolution operation, the convolution kernel size is 3×3, 5×5 and 7×7 respectively, and the step is 1. After each layer of convolution operation, a ReLU activation function is used for nonlinear transformation, and λ1, λ2, λ3 are set to 0.5, 0.3 and 0.2 respectively. The convolution features of different scales are fused by weighted fusion, so that the output high-dimensional feature tensor contains both detailed texture and global contour information. During implementation, GPU parallel processing is used for model operation, and the feature extraction time of single frame three-dimensional point cloud is controlled within 25ms, and the two-dimensional texture feature processing time is controlled within 20ms, which ensures synchronization with the image acquisition frame rate, provides higher quality visual feature input for subsequent multi-modal fusion, and significantly improves the perception accuracy of robot on scene geometric structure and texture features.
[0082] Preferably, in step S3, when the language parsing module of the visual-language-action model performs word segmentation processing on the natural language instruction, a word segmentation algorithm based on bidirectional long short-term memory network is adopted to determine the word segmentation boundary by calculating the probability distribution of each character in the context, and generate a word sequence containing part-of-speech tagging; in the process of constructing the semantic dependency tree, the subject-predicate, verb-object and bias-positive relationships are determined by calculating the dependency probability between words, and the dependency probability is calculated by the following formula:
[0083]
[0084] wherein, P dep (w i , w j , r) is the word wi with w j the probability of the dependency r existing between w r is a weight matrix corresponding to the dependency r, h i , h j are the word vector representations of w i , w j , is the Hadamard product of the vectors, and R is the set of all possible dependencies;
[0085] The entity mapping process in step S3 is achieved by calculating the semantic similarity between the terms in the industry term library and the semantics of the nodes in the semantic dependency tree, and the semantic similarity is calculated by the following formula:
[0086]
[0087] where S(w, t) is the semantic similarity between the word w and the term t, w k , t k are the projection values of the word w and the term t in the k-th dimensional semantic space, d is the dimension of the semantic space, and δ(w, t) is the domain relevance coefficient, which is 1 when the word w and the term t belong to the same industry domain, and 0.3 otherwise.
[0088] Specifically, the accuracy and industry adaptability of language instruction analysis are improved by the related formulas of semantic dependency probability calculation and semantic similarity calculation. In terms of semantic dependency tree construction, a word segmentation algorithm based on bidirectional long short-term memory network is used, and the context probability distribution of each character is calculated during word segmentation. The probability calculation window size is set to 5 characters before and after, and the part-of-speech tagging accuracy is above 97%. In dependency probability calculation, the weight matrix W_r is trained according to the industrial domain corpus, containing 512 neuron nodes, and the word vector dimension is 300. The semantic association between words is strengthened through Hadamard product operation, and the dependency set R covers 15 common relationship types such as subject-predicate, verb-object, and partial-positive, ensuring the integrity of semantic structure analysis. In the entity mapping process, the industry term library is stored in categories such as equipment, action, and space, each term contains 10-20 semantic feature dimensions, the semantic space dimension d is set to 200, the matching between words and terms is realized through cosine similarity calculation, the domain relevance coefficient δ is 1 in the same domain and 0.3 in different domains, and the matching threshold is set to 0.75. When the semantic similarity exceeds the threshold, entity replacement is performed. In implementation, the language analysis module uses CPU and GPU cooperative processing, the word segmentation and semantic dependency tree construction time is controlled within 30ms, the entity mapping time is controlled within 20ms, and the overall high-precision conversion from natural language instructions to machine recognizable sequences is ensured, so that the robot can more accurately understand the meaning of industry-specific instructions and reduce the interaction errors caused by semantic ambiguity.
[0089] Preferably, in step S4, the multi-modal attention mechanism determines the attention weight by calculating the mutual information entropy between the structured description vector, the high-dimensional feature tensor and the machine recognizable instruction sequence, and the mutual information entropy is calculated by the following formula:
[0090]
[0091] Wherein, I(X, Y) is the mutual information entropy of modal X and modal Y, P(x, y) is the joint probability distribution of x and y, P(x), P(y) are the marginal probability distribution of x, y respectively;
[0092] In the generation process of the interaction decision matrix, the fusion feature is calculated by the following formula:
[0093] D t,j =σ(W d ·[F geo,t ;T tex,t ;L cmd,t ]+b d )
[0094] Wherein, D t,j is the output value of the interaction decision matrix at time t, the jth decision dimension, σ is the sigmoid activation function, W d is the decision layer weight matrix, F geo,t is the structured description vector at time t, T tex,t is the high-dimensional feature tensor at time t, L cmd,t is the machine recognizable instruction sequence feature at time t, b d is the decision layer bias term.
[0095] Specifically, the effectiveness of multi-modal feature fusion and the accuracy of decision output are improved through the related formulas of mutual information entropy calculation and decision value calculation. The multi-modal attention mechanism determines the attention weight by calculating the mutual information entropy between the structured description vector, the high-dimensional feature tensor and the machine-recognizable instruction sequence. The mutual information entropy calculation is based on joint probability distribution and marginal probability distribution. The estimation of probability distribution adopts kernel density estimation method, and the bandwidth parameter is set to 0.1 to ensure the smoothness of the probability distribution. The time alignment of the three modal data adopts dynamic time warping algorithm, and the alignment error is controlled within 10ms. The value range of mutual information entropy is 0-5bits, and the larger the value is, the stronger the correlation between the modes is, and the higher the corresponding attention weight is. When the interactive decision matrix is generated, the dimension of the decision layer weight matrix W_d is 20x512, which is obtained by training through the stochastic gradient descent method. The learning rate is set to 0.001, the iteration number is 10000 times, and the bias term b_d is initialized to 0 and adjusted to the optimal value through iterative optimization. The output value is mapped to the 0-1 interval through the sigmoid activation function, and then converted to the actual decision parameter value through linear scaling. The scaling range of the joint motion parameters is set according to the physical limits of the robot joints, such as the joint angle range of-180° to 180°, the speed range of 0-60° / s, and the acceleration range of 0-50° / s 2 The position accuracy in the end effector operation parameter is ±0.05mm, the attitude accuracy is ±0.1°, the gripper opening range is 0-100mm, and the force range is 0-50N. During the implementation, the operation of the decision layer adopts batch processing, each batch processes 20 frames of data, and the single batch processing time is controlled within 15ms to ensure that the interactive decision matrix can reflect the scene changes and instruction requirements in real time, providing accurate decision basis for subsequent motion planning.
[0096] Preferably, in step S5, the inverse kinematics solving process is realized through iterative calculation. In each iteration process, the joint angle is corrected through the following formula:
[0097] Δθ=(J T J+λI) -1 J T ΔX
[0098] Where Δθ is the joint angle correction vector, J is the robot Jacobian matrix, λ is the damping coefficient, I is the unit matrix, and ΔX is the end effector pose error vector.
[0099] The dynamics parameter calculation includes the calculation of joint driving torque, which is realized through the following formula:
[0100]
[0101] Where τ is the joint driving torque vector, M(θ) is the inertia matrix, is the Coriolis and centrifugal force matrix, is the Coriolis and centrifugal force matrix, is the joint angular velocity vector, and G(θ) is the gravity term vector.
[0102] Specifically, the feasibility and safety of the trajectory planning are ensured by iterative correction formula and driving torque calculation. The inverse kinematics solution adopts an iterative calculation method, and the initial joint angle is set as the current actual joint angle of the robot. In each iteration process, the joint angle correction amount is calculated by pseudo-inverse of the Jacobian matrix, and the Jacobian matrix is generated according to the DH parameter modeling of the robot, with the parameter error controlled within ±0.1 mm and ±0.1°. The damping coefficient λ is dynamically adjusted according to the joint motion speed, with the value of 0.01 when the speed is lower than 20° / s and the value of 0.05 when the speed is higher than 20° / s, so as to ensure the stability of the iteration process. The calculation of the end effector pose error vector ΔX is based on the difference between the current pose and the target pose, and the position error and the attitude error are set to be 0.1 mm and 0.5° respectively. The iteration is stopped when the error is less than the threshold, and the iteration number is usually controlled within 20-50 times, with the single-step iteration time not exceeding 1 ms. In the calculation of the dynamic parameters, the inertia matrix M(θ) is calculated according to the mass, mass center position and moment of inertia of each link of the robot, with the mass measurement error not exceeding 5 g, the mass center position error not exceeding 0.5 mm and the moment of inertia error not exceeding 0.001 kg·m 2 ; the Coriolis and centrifugal force matrix C(θ, θ · ) is calculated by the Christoffel symbol, with the calculation accuracy of ±0.1 N·m; the gravity term vector G(θ) is calculated according to the gravitational acceleration (the value is 9.81 m / s 2 ) and the mass center position of each link, with the error controlled within ±0.5 N·m. The calculation of the driving torque vector τ takes into account the friction of the joint, and the friction model adopts a combined model of Coulomb friction and viscous friction, with the Coulomb friction torque value of 0.5-2 N·m and the viscous friction coefficient value of 0.01-0.1 N·m·s / rad. In implementation, the inverse kinematics solution and the dynamic parameter calculation are processed in parallel, and the planning time of a single trajectory is controlled within 30 ms. The generated joint angle displacement sequence, motion speed curve and acceleration threshold data ensure that the joint driving torque of the robot does not exceed 80% of the rated value and the power consumption does not exceed 90% of the rated power during the motion process, avoiding equipment damage caused by overloading, while ensuring the smoothness of the motion trajectory, with the change rates of the speed and the acceleration controlled within 50° / s 2 and 100° / s 3 respectively, so as to meet the operation requirements of high precision and high safety of the industrial robot.
[0103] Preferably, in step S6, the driving process of the robot joint is realized by using a model predictive control method, and the control amount is calculated by the following formula:
[0104]
[0105] wherein u(t) is the control quantity at time t, N is the prediction time domain, x(t+k|t) is the system state at time t+k based on the state prediction at time t, x r (t+j) is the reference state at time t+j, Q and R are the weight matrices of state error and control quantity respectively;
[0106] In the feedback process of the scene change image, the difference degree of the current image and the last period image is calculated by the following formula:
[0107]
[0108] wherein D img is the image difference degree, H and W are the height and width of the image respectively, I c (i,j) is the gray value of the current image at pixel point (i,j), I p (i,j) is the gray value of the last period image at pixel point (i,j). When D img is greater than the preset threshold, the next period of interaction cycle is triggered.
[0109] Specifically, the driving mode of the motion control unit and the feedback mechanism of the scene change image are refined. The precise control of the robot action and the closed-loop adjustment of the interaction process are realized through the model predictive control formula and the image difference degree calculation mechanism. The motion control unit adopts a model predictive control-based mode, the prediction time domain N is set to 10, the control time domain is the same as the prediction time domain, the system state x(t+k|t) is obtained according to the kinematics model of the robot, the model parameters are obtained through the system identification method, the identification error is controlled within 5%, the reference state x_r(t+k) is generated according to the trajectory planning data interpolation, and the time interval is 0.01s. The state error weight matrix Q is a diagonal matrix, and the diagonal elements corresponding to the position, velocity and acceleration weights are 100, 10 and 1 respectively. The control weight matrix R is also a diagonal matrix, and the weight values are set according to the importance of the control amount. The joint angle control weight is 1, and the velocity control weight is 0.1. The solution of the control amount adopts the interior point method, the iteration number is controlled within 10 times, the single-step control amount calculation time is not more than 5ms, and the output control signal is converted into digital-analog after the digital-analog conversion to drive the servo motor. The position loop proportional gain of the motor is set to 5-10, the speed loop proportional gain is set to 0.5-2, and the integral gain is set to 0.1-1, so as to ensure that the motor responds quickly and has no overshoot. The difference degree calculation of the scene change image adopts the average of the sum of the absolute difference of the gray value. The image gray value range is 0-255, and the threshold value of the difference degree D_img is dynamically adjusted according to the ambient light intensity. When the light intensity is lower than 1000lux, the threshold value is set to 0.1, and when the light intensity is higher than 1000lux, the threshold value is set to 0.05. When the difference degree exceeds the threshold value, the next cycle of interaction loop is triggered. The transmission of feedback data adopts UDP protocol, the transmission rate is 100Mbps, and the transmission delay is controlled within 10ms, so as to ensure that the perception layer of Gemini Robotics-ER model can obtain the scene change in the action execution process in time. In the implementation process, the control period of the motion control unit is 1ms, and the sampling period of the visual feedback is 33ms (consistent with the image acquisition frame rate). Through the cooperative work of the two, the high-precision control of the robot action and the real-time closed-loop adjustment of the interaction process are realized, and the reliability and adaptability of the multi-modal interaction are effectively improved.
[0110] Preferably, step S3 comprises the following sub-steps: S31, after receiving the externally input natural language instruction, converting the instruction text into a UTF-8 encoded character stream, performing sentence segmentation according to a preset punctuation symbol set to obtain a plurality of clauses, and marking the positions of numbers, letters and special symbols in each clause through character-level traversal; S32, calling a word segmentation interface of the vision-language-action model to perform word boundary detection on the segmented clauses, determining word segmentation points based on a bidirectional maximum matching algorithm, and generating a word segmentation result list containing word sequence numbers, word lengths and part-of-speech tags, wherein the part-of-speech tags include verbs, nouns, adjectives, adverbs and prepositions; S33, constructing a semantic dependency tree with a verb as a root node, determining the connection relationship between parent nodes and child nodes by calculating the co-occurrence frequency between words, taking nouns as action object nodes, adverbs as action modification nodes, and prepositions as spatial relationship nodes, and forming a tree diagram with a hierarchical structure; S34, performing string matching between each node of the semantic dependency tree and a term in a preset industry term library, calculating the machine code corresponding to the term with a matching degree exceeding a set threshold, replacing the text content of the corresponding node in the semantic dependency tree, and generating an identifiable instruction sequence containing machine code, node type and connection relationship.
[0111] Preferably, step S4 comprises the following sub-steps: S41, aligning the dimensions of the structured description vector output in step S2 and the high-dimensional feature tensor, keeping the time sequence lengths of the two consistent through zero padding, converting the machine recognizable instruction sequence into a feature matrix with the same time sequence length as the instruction sequence, and each time step corresponding to an operation instruction; S42, initializing the weight parameter matrix of the multi-modal attention mechanism, setting the number of attention heads and the feature dimension, and inputting the structured description vector, the high-dimensional feature tensor and the machine recognizable instruction feature matrix into different attention sub-layers; S43, each attention sub-layer performs linear transformation on the input features to generate query vectors, key vectors and value vectors, determines attention scores by calculating the dot product of the query vectors and the key vectors, and after normalization by a softmax function, multiplies the value vectors to obtain single-modal attention features; S44, splicing the single-modal attention features output by each attention sub-layer, performing feature fusion through a fully connected layer to generate an interaction decision matrix containing a time dimension, and each row of the matrix corresponding to a decision parameter at a time step and each column corresponding to a decision dimension.
[0112] Preferably, step S5 comprises the following sub-steps: S51, extracting the target pose of the end effector at each time step from the interaction decision matrix, including position coordinates and attitude angles, arranging the target poses in chronological order to form a pose sequence; S52, smoothing the pose sequence, calculating the transition points between adjacent target poses through a cubic spline interpolation algorithm to keep the pose change rate continuous, and generating a dense pose trajectory containing more intermediate points; S53, based on the dense pose trajectory, calculating the angle value of each joint at each time step through inverse kinematics to form a joint angle sequence, and performing velocity and acceleration constraint checking on the joint angle sequence, and adjusting the joint angle value of the corresponding time step when the preset upper limit of velocity or acceleration is exceeded; S54, according to the adjusted joint angle sequence, calculating the angular velocity and angular acceleration of each joint at each time step, and combining the inertia parameters, mass distribution and friction coefficient of the robot to determine the required driving torque and power consumption of each joint, and generating trajectory planning data containing joint angle, velocity, acceleration and torque.
[0113] As shown in Figure 2 A multi-modal interaction system applied to an intelligent robot, comprising:
[0114] A multi-spectral vision information acquisition and preprocessing unit, the input end of which is connected with a vision sensor group of an industrial robot, and the output end of which is connected with a perception layer unit of a Gemini Robotics-ER model through a data bus, for converting collected image data into three-dimensional point cloud and two-dimensional texture features;
[0115] A Gemini Robotics-ER model processing unit, containing a perception layer interface, a feature extraction module and a decision layer module, the perception layer interface being connected with the output end of the multi-spectral vision information acquisition and preprocessing unit, and the output end of the decision layer module being connected with an action planning unit of a vision-language-action model, for performing extraction and fusion decision of multi-modal features;
[0116] A natural language instruction analysis unit, the input end of which receives external language instructions, and the output end of which is connected with a semantic mapping unit of the vision-language-action model, for converting natural language into a semantic dependency tree;
[0117] A vision-language-action collaborative processing unit, connected with the Gemini Robotics-ER model processing unit, the natural language instruction analysis unit and a trajectory planning unit respectively, for performing associated mapping of language instructions and vision features;
[0118] A trajectory planning and dynamics calculation unit, the input end of which is connected with the decision layer module of the Gemini Robotics-ER model processing unit, and the output end of which is connected with a motion control unit of the industrial robot, for generating joint motion trajectory and dynamics parameters;
[0119] Motion control and feedback unit, input end is connected with trajectory planning and dynamics calculation unit, output end drives each joint actuator of industrial robot, and simultaneously through internal feedback interface, is connected with multispectral vision information acquisition and pretreatment unit, and carries out closed loop control of interactive process.
[0120] The full name of Gemini Robotics-ER in the application is Gemini Robotics-Embodied Reasoning, i.e., Gemini Robotics-Embodied Reasoning model. It was launched by Google Corporation of the United States on March 12, 2025, and is an artificial intelligence model based on Gemini 2.0.
[0121] As a visual-language model, Gemini Robotics-ER has strong spatial understanding ability, and can enable robot experts to run their own programs using the embodied reasoning ability of Gemini. The model focuses on spatial reasoning and environmental dynamic analysis, and can autonomously generate efficient solutions by interfacing with low-level control systems. In practical applications, such as in tasks like packing lunch boxes, the robot can comprehensively judge the placement position and operation sequence of the objects to complete the task. Its functions include object detection, which can identify and track the position and size of objects in 2D and 3D space; pointing function, which can identify objects and elements in objects for interaction; grasp prediction, which calculates the way to grasp the object and adjusts as needed; trajectory reasoning, which generates a plan of operations needed to complete the task; multi-view correspondence, which identifies objects in 3D space from different angles, etc.
[0122] A multi-modal interaction method and system applied to intelligent robots, through the structured processing of visual data by the perception layer of the Gemini Robotics-ER model, three-dimensional point cloud data is converted into a description vector containing geometric topological relationships, and a high-dimensional feature tensor is extracted from a two-dimensional texture feature matrix. The visual-language-action model performs semantic analysis on natural language instructions, generates a semantic dependency tree and maps it into machine recognizable instructions. Both cooperate through a multi-modal attention mechanism to achieve deep correlation, accurately interface visual features and language instructions, and effectively solve the problem of insufficient multi-modal fusion accuracy in the prior art, greatly reducing the object recognition deviation and action intention misjudgment.
[0123] In terms of real-time response, the method and system build an efficient closed-loop mechanism. The GeminiRobotics-ER model decision layer quickly generates an interaction decision matrix containing time series information, and the action planning module of the visual-language-action model efficiently completes inverse kinematics solving and trajectory planning based on this. The motion control unit drives the robot to execute actions in real time. At the same time, the visual acquisition device captures scene changes in real time and feeds back, triggering the next cycle of interaction, forming a complete closed loop of perception, decision-making, and execution, successfully solving the problem of response lag in dynamic scenarios and ensuring that the robot can adjust actions in a timely manner to adapt to environmental changes.
[0124] In addition, the method and system significantly improve the overall interaction efficiency of industrial robots. Through the coordinated operation of each module, the precision and stability of the robot when executing actions are effectively guaranteed. From the acquisition and processing of visual data, the analysis and conversion of language instructions, to the generation of decision matrices, trajectory planning, and action execution and feedback, each link is closely connected, forming an efficient interaction process that can meet the high-precision and high-flexibility industrial interaction requirements, providing strong support for natural and efficient interaction between robots and the environment and humans in industrial automation scenarios.
[0125] In the description of the present application, it should be noted that unless otherwise explicitly specified and limited, the terms "arrangement", "installation", "connection", "connection", "fixing" should be understood broadly, for example, it can be fixed connection, or detachable connection, or integrally connected; it can be mechanical connection, or electrical connection; it can be directly connected, or indirectly connected through an intermediate medium, or the internal communication of two elements. For those skilled in the art, the specific meaning of the above terms in the present application can be understood through specific circumstances.
[0126] Although embodiments of the present application have been shown and described, it will be understood by those skilled in the art that various equivalent changes, modifications, replacements and variations of these embodiments can be made without departing from the principles and spirit of the present application, and the scope of the present application is defined by the appended claims and their equivalent scope.
Claims
1. A multimodal interaction method for intelligent robots, characterized in that, Includes the following steps: Step S1: Using the multispectral vision acquisition device mounted on the industrial robot, dynamic image sequence capture of the interactive scene is performed, and the robot joint angle, end effector pose and ambient light intensity parameters at the moment of image acquisition are recorded simultaneously. The dynamic image sequence is decomposed into three-dimensional point cloud data containing depth information and two-dimensional texture feature matrix. Step S2: Call the perceptual layer interface of the Gemini Robotics-ER model to extract features from the 3D point cloud data, generate a structured description vector containing geometric topological relationships, and simultaneously perform multi-scale convolution operation on the 2D texture feature matrix to output a high-dimensional feature tensor containing color distribution, texture gradient and edge contour. Step S3: Receive natural language instructions from external input, perform word segmentation processing through the language parsing module of the vision-language-action model, generate a semantic dependency tree composed of verbs, nouns and spatial relation words, and perform entity mapping on the semantic dependency tree in combination with a preset industry terminology library, converting it into a machine-recognizable instruction sequence containing action type, operation object and spatial coordinates. Step S4: Input the structured description vector, high-dimensional feature tensor, and machine-recognizable instruction sequence into the decision layer of the GeminiRobotics-ER model. Use a multimodal attention mechanism to fuse the features of the three types of data to generate an interactive decision matrix containing time series information. Each element in the interactive decision matrix corresponds to the joint motion parameters and end effector operation parameters of the industrial robot at the calibration time. Step S5: Based on the interactive decision matrix, inverse kinematics is solved through the motion planning module of the vision-language-action model to generate trajectory planning data containing joint angular displacement sequence, motion velocity curve and acceleration threshold, and robot dynamic parameters at each moment during trajectory execution are calculated simultaneously. Step S6: The trajectory planning data and dynamic parameters are transmitted to the motion control unit of the industrial robot, driving each joint of the robot to perform corresponding actions according to the trajectory planning data. At the same time, the scene change images during the action execution process are captured in real time by the vision acquisition device and fed back to the perception layer of the Gemini Robotics-ER model to trigger the next cycle of interaction.
2. The method according to claim 1, characterized in that, In step S2, when the Gemini Robotics-ER model extracts features from the 3D point cloud data, it adopts a feature descriptor generation method based on spatial neighborhood search. By calculating the angle between the normal vectors of each sampling point in the point cloud and its k nearest neighbors, the Euclidean distance, and the rate of curvature change, a local geometric feature matrix is constructed. The local geometric feature matrix is then enhanced using the following formula: Among them, F geo The enhanced local geometric feature vector, where n is the total number of sampling points in the point cloud, k is the number of nearest neighbors of each sampling point, and θ ij Let d be the angle between the normal vectors of the i-th sampling point and the j-th nearest neighbor point. ij Let d be the Euclidean distance between the i-th sampling point and the j-th nearest neighbor. max κ is the maximum Euclidean distance between all point pairs in the point cloud. ij V represents the rate of curvature change between the i-th sampling point and the j-th nearest neighbor point, where α, β, and γ are the weighting coefficients of the angle between the normal vectors, the Euclidean distance, and the rate of curvature change, respectively. ij Let be the unit vector pointing from the i-th sampling point to the j-th nearest neighbor. In step S2, when performing multi-scale convolution operations on the two-dimensional texture feature matrix, feature extraction is performed by setting three convolution kernels of different sizes. The feature maps output by each convolution kernel are fused using the following formula: Among them, T tex The fused texture feature tensor is denoted by m, where m is the kernel number and λ is the finite texture feature tensor. m T represents the weight coefficients of the output features of the m-th convolutional kernel. in Let K be the input two-dimensional texture feature matrix. m Let s be the parameter matrix of the m-th convolutional kernel. m The stride parameter is the stride parameter of the m-th convolution kernel, and Conv(·) represents the convolution operation.
3. The method according to claim 1, characterized in that, In step S3, when the language parsing module of the vision-language-action model performs word segmentation on natural language instructions, it adopts a word segmentation algorithm based on a bidirectional long short-term memory network. It determines the word segmentation boundaries by calculating the probability distribution of each character in the context, generating a word sequence containing part-of-speech tags. During the construction of the semantic dependency tree, subject-predicate, verb-object, and modifier-head relationships are determined by calculating the dependency probabilities between words. The dependency probabilities are calculated using the following formula: Among them, P dep (w i w j ,r) is the word w i with w j The probability W of the existence of a dependency relationship r r h is the weight matrix corresponding to dependency relation r. i h j The words are respectively w i w j Word vector representation, Let R denote the Hadamard product of vectors, where R is the set of all possible dependencies. The entity mapping process described in step S3 is achieved by calculating the semantic similarity between terms in the industry terminology database and nodes in the semantic dependency tree. The semantic similarity is calculated using the following formula: Where S(w,t) is the semantic similarity between word w and term t, w k t k δ(w,t) represents the projection values of word w and term t in the k-th semantic space, respectively, where d is the dimension of the semantic space, and δ(w,t) is the domain relevance coefficient, which takes the value of 1 when word w and term t belong to the same industry domain, and 0.3 otherwise.
4. The method according to claim 1, characterized in that, In step S4, the multimodal attention mechanism determines the attention weights by calculating the mutual information entropy between the structured description vector, the high-dimensional feature tensor, and the machine-recognizable instruction sequence. The mutual information entropy is calculated using the following formula: Where I(X,Y) is the mutual information entropy between mode X and mode Y, P(x,y) is the joint probability distribution of x and y, and P(x) and P(y) are the marginal probability distributions of x and y, respectively. During the generation of the interactive decision matrix, the decision values of the fused features are calculated using the following formula: D t,j =σ(W d ·[F geo,t ;T tex,t ;L cmd,t ]+b d ) Among them, D t,j Let W be the output value of the interaction decision matrix at time t, the j-th decision dimension, σ be the sigmoid activation function, and W be the output value of the interaction decision matrix at time t, the j-th decision dimension. d Let F be the weight matrix of the decision-making layer. geo,t Let T be the structured description vector at time t. tex,t Let L be the high-dimensional feature tensor at time t. cmd,t b represents the machine-recognizable instruction sequence characteristics at time t. d This represents the decision-making bias.
5. The method according to claim 1, characterized in that, In step S5, the inverse kinematics solution process is achieved through iterative calculation, and the joint angle is corrected in each iteration using the following formula: Δθ=(J T J+λI) -1 J T ΔX Where Δθ is the joint angle correction vector, J is the robot Jacobian matrix, λ is the damping coefficient, I is the identity matrix, and ΔX is the end effector pose error vector. The calculation of the dynamic parameters includes the calculation of the joint driving torque, which is achieved through the following formula: Where τ is the joint driving torque vector, and M(θ) is the inertia matrix. The joint angular acceleration vector. The matrix of Coriolis force and centrifugal force. G(θ) is the joint angular velocity vector, and G(θ) is the gravity term vector.
6. The method according to claim 1, characterized in that, In step S6, the motion control unit drives the robot joints using model predictive control, and the control quantity is calculated using the following formula: Where u(t) is the control quantity at time t, N is the prediction time domain, and x(t+k|t) is the system state predicted at time t+k based on the state at time t. r (t+k) represents the reference state at time t+k, and Q and R are the weight matrices for the state error and the control quantity, respectively. During the feedback process of the scene change image, the difference between the current image and the previous period image is calculated using the following formula: Among them, D img For image difference, H and W represent the height and width of the image, respectively, and I... c (i, j) represents the grayscale value of the current image at pixel (i, j), I p (i, j) represents the gray value of pixel (i, j) in the previous period image. When D img When the value exceeds the preset threshold, the next interactive cycle is triggered.
7. The method according to claim 1, characterized in that, Step S3 includes the following sub-steps: S31, after receiving the natural language instruction from the external input, the instruction text is converted into a UTF-8 encoded character stream, and the sentence is segmented according to a preset punctuation set to obtain multiple clauses. Each clause is traversed at the character level, and the positions of numbers, letters, and special symbols are marked; S32, the word segmentation interface of the vision-language-action model is called to perform word boundary detection on the segmented clauses, determine word segmentation points based on the bidirectional maximum matching algorithm, and generate a word segmentation result list containing word sequence numbers, word lengths, and part-of-speech tags, where the part-of-speech tags include verbs, nouns, adjectives, etc. Adverbs and prepositions; S33, construct a semantic dependency tree based on the word segmentation result list, with verbs as the root node, determine the connection relationship between parent and child nodes by calculating the co-occurrence frequency between words, and use nouns as action object nodes, adverbs as action modifying nodes, and prepositions as spatial relationship nodes to form a hierarchical tree diagram; S34, perform string matching between each node of the semantic dependency tree and terms in the preset industry terminology library, calculate the machine code corresponding to the term with a matching degree exceeding a set threshold, replace the text content of the corresponding node in the semantic dependency tree, and generate a recognizable instruction sequence containing machine code, node type, and connection relationship.
8. The method according to claim 1, characterized in that, Step S4 includes the following sub-steps: S41, Align the structured description vector and high-dimensional feature tensor output from step S2 in terms of dimensions, and use zero-padding to ensure that the time series lengths of the two are consistent. Convert the machine-recognizable instruction sequence into a feature matrix with the same length as the time series, with each time step corresponding to an operation instruction in the instruction sequence; S42, Initialize the weight parameter matrix of the multimodal attention mechanism, set the number of attention heads and feature dimensions, and input the structured description vector, high-dimensional feature tensor, and machine-recognizable instruction feature matrix into different attention sub-layers respectively; S43, Each attention sub-layer performs a linear transformation on the input features to generate a query vector, key vector, and value vector. The attention score is determined by calculating the dot product of the query vector and the key vector. After normalization by the softmax function, the score is multiplied by the value vector to obtain the single-modal attention feature; S44, Concatenate the single-modal attention features output from each attention sub-layer, perform feature fusion through a fully connected layer, and generate an interaction decision matrix containing the time dimension. Each row of the matrix corresponds to the decision parameters of a time step, and each column corresponds to a decision dimension.
9. The method according to claim 1, characterized in that, Step S5 includes the following sub-steps: S51, extracting the target poses of the end effector at each time step from the interaction decision matrix, including position coordinates and attitude angles, and arranging these target poses in chronological order to form a pose sequence; S52, smoothing the pose sequence, calculating the transition points between adjacent target poses using a cubic spline interpolation algorithm to keep the pose change rate continuous, and generating a dense pose trajectory containing more intermediate points; S53, based on the dense pose trajectory, calculating the angle values of each joint at each time step through inverse kinematics to form a joint angle sequence, checking the velocity and acceleration constraints of the joint angle sequence, and adjusting the joint angle values at the corresponding time step when the preset velocity or acceleration upper limit is exceeded; S54, calculating the angular velocity and angular acceleration of each joint at each time step according to the adjusted joint angle sequence, and determining the required driving torque and power consumption of each joint by combining the robot's inertial parameters, mass distribution, and friction coefficient, generating trajectory planning data containing joint angles, velocities, accelerations, and torques.
10. A multimodal interaction system for intelligent robots, characterized in that, include: The multispectral visual information acquisition and preprocessing unit is connected to the visual sensor group of the industrial robot at the input end and to the perception layer unit of the Gemini Robotics-ER model at the output end through the data bus. It is used to convert the acquired image data into three-dimensional point clouds and two-dimensional texture features. The Gemini Robotics-ER model processing unit includes a perception layer interface, a feature extraction module, and a decision layer module. The perception layer interface is connected to the output of the multispectral visual information acquisition and preprocessing unit, and the output of the decision layer module is connected to the action planning unit of the vision-language-action model for multimodal feature extraction and fusion decision-making. The natural language instruction parsing unit receives external language instructions at its input end and is connected to the semantic mapping unit of the vision-language-action model at its output end, and is used to convert natural language into a semantic dependency tree; The vision-language-action collaborative processing unit is connected to the Gemini Robotics-ER model processing unit, the natural language command parsing unit, and the trajectory planning unit, respectively, and is used to perform the association mapping between language commands and visual features; The trajectory planning and dynamics calculation unit is connected to the decision layer module of the Gemini Robotics-ER model processing unit at its input end and to the motion control unit of the industrial robot at its output end, and is used to generate joint motion trajectories and dynamic parameters. The motion control and feedback unit connects to the trajectory planning and dynamics calculation unit at its input end and drives the actuators of each joint of the industrial robot at its output end. At the same time, it connects to the multispectral visual information acquisition and preprocessing unit through an internal feedback interface to perform closed-loop control of the interactive process.
Citation Information
Cited By
Autonomous navigation system and method for electric power inspection unmanned aerial vehicle
CN121275000A
GIS in-place robot coordination control method with posture correction function
CN121348923A
Robot control method based on multi-modal large model
CN121403367A