Mechanical arm control method and device based on multi-mode AI intelligent interaction
By employing a multi-modal AI intelligent interaction method, combining microphone arrays, depth cameras, inertial sensors, and RGB-D cameras to collect multi-modal information, and performing information fusion and structured analysis, the problem of limited operational precision of robotic arms in complex tasks is solved, achieving accurate conversion of user intent and improved operational efficiency.
Patent Information
- Application Number
- CN202511128919.3
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2025-08-13
- Publication Date
- 2025-10-21
- Estimated Expiration
- 2045-08-13
AI Technical Summary
Existing robotic arm control methods rely on a single input method or simple information superposition, which makes it difficult to accurately understand user intentions, resulting in limited operational accuracy in complex tasks.
A multi-modal AI intelligent interaction method is adopted, which collects multi-modal interaction information through microphone array, depth camera, inertial sensor and RGB-D camera, performs information fusion processing and structured analysis, and constructs interaction primitive sequence to achieve precise operation.
It improves the operational accuracy and efficiency of robotic arms in complex tasks, ensures that user intentions are accurately translated into specific operational stages, and enhances the control precision and task success rate of robotic arms.
Smart Images

Figure CN120620238B_ABST
Abstract
Description
Technical Field
[0001] The present application relates to the field of robotic arm control technology, and in particular to a robotic arm control method and device based on multi-mode AI intelligent interaction. Background Art
[0002] There are various ways to control robotic arms, including those based on teach-and-play, pre-programmed routines, and those that receive user commands through sensors (such as vision, force sensing, voice, and gestures). Currently, most robotic arm control methods still rely on a single input source, such as manually programming the robotic arm's motion trajectory or relying solely on simple voice or gesture commands. These methods struggle to fully and accurately capture the user's complex operational intent, especially when faced with multi-step tasks that require precise coordination or dynamically changing environments. This can easily lead to discrepancies between the user's expressed intent and the robotic arm's understanding.
[0003] Furthermore, while some robotic arm control methods attempt to integrate multiple interaction methods, such as using voice and gestures simultaneously, they simply overlay or process information from different sensors in parallel. They lack a deep information fusion mechanism and fail to fully exploit the complementarity and correlation between different modalities. This leads to persistent misunderstandings, forcing users to provide information repeatedly or redundantly, or even requiring more complex operations to compensate for the shortcomings of a single modality. This misunderstanding of user intent directly limits the accuracy of robotic arm operations. For example, in operations such as grasping, assembly, and placement, the robot cannot accurately locate the specific location of the target object or interact with it in the correct posture, affecting the efficiency and success rate of task execution.
[0004] In summary, the existing technology has a technical problem in that the robotic arm control method relies on a single input method or simple information superposition, which makes it difficult to accurately understand the user's intention when processing complex tasks, resulting in limited robotic arm operation accuracy. Summary of the Invention
[0005] The purpose of this application is to provide a robotic arm control method and device based on multi-mode AI intelligent interaction, so as to solve the technical problem in the prior art that the robotic arm control method relies on a single input method or simple information superposition, which makes it difficult to accurately understand the user's intention when processing complex tasks, resulting in limited robotic arm operation accuracy.
[0006] In view of the above problems, the present application provides a robotic arm control method and device based on multi-mode AI intelligent interaction.
[0007] In a first aspect, the present application provides a robotic arm control method based on multimodal AI intelligent interaction, which is implemented by a robotic arm control device based on multimodal AI intelligent interaction, wherein the robotic arm control method based on multimodal AI intelligent interaction includes: collecting multimodal interaction information, and fusing the multimodal interaction information to obtain a structured interaction instruction data sequence; performing user task intention analysis based on the structured interaction instruction data sequence to obtain a structured operation phase sequence of the target robotic arm, wherein each structured operation phase includes an active object, a passive object and an action type; combining the structured interaction instruction data sequence, for each structured operation phase in the structured operation phase sequence, constructing an interaction primitive centered on the active object to obtain an interaction primitive sequence, wherein each interaction primitive includes an interaction point and an interaction direction, wherein the interaction point is a key interaction position in the active object, and the interaction direction is a spatial direction vector combining the action type and the interaction scene; using the target robotic arm to transfer the active object to the passive object according to the interaction primitive sequence.
[0008] Optionally, user voice commands are collected through a microphone array, and the user voice commands are input into a voice recognition module for semantic analysis to extract a keyword information sequence; user gesture movements are captured through a depth camera or an inertial sensor, and spatial posture analysis and recognition are performed on the user gesture movements to obtain action intention information; three-dimensional scene image information of the current operation area is collected through an RGB-D camera to obtain visual feature information and spatial coordinate data of the operation object in space; the keyword information sequence, action intention information, visual feature information and spatial coordinate data are summarized as multimodal interaction information.
[0009] Optionally, the keyword information sequence is divided into front-to-back associations according to time sequence to obtain multiple divided keyword information subsequences; the timestamp of the action intention information is obtained, and the multiple divided keyword information subsequences are supplemented with action intentions to obtain multiple supplemented divided keyword information subsequences; the multiple supplemented divided keyword information subsequences are supplemented with scenes in combination with the visual feature information and spatial coordinate data to obtain a structured interactive instruction data sequence.
[0010] Optionally, according to a preset time interval constraint, the keyword information sequence is subjected to first-order division node identification to determine a first-order division node set; according to a preset semantic association constraint, the keyword information sequence is subjected to second-order division node identification to determine a second-order division node set; the first-order division node set and the second-order division node set are unioned to obtain a fused division node set; the keyword information sequence is divided using the fused division node set to obtain the multiple divided keyword information subsequences.
[0011] Optionally, the structured interactive instruction data sequence is traversed according to the active object name, passive object name and action trajectory to extract the user task intention, and the active object sequence, passive object sequence and action trajectory sequence are obtained; the action trajectory sequence is matched respectively in the action type library of the target robotic arm to obtain the action type sequence; the active object sequence, passive object sequence and action type sequence are mapped and associated to construct the structured operation phase sequence.
[0012] Optionally, the motion trajectory sequence is traversed to extract trajectory features to obtain a motion trajectory feature sequence; the motion trajectory feature sequence is matched with the historical motion trajectory features in the motion type library respectively, and the motion type corresponding to the maximum matching similarity is output to obtain the motion type sequence.
[0013] Optionally, the first interaction scene of the first structured operation stage is extracted from the structured interaction instruction data sequence: image information of the active object in the first structured operation stage is extracted based on the first interaction scene, interaction points are identified based on the image information, and the first interaction point is determined; path obstacle correction is performed on the action type of the first structured operation stage based on the first interaction scene to obtain a first interaction direction; the first interaction direction and the first interaction point are associated to construct a first interaction primitive; the structured interaction instruction data sequence and the structured operation stage sequence are traversed for analysis to construct the interaction primitive sequence.
[0014] Optionally, a path obstacle identifier is used to identify the first interaction scene and the action type to determine a first path obstacle identification result; based on the first path obstacle identification result and the motion tolerance range of the target robotic arm, the action type is corrected to obtain a first interaction direction.
[0015] Optionally, the spatial posture of the active object corresponding to each interactive primitive in the interactive primitive sequence is monitored in real time to obtain a sequence of active object spatial posture monitoring results; correction instructions are identified based on the active object spatial posture monitoring result sequence, and a correction instruction sequence is determined, and the correction instruction sequence is transmitted to the end effector of the target robotic arm for control correction.
[0016] In a second aspect, the present application also provides a robotic arm control device based on multimodal AI intelligent interaction, which is used to execute the robotic arm control method based on multimodal AI intelligent interaction as described in the first aspect, wherein the robotic arm control device based on multimodal AI intelligent interaction includes: an information processing module, which is used to collect multimodal interaction information and fuse the multimodal interaction information to obtain a structured interaction instruction data sequence; a task intention parsing module, which is used to perform user task intention parsing based on the structured interaction instruction data sequence to obtain a structured operation phase sequence of the target robotic arm, wherein each structured operation phase includes an active object, a passive object and an action type; an interaction primitive construction module, which is used to combine the structured interaction instruction data sequence and, for each structured operation phase in the structured operation phase sequence, construct an interaction primitive centered on the active object to obtain an interaction primitive sequence, wherein each interaction primitive includes an interaction point and an interaction direction, wherein the interaction point is a key interaction position in the active object, and the interaction direction is a spatial direction vector combining the action type and the interaction scene; a robotic arm control module, which is used to use the target robotic arm to transfer the active object to the passive object according to the interaction primitive sequence.
[0017] One or more technical solutions provided in this application have at least the following beneficial effects:
[0018] By collecting multimodal interaction information and fusing the multimodal interaction information, a structured interaction instruction data sequence is obtained; based on the structured interaction instruction data sequence, user task intention is parsed to obtain a structured operation phase sequence of the target robotic arm, wherein each structured operation phase includes an active object, a passive object and an action type; combined with the structured interaction instruction data sequence, for each structured operation phase in the structured operation phase sequence, an interaction primitive centered on the active object is constructed to obtain an interaction primitive sequence, wherein each interaction primitive includes an interaction point and an interaction direction, wherein the interaction point is a key interaction position in the active object, and the interaction direction is a spatial direction vector combining the action type and the interaction scene; and the target robotic arm is used to transfer the active object to the passive object according to the interaction primitive sequence. In other words, by fusing and processing multimodal interaction information, analyzing user task intentions, and converting the user's abstract intentions into specific operation stages, for each structured operation stage, an interaction primitive centered on the active object is constructed, including interaction points and interaction directions. The target robotic arm transfers the active object to the passive object according to the interaction primitive sequence, realizing the conversion from user intentions to specific robotic arm actions, and improving the accuracy and efficiency of robotic arm operations.
[0019] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, which can be implemented in accordance with the contents of the description, and to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are specifically listed below. It should be understood that the content described in this section is not intended to identify the key or important features of the embodiments of the present application, nor is it intended to limit the scope of the present application. Other features of the present application will become easy to understand through the following description. BRIEF DESCRIPTION OF THE DRAWINGS
[0020] In order to more clearly illustrate the technical solutions in this application or the prior art, the following briefly introduces the drawings required for use in the embodiments or the description of the prior art. Obviously, the drawings described below are merely exemplary, and a person of ordinary skill in the art can obtain other drawings based on the provided drawings without creative work.
[0021] Figure 1 This is a flow chart of the robotic arm control method based on multi-mode AI intelligent interaction in this application.
[0022] Figure 2 This is a structural diagram of the robotic arm control device based on multi-mode AI intelligent interaction in this application.
[0023] Explanation of the accompanying drawings: information processing module 11, task intention analysis module 12, interaction primitive construction module 13, robotic arm control module 14. DETAILED DESCRIPTION
[0024] This application provides a robotic arm control method and device based on multi-modal AI intelligent interaction, solving the technical problem in the prior art that the robotic arm control method relies on a single input method or simple information superposition, making it difficult to accurately understand the user's intention when handling complex tasks, resulting in limited robotic arm operation accuracy. By fusing multi-modal interaction information, analyzing the user's task intention, and converting the user's abstract intention into a specific operation stage, for each structured operation stage, an interaction primitive centered on the active object is constructed, including interaction points and interaction directions. The target robotic arm transfers the active object to the passive object according to the interaction primitive sequence, realizing the conversion from user intention to specific robotic arm action, and improving the accuracy and efficiency of robotic arm operation.
[0025] Below, the technical solutions in this application will be clearly and completely described with reference to the accompanying drawings. Obviously, the described embodiments are only part of the embodiments of this application, rather than all the embodiments of this application. It should be understood that this application is not limited to the example embodiments described herein. Based on the embodiments of this application, all other embodiments obtained by ordinary technicians in this field without making creative work are within the scope of protection of this application. It should also be noted that, for the convenience of description, only the parts related to this application, rather than all of them, are shown in the accompanying drawings.
[0026] For example, see the attached Figure 1 The present application provides a robotic arm control method based on multi-mode AI intelligent interaction, wherein the robotic arm control method based on multi-mode AI intelligent interaction is executed by a robotic arm control device based on multi-mode AI intelligent interaction, and the robotic arm control method based on multi-mode AI intelligent interaction specifically includes the following steps:
[0027] Multimodal interaction information is collected and fused to obtain a structured interaction instruction data sequence.
[0028] Furthermore, the present application further comprises the following steps:
[0029] User voice commands are collected through a microphone array, and the user voice commands are input into a voice recognition module for semantic analysis to extract a keyword information sequence; user gestures are captured through a depth camera or inertial sensor, and spatial posture analysis and recognition of the user gestures are performed to obtain action intention information; three-dimensional scene image information of the current operation area is collected through an RGB-D camera to obtain visual feature information and spatial coordinate data of the operation object in space; the keyword information sequence, action intention information, visual feature information and spatial coordinate data are summarized into multimodal interaction information.
[0030] Specifically, the microphone array captures user voice commands—operation instructions or expressed needs issued by the user through spoken language, such as "place a tool on the workbench." A microphone array is a sensor system composed of multiple microphone units arranged in a specific geometric structure. It uses the time, phase, or intensity differences between sound waves reaching different microphones to locate the direction of the sound source, enhance the target voice signal, and suppress noise and reverberation, thereby improving the quality of voice acquisition and its ability to resist interference. Users issue voice commands in natural language, and the microphone array receives them and filters background noise using beamforming technology before transmitting them to the speech recognition module.
[0031] The speech recognition module performs semantic analysis (e.g., ASR + NLU) on user voice commands. It converts user voice signals into text, extracts features, decodes the acoustic model, and understands the semantics of the sound signal. The speech recognition module processes the input voice signal. First, the preprocessing module performs denoising, framing, and feature extraction on the sound signal to obtain a digital signal containing speech features. This digital signal is then input into the acoustic model for decoding. The acoustic model uses training data (including vocabulary, speech features, and their combination patterns) to identify the user's spoken content and convert it into readable text. The semantic parsing module analyzes the converted text, identifies instructional information, and extracts keywords. Each sentence may contain one or more keywords, resulting in a keyword sequence. This keyword sequence extracts core semantic terms from the voice command, typically those critical to the task.
[0032] Use depth cameras or inertial sensors to capture user gestures in real time. For example, if a user points their finger from the left side of a tabletop to the right, the gesture is interpreted as an intended direction from left to right. A depth camera is a camera that can capture three-dimensional spatial data. It obtains depth information of the scene by calculating the distance between each pixel and the camera. Inertial sensors, such as accelerometers and gyroscopes, are typically used to detect and record dynamic information such as acceleration and angular velocity of an object. They are commonly used in hand tracking and gesture recognition, and can capture the motion trajectory and rotation of the hand or body. A depth camera captures three-dimensional image data of the user's hand and its surroundings, generating depth information for each pixel and identifying the hand's specific position, angle, and motion trajectory in space. Inertial sensors (such as accelerometers and gyroscopes in handheld devices) can detect dynamic changes in the hand, identifying motion states such as acceleration and rotation, and thus capturing rapid changes in user gestures. User gestures are movements of the hand or other body parts to express commands or intentions, such as pointing, waving, or grasping.
[0033] Spatial gesture analysis analyzes the hand's spatial posture and extracts key spatial features, such as finger pointing direction, hand angle, and motion trajectory. Spatial gesture analysis analyzes the spatial position and angle (e.g., posture, direction, and motion trajectory) of the user's gesture to identify its specific posture in three-dimensional space (e.g., finger pointing direction, bending angle, etc.). Action intention information is extracted from the spatial gesture analysis results. This means that spatial analysis of gestures can be used to extract information related to the user's intention. For example, a user standing approximately 2 meters from a depth camera makes a pointing gesture. The depth camera captures an image of the user's hand and generates 3D point cloud data. Based on the depth image, spatial gesture analysis identifies the location of the user's finger, such as the finger pointing coordinates (1.0, 1.2, 0.5). Simultaneously, the inertial sensor records the user's hand acceleration and rotation data, identifying the finger's pointing motion as pointing toward the center of the work surface.
[0034] The RGB-D camera is activated to capture three-dimensional scene image information of the current operating area. This information includes not only visual features such as color and shape of all objects in the operating area, but also the precise three-dimensional coordinates of each operating object in space. For example, a red cylindrical cup on a workbench has spatial coordinates of its center point of X = 1.2 meters, Y = 0.5 meters, and Z = 0.8 meters. An RGB-D camera simultaneously captures color images (RGB) and depth information (D) and is commonly used to acquire three-dimensional scene data. Visual feature information refers to the characteristics of the operating object in the image, such as color, shape, texture, and size, analyzed by the RGB-D camera. This information helps identify and distinguish different objects. Spatial coordinate data refers to the precise position of the operating object in three-dimensional space, typically represented by X, Y, and Z coordinate values, which define the object's position relative to the camera or a fixed reference frame.
[0035] The keyword sequence, action intention, visual feature information, and spatial coordinate data are aggregated to form multimodal interaction information. For example, a user instructs the robotic arm to pick up a blue screwdriver and screw it into a screw hole on a wooden board. The microphone array captures the following: "pick up the blue screwdriver." Voice analysis yields the keywords "pick up," "blue," and "screwdriver." Inertial sensors capture the user's gestures: "pick up an object on the table," along with the hand's closing and lifting movements. Posture analysis identifies the action intent: grasping. An RGB-D camera scans the workbench and identifies a long, blue object (a screwdriver) on the table. The gripping point coordinates of the handle are X=0.9m, Y=0.3m, and Z=0.7m. Visual features such as the blue color and long shape are extracted. Next, the microphone array captures the following: "screw into a screw hole on a wooden board." Voice analysis yields the keywords "screw into," "wood board," and "screw hole." The user points their finger to a specific location on the board, and the action intent is recognized as pointing to a specific area on the board. The RGB-D camera identifies the screw hole coordinates on the wooden board as X=1.1m, Y=0.4m, and Z=0.6m. The resulting multimodal interaction information includes the following: pick up, blue, screwdriver, screw to, wooden board, screw hole; the action intention sequence: grasping and pointing to the wooden board; visual features (description of the screwdriver and the screw hole in the board); and the corresponding spatial coordinates (screwdriver X=0.9m, Y=0.3m, Z=0.7m; screw hole X=1.1m, Y=0.4m, Z=0.6m).
[0036] Furthermore, the present application also includes the following steps: dividing the keyword information sequence into front-to-back associations according to the time sequence to obtain multiple divided keyword information subsequences; obtaining the timestamp of the action intention information, supplementing the multiple divided keyword information subsequences with action intentions to obtain multiple supplemented divided keyword information subsequences; combining the visual feature information and spatial coordinate data to supplement the multiple supplemented divided keyword information subsequences with scenes to obtain a structured interactive instruction data sequence.
[0037] Furthermore, the present application also includes the following steps: according to the preset time interval constraint, the first-order division node identification is performed on the keyword information sequence to determine the first-order division node set; according to the preset semantic association constraint, the second-order division node identification is performed on the keyword information sequence to determine the second-order division node set; the first-order division node set and the second-order division node set are obtained as the union to obtain the fused division node set; the keyword information sequence is divided using the fused division node set to obtain the multiple divided keyword information subsequences.
[0038] Specifically, keyword sequences, action intent information, visual feature information, and spatial coordinate data are aligned according to timestamps. Because devices such as microphone arrays, depth cameras / inertial sensors, and RGB-D cameras may have varying operating frequencies and data processing latencies, the time points at which they collect or process data naturally differ. Without alignment, directly mixing data from these three sources can lead to temporal misalignment: a voice keyword at one moment may be incorrectly associated with a gesture at a previous moment or visual scene data at a later moment, resulting in an erroneous understanding. Using the timestamp information associated with each data point, a shared timeline or time window is established. For each newly arriving data point (whether it's a keyword, action intent, or visual data), its timestamp is checked and placed at the corresponding position on this same timeline. If the arrival time of data from a data source deviates slightly from the current moment on the timeline (e.g., due to processing delays or imperfect sensor sampling), interpolation, sliding window matching, or nearest neighbor matching based on time differences are used to associate data from other modalities closest to that timestamp. For example, if the timestamp is t=1.235s and the keyword is "pick up", search for action intention information (such as a finger pointing to a certain area) and visual data (including the image of the area, object coordinates, etc.) that arrive around t=1.235s (such as within a window of ±50ms), and regard these three as information that is aligned in time and describes the user behavior and scene status at the same moment.
[0039] The preset timing interval constraint is a pre-set threshold used to determine whether two consecutive keyword messages are sufficiently close in time. If the time interval between two keyword messages exceeds this threshold, they are considered to belong to different instruction or intent stages. The time interval constraint (i.e., the time difference between adjacent keyword messages) is used to determine the keyword message splitting node. If the time interval between two keyword messages exceeds the preset timing interval constraint, they are considered the starting points of different subtasks or actions, and a first-order split node is set at this point.
[0040] The preset semantic association constraint means that the semantic similarity between adjacent keywords is calculated to determine whether they belong to the same task or action. Keywords with higher semantic similarity are considered to belong to the same action, while keywords with lower semantic similarity are considered to be different actions. By calculating the similarity of two keywords (such as using the cosine similarity of the word vector model), if the similarity of two adjacent keywords is lower than the set semantic similarity threshold, a second-order partition node is set. If the similarity is lower than a certain threshold, it is considered that they may belong to different instructions or intent stages. Traverse the keywords and their subsequent keywords in the keyword sequence in turn. If the similarity between the two words is low, it may be because the semantic relationship is far, resulting in a low similarity between the two, and thus they are divided into different subtasks, and a second-order partition node is set.
[0041] The first-order partition node set and the second-order partition node set are unioned to obtain a fused partition node set, which simultaneously considers information from two dimensions, time interval and semantic similarity, to accurately divide the user's semantic intent. The fused partition node set is the result of the union operation of the first-order and second-order partition node sets. It will more accurately divide the keyword information by considering both time interval and semantic similarity. By fusing the partition node set, the keyword information sequence is divided to obtain multiple subsequences of divided keyword information, that is, the subsequences obtained by segmenting the keyword information sequence through the fused partition node. Each subsequence contains keyword information for a specific action or task. For example, when placing object A on object B, the data keywords in this process are all related and can be divided into one subsequence; moving object B a certain distance is another action, so it needs to be divided into another subsequence.
[0042] For example, suppose a user issues a voice command. The speech recognition module extracts the following keyword information sequence (one keyword per sentence) and its timestamps: Keyword 1: "pick up," 0 seconds; Keyword 2: "blue screwdriver," 1.5 seconds; Keyword 3: "place," 4.0 seconds; Keyword 4: "screw hole in the wood," 5.5 seconds; Keyword 5: "tighten," 7.0 seconds. The preset temporal interval constraint is set to ΔT = 2.5 seconds; the preset semantic association constraint is set to a similarity threshold of 0.4. First-order partition node identification: The time interval between keywords 1 and 2 is 1.5s < 2.5s, so no partition node is set; the time interval between keywords 2 and 3 is 2.5s ≥ 2.5s, so a first-order partition node is set between "blue screwdriver" and "place"; the time interval between keywords 3 and 4 is 1.5s < 2.5s, so no partition node is set; the time interval between keywords 4 and 5 is 1.5s < 2.5s, so no partition node is set. Second-order partition node identification: The similarity between keywords 1 and 2 is 0.6 > 0.4, so no node is set; the similarity between keywords 2 and 3 is 0.2 < 0.4, so a second-order partition node is set between the blue screwdriver and placement; the similarity between keywords 3 and 4 is 0.7 > 0.4, so no node is set; the similarity between keywords 4 and 5 is 0.5 > 0.4, so no node is set. The union of the first-order partition node set and the second-order partition node set yields the fused partition node set: between keywords 2 and 3. Partitioning the keyword information sequence yields {pick up, screwdriver} and {place, screw hole on template, tighten}.
[0043] Obtain the timestamp of the action intent information, that is, the specific time at which each action occurred. Timestamps are used to synchronize voice commands and action intents, allowing for the supplementation of corresponding keyword information subsequences. The multiple previously divided keyword information subsequences are supplemented based on the timestamps and action intent information. Specifically, the keywords are supplemented based on gestures. For example, if the voice command instructs you to move object A, and the action intent information describes the specific location to move it to, then supplementation can be performed. This supplementation process relies primarily on synchronization of timestamps and accurate matching of action intent information. Find the action intent information that is closest in time to or matches the currently processed keyword information subsequence. If found, and this action intent information can provide more specific and accurate operational details for the current subsequence (especially information that may be missing or ambiguous in the voice keywords, such as the specific target location, direction, and force), this gesture intent information is supplemented to the corresponding keyword information subsequence. The multiple supplemented keyword information subsequences are subsequences supplemented with action intent information, containing more complete and specific task instructions.
[0044] Combining visual feature information and spatial coordinate data, the scene is supplemented for multiple supplementary keyword information subsequences. Specifically, the visual feature information and spatial coordinate data are searched for visual feature information and spatial coordinate data that are temporally aligned with the multiple supplementary keyword information subsequences. The visual feature information and spatial coordinate data are matched with each supplementary keyword information subsequence to obtain a structured interaction command data sequence. A structured interaction command data sequence refers to complete control data after semantic parsing, intent supplementation, and visual positioning completion. It typically includes fields such as the operation object, action type, target location, object visual features, and coordinate data. For example, for the subsequence [pick up, screwdriver, target position (0.8, 0.3, 0.6)], visual information collected at similar timestamps is searched to identify which objects in the scene match the visual features of a screwdriver, such as specific shape, color, and texture. Combined with the spatial coordinate data, the current position coordinates of the screwdriver, such as (0.9, 0.3, 0.7), and the precise coordinates and posture of the screw at the target position (0.8, 0.3, 0.6), such as the screw head facing up and the center point coordinates (0.8, 0.3, 0.6), are determined. This specific information is added to the subsequence, replacing or clarifying the original, potentially ambiguous description. The subsequence might be supplemented with [pick up, object ID3 (visual features: cylindrical shape, silver handle, spatial coordinates: (0.9, 0.3, 0.7)), target position: object ID5 (visual features: round, metal, spatial coordinates: (0.8, 0.3, 0.6, posture facing up)].
[0045] By combining abstract language and gesture intentions with precise visual recognition and spatial positioning, an instruction containing all necessary execution parameters is obtained: it clearly specifies which specific object and what action to perform at which exact location, avoiding ambiguous execution caused by incomplete semantics.
[0046] The user task intention is parsed based on the structured interaction instruction data sequence to obtain a structured operation phase sequence of the target robotic arm, where each structured operation phase includes an active object, a passive object and an action type.
[0047] Furthermore, the present application also includes the following steps: traversing the structured interactive instruction data sequence according to the active object name, passive object name and action trajectory to extract the user task intention, and obtain the active object sequence, passive object sequence and action trajectory sequence; matching the action trajectory sequence in the action type library of the target robotic arm respectively to obtain the action type sequence; mapping and associating the active object sequence, passive object sequence and action type sequence to construct the structured operation stage sequence.
[0048] Furthermore, the present application also includes the following steps: traversing the action trajectory sequence to extract trajectory features to obtain an action trajectory feature sequence; matching the action trajectory feature sequence with the historical action trajectory features in the action type library, and outputting the action type corresponding to the maximum matching similarity value to obtain the action type sequence.
[0049] Specifically, obtain the active object name, passive object name and action trajectory. The active object name refers to the subject or starting object that performs the action. In more complex scenarios, if multiple objects interact, the active object is the object that is operated or moved first. The passive object name refers to the object or target object of the action. For example, if you put a cup on the table, the active object is the cup and the passive object is the table. The action trajectory is the path that the active object moves in space, which can be a simple straight line or a complex curve. In the instruction data sequence, it is usually represented by a series of spatial coordinate points, which defines the movement path of the object from the starting point to the end point.
[0050] According to the names of active objects, passive objects and motion trajectories, the structured interaction instruction data sequence is traversed to extract the user's task intention. That is, the core task elements are identified from the structured instruction data sequence, including clarifying which objects are the main body of the operation (active objects), which are the objects of the operation (passive objects), and how the objects move (motion trajectory), to obtain the active object sequence, passive object sequence and motion trajectory sequence.
[0051] An action trajectory sequence is a continuous spatial path of an active object from its starting position to its target position. It typically consists of a series of consecutive coordinate points (x, y, z), representing the object's motion trajectory in three-dimensional space. Traversing the action trajectory sequence, we extract trajectory features. Specifically, we extract a series of representative features from each trajectory to describe the trajectory's spatial properties and dynamic behavior, including path length, direction vector, velocity, acceleration, curvature, and so on.
[0052] For each trajectory, the Euclidean distance from the starting point to the end point is calculated to obtain the path length. The velocity difference between each point in the trajectory is used to obtain the velocity change. The direction of each trajectory segment is obtained, represented as a unit vector, to obtain the direction vector. The rate of change of velocity between trajectory points is obtained to obtain the acceleration. The curvature of the trajectory path is calculated to obtain the curvature. The motion trajectory feature sequence is a set of parameters extracted from a series of motion trajectories that can represent the key characteristics of the trajectory. For example, if the first trajectory segment (grasp) is [(0.35, 0.12, 0.14), (0.36, 0.13, 0.14), ..., (0.38, 0.14, 0.14)], feature extraction results in a path length of 0.13 meters, a direction vector of (1, 0, 0), and a velocity of 0.2 m / s. The second trajectory (placed) is [(0.38, 0.14, 0.14), (0.40, 0.15, 0.14), ..., (0.48, 0.16, 0.14)]. Feature extraction is performed and the path length is 0.25 meters, the direction vector is (1, 0, 0), and the speed is 0.1 m / s.
[0053] Based on the action trajectory feature sequence, the current trajectory features are matched against the historical action trajectory features of each action type. The current trajectory features are compared with the features stored in the historical action type library, and the most matching action type is selected. The higher the similarity value, the higher the match between the current trajectory and the action type. The historical action type library has preset trajectory features for various actions. For example, the path of a grabbing action is short, the speed is fast, and it is a straight line; the path of a preventing action is relatively smooth and the speed is slow; and the trajectory of a rotating action is circular or arc-shaped. A similarity matching algorithm, such as cosine similarity, is used to calculate the similarity between the current trajectory and the historical action trajectory. The similarity is obtained by calculating the ratio of the product of the feature vectors and the modular product of the two. The current trajectory features are matched with the trajectory features of the placement action in the historical library. If the calculated similarity is 0.95, it indicates that the current trajectory is most similar to the placement action.
[0054] Through similarity calculation, the action type corresponding to the maximum similarity is output, and all action trajectories are traversed and the corresponding action type is output for each trajectory to generate an action type sequence, which includes the action type of each action trajectory. For example, the path length of the grasping action is <0.2 meters, and the speed is fast; the path length of the placement action is >0.2 meters, and the speed is slow. The calculated similarity between the first segment of the trajectory and the grasping is 0.92; the similarity between the second segment of the trajectory and the placement is 0.95, so the output action type sequence is {grab, place}. If the target position is missing or the target object is blurred in some instructions, semantic completion or ambiguity resolution can be performed based on the target pointed by the most recent gesture, the target object detected visually in the same frame, and the target reasoning of historical instructions in the context to improve the accuracy of intent parsing.
[0055] The active object sequence, passive object sequence, and action type sequence are aligned and combined one by one to generate a structured operation phase sequence. The active object sequence, passive object sequence, and action type sequence are combined based on their temporal or logical correspondence. For example, the first phase "pick up" corresponds to the active object, the blue screwdriver; the second phase "translation" corresponds to the movement between the active object, the blue screwdriver, and the passive object, the screw hole; and the third phase "placement" corresponds to the active object, the blue screwdriver, and the passive object, the screw hole. The originally scattered information is integrated into a structured operation phase sequence: [Phase 1: Active object = blue screwdriver, Action type = pick up; Phase 2: Active object = blue screwdriver, Passive object = screw hole, Action type = translation; Phase 3: Active object = blue screwdriver, Passive object = screw hole, Action type = placement].
[0056] By accurately converting the user's relatively natural and potentially ambiguous interactive instructions into a structured sequence of operation steps that the robotic arm can understand and execute, and by extracting core elements (active / passive objects, motion trajectories), analyzing trajectory features, and matching the action type library, the target actions in complex instructions can be accurately identified, thereby improving task execution accuracy.
[0057] Combined with the structured interaction instruction data sequence, for each structured operation stage in the structured operation stage sequence, an interaction primitive centered on the active object is constructed to obtain an interaction primitive sequence, wherein each interaction primitive includes an interaction point and an interaction direction, the interaction point being a key interaction position in the active object, and the interaction direction being a spatial direction vector combining the action type and the interaction scene.
[0058] Furthermore, the present application also includes the following steps: extracting the first interaction scene of the first structured operation stage from the structured interaction instruction data sequence: extracting the image information of the active object in the first structured operation stage based on the first interaction scene, identifying the interaction point based on the image information, and determining the first interaction point; performing path obstacle correction on the action type of the first structured operation stage based on the first interaction scene to obtain the first interaction direction; associating the first interaction direction and the first interaction point to construct a first interaction primitive; traversing the structured interaction instruction data sequence and the structured operation stage sequence for analysis to construct the interaction primitive sequence.
[0059] Furthermore, the present application also includes the following steps: using a path obstacle identifier to identify the first interaction scene and the action type to determine a first path obstacle identification result; based on the first path obstacle identification result and the motion tolerance range of the target robotic arm, correcting the action type to obtain a first interaction direction.
[0060] Specifically, the first structured operation phase is randomly extracted from the sequence of structured operation phases and matched against the structured interaction instruction data sequence to obtain the corresponding first interaction scene. The first interaction scene is the frame or set of environmental perception data (image, point cloud, target detection, etc.) that matches the extracted structured operation phase. It includes information such as the spatial layout and relative position of the active object, passive objects, and their surrounding environment when executing the first structured operation phase. Based on the first interaction scene, detailed visual details of the active object in the first structured operation phase are extracted, including the active object's color, shape, texture, etc. For example, the image information shows that the cup is 10 cm high and 5 cm in diameter, with the handle located on the right and a protruding edge at the base of the handle.
[0061] Based on detailed visual information, the robot locates the key point of contact between the robot arm and the active object, known as the first interaction point. A description of the first interaction scenario is obtained, identifying the object to be manipulated and its general position within the environment. An RGB-D camera is focused on the active object, capturing high-resolution color and depth images, providing rich image information including color, shape, texture, and 3D structure. The robot then identifies characteristic regions on the object that meet operational requirements such as grasping and twisting. For example, for a cup, the handle is of particular interest, as it is designed for gripping. The robot examines the handle's base edge, shape, and connection to the cup body to comprehensively determine this area as the optimal gripping point. Recognition results are presented as coordinates (such as offsets relative to the cup's rim or 3D coordinates in the global coordinate system) and a possible region range. The resulting first interaction point is the precise location where the robot arm's gripper should align and apply force, such as the center of the cup's handle.
[0062] For example, assume the first interaction scenario involves picking up a white ceramic mug located at (0.8, 0.2, 0.7) on a workbench. The mug is 12 cm tall, 6 cm in diameter, and has a handle on the right side. The RGB-D camera captures image information of the mug, including the clear white ceramic surface texture and the three-dimensional shape and position of the handle. The interaction point recognition algorithm analyzes the image, first locating the handle, then calculating the three-dimensional coordinates of the base of the handle, and finding that its center point is (0.82, 0.18, 0.75) in the global coordinate system. The surface of the mug is relatively flat, suitable for gripping with the gripper, and it is a safe grip point at a certain distance from the cup's rim. Therefore, the first interaction point determined is the center area of the base of the handle at coordinates (0.82, 0.18, 0.75).
[0063] The path obstacle detector is a module used to identify obstacles in the path between active and passive objects in an interactive scene. It analyzes the current operating environment (the first interactive scene) and the type of action being planned to determine whether there are objects or areas in the robot's path that could hinder its motion during the execution of the action. A large amount of training data is collected, consisting of two components: 1. 3D point cloud data or depth images captured by an RGB-D camera or other sensor, representing various scenes and reflecting the distribution of objects in the operating environment; 2. The robot's planned motion trajectory data, including the starting point, end point, and possible intermediate points, is collected. To ensure model generalization, it is necessary to cover a variety of scenarios (such as desktops, assembly lines, and warehouses), different object types, and different action types (such as linear movement, curved movement, grasping, and placing). Furthermore, the training data must be labeled to clearly indicate which points or areas are safe and which will collide with obstacles when executing a given motion trajectory in a given scenario. Labels are generated by combining precise collision detection algorithms (such as precise collision detection based on geometric models or distance field calculations based on point clouds). For example, a dataset can be generated where each item contains: a scene point cloud, an action trajectory (a series of coordinate points), and a list of collision labels (one label for each trajectory point: 0 for no collision, 1 for collision).
[0064] Machine learning, such as deep learning models, uses 3D convolutional neural networks to learn features directly from scene representations (point clouds or voxel representations). Combined with motion trajectory information (encoded as a sequence of trajectory points or a parameterized representation), the model outputs collision probabilities or directly predicts collision points. Scene and trajectory information is fed into a graph neural network (GNN), treating objects and trajectory points as nodes in the graph. Collision relationships are inferred through information transfer between nodes and edges. The choice of model depends on the trade-off between accuracy, speed, training cost, and required hardware resources. For example, if real-time requirements are extremely high, fast geometry-based algorithms may be preferred. However, if the scene is complex and higher generalization is required, deep learning models may be more appropriate and are generally chosen. The model learns the complex mapping from scene and trajectory information to collision predictions. A loss function (such as mean squared error) is used to calculate the loss, and the model parameters are adjusted to minimize the prediction error on the training data. After training, the model's performance should be evaluated on an independent test dataset. Metrics may include precision, recall, F1 score, and mean detection time. The test data should simulate the diversity of real-world application scenarios as closely as possible. If the model's performance meets the requirements, it can be deployed in the actual robotic arm control system. During deployment, it's important to consider whether the model's inference speed meets real-time requirements and whether further optimization (such as model pruning and quantization) is required to adapt it to the embedded hardware environment. After deployment, the model's performance must be continuously monitored during actual operation, and fine-tuning or retraining may be necessary based on new data or feedback.
[0065] The first interaction scene and action type are identified using the path obstacle identifier. Information about all objects in the first main scene is obtained. Combined with the action type, the robot arm is analyzed for collision risks along the straight line or pre-set path from its starting point to the interaction point, generating a first path obstacle identification result. The first path obstacle identification result is the analysis output by the path obstacle identifier, clearly indicating whether there are obstacles along the robot arm's path when executing the action type planned in the first structured operation phase, as well as the specific location or type of the obstacles, such as no obstacles, obstacle X located at position Y on the path, etc.
[0066] Obtain the target manipulator's motion tolerance range. This refers to the tolerance or flexibility within which the target manipulator's motion parameters (such as position, posture, velocity, and acceleration) can vary when performing an action. This is typically determined by the manipulator's physical structure, control accuracy, and safety requirements. For example, the manipulator may allow its end effector to adjust its position within a ±5 mm range around the grasping point, or allow a ±5-degree deviation in posture before reaching the target point. Based on the obstacle detection results of the first path and the target manipulator's motion tolerance range, the action type is modified, leveraging the manipulator's motion flexibility to modify the original action type. The goal of the modification is to avoid obstacles while minimizing the efficiency of the action and aligning it with the user's intent. For example, rather than completely changing the grasping point, the modification might involve changing the manipulator's motion path from a straight line to an arc, or raising the arm before lowering it to grasp. The corrected first interaction direction is no longer to (0.82, 0.18, 0.75). Instead, it moves upward to (0.82, 0.18, 0.85), then horizontally to (0.82, 0.18, 0.75), and finally descends to grasp. This first interaction direction ensures that the robot arm can safely circumvent obstacles and avoid collisions when performing a grasping action.
[0067] The first interaction direction and the first interaction point are associated to obtain the first interaction primitive. The first interaction direction is the motion direction vector that the robot arm needs to execute in the current operation phase. It is the spatial motion direction obtained by combining the user's operation intention and environmental constraints (such as obstacle avoidance), and is generally expressed in the form of a three-dimensional vector. The first interaction point is the key operation position on the active object, usually the most suitable part for grasping or applying force, such as the center of the handle or the balance center of gravity of a cup. The interaction primitive is a composite entity that describes the smallest operation unit of the robot arm. It contains a clear interaction point and the corresponding interaction direction. It constitutes the basic control instruction unit for the robot arm to perform actions and is the basic element for realizing complex actions (such as handling, splicing, and assembly).
[0068] Traverse the entire sequence of structured interaction instruction data and structured operation phases, repeat the above process, and generate a sequence of interaction primitives corresponding to each structured operation phase. Each interaction primitive consists of an interaction point and an interaction direction. The interaction point is the key interaction location of the active object, that is, the location where the robotic arm grasps. The interaction direction is the spatial direction vector that combines the action type and the interaction scene, that is, the interaction process after the action type is improved based on the scene. The interaction primitive sequence is a sequence formed by converting each stage in the structured operation phase sequence into a corresponding interaction primitive and arranging them in order. It fully expresses the decomposition and refinement of the user task and is the underlying instruction stream that the robotic arm controller can directly understand and execute.
[0069] For example, a sequence of structured manipulation stages involves picking up a blue cube located at (0.5, 0.3, 0.1) on a workbench. Specific visual features include blue, cube-shaped, 0.05-meter-edge, a top center 0.025 meters higher than the bottom center, and precise spatial coordinates: the center is located at (0.5, 0.3, 0.1). The active object is determined to be the blue cube. Based on the action type (pick up) and the cube's geometry (cube), the interaction point is determined to be the center of the cube's top surface, with coordinates of (0.5, 0.3, 0.1 + 0.025) = (0.5, 0.3, 0.125) meters. An obstacle (such as another small building block) is detected 0.08 meters in front of the cube (along the positive X-axis), and the target grasping direction is vertically downward. The interaction direction vector is corrected to point from the center of the block's top surface (0.5, 0.3, 0.125) to a safe point slightly behind (along the negative X-axis), such as (0.48, 0.3, 0.125), which can be represented as (-0.02, 0, 0). This direction vector instructs the robotic arm's end-effector to first approach the block from behind, then descend vertically to grasp it, avoiding collisions with obstacles ahead. An interaction primitive is constructed, consisting of the interaction point (0.5, 0.3, 0.125) and the interaction direction vector (-0.02, 0, 0).
[0070] By breaking down each structured operation phase into interaction primitives containing precise interaction points and interaction directions, the accuracy and safety of the robot arm's control instructions are improved. The interaction points ensure that the robot arm can accurately physically interact with key locations on the target object, avoiding failure or damage caused by inaccurate grasping or operation positions. The interaction direction combines the action type and real-time environmental information, allowing the robot arm to approach and perform operations in the safest and most reasonable manner, effectively avoiding potential collision risks. The primitive-based control method decomposes complex user tasks into a series of simple, clear, and executable action units, which not only reduces the complexity of the control system, but also improves the reliability and success rate of task execution. The interaction primitive sequence provides a direct and reliable instruction stream for the robot arm controller, enabling the robot arm to efficiently and accurately complete complex user tasks in complex and dynamic environments, significantly improving the level of human-machine collaboration.
[0071] The active object is transferred to the passive object using the target manipulator according to the sequence of interaction primitives.
[0072] Furthermore, the present application further comprises the following steps:
[0073] The spatial posture of the active object corresponding to each interactive primitive in the interactive primitive sequence is monitored in real time to obtain a sequence of active object spatial posture monitoring results; correction instructions are identified based on the active object spatial posture monitoring result sequence, a correction instruction sequence is determined, and the correction instruction sequence is transmitted to the end effector of the target robotic arm for control correction.
[0074] Specifically, an interaction primitive sequence is a chronological sequence of multiple interaction primitives. Each interaction primitive contains the most basic information required to execute a subtask, such as the coordinates of key interaction points on the active object and the spatial direction vector required to execute the action. These primitives serve as a series of basic motion instructions to guide the robotic arm to complete the user's overall task. For each active object in the interaction primitive sequence, an RGB-D camera continuously captures image information of the active object. The precise spatial pose of the active object (e.g., a cube to be grasped) is calculated in real time at each moment. This consists of the coordinates (x, y, z) of its center point and the orientation of its surface. Position is typically expressed in three-dimensional coordinates (x, y, z), while pose is typically expressed in Euler angles or quaternions, describing the object's rotation relative to a reference coordinate system. The active object spatial pose monitoring result sequence is the pose monitoring result of the active object in each interaction primitive during task execution. Each data point is associated with a timestamp, recording the precise position and pose of the active object at that moment.
[0075] The sequence of spatial pose monitoring results for the active object is compared with its trajectory state in the sequence of interaction primitives. If a significant deviation is found between the actual pose and the expected pose, such as a slight positional shift or slight rotation of the cube due to a minor collision during the grasping process, the correction command recognition module is activated. Correction command recognition automatically identifies control commands that require posture or path adjustments to the robotic arm based on the error between the spatial pose monitoring results and the target task plan to ensure that the active object is accurately placed at the target location.
[0076] The generated correction instruction sequence is sent to the end effector of the robot arm for control. The end effector controller will parse these instructions and adjust its motion trajectory and posture in real time. It may make slight adjustments to the joint angles to compensate for the actual offset of the block, ensuring that the end effector can accurately and stably grasp the block. Finally, according to the guidance of the interaction primitive sequence, the block is accurately transferred and placed on the passive object, such as a specified mark point on the workbench with coordinates (0.8, 0.3, 0.1).
[0077] By continuously monitoring the active object's actual state, the robot can promptly detect and respond to unexpected situations during execution, such as slight object movement, posture changes, and even external interference. Modifying its command recognition and execution mechanisms allows the robot arm to not only follow a preset path but also dynamically adjust its behavior to adapt to the dynamics and uncertainties of the real world. This significantly improves the accuracy and success rate of task completion, especially when handling fragile objects or operating in complex dynamic environments.
[0078] In summary, the robotic arm control method based on multi-modal AI intelligent interaction provided by this application has the following beneficial effects:
[0079] By collecting multimodal interaction information and fusing the multimodal interaction information, a structured interaction instruction data sequence is obtained; based on the structured interaction instruction data sequence, user task intention is parsed to obtain a structured operation phase sequence of the target robotic arm, wherein each structured operation phase includes an active object, a passive object and an action type; combined with the structured interaction instruction data sequence, for each structured operation phase in the structured operation phase sequence, an interaction primitive centered on the active object is constructed to obtain an interaction primitive sequence, wherein each interaction primitive includes an interaction point and an interaction direction, wherein the interaction point is a key interaction position in the active object, and the interaction direction is a spatial direction vector combining the action type and the interaction scene; and the target robotic arm is used to transfer the active object to the passive object according to the interaction primitive sequence. In other words, by fusing and processing multimodal interaction information, analyzing user task intentions, and converting the user's abstract intentions into specific operation stages, for each structured operation stage, an interaction primitive centered on the active object is constructed, including interaction points and interaction directions. The target robotic arm transfers the active object to the passive object according to the interaction primitive sequence, realizing the conversion from user intentions to specific robotic arm actions, and improving the accuracy and efficiency of robotic arm operations.
[0080] In the second embodiment, based on the same inventive concept as the robot arm control method based on multi-mode AI intelligent interaction in the aforementioned embodiment, the present application also provides a robot arm control device based on multi-mode AI intelligent interaction, please refer to the attached Figure 2 , the robotic arm control device based on multi-mode AI intelligent interaction includes:
[0081] The information processing module 11 is used to collect multimodal interaction information and fuse the multimodal interaction information to obtain a structured interaction instruction data sequence; the task intention parsing module 12 is used to perform user task intention parsing based on the structured interaction instruction data sequence to obtain a structured operation phase sequence of the target robotic arm, wherein each structured operation phase includes an active object, a passive object and an action type; the interaction primitive construction module 13 is used to combine the structured interaction instruction data sequence and, for each structured operation phase in the structured operation phase sequence, construct an interaction primitive centered on the active object to obtain an interaction primitive sequence, wherein each interaction primitive includes an interaction point and an interaction direction, wherein the interaction point is a key interaction position in the active object, and the interaction direction is a spatial direction vector combining the action type and the interaction scene; the robotic arm control module 14 is used to use the target robotic arm to transfer the active object to the passive object according to the interaction primitive sequence.
[0082] Furthermore, the information processing module 11 in the robotic arm control device based on multimodal AI intelligent interaction is also used to: collect user voice commands through a microphone array, and input the user voice commands into a voice recognition module for semantic analysis to extract a keyword information sequence; capture user gestures through a depth camera or an inertial sensor, perform spatial posture analysis and recognition on the user gestures, and obtain action intention information; collect three-dimensional scene image information of the current operation area through an RGB-D camera to obtain visual feature information and spatial coordinate data of the operation object in space; and summarize the keyword information sequence, action intention information, visual feature information, and spatial coordinate data into multimodal interaction information.
[0083] Furthermore, the information processing module 11 in the robotic arm control device based on multi-modal AI intelligent interaction is also used to: divide the keyword information sequence into front-to-back associations according to the time sequence to obtain multiple divided keyword information subsequences; obtain the timestamp of the action intention information, supplement the multiple divided keyword information subsequences with action intentions, and obtain multiple supplemented divided keyword information subsequences; combine the visual feature information and spatial coordinate data to supplement the multiple supplemented divided keyword information subsequences with scenes to obtain a structured interaction instruction data sequence.
[0084] Furthermore, the information processing module 11 in the robotic arm control device based on multi-modal AI intelligent interaction is also used to: perform first-order division node identification on the keyword information sequence according to a preset time interval constraint, and determine a first-order division node set; perform second-order division node identification on the keyword information sequence according to a preset semantic association constraint, and determine a second-order division node set; perform a union on the first-order division node set and the second-order division node set to obtain a fused division node set; and use the fused division node set to divide the keyword information sequence to obtain the multiple divided keyword information subsequences.
[0085] Furthermore, the task intention parsing module 12 in the robotic arm control device based on multi-modal AI intelligent interaction is also used to: traverse the structured interaction instruction data sequence according to the active object name, passive object name and action trajectory to extract the user task intention, and obtain the active object sequence, passive object sequence and action trajectory sequence; match the action trajectory sequence in the action type library of the target robotic arm respectively to obtain the action type sequence; map and associate the active object sequence, passive object sequence and action type sequence to construct the structured operation stage sequence.
[0086] Furthermore, the task intention analysis module 12 in the robotic arm control device based on multi-modal AI intelligent interaction is also used to: traverse the motion trajectory sequence to extract trajectory features to obtain a motion trajectory feature sequence; match the motion trajectory feature sequence with the historical motion trajectory features in the motion type library, and output the action type corresponding to the maximum matching similarity value to obtain the action type sequence.
[0087] Furthermore, the interaction primitive construction module 13 in the robotic arm control device based on multi-modal AI intelligent interaction is also used to: extract the first interaction scene of the first structured operation stage from the structured interaction instruction data sequence: extract the image information of the active object in the first structured operation stage based on the first interaction scene, perform interaction point recognition based on the image information, and determine the first interaction point; perform path obstacle correction on the action type of the first structured operation stage based on the first interaction scene to obtain the first interaction direction; associate the first interaction direction and the first interaction point to construct the first interaction primitive; traverse the structured interaction instruction data sequence and the structured operation stage sequence for analysis to construct the interaction primitive sequence.
[0088] Furthermore, the interaction primitive construction module 13 in the robotic arm control device based on multi-modal AI intelligent interaction is also used to: use a path obstacle identifier to identify the first interaction scene and the action type to determine a first path obstacle identification result; based on the first path obstacle identification result and the motion tolerance range of the target robotic arm, correct the action type to obtain a first interaction direction.
[0089] Furthermore, the robotic arm control module 14 in the robotic arm control device based on multi-modal AI intelligent interaction is also used to: perform real-time monitoring of the spatial posture of the active object corresponding to each interactive primitive in the interactive primitive sequence, and obtain a sequence of active object spatial posture monitoring results; perform correction instruction identification based on the active object spatial posture monitoring result sequence, determine a correction instruction sequence, and transmit the correction instruction sequence to the end effector of the target robotic arm for control correction.
[0090] The various embodiments in this specification are described in a progressive manner, and each embodiment focuses on the differences from other embodiments. Figure 1 The robotic arm control method based on multi-mode AI intelligent interaction and the specific examples in Example 1 are also applicable to the robotic arm control device based on multi-mode AI intelligent interaction in this embodiment. Through the above detailed description of the robotic arm control method based on multi-mode AI intelligent interaction, those skilled in the art can clearly understand the robotic arm control device based on multi-mode AI intelligent interaction in this embodiment, so for the sake of brevity of the specification, it will not be described in detail here.
[0091] The above description of the disclosed embodiments is intended to enable one skilled in the art to implement or use the present application. Various modifications to these embodiments will be readily apparent to one skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of the present application. Therefore, the present application is not limited to the embodiments shown herein, but is intended to conform to the widest scope consistent with the principles and novel features disclosed herein.
[0092] Obviously, for those skilled in the art, several improvements and modifications can be made to the present application without departing from the principles of the present application, and these improvements and modifications also fall within the scope of protection of the present application.
Claims
1. A robotic arm control method based on multi-mode AI intelligent interaction, characterized in that: include: Collecting multimodal interaction information and fusing the multimodal interaction information to obtain a structured interaction instruction data sequence; The user's task intention is parsed based on the structured interaction instruction data sequence to obtain a structured operation phase sequence of the target robotic arm, where each structured operation phase includes an active object, a passive object, and an action type. In combination with the structured interaction instruction data sequence, for each structured operation phase in the structured operation phase sequence, an interaction primitive centered on the active object is constructed to obtain an interaction primitive sequence, wherein each interaction primitive includes an interaction point and an interaction direction, the interaction point being a key interaction position in the active object, and the interaction direction being a spatial direction vector combining the action type and the interaction scene; transferring the active object to the passive object using the target manipulator according to the sequence of interaction primitives; The collection of multimodal interaction information includes: The microphone array collects user voice commands, and inputs the user voice commands into the speech recognition module for semantic analysis to extract keyword information sequences; Capturing user gestures through a depth camera or inertial sensor, performing spatial posture analysis and recognition on the user gestures to obtain action intention information; The RGB-D camera collects the three-dimensional scene image information of the current operation area and obtains the visual feature information and spatial coordinate data of the operation object in space; Aggregating the keyword information sequence, action intention information, visual feature information, and spatial coordinate data into multimodal interaction information; The multimodal interaction information is subjected to fusion processing to obtain a structured interaction instruction data sequence, including: Dividing the keyword information sequence into front-to-back associations according to the time sequence to obtain a plurality of divided keyword information subsequences; Obtaining a timestamp of the action intention information, performing action intention supplementation on the multiple division keyword information subsequences, and obtaining multiple supplemented division keyword information subsequences; Combining the visual feature information and the spatial coordinate data to perform scene supplementation on the multiple supplementary division keyword information subsequences to obtain a structured interactive instruction data sequence; The keyword information sequence is divided into a front-to-back correlation according to a time sequence to obtain a plurality of divided keyword information subsequences, including: According to a preset time interval constraint, the keyword information sequence is subjected to first-order partition node identification to determine a first-order partition node set; According to the preset semantic association constraints, the keyword information sequence is subjected to second-order partition node identification to determine a second-order partition node set; Calculating a union of the first-order partition node set and the second-order partition node set to obtain a fused partition node set; Using the fused partitioning node set to partition the keyword information sequence to obtain the plurality of partitioned keyword information subsequences; The user task intention is parsed based on the structured interaction instruction data sequence to obtain a structured operation phase sequence of the target robotic arm, wherein each structured operation phase includes an active object, a passive object, and an action type, including: Traversing the structured interaction instruction data sequence according to the active object name, the passive object name and the action trajectory to extract the user task intention, and obtaining the active object sequence, the passive object sequence and the action trajectory sequence; Matching the motion trajectory sequences in the motion type library of the target robotic arm to obtain a motion type sequence; Mapping and associating the active object sequence, the passive object sequence, and the action type sequence to construct the structured operation phase sequence; The step of combining the structured interaction instruction data sequence and constructing an interaction primitive centered on the active object for each structured operation phase in the structured operation phase sequence to obtain an interaction primitive sequence includes: Extracting a first interaction scenario of a first structured operation phase from the structured interaction instruction data sequence: extracting image information of the active object in the first structured operation phase based on the first interaction scene, performing interaction point recognition based on the image information, and determining a first interaction point; performing path obstacle correction on the action type of the first structured operation phase based on the first interaction scenario to obtain a first interaction direction; Associating the first interaction direction with the first interaction point to construct a first interaction primitive; The structured interactive instruction data sequence and the structured operation phase sequence are traversed for analysis to construct the interactive primitive sequence.
2. The robotic arm control method based on multi-mode AI intelligent interaction according to claim 1, characterized in that: Match the motion trajectory sequences in the motion type library of the target manipulator to obtain motion type sequences, including: Traversing the motion trajectory sequence to extract trajectory features and obtain a motion trajectory feature sequence; The action trajectory feature sequence is matched with the historical action trajectory features in the action type library respectively, and the action type corresponding to the maximum matching similarity is output to obtain the action type sequence.
3. The robotic arm control method based on multi-mode AI intelligent interaction according to claim 1, characterized in that: Performing path obstacle correction on the action type of the first structured operation phase based on the first interaction scenario to obtain a first interaction direction includes: Using a path obstacle identifier to identify the first interaction scenario and the action type, and determine a first path obstacle identification result; Based on the first path obstacle recognition result and the motion tolerance range of the target robotic arm, the action type is corrected to obtain a first interaction direction.
4. The robotic arm control method based on multi-mode AI intelligent interaction according to claim 1, characterized in that: Transferring the active object to the passive object using the target manipulator according to the interaction primitive sequence comprises: Performing real-time monitoring on the spatial posture of the active object corresponding to each interactive primitive in the interactive primitive sequence to obtain a sequence of active object spatial posture monitoring results; Correction instruction identification is performed according to the sequence of spatial posture monitoring results of the active object, a correction instruction sequence is determined, and the correction instruction sequence is transmitted to the end effector of the target robot arm for control correction.
5. A robotic arm control device based on multi-mode AI intelligent interaction, characterized in that: The steps for implementing the method for controlling a robotic arm based on multi-modal AI intelligent interaction according to any one of claims 1 to 4, wherein the robotic arm control device based on multi-modal AI intelligent interaction comprises: An information processing module is used to collect multimodal interaction information and perform fusion processing on the multimodal interaction information to obtain a structured interaction instruction data sequence; The task intention parsing module is used to parse the user's task intention based on the structured interaction instruction data sequence and obtain the structured operation phase sequence of the target robot arm, where each structured operation phase includes an active object, a passive object, and an action type; an interaction primitive construction module, configured to construct, in combination with the structured interaction instruction data sequence, an interaction primitive centered on the active object for each structured operation phase in the structured operation phase sequence, thereby obtaining an interaction primitive sequence, wherein each interaction primitive includes an interaction point and an interaction direction, wherein the interaction point is a key interaction position in the active object, and the interaction direction is a spatial direction vector that combines the action type and the interaction scene; A manipulator control module is configured to utilize the target manipulator to transfer the active object to the passive object according to the interaction primitive sequence.
Citation Information
Patent Citations
Mechanical arm control method and system based on multi-mode driving and storage medium
CN118752495A
Multi-modal feedback integration type knowledge-search enhanced robot control system
JP2025108598A