Industrial embodied intelligent assembly task processing method based on world action model
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-07-14
- Publication Date
- 2026-08-11
AI Technical Summary
[0003]然而,上述方案多依据当前观测或固定参数直接执行动作,缺乏对装配过程中状态连续演化的有效建模,难以及时识别位姿偏差、接触变化和异常工况,导致在动态扰动和多阶段接触约束下适应性不足,装配成功率与稳定性受限
[0040]第五方面,本申请实施例提供一种计算机程序产品,包括计算机程序,该计算机程序被处理器执行时实现上述提供的方法。
Smart Images

Figure CN122546953A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of embodied intelligence technology, and in particular to an industrial embodied intelligence assembly task processing method based on a world motion model. Background Technology
[0002] In the field of high-precision industrial assembly, existing embodied intelligent systems typically use teach programming, preset trajectories, and rule control to complete tasks such as grasping, alignment, insertion, and fixing.
[0003] However, the above-mentioned schemes mostly rely on current observations or fixed parameters to directly execute actions, lacking effective modeling of the continuous evolution of the state during the assembly process. It is difficult to identify pose deviations, contact changes and abnormal working conditions in a timely manner, resulting in insufficient adaptability under dynamic disturbances and multi-stage contact constraints, and limiting the assembly success rate and stability.
[0004] Therefore, how to improve the ability of embodied intelligence to predict the evolution of subsequent states of actions in complex industrial assembly processes, while taking into account assembly adaptability and closed-loop adjustment capabilities, has become an urgent technical problem to be solved. Summary of the Invention
[0005] This application provides an industrial embodied intelligent assembly task processing method based on a world action model to solve the aforementioned technical problems. This solution addresses the issues of insufficient perception of subsequent state evolution and difficulty in balancing assembly adaptability and closed-loop adjustment capabilities in complex assembly scenarios using world action models for industrial assembly task processing. It comprehensively considers the current multi-dimensional state information of the embodied intelligence, the task objective, and the overall impact after action execution to evaluate and select candidate actions, thereby improving the rationality of action decision-making and assembly control capabilities of the embodied intelligence under dynamic disturbances and multi-stage contact constraints.
[0006] In a first aspect, embodiments of this application provide an industrial embodied intelligent assembly task processing method based on a world action model, including:
[0007] Acquire the first multimodal feature data, which includes embodied intelligent visual state feature data, embodied intelligent ontological state feature data, and contact state feature data;
[0008] Obtain task objective information;
[0009] Based on the first multimodal feature data, task target information, and preset condition constraints, multiple first action commands are generated;
[0010] The first multimodal feature data, multiple first action commands and task target information are input into a preset world action prediction model to predict the changes in the embodied intelligence's visual state, part position and / or posture, contact state and force feedback after the embodied intelligence executes the first action command, and thus obtain the comprehensive score of each first action command.
[0011] Based on a comprehensive score, a second action instruction is determined from multiple first action instructions. The second action instruction is used to control the embodied intelligence to process industrial assembly tasks.
[0012] Output the second action command.
[0013] In one possible embodiment, the first multimodal feature data also includes historical action instruction information, and contact event information obtained after the embodied intelligence executes the historical action instruction information.
[0014] In one possible embodiment, acquiring first multimodal feature data includes:
[0015] Acquire multimodal data, which includes video data from multiple perspectives, part segmentation data, embodied intelligent joint state data, end effector position and attitude data, and actuator state data;
[0016] Embody intelligent visual state feature data is extracted from video data and part segmentation data from multiple perspectives. The embody intelligent visual state feature data includes scene global semantic features, part category and instance features, part initial position and posture features, and part relative position relationship features.
[0017] The embodied intelligent body state feature data is extracted from the embodied intelligent joint state data and the end effector position and posture data. The embodied intelligent body state feature data includes joint angle, joint angular velocity, joint torque, end effector six-degree-of-freedom pose, and end effector motion speed.
[0018] Contact state feature data is extracted from the actuator state data. The contact state feature data includes the three-dimensional contact force, three-dimensional torque of the end effector, the contact point position, the contact normal vector, and the contact duration.
[0019] By integrating embodied intelligent visual state feature data, embodied intelligent ontological state feature data, and contact state feature data, the first multimodal feature data is obtained.
[0020] In one possible embodiment, based on the first multimodal feature data, task target information, and preset condition constraints, a plurality of first action commands are generated, including:
[0021] Based on the first multimodal feature data, task target information and preset condition constraints, action condition information is generated. The action condition information includes at least one of the following: the relative relationship between the part and the task target, the current operation stage, the task advancement direction, and safety constraints.
[0022] The action condition information is input into the preset flow matching action generation model to obtain multiple first action instructions.
[0023] In one possible embodiment, the second action instruction includes a control parameter sequence and an execution action sequence. The control parameter sequence includes at least one of position control parameters, velocity control parameters, force control parameters, impedance control parameters, and compliance control parameters. The execution action sequence includes a plurality of first execution action instructions.
[0024] In one possible embodiment, it also includes:
[0025] The second multimodal feature data is acquired after the number of first execution action instructions in the embodied intelligence execution action sequence exceeds a first threshold.
[0026] Based on the second multimodal feature data and the world action prediction model, multiple second execution action instructions are determined;
[0027] Based on the unexecuted first execution action instruction and multiple second execution action instructions, multiple third execution action instructions are determined;
[0028] Output multiple third-party execution action instructions.
[0029] Secondly, embodiments of this application provide an industrial assembly task processing device based on a world motion model, comprising:
[0030] The acquisition module is used to acquire the first multimodal feature data, which includes embodied intelligent visual state feature data, embodied intelligent ontology state feature data, and contact state feature data.
[0031] The acquisition module is also used to acquire task target information;
[0032] The processing module is used to generate multiple first action commands based on the first multimodal feature data, task target information, and preset condition constraints;
[0033] The processing module is also used to input the first multimodal feature data, multiple first action commands and task target information into a preset world action prediction model, so as to predict the changes in the visual state of the embodied intelligence, the changes in the position and / or posture of the parts, the changes in the contact state and the changes in the force feedback after the embodied intelligence executes the first action command, and thus obtain the comprehensive score of each first action command.
[0034] The processing module is also used to determine a second action instruction from multiple first action instructions based on a comprehensive score. The second action instruction is used to control the embodied intelligent processing of industrial assembly tasks.
[0035] The output module is used to output the second action command.
[0036] Thirdly, embodiments of this application provide an electronic device, including: a memory and a processor;
[0037] The memory stores the instructions that the computer executes;
[0038] The processor executes computer execution instructions stored in memory, causing the processor to perform the methods described above.
[0039] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, are used to implement the methods provided above.
[0040] Fifthly, embodiments of this application provide a computer program product, including a computer program that, when executed by a processor, implements the method described above.
[0041] The industrial embodied intelligent assembly task processing method based on the world action model provided in this application acquires first multimodal feature data, including embodied intelligent visual state feature data, embodied intelligent body state feature data, and contact state feature data. It then generates multiple first action commands by combining task target information and preset condition constraints. The first multimodal feature data, multiple first action commands, and task target information are then input into a preset world action prediction model to predict changes in the embodied intelligent visual state, part position and / or posture, contact state, and force feedback after the embodied intelligent executes the first action command. The method obtains a comprehensive score for each first action command and can comprehensively evaluate the assembly state evolution that may be caused by candidate actions before the action is executed, and select a better second action command. This improves the adaptability, closed-loop adjustment capability, assembly success rate, and stability of embodied intelligence in complex assembly processes. Attached Figure Description
[0042] The accompanying drawings, which are incorporated in and form part of this specification, illustrate embodiments consistent with this application and, together with the description, serve to explain the principles of this application.
[0043] Figure 1 A flowchart illustrating the industrial embodied intelligent assembly task processing method based on a world motion model, provided in this application embodiment;
[0044] Figure 2A schematic diagram of the structure of an industrial embodied intelligent assembly task processing device based on a world motion model, provided in an embodiment of this application;
[0045] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0046] The accompanying drawings illustrate specific embodiments of this application, which will be described in more detail below. These drawings and descriptions are not intended to limit the scope of the concept in any way, but rather to illustrate the concept of this application to those skilled in the art through reference to particular embodiments. Detailed Implementation
[0047] Exemplary embodiments will now be described in detail, examples of which are illustrated in the accompanying drawings. When the following description relates to the drawings, unless otherwise indicated, the same numbers in different drawings denote the same or similar elements. The embodiments described in the following exemplary embodiments do not represent all embodiments consistent with this application. Rather, they are merely examples of apparatuses and methods consistent with some aspects of this application as detailed in the appended claims.
[0048] This technology relates to the fields of industrial embodied intelligent control and intelligent assembly, and is applicable to high-precision industrial assembly scenarios such as hole-shaft mating, connector insertion, and snap-fitting. Related systems typically include a vision sensing unit, an embodied intelligent body status acquisition unit, a contact or force feedback detection unit, and an embodied intelligent actuator to perform operations such as part gripping, alignment, insertion, and fixing.
[0049] In existing industrial production lines, embodied intelligence often employs teach-in programming, preset trajectories, or rule-based control methods, directly outputting execution actions based on current visual observations, pose information, or fixed control parameters. This type of solution can complete basic assembly processes in a structured environment. Its working method typically involves first identifying the target position, and then advancing the end effector according to a pre-set motion path to complete the assembly.
[0050] However, the assembly process is often accompanied by initial positional deviations of parts, abrupt changes in contact state, frictional variations, and environmental disturbances. Simply relying on current observations or fixed rules makes it difficult to characterize the continuous evolution of the state after the action is performed. Especially in tasks with strong contact constraints such as insertion and press-fitting, the system has difficulty in timely determining whether a certain action will cause lateral collisions, jamming, attitude mismatch, or abnormal force growth.
[0051] Furthermore, existing solutions generally lack the ability to compare the consequences of candidate actions, which means that when embodied intelligence deviates, it often relies on repeated trials or manual intervention to correct the error. This not only reduces the assembly cycle time but may also cause parts wear, assembly failure, or even equipment downtime, making it difficult to balance assembly success rate, stability, and dynamic adaptability.
[0052] In view of this, how to improve the ability of embodied intelligence to predict the subsequent state evolution of actions in complex assembly processes, and to select the best actions accordingly, has become an urgent technical problem to be solved. To address this problem, an industrial embodied intelligence assembly task processing method based on a world action model is proposed. After acquiring embodied intelligence visual state feature data, embodied intelligence body state feature data, and contact state feature data, multiple first action commands are generated by combining task target information and preset condition constraints. Then, the first multimodal feature data, multiple first action commands, and task target information are input into a preset world action prediction model to predict changes in the embodied intelligence's visual state, part position and / or posture, contact state, and force feedback after the embodied intelligence executes the first action commands. Finally, the second action command is determined and output based on the comprehensive score of each first action command, thereby improving the accuracy and adaptability of action selection in the assembly process.
[0053] The technical solution of this application and how the technical solution of this application solves the above-mentioned technical problems are described in detail below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments. The embodiments of this application will now be described with reference to the accompanying drawings.
[0054] Figure 1 This is a flowchart illustrating the industrial embodied intelligent assembly task processing method based on a world motion model, as provided in this application embodiment. The method provided in this application embodiment can be applied to any electronic device. Figure 1 As shown, the method includes:
[0055] S101. Obtain the first multimodal feature data, which includes embodied intelligent visual state feature data, embodied intelligent ontological state feature data, and contact state feature data.
[0056] In this step, the first multimodal feature data serves as a unified input representation for subsequent action generation, prediction, and scoring. It encompasses embodied intelligence visual state feature data, embodied intelligence ontological state feature data, and contact state feature data. The embodied intelligence visual state feature data characterizes the parts, assembly targets, end effectors, and their spatial relationships within the current assembly scenario; the embodied intelligence ontological state feature data characterizes the current motion state of the embodied intelligence; and the contact state feature data characterizes the contact relationship between the parts and the assembly targets. The first multimodal feature data is not simply a concatenation of raw data; rather, it integrates multi-source perception results through temporal alignment, spatial calibration, feature extraction, and fusion to form a unified feature representation that the model can process.
[0057] In practical implementation, the executing entity can be an industrial embodied intelligent controller, an edge computing unit, or a host computer communicating with the embodied intelligent controller. Embodied intelligent visual state feature data can be collected through visual perception modules positioned around the assembly station to obtain image information of the current part relative to the target assembly position, and to extract visual state features such as part position, target position, relative offset, and posture deviation. Embodied intelligent body state feature data can be directly read from the embodied intelligent controller bus, while contact state feature data can be obtained through contact perception modules installed on the end effector or tooling. After completing the acquisition of data from each source, unified time synchronization and feature fusion processing are required. Specifically, the system aligns visual data, embodied intelligent controller data, and contact perception data to the same time base, forming data slices at the same moment. Subsequently, the visual features, body features, and contact features are processed to form the first multimodal feature data, which is then used in subsequent steps. Based on the above processing, this step unifies the visual state, embodied intelligent body state, and contact state in the current assembly scenario into structured inputs, so that the generation of subsequent actions no longer depends solely on a single moment image or fixed control parameters, but can jointly characterize the continuous evolution basis of the assembly process based on multiple source states.
[0058] S102. Obtain task target information.
[0059] In this step, the task objective information indicates the assembly objective that the embodied intelligence needs to complete and provides directional constraints for subsequent action generation and selection. The task objective information can represent the target state, target position, target posture, or expected assembly result of the assembly task to be completed. It can also include the current step identifier, allowed assembly depth, insertion direction, pressing endpoint, target contact force range, and the threshold condition for determining completion. When combined with the first multimodal feature data, the task objective information enables the system not only to know the current state of the embodied intelligence but also to clearly define the next target result to be pursued, thus giving the subsequently generated candidate actions task-oriented guidance.
[0060] In practice, task target information can originate from the Manufacturing Execution System (MES), workstation control unit, teaching record database, or assembly recipe file. In one possible embodiment, after the workpiece enters the assembly station, the host scheduling system sends the current product model, assembly process number, and target parameter set to the embodied intelligent control unit. The target parameter set includes at least the target part identifier, assembly reference coordinate system, target insertion direction vector, target insertion depth, target posture tolerance, and maximum allowable contact force threshold. For hole-shaft mating tasks, the task target information can indicate that the shaft center should coincide with the hole center, the axial direction should be consistent with the hole axis direction, and the advancement should stop after reaching the specified depth. For connector mating tasks, the task target information can indicate that the plug normal and socket normal are aligned, the insertion stroke reaches a preset value, and the contact force change enters a stable range. For snap-fit pressing tasks, the task target information can indicate the pressing displacement, the snap-fit confirmation posture, and the mechanical feedback characteristics after locking.
[0061] Before being imported into the control flow, the task objective information can undergo normalization processing. Specifically, the target position can be represented as three-dimensional coordinates in the embodied intelligent base coordinate system, the target attitude can be represented as a quaternion, Euler angle, or rotation matrix, the target insertion direction can be represented as a unit vector, and the target contact force threshold can be represented as a scalar upper limit or a sub-direction threshold vector. For different product models, the system can convert discrete process numbers into a unified data structure through task template mapping for subsequent processing. If the current task has multiple objective constraints, such as requiring the assembly depth to reach a specified range, the contact force to not exceed the threshold, and the attitude error to remain within the tolerance, these objective parameters can be assigned weight coefficients in the task objective information for calculating the comprehensive score in the subsequent scoring stage. Through the above processing, this step outputs structured task objective information and sends it along with the current multimodal state into the subsequent action generation process, so that the formation of action candidates has clear process objectives and termination conditions.
[0062] S103. Based on the first multimodal feature data, task target information and preset condition constraints, generate multiple first action commands.
[0063] In this step, the first action instruction represents multiple candidate actions generated by the system, which are then evaluated one by one by the subsequent world action prediction model. Conditional constraints are used to limit the executability and safety of the generated first action instruction. The first action instruction can correspond to control intentions such as insertion, fine-tuning, retraction, rotation correction, advancement along the normal direction, compensation along the tangential direction, or remaining in standby. In terms of data format, it can be represented using an action control expression form suitable for embodied intelligence execution.
[0064] In practice, the system first reads the first multimodal feature data obtained in S101 and the task target information obtained in S102, and then retrieves the preset condition constraints corresponding to the current process from the constraint configuration library. Taking hole-shaft assembly as an example, the assembly direction constraint requires the principal component of the motion to be along the hole-shaft direction, the contact safety constraint requires the motion to meet the contact safety requirements, and the embodied intelligent motion range constraint requires each joint increment to meet the speed, acceleration, and workspace boundary limits. Based on this, the system constructs a candidate motion generation space, and the motion parameters include at least translational increments Δx, Δy, and Δz, and attitude increments Δrx, Δry, and Δrz.
[0065] In one possible embodiment, multiple first action commands are formed using a flow-matching-based generation strategy. Specifically, the system pre-trains an action generation network, enabling it to learn to evolve from an initial noisy action distribution to a target action distribution that satisfies the current state and task conditions. During online execution, the first multimodal feature data, task target information, and conditional constraints are encoded and input into the action generation network. The network outputs several sets of action parameter samples that satisfy the constraints. Each set of samples, after inverse normalization, corresponds to a first action command. For example, when visual detection indicates that the end of a part's shaft is laterally offset relative to the center of the hole and that the end posture has an angle of inclination, and the contact sensor detects a slight lateral force, the action generation network can generate multiple differentiated candidate actions. One set of commands involves lateral compensation and posture correction followed by small-step advancement; another set involves retraction followed by posture correction; and yet another set involves only tangential fine-tuning without advancement. This results in multiple sets of first action commands with different strategy characteristics.
[0066] In another possible embodiment, the system can also adjust the range of action generation based on the current task requirements, so that candidate actions have appropriate differences while ensuring execution safety and task relevance. The number of candidate actions can be set according to the control cycle and computing resources to balance prediction efficiency and action diversity.
[0067] To ensure the executable nature of the generated actions, the system performs constraint screening and correction before outputting the actions. Specifically, inverse kinematics is solved for each candidate action, eliminating actions that could lead to joint overreach, unreachable postures, or collisions between the end effector and known obstacles in the environment; actions that do not meet contact safety requirements or are clearly inconsistent with the task objective are pruned or discarded. The multiple first action instructions after screening are written into the candidate action queue and bound to corresponding state identifiers, then proceed to the subsequent prediction and scoring steps. Based on the above processing, this step, under the combined influence of the current multimodal state, task objective, and conditional constraints, generates multiple differentiated candidate actions that meet the executable requirements. This allows the system to compare the consequences of different behaviors, rather than blindly proceeding along a single fixed trajectory, thereby establishing a basis for action selection based on pose deviations, friction changes, and contact mutations during assembly.
[0068] S104. Input the first multimodal feature data, multiple first action commands and task target information into the preset world action prediction model to predict the changes in the embodied intelligence's visual state, part position and / or posture, contact state and force feedback after the embodied intelligence executes the first action command, and then obtain the comprehensive score of each first action command.
[0069] In this step, the world motion prediction model is used to jointly model the correspondence between the current state, candidate actions, and future results. Its output includes at least embodied intelligent visual state changes, part position and / or posture changes, contact state changes, and force feedback changes. Embodied intelligent visual state changes characterize the changes in the relative relationships between parts, assembly targets, and end effectors in the visual scene after the execution of candidate actions; part position and / or posture changes characterize the displacement and angular evolution of parts or end effectors in space; contact state changes describe whether there are signs of continuous contact, instantaneous collision, lateral contact, or jamming; force feedback changes represent the changing trends of contact force and torque after the execution of actions. The comprehensive score is used to quantify the quality of each first action command, which can be formed by combining multiple indicators such as assembly progress, pose error, contact force risk, jamming risk, and success probability.
[0070] In practical implementation, the pre-defined world action prediction model can be deployed on an edge GPU (Graphics Processing Unit) computing unit and trained offline using historical assembly data, simulation data, and trial operation data. In one possible embodiment, the model includes a visual evolution prediction subnetwork, a state transition prediction subnetwork, and a contact force prediction subnetwork. The visual evolution prediction subnetwork can be composed of a spatiotemporal prediction network, whose input is the current visual encoding features and candidate action encodings, and whose output is the visual features of the assembly area or the latent space representation of the future image under one or more future control steps. The state transition prediction subnetwork receives the first multimodal feature data and task target information, and outputs the part position increment, posture increment, and end-effector relative target deviation. The contact force prediction subnetwork outputs the contact force vector, torque vector, and contact state category probability within a future short time window based on the current contact state features and action parameters. For each first action command, the model performs forward inference once to obtain the corresponding future state prediction result.
[0071] The comprehensive score can be generated based on multiple computable indicators. For example, the system first calculates a target proximity score based on the predicted part position and / or posture changes. This score can be obtained by comparing the predicted pose error with the target pose in the task objective. Next, a risk score is calculated based on changes in contact state and force feedback. This risk score is related to the predicted peak contact force, lateral force ratio, jamming category probability, and abnormal torque growth rate. Simultaneously, an assembly progress score is calculated based on changes in embodied intelligent vision state, such as the degree of overlap between the predicted part boundary and the target area boundary, the change in insertion depth, and changes in occlusion relationships. An implementable scoring expression can be written as: Score = α × Pprogress - β × Epose - γ × Rcontact - δ × Rforce, where Pprogress represents the assembly progress indicator, Epose represents the predicted pose error, Rcontact represents the contact risk indicator, Rforce represents the force feedback risk indicator, and α, β, γ, and δ are weighting coefficients. These weights can be set by process experience or obtained through offline optimization using historical successful assembly data.
[0072] To ensure comparability of scoring results, the system can standardize various indicators of different candidate actions before synthesizing a comprehensive score. For example, pose error is normalized according to task tolerance, force feedback risk is normalized according to safety threshold, and jamming probability is directly mapped to the 0-1 range. If a candidate action performs well in assembly progress, but its predicted force feedback change shows a lateral force peak exceeding the threshold, the risk component of that action will increase significantly, thus lowering the comprehensive score. If the prediction result of a candidate action shows that the part position continuously approaches the target, the attitude deviation decreases, the contact state smoothly transitions from non-contact to target contact, and the force feedback change remains within the allowable range, then its comprehensive score will be higher than other candidate actions. Through the above-mentioned world action prediction and scoring process, the system conducts a forward-looking assessment of the consequences of each action before actual execution, thereby expanding the basis for action selection from current static observation to comparison of future state evolution, and directly addressing the problem of difficulty in timely judgment of lateral collisions, jamming, attitude mismatch, or abnormal force growth during assembly.
[0073] S105. Based on the comprehensive score, determine the second action instruction from multiple first action instructions. The second action instruction is used to control the embodied intelligent processing of industrial assembly tasks.
[0074] In this step, the second action instruction represents the target action selected from multiple candidate first action instructions, and is the action subsequently output to the embodied intelligence for execution. The determination process of the second action instruction is not a simple random selection, but rather a selection based on the comprehensive score obtained in S104 and necessary safety constraints, to ensure that the selected action simultaneously satisfies both task objective consistency and execution safety.
[0075] Industrial assembly tasks are core precision operation scenarios in intelligent manufacturing production processes, encompassing various contact-based precision operations such as hole-shaft docking, connector insertion, snap-locking, slide rail adaptation, and gear shaft assembly. These tasks are characterized by complex operating conditions, frequent contact and interaction, high assembly precision requirements, and low fault tolerance, directly determining the assembly accuracy, operational stability, and finished product quality of industrial products. To enable embodied intelligence to perform these tasks autonomously, accurately, and stably, a dedicated second action command can fully control the embodied intelligence to complete various industrial assembly operations, leveraging three core capabilities to upgrade and overcome the limitations of traditional assembly embodied intelligence. This action command first endows the embodied intelligence with predictive action capabilities, allowing the equipment to anticipate and evaluate the subsequent operational results of various alternative actions before actually executing the assembly operation. This avoids problems such as alignment deviations, part collisions, and assembly failures caused by blind operation, optimizing the operational logic from the source. Simultaneously, the command-driven embodied intelligence enables joint modeling of all-dimensional assembly information. It integrates diverse data such as visual image status during the operation, real-time part pose, component contact patterns, mechanical feedback trends, current assembly progress stage, and task success probability to construct a comprehensive and dynamic assembly scenario model, accurately perceiving changes in operating conditions. Based on this, the action commands specifically optimize the operational logic for various high-contact precision assembly tasks, effectively enhancing the adaptability of embodied intelligence in typical scenarios such as hole-shaft assembly, connector insertion, snap-fit assembly, slide rail assembly, and gear shaft assembly. This significantly improves the success rate of precision assembly tasks, strengthens operational stability under complex conditions, ensures efficient, precise, and continuous progress of industrial assembly processes, and meets the core needs of industrialized mass precision production.
[0076] In practice, the system first reads the comprehensive score corresponding to each first action instruction from the candidate action queue and sorts them from highest to lowest score. In one possible embodiment, the first action instruction with the highest comprehensive score is directly selected as the second action instruction. When the scores of multiple candidate actions are close, the system can also introduce additional decision logic, such as prioritizing actions with smaller predicted pose errors or actions with lower force feedback risks. This allows the action selection strategy to adapt to the requirements of the current task.
[0077] In another possible embodiment, the system performs a re-constraint screening on the comprehensive score before finally determining the second action instruction. Specifically, if a first action instruction has the highest comprehensive score, but its corresponding predicted contact force peak exceeds the maximum allowable threshold defined in the task target information, or its predicted attitude change will cause the end effector to approach the tooling restricted area, then the action is marked as unexecutable and removed from the candidate set; subsequently, the highest-scoring action is selected from the remaining actions as the second action instruction. The re-constraint screening can be performed using Boolean decision method, or a penalty function method can be used to further reduce the score of actions that violate safety constraints. For example, a safety constraint penalty term Penalty can be set. When the prediction result triggers any hard constraint, the corrected score Score′ = Score - Penalty, and the value of Penalty is greater than the upper limit of the current score interval to ensure that the violating action is not selected.
[0078] To improve continuous control stability, the system can also combine the direction and amplitude of actions executed in the previous control cycle to perform a smooth consistency check on the action with the highest score. If the difference in direction between the current optimal action and the action in the previous cycle exceeds a preset angle threshold, and the difference in scores between the two is lower than the stability determination threshold, the system can select a smoother candidate action as the second action command to reduce action jitter within short cycles. This consistency check is particularly suitable for scenarios requiring continuous fine-tuning, such as connector insertion and press-fitting. In one possible embodiment, the stability determination threshold can be set based on the score normalization interval, and the direction difference threshold can also be set according to control stability requirements to suppress unnecessary frequent reverse adjustments.
[0079] Through the above screening process, the system ultimately outputs a unique second action instruction and writes it into the execution buffer, awaiting issuance in the next step. Based on the above analysis, this step establishes a clear mechanism for determining the target action from multiple candidate actions through comprehensive scoring and ranking of candidate actions, safety re-constraint, and necessary smoothing and consistency screening. This enables embodied intelligence to select the execution action that matches the current assembly requirements even when there are pose deviations, contact uncertainties, and environmental disturbances, rather than relying on trial-and-error corrections.
[0080] S106, Output the second action instruction.
[0081] In this step, the output of the second action command indicates that the target execution action determined in S105 is sent to the embodied intelligent execution link to drive the embodied intelligent end effector to complete the assembly action of the current control cycle. The second action command can be output to the industrial embodied intelligent servo control interface, motion planning interface, or upper-level control unit, or it can be sent to the embodied intelligent controller after interface protocol conversion before execution. This step translates the prediction and screening results into actual control signals and constitutes the execution link of subsequent closed-loop control.
[0082] In practical implementation, the system converts the second action command from its internal unified representation into a control format recognizable by the embodied intelligent controller. For example, when the second action command is represented as an end-effector pose increment, the control unit first calculates the target end-effector pose based on the current embodied intelligent end-effector pose, then obtains the target increments for each joint through inverse kinematics, and generates the corresponding short-time trajectory in joint space via the trajectory interpolation module. When the second action command is represented as a velocity control quantity, the control unit directly generates end-effector linear velocity and angular velocity commands and sends them out through the embodied intelligent real-time bus. When the second action command simultaneously includes displacement and target contact force parameters, the controller can switch to impedance control or force-position hybrid control mode, ensuring that the propulsion process follows the mechanical constraints selected in the prediction phase. In one possible embodiment, the execution time window of the second action command can be set to one control cycle or multiple consecutive control cycles to balance assembly fine-tuning accuracy and control response speed.
[0083] During the output process, the control unit can also simultaneously record the issuance time, execution number, and associated status snapshot of the second action command to form a closed-loop feedback in the next control cycle. After the embodied intelligence executes the second action command, the visual perception unit, the embodied intelligence body status acquisition unit, and the contact detection unit continue to collect new status data, thereby forming new first multimodal feature data again, and re-entering the cycle from S101 to S106. For continuous action tasks such as insertion and pressing, this closed loop can be repeated in each cycle until the completion conditions in the task target information are met, such as the insertion depth reaching the threshold, the attitude error being less than the tolerance, and the contact force entering the target stable range. If the actual force feedback exceeds the protection limit or the visual status shows that the part is significantly deviating from the target area during execution, the control unit can suspend the current action execution and trigger the regeneration of candidate actions or execute a reversal action command to maintain the system operation within the controlled range.
[0084] In one possible embodiment, the second action command can be output to the upper-level control unit for confirmation before being sent to the embodied intelligent controller. The upper-level control unit is used to perform the review in debug mode, semi-automatic mode, or manual monitoring mode. After the review is passed, the second action command is then sent to the embodied intelligent controller for execution; in fully automatic mode, the control unit can directly complete the issuance. It should be understood that the above examples are merely illustrative and not limiting. The output interface, control format, and execution mode of the second action command can be adapted according to the industrial embodied intelligent control architecture, as long as its implementation method is consistent with the aforementioned action screening results.
[0085] Based on the above analysis, this application provides an industrial embodied intelligent assembly task processing method based on a world motion model, including acquiring first multimodal feature data, acquiring task target information, generating multiple first action commands based on the first multimodal feature data, task target information and preset condition constraints, inputting the first multimodal feature data, multiple first action commands and task target information into a preset world motion prediction model to predict changes in the embodied intelligent visual state, changes in part position and / or posture, changes in contact state and force feedback after the embodied intelligent executes the first action commands and obtaining a comprehensive score for each first action command, determining a second action command from multiple first action commands based on the comprehensive score, and outputting the second action command. In this embodiment, by fusing visual state, embodied intelligent body state, and contact state into unified first multimodal feature data, and generating multiple candidate actions under task objective information and conditional constraints, the system uses a world action prediction model to predict and score the future state evolution after each candidate action is executed. Before actual execution, the system can compare the impact of different actions on assembly progress, pose error, contact risk, and force feedback risk, and determine and execute the second action command accordingly. This allows embodied intelligent action selection to be based on the prediction of action consequences. For high-precision assembly scenarios such as hole-shaft mating, connector insertion, and snap-fit assembly, this method can continuously update action decisions based on the current assembly state, allowing lateral collisions, jamming, posture mismatch, and abnormal force growth during the assembly process to be evaluated before the action is issued, thereby improving the accuracy of action selection and dynamic adaptability during the assembly process. It should be understood that the above examples are merely illustrative and not limiting; other implementations consistent with the technical features in the claims can also be adopted in the embodiments of this application.
[0086] Based on the aforementioned embodiments, the first multimodal feature data further includes historical action instruction information and contact event information obtained after the embodied intelligence executes the historical action instruction information.
[0087] Historical action instruction information is used to record action instructions previously executed by the embodied intelligence in the current assembly task, and to establish associations with the corresponding execution time, execution sequence, and execution result to form a temporal context that can characterize the cumulative impact of previous actions. Contact event information is used to describe the contact state generated on the assembly object after the embodied intelligence executes historical action instructions. Contact event information can be generated by force / torque sensors, end effector tactile detection units, or motor current change signals, and is used to reflect whether contact has occurred, the duration of contact, and the risk of abnormal contact.
[0088] In its implementation, after executing historical action commands, the embodied intelligent controller continuously collects the three-dimensional contact force, three-dimensional torque, contact point position information, and contact duration of the end effector. This information is then time-aligned with the historical action command information to form a multimodal input with sequential dependencies. Subsequently, the controller converts the historical action command information into a computable sequence code and the contact event information into a contact state code. This code is then combined with the embodied intelligent visual state feature data and the embodied intelligent ontological state feature data, or mapped to a unified feature space, resulting in the first multimodal feature data containing assembly state features. If a recurrent neural network, a temporal Transformer, or a gating unit is used to encode the historical action command information, the sequential relationship of action execution and its cumulative impact on the current contact state can be preserved. If a graph neural network is used to model the relationship between the contact point and the assembly part, the changes in contact position and their association with the target assembly area can be further characterized. The specific model of the above encoding module can be selected based on the controller's computing power and task cycle time. In practical applications, other models of this component can also be selected, and this embodiment does not limit this selection.
[0089] This structure allows the first multimodal feature data to include not only the currently observed embodied intelligence state but also the contact consequences caused by previous actions. This enables the integration of historical probing, jamming, engagement, or disengagement processes into a unified state representation when generating the first action command and predicting world actions. Consequently, the system can determine subsequent action commands that better match the current contact conditions based on the continuous evolution of the assembly state, and the prediction model's inferences regarding collisions, changes in insertion force, and attitude shifts are more complete.
[0090] By adopting this specific implementation method, historical action instruction information and contact event information are simultaneously incorporated into multimodal feature expression, which can improve the completeness and temporal consistency of assembly state description, and provide a more sufficient input basis for action generation and action consequence prediction, thereby improving the accuracy of action selection and the reliability of state prediction in complex contact assembly scenarios.
[0091] Based on the foregoing embodiments, the first multimodal feature data is obtained, including:
[0092] Acquire multimodal data, which includes video data from multiple perspectives, part segmentation data, embodied intelligent joint state data, end effector position and attitude data, and actuator state data;
[0093] Embody intelligent visual state feature data is extracted from video data and part segmentation data from multiple perspectives. The embody intelligent visual state feature data includes scene global semantic features, part category and instance features, part initial position and posture features, and part relative position relationship features.
[0094] The embodied intelligent body state feature data is extracted from the embodied intelligent joint state data and the end effector position and posture data. The embodied intelligent body state feature data includes joint angle, joint angular velocity, joint torque, end effector six-degree-of-freedom pose, and end effector motion speed.
[0095] Contact state feature data is extracted from the actuator state data. The contact state feature data includes the three-dimensional contact force, three-dimensional torque of the end effector, the contact point position, the contact normal vector, and the contact duration.
[0096] By integrating embodied intelligent visual state feature data, embodied intelligent ontological state feature data, and contact state feature data, the first multimodal feature data is obtained.
[0097] The system utilizes multiple perspectives of video data acquired by industrial cameras mounted at different locations within the assembly station. Part segmentation data is obtained by a segmentation network that performs pixel-level separation of the workpiece, fixture, and end effector regions within the video frames. The embodied intelligent joint state data is acquired from feedback by joint encoders and actuators. End effector position and attitude data are output from forward kinematics calculations or external positioning devices. Actuator state data is acquired by torque sensors, six-dimensional force sensors, or tactile sensors. Scene-wide semantic features characterize the background structure, target station, and part layout within the assembly scene. Part category and instance features distinguish different parts and their current instances. Initial part position and attitude features describe the spatial state of the part after gripping or before assembly. Relative positional relationship features characterize the spatial relationships between parts and between parts and target holes / clamp positions. Joint angles, joint angular velocities, and joint torques characterize the kinematic and dynamic states of the embodied intelligent entity. The six-degree-of-freedom pose and motion velocity of the end effector characterize the current position, orientation, and motion trend of the end effector within the assembly space. The three-dimensional contact force, three-dimensional torque, contact point position, contact normal vector, and contact duration of the end effector are used to describe the force distribution, contact direction, and contact stability when the end effector contacts the workpiece.
[0098] In the specific fusion process, the embodied intelligent visual state feature data, embodied intelligent ontology state feature data, and contact state feature data can be mapped into vectors of a unified dimension via corresponding encoders, and then the first multimodal feature data can be formed by concatenation, weighted summation, or attention fusion. The encoder can be constructed using convolutional neural networks, temporal networks, or multilayer perceptual networks, and different weights can be assigned to different modalities according to the assembly task to maintain the consistency of the representation of various state information in a unified feature space. In practical applications, other encoder models can also be selected, and this application embodiment does not limit this.
[0099] Based on the above processing, the first multimodal feature data simultaneously contains three key types of information: assembly scenario, embodied intelligent entity, and contact interaction, which can serve as a unified input for subsequent action generation and state prediction. This data fusion method can link the spatial relationships of parts, the motion state of the end effector, and contact feedback, thereby making the state representation during the assembly process more complete. Adopting this method provides a more comprehensive state basis for subsequent action selection and improves the ability to represent complex assembly scenarios.
[0100] Based on the aforementioned embodiments, multiple first action commands are generated based on the first multimodal feature data, task target information, and preset condition constraints. This includes: generating action condition information based on the first multimodal feature data, task target information, and preset condition constraints. The action condition information includes at least one of the following: the relative relationship between the part and the task target, the current operation stage, the task advancement direction, and safety constraints. The action condition information is then input into a preset flow matching action generation model to obtain multiple first action commands.
[0101] Among them, the flow matching action generation model is based on the generation mechanism of continuous flow transformation. It maps the condition vector that satisfies the constraints into an action distribution and samples multiple candidate first action commands from it. The action commands can correspond to the displacement, attitude adjustment, insertion direction or contact correction amount of the embodied intelligent end effector.
[0102] In specific implementation, the first multimodal feature data may include at least one of the following: multi-view red-green-blue images, embodied intelligent joint state data, end effector pose data, and contact force sensor data. The action condition information is obtained by jointly encoding the first multimodal feature data and the task target information. Taking the hole-shaft mating scenario as an example, the system determines the relative deviation between the hole opening and the shaft end based on the multi-view images, determines the current posture reachability of the embodied intelligent device based on the joint state and end effector pose, and determines whether it is in the pre-contact, contact correction, or insertion advancement stage based on the contact force change. Then, the relative relationship between the part and the target, the current operation stage, the task advancement direction, and the safety constraints are encoded into a unified condition vector and input into the action generation model to output multiple first action commands that satisfy the constraints.
[0103] The aforementioned flow matching action generation model can be implemented using a neural network structure. Its input layer is connected to a conditional vector, and its output layer generates a sequence of action parameters, which may include the translational increment, rotational increment, and velocity limit parameters of the end effector. During training, the model learns the mapping relationship from the initial action distribution to the target action distribution. During inference, it generates multiple candidate actions that satisfy the current assembly semantics based on the action condition information, enabling embodied intelligence to obtain different action schemes within the same task context.
[0104] By unifying assembly status, task objectives, and conditional constraints into action condition information, and then having the flow matching action generation model output multiple first action instructions, the generated results can be kept consistent with the current relative relationship of parts, stage status, and safety boundaries, and provide diverse candidates for subsequent action screening, thereby improving the adaptability and stability of action generation in complex assembly scenarios.
[0105] In one possible implementation, the second action instruction includes a control parameter sequence and an execution action sequence. The control parameter sequence includes at least one of position control parameters, velocity control parameters, force control parameters, impedance control parameters, and compliance control parameters. The execution action sequence includes a plurality of first execution action instructions.
[0106] The second action command is used to uniformly represent the control method, control intensity, and action sequence of the embodied intelligent end effector. The control parameter sequence is used to map the target action into directly executable control quantities, and the execution action sequence is used to describe the specific action units implemented sequentially under the constraints of these control quantities. Position control parameters are used to limit the target pose of the end effector during insertion, alignment, or retraction; velocity control parameters are used to limit the motion propulsion rate; force control parameters are used to limit the upper limit of the force during the contact phase; impedance control parameters are used to set the equivalent stiffness, damping, and inertia; and compliance control parameters are used to impart a certain degree of compliance to the contact process. Multiple first execution action commands correspond to action units such as single-step movement, fine-tuning alignment, tentative insertion, contact retraction, or attitude correction, and are combined into an execution action sequence in a predetermined order.
[0107] In specific implementation, the second action command can be selected by the controller based on the comprehensive score output by the world action prediction model, and the control parameter sequence is bound to the execution action sequence to generate the final control message. The control parameter sequence can be written into the control command in the form of joint space parameters, Cartesian space parameters, or end effector force-position hybrid parameters. The execution action sequence identifies the start and end times, duration, and switching conditions of each first execution action command in chronological order. For hole-shaft mating or connector insertion scenarios, position control parameters can be used to set the insertion depth and axial offset compensation; speed control parameters can be used to set the low-speed search segment and fast retraction segment; force control parameters can be used to set the contact threshold and press-fit holding force; impedance control parameters can be used to set the contact stiffness curve; and compliance control parameters can be used to set the lateral disturbance absorption range. The above control parameters can be written from the internal parameter table of the embodied intelligent controller, servo driver parameters, or the upper-level task planning module. The controller can be a real-time control unit in an industrial embodied intelligent control cabinet. In practical applications, other models of the controller can also be selected; this application embodiment does not limit this.
[0108] In this structure, the control parameter sequence first constrains the execution boundary of the embodied intelligence, and the execution action sequence then provides the specific action unfolding method. Together, they constitute the second action command, enabling the embodied intelligence to complete continuous action switching and maintain action consistency during the assembly process according to the controlled parameters. Thus, when performing assembly tasks, the embodied intelligence can simultaneously obtain control strength information and action arrangement information, allowing the output command to directly drive the embodied intelligence to complete stable and high-precision assembly.
[0109] Based on the foregoing embodiments, it further includes:
[0110] The second multimodal feature data is acquired after the number of first execution action instructions in the embodied intelligence execution action sequence exceeds a first threshold.
[0111] Based on the second multimodal feature data and the world action prediction model, multiple second execution action instructions are determined;
[0112] Based on the unexecuted first execution action instruction and multiple second execution action instructions, multiple third execution action instructions are determined;
[0113] Output multiple third-party execution action instructions.
[0114] The second multimodal feature data is used to re-represent the current assembly state after the first execution action command from the embodied intelligent execution part. It can include updated embodied intelligent visual state feature data, embodied intelligent body state feature data, and contact state feature data, and can be obtained by synchronously acquiring and fusing data from the camera, force / torque sensor, joint encoder, and end effector state acquisition module. The first threshold is used to limit the triggering of re-acquisition and replanning when the number of executed first execution action commands reaches a preset critical value. This threshold can be set according to the complexity of the assembly task, target accuracy, and contact sensitivity.
[0115] In practical implementation, the world motion prediction model follows the aforementioned modeling approach based on multimodal input for state evolution prediction. Its input is second multimodal feature data, and its output consists of multiple second-execution action commands matching the current state. These second-execution action commands can be characterized using position control parameters, velocity control parameters, force control parameters, impedance control parameters, or compliance control parameters, and correspond to local corrections for subsequent remaining assembly actions. The unexecuted first-execution action commands and multiple second-execution action commands together constitute a candidate action set. The controller compares, filters, and recombines these two sets based on the remaining task objectives, the current contact state, and action reachability, obtaining multiple third-execution action commands, which are then output to the embodied intelligent execution end as a sequence of actions to be executed in subsequent closed-loop control. This processing method enables the embodied intelligence to dynamically update the remaining actions based on new multimodal states during execution, maintaining consistency between the assembly trajectory and the current actual working conditions, and improving its adaptability to mid-way deviations and contact disturbances.
[0116] Figure 2 A schematic diagram of the structure of the industrial embodied intelligent assembly task processing device based on a world motion model provided in this application embodiment is shown below. Figure 2 As shown, the industrial embodied intelligent assembly task processing device 20 based on the world motion model provided in this embodiment includes an acquisition module 201, a processing module 202, and an output module 203.
[0117] The acquisition module 201 is used to acquire the first multimodal feature data, which includes embodied intelligent visual state feature data, embodied intelligent ontology state feature data, and contact state feature data.
[0118] The acquisition module 201 is also used to acquire task target information;
[0119] The processing module 202 is used to generate multiple first action commands based on the first multimodal feature data, task target information and preset condition constraints;
[0120] The processing module 202 is also used to input the first multimodal feature data, multiple first action commands and task target information into a preset world action prediction model, so as to predict the changes in the visual state of the embodied intelligence, the changes in the position and / or posture of the parts, the changes in the contact state and the changes in the force feedback after the embodied intelligence executes the first action command, and thus obtain the comprehensive score of each first action command.
[0121] The processing module 202 is also used to determine a second action instruction from multiple first action instructions based on a comprehensive score. The second action instruction is used to control the embodied intelligent processing of industrial assembly tasks.
[0122] Output module 203 is used to output the second action command.
[0123] In one possible embodiment, the first multimodal feature data also includes historical action instruction information, and contact event information obtained after the embodied intelligence executes the historical action instruction information.
[0124] In one possible embodiment, acquiring first multimodal feature data includes:
[0125] Acquire multimodal data, which includes video data from multiple perspectives, part segmentation data, embodied intelligent joint state data, end effector position and attitude data, and actuator state data;
[0126] Embody intelligent visual state feature data is extracted from video data and part segmentation data from multiple perspectives. The embody intelligent visual state feature data includes scene global semantic features, part category and instance features, part initial position and posture features, and part relative position relationship features.
[0127] The embodied intelligent body state feature data is extracted from the embodied intelligent joint state data and the end effector position and posture data. The embodied intelligent body state feature data includes joint angle, joint angular velocity, joint torque, end effector six-degree-of-freedom pose, and end effector motion speed.
[0128] Contact state feature data is extracted from the actuator state data. The contact state feature data includes the three-dimensional contact force, three-dimensional torque of the end effector, the contact point position, the contact normal vector, and the contact duration.
[0129] By integrating embodied intelligent visual state feature data, embodied intelligent ontological state feature data, and contact state feature data, the first multimodal feature data is obtained.
[0130] In one possible embodiment, based on the first multimodal feature data, task target information, and preset condition constraints, a plurality of first action commands are generated, including:
[0131] Based on the first multimodal feature data, task target information and preset condition constraints, action condition information is generated. The action condition information includes at least one of the following: the relative relationship between the part and the task target, the current operation stage, the task advancement direction, and safety constraints.
[0132] The action condition information is input into the preset flow matching action generation model to obtain multiple first action instructions.
[0133] In one possible embodiment, the second action instruction includes a control parameter sequence and an execution action sequence. The control parameter sequence includes at least one of position control parameters, velocity control parameters, force control parameters, impedance control parameters, and compliance control parameters. The execution action sequence includes a plurality of first execution action instructions.
[0134] In one possible embodiment, it also includes:
[0135] The second multimodal feature data is acquired after the number of first execution action instructions in the embodied intelligence execution action sequence exceeds a first threshold.
[0136] Based on the second multimodal feature data and the world action prediction model, multiple second execution action instructions are determined;
[0137] Based on the unexecuted first execution action instruction and multiple second execution action instructions, multiple third execution action instructions are determined;
[0138] Output multiple third-party execution action instructions.
[0139] The industrial embodied intelligent assembly task processing device based on the world action model provided in this embodiment can execute the method provided in the above method embodiment. Its implementation principle and technical effect are similar, and will not be described in detail here.
[0140] Figure 3 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application. Figure 3 As shown, the electronic device 30 provided in this embodiment includes at least one processor 301 and a memory 302. Optionally, the device 30 further includes a communication component 303. The processor 301, memory 302, and communication component 303 are connected via a bus.
[0141] In a specific implementation, at least one processor 301 executes computer execution instructions stored in memory 302, causing at least one processor 301 to perform the above-described method.
[0142] The specific implementation process of processor 301 can be found in the above method embodiments, and its implementation principle and technical effect are similar. It will not be repeated here.
[0143] In the above embodiments, it should be understood that the processor can be a Central Processing Unit (CPU), or other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), etc. The general-purpose processor can be a microprocessor or any conventional processor. The steps of the method disclosed in this invention can be directly implemented by a hardware processor, or implemented by a combination of hardware and software modules within the processor.
[0144] The memory may include random access memory (RAM) and may also include non-volatile memory (NVM), such as at least one disk storage device.
[0145] The bus can be an Industry Standard Architecture (ISA) bus, a Peripheral Component Interconnect (PCI) bus, or an Extended Industry Standard Architecture (EISA) bus, etc. Buses can be categorized as address buses, data buses, control buses, etc. For ease of illustration, the buses shown in the accompanying drawings are not limited to a single bus or a single type of bus.
[0146] This application also provides a computer program product, including a computer program that, when executed by a processor, implements the above-described method.
[0147] This application also provides a computer-readable storage medium storing computer-executable instructions, which, when executed by a processor, implement the above-described method.
[0148] The aforementioned readable storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as static random access memory (SRAM), electrically erasable programmable read-only memory (EEPROM), erasable programmable read-only memory (EPROM), programmable read-only memory (PROM), read-only memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk. The readable storage medium can be any available medium accessible to a general-purpose or special-purpose computer.
[0149] An exemplary readable storage medium is coupled to a processor, enabling the processor to read information from and write information to the readable storage medium. Of course, the readable storage medium can also be a component of the processor. The processor and the readable storage medium can reside in an Application Specific Integrated Circuit (ASIC). Alternatively, the processor and the readable storage medium can exist as discrete components in the device.
[0150] The division of units is merely a logical functional division; in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be indirect coupling or communication connection through some interfaces, devices, or units, and may be electrical, mechanical, or other forms.
[0151] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0152] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0153] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0154] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0155] Finally, it should be noted that other embodiments of the invention will readily occur to those skilled in the art upon consideration of the specification and practice of the invention disclosed herein. This invention is intended to cover any variations, uses, or adaptations of the invention that follow the general principles of the invention and include common knowledge or customary techniques in the art not disclosed herein, and is not limited to the precise structures described above and shown in the accompanying drawings, and various modifications and changes can be made without departing from its scope. The scope of the invention is limited only by the appended claims.
Claims
1. A method for processing industrial embodied intelligent assembly tasks based on a world motion model, characterized in that, include: Acquire first multimodal feature data, the multimodal feature data including embodied intelligent visual state feature data, embodied intelligent ontological state feature data and contact state feature data; Obtain task objective information; Based on the first multimodal feature data, the task target information, and preset condition constraints, multiple first action commands are generated, including: Based on the first multimodal feature data, the task target information, and the preset condition constraints, action condition information is generated. The action condition information includes at least one of the following: the relative relationship between the part and the task target, the current operation stage, the task advancement direction, and safety constraints. The action condition information is input into a preset flow matching action generation model to obtain multiple first action instructions; The first multimodal feature data, the multiple first action commands, and the task target information are input into a preset world action prediction model to predict changes in the embodied intelligence's visual state, part position and / or posture, contact state, and force feedback after the embodied intelligence executes the first action command, thereby obtaining a comprehensive score for each first action command. Based on the comprehensive score, a second action instruction is determined from the plurality of first action instructions, and the second action instruction is used to control the embodied intelligent processing industrial assembly task. Output the second action command.
2. The method according to claim 1, characterized in that, The first multimodal feature data also includes historical action command information, and contact event information obtained after the embodied intelligence executes the historical action command information; acquiring the first multimodal feature data includes: Acquire multimodal data, which includes video data from multiple perspectives, part segmentation data, embodied intelligent joint state data, end effector position and attitude data, and actuator state data; Embody intelligent visual state feature data is extracted from the video data from the multiple perspectives and the part segmentation data. The embody intelligent visual state feature data includes scene global semantic features, part category and instance features, part initial position and posture features, and part relative position relationship features. The embodied intelligent body state feature data is extracted from the embodied intelligent joint state data and the end effector position and posture data. The embodied intelligent body state feature data includes joint angle, joint angular velocity, joint torque, end effector six-degree-of-freedom pose, and end effector motion speed. Contact state feature data is extracted from the actuator state data. The contact state feature data includes the end effector's three-dimensional contact force, three-dimensional torque, contact point position, contact normal vector, and contact duration. The first multimodal feature data is obtained by fusing the embodied intelligent visual state feature data, the embodied intelligent body state feature data, and the contact state feature data.
3. The method according to claim 1, characterized in that, The second action instruction includes a control parameter sequence and an execution action sequence. The control parameter sequence includes at least one of position control parameters, velocity control parameters, force control parameters, impedance control parameters, and compliance control parameters. The execution action sequence includes a plurality of first execution action instructions.
4. The method according to claim 3, characterized in that, Also includes: Acquire second multimodal feature data, which is collected after the number of first execution action instructions in the execution action sequence executed by the embodied intelligence exceeds a first threshold; Based on the second multimodal feature data and the world action prediction model, multiple second execution action commands are determined; Based on the unexecuted first execution action instruction and the plurality of second execution action instructions, a plurality of third execution action instructions are determined; Output the multiple third execution action instructions.
5. An industrial assembly task processing device based on a world motion model, characterized in that, include: The acquisition module is used to acquire first multimodal feature data, which includes embodied intelligent visual state feature data, embodied intelligent ontology state feature data, and contact state feature data. The acquisition module is also used to acquire task target information; The processing module is configured to generate multiple first action commands based on the first multimodal feature data, the task target information, and preset condition constraints, including: Based on the first multimodal feature data, the task target information, and the preset condition constraints, action condition information is generated. The action condition information includes at least one of the following: the relative relationship between the part and the task target, the current operation stage, the task advancement direction, and safety constraints. The action condition information is input into a preset flow matching action generation model to obtain multiple first action instructions; The processing module is further configured to input the first multimodal feature data, the multiple first action commands and the task target information into a preset world action prediction model, so as to predict the changes in the visual state of the embodied intelligence, the changes in the position and / or posture of the parts, the changes in the contact state and the changes in the force feedback after the embodied intelligence executes the first action command, and thus obtain a comprehensive score for each first action command. The processing module is further configured to determine a second action instruction from the plurality of first action instructions based on the comprehensive score; The output module is used to output the second action instruction.
6. An electronic device, characterized in that, include: Memory, processor; The memory stores computer-executed instructions; The processor executes computer execution instructions stored in the memory, causing the processor to perform the method as described in any one of claims 1-4.
7. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer-executable instructions, which, when executed by a processor, are used to implement the method as described in any one of claims 1-4.
8. A computer program product, characterized in that, Includes a computer program that, when executed by a processor, implements the method described in any one of claims 1-4.