Reliability analysis method and system of body model, medium, equipment and product

By analyzing the consistency and logical rationality of videos of target intelligent agents performing operational tasks generated by embodied models, the problem of reliability assessment in the field of embodied intelligence is solved, and the performance quality of intelligent agents in actual tasks is improved.

CN120449922BActive Publication Date: 2025-12-30AGIBOT INNOVATION (SHANGHAI) TECHNOLOGY CO LTD
View PDF 1 Cites 0 Cited by

Patent Information

Application Number
CN202510538299.4
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-04-25
Publication Date
2025-12-30
Estimated Expiration
2045-04-25

AI Technical Summary

Technical Problem

In the field of embodied intelligence, existing technologies struggle to effectively measure the reliability of embodied models, resulting in agents failing to meet performance requirements in real-world tasks, particularly in physical interactions or long-term tasks where logical breaks and insufficient modeling of physical laws exist.

Method used

By analyzing target videos of the target intelligent agent performing operational tasks generated by the embodied model, motion consistency, visual consistency, and logical rationality analyses are performed to obtain corresponding index values ​​for evaluating the reliability of the embodied model. Specific methods include trajectory consistency analysis, dynamic consistency analysis, agent consistency analysis, overall consistency analysis, and logical rationality analysis, and a multimodal large language model is used for evaluation.

Benefits of technology

It enables reliability assessment of embodied models in real-world task scenarios, identifies and optimizes potential problems in embodied models, improves the performance quality of intelligent agents in real-world tasks, and meets the needs of application scenarios.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120449922B_ABST
    Figure CN120449922B_ABST
Patent Text Reader

Abstract

The embodiment of the application provides a kind of reliability analysis method and system of embodied model, medium, equipment and product, the method comprises: receiving the target video of target agent execution operation task generated by the embodied model;One or more of motion consistency analysis, visual consistency analysis and logic rationality analysis are carried out on the target video, to obtain one or more of motion consistency index value, visual consistency index value and logic rationality index value;Based on one or more of the motion consistency index value, the visual consistency index value and the logic rationality index value, determine the reliability analysis result of the embodied model.The embodiment of the application can effectively evaluate the reliability of embodied model in actual task scene by analyzing the target video of target agent execution operation task generated by embodied model, to obtain key index.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of embodied intelligence technology, and in particular to methods, systems, media, devices and products for reliability analysis of embodied models. Background Technology

[0002] In the field of embodied AI, especially in applications involving physical interaction or long-term tasks, the relevant evaluation methods are difficult to measure the reliability of embodied models, which directly affects the behavior of intelligent agents in actual tasks, and may result in the behavior of intelligent agents in actual tasks failing to meet the requirements.

[0003] Based on this, embodiments of this application provide a reliability analysis method and system, medium, device and product for embodied models to improve related technologies. Summary of the Invention

[0004] The purpose of this application is to provide a method, system, medium, device and product for reliability analysis of embodied models. By analyzing the target video of the target intelligent agent generated by the embodied model performing the operation task, key indicators can be obtained, which can effectively evaluate the reliability of the embodied model in actual task scenarios.

[0005] The objective of this application embodiment is achieved using the following technical solutions:

[0006] In a first aspect, embodiments of this application provide a reliability analysis method for an embodied model. The method includes: receiving a target video generated by the embodied model of a target agent performing an operation task; performing one or more of motion consistency analysis, visual consistency analysis, and logical rationality analysis on the target video to obtain one or more of motion consistency index values, visual consistency index values, and logical rationality index values; and determining the reliability analysis result of the embodied model based on one or more of the motion consistency index values, visual consistency index values, and logical rationality index values.

[0007] In some embodiments, the process of performing motion consistency analysis on the target video includes: performing one or more of trajectory consistency analysis and dynamic consistency analysis on the target video to obtain one or more of trajectory consistency index values ​​and dynamic consistency index values; and determining the motion consistency index value based on one or more of the trajectory consistency index values ​​and the dynamic consistency index values.

[0008] In some embodiments, the process of performing trajectory consistency analysis on the target video includes: performing end effector detection on the target video and the reference video to obtain the target trajectory corresponding to the target agent and the reference trajectory corresponding to the reference agent, wherein the target trajectory and the reference trajectory are the spatial trajector trajector trajector of the corresponding agents; the reference video is a real video or a synthetic video of the reference agent performing the operation task, and the reference video meets the reliability requirements; and the trajectory consistency index value is determined based on the trajectory similarity between the target trajectory and the reference trajectory.

[0009] In some embodiments, the trajectory similarity is determined based on the bidirectional Hausdorff distance and / or normalized dynamic time warping distance between the target trajectory and the reference trajectory.

[0010] In some embodiments, the process of performing dynamic consistency analysis on the target video includes: processing the target trajectory based on a difference algorithm to obtain a target velocity sequence and a target acceleration sequence; processing the reference trajectory based on the difference algorithm to obtain a reference velocity sequence and a reference acceleration sequence; and calculating a dynamic consistency index value based on the velocity similarity between the target velocity sequence and the reference velocity sequence and / or the acceleration similarity between the target acceleration sequence and the reference acceleration sequence.

[0011] In some embodiments, the velocity similarity is determined based on the Wasserstein distance between the target velocity sequence and the reference velocity sequence; and / or, the acceleration similarity is determined based on the Wasserstein distance between the target acceleration sequence and the reference acceleration sequence.

[0012] In some embodiments, the process of performing visual consistency analysis on the target video includes: performing one or more of subject consistency analysis, overall consistency analysis, and appearance style analysis on the target video to obtain one or more of subject consistency index values, overall consistency index values, and appearance style index values; and determining the visual consistency index value based on one or more of the subject consistency index value, the overall consistency index value, and the appearance style index value.

[0013] In some embodiments, the process of performing subject consistency analysis on the target video includes: extracting subject image features of each frame in the target video; calculating the cosine similarity between the subject image features of the first frame and the subject image features of each frame to obtain a first similarity score; calculating the cosine similarity between the subject image features of adjacent frames to obtain a second similarity score; and determining the subject consistency index value based on the first similarity score and the second similarity score.

[0014] In some embodiments, the process of performing overall consistency analysis on the target video includes: extracting the overall image features of each frame in the target video; calculating the cosine similarity between the overall image features of the first frame and the overall image features of each frame to obtain an overall consistency index value.

[0015] In some embodiments, the process of performing appearance style analysis on the target video includes: inputting the target video into a specified model to generate free-format natural language description information; the specified model is a first multimodal large language model or an image-to-text generation model; inputting the natural language description information into a second multimodal large language model for style judgment, the style judgment including classification judgment based on a preset language question or numerical judgment mapping the natural language description information to a realism score; and determining the appearance style index value based on the style judgment result.

[0016] In some embodiments, the process of performing logical rationality analysis on the target video includes: using a third multimodal large language model to analyze the logical rationality of each operation step in the target video to obtain a rationality score for each operation step; and determining the logical rationality index value based on the rationality scores of all operation steps.

[0017] Secondly, embodiments of this application provide a reliability analysis system for embodied models, the system including a control module for performing the method described in any one of the first aspects.

[0018] Thirdly, embodiments of this application provide a computer-readable storage medium storing a computer program that, when executed by a processor, implements the method described in any one of the first aspects.

[0019] Fourthly, embodiments of this application provide a computer device, the computer device including a memory and a processor, the memory storing a computer program, and the processor executing the computer program to implement the method described in any one of the first aspects.

[0020] Fifthly, embodiments of this application provide a computer program product, the computer program product including a computer program, which, when executed by a processor, implements the method described in any one of the first aspects.

[0021] This application provides a method, system, medium, device, and product for reliability analysis of embodied models. It involves receiving a target video of a target agent performing an operation task generated by an embodied model; performing one or more of motion consistency analysis, visual consistency analysis, and logical rationality analysis on the target video to obtain one or more of these indices; and determining the reliability analysis result of the embodied model based on these indices. By performing one or more of these analyses on the target video of the target agent performing an operation task generated by the embodied model, the performance quality of the target agent performing the operation task in the target video can be evaluated. Based on one or more of these indices, the reliability of the embodied model is quantitatively evaluated, providing data support for optimizing the embodied model. As can be seen, the embodiments of this application can analyze one or more deviations in the physical laws, visual performance and task logic of the agent behavior (i.e. the behavior of the target agent in performing operational tasks) in the target video, realize the reliability assessment of the embodied model, thereby providing targeted guidance for the optimization of the embodied model, and thus improving the agent's ability to exhibit high-quality behavior in actual tasks, so as to better meet the needs of actual application scenarios. Attached Figure Description

[0022] The embodiments of this application are further described below with reference to the accompanying drawings and specific implementation details.

[0023] Figure 1 This is a flowchart illustrating a reliability analysis method for an embodied model provided in an embodiment of this application.

[0024] Figure 2 This is a schematic diagram of the framework of an embodied scene world model generation quality evaluation system provided in an embodiment of this application.

[0025] Figure 3 This is a schematic diagram comparing a reference trajectory and a target trajectory provided in an embodiment of this application.

[0026] Figure 4 This is a structural block diagram of a computer device provided in an embodiment of this application. Detailed Implementation

[0027] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of them. Based on the embodiments of this application, all other embodiments obtained by those skilled in the art without creative effort are within the protection scope of the embodiments of this application.

[0028] In the description of the embodiments of this application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Therefore, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of the embodiments of this application, "multiple" means two or more, unless otherwise explicitly specified.

[0029] Current evaluation benchmarks for generative models primarily focus on evaluating general content generation tasks, such as image, text, and video generation. These evaluations typically rely on specific metrics, such as FID (Frechet Inception Distance, used to assess the realism of generated images) and CLIP (Contrastive Language–Image Pre-training Score, which measures the alignment between text and images). These methods are highly effective in evaluating the quality of generated content, especially in static or non-interactive content generation.

[0030] In the field of embodied intelligence, related technologies have proposed specialized task-oriented model evaluation frameworks. These frameworks aim to assess the effectiveness of reinforcement learning strategies or low-level action control (such as a robotic arm grasping objects). For example, some frameworks evaluate the agent's performance under physical constraints by task completion efficiency (such as the rate of object handling); while others focus on how large multimodal models decompose complex instructions into specific action steps.

[0031] However, in some specific application scenarios, such as embodied intelligence applications involving physical interaction or long-term tasks, the above evaluation methods or evaluation frameworks are difficult to measure the reliability of embodied models, and have at least the following limitations: (1) Disruption in the generation logic of long-term tasks: When the embodied model drives the agent to perform multi-step operation tasks of the robotic arm, if the embodied model has logical planning defects, it is easy to cause the agent to have a disordered operation sequence. For example, the embodied model may incorrectly plan the operation sequence of "closing the refrigerator door first and then placing the items". As the execution carrier of the agent, the robotic arm will directly present the incoherence of the task logic, which will eventually lead to the failure of the task. (2) Insufficient modeling of physical laws: In the generated video, the motion trajectory of the object may violate Newtonian mechanics (such as rigid objects floating without support), and the relevant indicators cannot effectively detect such problems.

[0032] To address the shortcomings of existing embodied model evaluation methods, this application proposes a reliability analysis method for embodied models. This method analyzes target videos of target agents generated by embodied models performing operational tasks to obtain key indicators, which can effectively evaluate the reliability of embodied models in real-world task scenarios.

[0033] See Figure 1 , Figure 1 This is a flowchart illustrating a reliability analysis method for an embodied model provided in an embodiment of this application.

[0034] In order to improve the relevant technology, this application provides a reliability analysis method for embodied models, the method including steps S101 to S103.

[0035] Step S101: Receive the target video of the target intelligent agent performing the operation task generated by the embodied model.

[0036] Step S102: Perform one or more of motion consistency analysis, visual consistency analysis, and logical rationality analysis on the target video to obtain one or more of the motion consistency index value, visual consistency index value, and logical rationality index value.

[0037] Step S103: Determine the reliability analysis result of the embodied model based on one or more of the motion consistency index value, the visual consistency index value, and the logical rationality index value.

[0038] In some embodiments, the method can be executed by a reliability analysis system deployed on a cloud server, local computing device, or edge computing node. This method is applicable to a variety of practical application scenarios, such as embodied intelligence applications involving physical interaction or long-range tasks. Furthermore, the method is also applicable to other intelligent agent task scenarios requiring high reliability (or high quality), such as robot simulation, robot data acquisition, and industrial automation scenarios.

[0039] In some embodiments, an intelligent agent refers to an agent capable of perceiving the environment and taking actions to achieve a specific goal. The target intelligent agent can be a physical or digital intelligent agent. Physical intelligent agents may include, for example, bipedal robots (also known as humanoid robots), quadrupedal robots, or wheeled robots, which can interact with the environment through mechanical feedback and perform operational tasks in real-world scenarios (such as grasping or moving objects). Digital intelligent agents may include, for example, virtual robots, software agents, or NPC characters in game AI, which can perform operational tasks in virtual scenarios (such as multi-step path planning, virtual environment navigation, or performing robot operation tasks in a simulation environment).

[0040] In some embodiments, an embodied model can generate videos of a target agent performing tasks, with the video content typically derived from open-source datasets. For example, an embodied model can generate dynamic videos of a target agent performing tasks based on an input task description or instructions, showcasing the target agent's action sequences and its interactions with the environment. By analyzing the video content generated by the embodied model, its understanding of task logic, physical laws, and / or scene interactions can be evaluated.

[0041] In some embodiments, the embodied model can be an Embodied Scene WorldModel, which is a multimodal generative model that combines embodied intelligence with scene understanding capabilities, and is designed to simulate and predict the behavior of a target intelligent agent in a specific scene and its interaction with the environment.

[0042] In some embodiments, the video content generated by the embodied model can be derived from open-source datasets such as Agibot-World, covering various task contexts and types of manipulated objects to ensure data diversity. Manipulated objects can include, for example, furniture, appliances, rigid objects, and flexible objects. For instance, one or more of the following data selection criteria can be used to guide the embodied model in generating target videos of the target agent performing an operation task: significant difference between the task operation and the task context; each task scenario conforming to real-world applications; a clean task background; a clear sequence of sub-actions; and a unique manipulated object (to ensure clear task logic). Through this design, the generated videos not only demonstrate the target agent's behavioral process but also provide high-quality support for evaluating the embodied model's task understanding and execution capabilities. Furthermore, this method of video generation helps identify potential problems in task logic inference, physical law modeling, and dynamic interaction processing of the embodied model, thereby providing a better basis for optimizing the target agent's actual task performance.

[0043] See Figure 2 , Figure 2 This is a schematic diagram of the framework of an embodied scene world model generation quality evaluation system provided in this application embodiment. This system (also referred to as a reliability analysis system) can generate target videos of a target agent performing an operation task based on data selection criteria (e.g., data from the Agibot-World open-source dataset, task and background diversity, clean background, and clear sub-action execution order). The target videos are evaluated according to multiple evaluation dimensions to obtain evaluation indicators. These dimensions may include, for example, motion consistency, visual consistency, and logical rationality. Motion consistency can be measured by calculating the trajectory similarity and motion similarity between the reference trajectory (also referred to as the ground truth trajectory) and the target trajectory (also referred to as the generated trajectory) to determine whether the target trajectory conforms to general physical laws and task logic. Visual consistency can be achieved by fine-tuning the DINOv2backbone (a self-supervised learning-based visual feature extraction model) using the Agibot-World embodied operation dataset to give it the ability to understand the subject performing the operation task. Logical rationality can be evaluated by using a multimodal large language model to assess the embodied scene world model's ability to handle complex, long-term tasks. By evaluating the target agent's performance of tasks in the target video generated by the model according to multiple evaluation dimensions (such as motion consistency, visual consistency, and logical rationality), the quality of the target agent's behavior generated by the embodied model and its ability to understand the logic and physical laws of the task can be better measured.

[0044] In some embodiments, motion consistency analysis can be used to evaluate whether the behavior of a target agent performing an operational task in a target video conforms to physical laws. Motion consistency analysis may include one or more of trajectory consistency analysis and dynamic consistency analysis.

[0045] Trajectory consistency analysis can be used to measure the consistency of a target agent's motion trajectory by calculating the trajectory similarity. Dynamic consistency analysis can be used to measure the consistency of a target agent's dynamic characteristics by calculating the velocity similarity and / or acceleration similarity.

[0046] In some embodiments, visual consistency analysis can be used to evaluate whether the visual representation of a target agent and its environment in a target video is reasonable. Visual consistency analysis may include one or more of subject consistency analysis, overall consistency analysis, and appearance style analysis.

[0047] Subject consistency analysis can be used to assess whether the main image features of the target agent in a target video remain consistent over time. Overall consistency analysis can be used to assess whether the overall features in a target video remain consistent over time. Appearance style analysis can be used to evaluate the appearance style of a target video by assessing whether its appearance style meets realism requirements.

[0048] In some embodiments, logical rationality analysis can be used to evaluate whether the logic by which a target agent performs operational tasks in a target video is reasonable.

[0049] In some embodiments, features can be extracted from the target video and quantified using a deep learning model (e.g., a multimodal large language model) to obtain one or more of the following: motion consistency index, visual consistency index, and logical rationality index.

[0050] In some embodiments, if any one of the motion consistency index, visual consistency index, and logical rationality index is not within the corresponding index value range, a corresponding prompt message is output. The prompt message can indicate the index that has not met the standard (such as unreasonable motion, visual inconsistency, or logical error) to optimize the embodied model.

[0051] In some embodiments, methods such as weighted average, threshold comparison, or multidimensional scoring mapping can be used to process the corresponding index values ​​to obtain the reliability analysis results of the embodied model.

[0052] Reliability analysis results can be used to indicate the performance of the target agent in the target video generated by the embodied model in terms of motion consistency, visual consistency, and logical rationality, thereby assessing whether the target agent's performance in the target video conforms to physical laws, task logic, and scene realism requirements. The reliability analysis results can provide a reliable simulation basis for the actual deployment of the target agent and guide the continuous optimization of the embodied model.

[0053] For example, the values ​​of each indicator can be weighted and summed according to preset weights to obtain a comprehensive score, which can be used to characterize the reliability analysis results of the embodied model. The preset weights can be adjusted according to the importance of the specific application scenario. For example, in tasks involving physical interaction, the weight of the motion consistency indicator is higher than that of the visual consistency indicator, and also higher than that of the logical rationality indicator; while in tasks requiring high visual realism, the weight of the visual consistency indicator is higher than that of the motion consistency indicator, and also higher than that of the logical rationality indicator.

[0054] For example, each indicator value can be compared with its corresponding range, and the reliability analysis results of the embodied model can be obtained based on the comparison results. Specifically, the comparison can be performed according to a priority order set based on the importance of each indicator value. The selection of the indicator range can be set according to actual needs and is not specifically limited here.

[0055] As another example, the values ​​of each indicator can be mapped to a multi-dimensional scoring space, and the reliability level of the embodied model can be determined by classification or clustering algorithms to obtain the reliability analysis results. For example, the values ​​of motion consistency, visual consistency, and logical rationality indicators can be input into a pre-trained classification model to output a natural language description of the reliability level.

[0056] In the above embodiments, by performing one or more of motion consistency analysis, visual consistency analysis, and logical rationality analysis on the target video of the target agent performing the operation task generated by the embodied model, the performance quality of the target agent performing the operation task in the target video can be evaluated. Based on one or more of the motion consistency index value, visual consistency index value, and logical rationality index value, the reliability of the embodied model is quantitatively evaluated, and data support is provided for optimizing the embodied model. It can be seen that the embodiments of this application can analyze one or more deviations in the agent's behavior in the target video in terms of physical laws, visual performance, and task logic, realize the reliability assessment of the embodied model, thereby providing targeted guidance for the optimization of the embodied model, and thus improving the agent's ability to exhibit high-quality behavior in actual tasks, making it better meet the needs of actual application scenarios.

[0057] To effectively quantify and evaluate the motion consistency of agent behavior in a target video, in some embodiments, the process of performing motion consistency analysis on the target video may include: performing one or more of trajectory consistency analysis and dynamic consistency analysis on the target video to obtain one or more of trajectory consistency index values ​​and dynamic consistency index values; and determining the motion consistency index value based on one or more of the trajectory consistency index values ​​and the dynamic consistency index values.

[0058] For example, performing trajectory consistency analysis on a target video to obtain a trajectory consistency index value may include: determining the trajectory consistency index value by evaluating whether the trajectory corresponding to the target agent in the target video is consistent with a reference trajectory.

[0059] For example, performing dynamic consistency analysis on a target video to obtain a dynamic consistency index value may include: determining the dynamic consistency index value by evaluating whether the velocity distribution and / or acceleration distribution of the target agent in the target video is consistent with the reference velocity distribution and / or reference acceleration distribution.

[0060] For example, depending on the needs of the actual application scenario, one or more of the trajectory consistency index value and the dynamic consistency index value can be selected to determine the motion consistency index value. For instance, the trajectory consistency index value and the dynamic consistency index value can be weighted and summed according to preset weights to obtain a comprehensive motion consistency index value; or, in some cases, either the trajectory consistency index value or the dynamic consistency index value can be selected to determine the motion consistency index value. Furthermore, by setting a threshold, the motion consistency index value used to characterize whether the behavior of the target intelligent agent conforms to physical laws can be determined based on whether both the trajectory consistency index value and the dynamic consistency index value reach a preset threshold.

[0061] In the above embodiments, trajectory consistency analysis and dynamic consistency analysis are performed on the target video to quantitatively evaluate the rationality of its trajectory and motion dynamics, so as to obtain the motion consistency index value of the agent's behavior in the target video. This can more accurately quantify whether the behavior of the target agent conforms to the physical laws, and can effectively alleviate the problems caused by insufficient modeling of physical laws (such as rigid objects suspending without support). This quantitative evaluation provides reliable data support for the motion consistency analysis of agent behavior, thereby improving the reliability analysis effect of the embodied model.

[0062] To effectively quantify and evaluate the rationality of the trajectory of an agent's behavior in a target video, in some embodiments, the process of performing trajectory consistency analysis on the target video may include: performing end effector detection on the target video and a reference video to obtain the target trajectory corresponding to the target agent and the reference trajectory corresponding to the reference agent, respectively. The target trajectory is the spatial trajectory of the end effector of the target agent, and the reference trajectory is the spatial trajectory of the end effector of the reference agent. The reference video is a real video or a synthetic video of the reference agent performing the operation task, and the reference video meets reliability requirements. Based on the trajectory similarity between the target trajectory and the reference trajectory, the trajectory consistency index value is determined.

[0063] For example, the positions of the end effectors in the target video and the reference video can be extracted frame by frame to obtain the target trajectory corresponding to the end effector in the target agent and the reference trajectory corresponding to the end effector in the reference agent.

[0064] The end effector position can be extracted from video frames, for example, using object detection algorithms or keypoint detection algorithms. The target trajectory and reference trajectory can be stored separately in time-series format. The target trajectory reflects the motion path of the end effector of the target agent during task execution, while the reference trajectory reflects the motion path of the end effector of a reference agent in the same task or baseline scenario, and can be considered as the baseline trajectory. The end effector can be located at the end of the agent's robotic arm or other manipulator. Examples of end effectors include grippers, suction cups, or dexterous hands within a robotic arm.

[0065] A reference agent is an agent used to provide baseline behavior or trajectory. The reference agent can be, for example, a manually controlled robot, a rigorously trained humanoid robot, a robotic arm, or a virtual agent generated by a high-precision simulation environment. The trajectory data generated must conform to physical laws to ensure its reliability as a baseline.

[0066] Real-world video can be recorded by high-precision sensors (such as depth cameras or motion capture systems) showing a reference agent performing an operation task; synthetic video can be generated by a high-precision simulation environment, showing a simulated reference agent performing an operation task. Both real-world and synthetic videos of the reference agent performing an operation task must meet reliability requirements to provide a reliable benchmark for trajectory consistency analysis of the target agent.

[0067] See Figure 3 , Figure 3This is a schematic diagram comparing a reference trajectory and a target trajectory provided in an embodiment of this application. For example, by using a YOLO-World model fine-tuned based on the Agibot-World dataset, target detection can be performed on the end effector of the robotic arm in both a reference video (also known as the ground truth video) and a target video (also known as the generated video), thereby extracting the motion trajectories corresponding to the left and right hands of the robotic arm in the reference video as the reference trajectory (or ground truth trajectory), and the motion trajectories corresponding to the left and right hands of the robotic arm in the target video as the target trajectory (or generated trajectory). For ease of comparison, the reference trajectory and the target trajectory can be visualized in the same image. For example, in Figure 3 In the diagram, the dashed line represents the target trajectory, and the solid line represents the reference trajectory. The difference between the reference trajectory (solid line) and the target trajectory (dashed line) is displayed intuitively through visualization, which facilitates the analysis and evaluation of their consistency.

[0068] For example, the distance between the target trajectory and the reference trajectory can be calculated using one or more selected measurement methods (e.g., spatial distribution characteristic measurement and / or time series characteristic measurement), and the calculated distances can be standardized (e.g., by taking the reciprocal) to make the calculated distances positively correlated with trajectory consistency; the calculated distances are evaluated (assuming that more than one distance is calculated, a weighted summation or selective use of a single indicator can be used) to obtain the trajectory consistency index value.

[0069] Spatial distribution characteristics can be used to assess the similarity between a target trajectory and a reference trajectory in terms of spatial distribution. For example, the maximum deviation between the target trajectory and the reference trajectory in terms of spatial distribution can be quantified by calculating the Hausdorff distance. The Hausdorff distance defines the maximum distance between two sets, that is, the maximum distance from any point in one set to the nearest point in the other set.

[0070] Time series characteristic metrics can be used to evaluate the consistency of target and reference trajectories in terms of temporal distribution. For example, the similarity between target and reference trajectories in time series can be quantified by calculating Normalized Dynamic Time Warping (NDTW) distance or Dynamic Time Warping (DTW). NDTW is a distance metric method based on DTW, and the normalization process makes the results easier to interpret and compare.

[0071] In the above embodiments, by performing end effector detection on the target video and the reference video, the target trajectory and the reference trajectory are obtained respectively, and the trajectory consistency index value is determined based on the trajectory similarity between the two. This enables a precise quantitative evaluation of the trajectory rationality of the agent's behavior, and provides reliable data support for evaluating whether the motion trajectory of the target agent conforms to physical laws, thereby improving the reliability analysis effect of the embodied model.

[0072] To more accurately quantify the rationality of the agent's behavior trajectory in the target video, in some embodiments, the trajectory similarity can be determined based on the bidirectional Hausdorff distance and / or normalized dynamic time warping distance between the target trajectory and the reference trajectory.

[0073] For example, to more accurately assess the spatial similarity between a target trajectory and a reference trajectory, trajectory similarity can be determined based on the bidirectional Hausdorff distance (also known as the symmetric Hausdorff distance) between the two trajectories. The bidirectional Hausdorff distance is an extension of the traditional Hausdorff distance, simultaneously capturing both the maximum deviation of the target trajectory from the reference trajectory and the maximum deviation of the reference trajectory from the target trajectory. This allows for the discovery of anomalies or unreasonable portions in the target trajectory. By comprehensively considering the bidirectional maximum distance, the accuracy of the analysis is improved, facilitating the discovery of anomalous parts in the target agent's behavior.

[0074] For example, to more accurately assess the similarity between a target trajectory and a reference trajectory over time, trajectory similarity can be determined based on the NDTW distance between the two trajectories. NDTW, through non-linear alignment of two time series, allows for flexible matching of trajectories with different lengths or speeds, thereby minimizing the cumulative distance between them. For instance, when the action rhythms of the target trajectory and the reference trajectory are inconsistent, NDTW can effectively capture their deviations over time, facilitating the discovery of anomalous components in the target agent's behavior.

[0075] In the above embodiments, a single distance can be selectively used to determine the trajectory consistency index value for the bidirectional Hausdorff distance and NDTW distance between the target trajectory and the reference trajectory. Alternatively, a weighted sum of the bidirectional Hausdorff distance and NDTW distance can be performed to evaluate the consistency between the target trajectory and the reference trajectory from both spatial distribution characteristics and temporal series characteristics. This multi-dimensional comprehensive consideration not only effectively improves the accuracy of trajectory similarity analysis but also provides more reliable data support for assessing whether the behavior of the target agent conforms to physical laws, thereby improving the reliability analysis effect of the embodied model.

[0076] To effectively quantify and evaluate the motion dynamics of agent behavior in a target video, in some embodiments, the process of performing dynamic consistency analysis on the target video may include: processing the target trajectory based on a difference algorithm to obtain a target velocity sequence and a target acceleration sequence; processing the reference trajectory based on the difference algorithm to obtain a reference velocity sequence and a reference acceleration sequence; and calculating a dynamic consistency index value based on the velocity similarity between the target velocity sequence and the reference velocity sequence and / or the acceleration similarity between the target acceleration sequence and the reference acceleration sequence.

[0077] For example, the target trajectory and reference trajectory can be differentially analyzed using a difference algorithm to obtain the target velocity sequence and target acceleration sequence corresponding to the target trajectory, and the reference velocity sequence and reference acceleration sequence corresponding to the reference trajectory. A difference algorithm is a method for calculating the rate of change of a numerical sequence. In first-order difference, velocity information can be extracted by calculating the differences between adjacent positions of trajectory points; in second-order difference, further difference of the velocity sequence can extract acceleration information. In trajectory analysis, difference algorithms can be used to extract velocity and acceleration information from trajectory data (e.g., corresponding to one or more positional information) and quantify the similarity between the target trajectory and the reference trajectory in velocity or acceleration distribution. For example, given a series of trajectory points, first-order difference reveals the velocity at each step, while second-order difference shows the trend of velocity change (i.e., acceleration information).

[0078] In some embodiments, the velocity similarity may be determined based on the Wasserstein distance between the target velocity sequence and the reference velocity sequence; and / or, the acceleration similarity may be determined based on the Wasserstein distance between the target acceleration sequence and the reference acceleration sequence.

[0079] For example, the velocity similarity between the target velocity sequence and the reference velocity sequence, and / or the acceleration similarity between the target acceleration sequence and the reference acceleration sequence can be calculated based on the Wasserstein distance.

[0080] Wasserstein distance (also known as Earth Mover's Distance, EMD) is a method for measuring the difference between two probability distributions. It reflects the difference by calculating the minimum "work" required to transform one distribution into another, where the "work" is determined by both the quality of the movement and the distance traveled. In multidimensional space, Wasserstein distance can capture the geometric differences between different distributions, making it particularly suitable for distribution comparison tasks in fields such as image processing and natural language processing. For example, in dynamic consistency analysis, it can be used to quantify the similarity between a target trajectory and a reference trajectory in terms of velocity or acceleration distributions.

[0081] By using Wasserstein distance to measure the similarity between the target velocity sequence and the reference velocity sequence, and / or the similarity between the target acceleration sequence and the reference acceleration sequence, a dynamic consistency index value is obtained, which facilitates the discovery of unreasonable dynamic characteristics (such as sudden velocity changes, acceleration anomalies, etc.) in the target generated trajectory.

[0082] In the above embodiments, either velocity similarity or acceleration similarity can be selected to determine the dynamic consistency index value. Alternatively, the dynamic consistency index value can be determined based on both velocity similarity and acceleration similarity. For example, velocity similarity and acceleration similarity can be weighted and summed according to preset weights to obtain a comprehensive dynamic consistency index value. This improves the accuracy of the dynamic consistency assessment of the agent's behavior, thereby enhancing the reliability analysis effect of the embodied model.

[0083] To effectively quantify and evaluate the visual consistency of agent behavior in a target video, in some embodiments, the process of performing visual consistency analysis on the target video may include: performing one or more of subject consistency analysis, overall consistency analysis, and appearance style analysis on the target video to obtain one or more of subject consistency index values, overall consistency index values, and appearance style index values; and determining the visual consistency index value based on one or more of the subject consistency index value, the overall consistency index value, and the appearance style index value.

[0084] For example, performing subject consistency analysis on a target video to obtain a subject consistency index value may include: determining the subject consistency index value by extracting the subject image features of each frame in the target video and calculating the cosine similarity.

[0085] For example, performing an overall consistency analysis on the target video to obtain an overall consistency index value may include: determining the overall consistency index value by extracting the overall image features of each frame in the target video and calculating the cosine similarity.

[0086] For example, performing appearance style analysis on a target video to obtain appearance style index values ​​may include: evaluating the appearance style of the target video using a multimodal large language model (MLLM) to determine appearance style index values.

[0087] For example, depending on the needs of the actual application scenario, one or more of the following can be selected to determine the visual consistency index value: subject consistency index value, overall consistency index value, and appearance style index value. For instance, the subject consistency index value, overall consistency index value, and appearance style index value can be weighted and summed according to preset weights to obtain a comprehensive visual consistency index value; alternatively, one of the subject consistency index value, overall consistency index value, or appearance style index value can be selected to determine the visual consistency index value. Furthermore, the visual consistency index value used to characterize whether the target agent's behavior conforms to visual rationality can be determined based on whether the subject consistency index value, overall consistency index value, and appearance style index value all reach preset thresholds.

[0088] In the above embodiments, by determining the visual consistency index value based on one or more of the subject consistency index value, the overall consistency index value, and the appearance style index value, the rationality of the target agent's behavior can be quantitatively evaluated. This effectively alleviates problems caused by insufficient visual representation of the agent's behavior (such as excessive CG feel or lack of realism). This quantitative evaluation provides reliable data support for the visual consistency analysis of the agent, thereby improving the reliability analysis effect of the embodied model. Here, CG refers to Computer Graphics.

[0089] To effectively quantify and evaluate the subject consistency of agent behavior in a target video, in some embodiments, the process of performing subject consistency analysis on the target video may include: extracting the subject image features of each frame in the target video; calculating the cosine similarity between the subject image features of the first frame and the subject image features of each frame to obtain a first similarity score; calculating the cosine similarity between the subject image features of adjacent frames to obtain a second similarity score; and determining the subject consistency index value based on the first similarity score and the second similarity score.

[0090] For example, a visual feature extraction model can be used to extract subject image features from each frame of a target video. Subject image features can be used to reflect key visual information about the target agent. Subject image features may include, for example, key attributes such as the target agent's pose and visual representation. The visual feature extraction model could be, for example, DINOv2, which, after fine-tuning, focuses more on subject features within the embodied scene, thereby improving the accuracy of subject image feature extraction.

[0091] For example, a first similarity score, represented as a sequence, can be obtained by calculating the cosine similarity between the main image features of the first frame of the target video and the main image features of each subsequent frame. For instance, if the target video has N frames, the cosine similarity between the first frame and the i-th frame (i = 1, 2, ..., N) can be calculated, resulting in N first similarity scores. These N first similarity scores form a sequence according to the chronological order of the frames in the target video. The first similarity score can be used to reflect the global consistency of the target agent's main features throughout the entire video.

[0092] For example, a second similarity score, represented as a sequence, can be obtained by calculating the cosine similarity between the main image features of adjacent frames in a target video. For instance, if the target video has N frames, the cosine similarity between the i-th frame and the (i+1)-th frame (i = 1, 2, ..., N-1) can be calculated, resulting in N-1 second similarity scores. These N-1 scores form a sequence according to the chronological order of the frames in the target video. The second similarity score reflects the coherence of the target agent's main features between adjacent frames, avoiding unreasonable phenomena caused by frame jitter or abrupt changes.

[0093] The first similarity score, represented as a sequence, can be used to quantify the overall consistency of the target agent's subject features throughout the entire video time span. The second similarity score, also represented as a sequence, can be used to assess the temporal coherence of the target agent's subject features between adjacent frames. If either the first or second similarity score shows a significant decrease or abrupt change at certain frames of the target video, it may indicate that the subject image features of the target agent have undergone anomalous changes in the corresponding frames of the target video (e.g., viewpoint switching or incoherent actions).

[0094] In the above embodiments, the first similarity score and the second similarity score can be aggregated to determine the subject consistency index value. For example, the first and second similarity comprehensive scores can be generated by calculating the statistics (such as average, minimum, or weighted average) of the first and second similarity score sequences, respectively. The first and second similarity comprehensive scores are then weighted and summed using preset weights to obtain the subject consistency index value. It is evident that determining the subject consistency index value based on the first and second similarity scores provides crucial and accurate data support for the subject consistency analysis of the agent from both a global dimension (i.e., the overall consistency of the target agent's subject features throughout the entire video time span) and a local dimension (i.e., the temporal coherence of the target agent's subject features between adjacent frames), thereby improving the reliability analysis effect of the embodied model.

[0095] To effectively quantify and evaluate the overall consistency of agent behavior in a target video, in some embodiments, the process of performing an overall consistency analysis on the target video may include: extracting the overall image features of each frame in the target video; calculating the cosine similarity between the overall image features of the first frame and the overall image features of each frame to obtain an overall consistency index value.

[0096] For example, a visual language model can be used to extract overall image features from each frame of a target video. These overall image features can reflect the global visual representation of each frame in the target video. For instance, if the overall target video undergoes a sudden change (such as a scene transition), the overall image features will change significantly. Overall image features can include key attributes such as the overall layout of the scene, background information, and the relationship between the subject and its environment.

[0097] For example, a visual language model could be a CLIP (Contrastive Language–Image Pretraining) model. CLIP is a multimodal model that achieves cross-modal understanding between images and text by jointly training on image and text data. By training on a large number of image-text pairs, CLIP can perform various zero-shot transfer tasks, such as classification and retrieval, without additional training. It offers high flexibility, wide applicability, and reduces the need to tailor models for specific tasks.

[0098] For example, the overall image features of each frame in the target video can be extracted, and the cosine similarity between the overall image features of the first frame and the overall image features of each subsequent frame can be calculated to obtain multiple overall consistency scores. The multiple overall consistency scores are then aggregated (e.g., by taking the average or minimum value) to determine the overall consistency index value.

[0099] In the above embodiments, by extracting the overall image features of each frame in the target video and calculating the cosine similarity between the first frame and each subsequent frame, the overall consistency index value is finally determined. This quantitative evaluation provides key data support for the overall consistency analysis of the target video, provides a basis for the visual performance evaluation of the agent's behavior, and can thus improve the reliability analysis effect of the embodied model.

[0100] To effectively quantify and evaluate the appearance style of agent behavior in a target video, in some embodiments, the process of performing appearance style analysis on the target video may include: inputting the target video into a specified model to generate free-format natural language description information; the specified model is a first multimodal large language model or an image-to-text generation model; inputting the natural language description information into a second multimodal large language model for style judgment, the style judgment including classification judgment based on a preset language question or numerical judgment mapping the natural language description information to a realism score; and determining the appearance style index value based on the style judgment result.

[0101] For example, a target video can be input into a specified model, which then interprets the image content of each frame or keyframe in the target video to generate free-form natural language description information. For instance, if the target video contains a scene of a robotic arm grasping an object, the natural language description information generated by the specified model could be: "A robotic arm is grasping a showerhead; the background is a textured wall, but the image has a CG feel." The CG feel can be understood as the image being computer-generated.

[0102] The specified model can be either a Multimodal Large Language Model (MLLM) or an Image-to-Text Generation Model. MLLM is a large-scale language model that combines multiple modalities (such as text, images, and audio). It can not only understand textual information but also process and generate related data in various forms, such as images and sounds, achieving cross-modal information conversion and integration. An Image-to-Text Generation Model is a deep learning-based multimodal model that can convert input images into natural language text describing the image content.

[0103] The natural language description information generated by the specified model (e.g., MLLM) based on the target video can be in a free-format, meaning that the form and content of the natural language description information are not limited to a fixed structure or template, but express information in a flexible and natural way.

[0104] For example, classification based on preset language questions can be achieved by setting preset language questions (such as "Does it have a CG feel?") for the second multimodal large language model to classify and judge the natural language description information. For instance, if the natural language description information contains keywords such as "CG feel", "too perfect", or "lacking realism", the second multimodal large language model will judge that the target video has a CG feel.

[0105] For example, mapping the natural language description information to a numerical judgment of realism score can be achieved by using a second multimodal large language model. This model generates a realism score based on semantic features in the description information, combined with preset scoring rules or training data. The realism score can be used to quantify the appearance and style characteristics of the target video.

[0106] The appearance style index value is determined based on the style judgment result, and the style judgment result can be mapped to the corresponding appearance style index value. For example, the index value corresponding to the style judgment result of "meets the requirements of realism" is the first score, and the index value corresponding to the style judgment result of "has a CG feel" is the second score. The difference between the first score and the second score is greater than or equal to a preset difference threshold, such as 0.5 or 0.6. The first score is, for example, 0.8, and the second score is, for example, 0.2.

[0107] In the above embodiments, during the appearance style analysis of the target video, natural language description information is generated by the first multimodal large language model, and the second multimodal large language model is used to classify and judge or score the appearance style index value, thereby quantitatively evaluating the realism and style rationality of the target video, providing reliable data basis for evaluating the visual performance of the agent's behavior in the target video, and thus improving the reliability analysis effect of the embodied model.

[0108] In order to effectively perform logical rationality analysis on the behavior of agents in target videos, in some embodiments, the process of performing logical rationality analysis on the target video may include: using a third multimodal large language model to analyze the logical rationality of each operation step in the target video to obtain a rationality score for each operation step; and determining the logical rationality index value based on the rationality scores of all operation steps.

[0109] For example, the target video can be input into a third multimodal large language model. Based on the key information of each operation step performed by the target agent in the target video, the third multimodal large language model analyzes the logical rationality of each operation step to obtain a rationality score for each operation step.

[0110] For example, a third-mode multimodal large language model can be used to extract key information from the current operation step of a target agent performing an operation task in a target video. This information, combined with key information from preceding operation steps (i.e., one or more operation steps that occur before the current operation step), generates contextual information. Based on this contextual information, a logical rationality analysis is performed on the current operation step to obtain a rationality score. Key information about the operation step can include the target agent's actions, state, and / or the object being manipulated in that operation step. Contextual information can reflect the dependencies and logical coherence between operation steps.

[0111] For example, suppose the target video records a robotic arm grasping and moving a shower head. The execution sequence of the sub-actions of the operation task is as follows: the robotic arm moves from its initial position to near the shower head; the robotic arm grasps the shower head; and the robotic arm moves the shower head to the target position. If the operation steps are logically reasonable (e.g., the robotic arm moves smoothly from its initial position to near the shower head), the reasonableness score is the first score. If the operation steps are not logically reasonable (e.g., the robotic arm grasps the shower head without touching it), the reasonableness score is the second score. The difference between the first score and the second score is greater than or equal to a preset difference threshold, which can be, for example, 0.5 or 0.6. The first score is, for example, 0.9, and the second score is, for example, 0.1.

[0112] For example, the minimum value among the reasonableness scores of all operation steps can be selected as the logical reasonableness index value, or a weighted average of the reasonableness scores of all operation steps can be calculated according to different weights assigned based on the importance of the operation steps, and used as the logical reasonableness index value.

[0113] In the above embodiments, by analyzing the logical rationality of each operation step in the target video and determining the logical rationality index value based on the rationality score of all operation steps, the logical rationality of the agent's operation behavior in the target video can be effectively evaluated, thereby improving the reliability analysis effect of the embodied model.

[0114] This application also provides a reliability analysis system for embodied models, the system including a control module, the control module being used to execute the reliability analysis method for embodied models as described in any of the above embodiments.

[0115] This application also provides a computer-readable storage medium storing a computer program that, when executed by a processor, implements the reliability analysis method for the embodied model described in any of the above embodiments.

[0116] This application also provides a computer program product, which includes a computer program that, when executed by a processor, implements the reliability analysis method for the embodied model described in any of the above embodiments.

[0117] The computer program product may be a portable compact disc read-only memory (CD-ROM) and include program code, and may run on a terminal device, such as a personal computer. However, the computer program product of the embodiments of this application is not limited thereto, and the computer program product may be any combination of one or more computer-readable media.

[0118] See Figure 4 , Figure 4 This is a structural block diagram of a computer device provided in an embodiment of this application.

[0119] This application also provides a computer device, which includes a memory and a processor. The memory stores a computer program, and the processor executes the computer program to implement the reliability analysis method of any of the above-described embodied models.

[0120] The embodiments of this application do not limit the computer device, which may be, for example, a local computer device, a cloud computer device, a distributed computer device, etc.

[0121] The computer device may include: a memory 110, a processor 120, and a communication interface 130. The memory 110, the processor 120, and the communication interface 130 are connected through internal connection paths.

[0122] The memory 110 is used to store computer programs, which in some implementations may include code for implementing the methods of the embodiments of this application.

[0123] The processor 120 executes the computer program stored in the memory 110 to control the communication interface 130 to receive input data and information, and output operation results and other data. In some implementations, when the solutions of the embodiments of this application are implemented by software or firmware, the computer program used to implement the solutions of the embodiments of this application can be stored in the processor 120 and executed by the processor 120.

[0124] The memory 110 may be volatile or non-volatile, or may include both. The non-volatile memory may be read-only memory (ROM), programmable read-only memory (PROM), erasable programmable read-only memory (EPROM), electrically erasable programmable read-only memory (EEPROM), or flash memory. The volatile memory may be random access memory (RAM). It should be noted that the memory 110 described herein is intended to include, but is not limited to, any memory of these and other suitable types. As an example, the memory 110 includes random access memory (RAM), cache memory, and read-only memory (ROM). The memory 110 stores a computer program that can be executed by processor 120, causing processor 120 to implement the steps of the reliability analysis method for the embodied model described above.

[0125] The processor 120 can be a central processing unit (CPU), or it can be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, etc. The general-purpose processor can be a microprocessor, or the processor 120 can be any conventional processor.

[0126] In implementation, each step of the above method can be completed by the integrated logic circuitry of the hardware in the processor 120 or by instructions in software form. The method disclosed in the embodiments of this application can be directly implemented by the hardware processor, or by a combination of hardware and software modules in the processor 120. The software modules can be located in mature storage media in the art, such as random access memory, flash memory, read-only memory, programmable read-only memory, electrically erasable programmable memory, or registers. This storage medium is located in the memory 110, and the processor 120 reads the information in the memory 110 and, in conjunction with its hardware, completes the steps of the above method. To avoid repetition, detailed descriptions are not provided here.

[0127] In some implementations, in addition to the hardware units described above, computer devices may also include software modules, such as operating systems, basic input / output systems (BIOS), and application software.

[0128] An operating system is used to manage the hardware and / or software resources of a computer device; it is the kernel and foundation of the computer. The operating system handles fundamental tasks such as managing and configuring memory, determining the priority of system resource allocation, controlling input and output devices, operating the network, and managing the file system. To facilitate user operation, most operating systems provide a user interface for interaction with the system.

[0129] The BIOS is used to perform hardware initialization during the power-on boot phase and to provide runtime services for the operating system and applications. In some implementations, the BIOS can also monitor and display processor temperature and execute temperature protection strategies.

[0130] Application software, also known as an application program, can be understood as software written for a specific user application purpose, and is one of the main categories of computer software. For example, application software can be a program used to achieve purposes such as power control and temperature management.

[0131] It is understood that the specific examples in this application are only intended to help those skilled in the art better understand the implementation of this application, and are not intended to limit the scope of protection of this application.

[0132] It is understood that in the various embodiments of this application, the sequence number of each process does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of this application.

[0133] It is understood that the various implementation methods described in this application can be implemented individually or in combination, and this application does not limit them.

[0134] Unless otherwise stated, all technical and scientific terms used in this application have the same meaning as commonly understood by one of ordinary skill in the art. The terminology used in this application is for the purpose of describing particular embodiments only and is not intended to limit the scope of this application. The singular forms “a,” “the,” and “the” used in this application and the appended claims are also intended to include the plural forms unless the context clearly indicates otherwise.

[0135] Those skilled in the art will recognize that the units and algorithm steps of the various examples described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are implemented in hardware or software depends on the specific application and design constraints of the technical solution. Those skilled in the art can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0136] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the embodiments described above can be referred to the corresponding processes in other embodiments, and will not be repeated here.

[0137] In the embodiments provided in this application, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. For example, the apparatus embodiments described above are merely illustrative; for instance, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. For example, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Furthermore, the coupling or direct coupling or communication connection shown or discussed may be through some interfaces; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.

[0138] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of the technical solution in this application, depending on actual needs.

[0139] In addition, the functional units in the various embodiments of this application can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.

[0140] If a function is implemented as a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of this application, essentially, or the part that contributes to related technologies, or part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods described in the various embodiments of this application. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.

[0141] The above are merely specific embodiments of this application, but the scope of protection of this application is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the scope of the technology disclosed in this application should be included within the scope of protection of this application. Therefore, the scope of protection of this application should be determined by the scope of the claims.

Claims

1. A method of reliability analysis of a physical model, characterized by, The method for evaluating the understanding ability of the embodied model to task logic, physical laws and / or scene interaction in embodied intelligent application scenarios involving physical interaction or long-range tasks, the method comprises: receiving a target video of a target agent performing an operation task generated by the embodied model; the embodied model is used to generate a target video of a target agent performing an operation task according to an input task description or instruction to show the action sequence of the target agent and the interaction behavior of the target agent with the environment; performing multiple kinds of motion consistency analysis, visual consistency analysis and logical rationality analysis on the target video to obtain multiple kinds of motion consistency index values, visual consistency index values and logical rationality index values; determining the reliability analysis result of the embodied model based on the multiple kinds of motion consistency index values, visual consistency index values and logical rationality index values; if any one of the motion consistency index value, the visual consistency index value and the logical rationality index value is not within the corresponding index value range, outputting corresponding prompt information, the prompt information is used to indicate the unqualified index; wherein, in the case that the motion consistency analysis includes trajectory consistency analysis, the process of performing trajectory consistency analysis on the target video includes: performing end effector detection on the target video and a reference video, respectively obtaining a target trajectory corresponding to the target agent and a reference trajectory corresponding to a reference agent, the target trajectory and the reference trajectory being the spatial trajectory of the end effector of the corresponding agent; the reference video is a real video or a synthetic video of the reference agent performing the operation task, and the reference video meets the reliability requirement; based on the trajectory similarity between the target trajectory and the reference trajectory, a trajectory consistency index value is determined; the process of performing logical rationality analysis on the target video includes: using a third multi-modal large language model to analyze the logical rationality of each operation step in the target video to obtain a rationality score of each operation step; based on the rationality scores of all operation steps, the logical rationality index value is determined.

2. The body model reliability analysis method according to claim 1, wherein the process of performing motion consistency analysis on the target video includes: performing one or more of trajectory consistency analysis and dynamic consistency analysis on the target video to obtain one or more of trajectory consistency index value and dynamic consistency index value; determining the motion consistency index value according to one or more of the trajectory consistency index value and the dynamic consistency index value.

3. The body model reliability analysis method according to claim 2, wherein The trajectory similarity is determined based on the bidirectional Hausdorff distance and / or normalized dynamic time warping distance between the target trajectory and the reference trajectory.

4. The body model reliability analysis method according to claim 2, wherein the process of performing dynamic consistency analysis on the target video includes: processing the target trajectory based on a difference algorithm to obtain a target velocity sequence and a target acceleration sequence; processing the reference trajectory based on the difference algorithm to obtain a reference velocity sequence and a reference acceleration sequence; A dynamic consistency index value is calculated based on a velocity similarity between the target velocity sequence and the reference velocity sequence and / or an acceleration similarity between the target acceleration sequence and the reference acceleration sequence.

5. The body model reliability analysis method according to claim 4, wherein The velocity similarity is determined based on a Wasserstein distance between the target velocity sequence and the reference velocity sequence; and / or, The acceleration similarity is determined based on a Wasserstein distance between the target acceleration sequence and the reference acceleration sequence.

6. The somatic-model reliability analysis method of claim 1, wherein The process of performing visual consistency analysis on the target video comprises: performing one or more of subject consistency analysis, overall consistency analysis and appearance style analysis on the target video to obtain one or more of a subject consistency index value, an overall consistency index value and an appearance style index value; determining the visual consistency index value according to one or more of the subject consistency index value, the overall consistency index value and the appearance style index value.

7. The body model reliability analysis method according to claim 6, wherein The process of performing subject consistency analysis on the target video comprises: extracting subject image features of each frame in the target video; calculating cosine similarity between the subject image features of the first frame and the subject image features of each frame to obtain a first similarity score; calculating cosine similarity between the subject image features of adjacent frames to obtain a second similarity score; determining the subject consistency index value based on the first similarity score and the second similarity score.

8. The body model reliability analysis method according to claim 6, wherein The process of performing overall consistency analysis on the target video comprises: extracting overall image features of each frame in the target video; calculating cosine similarity between the overall image features of the first frame and the overall image features of each frame to obtain an overall consistency index value.

9. The body model reliability analysis method according to claim 6, wherein The process of performing appearance style analysis on the target video comprises: inputting the target video into a specified model to generate free-form natural language description information; the specified model is a first multi-modal large language model or an image-to-text generation model; inputting the natural language description information into a second multi-modal large language model for style judgment, the style judgment comprising classification judgment based on a preset language question or numerical judgment of mapping the natural language description information into a realism score; determining the appearance style index value based on the style judgment result.

10. A physical model reliability analysis system, characterized by, The system comprises a control module for performing the method of any one of claims 1 to 9.

11. A computer readable storage medium, characterized in that, The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the method of any one of claims 1 to 9.

12. A computer device, comprising: The computer device comprises a memory and a processor, the memory stores a computer program, and the processor implements the method of any one of claims 1 to 9 when executing the computer program.

13. A computer program product, characterised in that, The computer program product comprises a computer program, which, when executed by a processor, implements the method of any one of claims 1 to 9.

Citation Information

Patent Citations

  • Motion planning method, training method and system of intelligent body model

    CN119398105A