Reliability analysis method and system of body model, medium, equipment and product
By analyzing the motion consistency, visual consistency and logical rationality of the videos generated by the embodied model, the problem of difficulty in evaluating the reliability of the embodied model in the prior art is solved, and the performance quality of the agent in actual tasks is improved.
Patent Information
- Application Number
- CN202510538299.4
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-25
- Publication Date
- 2025-08-08
- Estimated Expiration
- 2045-04-25
AI Technical Summary
The prior art is difficult to effectively evaluate the reliability of embodied models in physical interaction or long-range tasks, resulting in the behavioral performance of agents that does not conform to task logic and physical laws.
By analyzing the video of the target agent performing operation tasks generated by the embodied model, the motion consistency, visual consistency and logical rationality analysis are performed, and the corresponding index values are obtained to evaluate the reliability of the embodied model.
The reliability evaluation of the embodied model in actual task scenarios is realized, and its deviations in task logic, physical laws and visual performance are discovered and optimized, so as to improve the quality of the agent's behavior.
Smart Images

Figure CN120449922A_ABST
Abstract
Description
Technical Field
[0001] The embodiments of the present application relate to the field of embodied intelligence technology, and in particular to a reliability analysis method, system, medium, device, and product for an embodied model. Background Art
[0002] In the field of embodied AI, especially in embodied AI applications involving physical interactions or long-range tasks, relevant evaluation methods have difficulty measuring the reliability of embodied models, which directly affects the behavioral performance of the intelligent agent in actual tasks, and may result in the intelligent agent's behavioral performance in actual tasks failing to meet the requirements.
[0003] Based on this, the embodiments of the present application provide a reliability analysis method and system, medium, equipment and product of an embodied model to improve related technologies. Summary of the Invention
[0004] The purpose of the embodiments of the present application is to provide a reliability analysis method and system, medium, equipment and product for an embodied model. By analyzing the target video of the target intelligent agent generated by the embodied model performing the operation task, key indicators are obtained, which can effectively evaluate the reliability of the embodied model in actual task scenarios.
[0005] The purpose of the embodiments of the present application is achieved by using the following technical solutions:
[0006] In the first aspect, an embodiment of the present application provides a reliability analysis method for an embodied model, the method comprising: receiving a target video of a target intelligent agent performing an operation task generated by the embodied model; performing one or more of motion consistency analysis, visual consistency analysis, and logical rationality analysis on the target video to obtain one or more of a motion consistency index value, a visual consistency index value, and a logical rationality index value; and determining the reliability analysis result of the embodied model based on one or more of the motion consistency index value, the visual consistency index value, and the logical rationality index value.
[0007] In some embodiments, the process of performing motion consistency analysis on the target video includes: performing one or more of trajectory consistency analysis and dynamic consistency analysis on the target video to obtain one or more of trajectory consistency index values and dynamic consistency index values; determining the motion consistency index value based on one or more of the trajectory consistency index values and the dynamic consistency index values.
[0008] In some embodiments, the process of performing trajectory consistency analysis on the target video includes: performing end effector detection on the target video and the reference video, respectively obtaining the target trajectory corresponding to the target agent and the reference trajectory corresponding to the reference agent, the target trajectory and the reference trajectory being the spatial trajectories of the end effectors of the corresponding agents; the reference video is a real video or a synthetic video of the reference agent performing the operation task, and the reference video meets the reliability requirements; based on the trajectory similarity between the target trajectory and the reference trajectory, determining the trajectory consistency index value.
[0009] In some embodiments, the trajectory similarity is determined based on a bidirectional Hausdorff distance and / or a normalized dynamic time warping distance between the target trajectory and the reference trajectory.
[0010] In some embodiments, the process of performing dynamic consistency analysis on the target video includes: processing the target trajectory based on a differential algorithm to obtain a target velocity sequence and a target acceleration sequence; processing the reference trajectory based on the differential algorithm to obtain a reference velocity sequence and a reference acceleration sequence; and calculating a dynamic consistency index value based on the velocity similarity between the target velocity sequence and the reference velocity sequence and / or the acceleration similarity between the target acceleration sequence and the reference acceleration sequence.
[0011] In some embodiments, the speed similarity is determined based on a Wasserstein distance between the target speed sequence and the reference speed sequence; and / or the acceleration similarity is determined based on a Wasserstein distance between the target acceleration sequence and the reference acceleration sequence.
[0012] In some embodiments, the process of performing visual consistency analysis on the target video includes: performing one or more of subject consistency analysis, overall consistency analysis, and appearance style analysis on the target video to obtain one or more of subject consistency index values, overall consistency index values, and appearance style index values; determining the visual consistency index value based on one or more of the subject consistency index value, the overall consistency index value, and the appearance style index value.
[0013] In some embodiments, the process of performing subject consistency analysis on the target video includes: extracting subject image features of each frame in the target video; calculating the cosine similarity between the subject image features of the first frame and the subject image features of each frame to obtain a first similarity score; calculating the cosine similarity between the subject image features of adjacent frames to obtain a second similarity score; and determining the subject consistency index value based on the first similarity score and the second similarity score.
[0014] In some embodiments, the process of performing an overall consistency analysis on the target video includes: extracting the overall image features of each frame in the target video; calculating the cosine similarity between the overall image features of the first frame and the overall image features of each frame to obtain an overall consistency index value.
[0015] In some embodiments, the process of performing appearance style analysis on the target video includes: inputting the target video into a designated model to generate free-format natural language description information; the designated model is a first multimodal large language model or an image-to-text generation model; inputting the natural language description information into a second multimodal large language model for style judgment, the style judgment including a classification judgment based on a preset language question or mapping the natural language description information into a numerical judgment of a realism score; and determining the appearance style index value based on the style judgment result.
[0016] In some embodiments, the process of performing a logical rationality analysis on the target video includes: using a third multimodal large language model to analyze the logical rationality of each operation step in the target video to obtain a rationality score for each operation step; and determining the logical rationality index value based on the rationality scores of all operation steps.
[0017] In a second aspect, an embodiment of the present application provides a reliability analysis system for an embodied model, the system comprising a control module, and the control module is configured to execute any one of the methods described in the first aspect.
[0018] In a third aspect, an embodiment of the present application provides a computer-readable storage medium, wherein the computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method described in any one of the first aspects is implemented.
[0019] In a fourth aspect, an embodiment of the present application provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements any one of the methods described in the first aspect when executing the computer program.
[0020] In a fifth aspect, an embodiment of the present application provides a computer program product, which includes a computer program, and when the computer program is executed by a processor, it implements any one of the methods in the first aspect.
[0021] The embodiments of the present application provide a reliability analysis method and system, medium, device and product for an embodied model, which receives a target video of a target agent performing an operation task generated by an embodied model; performs one or more of motion consistency analysis, visual consistency analysis and logical rationality analysis on the target video to obtain one or more of motion consistency index value, visual consistency index value and logical rationality index value; and determines the reliability analysis result of the embodied model based on one or more of the motion consistency index value, visual consistency index value and logical rationality index value. By performing one or more of motion consistency analysis, visual consistency analysis and logical rationality analysis on a target video of a target agent performing an operation task generated by an embodied model, the performance quality of the target agent performing the operation task in the target video can be evaluated. Based on one or more of the motion consistency index value, visual consistency index value and logical rationality index value, the reliability of the embodied model is quantitatively evaluated, and data support is provided for optimizing the embodied model. It can be seen that the embodiments of the present application can analyze one or more deviations in the physical laws, visual performance and task logic of the intelligent agent behavior in the target video (i.e., the behavior of the target intelligent agent in performing the operation task), and realize the reliability evaluation of the embodied model, so as to guide the optimization of the embodied model in a targeted manner, thereby improving the ability of the intelligent agent to perform high-quality behavior in actual tasks, so that it can better meet the needs of actual application scenarios. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] The embodiments of the present application are further described below with reference to the accompanying drawings and specific implementation methods.
[0023] Figure 1 This is a flow chart of a reliability analysis method of an embodied model provided in an embodiment of the present application.
[0024] Figure 2 This is a schematic diagram of the framework of an embodied scene world model generation quality assessment system provided in an embodiment of the present application.
[0025] Figure 3 3 is a schematic diagram comparing a reference trajectory and a target trajectory provided in an embodiment of the present application.
[0026] Figure 4 This is a structural block diagram of a computer device provided in an embodiment of the present application. DETAILED DESCRIPTION
[0027] The following will be combined with the drawings in the embodiments of the present application to clearly and completely describe the technical solutions in the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Based on the embodiments in the embodiments of the present application, all other embodiments obtained by those skilled in the art without making creative efforts are within the scope of protection of the embodiments of the present application.
[0028] In the description of the embodiments of the present application, it should be understood that the terms "first" and "second" are used for descriptive purposes only and should not be understood as indicating or implying relative importance or implicitly indicating the number of the indicated technical features. Thus, features defined as "first" and "second" may explicitly or implicitly include one or more of the features. In the description of the embodiments of the present application, the meaning of "plurality" is two or more, unless otherwise clearly and specifically defined.
[0029] Current generative model evaluation benchmarks primarily focus on evaluating general content generation tasks, such as image, text, and video generation. These evaluations typically rely on specific metrics, such as the Frechet Inception Distance (FID), which assesses the realism of generated images, and the Contrastive Language–Image Pre-training Score (CLIP), which measures the alignment between text and images. These methods are very effective in assessing the quality of generated content, especially for static or non-interactive content generation.
[0030] In the field of embodied intelligence, specialized task-based model evaluation frameworks have been proposed. These frameworks aim to assess the effectiveness of reinforcement learning strategies or low-level motion control (such as a robotic arm grasping an object). For example, some frameworks evaluate the performance of an agent under physical constraints by measuring task completion efficiency (such as object handling rate), while others focus on how large multimodal models decompose complex instructions into specific action steps.
[0031] However, in some specific application scenarios, such as embodied intelligence applications involving physical interactions or long-range tasks, the above evaluation methods or evaluation frameworks are difficult to measure the reliability of embodied models, and have at least the following limitations: (1) Logic breaks in long-range task generation: When the embodied model drives the intelligent agent to perform a multi-step operation task of the robotic arm, if the embodied model has logical planning defects, it is easy for the intelligent agent to have a problem of disordered operation sequence. For example, the embodied model may incorrectly plan the operation sequence of "closing the refrigerator door first and then placing the item". The robotic arm, as the execution carrier of the intelligent agent, will directly present the task logic incoherence, and ultimately lead to task failure. (2) Insufficient modeling of physical laws: In the generated video, the motion trajectory of the object may violate Newtonian mechanics (such as a rigid object floating without support), and the relevant indicators cannot effectively detect such problems.
[0032] In response to the shortcomings of the embodied model evaluation method, an embodiment of the present application proposes a reliability analysis method for the embodied model. This method obtains key indicators by analyzing the target video of the target intelligent agent generated by the embodied model performing the operation task, which can effectively evaluate the reliability of the embodied model in actual task scenarios.
[0033] See also Figure 1 , Figure 1 This is a flow chart of a reliability analysis method of an embodied model provided in an embodiment of the present application.
[0034] In order to improve the relevant technology, an embodiment of the present application provides a reliability analysis method of an embodied model, which includes steps S101 to S103.
[0035] Step S101: Receive a target video of a target agent performing an operation task generated by the embodied model.
[0036] Step S102: performing one or more of motion consistency analysis, visual consistency analysis and logical rationality analysis on the target video to obtain one or more of motion consistency index values, visual consistency index values and logical rationality index values.
[0037] Step S103: Determine a reliability analysis result of the embodied model based on one or more of the motion consistency index value, the visual consistency index value, and the logical rationality index value.
[0038] In some embodiments, the method can be performed by a reliability analysis system, which can be deployed on a cloud server, a local computing device, or an edge computing node. The method can be applied to a variety of practical application scenarios, such as embodied intelligence applications involving physical interactions or long-range tasks. Furthermore, the method is also applicable to other intelligent agent task scenarios requiring high reliability (or quality), such as robotic simulation, robotic data acquisition, and industrial automation scenarios.
[0039] In some embodiments, an intelligent agent refers to an agent that can perceive the environment and take actions to achieve specific goals. The target intelligent agent can be an intelligent agent in physical form or digital form. Physical intelligent agents may include, for example, bipedal robots (also known as humanoid robots, humanoid robots), quadruped robots or wheeled robots, etc., which can interact with the environment through mechanical feedback and perform operational tasks in real scenes (such as grabbing objects, moving objects, etc.); digital intelligent agents may include, for example, virtual robots, software agents or NPC characters in game AI, which can perform operational tasks in virtual scenes (such as multi-step path planning, virtual environment navigation or performing robot operation tasks in a simulated environment, etc.).
[0040] In some embodiments, the embodied model can generate a video of the target agent performing an operational task, with the video content typically sourced from an open-source dataset. For example, the embodied model can generate a dynamic video of the target agent performing the operational task based on the input task description or instructions, demonstrating the target agent's action sequence and its interaction with the environment. By analyzing the video content generated by the embodied model, one can assess the embodied model's ability to understand the task logic, physical laws, and / or scene interactions.
[0041] In some embodiments, the embodied model can be an embodied scene world model (Embodied Scene World Model), which is a multimodal generative model that combines embodied intelligence and scene understanding capabilities, and aims to simulate and predict the behavior of the target intelligent agent in a specific scene and its interaction process with the environment.
[0042] In some embodiments, the video content generated by the embodied model can be derived from an open source data set similar to Agibot-World, covering a variety of task backgrounds and types of objects being operated to ensure data diversity. The operated objects can include, for example, furniture, electrical appliances, rigid objects, and flexible objects. For example, one or more of the following data selection criteria can be used to guide the embodied model to generate a target video of the target agent performing the operation task: the task operation is quite different from the task background, each task scene conforms to the actual application, the task background is clean, the sub-action sequence is clear, and the operated object is unique (to ensure that the task logic is clear). Through the above design, the generated video can not only show the behavior process of the target agent, but also provide high-quality support for evaluating the task understanding and execution capabilities of the embodied model. In addition, this way of generating videos helps to discover potential problems of the embodied model in task logic inference, physical law modeling, and dynamic interaction processing, thereby providing a better basis for optimizing the actual task performance of the target agent.
[0043] See also Figure 2 , Figure 2 It is a schematic diagram of a framework of a quality evaluation system for embodied scene world model generation provided by an embodiment of the present application. The system (also referred to as a reliability analysis system) can generate a target video of a target agent performing an operation task based on data selection criteria (for example, data from the Agibot-World open source data set, tasks and backgrounds, and the diversity of the operated objects, a clean background, and a clear sub-action execution order), and obtain an evaluation index by evaluating the target video according to multiple evaluation dimensions. The multiple evaluation dimensions may include, for example, motion consistency, visual consistency, and logical rationality. Among them, motion consistency can be measured by calculating the trajectory similarity and motion similarity between the reference trajectory (also referred to as: true value trajectory) and the target trajectory (also referred to as: generated trajectory) to measure whether the target trajectory conforms to general physical laws and task logic. Visual consistency can be achieved by fine-tuning DINOv2backbone (a visual feature extraction model based on self-supervised learning) using the embodied operation data set Agibot-World so that it has the ability to understand the subject with the operation task. Logical rationality can be achieved by using a multimodal large language model to evaluate the ability of the embodied scene world model to handle complex and long-term tasks. By evaluating the target agent's operation tasks in the target video generated by the model according to multiple evaluation dimensions (such as motion consistency, visual consistency, and logical rationality), we can better measure the quality of the target agent's behavior generated by the embodied model and its ability to understand the task logic and physical laws.
[0044] In some embodiments, motion consistency analysis can be used to evaluate whether the behavior of the target agent performing the operation task in the target video conforms to physical laws. Motion consistency analysis can include one or more of trajectory consistency analysis and dynamic consistency analysis.
[0045] Trajectory consistency analysis can be used to measure the consistency of the target agent's motion trajectory by calculating the trajectory similarity. Dynamic consistency analysis can be used to measure the consistency of the target agent's dynamic characteristics by calculating the velocity similarity and / or acceleration similarity.
[0046] In some embodiments, visual consistency analysis can be used to evaluate whether the visual representation of the target agent and its environment in the target video is reasonable. Visual consistency analysis can include one or more of subject consistency analysis, overall consistency analysis, and appearance style analysis.
[0047] Subject consistency analysis can be used to assess whether the main image features of the target agent in the target video remain consistent across time. Overall consistency analysis can be used to assess whether the overall features of the target video remain consistent across time. Appearance style analysis can be used to evaluate the appearance style of the target video by assessing whether it meets the requirements of realism.
[0048] In some embodiments, logic rationality analysis can be used to evaluate whether the logic of the target agent performing the operation task in the target video is reasonable.
[0049] In some embodiments, features can be extracted from the target video and quantitatively evaluated using a deep learning model (e.g., a multimodal large language model) to obtain one or more of a motion consistency index value, a visual consistency index value, and a logical rationality index value.
[0050] In some embodiments, if any one of the motion consistency index value, visual consistency index value, and logical rationality index value is not within the corresponding index value range, a corresponding prompt message is output. The prompt message may indicate the unqualified indicator (such as unreasonable motion, visual incoherence, or logical error) for optimizing the embodied model.
[0051] In some embodiments, weighted averaging, threshold comparison method, or multi-dimensional scoring mapping methods may be used to process the corresponding indicator values to obtain reliability analysis results of the embodied model.
[0052] Reliability analysis results indicate the motion consistency, visual consistency, and logical rationality of the target agent performing the task in the target video generated by the embodied model. This allows assessment of whether the target agent's execution of the task in the target video conforms to physical laws, task logic, and scene realism. Reliability analysis results provide a reliable simulation basis for the actual deployment of the target agent and guide the continuous optimization of the embodied model.
[0053] For example, the values of each indicator can be weighted and summed according to preset weights to obtain a comprehensive score to characterize the reliability analysis results of the embodied model. The setting of the preset weights can be adjusted according to the importance of the specific application scenario. For example, in tasks involving physical interaction, the weight of the motion consistency index value is higher than the weight of the visual consistency index value, and higher than the weight of the logical rationality index value; while in tasks requiring high visual fidelity, the weight of the visual consistency index value is higher than the weight of the motion consistency index value, and higher than the weight of the logical rationality index value.
[0054] As another example, each indicator value can be compared with a corresponding indicator value range, and based on the comparison results, a reliability analysis result of the embodied model can be obtained. The comparison of each indicator value with the corresponding indicator value range can be performed in a priority order based on the importance of each indicator value. The selection of the indicator value range can be set according to actual needs and is not specifically limited here.
[0055] As another example, each indicator value can be mapped to a multidimensional scoring space, and the reliability level of the embodied model can be determined using a classification or clustering algorithm to obtain a reliability analysis result. For example, the values of the motion consistency, visual consistency, and logical rationality indicators can be input into a pre-trained classification model, which can then output a natural language description of the reliability level.
[0056] In the above embodiment, by performing one or more of motion consistency analysis, visual consistency analysis and logical rationality analysis on the target video of the target agent performing the operation task generated by the embodied model, the performance quality of the target agent performing the operation task in the target video can be evaluated. Based on one or more of the motion consistency index value, the visual consistency index value and the logical rationality index value, the reliability of the embodied model is quantitatively evaluated, and data support is provided for optimizing the embodied model. It can be seen that the embodiment of the present application can analyze one or more deviations of the agent behavior in the target video in physical laws, visual performance and task logic, and realize the reliability evaluation of the embodied model, so as to guide the optimization of the embodied model in a targeted manner, thereby enhancing the ability of the agent to perform high-quality behavior in actual tasks, so that it can better meet the needs of actual application scenarios.
[0057] In order to effectively quantify and evaluate the motion consistency of the intelligent agent behavior in the target video, in some embodiments, the process of performing motion consistency analysis on the target video may include: performing one or more of trajectory consistency analysis and dynamic consistency analysis on the target video to obtain one or more of trajectory consistency index values and dynamic consistency index values; determining the motion consistency index value based on one or more of the trajectory consistency index values and the dynamic consistency index values.
[0058] Illustratively, performing trajectory consistency analysis on the target video to obtain a trajectory consistency index value may include: determining the trajectory consistency index value by evaluating whether the trajectory corresponding to the target intelligent agent in the target video is consistent with a reference trajectory.
[0059] Exemplarily, performing a dynamic consistency analysis on a target video to obtain a dynamic consistency index value may include: determining the dynamic consistency index value by evaluating whether the speed distribution and / or acceleration distribution of the target intelligent body in the target video is consistent with a reference speed distribution and / or reference acceleration distribution.
[0060] For example, one or more of the trajectory consistency index value and the dynamic consistency index value can be selected to determine the motion consistency index value based on the needs of the actual application scenario. For example, the trajectory consistency index value and the dynamic consistency index value can be weighted and summed according to preset weights to obtain a comprehensive motion consistency index value; or, in some cases, the trajectory consistency index value or the dynamic consistency index value can be selected to determine the motion consistency index value. In addition, by setting a threshold, the motion consistency index value used to characterize whether the behavior of the target intelligent agent conforms to the laws of physics can be determined based on whether both the trajectory consistency index value and the dynamic consistency index value reach a preset threshold.
[0061] In the above embodiment, by performing trajectory consistency analysis and dynamic consistency analysis on the target video, the rationality of its trajectory and motion dynamics is quantitatively evaluated to obtain the motion consistency index value of the intelligent agent behavior in the target video. This can more accurately quantitatively evaluate whether the behavior of the target intelligent agent conforms to the laws of physics, and can effectively alleviate problems caused by insufficient modeling of physical laws (such as unsupported suspension of rigid objects, etc.). This quantitative evaluation provides reliable data support for the motion consistency analysis of the intelligent agent behavior, thereby improving the reliability analysis effect of the embodied model.
[0062] In order to effectively quantify and evaluate the rationality of the trajectory of the agent's behavior in the target video, in some embodiments, the process of performing trajectory consistency analysis on the target video may include: performing end effector detection on the target video and the reference video, and respectively obtaining the target trajectory corresponding to the target agent and the reference trajectory corresponding to the reference agent, wherein the target trajectory is the spatial trajectory of the end effector of the target agent, and the reference trajectory is the spatial trajectory of the end effector of the reference agent; the reference video is a real video or a synthetic video of the reference agent performing the operation task, and the reference video meets the reliability requirements; based on the trajectory similarity between the target trajectory and the reference trajectory, the trajectory consistency index value is determined.
[0063] For example, the end effector position in the target video and the end effector position in the reference video can be extracted frame by frame to obtain the target trajectory corresponding to the end effector in the target agent and the reference trajectory corresponding to the end effector in the reference agent.
[0064] The end effector position can be extracted from video frames using, for example, a target detection algorithm or a keypoint detection algorithm. The target trajectory and reference trajectory can be stored as time series. The target trajectory can be used to reflect the motion path of the target agent's end effector during task execution, while the reference trajectory can be used to reflect the motion path of the reference agent's end effector during the same task or benchmark scenario and can be considered a benchmark trajectory. The end effector can be located at the end of the agent's robotic arm or other manipulator. Examples of the end effector include a gripper, suction cup, or dexterous hand.
[0065] The reference agent is an agent used to provide benchmark behavior or trajectory. For example, the reference agent can be a manually controlled robot, a rigorously trained humanoid robot, a robotic arm, or a virtual agent generated by a high-precision simulation environment, etc. The generated trajectory data must conform to the laws of physics to ensure its reliability as a benchmark.
[0066] Real videos can be videos of a reference agent performing an operation task, recorded by high-precision sensors (such as depth cameras or motion capture systems). Synthetic videos can be videos of a reference agent performing an operation task, generated in a high-precision simulation environment. Both real and synthetic videos of the reference agent performing the operation task must meet reliability requirements to provide a reliable benchmark for trajectory consistency analysis of the target agent.
[0067] See also Figure 3 , Figure 3It is a schematic diagram of a comparison between a reference trajectory and a target trajectory provided by an embodiment of the present application. For example, by using the YOLO-World model fine-tuned based on the Agibot-World dataset, target detection can be performed on the end effector of the robotic arm in the reference video (also called the true value video) and the target video (also called the generated video), thereby extracting the motion trajectory corresponding to the left and right hands of the robotic arm in the reference video as the reference trajectory (or true value trajectory), and the motion trajectory corresponding to the left and right hands of the robotic arm in the target video as the target trajectory (or generated trajectory). For ease of comparison, the reference trajectory and the target trajectory can be visualized in the same image. For example, in Figure 3 In the figure, the dotted line represents the target trajectory and the solid line represents the reference trajectory. The difference between the reference trajectory (solid line) and the target trajectory (dashed line) is intuitively displayed in a visual way, which facilitates the analysis and evaluation of the consistency between the two.
[0068] Exemplarily, the distance between the target trajectory and the reference trajectory can be calculated using one or more selected measurement methods (e.g., spatial distribution characteristic measurement and / or time series characteristic measurement), and the calculated one or more distances are normalized (e.g., taking the inverse) so that the calculated distance is positively correlated with the trajectory consistency; the calculated distances are evaluated (assuming that the calculated distances are greater than one, a weighted summation or a single indicator can be selectively used) to obtain a trajectory consistency index value.
[0069] Spatial distribution characteristic metrics can be used to assess the similarity between the target trajectory and the reference trajectory in terms of spatial distribution. For example, the maximum deviation between the target trajectory and the reference trajectory in terms of spatial distribution can be quantified by calculating the Hausdorff distance. The Hausdorff distance defines the maximum distance between two sets, that is, the maximum value from any point in one set to the nearest point in the other set.
[0070] Time series feature metrics can be used to assess the temporal consistency of a target trajectory and a reference trajectory. For example, the Normalized Dynamic Time Warping (NDTW) distance or Dynamic Time Warping (DTW) distance can be calculated to quantify the temporal similarity between the target and reference trajectories. NDTW is a distance metric based on DTW that is normalized to make the results easier to interpret and compare.
[0071] In the above embodiment, by performing end-effector detection on the target video and the reference video, the target trajectory and the reference trajectory are obtained respectively, and the trajectory consistency index value is determined based on the trajectory similarity between the two, thereby achieving an accurate quantitative evaluation of the trajectory rationality of the intelligent agent's behavior. At the same time, it provides reliable data support for evaluating whether the motion trajectory of the target intelligent agent conforms to the laws of physics, thereby improving the reliability analysis effect of the embodied model.
[0072] In order to more accurately quantitatively evaluate the trajectory rationality of the agent's behavior in the target video, in some embodiments, the trajectory similarity can be determined based on the bidirectional Hausdorff distance and / or the normalized dynamic time warping distance between the target trajectory and the reference trajectory.
[0073] For example, to more accurately assess the similarity between the target trajectory and the reference trajectory in spatial distribution, trajectory similarity can be determined based on the bidirectional Hausdorff distance (also known as the symmetric Hausdorff distance) between the target trajectory and the reference trajectory. The bidirectional Hausdorff distance is an extension of the traditional Hausdorff distance. It can simultaneously capture the maximum deviation of the target trajectory from the reference trajectory, as well as the maximum deviation of the reference trajectory from the target trajectory, thereby discovering abnormal points or unreasonable parts in the target trajectory. This comprehensive consideration of the bidirectional maximum distance improves the accuracy of the analysis and facilitates the discovery of abnormal parts in the target agent's behavior.
[0074] For example, to more accurately assess the similarity between the target trajectory and the reference trajectory in time series, trajectory similarity can be determined based on the NDTW distance between the target trajectory and the reference trajectory. NDTW allows for flexible matching of trajectories of different lengths or speeds by nonlinearly aligning two time series, thereby minimizing the cumulative distance between them. For example, when the action rhythms of the target trajectory and the reference trajectory are inconsistent, NDTW can effectively capture the deviation between the two in time series, making it easier to detect abnormalities in the behavior of the target agent.
[0075] In the above embodiment, a single distance can be selectively used for the bidirectional Hausdorff distance and NDTW distance between the target trajectory and the reference trajectory to determine the trajectory consistency index value, or a weighted sum of the bidirectional Hausdorff distance and NDTW distance can be performed to evaluate the consistency of the target trajectory with the reference trajectory from two dimensions: spatial distribution characteristics and time series characteristics. This multi-dimensional comprehensive consideration not only effectively improves the accuracy of trajectory similarity analysis, but also provides more reliable data support for evaluating whether the target agent's behavior conforms to physical laws, thereby improving the reliability analysis effect of the embodied model.
[0076] In order to effectively quantify and evaluate the motion dynamics of the intelligent agent's behavior in the target video, in some embodiments, the process of performing dynamic consistency analysis on the target video may include: processing the target trajectory based on a differential algorithm to obtain a target velocity sequence and a target acceleration sequence; processing the reference trajectory based on the differential algorithm to obtain a reference velocity sequence and a reference acceleration sequence; and calculating a dynamic consistency index value based on the velocity similarity between the target velocity sequence and the reference velocity sequence and / or the acceleration similarity between the target acceleration sequence and the reference acceleration sequence.
[0077] Exemplarily, the target trajectory and the reference trajectory can be differentially operated based on the differential algorithm to obtain the target velocity sequence and target acceleration sequence corresponding to the target trajectory, as well as the reference velocity sequence and reference acceleration sequence corresponding to the reference trajectory. The differential algorithm is a method for calculating the rate of change of a numerical sequence. In the first-order differential, the velocity information can be extracted by calculating the difference between adjacent positions of the trajectory points; in the second-order differential, the velocity sequence is further differentiated to extract the acceleration information. In trajectory analysis, the differential algorithm can be used to extract velocity and acceleration information from trajectory data (e.g., corresponding to one or more position information) and quantify the similarity between the target trajectory and the reference trajectory in velocity or acceleration distribution. For example, given a series of trajectory points, the first-order differential can reveal the velocity of each step, while the second-order differential shows the trend of velocity change (i.e., acceleration information).
[0078] In some embodiments, the speed similarity may be determined based on a Wasserstein distance between the target speed sequence and the reference speed sequence; and / or, the acceleration similarity may be determined based on a Wasserstein distance between the target acceleration sequence and the reference acceleration sequence.
[0079] For example, the speed similarity between the target speed sequence and the reference speed sequence, and / or the acceleration similarity between the target acceleration sequence and the reference acceleration sequence may be calculated based on the Wasserstein distance.
[0080] The Wasserstein distance (also known as the Earth Mover's Distance, EMD) is a method for measuring the difference between two probability distributions. It reflects the difference between the two by calculating the minimum "work" required to transform one distribution into the other, where the "work" is determined by both the mass moved and the distance moved. In multidimensional space, the Wasserstein distance can capture the geometric differences between different distributions, making it particularly suitable for distribution comparison tasks in fields such as image processing and natural language processing. For example, in dynamic consistency analysis, it can be used to quantify the similarity between the target trajectory and the reference trajectory in terms of velocity or acceleration distribution.
[0081] By using the Wasserstein distance to measure the similarity between the target velocity sequence and the reference velocity sequence, and / or the similarity between the target acceleration sequence and the reference acceleration sequence, a dynamic consistency index value is obtained, which facilitates the discovery of possible unreasonable dynamic characteristics problems (such as sudden speed changes, abnormal acceleration, etc.) in the target generated trajectory.
[0082] In the above embodiment, either velocity similarity or acceleration similarity can be selected to determine the dynamic consistency index value. Alternatively, the dynamic consistency index value can be determined based on velocity similarity and acceleration similarity. For example, the velocity similarity and acceleration similarity can be weighted and summed according to preset weights to obtain a comprehensive dynamic consistency index value. This improves the accuracy of the dynamic consistency assessment of the intelligent agent's behavior, thereby improving the reliability analysis effect of the embodied model.
[0083] In order to effectively quantify and evaluate the visual consistency of the intelligent agent behavior in the target video, in some embodiments, the process of performing visual consistency analysis on the target video may include: performing one or more of subject consistency analysis, overall consistency analysis and appearance style analysis on the target video to obtain one or more of the subject consistency index value, the overall consistency index value and the appearance style index value; determining the visual consistency index value based on one or more of the subject consistency index value, the overall consistency index value and the appearance style index value.
[0084] Exemplarily, performing subject consistency analysis on the target video to obtain a subject consistency index value may include: extracting subject image features of each frame in the target video and calculating cosine similarity to determine the subject consistency index value.
[0085] Exemplarily, performing an overall consistency analysis on the target video to obtain an overall consistency index value may include: extracting overall image features of each frame in the target video and calculating cosine similarity to determine the overall consistency index value.
[0086] Exemplarily, performing appearance style analysis on a target video to obtain an appearance style index value may include: evaluating the appearance style of the target video using a multimodal large language model (MLLM) to determine the appearance style index value.
[0087] For example, one or more of the subject consistency index value, the overall consistency index value, and the appearance style index value can be selected to determine the visual consistency index value based on the needs of the actual application scenario. For example, the subject consistency index value, the overall consistency index value, and the appearance style index value can be weighted and summed according to preset weights to obtain a comprehensive visual consistency index value; or, one of the subject consistency index value, the overall consistency index value, or the appearance style index value can be selected to determine the visual consistency index value. In addition, the visual consistency index value used to characterize whether the target intelligent agent's behavior conforms to visual rationality can be determined based on whether the subject consistency index value, the overall consistency index value, and the appearance style index value all reach a preset threshold.
[0088] In the above embodiment, by determining the visual consistency index value based on one or more of the subject consistency index value, the overall consistency index value, and the appearance style index value, it is possible to quantitatively evaluate whether the behavior of the target intelligent agent conforms to the rationality of visual representation, which can effectively alleviate problems caused by insufficient visual representation of the intelligent agent's behavior (such as excessive CG sense, lack of realism, etc.). This quantitative evaluation provides reliable data support for the visual consistency analysis of the intelligent agent, thereby improving the reliability analysis effect of the embodied model. CG refers to Computer Graphics.
[0089] In order to effectively quantify and evaluate the subject consistency of the intelligent agent behavior in the target video, in some embodiments, the process of performing subject consistency analysis on the target video may include: extracting the subject image features of each frame in the target video; calculating the cosine similarity between the subject image features of the first frame and the subject image features of each frame to obtain a first similarity score; calculating the cosine similarity between the subject image features of adjacent frames to obtain a second similarity score; and determining the subject consistency index value based on the first similarity score and the second similarity score.
[0090] For example, a visual feature extraction model can be used to extract subject image features from each frame of the target video. Subject image features can be used to reflect key visual information about the target agent. For example, subject image features may include key attributes such as the target agent's posture and visual representation. An example visual feature extraction model could be DINOv2, which, after fine-tuning, can better focus on subject features in embodied scenes, thereby improving the accuracy of subject image feature extraction.
[0091] For example, the first similarity score expressed in the form of a sequence can be obtained by calculating the cosine similarity between the main image features of the first frame of the target video and the main image features of each subsequent frame. For example, if the target video has N frames, the cosine similarity between the first frame and the i-th frame can be calculated (i = 1, 2, ..., N), and N first similarity scores can be obtained. The N first similarity scores form a sequence according to the time order of the frames in the target video. The first similarity score can be used to reflect the global consistency of the main features of the target intelligent agent in the entire video.
[0092] For example, the second similarity score expressed in the form of a sequence can be obtained by calculating the cosine similarity between the main image features of adjacent frames in the target video. For example, if the target video has N frames, the cosine similarity between the i-th frame and the i+1-th frame can be calculated (i=1, 2, ..., N-1), and N-1 second similarity scores can be obtained. The N-1 second similarity scores form a sequence according to the time order of the frames in the target video. The second similarity score can reflect the coherence of the main features of the target intelligent agent between adjacent frames, and can avoid unreasonable phenomena caused by jitter or mutation between frames.
[0093] The first similarity score expressed in the form of a sequence can be used to quantify the overall consistency of the main features of the target intelligent agent over the entire time range of the video. The second similarity score expressed in the form of a sequence can be used to evaluate the temporal coherence of the main features of the target intelligent agent between adjacent frames. If the first similarity score or the second similarity score shows a significant drop or sudden change in certain frames of the target video, it may indicate that the main image features of the target intelligent agent have undergone abnormal changes in the corresponding frames of the target video (for example, perspective switching or discontinuous movements, etc.).
[0094] In the above embodiment, the first similarity score and the second similarity score can be aggregated to determine the subject consistency index value. For example, the first similarity comprehensive score and the second similarity comprehensive score can be generated respectively by calculating the statistics (such as the average, minimum or weighted average) of the first similarity score sequence and the second similarity score sequence, and the first similarity comprehensive score and the second similarity comprehensive score are weighted and summed by the preset weights to obtain the subject consistency index value. It can be seen that the subject consistency index value is determined based on the first similarity score and the second similarity score. This quantitative evaluation provides key and accurate data support for the subject consistency analysis of the intelligent agent from the global dimension (i.e., the overall consistency of the subject features of the target intelligent agent within the entire video time range) and the local dimension (i.e., the temporal coherence of the subject features of the target intelligent agent between adjacent frames), thereby improving the reliability analysis effect of the embodied model.
[0095] In order to effectively quantify and evaluate the overall consistency of the intelligent agent's behavior in the target video, in some embodiments, the process of performing an overall consistency analysis on the target video may include: extracting the overall image features of each frame in the target video; calculating the cosine similarity between the overall image features of the first frame and the overall image features of each frame to obtain an overall consistency index value.
[0096] For example, a visual language model can be used to extract overall image features from each frame of the target video. Overall image features can be used to reflect the global visual representation of each frame in the target video. For example, if the target video undergoes a sudden change (such as a scene switch), the overall image features will change significantly. Overall image features can include key attributes such as the overall layout of the scene, background information, and the relationship between the subject and the environment.
[0097] For example, a visual language model could be a CLIP (Contrastive Language–Image Pretraining) model. The CLIP model is a multimodal model that enables cross-modal understanding between images and text by jointly training on image and text data. By training on a large number of image-text pairs, CLIP can perform various zero-shot transfer tasks, such as classification and retrieval, without requiring additional training. This flexibility allows for a wide range of applications, reducing the need to tailor the model to specific tasks.
[0098] For example, the overall image features of each frame in the target video can be extracted, and the cosine similarity between the overall image features of the first frame and the overall image features of each subsequent frame can be calculated to obtain multiple overall consistency scores. The multiple overall consistency scores are aggregated (e.g., taking the average or minimum value) to determine the overall consistency index value.
[0099] In the above embodiment, by extracting the overall image features of each frame in the target video and calculating the cosine similarity between the first frame and each subsequent frame, the overall consistency index value is finally determined. This quantitative evaluation provides key data support for the overall consistency analysis of the target video, provides a basis for the visual performance evaluation of the intelligent agent's behavior, and can further improve the reliability analysis effect of the embodied model.
[0100] In order to effectively quantify and evaluate the appearance style of the intelligent agent behavior in the target video, in some embodiments, the process of performing appearance style analysis on the target video may include: inputting the target video into a specified model to generate free-format natural language description information; the specified model is a first multimodal large language model or an image-to-text generation model; inputting the natural language description information into a second multimodal large language model for style judgment, and the style judgment includes a classification judgment based on a preset language question or mapping the natural language description information into a numerical judgment of a realism score; and determining the appearance style index value based on the style judgment result.
[0101] For example, a target video can be input into a designated model, which then interprets the image content of each frame or keyframe in the target video to generate a free-form natural language description. For example, if the target video contains a scene of a robotic arm grasping an object, the natural language description generated by the designated model might be: "A robotic arm grasps a shower head against a textured wall, but the image has a CG feel." This CG feel can be understood as indicating that the image is computer-generated.
[0102] The specified model can be a Multimodal Large Language Model (MLLM) or an Image-to-Text Generation Model. MLLM is a large language model that combines information from multiple modalities (such as text, images, and audio). It not only understands text information but also processes and generates related data in various forms, such as images and sounds, enabling cross-modal information conversion and integration. The Image-to-Text Generation Model is a multimodal model based on deep learning that converts an input image into natural language text describing the image content.
[0103] The natural language description information generated by a specified model (e.g., MLLM) based on the target video can be in free format, that is, the form and content of the natural language description information are not limited to a fixed structure or template, but can be expressed in a flexible and natural way.
[0104] For example, the classification judgment based on the preset language question can be performed by setting a preset language question (such as "whether it has a CG feel") for the second multimodal large language model to classify and judge the natural language description information. For example, if the natural language description information contains keywords such as "CG feel", "too perfect", or "lack of realism", the second multimodal large language model will determine that the target video has a CG feel.
[0105] For example, the natural language description information is mapped to a numerical judgment of the realism score. A second multimodal large language model can generate a numerical realism score based on the semantic features in the description information and combined with preset scoring rules or training data. The numerical realism score can be used to quantify the appearance and style characteristics of the target video.
[0106] The appearance style index value is determined based on the style judgment result, and the style judgment result can be mapped to a corresponding appearance style index value. For example, the index value corresponding to the style judgment result of "meeting the requirements of realism" is a first score, and the index value corresponding to the style judgment result of "having a CG feel" is a second score. The difference between the first score and the second score is greater than or equal to a preset difference threshold, which can be, for example, 0.5 or 0.6. The first score is, for example, 0.8, and the second score is, for example, 0.2.
[0107] In the above embodiment, in the process of performing appearance style analysis on the target video, natural language description information is generated by the first multimodal large language model, and the second multimodal large language model performs classification judgment or numerical scoring on it, and finally determines the appearance style index value, thereby quantitatively evaluating the realism and style rationality of the target video, providing a reliable data basis for the visual performance evaluation of the intelligent body behavior in the target video, and thus improving the reliability analysis effect of the embodied model.
[0108] In order to effectively perform a logical rationality analysis on the behavior of the intelligent agent in the target video, in some embodiments, the process of performing a logical rationality analysis on the target video may include: using a third multimodal large language model to analyze the logical rationality of each operation step in the target video to obtain a rationality score for each operation step; and determining the logical rationality index value based on the rationality scores of all operation steps.
[0109] Exemplarily, the target video can be input into a third multimodal large language model, and the third multimodal large language model can analyze the logical rationality of each operation step based on the key information of each operation step of the target intelligent agent performing the operation task in the target video to obtain a rationality score for each operation step.
[0110] For example, the third multimodal large language model can be used to extract key information from the current operation step of the target agent performing the operation task in the target video, and combine the key information of the previous operation step (i.e., one or more operation steps that occur before the current operation step) to generate context information. Based on the context information, the current operation step is logically rationally analyzed to obtain the rationality score of the current operation step. The key information of the operation step may include information such as the action, state and / or operated object of the target agent in the operation step. The context information can reflect the dependency and logical coherence between the operation steps.
[0111] For example, suppose the target video records the process of a robotic arm grabbing and moving a shower head. The sub-action execution order of the operation task is: the robotic arm moves from the initial position to the vicinity of the shower head, the robotic arm grabs the shower head, and the robotic arm moves the shower head to the target position. If the operation steps are logically reasonable (for example, the robotic arm moves smoothly from the initial position to the vicinity of the shower head), the rationality score is the first score; if the operation steps are not logically reasonable (for example, the robotic arm grabs the shower head without touching it), the rationality score is the second score. The difference between the first score and the second score is greater than or equal to a preset difference threshold, which can be, for example, 0.5 or 0.6. The first score is, for example, 0.9, and the second score is, for example, 0.1.
[0112] For example, the minimum value of the rationality scores of all operation steps can be selected as the logical rationality index value, or the weighted average value of the rationality scores of all operation steps can be calculated according to different weights assigned based on the importance of the operation steps to serve as the logical rationality index value.
[0113] In the above embodiment, by analyzing the logical rationality of each operation step in the target video and determining the logical rationality index value based on the rationality scores of all operation steps, the logical rationality of the intelligent agent's operation behavior in the target video can be effectively evaluated, thereby improving the reliability analysis effect of the embodied model.
[0114] An embodiment of the present application further provides a reliability analysis system for an embodied model, the system comprising a control module, the control module being configured to execute the reliability analysis method for an embodied model described in any one of the above embodiments.
[0115] An embodiment of the present application further provides a computer-readable storage medium storing a computer program. When the computer program is executed by a processor, the reliability analysis method of the embodied model described in any one of the above embodiments is implemented.
[0116] An embodiment of the present application further provides a computer program product, which includes a computer program. When the computer program is executed by a processor, the reliability analysis method of the embodied model described in any one of the above embodiments is implemented.
[0117] The computer program product may be a portable compact disc read-only memory (CD-ROM) and include program code, and may be run on a terminal device, such as a personal computer. However, the computer program product of the embodiments of the present application is not limited thereto, and the computer program product may be any combination of one or more computer-readable media.
[0118] See also Figure 4 , Figure 4 This is a structural block diagram of a computer device provided in an embodiment of the present application.
[0119] An embodiment of the present application further provides a computer device, comprising a memory and a processor, wherein the memory stores a computer program, and the processor implements any of the above-mentioned reliability analysis methods of the embodied model when executing the computer program.
[0120] The embodiments of the present application do not limit the computer device, which may be, for example, a local computer device, a cloud computer device, a distributed computer device, etc.
[0121] The computer device may include: a memory 110, a processor 120, and a communication interface 130. The memory 110, the processor 120, and the communication interface 130 are connected via an internal connection path.
[0122] The memory 110 is used to store computer programs. In some implementations, the computer programs may include codes for implementing the methods of the embodiments of the present application.
[0123] The processor 120 is configured to execute the computer program stored in the memory 110 to control the communication interface 130 to receive input data and information and output data such as operation results. In some implementations, when the solutions of the embodiments of the present application are implemented through software or firmware, the computer program for implementing the solutions of the embodiments of the present application may be stored in the processor 120 and executed by the processor 120.
[0124] The memory 110 may be a volatile memory or a non-volatile memory, or may include both volatile and non-volatile memories. The non-volatile memory may be a read-only memory (ROM), a programmable read-only memory (PROM), an erasable programmable read-only memory (EPROM), an electrically erasable programmable read-only memory (EEPROM), or a flash memory. The volatile memory may be a random access memory (RAM). It should be noted that the memory 110 described herein is intended to include, but is not limited to, any memory of these and other suitable types. As an example, the memory 110 includes a random access memory (RAM), a cache memory, and a read-only memory (ROM). The memory 110 stores a computer program, which can be executed by the processor 120, so that the processor 120 implements the steps of the reliability analysis method of the embodied model described in any one of the above.
[0125] The processor 120 may be a central processing unit (CPU), or other general-purpose processors, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA), or other programmable logic devices, discrete gate or transistor logic devices, or discrete hardware components. A general-purpose processor may be a microprocessor, or the processor 120 may be any conventional processor.
[0126] During implementation, each step of the above method can be completed by an integrated logic circuit of the hardware in the processor 120 or by instructions in the form of software. The method disclosed in conjunction with the embodiments of the present application can be directly embodied as being executed by a hardware processor, or can be executed by a combination of hardware and software modules in the processor 120. The software module can be located in a mature storage medium in the art, such as a random access memory, a flash memory, a read-only memory, a programmable read-only memory, or an electrically erasable programmable memory, a register, etc. The storage medium is located in the memory 110, and the processor 120 reads the information in the memory 110 and completes the steps of the above method in combination with its hardware. To avoid repetition, it will not be described in detail here.
[0127] In some implementations, in addition to the hardware units described above, the computer device may also include software modules, where the software modules may be, for example, an operating system, a basic input and output system (BIOS), application software, etc.
[0128] An operating system manages the hardware and / or software resources of a computer device and is the core and cornerstone of the computer. It handles basic tasks such as managing and allocating memory, prioritizing the supply and demand of system resources, controlling input and output devices, operating the network, and managing the file system. To facilitate user operation, most operating systems provide an interface for users to interact with the system.
[0129] The BIOS is used to run hardware initialization during the power-on boot phase and provide runtime services for the operating system and applications. In some implementations, the BIOS can also monitor and display the processor temperature and execute functions such as adjusting temperature protection strategies.
[0130] Application software, also known as an application program, is software written for a specific user purpose. It is a major category of computer software. For example, application software might be a program used for power control, temperature management, and other purposes.
[0131] It should be understood that the specific examples in this application are only intended to help those skilled in the art better understand the implementation methods of this application, rather than to limit the scope of protection of this application.
[0132] It can be understood that in various implementations of the present application, the size of the serial number of each process does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the present application.
[0133] It can be understood that the various embodiments described in this application can be implemented individually or in combination, and this application is not limited to this.
[0134] Unless otherwise indicated, all technical and scientific terms used in this application have the same meaning as those commonly understood by those skilled in the art in the art of this application. The terms used in this application are only for the purpose of describing specific embodiments and are not intended to limit the scope of this application. The singular forms "a," "above," and "the" used in this application and the appended claims are also intended to include plural forms unless the context clearly indicates otherwise.
[0135] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.
[0136] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the above-described embodiments may refer to the corresponding processes in other embodiments and will not be repeated here.
[0137] In the embodiments provided in this application, it should be understood that the disclosed systems, devices and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of units is only a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.
[0138] Units described as separate components may or may not be physically separate, and components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected according to actual needs to achieve the objectives of the technical solutions of this application.
[0139] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.
[0140] If the function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the relevant technology or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes a number of instructions for enabling a computer device (which can be a personal computer, server, or network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes: various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.
[0141] The above are only specific embodiments of the present application, but the scope of protection of this application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.
Claims
1. A reliability analysis method for an embodied model, characterized in that: The method comprises: Receiving a target video of a target agent performing an operation task generated by the embodied model; Performing one or more of motion consistency analysis, visual consistency analysis, and logical rationality analysis on the target video to obtain one or more of a motion consistency index value, a visual consistency index value, and a logical rationality index value; Based on one or more of the motion consistency index value, the visual consistency index value, and the logical rationality index value, a reliability analysis result of the embodied model is determined.
2. The reliability analysis method of the embodied model according to claim 1, characterized in that: The process of performing motion consistency analysis on the target video includes: Performing one or more of a trajectory consistency analysis and a dynamic consistency analysis on the target video to obtain one or more of a trajectory consistency index value and a dynamic consistency index value; The motion consistency index value is determined according to one or more of the trajectory consistency index value and the dynamic consistency index value.
3. The reliability analysis method of the embodied model according to claim 2, characterized in that: The process of performing trajectory consistency analysis on the target video includes: Performing end-effector detection on the target video and the reference video to obtain a target trajectory corresponding to the target agent and a reference trajectory corresponding to the reference agent, respectively. The target trajectory and the reference trajectory are spatial trajectories of the end-effectors of the corresponding agents; the reference video is a real video or a synthesized video of the reference agent performing the operation task, and the reference video meets reliability requirements; The trajectory consistency index value is determined based on the trajectory similarity between the target trajectory and the reference trajectory.
4. The reliability analysis method of the embodied model according to claim 3, characterized in that: The trajectory similarity is determined based on a bidirectional Hausdorff distance and / or a normalized dynamic time warping distance between the target trajectory and the reference trajectory.
5. The reliability analysis method of the embodied model according to claim 3, characterized in that: The process of performing dynamic consistency analysis on the target video includes: Processing the target trajectory based on a differential algorithm to obtain a target velocity sequence and a target acceleration sequence; Processing the reference trajectory based on the differential algorithm to obtain a reference velocity sequence and a reference acceleration sequence; A dynamic consistency index value is calculated based on the speed similarity between the target speed sequence and the reference speed sequence and / or the acceleration similarity between the target acceleration sequence and the reference acceleration sequence.
6. The reliability analysis method of the embodied model according to claim 5, characterized in that: The speed similarity is determined based on the Wasserstein distance between the target speed sequence and the reference speed sequence; and / or, The acceleration similarity is determined based on a Wasserstein distance between the target acceleration sequence and the reference acceleration sequence.
7. The reliability analysis method of the embodied model according to claim 1, characterized in that: The process of performing visual consistency analysis on the target video includes: Performing one or more of a subject consistency analysis, an overall consistency analysis, and an appearance style analysis on the target video to obtain one or more of a subject consistency index value, an overall consistency index value, and an appearance style index value; The visual consistency index value is determined according to one or more of the main consistency index value, the overall consistency index value, and the appearance style index value.
8. The reliability analysis method of the embodied model according to claim 7, characterized in that: The process of performing subject consistency analysis on the target video includes: Extracting subject image features of each frame in the target video; Calculating the cosine similarity between the subject image feature of the first frame and the subject image feature of each frame to obtain a first similarity score; Calculating the cosine similarity between the subject image features of adjacent frames to obtain a second similarity score; The subject consistency index value is determined based on the first similarity score and the second similarity score.
9. The reliability analysis method of the embodied model according to claim 7, characterized in that: The process of performing overall consistency analysis on the target video includes: Extracting overall image features of each frame in the target video; The cosine similarity between the overall image features of the first frame and the overall image features of each frame is calculated to obtain the overall consistency index value.
10. The reliability analysis method of the embodied model according to claim 7, characterized in that: The process of performing appearance style analysis on the target video includes: Inputting the target video into a designated model to generate free-format natural language description information; the designated model is a first multimodal large language model or an image-to-text generation model; Inputting the natural language description information into a second multimodal large language model for style judgment, wherein the style judgment includes a classification judgment based on a preset language question or a numerical judgment of mapping the natural language description information into a realism score; The appearance style index value is determined based on the style judgment result.
11. The reliability analysis method of the embodied model according to claim 1, characterized in that: The process of performing logical rationality analysis on the target video includes: Using the third multimodal large language model to analyze the logical rationality of each operation step in the target video to obtain a rationality score for each operation step; The logic rationality index value is determined based on the rationality scores of all operation steps.
12. A reliability analysis system for an embodied model, characterized in that: The system comprises a control module configured to execute the method according to any one of claims 1 to 11.
13. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.
14. A computer device, characterized in that: The computer device includes a memory and a processor, the memory stores a computer program, and the processor implements the method according to any one of claims 1 to 11 when executing the computer program.
15. A computer program product, characterized in that The computer program product comprises a computer program, and when the computer program is executed by a processor, the method according to any one of claims 1 to 11 is implemented.
Citation Information
Patent Citations
Dynamic spatial position simulation method of multi-target infrared simulation system
CN118211405A
Motion planning method, training method and system of intelligent body model
CN119398105A
Behavior instruction generation method and device, equipment, robot, medium and product
CN119610099A
Data processing method and apparatus, electronic device, storage medium, and program product
US20240193790A1
Image generation processing method and electronic device
WO2025066457A1
Cited By
Code verification method and device
CN120929354A
A code verification method and apparatus
CN120929354B
Personal intelligent multi-source data quality evaluation and verification method, device, medium and product
CN121188440A
Body intelligence multi-source data quality evaluation and verification method, device, medium and product
CN121188440B