Multi - perspective Video Reward Mechanism Learning System and Its Construction Method

Through the learning system of multi-view video reward mechanism, the multi-view video learning framework MVR combines multi-view video and task rewards, and solves the problems of setting states in the existing technology, such as relying on single-view images, and achieving more accurate visual feedback and more efficient robot learning effects.

CN119830993BActive Publication Date: 2025-06-20BEIJING INSTITUTE FOR GENERAL ARTIFICIAL INTELLIGENCE
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510299684.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-03-14
Publication Date
2025-06-20
Estimated Expiration
2045-03-14

AI Technical Summary

Technical Problem

The existing reinforcement learning technology based on visual language models has problems in complex robot tasks such as biased setting states, relying on single-view images, inability to adapt to dynamic environment changes, and lack of task semantic understanding.

Method used

A learning system for multi-view video reward mechanism is proposed. The multi-view video learning framework MVR combines multi-view video and task rewards, generate visual feedback through visual language big models, dynamically adjust the relative size between task rewards and visual language model rewards, and achieve a balance between visual feedback and task rewards.

Benefits of technology

The system can provide more precise visual feedback, a comprehensive understanding of robot motion, improve the accuracy and effectiveness of rewards, and help the robot learn the required motor skills faster and more accurately in complex tasks.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN119830993B_ABST
    Figure CN119830993B_ABST
Patent Text Reader

Abstract

The present invention provides a multi-view video reward mechanism learning system and method, including: a multi-view video evaluation subsystem that uses the multi-view video learning framework MVR to evaluate robot behavior based on multi-view videos; a visual feedback reward feedback strategy subsystem that generates visual feedback according to task text descriptions through a vision-language large model; obtaining accurate reward feedback based on the latest state correlation evaluation, and then more effectively adjusting the strategy; a visual feedback task reward balance subsystem that analyzes the importance degree of task reward feedback according to the degree to which the robot behavior approaches the expected goal through a task reward model, and dynamically adjusts the relative magnitude between task rewards and vision-language model rewards according to state correlation to balance visual feedback and task rewards; a multi-view video reward combination subsystem that combines multi-view videos and task rewards to provide more accurate visual feedback and learning effects in complex robot motion tasks.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of intelligent large model information processing for humanoid robot data analysis, and more specifically, to a multi-view video reward mechanism learning system and a method for constructing the same. Background Art

[0002] Currently, reinforcement learning (RL) has shown great potential in solving complex tasks, and reward design is a core part of it. In the field of robot tasks, traditional reward setting methods mainly encode quantitative goals such as motion and distance. Taking the running task of a robot as an example, the task reward will prompt the robot to track the set forward speed. To further optimize the robot's motion performance, researchers often draw on animal kinematics principles and combine task rewards with energy consumption penalties; this method can cultivate the robot's agile motion ability to a certain extent. For example, in some studies, by reasonably adjusting the weights of rewards and penalties, the robot can use energy more efficiently during motion and achieve more flexible movements.

[0003] The skill learning process of humans is significantly different from the traditional robot task reward design. When a human coach guides a learner, they often first evaluate the learner's motion pattern visually and give corresponding feedback, and then introduce quantitative goals to further improve the learner's skill level. This way of relying on visual guidance first and then conducting quantitative training has inspired researchers to explore integrating visual feedback into robot task rewards to enhance the robot's skill learning ability. With the continuous development of visual language models (VLMs), their powerful image-text processing capabilities have provided a new opportunity for this research direction.

[0004] Reward methods based on VLMs have gradually become a research hotspot. Some studies use the ability of VLMs to calculate image-text similarity to guide the robot towards a state that matches the task description. For example, by constructing a reward mechanism based on image-text similarity, the robot can understand the visual requirements of the task to a certain extent and make corresponding actions. However, these methods have obvious defects. On the one hand, they tend to select the state with the highest image-text similarity score, and this preference may lead to the inability to comprehensively consider various relevant states in actual tasks. In the robot running task, observing the robot's running posture from different angles will result in different image-text similarity scores. Due to the bias of the method itself towards high-score states, the robot may over-pursue a certain set posture (such as the set leg-lifting posture) and ignore other equally important postures in actual running, which is not conducive to the robot learning comprehensive and reasonable running actions.

[0005] On the other hand, most existing methods rely on images captured from a single perspective. However, in practical scenarios, a single perspective cannot cover all the key aspects of a robot's movement. For example, in complex operation tasks of a robot, from a fixed perspective, it may not be possible to observe all the detailed movements of the robot's arm and the interaction between the robot and the surrounding environment. This limits the comprehensive understanding and accurate evaluation of the robot's movement, thereby affecting the accuracy and effectiveness of the reward, making it difficult for the robot to obtain comprehensive and accurate feedback during the learning process and unable to achieve the optimal learning effect.

[0006] In the field of reinforcement learning (RL), existing reward techniques based on vision-language models (VLM) have various defects, which seriously restrict the learning and decision-making abilities of robots in complex tasks.

[0007] Bias towards the set state: Existing techniques rely on image-text similarity to generate reward signals, which makes them generally biased towards the set state. In practical applications, tasks often involve multiple visually different but task-related states. Take the example of a robot performing a complex dance movement task. Different dance poses and action sequences are all important for task completion. However, the reward mechanism based on image-text similarity will tend to favor those states with the highest similarity to the given text description in the image. This means that if a set dance pose has the highest match with the description in the image, even if other poses are equally crucial for a coherent dance movement sequence, the robot will overly pursue the set pose, resulting in single and inflexible movements and being unable to learn a complete and smooth dance movement sequence. This bias not only affects the robot's comprehensive understanding and handling of complex tasks but may also cause the robot to fall into a local optimal solution and be unable to find the true optimal strategy.

[0008] Limitations of single - perspective image information: Currently, most methods rely only on images captured from a single perspective, which leads to severely insufficient information acquisition. In real - world task scenarios, the movement of a robot is multi - dimensional and all - around, and a single perspective cannot capture all the important aspects of the robot's movement. Take the navigation task of a robot in an indoor environment as an example. A single perspective may only be able to see part of the scene in front of the robot and cannot obtain information about its sides and rear, such as whether there are obstacles approaching and the overall layout of the surrounding environment. This makes it difficult for the robot to make accurate and comprehensive judgments when making decisions due to lack of sufficient basis. In addition, different perspectives may present completely different visual information. Relying only on single - perspective images may miss key clues, resulting in inaccurate assessment of the robot's movement. For example, in a robot operation task, the subtle contact between the robot's arm and an object may not be observable from a certain angle, making it impossible to accurately judge whether the operation is successful, thus affecting the accurate awarding of rewards. As a result, the robot cannot obtain effective feedback during the learning process, hindering its effective learning and execution of tasks.

[0009] Inability to adapt to dynamic environmental changes: Due to the fact that existing technologies are mainly based on fixed image - text matching patterns, it is difficult to adapt to dynamically changing environments. In practical applications, the environment is often complex and changeable. The attributes of objects such as position, shape, and color may change at any time, and the requirements of tasks may also change over time. The reward mechanism based on fixed patterns cannot adjust the reward strategy for the robot in a timely manner. In a rescue task participated by a robot, the environment at the disaster site may change continuously. New obstacles may appear, and the location of the rescue target may also change. At this time, the reward technology relying on fixed image - text similarity cannot quickly adapt to these changes, resulting in the reward signal received by the robot being out of touch with the actual task requirements, unable to effectively guide the robot to make correct decisions, and reducing its response ability and task execution efficiency in dynamic environments.

[0010] Lack of in-depth understanding of task semantics: The existing technologies have obvious deficiencies in understanding task semantics. Although image-text similarity can reflect the matching degree between visual information and text description to a certain extent, it cannot deeply understand the semantic logic and context information behind the task. In some tasks that require reasoning and planning, this defect is particularly prominent. In the home service task of a robot, the robot is required to find the red cup on the table and take it to the kitchen. The reward technology based on image-text similarity may only focus on the matching of the appearance features of the cup with "red cup", and cannot understand the semantics and sequence requirements of actions such as "find" and "take to the kitchen", nor can it consider context information such as the spatial relationship between the table and the kitchen. This may cause confusion when the robot executes the task, and it cannot complete the task according to the correct logic, affecting the quality and efficiency of task completion. Such problems remain to be solved; therefore, it is necessary to propose a multi-perspective video reward mechanism learning system and its construction method to at least partially solve the problems existing in the existing technologies. Summary of the Invention

[0011] A series of simplified concepts are introduced in the Summary of the Invention section, which will be further described in detail in the Detailed Implementation section; the Summary of the Invention section of the present invention does not mean to attempt to define the key features and essential technical features of the claimed technical solution, nor does it mean to attempt to determine the protection scope of the claimed technical solution.

[0012] To at least partially solve the above problems, the present invention provides a multi-perspective video reward mechanism learning system, including:

[0013] A multi-perspective video evaluation sub-system, which uses the multi-perspective video learning framework MVR to evaluate the robot's behavior according to multi-perspective videos;

[0014] A visual feedback reward feedback strategy sub-system, which generates visual feedback according to the task text description through a vision-language large model to improve the robot's motion skill learning; according to the latest state correlation evaluation, accurate reward feedback is obtained, and then the strategy is adjusted more effectively;

[0015] A visual feedback task reward balance sub-system, which, through a task reward model, analyzes the importance degree of task reward feedback according to the degree to which the robot's behavior approaches the expected goal when the robot's behavior approaches the expected goal, and dynamically adjusts the relative magnitude between the task reward and the vision-language model reward according to the state correlation to balance visual feedback and task reward;

[0016] A multi-perspective video reward combination sub-system, which combines multi-perspective videos and task rewards to provide more accurate visual feedback and learning effects in complex robot motion tasks.

[0017] Preferably, the multi - perspective video evaluation sub - system includes:

[0018] Sequence acquisition dataset sub - system. During the process of the intelligent agent robot performing tasks in the environment, it collects the robot state sequence at a set frequency, and converts the robot state sequence into a multi - perspective video; at the same time, it refines the target task into a concise text description; calculates the similarity between the video and the text through VLM; constructs a dataset D; the dataset D includes: the robot state sequence, the multi - perspective video corresponding to the robot state sequence, and the similarity score between the video and the text;

[0019] Matching paired comparison sub - system. The multi - perspective video learning framework MVR, through the matching paired comparison method, uses the sports game statistical model to keep the comparison results of the robot state sequence and the video comparison results as consistent as possible, and more accurately learns the state correlation;

[0020] Regularization interference reduction sub - system. The multi - perspective video learning framework MVR performs regularization state representation to reduce the interference caused by perspective changes in the multi - view video.

[0021] Preferably, the visual feedback reward feedback strategy sub - system includes:

[0022] Trajectory rendering behavior vision sub - system. During the entire online reinforcement learning process, the multi - perspective video learning framework MVR renders the trajectory generated by the robot into a video at a set rendering frequency TRender; and timely obtains the behavior vision information of the robot;

[0023] Similarity score video embedding sub - system. The rendered video is input into the VLM to obtain the video - text similarity score and the video embedding; these new information are used to update the dataset D to keep the latest robot behavior and task matching situation in the dataset;

[0024] Buffer state sequence selection and storage sub - system. The multi - perspective video learning framework MVR maintains a reference buffer Dref, which is specifically used to store those robot state sequences with higher video - text similarity scores; these robot state sequences are considered to be more relevant to the task goal;

[0025] The multi - perspective video learning framework MVR updates the model used to measure the state correlation with the samples in the dataset at a set update frequency TUpdate, continuously optimizing the understanding and judgment of the state correlation; uses the off - policy strategy algorithm. After each model update, the rewards in the robot replay buffer need to be recalculated to maintain the consistency of the rewards, so that the robot can obtain accurate reward feedback based on the latest state correlation evaluation, and then adjust the strategy more effectively.

[0026] Preferably, the visual feedback task reward balance sub - system includes:

[0027] Multi-class reward flexible balancing subsystem. The reward model designed by the multi-view video learning framework MVR realizes the flexible balance between the robot task reward rtask and the visual guidance reward rVLM;

[0028] Visual guidance reward adjustment subsystem. First, the visual guidance reward rVLM enables the robot to obtain the correct action mode, and then fine-tuning is carried out relying on the quantitative target robot task reward rtask;

[0029] The flexible balance between the robot task reward rtask and the visual guidance reward rVLM realized by the reward model designed by the multi-view video learning framework MVR includes: The visual guidance reward rVLM encourages the robot to explore states related to the task; However, when the robot gradually approaches these states, the influence of rVLM should gradually decrease, and at this time, the robot task reward rtask should play a dominant role.

[0030] Preferably, the multi-view video reward combination subsystem includes:

[0031] Policy relevance matching measurement subsystem. The multi-view video learning framework MVR introduces policy relevance, comprehensively considering the relevance of states and the access frequency of these states by the policy; More comprehensively measure the matching degree between the policy and the task;

[0032] Reward-guided exploration learning subsystem. The multi-view video learning framework MVR uses approximate calculation to provide the robot with a reward signal based on state relevance, guiding the robot to explore and learn towards states more consistent with the task goals;

[0033] Policy relevance includes: When the multi-view video learning framework MVR learns a policy, it will strive to maximize a set objective function; This objective function contains two key parts. One part is the value function of the policy, which reflects the expected return that can be obtained by performing actions according to this policy; The other part is related to policy relevance, and its influence degree is adjusted through a special function; There is a hyperparameter w, which can be adjusted according to the specific characteristics and requirements of the task to balance the roles of these two parts; As the robot's policy gets closer and closer to the policy that best matches the task description, the influence of this part related to state relevance will gradually weaken, realizing the dynamic transformation of the reward function from mainly based on rVLM to mainly based on rtask.

[0034] The present invention provides a multi-view video reward mechanism learning system, including:

[0035] S1, using the multi-view video learning framework MVR to evaluate the robot's behavior according to multi-view video learning;

[0036] S2. Generate visual feedback through a vision-language large model according to the task text description to improve the learning of robot motion skills; obtain accurate reward feedback based on the latest state correlation evaluation, and then adjust the strategy more effectively;

[0037] S3. Through the task reward model, when the robot's behavior is close to the expected goal, analyze the importance degree of the task reward feedback according to the degree of the robot's behavior approaching the expected goal, dynamically adjust the relative size between the task reward and the vision-language model reward according to the state correlation, and balance the visual feedback and the task reward;

[0038] S4. Combine multi-view videos and task rewards to provide more accurate visual feedback and learning effects in complex robot motion tasks.

[0039] Preferably, S1 includes:

[0040] S11. During the process of the intelligent agent robot performing tasks in the environment, collect the robot state sequence at a set frequency and convert the robot state sequence into a multi-view video; at the same time, refine the target task into a concise text description; calculate the similarity between the video and the text through the VLM; construct a dataset D; the dataset D includes: the robot state sequence, the multi-view video corresponding to the robot state sequence, and the similarity score between the video and the text;

[0041] S12. The multi-view video learning framework MVR uses the sports game statistical model through the paired comparison method to keep the comparison results of the robot state sequence as consistent as possible with the video comparison results, and more accurately learn the state correlation;

[0042] S13. The regularization interference reduction subsystem, the multi-view video learning framework MVR performs regularized state representation to reduce the interference caused by the perspective change in the multi-view video.

[0043] Preferably, S2 includes:

[0044] S21. During the entire online reinforcement learning process, the multi-view video learning framework MVR renders the trajectory generated by the robot into a video at the set rendering frequency TRender; obtain the visual information of the robot's behavior in a timely manner;

[0045] S22. The rendered video will be input into the VLM to obtain the video-text similarity score and the video embedding; these new information will be used to update the dataset D to keep the latest robot behavior and task matching situation in the dataset;

[0046] In S23, the multi-view video learning framework MVR maintains a reference buffer Dref that specifically stores those robot state sequences with relatively high video-text similarity scores; these robot state sequences are considered to be more relevant to the task objectives.

[0047] The multi-view video learning framework MVR updates the model for measuring state relevance using samples in the dataset at a set update frequency TUpdate, continuously optimizing the understanding and judgment of state relevance; the off-policy strategy algorithm is used. After each model update, the rewards in the robot replay buffer need to be recalculated to maintain the consistency of the rewards, enabling the robot to obtain accurate reward feedback based on the latest state relevance assessment, and thus adjust the strategy more effectively.

[0048] Preferably, S3 includes:

[0049] In S31, the reward model designed by the multi-view video learning framework MVR achieves a flexible balance between the robot task reward rtask and the visual guidance reward rVLM.

[0050] In S32, first, the robot obtains the correct action mode through the visual guidance reward rVLM, and then makes fine adjustments relying on the quantitative target robot task reward rtask.

[0051] The reward model designed by the multi-view video learning framework MVR achieving a flexible balance between the robot task reward rtask and the visual guidance reward rVLM includes: the visual guidance reward rVLM encourages the robot to explore states related to the task; however, when the robot gradually approaches these states, the influence of rVLM should gradually decrease, and at this time, the robot task reward rtask should play a dominant role.

[0052] Preferably, S4 includes:

[0053] In S41, the multi-view video learning framework MVR introduces policy relevance, comprehensively considering the relevance of states and the access frequency of these states by the policy; more comprehensively measuring the matching degree between the policy and the task.

[0054] In S42, the multi-view video learning framework MVR uses approximate calculation to provide the robot with a reward signal based on state relevance, guiding the robot to explore and learn towards states more consistent with the task objectives.

[0055] Policy relevance includes: When the multi-view video learning framework MVR learns a policy, it tries to maximize a set objective function; this objective function consists of two key parts. One part is the value function of the policy, which reflects the expected reward that can be obtained by performing actions according to this policy; the other part is related to policy relevance and its influence degree is adjusted through a special function; there is a hyperparameter w, which can be adjusted according to the specific characteristics and requirements of the task to balance the roles of these two parts; as the robot's policy gets closer and closer to the policy that best matches the task description, this part of the influence related to state relevance will gradually weaken, realizing the dynamic transformation of the reward function from mainly relying on rVLM to mainly relying on rtask.

[0056] Compared with the prior art, the present invention has at least the following beneficial effects:

[0057] The present invention provides a multi-view video reward mechanism learning system and a construction method thereof. Through a multi-view video evaluation subsystem, using a multi-view video learning framework MVR, the behavior of a robot is evaluated based on multi-view videos; a visual feedback reward feedback strategy subsystem, through a visual language large model, generates visual feedback according to the task text description to improve the learning of the robot's motion skills; according to the latest state correlation evaluation, accurate reward feedback is obtained, and then the strategy is adjusted more effectively; a visual feedback task reward balance subsystem, through a task reward model, when the robot's behavior approaches the expected goal, analyzes the importance degree of the task reward feedback according to the degree to which the robot's behavior approaches the expected goal, and dynamically adjusts the relative magnitude between the task reward and the visual language model reward according to the state correlation to balance the visual feedback and the task reward; a multi-view video reward combination subsystem, which combines multi-view videos and task rewards to provide more accurate visual feedback and learning effects in complex robot motion tasks; can evaluate the behavior of the robot using multi-view videos through the multi-view video learning framework MVR, generate visual feedback according to the task text description to improve the learning effect of the robot's motion skills; through the reward function, when the robot's behavior approaches the expected goal, analyze the importance degree of the task reward feedback according to the degree to which the robot's behavior approaches the expected goal, and solve the balance problem between visual feedback and task reward; by combining multi-view videos and task rewards, the multi-view video learning framework MVR can provide more accurate visual feedback in complex robot complex motion tasks and improve the learning effect; a reinforcement learning framework based on multi-view videos and a visual language model (VLM) enables the robot to master complex motion skills on a humanoid robot simulation platform; the reward function dynamically adjusts the relative magnitude between the task reward and the visual language model reward according to the state correlation to achieve the balance between the two; improves the learning effect of the robot in complex tasks and enables the robot to learn the required motion skills faster and more accurately on a humanoid robot simulation platform; effectively solves the problems of bias towards set states and dependence on single-view images existing in the reward method based on the visual language model, comprehensively understands the robot's motion, and improves the accuracy and effectiveness of the reward; has great significance and remarkable effects.

[0058] Regarding the multi-view video reward mechanism learning system and the construction method thereof of the present invention, other advantages, objectives, and features of the present invention will be partially reflected by the following description and partially understood by those skilled in the art through the research and practice of the present invention. BRIEF DESCRIPTION OF THE DRAWINGS

[0059] The drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation to the present invention. In the drawings:

[0060] Figure 1This is a diagram of an embodiment of the multi-view video reward mechanism learning system of the present invention.

[0061] Figure 2 This is a diagram of another embodiment of the multi-view video reward mechanism learning system of the present invention.

[0062] Figure 3 This is a diagram of an embodiment of the method for building the multi-view video reward mechanism learning system of the present invention. Detailed implementation manners

[0063] The following further elaborates on the present invention in conjunction with the accompanying drawings and embodiments, so that those skilled in the art can implement it with reference to the specification; as Figures 1 - 3 shown, the present invention provides a multi-view video reward mechanism learning system, including:

[0064] A multi-view video evaluation subsystem, which uses the multi-view video learning framework MVR to evaluate the behavior of the robot based on multi-view videos.

[0065] A visual feedback reward feedback strategy subsystem, which generates visual feedback according to the task text description through a visual language large model to improve the learning of the robot's motion skills; obtains accurate reward feedback based on the latest state correlation evaluation, and then adjusts the strategy more effectively.

[0066] A visual feedback task reward balance subsystem, which analyzes the importance degree of the task reward feedback according to the degree of the robot's behavior approaching the expected goal through a task reward model when the robot's behavior approaches the expected goal, and dynamically adjusts the relative size between the task reward and the visual language model reward according to the state correlation to balance the visual feedback and the task reward.

[0067] A multi-view video reward combination subsystem, which combines multi-view videos and task rewards to provide more accurate visual feedback and learning effects in complex robot motion tasks.

[0068] The principles and effects of the above technical solution are as follows: The present invention provides a multi-view video reward mechanism learning system, including: a multi-view video evaluation subsystem that uses a multi-view video learning framework MVR to evaluate robot behavior based on multi-view videos; a visual feedback reward feedback strategy subsystem that generates visual feedback according to task text descriptions through a visual language large model to improve robot motion skill learning; obtains accurate reward feedback based on the latest state correlation evaluation, and then adjusts the strategy more effectively; a visual feedback task reward balance subsystem that, through a task reward model, analyzes the importance degree of task reward feedback according to the degree of the robot's behavior approaching the expected goal when the robot's behavior is close to the expected goal, and dynamically adjusts the relative magnitude between task rewards and visual language model rewards according to state correlation to balance visual feedback and task rewards; a multi-view video reward combination subsystem that combines multi-view videos and task rewards to provide more accurate visual feedback and learning effects in complex robot motion tasks; can use multi-view videos to evaluate robot behavior through the multi-view video learning framework MVR, and generate visual feedback according to task text descriptions to improve the learning effect of robot motion skills; through the reward function, when the robot's behavior is close to the expected goal, the importance degree of task reward feedback can be analyzed according to the degree of the robot's behavior approaching the expected goal, and the balance problem between visual feedback and task rewards can be solved; by combining multi-view videos and task rewards, the multi-view video learning framework MVR can provide more accurate visual feedback in complex robot complex motion tasks and improve the learning effect; a reinforcement learning framework based on multi-view videos and a visual language model (VLM) enables the robot to master complex motion skills on a humanoid robot simulation platform; the reward function dynamically adjusts the relative magnitude between task rewards and visual language model rewards according to state correlation to achieve the balance between the two; among them, Top-k sampling samples from the top-k basic unit tokens, including randomly sampling only from the k words with the highest probability, and the randomness allows other basic unit tokens with higher scores or probabilities to have the opportunity to be selected; improves the learning effect of the robot in complex tasks, and enables the robot to learn the required motion skills faster and more accurately on a humanoid robot simulation platform; effectively solves the problems of bias setting states and dependence on single-view images existing in the reward method based on the visual language model, comprehensively understands robot motion, and improves the accuracy and effectiveness of rewards.

[0069] In one embodiment, the multi-view video evaluation subsystem includes:

[0070] Sequence acquisition dataset subsystem. During the process of the agent robot performing tasks in the environment, it acquires the robot state sequence at a set frequency and converts the robot state sequence into a multi-view video. At the same time, it refines the target task into a concise text description. It calculates the similarity between the video and the text through VLM and constructs a dataset D. The dataset D includes: the robot state sequence, the multi-view video corresponding to the robot state sequence, and the similarity score between the video and the text.

[0071] Matching pair comparison subsystem. The multi-view video learning framework MVR uses the matching pair comparison method and utilizes the sports game statistical model to keep the comparison results of the robot state sequence and the video comparison results as consistent as possible, and more accurately learn the state correlation.

[0072] Regularization to reduce interference subsystem. The multi-view video learning framework MVR performs regularization state representation to reduce the interference caused by perspective changes in the multi-view video.

[0073] The principle and effect of the above technical solution are as follows: The multi-view video evaluation subsystem includes: the sequence acquisition dataset subsystem. During the process of the agent robot performing tasks in the environment, it acquires the robot state sequence at a set frequency and converts the robot state sequence into a multi-view video. At the same time, it refines the target task into a concise text description. It calculates the similarity between the video and the text through VLM and constructs a dataset D. The dataset D includes: the robot state sequence, the multi-view video corresponding to the robot state sequence, and the similarity score between the video and the text. The matching pair comparison subsystem. The multi-view video learning framework MVR uses the matching pair comparison method and utilizes the sports game statistical model to keep the comparison results of the robot state sequence and the video comparison results as consistent as possible, and more accurately learn the state correlation. The sports game statistical model includes the Bradley-Terry model. The regularization to reduce interference subsystem. The multi-view video learning framework MVR performs regularization state representation to reduce the interference caused by perspective changes in the multi-view video.

[0074] Keeping the comparison results of the robot state sequence and the video comparison results as consistent as possible and more accurately learning the state correlation includes: In the actual task, the matching degree between any two videos and the task description will form a probabilistic comparison relationship, and the multi-view video learning framework MVR can also reflect this relationship at the robot state sequence level. For example, if one video is more in line with the task description than another video, then a similar relationship should also be presented between the corresponding robot state sequences. By continuous adjustment, the similarity between the robot state sequence and the video sequence is made as aligned as possible, so as to avoid the bias towards the set state caused by the calculation of the image-text similarity and more comprehensively consider the correlation between different states and tasks.

[0075] The multi-view video learning framework MVR regularizes the state representation to reduce the interference caused by perspective changes in multi-view videos, including: the similarity scores of videos from different perspectives may fluctuate, which may affect the accurate judgment of state relevance; to reduce this interference, the multi-view video learning framework MVR keeps the similarity structure of the robot state sequence and the video embedding matched. Specifically, the state is processed through a set function to map the state in the state space to a suitable vector space, and then some operations are used to measure the relevance between the state and the task; during the calculation process, different perspectives are randomly and uniformly sampled to calculate the video embedding, and the state representation is adjusted with this as a reference, so that the state representation can reflect the common information from different perspectives, reduce the interference of perspective setting factors, and thus learn the state relevance more stably and accurately.

[0076] In one embodiment, the visual feedback reward feedback strategy subsystem includes:

[0077] The trajectory rendering behavior vision subsystem. During the entire online reinforcement learning process, the multi-view video learning framework MVR renders the trajectory generated by the robot into a video according to the set rendering frequency TRender; and timely obtains the behavior vision information of the robot.

[0078] The similarity score video embedding subsystem. The rendered video is input into the VLM to obtain the video-text similarity score and the video embedding; these new information are used to update the dataset D to keep the latest robot behavior and task matching situation in the dataset.

[0079] The buffer state sequence selection and storage subsystem. The multi-view video learning framework MVR maintains a reference buffer Dref, which is specifically used to store those robot state sequences with relatively high video-text similarity scores; these robot state sequences are considered to be more relevant to the task objectives.

[0080] The multi-view video learning framework MVR updates the model used to measure the state relevance with the samples in the dataset at the set update frequency TUpdate, continuously optimizing the understanding and judgment of the state relevance; the offline policy algorithm is used. After each update of the model, the rewards in the robot replay buffer need to be recalculated to maintain the consistency of the rewards, so that the robot can obtain accurate reward feedback based on the latest state relevance evaluation, and then adjust the strategy more effectively.

[0081] The principles and effects of the above technical solution are as follows: The visual feedback reward feedback strategy subsystem includes: a trajectory rendering behavior vision subsystem. During the entire online reinforcement learning process, the multi-view video learning framework MVR renders the trajectory generated by the robot into a video according to the set rendering frequency TRender; timely obtains the behavior vision information of the robot; a similarity score video embedding subsystem. The rendered video is input into the VLM to obtain the video-text similarity score and video embedding; these new information are used to update the dataset D to keep the latest robot behavior and task matching situation in the dataset; a buffer state sequence selection and storage subsystem. The multi-view video learning framework MVR maintains a reference buffer Dref, which is specifically used to store those robot state sequences with relatively high video-text similarity scores; these robot state sequences are considered to be more relevant to the task objectives; the multi-view video learning framework MVR updates the model for measuring state relevance using the samples in the dataset at the set update frequency TUpdate, continuously optimizing the understanding and judgment of state relevance; an off-policy algorithm is used. After each update of the model, the rewards in the robot replay buffer need to be recalculated to maintain the consistency of the rewards, enabling the robot to obtain accurate reward feedback based on the latest state relevance evaluation, and then adjusting the strategy more effectively;

[0082] MVR design idea: On the one hand, through the way of pairwise comparison and using the sports competition statistical model, make the comparison results of the state sequences as consistent as possible with the comparison results of the videos; in the actual task, the matching degree of any two videos with the task description will form a probabilistic comparison relationship.

[0083] The probability that a video o better matches the task description l than another video o′ is given by the following formula:

[0084] where, is the sigmoid function, represents the cosine similarity of the VLM model.

[0085] MVR expects this relationship to be reflected at the state sequence level as well. For example, if a video better matches the task description than another video, then a similar relationship should also be presented between the corresponding state sequences. Similarly, the probability that a state sequence s better matches l than another state sequence s′ is:

[0086]

[0087] where, n(s) is the length of s, represents the state relevance function; it should be noted that the state relevance is determined based on the video, thus overcoming the deviation from the set state; given two samples in the dataset D and , minimize the following loss function:

[0088]

[0089] Here, we utilize the property 1 - σ(x) = σ(-x); by continuously adjusting, we align the similarity between the state sequence and the video sequence as much as possible, thereby avoiding the bias towards the set state caused by the calculation of image - text similarity, and more comprehensively considering the relevance between different states and tasks.

[0090] On the other hand, the interference of the perspective setting is independent of the similarity between the underlying states and states, because the change of perspective does not affect the state itself. Therefore, this interference can be eliminated by keeping the similarity structure between the state sequences matched with the similarity structure of the video embedding, where the perspective is uniformly randomly sampled. Using the state encoder and the vector to parameterize the correlation function, where d is the dimension of the state embedding; and are both normalized; the correlation calculation of state s is as follows:

[0091] where 〈·〉 represents the dot - product operation; for the state sequence s, the following formula is used as its representation:

[0092] where o represents the video stream; since the perspective is uniformly randomly sampled, this regularization will generate an average state representation, thereby reducing the interference of the perspective setting. Finally, the overall objective function for learning state correlation combines the matching term and the regularization term:

[0093] Lrel is the overall objective function of state correlation, Lmatching is the matching term, and Lreg is the regularization term.

[0094] In one embodiment, the visual feedback task reward balance subsystem includes:

[0095] A multi - class reward flexible balance subsystem, where the reward model designed by the multi - view video learning framework MVR realizes the flexible balance between the robot task reward rtask and the visual guidance reward rVLM;

[0096] A visual guidance reward adjustment subsystem, which first enables the robot to obtain the correct action mode through the visual guidance reward rVLM, and then makes fine - tuning relying on the quantitative target robot task reward rtask;

[0097] The reward model designed by the multi-view video learning framework MVR for the implementation of the Multi-view Video Learning (MVR) framework enables a flexible balance between the robot task reward rtask and the vision-guided reward rVLM, including: the vision-guided reward rVLM encourages the robot to explore task-related states; however, when the robot gradually approaches these states, the influence of rVLM should gradually decrease, and at this time, the robot task reward rtask should play a dominant role.

[0098] The principle and effect of the above technical solution are as follows: The visual feedback task reward balancing subsystem includes: a multi-class reward flexible balancing subsystem, and the reward model designed by the multi-view video learning framework MVR enables a flexible balance between the robot task reward rtask and the vision-guided reward rVLM; a vision-guided reward adjustment subsystem, which first enables the robot to obtain the correct action pattern through the vision-guided reward rVLM, and then relies on the quantitative target robot task reward rtask for fine adjustment; the reward model designed by the multi-view video learning framework MVR enables a flexible balance between the robot task reward rtask and the vision-guided reward rVLM, including: the vision-guided reward rVLM encourages the robot to explore task-related states; however, when the robot gradually approaches these states, the influence of rVLM should gradually decrease, and at this time, the robot task reward rtask should play a dominant role.

[0099] The reward function designed by MVR aims to achieve a flexible balance between the task reward rtask and the VLM-based reward rVLM. Specifically, the role of rVLM is to encourage the agent to explore task-related states. However, when the agent gradually approaches these states, the influence of rVLM should gradually decrease, and at this time, rtask should play a dominant role. This design concept is similar to the human skill learning process, where the agent first obtains the correct action pattern through rVLM and then relies on rtask for fine adjustment.

[0100] To achieve this balance, MVR introduces the concept of policy relevance. Briefly, it comprehensively considers the relevance of states and the access frequency of these states by the policy to more comprehensively measure the matching degree between the policy and the task. Based on policy relevance, when MVR learns a policy, it will strive to maximize a set target function. This target function consists of two key parts. One part is the value function of the policy, which reflects the expected return obtained by performing actions according to this policy; the other part is related to policy relevance and adjusts its influence degree through a special function. The policy relevance function can be adjusted according to the specific characteristics and requirements of the task to balance the roles of these two parts. As the agent's policy gets closer and closer to the policy that best matches the task description, the influence of this part related to state relevance will gradually weaken, realizing the dynamic transformation of the reward function from being mainly based on rVLM to being mainly based on rtask:

[0101] Among them, refers to the value function of reinforcement learning under the policy , and refers to the policy that best matches the language . To maximize the log-likelihood and make better match under the BT model than , is introduced to adjust the scales of the two terms. The function log(σ(x)) is monotonically decreasing when x ≤ 0, and the influence of weakens as gets closer and closer to .

[0102] When specifically calculating rVLM, MVR adopts an approximate method. Since it is difficult to directly sample from the ideal policy, MVR uses the previously mentioned reference set Dref for approximate sampling. The state sequences in the reference set are screened according to the video-text similarity scores and are considered to be highly correlated with the target policy. By sampling from the reference set, rVLM can be approximately calculated, providing a reward signal based on state correlation for the agent to guide the agent to explore and learn towards states that are more consistent with the task objectives:

[0103] Among them, E[] represents the expectation, represents the state sampled from the policy , represents the correlation function, σ and

[0104] is the sigmoid function.

[0105] In one embodiment, the multi-view video reward combination subsystem includes:

[0106] The policy correlation matching measurement subsystem. The multi-view video learning framework MVR introduces policy correlation, comprehensively considering the correlation of states and the access frequency of the policy to these states; more comprehensively measuring the matching degree of the policy and the task;

[0107] Policy relevance includes: When the multi-view video learning framework MVR learns a policy, it tries to maximize a set objective function. This objective function consists of two key parts. One part is the value function of the policy, which reflects the expected return obtained by performing actions according to this policy. The other part is related to policy relevance and its influence degree is adjusted through a special function. The policy relevance function can be adjusted according to the specific characteristics and requirements of the task to balance the roles of these two parts. As the robot's policy gets closer and closer to the policy that best matches the task description, the influence of this part related to state relevance will gradually weaken, realizing the dynamic transformation of the reward function from mainly relying on rVLM to mainly relying on rtask.

[0108] The principle and effect of the above technical solution are as follows: The multi-view video reward combination subsystem includes:

[0109] The policy relevance matching measurement subsystem. The multi-view video learning framework MVR introduces policy relevance, comprehensively considering the relevance of states and the access frequency of these states by the policy, and more comprehensively measures the matching degree between the policy and the task.

[0110] The reward-guided exploration learning subsystem. The multi-view video learning framework MVR uses approximate calculation to provide the robot with a reward signal based on state relevance, guiding the robot to explore and learn towards states that are more consistent with the task goals.

[0111] Policy relevance includes: When the multi-view video learning framework MVR learns a policy, it tries to maximize a set objective function. This objective function consists of two key parts. One part is the value function of the policy, which reflects the expected return obtained by performing actions according to this policy. The other part is related to policy relevance and its influence degree is adjusted through a special function. There is a hyperparameter w, which can be adjusted according to the specific characteristics and requirements of the task to balance the roles of these two parts. As the robot's policy gets closer and closer to the policy that best matches the task description, the influence of this part related to state relevance will gradually weaken, realizing the dynamic transformation of the reward function from mainly relying on rVLM to mainly relying on rtask.

[0112] The multi-view video learning framework MVR uses approximate calculations to provide the robot with a state-correlated reward signal, guiding the robot to explore and learn towards a state that is more consistent with the task goal, including: The multi-view video learning framework MVR uses an approximate calculation method to calculate rVLM, providing the robot with a state-correlated reward signal, guiding the robot to explore and learn towards a state that is more consistent with the task goal; The multi-view video learning framework MVR uses the previously mentioned reference set Dref for approximate sampling; The robot state sequences in the reference set are filtered according to the video-text similarity score and are considered to be highly correlated with the target policy; By sampling from the reference set, rVLM can be approximately calculated, providing the robot with a state-correlated reward signal, guiding the robot to explore and learn towards a state that is more consistent with the task goal;

[0113] Select multiple complex motion tasks from the simulation humanoid robot benchmark for whole - body motion and manipulation, including: dynamic motion tasks and static posture tasks, to test the complex skills of the Unitree H1 robot; the dynamic motion tasks include: running, walking, climbing stairs, and sliding dynamic tasks; the static posture tasks include: standing, balancing, sitting static tasks; the robot can only obtain proprioceptive data such as joint positions and angular velocities during task execution, which is closer to the real - world application scenario; during the experiment, one trajectory out of every nine is selected and rendered as a video to reduce the computational cost and maintain data diversity; for each video, a segment with a length of 64 is intercepted, and the ViCLIP model is used to calculate the video - text similarity score. The reference model is updated every 100,000 environment steps, and an early - stopping strategy is adopted to prevent overfitting. The hyperparameter w is determined to be the optimal value through grid search from {0.01, 0.1, 0.5}; based on the TQC algorithm, the JAX version of TQC is used to conduct experiments in the SBX framework, and its parameters are finely tuned for complex humanoid robot control tasks; each task is trained for 10 million steps, samples are collected from 8 parallel environments during the training process, and the learning rate gradually decays from the initial 6e - 4 to the final 5e - 5 to maintain training stability. Each time, 256 samples are sampled from the replay buffer for updating the policy, and 16 gradient updates are performed every 16 steps; to enhance the exploration ability of the algorithm, state - dependent exploration (SDE) is enabled; among them, the neural networks are all three - layer neural networks with a width of 256 for each layer, the quantile count is set to 50, and 5 top quantiles are discarded for each network; multiple cutting - edge methods are selected for comparison with the multi - view video learning framework MVR; TQC, as a classic algorithm, only relies on task rewards for policy learning; VLM - RM combines task rewards and the image - text similarity score generated by the CLIP model; FuRL is based on image - text similarity and solves the reward - sparsity problem through online fine - tuning; RoboCLIP uses video - text similarity to provide rewards for online RL robots, but only gives a single reward at the end of the trajectory; in addition, the experimental results of DreamerV3 are also included for comparative analysis; the results show that the multi - view video learning framework MVR outperforms all baseline methods in the average performance of multiple tasks;

[0114] The multi-view video learning framework MVR can also identify high-reward but sub-optimal states, including: As shown in Figure 2, the multi-view video learning framework MVR identifies states with high task rewards but sub-optimal; visualizes the states generated by the TQC agent for the sit_hard task; the left part illustrates the task rewards. The right part shows the combination of task rewards and visually-guided rewards calculated by the trained multi-view video learning framework MVR agent; the multi-view video learning framework MVR identifies several parts with high task reward values in the state space as sub-optimal, mainly corresponding to unstable sitting postures; the multi-view video learning framework MVR can also identify high-reward but sub-optimal states.

[0115] The present invention provides a method for constructing a multi-view video reward mechanism learning system, including:

[0116] S1, using the multi-view video learning framework MVR to evaluate the robot's behavior according to multi-view video learning;

[0117] S2, through the vision-language large model, generating visual feedback according to the task text description to improve the robot's motion skill learning; obtaining accurate reward feedback according to the latest state correlation evaluation, and then adjusting the strategy more effectively;

[0118] S3, through the task reward model, when the robot's behavior approaches the expected goal, analyzing the importance degree of the task reward feedback according to the degree of the robot's behavior approaching the expected goal, dynamically adjusting the relative magnitude between the task reward and the vision-language model reward according to the state correlation, and balancing the visual feedback and the task reward;

[0119] S4, combining the multi-view video and the task reward to provide more accurate visual feedback and learning effects in complex robot motion tasks.

[0120] The principles and effects of the above technical solution are as follows: The present invention provides a method for constructing a multi-view video reward mechanism learning system, including: using a multi-view video learning framework MVR to evaluate the behavior of a robot based on multi-view videos; through a vision-language large model, generating visual feedback according to the task text description to improve the learning of the robot's motion skills; obtaining accurate reward feedback based on the latest state correlation evaluation, and then more effectively adjusting the strategy; through a task reward model, when the robot's behavior approaches the expected goal, analyzing the importance degree of the task reward feedback according to the degree to which the robot's behavior approaches the expected goal, and dynamically adjusting the relative magnitude between the task reward and the vision-language model reward according to the state correlation to balance the visual feedback and the task reward; combining multi-view videos and task rewards to provide more accurate visual feedback and learning effects in complex robot motion tasks; being able to use multi-view videos to evaluate the behavior of a robot through the multi-view video learning framework MVR, and generating visual feedback according to the task text description to improve the learning effect of the robot's motion skills; through a reward function, when the robot's behavior approaches the expected goal, analyzing the importance degree of the task reward feedback according to the degree to which the robot's behavior approaches the expected goal, and solving the problem of balancing visual feedback and task rewards; by combining multi-view videos and task rewards, the multi-view video learning framework MVR can provide more accurate visual feedback in complex robot complex motion tasks and improve the learning effect; based on a reinforcement learning framework of multi-view videos and a vision-language model (VLM), enabling the robot to master complex motion skills on a humanoid robot simulation platform; the reward function dynamically adjusts the relative magnitude between the task reward and the vision-language model reward according to the state correlation to achieve the balance between the two; among them, Top-k sampling samples from the top-k basic unit tokens, including randomly sampling only from the k words with the highest probability, and the randomness allows other basic unit tokens with higher scores or probabilities to have the opportunity to be selected; improving the learning effect of the robot in complex tasks, enabling the robot to learn the required motion skills faster and more accurately on a humanoid robot simulation platform; effectively solving the problems of bias towards set states and dependence on single-view images existing in the reward method based on the vision-language model, comprehensively understanding the robot's motion, and improving the accuracy and effectiveness of the reward.

[0121] In one embodiment, S1 includes:

[0122] S11. During the process of the intelligent agent robot performing tasks in the environment, collecting the robot state sequence at a set frequency and converting the robot state sequence into a multi-view video; at the same time, refining the target task into a concise text description; calculating the similarity between the video and the text through the VLM; constructing a dataset D; the dataset D includes: the robot state sequence, the multi-view video corresponding to the robot state sequence, and the similarity score between the video and the text.

[0123] S12. In the multi-view video learning framework MVR, through the paired comparison method, by using the sports game statistical model, the comparison result of the robot state sequence is kept as consistent as possible with the video comparison result, and the state correlation is learned more accurately.

[0124] S13. Regularization interference reduction subsystem. The multi-view video learning framework MVR performs regularized state representation to reduce the interference caused by the view change in the multi-view video.

[0125] The principles and effects of the above technical solutions are as follows: During the process of the intelligent robot performing tasks in the environment, the robot state sequence is collected at a set frequency and converted into a multi-view video. At the same time, the target task is refined into a concise text description. The similarity between the video and the text is calculated through VLM. A dataset D is constructed. The dataset D includes: the robot state sequence, the multi-view video corresponding to the robot state sequence, and the similarity score between the video and the text. In the multi-view video learning framework MVR, through the paired comparison method, by using the sports game statistical model, the comparison result of the robot state sequence is kept as consistent as possible with the video comparison result, and the state correlation is learned more accurately. The sports game statistical model includes the Bradley-Terry model. Regularization interference reduction subsystem. The multi-view video learning framework MVR performs regularized state representation to reduce the interference caused by the view change in the multi-view video.

[0126] Keeping the comparison result of the robot state sequence as consistent as possible with the video comparison result and learning the state correlation more accurately includes: In the actual task, the matching degree between any two videos and the task description will form a probabilistic comparison relationship, and the multi-view video learning framework MVR can also reflect this relationship at the robot state sequence level. For example, if one video is more in line with the task description than another video, then a similar relationship should also be presented between the corresponding robot state sequences. By continuous adjustment, the similarity between the robot state sequence and the video sequence is aligned as much as possible, so as to avoid the bias towards the set state caused by the image-text similarity calculation and consider the correlation between different states and tasks more comprehensively.

[0127] The multi-view video learning framework MVR regularizes the state representation to reduce the interference caused by perspective changes in multi-view videos, including: the similarity scores of videos from different perspectives may fluctuate, which may affect the accurate judgment of state relevance; to reduce this interference, the multi-view video learning framework MVR keeps the similarity structure of the robot state sequence and the video embedding matched. Specifically, the state is processed through a set function to map the state in the state space to a suitable vector space, and then some operations are used to measure the relevance between the state and the task; during the calculation process, different perspectives are randomly and uniformly sampled to calculate the video embedding, and the state representation is adjusted with this as a reference, so that the state representation can reflect the common information from different perspectives, reduce the interference of perspective setting factors, and thus learn the state relevance more stably and accurately.

[0128] In one embodiment, S2 includes:

[0129] S21, during the entire online reinforcement learning process, the multi-view video learning framework MVR renders the trajectory generated by the robot into a video according to the set rendering frequency TRender; and timely obtains the behavioral visual information of the robot;

[0130] S22, the rendered video is input into the VLM to obtain the video-text similarity score and the video embedding; these new information are used to update the dataset D to keep the latest robot behavior and task matching situation in the dataset;

[0131] S23, the multi-view video learning framework MVR maintains a reference buffer Dref specifically for storing those robot state sequences with relatively high video-text similarity scores; these robot state sequences are considered to be more relevant to the task objective;

[0132] The multi-view video learning framework MVR updates the model for measuring state relevance using the samples in the dataset at the set update frequency TUpdate, continuously optimizing the understanding and judgment of state relevance; an off-policy algorithm is used, and after each model update, the rewards in the robot replay buffer need to be recalculated to maintain the consistency of the rewards, so that the robot can obtain accurate reward feedback based on the latest state relevance evaluation, and then adjust the policy more effectively.

[0133] The principle and effects of the above technical solution are as follows: In the entire online reinforcement learning process, the multi-view video learning framework MVR renders the trajectory generated by the robot into a video according to the set rendering frequency TRender, and timely obtains the visual information of the robot's behavior. The rendered video will be input into the VLM to obtain the video-text similarity score and video embedding. These new pieces of information will be used to update the dataset D to keep the latest robot behavior and task matching situation in the dataset. For the buffer state sequence selection and storage subsystem, the multi-view video learning framework MVR maintains a reference buffer Dref, which is specifically used to store those robot state sequences with relatively high video-text similarity scores. These robot state sequences are considered to be more relevant to the task objectives. The multi-view video learning framework MVR updates the model for measuring state relevance using the samples in the dataset at the set update frequency TUpdate, continuously optimizing the understanding and judgment of state relevance. The off-policy strategy algorithm is used. After each model update, the rewards in the robot replay buffer need to be recalculated to maintain the consistency of the rewards, enabling the robot to obtain accurate reward feedback based on the latest state relevance assessment, and thus adjusting the strategy more effectively.

[0134] MVR design idea: On the one hand, through the method of pairwise comparison and using the sports game statistical model, the comparison results of state sequences are made to be as consistent as possible with the comparison results of videos. In actual tasks, the matching degree between any two videos and the task description will form a probabilistic comparison relationship.

[0135] The probability that a video o better matches the task description l than another video o′ is given by the following formula:

[0136] where is the sigmoid function, represents the cosine similarity of the VLM model.

[0137] MVR expects this relationship to be reflected at the state sequence level as well. For example, if a video better matches the task description than another video, then a similar relationship should also exist between the corresponding state sequences. Similarly, the probability that a state sequence s better matches l than another state sequence s′ is:

[0138]

[0139] where n(s) is the length of s, represents the state relevance function; it should be noted that the state relevance is determined based on the video, thus overcoming the deviation from the set state; given two samples and in the dataset D, minimize the following loss function:

[0140]

[0141] Here, the property 1 - σ(x) = σ(-x) is utilized; through continuous adjustment, the similarity between the state sequence and the video sequence is made as aligned as possible, thereby avoiding the bias towards the set state caused by the calculation of the image - text similarity and considering the relevance between different states and tasks more comprehensively.

[0142] On the other hand, the interference of the perspective setting is independent of the similarity between the underlying states and states because the change of perspective does not affect the states themselves. Therefore, this interference can be eliminated by keeping the similarity structure between the state sequences matched with the similarity structure of the video embeddings, where the perspective is uniformly randomly sampled. Using the state encoder and the vector to parameterize the correlation function, where d is the dimension of the state embedding; and are both normalized; the correlation calculation of state s is as follows:

[0143] where 〈·〉 represents the dot - product operation; for the state sequence s, the following formula is used as its representation:

[0144] where o represents the video stream; since the perspective is uniformly randomly sampled, this regularization will generate an average state representation, thereby reducing the interference of the perspective setting. Finally, the overall objective function for learning state correlation combines the matching term and the regularization term:

[0145] Lrel is the overall objective function for state correlation, Lmatching is the matching term, and Lreg is the regularization term.

[0146] In one embodiment, the visual feedback task reward balance sub - system includes:

[0147] A multi - class reward flexible balance sub - system, where the reward model designed by the multi - view video learning framework MVR realizes the flexible balance between the robot task reward rtask and the visual guidance reward rVLM;

[0148] A visual guidance reward adjustment sub - system, which first enables the robot to obtain the correct action mode through the visual guidance reward rVLM and then makes fine - tuning relying on the quantitative target robot task reward rtask;

[0149] The reward model designed by the multi-view video learning framework MVR for the implementation of the multi-view video learning framework MVR achieves a flexible balance between the robot task reward rtask and the vision-guided reward rVLM, including: the vision-guided reward rVLM encourages the robot to explore task-related states; however, when the robot gradually approaches these states, the influence of rVLM should gradually decrease, and at this time, the robot task reward rtask should play a dominant role.

[0150] The principle and effect of the above technical solution are as follows: The visual feedback task reward balancing subsystem includes: a multi-class reward flexible balancing subsystem, and the reward model designed by the multi-view video learning framework MVR achieves a flexible balance between the robot task reward rtask and the vision-guided reward rVLM; the vision-guided reward adjustment subsystem first enables the robot to obtain the correct action pattern through the vision-guided reward rVLM, and then relies on the quantitative target robot task reward rtask for fine adjustment; the reward model designed by the multi-view video learning framework MVR achieves a flexible balance between the robot task reward rtask and the vision-guided reward rVLM, including: the vision-guided reward rVLM encourages the robot to explore task-related states; however, when the robot gradually approaches these states, the influence of rVLM should gradually decrease, and at this time, the robot task reward rtask should play a dominant role.

[0151] The reward function designed by MVR aims to achieve a flexible balance between the task reward rtask and the VLM-based reward rVLM; specifically, the role of rVLM is to encourage the agent to explore task-related states, but when the agent gradually approaches these states, the influence of rVLM should gradually decrease, and at this time, rtask should play a dominant role; this design concept is similar to the human skill learning process, first enabling the agent to obtain the correct action pattern through rVLM, and then relying on rtask for fine adjustment.

[0152] To achieve this balance, MVR introduces the concept of policy relevance; simply put, it comprehensively considers the relevance of states and the access frequency of these states by the policy to more comprehensively measure the matching degree between the policy and the task; based on policy relevance, when MVR learns a policy, it will strive to maximize a set target function; this target function contains two key parts, one part is the value function of the policy, which reflects the expected return obtained by executing actions according to this policy; the other part is related to policy relevance, and its influence degree is adjusted through a special function; the policy relevance function can be adjusted according to the specific characteristics and requirements of the task to balance the roles of these two parts; as the agent's policy gets closer and closer to the policy that best matches the task description, the influence of this part related to state relevance will gradually weaken, realizing the dynamic transformation of the reward function from being mainly based on rVLM to being mainly based on rtask.

[0153] Among them, refers to the value function of reinforcement learning under the policy , and refers to the policy that best matches the language . To maximize the log-likelihood, make match better than under the BT model , and is introduced to adjust the scales of the two terms. The function log(σ(x)) is monotonically decreasing when x ≤ 0, and the influence of weakens as gets closer and closer to .

[0154] When specifically calculating rVLM, MVR adopts an approximate method. Since it is difficult to directly sample from the ideal policy, MVR uses the previously mentioned reference set Dref for approximate sampling. The state sequences in the reference set are selected according to the video-text similarity scores and are considered to be highly correlated with the target policy. By sampling from the reference set, rVLM can be approximately calculated, providing a reward signal based on state correlation for the agent to guide the agent to explore and learn towards states that are more consistent with the task goals:

[0155] Among them, E[] represents the expectation, represents the state sampled from the policy , represents the correlation function, σ and

[0156] is the sigmoid function.

[0157] In one embodiment, S4 includes:

[0158] S41, the multi-view video learning framework MVR introduces policy correlation, comprehensively considering the correlation of states and the access frequency of the policy to these states; more comprehensively measuring the matching degree of the policy and the task;

[0159] Policy relevance includes: When the multi-view video learning framework MVR learns a policy, it tries to maximize a set objective function; this objective function consists of two key parts. One part is the value function of the policy, which reflects the expected return obtained by executing actions according to this policy; the other part is related to policy relevance and its influence degree is adjusted through a special function; there is a hyperparameter w that can be adjusted according to the specific characteristics and requirements of the task to balance the roles of these two parts; as the robot's policy gets closer and closer to the policy that best matches the task description, this part of the influence related to state relevance will gradually weaken, realizing the dynamic transformation of the reward function from mainly based on rVLM to mainly based on rtask.

[0160] The principle and effect of the above technical solution are: The multi-view video learning framework MVR introduces policy relevance, comprehensively considering the relevance of states and the access frequency of these states by the policy; more comprehensively measuring the matching degree between the policy and the task; the multi-view video learning framework MVR adopts approximate calculation to provide the robot with a reward signal based on state relevance, guiding the robot to explore and learn towards states that are more in line with the task objectives; Policy relevance includes: When the multi-view video learning framework MVR learns a policy, it tries to maximize a set objective function; this objective function consists of two key parts. One part is the value function of the policy, which reflects the expected return obtained by executing actions according to this policy; the other part is related to policy relevance and its influence degree is adjusted through a special function; there is a hyperparameter w that can be adjusted according to the specific characteristics and requirements of the task to balance the roles of these two parts; as the robot's policy gets closer and closer to the policy that best matches the task description, this part of the influence related to state relevance will gradually weaken, realizing the dynamic transformation of the reward function from mainly based on rVLM to mainly based on rtask;

[0161] The multi-view video learning framework MVR adopts approximate calculation to provide the robot with a reward signal based on state relevance, guiding the robot to explore and learn towards states that are more in line with the task objectives, including: The multi-view video learning framework MVR uses an approximate calculation method to calculate rVLM, providing the robot with a reward signal based on state relevance, guiding the robot to explore and learn towards states that are more in line with the task objectives; The multi-view video learning framework MVR uses the previously mentioned reference set Dref for approximate sampling; the robot state sequences in the reference set are selected according to the video-text similarity scores and are considered to be highly related to the target policy; by sampling from the reference set, rVLM can be approximately calculated, providing the robot with a reward signal based on state relevance, guiding the robot to explore and learn towards states that are more in line with the task objectives;

[0162] Select multiple complex motion tasks from the simulation humanoid robot benchmarks for whole-body motion and manipulation, including: dynamic motion tasks and static posture tasks, to test the complex skills of the Unitree H1 robot; the dynamic motion tasks include: running, walking, climbing stairs, and sliding dynamic tasks; the dynamic motion tasks include: standing, balancing, sitting static tasks; the robot can only obtain proprioceptive data during task execution, such as joint positions and angular velocities, which is closer to the real application scenario; during the experiment, one trajectory out of every nine is selected and rendered as a video to reduce the computational cost and maintain data diversity; for each video, a segment with a length of 64 is intercepted, and the ViCLIP model is used to calculate the video-text similarity score. The reference model is updated every 100,000 environment steps, and an early stopping strategy is adopted to prevent overfitting. The hyperparameter w is determined by grid search from {0.01, 0.1, 0.5}; based on the TQC algorithm, the JAX version of TQC is used to conduct experiments in the SBX framework, and its parameters are finely tuned for complex humanoid robot control tasks. Each task is trained for 10 million steps, and samples are collected from 8 parallel environments during the training process. The learning rate gradually decays from the initial 6e - 4 to the final 5e - 5 to maintain training stability. Each time, 256 samples are sampled from the replay buffer for policy update, and 16 gradient updates are performed every 16 steps; to enhance the exploration ability of the algorithm, state-dependent exploration (SDE) is enabled; among them, the neural networks are all three-layer neural networks with a width of 256 for each layer, the quantile count is set to 50, and 5 top quantiles are discarded for each network; multiple cutting-edge methods are selected for comparison with the multi-view video learning framework MVR; TQC, as a classic algorithm, only relies on task rewards for policy learning; VLM - RM combines task rewards and the image-text similarity score generated by the CLIP model; FuRL is based on image-text similarity and solves the reward sparsity problem through online fine-tuning; RoboCLIP uses video-text similarity to provide rewards for online RL robots, but only gives a single reward at the end of the trajectory. In addition, the experimental results of DreamerV3 are also included for comparative analysis; the results show that the multi-view video learning framework MVR outperforms all baseline methods in the average performance of multiple tasks;

[0163] The perspective video learning framework MVR can also identify highly rewarding but sub-optimal states, including: As shown in Figure 2, the multi-view video learning framework MVR identifies states with high task rewards but sub-optimal; visualizes the states generated by the TQC agent for the sit_hard task; the left part illustrates the task rewards. The right part shows the combination of task rewards and visually-guided rewards calculated by the trained multi-view video learning framework MVR agent; the multi-view video learning framework MVR identifies several parts of the state space with high task reward values as sub-optimal, mainly corresponding to unstable sitting postures; the multi-view video learning framework MVR can also identify highly rewarding but sub-optimal states.

[0164] Although the embodiments of the present invention have been disclosed as above, it is not limited to the applications listed in the specification and embodiments. It can be fully applied to various fields suitable for the present invention. For those familiar with the field, additional modifications can be easily made. Therefore, without departing from the general concept defined by the claims and the equivalent scope, the present invention is not limited to the specific details and the illustrated and described examples here.

Claims

1. A multi-view video reward mechanism learning system, characterized by: include: Multi-view video evaluation subsystem, using the multi-view video learning framework MVR, to evaluate robot behavior based on multi-view video learning; The visual feedback reward feedback strategy subsystem generates visual feedback based on the task text description through the visual language large model to improve the robot's motor skill learning; based on the latest state relevance evaluation, accurate reward feedback is obtained to adjust the strategy more effectively; The visual feedback task reward balance subsystem uses the task reward model to analyze the importance of task reward feedback based on the degree to which the robot's behavior approaches the expected goal. It dynamically adjusts the relative size between the task reward and the visual language model reward based on the state correlation to balance the visual feedback and task reward. Multi-view video reward combination subsystem, which combines multi-view video and task rewards to provide more accurate visual feedback and learning effects in complex robot motion tasks; Visual feedback task reward balance subsystem, including: The reward model designed by the multi-view video learning framework MVR realizes the flexible balance between the robot task reward rtask and the visual guidance reward rVLM. The visual guidance reward adjustment subsystem first uses the visual guidance reward rVLM to let the robot obtain the correct action mode, and then relies on the quantitative target robot task reward rtask for fine adjustment; The reward model designed by the multi-view video learning framework MVR achieves a flexible balance between the robot task reward rtask and the vision-guided reward rVLM, including: the vision-guided reward rVLM encourages the robot to explore task-related states; but when the robot gradually approaches these states, the influence of rVLM should gradually decrease, and the robot task reward rtask should play a dominant role.

2. The multi-view video reward mechanism learning system according to claim 1, characterized in that: Multi-view video evaluation subsystem, including: Sequence acquisition dataset subsystem: When the intelligent robot performs tasks in the environment, it collects robot state sequences at a set frequency and converts the robot state sequences into multi-view videos. At the same time, it extracts the target tasks into concise text descriptions. It calculates the similarity between the video and the text through VLM and constructs a dataset D. The dataset D contains: robot state sequences, multi-view videos corresponding to the robot state sequences, and similarity scores between the videos and the text. The matching pairwise comparison subsystem, the multi-view video learning framework MVR uses the sports competition statistical model to keep the robot state sequence comparison results as consistent as possible with the video comparison results, and learn the state correlation more accurately; Regularization interference reduction subsystem,The multi-view video learning framework MVR performs regularized state representation to reduce the interference caused by perspective changes in multi-view videos.

3. The multi-view video reward mechanism learning system according to claim 1, characterized in that: Visual feedback reward feedback strategy subsystem, including: Trajectory rendering behavior visual subsystem, multi-view video learning framework MVR During the entire online reinforcement learning process, the trajectory generated by the robot is rendered into a video according to the set rendering frequency TRender; the robot's behavior visual information is obtained in a timely manner; Similarity score video embedding subsystem, the rendered video will be input into VLM to obtain the video-text similarity score and video embedding; this new information will be used to update the dataset D to keep the dataset containing the latest robot behavior and task matching; Buffer state sequence selection subsystem,The multi-view video learning framework MVR maintains a reference buffer Dref, which specifically stores those robot state sequences with high video-text similarity scores; these robot state sequences are considered to be more relevant to the task objectives; The multi-view video learning framework MVR uses samples in the dataset to update the model used to measure state relevance at a set update frequency TUpdate, continuously optimizing the understanding and judgment of state relevance; an offline strategy algorithm is used, and after each model update, the rewards in the robot's playback buffer need to be recalculated to maintain the consistency of the rewards, so that the robot can obtain accurate reward feedback based on the latest state relevance evaluation, and thus adjust the strategy more effectively.

4. The multi-view video reward mechanism learning system according to claim 1, characterized in that: Multi-view video reward combined with subsystem, including: Policy relevance matching measurement subsystem, the multi-view video learning framework MVR introduces policy relevance, comprehensively considers the relevance of states and the frequency of policy visits to these states; more comprehensively measures the degree of match between policy and task; Reward-guided exploration and learning subsystem: The multi-view video learning framework MVR uses approximate computing to provide the robot with a reward signal based on state relevance, guiding the robot to explore and learn in a state that is more consistent with the task goal; Policy relevance includes: when learning a strategy, the multi-view video learning framework MVR strives to maximize a set objective function; this objective function contains two key parts, one is the value function of the strategy, which reflects the expected return that can be obtained by performing actions according to the strategy; the other part is related to policy relevance, and its influence is adjusted by a special function; there is a hyperparameter w, which can be adjusted according to the specific characteristics and requirements of the task to balance the effects of these two parts; as the robot's strategy gets closer and closer to the strategy that best matches the task description, the influence of this part related to state relevance will gradually weaken, realizing a dynamic transformation of the reward function from rVLM-based to rtask-based.

5. A method for constructing a multi-view video reward mechanism learning system, characterized in that: include: S1, using the multi-view video learning framework MVR, evaluates robot behavior based on multi-view video learning; S2, through the visual language large model, generates visual feedback based on the task text description to improve the robot's motor skill learning; based on the latest state relevance evaluation, accurate reward feedback is obtained to adjust the strategy more effectively; S3, through the task reward model, when the robot behavior is close to the expected goal, according to the degree of closeness of the robot behavior to the expected goal, analyze the importance of the task reward feedback, dynamically adjust the relative size between the task reward and the visual language model reward according to the state correlation, and balance the visual feedback and task reward; S4, combining multi-view videos and task rewards to provide more accurate visual feedback and learning effects in complex robot motion tasks; S3 includes: S31, the reward model designed in the multi-view video learning framework MVR achieves a flexible balance between the robot task reward rtask and the visual guidance reward rVLM; S32, first use the visual guidance reward rVLM to let the robot acquire the correct action mode, and then rely on the quantitative target robot task reward rtask for fine adjustment; The reward model designed by the multi-view video learning framework MVR achieves a flexible balance between the robot task reward rtask and the vision-guided reward rVLM, including: the vision-guided reward rVLM encourages the robot to explore task-related states; but when the robot gradually approaches these states, the influence of rVLM should gradually decrease, and the robot task reward rtask should play a dominant role.

6. The method for constructing a multi-view video reward mechanism learning system according to claim 5, characterized in that: S1 includes: S11, when the intelligent robot performs tasks in the environment, it collects robot state sequences at a set frequency and converts the robot state sequences into multi-view videos; at the same time, it extracts the target tasks into concise text descriptions; calculates the similarity between the video and the text through VLM; and constructs a data set D; the data set D includes: robot state sequences, multi-view videos corresponding to the robot state sequences, and similarity scores between the videos and the texts; S12, the multi-view video learning framework MVR uses a sports game statistical model to keep the robot state sequence comparison results and the video comparison results as consistent as possible through matching pairwise comparisons, and learn state correlations more accurately; S13, regularized interference reduction subsystem, the multi-view video learning framework MVR performs regularized state representation to reduce the interference caused by perspective changes in multi-view videos.

7. The method for constructing a multi-view video reward mechanism learning system according to claim 5, characterized in that S2 include: S21, the multi-view video learning framework MVR, renders the trajectory generated by the robot into a video according to the set rendering frequency TRender during the entire online reinforcement learning process; and obtains the robot's behavioral visual information in a timely manner; S22, the rendered video will be input into VLM to obtain the video-text similarity score and video embedding; this new information will be used to update the dataset D to keep the dataset containing the latest robot behavior and task matching; S23, the multi-view video learning framework MVR maintains a reference buffer Dref, which specifically stores robot state sequences with higher video-text similarity scores; these robot state sequences are considered to be more relevant to the task goal; The multi-view video learning framework MVR uses samples in the dataset to update the model used to measure state relevance at a set update frequency TUpdate, continuously optimizing the understanding and judgment of state relevance; it uses an off policy strategy algorithm, and each time the model is updated, the rewards in the robot's playback buffer need to be recalculated to maintain the consistency of the rewards, so that the robot can obtain accurate reward feedback based on the latest state relevance evaluation, and thus adjust the strategy more effectively.

8. The method for constructing a multi-view video reward mechanism learning system according to claim 5, characterized in that S4 include: S41, the multi-view video learning framework MVR introduces policy relevance, which comprehensively considers the relevance of states and the frequency of policy visits to these states; it more comprehensively measures the degree of match between policy and task; S42, the multi-view video learning framework MVR uses approximate computing to provide the robot with a reward signal based on state relevance, guiding the robot to explore and learn in a state that is more consistent with the task goal; Policy relevance includes: when learning a strategy, the multi-view video learning framework MVR strives to maximize a set objective function; this objective function contains two key parts, one is the value function of the strategy, which reflects the expected return that can be obtained by performing actions according to the strategy; the other part is related to policy relevance, and its influence is adjusted by a special function; there is a hyperparameter w, which can be adjusted according to the specific characteristics and requirements of the task to balance the effects of these two parts; as the robot's strategy gets closer and closer to the strategy that best matches the task description, the influence of this part related to state relevance will gradually weaken, realizing a dynamic transformation of the reward function from rVLM-based to rtask-based.

Citation Information

Patent Citations

  • Complex skill dynamic enhancement method and system with uncertain production line

    CN118862626A

  • Multi-source video data robot skill learning method and system

    CN119204085A