Control strategy determination method and device based on comparative learning, equipment and medium
By constructing a trajectory embedding space through comparative learning, the problem of inefficient annotation caused by ambiguous queries in robotic arm control is solved, the annotation accuracy and performance of robotic arm control are improved, and a more efficient robotic arm control strategy optimization is achieved.
Patent Information
- Application Number
- CN202510872079.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-06-26
- Publication Date
- 2025-10-28
AI Technical Summary
Existing preference-based learning methods for robotic arm control suffer from low labeling efficiency when the robotic arm executes similar trajectories, limiting the reliability and accuracy of robotic arm control.
By using a contrastive learning method, historical trajectory data of the robotic arm is acquired and a trajectory embedding space is constructed. The embedding space is optimized using a contrastive learning loss function, the discriminativeness of trajectory segment pairs is quantified, trajectory pairs with high discriminativeness are selected, and a reward model is trained. Finally, an offline reinforcement learning algorithm is used to optimize the motion control strategy.
It improves the labeling accuracy in robotic arm control, reduces labor costs, learns a more accurate reward model, and enhances the performance of robotic arm control.
Smart Images

Figure CN120848178A_ABST
Abstract
Description
Technical Field
[0001] This application belongs to the field of artificial intelligence technology, and in particular relates to a method, apparatus, device and medium for determining control strategies based on contrastive learning. Background Technology
[0002] Reinforcement learning (RL) is an important machine learning paradigm that learns optimal policies through the interaction of an agent with its environment to maximize cumulative rewards. RL is widely used in robotic arm control (such as grasping, assembly, and tool manipulation), but its core challenge lies in designing accurate reward functions to align with human intentions (e.g., "gently close the box" vs. "quickly close"). Traditional RL relies on manually designing reward functions, which is time-consuming and prone to deviating from real-world requirements.
[0003] To overcome the challenges of reward function design, several methods have emerged that learn directly from human feedback or demonstrations. Preference-based reinforcement learning (PbRL) learns a reward model by observing human preferences for pairs of robotic arm trajectory segments, bypassing explicit reward engineering. This pairwise comparison-based framework is easy to implement and effectively captures human intent. For example, in MetaWorld tasks (such as Box-close and Dial-turn), human annotators can compare the completeness of the robotic arm closing a box or the angular accuracy of rotating a dial.
[0004] However, current PbRL methods struggle to distinguish and provide clear preference signals when robotic arms execute similar trajectories (such as slight angular differences when closing a box), leading to inefficient labeling, limiting the reliability of the robotic arm control system, and consequently resulting in poor control accuracy. Summary of the Invention
[0005] This application provides a control strategy determination method, apparatus, device, and medium based on contrastive learning, which at least solves the problem in the related art that when a robotic arm executes similar trajectories, it is difficult for humans to distinguish them and give clear preference signals, resulting in low label efficiency and limiting the reliability of robotic arm control.
[0006] In a first aspect, embodiments of this application provide a control strategy determination method based on contrastive learning, including: A first trajectory dataset and its corresponding first preference dataset of the robot arm's historical trajectory are obtained. The first preference dataset includes a first preference label corresponding to a plurality of first trajectory segments. The plurality of first trajectory segments are randomly sampled from the first trajectory dataset. Each first trajectory segment pair includes two trajectory segments of arbitrary length. Based on the first preference dataset, the initial trajectory encoder and initial decoder are obtained; The initial trajectory encoder maps multiple first trajectory segments to the trajectory embedding space to obtain corresponding embedding vectors, which are then input to the initial decoder. The pre-constructed contrastive learning loss function is used for iterative training to obtain the trajectory embedding space corresponding to the trained trajectory encoder and decoder. Based on the embedding distance of multiple first trajectory segment pairs in the trajectory embedding space, multiple second trajectory segment pairs are determined from the multiple first trajectory segment pairs, wherein two trajectory segments in the second trajectory segment pairs satisfy a preset discrimination condition; Obtain the second preference labels corresponding to each of the multiple second trajectory segments to obtain the second preference dataset; The reward model is trained using the second preference dataset to obtain the trained reward model; The reward value is labeled on the first trajectory dataset using the trained reward model to obtain the second trajectory dataset; Based on the second trajectory dataset, a motion control strategy is trained using an offline reinforcement learning algorithm. The motion control strategy is migrated and deployed to the robotic arm to control its motion.
[0007] Optionally, determining multiple second trajectory segment pairs from multiple first trajectory segment pairs based on their embedding distances in the trajectory embedding space includes: for each first trajectory segment pair, using a trained trajectory encoder to map the two trajectory segments in the first trajectory segment pair to the trajectory embedding space to obtain corresponding embedding vectors; calculating the embedding distance between the two trajectory segments using the trained trajectory encoder based on their respective embedding vectors; constructing a density function based on the embedding distance; assigning sampling weights to the embedding distance based on the density function; and determining multiple second trajectory segment pairs that satisfy a preset discrimination condition from multiple first trajectory segment pairs based on the embedding distance and the sampling weights.
[0008] Optionally, the contrastive learning loss function includes an ambiguity loss function and a quadrilateral loss function; The optimization objective of the ambiguity loss function is: ; Where p represents a trajectory segment pair The preference label, p = 0 indicates Superior p = 1 means Superior p = express and Their performance is too similar to distinguish them; Represents a preference dataset; This indicates the use of a trajectory encoder. trajectory segment The pre-defined embedding vector obtained by mapping ; This indicates the use of a trajectory encoder. trajectory segment The pre-defined embedding vector obtained by mapping ; Let l represent any segment of the trajectory, and l represent the distance. The optimization objective of the quadrilateral loss function is: ; Among them, for two clearly distinguishable trajectory segments... and , Superior and Superior , and its corresponding embedding vector They form a quadrilateral relationship.
[0009] Optionally, the contrastive learning loss function may further include a reconstruction loss function; The optimization objective of the reconstruction loss function is: ; in, Represents the state-action pairs in the trajectory dataset D. This indicates the decoder.
[0010] Optionally, the contrastive learning loss function may further include a norm constraint function; The optimization objective of the norm constraint function is: .
[0011] Optionally, the contrastive learning loss function is a weighted sum of the ambiguity loss function, the quadrilateral loss function, the reconstruction loss function, and the norm constraint function.
[0012] Secondly, embodiments of this application provide a control strategy determination device based on contrastive learning, the device comprising: The first acquisition module is used to acquire a first trajectory dataset of the robot arm's historical trajectory and its corresponding first preference dataset. The first preference dataset includes a first preference label corresponding to a plurality of first trajectory segment pairs. The plurality of first trajectory segment pairs are randomly sampled from the first trajectory dataset. Each first trajectory segment pair includes two trajectory segments of arbitrary length. An initial module is used to obtain an initial trajectory encoder and an initial decoder based on the first preference dataset; The first training module is used to map multiple first trajectory segment pairs into the trajectory embedding space using the initial trajectory encoder to obtain the corresponding embedding vectors, and then input them into the initial decoder. Iterative training is performed using a pre-constructed contrastive learning loss function to obtain the trajectory embedding space corresponding to the trained trajectory encoder and decoder. The determining module is used to determine multiple second trajectory segment pairs from multiple first trajectory segment pairs based on the embedding distance of each pair in the trajectory embedding space, wherein two trajectory segments in the second trajectory segment pairs satisfy a preset discrimination condition. The second acquisition module is used to acquire the second preference labels corresponding to each of the multiple second trajectory segments to obtain the second preference dataset; The second training module is used to train the reward model using the second preference dataset to obtain the trained reward model. The annotation module is used to annotate the reward values of the first trajectory dataset using the trained reward model to obtain the second trajectory dataset; The third training module is used to train a motion control strategy based on the second trajectory dataset using an offline reinforcement learning algorithm. The migration module is used to migrate and deploy the motion control strategy to the robotic arm to perform motion control on the robotic arm.
[0013] Thirdly, embodiments of this application provide an electronic device, which includes: a processor and a memory storing computer program instructions; when the processor executes the computer program instructions, it implements the steps of the control strategy determination method based on contrastive learning as described in any embodiment of the first aspect.
[0014] Fourthly, embodiments of this application provide a computer-readable storage medium storing computer program instructions, which, when executed by a processor, implement the steps of the control strategy determination method based on contrastive learning as described in any embodiment of the first aspect.
[0015] Fifthly, embodiments of this application provide a computer program product, which is stored in a storage medium and executed by at least one processor to implement the steps of the control strategy determination method based on contrastive learning provided in the first aspect of embodiments of this application.
[0016] The control strategy determination method, apparatus, device, and medium based on contrastive learning in this application embodiment map trajectory segments of a robotic arm performing complex tasks into embedding vectors by training a trajectory encoder. The trajectory embedding space is optimized based on a pre-constructed contrastive loss function of the robotic arm's task characteristics. The discriminative power of trajectory segment pairs is quantified by the distance between them in the embedding space, prioritizing robotic arm trajectory pairs with high discriminative power. This solves the problem of inefficient labeling caused by ambiguous queries in offline PbRL, and reduces labor costs while improving labeling accuracy. Furthermore, a more accurate reward model is learned, ultimately training a superior strategy. The deep integration of contrastive learning and robotic arm task characteristics enhances the control performance of the robotic arm. Attached Figure Description
[0017] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments of the present application. For ordinary technicians in this field, other drawings can be obtained based on these drawings without any creative work.
[0018] Figure 1 This is a flowchart illustrating a control strategy determination method based on contrastive learning provided in an embodiment of this application; Figure 2 This is a schematic diagram of the embedded visualization provided in the embodiments of this application; Figure 3 This is a flowchart illustrating another control strategy determination method based on contrastive learning provided in an embodiment of this application; Figure 4 This is a schematic diagram of a control strategy determination device based on contrastive learning provided in an embodiment of this application; Figure 5 This is a schematic diagram of the structure of an electronic device provided in an embodiment of this application.
[0019] Figure label: A control strategy determination device 400 based on contrastive learning includes a first acquisition module 401, an initialization module 402, a first training module 403, a determination module 404, a second acquisition module 405, a second training module 406, a labeling module 407, a third training module 408, and a transfer learning module 409. Electronic device 500, processor 501, memory 502, communication interface 503, bus 510. Detailed Implementation
[0020] The features and exemplary embodiments of various aspects of the present application will be described in detail below. In order to make the purpose, technical solutions and advantages of the present application clearer, the present application will be further described in detail below in conjunction with the accompanying drawings and specific embodiments. It should be understood that the specific embodiments described herein are only intended to explain the present application, rather than to limit the present application. For those skilled in the art, the present application can be implemented without the need for some of these specific details. The following description of the embodiments is merely to provide a better understanding of the present application by illustrating the examples of the present application.
[0021] It should be noted that, in this document, relational terms such as "first" and "second" are used merely to distinguish one entity or operation from another, and do not necessarily require or imply any such actual relationship or order between these entities or operations. Furthermore, the terms "comprising," "including," or any other variations thereof are intended to cover non-exclusive inclusion, such that a process, method, article, or apparatus that comprises a list of elements includes not only those elements but also other elements not expressly listed, or elements inherent to such a process, method, article, or apparatus. Without further limitations, an element defined by the phrase "comprising..." does not exclude the presence of additional identical elements in the process, method, article, or apparatus that includes said element.
[0022] Reinforcement learning (RL) is an important machine learning paradigm that learns optimal policies through the interaction of an agent with its environment to maximize cumulative rewards. RL is widely used in robotic arm control (such as grasping, assembly, and tool manipulation), but its core challenge lies in designing accurate reward functions to align with human intentions (e.g., "gently close the box" vs. "quickly close"). Traditional RL relies on manually designing reward functions, which is time-consuming and prone to deviating from real-world requirements.
[0023] To overcome the challenges of reward function design, several methods have emerged that learn directly from human feedback or demonstrations. Preference-based reinforcement learning (PbRL) learns a reward model by observing human preferences for pairs of robotic arm trajectory segments, bypassing explicit reward engineering. This pairwise comparison-based framework is easy to implement and effectively captures human intent. For example, in MetaWorld tasks (such as Box-close and Dial-turn), human annotators can compare the completeness of the robotic arm closing a box or the angular accuracy of rotating a dial.
[0024] However, current PbRL methods have the following significant drawbacks in the field of robotic arm control (such as complex tasks like grasping, assembly, and tool manipulation): First, there is the problem of ambiguous queries, i.e., queries are indivisible. When the robotic arm performs similar trajectories (such as slight angle differences when closing a box), humans have difficulty distinguishing them and giving clear preference signals, leading to low labeling efficiency. Second, there is a risk of representation collapse. With limited preference data, relying solely on discriminative information may lead to overfitting of the embedding space, making it impossible to effectively distinguish trajectories with different performance. Third, practical applications are limited. Existing methods (such as OPRL and LiRE) experience performance degradation in scenarios with non-ideal experts (such as labeling errors or fuzzy feedback), limiting the reliability of robotic arm control systems.
[0025] Specifically, in robotic arm control tasks (such as the Box-close task in MetaWorld), when trajectory segments have high similarity (e.g., the angle difference in closing the boxes is <5°), human annotators struggle to distinguish between superior and inferior performance, leading to a large number of queries being marked as "undetermined" (ambiguous queries). Ambiguous queries require repeated annotation or filtering, directly reducing the efficiency of human annotation and increasing annotation costs. This technical challenge remains largely unresolved, especially in offline PbRL scenarios. Furthermore, some existing methods attempting to learn trajectory representations may overfit when the preference dataset is small, or cause the learned representation space to collapse, failing to effectively distinguish trajectory segments with different performance levels, if they fail to effectively utilize preference information or employ appropriate mechanisms. Moreover, traditional methods (such as LIRE) rely on complex feedback mechanisms (e.g., multi-trajectory sorting), which are costly to implement and have poor generalization in robotic arm tasks.
[0026] Traditional offline PbRL methods typically involve two stages: first, a reward model is trained from the preference feedback, and then offline RL is performed using that reward model. The preference model is usually based on the Bradley-Terry model and is used to estimate the probability that one segment is better than another.
[0027] Contrastive learning is a core technique of self-supervised learning. It constructs a structured embedding space by narrowing the distance between similar samples (positive samples) and widening the distance between dissimilar samples (negative samples). In robotic arm control, contrastive learning has been used for trajectory representation learning (such as distinguishing between high-value and low-value trajectories), but it has not yet systematically solved the problem of ambiguous queries.
[0028] Among related technologies, OPRL, OPPO, PT and LIRE are representative methods for solving the problems of PbRL feedback efficiency and trajectory representation learning, but none of them have effectively solved the problem of ambiguous queries in robotic arm control.
[0029] OPRL (Offline Preference-based Reward Learning) is an offline PbRL method that utilizes reward ensembles and selects the query with the greatest divergence. This means that OPRL considers query selection to some extent to improve feedback efficiency. Its core steps include training an ensemble of reward models, selecting preferred queries based on the prediction discrepancies of these models, and then using these queries to update the reward models and train the policy.
[0030] OPPO (Offline Preference-Guided Policy Optimization) is also an offline PbRL method that learns trajectory embeddings, but its policy optimization is performed directly in the embedding space. This means that OPPO focuses on learning trajectory representations and attempts to use these representations to guide policy learning. However, OPPO more tightly integrates policy optimization with embedding learning, and its focus is not on using the embedding space to solve ambiguous query problems to improve annotation efficiency.
[0031] The PT (PreferenceTransformer) method uses a Transformer model for reward modeling, replacing the traditional MLP structure in the reward learning phase to better capture the features and preference information of trajectory segments. PT's main contribution lies in the architectural improvement of the reward model, rather than optimizing specifically for ambiguous queries or the geometric characteristics of the embedding space.
[0032] The LIRE (LIstwiseRewardEstimation) method enhances feedback efficiency by using list-based comparisons, meaning it processes the ordering information of multiple trajectory segments at once, rather than just pairwise comparisons. This approach aims to learn from richer forms of feedback, but its core mechanism does not involve optimizing the embedding space through contrastive learning to handle ambiguous query differences.
[0033] To address the problems in related technologies, this application provides a control strategy determination method, apparatus, device, and medium based on contrastive learning. Combining contrastive learning with robotic arm control, it is an offline preference-based reinforcement learning (OfflinePbRL) method for resolving ambiguous feedback. It aims to improve human annotation efficiency, reward model accuracy, and robotic arm task performance.
[0034] The control strategy determination method based on contrastive learning provided in this application will be described in detail below with reference to the accompanying drawings, through specific embodiments and application scenarios.
[0035] Figure 1 A flowchart illustrating a control strategy determination method based on contrastive learning, according to an embodiment of this application, is shown. Figure 1 As shown, the control policy determination method based on contrastive learning may specifically include the following steps: S101. Obtain the first trajectory dataset of the robot arm's historical trajectory and its corresponding first preference dataset. The first preference dataset includes multiple first trajectory segment pairs and their corresponding first preference labels. The multiple first trajectory segment pairs are randomly sampled from the first trajectory dataset. Each first trajectory segment pair includes two trajectory segments of arbitrary length. S102. Based on the first preference dataset, obtain the initial trajectory encoder and the initial decoder; S103. Using the initial trajectory encoder, multiple first trajectory segment pairs are mapped to the trajectory embedding space to obtain the corresponding embedding vectors, which are then input to the initial decoder. The pre-constructed contrastive learning loss function is used for iterative training to obtain the trajectory embedding space corresponding to the trained trajectory encoder and decoder. S104. Based on the embedding distance of multiple first trajectory segment pairs in the trajectory embedding space, determine multiple second trajectory segment pairs from the multiple first trajectory segment pairs, wherein two trajectory segments in the second trajectory segment pairs satisfy a preset discrimination condition. S105. Obtain the second preference labels corresponding to each of the multiple second trajectory segments to obtain the second preference dataset; S106. Train the reward model using the second preference dataset to obtain the trained reward model; S107. Use the trained reward model to label the first trajectory dataset with reward values to obtain the second trajectory dataset; S108. Based on the second trajectory dataset, a motion control strategy is obtained by training using an offline reinforcement learning algorithm; S109. The motion control strategy is migrated and deployed to the robotic arm to perform motion control on the robotic arm.
[0036] Therefore, by training a trajectory encoder to map trajectory segments of a robotic arm performing complex tasks into embedding vectors, and optimizing the trajectory embedding space based on a pre-built contrastive loss function of the robotic arm's task characteristics, the discriminative power of trajectory segment pairs is quantified by the distance between them in the embedding space. High-discriminative robotic arm trajectory pairs are prioritized, thus solving the problem of inefficient labeling caused by ambiguous queries in offline PbRL, and reducing labor costs and improving labeling accuracy. Furthermore, a more accurate reward model is learned, and ultimately a better-performing strategy is trained. Through the deep integration of contrastive learning and the characteristics of the robotic arm's tasks, the control performance of the robotic arm is improved.
[0037] The specific implementation methods for each of the above steps are described below.
[0038] In some embodiments, in S101, a first trajectory dataset of the robot arm's historical trajectories is acquired; M first trajectory segment pairs are randomly sampled from the first trajectory dataset. Each of the first trajectory segment pairs includes two trajectory segments of arbitrary length, such as the trajectory of the robotic arm closing the box at different angles in the Box-close task; initialize the preference dataset. The first preference dataset is empty; then, the first preference dataset is obtained by obtaining the first preference label p corresponding to each of the M first trajectory segments. .
[0039] It is understandable that the first preference label p represents the human experts' preferences for these first trajectory segments. Provided preference feedback, and Where p = 0 indicates Superior ,For example Closer to completely closing the box; p = 1 indicates Superior p = express and Their performance is too similar to distinguish them.
[0040] It should be noted that the first trajectory dataset of the robotic arm's historical trajectories is an offline trajectory dataset of the robotic arm, without environmental interaction. Thus, using the offline PbRL method, the agent learns its policy from the offline trajectory dataset without labeled rewards, and combines this with feedback on preferences based on fragments provided by humans.
[0041] In some embodiments, in S102, the first preference dataset is used. For the trajectory encoder and reward model Preliminary optimizations were performed. Specifically, the trajectory encoder... Configured to extend the trajectory segment of the robotic arm (e.g., end-effector position, joint angle sequence) are mapped to a fixed-dimensional embedding vector z; reward model It is configured to predict the reward value of a trajectory segment based on preference data, where For encoder parameters, For reward model parameters, For state-action pairs.
[0042] Furthermore, in some embodiments, in S103, a structured embedding space is learned to separate robotic arm trajectory segments according to performance differences. The goal is to train a trajectory encoder. To determine the trajectory of a robotic arm of arbitrary length Mapped to a fixed-dimensional embedding vector .
[0043] refer to Figure 2 This is a schematic diagram illustrating the embedding visualization in this embodiment. It should be understood that the learned embedding space should possess the following characteristics: Figure 2 The following characteristics are shown: Figure 2 As shown in (b) and (c), clearly distinguishable trajectory segments (i.e., those with significant performance differences) are far apart in the embedding space, while ambiguous segments (i.e., those with minor or indistinguishable performance differences) are closer together. Ideally, trajectories with similar performance should cluster together, while trajectories with significant performance differences should have noticeable gaps between them. In this way, by constructing a structured embedding space through contrastive learning (such as high / low performance trajectory clustering in the Box-close task), clearly distinguishable robotic arm trajectory pairs can be separated in the space, quantifying the degree of ambiguity.
[0044] Optionally, a Bi-directional Decision Transformer (BDT) architecture can be used as the trajectory encoder. and decoder BDT is a Transformer-based model that can efficiently process sequence data. The encoder... Employing the BERT (Bidirectional Encoder Representations from Transformers) structure, contextual information can be extracted from robotic arm trajectory sequences; the decoder... A GPT (Generative Pre-trained Transformer) structure is employed to predict robotic arm motion sequences given a state and trajectory embedding. This leverages the bidirectional nature of BDT to better understand the overall trajectory information, facilitating the learning of high-quality trajectory representations.
[0045] As an optional implementation, two contrastive learning loss functions are introduced to train the trajectory encoder. Furthermore, preference information is incorporated into the trajectory representation, including the ambiguity loss function and the quadrilateral loss function.
[0046] Ambiguity loss ( The function directly optimizes the objective of distinguishing distinguishable segments and narrowing ambiguity. For the preference dataset... Each robotic arm trajectory segment in the data and their preference tags ,make and The corresponding embedding vectors are used, and L2 distance is used as the distance metric. Differences in reaction trajectory performance.
[0047] Specifically, the optimization objective of ambiguity loss is: ; The first item contributes to clearly distinguishable fragment pairs ( The second term, being farther apart in the embedding space, leads to ambiguous fragment pairs ( (Closer distance)
[0048] It can be seen that this ambiguity loss function is effective on the preference dataset. In the process, the embedding distance is maximized for robotic arm trajectory pairs that are clearly labeled by humans (such as successfully closing a box vs. not closing it), and the embedding distance is minimized for ambiguous trajectory pairs (such as angle differences < 2°).
[0049] However, relying solely on ambiguity loss can lead to overfitting and representation collapse because it only utilizes information about whether segments are distinguishable, ignoring the specific preference relationships between segments. Therefore, to address the shortcomings of relying solely on ambiguity loss and better capture preference structure, quadrilateral loss (Quadrilateral Loss) is further introduced. This loss function directly models the preference relationship between fragment pairs.
[0050] Specifically, such as Figure 2 As shown in (a), for two clearly distinguishable queries and ,in Superior and Superior , their embedding vectors This forms a quadrilateral relationship. Quadrilateral loss aims to encourage "positive samples" (preferred fragments). The sum of the distances between the samples and the "negative samples" (unfavored fragments) The sum of the distances between positive and negative samples is less than the sum of the distances between positive and negative sample cross-pairings.
[0051] Specifically, this can be achieved through the following optimization goals: .
[0052] It is evident that the quadrilateral loss function utilizes the preference relationship between robotic arm trajectory pairs to construct a geometric quadrilateral constraint, which can alleviate overfitting and representation collapse problems under finite preference data, ensuring that the learned embedding space is stable and meaningful.
[0053] In other words, by utilizing the combination of query pairs, the quadrilateral loss effectively increases the scale of the training data, which helps to alleviate the overfitting problem caused by limited preference data. It also serves as a regularization term for the ambiguity loss, ensuring that the embedding space better captures the complete preference structure and improves the representation quality.
[0054] Furthermore, in some embodiments, an additional loss term, namely reconstruction loss, is introduced to stabilize the training process and is used to train the BDT architecture, ensuring that the decoder can reconstruct the action in the trajectory given the trajectory embedding.
[0055] Specifically, the optimization objective of the reconstruction loss function is: ; in, Represents the state-action pairs in the trajectory dataset D. This indicates the decoder.
[0056] In some optional embodiments, to prevent the norm of the embedding vector from growing indefinitely or collapsing to the origin, a constraint is imposed on the L2 norm of the embedding vector, i.e., the norm constraint is: .
[0057] Therefore, it is ultimately used to train the trajectory encoder. and decoder The total loss function is the weighted sum of the losses mentioned above. That is: ; in It is a hyperparameter used to balance the weights of various loss parameters. Distance metric The L2 distance is used.
[0058] In this way, by jointly optimizing these loss functions, a trajectory embedding space that can both distinguish ambiguous segments and capture preference relationships can be learned, laying the foundation for subsequent ambiguous query filtering.
[0059] Furthermore, in some embodiments, in S104, for each first trajectory segment pair, the trained trajectory encoder maps the two trajectory segments in the first trajectory segment pair to the trajectory embedding space to obtain corresponding embedding vectors; based on the embedding vectors corresponding to the two trajectory segments respectively, the trained trajectory encoder calculates the embedding distance between them; based on the embedding distance, a density function is constructed; based on the density function, a sampling weight is assigned to the embedding distance; based on the embedding distance and the sampling weight, multiple second trajectory segment pairs that satisfy the preset discrimination condition are determined from multiple first trajectory segment pairs.
[0060] In practice, for the queries to be selected (i.e., pairs of trajectory segments) Using the currently learned trajectory encoder Calculate its embedding vector and Distance between .
[0061] In practical implementation, in order to handle continuous embedding distances... Discretize it into Each interval. Based on the current preference dataset. Estimate the distance distribution density of clearly distinguishable fragment pairs Distance distribution density of ambiguous fragment pairs .
[0062] In practice, a new density function is constructed. This function aims to highlight the distance intervals where clearly distinguishable fragment pairs occur more frequently than ambiguous fragment pairs.
[0063] Optional, density function Calculated using the following formula: First, calculate the two intermediate density functions: ; ; Then, averaging the two yields the final density function: .
[0064] It should be noted that, in In the calculation, if If the value is zero or close to zero, smoothing or other numerical stabilization measures may be necessary.
[0065] Finally, the original embedding distance distribution Multiplying the distance distribution of all candidate queries by the density function yields the distribution used for rejection sampling. : .
[0066] It is understandable that this rejection sampling distribution assigns higher sampling weights to the embedding distances of those corresponding to clearly distinguishable fragment pairs.
[0067] In practice, based on the rejection sampling distribution The candidate query is sampled. This means that the embedding distance is related to... Queries with density values proportional to the target density are more likely to be selected. This approach prioritizes queries that are geographically distant in the embedding space and are more likely to be clearly distinguishable (e.g., those with significant differences in the end-effector position). In other words, the query selection mechanism based on rejection sampling prioritizes labeling highly discriminative trajectories (e.g., trajectory pairs with significant differences in end-effector position), reducing labor costs and improving labeling accuracy.
[0068] In addition, for queries that fail to reach the required number through the rejection sampling mechanism, strategies from related technologies can be used for supplementary selection, such as selecting those queries that have the greatest divergence in the reward model integration.
[0069] In this way, by rejecting sampling based on the learned embedding space, we can actively filter out ambiguous queries in robotic arm tasks that are difficult for humans to give clear preferences for, thereby using the limited human feedback budget on the most valuable queries that provide the clearest signals, significantly improving annotation efficiency.
[0070] Therefore, in S105, after selecting M new trajectory segment pairs as queries using the currently learned trajectory embedding space and based on the ambiguous query filtering method, human experts provide preference feedback on the newly selected queries, and this feedback is added to the preference dataset, thus obtaining the second preference dataset; and the updated preference dataset is then used... Further optimize the trajectory encoder .
[0071] Furthermore, in some embodiments, in S106, after collecting human feedback on their preferences for the query, this feedback is used to train a reward model. Reward models are typically based on the Bradley-Terry model, estimating the probability that one segment is better than another by minimizing the cross-entropy loss.
[0072] Optionally, the reward model in this embodiment The structure can be based on a multilayer perceptron (MLP) or a Transformer.
[0073] Optional, used for training the reward model The loss function can be shown below: ; in, It is estimated by the reward model. Superior The probability of is usually expressed as: .
[0074] Furthermore, in some embodiments, in S107, after the reward learning phase ends, a trained reward model is obtained. And use the reward model on the original offline dataset. Each state-action pair Rewards are labeled to generate an offline dataset with reward signals, namely the second trajectory dataset.
[0075] Furthermore, in S108, a standard offline reinforcement learning algorithm, such as Implicit Q-Learning (IQL), is used to learn a policy that maximizes the cumulative prediction reward on the labeled offline dataset, i.e., to train the final policy. Among them, IQL is an algorithm that can perform offline policy evaluation and improvement without explicit estimation of the behavior policy, and is suitable for offline RL settings. Therefore, the motion control policy is transferred and deployed to the robotic arm to perform motion control on the robotic arm, that is, to execute S109.
[0076] In summary, the control strategy determination method based on contrastive learning in this application addresses the ambiguous query problem in robotic arm control by introducing two contrastive losses (ambiguity loss and quadrilateral loss) to learn a structured trajectory embedding space. This embedding space allows clearly distinguishable segments to be geometrically separated, while similar segments remain close. Furthermore, a query selection method based on rejection sampling is employed to prioritize queries that are more explicit and easier to distinguish in the learned embedding space, thereby improving the efficiency of human annotation in robotic arm control. Unlike methods that only perform reward modeling or directly optimize the strategy in the embedding space, this application optimizes the structure of the embedding space to directly serve the identification and filtering of ambiguous queries, thus optimizing the robotic arm strategy.
[0077] Therefore, the embodiments of this application optimize the robotic arm trajectory embedding space through comparative learning, directly solving the problem of inefficient annotation caused by ambiguous queries in offline PbRL, and realizing efficient query selection and strategy optimization based on this space, significantly improving the practical application capability of robotic arm control.
[0078] In addition, refer to Figure 3 This application also provides a flowchart illustrating another control strategy determination method based on contrastive learning. The advantages of this method are mainly reflected in resolving ambiguity in human feedback during robotic arm control and improving strategy performance. Through deep integration of contrastive learning and the characteristics of the robotic arm's tasks, it overcomes the bottlenecks of existing technologies in annotation efficiency, strategy performance, and practical deployment, providing a feasible and efficient solution for high-precision robot control.
[0079] like Figure 3As shown, the entire process of implementing the control policy determination method based on contrastive learning can be divided into two main stages: the reward learning stage and the policy learning stage. The reward learning stage is the key step, especially the learning of the trajectory embedding space and the query selection mechanism based on this space.
[0080] The reward learning phase iteratively learns trajectory embeddings and reward models, and uses the learned embedding space to select high-quality (unambiguous) queries to continuously optimize the preference dataset; the policy learning phase uses the finally learned reward model to label the offline dataset and train a reinforcement learning policy, which aims to plan the robotic arm's movements.
[0081] In other words, this method aims to: effectively identify and process ambiguous queries, thereby improving the annotation efficiency in robotic arm control scenarios; learn a meaningful and coherent trajectory embedding space that can clearly distinguish trajectory segments with large performance differences and alleviate the problems of overfitting and representation collapse; and optimize the offline query selection mechanism based on the learned embedding space to improve the performance of robotic arm strategies.
[0082] In some embodiments, a bidirectional decision transformer (BDT) trajectory encoder is trained to map trajectory segments (such as grasping, rotating, and pushing motion sequences) of a robotic arm performing complex tasks (such as box-close and dial-turn) into fixed-dimensional embedding vectors.
[0083] In some embodiments, a contrastive loss function based on the characteristics of the robotic arm task is designed to optimize the trajectory embedding space; the distance between trajectory fragment pairs in the embedding space is used to quantify their discriminativeness; and a rejection sampling mechanism based on distance distribution is constructed to preferentially select robotic arm trajectory pairs with high discriminativeness, thereby improving the efficiency of human annotation.
[0084] In some embodiments, iterative query selection and model optimization includes repeatedly performing the following steps until a total is collected. Feedback: Query selection utilizes the currently learned trajectory embedding space and selects based on an ambiguous query filtering method. A new trajectory segment is selected as the query; feedback is obtained and the dataset is updated, with human experts providing preference feedback on the newly selected queries and adding this feedback to the preference dataset. ; Optimize the trajectory encoder using the updated preference dataset Further optimize the trajectory encoder ; Optimize the reward model using the updated preference dataset. Further optimize the reward model .
[0085] In some embodiments, the policy learning phase specifically includes the following steps: labeling offline datasets and using the finally learned reward model. For the entire offline dataset State-action pairs Label the reward values; train the reinforcement learning policy using a standard offline reinforcement learning algorithm (e.g., Implicit Q-Learning, IQL) on an offline dataset with reward labels. Maximize accumulated rewards.
[0086] In this way, by iteratively optimizing the trajectory encoder, reward model (based on the Bradley-Terry model), and IQL strategy, the performance of the robotic arm can be gradually improved.
[0087] It is evident that, based on the robotic arm control task, by introducing trajectory embedding based on contrastive learning and ambiguity query filtering mechanism based on embedding space during the reward learning stage, the ambiguity problem of human feedback in offline PbRL is effectively solved, the annotation efficiency and feedback quality are improved, and a more accurate reward model is learned, ultimately training a better-performing strategy and improving the performance of the robotic arm task.
[0088] It should be noted that the above description describes some embodiments of this application. Other embodiments are within the scope of the appended claims. In some cases, the actions or steps recorded in the claims can be performed in a different order than that shown in the above embodiments and still achieve the desired result. Furthermore, the processes depicted in the drawings do not necessarily require a specific or sequential order to achieve the desired result. In some embodiments, multitasking and parallel processing are also possible or may be advantageous.
[0089] Based on the same technical concept, corresponding to any of the above embodiments, this application also provides a control strategy determination device 400 based on contrastive learning.
[0090] like Figure 4 As shown, the control strategy determination device 400 based on contrastive learning may include: The first acquisition module 401 is used to acquire a first trajectory dataset of the robot arm's historical trajectory and its corresponding first preference dataset. The first preference dataset includes a plurality of first trajectory segment pairs and their respective first preference labels. The plurality of first trajectory segment pairs are randomly sampled from the first trajectory dataset, and each first trajectory segment pair includes two trajectory segments of arbitrary length. The initial module 402 is used to obtain an initial trajectory encoder and an initial decoder based on the first preference dataset; The first training module 403 is used to map multiple first trajectory segment pairs into the trajectory embedding space using the initial trajectory encoder to obtain the corresponding embedding vectors, and then input them into the initial decoder. Iterative training is performed using a pre-constructed contrastive learning loss function to obtain the trajectory embedding space corresponding to the trained trajectory encoder and decoder. The determining module 404 is used to determine a plurality of second trajectory segment pairs from the plurality of first trajectory segment pairs based on the embedding distance of the plurality of first trajectory segment pairs in the trajectory embedding space, wherein two trajectory segments in the second trajectory segment pairs satisfy a preset discrimination condition. The second acquisition module 405 is used to acquire the second preference labels corresponding to each of the multiple second trajectory segments to obtain the second preference dataset; The second training module 406 is used to train the reward model using the second preference dataset to obtain the trained reward model. The annotation module 407 is used to annotate the first trajectory dataset with reward values using the trained reward model to obtain the second trajectory dataset; The third training module 408 is used to train a motion control strategy based on the second trajectory dataset using an offline reinforcement learning algorithm. The migration module 409 is used to migrate and deploy the motion control strategy to the robotic arm to perform motion control on the robotic arm.
[0091] In some embodiments, the determining module 404 is specifically configured to, for each first trajectory segment pair, use a trained trajectory encoder to map the two trajectory segments in the first trajectory segment pair to the trajectory embedding space to obtain corresponding embedding vectors; calculate the embedding distance between the two trajectory segments based on their respective embedding vectors using the trained trajectory encoder; construct a density function based on the embedding distance; assign sampling weights to the embedding distance based on the density function; and determine multiple second trajectory segment pairs that satisfy a preset discrimination condition from multiple first trajectory segment pairs based on the embedding distance and the sampling weights.
[0092] Optionally, the contrastive learning loss function is a weighted sum of the ambiguity loss function, the quadrilateral loss function, the reconstruction loss function, and the norm constraint function.
[0093] The optimization objective of the ambiguity loss function is: ; Where p represents a trajectory segment pair The preference label, p = 0 indicates Superior p = 1 means Superior p = express and Their performance is too similar to distinguish them; Represents a preference dataset; This indicates the use of a trajectory encoder. trajectory segment The pre-defined embedding vector obtained by mapping ; This indicates the use of a trajectory encoder. trajectory segment The pre-defined embedding vector obtained by mapping ; Let l represent any segment of the trajectory, and l represent the distance.
[0094] The optimization objective of the quadrilateral loss function is: ; Among them, for two clearly distinguishable query trajectory segments... and , Superior and Superior , and its corresponding embedding vector They form a quadrilateral relationship.
[0095] The optimization objective of the reconstruction loss function is: ; in, Represents the state-action pairs in the trajectory dataset D. This indicates the decoder.
[0096] The optimization objective of the norm constraint function is: .
[0097] In addition, this application also provides another control policy determination device based on contrastive learning, including the following functional modules: an embedding learning module, configured to input the robotic arm trajectory and output a structured embedding space (such as high / low performance trajectory clustering in a Box-close task); a query selection module, configured to filter clear queries based on embedding distance (such as rejecting sampling and prioritizing trajectory pairs with large end-position differences); a reward modeling module, configured to train a reward model based on preference feedback; and a policy optimization module, configured to generate the final policy using the IQL algorithm.
[0098] It should be noted that, for ease of description, the above devices are described in terms of function, divided into various modules. Of course, in implementing this application, the functions of each module can be implemented in one or more software and / or hardware.
[0099] The apparatus of the above embodiments is used to implement the corresponding control strategy determination method based on contrastive learning in any of the foregoing embodiments, and has the beneficial effects of the corresponding method embodiments, which will not be repeated here.
[0100] Based on the same technical concept, corresponding to any of the above embodiments, this application also provides an electronic device.
[0101] Figure 5 A schematic diagram of a more specific electronic device hardware structure provided in this embodiment is shown.
[0102] The electronic device 500 may include a processor 501 and a memory 502 storing computer program instructions.
[0103] Specifically, the processor 501 may include a central processing unit (CPU), an application-specific integrated circuit (ASIC), or one or more integrated circuits that can be configured to implement the embodiments of this application.
[0104] Memory 502 may include mass storage for data or instructions. For example, and not limitingly, memory 502 may include a hard disk drive (HDD), floppy disk drive, flash memory, optical disk, magneto-optical disk, magnetic tape, or Universal Serial Bus (USB) drive, or a combination of two or more of these. Where appropriate, memory 502 may include removable or non-removable (or fixed) media. Where appropriate, memory 502 may be internal or external to the integrated gateway disaster recovery device. In a particular embodiment, memory 502 is non-volatile solid-state memory.
[0105] In certain embodiments, the memory may include read-only memory (ROM), random access memory (RAM), disk storage media devices, optical storage media devices, flash memory devices, and electrical, optical, or other physical / tangible memory storage devices. Thus, generally, memory includes one or more tangible (non-transitory) computer-readable storage media (e.g., memory devices) encoded with software including computer-executable instructions, and when the software is executed (e.g., by one or more processors), it is operable to perform the operations described with reference to the method according to one aspect of this application.
[0106] The processor 501 reads and executes computer program instructions stored in the memory 502 to implement any of the control strategy determination methods based on contrastive learning in the above embodiments.
[0107] In some examples, the electronic device 500 may also include a communication interface 503 and a bus 510. For example, Figure 5 As shown, the processor 501, memory 502, and communication interface 503 are connected through bus 510 and complete communication with each other.
[0108] The communication interface 503 is mainly used to realize communication between various modules, devices, units and / or equipment in the embodiments of this application.
[0109] Bus 510 includes hardware, software, or both, that couples components of an online data traffic metering device together. For example, and not as a limitation, bus 510 may include an Accelerated Graphics Port (AGP) or other graphics bus, an Enhanced Industry Standard Architecture (EISA) bus, a Front Side Bus (FSB), HyperTransport (HT) interconnect, an Industry Standard Architecture (ISA) bus, an Infinite Bandwidth Interconnect, a Low Pin Count (LPC) bus, a memory bus, a Microchannel Architecture (MCA) bus, a Peripheral Component Interconnect (PCI) bus, a PCI-Express (PCI-X) bus, a Serial Advanced Technology Attachment (SATA) bus, a Video Electronics Standards Association Local (VLB) bus, or other suitable buses, or combinations of two or more of these. Where appropriate, bus 510 may include one or more buses. Although specific buses are described and illustrated in embodiments of this application, this application contemplates any suitable bus or interconnect.
[0110] For example, the electronic device 500 can be a mobile phone, tablet computer, laptop computer, handheld computer, in-vehicle electronic device, ultra-mobile personal computer (UMPC), netbook, or personal digital assistant (PDA), etc.
[0111] Based on the same technical concept, corresponding to any of the methods in the above embodiments, this application also provides a non-transitory computer-readable storage medium. This computer-readable storage medium stores computer program instructions; when executed by a processor, these computer program instructions implement any of the control strategy determination methods based on contrastive learning in the above embodiments. Examples of computer-readable storage media include non-transitory computer-readable storage media, such as portable disks, hard disks, random access memory (RAM), read-only memory (ROM), erasable programmable read-only memory (EPROM or flash memory), portable compact disk read-only memory (CD-ROM), optical storage devices, magnetic storage devices, etc.
[0112] Based on the same technical concept, corresponding to any of the above embodiments, this application also provides a computer program product, which includes computer program instructions. In some embodiments, the computer program instructions can be executed by one or more processors of a computer to cause the computer and / or the processors to perform the control policy determination method based on contrastive learning. Corresponding to the execution entity for each step in each embodiment of the control policy determination method based on contrastive learning, the processor executing the corresponding step may belong to the corresponding execution entity.
[0113] It should be clarified that this application is not limited to the specific configurations and processes described above and shown in the figures. For the sake of brevity, detailed descriptions of known methods are omitted here. In the above embodiments, several specific steps are described and shown as examples. However, the method process of this application is not limited to the specific steps described and shown. Those skilled in the art can make various changes, modifications, and additions, or change the order of steps, after understanding the spirit of this application.
[0114] The functional blocks shown in the above-described structural diagram can be implemented as hardware, software, firmware, or a combination thereof. When implemented in hardware, they can be, for example, electronic circuits, application-specific integrated circuits (ASICs), appropriate firmware, plug-ins, function cards, etc. When implemented in software, the elements of this application are programs or code segments used to perform the required tasks. Programs or code segments can be stored on a machine-readable medium or transmitted over a transmission medium or communication link via data signals carried on a carrier wave. "Machine-readable medium" can include any medium capable of storing or transmitting information. Examples of machine-readable media include electronic circuits, semiconductor memory devices, ROM, flash memory, erasable ROM (EROM), floppy disks, CD-ROMs, optical disks, hard disks, fiber optic media, radio frequency (RF) links, etc. Code segments can be downloaded via computer networks such as the Internet, intranets, etc.
[0115] It should also be noted that the exemplary embodiments mentioned in this application describe methods or systems based on a series of steps or apparatus. However, this application is not limited to the order of the above steps; that is, the steps can be performed in the order mentioned in the embodiments, or in a different order, or several steps can be performed simultaneously.
[0116] The aspects of this application have been described above with reference to flowchart illustrations and / or block diagrams of methods, apparatus (systems), and computer program products according to embodiments of this application. It should be understood that each block in the flowchart illustrations and / or block diagrams, and combinations of blocks in the flowchart illustrations and / or block diagrams, can be implemented by computer program instructions. These computer program instructions can be provided to a processor of a general-purpose computer, a special-purpose computer, or other programmable data processing apparatus to produce a machine such that these instructions, executable via the processor of the computer or other programmable data processing apparatus, enable the implementation of the functions / actions specified in one or more blocks of the flowchart illustrations and / or block diagrams. Such a processor can be, but is not limited to, a general-purpose processor, a special-purpose processor, a special application processor, or a field-programmable logic circuit. It is also understood that each block in the block diagrams and / or flowcharts, and combinations of blocks in the block diagrams and / or flowcharts, can also be implemented by dedicated hardware performing the specified functions or actions, or can be implemented by a combination of dedicated hardware and computer instructions.
[0117] The above description is merely a specific implementation of this application. Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the specific working processes of the systems, modules, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here. It should be understood that the protection scope of this application is not limited thereto. Any person skilled in the art can easily conceive of various equivalent modifications or substitutions within the technical scope disclosed in this application, and these modifications or substitutions should all be covered within the protection scope of this application.
Claims
1. A method for determining control strategies based on contrastive learning, characterized in that, include: A first trajectory dataset and its corresponding first preference dataset of the robot arm's historical trajectory are obtained. The first preference dataset includes a first preference label corresponding to a plurality of first trajectory segments. The plurality of first trajectory segments are randomly sampled from the first trajectory dataset. Each first trajectory segment pair includes two trajectory segments of arbitrary length. Based on the first preference dataset, the initial trajectory encoder and initial decoder are obtained; The initial trajectory encoder maps multiple first trajectory segments to the trajectory embedding space to obtain corresponding embedding vectors, which are then input to the initial decoder. The pre-constructed contrastive learning loss function is used for iterative training to obtain the trajectory embedding space corresponding to the trained trajectory encoder and decoder. Based on the embedding distance of multiple first trajectory segment pairs in the trajectory embedding space, multiple second trajectory segment pairs are determined from the multiple first trajectory segment pairs, wherein two trajectory segments in the second trajectory segment pairs satisfy a preset discrimination condition; Obtain the second preference labels corresponding to each of the multiple second trajectory segments to obtain the second preference dataset; The reward model is trained using the second preference dataset to obtain the trained reward model; The reward value is labeled on the first trajectory dataset using the trained reward model to obtain the second trajectory dataset; Based on the second trajectory dataset, a motion control strategy is trained using an offline reinforcement learning algorithm. The motion control strategy is migrated and deployed to the robotic arm to control its motion.
2. The method according to claim 1, characterized in that, The step of determining multiple second trajectory segment pairs from multiple first trajectory segment pairs based on their embedding distances in the trajectory embedding space includes: For each first trajectory segment pair, the trained trajectory encoder is used to map the two trajectory segments in the first trajectory segment pair to the trajectory embedding space to obtain the corresponding embedding vector; Based on the embedding vectors corresponding to the two trajectory segments, the embedding distance between them is calculated using the trained trajectory encoder. Construct a density function based on the embedding distance; According to the density function, the embedding distance is assigned a sampling weight; Based on the embedding distance and sampling weight, multiple second trajectory segment pairs that satisfy the preset discrimination condition are determined from multiple first trajectory segment pairs.
3. The method according to claim 1, characterized in that, The contrastive learning loss function includes the ambiguity loss function and the quadrilateral loss function; The optimization objective of the ambiguity loss function is: ; Where p represents a trajectory segment pair The preference label, p = 0 indicates Superior p = 1 means Superior p = express and Their performance is too similar to distinguish them; Represents a preference dataset; This indicates the use of a trajectory encoder. trajectory segment The pre-defined embedding vector obtained by mapping ; This indicates the use of a trajectory encoder. trajectory segment The pre-defined embedding vector obtained by mapping ; Let l represent any segment of the trajectory, and l represent the distance. The optimization objective of the quadrilateral loss function is: ; Among them, for two clearly distinguishable trajectory segments... and , Superior and Superior , and its corresponding embedding vector They form a quadrilateral relationship.
4. The method according to claim 3, characterized in that, The contrastive learning loss function also includes a reconstruction loss function; The optimization objective of the reconstruction loss function is: ; in, Represents the state-action pairs in the trajectory dataset D. This indicates the decoder.
5. The method according to claim 4, characterized in that, The contrastive learning loss function also includes a norm constraint function; The optimization objective of the norm constraint function is: 。 6. The method according to claim 5, characterized in that, The contrastive learning loss function is a weighted sum of the ambiguity loss function, the quadrilateral loss function, the reconstruction loss function, and the norm constraint function.
7. A control strategy determination device based on contrastive learning, characterized in that, include: The first acquisition module is used to acquire a first trajectory dataset of the robot arm's historical trajectory and its corresponding first preference dataset. The first preference dataset includes a first preference label corresponding to a plurality of first trajectory segment pairs. The plurality of first trajectory segment pairs are randomly sampled from the first trajectory dataset. Each first trajectory segment pair includes two trajectory segments of arbitrary length. An initial module is used to obtain an initial trajectory encoder and an initial decoder based on the first preference dataset; The first training module is used to map multiple first trajectory segment pairs into the trajectory embedding space using the initial trajectory encoder to obtain the corresponding embedding vectors, and then input them into the initial decoder. Iterative training is performed using a pre-constructed contrastive learning loss function to obtain the trajectory embedding space corresponding to the trained trajectory encoder and decoder. The determining module is used to determine multiple second trajectory segment pairs from multiple first trajectory segment pairs based on the embedding distance of each pair in the trajectory embedding space, wherein two trajectory segments in the second trajectory segment pairs satisfy a preset discrimination condition. The second acquisition module is used to acquire the second preference labels corresponding to each of the multiple second trajectory segments to obtain the second preference dataset; The second training module is used to train the reward model using the second preference dataset to obtain the trained reward model. The annotation module is used to annotate the reward values of the first trajectory dataset using the trained reward model to obtain the second trajectory dataset; The third training module is used to train a motion control strategy based on the second trajectory dataset using an offline reinforcement learning algorithm. The migration module is used to migrate and deploy the motion control strategy to the robotic arm to perform motion control on the robotic arm.
8. An electronic device, characterized in that, The device includes: a processor and a memory storing computer program instructions; when the processor invokes the computer program instructions, it implements the control strategy determination method based on contrastive learning as described in any one of claims 1-6.
9. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions, which, when invoked by a processor, implement the control strategy determination method based on contrastive learning as described in any one of claims 1-6.
10. A computer program product, characterized in that, When the instructions in the computer program product are executed by the processor of the electronic device, the electronic device performs the control strategy determination method based on contrastive learning as described in any one of claims 1-6.
Citation Information
Cited By
Mechanical arm control method and system based on neutral logic parallel neural network
CN121535764A