Physical man-machine cooperation intention estimation method and related device
Through the accessibility and obstacle trajectory prediction model combined with differential cooperative game theory, the problem of intention estimation in the existing technology relies on short-term data, realize accurate intention prediction and reasonable role allocation, and improve the performance of physical human-computer collaboration.
Patent Information
- Application Number
- CN202510501366.5
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-04-21
- Publication Date
- 2025-08-05
AI Technical Summary
In the prior art, intention estimation mainly relies on short-term motion data, affecting the accuracy and safety of predictions. The effectiveness of long-term collaboration human intention estimation is reduced, and the complexity of human-machine role allocation has not been effectively solved.
Accessible and obstacle-free trajectory prediction models are used to process human motion trajectory data separately, and role allocation is performed in combination with differential cooperative game theory, and accurate intention prediction and role allocation is performed through the transformer's conditional variational autoencoder and future-past attention mechanism.
It improves the prediction accuracy and security of physical human-machine collaboration, realizes reasonable human-machine role allocation, and enhances the autonomy and operability of the robot.
Smart Images

Figure CN120422221A_ABST
Abstract
Description
Technical Field
[0001] The present application relates to an estimation method, and specifically to an intention estimation method and related device for physical human-machine collaboration. Background Art
[0002] In physical human-robot collaboration, effective strategies are needed to ensure seamless collaboration between robots and humans, accurately estimate intentions, dynamically adjust behaviors, and assist humans with minimal effort. Therefore, accurate human intention estimation and reasonable human-robot role allocation are key challenges to improving collaborative performance.
[0003] In complex environments with potential dangers, such as unknown obstacles for robots, the rapid changes in human intentions pose a major challenge to intention estimation. Existing technologies mainly rely on short-term motion data, such as position and velocity, which limits the ability to detect changes in human intentions and affects the accuracy and safety of predictions. At the same time, short-term data reduces the effectiveness of long-term collaborative human intention estimation. In addition, human-machine role allocation involves a complex task control allocation mechanism between humans and robots. This process coordinates the human-machine relationship in real time, reduces disagreements, and improves the level of robot assistance. However, ensuring that the robot's behavior is consistent with human intentions is a complex process. Figure 1 While maintaining autonomy and operability remains a major challenge. Summary of the Invention
[0004] This application addresses the technical problems in the prior art where intention estimation mainly relies on short-term motion data, which affects the accuracy and safety of the prediction and reduces the effectiveness of long-term collaborative human intention estimation. It provides a physical human-machine collaborative intention estimation method and related devices.
[0005] In order to achieve the above objectives, this application adopts the following technical solutions: In a first aspect, the present application proposes a method for estimating intention in physical human-machine collaboration, comprising: Obtain obstacle-free past trajectories and obstacle-containing past trajectories respectively; Inputting the obstacle-free past trajectory and the obstacle-containing past trajectory into the obstacle-free trajectory prediction model and the obstacle-containing trajectory prediction model respectively to obtain the obstacle-free predicted trajectory and the obstacle-containing predicted trajectory; The obstacle-free trajectory prediction model and the obstacle-prone trajectory prediction model have the same structure, and both processing methods include: The past trajectory and the position encoding results of the past trajectory are embedded and connected and then input into the encoder, and then pass through the multi-layer perceptron and sampling in sequence to obtain the sampling result; After performing temporal convolution on the past trajectory, it is combined with the sampling result and input into the decoder; The output of the decoder and the output of the encoder are subjected to the future-past attention mechanism to obtain updated features; The updated features are sequentially passed through the multi-layer perceptron, Gaussian mixture model and sampling to obtain the predicted trajectory; Roles are assigned based on the obstacle-free predicted trajectory and the obstacle predicted trajectory.
[0006] Furthermore, the method for performing role allocation includes: Differential cooperative game theory is used to allocate human and machine roles.
[0007] Furthermore, the method of allocating human-machine roles using differential cooperative game theory includes: Set the reference model of the robot; Introducing state variables, the reference model of the robot is rewritten as a linear state space system; Considering humans and robots as two actors, define the cost functions for humans and robots; Based on differential cooperative game theory, define the shared cost function in cooperation; Substituting the cost functions of humans and robots into the shared cost function in cooperation, we obtain the equation to be solved; Obtain the solution that minimizes the equation to be solved and obtain the control signal for input to the robot.
[0008] Furthermore, the expression of the reference model of the robot includes:
[0009] in, is the inertia matrix, is the Coriolis / centrifugal force matrix, is gravity, is the trajectory tracking error, is the acceleration error, is the speed error, The force exerted on humans, is the force of the virtual robot; The expression of the linear state space system includes:
[0010] in, , , ; is a linear state-space system, is the introduced state variable.
[0011] Furthermore, the cost functions of humans and robots include:
[0012] in, is the cost function for humans, For human reference trajectory, is the tracking error weight of the human cost function compared to the human intended trajectory, is the robot reference trajectory, is the tracking error weight of the human cost function compared to the robot trajectory, The control input weights for the forces applied by humans, is the cost function of the robot, is the tracking error weight of the robot cost function compared to the robot trajectory, is the tracking error weight of the robot cost function compared to the human’s intended trajectory, The control input weights for the robot's control forces; The shared cost function in the cooperation includes:
[0013] in, is the shared cost function in cooperation, To control the role allocation coefficient of each cost weight, is a weighted combination of the human and robot reference trajectories, , Construct a diagonal matrix; The equation to be solved includes:
[0014] in, is the human intention trajectory weight, is the robot trajectory weight.
[0015] Furthermore, the processing method in the encoder includes: Given the end-effector data at T time steps, extract features from the past trajectory and the concatenation of the position encoding results of the past trajectory, and perform linear projection to obtain the key, query and value; The key, query, and value are fed into a self-attention mechanism to learn a latent distribution.
[0016] Furthermore, the future-past attention mechanism includes: Combining the keys and values of features from past data, and queries from future data, follows the standard feedforward and norm layer implementation.
[0017] In a second aspect, the present application proposes a physical human-machine collaborative intention estimation system, comprising: The data module is used to obtain the obstacle-free past trajectory and the obstacle-containing past trajectory respectively; A prediction module is used to input the obstacle-free past trajectory and the obstacle-involved past trajectory into an obstacle-free trajectory prediction model and an obstacle-involved trajectory prediction model, respectively, to obtain an obstacle-free predicted trajectory and an obstacle-involved predicted trajectory; The obstacle-free trajectory prediction model and the obstacle-prone trajectory prediction model have the same structure, and both processing methods include: The past trajectory and the position encoding results of the past trajectory are embedded and connected and then input into the encoder, and then pass through the multi-layer perceptron and sampling in sequence to obtain the sampling result; After performing temporal convolution on the past trajectory, it is combined with the sampling result and input into the decoder; The output of the decoder and the output of the encoder are subjected to the future-past attention mechanism to obtain updated features; The updated features are sequentially passed through the multi-layer perceptron, Gaussian mixture model and sampling to obtain the predicted trajectory; The role allocation module is used to allocate roles according to the obstacle-free predicted trajectory and the obstacle predicted trajectory.
[0018] In the third aspect, the present application proposes an electronic device comprising: a memory, one or more processors; the memory is coupled to the processor; wherein computer program code is stored in the memory, and the computer program code includes computer instructions, and when the computer instructions are executed by the processor, the electronic device performs the steps of the above-mentioned physical human-computer collaboration intention estimation method.
[0019] In a fourth aspect, the present application proposes a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned method for estimating the intention of physical human-computer collaboration are implemented.
[0020] Compared with the prior art, this application has the following beneficial effects: This application proposes a method for estimating intentions in physical human-machine collaboration, which obtains obstacle-free past trajectories and obstacle-filled past trajectories respectively, and then inputs the obstacle-free past trajectories and obstacle-filled past trajectories into obstacle-free trajectory prediction models and obstacle-filled trajectory prediction models respectively to obtain obstacle-free predicted trajectories and obstacle-filled predicted trajectories, and then performs role assignment based on the obstacle-free predicted trajectories and obstacle-filled predicted trajectories. Human intention estimation uses two transformer-based conditional variational autoencoders to combine the robot's motion data in the absence of obstacles with the human-guided trajectory and obstacle avoidance force. In addition, role assignment is performed based on the obstacle-free predicted trajectory and obstacle-filled predicted trajectory to ensure that the robot's behavior is consistent with human intentions. Figure 1 This application incorporates human dynamics into long-term predictions, provides accurate intention understanding, enables reasonable role allocation, and realizes robot autonomy and operability.
[0021] The present application also proposes a physical human-machine collaborative intention estimation system, an electronic device, and a computer-readable storage medium, which have all the advantages of the aforementioned physical human-machine collaborative intention estimation method. BRIEF DESCRIPTION OF THE DRAWINGS
[0022] In order to more clearly illustrate the technical solutions of the embodiments of the present application, the following is a brief introduction to the drawings required for use in the embodiments. It should be understood that the following drawings only illustrate certain embodiments of the present application and therefore should not be regarded as limiting the scope. For ordinary technicians in this field, other relevant drawings can be obtained based on these drawings without paying any creative work.
[0023] Figure 1 A flowchart of the intention estimation method for physical human-machine collaboration in this application; Figure 2 This is a schematic diagram of an implementation of the intention estimation method for physical human-machine collaboration of this application; Figure 3 Schematic diagram of the overall structure of the obstacle-free trajectory prediction model and the obstacle trajectory prediction model in the embodiment of the present application; Figure 4 This is a structural diagram of the training of the obstacle-free trajectory prediction model and the obstacle trajectory prediction model in the embodiment of the present application; Figure 5 This is a schematic diagram of the structure of the obstacle-free trajectory prediction model and the obstacle trajectory prediction model during testing / actual prediction in the embodiment of the present application; Figure 6 A schematic diagram of the principle of role allocation in an embodiment of the present application; Figure 7 This is a schematic diagram of the intention estimation system for physical human-machine collaboration in this application. DETAILED DESCRIPTION
[0024] To make the objectives, technical solutions, and advantages of the embodiments of the present application more clear, the technical solutions in the embodiments of the present application will be clearly and completely described below in conjunction with the accompanying drawings of the embodiments of the present application. Obviously, the described embodiments are only part of the embodiments of the present application, not all of the embodiments. Generally, the components of the embodiments of the present application described and shown in the drawings herein can be arranged and designed in various different configurations.
[0025] Therefore, the following detailed description of the embodiments of the present application provided in the accompanying drawings is not intended to limit the scope of the present application for protection, but merely represents selected embodiments of the present application. All other embodiments obtained by persons of ordinary skill in the art based on the embodiments in the present application without making any creative efforts shall fall within the scope of protection of the present application.
[0026] It should be noted that similar reference numerals and letters denote similar items in the following drawings, and therefore, once an item is defined in one drawing, it does not need to be further defined or explained in subsequent drawings.
[0027] In the description of the embodiments of the present application, it should be noted that if the terms "upper", "lower", "horizontal", "inner", etc. appear, the orientation or position relationship indicated is based on the orientation or position relationship shown in the accompanying drawings, or the orientation or position relationship in which the product of the invention is usually placed when in use. This is only for the convenience of describing the present application and simplifying the description, and does not indicate or imply that the device or element referred to must have a specific orientation, be constructed and operate in a specific orientation. Therefore, it should not be understood as a limitation on the present application. In addition, the terms "first", "second", etc. are only used to distinguish the description and should not be understood as indicating or implying relative importance.
[0028] In addition, if the term "horizontal" appears, it does not mean that the component must be absolutely horizontal, but can be slightly tilted. For example, "horizontal" only means that its direction is more horizontal than "vertical", and does not mean that the structure must be completely horizontal, but can be slightly tilted.
[0029] In the description of the embodiments of this application, it should also be noted that, unless otherwise expressly specified or limited, the terms "disposed," "installed," "connected," and "connected" should be understood in a broad sense. For example, they can refer to fixed connections, detachable connections, or integral connections; they can refer to mechanical connections or electrical connections; they can refer to direct connections or indirect connections through an intermediate medium; and they can refer to internal connections between two components. Those skilled in the art can understand the specific meanings of the above terms in this application based on specific circumstances.
[0030] In physical human-robot collaboration scenarios, achieving seamless collaboration between robots and humans requires effective strategies. This requires robots to accurately estimate human intentions and dynamically and flexibly adjust their behavior based on these estimates, providing effective assistance to humans with minimal resource investment. In this process, accurate estimation of human intentions and appropriate allocation of human and machine roles become critical and challenging tasks for improving collaborative effectiveness.
[0031] In complex, potentially dangerous environments, numerous factors pose significant challenges to intent estimation. For example, robots may encounter unknown obstacles, while human intentions can also change rapidly. Currently, mainstream technologies rely on short-term motion data, such as position and velocity, to estimate intent. However, this approach has significant limitations, weakening the system's ability to detect changes in human intent, which in turn negatively impacts prediction accuracy and safety. Furthermore, relying solely on short-term data reduces the effectiveness of intent estimation in long-term collaborative scenarios.
[0032] Furthermore, human-robot role allocation involves a complex mechanism for allocating task control between humans and robots. This mechanism requires real-time coordination between humans and robots, minimizing human-robot disagreements and enhancing the robot's ability to assist humans. However, ensuring that robots' behavior aligns with human intent while maintaining their autonomy and operability remains a significant challenge.
[0033] Based on the above situation, the present application proposes a method for estimating intentions of physical human-machine collaboration and related devices, which will be described in detail below with reference to embodiments and accompanying drawings.
[0034] like Figure 1 FIG. 1 is a flow chart of a method for estimating intentions in physical human-machine collaboration of the present application, which may include: S101, obtaining the obstacle-free past trajectory and the obstacle-containing past trajectory respectively.
[0035] In physical human-robot collaboration scenarios, accurate estimation of human intention requires collecting relevant data. Past trajectories refer to the motion trajectories generated by the collaboration between humans and robots over a period of time. Unobstructed past trajectories use sensors (such as motion capture devices and inertial measurement units) to continuously monitor and record the movements of humans or robots without encountering significant obstacles. For example, in an empty factory workshop, a worker and a collaborative robot collaborate normally to complete an assembly task. The recorded motion trajectories of the worker and robot are known as unobstructed past trajectories. This trajectory data typically exists in the form of a time series, containing information such as the position and velocity of the moving objects at each point in time. Obstacled past trajectories also utilize these sensors to collect data in the presence of obstacles during collaboration. For example, if goods are suddenly temporarily stacked in the workshop, the resulting motion trajectories of the worker and robot as they navigate around them to continue their task contain information such as the position, velocity, and force of the moving objects at each point in time. Collecting obstacle-prone past trajectories allows the system to learn the changing characteristics of motion patterns when faced with different obstacles, which is crucial for accurately estimating human intention in complex environments.
[0036] S102 , inputting the obstacle-free past trajectory and the obstacle-containing past trajectory into an obstacle-free trajectory prediction model and an obstacle-containing trajectory prediction model respectively to obtain an obstacle-free predicted trajectory and an obstacle-containing predicted trajectory.
[0037] Although the two models are designed for different scenarios, they have the same structure and consistent processing methods. The following details their internal workflow to explain how the predicted trajectory is obtained.
[0038] The obstacle-free trajectory prediction model and the obstacle-prone trajectory prediction model have the same structure, and both processing methods include: (1) The past trajectory and the position encoding results of the past trajectory are embedded and connected and then input into the encoder, and then pass through the multi-layer perceptron and sampling in sequence to obtain the sampling result.
[0039] Since past trajectory data is a chronological sequence, positional encoding is required to enable the model to capture the positional information (i.e., temporal order) in the data. For example, the trajectory position information corresponding to each time point is converted into an encoded vector using a specific mathematical function. This vector not only contains the positional information but also allows the model to distinguish elements at different positions when processing the sequence. The positionally encoded result is embedded and concatenated with the original past trajectory data, effectively adding an additional dimension of positional information to the original trajectory data. This allows the model to better understand the positional relationship of the data in the time series during subsequent processing. The concatenated result is then input into the encoder. The encoder extracts and encodes features from the input trajectory sequence, converting the original trajectory data into a more abstract feature representation that is easier to process. The features output by the encoder are fed into a multilayer perceptron (MLP). The MLP is a feedforward neural network composed of multiple fully connected layers. It performs further nonlinear transformations on the features output by the encoder. By adjusting the weight parameters between different layers, the features are filtered and combined to highlight features useful for subsequent predictions and suppress noise and irrelevant information. The results processed by the multilayer perceptron are then sampled. The purpose of sampling is to randomly extract samples from the feature distribution output by the multilayer perceptron to increase the robustness and generalization ability of the model. For example, methods such as Monte Carlo sampling can be used to randomly select a number of sample points from the feature distribution. These sample points will serve as the basis for subsequent processing to capture the uncertainty and diversity in the feature distribution.
[0040] (2) After performing temporal convolution on the past trajectory, it is combined with the sampling results and input into the decoder.
[0041] Temporal convolutional neural networks can effectively capture local and long-term dependencies in time series data. By designing convolution kernels of varying sizes and numbers of convolution layers, convolution operations are performed on past trajectories in the temporal dimension to extract local and time series features from the trajectory data. The results of the time series convolution are combined with the previously obtained sampling results. This combination fully utilizes the information obtained from the two different processing methods, including both a global abstract representation of trajectory features (the sampling results) and the extraction of local and time series features of the trajectory (the time series convolution results). This combined information is input into the decoder, which also consists of multiple neural network layers. Its task is to convert the input feature information into trajectory data that can be used for prediction.
[0042] (3) The output of the decoder and the output of the encoder are subjected to the future-past attention mechanism to obtain updated features.
[0043] Attention mechanisms are widely used in deep learning to enable models to focus on the importance of different parts of data when processing them. The future-past attention mechanism is used to establish a correlation between the decoder output and the encoder output. By calculating the similarity between the decoder output features and the encoder output features, it determines which parts of the encoder output are more important for predicting the decoder's current output, thereby giving these parts higher weights. For example, this mechanism allows the model to use the currently predicted future trajectory to find the most relevant features from the encoded information of past trajectories for reference, thereby revising and improving the current prediction to obtain an updated feature representation. This updated feature representation combines past trajectory information with the current prediction, making it more accurate for predicting future trajectories.
[0044] (4) The updated features are sequentially passed through a multi-layer perceptron, a Gaussian mixture model, and sampling to obtain the predicted trajectory.
[0045] The updated features are again processed by the multilayer perceptron (MLP) for further adjustment and transformation to better suit subsequent processing requirements. This MLP performs final feature optimization on the updated features, ensuring they more closely match the characteristic distribution of the predicted trajectory. The features processed by the MLP are then input into a Gaussian mixture model (MM). The MM is a probabilistic model that assumes data is a mixture of multiple Gaussian distributions. In trajectory prediction, the MM is used to model features and estimate the parameters (mean, covariance, etc.) of different Gaussian distributions. These parameters enable the generation of multiple possible trajectory predictions, as each Gaussian distribution represents a possible motion pattern. Finally, sampling is performed from the multiple possible trajectory predictions generated by the MM to produce the final predicted trajectory. This sampling process allows the model to select one of the multiple possible motion patterns as the output for the future trajectory prediction. This predicted trajectory is the model's estimate of the likely future trajectory based on the input past trajectory and the learned motion patterns.
[0046] S103: Perform role allocation based on the obstacle-free predicted trajectory and the obstacle predicted trajectory.
[0047] Once the obstacle-free predicted trajectory and the obstacle-predicted trajectory are obtained, human-machine role assignment can be performed based on these prediction results. Regarding the role assignment strategy, one possible strategy is to match the characteristics of the predicted trajectory with the requirements of the current collaborative task. For example, if the obstacle-free predicted trajectory shows that humans can complete a subtask more efficiently (for example, when there are no obstacles, humans can quickly reach the target location and perform fine operations), and the obstacle-predicted trajectory shows that the robot has better motion planning capabilities when dealing with obstacles (for example, the robot can more accurately calculate the path to bypass obstacles), then in actual collaboration, tasks that require quickly reaching the target location can be assigned to humans, and path planning and movement tasks related to handling obstacles can be assigned to robots.
[0048] The application explores scenarios where humans and robots collaborate in potentially dangerous environments, such as obstacles that the robot cannot avoid autonomously. In these situations, the human's intentions change to avoid the hazard, prompting the robot to adjust its behavior accordingly. These scenarios can be categorized into two distinct cases: 1. Obstacle-free: In this case, the force applied by the human is close to 0, and the robot automatically takes the lead. The robot acts as a leader, avoiding measurement errors in the force applied by the human and environmental interference.
[0049] 2. Obstacle avoidance: The human guides the robot to adjust its trajectory to avoid hazards. The human takes the lead, and the robot provides assistance. By incorporating the forces applied by the human into its predictions, the robot can better understand the human's intentions and avoid hazards in the environment.
[0050] This application can effectively utilize past trajectory data to predict future motion trajectories, and make reasonable human-machine role allocations based on these predictions, thereby improving the performance of physical human-machine collaboration.
[0051] The present application is further described in detail below through a more detailed embodiment of the method for estimating intentions of physical human-machine collaboration of the present application: like Figure 2 The figure shows an implementation diagram of the intention estimation method for physical human-machine collaboration of the present application. The present application proposes a dual Transformer-based robot trajectory generator for physical human-machine collaboration, which has a hierarchical structure and uses human-guided motion and force data to quickly capture changes in human intentions, thereby achieving accurate trajectory prediction and dynamic robot behavior adjustment, thereby achieving effective collaboration. Specifically, the human intention estimation in the trajectory generator uses two transformer-based conditional variational autoencoders to combine the robot's motion data in the absence of obstacles with the human-guided trajectory and obstacle avoidance force. In addition, differential cooperative game theory is used to integrate predictions based on human-applied forces to ensure that the robot's behavior is consistent with human intentions. Figure 1The robot trajectory generator based on the dual Transformer incorporates human dynamics into long-term predictions, provides accurate intention understanding, achieves reasonable role allocation, and realizes the autonomy and operability of the robot. Figure 3 As shown in Figure 1, it is a schematic diagram of the overall structure of the obstacle-free trajectory prediction model and the obstacle trajectory prediction model. Figure 4 As shown in Figure 1, it is a structural diagram of the training of obstacle-free trajectory prediction model and obstacle trajectory prediction model. Figure 5 Figure 2 shows the structure of the obstacle-free trajectory prediction model and the obstacle-involved trajectory prediction model during testing / actual prediction. It should be noted that during training, both past and future data are known, forming data pairs, while during testing, only past data is available. The overall structure is essentially the same, with slight differences.
[0052] 1. Intent estimation.
[0053] This application is to learn multimodal probability distribution function The obstacle-free trajectory prediction model and the obstacle-prone trajectory prediction model of this application adopt the CVAE (Conditional Variational Autoencoder) framework and introduce Gaussian latent variables. The potential distribution of the data can be expressed as .
[0054] Through the Transformer and linear layer probability distribution ( ) is encoded and the loss is minimized by To optimize the probability distribution, loss It consists of a weighted negative variational lower bound and a prediction error.
[0055] Before learning features, the Transformer position encoding method is used to represent the temporal information of the data. This encoding is connected with the data input embedding and fed into the encoder. The temporal dependency of the data is learned through a multi-head self-attention mechanism. Specifically, given The end effector data of time steps is first extracted from the input and linearly projected to obtain the key , query Sum The keys, queries, and values of these mappings are fed into a self-attention mechanism to learn these features.
[0056] The past encoder uses a multi-head self-attention mechanism combined with a standard feedforward and norm layer to capture past features and encode past data into a potential distribution. It is then fed into a multilayer perceptron to generate Gaussian parameters. Similarly, the future encoder encodes future data into a potential distribution to capture features. This application combines a future-past attention module to update features by learning the relationship between past and future data. This is done by integrating past data features into the future. and and future data characteristics Combined, it is implemented by following the standard feedforward and norm layers. Finally, it is input into a multilayer perceptron to generate Gaussian parameters.
[0057] In the decoder, the sampling result Combined with the time convolution result of past data as input. During training, Generated by the Future Encoder Get, such as Figure 4 As shown. When testing, it is generated by the past encoder Obtained, such as Figure 5 As shown in Figure 2, the decoder structure is similar to the future encoder. Furthermore, appropriate masks are used during prediction to prevent the model from influencing subsequent data. Finally, the output is used as the parameters of the Gaussian mixture model, from which prediction results are sampled.
[0058] In this process, this application introduces a hierarchical structure to predict future trajectories in the long term. Two branches handle human-guided obstacle avoidance and obstacle-free trajectories respectively, and integrate them into role allocation.
[0059] 2. Role allocation.
[0060] like Figure 6 The figure shows the principle of role allocation. This application introduces differential cooperative game theory to reasonably allocate human and machine roles and reduce the differences based on human intention estimation.
[0061] The reference model of the robot can be designed as
[0062] in, are the required inertia matrix, Coriolis / centrifugal force matrix, and gravity, respectively. , is the trajectory tracking error, is the reference trajectory, i.e. the result of intention estimation, is the actual trajectory. It is the virtual robot force, It is the force exerted by humans. is the acceleration error, is the speed error.
[0063] First, introduce the state variable , and rewrite (1) as a linear state space system as .in, , represents the identity matrix, is the control input.
[0064] Next, consider humans and robots as two actors and define their cost function as:
[0065] in, is the tracking error weight of the human cost function compared to the human intended trajectory, is the tracking error weight of the human cost function compared to the robot trajectory, is the tracking error weight of the robot cost function compared to the robot trajectory, is the tracking error weight of the robot cost function compared to the human’s intended trajectory, The control input weights for the forces applied by humans, is the control input weight of the robot's control force. Based on differential cooperative game theory, the shared cost function in cooperation is defined as:
[0066] in, , is the role allocation coefficient that controls the weight of each cost, Construct a diagonal matrix, is a weighted combination of the human and robot reference trajectories. When the prediction is accurate, substituting (2) into (3) yields:
[0067] in, is the human intention trajectory weight, is the robot trajectory weight.
[0068] Minimizing the solution of (3) is a classic linear quadratic optimal control (LQR) problem, and its optimal control input is , is the final control signal input to the robot. , which can be obtained by solving the algebraic Riccati equation Obviously, the robot's assistance level can be adjusted by adjusting the role allocation coefficient In this application, the force applied by humans can be adjusted based on environmental factors and task requirements. To optimize robot behavior.
[0069] This application proposes a robot trajectory generator based on a dual Transformer model, which adopts a hierarchical structure of two transformer-based conditional variational autoencoders and introduces human-applied forces in long-term predictions, thereby realizing reasonable human-machine role allocation based on differential cooperative game theory and reducing human-machine disagreements. This application combines human intention estimation with role allocation to detect intention changes and reduce disagreements, effectively improving collaborative performance in complex and prone environments. The hierarchical structure of human intention estimation can simultaneously process motion and force in physical human-machine collaboration, improves prediction accuracy, and provides an accurate understanding of human intentions. Role allocation based on differential cooperative game theory realizes adaptive leader switching based on human-applied forces, ensuring that the robot's behavior is consistent with human intentions. Figure 1 to reduce disagreements while maintaining robot autonomy.
[0070] like Figure 7 FIG. 1 is a schematic diagram of an intention estimation system for physical human-machine collaboration, which may include: The data module is used to obtain the obstacle-free past trajectory and the obstacle-containing past trajectory respectively; A prediction module is used to input the obstacle-free past trajectory and the obstacle-involved past trajectory into an obstacle-free trajectory prediction model and an obstacle-involved trajectory prediction model, respectively, to obtain an obstacle-free predicted trajectory and an obstacle-involved predicted trajectory; The obstacle-free trajectory prediction model and the obstacle-prone trajectory prediction model have the same structure, and both processing methods include: The past trajectory and the position encoding results of the past trajectory are embedded and connected and then input into the encoder, and then pass through the multi-layer perceptron and sampling in sequence to obtain the sampling result; After performing temporal convolution on the past trajectory, it is combined with the sampling result and input into the decoder; The output of the decoder and the output of the encoder are subjected to the future-past attention mechanism to obtain updated features; The updated features are sequentially passed through the multi-layer perceptron, Gaussian mixture model and sampling to obtain the predicted trajectory; The role allocation module is used to allocate roles according to the obstacle-free predicted trajectory and the obstacle predicted trajectory.
[0071] It should be noted that in the several embodiments provided in this application, it should be understood that the disclosed devices and methods can be implemented in other ways. For example, the system embodiments described above are merely schematic. For example, the division of each module is only a logical function division. There may be other division methods in actual implementation. For example, multiple modules can be combined or integrated into another device, or some features can be ignored or not executed. The modules described as separate components may or may not be physically separated. The components displayed as modules may be one physical unit or multiple physical units, that is, they may be located in one place, or they may be distributed in multiple different places. Some or all of the modules may be selected according to actual needs to achieve the purpose of the scheme of this embodiment.
[0072] In addition, the modules in the various embodiments of the present invention may be integrated into a single processing unit, each module may exist physically separately, or two or more modules may be integrated into a single unit. The aforementioned integrated units may be implemented in the form of hardware or software functional units.
[0073] An embodiment of the present application also provides an electronic device, which may include one or more processors, memories, and communication interfaces.
[0074] The memory, the communication interface, and the processor are coupled together. For example, the memory, the communication interface, and the processor may be coupled together via a bus.
[0075] The communication interface is used to transmit data with other devices. The memory stores computer program code. The computer program code includes computer instructions that, when executed by a processor, cause the electronic device to perform the steps of the above-mentioned method for estimating intentions for physical human-machine collaboration.
[0076] Among them, the processor can be a processor or a controller, for example, a central processing unit (CPU), a general-purpose processor, a digital signal processor (DSP), an application-specific integrated circuit (ASIC), a field programmable gate array (FPGA) or other programmable logic devices, transistor logic devices, hardware components or any combination thereof. It can implement or execute the various exemplary logic blocks, modules and circuits described in conjunction with the contents of this disclosure. The processor can also be a combination that implements computing functions, such as a combination of one or more microprocessors, a combination of a DSP and a microprocessor, and the like. The processor can be used to support electronic devices in executing the method steps provided in the above embodiments.
[0077] The bus may be a Peripheral Component Interconnect (PCI) bus or an Extended Industry Standard Architecture (EISA) bus, etc. The above buses may be divided into an address bus, a data bus, a control bus, etc.
[0078] An embodiment of the present application provides a computer-readable storage medium, in which a computer program is stored. When the computer program is executed by a processor, the steps of the above-mentioned method for estimating the intention of physical human-computer collaboration are implemented.
[0079] The computer-readable storage medium involved in this application includes random access memory (RAM), internal memory, read-only memory (ROM), electrically programmable ROM, electrically erasable programmable ROM, registers, hard disks, removable disks, CD ROMs, or any other form of storage medium known in the technical field.
[0080] The above are merely preferred embodiments of the present application and are not intended to limit the present application. Those skilled in the art will readily appreciate that various modifications and variations are possible. Any modifications, equivalent substitutions, or improvements made within the spirit and principles of the present application shall be included within the scope of protection of the present application.
Claims
1. A method for estimating intention in physical human-machine collaboration, characterized in that: include: Obtain obstacle-free past trajectories and obstacle-containing past trajectories respectively; Inputting the obstacle-free past trajectory and the obstacle-containing past trajectory into the obstacle-free trajectory prediction model and the obstacle-containing trajectory prediction model respectively to obtain the obstacle-free predicted trajectory and the obstacle-containing predicted trajectory; The obstacle-free trajectory prediction model and the obstacle-prone trajectory prediction model have the same structure, and both processing methods include: The past trajectory and the position encoding results of the past trajectory are embedded and connected and then input into the encoder, and then pass through the multi-layer perceptron and sampling in sequence to obtain the sampling result; After performing temporal convolution on the past trajectory, it is combined with the sampling result and input into the decoder; The output of the decoder and the output of the encoder are subjected to the future-past attention mechanism to obtain updated features; The updated features are sequentially passed through the multi-layer perceptron, Gaussian mixture model and sampling to obtain the predicted trajectory; Roles are assigned based on the obstacle-free predicted trajectory and the obstacle predicted trajectory.
2. The method for estimating intention of physical human-machine collaboration according to claim 1, characterized in that: The method for performing role allocation includes: Differential cooperative game theory is used to allocate human and machine roles.
3. The method for estimating intention of physical human-machine collaboration according to claim 2, characterized in that: The method for allocating human-machine roles using differential cooperative game theory includes: Set the reference model of the robot; Introducing state variables, the reference model of the robot is rewritten as a linear state space system; Considering humans and robots as two actors, define the cost functions for humans and robots; Based on differential cooperative game theory, define the shared cost function in cooperation; Substituting the cost functions of humans and robots into the shared cost function in cooperation, we obtain the equation to be solved; Obtain the solution that minimizes the equation to be solved and obtain the control signal for input to the robot.
4. The method for estimating intention of physical human-machine collaboration according to claim 3, characterized in that: The expression of the reference model of the robot includes: in, is the inertia matrix, is the Coriolis / centrifugal force matrix, is gravity, is the trajectory tracking error, is the acceleration error, is the speed error, The force exerted on humans, is the force of the virtual robot; The expression of the linear state space system includes: in, , , ; is a linear state-space system, is the introduced state variable.
5. The method for estimating intention of physical human-machine collaboration according to claim 3, characterized in that: The cost functions for humans and robots include: in, is the cost function for humans, For human reference trajectory, is the tracking error weight of the human cost function compared to the human intended trajectory, is the robot reference trajectory, is the tracking error weight of the human cost function compared to the robot trajectory, The control input weights for the forces applied by humans, is the cost function of the robot, is the tracking error weight of the robot cost function compared to the robot trajectory, is the tracking error weight of the robot cost function compared to the human’s intended trajectory, The control input weights for the robot's control forces; The shared cost function in the cooperation includes: in, is the shared cost function in cooperation, To control the role allocation coefficient of each cost weight, is a weighted combination of the human and robot reference trajectories, , Construct a diagonal matrix; The equation to be solved includes: in, is the human intention trajectory weight, is the robot trajectory weight.
6. The method for estimating intention of physical human-machine collaboration according to claim 1, characterized in that: The processing method in the encoder includes: Given the end-effector data at T time steps, extract features from the past trajectory and the concatenation of the position encoding results of the past trajectory, and perform linear projection to obtain the key, query and value; The key, query, and value are fed into a self-attention mechanism to learn a latent distribution.
7. The method for estimating intention of physical human-machine collaboration according to claim 6, characterized in that: The future-past attention mechanism includes: Combining the keys and values of features from past data, and queries from future data, follows the standard feedforward and norm layer implementation.
8. A physical human-machine collaborative intention estimation system, characterized in that: include: The data module is used to obtain the obstacle-free past trajectory and the obstacle-containing past trajectory respectively; A prediction module is used to input the obstacle-free past trajectory and the obstacle-involved past trajectory into an obstacle-free trajectory prediction model and an obstacle-involved trajectory prediction model, respectively, to obtain an obstacle-free predicted trajectory and an obstacle-involved predicted trajectory; The obstacle-free trajectory prediction model and the obstacle-prone trajectory prediction model have the same structure, and both processing methods include: The past trajectory and the position encoding results of the past trajectory are embedded and connected and then input into the encoder, and then pass through the multi-layer perceptron and sampling in sequence to obtain the sampling result; After performing temporal convolution on the past trajectory, it is combined with the sampling result and input into the decoder; The output of the decoder and the output of the encoder are subjected to the future-past attention mechanism to obtain updated features; The updated features are sequentially passed through the multi-layer perceptron, Gaussian mixture model and sampling to obtain the predicted trajectory; The role allocation module is used to allocate roles according to the obstacle-free predicted trajectory and the obstacle predicted trajectory.
9. An electronic device, characterized in that: include: A memory, one or more processors; the memory is coupled to the processor; wherein computer program code is stored in the memory, and the computer program code includes computer instructions, and when the computer instructions are executed by the processor, the electronic device performs the steps of the intention estimation method for physical human-computer collaboration as described in any one of claims 1 to 7.
10. A computer-readable storage medium, characterized in that The computer-readable storage medium stores a computer program, which, when executed by a processor, implements the steps of the method for estimating intention of physical human-machine collaboration according to any one of claims 1 to 7.