Robot interaction method, model training method, device, equipment and storage medium
Patent Information
- Application Number
- CN202610957229.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-29
- Publication Date
- 2026-09-25
AI Technical Summary
[0003]本申请实施例提供一种机器人交互方法、装置、设备和存储介质,旨在解决现有技术中机器人交互过程中拟人化程度和生动程度较低导致用户沉浸程度较低的技术问题
[0023]本申请中通过响应针对目标机器人的机器人交互请求,获取所述机器人交互请求对应的关联交互数据和交互环境信息;驱动目标处理模型基于所述关联交互数据、交互环境信息和噪声潜向量生成交互潜在特征;基于所述关联交互数据中的历史运动帧和所述交互潜在特征进行解码,得到所述目标机器人的交互运动数据,所述交互运动数据包括关节运动数据和表情特效数据;按照所述交互运动数据控制所述目标机器人的目标硬件进行协同交互处理,得到机器人交互结果。实现在与目标机器人的人机交互过程中,能够通过目标处理模型基于关联交互数据和交互环境数据生成表征完整运动序列的交互潜在特征,并通过关联交互数据的历史运动帧和该交互潜在特征进行解码从而生成指导目标机器人目标硬件进行协同运动的交互运动数据,其中,该交互运动数据包括关节运动数据和表情特效数据,并按照该交互控制数据中的关节运动数据和表情特效数据控制目标机器人进行协同交互,实现在交互过程中目标机器人能够预测与本次交互对应的关节运动和表情特效并协同展示交互,从而提高目标机器人的交互灵活度和交互沉浸感。
Smart Images

Figure CN122816458A_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of computer technology, specifically to a robot interaction method, apparatus, device, and storage medium. Background Technology
[0002] Currently, with the rapid development of artificial intelligence and robotics, more and more robots are able to carry artificial intelligence models to interact with users in ways such as emotional companionship and care. However, existing robot interaction methods usually involve controlling specific joints in the robot to perform specific actions, and keeping the robot with a fixed expression or a simple preset expression during the interaction process. This results in existing robots having a low degree of anthropomorphism and vividness during the interaction process, thereby reducing the user's level of immersion in the interaction. Summary of the Invention
[0003] This application provides a robot interaction method, apparatus, device, and storage medium, aiming to solve the technical problem that the low degree of anthropomorphism and vividness in robot interaction in the prior art leads to a low degree of user immersion.
[0004] On one hand, embodiments of this application provide a robot interaction method, which includes the following steps: In response to a robot interaction request for a target robot, obtain the associated interaction data and interaction environment information corresponding to the robot interaction request; The target processing model generates latent interaction features based on the associated interaction data, interaction environment information, and noise latent vectors. Based on the historical motion frames in the associated interaction data and the potential interaction features, the interaction motion data of the target robot is obtained by decoding. The interaction motion data includes joint motion data and facial expression effect data. The target hardware of the target robot is controlled to perform collaborative interaction processing according to the interactive motion data to obtain the robot interaction result.
[0005] In one possible implementation of this application, the driving target processing model generates interaction latent features based on the associated interaction data, interaction environment information, and noise latent vectors, including: Analyze the associated interaction data to determine the historical motion frames corresponding to the associated interaction data; The target policy network in the target processing model is used to generate a noise latent vector according to the associated interaction data, historical motion frames and interaction environment information. The noise potential vector is sampled based on the target sampling strategy to obtain sampled noise features; The sampling noise features are denoised using the interactive environment information to obtain the interactive potential features.
[0006] In one possible implementation of this application, the step of decoding based on historical motion frames in the associated interaction data and the latent interaction features to obtain the interaction motion data of the target robot includes: Obtain the model decoding data corresponding to the target processing model; Based on the historical motion frames of the associated interaction data and the potential interaction features, the model decoding data is decoded to obtain the interactive motion data of the target robot.
[0007] In one possible implementation of this application, the step of decoding the model decoding data based on the historical motion frames of the associated interaction data and the latent interaction features to obtain the interactive motion data of the target robot includes: Obtain the historical motion frames of the target robot from the associated interaction data, the historical motion frames including joint historical motion frames and facial expression historical motion frames; The joint motion data is obtained by decoding the joint historical motion frames, the interaction potential features, and the model decoding data using the first decoding module; The second decoding module is used to decode the expression history motion frames, the interaction potential features, and the model decoding data to obtain expression effect data; Interactive motion data is generated based on the joint motion data and facial expression effects data.
[0008] In one possible implementation of this application, the step of collaboratively controlling the target hardware of the target robot according to the joint motion data and the facial expression effect data to perform collaborative interaction processing and obtain robot interaction results includes: The target joint hardware in the target hardware of the target robot is driven to perform the first interaction processing according to the joint motion data to obtain the action interaction result. The facial expression effect data drives the appearance display hardware in the target hardware of the target robot to perform a second interaction process to obtain the effect interaction result. The appearance display component is a hardware component used to display the appearance changes of the target robot during the interaction process. The robot interaction results are generated based on the action interaction results and the special effects interaction results.
[0009] On the other hand, this application provides a model training method, the model training method comprising: Obtain the initial model corresponding to the target robot, and the model training data corresponding to the target robot; The initial model is driven to generate training latent features based on the model training data and the distribution data to be trained. The initial model is driven to perform decoding processing based on the model training data and the training latent features to obtain the first predicted motion data; The initial model is updated based on the model training data, the training latent features, and the first predicted motion data to obtain the target processing model.
[0010] In one possible implementation of this application, driving the initial model to generate training latent features based on the model training data and the distribution data to be trained includes: Obtain historical frame training data and supervised training data from the model training data; The motion coding module in the initial model is used to perform prediction training on the historical frame training data, supervised training data, and training distribution data to obtain the prediction distribution parameters. The predicted distribution parameters are sampled based on the training sampling strategy to obtain training latent features.
[0011] In one possible implementation of this application, the decoding process based on the model training data and the training latent features to obtain the first predicted motion data includes: Obtain the training placeholder data for the initial model; The motion decoding module in the initial model is used to decode the training placeholder data based on the training latent features, historical frame training data, and training distribution data to obtain the first predicted motion data.
[0012] In one possible implementation of this application, updating the initial model based on the model training data, the training latent features, and the first predicted motion data to obtain the target processing model includes: Freeze the motion encoding module and motion decoding module in the initial model to obtain the model to be trained; The training data of the model and the first noise latent vector are used to generate second predicted motion data; The model to be trained is used to generate predicted trajectory information based on the model training data and target training conditions. The initial model is updated based on any one or more of the first predicted motion data, the second predicted motion data, and the predicted trajectory information to obtain the target processing model.
[0013] In one possible implementation of this application, the step of generating second predicted motion data using the model training data and the first noise latent vector based on the model to be trained includes: The historical frame training data and supervised training data in the model training data are encoded to obtain the potential features to be processed. The potential features to be processed are subjected to noise processing using preset noise data to obtain a first noise potential vector; The initial denoising module is driven to perform denoising prediction on the first noise potential vector to obtain the predicted potential features and the corresponding first potential loss data. The second predicted motion data is obtained by decoding the predicted latent features and historical frame training data using the frozen decoder in the model to be trained.
[0014] In one possible implementation of this application, generating predicted trajectory information using the model to be trained based on the model training data and target training conditions includes: Obtain the initial policy network in the initial model, and drive the initial policy network to generate a second noise latent vector according to the model training data and target training conditions; The second noise latent vector is predicted using the frozen target denoising module and the frozen motion decoding module in the model to be trained, thereby obtaining training trajectory data; The training trajectory data is simulated and predicted to obtain the predicted trajectory information.
[0015] In one possible implementation of this application, updating the initial model based on any one or more of the first predicted motion data, the second predicted motion data, and the predicted trajectory information to obtain the target processing model includes: The error is measured on the first predicted motion data to obtain the reconstruction loss data corresponding to the first predicted motion data. And / or, determine the divergence loss data of the first predicted motion data based on the training latent features and prior latent features corresponding to the first predicted motion data; And / or, determine auxiliary loss data for the first predicted motion data based on the first predicted motion data and preset constraints; Calculate the first training loss data corresponding to the first predicted motion data based on any one or more of the reconstructed loss data, divergence loss data, and auxiliary loss data and the coefficient of the first loss data. The initial model is updated using the first training loss data to obtain the target processing model.
[0016] In one possible implementation of this application, updating the initial model based on any one or more of the first predicted motion data, the second predicted motion data, and the predicted trajectory information to obtain the target processing model includes: Obtain the first potential loss data corresponding to the predicted potential features and the second potential loss data corresponding to the second predicted motion data; The initial model is updated using the first training loss data, the first potential loss data, and the second potential loss data corresponding to the first predicted motion data to obtain the target processing model.
[0017] In one possible implementation of this application, updating the initial model based on any one or more of the first predicted motion data, the second predicted motion data, and the predicted trajectory information to obtain the target processing model includes: Based on the predicted trajectory information and the target trajectory constraint data, determine the trajectory training information corresponding to the predicted trajectory information; The initial model is updated based on the trajectory training information to obtain the target processing model.
[0018] In one possible implementation of this application, obtaining the model training data corresponding to the target robot includes: Obtain the initial motion data corresponding to the target robot; The initial motion data is augmented to obtain model training data.
[0019] In one possible implementation of this application, the step of data augmentation of the initial motion data to obtain model training data includes: The target robot is subjected to posture recognition to obtain the target hardware corresponding to the target robot. Based on the target hardware and the initial motion data, the model training data is obtained by retargeting. And / or, perform data augmentation on the initial motion data based on a predictive data augmentation strategy to obtain model training data; And / or, generate motion data associated with the initial motion data based on the animation generation model, and redirect the initial motion data and the generated motion data with the target hardware to obtain model training data.
[0020] On the other hand, this application also provides a robot interaction device, the robot interaction device comprising: The information acquisition module is configured to respond to a robot interaction request for a target robot and acquire the associated interaction data and interaction environment information corresponding to the robot interaction request. The feature processing module is configured to drive the target processing model to generate latent interaction features based on the associated interaction data, interaction environment information, and noise latent vectors; The multidimensional motion generation module is configured to decode the historical motion frames and the potential interaction features in the associated interaction data to obtain the interactive motion data of the target robot, wherein the interactive motion data includes joint motion data and facial expression effect data. The collaborative interaction module is configured to control the target hardware of the target robot to perform collaborative interaction processing according to the interactive motion data, so as to obtain the robot interaction result.
[0021] On the other hand, this application also provides a robot interaction device, the robot interaction device comprising: One or more processors; Memory; and One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the steps of the robot interaction method.
[0022] On the other hand, this application also provides a computer-readable storage medium having a computer program stored thereon, the computer program being loaded by a processor to execute the steps in the robot interaction method.
[0023] In this application, in response to a robot interaction request for a target robot, associated interaction data and interaction environment information corresponding to the robot interaction request are obtained; a target processing model is driven to generate latent interaction features based on the associated interaction data, interaction environment information, and noise latent vectors; based on the historical motion frames in the associated interaction data and the latent interaction features, the interaction motion data of the target robot is decoded to obtain the interaction motion data of the target robot, the interaction motion data including joint motion data and facial expression effect data; the target hardware of the target robot is controlled to perform collaborative interaction processing according to the interaction motion data to obtain the robot interaction result. During human-robot interaction with a target robot, a target processing model can generate potential interaction features representing a complete motion sequence based on associated interaction data and interaction environment data. By decoding historical motion frames of associated interaction data and these potential interaction features, interactive motion data guiding the target robot's hardware to perform coordinated movements can be generated. This interactive motion data includes joint motion data and facial expression effect data. The target robot is controlled to perform coordinated interactions according to the joint motion data and facial expression effect data in the interactive control data. During the interaction, the target robot can predict the joint motion and facial expression effect corresponding to the current interaction and coordinate the interaction, thereby improving the interactive flexibility and immersion of the target robot. Attached Figure Description
[0024] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0025] Figure 1This is a schematic diagram of a scenario illustrating the robot interaction method in an embodiment of this application; Figure 2 This is a flowchart illustrating one embodiment of the robot interaction method in this application. Figure 3a A hardware schematic diagram of one embodiment of the target hardware of the target robot provided in this application; Figure 3b A timing diagram illustrating collaborative interaction in the robot interaction method provided in the embodiments of this application; Figure 4 A flowchart illustrating one embodiment of the model training method provided in this application; Figure 5a This is a schematic diagram of a real-world example of training a motion coding module in the model training method provided in this application embodiment; Figure 5b A schematic diagram of the structure of an embodiment of the model training method provided in this application for training the initial denoising module; Figure 5c A schematic diagram of a scenario in which an initial policy network is trained in the model training method provided in this application; Figure 6 A schematic diagram of the structure of one embodiment of the robot interaction device provided in this application; Figure 7 This is a schematic diagram of the structure of one embodiment of the robot interaction device provided in this application. Detailed Implementation
[0026] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0027] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.
[0028] In this application, the term "exemplary" is used to mean "serving as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.
[0029] Currently, with the rapid development of artificial intelligence and robotics, more and more robots are able to carry artificial intelligence models to interact with users in ways such as emotional companionship and care. However, existing robot interaction methods usually involve controlling specific joints in the robot to perform specific actions, and keeping the robot with a fixed expression or a simple preset expression during the interaction process. This results in existing robots having a low degree of anthropomorphism and vividness during the interaction process, thereby reducing the user's level of immersion in the interaction.
[0030] Based on this, this application proposes a robot interaction method, apparatus, device, and computer-readable storage medium to solve the technical problem that the low degree of anthropomorphism and vividness in robot interaction in the prior art leads to a low degree of user immersion.
[0031] The robot interaction method in this embodiment of the invention is applied to a robot interaction device, which is set in a robot interaction equipment. The robot interaction equipment is provided with one or more processors, a memory, and one or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the robot interaction method. The robot interaction equipment can be an intelligent robot terminal capable of human-computer interaction with users, such as any intelligent robot such as a companion robot or a sweeping robot. Optionally, the robot interaction equipment can also be a server or a service cluster composed of multiple servers.
[0032] like Figure 1 As shown, Figure 1 This is a schematic diagram of a robot interaction method according to an embodiment of the present application. The robot interaction scenario in this embodiment includes a robot interaction device 100 (the robot interaction device 100 integrates a robot interaction apparatus), and the robot interaction device 100 runs a computer-readable storage medium corresponding to the robot interaction method to execute the steps of the robot interaction method.
[0033] Understandable Figure 1 The robot interaction device in the robot interaction method scenario shown, or the device included in the robot interaction device, does not constitute a limitation on the embodiments of the present invention. That is, the number or type of robot interaction device included in the robot interaction method scenario, or the number or type of device included in each device, does not affect the overall implementation of the technical solution in the embodiments of the present invention, and can all be considered as equivalent substitutions or derivatives of the technical solutions claimed in the embodiments of the present invention.
[0034] In this embodiment of the invention, the robot interaction device 100 is mainly used to: respond to a robot interaction request for a target robot, and obtain the associated interaction data and interaction environment information corresponding to the robot interaction request; The target processing model generates latent interaction features based on the associated interaction data, interaction environment information, and noise latent vectors. Based on the historical motion frames in the associated interaction data and the potential interaction features, the interaction motion data of the target robot is obtained by decoding. The interaction motion data includes joint motion data and facial expression effect data. The target hardware of the target robot is controlled to perform collaborative interaction processing according to the interactive motion data to obtain the robot interaction result.
[0035] The robot interaction device 100 in this embodiment of the invention can be an independent robot interaction device. For example, the robot interaction device can be an intelligent robot terminal that can interact with users, such as any intelligent robot like a companion robot or a sweeping robot. It can also be a robot interaction network or robot interaction cluster composed of multiple robot interaction devices.
[0036] This application provides a robot interaction method, apparatus, device, and computer-readable storage medium, which will be described in detail below.
[0037] It will be understood by those skilled in the art that Figure 1 The application environment shown is only one application scenario related to the solution of this application and does not constitute a limitation on the application scenario of this application. Other application environments may include more than one application scenario. Figure 1 The diagram shows more or fewer robot interaction devices, or robot interaction network connections, for example... Figure 1 Only one robot interaction device is shown in the figure. It can be understood that the scenario of this robot interaction method may also include one or more robot interaction devices, which are not limited here. The robot interaction device 100 may also include a memory for storing interactive motion data and other data.
[0038] It should be noted that, Figure 1 The schematic diagram of the robot interaction method shown is merely an example. The scenarios of the robot interaction method described in the embodiments of the present invention are intended to more clearly illustrate the technical solutions of the embodiments of the present invention and do not constitute a limitation on the technical solutions provided in the embodiments of the present invention.
[0039] Based on the scenarios described above regarding robot interaction methods, various embodiments of the robot interaction method disclosed in this invention are proposed.
[0040] like Figure 2 As shown, Figure 2 This is a flowchart illustrating one embodiment of the robot interaction method in this application. The robot interaction method includes the following steps 201 to 204: 201. Respond to a robot interaction request for the target robot, and obtain the associated interaction data and interaction environment information corresponding to the robot interaction request; The robot interaction method in this embodiment is applied to a robot interaction device. The type and number of robot interaction devices are not specifically limited. That is, the robot interaction device can be one or more intelligent robots or intelligent robot control terminals that can interact with users. In a specific embodiment, the robot interaction device is a target robot or a control terminal corresponding to the target robot.
[0041] Optionally, the target robot is an intelligent robot that responds to robot interaction requests and can generate interactive motion data in real time to control the synchronous and coordinated movement of various target hardware components for human-computer interaction. Optionally, the target robot can be a companion robot used for interactions such as caregiving, companionship, and monitoring. The target hardware is robot hardware used to output interactive actions and special effects of the target robot. In a specific embodiment, the target hardware includes joint hardware, special effects display hardware (eye components, blush components, etc.), and audio components, etc.
[0042] Optionally, a robot interaction request is an operation event that drives the target robot to generate corresponding interactive motion data and controls the target hardware to collaboratively display the interactive motion data to achieve interaction with the target object. This robot interaction request can be triggered by the target object interacting in any way: for example, in one specific embodiment, the robot interaction request can be triggered by the target object inputting text interaction data to the target robot via text interaction. In another specific embodiment, the robot interaction request can also be triggered by the target object inputting voice interaction data to the target robot via voice interaction. Yet another specific embodiment, the robot interaction request can also be triggered by the target object displaying a specified action to the target robot via gesture interaction. Furthermore, the robot interaction request can also be an interaction event triggered by the target robot's real-time detection of the target object or the environmental area where the target object is located. The target object can be a user or other interaction object.
[0043] Optionally, the associated interaction data is an interaction information stream between the target robot and the target object during interaction or dialogue. This associated interaction data includes the sum of current and historical interaction information between the target robot and the target object. In one specific embodiment, the associated interaction data includes the target object's historical interaction commands and the historical motion frames and / or historical dialogue streams executed by the target robot based on these historical interaction commands.
[0044] Optionally, the interactive environment information is environmental information obtained by the target robot through data collection of the environmental area where the target object is located. Optionally, in one specific embodiment, the interactive environment information can be environmental image information of the environmental area where the target object is located. Optionally, the interactive environment information also includes sensor data obtained by other sensors in the target robot (such as temperature and humidity sensors) through data collection of the environmental area where the target object is located.
[0045] Optionally, during operation, the target robot or robot interaction device responds to robot interaction requests for the target robot. After receiving the robot interaction request, the robot interaction device reads the current interaction data and historical interaction data corresponding to the robot interaction request as associated interaction data. In addition, the target robot can also call environmental acquisition components (such as image acquisition components or other sensor components) to collect environmental images or other sensor data of the environmental area where the target object that triggered the robot interaction request is located as interaction environment information.
[0046] 202. The target processing model generates interactive latent features based on the associated interaction data, interaction environment information, and noise latent vectors; Optionally, after responding to the robot interaction request and obtaining the interaction data and interaction environment information associated with the robot interaction request, the target robot also drives the target processing model to perform data processing based on the associated interaction data, interaction environment information and pre-generated noise latent vectors to generate interaction latent features.
[0047] Optionally, the target processing model is an artificial intelligence model used to generate interactive motion data that drives the coordinated movement of various target hardware components of the target robot. Here, coordinated movement of the target hardware components refers to the fact that during human-computer interaction, each target hardware component jointly outputs corresponding interactive motion data to achieve integrated expression of the target robot's actions and facial expressions, thereby improving the target robot's interactive performance. Optionally, in a specific embodiment, the target processing model includes a motion decoding module, a target denoising module, and a target policy network.
[0048] The motion coding module has its input terminal as the first input terminal of the target processing model, used to receive the interaction feature data and the interaction environment information; the output terminal of the motion coding module is connected to the first input terminal of the motion decoding module, and the output terminal of the motion decoding module is the output terminal of the target processing model; the first denoising input terminal of the target denoising module is the second input terminal of the target processing model, and the output terminal of the target denoising module is connected to the second input terminal of the motion decoding module; the input terminal of the target policy network is the third input terminal of the target processing model, and the output terminal of the target policy network is connected to the second denoising input terminal of the target denoising module.
[0049] Optionally, this motion decoding module is used to decode historical motion frames and latent vector data to obtain interactive motion data.
[0050] Optionally, the target denoising module is a denoising module for generating interaction latent features based on the noise latent vector, the interaction feature data, and the interaction environment information.
[0051] Optionally, the target policy network is a policy network module for generating noisy latent vectors based on historical motion sequences, the interaction feature data, and the interaction environment information.
[0052] Optionally, the noise latent vector is a noise-carrying latent vector generated by the target processing model. In one specific embodiment, the noise latent vector is a noise latent vector generated by the target policy network in the target processing model.
[0053] Optionally, the interaction latent features are low-noise latent vectors characterizing the motion sequences to be executed by the target robot. That is, the interaction latent features contain the motion sequences (including joint motion features and facial expression motion features) to be executed by each target hardware component of the target robot. Optionally, the interaction latent features can be generated and diffused by the target processing model from historical motion sequences and interaction features to obtain interaction latent features and interaction motion data.
[0054] Optionally, the terminal parses the associated interaction data to determine the corresponding historical motion frame. This historical motion frame contains motion data information already performed by the target robot within the associated interaction data. It provides contextual information about the joint movements and facial expressions that the target robot has performed, thus describing the target robot's current posture and motion trend.
[0055] Optionally, the terminal utilizes the target policy network in the target processing model to generate a noise latent vector based on the associated interaction data, historical motion frames, and interaction environment information. That is, the terminal uses the already trained target policy network in the target processing model to generate a noise latent vector using the historical motion frames. In other words, the target policy network can process the model input based on the historical motion frames, interaction environment information, and associated interaction data to generate a noise latent vector associated with the current interaction scenario of the target robot, ensuring that the subsequently generated interaction motion data can be safely executed in the target hardware of the target robot. Here, the target policy network is a policy network pre-trained using reward functions corresponding to various physical constraints (e.g., velocity limits, acceleration smoothness, joint torque limits, etc.). In a specific embodiment, the target policy network is a policy network trained based on the PPO (Proximal Policy Optimization) algorithm.
[0056] Optionally, the terminal generates multiple motion trajectories corresponding to historical motion frames, interaction environment information, and associated interaction data based on the target generation strategy of the target strategy network, and obtains the conditional state corresponding to the motion trajectory. This conditional state is then input into the target strategy network to generate a noise latent vector. Optionally, in a specific embodiment, the expressions for the target generation strategy and the conditional state are as follows:
[0057]
[0058] in, The target generation strategy is defined by Z, where Z is the potential noise vector. Let θ be the conditional state, and θ be the policy network parameters obtained from training the target policy network.
[0059] Optionally, after acquiring the noise latent vector, the terminal samples the noise latent vector to obtain sampled noise features. That is, the terminal samples the noise latent vector from a standard normal distribution to generate sampled noise features, and then uses the interaction environment information to denoise these sampled noise features to obtain interaction latent features. In other words, the terminal combines the interaction environment information and associated interaction data to perform multi-step denoising on the sampled noise features to generate interaction latent features.
[0060] 203. Based on the historical motion frames in the associated interaction data and the potential interaction features, decode to obtain the interaction motion data of the target robot; Optionally, after acquiring the interaction potential features that characterize the motion sequence to be executed by the target robot, the terminal also decodes the historical motion frames and interaction potential features in the associated interaction data to generate the interaction motion data of the target robot.
[0061] Optionally, the interactive motion data is a sequence of interactive motions to be executed by the target hardware in the target robot. This interactive motion data includes joint motion data and facial expression effect data. Specifically, the key motion data is a motion sequence used to control the movement of the target joint hardware in the target hardware. The facial expression effect data is used to control the appearance display components (e.g., eye components, display screen components, LED components, etc.) in the target hardware to display corresponding facial expression effect animations. In one specific embodiment, the interactive motion data may include robot motion data such as the displacement and rotation of the root joints, the rotation of the body joints, the rotation and scaling of the eyes, and changes in the motion attributes (displacement, rotation, scaling) and material attributes (e.g., color, transparency, etc.) of other motion expression components (e.g., nose, blush, heartbeat, etc.).
[0062] Optionally, the terminal acquires the model decoding data corresponding to the target processing model. This model decoding data is placeholder data for the model to generate interactive motion data.
[0063] Optionally, the terminal decodes the model decoding data based on the historical motion frames of the associated interaction data and the interaction latent features to obtain the interactive motion data of the target robot. That is, the terminal combines the historical motion frames and interaction latent features to obtain a combined latent vector, uses the combined latent vector as the module input of the motion decoding module in the target processing model, and uses the motion decoding module in the target processing model to decode the model decoding data based on the combined latent vector to obtain the joint motion data and facial expression effect data to be executed by the target robot from the model decoding data, and determines the joint motion data and facial expression effect data as the interactive motion data of the target robot.
[0064] Optionally, in one specific embodiment, the target processing model can set up a motion decoding model to uniformly generate interactive motion data, thereby improving data generation efficiency. Optionally, in other embodiments, the target processing model can also set up multiple motion decoders corresponding to different target hardware to decode the interactive motion data corresponding to different target hardware, thereby ensuring the detail quality of the predicted interactive motion data.
[0065] Optionally, in one specific embodiment, the motion decoder in the target processing model includes a first decoding module and a second decoding module. The first decoding module is used to predict the motion data corresponding to the target joint hardware in the target hardware of the target robot. The second decoding module is used to predict the facial expression effect data corresponding to the appearance display hardware in the target hardware of the target robot.
[0066] Optionally, the terminal may acquire the historical motion frames of the target robot in the associated interaction data, including joint historical motion frames and facial expression historical motion frames.
[0067] Optionally, the terminal uses the first decoding module to decode the joint's historical motion frames, the interaction potential features, and the model decoding data to obtain the joint motion data to be executed by the target joint hardware.
[0068] Optionally, the terminal uses the second decoding module to decode the historical motion frame of the facial expression, the potential interaction features, and the model decoding data to obtain the facial expression effect data to be executed by the appearance display hardware. Based on the joint motion data and the facial expression effect data, the terminal generates interactive motion data. That is, the terminal determines the joint motion data and the facial expression effect data as the interactive motion data of the target robot.
[0069] 204. Control the target hardware of the target robot to perform collaborative interaction processing according to the interactive motion data to obtain the robot interaction result.
[0070] Optionally, after acquiring the interactive motion data corresponding to each target hardware, the terminal also controls the target hardware of the target robot to perform collaborative interactive processing according to the interactive motion data, and obtains the robot interaction result.
[0071] Optionally, the target hardware is robot hardware that enables the target robot to move or display during human-computer interaction. That is, the target hardware is a component used for motion design of the target robot. Optionally, in one specific embodiment, the target hardware includes movable target joint hardware of the target robot and an appearance display component capable of displaying content. The target joint hardware is robot hardware used to perform corresponding interactive actions through movements such as displacement and rotation. The appearance display component is a hardware component used to display changes in the appearance effects of the target robot during interaction. In one specific embodiment, the target joint hardware includes a motion joint bound to a skeletal structure. For example, such as... Figure 3a As shown, Figure 3a This is a hardware schematic diagram of one embodiment of the target hardware for the target robot provided in this application. Figure 3a In the illustrated embodiment, the target joint hardware includes root joints, leg joints, and shoulder joints. Each target joint hardware component has its motion angle limited according to actual needs to ensure that the target robot can move effectively under the physical constraints set during the design phase. Furthermore, the appearance display components include the target robot's eye components, blush components, nose components, and skin tone components. Optionally, the terminal can control the deformation and color effect changes of each display component of the target robot through geometric properties. Optionally, the terminal can also configure other appearance display components of the target robot, such as blush, nose, and heartbeat, using keyframes.
[0072] Optionally, the terminal can utilize interactive motion data to control different target hardware within the target robot to execute corresponding actions or effects, thereby enabling the coordinated output of interactive actions and effects by each target hardware component. For example... Figure 3b As shown, Figure 3b This is a timing diagram illustrating the collaborative interaction in the robot interaction method provided in the embodiments of this application. Figure 3b In the illustrated embodiment, the target robot, based on interactive motion data, performs an interactive operation where it first raises its hand and then kicks its leg, with its eyes gradually flattening and then rounding, and its eye color changing from black to red. The target hardware of the target robot consists of shoulder joints, leg joints, and an eye display component. The target robot can collaboratively drive the shoulder and leg joints to swing based on joint motion data and facial expression effect data, and drive the eye display component to display animation effects (blinking, changes in pupil color and size, etc.), thereby improving the flexibility and immersion of the interaction.
[0073] Optionally, in one specific embodiment, the terminal drives the target joint hardware (including root joints, leg joints, and shoulder joints, etc.) of the target robot to perform a first interaction process according to the joint motion data, thereby obtaining a motion interaction result. Additionally, the terminal drives the appearance display hardware of the target robot to perform a second interaction process according to the facial expression effect data, thereby obtaining an effect interaction result, and generates a robot interaction result based on the motion interaction result and the effect interaction result.
[0074] In this embodiment, the robot interaction device responds to a robot interaction request for a target robot, acquires associated interaction data and interaction environment information corresponding to the robot interaction request, drives the target processing model to generate interaction latent features based on the associated interaction data, interaction environment information, and noise latent vectors, decodes historical motion frames in the associated interaction data and the interaction latent features to obtain the interaction motion data of the target robot, the interaction motion data including joint motion data and facial expression effect data, and controls the target hardware of the target robot to perform collaborative interaction processing according to the interaction motion data to obtain the robot interaction result. During human-robot interaction with a target robot, the system can generate potential interaction features that characterize a complete motion sequence by associating interaction data and interaction environment data. It can then decode historical motion frames of the associated interaction data and these potential interaction features to generate interactive motion data that guides the target robot's hardware to perform coordinated movements. This interactive motion data includes joint motion data and facial expression effect data. The system controls the target robot to perform coordinated interactions according to the joint motion data and facial expression effect data in the interaction control data. This allows the target robot to predict and coordinate the joint movements and facial expression effects corresponding to the current interaction, thereby improving the target robot's interactive flexibility and immersive experience.
[0075] like Figure 4 As shown, Figure 4 This is a flowchart illustrating one embodiment of the model training method provided in this application. Figure 4 In the illustrated embodiment, the model training method includes steps 301 to 304: 301. Obtain the initial model corresponding to the target robot, and the model training data corresponding to the target robot; 302. Drive the initial model to generate training latent features based on the model training data and the distribution data to be trained; 303. Drive the initial model to perform decoding processing based on the model training data and the training latent features to obtain the first predicted motion data; 304. Update the initial model based on the model training data, the training latent features, and the first predicted motion data to obtain the target processing model.
[0076] Based on the above embodiments, in this embodiment, before the control terminal corresponding to the target robot generates interactive motion data for coordinating the interaction of various target hardware components of the target robot using the target processing model, it also trains the initial model to generate the target processing model.
[0077] Optionally, the initial model is an artificial intelligence model to be trained for the target robot, used to predict the interactive motion data to be performed by the target robot. This initial model can be a local artificial intelligence model deployed on the target robot or a cloud model deployed on a cloud server. Optionally, in one specific embodiment, the initial model includes a motion encoding module, a motion decoding module, an initial denoising module, and an initial policy network. The motion coding module has its input terminal as the first input terminal of the initial model, used to receive the interaction feature data and the interaction environment information; the output terminal of the motion coding module is connected to the first input terminal of the motion decoding module, and the output terminal of the motion decoding module is the output terminal of the initial model; the first denoising input terminal of the initial denoising module is the second input terminal of the initial model, and the output terminal of the initial denoising module is connected to the second input terminal of the motion decoding module; the input terminal of the initial policy network is the third input terminal of the initial model, and the output terminal of the initial policy network is connected to the second denoising input terminal of the initial denoising module.
[0078] Optionally, after acquiring the initial model corresponding to the target robot, the terminal further acquires the model training data corresponding to the target robot. The model training data is the training data used to train the initial model. In one specific embodiment, the model training data is keyframe animation of motion data associated with the target robot's motion.
[0079] Optionally, in one specific embodiment, the terminal acquires the initial motion data corresponding to the target robot. That is, after acquiring the target hardware of the target robot as a motion expression component, the terminal also acquires the initial motion data corresponding to each target hardware. This initial motion data is a keyframe animation corresponding to pre-drawn motion data. In other words, this initial motion data can be pre-drawn motion data animations corresponding to the target robot by animators for different interactive scenarios. This initial motion data can be a continuously changing quantity (e.g., the target robot's pupils gradually dilate, joint positions smoothly change), or a sudden change (e.g., a sudden change in the value of a heartbeat between two keyframes). This initial motion data can be saved in a 3D asset exchange file format.
[0080] Optionally, in order to increase the amount of training data, the terminal performs data augmentation on the initial motion data after acquiring it to obtain model training data, thereby expanding the scale of training data and improving the subsequent training effect.
[0081] Optionally, in other embodiments, the terminal performs posture recognition on the target robot to obtain the target hardware corresponding to the target robot, and redirects the initial motion data based on the target hardware and the initial motion data to obtain model training data. That is, the terminal performs posture recognition on the robot to determine the target hardware and corresponding target skeleton information of the target robot, and redirects the initial motion data using the target skeleton information corresponding to the target hardware to generate model training data.
[0082] Optionally, the terminal can also perform data augmentation on the initial motion data based on a predictive data augmentation strategy to obtain model training data. That is, the terminal can perform data transformation operations such as frame interpolation, frame dropping, motion noise addition, smoothing, and playback time transformation on each initial motion data based on the data augmentation strategy to obtain expanded model training data. Here, the data augmentation strategy is a control strategy that selects corresponding data transformation operations to transform the initial motion data. Optionally, in a specific embodiment, the data transformation operations in the data augmentation strategy include any one or more of frame interpolation, frame dropping, motion noise addition, smoothing, and playback time transformation.
[0083] Optionally, the terminal can also generate generated motion data associated with the initial motion data based on the animation generation model, and redirect the initial motion data and the generated motion data with the target hardware to obtain model training data. That is, the terminal uses the initial motion data of the target robot as model input, uses the animation generation model to generate new generated motion data from the initial motion data, and redirects the initial motion data and the generated motion data with the target hardware and target skeleton information to obtain model training data. The animation generation model is a production model used to generate new image data or video data. Optionally, in a specific embodiment, the animation generation model can be any video generation model such as the Stable VideoDiffusion model or the JiMeng video generation model.
[0084] Optionally, after acquiring the augmented model training data, the terminal also uses a data annotation model to score and annotate the training data. Specifically, the terminal uses the data annotation model to annotate the action semantics and emotional information corresponding to each training data point, generating sufficient, high-quality motion data with task-related text or visual annotations as the annotated model training data. This model training data consists of a time-series containing motion data. Each frame includes the pose of each joint of the robot (position, rotation, scaling, etc.), expression-related parameters (such as eyes, mouth shape, and facial control parameters), and possible material or appearance attributes, used to fully characterize the action and facial expression state at that moment. In addition, each motion segment is accompanied by label information jointly generated by a multimodal model and humans, including action semantics, emotional expression, and other task-related text or visual labels.
[0085] Optionally, after acquiring the model training data, the terminal also uses the model training data to train the initial model. That is, the terminal drives the initial model to generate training latent features based on the model training data and the distribution data to be trained. Optionally, the terminal uses the model training data and the distribution data to be trained to train the motion coding module to generate training latent features and corresponding loss data information, and uses this loss data information to update the parameters of the initial model in subsequent steps. The distribution data to be trained is a pre-set trainable distribution Token (Tμ, Tσ) in the initial model.
[0086] Optionally, the terminal acquires historical frame training data and supervised training data from the model's training data. Historical frame training data consists of motion frames that have already occurred within the model's training data, providing context for past actions and facial expressions, and describing the current posture and motion trend. Supervised training data consists of motion frames from the current moment onwards, used as supervisory signals in latent distribution learning, enabling the generated training latent features to represent the complete motion pattern.
[0087] Optionally, the terminal uses the motion coding module in the initial model to perform prediction training on the historical frame training data, supervised training data, and training distribution data to obtain prediction distribution parameters. The motion coding module is an encoding module used to generate training latent features. In one specific embodiment, this motion coding module is a variational autoencoder based on a Transformer structure.
[0088] That is, such as Figure 5a As shown, Figure 5a This is a schematic diagram illustrating a real-world example of training a motion coding module in the model training method provided in this application embodiment. Figure 5aIn the illustrated embodiment, the terminal uses historical frame training data, supervised training data, and training distribution data as model inputs to the motion encoder. The motion encoder predicts encoding distribution parameters based on these historical frame training data, supervised training data, and training distribution data to obtain predicted distribution parameters. These predicted distribution parameters are then sampled based on a training sampling strategy to obtain training latent features. Optionally, in one specific embodiment, this training sampling strategy is reparameterization trick sampling. The training latent features are latent vectors representing the complete motion sequence, learned from the predicted distribution parameters predicted by the motion coding module. Optionally, these training latent features follow an easily sampled prior distribution, thus laying the foundation for subsequent generation and diffusion modeling in the latent space.
[0089] Optional, such as Figure 5a As shown, after generating training latent features, the motion encoding module inputs these features into the motion decoding module for decoding training. That is, the terminal obtains the training placeholder data for the initial model. The training placeholder data refers to the placeholder data information to be decoded to generate predicted motion data. Optionally, in a specific embodiment, the training placeholder data is a placeholder token in the initial model. Optionally, the terminal uses the training latent features, historical frame training data, and the training distribution data as model inputs to the motion decoding module. The motion decoding module in the initial model decodes the training placeholder data based on the training latent features, historical frame training data, and the training distribution data to obtain first predicted motion data and corresponding first training loss data. The first predicted motion data is the predicted motion data obtained by training the motion encoding module and the motion decoding module to predict the future movements and expressions of the target robot.
[0090] Optionally, the first training loss data is loss data used to measure and evaluate the generation quality of the first predicted motion data. This first training loss data is calculated from any one or more of the reconstruction loss data, divergence loss data, and auxiliary loss data corresponding to the first predicted motion data.
[0091] Optionally, in one specific embodiment, the reconstruction loss data is loss data obtained by using supervised training data to measure the error of the first predicted motion data. Optionally, in one specific embodiment, the reconstruction loss data is calculated as follows:
[0092] in, This refers to the divergence loss data corresponding to the first predicted motion data. The value of the first predicted motion data at time step t and dimension d. To supervise the training data at time step t and dimension d. These are model parameters used to emphasize the importance of certain target hardware.
[0093] Optionally, in one specific embodiment, the divergence loss data is a loss data function used to align the first predicted motion data and the supervised training data. Optionally, in one specific embodiment, the divergence loss data is calculated as follows:
[0094] in, This is the divergence loss data corresponding to the first predicted motion data.
[0095] Optionally, in one specific embodiment, the auxiliary loss data is a loss data function used to penalize specified motion data in the first predicted motion data. For example, the auxiliary loss data can be a loss data function that penalizes joint rotation in the first predicted motion data that exceeds a predetermined angle; that is, when a joint rotation angle in the first predicted motion data exceeds a preset upper or lower angle limit, a penalty is applied to the first predicted motion data. For example, in one specific embodiment, the auxiliary loss data can be:
[0096] in, To constrain the auxiliary loss data of joint rotation angles in the first predicted motion data, The joint rotation angle is the first predicted motion data.
[0097] Optionally, in another specific embodiment, the auxiliary loss data is a loss data function used to penalize the acceleration in the first predicted motion data, thereby constraining the degree of change between adjacent frames in the first predicted motion data and preventing the changes between predicted adjacent frames from being too abrupt. Optionally, in one specific embodiment, the auxiliary loss data can be:
[0098] Optional, To constrain the auxiliary loss data of joint rotation angles in the first predicted motion data, and This represents the acceleration of adjacent frames.
[0099] Optionally, after acquiring any one or more of the reconstruction loss data, divergence loss data, and auxiliary loss data corresponding to the first predicted motion data, the terminal calculates the first training loss data corresponding to the first predicted motion data based on the first loss data coefficients and any one or more of the reconstruction loss data, divergence loss data, and auxiliary loss data corresponding to the first predicted motion data. In a specific embodiment, the calculation method of the first training loss data is as follows:
[0100] in, This is the first training loss data. and The coefficient for the first loss data.
[0101] Optionally, when training the motion encoder and motion decoder, the terminal can minimize the first training loss data using a training optimizer to train the motion encoder and motion decoder. To improve convergence speed, the terminal can first train using the reconstruction loss data and then gradually increase the first loss data coefficients corresponding to the divergence loss data. The first loss data coefficients are calculated coefficients used to weight the divergence loss data and / or auxiliary loss data.
[0102] Optionally, after obtaining the minimized first training loss data, the terminal uses the first training loss data to configure the parameters of the motion encoder and motion decoder in the initial model in subsequent steps to obtain the trained motion encoder and trained motion decoder.
[0103] Optionally, the terminal also performs diffusion model training on the initial denoising module in the initial model to generate a target processing model containing the target denoising network obtained by training the initial denoising module.
[0104] Optionally, during the diffusion model training phase of the initial model, the terminal freezes the motion coding module and motion decoding module in the initial model to obtain the model to be trained. The model to be trained is the model structure obtained by freezing the motion coding module and motion decoding module of the initial model.
[0105] Optional, such as Figure 5b As shown, Figure 5b This is a schematic diagram of the structure of one embodiment of the model training method provided in this application for training the initial denoising module. Figure 5b In the embodiment shown, the terminal inputs the model training data, the number of noise addition steps, and the first noise latent vector into the initial denoising module, and uses the model to be trained to generate second predicted motion data from the model training data and the first noise latent vector.
[0106] Optionally, after generating the model to be trained, the terminal acquires historical frame training data and corresponding supervised training data from the model training data, and encodes the complete motion sequences of the historical frame training data and supervised training data to obtain the latent features to be processed. Optionally, the latent features to be processed are noise-free latent vectors representing the encoded information of the complete motion sequence obtained during the initial model training process.
[0107] Optionally, after acquiring the latent feature to be processed, the terminal adds noise to the latent feature using preset noise data to obtain a first noise latent vector. Optionally, in one specific embodiment, the terminal uses preset noise data to forward-add noise to the latent feature to be processed to construct a noise-carrying latent vector, obtaining a first noise latent vector. The first noise latent vector is a latent vector containing preset noise. Optionally, in one specific embodiment, the preset noise data can be Gaussian noise.
[0108] Optionally, the noise-adding process for the latent features to be processed is as follows:
[0109] in, Let be the first noise potential vector. This is a noise scheduling hyperparameter.
[0110] Optionally, after obtaining the first noise latent vector, the terminal drives the initial denoising module to perform denoising prediction on the first noise latent vector, obtaining predicted latent features and corresponding first latent loss data. That is, the terminal inputs the first noise latent vector, the number of diffusion steps, and the label information in the model training data as training conditions and historical frame training data into the initial denoising module to drive the initial denoising module to perform denoising prediction on the first noise latent vector, outputting the denoised predicted latent features and the corresponding first latent loss data. The first latent loss data is loss data characterizing and measuring the denoising quality of the predicted latent features. Optionally, in a specific embodiment, the prediction process of the predicted latent features is as follows:
[0111]
[0112] in, To predict latent features, H represents historical frame training data, and c represents training conditions generated from the label information corresponding to the historical frame training data.
[0113] Optionally, in one specific embodiment, the first potential loss data is calculated as follows:
[0114] This is the first potential loss data. To predict latent features, These are potential features to be processed.
[0115] Optionally, after generating the predicted latent features, the terminal further decodes the predicted latent features and historical frame training data using the frozen decoder in the model to be trained, to obtain the second predicted motion data. That is, the terminal uses the predicted latent features as model input to the frozen decoder, and uses the frozen decoder to decode the training placeholder data based on the predicted latent features and historical frame training data to obtain the second predicted motion data.
[0116] Optionally, the terminal further calculates second potential loss data corresponding to the second predicted motion data. The second potential loss data is a loss function used to measure the difference between the second predicted motion data and the corresponding supervised training data representing the true future frame. Optionally, in a specific embodiment, the second potential loss data is reconstruction loss data, and the calculation method of the second potential loss data is as follows:
[0117] in, This is the second potential loss data. For the second predicted motion data, To monitor the training data.
[0118] Optionally, in other embodiments, the terminal can also calculate auxiliary loss data of the second predicted motion data based on specified constraints, and calculate the sum of the auxiliary loss data and the reconstruction loss data as the second potential loss data.
[0119] Optionally, after obtaining the first potential loss data and the second potential loss data, the terminal performs backpropagation to update the model parameters of the initial denoising module in the model to be trained according to the first potential loss data and the second potential loss data, so as to obtain the target denoising module.
[0120] Optionally, in one specific embodiment, in order to ensure that the motion data predicted by the model can be executed stably and accurately on the target robot, the terminal also performs reinforcement learning on the model to be trained, thereby encoding the dynamic constraints that can be safely executed on the real target robot into the policy optimization objective, so as to train the initial policy network in the model to be trained, and using the initial policy network in the model to be trained to generate predicted trajectory information based on the model training data and target training conditions.
[0121] Optionally, the initial policy network in the initial model is obtained, and this initial policy network is driven to generate a second noisy latent vector according to the model training data and the target training conditions. That is, as shown... Figure 5c As shown, Figure 5c This is a schematic diagram illustrating a scenario of training an initial policy network in the model training method provided in this application. Figure 5c In the illustrated embodiment, the terminal uses label information (e.g., label text and visual labels) from the model training data as the target training condition, and historical frame data from the model training data as the model input to the initial policy network. This drives the initial policy network to process the historical frame data and the target training condition according to the initialized policy parameters and the value network parameters to generate a second noisy latent vector. Optionally, in one specific embodiment, the initial policy network is defined as follows:
[0122] in, Let z be the initial policy network, and z be the second noise potential vector.
[0123] Optionally, after generating the second noise latent vector, the initial policy network uses the frozen target denoising module and frozen motion decoding module in the model to be trained to predict the second noise latent vector, obtaining multiple training trajectory data. That is, under the current policy, the initial policy network interacts with the model to be trained and the simulation environment using the second noise latent vector to obtain multiple training trajectory data, wherein the training trajectory data represents the joint motion trajectory information predicted by the model to be trained based on the second noise latent vector.
[0124] Optionally, after acquiring the training trajectory data, the terminal inputs the training trajectory data into the robot dynamics simulator to perform simulation prediction on the training trajectory data, obtaining predicted trajectory information representing the joint simulation state and the next simulation state for executing the training trajectory data. Based on the predicted trajectory information and the target trajectory constraint data, the terminal calculates the trajectory training information corresponding to each predicted trajectory. This trajectory training information represents the loss data during the training of the initial policy network.
[0125] That is, the terminal calculates the instant reward data based on the predicted trajectory information and target constraint data, and calculates the training advantage data and target reward data at each time step. It then uses the instant reward data, training advantage data, and target reward data to update the policy network parameters of the initial policy network. The resulting trajectory training information is shown below:
[0126]
[0127]
[0128] in, For trajectory training information, To reward data in a timely manner, The cut strategy loss data, For data loss in the value network, For policy entropy, and These are the policy network parameters and value network parameters of the initial policy network, respectively. and Calculate coefficients for the trajectory.
[0129] Optionally, the strategy loss data can be calculated as follows:
[0130] in, This represents the probability ratio.
[0131] Optionally, the value network loss data can be calculated as follows:
[0132] Optionally, the policy entropy can be calculated as follows:
[0133] Optionally, the terminal updates the initial model based on any one or more of the first predicted motion data, the second predicted motion data, and the predicted trajectory information to obtain a target processing model. Specifically, the terminal updates the parameters of the motion encoding and motion decoding modules using the acquired first predicted motion data and the corresponding first training loss data; updates the parameters of the initial denoising module using the acquired first and second potential loss data corresponding to the second predicted motion data; and updates the initial policy network using the acquired predicted trajectory information to generate a target processing model for interactive prediction of the target robot.
[0134] In this embodiment, the terminal acquires an initial model corresponding to the target robot and its corresponding model training data; drives the initial model to generate training latent features based on the model training data and the distribution data to be trained; drives the initial model to perform decoding processing based on the model training data and the training latent features to obtain first predicted motion data; and updates the initial model based on the model training data, the training latent features, and the first predicted motion data to obtain a target processing model. This achieves multi-dimensional training of the initial model associated with the target robot to generate a target processing model capable of accurately generating interactive motion data for controlling the collaborative interaction of various target hardware components of the target robot, thereby improving the interactivity and accuracy of the target robot.
[0135] To better implement the robot interaction method in the embodiments of this application, based on the robot interaction method, the embodiments of this application also provide a robot interaction device, such as... Figure 6 As shown, Figure 6 This is a schematic diagram of the structure of one embodiment of the robot interaction device provided in this application. Specifically, the robot interaction device 400 includes: The information acquisition module 401 is configured to respond to a robot interaction request for a target robot and acquire associated interaction data and interaction environment information corresponding to the robot interaction request. Feature processing module 402 is configured to drive the target processing model to generate interaction potential features based on the associated interaction data, interaction environment information and potential vectors to be processed; The multidimensional motion generation module 403 is configured to decode based on the historical motion frames in the associated interaction data and the interaction potential features to obtain the interaction motion data of the target robot, wherein the interaction motion data includes joint motion data and facial expression effect data; The collaborative interaction module 404 is configured to control the target hardware of the target robot to perform collaborative interaction processing according to the interactive motion data, so as to obtain the robot interaction result.
[0136] In one possible implementation of this embodiment, the robot interaction device drives the target processing model to generate interaction latent features based on the associated interaction data, interaction environment information, and the latent vector to be processed, including: Analyze the associated interaction data to determine the historical motion frames corresponding to the associated interaction data; The target policy network in the target processing model is used to generate a noise latent vector according to the associated interaction data, historical motion frames and interaction environment information. The noise potential vector is sampled based on the target sampling strategy to obtain sampled noise features; The sampling noise features are denoised using the interactive environment information to obtain the interactive potential features.
[0137] In one possible implementation of this embodiment, the robot interaction device decodes the historical motion frames and potential interaction features in the associated interaction data to obtain the interaction motion data of the target robot, including: Obtain the model decoding data corresponding to the target processing model; Based on the historical motion frames of the associated interaction data and the potential interaction features, the model decoding data is decoded to obtain the interactive motion data of the target robot.
[0138] In one possible implementation of this embodiment, the robot interaction device decodes the model decoding data based on the historical motion frames of the associated interaction data and the latent interaction features to obtain the interactive motion data of the target robot, including: Obtain the historical motion frames of the target robot from the associated interaction data, the historical motion frames including joint historical motion frames and facial expression historical motion frames; The joint motion data is obtained by decoding the joint historical motion frames, the interaction potential features, and the model decoding data using the first decoding module; The second decoding module is used to decode the expression history motion frames, the interaction potential features, and the model decoding data to obtain expression effect data; Interactive motion data is generated based on the joint motion data and facial expression effects data.
[0139] In one possible implementation of this embodiment, the robot interaction device coordinates the target hardware of the target robot to perform collaborative interaction processing according to the joint motion data and the facial expression effect data, and obtains the robot interaction result, including: The target joint hardware in the target hardware of the target robot is driven to perform the first interaction processing according to the joint motion data to obtain the action interaction result. The facial expression effect data drives the appearance display hardware in the target hardware of the target robot to perform a second interaction process to obtain the effect interaction result. The appearance display component is a hardware component used to display the appearance changes of the target robot during the interaction process. The robot interaction results are generated based on the action interaction results and the special effects interaction results.
[0140] In one possible implementation of this embodiment, the robot interaction device is further used for: Obtain the initial model corresponding to the target robot, and the model training data corresponding to the target robot; The initial model is driven to generate training latent features based on the model training data and the distribution data to be trained. The initial model is driven to perform decoding processing based on the model training data and the training latent features to obtain the first predicted motion data; The initial model is updated based on the model training data, the training latent features, and the first predicted motion data to obtain the target processing model.
[0141] In one possible implementation of this embodiment, the robot interaction device drives the initial model to generate training latent features based on the model training data and the distribution data to be trained, including: Obtain historical frame training data and supervised training data from the model training data; The motion coding module in the initial model is used to perform prediction training on the historical frame training data, supervised training data, and training distribution data to obtain the prediction distribution parameters. The predicted distribution parameters are sampled based on the training sampling strategy to obtain training latent features.
[0142] In one possible implementation of this embodiment, the robot interaction device performs decoding processing based on the model training data and the training latent features to obtain first predicted motion data, including: Obtain the training placeholder data for the initial model; The motion decoding module in the initial model is used to decode the training placeholder data based on the training latent features, historical frame training data, and training distribution data to obtain the first predicted motion data.
[0143] In one possible implementation of this embodiment, the robot interaction device updates the initial model based on the model training data, the training latent features, and the first predicted motion data to obtain a target processing model, including: Freeze the motion encoding module and motion decoding module in the initial model to obtain the model to be trained; The training data of the model and the first noise latent vector are used to generate second predicted motion data; The model to be trained is used to generate predicted trajectory information based on the model training data and target training conditions. The initial model is updated based on any one or more of the first predicted motion data, the second predicted motion data, and the predicted trajectory information to obtain the target processing model.
[0144] In one possible implementation of this embodiment, the robot interaction device uses the model to be trained to generate second predicted motion data from the model training data and the first noise latent vector, including: The historical frame training data and supervised training data in the model training data are encoded to obtain the potential features to be processed. The potential features to be processed are subjected to noise processing using preset noise data to obtain a first noise potential vector; The initial denoising module is driven to perform denoising prediction on the first noise potential vector to obtain the predicted potential features and the corresponding first potential loss data. The second predicted motion data is obtained by decoding the predicted latent features and historical frame training data using the frozen decoder in the model to be trained.
[0145] In one possible implementation of this embodiment, the robot interaction device uses the model to be trained to generate predicted trajectory information based on the model training data and target training conditions, including: Obtain the initial policy network in the initial model, and drive the initial policy network to generate a second noise latent vector according to the model training data and target training conditions; The second noise latent vector is predicted using the frozen target denoising module and the frozen motion decoding module in the model to be trained, thereby obtaining training trajectory data; The training trajectory data is simulated and predicted to obtain the predicted trajectory information.
[0146] In one possible implementation of this embodiment, the robot interaction device updates the initial model based on any one or more of the first predicted motion data, the second predicted motion data, and the predicted trajectory information to obtain a target processing model, including: The error is measured on the first predicted motion data to obtain the reconstruction loss data corresponding to the first predicted motion data. And / or, determine the divergence loss data of the first predicted motion data based on the training latent features and prior latent features corresponding to the first predicted motion data; And / or, determine auxiliary loss data for the first predicted motion data based on the first predicted motion data and preset constraints; Calculate the first training loss data corresponding to the first predicted motion data based on any one or more of the reconstructed loss data, divergence loss data, and auxiliary loss data and the coefficient of the first loss data. The initial model is updated using the first training loss data to obtain the target processing model.
[0147] In one possible implementation of this embodiment, the robot interaction device updates the initial model based on any one or more of the first predicted motion data, the second predicted motion data, and the predicted trajectory information to obtain a target processing model, including: Obtain the first potential loss data corresponding to the predicted potential features and the second potential loss data corresponding to the second predicted motion data; The initial model is updated using the first training loss data, the first potential loss data, and the second potential loss data corresponding to the first predicted motion data to obtain the target processing model.
[0148] In one possible implementation of this embodiment, the robot interaction device updates the initial model based on any one or more of the first predicted motion data, the second predicted motion data, and the predicted trajectory information to obtain a target processing model, including: Based on the predicted trajectory information and the target trajectory constraint data, determine the trajectory training information corresponding to the predicted trajectory information; The initial model is updated based on the trajectory training information to obtain the target processing model.
[0149] In one possible implementation of this embodiment, the robot interaction device acquires the model training data corresponding to the target robot, including: Obtain the initial motion data corresponding to the target robot; The initial motion data is augmented to obtain model training data.
[0150] In one possible implementation of this embodiment, the robot interaction device performs data augmentation on the initial motion data to obtain model training data, including: The target robot is subjected to posture recognition to obtain the target hardware corresponding to the target robot. Based on the target hardware and the initial motion data, the model training data is obtained by retargeting. And / or, perform data augmentation on the initial motion data based on a predictive data augmentation strategy to obtain model training data; And / or, generate motion data associated with the initial motion data based on the animation generation model, and redirect the initial motion data and the generated motion data with the target hardware to obtain model training data.
[0151] In this embodiment, the robot interaction device responds to a robot interaction request for the target robot, acquires the associated interaction data and interaction environment information corresponding to the robot interaction request, drives the target processing model to generate interaction potential features based on the associated interaction data, interaction environment information, and potential vectors to be processed, decodes the historical motion frames in the associated interaction data and the interaction potential features to obtain the interaction motion data of the target robot, the interaction motion data including joint motion data and facial expression effect data, and controls the target hardware of the target robot to perform collaborative interaction processing according to the interaction motion data to obtain the robot interaction result. During human-robot interaction with a target robot, a target processing model can generate potential interaction features representing a complete motion sequence based on associated interaction data and interaction environment data. By decoding historical motion frames of associated interaction data and these potential interaction features, interactive motion data guiding the target robot's hardware to perform coordinated movements can be generated. This interactive motion data includes joint motion data and facial expression effect data. The target robot is controlled to perform coordinated interactions according to the joint motion data and facial expression effect data in the interactive control data. During the interaction, the target robot can predict the joint motion and facial expression effect corresponding to the current interaction and coordinate the interaction, thereby improving the interactive flexibility and immersion of the target robot.
[0152] This invention also provides a robot interaction device, such as... Figure 7 As shown, Figure 7 This is a schematic diagram of one embodiment of the robot interaction device provided in this application.
[0153] The robot interaction device integrates any one of the robot interaction devices provided in the embodiments of the present invention, and the robot interaction device includes: One or more processors; Memory; and One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor in the steps of the robot interaction method described in any of the embodiments of the above robot interaction method.
[0154] Specifically, a robot interaction device may include components such as a processor 501 with one or more processing cores, a memory 502 with one or more computer-readable storage media, a power supply 503, and an input unit 504. Those skilled in the art will understand that... Figure 7 The robot interaction device structure shown does not constitute a limitation on the robot interaction device. It may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein: The processor 501 is the control center of the robot interaction device. It connects various parts of the robot interaction device via various interfaces and lines. By running or executing software programs and / or modules stored in the memory 502, and by calling data stored in the memory 502, it performs various functions and processes data of the robot interaction device, thereby providing overall monitoring of the robot interaction device. Optionally, the processor 501 may include one or more processing cores; preferably, the processor 501 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, and the modem processor mainly handles wireless communication. It is understood that the modem processor may not be integrated into the processor 501.
[0155] The memory 502 can be used to store software programs and modules. The processor 501 executes various functional applications and robot interactions by running the software programs and modules stored in the memory 502. The memory 502 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, at least one application program required for a function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created based on the use of the robot interaction device, etc. In addition, the memory 502 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 502 may also include a memory controller to provide the processor 501 with access to the memory 502.
[0156] The robot interaction device also includes a power supply 503 that supplies power to the various components. Preferably, the power supply 503 can be logically connected to the processor 501 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 503 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.
[0157] The robot interaction device may also include an input unit 504, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.
[0158] Although not shown, the robot interaction device may also include a display unit, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 501 in the robot interaction device loads the executable files corresponding to the processes of one or more applications into the memory 502 according to the following instructions, and the processor 501 runs the applications stored in the memory 502 to realize various functions, as follows: In response to a robot interaction request for a target robot, obtain the associated interaction data and interaction environment information corresponding to the robot interaction request; The target processing model generates interaction latent features based on the associated interaction data, interaction environment information, and potential vectors to be processed. Based on the historical motion frames in the associated interaction data and the potential interaction features, the interaction motion data of the target robot is obtained by decoding. The interaction motion data includes joint motion data and facial expression effect data. The target hardware of the target robot is controlled to perform collaborative interaction processing according to the interactive motion data to obtain the robot interaction result.
[0159] Therefore, embodiments of the present invention provide a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), a disk, or an optical disk, etc. A computer program is stored thereon, which is loaded by a processor to execute the steps in any of the robot interaction methods provided in the embodiments of the present invention. For example, the computer program loaded by the processor can execute the following steps: In response to a robot interaction request for a target robot, obtain the associated interaction data and interaction environment information corresponding to the robot interaction request; The target processing model generates interaction latent features based on the associated interaction data, interaction environment information, and potential vectors to be processed. Based on the historical motion frames in the associated interaction data and the potential interaction features, the interaction motion data of the target robot is obtained by decoding. The interaction motion data includes joint motion data and facial expression effect data. The target hardware of the target robot is controlled to perform collaborative interaction processing according to the interactive motion data to obtain the robot interaction result.
[0160] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the detailed descriptions of other embodiments above, which will not be repeated here.
[0161] In practice, each of the above units or structures can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units or structures, please refer to the previous method embodiments, which will not be repeated here.
[0162] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.
[0163] The robot interaction method provided by the embodiments of this application has been described in detail above. Specific embodiments have been used to illustrate the principle and implementation of the present invention. The description of the above embodiments is only for the purpose of helping to understand the method and core idea of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation and application scope based on the idea of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.
Claims
1. A robot interaction method, characterized in that, The robot interaction method includes: In response to a robot interaction request for a target robot, obtain the associated interaction data and interaction environment information corresponding to the robot interaction request; The target processing model generates latent interaction features based on the associated interaction data, interaction environment information, and noise latent vectors. Based on the historical motion frames in the associated interaction data and the potential interaction features, the interaction motion data of the target robot is obtained by decoding. The interaction motion data includes joint motion data and facial expression effect data. The target hardware of the target robot is controlled to perform collaborative interaction processing according to the interactive motion data to obtain the robot interaction result.
2. The robot interaction method according to claim 1, characterized in that, The driving target processing model generates latent interaction features based on the associated interaction data, interaction environment information, and noise latent vectors, including: Analyze the associated interaction data to determine the historical motion frames corresponding to the associated interaction data; The target policy network in the target processing model is used to generate a noise latent vector according to the associated interaction data, historical motion frames and interaction environment information. The noise potential vector is sampled to obtain sampled noise features; The sampling noise features are denoised using the interactive environment information to obtain the interactive potential features.
3. The robot interaction method according to claim 1, characterized in that, The step of decoding based on historical motion frames and latent interaction features in the associated interaction data to obtain the interaction motion data of the target robot includes: Obtain the model decoding data corresponding to the target processing model; Based on the historical motion frames of the associated interaction data and the potential interaction features, the model decoding data is decoded to obtain the interactive motion data of the target robot.
4. The robot interaction method according to claim 3, characterized in that, The process of decoding the model decoding data based on the historical motion frames of the associated interaction data and the latent interaction features to obtain the interactive motion data of the target robot includes: Obtain the historical motion frames of the target robot from the associated interaction data, the historical motion frames including joint historical motion frames and facial expression historical motion frames; The joint motion data is obtained by decoding the joint historical motion frames, the interaction potential features, and the model decoding data using the first decoding module; The second decoding module is used to decode the expression history motion frames, the interaction potential features, and the model decoding data to obtain expression effect data; Interactive motion data is generated based on the joint motion data and facial expression effects data.
5. The robot interaction method according to any one of claims 1-4, characterized in that, The process of collaboratively controlling the target hardware of the target robot according to the joint motion data and the facial expression effect data to perform collaborative interaction processing and obtain robot interaction results includes: The target joint hardware in the target hardware of the target robot is driven to perform the first interaction processing according to the joint motion data to obtain the action interaction result. The facial expression effect data drives the appearance display hardware in the target hardware of the target robot to perform a second interaction process to obtain the effect interaction result. The appearance display hardware is a hardware component used to display the appearance changes of the target robot during the interaction process. The robot interaction results are generated based on the action interaction results and the special effects interaction results.
6. A model training method, characterized in that, The model training method includes: Obtain the initial model corresponding to the target robot, and the model training data corresponding to the target robot; The initial model is driven to generate training latent features based on the model training data and the distribution data to be trained. The initial model is driven to perform decoding processing based on the model training data and the training latent features to obtain the first predicted motion data; The initial model is updated based on the model training data, the training latent features, and the first predicted motion data to obtain the target processing model.
7. The model training method according to claim 6, characterized in that, The process of driving the initial model to generate training latent features based on the model training data and the distribution data to be trained includes: Obtain historical frame training data and supervised training data from the model training data; The motion coding module in the initial model is used to perform prediction training on the historical frame training data, supervised training data, and training distribution data to obtain the prediction distribution parameters. The predicted distribution parameters are sampled based on the training sampling strategy to obtain training latent features.
8. The model training method according to claim 6, characterized in that, The decoding process based on the model training data and the training latent features to obtain the first predicted motion data includes: Obtain the training placeholder data for the initial model; The motion decoding module in the initial model is used to decode the training placeholder data based on the training latent features, historical frame training data, and training distribution data to obtain the first predicted motion data.
9. The model training method according to claim 6, characterized in that, The step of updating the initial model based on the model training data, the training latent features, and the first predicted motion data to obtain the target processing model includes: Freeze the motion encoding module and motion decoding module in the initial model to obtain the model to be trained; The training data of the model and the first noise latent vector are used to generate second predicted motion data; The model to be trained is used to generate predicted trajectory information based on the model training data and target training conditions; The initial model is updated based on any one or more of the first predicted motion data, the second predicted motion data, and the predicted trajectory information to obtain the target processing model.
10. The model training method according to claim 9, characterized in that, The step of generating second predicted motion data using the model training data and the first noise latent vector through the model to be trained includes: The historical frame training data and supervised training data in the model training data are encoded to obtain the potential features to be processed. The potential features to be processed are subjected to noise processing using preset noise data to obtain a first noise potential vector; The initial denoising module is driven to perform denoising prediction on the first noise potential vector to obtain the predicted potential features and the corresponding first potential loss data. The second predicted motion data is obtained by decoding the predicted latent features and historical frame training data using the frozen decoder in the model to be trained.
11. The model training method according to claim 9, characterized in that, The step of generating predicted trajectory information using the model to be trained based on the model training data and target training conditions includes: Obtain the initial policy network in the initial model, and drive the initial policy network to generate a second noise latent vector according to the model training data and target training conditions; The second noise latent vector is predicted using the frozen target denoising module and the frozen motion decoding module in the model to be trained, thereby obtaining training trajectory data; The training trajectory data is simulated and predicted to obtain the predicted trajectory information.
12. The model training method according to claim 9, characterized in that, The step of updating the initial model based on any one or more of the first predicted motion data, the second predicted motion data, and the predicted trajectory information to obtain the target processing model includes: The error is measured on the first predicted motion data to obtain the reconstruction loss data corresponding to the first predicted motion data. And / or, determine the divergence loss data of the first predicted motion data based on the training latent features and prior latent features corresponding to the first predicted motion data; And / or, determine auxiliary loss data for the first predicted motion data based on the first predicted motion data and preset constraints; Calculate the first training loss data corresponding to the first predicted motion data based on any one or more of the reconstructed loss data, divergence loss data, and auxiliary loss data and the coefficient of the first loss data. The initial model is updated using the first training loss data to obtain the target processing model.
13. The model training method according to claim 9, characterized in that, The step of updating the initial model based on any one or more of the first predicted motion data, the second predicted motion data, and the predicted trajectory information to obtain the target processing model includes: Obtain the first potential loss data corresponding to the predicted potential features and the second potential loss data corresponding to the second predicted motion data; The initial model is updated using the first training loss data, the first potential loss data, and the second potential loss data corresponding to the first predicted motion data to obtain the target processing model.
14. The model training method according to claim 9, characterized in that, The step of updating the initial model based on any one or more of the first predicted motion data, the second predicted motion data, and the predicted trajectory information to obtain the target processing model includes: Based on the predicted trajectory information and the target trajectory constraint data, determine the trajectory training information corresponding to the predicted trajectory information; The initial model is updated based on the trajectory training information to obtain the target processing model.
15. The model training method according to claim 6, characterized in that, The step of obtaining the model training data corresponding to the target robot includes: Obtain the initial motion data corresponding to the target robot; The initial motion data is augmented to obtain model training data.
16. The model training method according to claim 15, characterized in that, The step of augmenting the initial motion data to obtain model training data includes: The target robot is subjected to posture recognition to obtain the target hardware corresponding to the target robot. Based on the target hardware and the initial motion data, the model training data is obtained by retargeting. And / or, perform data augmentation on the initial motion data based on a predictive data augmentation strategy to obtain model training data; And / or, generate motion data associated with the initial motion data based on the animation generation model, and redirect the initial motion data and the generated motion data with the target hardware to obtain model training data.
17. A robot interaction device, characterized in that, The robot interaction device includes: The information acquisition module is configured to respond to a robot interaction request for a target robot and acquire the associated interaction data and interaction environment information corresponding to the robot interaction request. The feature processing module is configured to drive the target processing model to generate latent interaction features based on the associated interaction data, interaction environment information, and noise latent vectors; The multidimensional motion generation module is configured to decode the historical motion frames and the potential interaction features in the associated interaction data to obtain the interactive motion data of the target robot, wherein the interactive motion data includes joint motion data and facial expression effect data. The collaborative interaction module is configured to control the target hardware of the target robot to perform collaborative interaction processing according to the interactive motion data, so as to obtain the robot interaction result.
18. A robot interaction device, characterized in that, The robot interaction device includes: One or more processors; Memory; and One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the steps of the robot interaction method of any one of claims 1 to 5 and / or the steps of the model training method of any one of claims 6 to 16.
19. A computer-readable storage medium, characterized in that, It stores a computer program, which is loaded by a processor to execute the steps of the robot interaction method according to any one of claims 1 to 5 and / or the steps of the model training method according to any one of claims 6 to 16.