Self-adaptive bowing control method for violin playing robot
By combining a contact state encoder and a diffusion strategy network, adaptive bowing control of a violin playing robot was achieved, solving the problem that traditional methods cannot adapt to changes in string and bowing techniques in real time, thus improving the accuracy of control and the stability of performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- HARBIN INST OF TECH
- Filing Date
- 2026-04-02
- Publication Date
- 2026-05-12
AI Technical Summary
Existing bow control methods for violin playing robots cannot adapt to changes in string type or bowing technique in real time, resulting in poor control accuracy and difficulty in achieving high-quality continuous performance under multiple working conditions.
By combining a contact state encoder and a diffusion strategy network, an adaptive bowing control strategy is generated by training the contact state encoder and diffusion strategy network with expert performance data, and real-time adjustments are made using visual information and force perception.
It achieves adaptive control for different string types, bowing techniques, and dynamic conditions, improving the tone quality and environmental adaptability of the performance, reducing motion stuttering and sudden changes in dynamics, and enhancing the robustness and adaptability of the system.
Smart Images

Figure CN122008242A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent robot control. Specifically, it is an adaptive bowing control method for a violin playing robot based on a diffusion strategy and bow string contact encoding. Background Technology
[0002] As robotics technology continues to expand from industrial applications to service, education, and art, performing robots with artistic expression have become a frontier of interdisciplinary research. Among these, the violin, requiring the simulation of complex bowing techniques and precise dynamic control of the human arm, places extremely high demands on the robot's compliant operation and environmental interaction capabilities, making it an ideal platform for verifying fine motor skills.
[0003] The core operation of violin playing—bowing—is essentially a typical task involving rich and smooth contact. During performance, there is a continuous and varied interaction of contact forces between the bow and the strings. The precise coordination of bow pressure, bow speed, and the point of contact directly determines the tone quality. Different strings (G, D, A, E strings) exhibit significantly different mechanical properties due to differences in string diameter, tension, and material; different bowing techniques (split bowing, legato bowing, staccato bowing, spiccato, etc.) correspond to drastically different force-velocity-displacement coupling patterns. Therefore, bowing control strategies need to be able to sense the current bow-string contact state and adaptively adjust the bowing trajectory and pressure accordingly.
[0004] However, existing bow control methods for violin playing robots have significant technical flaws and are difficult to meet actual performance needs. Traditional PID control sets a target bow pressure / speed and uses proportional-integral-derivative feedback loops to correct joint torque in real time, ensuring the actual contact force tracks the target value. Traditional impedance control strategies set desired stiffness and damping parameters to make the robotic arm's end effector exhibit a compliant response when contacting the strings. However, both of these traditional control methods require manual pre-tuning of control parameters, including PID gain, target bow pressure, desired stiffness, and damping. These parameters need to be manually tuned for specific string types, bowing techniques, and force conditions. Therefore, a set of parameters can only adapt to a specific playing condition (including: string type (different diameters, tensions, and materials of G / D / A / E strings), bowing technique (different requirements for bow pressure and speed for separate bowing, legato bowing, staccato bowing, spiccato, etc.), dynamic level (different requirements for bow pressure for soft and loud playing) etc.). When the string type or bowing technique changes, the original parameters are no longer applicable and need to be manually readjusted to achieve accurate control. Therefore, fixed parameter combinations cannot be adapted in real time during performance. However, in actual performance, the switching of strings and bowing techniques is frequent and continuous, and traditional methods cannot automatically adapt to these changes during performance. Adjusting parameters by modeling is also extremely difficult because there are nonlinear frictions and flexible deformation of bow hair during bow-string contact, which often leads to inaccurate parameter adjustments.
[0005] In recent years, imitation learning has made significant progress in the field of robotics manipulation. Among them, diffusion policy, which models the probability distribution of action sequences through conditional diffusion models, has demonstrated powerful multimodal action distribution modeling capabilities and generalization performance in contact-rich manipulation tasks. However, diffusion policy has not yet been applied in fine manipulation scenarios requiring force perception and compliant control, such as music performance robots.
[0006] In summary, existing bow control methods for violin playing robots suffer from problems such as difficult modeling, fixed parameters, and lack of adaptive capabilities. These methods cannot adapt to the dynamic changes in the bow string contact state, resulting in poor control accuracy and difficulty in achieving high-quality continuous performance under multiple working conditions. There is an urgent need for a new adaptive bow control method to solve the above-mentioned technical pain points. Summary of the Invention
[0007] The purpose of this invention is to solve the problem of poor control accuracy in existing bow control methods for violin playing robots, and to propose an adaptive bow control method for violin playing robots.
[0008] An adaptive bowing control method for a violin playing robot, wherein the violin playing robot has an anthropomorphic configuration and is used to simulate a human holding and playing a violin;
[0009] The control method includes the following:
[0010] Step 1: Collect historical action samples of expert performances, extract visual information, three-axis horizontal force during bowing, three-axis rotational torque during bowing, arm joint angle, arm end position and arm end posture at the same moment in the sample, obtain contact state latent features based on the three-axis horizontal force and three-axis rotational torque at each moment, and label the string type and bowing technique type at each moment.
[0011] The three-axis horizontal force and three-axis rotational torque during the bowing process at each moment are used as the input data of sample 1, and the string type and bowing technique type label at that moment are used as the output data of sample 1.
[0012] The contact state latent features, three-axis horizontal force, three-axis rotational torque, arm end position, arm end posture, arm joint angle, and visual information at each moment are used as input data for one sample No. 2, and the arm joint angles at multiple consecutive future moments are used as output data for one sample No. 2.
[0013] Step 2: Train the contact state encoder using the input and output data of sample 1 to obtain the pre-trained contact state encoder.
[0014] The diffusion policy network is trained using the input and output data of sample number 2 to obtain the pre-trained diffusion policy network.
[0015] Step 3, in the At any moment, collect three-axis horizontal force, three-axis rotational torque, arm end position, arm end posture, arm joint angle and visual information during the bowing process of the robot playing music;
[0016] The collected triaxial horizontal force and triaxial rotational torque are input into the pre-trained contact state encoder to predict the string type and bowing technique. The latent contact state features generated by the contact state encoder during the prediction process are extracted. The initial value is 1;
[0017] The contact state hidden features, and in the first The three-axis horizontal force, three-axis rotational torque, arm end position, arm end posture, arm joint angle, and visual information collected at all times are input into the pre-trained diffusion strategy network to predict the robot's bowing action sequence at multiple consecutive moments in the future. The bowing action sequence is composed of the joint angles of the robotic arm at multiple consecutive moments.
[0018] Step 4: Generate control commands based on the robot's bowing action sequence predicted for future moments, and adjust the robot's playing actions at the corresponding moments;
[0019] Step 5, let = +1, collect three-axis horizontal force, three-axis rotational torque, arm end position, arm end posture, arm joint angle and visual information during the bowing process of the robot's performance;
[0020] The collected triaxial horizontal force and triaxial rotational torque are input into the pre-trained contact state encoder to predict the string type and bowing method, and the contact state latent features generated by the contact state encoder during the prediction process are extracted.
[0021] The contact state hidden features, and in the first The three-axis horizontal force, three-axis rotational torque, arm end position, arm end posture, arm joint angle, and visual information collected at all times are input into the pre-trained diffusion strategy network to predict the robot's bowing action sequence at multiple consecutive moments in the future. The bowing action sequence is composed of the joint angles of the robotic arm at multiple consecutive moments.
[0022] Extract the first The bowing action at time +1, and extract the first bowing action from the bowing action sequence predicted from the previous time. The bowing action at time +1 will extract all the first... The bowing actions at time +1 are weighted and fused to obtain the fused result. At time +1, the bowing motion is used to generate control commands, which adjust the robot's playing motion at that moment.
[0023] Step 6, Judgment Is it equal to the preset time? If yes, stop; if no, proceed to step 5.
[0024] The beneficial effects of this invention are:
[0025] This invention thoroughly solves the technical pain points of traditional control methods by focusing on four core aspects: perception accuracy, adaptability, motion smoothness, and scene robustness. It achieves adaptive bow control for a violin-playing robot, as detailed below:
[0026] This invention extracts features from the timing force / torque signals during bowing, enabling the extraction of stable state representations from the bowstring contact process. This provides a contact perception basis for subsequent strategy generation, allowing the control method to differentiate and respond to different contact states. For example, it can accurately distinguish the contact features of different strings (G, D, A, E) and different bowing techniques such as separate bowing, legato bowing, staccato bowing, and spiccato bowing. The accuracy of contact state recognition is greatly improved, providing a reliable perception basis for bowing control and achieving precise bowstring contact state perception.
[0027] This invention integrates multimodal observations such as force, torque, and vision, enabling the strategy network to make decisions simultaneously using the robotic arm's motion state, bow string contact state, and spatial position information. This improves the bow control's adaptability to changes in working conditions and positional deviations. Without the need for manual parameter readjustment, it can adaptively match the working conditions of different strings, bowing techniques, and playing strengths, achieving adaptive control during the performance process.
[0028] This invention employs a conditional diffusion strategy to generate multi-step bowing action sequences, enabling sequence modeling of complex and continuous bowing actions. It generates multi-step control actions that match the current observation under different string types and bowing techniques, improving the flexibility and adaptability of action generation. It is suitable for bowing control tasks with complex contact relationships and high requirements for action continuity.
[0029] This invention establishes a teaching data acquisition method for different string types, bowing techniques, and force conditions, enabling training samples to cover multiple performance conditions. This provides a unified data foundation for contact state encoder training and strategy network training, thereby improving the applicability of the method to bowing tasks under multiple conditions.
[0030] This invention employs a phased training process that combines pre-training of the contact state encoder with training of the diffusion strategy network. This facilitates obtaining stable contact state feature representations first, and then completing action strategy learning on this basis, thereby improving the training feasibility and feature utilization efficiency of the entire control method and enhancing robustness.
[0031] This invention employs an online motion output method that combines rolling execution with weighted smoothing, which can alleviate the problem of sudden motion changes between adjacent predicted sequences, improve the continuity and smoothness of the bowing trajectory, make the online execution process of the robotic arm more stable, avoid problems such as motion stuttering and sudden changes in force that are prone to occur in traditional control, and ensure the quality of the performance tone.
[0032] This invention acquires visual information to obtain the spatial relative position information between the bow and the violin. Based on an image acquisition device, it obtains environmental images and, through target detection and depth information fusion, extracts the relative positional deviation between the bow and strings, the spatial position of the contact point, and posture-related visual features, which are then used as visual observation inputs to a policy network. This visual information can provide spatial correction references, enabling accurate adjustment of bowing movements even in scenarios with non-fixed supports, target position offsets, and slight environmental disturbances. This improves environmental adaptability and anti-interference capabilities, breaking away from the traditional reliance on fixed scenes; thus enhancing the system's adaptability to environmental disturbances and changes in target position. Attached Figure Description
[0033] Figure 1 A diagram illustrating the overall system framework of an adaptive bowing control method for a violin playing robot.
[0034] Figure 2 This is a schematic diagram of the network structure of an encoder for bowstring contact state.
[0035] Figure 3 This is a schematic diagram of the teaching data collection and two-stage training process.
[0036] Figure 4 This is a schematic diagram of the training and inference process of a diffusion strategy network. Detailed Implementation
[0037] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.
[0038] It should be noted that, unless otherwise specified, the embodiments and features described in the present invention can be combined with each other. The present invention will be further described below with reference to the accompanying drawings and specific embodiments, but this is not intended to limit the scope of the invention.
[0039] Example:
[0040] This embodiment can adaptively adjust the bowing action under different string types, bowing techniques, force conditions, and target position deviations, thereby maintaining the continuity, stability, and environmental adaptability of the bowing trajectory.
[0041] Combination Figure 1 and Figure 3 This embodiment describes an adaptive bowing control method for a violin playing robot. The violin playing robot has an anthropomorphic configuration and is used to simulate a human holding and playing a violin.
[0042] The control method includes the following:
[0043] Step 1: Collect historical action samples of expert performances, extract visual information, three-axis horizontal force during bowing, three-axis rotational torque during bowing, arm joint angle, arm end position and arm end posture at the same moment in the sample, obtain contact state latent features based on the three-axis horizontal force and three-axis rotational torque at each moment, and label the string type and bowing technique type at each moment.
[0044] The three-axis horizontal force and three-axis rotational torque during the bowing process at each moment are used as the input data of sample 1, and the string type and bowing technique type label at that moment are used as the output data of sample 1.
[0045] The contact state latent features, three-axis horizontal force, three-axis rotational torque, arm end position, arm end posture, arm joint angle, and visual information at each moment are used as input data for one sample No. 2, and the arm joint angles at multiple consecutive future moments are used as output data for one sample No. 2.
[0046] Step 2: Train the contact state encoder using the input and output data of sample 1 to obtain the pre-trained contact state encoder.
[0047] The diffusion policy network is trained using the input and output data of sample number 2 to obtain the pre-trained diffusion policy network.
[0048] Step 3, in the At any moment, collect three-axis horizontal force, three-axis rotational torque, arm end position, arm end posture, arm joint angle and visual information during the bowing process of the robot playing music;
[0049] The collected triaxial horizontal force and triaxial rotational torque are input into the pre-trained contact state encoder to predict the string type and bowing technique. The latent contact state features generated by the contact state encoder during the prediction process are extracted. The initial value is 1;
[0050] The contact state hidden features, and in the first The three-axis horizontal force, three-axis rotational torque, arm end position, arm end posture, arm joint angle, and visual information collected at all times are input into the pre-trained diffusion strategy network to predict the robot's bowing action sequence at multiple consecutive moments in the future. The bowing action sequence is composed of the joint angles of the robotic arm at multiple consecutive moments.
[0051] Step 4: Generate control commands based on the robot's bowing action sequence predicted for future moments, and adjust the robot's playing actions at the corresponding moments;
[0052] Step 5, let = +1, collect three-axis horizontal force, three-axis rotational torque, arm end position, arm end posture, arm joint angle and visual information during the bowing process of the robot's performance;
[0053] The collected triaxial horizontal force and triaxial rotational torque are input into the pre-trained contact state encoder to predict the string type and bowing method, and the contact state latent features generated by the contact state encoder during the prediction process are extracted.
[0054] The contact state hidden features, and in the first The three-axis horizontal force, three-axis rotational torque, arm end position, arm end posture, arm joint angle, and visual information collected at all times are input into the pre-trained diffusion strategy network to predict the robot's bowing action sequence at multiple consecutive moments in the future. The bowing action sequence is composed of the joint angles of the robotic arm at multiple consecutive moments.
[0055] Extract the first The bowing action at time +1, and extract the first bowing action from the bowing action sequence predicted from the previous time. The bowing action at time +1 will extract all the first... The bowing actions at time +1 are weighted and fused to obtain the fused result. The bowing action at time +1 generates control commands to adjust the robot's performance at that time.
[0056] Step 6, Judgment Is it equal to the preset time? If yes, stop; if no, proceed to step 5.
[0057] Further, a six-dimensional force / torque sensor is used to collect the three-axis horizontal force and the three-axis rotational torque during the bowing process.
[0058] To further refine the approach, a depth camera is used to collect visual information.
[0059] Further specifying, the fused first The bowing action at time +1 is represented as:
[0060] ,
[0061] In the formula, For the post-fusion moment +1 to bowing action For the first The number of predictions that contribute at time +1 For the first The weight of each prediction result, For the first The corresponding prediction Momentary actions , This is the smoothing coefficient.
[0062] Specifically, this embodiment determines the joint angles for multiple consecutive future moments based on motion information such as the joint angle at the current moment. That is, when the robot plays a note, it predicts multiple future motion curves.
[0063] Example of weighted fusion in step 6: Assuming a control period of 10ms, each prediction generates a sequence of actions for the next 5 steps (covering times t, t+1, t+2, t+3, t+4); only the first 2 steps are executed each time, and then the timeline is re-predicted using new observations.
[0064] 1. First prediction (based on t0 observation): Generate action execution sequence: , , , , ,implement: , ;
[0065] 2. Second prediction (based on t2 observation): Generate action execution sequence: , , , , ,implement: , ;(here , (When it appears in both the first and second predictions, it's an overlapping action moment).
[0066] 3. Third prediction (based on t4 observation): Generate action execution sequence: , , , , ,implement: , ;
[0067] Example of a repetition time:
[0068] : The first prediction ( ) and the second prediction ( Coverage, i.e., overlap;
[0069] : The first prediction ( ) and the second prediction ( Coverage, i.e., overlap;
[0070] : Covered by the 1st, 2nd, and 3rd predictions, i.e., overlapping;
[0071] The final instruction is obtained by averaging the multiple predicted actions at overlapping times according to their weights.
[0072] The execution system of the violin-playing robot consists of the following parts:
[0073] (1) Multi-degree-of-freedom robotic arm: It adopts an anthropomorphic configuration design and has multiple rotational joints to simulate the human body's "torso-upper arm-forearm-wrist" movement mechanism.
[0074] (2) End bow clamping mechanism: It adopts a semi-enclosed design and is made of 3D printed engineering plastic. The contact surface is made of elastic material with a high coefficient of friction to ensure the clamping stability of the bow during bowing. At the same time, it reserves space for bow hair vibration so as not to affect the acoustic effect of the performance.
[0075] (3) Six-dimensional force / torque sensor: installed between the end of the robotic arm and the bow clamping mechanism, it can measure force and torque along three orthogonal axes and collect the bow string contact force signal in real time during the bowing process.
[0076] (4) Depth camera: A depth camera is used to acquire the three-dimensional spatial coordinates and posture information of the violin body, strings and bow in real time.
[0077] (5) Control and communication module: used to realize data interaction between the vision perception module, force perception module, target detection module and robotic arm control module, and to support the transmission of control commands between the strategy network and the robotic arm controller.
[0078] Analyze the correspondence between the bowstring contact state and the force sensory signal:
[0079] During bowing, there is a continuous and changing contact between the bow and the strings. Differences in the contact state under different string types, bowing techniques, and bowing conditions are directly reflected in the force / torque signals acquired by the robotic arm. Specifically, differences in string diameter, tension, and material will lead to different force responses during bow-string contact; different bowing techniques will result in different bow pressure changes, relative motion patterns, and contact durations, which will also exhibit different waveform characteristics in the force / torque timing signals.
[0080] Because the end effector's six-dimensional force / torque sensor can simultaneously measure force and torque in three directions, it can comprehensively reflect information such as normal contact, tangential friction, and attitude disturbances during bowing. The six-dimensional force / torque timing signal can serve as important sensing information characterizing the bowstring contact state and as input to the subsequent bowstring contact state encoder.
[0081] The following is combined Figure 2 The structure of the contact state encoder is described. The contact state encoder includes a convolutional feature extraction layer, a temporal attention layer, a global average pooling layer, a fully connected mapping layer, a fully connected layer, and an activation function layer.
[0082] The convolutional feature extraction layer is used to extract the local temporal features of the horizontal force and the rotational torque during the bowing process, and send them to the temporal attention layer;
[0083] The temporal attention layer is used to assign different weights to the extracted local temporal features and then send the weighted local temporal features to the global average pooling.
[0084] Global average pooling is used to reduce the dimensionality of weighted local temporal features and fuse them before sending them to the fully connected mapping layer.
[0085] The fully connected mapping layer is used to map the fused features into low-dimensional contact state latent features, which are then sent to the fully connected layer.
[0086] The fully connected layer is used to map the latent features of the contact state into a multi-dimensional vector, which is then sent to the activation function layer.
[0087] The activation function layer is used to normalize the multidimensional vector into a probability distribution and output the prediction results of string type and bowing technique.
[0088] Further specifying, the latent features of the contact state are represented as:
[0089] ,
[0090] In the formula, The parameter is The encoder for bowstring contact state. Indicates time The corresponding contact state latent feature vector The dimension representing the latent features of the contact state. For time The three-axis horizontal force and three-axis rotational torque are within the time-series observation window at the end.
[0091] Specifically, according to Figure 2 Explain the input and output data used during the training and prediction of the contact state encoder:
[0092] The bowstring contact state encoder is used to extract low-dimensional latent feature representations of the bowstring contact state from a six-dimensional force / torque time-series signal. Therefore, the input data used during training and prediction is force / torque, and the time interval is set to... The collected single-step force / torque observations are as follows:
[0093] (1)
[0094] In the formula, There are three horizontal forces. There are three rotational torques.
[0095] The time series input window is constructed using the observations from the most recent L times:
[0096] (2)
[0097] in, Indicated by time For the end of the time series observation window, Indicates the length of the timing window.
[0098] The timing input window is input into the bowstring contact state encoder to obtain the contact state hidden features:
[0099] (3)
[0100] in, The parameter is The encoder for bowstring contact state. Indicates time The corresponding contact state latent feature vector The dimension represents the latent features of the contact state. The encoder can employ a structure combining a one-dimensional convolutional neural network and a temporal attention mechanism to extract local temporal patterns and long-range dependent features from the force / torque signal.
[0101] To enable the encoder to distinguish between different string types and bowing techniques, a classification head is connected to the encoder output during the pre-training phase to obtain contact state prediction results.
[0102] (4)
[0103] in, This represents the weight matrix of the classification heads. This represents the bias vector of the classification head. This indicates the contact state prediction result. This contact state prediction result refers to a probability density; the one with the highest probability is selected as the label for string type and bowing technique type.
[0104] The encoder and classifier head are trained under supervised conditions using the cross-entropy loss function.
[0105] (5)
[0106] in, Indicates the total number of contact status categories. Indicates the true category label, This represents the corresponding predicted probability.
[0107] The encoder parameters are obtained by minimizing the classification loss function:
[0108] (6)
[0109] After the encoder pre-training is complete, the classification head is removed, and the encoder backbone network is retained as the contact state feature extractor. During subsequent policy network training and online execution, the latent features of the contact states are output. .
[0110] Therefore, the training and prediction process of the contact state encoder is as follows: Force / torque time-series signals and corresponding contact state labels are collected from the teaching data, and the bowstring contact state encoder is trained under supervision. During training, the encoder parameters are optimized by minimizing the classification loss function, enabling the encoder to learn the temporal force perception feature representations corresponding to different string types, bowing techniques, and contact state changes. After training, the classification head is removed, and the encoder backbone network is retained as the contact state feature extractor, outputting the latent contact state features during subsequent diffusion strategy network training and online execution phases.
[0111] Further defining the steps, the relationship between step 1 (using the arm joint angles at multiple consecutive future moments as the output data of sample 2) and step 2 (using the input and output data of sample 2 to train the diffusion strategy network and obtain the pre-trained diffusion strategy network) includes: adding Gaussian noise to the arm joint angles at multiple consecutive future moments to obtain the noisy action sequence.
[0112] Further specifying, after predicting the robot's bowing action sequence at multiple consecutive future moments in steps 3 and 5, the following steps are included: denoising the predicted robot bowing action sequence at multiple consecutive future moments to obtain a clean action sequence.
[0113] Further specifying, the noisy action sequence Represented as:
[0114] ,
[0115] In the formula, Indicates the number of diffusion steps. Indicates the first The cumulative noise scheduling coefficient corresponding to the step, This represents standard Gaussian noise of the same dimension as the action sequence. It is a sequence composed of the arm joint angles extracted from multiple consecutive future moments. It is an identity matrix.
[0116] Specifically, the following is combined with Figure 4 Introducing the training and online prediction process of the diffusion strategy network:
[0117] Expert performance demonstration data was collected using a teach-and-playback method. For different string types, bowing techniques, and force conditions, the operator guided the robotic arm's end effector to perform bowing movements, simultaneously recording the robotic arm's joint angles, end effector position, end effector posture, six-dimensional force / torque signals, and visual perception information. For each set of teaching data, the corresponding string type and bowing technique were labeled for pre-training of the bow-string contact state encoder.
[0118] Furthermore, the observation sequence, action sequence, and contact state labels are combined to form a teaching dataset:
[0119] (7)
[0120] in, Indicates the total number of teaching samples. This represents the observation sequence corresponding to the m-th sample group. This represents the corresponding action sequence. This indicates the corresponding contact status label.
[0121] Define the observation space and the action space:
[0122] This invention fuses the robot arm's body state, end-effector pose information, force perception information, contact state latent features, and visual features to form conditional observations within a policy network. A single-moment observation can be represented as:
[0123] (8)
[0124] in, Indicates the joint angle of the robotic arm. Indicates the end position. Indicates the end-effector attitude. Indicates triaxial force. Indicates triaxial torque. This represents the hidden features output by the contact state encoder. This represents the spatial characteristics output by the visual perception module.
[0125] Furthermore, with the most recent The observations at each time point constitute the input observation window of the strategy network:
[0126] (9)
[0127] in, Indicates the length of the observation window.
[0128] The motion space is defined as the sequence of target joint angles of the robotic arm at the current moment and several subsequent moments. A single step can be represented as:
[0129] (10)
[0130] in, Indicates time Output the target joint angle command.
[0131] The corresponding action sequence can be represented as:
[0132] (11)
[0133] in, Indicates the number of degrees of freedom of the robotic arm. This represents the length of the action sequence predicted by the policy network. This represents the target action sequence without added noise.
[0134] Training the diffusion policy network: The diffusion policy network uses an observation window As a condition, the expert demonstration sequence of actions Conditional generation modeling is performed. During the training phase, Gaussian noise is gradually added to the action sequences to obtain noisy action sequences corresponding to different diffusion steps:
[0135] (12)
[0136] in, Indicates the number of diffusion steps. This represents the cumulative noise scheduling coefficient corresponding to step k. This represents standard Gaussian noise of the same dimension as the action sequence. It is an identity matrix.
[0137] Subsequently, a noise prediction network was constructed. The network predicts the added noise using the current noise action sequence, the diffusion step, and the observation window as input. The training objective is to minimize the noise prediction error.
[0138] (13)
[0139] in, The parameter is A noise prediction network is used to predict the added noise based on the noise action sequence, the number of diffusion steps, and the observation window.
[0140] The diffusion policy network parameters are obtained by minimizing the loss function:
[0141] (14)
[0142] After training, the diffusion strategy network is able to generate a bowing action sequence that is adapted to the current bowstring contact state, given the current multimodal observation conditions.
[0143] While the invention has been described herein with reference to specific embodiments, it should be understood that these embodiments are merely examples of the principles and applications of the invention. Therefore, it should be understood that many modifications can be made to the exemplary embodiments, and other arrangements can be designed without departing from the spirit and scope of the invention as defined by the appended claims. It should be understood that different dependent claims and features described herein can be combined in ways different from those described in the original claims. It is also understood that features described in conjunction with individual embodiments can be used in other described embodiments.
Claims
1. An adaptive bowing control method for a violin playing robot, wherein the violin playing robot has an anthropomorphic configuration and is used to simulate a human holding and playing a violin; Its features are, The control method includes the following: Step 1: Collect historical action samples of expert performances, extract visual information, three-axis horizontal force during bowing, three-axis rotational torque during bowing, arm joint angle, arm end position and arm end posture at the same moment in the sample, obtain contact state latent features based on the three-axis horizontal force and three-axis rotational torque at each moment, and label the string type and bowing technique type at each moment. The three-axis horizontal force and three-axis rotational torque during the bowing process at each moment are used as the input data of sample 1, and the string type and bowing technique type label at that moment are used as the output data of sample 1. The contact state latent features, three-axis horizontal force, three-axis rotational torque, arm end position, arm end posture, arm joint angle, and visual information at each moment are used as input data for one sample No. 2, and the arm joint angles at multiple consecutive future moments are used as output data for one sample No.
2. Step 2: Train the contact state encoder using the input and output data of sample 1 to obtain the pre-trained contact state encoder. The diffusion policy network is trained using the input and output data of sample number 2 to obtain the pre-trained diffusion policy network. Step 3, in the At any moment, collect three-axis horizontal force, three-axis rotational torque, arm end position, arm end posture, arm joint angle and visual information during the bowing process of the robot playing music; The collected triaxial horizontal force and triaxial rotational torque are input into the pre-trained contact state encoder to predict the string type and bowing technique. The latent contact state features generated by the contact state encoder during the prediction process are extracted. The initial value is 1; The contact state hidden features, and in the first The three-axis horizontal force, three-axis rotational torque, arm end position, arm end posture, arm joint angle, and visual information collected at all times are input into the pre-trained diffusion strategy network to predict the robot's bowing action sequence at multiple consecutive moments in the future. The bowing action sequence is composed of the joint angles of the robotic arm at multiple consecutive moments. Step 4: Generate control commands based on the robot's bowing action sequence predicted for future moments, and adjust the robot's playing actions at the corresponding moments; Step 5, let = +1, collect three-axis horizontal force, three-axis rotational torque, arm end position, arm end posture, arm joint angle and visual information during the bowing process of the robot's performance; The collected triaxial horizontal force and triaxial rotational torque are input into the pre-trained contact state encoder to predict the string type and bowing method, and the contact state latent features generated by the contact state encoder during the prediction process are extracted. The contact state hidden features, and in the first The three-axis horizontal force, three-axis rotational torque, arm end position, arm end posture, arm joint angle, and visual information collected at all times are input into the pre-trained diffusion strategy network to predict the robot's bowing action sequence at multiple consecutive moments in the future. The bowing action sequence is composed of the joint angles of the robotic arm at multiple consecutive moments. Extract the first The bowing action at time +1, and extract the first bowing action from the bowing action sequence predicted from the previous time. The bowing action at time +1 will extract all the first... The bowing actions at time +1 are weighted and fused to obtain the fused result. The bowing action at time +1 generates control commands to adjust the robot's performance at that time. Step 6, Judgment Is it equal to the preset time? If yes, stop; if no, proceed to step 5.
2. The adaptive bowing control method for a violin playing robot according to claim 1, characterized in that, The fusion of the first The bowing action at time +1 is represented as: , In the formula, For the post-fusion moment +1 to the bowing action For the first The number of predictions that contribute at time +1 For the first The weight of each prediction result, For the first The corresponding prediction Momentary actions , This is the smoothing coefficient.
3. The adaptive bowing control method for a violin playing robot according to claim 1, characterized in that, The contact state encoder includes a convolutional feature extraction layer, a temporal attention layer, a global average pooling layer, a fully connected mapping layer, a fully connected layer, and an activation function layer; The convolutional feature extraction layer is used to extract the local temporal features of the horizontal force and the rotational torque during the bowing process, and send them to the temporal attention layer; The temporal attention layer is used to assign different weights to the extracted local temporal features and then send the weighted local temporal features to the global average pooling. Global average pooling is used to reduce the dimensionality of weighted local temporal features and fuse them before sending them to the fully connected mapping layer. The fully connected mapping layer is used to map the fused features into low-dimensional contact state latent features, which are then sent to the fully connected layer. The fully connected layer is used to map the latent features of the contact state into a multi-dimensional vector, which is then sent to the activation function layer. The activation function layer is used to normalize the multidimensional vector into a probability distribution and output the prediction results of string type and bowing technique.
4. The adaptive bowing control method for a violin playing robot according to claim 1, characterized in that, The latent features of the contact state are represented as follows: , In the formula, The parameter is The encoder for bowstring contact state. Indicates time The corresponding contact state latent feature vector The dimension representing the latent features of the contact state. For time The three-axis horizontal force and three-axis rotational torque are within the time-series observation window at the end.
5. The adaptive bowing control method for a violin playing robot according to claim 1, characterized in that, The process between step 1, which uses the arm joint angles at multiple consecutive future moments as the output data of sample 2, and step 2, which uses the input and output data of sample 2 to train the diffusion strategy network and obtain the pre-trained diffusion strategy network, includes: adding Gaussian noise to the arm joint angles at multiple consecutive future moments to obtain the noisy action sequence.
6. The adaptive bowing control method for a violin playing robot according to claim 5, characterized in that, After predicting the robot's bowing motion sequence for multiple consecutive future moments in steps 3 and 5, the following steps are performed: denoising the predicted robot bowing motion sequence for multiple consecutive future moments to obtain a clean motion sequence.
7. The adaptive bowing control method for a violin playing robot according to claim 5, characterized in that, Noisy action sequences Represented as: , In the formula, Indicates the number of diffusion steps. Indicates the first The cumulative noise scheduling coefficient corresponding to the step, This represents standard Gaussian noise of the same dimension as the action sequence. It is a sequence consisting of arm joint angles at multiple consecutive future moments. It is an identity matrix.
8. The adaptive bowing control method for a violin playing robot according to claim 1, characterized in that, A six-dimensional force / torque sensor was used to collect the three-axis horizontal force and the three-axis rotational torque during the bow movement process.
9. The adaptive bowing control method for a violin playing robot according to claim 1, characterized in that, Visual information is acquired using a depth camera.