Multi-degree-of-freedom double-machine cooperative cloth bag play system
Patent Information
- Application Number
- CN202610900678.8
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-22
- Publication Date
- 2026-09-29
AI Technical Summary
[0004]然而,现有的相关技术方案在实际应用中仍存在明显不足
[0005]为了解决现有技术中所存在的上述问题,本发明提供了一种多自由度双机协同布袋戏表演系统。
Smart Images

Figure CN122837174A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of intelligent performance and robot collaboration technology, specifically to a multi-degree-of-freedom dual-machine collaborative puppet show performance system. Background Technology
[0002] Traditional glove puppetry relies on puppeteers manually manipulating the puppets to perform actions, and its quality is highly dependent on the puppeteer's skill and improvisation. This manual puppetry method has technical limitations, including poor performance consistency, difficulty in maintaining performance for extended periods, and the inability to simultaneously control multiple puppets to perform complex coordinated movements. Therefore, how to replace manual puppetry with automation and intelligent technologies to achieve accurate reproduction and flexible control of traditional glove puppetry performance actions has become a key research focus in related technical fields.
[0003] Currently, some technological explorations have been undertaken in the digital preservation and display of traditional performing arts. In terms of motion recording and reproduction, motion capture technology based on optical or inertial sensors exists, which can be used to record the movement trajectories of people or objects. Regarding enhanced experiences, methods exist for generating virtual puppets and rendering interactive elements using Augmented Reality (AR) or Virtual Reality (VR) technologies. These technologies offer some degree of possibility for the digital transformation or auxiliary display of traditional performance forms.
[0004] However, existing related technical solutions still have significant shortcomings in practical applications. First, existing interactive experience methods are mostly limited to the presentation of virtual avatars, or only support one-way, simple motion synchronization, making it difficult to achieve real-time, coordinated performance interaction between puppets, thus limiting interactivity and immersion. Second, existing systems that can achieve some complex functions are often structurally complex and costly, which is not conducive to deployment and promotion in scenarios that require widespread application, such as museums, campuses, and cultural venues. Summary of the Invention
[0005] To address the aforementioned problems in the existing technology, this invention provides a multi-degree-of-freedom dual-machine collaborative puppet show performance system.
[0006] The technical problem to be solved by this invention is achieved through the following technical solution: This invention provides a multi-degree-of-freedom dual-machine collaborative puppet show performance system, comprising: a perception layer, a control layer, an execution layer, and a feedback layer; The perception layer is used to collect external perception information, including: preset puppet show performance videos, audience action images, and gesture command information. The perception layer is equipped with a visual camera module. The control layer, connected to the perception layer, execution layer, and feedback layer, receives external sensing information and operational status data from the feedback layer, generates control commands, and sends these commands to the execution layer. The control layer includes a host computer and an embedded main controller. The execution layer is used to receive and execute control commands; the execution layer includes a multi-degree-of-freedom robotic arm, a dexterous hand, and a gong and drum sound effect module. The feedback layer is used to collect real-time operating status data from the execution layer and transmit it back to the control layer; the feedback layer includes force sensors and servo torque detection sensors.
[0007] This invention provides a multi-degree-of-freedom dual-machine collaborative puppet show performance system, comprising: a perception layer, a control layer, an execution layer, and a feedback layer. The perception layer is used to collect external sensing information, including preset puppet show performance videos, audience action images, and gesture command information. The perception layer is equipped with a visual camera module. The control layer is connected to the perception layer, execution layer, and feedback layer, respectively, and is used to receive external sensing information and operational status data returned from the feedback layer, generate control commands, and send the control commands to the execution layer. The control layer includes a host computer and an embedded main controller. The execution layer is used to receive and execute control commands. The execution layer includes a multi-degree-of-freedom robotic arm, a dexterous hand, and a gong and drum sound effect module. The feedback layer is used to collect operational status data from the execution layer in real time and return it to the control layer. The feedback layer includes a force sensor and a servo torque detection sensor. In this invention, firstly, by integrating a visual camera module into the perception layer, multi-source information such as preset performance videos, audience actions, and gesture commands is collected in real time, solving the problem of unidirectional interaction and limited input in existing technologies, and realizing bidirectional, real-time interactive response. Then, relying on the intelligent processing of multi-source information at the control layer, the multi-degree-of-freedom robotic arm and dexterous hand at the execution layer are driven to perform precise and coordinated synchronized movements. Combined with closed-loop optimization of sensor data at the feedback layer, the problems of simple performance movements and insufficient coordination are solved, achieving an immersive experience for complex puppet performances. Finally, the adoption of a layered modular architecture and embedded main control reduces system complexity and hardware costs. Simultaneously, the integrated design of gong and drum sound effects solves the problems of traditional systems being complex, expensive, and difficult to popularize, enabling efficient deployment and promotion in museums, campuses, and other scenarios. In summary, this invention, through the collaborative efforts of four layers—perception, control, execution, and feedback—achieves a highly interactive, highly coordinated, and low-cost puppet show performance system, effectively overcoming the limitations of existing technologies in interactive experience and widespread application.
[0008] The present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0009] Figure 1 This is a schematic diagram of a multi-degree-of-freedom dual-machine collaborative puppet show performance system provided in an embodiment of the present invention. Detailed Implementation
[0010] The present invention will be further described in detail below with reference to specific embodiments, but the implementation of the present invention is not limited thereto.
[0011] To achieve a highly interactive, highly coordinated, and low-cost puppet show performance system, embodiments of the present invention provide a multi-degree-of-freedom dual-machine collaborative puppet show performance system. Figure 1 This is a schematic diagram of the structure of a multi-degree-of-freedom dual-machine collaborative glove puppetry performance system provided in an embodiment of the present invention, as shown below. Figure 1 As shown, it includes: perception layer, control layer, execution layer and feedback layer; The perception layer is used to collect external perception information, including: preset puppet show performance videos, audience action images, and gesture command information. The perception layer is equipped with a visual camera module.
[0012] Optionally, the multi-degree-of-freedom dual-machine collaborative puppet show performance system includes: a conventional performance mode and a puppet-following-human interaction mode; The visual camera module is used to capture preset puppet show performance videos in the regular performance mode, and to capture audience movement images in real time in the puppet-following-the-person interactive mode. The visual camera module is also used to collect gesture command information in both regular performance mode and puppet-following interactive mode.
[0013] Optionally, the gesture command information includes start, pause, stop, and mode switching commands.
[0014] In this invention, the activation command in the gesture instruction information is a hand gesture by the operator, used to activate the multi-degree-of-freedom dual-machine collaborative puppet show performance system, enabling the system to enter the performance operation state from the standby state; the pause command is a palm-forward gesture by the operator, used to temporarily interrupt the current performance action, enabling the system to enter the hold state; the stop command is a fist-clenching gesture by the operator, used to completely terminate the current performance, resetting the system to the standby state; the mode switching command is an extended index finger and circle gesture by the operator, used to switch between the regular performance mode and the puppet-following-human interaction mode. All gesture instructions are captured by the visual camera module and parsed by the gesture recognition model to generate control commands with the highest priority.
[0015] The control layer, which is connected to the perception layer, execution layer and feedback layer respectively, is used to receive external perception information and operation status data returned by the feedback layer, generate control commands and send the control commands to the execution layer; the control layer includes a host computer and an embedded main controller.
[0016] Optionally, the control commands include a first control command, a second control command, and a third control command; the host computer is used for: In the normal performance mode, a preset puppet show performance video is acquired, and the preset puppet show performance video is processed for motion extraction and trajectory planning based on the human pose estimation algorithm to generate the first control command used to drive the execution layer to reproduce the performance action; In the human-following interactive mode, the system acquires images of the audience's movements and performs motion extraction and motion parameter conversion on the images of the audience's movements based on the human posture estimation algorithm and motion mapping model, generating a second control command to drive the execution layer to synchronously follow the audience's movements. In both the regular performance mode and the puppet-following-human interaction mode, gesture command information is acquired, and the gesture command information is recognized based on the gesture recognition model to generate a third control command; the third control command is used to control the operating status and mode switching of the multi-degree-of-freedom dual-machine collaborative puppet show performance system. Among them, during the operation of the multi-degree-of-freedom dual-machine collaborative glove puppetry performance system, the third control command has the highest priority; The human pose estimation algorithm is a key point detection model based on deep learning, including one of the YOLO-Pose model, OpenPose model, or HRNet (High-Resolution Network) model; the gesture recognition model adopts the MediaPipe Hands model.
[0017] Specifically, the key point detection model is deployed on a host computer equipped with an Intel Core i7-12700 CPU and an NVIDIA GeForce RTX 3060 GPU.
[0018] The input image resolution for the keypoint detection model was uniformly set to 640×480 pixels, and the inference frame rate was consistently above 30 FPS. After keypoint detection, the system filtered all keypoints based on their confidence scores, retaining only those with a confidence score greater than 0.6. Subsequently, a Kalman filter algorithm was used to smooth the keypoint coordinates in consecutive frames after the filtering process to suppress jitter noise.
[0019] For the MediaPipe Hands gesture recognition model, the gesture recognition confidence threshold is set to 0.7, meaning that the gesture is considered valid only when the model outputs a gesture classification probability exceeding 0.7.
[0020] It should be noted that, in any mode, once the gesture recognition model outputs a valid third control command, that command is immediately assigned the highest priority. The host computer's main control program uses an event-triggered mechanism. When it receives a third control command, it immediately pauses the currently executing regular or follow task thread, saves the current task's context (including the currently executed action frame number, joint angles, and other states), and jumps to the gesture command processing subroutine. For example, if a "pause" command is received, the host computer stops sending new control commands and notifies the embedded main controller to maintain the current position of the execution layer; if a "stop" command is received, the host computer closes all task threads and sends a reset command to the embedded main controller, causing the execution layer to return to a safe initial pose.
[0021] Furthermore, at the hardware level, the host computer and the embedded main controller are connected via a USB bus and enumerated as a virtual COM Port (VCP) device on the host side. Both parties communicate using serial data frames based on the USB CDC-ACM protocol. The gesture recognition model runs on the host computer, and its generated third control command is sent through a separate USB interrupt endpoint to ensure that the command receives high priority transmission in bus contention. The embedded main controller's firmware configures a hardware interrupt for the USB interface receiving data from this interrupt endpoint. When data arrives, the CPU responds to the hardware interrupt request, saves the current task context (including the state of the ongoing motor closed-loop control calculation), and jumps to the Interrupt Service Routine (ISR) to perform command parsing and emergency processing. After processing, the original task context is restored, and execution continues.
[0022] At the software level, the ISR functions as follows: First, it reads and parses the type of the third control command (start, pause, stop, switch). Second, it executes the corresponding operation based on the command type: If it is "pause," the ISR freezes the pointer of the currently executing motor control command queue and locks the PWM output of all motors, keeping them in their current position; if it is "stop," the ISR clears all control command queues and calls the system reset function to return each joint motor to its zero-point pose at a predetermined speed; if it is "switch," the ISR sends an acknowledgment signal to the host computer. Upon receiving the acknowledgment, the host computer switches its internal operating mode flag and begins generating control commands for the new mode. To ensure the recoverability of the original motion trajectory, during the "pause" operation, the ISR backs up the currently unexecuted command queue and the index of the last executed command to a specific area of Random Access Memory (RAM). When a "start" command is received, the system resumes from the interrupt point and continues executing the backed-up command queue.
[0023] In the conventional performance mode, the multi-degree-of-freedom dual-machine collaborative puppet show performance system analyzes the preset puppet show performance video frame by frame using a human pose estimation algorithm. It identifies and extracts the pixel coordinates and corresponding confidence scores of multiple body key points (such as the head, shoulders, elbows, and wrists) of the puppet in each frame. Subsequently, the multi-degree-of-freedom dual-machine collaborative puppet show performance system filters and verifies the pixel coordinates of the extracted body key points based on the confidence scores. It then concatenates the pixel coordinates of the same body key point in consecutive frames in chronological order to form a reliable pixel-level motion trajectory for that body key point. Next, the multi-degree-of-freedom dual-machine collaborative puppet show performance system performs coordinate space transformation, inverse kinematics solution, and smooth trajectory planning on this pixel-level motion trajectory. Finally, it generates a sequence of joint angle and position commands that can directly drive the joints of the multi-degree-of-freedom robotic arm and dexterous hand in the execution layer to perform precise movements, i.e., the first control command.
[0024] In some implementations, the full execution process of the conventional performance mode is as follows: S1, the system powers on and performs a self-test, all joints return to zero, and enters standby mode; S2, the host computer loads a preset puppet show performance video file; S3, the host computer estimates the human posture frame by frame in the video, extracts the key point trajectory, and plans the joint angle sequence of the robotic arm and dexterous hand (i.e., the first control command); S4, the host computer sends the first control command to the embedded main controller via the TCP / IP protocol; S5, after receiving the command, the embedded main controller stores it in a buffer and sends it to the servo driver of the execution layer in time stamp order; S6, the execution layer drives the robotic arm and dexterous hand to move, while the feedback layer collects joint angles and fingertip pressure in real time; S7, the embedded main controller performs fuzzy PID closed-loop correction based on the feedback data and corrects the control command for the next cycle; S8, repeats S5-S7 until the actions corresponding to all video frames are completed; S9, the system automatically returns to standby mode.
[0025] In the puppet-following-human interaction mode, the multi-degree-of-freedom dual-machine collaborative puppet show performance system analyzes the real-time captured audience movement images frame by frame. Based on the human posture estimation algorithm, it identifies and extracts the pixel coordinates and corresponding confidence scores of key points of the audience's body (such as shoulders, elbows, and wrists) in real time. Subsequently, the pixel coordinates of the audience's key points are verified and filtered according to the confidence scores to reduce noise and jitter, resulting in processed key point coordinates. The processed key point coordinates are then input into a predefined motion mapping model. This motion mapping model, based on a pre-set human-robotic arm kinematic mapping relationship, converts the processed key point coordinates into target motion parameters for controlling the joint movements of the multi-degree-of-freedom robotic arm and dexterous hand in real time. Simultaneously, the physical motion limits of the multi-degree-of-freedom robotic arm are combined to constrain and smooth these target motion parameters, ultimately generating a joint control command sequence that can drive the execution layer to reproduce the audience's movements in real time and synchronously, i.e., the second control command.
[0026] Optionally, the timing of the audience-following interactive mode is as follows: Step T1, the system is in the normal performance mode or standby state, and the visual camera module captures images in real time; Step T2, the gesture recognition model detects the "mode switching" gesture, generates a third control command, the embedded main control responds, and the system switches to the audience-following interactive mode; Step T3, the visual camera module begins to capture images of the audience's movements in real time; Step T4, the host computer performs human posture estimation and motion mapping on the images, and generates a second control command; Step T5, the subsequent process is the same as S5-S7 in the normal mode, realizing real-time synchronous following of the audience's movements; Step T6, when the gesture recognition model detects the "mode switching" or "stop" gesture again, it exits the mode.
[0027] Furthermore, the specific workflow of the motion mapping model is as follows: Upon receiving the processed keypoint coordinates, the two-dimensional pixel coordinates corresponding to the processed keypoint coordinates are first transformed to a three-dimensional spatial coordinate system with the robotic arm base as the origin through camera calibration and spatial transformation. Subsequently, the human joint angles in the three-dimensional spatial coordinate system are converted in real time to the corresponding robotic arm joint angles according to a preset proportional mapping formula. The proportional mapping formula is: ; in, For the joint angle of the robotic arm, These are the angles of human joints in a three-dimensional coordinate system. This is an angle scaling factor determined based on the ratio of the range of motion of the human body to that of a multi-degree-of-freedom robotic arm joint. This represents the angular offset.
[0028] In this embodiment, the angle scaling factor The range of motion is determined by pre-measuring the ratio of the maximum range of motion of a human joint to the maximum range of motion of the corresponding joint of a multi-degree-of-freedom robotic arm. For example, for the shoulder joint, the maximum horizontal abduction angle of the human shoulder joint is approximately 180°, while the maximum horizontal rotation angle of the selected robotic arm's shoulder joint is 270°. The value is 270 / 180 = 1.5. (Angle offset) This is used to align the differences in joint angles between the human body and the robotic arm in their initial postures, and is obtained through calibration during the system initialization phase.
[0029] The specific calibration method is as follows: Have the operator assume a standard T-pose posture, record the angle values of the shoulder and elbow joints at this time, and simultaneously read the corresponding joint angle values when the robotic arm is in its initial safe position. The difference between the two is the calibration result. For example, if the angle of the human right shoulder joint is 0° in the T-pose, while the corresponding joint angle in the initial pose of the robotic arm is 10°, then... The value is 10°. and The values must be calibrated on-site based on the actual robotic arm model and operator's body size, and stored in the host computer configuration file.
[0030] Next, the motion mapping model calculates the joint angles of the multi-degree-of-freedom robotic arm based on the physical motion limits of each joint. Real-time truncation and dynamic smoothing filtering are performed to finally output a set of target joint angle sequences that are proportional to the audience's movements.
[0031] Optionally, the embedded main controller is used to receive the first control command, the second control command, or the third control command generated by the host computer, and to perform real-time scheduling and control of the execution layer based on the first control command, the second control command, or the third control command. The embedded master controller is also used to receive operating status data and perform closed-loop error correction on the first control command or the second control command based on the operating status data.
[0032] Optionally, the embedded master controller is used to calculate the deviation information between the operating status data and the first control command or the second control command; the deviation information includes: deviation value and deviation change rate; Based on the deviation information, a fuzzy PID (Proportional-Integral-Derivative) controller is used to adjust the first or second control command to form a corrected control command; The execution layer is scheduled and controlled in real time based on the revised control commands.
[0033] The embedded master controller's processing flow is as follows: First, it receives real-time operational status data from the feedback layer. This data includes the actual pressure applied to the dexterous fingertip and the actual angles and torques of each joint of the multi-degree-of-freedom robotic arm, collected by force sensors and servo torque detection sensors. Then, the embedded master controller compares this actual data with the target state set by the control commands from the host computer, calculating the deviation and rate of change between the two. Based on this deviation information, the embedded master controller uses a fuzzy PID controller to dynamically adjust the control commands, generating corrected control commands. Finally, based on these corrected control commands, the embedded master controller performs real-time scheduling and control of the multi-degree-of-freedom robotic arm and dexterous hand at the execution layer, completing closed-loop correction.
[0034] The dynamic adjustment of the fuzzy PID controller is key to improving the system's adaptive performance. Its process is a continuous flow involving fuzzification, rule-based reasoning, and defuzzification. Specifically, the embedded controller transforms the input deviation value and deviation rate of change into seven membership degrees (positive large, positive medium, positive small, zero, negative small, negative medium, and negative large) using a predefined membership function, thus completing the fuzzification. Subsequently, reasoning is performed based on a pre-built fuzzy rule base containing multiple rules. This rule base defines the appropriate adjustments to the original parameters of the fuzzy PID controller, including the proportional coefficient, under all possible combinations of fuzzy states of the deviation value and deviation rate of change. Integral coefficient With differential coefficients The direction of the fuzzy adjustment. For example, when both the deviation value and the rate of change of deviation are evaluated as negative, the fuzzy rule base will output a proportional coefficient. The conclusion is that adjustments should be made in a more positive direction.
[0035] By defuzzifying the data, the above conclusions are converted into precise parameter adjustment values, which are then used to update the parameters of the fuzzy PID controller online. The fuzzy PID controller uses the deviation information and the updated parameters to calculate new control quantities, thereby generating corrected control commands.
[0036] Ultimately, based on this corrected control command, the embedded master controller performs real-time scheduling and driving of the multi-degree-of-freedom robotic arm and dexterous hand in the execution layer, completing a highly adaptive closed-loop correction cycle.
[0037] The fuzzy adjustment process of the original parameters of the above fuzzy PID controller can be performed according to the following formula: ; ;
[0038] in, This represents the updated scaling factor. This represents the updated integral coefficient. Indicates the updated differential coefficients. Represents the original scaling factor. Represents the original integral coefficients. Represents the original differential coefficients. The parameter adjustment amount represents the proportionality coefficient. The parameter adjustment amount represents the integral coefficient. The parameter adjustment amount represents the differential coefficient. This represents the weighting coefficient.
[0039] In this embodiment, the weighting coefficient The value range is from 0 to 1, and its function is to balance the influence of the original parameter and the adjustment amount on the final parameter. In this embodiment, preferably... A value of 0.8 is chosen to prioritize the stability of the original parameters. Original scaling factor. Original integral coefficients Original differential coefficients The initial calibration values were obtained using the Ziegler-Nichols tuning method, specifically: first, the initial calibration values were... and Set to 0, then gradually increase. Continue until the system produces constant-amplitude oscillations, and record the critical gain at this point. and oscillation period Then, calculations were performed based on empirical formulas: ; ; ; The seven membership functions employ triangular membership functions, with their universe of discourse set as follows: These correspond to negative large (NB), negative medium (NM), negative small (NS), zero (ZO), positive small (PS), positive medium (PM), and positive large (PB), respectively. The fuzzy rule base contains 49 rules. For example, when the deviation value e is NB and the deviation change rate ec is NB, For PB, For NB, For PS; when e is ZO and ec is ZO, For ZO, For ZO, ZO is used. This rule base is built based on expert experience and aims to achieve a balance between fast response and overshoot suppression.
[0040] The execution layer is used to receive and execute control commands; the execution layer includes a multi-degree-of-freedom robotic arm, a dexterous hand, and a gong and drum sound effect module.
[0041] Optionally, the gong and drum sound effect module is connected to the control layer for synchronously generating or triggering corresponding gong and drum sound effects based on the action timing information in the control command.
[0042] Optionally, the operational status data collected by the feedback layer includes the actual angles and torques of each joint of the multi-degree-of-freedom robotic arm, as well as the actual pressure of the dexterous fingertips.
[0043] Optionally, the multi-degree-of-freedom robotic arm and the dexterous hand form a collaborative dual-machine system for jointly holding and manipulating the puppet; the host computer and the embedded main controller work together to perform kinematic calculations and trajectory coordination for the dual-machine system.
[0044] The feedback layer is used to collect real-time operating status data from the execution layer and transmit it back to the control layer; the feedback layer includes force sensors and servo torque detection sensors.
[0045] In this embodiment, the kinematic calculation of the dual-machine system by the host computer and the embedded main controller includes the following steps: First, the host computer uses the damped least squares (DLS) method to perform inverse kinematics calculation on a single robotic arm based on the desired puppet end-effector pose, obtaining a set of joint angle solutions. For the dual-machine system, where the left and right robotic arms each hold different parts of the puppet (such as the torso and arms), the relative poses of the end-effectors (i.e., the dexterous hand gripping points) of the two robotic arms must satisfy the puppet's structural constraints. To this end, a cooperative constraint equation is introduced, using the joint angles of the left and right robotic arms as joint variables to construct an augmented Jacobian matrix that includes the kinematic models of the two robotic arms and the rigid body constraints of the puppet. The DLS method is then used again for joint solution to obtain a unique solution that satisfies the cooperative requirements.
[0046] For trajectory coordination, a bounding box-based collision detection algorithm is used to avoid collisions during the movement of the two robotic arms. An axis-aligned bounding box is established for each link of each robotic arm, and the minimum distance between all bounding boxes is calculated in each control cycle. If the distance is less than a preset safety threshold (e.g., 5 cm), a collision avoidance strategy is triggered, including: reducing the movement speed of the faster robotic arm, or applying a small repulsive force vector along the normal direction of the original trajectory to separate the two trajectories until the distance recovers to above the safety threshold.
[0047] Furthermore, the gripping force distribution logic of the dexterous hand holding the puppet is as follows: the gripping force is detected in real time by force sensors installed at the tips of the dexterous fingers, and a desired gripping force range (e.g., 3-5 Newtons) is set. When the force of a fingertip is detected to be below the lower limit, the duty cycle of the pulse width modulation (PWM) of the corresponding finger servo is increased to increase the clamping force; conversely, when it is above the upper limit, the PWM duty cycle is decreased, thereby ensuring a stable grip without damaging the puppet.
[0048] To achieve synchronized performance between the sound puppets and the puppets, some implementation methods also include motion timing marking, sound effect library binding, and delay calibration processing, specifically including: During motion trajectory planning, the host computer simultaneously parses the audio track of the preset performance video and extracts the audio beat points. Each beat point corresponds to a specific motion keyframe (e.g., the peak frame of the puppet raising its leg or waving its hand). In the first control command generated, the host computer adds a timestamp tag to each keyframe, which serves as the motion timing marker.
[0049] Additionally, a sound effects database is pre-stored in the system, containing various drum and gong sound files (such as "dong," "qiang," and "da"). Each action sequence marker is associated with one or more sound effect IDs. For example, an action frame marked as "raising a leg" is associated with the sound effect ID "dong"; an action frame marked as "waving" is associated with the sound effect ID "qiang." This binding rule is stored in an XML configuration file and can be customized by the user.
[0050] Because the robotic arm's movements have a mechanical inertia delay, while audio playback has almost no delay, direct triggering would cause audio-visual desynchronization. Therefore, when generating the first control command, the host computer estimates the time delay from issuing the command to the completion of the action based on the robotic arm's dynamic model. While sending action commands to the embedded main controller, the host computer also sends trigger commands to the gong and drum sound effect module, but the trigger command is sent earlier. Specifically, suppose the first Frame action is expected in When the timing is right, the sound effect trigger command will be in effect. It will be sent out at any time. The value can be obtained through offline identification experiments, such as measuring the time from sending a step angle command to the encoder feedback reaching 90% of the target angle, and writing it as a fixed compensation value into the system parameters.
[0051] This invention provides a multi-degree-of-freedom dual-machine collaborative puppet show performance system. First, by integrating a visual camera module into the perception layer, it collects multi-source information in real time, including preset performance videos, audience movements, and gesture commands. This solves the problems of unidirectional interaction and limited input in existing technologies, achieving bidirectional, real-time interactive response. Then, relying on the intelligent processing of multi-source information by the control layer, it drives the multi-degree-of-freedom robotic arm and dexterous hand in the execution layer to perform precise and coordinated synchronized movements. Combined with closed-loop optimization of sensor data in the feedback layer, it solves the problems of simple performance movements and insufficient coordination, achieving an immersive experience for complex puppet performances. Finally, by adopting a layered modular architecture and embedded main control, it reduces system complexity and hardware costs. Simultaneously, it integrates gong and drum sound effects into a unified design, solving the problems of traditional systems being complex, expensive, and difficult to popularize, enabling efficient deployment and promotion in museums, campuses, and other scenarios. In summary, this invention, through four layers of collaboration—perception, control, execution, and feedback—achieves a highly interactive, highly coordinated, and low-cost puppet show performance system, effectively overcoming the limitations of existing technologies in interactive experience and widespread application.
[0052] To verify the technical effectiveness of the multi-degree-of-freedom dual-machine collaborative puppet show performance system provided by this invention in terms of interactivity, collaboration, and cost, the applicant built a prototype system and conducted comparative experiments. The experimental subjects included: the system of this invention (denoted as Scheme A) and a traditional automated puppet show performance system based on a single robotic arm and pre-programmed control (denoted as Scheme B, using industrial PLC control, a single six-axis robotic arm). The experimental conditions were kept consistent: the test environment was a standard laboratory at 25℃, and the puppets used were puppet shows of the same standard size. Specific experimental results are as follows: I. The following is a comparative experiment on interactive capabilities: Experimental Method: Ten volunteers with no relevant operational experience were invited to complete three tasks using scheme A and scheme B respectively: (1) start the system and enter the performance state; (2) pause the performance and resume it midway; (3) switch the performance mode. The success rate of each volunteer's operation and the time required to complete the task were recorded.
[0053] Experimental Results: In Scheme A, all volunteers successfully completed all three tasks on their first attempt, with an average completion time of 8.2 seconds. In Scheme B, because it only supported a single pre-programmed mode and could not implement pause, resume, or mode switching functions, the corresponding tasks could not be completed. Furthermore, Scheme A supported three interaction modes: regular performance, puppet-following-person movement, and gesture control, while Scheme B only supported the regular performance mode. The number of user interaction modes in Scheme A was 300% of that in Scheme B.
[0054] II. The following is a comparative experiment on synergistic performance: Experimental Methods: A standard puppet show sequence of movements lasting 30 seconds was performed using both schemes A and B. This sequence included coordinated movements of the puppet's two arms (such as clasping hands, bowing, and alternating waving of the left and right hands). The actual motion trajectories of the end effectors (i.e., the dexterous hand gripping points) of the two robotic arms were recorded using a high-speed camera system (sampling frequency 1000Hz) and compared with the theoretical target trajectories. Evaluation indicators included: (1) relative pose error of the two end effectors (unit: mm); (2) completion rate of the movements (i.e., the percentage of the number of movements successfully completed).
[0055] Experimental results: Scheme A achieved an average relative pose error of 1.8mm between the two robotic arms, with a maximum error not exceeding 3.5mm, and a 100% completion rate for the motion execution. Scheme B, using only a single robotic arm, could not achieve coordinated motion between the two arms and therefore could not complete the coordinated motion sequence; its motion execution completion rate was 40% (only the single-arm motion portion could be completed). Scheme A improved the coordinated motion completion rate by 150% compared to Scheme B.
[0056] III. Hardware cost comparison experiment is as follows: Experimental Methodology: A list of core hardware components and their procurement costs for both Scheme A and Scheme B were compiled, including but not limited to controllers, actuators, sensors, and communication modules. Cost data was obtained from publicly available quotations from major industrial component suppliers during the same period.
[0057] Experimental Results: Scheme A adopts a layered modular architecture. The core control unit uses a commercial embedded motherboard (costing approximately 800 RMB) and an STM32 series microcontroller (costing approximately 120 RMB). The actuators consist of two lightweight six-axis collaborative robotic arms (costing approximately 6000 RMB per unit) and a customized dexterous hand (costing approximately 1500 RMB). Including the vision module and sensors (costing approximately 500 RMB), the total core hardware cost is approximately 14920 RMB. Scheme B uses an industrial PLC controller (costing approximately 3500 RMB), a dedicated motion control card (costing approximately 2500 RMB), and an industrial six-axis robotic arm (costing approximately 12000 RMB), with a total core hardware cost of approximately 18000 RMB. While achieving richer interactive functions and dual-machine collaboration capabilities, Scheme A's total hardware cost is approximately 17.1% lower than Scheme B. Considering that Scheme A achieves multi-user interaction and dual-machine collaboration functions that Scheme B lacks at the same cost, its cost-performance advantage is even more significant.
[0058] In summary, the experimental data fully demonstrates that the system of the present invention has substantial technical advantages in terms of interaction mode richness, dual-machine collaboration accuracy and action completion, as well as hardware cost control, achieving the design goal of "high interaction, high collaboration, and low cost," rather than subjective textual assertions.
[0059] The above are merely specific embodiments of the present invention, but the scope of protection of the present invention is not limited thereto. Any variations or substitutions that can be easily conceived by those skilled in the art within the technical scope disclosed in the present invention should be included within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
Claims
1. A multi-degree-of-freedom dual-machine collaborative glove puppetry performance system, characterized in that, include: The layers are: perception layer, control layer, execution layer, and feedback layer. The perception layer is used to collect external sensing information; The external sensing information includes: preset puppet show performance videos, audience action images, and gesture command information; the sensing layer is equipped with a visual camera module. The control layer is connected to the perception layer, the execution layer, and the feedback layer, respectively, and is used to receive external perception information and operation status data returned by the feedback layer, generate control commands, and send the control commands to the execution layer; the control layer includes a host computer and an embedded main controller; The execution layer is used to receive and execute the control commands; the execution layer includes a multi-degree-of-freedom robotic arm, a dexterous hand, and a gong and drum sound effect module; The feedback layer is used to collect the operating status data of the execution layer in real time and transmit it back to the control layer; the feedback layer includes a force sensor and a servo torque detection sensor.
2. The multi-degree-of-freedom dual-machine collaborative glove puppetry performance system according to claim 1, characterized in that, The multi-degree-of-freedom dual-machine collaborative puppet show performance system includes: a conventional performance mode and a puppet-following-human interactive mode; The visual camera module is used to capture the preset puppet show performance video in the normal performance mode, and to capture the audience's movements in real time in the puppet-following-the-person interactive mode. The visual camera module is also used to collect the gesture command information in the regular performance mode and the puppet-following human interaction mode.
3. The multi-degree-of-freedom dual-machine collaborative glove puppetry performance system according to claim 1, characterized in that, The gesture command information includes start, pause, stop, and mode switching commands.
4. A multi-degree-of-freedom dual-machine collaborative glove puppetry performance system according to claim 2, characterized in that, The control commands include a first control command, a second control command, and a third control command; the host computer is used for: In the normal performance mode, the preset puppet show performance video is acquired, and the preset puppet show performance video is processed for motion extraction and trajectory planning based on the human posture estimation algorithm to generate the first control command for driving the execution layer to reproduce the performance action; In the puppet-following human interaction mode, the audience's action image is acquired, and the audience's action image is processed by action extraction and action parameter conversion based on the human posture estimation algorithm and action mapping model to generate the second control command for driving the execution layer to synchronously follow the audience's action. In the conventional performance mode and the puppet-following-human interaction mode, the gesture instruction information is acquired, and the gesture instruction information is recognized based on the gesture recognition model to generate the third control command; The third control command is used to control the operating status and mode switching of the multi-degree-of-freedom dual-machine collaborative glove puppetry performance system. During the operation of the multi-degree-of-freedom dual-machine collaborative puppet show performance system, the third control command has the highest priority. The human pose estimation algorithm is a key point detection model based on deep learning, including one of the YOLO-Pose model, OpenPose model, or HRNet model; the gesture recognition model adopts the MediaPipe Hands model.
5. A multi-degree-of-freedom dual-machine collaborative glove puppetry performance system according to claim 4, characterized in that, The embedded main controller is used to receive the first control command, the second control command, or the third control command generated by the host computer, and to perform real-time scheduling and control of the execution layer based on the first control command, the second control command, or the third control command. The embedded main controller is also used to receive the operating status data and perform closed-loop error correction on the first control command or the second control command based on the operating status data.
6. A multi-degree-of-freedom dual-machine collaborative glove puppetry performance system according to claim 5, characterized in that, The embedded main controller is used to calculate the deviation information between the operating status data and the first control command or the second control command; The deviation information includes: the deviation value and the rate of change of deviation; Based on the deviation information, a fuzzy PID controller is used to adjust the first control command or the second control command to form a corrected control command. The execution layer is scheduled and controlled in real time based on the modified control commands.
7. A multi-degree-of-freedom dual-machine collaborative glove puppetry performance system according to claim 1, characterized in that, The gong and drum sound effect module is communicatively connected to the control layer and is used to synchronously generate or trigger the corresponding gong and drum sound effects according to the action timing information in the control command.
8. A multi-degree-of-freedom dual-machine collaborative glove puppetry performance system according to claim 1, characterized in that, The operational status data collected by the feedback layer includes the actual angles and torques of each joint of the multi-degree-of-freedom robotic arm, as well as the actual pressure of the dexterous fingertips.
9. A multi-degree-of-freedom dual-machine collaborative glove puppetry performance system according to claim 1, characterized in that, The multi-degree-of-freedom robotic arm and the dexterous hand form a collaborative dual-machine system for jointly holding and manipulating the puppet; the host computer and the embedded main controller work together to perform kinematic calculations and trajectory coordination for the dual-machine system.