Shadow puppet digital human interaction method, system, equipment and medium
By acquiring user command information and utilizing motion mapping tables and baseline joint translation alignment technology, the problems of insufficient interaction and disjointed movements in traditional digital shadow puppet displays have been solved. Stable interaction between the digital shadow puppet and the user and the presentation of artistic style have been achieved, thus enhancing the viewing experience.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- WUYI UNIV
- Filing Date
- 2025-12-12
- Publication Date
- 2026-05-05
AI Technical Summary
The digital display of traditional shadow puppetry suffers from the inability to achieve deep interaction with users, the lack of effective position calibration in motion generation, resulting in positional deviations and disjointed transitions when switching between segments, and the failure to fully present the unique artistic style of shadow puppetry, thus affecting the viewing experience.
By acquiring user command information, querying the shadow puppet action sequence using the motion mapping table, and aligning it with the target reference joint as the baseline, the joint coordinates are corrected to achieve a smooth and stable presentation of the action segments. Combined with joint interpolation calculations and rendering technology, the continuity of the action and the presentation of the artistic style are ensured.
It achieves seamless and stable interaction between the digital shadow puppet and the user, as well as cross-segment switching of movements, enhancing the viewing experience and ensuring smooth movements and faithful expression of artistic style.
Smart Images

Figure CN121982173A_ABST
Abstract
Description
Technical Field
[0001] This application relates to, but is not limited to, the field of shadow puppet design, and particularly to a method, system, device, and medium for interactive digital human shadow puppetry. Background Technology
[0002] Traditional shadow puppetry, an important intangible cultural heritage of my country, has long been limited in its transmission and dissemination due to its reliance on manual operation and the constraints of performance venues, making large-scale promotion difficult. With the gradual application of digital technology in the cultural field, digital display of shadow puppetry has become an important direction for the protection and dissemination of this intangible cultural heritage; however, existing solutions still have many shortcomings. Some digital solutions only remain at the level of static material display or pre-set animation playback, failing to achieve deep interaction with users. Furthermore, the current digital shadow puppetry motion generation lacks an effective position calibration mechanism, leading to problems such as positional shifts and disjointed transitions during segment transitions. Moreover, most of these solutions do not fully consider the unique artistic performance logic of shadow puppetry, making it difficult to present its distinctive artistic style, seriously affecting the viewing experience and hindering the practical implementation and widespread dissemination of digital shadow puppetry. Summary of the Invention
[0003] The following is an overview of the subject matter described in detail herein. This overview is not intended to limit the scope of the claims.
[0004] This application provides a method, system, device, and medium for interactive shadow puppet digital human, which realizes the interaction between the shadow puppet digital human and the user and the smooth and stable presentation of cross-segment switching of actions, effectively improving the viewing experience.
[0005] In a first aspect, embodiments of this application provide a shadow puppet digital human interaction method, comprising: acquiring instruction information input by a user; querying a corresponding shadow puppet action sequence from a preset action mapping table according to the instruction information, the shadow puppet action sequence including an initial action segment and at least one subsequent action segment, each action segment containing multiple consecutive action frames; using the target reference joint of the first frame in the shadow puppet action sequence as a baseline; determining the target reference joint coordinate difference between the start and end frames of the initial action segment and the subsequent action segment; translating and aligning the joint coordinates of the action segment according to the baseline and the corresponding target reference joint coordinate difference to obtain the translated and aligned joint coordinates; obtaining a corrected shadow puppet action sequence according to the translated and aligned joint coordinates, and rendering the corrected shadow puppet action sequence to obtain a dynamic performance scene of the shadow puppet digital human.
[0006] In conjunction with the first aspect, in one embodiment of this application, the step of constructing the action mapping table includes: acquiring a human performance action sequence, the performance action sequence including multiple performance frames, each performance frame corresponding to multiple skeletal key points of the human body; filtering out the core skeletal key points of the shadow puppet digital human from the multiple skeletal key points, and performing coordinate normalization processing on the core skeletal key points using the target reference joint coordinates as reference points to obtain the processed key point coordinates; obtaining a custom performance action sequence based on the processed key point coordinates; and adding the custom performance action sequence to the initial mapping table after associating it with the corresponding semantic instructions to obtain the action mapping table.
[0007] In one embodiment of this application, the core skeletal key points include at least one of the following: head key points, left shoulder key points, right shoulder key points, left elbow key points, right elbow key points, left wrist key points, right wrist key points, left hip key points, right hip key points, left knee key points, right knee key points, left ankle key points, and right ankle key points.
[0008] In one embodiment of this application, before rendering the modified shadow puppet action sequence, the method further includes: performing joint interpolation on the modified shadow puppet action sequence to obtain a shadow puppet action sequence with continuous and smooth movements.
[0009] In one embodiment of this application, rendering the modified shadow puppet action sequence includes: loading the materials of each component of the shadow puppet digital human; mapping the joint coordinates of the action frames in the modified shadow puppet action sequence to the corresponding shadow puppet components; correcting the shoulder midpoint position, shoulder offset and elbow height of the shadow puppet components, and rendering each shadow puppet component frame by frame based on a preset frame rate.
[0010] In one embodiment of this application, the step of querying the corresponding shadow puppet action sequence from a preset action mapping table based on the instruction information includes: extracting keywords and performing semantic parsing on the instruction information to obtain the corresponding text instruction; and querying the corresponding shadow puppet action sequence from the preset action mapping table based on the text instruction.
[0011] In one embodiment of this application, after obtaining the instruction information input by the user, the method further includes: generating adapted response information based on the instruction information; and displaying the response information.
[0012] Secondly, embodiments of this application provide a shadow puppet digital human interaction system, comprising: an input module for acquiring user-input instruction information; an action query module for querying a corresponding shadow puppet action sequence from a preset action mapping table based on the instruction information, the shadow puppet action sequence including an initial action segment and at least one subsequent action segment, each action segment containing multiple consecutive action frames; a joint correction module for using the target reference joint of the first frame in the shadow puppet action sequence as a baseline; determining the target reference joint coordinate difference between the start and end frames of the initial action segment and the subsequent action segment; translating and aligning the joint coordinates of the action segment according to the baseline and the corresponding target reference joint coordinate difference to obtain the translated and aligned joint coordinates; and an action rendering module for obtaining the corrected shadow puppet action sequence based on the translated and aligned joint coordinates, and rendering the corrected shadow puppet action sequence to obtain a dynamic performance image of the shadow puppet digital human.
[0013] On the other hand, embodiments of this application provide an electronic device, which includes a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein when the computer program is executed by the processor, it implements the method described above.
[0014] On the other hand, embodiments of this application provide a computer-readable storage medium storing a processor-executable program, which, when executed by a processor, is used to perform the methods described above.
[0015] This application provides a shadow puppet digital human interaction method, interaction system, electronic device, and computer-readable storage medium. The method first acquires user-input command information; then, based on the command information, it queries a preset action mapping table for the corresponding shadow puppet action sequence. The shadow puppet action sequence includes an initial action segment and at least one subsequent action segment, each action segment containing multiple consecutive action frames. Using the target reference joint of the first frame in the shadow puppet action sequence as a baseline, after determining the target reference joint coordinate difference between the start and end frames of the initial action segment and the subsequent action segments, the joint coordinates of the action segments are translated and aligned according to the baseline and the corresponding target reference joint coordinate difference, resulting in translated and aligned joint coordinates. Then, a corrected shadow puppet action sequence is obtained based on the translated and aligned joint coordinates, and the corrected shadow puppet action sequence is rendered to obtain a dynamic performance of the shadow puppet digital human. This application embodiment, through command-driven action matching and action segment translation alignment based on reference joints, achieves a coherent and stable presentation of interaction between the shadow puppet digital human and the user, as well as cross-segment action switching, effectively improving the viewing experience. Attached Figure Description
[0016] Figure 1 This is a flowchart of the shadow puppet digital human interaction method provided in the embodiments of this application; Figure 2 This is a schematic diagram illustrating the construction process of the action mapping table provided in an embodiment of this application; Figure 3 This is a flowchart illustrating the specific process of rendering the action sequence provided in the embodiments of this application; Figure 4 This is a block diagram of the shadow puppet digital human interaction system provided in the embodiments of this application; Figure 5 This is a system overall structure diagram provided in a specific embodiment of this application; Figure 6 This is a general flowchart of the interactive system provided in one embodiment of this application; Figure 7 This is a flowchart of a speech recognition processing embodiment provided in this application; Figure 8 This is a flowchart of the semantic parsing module processing provided in one embodiment of this application; Figure 9 This is a flowchart of the motion splicing and rendering process provided in one embodiment of this application; Figure 10 This is a schematic diagram of human skeleton recognition feature points provided in one embodiment of this application; Figure 11 This is a text instruction and action path mapping diagram provided in one embodiment of this application; Figure 12 This is a schematic diagram of a mapping process provided in one embodiment of this application; Figure 13 This is a graphical interface demonstration diagram provided in one embodiment of this application. Detailed Implementation
[0017] To make the objectives, technical solutions, and advantages of this application clearer, the following detailed description is provided in conjunction with the accompanying drawings and embodiments. It should be understood that the specific embodiments described herein are merely illustrative and not intended to limit the scope of this application.
[0018] It should be noted that although the flowchart shows a logical order, in some cases, the steps shown or described may be performed in a different order than that shown in the flowchart. The terms "first," "second," etc., used in the specification, claims, and the aforementioned drawings are used to distinguish similar objects and are not necessarily used to describe a specific order or sequence. It should be understood that the structures, proportions, sizes, etc., depicted in the drawings are only used to complement the content disclosed in the specification for those skilled in the art to understand and read, and are not intended to limit the implementation conditions of this application. Therefore, they have no substantial technical significance. Any modifications to the structure, changes in proportions, or adjustments to size, without affecting the effects and purposes achieved by this application, should still fall within the scope of the technical content disclosed in this application. Similarly, the terms such as "upper," "lower," "left," "right," "middle," and "one" used in this specification are only for clarity of description and are not used to limit the scope of implementation of this application. Changes or adjustments in their relative relationships, without substantially altering the technical content, should also be considered within the scope of implementation of this application.
[0019] Unless otherwise defined, all technical and scientific terms used herein have the same meaning as commonly understood by one of ordinary skill in the art to which this application belongs. The terminology used herein is for the purpose of describing embodiments of this application only and is not intended to limit this application.
[0020] Traditional shadow puppetry, an important intangible cultural heritage of my country, has long been limited in its transmission and dissemination due to its reliance on manual operation and the constraints of performance venues, making large-scale promotion difficult. With the gradual application of digital technology in the cultural field, digital display of shadow puppetry has become an important direction for the protection and dissemination of this intangible cultural heritage; however, existing solutions still have many shortcomings. Some digital solutions only remain at the level of static material display or pre-set animation playback, failing to achieve deep interaction with users. Furthermore, the current digital shadow puppetry motion generation lacks an effective position calibration mechanism, leading to problems such as positional shifts and disjointed transitions during segment transitions. Moreover, most of these solutions do not fully consider the unique artistic performance logic of shadow puppetry, making it difficult to present its distinctive artistic style, seriously affecting the viewing experience and hindering the practical implementation and widespread dissemination of digital shadow puppetry.
[0021] In view of this, embodiments of this application provide a shadow puppet digital human interaction method, interaction system, electronic device, and computer-readable storage medium. In this method, firstly, user-input instruction information is acquired; then, based on the instruction information, the corresponding shadow puppet action sequence is queried from a preset action mapping table. The shadow puppet action sequence includes an initial action segment and at least one subsequent action segment, each action segment containing multiple consecutive action frames. Using the target reference joint of the first frame in the shadow puppet action sequence as a baseline, after determining the target reference joint coordinate difference between the start and end frames of the initial action segment and the subsequent action segments, the joint coordinates of the action segments are translated and aligned according to the baseline and the corresponding target reference joint coordinate difference, resulting in translated and aligned joint coordinates. Then, a corrected shadow puppet action sequence is obtained based on the translated and aligned joint coordinates, and the corrected shadow puppet action sequence is rendered to obtain a dynamic performance of the shadow puppet digital human. Embodiments of this application, through instruction-driven action matching and action segment translation alignment design based on reference joints, achieve a coherent and stable presentation of interaction between the shadow puppet digital human and the user, as well as cross-segment action switching, effectively improving the viewing experience.
[0022] The embodiments of this application will be further described below with reference to the accompanying drawings.
[0023] Reference Figure 1 , Figure 1 This is a flowchart of the shadow puppet digital human interaction method provided in the embodiments of this application. The process may specifically include, but is not limited to, steps 110 to 160.
[0024] Step 110: Obtain the user-inputted instruction information; Step 120: Query the corresponding shadow puppet action sequence from the preset action mapping table according to the instruction information. The shadow puppet action sequence includes an initial action segment and at least one subsequent action segment. Each action segment contains multiple consecutive action frames. Step 130: Use the target reference joint of the first frame in the shadow puppet action sequence as the baseline; Step 140: Determine the target reference joint coordinate difference between the start and end frames of the initial action segment and the subsequent action segment; Step 150: Based on the difference between the baseline and the corresponding target reference joint coordinates, translate and align the joint coordinates of the motion segment to obtain the translated and aligned joint coordinates; Step 160: Obtain the corrected shadow puppet action sequence based on the joint coordinates after translation and alignment, and render the corrected shadow puppet action sequence to obtain the dynamic performance of the shadow puppet digital human.
[0025] Steps 110 to 160 will be described in detail below.
[0026] Understandably, digital shadow puppets refer to virtual character systems that digitally reconstruct the artistic form and performance style of traditional shadow puppet characters. These systems can be driven by external commands, provide interactive responses, and reproduce characteristic shadow puppet movements, making them suitable for various applications such as exhibitions, cultural performances, and intangible cultural heritage education. For example, in a museum's intangible cultural heritage exhibition area, this system can present interactive shadow puppet performances to visitors.
[0027] In one feasible embodiment, in step 110, the user-inputted instruction information may include text instructions, voice instructions, or both. In practical applications, users can type text information through the input box on the interactive device's screen. The text information can correspond to preset action instructions for the shadow puppet digital human, such as "dance," "bow," "walk forward," "wave," etc. After typing, pressing Enter or Tab will submit the text instruction information. In addition, users can also input voice information through the voice acquisition device (such as a microphone) on the interactive device.
[0028] In one feasible embodiment, if the information input by the user is a voice command, the voice command can be converted into a text command. Specifically, the user's voice signal can be collected through a microphone, and then the collected voice signal can be recognized and processed by the offline speech recognition engine Vosk. After recognition and parsing, the text information corresponding to the voice signal is obtained, so as to perform semantic understanding and intent extraction on the user's command information in the future.
[0029] In a feasible embodiment, before executing step 120, keyword extraction and semantic parsing can be performed on the instruction information to obtain the corresponding text instructions; then, the corresponding shadow puppet action sequence can be queried from the preset action mapping table according to the text instructions.
[0030] In a feasible embodiment, in step 120, the corresponding shadow puppet action sequence can be queried from a preset action mapping table according to the text command. This action mapping table can pre-store the correspondence between various text commands and shadow puppet action sequences, enabling the query process to be completed quickly and accurately. The queried shadow puppet action sequence may include an initial action segment and at least one subsequent action segment, wherein the initial action segment corresponds to the beginning part of the shadow puppet digital human's action, and the subsequent action segment corresponds to the continuous action part after the initial action. Each action segment may contain multiple consecutive action frames, which are sequentially connected to fully present the dynamic process of the corresponding action segment.
[0031] In a feasible embodiment, in step 130, the target reference joint of the first frame refers to the pelvic joint of the first frame in the shadow puppet action sequence. Specifically, after obtaining the corresponding shadow puppet action sequence, the pelvic joint coordinates of the first frame in the shadow puppet action sequence can be selected as a global anchor point. Based on the global anchor point, the first frame is centered so that the shadow puppet action corresponding to the first frame can be in a preset display reference position. After the processing is completed, the centered pelvic joint coordinates can be used as a baseline. This baseline can be used as a reference for the position calibration of subsequent action frames to ensure that the display position of the entire shadow puppet action sequence remains uniform and standardized.
[0032] In a feasible embodiment, in step 140, the start and end frames refer to the start frame (i.e., the first frame of the action segment) and the end frame (i.e., the last frame of the action segment) of each action segment. The target reference joint of the start and end frames refers to the pelvic joint of the start and end frames of each action segment. Each action frame contains the corresponding coordinate data of the pelvic joint of the shadow puppet digital human. For the initial action segment, the pelvic joint coordinates of its start and end frames can be extracted, and the coordinate difference between the two can be calculated. Similarly, for each subsequent action segment, the pelvic joint coordinates of its own start and end frames can be extracted, and then the pelvic joint coordinate difference corresponding to each subsequent action segment can be calculated. These coordinate differences can reflect the positional change of the pelvic joint within each action segment, providing data support for the positional calibration of subsequent action frames.
[0033] In a feasible embodiment, in step 150, after obtaining the baseline and the pelvic joint coordinate differences corresponding to each action segment, the baseline can be used as a unified position reference. Combined with the pelvic joint coordinate differences of each action segment, the joint coordinates of all action frames within that action segment are translated and aligned. Translation alignment refers to adjusting the overall position of the entire action segment while maintaining the relative positional relationship of each joint within the action segment, so that the position of the action segment matches the baseline. Specifically, the reference placement position of the action can be determined based on the baseline. Then, based on the pelvic joint coordinate differences of the corresponding action segment, the required translation adjustment amount for that action segment can be calculated. Then, according to this adjustment amount, all joint coordinates of each action frame within the action segment are synchronously translated, ensuring that the action within the action segment retains its original posture logic while maintaining consistency with the global reference position, ultimately obtaining the translated and aligned joint coordinates.
[0034] In a feasible embodiment, in step 160, after obtaining the joint coordinates of each action segment after translation and alignment, the translation-aligned action frames in all action segments can be sequentially connected and integrated according to their original time order to form a corrected shadow puppet action sequence. This sequence not only fully preserves the original action posture and coherent logic of each action segment, but also ensures the uniformity and stability of the entire set of actions in the display position through global baseline calibration.
[0035] In a feasible embodiment, during the process of extracting keywords and performing semantic parsing on instruction information to obtain corresponding text instructions, keyword extraction and semantic parsing can be performed first to extract keywords related to the core needs from the instruction information. Then, these keywords are combined to interpret the overall semantics of the instruction information, thereby obtaining text instructions that accurately correspond to the user's instruction information. If a processing timeout occurs during the semantic understanding process, or if the intent corresponding to the instruction information is not successfully matched, the text instruction can be automatically set as a normal action instruction, that is, a pre-set instruction corresponding to basic shadow puppet actions applicable to most common scenarios.
[0036] In one feasible embodiment, after acquiring the user's input instruction information, appropriate response information can be generated and displayed based on the instruction information. Specifically, while parsing the core meaning of the instruction information and the user's request, and clarifying the textual instructions corresponding to the actions the user expects the shadow puppet digitizer to perform, response information that conforms to the dialogue logic and fits the intangible cultural heritage dissemination scenario can be generated by combining the cultural context and interactive scenario of shadow puppet performance. After generating the response information, it can be output to the user through the output module of the interactive device (such as a display screen or speaker), which can not only provide timely feedback on the user's input request, but also realize real-time interaction with the user, making the interaction process more coherent and immersive. If a processing timeout occurs during semantic understanding, or if the intent corresponding to the instruction information is not successfully matched, preset response information can be generated. This preset response information is a pre-configured standardized feedback content that can promptly provide feedback on the processing result to the user, ensuring that the interaction process is not interrupted and guaranteeing the user experience.
[0037] In one feasible embodiment, such as Figure 2 As shown, the process of constructing the action mapping table may include at least steps 210 to 240.
[0038] Step 210: Collect the sequence of human performance movements, which includes multiple performance frames, each corresponding to multiple skeletal key points of the human body; Step 220: Select the core skeletal key points of the shadow puppet digital human from multiple skeletal key points, and perform coordinate normalization on the core skeletal key points using the target reference joint coordinates as the reference point to obtain the processed key point coordinates. Step 230: Obtain the custom performance action sequence based on the processed keypoint coordinates; Step 240: Associate the custom performance action sequence with the corresponding semantic instructions and add it to the initial mapping table to obtain the action mapping table.
[0039] In a feasible embodiment, in step 210, the motion sequence of the actual human performance can be captured by a motion acquisition device. The motion sequence consists of multiple consecutive performance frames. Each performance frame records the position information of multiple skeletal key points (including 33 points such as head, shoulder, elbow, hip, and knee) of the human body. The position data of these skeletal key points can fully reflect the human body's posture in that frame.
[0040] In a feasible embodiment, in step 220, after obtaining the position information of multiple skeletal key points of the human body, the core skeletal key points required for the shadow puppet digital human can be selected. These core skeletal key points are the key parts that can support the shadow puppet's movements and accurately reproduce the core posture of the shadow puppet performance. After the selection is completed, the coordinates of the target reference joint (i.e., the pelvic joint) are used as a unified reference point, and the coordinates of these core skeletal key points are processed using a coordinate normalization algorithm. This processing can eliminate the coordinate differences caused by the movements of performers of different body types, ensuring that the movements of users of different body types can achieve a consistent mapping, and finally obtaining the coordinates of the key points after coordinate normalization.
[0041] In a feasible embodiment, in step 230, after obtaining the coordinates of the key points after coordinate normalization, these coordinate data can be further organized and adapted to convert them into a format that conforms to the logic of shadow puppet digital human action presentation, thereby obtaining a custom performance action sequence. This custom performance action sequence can restore the actions of human performance and adapt to the display requirements of shadow puppets.
[0042] In a feasible embodiment, in step 240, the custom performance action sequence obtained above can be associated with the corresponding semantic instructions, that is, the user input semantic instructions (such as "dance" or "salute") corresponding to the custom performance action sequence can be identified. Then, the association is added to the initial mapping table. By continuously adding the association data between various custom performance action sequences and corresponding semantic instructions, a complete action mapping table is finally formed. This action mapping table can quickly query the corresponding shadow puppet action sequence when the user inputs semantic instructions in the future.
[0043] In one feasible embodiment, the core skeletal key points include at least one of the following: head key point, left shoulder key point, right shoulder key point, left elbow key point, right elbow key point, left wrist key point, right wrist key point, left hip key point, right hip key point, left knee key point, right knee key point, left ankle key point, and right ankle key point.
[0044] In a feasible embodiment, during the baseline determination process for the shadow puppet action sequence, if it is detected that the first frame fails to acquire valid pelvic joint coordinates (i.e., the first frame lacks pelvic joint anchor points), the first valid frame appearing in the shadow puppet action sequence, i.e., the action frame that can completely acquire the joint coordinate data of each joint of the shadow puppet digitizer, can be selected as the basis. Based on this joint set, an approximate baseline can be determined through a preset calculation method. After determining the approximate baseline, the first valid frame is then centered according to the same logic as when the first frame has pelvic joint anchor points.
[0045] In one feasible embodiment, before rendering the corrected shadow puppet action sequence, joint interpolation can be performed on the corrected shadow puppet action sequence to obtain a smooth and continuous shadow puppet action sequence. Specifically, for the joint coordinates of adjacent action frames in the action sequence, a preset interpolation calculation method can be used to supplement the joint coordinate data of the intermediate transition, fill the gaps in the joint positions between adjacent action frames, and make the action that may have had a slight stutter become more coherent and smooth. Finally, a smooth and continuous shadow puppet action sequence without any stutters is obtained, improving the smoothness and comfort of watching the performance.
[0046] In one feasible embodiment, such as Figure 3 As shown, the process of rendering the corrected shadow puppet action sequence may include, but is not limited to, steps 310 to 330.
[0047] Step 310: Load the materials for each component of the shadow puppet digital human; Step 320: Map the joint coordinates of the motion frames in the corrected shadow puppet motion sequence to the corresponding shadow puppet parts; Step 330: Correct the midpoint position of the shoulder, the shoulder offset, and the elbow height of the shadow puppet parts, and render each shadow puppet part frame by frame based on the preset frame rate.
[0048] In a feasible embodiment, in step 310, before performing the shadow puppet digital human interpretation rendering, the various component materials of the shadow puppet digital human can be preloaded. These component materials may include the corresponding split digital materials of the shadow puppet character's head, torso, limbs, clothing and decorations, etc. During the loading process, the component materials can be verified to ensure that the materials can be used normally for subsequent motion mapping and rendering processing.
[0049] In a feasible embodiment, in step 320, after loading the materials of each component, the joint coordinates of each action frame in the corrected shadow puppet action sequence can be mapped to the corresponding shadow puppet component. That is, according to the preset association between the joint and the shadow puppet component, the coordinate position of each joint is used as the placement reference of the corresponding shadow puppet component, so that the shadow puppet component can present the corresponding posture according to the change of the joint coordinates, ensuring that the position of the shadow puppet component corresponding to each action frame is consistent with the joint movement logic of the action sequence.
[0050] In a feasible embodiment, in step 330, after completing the mapping of joint coordinates to the shadow puppet parts, the midpoint position of the shoulder, the shoulder offset, and the elbow height of the shadow puppet parts can be specifically corrected. By fine-tuning the parameters of these key parts, the limb posture of the shadow puppet is made more in line with the performance form of traditional shadow puppetry. After the correction is completed, each shadow puppet part is rendered frame by frame based on a preset frame rate. During the rendering process, the connection and coordination of each part can be maintained, so that the final shadow puppet performance movements are smooth and natural.
[0051] In one feasible embodiment, this application provides a shadow puppet digital human interaction system, such as... Figure 4 As shown, the interactive system 400 includes: Input module 410 is used to acquire instruction information input by the user; The action query module 420 is used to query the corresponding shadow puppet action sequence from the preset action mapping table according to the instruction information. The shadow puppet action sequence includes an initial action segment and at least one subsequent action segment, and each action segment contains multiple consecutive action frames. The joint correction module 430 is used to take the target reference joint of the first frame in the shadow puppet action sequence as the baseline; determine the target reference joint coordinate difference between the start and end frames of the initial action segment and the subsequent action segment; and translate and align the joint coordinates of the action segment according to the baseline and the corresponding target reference joint coordinate difference to obtain the translated and aligned joint coordinates. The motion rendering module 440 is used to obtain the corrected shadow puppet motion sequence based on the joint coordinates after translation and alignment, and to render the corrected shadow puppet motion sequence to obtain the dynamic performance of the shadow puppet digital human.
[0052] In one feasible embodiment, the interactive system also includes a semantic parsing module, which is used to extract keywords and perform semantic parsing on the instruction information to obtain the corresponding text instructions; and to generate appropriate response information based on the instruction information.
[0053] In one feasible embodiment, the motion rendering module can also display response information (such as reply text).
[0054] In one feasible embodiment, the interactive system also includes a motion sampling module. This module is based on the MediaPipe open-source algorithm and is specifically improved and optimized to meet the particular needs of shadow puppet digital human posture recognition. As a pre-processing module of the entire shadow puppet digital human interactive system, this module can capture the motion trajectory and posture data of various joints of the human body before the subsequent core processes such as semantic parsing and motion generation are initiated. The collected motion information can provide raw data support for the expansion of the motion mapping table and the generation of custom motion sequences.
[0055] It should be noted that the functional execution logic of each module in this interactive system is consistent with the processing logic of the corresponding embodiments of the aforementioned interactive methods. For the specific workflows and related technical details of each module, please refer to the corresponding descriptions in the aforementioned interactive methods; they will not be repeated here.
[0056] In one feasible embodiment, the shadow puppet digital human interaction system is based on multi-agent technology and edge computing architecture. Through the Autogen multi-agent framework, it coordinates text-generating agents, action-generating agents, and user proxy agents to efficiently achieve core functions such as semantic script parsing, action command generation, and character interaction responses. The system incorporates SLM (Small Language Model) and Vosk offline speech recognition modules, enabling real-time parsing of user voice input and action requests even in weak or offline network environments, effectively ensuring the continuity and stability of the shadow puppet performance. It also features the MediaPipePose posture recognition component, which can accurately recognize single-person postures and capture collaborative postures of multiple users. Furthermore, it can quickly map the recognized posture data to the shadow puppet component skeleton, achieving precise action-driven operation of the shadow puppet digital human. Internally, the system can also predefine slot structures and style / action segment association mechanisms to construct a complete and unified link from semantic parsing to visual rendering. In addition, the system provides scalable interfaces and standardized data formats (such as JSON, PNG, and FPS standards), which can be flexibly adapted to various application scenarios such as exhibition hall display, cultural performances, and intangible cultural heritage education.
[0057] Understandably, the text generation agent is a core functional module focused on semantic parsing and text feedback generation. It receives text commands from the input management module and generates appropriate response text by analyzing the user's core input needs, contextual logic, and interaction scenarios. During its operation, it can optimize language expression by incorporating the background of shadow puppetry culture and the scenario of intangible cultural heritage dissemination, ensuring that the response text conforms to dialogue logic while conveying information related to traditional shadow puppetry. It can also generate guiding text to confirm user needs and ensure the accuracy of subsequent interactions for semantically ambiguous or unclear input content. The final generated response text can be output to the user through the system's audio components or display screen, achieving real-time interactive feedback. The action generation agent is a functional module responsible for action command conversion and action sequence planning. It can work collaboratively with the text generation agent, determining the types of actions the shadow puppet digital human needs to perform based on the semantically parsed core user needs. This intelligent agent can quickly query the corresponding shadow puppet action sequence from a preset action mapping table. Simultaneously, it adapts the action sequence based on action continuity requirements (such as optimizing action transitions and duration). If a completely matching action sequence cannot be found, it can also combine basic action fragments into a new action sequence based on preset action combination rules, ensuring accurate response to user action commands and providing complete action data support for subsequent joint correction and rendering. The user agent intelligent agent can serve as an interaction bridge between the user and various functional modules within the system. It can receive text / voice input signals from the input module, standardize the signal format, and distribute it to the text generation agent and action generation agent. It can also monitor the working status of each module in real time. If abnormal situations such as semantic parsing timeout or action sequence query failure occur, it can automatically trigger the fault tolerance mechanism (such as calling the preset response text or enabling the basic action sequence) to ensure the stable operation of the system. At the same time, it can also integrate the response text of the text generation agent and the action sequence data of the action generation agent and transmit them synchronously to the subsequent action driving and rendering modules to achieve full-process collaboration of semantic parsing, text generation, action planning, and rendering display, ensuring smooth and natural interaction between humans and shadow puppet digital humans.
[0058] In one feasible embodiment, the system's interaction can collect various user commands through input modules (including text input and voice input modules). After semantic parsing of the commands by the Autogen architecture system, corresponding behavior sequences are generated. Finally, the action-driven module and rendering module realize the shadow puppet character's action performance and voice feedback, enabling natural and smooth interaction between humans and digital shadow puppet characters. The overall system structure is as follows: Figure 5 As shown, its composition includes, but is not limited to, the following components: Camera: Installed on the "eyebrow" crossbeam, the tilt angle can be adjusted within the range of -10° to +25°, with 5° scale markings and positioning holes for easy and precise angle fixing; Display screen: Sizes are available from 27 to 43 inches, compatible with VESA 100×100 (also compatible with 200×100) installation standards, and can be flexibly selected according to the application scenario; Processor / Memory: Employs Jetson series edge devices or equivalent hardware, with built-in local inference, rendering, and ASR (speech recognition) modules; Power supply and main switch: Supports AC220V mains power input, with an independent main switch and fuse holder on the side of the base; Uprights and cable trays: The uprights have vertically built-in cable trays with a cross-sectional dimension of ≥20×30mm, and cable outlets are provided at both the top and bottom. Main unit tray and ventilation hole array: The main unit tray is set on the back of the column, and the ventilation hole array at the corresponding position has an opening rate of ≥20%, which can ensure the heat dissipation effect when the equipment is running. Base: Equipped with a counterweight of ≥5kg to ensure the system is stable and does not tip over. The corners of the base are rounded with R≥2mm to improve safety during use. Audio components: include speakers and microphones. The speakers can clearly output response voices and performance sound effects, and the microphones can accurately capture user voice input. Form factor: The overall design adopts a floor-standing, vertical design, which occupies a small area and is easy to deploy flexibly in various scenarios.
[0059] The following is combined with Figure 6 The overall workflow of the shadow puppet digital human interaction system is described using a specific embodiment.
[0060] The system's input module supports both text input and speech input. The speech recognition component uses the Vosk Model for Chinese Speech Recognition, accurately capturing the user's speech information and converting it into text. The semantic parsing module consists of a text generation agent and an action generation agent, specifically designed to perform core tasks such as text understanding, action generation, and response generation. It can accurately interpret the semantic requests of the user input and generate action commands and responses adapted to the scenario. The action query module can quickly index files using a predefined action mapping table. This table stores the association between various semantic commands and corresponding action files. Each action file contains multiple consecutive frames of joint coordinate sequences, which can be used to explicitly define the joint coordinates. The system tracks the two-dimensional motion trajectories of various parts of the shadow puppet (such as the head, shoulders, elbows, knees, and ankles). The joint correction module works as follows: when processing motion sequences, the pelvic joint of the initial frame can be set as the global reference point to ensure the center alignment and pose consistency of the whole body's movements. At the same time, the system can automatically detect detailed parameters such as shoulder offset and elbow height of the shadow puppet's movements and make targeted corrections to these parameters, so that different motion sequences can maintain natural continuity when splicing together, avoiding stuttering or posture imbalance. The motion rendering module is implemented based on the Pygame library, which can not only render and present the various parts of the shadow puppet character, but also perform joint interpolation calculations and frame animation control, making the motion transitions smoother.
[0061] The complete workflow of the entire system is as follows: (1) Receive user input signals (text input or voice input) through the input module; (2) Transmit the input signals to the semantic parsing module for semantic parsing to clarify the user's core needs; (3) Generate appropriate response text by the semantic parsing module; (4) Perform action planning based on the parsing results to determine the shadow puppet actions to be performed; (5) Read the corresponding action file from the action mapping table through the action query module; (6) Generate continuous action frames based on the joint coordinate sequence in the action file; (7) Render and display the action frames through the action rendering module, and output the response text as feedback in a synchronous manner.
[0062] In one implementation, such as Figure 7As shown, the input module provides users with two convenient input modes, which can flexibly adapt to different usage scenarios and user habits. Specifically: Text input mode: Users can directly type text commands through the on-screen input box of the interactive system. Text commands can include various requests related to the shadow puppet digital human's actions, such as "dance," "bow," and "walk forward." After typing, pressing the Enter or Tab key will submit the text command. The system will transmit the text string entered by the user to the Agent interface in real time, providing basic data for subsequent semantic analysis. Voice input mode: Users can record voice signals through the microphone provided by the system. After recording, the system will use the offline speech recognition engine Vosk to accurately recognize and process the collected voice signals. After recognition and parsing, the system will obtain text commands that correspond one-to-one with the voice content, and then transmit the text commands to the Agent interface simultaneously. Both input methods are encapsulated and scheduled by a unified input management module. Users can freely switch between input modes using the "Space" button or the mode switching button on the interface, achieving seamless integration between the two input methods and ensuring convenient and efficient operation. If recognition fails or no valid speech content is collected during the speech recognition process, the system will briefly display the message "No content recognized" in the upper right corner of the interface, promptly providing feedback to the user on the current processing result. This allows users to quickly adjust their input method and resubmit the command, ensuring a smooth interactive process.
[0063] In one implementation, such as Figure 8As shown, the semantic parsing module is built on the Autogen framework and achieves its core functions through efficient communication between agents. In the semantic processing stage, the instructions transmitted by the user agent are passed to the text generation agent for high-level semantic parsing. During parsing, the core intent of the user can be accurately extracted, and keywords related to the request can be extracted. After parsing, the results are structured into a unified JSON format, ensuring that subsequent modules can easily read and process them. In the action generation stage, the structured data after semantic parsing can be passed to the action generation agent. This agent can retrieve matching shadow puppet action sequences from a predefined action mapping table. Action sequences can include various suitable actions such as "bowing," "dancing," "bowing," and "normal state." Simultaneously, the system can call the default fallback mechanism. If semantic parsing timeouts or action sequence retrieval failures occur, the system can automatically switch to a "normal" or "idle" state to ensure stable system operation. The text generation agent and the action generation agent communicate locally through the Autogen framework. This communication method significantly improves the real-time performance of semantic processing and action generation, while enhancing the system's independence and reducing dependence on the external environment. The output of this semantic parsing layer can be passed to the action sequence generation module in an orderly manner via a queue, efficiently achieving a precise mapping from semantics to actions. After collaborative processing by two intelligent agents, a structured result of (acts, reply) is returned. For example, if a user inputs, "What are the highlights of this digital museum exhibition?", after processing by the semantic parsing module, the output will be: ([waving arms and bowing], the exhibition hall focuses on the comparison between Song Dynasty shadow puppetry and modern art creation, highlighting the cultural style of ancient China). Here, [waving arms and bowing] is the action sequence that the digital shadow puppet needs to perform, output by the module; "the exhibition hall focuses on the comparison between Song Dynasty shadow puppetry and modern art creation, highlighting the cultural style of ancient China" is the returned response text. This response text will be displayed on the interface at the front end of the interactive system through the rendering module for the user to view intuitively, while simultaneously achieving complete interactive feedback in conjunction with the action performance.
[0064] In one implementation, such as Figure 9As shown in the diagram, this illustrates the complete motion processing flow comprised of the motion sampling module, motion query module, joint correction module, and motion rendering module. After the process starts, a JSON-formatted motion file is generated via a sampling script. Each motion file contains multiple JSON-formatted motion segments, each corresponding to a specific action of the shadow puppet digitizer. Subsequently, based on the keypoint frame data within the motion segments, the system sequentially performs position offset calculations and rotational smoothing, while simultaneously performing fixed alignment operations using the pelvic anchor point as a reference to ensure the uniformity of position and the continuity of posture among the motion segments. After the above processing, a continuous and smooth motion sequence is formed. Finally, the motion rendering module renders each frame in the motion sequence.
[0065] In one implementation, the motion sampling module uses the MediaPipe algorithm as its core technology foundation and makes targeted improvements and optimizations based on the actual needs of the shadow puppet digital human's motion performance. From the joints of the entire human body identified by this algorithm, 13 core key points are precisely selected as the key basis for driving the shadow puppet digital human's motion. Figure 10 As shown, the key joints of the human body include: 0. Nose; 1. Inner corner of left eye; 2. Left eye; 3. Outer side of left eyeball; 4. Inner corner of right eye; 5. Right eye; 6. Outer side of right eyeball; 7. Left ear; 8. Right ear; 9. Left side of mouth; 10. Right side of mouth; 11. Left shoulder; 12. Right shoulder; 13. Left elbow; 14. Right elbow; 15. Left wrist; 16. Right wrist; 17. Left little finger; 18. Right little finger; 19. Left index finger; 20. Right index finger; 21. Left thumb; 22. Right thumb; 23. Left hip; 24. Right hip; 25. Left knee; 26. Right knee; 27. Left ankle; 28. Right ankle; 29. Left heel; 30. Right heel; 31. Left index toe; 32. Right index toe. The 13 core key points specifically include: nose, left shoulder, right shoulder, left elbow, right elbow, left wrist, right wrist, left hip, right hip, left knee, right knee, left ankle, and right ankle. By selecting these core key points, the system's computational power requirements for hardware devices can be significantly reduced, lowering hardware deployment costs. Simultaneously, the crucial posture information required for the shadow puppet's movements can be fully preserved, effectively ensuring the accuracy and consistency of the digital shadow puppet's movements. After sampling, the module can generate a standardized JSON skeleton file from the collected key point coordinate data. This file can be directly incorporated into the motion library (motion mapping table) of the digital shadow puppet interaction system, providing raw motion data support for subsequent motion queries, joint corrections, and rendering feedback.
[0066] The motion sampling process is illustrated below with a specific example.
[0067] To meet the needs of users to customize actions, this embodiment provides an action sampling extension module based on the MediaPipe algorithm. Users can collect their own performance actions and convert them into action frame data that can be recognized by the shadow puppet digital human, thereby realizing personalized shadow puppet mapping.
[0068] The sampling process is as follows: (1) Sampling Start: When the user selects the "Motion Sampling" mode in the system interface, the system calls the camera module to track key points of the human body and uses the improved MediaPipe Pose Estimation algorithm to obtain skeletal key points (33 points, including head, shoulder, elbow, hip, knee, etc.). (2) Data processing and filtering: To ensure the smoothness and consistency of the shadow puppet animation, this embodiment adds the following targeted optimization steps based on the original MediaPipe pose detection results: The sliding window averaging algorithm can be used to smooth the detected joint coordinates, effectively eliminating jitter or data fluctuations that may occur during motion capture, making the transition of shadow puppet movements more natural and smooth. The coordinates of all joints in the body can be normalized using the coordinates of the pelvic joint as the center. This operation can eliminate the coordinate differences caused by the movements of users of different body types, ensuring that the performance movements of users of different body types can be consistently mapped to the shadow puppet digital human. The collected three-dimensional coordinates (x, y, z) can be mapped to the two-dimensional plane (x, y) required for the shadow puppet display using a coordinate compression algorithm, while also matching the scale to the traditional proportions of the shadow puppet.
[0069] (3) Motion data generation: The system saves the sampled action sequences as JSON files, with an example structure as follows: { "video_info":{"resolution":[1920,1080]}, "frames":[ {"joints":{"pelvis":{"x":512,"y":384},"left_arm":{...}, ...}} ] } Each action frame corresponds to a set of joint coordinates, which can be directly stored in the actions / directory for the main program to call.
[0070] (4) Shadow puppet character mapping and display: Newly generated motion files (such as JSON files generated after human motion is collected by the motion sampling generation module) can be directly loaded by the main program of the interactive device. During the loading process, the system can automatically perform offset calculations and precise alignment processing on the joint data in the motion file to ensure that the new motion is compatible with the skeletal structure and component association logic of the digital shadow puppet. After receiving the loaded new motion data, the character rendering module can completely map the new motion onto the digital shadow puppet character. It can not only restore the core posture and movement trajectory of the new motion, but also optimize the details in combination with the artistic style of shadow puppetry, so that the new motion and the performance form of the shadow puppet character are highly consistent, ultimately realizing a shadow puppet performance with personalized characteristics.
[0071] In one implementation, users can record personalized actions such as "bowing," "waving," and "clapping" based on their own performance, which are then immediately mapped onto the digital shadow puppet character. In museums, visitors can use this function to record their own actions and watch the digital shadow puppet "reproduce" the performance in the same pose, thus achieving an interactive intangible cultural heritage experience.
[0072] In one implementation, such as Figure 11 As shown, the action query module can directly map action identifiers (i.e., text commands) to the corresponding action file paths, which are stored in JSON format. For example, the action identifier "bow" can be mapped to the file path "actions / bow.json", and the action identifier "dance" can be mapped to the file path "actions / dance.json". Through a predefined action mapping table, these action files can be quickly indexed and retrieved. Each action file contains multiple frames of continuous joint coordinate sequences. These sequences can clearly define the two-dimensional motion trajectories of various parts of the shadow puppet (such as the head, shoulders, elbows, knees, and ankles), and also completely reconstruct the posture change logic of the corresponding action.
[0073] In one implementation, such as Figure 12 As shown, during the mapping process, a "joint-part" mapping relationship is established by parsing each frame of the motion file, and coordinate translation and interpolation calculations are performed to ensure smooth and continuous motion. The initial frame sets the pelvic joint as the reference point to ensure center alignment and pose consistency of the whole-body motion. The system automatically detects and corrects details such as shoulder offset and elbow height to ensure that different motion sequences remain natural and coherent when stitched together.
[0074] In one implementation, such as Figure 13As shown, the Pygame library is used to implement real-time rendering of the shadow puppet digitizer. This module can load the corresponding material images of each part of the shadow puppet from the part texture folder, and then draw the complete shadow puppet character in real time on the interactive interface screen through coordinate mapping. During the rendering process, the system can use a fixed frame rate control mechanism, setting the frame rate to 45 FPS (i.e., 45 Hz), which ensures both the smoothness of the shadow puppet's movements and the stability of the visual effects, avoiding screen stuttering or jitter. At the same time, user-inputted text replies or various prompts generated by the system can be displayed in a semi-transparent text box above the shadow puppet character, which neither obscures the display of the shadow puppet's movements nor prevents the user from clearly viewing the interactive content.
[0075] The following describes an embodiment of this application using a specific application scenario.
[0076] In the exhibition hall of an intangible cultural heritage museum, the digital shadow puppet interactive system of this embodiment can be deployed as an interactive display device. Visitors can interact naturally with the digital shadow puppets through voice or text, thereby achieving an immersive experience and dissemination of intangible cultural heritage.
[0077] The system operation process is as follows: Upon startup, the main program loads the speech recognition model (Vosk) and the digital human intelligent agent module. The screen displays the initial normal posture of the shadow puppet and prompts the user: "Press the space bar to speak or press the Tab key to enter a question." When the audience issues a command via voice (such as "Please dance a short piece"), the system executes the following steps: (1) Speech acquisition and recognition: The microphone module initiates a recording stream to acquire real-time audio signals. After sampling, the audio is fed into the Vosk model for offline recognition and outputs the recognized text result. For example, if the word "dance" is displayed, the system will briefly show "Recognition: dancing" in the upper right corner of the screen. (2) Semantic understanding and intent recognition: The system will identify the text input data processing agent module. This module will determine the input intent as a "performance request" based on the text-generated agent and match it with the "dance" action from the action library; (3) Motion generation and visual presentation: The system loads the frame sequence from the actions / dance.json file, reads the coordinates of each joint of the shadow puppet character (such as pelvis, shoulder, elbow, etc.), and performs smooth offset processing on the bone position to ensure that the dance movements are naturally connected with the original position. (4) Rendering module: The system refreshes at 45 frames per second, mapping the motion frame by frame onto the shadow puppet character textures to achieve dynamic dance performances.
[0078] (5) Semantic feedback and user response: Simultaneously, the text generation agent generates the semantic response text "Okay, I'll perform a shadow puppet dance," which is displayed in the dialog box at the top of the screen. After the action is completed, the system automatically returns to the "normal" state.
[0079] Understandably, compared with existing virtual human systems that rely on large cloud models and pre-made animations (such as Tencent's "Smart Shadow Digital Human" and ByteDance's "OmniHuman"), the shadow puppet digital human interactive device proposed in this application adopts a multi-agent local autonomous architecture and edge computing deployment, which can realize the offline operation of the entire process of semantic understanding, action generation and rendering. With Vosk offline speech recognition and SLM small language model, it can stably complete semantic parsing and interactive response even in the absence of network or weak network environment, which significantly improves the stability and real-time performance of the device in edge scenarios such as exhibition halls, touring exhibitions, and schools. Compared with solutions that rely on cloud inference, it has lower latency, higher privacy and stronger deployability. Furthermore, this application possesses significant advantages in system structure and interaction mechanisms: through a dual-agent collaborative mechanism of Text-Agent and Action-Agent, the system can accurately generate matching action sequences based on user semantics, constructing a unified "semantic-action-rendering" link; with the improved MediaPipe algorithm and pelvic anchor point alignment mechanism, it can ensure natural splicing and smooth performance of different action segments; the standardized JSON action file format and ACTION_MAP mapping mechanism support personalized action extensions by users, forming an open and evolvable shadow puppet action library. Overall, this application comprehensively surpasses existing technologies in terms of real-time performance, cultural adaptability, scalability, and user engagement, effectively realizing the intelligent and sustainable digital interpretation of intangible cultural heritage shadow puppetry art.
[0080] In one feasible embodiment, the core engine of the speech recognition module can be Whisper local inference, small model Kaldi, self-trained Conformer small model, or domestic offline engines such as small ASR for mixed Chinese and English reading, to meet the technical selection needs of different scenarios. In environments with strong noise interference, the system can add VAD (voice endpoint detection) function, or use noise reduction preprocessing methods such as spectrum subtraction, thresholding, and lightweight neural noise reduction to effectively improve the accuracy of speech recognition. If there are restrictions on voice use in the application venue, users can switch the input method to touch screen input, gamepad operation, QR code H5 page input, or use gesture recognition as the command input method to ensure that the interactive function is not restricted by the venue conditions and adapts to more practical application scenarios.
[0081] In a feasible embodiment, the dual-agent architecture of Text-Agent and Action-Agent can be adapted as needed to implement the following: a. a single small language model (SLM) combined with a rule fallback mechanism (such as regular expressions / keyword lists); b. a finite state machine (FSM) / behavior tree (BT) combined with small model error correction functionality; c. a lightweight classifier (for intent / emotion / topic recognition) used in conjunction with an action retrieval unit.
[0082] In a feasible embodiment, the action matching method between semantics and action mapping table can be replaced by the implementation scheme of "semantics → embedding similarity retrieval (based on vector library) → candidate action ranking", or a lightweight policy network (such as a small GRU / LSTM / Transformer) can be used to plan the splicing of existing action segments. Both schemes support local deployment and can run stably on low computing power devices, while still retaining the fault tolerance mechanism of "idle / normal" rollback when timeout / miss.
[0083] In one feasible embodiment, the MediaPipe Pose pose recognition scheme can be replaced with BlazePose, MoveNet, RTMPose, OpenPose (adapted to multi-person scenarios), or a depth camera (such as Kinect / RealSense) paired with a lightweight skeleton fitting algorithm; in single-person scenarios, a 13-point reduction scheme can be used to control computing power consumption, while in multi-person scenarios or complex occlusion environments, the model can be upgraded or a skeleton tracker (such as Kalman / OneEuro) can be introduced to improve recognition stability.
[0084] In one feasible embodiment, the motion data acquisition method can be changed from camera sampling to input from IMU, controller, Leap Motion, inertial capture device, or by importing a third-party dataset and obtaining motion data after secondary normalization and format conversion.
[0085] In one feasible embodiment, the storage format of motion data can be replaced with CSV or binary protobuf / flatbuffers, and the skeleton representation can be replaced with BVH (which requires additional 2D projection processing). The above adjustments do not change the overall design concept of standardized motion files and runtime hot loading.
[0086] In a feasible embodiment, the translational alignment method based on the pelvic anchor point can be replaced by the following schemes: a. using the midpoint of the chest / shoulder as the anchor point; b. rigid body alignment based on Procrustes analysis (including translation + proportional scaling + small-angle rotation); c. inter-frame root node trajectory constraint alignment; when anchor point data is missing, the "most recent valid frame" or "multi-joint weighted centroid" can be used as a substitute.
[0087] In a feasible embodiment, the existing frame-by-frame translation and fixed frame rate processing methods can be replaced by a scheme combining OneEuro / Kalman filtering with Catmull-Rom / B spline interpolation; if a more natural end-effector motion effect is required, a light FK / IK constraint (forward / inverse kinematics) can be added to the joint layer, while always maintaining the core artistic style of 2D shadow puppetry.
[0088] In one feasible embodiment, the Pygame graphics library can be replaced with Godot / Unity (2D), OpenGL / SDL2, Qt / QML, or a web frontend (such as Canvas / SVG / WebGL / Three.js). If cross-platform and remote screen display is required, the implementation logic of locally calculated skeletal data, streaming point information via WebSocket / UDP, and frontend H5 / SVG assembly rendering can be used.
[0089] In one feasible embodiment, PNG / JPG format texture materials can be replaced with SVG vectors or layered Spine / SpritAtlas; if visual effects need to be optimized, paper textures, projection feathering and other light and shadow / material processing can be added through shaders without changing the inherent aesthetic style of 2D shadow puppets.
[0090] In one feasible embodiment, the hardware deployment platform can replace Jetson / local PC with x86 mini-PC, Raspberry Pi + NPU accelerator stick, Android all-in-one machine (with built-in local model), etc.; the microphone / camera can be replaced with array microphone and directional microphone, 4K / wide-angle camera or binocular / ToF depth camera according to the site requirements, to adapt to the hardware configuration requirements of different environments.
[0091] In one feasible embodiment, in addition to existing touch and motion-sensing interaction methods, interactive components such as foot switches, physical buttons, NFC tags (for triggering fixed storylines), and laser rangefinders (for entry triggering) can be added.
[0092] In one feasible embodiment, the system operates primarily in a fully local processing mode. If cloud capabilities are required (such as updating the story knowledge base), a hybrid mode of "offline as the default and online as an advantage" can be adopted, and local caching and offline rollback strategies can be configured to ensure that the core purpose of the invention, "usable in weak network conditions," is not deviated from.
[0093] It should be noted that none of the above alternatives change the core technical concept of this application: to achieve real-time interactive performance of the shadow puppet digital human through semantic-driven motion choreography and skeletal mapping at the edge. Any equivalent changes made in input methods, semantic / motion planning, pose acquisition, alignment / interpolation processing, rendering framework, data format, and hardware platform should be considered equivalent replacements or technically equivalent solutions of this application and fall within the protection scope of this application.
[0094] This application also discloses an electronic device, which includes a processor, a memory, and a computer program stored in the memory and executable by the processor. When the computer program is executed by the processor, it implements the shadow puppet digital human interaction method described above.
[0095] This application also discloses a computer-readable storage medium storing a processor-executable program, which, when executed by a processor, is used to perform the shadow puppet digital human interaction method described above.
[0096] The above description of the disclosed embodiments enables those skilled in the art to make or use this application. Various modifications to these embodiments will be readily apparent to those skilled in the art, and the general principles defined herein may be implemented in other embodiments without departing from the spirit or scope of this application. Therefore, this application is not to be limited to the embodiments shown herein, but is to be accorded the widest scope consistent with the principles and novel features disclosed herein.
Claims
1. A method for interactive use of a shadow puppet digital human, characterized in that, include: Obtain user input instructions; According to the instruction information, the corresponding shadow puppet action sequence is queried from the preset action mapping table. The shadow puppet action sequence includes an initial action segment and at least one subsequent action segment. Each action segment contains multiple consecutive action frames. The target reference joint in the first frame of the shadow puppet action sequence is used as the baseline; Determine the target reference joint coordinate difference between the start and end frames of the initial action segment and the subsequent action segment; Based on the baseline and the corresponding target reference joint coordinate difference, the joint coordinates of the motion segment are translated and aligned to obtain the translated and aligned joint coordinates; The corrected shadow puppet action sequence is obtained based on the joint coordinates after translation and alignment, and the corrected shadow puppet action sequence is rendered to obtain a dynamic performance of the shadow puppet digital human.
2. The method according to claim 1, characterized in that, The steps for constructing the action mapping table include: Collect human performance action sequences, the performance action sequences include multiple performance frames, each performance frame corresponds to multiple skeletal key points of the human body; The core skeletal key points of the shadow puppet digital human are selected from the multiple skeletal key points, and the coordinates of the core skeletal key points are normalized using the target reference joint coordinates as the reference points to obtain the processed key point coordinates. Based on the processed keypoint coordinates, a custom performance action sequence is obtained; The custom performance action sequence is associated with the corresponding semantic instructions and added to the initial mapping table to obtain the action mapping table.
3. The method according to claim 2, characterized in that, The core skeletal key points include at least one of the following: head key point, left shoulder key point, right shoulder key point, left elbow key point, right elbow key point, left wrist key point, right wrist key point, left hip key point, right hip key point, left knee key point, right knee key point, left ankle key point, and right ankle key point.
4. The method according to claim 1, characterized in that, Before rendering the corrected shadow puppet action sequence, the method further includes: performing joint interpolation on the corrected shadow puppet action sequence to obtain a shadow puppet action sequence with continuous and smooth movements.
5. The method according to claim 1, characterized in that, The step of rendering the corrected shadow puppet action sequence includes: Load the materials for each component of the shadow puppet digital human; Map the joint coordinates of the motion frames in the corrected shadow puppet motion sequence to the corresponding shadow puppet parts; The midpoint position of the shoulder, shoulder offset, and elbow height of the shadow puppet parts are corrected, and each shadow puppet part is rendered frame by frame based on the preset frame rate.
6. The method according to claim 1, characterized in that, The step of querying the corresponding shadow puppet action sequence from a preset action mapping table according to the instruction information includes: The instruction information is subjected to keyword extraction and semantic parsing to obtain the corresponding text instructions; The corresponding shadow puppet action sequence is retrieved from the preset action mapping table according to the text command.
7. The method according to claim 1, characterized in that, After obtaining the user-input instruction information, the method further includes: Generate appropriate response information based on the instruction information; The response information will be displayed.
8. A shadow puppet digital human interactive system, characterized in that, include: The input module is used to obtain user-input instructions. The action query module is used to query the corresponding shadow puppet action sequence from a preset action mapping table according to the instruction information. The shadow puppet action sequence includes an initial action segment and at least one subsequent action segment, and each action segment contains multiple consecutive action frames. The joint correction module is used to take the target reference joint of the first frame in the shadow puppet action sequence as the baseline; determine the target reference joint coordinate difference between the start and end frames of the initial action segment and the subsequent action segment; and translate and align the joint coordinates of the action segment according to the baseline and the corresponding target reference joint coordinate difference to obtain the translated and aligned joint coordinates. The motion rendering module is used to obtain the corrected shadow puppet motion sequence based on the joint coordinates after translation and alignment, and to render the corrected shadow puppet motion sequence to obtain a dynamic performance of the shadow puppet digital human.
9. An electronic device, wherein, The electronic device includes a processor, a memory, and a computer program stored in the memory and executable by the processor, wherein the computer program, when executed by the processor, implements the method as described in any one of claims 1 to 7.
10. A computer-readable storage medium storing a processor-executable program, characterized in that, The processor-executable program, when executed by the processor, is used to perform the method as described in any one of claims 1 to 7.