Exhibition narrative adaptive regulation method based on gesture trajectory recognition and real-time rendering of multimedia elements

By extracting key skeletal node sequences of the hand from the exhibition environment and performing spatial resampling and scale normalization, combined with a hybrid gesture recognition model and a hidden Markov model, the problems of low recognition rate of complex gestures and misrecognition of transitional gestures are solved, realizing adaptive control of exhibition content and immersive experience.

CN122454645APending Publication Date: 2026-07-24RESEARCH INSTITUTE OF TSINGHUA UNIVERSITY IN SHENZHEN +1
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
RESEARCH INSTITUTE OF TSINGHUA UNIVERSITY IN SHENZHEN
Filing Date
2026-06-24
Publication Date
2026-07-24

AI Technical Summary

Technical Problem

Existing gesture recognition technology struggles to understand users' deeper narrative exploration intentions in complex exhibition environments, and has a low recognition rate for complex gestures, with frequent misrecognition of transitional gestures, resulting in a stiff exhibition experience and a decrease in immersion.

Method used

By acquiring video data streams, depth images, and electromagnetic radiation signals from the exhibition space, key skeletal node sequences of the hand are extracted, and spatial resampling and scale normalization are performed. Combined with a hybrid gesture recognition model and a hidden Markov model, erroneous operations are filtered out, achieving accurate recognition and adaptive control of gestures.

Benefits of technology

It improves the accuracy of gesture recognition and the continuity of exhibition content, allowing users to explore exhibits intuitively, maintain an immersive experience, block invalid gestures, and enhance the autonomy and story extension of the exhibition.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122454645A_ABST
    Figure CN122454645A_ABST
Patent Text Reader

Abstract

The present application relates to the technical field of space computing, in particular to an exhibition narrative adaptive control method based on gesture trajectory recognition and real-time rendering of multimedia elements. In view of the problems of lack of continuous trajectory and narrative depth coupling, easy interference of operation deformation and transition action mis-triggering in existing digital exhibition interaction, the scheme obtains three-dimensional skeleton trajectory of the exhibition space through multi-modal sensing equipment; utilizes mathematical geometric transformation for space resampling and scale normalization; then inputs the standardized trajectory feature sequence into a hybrid model composed of dynamic time warping algorithm and hidden Markov model for time series inference, and filters the unconscious transition gestures; finally, the extracted intention features are input into the exhibition narrative control engine to trigger state machine node transition, generate narrative advancing instructions to control holographic and other multimedia devices. The present application improves the robustness of three-dimensional space continuous gesture recognition, and realizes smooth interactive deduction of digital storylines in the exhibition space.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of spatial computing technology, specifically to an adaptive control method for exhibition narrative based on gesture trajectory recognition and real-time rendering of multimedia elements. Background Technology

[0002] With the rapid advancement of display technology, the lag in human-computer interaction methods has gradually become a significant bottleneck restricting the upgrading of exhibition experiences. Early digital exhibitions relied heavily on mice, keyboards, or physical touchscreens for navigation and control. As the demand for contactless interaction increased, gesture recognition-based interaction methods emerged. Current typical gesture recognition applications mostly acquire user gestures, input them into a mapping model to obtain single interaction commands (such as zoom in, zoom out, or next page), and then execute the corresponding operation on the terminal device's interface.

[0003] Although this discrete gesture command mapping mechanism frees up the user's hands to some extent and completes basic gesture recognition on the display interface by building a key skeletal model of the human hand, the following problems still exist when faced with complex "narrative presentations": First, traditional interaction is often a superficial "trigger-execution" mechanism. However, in large museums or amusement park exhibits, there are vast, non-linear story trees behind the exhibits. If users can only use discrete "waving" gestures to turn pages, the entire visit will become rigid and fragmented. The system cannot understand the deep narrative exploration intentions represented by the user's continuously drawn complex trajectories (such as drawing a symbol with specific historical meaning in the air).

[0004] Secondly, in open three-dimensional space, when different users draw the same trajectory, their hand movement speed, the absolute spatial size of the trajectory, and the initial deflection angle all vary greatly. Traditional static image classification neural networks cannot effectively handle this feature that includes temporal dynamic changes. Although the Dynamic Time Warping (DTW) algorithm has advantages in handling sequence matching with misaligned time axes, in practical applications, the original DTW algorithm, without rigorous spatial normalization and structuring, has extremely high computational complexity and is highly susceptible to interference from spatial displacement and scale changes, resulting in low real-time recognition rates for complex character gestures and multi-stroke trajectories.

[0005] Third, in a real exhibition environment, user actions are highly unpredictable. Between two effective interactive actions, users often exhibit natural physical shifts; these meaningless connecting movements are called "transitional gestures." Existing systems, lacking contextual analysis mechanisms, are highly susceptible to misinterpreting these transitional gestures as control commands, thereby interrupting the current narrative playback and severely damaging the sense of immersion. Summary of the Invention

[0006] To address the aforementioned issues, this invention provides an adaptive control method for exhibition narrative based on gesture trajectory recognition and real-time rendering of multimedia elements.

[0007] An adaptive control method for exhibition narrative based on gesture trajectory recognition and real-time rendering of multimedia elements includes the following steps: S1. Acquire video data streams, depth image data, and electromagnetic radiation reflection signals within the exhibition space, and extract the humanoid region and key hand skeletal node sequences of the target user from them; S2. Based on the extracted key hand bone node sequence, continuously track the target user's hand movement within a preset reduced interaction volume, and record a continuous three-dimensional spatial coordinate sequence as the original gesture trajectory data. S3. Perform spatial resampling and scale normalization operations on the original gesture trajectory data, extract the six-degree-of-freedom motion feature vector in Eulerian space, and obtain a standardized trajectory feature vector sequence. S4. Input the standardized trajectory feature vector sequence into the pre-trained hybrid gesture recognition model. The hybrid gesture recognition model outputs the classification result of the target gesture through joint inference of dynamic time warping algorithm and hidden Markov model. S5. Input the classification results into the exhibition narrative control engine. The exhibition narrative control engine generates corresponding narrative advancement instructions based on the context node of the current exhibition narrative, the multimedia operation status of the exhibits, and the classification results. S6. According to the narrative progression instructions, the multimedia playback module, holographic imaging equipment and audio-visual interactive environment of the exhibition site are synchronously controlled to realize the dynamic interpretation of the exhibition content, timeline retracing or spatial hierarchy deconstruction.

[0008] Preferably, step S1 further includes: Electromagnetic waves are emitted towards the target user area by an electromagnetic radiation emitter, and electromagnetic radiation signals reflected back by retroreflective materials on items worn or held by the target user are received by a sensing device. The received reflection signal is aligned with the depth image stream, and the human bounding box of the target user is confirmed by combining it with a human detection network. By performing gesture detection on the human-shaped bounding box, the focus area is determined based on the gesture detection results. If the target area is set as the human-shaped area of ​​the speaker or core participant, the interactive focus and light tracking of the system are locked to the speaker in real time.

[0009] Preferably, the step of continuously tracking the target user's hand movements within a preset reduced interaction volume specifically includes: In augmented reality or virtual reality three-dimensional space, define a local interactive volume that is fixed relative to the user's body torso or relative to the physical display stand; The first neural network is provided with raw sensor data without preprocessing and preprocessed data frames filtered by the interaction volume to identify prior features characterizing subsequent gesture interactions. The system detects the user's waving and bending gestures, and records the bending trajectory of each of the five fingers and the global waving trajectory of the palm.

[0010] Preferably, the spatial resampling and scale normalization operation specifically includes: The original three-dimensional coordinate point set of the gesture trajectory is resampled to obtain a resampled point set, so that the physical curve distance between any two consecutive points along the trajectory is strictly equal. The resampled point set is translated and adjusted in the three-dimensional coordinate system so that the centroid of the translated gesture trajectory point cloud is located at the origin of the three-dimensional coordinate system. Extract the argument value of the first point of the trajectory after translation, and rotate the entire point set in a specified direction according to the argument value to eliminate the spatial difference of the user's initial hand posture; The rotated point set is scaled so that all the scaled coordinate points are within a square or cube bounding box of a specified size. Each trajectory after the above processing is converted into a vector representation in six Euler degrees of freedom in Euler space, and the vector with the largest amplitude is extracted as the dominant motion attribute characterizing the trajectory.

[0011] Preferably, the hybrid gesture recognition model outputs the classification result of the target gesture through joint inference of the dynamic time warping algorithm and the hidden Markov model, specifically including: The standardized trajectory feature vector sequence is compared with the continuous character gesture template and dynamic trajectory style template in the template library using the dynamic time warping algorithm. The candidate template sequence with the smallest dynamic time warping distance is selected as the initial screening result. The initial screening results are input into the Hidden Markov Model. The hidden state transition probabilities of hand shape and hand movement are combined to calculate the matching degree between the input and various known gesture action models, and the gesture category recognition result with the maximum posterior probability is output.

[0012] Preferably, step S5 includes transition gesture recognition logic to filter out erroneous operations: Between any two explicit custom gestures with clear mapping instructions, extract gesture action sequence features and determine whether the current action constitutes a transition gesture without mapping relationship. If the current gesture recognition result is determined to be a transitional gesture, or if two consecutive gesture recognition results are the same, the interaction command corresponding to the current gesture recognition result will be abandoned before the execution time of the corresponding gesture reaches a preset threshold, so as to avoid executing the same gesture continuously in a short period of time and avoid executing unnecessary transitional actions. If the current gesture recognition result is a non-transitional gesture, and the recognition results are different in two consecutive times, then a narrative progression instruction will be output.

[0013] Compared with the prior art, the advantages of this invention are: By introducing geometric resampling and centroid offset alignment algorithms, this scheme has a strong tolerance for changes in the amplitude of the interactor's gestures, waving speed, and relative distance, and its recognition accuracy is higher than that of traditional schemes that rely on image template matching.

[0014] By utilizing HMM and DTW to decode long-term complex gesture intentions, the internal narrative state graph transitions are driven, making digital interactions in the exhibition hall more intuitive and allowing visitors to explore exhibits in a more intuitive way.

[0015] Relying on a transitional gesture filtering mechanism and real-time area focusing technology based on human detection, this solution can stably track the speaker's trajectory and block invalid waving in venues with high traffic and severe background interference, thus maintaining the continuity and user experience of the exhibition content display process. Attached Figure Description

[0016] Figure 1 This is a flowchart of the adaptive control method for exhibition narrative based on gesture trajectory recognition and real-time rendering of multimedia elements proposed in this invention. Detailed Implementation

[0017] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0018] refer to Figure 1 This solution is based on an adaptive control method for exhibition narrative using gesture trajectory recognition and real-time rendering of multimedia elements, including: S1. Acquire video data streams, depth image data, and electromagnetic radiation reflection signals within the exhibition space, and extract the humanoid region and key skeletal node sequence of the target user's hand from them.

[0019] S2. Based on the extracted key hand bone node sequence, continuously track the target user's hand movement within a preset reduced interaction volume, and record a continuous three-dimensional spatial coordinate sequence as the original gesture trajectory data.

[0020] S3. Perform spatial resampling and scale normalization operations on the original gesture trajectory data to extract the six-degree-of-freedom motion feature vector in Eulerian space and obtain a standardized trajectory feature vector sequence.

[0021] S4. Input the standardized trajectory feature vector sequence into the pre-trained hybrid gesture recognition model. The hybrid gesture recognition model outputs the classification result of the target gesture through joint inference of dynamic time warping algorithm and hidden Markov model.

[0022] S5. Input the classification results into the exhibition narrative control engine. The exhibition narrative control engine generates corresponding narrative advancement instructions based on the context node of the current exhibition narrative, the multimedia running status of the exhibits, and the classification results.

[0023] S6. According to the narrative progression instructions, the multimedia playback module, holographic imaging equipment and audio-visual interactive environment of the exhibition site are synchronously controlled to realize the dynamic interpretation of the exhibition content, timeline retracing or spatial hierarchy deconstruction.

[0024] Example 1

[0025] This solution incorporates multi-dimensional data acquisition equipment. Firstly, it employs a high-performance binocular infrared depth camera array, which overcomes the complex and variable lighting conditions within the exhibition hall, accurately extracting environmental noise from the video data stream and acquiring high-contrast image frames. Secondly, for specific interactive exhibits (such as the pointer in the guide's hand), this embodiment deploys electromagnetic radiation emitters. The emitted electromagnetic waves of a specific wavelength illuminate the retroreflective material on the exhibit's surface, and the extremely strong reflected signals are captured by the corresponding sensing receivers. This combined sensing mode can not only record purely manual movements but also capture the precise trajectory displacement of objects, thereby extracting reliable three-dimensional spatial positioning data.

[0026] In addition, multiple data preprocessors were deployed. Some raw sensor data can be directly bypassed and input into the first neural network, which is specifically trained to quickly capture potential motion prior features without sacrificing any physical dimensional accuracy. The video frame sequence, after noise reduction and filtering, is then fed into the skeleton extraction network to locate the 3D coordinates of up to 21 key hand nodes, including the wrist joint and five fingertips. To reduce false triggers and improve the subsequent interactive experience, the system executes a human detection algorithm to confirm the position of all people within the image frame. Only when a user enters the preset interactive focus area, faces the booth, and raises their hand, is their 3D coordinate stream allowed to enter the deep processing queue. At this point, the lights will automatically fine-tune and focus on the speaker.

[0027] Example 2

[0028] After obtaining continuous gesture coordinates, due to the highly arbitrary nature of user operations, the unprocessed raw trajectory exhibits a chaotic state of random distribution across both the time and spatial scales. This solution processes it as follows: S201. Uniform Resampling. In actual hand waving, the hand typically undergoes a physical process of "acceleration-uniform speed-deceleration." This results in very sparse coordinate points in the middle of the trajectory, while the coordinate points at both ends are abnormally dense, under a fixed sensor sampling rate. To smooth out this uneven sampling distribution caused by speed, the total physical length of the entire three-dimensional curve is first calculated. Then, a fixed total number of sampling points N is set (e.g., N=64), and the standard step size is calculated. The algorithm uses a cubic spline interpolation function to re-trim N points on the original trajectory curve. This process ensures that the length of the curve segment between any two adjacent points in the sequence remains strictly consistent (i.e., ...). ).

[0029] S202, Centroid Translation Positioning. Users of different heights wave their hands from different distances from the screen, and their coordinates are incomparable in absolute space. The system calculates the centroid coordinates of 64 resampled points. Then, a translation vector is constructed to subtract the centroid coordinates from the entire set of points, thus re-anchoring it at the origin in Euclidean space.

[0030] S203, Argument Rotation. Different people start with different wrist deflection angles. The system extracts the microscopic tangent direction formed by the first three points in the initial stage of the trajectory and calculates its argument offset value in the local plane. Based on this argument, a three-dimensional rotation matrix is ​​generated and applied to the entire trajectory data. Thus, even circles drawn by the user with a horizontal wrist and circles drawn with a vertical wrist are mapped to the same standard geometric shape.

[0031] S204. Eulerian Space Vectorization Transformation. The extreme values ​​of the coordinates of the point set after translation and rotation are scanned and compressed proportionally into a cube space with a unit volume of 1. At this point, the absolute size of the gesture is completely eliminated. Finally, each stage of local trajectory change is converted into a state vector along six degrees of freedom in Eulerian space, from which the component with the most representative amplitude is selected as the identifier matrix.

[0032] Through the above operations, whether an adult quickly draws a half-meter-long arc in the air or a child slowly slides their finger within a 10-centimeter range, as long as their motion trajectory topology is similar, the final generated standard feature sequences are almost identical, which provides high-quality input data for subsequent algorithms.

[0033] Example 3

[0034] When a user attempts to input a complex, continuous instruction (such as drawing consecutive letter symbols in the air to represent the abbreviation of a component of an exhibit), the trajectory is not only long but also lacks clear breaks between strokes.

[0035] To address this problem, this solution first introduces the Structured Dynamic Temporal Warping (DTW) algorithm. The core of this algorithm lies in constructing a cumulative distance matrix. Given a standardized input sequence X and a pre-defined dictionary template sequence Y, the DTW algorithm can find an optimal matching path that minimizes the sum of local distances, while tolerating a certain degree of time axis distortion. Since the preceding steps have perfectly handled the spatial scale, the DTW algorithm can now efficiently segment the real-time trajectory from the long sequence and compare it with the dictionary, quickly selecting the top three most matching coarse-grained categories such as "spiral ascent" and "Z-shaped return."

[0036] Subsequently, the coarse screening results from DTW are input into a Hidden Markov Model (HMM). The HMM is defined by three key probability matrices: the initial state probability matrix, the state transition probability matrix, and the observation emission probability matrix. During the training phase, the HMM utilizes massive amounts of data to establish a hidden state transition model. In real-time inference, the HMM executes the Viterbi decoding algorithm to find the optimal path with the highest posterior probability among multiple hidden state paths, generating the current mixed observation sequence. Compared to using DTW alone, this hybrid architecture allows the system to not only compare gesture similarity but also evaluate the rationality of the action's evolution in temporal statistics. This enables the scheme to accurately recognize dynamic gestures with extremely low computational consumption and possesses strong generalization ability; it can even be directly used to determine the intentions of new gestures that did not appear in the original training set but exhibit the same statistical evolution patterns.

[0037] Example 4

[0038] When visitors are reading digital display panels or viewing holographic content, if they casually lower their arms or adjust their hair, the traditional single-map system (which immediately translates the acquired gestures into commands) will instantly trigger the interface to jump or close, severely disrupting the immersive experience.

[0039] Therefore, this solution also includes transition gesture recognition logic to filter out erroneous operations. Specifically, a class of "transition gestures" without actual operation mapping is predefined. When the current recognition result is determined to be this type of transition gesture, the transmission of the instruction to the downstream is interrupted.

[0040] Even worse, if a user is hesitant and their hand trembles frequently and slightly within a short period, the system might mistakenly identify the actual result as the same gesture twice (e.g., triggering "turn right page" twice in a row). To address this issue, this solution stipulates that the current result must be abandoned until the duration of the corresponding action reaches a certain set confidence threshold. Only when the gesture is determined to be non-transitional, and the inference results of two consecutive actions show a significant and stable difference, is the system considered to be acting on the user's active interaction intent and responds immediately. This method completely avoids repetitive operations and redundant actions, ensuring the smoothness of digital reading and display.

[0041] Example 5

[0042] The aforementioned classification results are ultimately input into the exhibition narrative control engine, which internally maintains a logical topology tree composed of multiple scene nodes. Each node carries the corresponding 3D assets of the exhibit, lighting parameters, audio tracks, and triggerable transition edges. Taking the display of the ancient architectural "mortise and tenon structure" as an example: Initial state S0 (macroscopic appearance display): The holographic projection device projects the complete wooden tower.

[0043] Action Response -> State S1 (Structural Perspective): When the system recognizes that the user's hands have performed a specific standardized "inward squeezing and flipping" gesture, the control engine does not simply scale the model, but consults the state table and determines that the operation has touched the boundary of entering the internal structure. Subsequently, the engine outputs a narrative advancement instruction package, ordering the building's tiles and exterior walls to become semi-transparent, the wooden frame to be highlighted, and the audio to switch to structural mechanics narration.

[0044] Action Response -> State S2 (Historical Evolution Deduction): At this point, if the user draws a "reverse spiral trajectory" representing time with one hand, the engine captures this narrative-level temporal instruction, immediately pauses the current physical structure disassembly, and jumps the state tree to the "Era Construction Tracing Branch". The holographic projection then changes, showing step-by-step animations of each historical stage from laying the foundation to raising the beam.

[0045] Through this engine design, the solution breaks down the technical barriers between micro-manipulation of the hands and digital environment display, giving the exhibition narrative a high degree of autonomy and story extensibility.

[0046] In the description of this specification, references to terms such as "an embodiment," "example," "specific example," etc., indicate that a specific feature, structure, material, or characteristic described in connection with that embodiment or example is included in at least one embodiment or example of the invention. In this specification, illustrative expressions of the above terms do not necessarily refer to the same embodiment or example. Furthermore, the specific features, structures, materials, or characteristics described may be combined in any suitable manner in one or more embodiments or examples.

[0047] The preferred embodiments of the present invention disclosed above are merely illustrative of the invention. These preferred embodiments do not exhaustively describe all details, nor do they limit the invention to the specific implementations described. Clearly, many modifications and variations can be made based on the content of this specification. This specification selects and specifically describes these embodiments to better explain the principles and practical applications of the invention, thereby enabling those skilled in the art to better understand and utilize the invention. The invention is limited only by the claims and their full scope and equivalents.

Claims

1. An adaptive control method for exhibition narrative based on gesture trajectory recognition and real-time rendering of multimedia elements, characterized in that, Includes the following steps: S1. Acquire video data streams, depth image data, and electromagnetic radiation reflection signals within the exhibition space, and extract the humanoid region and key hand skeletal node sequences of the target user from them; S2. Based on the extracted key hand bone node sequence, continuously track the target user's hand movement within a preset reduced interaction volume, and record a continuous three-dimensional spatial coordinate sequence as the original gesture trajectory data. S3. Perform spatial resampling and scale normalization operations on the original gesture trajectory data, extract the six-degree-of-freedom motion feature vector in Eulerian space, and obtain a standardized trajectory feature vector sequence. S4. Input the standardized trajectory feature vector sequence into the pre-trained hybrid gesture recognition model. The hybrid gesture recognition model outputs the classification result of the target gesture through joint inference of dynamic time warping algorithm and hidden Markov model. S5. Input the classification results into the exhibition narrative control engine. The exhibition narrative control engine generates corresponding narrative advancement instructions based on the context node of the current exhibition narrative, the multimedia operation status of the exhibits, and the classification results. S6. According to the narrative progression instructions, the multimedia playback module, holographic imaging equipment and audio-visual interactive environment of the exhibition site are synchronously controlled to realize the dynamic interpretation of the exhibition content, timeline retracing or spatial hierarchy deconstruction.

2. The adaptive control method for exhibition narrative based on gesture trajectory recognition and real-time rendering of multimedia elements according to claim 1, characterized in that, Step S1 further includes: Electromagnetic waves are emitted towards the target user area by an electromagnetic radiation emitter, and electromagnetic radiation signals reflected back by retroreflective materials on items worn or held by the target user are received by a sensing device. The received reflection signal is aligned with the depth image stream, and the human bounding box of the target user is confirmed by combining it with a human detection network. By performing gesture detection on the human-shaped bounding box, the focus area is determined based on the gesture detection results. If the target area is set as the human-shaped area of ​​the speaker or core participant, the interactive focus and light tracking of the system are locked to the speaker in real time.

3. The adaptive control method for exhibition narrative based on gesture trajectory recognition and real-time rendering of multimedia elements according to claim 1, characterized in that, The continuous tracking of the target user's hand movements within a preset reduced interaction volume specifically includes: In augmented reality or virtual reality three-dimensional space, define a local interactive volume that is fixed relative to the user's body torso or relative to the physical display stand; The first neural network is provided with raw sensor data without preprocessing and preprocessed data frames filtered by the interaction volume to identify prior features characterizing subsequent gesture interactions. The system detects the user's waving and bending gestures, and records the bending trajectory of each of the five fingers and the global waving trajectory of the palm.

4. The adaptive control method for exhibition narrative based on gesture trajectory recognition and real-time rendering of multimedia elements according to claim 1, characterized in that, The spatial resampling and scale normalization operations specifically include: The original three-dimensional coordinate point set of the gesture trajectory is resampled to obtain a resampled point set, so that the physical curve distance between any two consecutive points along the trajectory is strictly equal. The resampled point set is translated and adjusted in the three-dimensional coordinate system so that the centroid of the translated gesture trajectory point cloud is located at the origin of the three-dimensional coordinate system. Extract the argument value of the first point of the trajectory after translation, and rotate the entire point set in a specified direction according to the argument value to eliminate the spatial difference of the user's initial hand posture; The rotated point set is scaled so that all the scaled coordinate points are within a square or cube bounding box of a specified size. Each trajectory after the above processing is converted into a vector representation in six Euler degrees of freedom in Euler space, and the vector with the largest amplitude is extracted as the dominant motion attribute characterizing the trajectory.

5. The adaptive control method for exhibition narrative based on gesture trajectory recognition and real-time rendering of multimedia elements according to claim 1, characterized in that, The hybrid gesture recognition model outputs the classification result of the target gesture through joint inference using a dynamic time warping algorithm and a hidden Markov model, specifically including: The standardized trajectory feature vector sequence is compared with the continuous character gesture template and dynamic trajectory style template in the template library using the dynamic time warping algorithm. The candidate template sequence with the smallest dynamic time warping distance is selected as the initial screening result. The initial screening results are input into the Hidden Markov Model. The hidden state transition probabilities of hand shape and hand movement are combined to calculate the matching degree between the input and various known gesture action models. The gesture category recognition result with the maximum posterior probability is output.

6. The adaptive control method for exhibition narrative based on gesture trajectory recognition and real-time rendering of multimedia elements according to claim 1, characterized in that, Step S5 includes transition gesture recognition logic to filter out erroneous operations: Between any two explicit custom gestures with clear mapping instructions, extract gesture action sequence features and determine whether the current action constitutes a transition gesture without mapping relationship. If the current gesture recognition result is determined to be a transitional gesture, or if two consecutive gesture recognition results are the same, the interaction command corresponding to the current gesture recognition result will be abandoned before the execution time of the corresponding gesture reaches a preset threshold, so as to avoid executing the same gesture continuously in a short period of time and avoid executing unnecessary transitional actions. If the current gesture recognition result is a non-transitional gesture, and the recognition results are different in two consecutive times, then a narrative progression instruction will be output.