Exhibition gesture and voice multi-modal adaptive interaction method based on digital twinning

By constructing an inter-frame transformation matrix and an improved VL-BERT model in the exhibition hall interactive system for Lie algebra decoupling and joint coding, the stability and accuracy issues of gesture and speech recognition were resolved, enabling synchronous response between the virtual exhibition hall and the physical exhibition booth, thereby improving the intelligent interactive level of the exhibition hall and the immersive experience of the audience.

CN122387315APending Publication Date: 2026-07-14ANHUI EASTDOLT ELECTRONICS TECH
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
ANHUI EASTDOLT ELECTRONICS TECH
Filing Date
2026-04-17
Publication Date
2026-07-14

AI Technical Summary

Technical Problem

The existing interactive system for exhibition halls suffers from insufficient stability in gesture recognition and voice recognition. It is difficult to accurately represent changes in hand orientation and position in three-dimensional space. Furthermore, voice recognition is unstable in noisy environments and lacks global structured alignment capabilities, resulting in inaccurate interaction results and insufficient control over virtual and real interaction.

Method used

By constructing an inter-frame transformation matrix and mapping it to a special Euclidean group manifold for Lie algebra decoupling, and combining it with an improved VL-BERT model for joint coding and optimal transmission alignment, cross-modal alignment results are generated, which in turn drive the linkage control between the virtual engine and the physical booth.

Benefits of technology

It improves the stability and spatial positioning accuracy of gesture recognition, enhances the robustness of speech recognition and the accuracy of interactive intent parsing, realizes synchronous response between virtual exhibition halls and physical exhibition stands, and improves the level of intelligent interaction of exhibition halls and the immersive experience of visitors.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122387315A_ABST
    Figure CN122387315A_ABST
Patent Text Reader

Abstract

The application discloses an exhibition hall gesture and voice multimodal adaptive interaction method based on digital twinning, which comprises the following steps: obtaining a user hand three-dimensional coordinate sequence and generating a manifold gesture mark through special Euclidean group flow mapping and Lie algebra decoupling; completing an exhibition product intersection positioning and generating a manifold anchor point feature according to a rotation component; collecting a voice signal, a signal-to-noise ratio and a reverberation time, generating a text mark and an environment mark; inputting the environment mark, the text mark and the manifold gesture mark into an improved VL-BERT model for joint coding, and performing an entropy regularization optimal transmission iteration with environment damping to generate a cross-modal alignment result and an optimal transmission matrix; performing two-way tensor decoding according to a confidence degree to output a digital twinning action, an Internet of Things control instruction and a natural language explanation result; and driving the digital twinning body and the physical exhibition stand to be linked to execute. The application improves the semantic alignment accuracy, the adaptive capability and the linkage control effect of the exhibition hall interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of digital twin interactive control technology, and in particular to a multimodal adaptive interaction method for exhibition hall gestures and voice based on digital twins. Background Technology

[0002] With the increasing demand for digital twin technology, multimodal interaction technology, and smart exhibition hall construction, adaptive interaction methods that combine user gestures and voice input in exhibition hall scenarios have received widespread attention. Existing exhibition hall interaction systems mainly rely on ordinary two-dimensional gesture recognition, voice command recognition, or simple rule fusion methods to achieve exhibit retrieval and content broadcasting. However, these systems commonly suffer from the following problems in practical applications: Hand movements are typically described directly using Euclidean coordinates or image features, making it difficult to stably represent changes in hand orientation and position in three-dimensional space. This leads to drift in the user's pointing and positioning of exhibits under different viewpoints and amplitudes of movement. The exhibition environment is characterized by significant noise, reverberation, and crowd interference. Existing speech processing methods mostly only perform front-end noise reduction on the speech signal, failing to effectively transmit changes in sound field quality to the subsequent multimodal fusion process. This results in unstable speech recognition results that may still cause errors in interactive decision-making. Existing gesture and speech fusion methods often employ temporal order judgment, weighted concatenation, or local similarity matching, lacking the ability to achieve global structured alignment between textual semantics and spatial gestures, making it difficult to accurately complete exhibit semantic positioning and interactive intent parsing. Furthermore, existing solutions often remain at the level of virtual display or single broadcast, lacking the ability to further map interaction results into closed-loop execution capabilities that link digital twin actions with physical booth control.

[0003] Therefore, how to provide a digital twin-based multimodal adaptive interaction method for exhibition hall gestures and voice is a problem that urgently needs to be solved by those skilled in the art. Summary of the Invention

[0004] One objective of this invention is to propose a digital twin-based multimodal adaptive interaction method for exhibition hall gestures and voice. This invention fully utilizes technologies such as digital twin modeling, Lie algebra decoupling, multimodal joint coding, and optimal transmission alignment. It describes in detail the implementation process of gesture recognition, voice understanding, cross-modal semantic alignment, and virtual-real linkage control in exhibition hall scenarios. It has the advantages of strong natural interaction, high semantic matching accuracy, good environmental adaptability, and stable exhibition item linkage control effect.

[0005] The exhibition hall gesture and voice multimodal adaptive interaction method based on digital twin according to an embodiment of the present invention includes the following steps: Step 1: Obtain the 3D coordinate sequence of the user's hand, construct the inter-frame transformation matrix and map it to a special Euclidean group manifold, then perform Lie algebra decoupling to obtain rotational and translational components, and generate manifold gesture markers; Step 2: Construct a pointing ray based on the rotation component, and perform intersection operation with the 3D bounding box of each exhibit in the digital twin engine. Concatenate the semantic identifier of the hit exhibit with the corresponding Lie algebra motion vector to generate manifold anchor point features and update the manifold gesture marker. Step 3: Collect user voice signals, signal-to-noise ratio and reverberation time of the local sound field of the exhibition hall, perform automatic speech recognition on the voice signals to generate text tags, and map the signal-to-noise ratio and reverberation time to the optimal transmission cost damping coefficient to generate environmental tags; Step 4: Input the environment tag, text tag, and manifold gesture tag into the improved VL-BERT model for joint encoding. The improved VL-BERT model introduces an environment damping encoding mechanism, a manifold gesture embedding mechanism, and a cross-modal optimal transmission enhancement alignment mechanism. Step 5: Perform an entropy-regularized optimal transport iteration with environmental damping in the improved VL-BERT model to generate cross-modal alignment results and the optimal transport matrix; Step 6: Extract the edge distribution of the optimal transfer matrix as the confidence level, and perform dual-channel tensor decoding when the confidence level reaches the preset threshold to output digital twin actions, IoT control commands, and natural language interpretation results; Step 7: Map the digital twin actions and IoT control commands into affine transformation matrices and physical electrical signal parameters, drive the digital twin in the virtual engine to perform smooth manifold transformations, and drive the motors and lights of the physical display stand to perform linked control.

[0006] Optionally, step one specifically includes: The three-dimensional coordinate sequence of the user's hand is arranged according to the sampling time sequence, and the same hand region in adjacent sampling times is temporally correlated to obtain continuous sampling frames; Based on the center point of the palm, the wrist point, and the fingertip point in each consecutive sampling frame, a corresponding local coordinate system of the hand is constructed to determine the spatial pose of the local coordinate system of the hand. Rigid body registration is performed using the local coordinate system of the hand corresponding to adjacent consecutive sampling frames. The spatial pose change and spatial position change of the subsequent consecutive sampling frame relative to the previous consecutive sampling frame are calculated to generate the corresponding inter-frame transformation matrix. Map the inter-frame transformation matrix to a special Euclidean group manifold, and then convert the mapped inter-frame transformation matrix into its corresponding Lie algebra representation; The Lie algebra representation is decoupled, and the part representing the change in the spatial orientation of the hand is determined as the rotation component, and the part representing the change in the spatial displacement of the hand is determined as the translation component. The rotation and translation components corresponding to multiple consecutive sampling times are sequentially combined to generate manifold gesture markers.

[0007] Optionally, step two specifically includes: The three-dimensional models of each exhibit in the digital twin engine are spatially registered, and a corresponding three-dimensional bounding box is constructed based on the outer contour range of each exhibit's three-dimensional model in the digital twin scene. At the same time, the semantic identifier of the corresponding exhibit is written into each three-dimensional bounding box. Based on the rotation component and the hand's spatial position when generating the manifold gesture mark, the current pointing direction of the hand is determined, and a pointing direction ray is constructed; The pointing ray is input into the digital twin engine, and the pointing ray is made to intersect with each 3D bounding box along the ray extension direction to obtain the candidate hit results corresponding to each pointing ray. The candidate hit results are filtered, and the three-dimensional bounding box that is closest to the spatial position of the hand along the ray extension direction of the pointing direction is retained as the target hit bounding box; The rotation and translation components at the sampling time corresponding to the bounding box of the target are concatenated to generate the Lie algebra motion vector. The semantic identifiers of the hit exhibits are combined with the Lie algebra motion vectors to generate manifold anchor features, which are then written into the manifold gesture tags to obtain the updated manifold gesture tags.

[0008] Optionally, step three specifically includes: Speech signals and local sound field signals are acquired and processed by frame segmentation to obtain speech analysis frame sequences and sound field analysis frame sequences arranged in chronological order. Extract the target speech segment from the speech analysis frame sequence based on the speech start frame and speech end frame; Input the target speech segment into the Conformer speech recognition model, output the text result, and generate text tags according to the word order and the start and end times of the speech; During the preset background sampling period before the speech start frame, a background reference frame sequence is extracted from the sound field analysis frame sequence corresponding to the target speech segment, and the average frame energy of the background reference frame sequence is calculated as the background noise energy. Between the speech start frame and the speech end frame, a speech sampling frame sequence is extracted from the sound field analysis frame sequence corresponding to the target speech segment. The average frame energy of the speech sampling frame sequence is calculated as the speech energy, and the signal-to-noise ratio is determined based on the ratio between the speech energy and the background noise energy. Within a preset reverberation sampling period after the speech termination frame, a reverberation decay frame sequence is extracted from the sound field analysis frame sequence. An energy decay process is constructed from the frame energy of the reverberation decay frame sequence, and the reverberation time is determined. The signal-to-noise ratio is compared with a preset signal-to-noise ratio range to determine the signal-to-noise ratio classification result, and the reverberation time is compared with a preset reverberation time range to determine the reverberation time classification result. The signal-to-noise ratio classification result is used as the row index of the preset damping coefficient mapping table, and the reverberation time classification result is used as the column index of the preset damping coefficient mapping table. The damping value of the corresponding table entry is read, and the damping value is determined as the optimal transmission cost damping coefficient. The optimal transmission cost damping coefficient is bound to the start and end times of the target speech segment to generate an environmental label.

[0009] Optionally, the improved VL-BERT model includes a word embedding layer, a positional encoding layer, multiple Transformer hidden layers, and an output layer: The word embedding layer inputs environment tags, text tags, and updated manifold gesture tags into a unified embedding space. It performs an embedding lookup table on the text tags to obtain text word embedding vectors. It introduces an environment damping coding mechanism to perform a linear transformation on the optimal transmission cost damping coefficient in the environment tags and generates environment tag embeddings in combination with the interaction time. It introduces a manifold gesture embedding mechanism to perform embedding mapping and fusion on the semantic identifiers, rotation components, and translation components in the manifold gesture tags to obtain manifold gesture embedding vectors. The location encoding layer writes the environment tag embedding, text word embedding vector, and manifold gesture embedding vector into the location encoding and type encoding according to the interaction time sequence, and concatenates them in the order of environment tag, text tag, and manifold gesture tag to obtain the joint input sequence; The multi-layer Transformer hidden layer inputs the joint input sequence into each Transformer hidden layer, performing query transformation, key transformation, and value transformation respectively to obtain query matrix, key matrix, and value matrix. The basic attention score is calculated based on the matrix product of the query matrix and the key matrix, and the basic attention weight is obtained after scaling and Softmax normalization. The value matrix is ​​weighted and summed to obtain the basic attention output. A cross-modal optimal transmission enhancement alignment mechanism is introduced to construct cross-modal feature pairs between text word embedding vectors and manifold gesture embedding vectors. The initial transmission cost is calculated and modulated in combination with the environmental damping embedding vector to obtain the enhanced attention score. After Softmax normalization, the enhanced attention weight is obtained. The enhanced attention output is then added to the joint input sequence by residual summation and layer normalization, and then input into the feedforward network for processing to obtain the joint encoding result of the current Transformer hidden layer. The output layer inputs the joint encoding result output by the last Transformer hidden layer into the output layer, and performs linear projection on the environment marker position, text marker position and manifold gesture marker position in the joint encoding result to obtain the environment encoding result, text encoding result and gesture encoding result.

[0010] Optionally, step five specifically includes: Read the environment encoding result, text encoding result, and gesture encoding result, and perform temporal alignment of the text encoding result and the gesture encoding result according to the same interaction time to construct a cross-modal feature pair between text tags and gesture tags; For each cross-modal feature pair, the feature difference between the text encoding result and the gesture encoding result is calculated, and the absolute value operation, square operation and weighted summation are performed on the feature difference respectively to generate the basic transmission cost between each text tag and each gesture tag; Each basic transmission cost is written into the initial cost matrix according to the arrangement of text tags in rows and gesture tags in columns. Then, the initial cost matrix is ​​modulated by element-wise multiplication based on the damping intensity corresponding to each interaction moment in the environment coding result to generate the damping cost matrix. An entropy-regularized transmission kernel matrix is ​​constructed based on the damping cost matrix. The entropy-regularized transmission kernel matrix is ​​a transmission weight matrix obtained by performing negative scaling and exponential operations on each element in the damping cost matrix. The scaling magnitude of the negative scaling is determined by a preset regularization parameter. The optimal transfer matrix is ​​obtained by performing alternating normalization iterations on the entropy-regularized transfer kernel matrix. The text encoding results and gesture encoding results are bidirectionally weighted and aggregated using the optimal transfer matrix. The gesture encoding results are weighted and summed according to the weight distribution of each row in the optimal transfer matrix to obtain the text alignment features. The text encoding results are weighted and summed according to the weight distribution of each column in the optimal transfer matrix to obtain the gesture alignment features. The text alignment features, gesture alignment features, and environment coding results are concatenated, and the concatenated results are subjected to linear transformation, residual addition, and layer normalization to generate cross-modal alignment results.

[0011] Optionally, step six specifically includes: The optimal transfer matrix is ​​summed row by row to obtain the text edge distribution corresponding to each text marker, and the optimal transfer matrix is ​​summed column by column to obtain the gesture edge distribution corresponding to each manifold gesture marker. Based on the text edge distribution and gesture edge distribution, the maximum transmission allocation value corresponding to each text marker, the maximum transmission allocation value corresponding to each manifold gesture marker, and the concentrated area range of high transmission allocation values ​​are statistically analyzed in the optimal transmission matrix. The text edge distribution, gesture edge distribution, maximum transmission allocation value corresponding to each text marker, maximum transmission allocation value corresponding to each manifold gesture marker, and concentrated area range are combined to generate a confidence judgment result. The confidence score is normalized to obtain a confidence level, which is then compared with a preset confidence threshold. When the confidence level reaches the preset threshold, the cross-modal alignment result is read and dual-channel tensor decoding is performed. In the control output path of dual-path tensor decoding, the cross-modal alignment result is concatenated with the manifold anchor point feature. The concatenation result is linearly transformed and the residual is updated to generate a control semantic tensor. Based on the control semantic tensor, the target exhibit identifier, target action category, action direction, action amplitude and action duration are determined to generate a digital twin action. Based on the target exhibit identifier and the control semantic tensor, interval mapping and threshold determination are performed on each control component in the control semantic tensor. The control components that satisfy the motor control interval are converted into motor control parameters, and the control components that satisfy the lighting control interval are converted into lighting control parameters. The motor control parameters and lighting control parameters are combined to generate IoT control commands. The dual-path tensor decoding interpreter output path concatenates the text alignment features, gesture alignment features, and environment coding results from the cross-modal alignment results. The concatenated results are then processed by linear transformation, Gaussian error linear unit activation function, and linear transformation to generate an interpreter semantic tensor. Based on the interpreter semantic tensor, natural language interpreter results are output in word order.

[0012] Optionally, step seven specifically includes: Read digital twin actions and IoT control commands to extract target exhibit identifiers, target action categories, action directions, action amplitudes, action durations, as well as motor control parameters and lighting control parameters; Based on the target exhibit identifier, locate the corresponding digital twin in the virtual engine, read the initial spatial position and initial spatial posture of the digital twin, and generate an affine transformation matrix sequence based on the target action type, action direction, action amplitude and action duration; Based on the motor control parameters and the lighting control parameters, generate the physical electrical signal parameters for motor drive and lighting drive respectively; The execution time of the affine transformation matrix sequence is aligned with the output time of the physical electrical signal parameters to form a linkage control timing sequence, and the digital twin in the virtual engine and the motors and lights in the physical display stand are driven synchronously according to the linkage control timing sequence.

[0013] The beneficial effects of this invention are: This invention addresses the problem of unstable pointing and positioning in existing technologies due to the susceptibility of gesture representation to changes in viewpoint, motion jitter, and spatial pose, by constructing the user's three-dimensional hand coordinate sequence as an inter-frame transformation matrix and mapping it to a special Euclidean group manifold followed by Lie algebra decoupling. It achieves joint modeling of hand orientation and position changes, thereby improving the stability and spatial positioning accuracy of exhibit pointing recognition. Furthermore, by constructing the rotation component as a pointing ray and performing intersection calculations with the exhibit's three-dimensional bounding box, it addresses the difficulty of accurately associating traditional gesture interactions with specific exhibit objects, achieving direct binding between gesture actions and exhibit semantic identifiers, thus improving the accuracy of target exhibit hits. Finally, by mapping the signal-to-noise ratio and reverberation time to an optimal transmission cost damping system... This paper introduces an environmental damping coding mechanism, a manifold gesture embedding mechanism, and a cross-modal optimal transmission enhancement alignment mechanism into the improved VL-BERT model. Addressing the issues of unstable speech recognition in noisy exhibition hall environments and the weak anti-interference capability of traditional multimodal fusion, it achieves adaptive joint coding of speech, environmental, and gesture information, thereby improving the accuracy and robustness of interactive intent parsing. Furthermore, by decoding the cross-modal alignment results into digital twin actions, IoT control commands, and natural language explanations, it addresses the lack of closed-loop control for virtual-physical interaction in existing solutions, achieving synchronous response between the virtual exhibition hall and the physical exhibition booth. This has significant implications for enhancing the intelligent interactive level of exhibition halls, strengthening the immersive experience for visitors, and promoting the practical application of digital twin exhibition technology. Attached Figure Description

[0014] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings: Figure 1 The flowchart shows the multimodal adaptive interaction method for exhibition hall gestures and voice based on digital twins proposed in this invention. Figure 2 This is a schematic diagram of the exhibition hall gesture and voice multimodal adaptive interaction method based on digital twin proposed in this invention; Figure 3 This is a framework diagram of the improved VL-BERT model in the digital twin-based multimodal adaptive interaction method for exhibition hall gestures and voice proposed in this invention. Detailed Implementation

[0015] The present invention will now be described in further detail with reference to the accompanying drawings. These drawings are simplified schematic diagrams, illustrating only the basic structure of the invention, and therefore only show the components relevant to the invention.

[0016] refer to Figures 1-3 A digital twin-based multimodal adaptive interaction method for exhibition hall gestures and voice includes the following steps: Step 1: Obtain the 3D coordinate sequence of the user's hand, construct the inter-frame transformation matrix and map it to a special Euclidean group manifold, then perform Lie algebra decoupling to obtain rotational and translational components, and generate manifold gesture markers; Step 2: Construct a pointing ray based on the rotation component, and perform intersection operation with the 3D bounding box of each exhibit in the digital twin engine. Concatenate the semantic identifier of the hit exhibit with the corresponding Lie algebra motion vector to generate manifold anchor point features and update the manifold gesture marker. Step 3: Collect user voice signals, signal-to-noise ratio and reverberation time of the local sound field of the exhibition hall, perform automatic speech recognition on the voice signals to generate text tags, and map the signal-to-noise ratio and reverberation time to the optimal transmission cost damping coefficient to generate environmental tags; Step 4: Input the environment tag, text tag, and manifold gesture tag into the improved VL-BERT model for joint encoding. The improved VL-BERT model introduces an environment damping encoding mechanism, a manifold gesture embedding mechanism, and a cross-modal optimal transmission enhancement alignment mechanism. Step 5: Perform an entropy-regularized optimal transport iteration with environmental damping in the improved VL-BERT model to generate cross-modal alignment results and the optimal transport matrix; Step 6: Extract the edge distribution of the optimal transfer matrix as the confidence level, and perform dual-channel tensor decoding when the confidence level reaches the preset threshold to output digital twin actions, IoT control commands, and natural language interpretation results; Step 7: Map the digital twin actions and IoT control commands into affine transformation matrices and physical electrical signal parameters, drive the digital twin in the virtual engine to perform smooth manifold transformations, and drive the motors and lights of the physical display stand to perform linked control.

[0017] In this embodiment, step one specifically includes: The three-dimensional coordinate sequence of the user's hand is arranged according to the sampling time sequence, and the same hand region in adjacent sampling times is temporally correlated to obtain continuous sampling frames; A corresponding local coordinate system for the hand is constructed based on the center point of the palm, the wrist point, and the fingertip point in each consecutive sampling frame. The center point of the palm is used as the origin of the coordinate system, the direction from the center point of the palm to the fingertip point is used as the pointing reference direction, and the direction from the wrist point to the center point of the palm point is used as the motion reference direction. The spatial pose of the local coordinate system of the hand is determined based on the pointing reference direction and the motion reference direction. Rigid body registration is performed using the local coordinate system of the hand corresponding to adjacent consecutive sampling frames. The spatial pose change and spatial position change of the subsequent consecutive sampling frame relative to the previous consecutive sampling frame are calculated to generate the corresponding inter-frame transformation matrix. The inter-frame transformation matrix is ​​mapped to a special Euclidean group manifold, which is a joint pose space used to characterize the rotational and translational motions of the hand in three-dimensional space. The mapped inter-frame transformation matrix is ​​then converted into its corresponding Lie algebra representation. The Lie algebra representation is decoupled, and the part representing the change in the spatial orientation of the hand is determined as the rotation component, and the part representing the change in the spatial displacement of the hand is determined as the translation component. The rotation and translation components corresponding to multiple consecutive sampling times are sequentially combined according to the sampling time sequence to generate manifold gesture markers.

[0018] In this embodiment, step two specifically includes: The three-dimensional models of each exhibit in the digital twin engine are spatially registered, and a corresponding three-dimensional bounding box is constructed based on the outer contour range of each exhibit's three-dimensional model in the digital twin scene. At the same time, the semantic identifier of the corresponding exhibit is written into each three-dimensional bounding box. Based on the rotation component and the hand spatial position corresponding to the generation of manifold gesture mark, the current pointing direction of the hand is determined, and a pointing direction ray is constructed based on the current pointing direction of the hand and the hand spatial position; The pointing ray is input into the digital twin engine, and the pointing ray is made to intersect with each 3D bounding box along the ray extension direction to obtain the candidate hit results corresponding to each pointing ray. The candidate hit results are filtered, and the three-dimensional bounding box that is closest to the spatial position of the hand along the ray extension direction of the pointing direction is retained as the target hit bounding box; The rotation and translation components at the sampling time corresponding to the bounding box of the target are concatenated to generate the Lie algebra motion vector. The semantic identifiers of the hit exhibits are combined with the Lie algebra motion vectors to generate manifold anchor features, which are then written into the manifold gesture tags to obtain the updated manifold gesture tags.

[0019] In this embodiment, step three specifically includes: Speech signals and local sound field signals are acquired and processed by frame segmentation to obtain speech analysis frame sequences and sound field analysis frame sequences arranged in chronological order. Calculate the frame energy and zero-crossing rate for each frame in the speech analysis frame sequence. Determine the speech start frame as a speech analysis frame with multiple consecutive frames whose energy is higher than the preset start energy threshold and whose zero-crossing rate is within the preset speech zero-crossing rate range. Determine the speech stop frame as a speech analysis frame with multiple consecutive frames whose energy is lower than the preset stop energy threshold. Extract the target speech segment based on the speech start frame and the speech stop frame. Input the target speech segment into the Conformer speech recognition model, output the text result corresponding to the target speech segment, and generate text tags according to the word order and the start and end time of the speech; During the preset background sampling period before the speech start frame, a background reference frame sequence is extracted from the sound field analysis frame sequence corresponding to the target speech segment, and the average frame energy of the background reference frame sequence is calculated as the background noise energy. Between the speech start frame and the speech end frame, a speech sampling frame sequence is extracted from the sound field analysis frame sequence corresponding to the target speech segment. The average frame energy of the speech sampling frame sequence is calculated as the speech energy, and the signal-to-noise ratio is determined based on the ratio between the speech energy and the background noise energy. Within the preset reverberation sampling period after the speech termination frame, the reverberation attenuation frame sequence is extracted from the sound field analysis frame sequence. The frame energy of the reverberation attenuation frame sequence is constructed in chronological order to form an energy attenuation process. The duration corresponding to the energy attenuation process decreasing by 60 decibels from the attenuation starting point is determined as the reverberation time. The signal-to-noise ratio is compared with a preset signal-to-noise ratio range to determine the signal-to-noise ratio classification result, and the reverberation time is compared with a preset reverberation time range to determine the reverberation time classification result. The signal-to-noise ratio classification result is used as the row index of the preset damping coefficient mapping table, and the reverberation time classification result is used as the column index of the preset damping coefficient mapping table. The damping value of the corresponding table entry is read, and the damping value is determined as the optimal transmission cost damping coefficient. The preset damping coefficient mapping table is a two-dimensional parameter table used to store the damping values ​​corresponding to the signal-to-noise ratio classification result and the reverberation time classification result. The optimal transmission cost damping coefficient is bound to the start and end times of the target speech segment to generate an environmental label.

[0020] In this embodiment, the improved VL-BERT model includes a word embedding layer, a positional encoding layer, multiple Transformer hidden layers, and an output layer: The word embedding layer inputs environment tags, text tags, and updated manifold gesture tags into a unified embedding space. It performs embedding lookup on each lexical identifier in the text tag to obtain the text word embedding vector. An environment damping coding mechanism is introduced, performing scalar expansion and linear transformation on the optimal transmission cost damping coefficient in the environment tag to obtain the environment damping embedding vector. Temporal embedding is performed on the interaction time corresponding to the environment tag to obtain the environment temporal embedding vector. The environment damping embedding vector and the environment temporal embedding vector are then added element-wise to obtain the environment tag embedding. A manifold gesture embedding mechanism is introduced, performing embedding lookup on the semantic identifier in the manifold gesture tag to obtain the semantic embedding vector. Linear transformations are performed on the rotation and translation components in the manifold gesture tag to obtain rotation and translation embedding vectors, respectively. The rotation and translation embedding vectors are concatenated, and the concatenation result is linearly compressed to obtain the motion embedding vector. Finally, the semantic embedding vector and the motion embedding vector are weighted and added to obtain the manifold gesture embedding vector. The weighting coefficients are determined based on the normalized amplitudes of the rotation and translation components. The location encoding layer inputs the environment tag embedding, text word embedding vector, and manifold gesture embedding vector into the location encoding layer according to the interaction sequence. Each input tag is written into the location encoding, and then into the environment type encoding, text type encoding, and gesture type encoding according to the source of the tag. The environment tag embedding, text word embedding vector, and manifold gesture embedding vector after writing the location encoding and type encoding are added element by element and concatenated in the order of environment tag first, text tag in the middle, and manifold gesture tag last to obtain the joint input sequence. The multi-layered Transformer hidden layers input the joint input sequence into each Transformer hidden layer, performing query transformation, key transformation, and value transformation respectively to obtain query matrix, key matrix, and value matrix. A basic attention score is calculated based on the matrix product of the query matrix and key matrix, and this score is scaled according to the key vector dimension. The scaled result is then subjected to Softmax normalization to obtain the basic attention weights. These weights are used to perform a weighted summation of the value matrix to obtain the basic attention output. A cross-modal optimal transmission enhancement alignment mechanism is introduced, constructing cross-modal feature pairs based on the pairwise correspondence between text word embedding vectors and manifold gesture embedding vectors. The absolute value operation and square operation are performed on the differences between each cross-modal feature pair, and the results are weighted and summed to obtain the initial transmission cost. The environmental damping embedding vector is then extended to the initial transmission cost. For dimensions with the same value, element-wise multiplication modulation is performed on the initial transmission cost value to obtain the damped transmission cost value. The damped transmission cost value is then written into the basic attention score to obtain the enhanced attention score. Softmax normalization is performed on the enhanced attention score to obtain the enhanced attention weight. The enhanced attention weight is then used to perform a weighted summation on the value matrix to obtain the enhanced attention output. The enhanced attention output is added to the joint input sequence by residual addition, and layer normalization is performed on the residual addition result. The layer normalization result is then input into the feedforward network, where linear transformation, Gaussian error linear unit activation function processing, and linear transformation are performed sequentially to obtain the feedforward output. The feedforward output is added to the layer normalization result by residual addition, and layer normalization is performed on the residual addition result to obtain the joint encoding result of the current Transformer hidden layer. The joint encoding result is then passed sequentially through multiple Transformer hidden layers. The output layer takes the joint encoding result output from the last Transformer hidden layer as input to the output layer, and performs linear projection on the environment marker position, text marker position and manifold gesture marker position in the joint encoding result to obtain the environment encoding result, text encoding result and gesture encoding result; The improved VL-BERT model retains the basic structure of the original VL-BERT model, including word embedding layer, position encoding layer, multi-layer Transformer hidden layer and output layer, and follows the basic processing flow of the original VL-BERT model to perform embedding representation, position modeling, self-attention encoding and output projection on input tokens; Unlike the original VL-BERT model, this implementation introduces an environment-damped coding mechanism and a manifold gesture embedding mechanism in the word embedding layer, enabling the optimal transmission cost damping coefficient in the environment tag, the semantic identifier, rotation components, and translation components in the manifold gesture tag to enter a unified embedding space. At the same time, a cross-modal optimal transmission enhancement alignment mechanism is introduced in the multi-layer Transformer hidden layer, so that the text word embedding vector and the manifold gesture embedding vector are no longer associated solely by the basic attention score, but rather the initial transmission cost value is modulated by the environment-damped embedding vector to complete the enhancement alignment. Through the above improvements, the technical pain points of the original model, such as difficulty in directly processing the sound field disturbance information of the exhibition hall, difficulty in expressing the spatial pose changes of gestures, and unstable alignment between text and gestures, are solved. The model can still maintain high cross-modal alignment accuracy under the conditions of exhibition hall noise, reverberation and complex pointing actions, thereby improving the accuracy of exhibit positioning, the accuracy of interactive intent parsing and the overall interaction robustness.

[0021] In this embodiment, step five specifically includes: Read the environment encoding results, text encoding results, and gesture encoding results, and align the text encoding results and gesture encoding results temporally according to the same interaction time to construct cross-modal feature pairs between text tags and gesture tags; For each cross-modal feature pair, the feature difference between the text encoding result and the gesture encoding result is calculated, and the absolute value operation, square operation and weighted summation are performed on the feature difference respectively to generate the basic transmission cost between each text tag and each gesture tag. The weight in the weighted summation is determined by the environment encoding result output in step four after linear mapping and normalization. Each basic transmission cost is written into the initial cost matrix according to the arrangement of text tags in rows and gesture tags in columns. Then, the initial cost matrix is ​​modulated by element-wise multiplication based on the damping intensity corresponding to each interaction moment in the environment coding result to generate the damping cost matrix. An entropy-regularized transmission kernel matrix is ​​constructed based on the damping cost matrix. The entropy-regularized transmission kernel matrix is ​​a transmission weight matrix obtained by performing negative scaling and exponential operations on each element in the damping cost matrix. The scaling magnitude of the negative scaling is determined by a preset regularization parameter. Alternating normalization iterations are performed on the entropy-regularized transmission kernel matrix. First, the elements of each row are normalized according to the target distribution corresponding to each text tag. Then, the elements of each column are normalized according to the target distribution corresponding to each gesture tag. The row normalization and column normalization processes are repeated until the matrix obtained by two adjacent iterations remains stable within a preset convergence threshold, thus obtaining the optimal transmission matrix. The target distribution is a normalized discrete distribution constructed based on the number of text tags and the number of gesture tags. The text encoding results and gesture encoding results are bidirectionally weighted and aggregated using the optimal transfer matrix. The gesture encoding results are weighted and summed according to the weight distribution of each row in the optimal transfer matrix to obtain the text alignment features. The text encoding results are weighted and summed according to the weight distribution of each column in the optimal transfer matrix to obtain the gesture alignment features. The text alignment features, gesture alignment features, and environment coding results are concatenated, and the concatenated results are subjected to linear transformation, residual addition, and layer normalization to generate cross-modal alignment results.

[0022] In this embodiment, step six specifically includes: The optimal transfer matrix is ​​summed row by row to obtain the text edge distribution corresponding to each text marker, and the optimal transfer matrix is ​​summed column by column to obtain the gesture edge distribution corresponding to each manifold gesture marker. Based on the text edge distribution and gesture edge distribution, the maximum transmission allocation value corresponding to each text marker, the maximum transmission allocation value corresponding to each manifold gesture marker, and the concentrated area range of high transmission allocation values ​​are statistically analyzed in the optimal transmission matrix. The text edge distribution, gesture edge distribution, maximum transmission allocation value corresponding to each text marker, maximum transmission allocation value corresponding to each manifold gesture marker, and concentrated area range are combined to generate a confidence judgment result. The confidence score is normalized to obtain a confidence level, which is then compared with a preset confidence threshold. When the confidence level reaches the preset threshold, the cross-modal alignment result is read and dual-channel tensor decoding is performed. In the control output path of dual-path tensor decoding, the cross-modal alignment result is concatenated with the manifold anchor point feature. The concatenation result is linearly transformed and the residual is updated to generate a control semantic tensor. Based on the control semantic tensor, the target exhibit identifier, target action category, action direction, action amplitude and action duration are determined to generate a digital twin action. Based on the target exhibit identifier and the control semantic tensor, interval mapping and threshold determination are performed on each control component in the control semantic tensor. The control components that satisfy the motor control interval are converted into motor control parameters, and the control components that satisfy the lighting control interval are converted into lighting control parameters. The motor control parameters and lighting control parameters are combined to generate IoT control commands. The dual-path tensor decoding interpreter output path concatenates the text alignment features, gesture alignment features, and environment coding results from the cross-modal alignment results. The concatenated results are then processed by linear transformation, Gaussian error linear unit activation function, and linear transformation to generate an interpreter semantic tensor. Based on the interpreter semantic tensor, natural language interpreter results are output in word order.

[0023] In this embodiment, step seven specifically includes: Read the digital twin action and IoT control instructions output in step six, and extract the target exhibit identifier, target action category, action direction, action amplitude, action duration, as well as motor control parameters and lighting control parameters; Based on the target exhibit identifier, the corresponding digital twin is located in the virtual engine. The initial spatial position and initial spatial posture of the digital twin are read. Based on the target action type, action direction, action amplitude and action duration, an affine transformation matrix sequence is generated. The affine transformation matrix sequence is a combination of translation transformation matrix, rotation transformation matrix or scaling transformation matrix arranged in the action execution order, which is used to drive the digital twin to perform manifold smooth transformation. Based on the motor control parameters and the lighting control parameters, physical electrical signal parameters for motor drive and physical electrical signal parameters for lighting drive are generated respectively. The physical electrical signal parameters are level parameters and timing parameters used to characterize motor start / stop, rotation direction, speed level, as well as lighting on / off, brightness level and duration. The execution time of the affine transformation matrix sequence is aligned with the output time of the physical electrical signal parameters to form a linkage control timing sequence. The linkage control timing sequence is a timing table used to synchronize the actions of the digital twin in the virtual engine with the actions of the motors and lights in the physical exhibition stand. The digital twin in the virtual engine and the motors and lights in the physical exhibition stand are driven synchronously according to the linkage control timing sequence.

[0024] Example 1: To verify the feasibility of this invention in practice, it was applied to a multimodal interactive scenario of a digital twin exhibition hall in the natural sciences category. The hall includes a dinosaur skeleton exhibition area, a mineral profile exhibition area, and a mechanical transmission demonstration exhibition area. It also features a virtual explanation screen, a digital twin engine for the exhibits, and motor and lighting control modules connected to the physical exhibits. Visitors stand approximately 1.5 to 2.5 meters in front of the exhibits and can make requests such as "play the disassembly demonstration," "explain the formation process," or "switch the lights to observe details" by gesturing towards the target exhibit and using voice commands. The system generates digital twin actions, IoT control commands, and natural language explanations based on the user's hand 3D coordinate sequence, voice signals, and local sound field status. To ensure stable implementation, in this embodiment, the depth camera frame rate is set to 30 frames per second, the spatial resolution to 1280×720, and the hand 3D coordinate sampling duration is set to 1.2 seconds continuously. The speech sampling frequency is set to 16 kHz, the speech analysis frame length to 25 milliseconds, the frame shift to 10 milliseconds, the preset background sampling period to 300 milliseconds before the speech start frame, and the preset reverberation sampling period to 500 milliseconds after the speech end frame. The speech start energy threshold is set to 0.12, the end energy threshold is set to 0.05, and the speech zero-crossing rate range is set to 0.06 to 0.21. The preset signal-to-noise ratio (SNR) range is divided into four levels: greater than 20 dB, 10 dB to 20 dB, 0 dB to 10 dB, and less than 0 dB. The preset reverberation time range is divided into four levels: less than 0.4 seconds, 0.4 seconds to 0.8 seconds, 0.8 seconds to 1.2 seconds, and greater than 1.2 seconds. The damping coefficient mapping table has damping values ​​set to discrete values ​​of 0.15, 0.30, 0.45, 0.60, 0.75, and 0.90, respectively. Larger damping values ​​are read when the SNR is low and the reverberation time is long. The word embedding dimension of the improved VL-BERT model is set to 256, the multi-layer Transformer hidden layer is set to 6 layers, the number of attention heads is set to 8, the hidden dimension of the feedforward network is set to 1024, the optimal transmission regularization parameter is set to 0.08, and the preset signal threshold is set to 0.78.

[0025] In practical application within the exhibition hall, the system first continuously acquires the 3D coordinate sequence of the audience's hands using a depth camera, temporally associating the center point of the palm, the wrist point, and the fingertip point to generate continuous sampling frames. Then, a local coordinate system for the hand is constructed, and an inter-frame transformation matrix is ​​generated from the spatial pose changes of adjacent continuous sampling frames. This matrix is ​​then mapped to a special Euclidean group manifold and converted to a Lie algebra representation. Rotation and translation components are decoupled from this representation, ultimately forming a manifold gesture marker. Since audiences often stand at an angle and point while moving their arms within the exhibition hall, misjudgments are easily made using only ordinary 2D coordinates or single-frame pose determination. This embodiment uses rotation components to construct pointing rays, which are then intersected with the 3D bounding boxes in the digital twin engine. This allows for direct targeting of the pointed exhibit, and the semantic identifier of the exhibit is combined with the Lie algebra motion vector to generate manifold anchor point features, further improving the accuracy of subsequent semantic alignment.

[0026] In speech processing, the system simultaneously acquires speech signals and local sound field signals, determines the speech start and end frames using frame energy and zero-crossing rate, and extracts the target speech segment. This embodiment uses the Conformer speech recognition model to output text results, then generates text tags according to word order and speech start and end times. Simultaneously, the system calculates the signal-to-noise ratio (SNR) based on the background reference frame sequence and the speech sampling frame sequence, and calculates the reverberation time based on the reverberation decay frame sequence. For example, when there are fewer people in the venue, the measured average SNR can reach 22.6 dB, and the reverberation time is approximately 0.38 seconds; during peak weekend hours, the average SNR drops to 7.9 dB, and the reverberation time increases to 0.94 seconds. Based on the above classification results, the system reads the optimal transmission cost damping coefficient from a preset damping coefficient mapping table and binds it to the speech start and end times of the target speech segment to generate environmental tags. After this processing, the environmental acoustic state no longer only remains in the front-end noise reduction stage but directly affects the subsequent multimodal alignment process.

[0027] During the joint encoding stage, environmental tags, text tags, and updated manifold gesture tags are simultaneously input into the improved VL-BERT model. Compared to the original VL-BERT model, this embodiment does not change the main structure of its word embedding layer, positional encoding layer, multi-layer Transformer hidden layers, and output layer. Instead, it introduces an environmental damping encoding mechanism and a manifold gesture embedding mechanism in the word embedding layer, and a cross-modal optimal transmission enhancement alignment mechanism in the multi-layer Transformer hidden layers. This preserves the original model's basic modeling ability for cross-modal inputs while unifying environmental noise information, exhibit semantic information, and gesture spatial motion information into the same encoding framework. Especially in high-noise, long-reverberation scenarios, a larger optimal transmission cost damping coefficient increases the unreliable transmission cost between text and gestures, making the system more inclined to rely on stable manifold gesture tags and manifold anchor features to complete exhibit localization and action parsing, thereby avoiding speech recognition errors from being directly amplified into erroneous control commands.

[0028] To verify the practical effectiveness of this invention, comparative tests were conducted on the rule fusion method, the original VL-BERT method, the method that only introduces a manifold gesture embedding mechanism but not a cross-modal optimal transmission enhancement alignment mechanism, and the method of this invention, under the same exhibition hall and the same batch of 1200 interaction samples. The ratio of quiet environment to noisy environment in the test samples was approximately 1:1, and the test content covered four common interactive tasks: exhibit pointing, explanation requests, lighting linkage, and disassembly demonstration. The comprehensive test results are shown in Table 1: Table 1. Comprehensive Comparison Data of Multimodal Interaction Implementation Examples in Exhibition Halls

[0029] As shown in Table 1, although the rule-based fusion method has a shorter response time, it relies mainly on temporal judgment and simple weight concatenation. This results in a significant decrease in exhibit positioning accuracy and noise environment accuracy when there is increased noise in the exhibition hall, oblique pointing by visitors, or continuous changes in actions. It also has the highest false trigger rate. The original VL-BERT method outperforms the rule-based fusion method in text semantic modeling, thus improving instruction parsing accuracy and linkage control success rate. However, it is insufficient in expressing spatial pose changes in gestures, and its handling of environmental sound field disturbances still mainly relies on the front-end speech recognition layer. Therefore, in complex interactive scenarios, situations such as "correct text matching but exhibit pointing deviation" or "correct exhibit positioning but semantic control error" still occur.

[0030] The improved method, which only introduces a manifold gesture embedding mechanism, improves the exhibit localization accuracy by 3.7 percentage points and the accuracy in noisy environments by 5.3 percentage points compared to the original VL-BERT method. This indicates that mapping the inter-frame transformation matrix to a special Euclidean group manifold and obtaining rotation and translation components through Lie algebra decoupling can more stably describe the pointing changes and movement trends of the user's hand in three-dimensional space, solving the problem of traditional gesture representation being sensitive to changes in viewpoint. However, the improvement in command parsing accuracy and false trigger rate is still limited, indicating that gesture space modeling alone is insufficient to completely suppress interference caused by speech misrecognition in noisy environments.

[0031] The method of this invention achieves the highest accuracy in four key indicators: exhibit positioning, command parsing, linkage control success rate, and noise environment accuracy. Specifically, the exhibit positioning accuracy reaches 97.1%, an improvement of 12.5 percentage points compared to the rule-based fusion method and 7.4 percentage points compared to the original VL-BERT method; the command parsing accuracy reaches 95.8%, an improvement of 7.4 percentage points compared to the original VL-BERT method; the noise environment accuracy reaches 93.8%, an improvement of 12.2 percentage points compared to the original VL-BERT method; and the false trigger rate is reduced to 2.6%, only about one-third of that of the rule-based fusion method. This demonstrates that the introduction of the environmental damping coding mechanism and the cross-modal optimal transmission enhancement alignment mechanism enables the system to dynamically adjust the transmission relationship between the text modality and the gesture modality based on the signal-to-noise ratio and reverberation time, more effectively suppressing the disturbance of unreliable text information to the control results in high-noise environments. Furthermore, although the average response time of the method of this invention is 514 milliseconds, slightly higher than the rule-based fusion method, it is still within the acceptable range for real-time interaction in exhibition halls, and in return, it achieves higher positioning accuracy, control stability, and visitor satisfaction. Audience satisfaction improved from 3.88 points for the rule fusion method to 4.71 points, which also confirms the application value of this invention in real-world scenarios from the perspective of user experience.

[0032] As can be seen from the above embodiments, this invention achieves stable representation of gesture spatial pose through decoupling of a special Euclidean group manifold and Lie algebra, accurately binds gesture pointing to exhibit semantic identifiers through manifold anchor point features, explicitly introduces the local sound field state of the exhibition hall into the cross-modal alignment process through the optimal transmission cost damping coefficient, and completes the joint encoding of environmental labels, text labels, and manifold gesture labels through an improved VL-BERT model. Finally, the cross-modal alignment results are stably mapped into digital twin actions, IoT control commands, and natural language interpretation results. Compared with the prior art, this invention not only improves the accuracy of exhibit positioning and interactive intent parsing, but also significantly reduces the false trigger rate and increases the success rate of linkage control between virtual exhibition halls and physical exhibition stands, demonstrating strong engineering feasibility and application value.

[0033] The above description is only a preferred embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any equivalent substitutions or modifications made by those skilled in the art within the scope of the technology disclosed in the present invention, based on the technical solution and inventive concept of the present invention, should be covered within the scope of protection of the present invention.

Claims

1. A multimodal adaptive interaction method for exhibition hall gestures and voice based on digital twins, characterized in that, Includes the following steps: Step 1: Obtain the 3D coordinate sequence of the user's hand, construct the inter-frame transformation matrix and map it to a special Euclidean group manifold, then perform Lie algebra decoupling to obtain rotational and translational components, and generate manifold gesture markers; Step 2: Construct a pointing ray based on the rotation component, and perform intersection operation with the 3D bounding box of each exhibit in the digital twin engine. Concatenate the semantic identifier of the hit exhibit with the corresponding Lie algebra motion vector to generate manifold anchor point features and update the manifold gesture marker. Step 3: Collect user voice signals, signal-to-noise ratio and reverberation time of the local sound field of the exhibition hall, perform automatic speech recognition on the voice signals to generate text tags, and map the signal-to-noise ratio and reverberation time to the optimal transmission cost damping coefficient to generate environmental tags; Step 4: Input the environment tag, text tag, and manifold gesture tag into the improved VL-BERT model for joint encoding. The improved VL-BERT model introduces an environment damping encoding mechanism, a manifold gesture embedding mechanism, and a cross-modal optimal transmission enhancement alignment mechanism. Step 5: Perform an entropy-regularized optimal transport iteration with environmental damping in the improved VL-BERT model to generate cross-modal alignment results and the optimal transport matrix; Step 6: Extract the edge distribution of the optimal transfer matrix as the confidence level, and perform dual-channel tensor decoding when the confidence level reaches the preset threshold to output digital twin actions, IoT control commands, and natural language interpretation results; Step 7: Map the digital twin actions and IoT control commands into affine transformation matrices and physical electrical signal parameters, drive the digital twin in the virtual engine to perform smooth manifold transformations, and drive the motors and lights of the physical display stand to perform linked control.

2. The exhibition hall gesture and voice multimodal adaptive interaction method based on digital twin as described in claim 1, characterized in that, Step one specifically includes: The three-dimensional coordinate sequence of the user's hand is arranged according to the sampling time sequence, and the same hand region in adjacent sampling times is temporally correlated to obtain continuous sampling frames; Based on the center point of the palm, the wrist point, and the fingertip point in each consecutive sampling frame, a corresponding local coordinate system of the hand is constructed to determine the spatial pose of the local coordinate system of the hand. Rigid body registration is performed using the local coordinate system of the hand corresponding to adjacent consecutive sampling frames. The spatial pose change and spatial position change of the subsequent consecutive sampling frame relative to the previous consecutive sampling frame are calculated to generate the corresponding inter-frame transformation matrix. Map the inter-frame transformation matrix to a special Euclidean group manifold, and then convert the mapped inter-frame transformation matrix into its corresponding Lie algebra representation; The Lie algebra representation is decoupled, and the part representing the change in the spatial orientation of the hand is determined as the rotation component, and the part representing the change in the spatial displacement of the hand is determined as the translation component. The rotation and translation components corresponding to multiple consecutive sampling times are sequentially combined to generate manifold gesture markers.

3. The exhibition hall gesture and voice multimodal adaptive interaction method based on digital twin as described in claim 1, characterized in that, Step two specifically includes: The three-dimensional models of each exhibit in the digital twin engine are spatially registered, and a corresponding three-dimensional bounding box is constructed based on the outer contour range of each exhibit's three-dimensional model in the digital twin scene. At the same time, the semantic identifier of the corresponding exhibit is written into each three-dimensional bounding box. Based on the rotation component and the hand's spatial position when generating the manifold gesture mark, the current pointing direction of the hand is determined, and a pointing direction ray is constructed; The pointing ray is input into the digital twin engine, and the pointing ray is made to intersect with each 3D bounding box along the ray extension direction to obtain the candidate hit results corresponding to each pointing ray. The candidate hit results are filtered, and the three-dimensional bounding box that is closest to the spatial position of the hand along the ray extension direction of the pointing direction is retained as the target hit bounding box; The rotation and translation components at the sampling time corresponding to the bounding box of the target are concatenated to generate the Lie algebra motion vector. The semantic identifiers of the hit exhibits are combined with the Lie algebra motion vectors to generate manifold anchor features, which are then written into the manifold gesture tags to obtain the updated manifold gesture tags.

4. The exhibition hall gesture and voice multimodal adaptive interaction method based on digital twin as described in claim 1, characterized in that, Step three specifically includes: Speech signals and local sound field signals are acquired and processed by frame segmentation to obtain speech analysis frame sequences and sound field analysis frame sequences arranged in chronological order. Extract the target speech segment from the speech analysis frame sequence based on the speech start frame and speech end frame; Input the target speech segment into the Conformer speech recognition model, output the text result, and generate text tags according to the word order and the start and end times of the speech; During the preset background sampling period before the speech start frame, a background reference frame sequence is extracted from the sound field analysis frame sequence corresponding to the target speech segment, and the average frame energy of the background reference frame sequence is calculated as the background noise energy. Between the speech start frame and the speech end frame, a speech sampling frame sequence is extracted from the sound field analysis frame sequence corresponding to the target speech segment. The average frame energy of the speech sampling frame sequence is calculated as the speech energy, and the signal-to-noise ratio is determined based on the ratio between the speech energy and the background noise energy. Within a preset reverberation sampling period after the speech termination frame, a reverberation decay frame sequence is extracted from the sound field analysis frame sequence. An energy decay process is constructed from the frame energy of the reverberation decay frame sequence, and the reverberation time is determined. The signal-to-noise ratio is compared with a preset signal-to-noise ratio range to determine the signal-to-noise ratio classification result, and the reverberation time is compared with a preset reverberation time range to determine the reverberation time classification result. The signal-to-noise ratio classification result is used as the row index of the preset damping coefficient mapping table, and the reverberation time classification result is used as the column index of the preset damping coefficient mapping table. The damping value of the corresponding table entry is read, and the damping value is determined as the optimal transmission cost damping coefficient. The optimal transmission cost damping coefficient is bound to the start and end times of the target speech segment to generate an environmental label.

5. The exhibition hall gesture and voice multimodal adaptive interaction method based on digital twin as described in claim 1, characterized in that, The improved VL-BERT model includes a word embedding layer, a positional encoding layer, multiple Transformer hidden layers, and an output layer: The word embedding layer inputs environment tags, text tags, and updated manifold gesture tags into a unified embedding space. It performs an embedding lookup table on the text tags to obtain text word embedding vectors. It introduces an environment damping coding mechanism to perform a linear transformation on the optimal transmission cost damping coefficient in the environment tags and generates environment tag embeddings in combination with the interaction time. It introduces a manifold gesture embedding mechanism to perform embedding mapping and fusion on the semantic identifiers, rotation components, and translation components in the manifold gesture tags to obtain manifold gesture embedding vectors. The location encoding layer writes the environment tag embedding, text word embedding vector, and manifold gesture embedding vector into the location encoding and type encoding according to the interaction time sequence, and concatenates them in the order of environment tag, text tag, and manifold gesture tag to obtain the joint input sequence; The multi-layer Transformer hidden layer inputs the joint input sequence into each Transformer hidden layer, performing query transformation, key transformation, and value transformation respectively to obtain query matrix, key matrix, and value matrix. The basic attention score is calculated based on the matrix product of the query matrix and the key matrix, and the basic attention weight is obtained after scaling and Softmax normalization. The value matrix is ​​weighted and summed to obtain the basic attention output. A cross-modal optimal transmission enhancement alignment mechanism is introduced to construct cross-modal feature pairs between text word embedding vectors and manifold gesture embedding vectors. The initial transmission cost is calculated and modulated in combination with the environmental damping embedding vector to obtain the enhanced attention score. After Softmax normalization, the enhanced attention weight is obtained. The enhanced attention output is then added to the joint input sequence by residual summation and layer normalization, and then input into the feedforward network for processing to obtain the joint encoding result of the current Transformer hidden layer. The output layer inputs the joint encoding result output by the last Transformer hidden layer into the output layer, and performs linear projection on the environment marker position, text marker position and manifold gesture marker position in the joint encoding result to obtain the environment encoding result, text encoding result and gesture encoding result.

6. The exhibition hall gesture and voice multimodal adaptive interaction method based on digital twin as described in claim 1, characterized in that, Step five specifically includes: Read the environment encoding result, text encoding result, and gesture encoding result, and perform temporal alignment of the text encoding result and the gesture encoding result according to the same interaction time to construct a cross-modal feature pair between text tags and gesture tags; For each cross-modal feature pair, the feature difference between the text encoding result and the gesture encoding result is calculated, and the absolute value operation, square operation and weighted summation are performed on the feature difference respectively to generate the basic transmission cost between each text tag and each gesture tag; Each basic transmission cost is written into the initial cost matrix according to the arrangement of text tags in rows and gesture tags in columns. Then, the initial cost matrix is ​​modulated by element-wise multiplication based on the damping intensity corresponding to each interaction moment in the environment coding result to generate the damping cost matrix. An entropy-regularized transmission kernel matrix is ​​constructed based on the damping cost matrix. The entropy-regularized transmission kernel matrix is ​​a transmission weight matrix obtained by performing negative scaling and exponential operations on each element in the damping cost matrix. The scaling magnitude of the negative scaling is determined by a preset regularization parameter. The optimal transfer matrix is ​​obtained by performing alternating normalization iterations on the entropy-regularized transfer kernel matrix. The text encoding results and gesture encoding results are bidirectionally weighted and aggregated using the optimal transfer matrix. The gesture encoding results are weighted and summed according to the weight distribution of each row in the optimal transfer matrix to obtain the text alignment features. The text encoding results are weighted and summed according to the weight distribution of each column in the optimal transfer matrix to obtain the gesture alignment features. The text alignment features, gesture alignment features, and environment coding results are concatenated, and the concatenated results are subjected to linear transformation, residual addition, and layer normalization to generate cross-modal alignment results.

7. The exhibition hall gesture and voice multimodal adaptive interaction method based on digital twin as described in claim 1, characterized in that, Step six specifically includes: The optimal transfer matrix is ​​summed row by row to obtain the text edge distribution corresponding to each text marker, and the optimal transfer matrix is ​​summed column by column to obtain the gesture edge distribution corresponding to each manifold gesture marker. Based on the text edge distribution and gesture edge distribution, the maximum transmission allocation value corresponding to each text marker, the maximum transmission allocation value corresponding to each manifold gesture marker, and the concentrated area range of high transmission allocation values ​​are statistically analyzed in the optimal transmission matrix. The text edge distribution, gesture edge distribution, maximum transmission allocation value corresponding to each text marker, maximum transmission allocation value corresponding to each manifold gesture marker, and concentrated area range are combined to generate a confidence judgment result. The confidence score is normalized to obtain a confidence level, which is then compared with a preset confidence threshold. When the confidence level reaches the preset threshold, the cross-modal alignment result is read and dual-channel tensor decoding is performed. In the control output path of dual-path tensor decoding, the cross-modal alignment result is concatenated with the manifold anchor point feature. The concatenation result is linearly transformed and the residual is updated to generate a control semantic tensor. Based on the control semantic tensor, the target exhibit identifier, target action category, action direction, action amplitude and action duration are determined to generate a digital twin action. Based on the target exhibit identifier and the control semantic tensor, interval mapping and threshold determination are performed on each control component in the control semantic tensor. The control components that satisfy the motor control interval are converted into motor control parameters, and the control components that satisfy the lighting control interval are converted into lighting control parameters. The motor control parameters and lighting control parameters are combined to generate IoT control commands. The dual-path tensor decoding interpreter output path concatenates the text alignment features, gesture alignment features, and environment coding results from the cross-modal alignment results. The concatenated results are then processed by linear transformation, Gaussian error linear unit activation function, and linear transformation to generate an interpreter semantic tensor. Based on the interpreter semantic tensor, natural language interpreter results are output in word order.

8. The exhibition hall gesture and voice multimodal adaptive interaction method based on digital twin as described in claim 1, characterized in that, Step seven specifically includes: Read digital twin actions and IoT control commands to extract target exhibit identifiers, target action categories, action directions, action amplitudes, action durations, as well as motor control parameters and lighting control parameters; Based on the target exhibit identifier, locate the corresponding digital twin in the virtual engine, read the initial spatial position and initial spatial posture of the digital twin, and generate an affine transformation matrix sequence based on the target action type, action direction, action amplitude and action duration; Based on the motor control parameters and the lighting control parameters, generate the physical electrical signal parameters for motor drive and lighting drive respectively; The execution time of the affine transformation matrix sequence is aligned with the output time of the physical electrical signal parameters to form a linkage control timing sequence, and the digital twin in the virtual engine and the motors and lights in the physical display stand are driven synchronously according to the linkage control timing sequence.