Adaptive human-computer interaction method and system based on multi-modal perception

By collecting and processing multimodal perception information streams, detecting and handling modal conflicts in mixed reality environments, the problems of inaccurate and choppy interaction are solved, adaptive human-computer interaction is achieved, and the user experience is improved.

CN121934723BActive Publication Date: 2026-05-29FOSHAN WABON ELECTRONICS TECH
View PDF 2 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
FOSHAN WABON ELECTRONICS TECH
Filing Date
2026-03-27
Publication Date
2026-05-29

Smart Images

  • Figure CN121934723B_ABST
    Figure CN121934723B_ABST
Patent Text Reader

Abstract

The application provides a kind of adaptive human-computer interaction method and system based on multi-modal perception, it is related to human-computer interaction technical field, acquires the multi-modal original perception information flow that user obtains in mixed reality space through depth vision, bone conduction and flexible tactile sensor, respectively to its penetration depth calculation, voiceless speech decoding and mechanical response matching processing are handled, obtain interactive penetration depth sequence, internal speech instruction coding sequence and virtual tactile feedback expected intensity sequence.Multiple modal intention conflict detector detects cross-modal intention conflict, generates intention conflict representation vector, finally generates interactive mode smooth switching instruction according to the intention conflict representation vector, triggers mixed reality space interaction paradigm conversion, realizes the adaptive conversion of interaction paradigm in mixed reality space, effectively solves the problem of multi-modal interaction intention conflict processing difficulty in traditional method, greatly improves the accuracy, fluency and naturalness of human-computer interaction.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of human-computer interaction technology, and more specifically, to an adaptive human-computer interaction method and system based on multimodal perception. Background Technology

[0002] In the field of human-computer interaction, with the rapid development of mixed reality technology, users' demand for more natural, efficient, and immersive interactive experiences is growing. Traditional human-computer interaction methods are often limited to single-modal information input, such as operation via keyboard, mouse, or touchscreen. These interaction methods are difficult to meet the complex and diverse interactive needs of users in mixed reality spaces.

[0003] Currently, although there are some studies on human-computer interaction based on multimodal perception, most of them simply overlay information from different modalities, lacking precise analysis of the deep-seated relationships and contradictions between these modalities. For example, in mixed reality environments, users may interact with virtual objects simultaneously through hand gestures, voice, and touch. However, the intentions expressed by different modalities may conflict. Existing technologies cannot effectively detect and handle these conflicts, leading to inaccurate, choppy, or even erroneous interactions, severely impacting the user's interactive experience and task execution efficiency in mixed reality spaces. Summary of the Invention

[0004] In view of the aforementioned problems, and in conjunction with the first aspect of the present invention, embodiments of the present invention provide an adaptive human-computer interaction method based on multimodal perception, the method comprising:

[0005] The system collects multimodal raw sensory information streams of users in a mixed reality space. These multimodal raw sensory information streams include a point cloud sequence of the dynamic motion trajectory of the user's hand skeleton obtained by a depth vision sensor, a sequence of pre-vibration signals of the user's throat vocalization obtained by a bone conduction sensor, and a sequence of simulated pressing deformation distribution of the user on virtual objects obtained by a flexible tactile sensor.

[0006] The point cloud sequence of the dynamic motion trajectory of the user's hand skeleton is processed by the penetration depth calculation of the preset set of virtual space object bounding boxes to obtain the interaction penetration depth sequence between the user's hand and the virtual space object.

[0007] The user's laryngeal vocalization pre-vibration signal sequence is input into a pre-trained laryngeal vibration decoding network for silent speech decoding processing to obtain the user's internal speech command encoding sequence.

[0008] The simulated pressing deformation distribution sequence of the user on the virtual object is matched with the set of physical simulation attribute parameters of the virtual object to obtain the expected intensity sequence of the virtual tactile feedback of the user on the virtual object.

[0009] The interaction penetration depth sequence, the internal speech instruction encoding sequence, and the virtual haptic feedback expectation intensity sequence are input into a multimodal intent conflict detector for cross-modal intent conflict detection processing to obtain the user's intent conflict representation vector in the mixed reality space.

[0010] Based on the intent contradiction representation vector, the preset interaction mode switching rule base is matched to generate an interaction mode smooth switching instruction corresponding to the intent contradiction representation vector, and the interaction mode smooth switching instruction is sent to the mixed reality rendering engine to trigger the interaction paradigm shift operation of the mixed reality space.

[0011] Furthermore, embodiments of the present invention also provide an adaptive human-computer interaction system based on multimodal perception, characterized in that it includes:

[0012] A processor; a machine-readable storage medium for storing machine-executable instructions of the processor; wherein the processor is configured to execute the above-described adaptive human-computer interaction method based on multimodal perception by executing the machine-executable instructions.

[0013] In another aspect, embodiments of the present invention also provide a computer program product, the computer program product including machine-executable instructions, the machine-executable instructions being stored in a computer-readable storage medium, a processor of a computer device reading the machine-executable instructions from the computer-readable storage medium, the processor executing the machine-executable instructions, causing the computer device to execute the above-described adaptive human-computer interaction method based on multimodal perception.

[0014] Based on the above, by collecting the user's multimodal raw perceptual information stream in the mixed reality space, comprehensive interactive information from various aspects such as hand movements, vocalizations, and touch was acquired. Penetration depth calculation, silent speech decoding, and mechanical response matching were performed on each modality of information to deeply explore the user intent contained in different modalities. A multimodal intent conflict detector was used to detect cross-modal intent conflicts, enabling timely identification of conflicts between different user interaction intents and generating intent conflict representation vectors. Rule matching was performed based on these intent conflict representation vectors to generate smooth switching instructions for interaction modes, achieving adaptive transformation of interaction paradigms in the mixed reality space. This effectively solves the problem of difficult multimodal interaction intent conflict handling in traditional methods, greatly improving the accuracy, fluency, and naturalness of human-computer interaction, and providing users with a higher quality and more efficient mixed reality interaction experience. Attached Figure Description

[0015] Figure 1 This is a schematic diagram of the execution flow of the adaptive human-computer interaction method based on multimodal perception provided in the embodiments of the present invention.

[0016] Figure 2 This is a schematic diagram of exemplary hardware and software components of an adaptive human-computer interaction system based on multimodal perception provided in an embodiment of the present invention. Detailed Implementation

[0017] Figure 1 This is a flowchart illustrating an adaptive human-computer interaction method based on multimodal perception provided in one embodiment of the present invention. The following is a detailed description of this adaptive human-computer interaction method based on multimodal perception.

[0018] Step S110: Collect the user's multimodal raw sensory information stream in the mixed reality space. The multimodal raw sensory information stream includes the point cloud sequence of the user's hand skeleton dynamic motion trajectory obtained by the depth vision sensor, the sequence of the user's throat vocalization pre-vibration signal obtained by the bone conduction sensor, and the sequence of the user's simulated pressing deformation distribution on virtual objects obtained by the flexible tactile sensor.

[0019] In this embodiment, a depth vision sensor array installed on the top of the elevator car's inner wall first continuously acquires depth image data of the user's hand at a frame rate of 30 frames per second. For each frame of depth image, a pre-trained convolutional neural network hand pose estimation model is used to extract the coordinates of 25 three-dimensional key points of the hand with clear anatomical semantics, including the wrist, metacarpophalangeal joints, interphalangeal joints, and fingertips. These coordinates constitute a sparse point cloud describing hand pose and movement. Arranging these key point coordinates in chronological order forms a point cloud sequence of the user's hand skeleton dynamic motion trajectory, denoted as P_hand_seq={P_t|t=1, 2, ..., T}, where P_t is a 25x3 matrix storing the x, y, z spatial coordinates of all key points in the t-th frame in the elevator car coordinate system.

[0020] Simultaneously, a bone conduction microphone sensor array integrated near the call panel on the elevator car wall or into the user's handheld device continuously collects weak mechanical vibration signals from the throat region when the user attempts to issue floor commands via subtle throat movements in a quiet environment. This sensor can capture the pre-vibration of the throat muscles and bones caused by nerve impulses when the user is completely silent or speaking in a very low voice. The acquired raw signal is a one-dimensional time series, denoted as V_larynx_seq={v_tau|tau=1,2,…,Tau}, with a sampling rate of 16 kHz. Each sampling point v_tau represents the amplitude value of the vibration at that moment. To ensure user privacy, the acquired raw vibration signal is immediately feature-extracted on the local device; the raw waveform data is not stored or transmitted.

[0021] Furthermore, a flexible tactile sensor grid integrated below the projection area of ​​the virtual button panel inside the elevator car is used to sense the deformation of the user's fingertip skin when simulating pressing a virtual button in real time. When a user presses a virtual floor button on the elevator's projection interface, the skin on their fingertip undergoes a slight physical deformation due to the simulated tactile feedback. This deformation is captured by the flexible sensor grid as a change in capacitance or resistance. The output of each sensor represents the normal deformation displacement of its covered area. Arranging the readings of all sensors according to their spatial location forms a frame of simulated pressure deformation distribution data of the user pressing a virtual object, denoted as D_tactile_t. This is a matrix of dimension M multiplied by N, where M and N are the number of rows and columns of the sensor grid, respectively. Stacking the above matrix in chronological order yields the simulated pressure deformation distribution sequence D_tactile_seq={D_tactile_t|t=1,2,…,T}.

[0022] Step S120: Perform penetration depth calculation on the point cloud sequence of the dynamic motion trajectory of the user's hand skeleton and the preset set of virtual space object bounding boxes to obtain the penetration depth sequence of the interaction between the user's hand and the virtual space object.

[0023] This step aims to quantify the degree of physical interaction when a user's hand interacts with virtual objects such as virtual buttons. Specifically, the point cloud sequence P_hand_seq of the user's hand skeleton dynamic motion trajectory obtained in step S110 is input into the penetration depth calculation module. This penetration depth calculation module first loads the preset bounding box information of all virtual buttons in the elevator virtual interaction interface. These bounding boxes are usually axis-aligned bounding boxes, used for fast and coarse collision detection. Subsequently, it traverses the point cloud P_t of each frame and calculates its spatial inclusion relationship with the bounding boxes of all virtual buttons.

[0024] Step S121: Perform hand topology reconstruction processing on each frame of the user hand skeleton dynamic motion trajectory point cloud sequence to obtain the user hand triangular mesh surface model corresponding to each frame of the user hand skeleton dynamic motion trajectory point cloud. The user hand triangular mesh surface model includes multiple triangular facet units and the shared vertex connection relationship between the triangular facet units.

[0025] For each frame point cloud P_t in P_hand_seq, a radial basis function-based interpolation algorithm is first used to reconstruct a continuous triangular mesh surface model of the hand with geometric and topological information from the sparse 25 keypoint cloud data. The input is the 3D coordinates of all keypoints in P_t and their predefined adjacency relationships (e.g., which keypoints are connected by bones). The algorithm fits a smooth surface passing through all keypoints by solving a linear system of radial basis functions. Then, triangulation is performed on this surface to generate a mesh composed of multiple triangular facet units, denoted as Mesh_hand_t. This mesh model not only contains the spatial positions of the vertices of the triangular facets but also stores the shared vertex connections between facets, forming a complete graph structure data for subsequent accurate collision detection.

[0026] Step S122: Extract the spatial normal vector direction data of each triangular facet unit in the user's hand triangular mesh surface model, and identify the point cloud clusters in the fingertip region, the knuckle region, and the palm region in the user's hand triangular mesh surface model based on the spatial normal vector direction data.

[0027] For the generated `Mesh_hand_t`, all its triangular facets are traversed, and the normal vector of each facet is calculated. The normal vector is obtained by normalizing the cross product of the two edge vectors constituting the facet. Next, based on the normal vector directions and spatial positions of these facets, a density-based spatial clustering algorithm is used to cluster the mesh vertices. The algorithm clusters vertices with normal vectors pointing outwards from the hand, located at the extremities, and exhibiting drastic curvature changes into a point cloud cluster `C_fingertip` for the fingertip region; vertices with normal vectors parallel to the long axis of the fingers, located in the middle of the fingers, and exhibiting relatively smooth curvature into a point cloud cluster `C_knuckle` for the knuckle region; and vertices with normal vectors approximately perpendicular to the palm plane and located in the center of the palm into a point cloud cluster `C_palm` for the palm region. These point cloud clusters are labeled and stored separately for subsequent differentiated penetration depth calculations for different hand regions.

[0028] Step S123: Perform spatial inclusion relationship detection processing on the point cloud clusters of the fingertip region, the point cloud clusters of the knuckle region, and the point cloud clusters of the palm region with each virtual space object bounding box in the set of virtual space object bounding boxes, and determine at least one target virtual space object bounding box that has contacted and interacted with the user's hand triangular mesh surface model.

[0029] The coordinates of all vertices in the three point cloud clusters C_fingertip, C_knuckle, and C_palm obtained in step S122 are compared with the spatial coordinates of all bounding boxes of the elevator virtual button panel (e.g., bounding box B_A for button A, bounding box B_B for button B, etc.). For a given bounding box, it is defined by the minimum and maximum point coordinates. If the x, y, and z values ​​of a vertex coordinate are within the intervals defined by these two points, it is determined that the point is contained in the bounding box. After traversing all vertices and all bounding boxes, the bounding boxes containing more than a preset threshold (e.g., 3 points) are counted. The bounding boxes that are "hit" are determined as the set of bounding boxes of the target virtual space objects that interact with the user's hand, B_target={B_k|k=1, 2, ..., K}.

[0030] Step S124: Obtain the surface fine mesh model of the virtual space object corresponding to the bounding box of the at least one target virtual space object, and perform signed distance field pre-calculation processing on the surface fine mesh model to generate the external space distance field distribution data of the virtual space object corresponding to the bounding box of the at least one target virtual space object.

[0031] For each bounding box B_k in B_target, a fine mesh model of the corresponding virtual button surface, Mesh_button_k, is loaded from system memory. This fine mesh model contains millions of vertices and triangular faces on the button surface, accurately describing the button's geometry (e.g., slightly convex surfaces, rounded edges). To quickly query the distance and relative position of any point in space to the button surface, a signed distance field pre-computation is performed on Mesh_button_k. This pre-computation process discretizes the local space of Mesh_button_k into a three-dimensional voxel mesh (resolution 256 x 256 x 256). For the center point of each voxel, the nearest distance to the Mesh_button_k surface is calculated, and the distance value is assigned a positive or negative sign depending on whether the point is outside or inside the model. Finally, a corresponding signed distance field data field SDF_k is generated for each B_k. This signed distance field data field is essentially a three-dimensional array storing the signed distance values ​​for each voxel location.

[0032] Step S125: Based on the three-dimensional spatial coordinates of each point in the point cloud cluster of the fingertip region, the point cloud cluster of the knuckle region, and the point cloud cluster of the palm region, query the external spatial distance field distribution data of the virtual space object to obtain the set of external signed distance values ​​of each point on the user's hand triangular mesh surface model relative to the surface of the virtual space object.

[0033] For each vertex v on the hand mesh Mesh_hand_t, with coordinates (x, y, z), it is first transformed from the world coordinate system to the local coordinate system of the target button B_k, obtaining local coordinates (x', y', z'). Then, trilinear interpolation is performed on the voxel mesh of SDF_k based on (x', y', z') to obtain an accurate signed distance value sdf_v. If the signed distance value is positive, it indicates that vertex v is outside the button surface; if it is negative, it indicates that it has penetrated into the button's interior; the magnitude of the absolute value represents the distance from the surface. This query operation is performed on all vertices to obtain the set of signed external distance values ​​S_sdf={sdf_v|vinMesh_hand_t} for all vertices on the hand mesh relative to the virtual object they contact.

[0034] Step S126: Perform absolute value operation on the distance values ​​with values ​​less than zero in the set of external signed distance values ​​to extract the absolute value of the penetration depth of each penetration point on the user's hand triangular mesh surface model that has penetrated the surface of the virtual space object; assign a corresponding weight coefficient to each penetration point according to the hand region type to which each penetration point belongs; use the weight coefficient to perform weighted average processing on the absolute value of the penetration depth to obtain the comprehensive penetration depth characterization parameter of the user's hand triangular mesh surface model relative to the virtual space object.

[0035] From the set S_sdf, all values ​​of sdf_v less than 0 are selected, i.e., penetration points. The absolute value of these negative values ​​is then calculated to obtain the absolute penetration depth d_v = |sdf_v| for each penetration point. Next, based on the hand region type to which each penetration point v belongs (determined by the clustering results in step S122), a predefined weight coefficient is assigned to it. For example, the weight coefficient w_fingertip in the fingertip region point cloud cluster is set to 0.7, the weight coefficient w_knuckle in the knuckle region point cloud cluster is set to 0.2, and the weight coefficient w_palm in the palm region point cloud cluster is set to 0.1. Then, the weighted penetration depth composite value D_pen_t = Σ(w_v*d_v) / Σw_v is calculated for all penetration points. The summation is performed only on the penetration points, where w_v represents a preset weight coefficient assigned to the hand region type (fingertip, knuckle, or palm) of the penetration point v. This preset weight coefficient reflects the difference in importance of different hand regions during the interaction process; for example, the fingertip region contributes more to the penetration depth than the palm region. This D_pen_t is the composite penetration depth representation parameter of the user's hand relative to the virtual button in the current frame t. It is not a simple maximum penetration depth, but rather reflects the overall intensity of the penetration effect across the entire hand's penetration area.

[0036] Step S127: Perform nonlinear mapping processing on the comprehensive penetration depth characterization parameters and the material stiffness property parameters of the virtual space object to generate the interaction force feedback magnitude parameters between the user's hand and the virtual space object in the current frame.

[0037] The material stiffness parameter k_button is obtained from the properties of the virtual button B_k. This material stiffness parameter is a scalar describing the hardness or softness of the object. The combined penetration depth D_pen_t obtained in step S126 and k_button are fed into a nonlinear mapping function f_map, such as a modified Hooke's Law model: F_t = k_button * (D_pen_t)^β, where β is a nonlinear exponent greater than 1, used to simulate the feeling of force increasing nonlinearly with depth when pressed (e.g., the resistance increases sharply when pressing the bottom of a hard plastic button). The calculated F_t is the magnitude parameter of the force feedback between the user and the button in the current frame, which will be used for subsequent haptic rendering.

[0038] Step S128: Construct a time series of interaction force changes between the user's hand and the virtual space object based on the interaction force feedback magnitude parameters corresponding to the point cloud of the dynamic motion trajectory of the user's hand skeleton in multiple consecutive frames, and extract the force loading rate feature and force unloading rate feature from the interaction force change time series.

[0039] The interaction force feedback magnitude parameters {F_{t-14}, F_{t-13}, ..., F_t} calculated in step S127 within a time window (e.g., the past 0.5 seconds, corresponding to 15 frames) are arranged in chronological order to construct an interaction force change time series F_seq. Then, numerical differentiation is performed on this series. For example, the force difference ΔF_i = F_{i+1} - F_i between two adjacent frames is calculated. Statistical analysis is performed on the positive values ​​of ΔF_i, and its average value is taken as the force loading rate feature LF_rate, representing the speed at which the user presses the button; statistical analysis is performed on the negative values ​​of ΔF_i, and its average absolute value is taken as the force unloading rate feature ULF_rate, representing the speed at which the user releases the button.

[0040] Step S129: Perform dynamic time warping matching on the interaction force change time series, the force loading rate feature, and the force unloading rate feature with a preset virtual object grasping operation mode template library to obtain the user's current grasping operation mode recognition result for the virtual space object.

[0041] The system pre-stores a virtual object grasping operation mode template library, which includes various operation mode templates for objects such as buttons, such as the "fast click mode template" T_tap (characterized by fast force loading and unloading rates and high peak force), the "slow press mode template" T_press (characterized by slow force loading and unloading rates and long peak force holding time), and the "hover touch mode template" T_hover (characterized by force values ​​always close to 0, but with hand contact with the bounding box). The F_seq, LF_rate, and ULF_rate obtained in step S128 are combined into a multi-dimensional feature vector V_interact. Then, a dynamic time warping algorithm is used to calculate the similarity distance between V_interact and each template T in the template library. The dynamic time warping algorithm allows two sequences to be non-linearly aligned on the time axis, thus matching patterns of different lengths but similar shapes. The operation mode corresponding to the template with the smallest similarity distance to V_interact is selected as the current grasping operation mode recognition result M_rec, for example, M_rec = "fast click".

[0042] Step S1210: Perform feature association encoding processing on the current grasping operation mode recognition result and the comprehensive penetration depth characterization parameter to generate the interaction penetration depth sequence between the user's hand and the virtual space object. The interaction penetration depth sequence includes the penetration depth values ​​on consecutive time frames and the corresponding grasping operation mode labels.

[0043] The operation pattern recognition result M_rec identified in step S129 (e.g., a one-hot encoded vector, such as [1, 0, 0] representing "quick click") is associated with the comprehensive penetration depth representation parameter D_pen_t calculated in step S126 for feature association encoding. Specifically, the one-hot encoded vector of M_rec is concatenated with the scalar D_pen_t to form a new feature vector with dimension 3+1, denoted as P_enc_t=[D_pen_t, M_rec_vector], where M_rec_vector represents the vector representation formed after converting the operation pattern recognition result M_rec into one-hot encoded vectors. For example, if the operation pattern is "quick click", the corresponding one-hot encoded vector is [1, 0, 0]. This P_enc_t is an element of the interaction penetration depth sequence of the current frame. Repeat this process for multiple consecutive frames to obtain the final interaction penetration depth sequence P_enc_seq={P_enc_t|t=1,2,…,T}. Each frame of data in this interaction penetration depth sequence not only contains the penetration depth value, but also explicitly associates it with the current grabbing operation mode.

[0044] Step S130: Input the user's laryngeal vocalization pre-vibration signal sequence into a pre-trained laryngeal vibration decoding network for silent speech decoding processing to obtain the user's internal speech instruction encoding sequence.

[0045] This step aims to decode the weak vibration signals from the user's throat into semantic instructions. The pre-vibration signal sequence V_larynx_seq of the user's throat, obtained in step S110, is input into a pre-trained deep learning network, which is called the throat vibration decoding network.

[0046] Step S131: Perform time-frequency domain transformation processing on each segment of the user's laryngeal pre-vibration signal sequence to obtain the laryngeal vibration spectrum corresponding to each segment of the user's laryngeal pre-vibration signal. The laryngeal vibration spectrum contains information on the distribution pattern of frequency component intensity changing with time.

[0047] First, V_larynx_seq is divided into frames, each 25 milliseconds long with a frame shift of 10 milliseconds. For each frame, v_frame, a Short-Time Fourier Transform (SFT) is applied. The SFT is implemented by multiplying v_frame by a Hanning window function, followed by a Fast Fourier Transform (FFT) to convert the time-domain signal to the frequency domain. The result is the spectrum of the frame, with frequency on the horizontal axis and amplitude of the frequency component on the vertical axis. The spectra of all frames are stacked in chronological order to form a two-dimensional time-spectrum, or throat vibration spectrum, denoted as S_spec. In this throat vibration spectrum, the horizontal axis represents the time frame, the vertical axis represents the frequency, and the brightness of each pixel represents the energy intensity of a specific frequency at a given moment.

[0048] Step S132: Perform spectral energy centroid calculation processing on the throat vibration spectrum diagram, and extract the resonance peak frequency trajectory features in the throat vibration spectrum diagram. The resonance peak frequency trajectory features include the center frequency change curve of the first resonance peak, the center frequency change curve of the second resonance peak, and the center frequency change curve of the third resonance peak.

[0049] For each frame (i.e., each time slice) of the laryngeal vibration spectrum S_spec, linear predictive coding analysis or peak detection algorithms are used to locate local maxima in the spectral envelope. The frequencies corresponding to these maxima are the formants. Specifically, in the spectrum of each frame, a search is performed from low to high frequencies to find the centers of the first three energy concentration bands, denoted as F1_t, F2_t, and F3_t, respectively. Connecting the F1_t values ​​of multiple consecutive frames yields the frequency change curve F1_curve of the first formant center; similarly, F2_curve and F3_curve are obtained. These curves together constitute the formant frequency trajectory characteristics of laryngeal vibration, which are key acoustic features for distinguishing different vowels and consonants.

[0050] Step S133: Perform throat vibration waveform envelope extraction processing on the user's throat vocalization pre-vibration signal sequence to obtain the throat vibration amplitude envelope curve corresponding to the user's throat vocalization pre-vibration signal, and detect the start time and stop time of the user's throat vocalization based on the distribution of local maxima of the throat vibration amplitude envelope curve.

[0051] The V_larynx_seq signal is subjected to a Hilbert transform to calculate its analytic signal, thus obtaining the instantaneous amplitude, i.e., the signal envelope. This envelope curve is then smoothed using a low-pass filter to obtain the smoothed throat vibration amplitude envelope curve A_env. An energy threshold is then set on curve A_env. When the envelope value rises from below the threshold to above the threshold, the point is marked as the start-up time point T_start; when the envelope value falls from above the threshold to below the threshold, the point is marked as the end-up time point T_end. By detecting all these start-end point pairs, the continuous vibration signal can be segmented into individual vocal units.

[0052] Step S134: Based on the vibration start time and vibration stop time, the user's throat vocalization pre-vibration signal sequence is processed into silent speech units to obtain multiple continuous silent speech unit segments, each silent speech unit segment corresponding to a complete internal speech pronunciation action cycle.

[0053] Based on all the oscillation start time points T_start_i and oscillation stop time points T_end_i detected in step S133, the original V_larynx_seq signal is divided into multiple continuous segments. Each segment V_seg_i = V_larynx_seq[T_start_i:T_end_i] corresponds to the complete pronunciation cycle of a syllable or word completed by the user. These segments are the basic units for subsequent decoding.

[0054] Step S135: Input the laryngeal vibration spectrum corresponding to each silent speech unit segment into the feature encoding layer of the pre-trained laryngeal vibration decoding network for deep feature extraction processing to obtain the deep feature vector of laryngeal vibration corresponding to each silent speech unit segment.

[0055] For each segmented silent speech unit V_seg_i, its corresponding laryngeal vibration spectrogram S_spec_i is first generated according to the method in step S131. Then, S_spec_i is input to the feature encoding layer of the laryngeal vibration decoding network. This feature encoding layer is a convolutional neural network composed of multiple stacked two-dimensional convolutional layers, batch normalization layers, and max-pooling layers. S_spec_i serves as the input tensor, passing through each convolutional layer sequentially. Each convolutional layer performs a convolution operation with the input feature map using a learnable convolutional kernel to extract local time-frequency patterns. Through multiple stacked layers, the network gradually abstracts high-level semantic features related to articulation from low-level features such as edges and textures. The output of the last convolutional layer is flattened to obtain a high-dimensional feature vector, namely the deep laryngeal vibration feature vector F_deep_i.

[0056] Step S136: Input the deep feature vector of laryngeal vibration into the attention mechanism layer of the pre-trained laryngeal vibration decoding network for keyframe attention weight allocation processing, and generate the contribution weight distribution of different time frames in the deep feature vector of laryngeal vibration to silent speech decoding.

[0057] The deep feature vector F_deep_i of throat vibration is fed into a multi-head self-attention mechanism layer. This layer first transforms F_deep_i through three different linear transformations to generate a query matrix Q, a key matrix K, and a value matrix V. Then, it calculates the attention score matrix Attention_Score = softmax(Q*K^T / √d_k), where K^T represents the transpose of the key matrix K, used to calculate the dot product similarity between the query matrix Q and the key matrix K in the attention mechanism, thereby determining the attention weight allocation between different time frames. Here, d_k is the dimension of the key vector. Each element in this score matrix represents the attention weight of one time frame (or feature position) to another time frame. By performing a row-weighted summation of the score matrix, i.e., Attention_Output = Attention_Score * V, a new, context-enhanced feature representation is obtained. The attention mechanism enables the network to automatically focus on the most important parts of the input sequence for the current decoding task, such as the pronunciation time of certain key phonemes, and output a weight distribution that indicates the contribution of each time step or feature dimension in F_deep_i to the final decoding.

[0058] Step S137: Perform weighted summation and aggregation on the deep feature vector of laryngeal vibration according to the contribution weight distribution to obtain the semantic coding feature vector corresponding to each silent speech unit segment.

[0059] The contribution weight distribution output from the attention mechanism layer in step S136 (e.g., a vector A_att with the same time dimension as F_deep_i) is multiplied element-wise with the original deep laryngeal vibration feature vector F_deep_i to achieve feature weighting. Then, the weighted features are summed and pooled along the time dimension to obtain a fixed-dimensional semantic encoding feature vector F_semantic_i that aggregates information from the entire articulatory unit. This vector is a compact semantic representation of the silent speech unit segment.

[0060] Step S138: Input the semantic encoding feature vector into the classification output layer of the pre-trained laryngeal vibration decoding network to perform phoneme probability distribution prediction processing, and obtain the internal speech phoneme sequence probability distribution corresponding to each silent speech unit segment.

[0061] The semantically encoded feature vector F_semantic_i is fed into the final part of the laryngeal vibration decoding network, namely the classification output layer. This classification output layer typically consists of one or more fully connected layers and a Softmax activation function. F_semantic_i is first linearly transformed through a fully connected layer, mapping it to a space with a dimension equal to the total number of phoneme categories, obtaining the raw score for each phoneme category. Then, the Softmax function converts these scores into probability values, with the sum of all probability values ​​being 1. Finally, the probability distribution P_phoneme_i={p_1, p_2, ..., p_C} of the internal speech phoneme sequence corresponding to the silent speech unit segment is obtained, where C is the total number of phonemes (e.g., for Chinese, it could be vowels and consonants with tones). This probability distribution represents the likelihood that the segment belongs to each of the phonemes.

[0062] Step S139: Input the probability distribution of the internal speech phoneme sequence corresponding to multiple consecutive silent speech unit segments into the Viterbi decoder for optimal path search processing to generate the phoneme-level coding unit in the internal speech instruction coding sequence.

[0063] The phoneme probability distribution sequence {P_phoneme_1, P_phoneme_2, ..., P_phoneme_N} obtained from multiple consecutive segments is input into a Viterbi decoder. This decoder internally maintains a language model that statistically analyzes the probabilities of transitions between phonemes (i.e., a bigram grammar model). Based on the observed probability distribution and the phoneme transition probabilities, the Viterbi decoding algorithm dynamically searches for the path with the highest probability in a grid of all possible phoneme sequences. The phoneme sequence corresponding to this path is the most probable phoneme-level coding unit sequence, denoted as Phone_seq=[phone_1, phone_2, ..., phone_M]. For example, the decoded phoneme sequence might be / b / , / ei / , / 3 / , / j / , / i1 / , corresponding to the silent pronunciation of the word "Beijing".

[0064] Step S1310: Perform sequence matching processing between the phoneme-level encoding unit and the standard instruction phoneme template in the preset instruction lexicon to obtain the word-level encoding unit in the internal speech instruction encoding sequence. The internal speech instruction encoding sequence is used to represent the interactive operation intention content expressed by the user through silent speech.

[0065] The system pre-defines a dedicated elevator instruction vocabulary, including terms like "first floor," "second floor," "open door," "close door," and "alarm." Each instruction pre-stores its corresponding standard phoneme sequence template; for example, "second floor" corresponds to the templates ['e'r', 'l'ou'] (these are just illustrative phonemes). The phoneme-level encoding unit sequence Phone_seq obtained in step S139 is matched against each template in the vocabulary. The matching algorithm can employ a dynamic programming algorithm based on edit distance to calculate the similarity score between Phone_seq and each template. The instruction corresponding to the template with the highest similarity score exceeding a preset threshold is selected as the word-level encoding unit for the current internal speech instruction, denoted as Word_i. For example, if Phone_seq best matches the template for "second floor," then Word_i = "second floor." The above matching is performed on all decoded speech segment sequences to finally obtain the internal speech instruction encoding sequence Word_seq=[Word_1, Word_2, ...]. This internal speech instruction encoding sequence represents the user's continuous interactive operation intention expressed through silent speech, such as ["second floor", "open the door"].

[0066] Step S140: Perform mechanical response matching processing on the simulated pressing deformation distribution sequence of the user on the virtual object and the physical simulation attribute parameter set of the virtual object to obtain the expected intensity sequence of the user's virtual tactile feedback on the virtual object.

[0067] This step aims to infer the intensity of tactile feedback that the user expects to receive based on the user's physical pressing behavior of the virtual button. The simulated pressing deformation distribution sequence D_tactile_seq of the virtual object obtained in step S110 is input into the mechanical response matching processing module.

[0068] Step S141: Perform spatial interpolation processing on each frame of simulated pressing deformation distribution data in the simulated pressing deformation distribution sequence of the user on the virtual object to generate a continuous deformation displacement field corresponding to each frame of simulated pressing deformation distribution data. The continuous deformation displacement field includes the normal displacement of each sampling point on the surface of the virtual object.

[0069] For each frame of data D_tactile_seq, D_tactile_t is a sparse grid representing the deformation displacement at discrete sensor positions. To obtain continuous deformation information across the entire finger contact area, a bicubic spline interpolation algorithm is used, with the grid points of D_tactile_t and their displacement values ​​as control points, to fit a smooth two-dimensional surface function. This function can then be used to calculate the normal displacement at any continuous coordinate point (x, y) within the entire rectangular area covered by the sensor grid. Discretizing this interpolation result onto a finer virtual sampling point grid (e.g., 256 x 256 points) generates a continuous deformation displacement field Field_def_t.

[0070] Step S142: Extract the local deformation regions in the continuous deformation displacement field where the deformation displacement exceeds the preset deformation sensing threshold, and perform curvature analysis on the boundary contour of the local deformation regions to obtain the boundary curvature distribution characteristics and the region area expansion rate characteristics of the local deformation regions.

[0071] In the deformation displacement field Field_def_t, a deformation sensing threshold Th_def (e.g., 0.1 mm) is set. All sampling points with displacement greater than Th_def are marked, forming one or more connected regions, i.e., local deformation regions R_def. The boundary contour point set of each connected region is extracted. Curve fitting is performed on the above boundary points, and the curvature of each point on the boundary is calculated to obtain the boundary curvature distribution feature Curv_dist of the region. At the same time, the area change of the same local deformation region in multiple consecutive frames is tracked, and its area change rate over time is calculated, i.e., the region area expansion rate feature V_area_exp. A rapidly expanding region area may indicate that the user is pressing a button with force.

[0072] Step S143: Extract curvature feature values ​​from the boundary curvature distribution characteristics of the local deformation region to characterize the pressing operation characteristics; extract tension feature values ​​from the surface tension attribute parameters in the physical simulation attribute parameter set of the virtual object to characterize the material response characteristics; perform matching degree calculation after standardizing the curvature feature values ​​and the tension feature values ​​to generate the surface tension response matching degree parameter generated by the user's pressing operation on the virtual object in the local deformation region.

[0073] From the boundary curvature distribution feature Curv_dist obtained in step S142, its mean and variance are extracted as curvature feature values, denoted as C_op, to characterize the pressing operation. From the set of physical simulation attribute parameters of the virtual button, its surface tension attribute parameters, such as the surface tension coefficient sigma, are obtained as tension feature values, denoted as S_mat, to characterize the material response. To compare the characteristics of these two different physical quantities, they are first Z-score standardized to a mean of 0 and a variance of 1, resulting in C_op_norm and S_mat_norm. Then, their matching degree is calculated, for example, by calculating the Euclidean distance or Pearson correlation coefficient between them. Here, a Gaussian similarity function can be used: M_tension=exp(-||C_op_norm-S_mat_norm||^2 / (2*δ^2)), where δ is a width parameter. The calculated M_tension is the surface tension response matching degree parameter, which quantifies the degree of fit between the user-induced deformation curvature and the button's surface tension attribute.

[0074] Step S144: Based on the elastic modulus attribute parameters and simulated compression deformation distribution data, calculate the theoretical area expansion rate reference value of the virtual object under the compression operation; after standardizing the area expansion rate characteristics of the local deformation region and the theoretical area expansion rate reference value, calculate the matching degree to generate the elastic deformation response matching degree parameter generated by the user's compression operation on the virtual object in the local deformation region.

[0075] The elastic modulus attribute parameter E_button is obtained from the set of physical simulation attribute parameters of the virtual button. This elastic modulus attribute parameter is used to characterize the ability of the virtual object to resist elastic deformation under force and is a key mechanical parameter in the set of physical simulation attribute parameters of the virtual object. Combined with the deformation displacement field Field_def_t obtained in step S141, based on the Hertzian contact theory model in elasticity, the theoretically expected area expansion rate can be estimated under a given deformation depth and contact area. This theoretical value is denoted as V_area_theo. The area expansion rate feature V_area_exp actually observed in step S142 and the theoretical reference value V_area_theo are respectively Z-score normalized to obtain V_area_exp_norm and V_area_theo_norm. Then, the matching degree between the two is calculated using the Gaussian similarity function: M_elastic=exp(-||V_area_exp_norm-V_area_theo_norm||^2 / (2*δ^2)). The calculated M_elastic is the elastic deformation response matching parameter, which reflects whether the expansion of the deformation area caused by the user's pressure conforms to the elastic behavior that the object should have.

[0076] Step S145: Calculate the stress concentration region distribution map inside the virtual object based on the spatial gradient distribution of the deformation displacement in the continuous deformation displacement field, and extract the stress peak coordinates and stress peak intensity of each stress concentration region in the stress concentration region distribution map.

[0077] The spatial gradient of the deformation displacement field Field_def_t is calculated to obtain the displacement gradient tensor field. Based on the geometric and physical equations of elasticity, the stress distribution inside the virtual object can be approximately estimated using the displacement gradient tensor field. Specifically, at each sampling point, the displacement gradient tensor is multiplied by the elastic modulus E_button to obtain an approximate stress tensor. The principal stress values ​​of the stress tensor are extracted to form a stress concentration region distribution map Map_stress inside the virtual button. On this distribution map, by finding local maxima, the center coordinates P_stress_i and the corresponding peak stress intensity S_peak_i of the stress concentration region are located.

[0078] Step S146: Calculate the spatial distance between the stress peak coordinates and the coordinates of each vulnerable point in the internal structural vulnerability distribution map; determine the proximity between the stress concentration point and the structural vulnerability point based on the spatial distance, and generate the structural response matching degree parameter generated by the user's pressing operation on the virtual object within the virtual object.

[0079] The system pre-stores a map (Map_weak) of the internal structural vulnerabilities of the virtual button. This map identifies the coordinates of vulnerable or critical locations within the button's internal structure, such as circuit contact locations (P_weak_j), and is used to assess the impact of pressing operations on the internal structure of the virtual object. For each stress concentration point (P_stress_i) found in step S145, the spatial Euclidean distance d_ij = ||P_stress_i - P_weak_j|| from it to all structural vulnerabilities (P_weak_j) is calculated. The minimum value among all d_ij is taken as the proximity d_min between the stress concentration point and the structural vulnerabilities. Then, the distance d_min is converted into a matching degree using a monotonically decreasing function (e.g., a Gaussian function): M_structure = exp(-d_min^2 / (2*σ_struct^2)). This M_structure is the structural response matching degree parameter, which quantifies whether the user's pressing action caused stress concentration in the critical vulnerable parts of the object.

[0080] Step S147: Construct a multi-dimensional mechanical response feature vector of the user's pressing operation on the virtual object in the current frame based on the surface tension response matching parameter, the elastic deformation response matching parameter, and the structural response matching parameter.

[0081] The three matching parameters M_tension, M_elastic, and M_structure calculated in steps S143, S144, and S146 are combined to form a three-dimensional feature vector, denoted as V_mech_t=[M_tension, M_elastic, M_structure]. This vector comprehensively characterizes the degree of matching between the user's pressing operation in the current frame and the physical properties of the virtual object in three mechanical dimensions: surface tension, elastic deformation, and internal structure.

[0082] Step S148: Input the multi-dimensional mechanical response feature vector of the pressing operation into the pre-trained tactile feedback expectation prediction network for expectation intensity regression processing to obtain the preliminary prediction value of the user's virtual tactile feedback expectation intensity for the virtual object in the current frame.

[0083] The multi-dimensional mechanical response feature vector V_mech_t obtained in step S147 is input into a pre-trained haptic feedback expectation prediction network. This haptic feedback expectation prediction network is a multi-layer fully connected neural network with 3 nodes in the input layer, corresponding to the three dimensions of V_mech_t. The network contains two hidden layers, each with 128 neurons and the ReLU activation function. The output of the last hidden layer is connected to the output layer, which has only one neuron and no activation function, used to regress a continuous value. V_mech_t propagates forward in the network: first, the input layer receives the data, then multiplies it with the weight matrix of the first hidden layer and adds a bias, then performs ReLU activation to obtain the output of the first hidden layer; this output is then operated on with the weight matrix of the second hidden layer, and again performed ReLU activation to obtain the output of the second hidden layer; finally, it is operated on with the weight matrix of the output layer to obtain the final scalar output. This output is the preliminary prediction value of the user's virtual haptic feedback expectation intensity in the current frame, denoted as E_pre_t.

[0084] Step S149: Obtain the historical sequence of the user's expected virtual haptic feedback intensity in historical interaction frames, and perform temporal trend extrapolation processing on the historical sequence of the expected virtual haptic feedback intensity to generate the temporal correction factor of the user's expected virtual haptic feedback intensity in the current frame.

[0085] The system reads preliminary predicted values ​​of the expected intensity of virtual haptic feedback within a sliding time window (e.g., the past 1 second, 30 frames) from system memory, forming a historical sequence E_pre_seq={E_pre_{t-29}, E_pre_{t-28}, ..., E_pre_{t-1}}. Linear regression or simple exponential smoothing is applied to this sequence to fit a trend line. For example, using linear regression with the time step as the independent variable and E_pre as the dependent variable, a linear equation E=a*t+b is fitted. Substituting the current time step t into this equation yields the extrapolated value E_trend_t based on the historical trend. Then, the time-series correction factor α_t is calculated. This can be defined as the ratio of the extrapolated value to the mean of the historical sequence, or calculated using a function such as α_t=sigmoid(E_trend_t-E_pre_t). The purpose is to smooth the predicted values ​​and avoid severe haptic feedback jitter caused by single-frame noise.

[0086] Step S1410: Perform weighted correction processing on the preliminary predicted value of the virtual haptic feedback expectation intensity according to the time sequence correction factor to obtain the accurate value of the virtual haptic feedback expectation intensity of the user to the virtual object in the current frame, and perform serialization assembly processing on the accurate values ​​of the virtual haptic feedback expectation intensity of multiple consecutive frames to generate the virtual haptic feedback expectation intensity sequence of the user to the virtual object.

[0087] The preliminary predicted value E_pre_t obtained in step S148 is weighted and corrected by the temporal correction factor α_t obtained in step S149. The correction formula is E_refine_t = α_t * E_pre_t + (1 - α_t) * E_trend_t. This E_refine_t is the final determined precise value of the user's expected virtual haptic feedback intensity in the current frame. Steps S141 to S1410 are repeated for each frame to obtain precise values ​​E_refine_t for multiple consecutive frames. The above values ​​are assembled in chronological order to finally generate the user's expected virtual haptic feedback intensity sequence E_refine_seq = {E_refine_1, E_refine_2, ..., E_refine_T}. This sequence reflects the change in the haptic feedback force that the user expects to receive from the virtual button throughout the entire interaction process.

[0088] Step S150: Input the interaction penetration depth sequence, the internal speech instruction encoding sequence, and the virtual haptic feedback expectation intensity sequence into the multimodal intent conflict detector for cross-modal intent conflict detection processing to obtain the user's intent conflict representation vector in the mixed reality space.

[0089] This step aims to fuse information from three modalities: hand gestures, silent speech, and tactile expectations, to detect any contradictions between user intentions. The interaction penetration depth sequence P_enc_seq obtained in step S1210, the internal speech instruction encoding sequence Word_seq obtained in step S1310, and the virtual tactile feedback expectation intensity sequence E_refine_seq obtained in step S1410 are fed as three parallel inputs into a neural network model called a multimodal intention contradiction detector.

[0090] Step S151: Perform temporal feature extraction processing on the interaction penetration depth sequence to obtain the first intention trend feature vector of the user's hand interaction operation. The first intention trend feature vector includes operation direction consistency and operation force change stability parameters.

[0091] The interaction penetration depth sequence P_enc_seq is input into a Long Short-Term Memory (LSTM) network layer. This LSM layer contains 64 hidden units. Each frame of data in P_enc_seq is used as a time step input and sequentially fed into the LSM network units. Each unit updates its current hidden state h_t based on the current input P_enc_t and the hidden state h_{t-1} from the previous time step, through its internal input gate, forget gate, and output gate structure. After the entire sequence is processed, the hidden state h_T from the last time step is taken as the aggregated feature of the entire sequence. Then, this aggregated feature is input into a fully connected layer for dimensionality reduction and transformation, finally outputting a fixed-dimensional first intention trend feature vector F_hand. Some dimensions of this first intention trend feature vector are trained to encode operation direction consistency parameters (e.g., whether the hand movement trajectory is consistent with the direction of pointing to the target button) and operation force change smoothness parameters (e.g., whether the change in penetration depth is smooth, and whether there is violent shaking).

[0092] Step S152: Perform semantic feature extraction processing on the internal speech instruction encoding sequence to obtain the second intention trend feature vector of the user's internal speech expression. The second intention trend feature vector includes the instruction semantic clarity and instruction sentiment intensity parameters.

[0093] The internal speech instruction encoding sequence Word_seq is a sequence of words. First, each word Word_i is mapped to a low-dimensional dense word vector e_i through a word embedding layer. Then, the resulting word vector sequence [e_1, e_2, ...] is input into a Transformer encoder. The Transformer encoder captures long-distance dependencies between words through its multi-head self-attention mechanism and outputs a context-aware sequence representation. The output sequence is average-pooled to obtain a fixed-length vector. Finally, this vector is input into a fully connected layer to obtain the second intention trend feature vector F_speech. Part of the dimensions of this second intention trend feature vector are trained to encode the semantic clarity parameter of the instruction (e.g., measured by calculating the distribution entropy of the word vectors in the semantic space) and the sentiment intensity parameter of the instruction (e.g., outputting the probability of the user instruction being urgent or peaceful through an additional sentiment classification head).

[0094] Step S153: Perform expectation feature extraction processing on the virtual haptic feedback expectation intensity sequence to obtain the third intention trend feature vector of the user's haptic expectation expression. The third intention trend feature vector includes the expectation intensity change rate and expectation spatial distribution concentration parameters.

[0095] The virtual haptic feedback expected intensity sequence E_refine_seq is also input into another long short-term memory network layer with the same structure but independent weights. The processing flow is similar to step S151. The temporal dynamic characteristics of the sequence are captured through the long short-term memory network, the hidden state at the last moment is taken, and then it is passed through a fully connected layer to finally obtain the third intention trend feature vector F_touch. Some dimensions of this third intention trend feature vector are trained to encode the expected intensity change rate parameter (e.g., the first-order difference mean and variance of the sequence) and the expected spatial distribution concentration parameter (this parameter may not be directly applicable here, but can be understood as the degree of concentration of expected intensity in the finger contact area. This information can be extracted and fused from the deformation displacement field in step S141. For the sake of simplicity, it is assumed that it is already implicit in E_refine_seq).

[0096] Step S154: Input the first intention trend feature vector, the second intention trend feature vector, and the third intention trend feature vector into the feature alignment layer of the multimodal intention contradiction detector for cross-modal feature space mapping processing to obtain the first mapped feature vector, the second mapped feature vector, and the third mapped feature vector under a unified feature space.

[0097] The three feature vectors, F_hand, F_speech, and F_touch, from different modalities, are input into the feature alignment layer. This feature alignment layer consists of three independent fully connected layers with the same output dimension, each used to map the features of its respective modality to a unified D-dimensional feature space (e.g., D=128). Specifically, F_hand undergoes a linear transformation through the fully connected layers W_hand and b_hand to obtain the first mapped feature vector M_hand; F_speech is transformed through W_speech and b_speech to obtain M_speech; and F_touch is transformed through W_touch and b_touch to obtain M_touch. These three mapped feature vectors now reside in the same D-dimensional space, making cross-modal comparisons possible.

[0098] Step S155: Calculate the cosine similarity of the feature space between the first mapped feature vector and the second mapped feature vector to obtain the first modal intent similarity; calculate the Euclidean distance of the feature space between the first mapped feature vector and the third mapped feature vector to obtain the second modal intent difference; calculate the Manhattan distance of the feature space between the second mapped feature vector and the third mapped feature vector to obtain the third modal intent difference; process the first modal intent similarity, the second modal intent difference, and the third modal intent difference through a preset normalization function, mapping them to a comparable scale within the same numerical range to obtain the first modal intent difference between limb movement intent and verbal expression intent. Figure 1 The second modality of consistency coefficient, limb movement intention and tactile expectation intention Figure 1 The consistency coefficient, and the third modality between verbal expression intention and tactile expectation intention. Figure 1 Consistency coefficient.

[0099] In a unified feature space, the cosine similarity between M_hand and M_speech is calculated as S_cos = (M_hand·M_speech) / (||M_hand|| * ||M_speech||). The closer this value is to 1, the more semantically consistent the hand gesture intention and the speech intention are. The Euclidean distance between M_hand and M_touch is calculated as D_euc = sqrt(Σ(M_hand_i - M_touch_i)^2). The larger this distance, the greater the difference between the hand gesture and the tactile expectation. The Manhattan distance between M_speech and M_touch is calculated as D_man = Σ|M_speech_i - M_touch_i|. This distance is also used to measure the difference. Since S_cos, D_euc, and D_man have different dimensions and ranges, they need to be normalized to the same comparable scale. For example, S_cos, which is itself in the range [-1, 1], can be mapped to a consistency coefficient C_hand_speech = (S_cos + 1) / 2, making it fall within the interval [0, 1]. For distance metrics, this can be converted to similarity using a Gaussian kernel function: C_hand_touch = exp(-D_euc^2 / σ^2), C_speech_touch = exp(-D_man^2 / σ^2). This yields three consistency coefficients, each ranging from 0 to 1, representing the intention between the three modal pairs. Figure 1 Degree of consistency.

[0100] Step S156: According to the first modal intention Figure 1 Consistency coefficient, second mode meaning Figure 1 Consistency coefficient and the third mode intention Figure 1 The construction intention of the consistency coefficient Figure 1 Consistency coefficient matrix, and for the meaning of the above. Figure 1 Principal component analysis was performed on the consistency coefficient matrix to extract the meaning. Figure 1 The principal component with the largest eigenvalue in the consistency coefficient matrix is ​​used as the global meaning. Figure 1 Consistency characterization scalar.

[0101] The three consistency coefficients C_hand_speech, C_hand_touch, and C_speech_touch obtained in step S155 are combined into a 3-dimensional vector [C_hand_speech, C_hand_touch, C_speech_touch]. This is considered a 1x3 matrix. Principal component analysis (PCA) is then performed on this vector. First, the mean of the vector is subtracted to center it. Then, its covariance matrix is ​​calculated (the low dimension of the vector simplifies this process). In fact, the goal of PCA is to find the direction with the largest data variance. For this 3D vector, its first principal component can be obtained by solving for the eigenvalues ​​and eigenvectors of the covariance matrix. The eigenvector corresponding to the largest eigenvalue is extracted, and then the centered vector is multiplied by this eigenvector to obtain a scalar value, i.e., the global significance. Figure 1 The consistency scalar is G_consistency. This G_consistency integrates the consistency information between each pair of the three modalities, representing the degree of coordination of the current user's overall intent.

[0102] Step S157: Combine the first mapping feature vector, the second mapping feature vector, the third mapping feature vector, and the global meaning... Figure 1 The consistency representation scalar is input into the contradiction coding layer of the multimodal intent contradiction detector and subjected to nonlinear fusion processing to generate the user's initial intent contradiction representation vector in the mixed reality space.

[0103] The three mapped feature vectors M_hand, M_speech, and M_touch, along with the global consistency scalar G_consistency, are concatenated to form a fused feature vector F_fused with a dimension of 3D+1. F_fused is then input into the conflict encoding layer. This conflict encoding layer is a multi-layer fully connected network, for example, consisting of two hidden layers, each with 256 neurons and the ReLU activation function. The output of the last hidden layer is then passed through a linear layer, mapping to a K-dimensional output vector (e.g., K=64), denoted as V_conflict_raw. This V_conflict_raw is the initial intent conflict representation vector, where each dimension encodes different types or dimensions of intent conflict information through non-linear learning by the network.

[0104] Step S158: Perform sparsification encoding on the initial intention conflict representation vector, retain the dimension components in the initial intention conflict representation vector that exceed the preset sparsification threshold, and set the dimension components that are below the preset sparsification threshold to zero, to obtain the user's intention conflict representation vector in the mixed reality space. The intention conflict representation vector is used to indicate the specific dimensional direction and conflict intensity level of the conflict between different modal intentions.

[0105] For each dimension component v_i of the initial intention conflict representation vector V_conflict_raw, a soft thresholding function is applied. A sparsity threshold λ is set. If |v_i| < λ, the dimension component is set to 0; if |v_i| >= λ, the original value is retained or it is shrunk, for example, v_i' = sign(v_i) * (|v_i| - λ). This operation is called soft thresholding, which makes the vector sparse, retaining only those components with significant conflict intensity. The processed vector V_conflict_sparse is the final output intention conflict representation vector. The non-zero elements of this vector indicate the specific dimensional direction of the intention conflict, while the magnitude of the non-zero element represents the level of conflict intensity in that dimension. For example, a value of 0.8 for the 5th dimension of V_conflict_sparse might represent the spatial directional intention conflict described in the example of "the hand points to the first floor button, but the voice says the second floor."

[0106] Step S160: Perform rule matching processing on the preset interaction mode switching rule base according to the intent contradiction representation vector, generate an interaction mode smooth switching instruction corresponding to the intent contradiction representation vector, and send the interaction mode smooth switching instruction to the mixed reality rendering engine to trigger the interaction paradigm transformation operation of the mixed reality space.

[0107] This step aims to dynamically adjust the interaction mode to resolve the conflict or adapt to the user's state based on the detected intent conflicts. The intent conflict representation vector V_conflict_sparse obtained in step S158 is input into the interaction mode switching rule base for querying and matching.

[0108] Step S161: Perform contradiction type classification processing on the intention contradiction representation vector, identify the contradiction dominant mode corresponding to the intention contradiction representation vector, and obtain the contradiction dominant mode identifier corresponding to the intention contradiction representation vector. The contradiction dominant mode identifier is used to indicate the main perceptual mode source that leads to the intention contradiction.

[0109] The sparse intent conflict representation vector V_conflict_sparse is input into a pre-trained conflict type classifier. This conflict type classifier can be a simple rule-based method or a lightweight neural network. For example, the classifier can examine the dimensional range of the non-zero elements in V_conflict_sparse. If the non-zero elements are predominantly concentrated in dimensions related to hand movements (e.g., vector indices 1 to 20), the dominant conflict mode is determined to be "limb movement"; if concentrated in dimensions related to speech (indices 21 to 40), the dominant mode is "silent speech"; and if concentrated in dimensions related to tactile expectation (indices 41 to 60), the dominant mode is "tactile expectation". The classifier outputs a one-hot encoded conflict dominant mode identifier ID_dominant, for example, [1, 0, 0] represents limb movement dominance.

[0110] Step S162: Query the modality priority mapping table in the preset interaction mode switching rule base according to the contradiction-dominant modality identifier, and obtain the target interaction mode switching strategy code corresponding to the contradiction-dominant modality identifier.

[0111] The system pre-stores a modality priority mapping table, which defines which interaction mode the system should prioritize switching to under different conflict-dominant modalities. For example, the mapping table might define: when the conflict is dominated by "silent speech," it indicates that the user's voice commands may be unclear or misunderstood, and the system should switch to an interaction mode dominated by "hand operation" (policy code P_hand_priority); when the conflict is dominated by "limb movement," it indicates that the user's gestures may be unstable or inconsistent with their intentions, and the system should switch to an interaction mode dominated by "voice commands" (policy code P_speech_priority); when the conflict is dominated by "tactile expectation," it indicates that the user's expectation of tactile feedback does not match the system's current provision, and the system should switch to a mode dominated by "adaptive tactile feedback" (policy code P_tactile_priority). Based on the ID_dominant obtained in step S161, the corresponding target interaction mode switching policy code P_target is retrieved from the mapping table.

[0112] Step S163: Parse the conflict intensity level parameter in the intent conflict representation vector, and calculate the interaction mode switching speed factor corresponding to the target interaction mode switching strategy encoding based on the conflict intensity level parameter. The interaction mode switching speed factor is used to control the speed of the interaction mode switching process.

[0113] The conflict intensity level parameter is extracted from the intent conflict representation vector V_conflict_sparse. This can be achieved by calculating the sum of squares of all non-zero elements or taking the maximum value, resulting in a scalar C_strength. The interaction mode switching speed factor V_switch is designed as a function related to conflict intensity. For example, a function V_switch = f(C_strength) = V_base + α * C_strength can be defined, where V_base is the base switching speed and α is an adjustment coefficient. The stronger the conflict, the faster the switching speed to quickly alleviate the conflict; the weaker the conflict, the smoother the switching to avoid mode oscillation. Depending on P_target, a policy-related coefficient may also need to be multiplied. Finally, a scalar V_switch is obtained.

[0114] Step S164: Obtain the set of configuration parameters for the source interaction mode currently running in the mixed reality space. The set of configuration parameters for the source interaction mode includes visual feedback rendering parameters, auditory feedback rendering parameters, and tactile feedback rendering parameters in the source interaction mode.

[0115] From the runtime status of the mixed reality rendering engine, the set of configuration parameters Params_src for the currently used source interaction mode is read. For example, for visual feedback rendering parameters, these might include the highlight color of the virtual button (e.g., blue), transparency (e.g., 0.8), and whether it has an outline. For auditory feedback rendering parameters, these might include the volume (0.5) and pitch (440 Hz) of the "click" sound played when the button is pressed. For haptic feedback rendering parameters, these might include the amplitude (0.3), frequency (100 Hz), and duration (50 milliseconds) of the vibration motor. These parameters are encapsulated into a structured data set.

[0116] Step S165: Extract the target interaction mode configuration parameter template corresponding to the target interaction mode switching strategy code from the preset interaction mode switching rule library. The target interaction mode configuration parameter template includes preset values ​​for visual feedback rendering parameters, auditory feedback rendering parameters, and tactile feedback rendering parameters under the target interaction mode.

[0117] Based on P_target determined in step S162, the corresponding target interaction mode configuration parameter template Params_target is extracted from the rule base. For example, if P_target is P_hand_priority (hand operation priority mode), its visual parameters may be changed to a green highlight color, an opacity of 1.0, and a gesture guide cursor may be displayed; auditory feedback may be disabled or the volume reduced to 0.1 to avoid interference; haptic feedback parameters may be changed to an amplitude of 0.8 to emphasize the tactile sensation of hand operations.

[0118] Step S166: Calculate the parameter difference vector between the source interaction mode configuration parameter set and the target interaction mode configuration parameter template, and determine the list of interaction parameters to be transitioned that need to be smoothed based on the parameter difference vector.

[0119] For each parameter requiring transition (such as visual color, tactile amplitude, etc.), calculate the difference between its source value and target value. For example, if the tactile amplitude parameter changes from 0.3 to 0.8, the difference ΔA = 0.5. Record all parameters with non-zero differences and their difference values ​​Δ, forming a parameter difference vector Diff = {(param_id, Δ)}. All parameters appearing in Diff constitute the list of parameters to be transitioned, L_trans.

[0120] Step S167: Generate a corresponding parameter interpolation curve function for each interaction parameter in the list of interaction parameters to be transitioned according to the interaction mode switching speed factor. The parameter interpolation curve function is used to describe the trajectory of change from the source interaction mode parameter value to the target interaction mode parameter value within the switching time window.

[0121] Based on the switching speed factor V_switch, a switching time window T_switch can be determined (e.g., T_switch = 1 / V_switch seconds). For each parameter in L_trans, an interpolation curve function is generated. For example, linear interpolation can be used: P_value(t) = P_src + (P_target - P_src) * (t / T_switch), where t is the time elapsed since the start of the switch, ranging from 0 to T_switch. A smoother S-curve interpolation can also be used. Thus, each parameter corresponds to a function with time as the independent variable.

[0122] Step S168: Bind the parameter interpolation curve function to the current system clock to generate a dynamic update instruction sequence of interactive mode parameters associated with the time axis.

[0123] The interpolation curve function for each parameter generated in step S167 is bound to the timestamp T_start at the moment of system startup switching. This forms a dynamic update instruction whose content is "starting from T_start, within the future T_switch time, the value of parameter param_id should follow the curve function P_value(t)". ​​Organizing the above instructions for different parameters in chronological order forms an interactive mode parameter dynamic update instruction sequence Update_cmd_seq.

[0124] Step S169: Associate and encapsulate the dynamic update instruction sequence of the interaction mode parameters with the contradictory dominant modality identifier to generate the instruction header of the smooth switching instruction of the interaction mode.

[0125] The Update_cmd_seq generated in step S168 and the contradictory dominant modality identifier ID_dominant obtained in step S161 are encapsulated to form a command header. This header information can be used by the rendering engine to understand the reason for the switch and the modality of focus when performing a switch.

[0126] Step S1610: Combine the instruction header with the parameter identifier in the list of parameters to be transitioned and encode them to generate a complete smooth switching instruction for the interaction mode, which includes the switching target mode identifier, the switching speed factor and the parameter dynamic update function.

[0127] The instruction header from step S169, the switching speed factor V_switch from step S163, and the detailed parameter dynamic update functions (or pointers to these functions) generated in step S168 are finally combined and encoded to generate a complete interactive mode smooth switching instruction C_switch that can be parsed and executed by the mixed reality rendering engine. This interactive mode smooth switching instruction contains all the information for the switch.

[0128] Step S170: Send the smooth switching command of the interaction mode to the mixed reality rendering engine to trigger the interaction paradigm shift operation of the mixed reality space.

[0129] This step is responsible for executing instructions to achieve a seamless and smooth transition between interactive modes at the rendering level. The interactive mode smooth transition instruction C_switch generated in step S1610 is sent to the mixed reality rendering engine.

[0130] Step S171: Parse the switching target modality identifier in the smooth switching instruction of the interaction mode, and determine the target rendering pipeline unit to be activated in the mixed reality rendering engine according to the switching target modality identifier. The target rendering pipeline unit includes at least one of the visual rendering pipeline branch, the auditory rendering pipeline branch and the haptic rendering pipeline branch.

[0131] Upon receiving the instruction C_switch, the mixed reality rendering engine first parses its header to obtain the target modality identifier ID_dominant. Internally, the engine maintains multiple parallel rendering pipelines, each responsible for generating visual, auditory, and haptic feedback. Based on ID_dominant, the engine determines the target rendering pipeline unit that needs to be activated or prioritized. For example, if ID_dominant indicates "hand operation priority," then the target rendering pipeline units to be activated are the haptic rendering pipeline branch and the visual rendering pipeline branch that enhances visual hand feedback.

[0132] Step S172: Extract the switching speed factor from the smooth switching instruction of the interactive mode, and input the switching speed factor into the frame rate control module of the mixed reality rendering engine for dynamic adjustment of the rendering frame rate, so as to obtain the target rendering frame rate parameter that matches the switching speed factor.

[0133] The engine extracts the switching speed factor V_switch from C_switch and inputs it into the frame rate control module. This module can dynamically adjust the rendering frame rate based on the switching speed. For example, if the switching speed is fast (V_switch is large), to ensure a smooth switching process, the frame rate control module may temporarily increase the target rendering frame rate from 60 frames per second to 90 frames per second, resulting in a target rendering frame rate parameter FPS_target.

[0134] Step S173: Reorder the rendering task scheduling queue of the mixed reality rendering engine according to the target rendering frame rate parameter, and increase the execution priority of the target rendering pipeline unit to be activated in the rendering task scheduling queue.

[0135] The rendering task scheduling queue inside the engine manages the execution order of all rendering tasks. To achieve smooth switching, the scheduler reorders the queue based on FPS_target and ID_dominant. The rendering tasks corresponding to the target rendering pipeline units to be activated as determined in step S171 (such as force feedback calculation tasks for haptic rendering and outlining tasks for hand visual enhancement) are given higher priority in the queue to ensure that they can obtain more computing resources and be processed first.

[0136] Step S174: Obtain the parameter dynamic update function in the smooth switching instruction of the interaction mode, and inject the parameter dynamic update function into the shader program of the target rendering pipeline unit to be activated for real-time compilation processing to generate an updated shader program containing parameter dynamic update logic.

[0137] The engine parses the C_switch statement and extracts the dynamic update functions for each rendering parameter. Then, these functions are dynamically injected into the shader code of the corresponding rendering pipeline. For example, the interpolation function for visual color is injected into the pixel shader, and the interpolation function for haptic amplitude is injected into the force feedback calculation module for haptic rendering. After injection, the affected shader programs are recompiled in real-time to generate updated shader programs containing the logic for dynamic parameter updates.

[0138] Step S175: Run the updated shader program to perform pixel-by-pixel parameter adjustment processing on the current rendering frame of the mixed reality space, and gradually transition the source interaction mode parameter value of the current rendering frame to the target interaction mode parameter value within the switching time window.

[0139] Starting from the switching start time T_start, the engine begins executing the updated shader program. During the rendering of each frame, the shader program calls the parameter dynamic update function based on the built-in current time t and the switching time window T_switch to calculate the instantaneous values ​​of each parameter for the current frame. For example, at 0.1 seconds into the switching, the haptic amplitude is calculated as P_value(0.1) = 0.3 + 0.5 * (0.1 / 1.0) = 0.35. In this way, the visual, auditory, and haptic feedback of each frame smoothly changes towards the target value, achieving a seamless transition.

[0140] Step S176: Monitor the user's physiological feedback signal for each frame of rendered image within the switching time window. The user's physiological feedback signal includes user pupil diameter change data and user skin conductance response data. Normalize the user pupil diameter change data and the user skin conductance response data to obtain normalized pupil change index and normalized skin conductance response index. Calculate the cognitive load change curve of the user during the interaction paradigm shift based on the weighted sum of the normalized pupil change index and the normalized skin conductance response index, and extract the peak occurrence time point of the cognitive load change curve.

[0141] During the switching process, the user's physiological feedback signals are collected in real time through the eye-tracking sensor and skin conductance sensor built into the headset. Pupil diameter data (P_dilate) and skin conductance response data (GSR) are acquired for each frame. Within a sliding time window (e.g., 10 frames), these two sequences are Z-score normalized to obtain the normalized pupil diameter change index (P_norm) and the normalized skin conductance response index (G_norm). Then, their weighted sum is calculated, for example, the cognitive load index C_load = w1 * P_norm + w2 * G_norm. The C_load values ​​for multiple consecutive frames are plotted as a cognitive load change curve. The maximum value of this curve within the switching time window T_switch is detected, and the time point at which it occurs is recorded as T_load_peak.

[0142] Step S177: Compare the peak occurrence time with the preset parameter transition completion time in the parameter dynamic update function to obtain the execution lead-lag deviation of the interaction mode smooth switching instruction.

[0143] The parameter dynamic update function presets the parameter transition completion time point as T_start + T_switch. The cognitive load peak occurrence time point T_load_peak obtained in step S176 is compared with this. If T_load_peak is significantly earlier than the completion time point, it indicates that the user felt confused or uncomfortable before the switch was completed, and the switch may have been too slow; if it is significantly later, it indicates that the switch process may have been too fast, and the user's attention has lagged. The lead-lag deviation ΔT = T_load_peak - (T_start + T_switch).

[0144] Step S178: Fine-tune the switching speed factor according to the lead-lag deviation to generate a corrected switching speed factor, and use the corrected switching speed factor to update the transition rate parameter of the parameter dynamic update function.

[0145] The switching speed factor V_switch is fine-tuned based on the lead-lag deviation ΔT. For example, if ΔT is positive (peak lag), the switching is too fast, and the speed factor should be decreased: V_switch_new = V_switch - β * ΔT; if ΔT is negative (peak lead), the switching is too slow, and the speed factor should be increased: V_switch_new = V_switch + β * |ΔT|. This corrected speed factor is then fed back into the parameter dynamic update function to update its internal transition rate parameter, which is used to guide subsequent possible switching or to make online adjustments to the current switching.

[0146] Step S179: Lock the target interaction mode configuration parameter set after the parameter transition is completed as the current working configuration parameters of the mixed reality rendering engine, and complete the interaction paradigm transformation operation of the mixed reality space.

[0147] Once the system time exceeds T_start + T_switch, the parameter transition is complete. At this point, the mixed reality rendering engine locks all parameter values ​​in the target interaction mode configuration parameter template Params_target and sets them as the current working configuration parameters. The original source interaction mode configuration parameters are replaced. Thus, a complete, smooth, adaptive interaction paradigm shift triggered by a conflicting user intent is completed.

[0148] Step S180: Before sending the smooth switching instruction of the interaction mode to the mixed reality rendering engine, the method further includes: inputting the intent contradiction representation vector into the preloading timing prediction network to perform switching instruction preloading time point prediction processing to obtain the interaction mode smooth switching instruction preloading time point corresponding to the intent contradiction representation vector, and performing cache queue preloading processing on the interaction mode smooth switching instruction according to the interaction mode smooth switching instruction preloading time point.

[0149] To further reduce handover latency, the system can predict when the handover will occur before officially sending the handover command and load the command into the cache in advance.

[0150] Step S181: Obtain the historical sample set of intent conflict representation vectors stored internally in the preloading timing prediction network. The historical sample set of intent conflict representation vectors includes intent conflict representation vectors generated in multiple historical interaction periods and the actual switching time point of the interaction mode switching after each intent conflict representation vector is generated.

[0151] The preloading timing prediction network maintains an experience replay pool internally, which stores a large number of historical data samples. Each sample is a tuple containing an intent conflict representation vector V_conflict_hist generated at a certain time in the past, and the time difference Δt_real between the actual time of the next interaction mode switch after the vector was generated and the current time (i.e., the actual switch time).

[0152] Step S182: Input the intention contradiction representation vector generated at the current time into the feature comparison layer of the preload timing prediction network, and perform feature similarity matching processing with each historical intention contradiction representation vector in the historical sample set of the intention contradiction representation vector to obtain the feature similarity coefficient between the intention contradiction representation vector generated at the current time and each historical intention contradiction representation vector.

[0153] The intent conflict representation vector V_conflict_current generated at the current moment is input into the network. The feature comparison layer calculates the cosine similarity between V_conflict_current and each V_conflict_hist in the historical sample set, resulting in a similarity coefficient sequence S_sim={s_1, s_2, ..., s_N}.

[0154] Step S183: Select multiple similar historical intention contradiction representation vectors from the historical sample set of the intention contradiction representation vectors, whose feature similarity coefficient with the intention contradiction representation vector generated at the current time exceeds a preset similarity threshold, and form a set of similar historical intention contradiction representation vectors.

[0155] Set a similarity threshold θ_sim (e.g., 0.85). Filter out all historical intent conflict representation vectors corresponding to s_i > θ_sim from S_sim, forming a set of similar historical intent conflict representation vectors V_conflict_sim_set.

[0156] Step S184: Extract the historical real switching time point corresponding to each similar historical intention contradiction representation vector in the similar historical intention contradiction representation vector set from the storage unit of the preloading timing prediction network, and obtain the set of historical real switching time points corresponding to the similar historical intention contradiction representation vector set.

[0157] For each vector in V_conflict_sim_set, extract its corresponding historical real switching time point Δt_real from the storage unit to form a time point set T_real_set.

[0158] Step S185: Sort all historical real switching time points in the set of historical real switching time points, and identify the time point interval with the highest frequency in the set of historical real switching time points as candidate preloading time point intervals.

[0159] Perform statistical analysis on all time points in T_real_set. For example, kernel density estimation can be used to estimate the probability density distribution curve of each time point. Then, find the time interval where the peak of the curve is located, denoted as the candidate preloading time point interval I_candidate=[t_low, t_high]. For example, in most similar historical scenarios, the switching occurs between 1.5 and 1.8 seconds after the current time.

[0160] Step S186: Based on the distribution density of the intent contradiction representation vector generated at the current time in the historical sample set of intent contradiction representation vectors, perform offset correction processing on the candidate preloading time point interval to obtain the interaction mode smooth switching instruction preloading time point corresponding to the intent contradiction representation vector generated at the current time.

[0161] Calculate the local density ρ_local of the current V_conflict_current in the historical sample space. If ρ_local is high (indicating the current scenario is very common), the preloading time point can be directly taken as the center value of I_candidate, T_pre = (t_low + t_high) / 2. If ρ_local is low (indicating the current scenario is relatively novel), it means the reliability of historical experience has decreased, and a conservative correction is needed for the preloading time point, such as shifting to a larger value to reserve more reaction time, T_pre = t_high + δ. Finally, a definite preloading time point T_pre is obtained.

[0162] Step S187: Calculate the time interval between the current moment and the preloading time of the interaction mode smooth switching instruction based on the preloading time of the interaction mode smooth switching instruction, and obtain the preset waiting time of the interaction mode smooth switching instruction in the cache queue.

[0163] Calculate the waiting time Δt_wait = T_pre - T_current, where T_current is the current time.

[0164] Step S188: Encapsulate the smooth switching instruction of the interaction mode into a preload instruction unit with a preset waiting time label, and send the preload instruction unit to the instruction prefetch cache queue of the mixed reality rendering engine for queue writing operation.

[0165] The complete interactive mode smooth switching instruction C_switch generated in step S1610, together with the preset wait duration label Δt_wait calculated in step S187, is encapsulated into a preload instruction unit U_preload. Then, U_preload is sent to a dedicated instruction prefetch cache queue in the mixed reality rendering engine.

[0166] Step S189: In the instruction prefetch cache queue, the preloaded instruction units are sorted according to the preset wait duration tag to ensure that the storage location of the preloaded instruction units in the instruction prefetch cache queue matches the wake-up time order indicated by the preset wait duration tag.

[0167] The instruction prefetch cache queue is sorted according to the Δt_wait label of each preloaded instruction unit. The unit with the smallest Δt_wait is placed at the front of the queue, waiting to be woken up first.

[0168] Step S1810: When the system clock reaches the preloading time point of the interaction mode smooth switching instruction, wake up the preloaded instruction unit from the instruction prefetch cache queue, and convert the woken preloaded instruction unit into an interaction mode smooth switching instruction to be executed for the mixed reality rendering engine to obtain and execute.

[0169] The system clock runs continuously. When the time reaches T_pre, the corresponding preload instruction unit U_preload is awakened from the cache queue. Its internal C_switch is retrieved, converted into a high-priority instruction to be executed, and inserted into the rendering engine's main instruction stream, awaiting execution. In this way, when the system actually needs to perform a switch, the instruction is already ready, greatly reducing response latency.

[0170] For example, in step S190: after sending the smooth switching instruction of the interaction mode to the mixed reality rendering engine to trigger the interaction paradigm shift operation, the method further includes: collecting the post-switching multimodal original perception information stream within a preset time period after the interaction paradigm shift, and performing implicit feedback parsing processing on the post-switching multimodal original perception information stream to obtain the user's acceptance feedback representation vector after the interaction paradigm shift, and then inputting the acceptance feedback representation vector into the multimodal intent contradiction detector for adaptive optimization of detection parameters.

[0171] In order to form a closed-loop learning system, after the switch is completed, the system will evaluate the user's acceptance of the new interaction mode and optimize the parameters of the contradiction detector itself accordingly.

[0172] Step S191: Within the preset acquisition time window after the completion of the interaction paradigm conversion operation, continuously acquire the point cloud sequence of the dynamic motion trajectory of the user's hand skeleton through the visual sensor array, acquire the pre-vibration signal sequence of the user's throat vocalization through the bone conduction sensor, and acquire the simulated pressing deformation distribution sequence of the user's virtual object through the flexible tactile sensor, which together constitute the original perceptual information flow of the post-switching multimodal sensor.

[0173] After the interaction paradigm shift is completed in step S179, a preset acquisition time window (e.g., lasting 5 seconds) is immediately opened. Within this window, the operation of step S110 is repeated to continuously acquire the user's multimodal perception information in the new interaction mode, denoted as P_hand_post_seq, V_larynx_post_seq, and D_tactile_post_seq, collectively referred to as the post-switching multimodal raw perception information stream. Specifically, P_hand_post_seq represents the point cloud sequence of the dynamic motion trajectory of the user's hand skeleton after the switch, V_larynx_post_seq represents the pre-vibration signal sequence of the user's throat vocalization after the switch, and D_tactile_post_seq represents the distribution sequence of the simulated pressing deformation of the virtual object by the user after the switch.

[0174] Step S192: Perform motion hesitation analysis on the point cloud sequence of the dynamic motion trajectory of the user's hand skeleton after switching, and extract the time period length and frequency of the stagnation of the key points of the hand skeleton in the point cloud sequence of the dynamic motion trajectory of the user's hand skeleton after switching, as the original feature parameters of limb motion hesitation.

[0175] Analyze P_hand_post_seq to calculate the instantaneous velocity of each hand keypoint. When the velocity of a keypoint remains below a very low threshold (e.g., 1 mm / s) for multiple consecutive frames, that point is considered to be in a stagnant state. Calculate the total duration of stagnant states for all keypoints and the number of stagnant events to obtain the original feature parameters of limb movement hesitation, such as the total stagnant duration T_hesitate and the stagnant frequency F_hesitate.

[0176] Step S193: Perform silent speech fluency analysis on the pre-vibration signal sequence of the laryngeal vocalization of the user after switching, and extract the average time interval and variance of the time interval between adjacent silent speech units in the pre-vibration signal sequence of the laryngeal vocalization of the user after switching as the original feature parameters of speech expression fluency.

[0177] Applying steps S133 and S134 to V_larynx_post_seq, silent speech units are segmented. The duration of the silence interval between adjacent units is calculated. Then, the mean μ_interval and variance σ_interval of these intervals are calculated. A high mean and variance generally indicate that the user is hesitant and lacks fluency in speech. These two statistical values ​​are the original feature parameters of speech fluency.

[0178] Step S194: Perform pressure decisiveness analysis on the simulated pressing deformation distribution sequence of the virtual object by the user who switched later, and extract the average time length from the start of pressing to the deformation reaching a stable state in the simulated pressing deformation distribution sequence of the virtual object by the user who switched later, as the original feature parameter of the pressure operation decisiveness.

[0179] Analyze D_tactile_post_seq. For each press event, detect the time elapsed from the start of deformation (displacement exceeding the perception threshold) to the point where deformation stabilizes (displacement rate falls below the stabilization threshold). Calculate the average of this time length, T_decisive, for all press events. A smaller T_decisive indicates a decisive and confident press; conversely, a larger T_decisive indicates hesitation. T_decisive is the original feature parameter representing the decisiveness of the press operation.

[0180] Step S195: Input the original feature parameters of the limb movement hesitation, the original feature parameters of the speech expression fluency, and the original feature parameters of the pressing operation decisiveness into the pre-trained implicit encoder of user acceptance for feature embedding processing to obtain the user's acceptance feedback representation vector after the interaction paradigm transformation.

[0181] The multiple raw feature parameters obtained in steps S192 to S194 (e.g., [T_hesitate, F_hesitate, μ_interval, σ_interval, T_decisive]) are concatenated into a one-dimensional feature vector. This vector is then input into a pre-trained implicit user acceptance encoder. This encoder can be an autoencoder, whose encoding part maps the low-level statistical features of the input to a low-dimensional, semantically rich embedding space, outputting an acceptance feedback vector F_accept. This acceptance feedback vector quantifies the user's implicit, nonverbal acceptance of the new pattern.

[0182] Step S196: Obtain the contradiction detection sensitivity threshold parameter and contradiction detection response time window parameter currently used by the multimodal intent contradiction detector, as the detector operating parameters to be optimized.

[0183] From the configuration of the multimodal intent conflict detector, read its two currently used key operating parameters: the conflict detection sensitivity threshold θ_sens (e.g., used to determine intent). Figure 1 The threshold for whether the consistency coefficient is too low) and the contradiction detection response time window T_window (e.g., the length of the time window used to smooth detection results and avoid instantaneous false alarms).

[0184] Step S197: Input the acceptance feedback representation vector into the parameter adaptive tuning interface of the multimodal intent contradiction detector, and compare and analyze it with the historical acceptance feedback representation vector sample set stored inside the multimodal intent contradiction detector to determine the distribution position of the acceptance feedback representation vector in the historical acceptance feedback representation vector sample set.

[0185] The parameter adaptive tuning interface of the multimodal intent conflict detector, F_accept, is input into the interface. This interface internally stores a large set of acceptance feedback representation vector samples, Set_accept_hist, recorded from historical interactions. The interface calculates the position of F_accept in the high-dimensional space formed by Set_accept_hist, for example, by calculating its distance to the centers of all historical samples, or by calculating its landing region in a reduced-dimensional visualization (such as a t-distributed random neighborhood embedding graph).

[0186] Step S198: Based on the distribution position of the acceptance feedback representation vector in the historical acceptance feedback representation vector sample set, query the detector parameter adjustment direction indicator corresponding to the distribution position from the preset parameter adjustment mapping rule base.

[0187] The pre-defined parameter adjustment mapping rule base defines parameter adjustment strategies corresponding to different user acceptance distribution areas. For example, the rule base may define: if F_accept falls in the "high acceptance" area (i.e., the user behaves smoothly and decisively), the detector parameter adjustment direction indicator is "maintain or slightly reduce sensitivity to save computing resources"; if it falls in the "low acceptance" area (i.e., the user is hesitant and not smooth), the indicator is "increase sensitivity and shorten response time to detect contradictions and switch again more quickly".

[0188] Step S199: If the detector parameter adjustment direction indicator output by the parameter adjustment mapping rule base indicates that the contradiction detection sensitivity needs to be increased, then a decrease adjustment operation is performed on the contradiction detection sensitivity threshold parameter to obtain an updated contradiction detection sensitivity threshold parameter; if the detector parameter adjustment direction indicator output by the parameter adjustment mapping rule base indicates that the contradiction detection response time needs to be extended, then an extension adjustment operation is performed on the contradiction detection response time window parameter to obtain an updated contradiction detection response time window parameter.

[0189] Based on the queried indicator, the parameters to be tuned obtained in step S196 are modified. If the indicator is "increase sensitivity", the current threshold θ_sens is multiplied by a factor less than 1 (e.g., 0.95) to obtain a new, lower sensitivity threshold θ_sens_new, making the detector easier to trigger. If the indicator is "extend response time", the current time window T_window is multiplied by a factor greater than 1 (e.g., 1.1) to obtain a new, longer time window T_window_new, used to smooth the detection results for a longer period and avoid frequent switching.

[0190] Step S1910: Write the updated contradiction detection sensitivity threshold parameter and the updated contradiction detection response time window parameter back to the parameter storage unit of the multimodal intent contradiction detector, replacing the original contradiction detection sensitivity threshold parameter and contradiction detection response time window parameter, for cross-modal intent conflict detection processing of the subsequently acquired interaction penetration depth sequence, internal speech command encoding sequence and virtual haptic feedback expectation intensity sequence.

[0191] Finally, the updated parameters θ_sens_new and T_window_new are written back to the parameter storage unit of the multimodal intent conflict detector, overwriting the old values. From then on, the detector will use these new, adaptively adjusted parameters based on user feedback to detect intent conflicts in subsequent interaction data, forming a continuously self-optimizing closed-loop system.

[0192] Step S200: Before inputting the interaction penetration depth sequence, the internal speech instruction encoding sequence, and the virtual haptic feedback expectation intensity sequence into the multimodal intent conflict detector for cross-modal intent conflict detection processing, the method further includes: performing cross-modal semantic consistency pre-analysis processing on the interaction penetration depth sequence, the internal speech instruction encoding sequence, and the virtual haptic feedback expectation intensity sequence to obtain cross-modal semantic consistency pre-analysis results, and dynamically adjusting the conflict detection triggering timing of the multimodal intent conflict detector according to the cross-modal semantic consistency pre-analysis results.

[0193] To further improve detection efficiency, a lightweight pre-analysis can be performed before conducting complex contradiction detection to determine whether the detection needs to be started immediately or can be delayed.

[0194] For example, step S210: Obtain a preset cross-modal semantic alignment dictionary, which records the first semantic mapping relationship between limb movement interaction mode and internal speech command, the second semantic mapping relationship between limb movement interaction mode and virtual haptic feedback expectation, and the third semantic mapping relationship between internal speech command and virtual haptic feedback expectation.

[0195] The system preloads a cross-modal semantic alignment dictionary. This cross-modal semantic alignment dictionary is a knowledge base. For example, its first entry (first semantic mapping relation) may define that "limb movement pointing to the second floor button" and the internal speech instruction "second floor" are semantically mapped to each other; the second entry (second semantic mapping relation) defines that "limb movement pointing to the second floor button" and the tactile expectation "feeling a crisp click at the location of the second floor button" are mapped to each other; the third entry (third semantic mapping relation) defines that the internal speech instruction "second floor" and the tactile expectation "feeling a crisp click" are mapped to each other.

[0196] Step S220: Perform semantic tagging processing on the interaction penetration depth sequence to identify the limb movement interaction mode category label corresponding to each frame of interaction penetration depth data in the interaction penetration depth sequence, and obtain the limb movement interaction mode label sequence.

[0197] Analyze P_enc_seq. The grasping operation pattern recognition results contained within can be utilized (from step S129). For example, if M_rec is "quick click", it is mapped to a semantic label L_hand = "press button". Combined with hand spatial location information, it can be further refined, for example, L_hand = "point and click the second-floor button". This yields the limb movement interaction pattern label sequence L_hand_seq.

[0198] Step S230: Perform instruction semantic tagging processing on the internal speech instruction encoding sequence, identify the speech instruction intent category tag corresponding to each internal speech instruction encoding unit in the internal speech instruction encoding sequence, and obtain a speech instruction intent tag sequence.

[0199] Analyze Word_seq and map its word-level encoding units (such as "second floor") to a unified semantic label L_speech = "select_second floor". This yields the speech instruction intent label sequence L_speech_seq.

[0200] Step S240: Perform tactile expectation semantic tagging processing on the virtual tactile feedback expectation intensity sequence, identify the tactile feedback expectation mode category label corresponding to each frame of virtual tactile feedback expectation intensity data in the virtual tactile feedback expectation intensity sequence, and obtain a tactile feedback expectation mode label sequence.

[0201] The E_refine_seq is analyzed. Based on its intensity value and spatial distribution, it can be quantized into different semantic labels. For example, when the desired intensity is high and concentrated in the fingertip area, it can be mapped to L_touch = "desired strong click sensation"; when the intensity is low and dispersed, it can be mapped to L_touch = "desired gentle touch". This yields the haptic feedback desired pattern label sequence L_touch_seq.

[0202] Step S250: Perform frame-by-frame semantic matching degree verification on the limb movement interaction mode label sequence and the speech command intent label sequence according to the first semantic mapping relationship entry to obtain the first modality semantic matching degree sequence between the limb movement modality and the speech command modality.

[0203] The L_hand_seq and L_speech_seq are aligned frame by frame. For each frame t, the standard semantics corresponding to L_hand_t are searched in the cross-modal semantic alignment dictionary. If L_speech_t matches this standard semantics, the matching degree M_hand_speech_t = 1; otherwise, it is 0. This yields a binary sequence of first-modal semantic matching degrees, M_hand_speech_seq.

[0204] Step S260: Perform frame-by-frame semantic matching degree verification on the limb movement interaction mode label sequence and the tactile feedback expectation mode label sequence according to the second semantic mapping relationship entry to obtain the second modality semantic matching degree sequence between the limb movement modality and the tactile feedback expectation mode.

[0205] Similarly, L_hand_seq and L_touch_seq are aligned frame by frame. The expected haptic pattern mapped by L_hand_t is searched in the dictionary. If L_touch_t matches, M_hand_touch_t = 1; otherwise, it is 0. This yields the second modality semantic matching degree sequence M_hand_touch_seq.

[0206] Step S270: Perform frame-by-frame semantic matching degree verification on the speech instruction intent label sequence and the tactile feedback expectation mode label sequence according to the third semantic mapping relationship entry to obtain the third modality semantic matching degree sequence between the speech instruction modality and the tactile feedback expectation mode.

[0207] Perform frame-by-frame alignment of L_speech_seq and L_touch_seq. Search the dictionary for the expected haptic pattern mapped to L_speech_t. If L_touch_t matches, then M_speech_touch_t=1; otherwise, it is 0. This yields the third-modal semantic matching degree sequence M_speech_touch_seq.

[0208] Step S280: Mark the frame positions in the first modality semantic matching degree sequence, the second modality semantic matching degree sequence and the third modality semantic matching degree sequence whose values ​​are lower than the preset semantic matching degree threshold as potential contradictory frame positions, and obtain a set of potential contradictory frame positions.

[0209] Set a semantic matching threshold; since the matching degree is either 0 or 1, the threshold can be set to 0.5. Mark the frame positions with a value of 0 in all matching degree sequences. For example, in a frame t, if M_hand_speech_t is 0, then t is added to the potential conflicting frame position set Set_conflict_candidate.

[0210] Step S290: Calculate the temporal distribution density of frame positions in the potential conflict frame position set, and calculate the conflict detection trigger urgency level of the multimodal intent conflict detector within the current time window based on the temporal distribution density.

[0211] This involves counting the number of frames in `Set_conflict_candidate` and their clustering along the timeline. For example, it can calculate the frequency of potentially conflicting frames within a sliding window of size N frames. The higher the frequency, the more likely a conflict is about to occur or has already occurred, and the higher the urgency level (U_urgency) for conflict detection. U_urgency can be quantified into three levels: 1, 2, and 3.

[0212] Step S2100: Query the preset trigger timing adjustment mapping table according to the conflict detection trigger urgency level, and obtain the trigger timing advance parameter corresponding to the conflict detection trigger urgency level.

[0213] The preset trigger timing adjustment mapping table defines the mapping relationship between the urgency level and the trigger timing advance Δt_advance. For example, for level 1 (low urgency), the advance is set to 0 seconds (no advance); for level 2 (medium urgency), the advance is set to 0.2 seconds; and for level 3 (high urgency), the advance is set to 0.5 seconds.

[0214] Step S2110: Based on the trigger timing advance parameter, the fixed detection trigger time interval currently used by the multimodal intent conflict detector is reduced and adjusted to obtain the dynamically adjusted conflict detection trigger time interval.

[0215] The multimodal intent conflict detector might have originally had a fixed detection trigger interval T_fixed (e.g., once every 1 second). Now, based on the urgency level, the next detection trigger time is advanced. For example, the new detection trigger interval T_dynamic = max(0, T_fixed - Δt_advance).

[0216] Step S2120: Reset the detection trigger clock of the multimodal intent conflict detector according to the dynamically adjusted conflict detection trigger time interval, so that after receiving the interaction penetration depth sequence, the internal speech instruction encoding sequence and the virtual haptic feedback expectation intensity sequence, the multimodal intent conflict detector starts cross-modal intent conflict detection processing at a shortened time interval.

[0217] Finally, the internal timer of the multimodal intent conflict detector is updated with T_dynamic. This means that if the pre-analysis finds a potential conflict with high urgency in the current interaction, the detector will initiate a complex and accurate cross-modal intent conflict detection process (i.e., step S150) after a shorter time interval (or even immediately), thereby responding more quickly to potential user intent conflicts.

[0218] Based on the same inventive concept, please refer to Figure 2 The diagram shows a schematic block diagram of the structure of an adaptive human-computer interaction system 100 based on multimodal perception provided in this application embodiment, including a communication unit 110, a machine-readable storage medium 120, and a processor 130.

[0219] In this embodiment, the machine-readable storage medium 120 can also be integrated into the processor 130 and can communicate and interact with external systems through the communication unit 110. The machine-readable storage medium 120 stores machine-executable instructions for executing the scheme of this application, and the processor 130 executes the machine-executable instructions stored in the machine-readable storage medium 120 to implement the inspection video stream processing method provided in the aforementioned method embodiments.

[0220] It should be noted that, in order to simplify the description of the present invention and thus help to understand one or more embodiments of the invention, multiple features may sometimes be grouped into one embodiment, drawing or description thereof in the foregoing description of the embodiments of the present invention.

Claims

1. An adaptive human-computer interaction method based on multimodal perception, characterized in that, The method includes: The system collects multimodal raw sensory information streams of users in a mixed reality space. These multimodal raw sensory information streams include a point cloud sequence of the dynamic motion trajectory of the user's hand skeleton obtained by a depth vision sensor, a sequence of pre-vibration signals of the user's throat vocalization obtained by a bone conduction sensor, and a sequence of simulated pressing deformation distribution of virtual objects obtained by a flexible tactile sensor. The point cloud sequence of the dynamic motion trajectory of the user's hand skeleton is processed by the penetration depth calculation of the preset set of virtual space object bounding boxes to obtain the interaction penetration depth sequence between the user's hand and the virtual space object. The user's laryngeal vocalization pre-vibration signal sequence is input into a pre-trained laryngeal vibration decoding network for silent speech decoding processing to obtain the user's internal speech command encoding sequence. The simulated pressing deformation distribution sequence of the user on the virtual object is matched with the set of physical simulation attribute parameters of the virtual object to obtain the expected intensity sequence of the virtual tactile feedback of the user on the virtual object. The interaction penetration depth sequence, the internal speech instruction encoding sequence, and the virtual haptic feedback expectation intensity sequence are input into a multimodal intent conflict detector for cross-modal intent conflict detection processing to obtain the user's intent conflict representation vector in the mixed reality space. Based on the intent contradiction representation vector, the preset interaction mode switching rule base is matched to generate an interaction mode smooth switching instruction corresponding to the intent contradiction representation vector, and the interaction mode smooth switching instruction is sent to the mixed reality rendering engine to trigger the interaction paradigm shift operation of the mixed reality space.

2. The adaptive human-computer interaction method based on multimodal perception according to claim 1, characterized in that, The step of performing penetration depth calculation on the point cloud sequence of the dynamic motion trajectory of the user's hand skeleton and the preset set of bounding boxes of virtual space objects to obtain the penetration depth sequence of the interaction between the user's hand and the virtual space objects includes: The user hand skeleton dynamic motion trajectory point cloud of each frame in the user hand skeleton dynamic motion trajectory point cloud sequence is processed by hand topology reconstruction to obtain the user hand triangular mesh surface model corresponding to each frame of user hand skeleton dynamic motion trajectory point cloud. The user hand triangular mesh surface model includes multiple triangular facet units and the shared vertex connection relationship between the triangular facet units. Extract the spatial normal vector direction data of each triangular facet unit in the user's hand triangular mesh surface model, and identify the point cloud clusters in the fingertip region, knuckle region, and palm region in the user's hand triangular mesh surface model based on the spatial normal vector direction data; The point cloud clusters of the fingertip region, the point cloud clusters of the knuckle region, and the point cloud clusters of the palm region are respectively subjected to spatial inclusion relationship detection processing with each virtual space object bounding box in the set of virtual space object bounding boxes to determine at least one target virtual space object bounding box that has contacted and interacted with the user's hand triangular mesh surface model. Obtain a fine mesh model of the surface of the virtual space object corresponding to the bounding box of the at least one target virtual space object, and perform signed distance field pre-calculation processing on the fine mesh model to generate the external spatial distance field distribution data of the virtual space object corresponding to the bounding box of the at least one target virtual space object. Based on the three-dimensional spatial coordinates of each point in the point cloud clusters of the fingertip region, the point cloud clusters of the knuckle region, and the point cloud clusters of the palm region, query the external spatial distance field distribution data of the virtual space object to obtain the set of external signed distance values ​​of each point on the user's hand triangular mesh surface model relative to the surface of the virtual space object. The absolute value of the distance values ​​less than zero in the set of external signed distance values ​​is processed by performing an absolute value operation to extract the absolute value of the penetration depth of each penetration point on the user's hand triangular mesh surface model that has penetrated the surface of the virtual space object; according to the hand region type to which each penetration point belongs, a corresponding weight coefficient is assigned to each penetration point; the absolute value of the penetration depth is weighted and averaged using the weight coefficient to obtain the comprehensive penetration depth characterization parameter of the user's hand triangular mesh surface model relative to the virtual space object; The comprehensive penetration depth characterization parameters are nonlinearly mapped to the material stiffness property parameters of the virtual space object to generate the magnitude parameters of the interaction force feedback between the user's hand and the virtual space object in the current frame. Based on the interaction force feedback magnitude parameters corresponding to the point cloud of the dynamic motion trajectory of the user's hand skeleton in multiple consecutive frames, a time series of interaction force changes between the user's hand and the virtual space object is constructed, and the force loading rate feature and force unloading rate feature in the time series of interaction force changes are extracted. The interaction force change time series, the force loading rate feature, and the force unloading rate feature are dynamically time-warped and matched with a preset virtual object grasping operation mode template library to obtain the user's current grasping operation mode recognition result for the virtual space object. The current grasping operation mode recognition result is combined with the comprehensive penetration depth characterization parameter to perform feature association encoding processing, thereby generating the interaction penetration depth sequence between the user's hand and the virtual space object. The interaction penetration depth sequence includes the penetration depth values ​​on consecutive time frames and the corresponding grasping operation mode labels.

3. The adaptive human-computer interaction method based on multimodal perception according to claim 1, characterized in that, The step of inputting the user's laryngeal vocalization pre-vibration signal sequence into a pre-trained laryngeal vibration decoding network for silent speech decoding processing to obtain the user's internal speech command encoding sequence includes: The user's laryngeal pre-vibration signal in the user's vocalization pre-vibration signal sequence is subjected to time-frequency domain transformation processing to obtain the laryngeal vibration spectrum corresponding to each segment of the user's laryngeal pre-vibration signal. The laryngeal vibration spectrum contains the distribution pattern information of the frequency component intensity changing with time. The spectral energy centroid calculation process is performed on the throat vibration spectrum to extract the resonance peak frequency trajectory features in the throat vibration spectrum. The resonance peak frequency trajectory features include the center frequency change curve of the first resonance peak, the center frequency change curve of the second resonance peak, and the center frequency change curve of the third resonance peak. The user's laryngeal pre-vibration signal sequence is processed by extracting the laryngeal vibration waveform envelope to obtain the laryngeal vibration amplitude envelope curve corresponding to the user's laryngeal pre-vibration signal. The onset time and cessation time of the user's laryngeal voice are detected based on the distribution of local maxima of the laryngeal vibration amplitude envelope curve. Based on the start time and stop time, the user's laryngeal vocalization pre-vibration signal sequence is segmented into silent speech units to obtain multiple continuous silent speech unit segments, each of which corresponds to a complete internal speech articulation cycle. The laryngeal vibration spectrogram corresponding to each silent speech unit segment is input into the feature encoding layer of the pre-trained laryngeal vibration decoding network for deep feature extraction processing to obtain the deep feature vector of laryngeal vibration corresponding to each silent speech unit segment. The deep feature vector of laryngeal vibration is input into the attention mechanism layer of the pre-trained laryngeal vibration decoding network for keyframe attention weight allocation processing, thereby generating the contribution weight distribution of different time frames in the deep feature vector of laryngeal vibration to silent speech decoding. Based on the contribution weight distribution, the deep feature vector of laryngeal vibration is weighted and aggregated to obtain the semantic coding feature vector corresponding to each silent speech unit segment. The semantic encoding feature vector is input into the classification output layer of the pre-trained laryngeal vibration decoding network for phoneme probability distribution prediction processing to obtain the internal speech phoneme sequence probability distribution corresponding to each silent speech unit segment. The probability distribution of the internal speech phoneme sequence corresponding to multiple consecutive silent speech unit segments is input into the Viterbi decoder for optimal path search processing to generate the phoneme-level coding unit in the internal speech instruction coding sequence. The phoneme-level encoding unit is matched with the standard instruction phoneme template in the preset instruction lexicon to obtain the word-level encoding unit in the internal speech instruction encoding sequence. The internal speech instruction encoding sequence is used to represent the interactive operation intention expressed by the user through silent speech.

4. The adaptive human-computer interaction method based on multimodal perception according to claim 1, characterized in that, The step of performing mechanical response matching processing on the simulated pressure deformation distribution sequence of the user on the virtual object and the physical simulation attribute parameter set of the virtual object to obtain the expected intensity sequence of the user's virtual tactile feedback on the virtual object includes: Spatial interpolation processing is performed on each frame of simulated pressing deformation distribution data in the simulated pressing deformation distribution sequence of the virtual object to generate a continuous deformation displacement field corresponding to each frame of simulated pressing deformation distribution data. The continuous deformation displacement field includes the normal displacement of each sampling point on the surface of the virtual object. Local deformation regions in the continuous deformation displacement field where the deformation displacement exceeds a preset deformation sensing threshold are extracted, and the boundary contours of the local deformation regions are subjected to curvature analysis to obtain the boundary curvature distribution characteristics and the region area expansion rate characteristics of the local deformation regions. From the boundary curvature distribution characteristics of the local deformation region, curvature feature values ​​for characterizing the pressing operation characteristics are extracted; from the surface tension attribute parameters in the physical simulation attribute parameter set of the virtual object, tension feature values ​​for characterizing the material response characteristics are extracted; after standardizing the curvature feature values ​​and the tension feature values, the matching degree is calculated to generate the surface tension response matching degree parameter generated by the user's pressing operation on the virtual object in the local deformation region; Based on the elastic modulus attribute parameters and simulated compression deformation distribution data, the theoretical area expansion rate reference value of the virtual object under the pressing operation is calculated; the area expansion rate characteristics of the local deformation region are standardized and the theoretical area expansion rate reference value are then matched to calculate the matching degree parameter of the elastic deformation response generated by the user's pressing operation on the virtual object in the local deformation region. The stress concentration region distribution map inside the virtual object is calculated based on the spatial gradient distribution of the deformation displacement in the continuous deformation displacement field, and the stress peak coordinates and stress peak intensity of each stress concentration region in the stress concentration region distribution map are extracted. Calculate the spatial distance between the stress peak coordinates and the coordinates of each vulnerable point in the internal structural vulnerability distribution map; determine the proximity between the stress concentration point and the structural vulnerability point based on the spatial distance, and generate the structural response matching degree parameter generated by the user's pressing operation on the virtual object within the virtual object; Based on the surface tension response matching parameter, the elastic deformation response matching parameter, and the structural response matching parameter, a multi-dimensional mechanical response feature vector of the user's pressing operation on the virtual object in the current frame is constructed. The multi-dimensional mechanical response feature vector of the pressing operation is input into a pre-trained tactile feedback expectation prediction network for expectation intensity regression processing to obtain the preliminary prediction value of the user's virtual tactile feedback expectation intensity to the virtual object in the current frame. Obtain the historical sequence of the user's expected virtual haptic feedback intensity in historical interaction frames, and perform temporal trend extrapolation processing on the historical sequence of the expected virtual haptic feedback intensity to generate the temporal correction factor of the user's expected virtual haptic feedback intensity in the current frame; The preliminary predicted value of the virtual haptic feedback intensity is weighted and corrected according to the time-series correction factor to obtain the accurate value of the user's expected virtual haptic feedback intensity for the virtual object in the current frame. The accurate values ​​of the expected virtual haptic feedback intensity for multiple consecutive frames are then serialized and assembled to generate the sequence of the user's expected virtual haptic feedback intensity for the virtual object.

5. The adaptive human-computer interaction method based on multimodal perception according to claim 1, characterized in that, The process of inputting the interaction penetration depth sequence, the internal speech command encoding sequence, and the virtual haptic feedback expectation intensity sequence into a multimodal intent conflict detector for cross-modal intent conflict detection processing yields the user's intent conflict representation vector in the mixed reality space, including: Temporal feature extraction processing is performed on the interaction penetration depth sequence to obtain the first intention trend feature vector of the user's hand interaction operation. The first intention trend feature vector includes operation direction consistency and operation force change stability parameters. Semantic feature extraction processing is performed on the internal speech instruction encoding sequence to obtain a second intention trend feature vector of the user's internal speech expression. The second intention trend feature vector includes parameters of instruction semantic clarity and instruction sentiment intensity. The expected feature extraction process is performed on the virtual haptic feedback expected intensity sequence to obtain the third intention trend feature vector of the user's haptic expectation expression. The third intention trend feature vector includes the expected intensity change rate and the expected spatial distribution concentration parameter. The first intention trend feature vector, the second intention trend feature vector, and the third intention trend feature vector are input into the feature alignment layer of the multimodal intention contradiction detector for cross-modal feature space mapping processing to obtain the first mapped feature vector, the second mapped feature vector, and the third mapped feature vector under a unified feature space. Calculate the cosine similarity of the feature space between the first mapped feature vector and the second mapped feature vector to obtain the first modal intent similarity; calculate the Euclidean distance of the feature space between the first mapped feature vector and the third mapped feature vector to obtain the second modal intent difference; calculate the Manhattan distance of the feature space between the second mapped feature vector and the third mapped feature vector to obtain the third modal intent difference; process the first modal intent similarity, the second modal intent difference, and the third modal intent difference through a preset normalization function, and map them to a comparable scale within the same numerical range to obtain the first modal intent consistency coefficient between limb movement intent and verbal expression intent, the second modal intent consistency coefficient between limb movement intent and tactile expectation intent, and the third modal intent consistency coefficient between verbal expression intent and tactile expectation intent; An intent consistency coefficient matrix is ​​constructed based on the first modality intent consistency coefficient, the second modality intent consistency coefficient, and the third modality intent consistency coefficient. Principal component analysis is then performed on the intent consistency coefficient matrix to extract the principal component with the largest eigenvalue as a global intent consistency representation scalar. The first mapping feature vector, the second mapping feature vector, the third mapping feature vector, and the global intent consistency representation scalar are input into the contradiction coding layer of the multimodal intent contradiction detector for nonlinear fusion processing to generate the user's initial intent contradiction representation vector in the mixed reality space. The initial intention conflict representation vector is subjected to sparse encoding processing. The dimensional components in the initial intention conflict representation vector that exceed the preset sparsity threshold are retained, and the dimensional components that are below the preset sparsity threshold are set to zero. The user's intention conflict representation vector in the mixed reality space is obtained. The intention conflict representation vector is used to indicate the specific dimensional direction and conflict intensity level of the conflict between different modal intentions.

6. The adaptive human-computer interaction method based on multimodal perception according to claim 1, characterized in that, The step of performing rule matching processing on a preset interaction mode switching rule base based on the intent contradiction representation vector to generate an interaction mode smooth switching instruction corresponding to the intent contradiction representation vector includes: The intention contradiction representation vector is subjected to contradiction type classification processing to identify the contradiction dominant mode corresponding to the intention contradiction representation vector, and the contradiction dominant mode identifier corresponding to the intention contradiction representation vector is obtained. The contradiction dominant mode identifier is used to indicate the main perceptual mode source that leads to the intention contradiction. Based on the conflict-dominant modality identifier, query the modality priority mapping table in the preset interaction mode switching rule base to obtain the target interaction mode switching strategy code corresponding to the conflict-dominant modality identifier; The conflict intensity level parameter in the intent contradiction representation vector is analyzed, and the interaction mode switching speed factor corresponding to the target interaction mode switching strategy code is calculated based on the conflict intensity level parameter. The interaction mode switching speed factor is used to control the speed of the interaction mode switching process. Obtain the set of configuration parameters for the source interaction mode currently running in the mixed reality space. The set of configuration parameters for the source interaction mode includes visual feedback rendering parameters, auditory feedback rendering parameters, and tactile feedback rendering parameters in the source interaction mode. Extract the target interaction mode configuration parameter template corresponding to the target interaction mode switching strategy code from the preset interaction mode switching rule base. The target interaction mode configuration parameter template includes preset values ​​for visual feedback rendering parameters, auditory feedback rendering parameters, and tactile feedback rendering parameters under the target interaction mode. Calculate the parameter difference vector between the source interaction mode configuration parameter set and the target interaction mode configuration parameter template, and determine the list of interaction parameters to be transitioned that need to be smoothed based on the parameter difference vector; Based on the interaction mode switching speed factor, a corresponding parameter interpolation curve function is generated for each interaction parameter in the list of interaction parameters to be transitioned. The parameter interpolation curve function is used to describe the trajectory of change from the source interaction mode parameter value to the target interaction mode parameter value within the switching time window. The parameter interpolation curve function is bound to the current system clock to generate a dynamic update instruction sequence for interactive mode parameters associated with the time axis; The interaction mode parameter dynamic update instruction sequence is associated and encapsulated with the contradiction-dominant modality identifier to generate the instruction header of the interaction mode smooth switching instruction; The instruction header is combined with the parameter identifiers in the list of parameters to be transitioned and encoded to generate a complete smooth switching instruction for the interaction mode, which includes the switching target mode identifier, the switching speed factor, and the parameter dynamic update function.

7. The adaptive human-computer interaction method based on multimodal perception according to claim 1, characterized in that, The step of sending the smooth switching instruction of the interaction mode to the mixed reality rendering engine to trigger the interaction paradigm shift operation of the mixed reality space includes: The target modality identifier in the smooth switching instruction of the interaction mode is parsed, and the target rendering pipeline unit to be activated in the mixed reality rendering engine is determined according to the target modality identifier. The target rendering pipeline unit includes at least one of the visual rendering pipeline branch, the auditory rendering pipeline branch and the haptic rendering pipeline branch. The switching speed factor is extracted from the smooth switching instruction of the interaction mode, and the switching speed factor is input into the frame rate control module of the mixed reality rendering engine for dynamic adjustment of the rendering frame rate, so as to obtain the target rendering frame rate parameter that matches the switching speed factor. The rendering task scheduling queue of the mixed reality rendering engine is reordered according to the target rendering frame rate parameter to improve the execution priority of the target rendering pipeline unit to be activated in the rendering task scheduling queue. Obtain the parameter dynamic update function from the smooth switching instruction of the interaction mode, and inject the parameter dynamic update function into the shader program of the target rendering pipeline unit to be activated for real-time compilation processing to generate an updated shader program containing parameter dynamic update logic. The updated shader program is run to perform pixel-by-pixel parameter adjustment on the current rendering frame of the mixed reality space, and the source interaction mode parameter value of the current rendering frame is gradually transitioned to the target interaction mode parameter value within the switching time window. The system monitors the user's physiological feedback signals for each frame of rendered image within the switching time window. These signals include user pupil diameter change data and user skin conductance response data. The user pupil diameter change data and the user skin conductance response data are then normalized to obtain normalized pupil change indices and normalized skin conductance response indices. Based on the weighted sum of these indices, the system calculates the user's cognitive load change curve during the interaction paradigm shift and extracts the peak time point of the cognitive load change curve. The peak occurrence time is compared with the preset parameter transition completion time in the parameter dynamic update function to obtain the execution lead-lag deviation of the interaction mode smooth switching instruction. The switching speed factor is fine-tuned and corrected based on the lead-lag deviation to generate a corrected switching speed factor, and the transition rate parameter of the parameter dynamic update function is updated using the corrected switching speed factor. After the parameter transition is completed, the set of target interaction mode configuration parameters is locked as the current working configuration parameters of the mixed reality rendering engine, thus completing the interaction paradigm transformation operation of the mixed reality space.

8. The adaptive human-computer interaction method based on multimodal perception according to claim 1, characterized in that, Before sending the smooth switching command for the interaction mode to the mixed reality rendering engine, the method further includes: The intent conflict representation vector is input into a preloading timing prediction network for switching instruction preloading time point prediction processing to obtain the interaction mode smooth switching instruction preloading time point corresponding to the intent conflict representation vector. Then, based on the interaction mode smooth switching instruction preloading time point, the interaction mode smooth switching instruction is pre-set in a cache queue, including: Obtain the historical sample set of intent conflict representation vectors stored inside the preloading timing prediction network. The historical sample set of intent conflict representation vectors includes intent conflict representation vectors generated in multiple historical interaction periods and the actual switching time point of the interaction mode switching after each intent conflict representation vector is generated. The intention contradiction representation vector generated at the current moment is input into the feature comparison layer of the preload timing prediction network, and feature similarity matching is performed with each historical intention contradiction representation vector in the historical sample set of the intention contradiction representation vector to obtain the feature similarity coefficient between the intention contradiction representation vector generated at the current moment and each historical intention contradiction representation vector. From the historical sample set of the intention contradiction representation vectors, select multiple similar historical intention contradiction representation vectors whose feature similarity coefficient with the intention contradiction representation vector generated at the current time exceeds a preset similarity threshold, and form a set of similar historical intention contradiction representation vectors. Extract the historical real switching time point corresponding to each similar historical intention contradiction representation vector in the similar historical intention contradiction representation vector set from the storage unit of the preloading timing prediction network, and obtain the set of historical real switching time points corresponding to the similar historical intention contradiction representation vector set. All historical real switch time points in the set of historical real switch time points are sorted, and the time point interval with the highest frequency in the set of historical real switch time points is identified as the candidate preloading time point interval. Based on the distribution density of the intent contradiction representation vector generated at the current moment in the historical sample set of intent contradiction representation vectors, the candidate preloading time point interval is offset correction processed to obtain the interaction mode smooth switching instruction preloading time point corresponding to the intent contradiction representation vector generated at the current moment. Calculate the time interval between the current moment and the preloading time of the smooth switching instruction based on the preloading time of the smooth switching instruction for the interaction mode, and obtain the preset waiting time of the smooth switching instruction for the interaction mode in the cache queue. The smooth switching instruction of the interaction mode is encapsulated into a preloaded instruction unit with a preset waiting time label, and the preloaded instruction unit is sent to the instruction prefetch cache queue of the mixed reality rendering engine for queue writing operation; In the instruction prefetch cache queue, the preloaded instruction units are sorted according to the preset wait duration tag to ensure that the storage location of the preloaded instruction units in the instruction prefetch cache queue matches the wake-up time order indicated by the preset wait duration tag. When the system clock reaches the preloading time point of the interaction mode smooth switching instruction, the preloaded instruction unit is woken up from the instruction prefetch cache queue, and the woken-up preloaded instruction unit is converted into an interaction mode smooth switching instruction to be executed for the mixed reality rendering engine to obtain and execute.

9. An adaptive human-computer interaction system based on multimodal perception, characterized in that, include: processor; A machine-readable storage medium for storing machine-executable instructions of the processor; The processor is configured to execute the adaptive human-computer interaction method based on multimodal perception as described in any one of claims 1 to 8 by executing the machine-executable instructions.

10. A computer program product, characterized in that, The computer program product includes machine-executable instructions stored in a computer-readable storage medium. A processor of a computer device reads the machine-executable instructions from the computer-readable storage medium and executes the machine-executable instructions, causing the computer device to perform the adaptive human-computer interaction method based on multimodal perception as described in any one of claims 1 to 8.