Multi-mode virtual-real fusion interaction method and system
Data acquisition through the depth-aware camera and microphone array, combined with a lightweight convolutional neural network to recognize gestures and generate haptic feedback, the problem of high interaction delay and error rate in virtual reality interactive systems is solved, and the precise linkage between gesture operation and device status is achieved and the coordination of multimodal instructions is improved.
Patent Information
- Application Number
- CN202510387510.7
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-03-31
- Publication Date
- 2025-07-04
AI Technical Summary
In the prior art, virtual reality and augmented reality interactive systems have problems such as high interaction delay, disconnection of tactile feedback from the device state, and poor coordination of multimodal instructions, resulting in insufficient immersion and operational authenticity of the user experience, especially in scenarios where refined operations are required.
Through the depth perception camera, the hand three-dimensional coordinate data is collected and the microphone array is collected. The lightweight convolutional neural network model is used to identify the gesture operation intention, and the tactile feedback parameters are generated in combination with the virtual device status to realize the spatiotemporal alignment and synchronization control of gestures and voice commands, reducing the error operation rate.
It realizes accurate haptic feedback that links low-cost and low-latency gesture operation with device status, significantly improves multimodal command coordination capabilities and reduces the error operation rate.
Smart Images

Figure CN120255698A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the field of human-computer interaction technology, and more particularly, to a multi-modal virtual-real fusion interaction method and system. Background Art
[0002] In recent years, virtual reality (VR) and augmented reality (AR) technologies have been rapidly popularized in fields such as education, training, entertainment, and remote operation, and users' demand for immersive real experiences has been continuously increasing. However, existing technologies generally face problems such as high interaction latency, disconnection between tactile feedback and device status, and poor coordination of multi-modal instructions, resulting in insufficient immersion and operational authenticity of the user experience. Especially in scenarios that require fine operations (such as virtual training of power equipment), the error operation rate remains high, seriously restricting the depth and breadth of technology applications.
[0003] In traditional solutions, a handle-based interaction system (such as patent CN202098765432) relies on button triggers for operations. Although it can provide simple feedback through vibration, the operation logic is significantly different from that of real devices, resulting in tactile distortion and lack of dynamic association with the device status. The high-end force feedback glove combined with eye tracking solution (such as the literature "Research on Multi-modal Interaction in Power Virtual Training") can simulate operation resistance and optimize target switching, but it is difficult to be scaled up due to high hardware costs (over 50,000 yuan per set) and data processing latency (>80 ms). None of the above solutions effectively balance the core requirements of low cost, low latency, and natural interaction.
[0004] In summary, how to achieve tactile feedback with gesture operations linked to device status through low-cost and low-latency natural interaction methods, and improve the multi-modal instruction coordination ability to reduce the error operation rate is a technical problem that urgently needs to be solved currently. Summary of the Invention
[0005] The main objective of the present invention is to provide a multi-modal virtual-real fusion interaction method and system to solve the technical problem of how to achieve tactile feedback with gesture operations linked to device status through low-cost and low-latency natural interaction methods, and improve the multi-modal instruction coordination ability to reduce the error operation rate, thereby realizing precise tactile feedback with gesture operations linked to the status of virtual devices, improving the multi-modal instruction coordination ability, and significantly reducing the error operation rate.
[0006] To achieve the above objective, the present invention provides a multi-modal virtual-real fusion interaction method and system.
[0007] In a first aspect, the present invention provides a multi-modal virtual-real fusion interaction method, the method comprising:
[0008] Collecting three-dimensional coordinate data of the hand through a depth perception camera and collecting voice data through a microphone array;
[0009] Input the three-dimensional coordinate data of the hand into the gesture recognition model, and output the operation intention parameters through the gesture recognition model;
[0010] Obtain the real-time status data of the virtual device corresponding to the operation intention parameters, generate the tactile feedback parameters based on the real-time status data of the virtual device, and associate the tactile feedback parameters with the real-time status data of the virtual device through a linear mapping relationship;
[0011] Generate a vibration waveform according to the tactile feedback parameters and output it to the tactile execution unit;
[0012] Parse the voice data to obtain a voice command, perform spatio-temporal alignment processing on the voice command and the operation intention parameters, and generate a virtual-real synchronization control command.
[0013] Optionally, the step of inputting the three-dimensional coordinate data of the hand into the gesture recognition model and outputting the operation intention parameters through the gesture recognition model includes:
[0014] The gesture recognition model is a lightweight convolutional neural network model, and the three-dimensional coordinate data of the hand is received through the input layer of the lightweight convolutional neural network model;
[0015] Output the operation intention parameters including the knob rotation angle and the pressing force through the output layer of the lightweight convolutional neural network model.
[0016] Optionally, the linear mapping relationship is f = K × I, where f represents the vibration frequency in the tactile feedback parameters, K is a preset mapping coefficient, and I represents the current value in the real-time status data of the virtual device. The tactile feedback parameters are used to generate the vibration waveform.
[0017] Optionally, the step of performing spatio-temporal alignment processing on the voice command and the operation intention parameters to generate a virtual-real synchronization control command includes:
[0018] Extract the timestamp of the voice command and the generation time of the operation intention parameters;
[0019] When the difference between the timestamp and the generation time is less than 50 ms, trigger the generation of the virtual-real synchronization control command.
[0020] Optionally, after generating the vibration waveform according to the tactile feedback parameters and outputting it to the tactile execution unit, it further includes:
[0021] The tactile execution unit includes an array of vibration motors, and the drive circuit of the array of vibration motors receives the vibration waveform and converts it into a pulse width modulation signal.
[0022] Optionally, generating the haptic feedback parameters based on the real-time status data of the virtual device includes:
[0023] When the real-time status data of the virtual device contains a fault code, the haptic feedback parameters are generated in a gradient change mode, and the gradient change mode is used to indicate a continuous change process in which the vibration frequency linearly increases from 10 Hz to 100 Hz.
[0024] In a second aspect, the present invention provides a multi-modal virtual-real fusion interaction system, and the system is applied to the method in the first aspect. The system includes:
[0025] A data acquisition unit, which is used to collect three-dimensional hand coordinate data through a depth perception camera and collect voice data through a microphone array;
[0026] A gesture recognition unit, which is connected to the data acquisition unit. The gesture recognition unit is used to input the three-dimensional hand coordinate data into a gesture recognition model and output operation intention parameters through the gesture recognition model;
[0027] A virtual-real state haptic mapping unit, which is connected to the gesture recognition unit. The virtual-real state haptic mapping unit is used to obtain corresponding real-time status data of the virtual device according to the operation intention parameters, generate haptic feedback parameters based on the real-time status data of the virtual device, and associate the haptic feedback parameters with the real-time status data of the virtual device through a linear mapping relationship;
[0028] A haptic waveform driving unit, which is connected to the virtual-real state haptic mapping unit. The haptic waveform driving unit is used to generate a vibration waveform according to the haptic feedback parameters and output it to a haptic execution unit;
[0029] A multi-modal instruction spatio-temporal alignment unit, which is connected to the haptic waveform driving unit. The multi-modal instruction spatio-temporal alignment unit is used to parse the voice data to obtain a voice instruction, perform spatio-temporal alignment processing on the voice instruction and the operation intention parameters, and generate a virtual-real synchronization control instruction.
[0030] Optionally, the depth perception camera is an RGB-D camera, and the depth sensor of the RGB-D camera acquires the three-dimensional hand coordinate data at a sampling frequency not lower than 30 Hz.
[0031] Optionally, the gesture recognition unit includes:
[0032] A lightweight convolutional neural network module, which is used to receive the three-dimensional hand coordinate data through the input layer of the lightweight convolutional neural network model;
[0033] An operation intention parameter output module, which is connected to the lightweight convolutional neural network module, and is used to output operation intention parameters including the knob rotation angle and the pressing force through the output layer of the lightweight convolutional neural network model.
[0034] Optionally, the multi-modal instruction spatio-temporal alignment unit includes:
[0035] A time extraction module, which is used to extract the timestamp of the voice instruction and the generation time of the operation intention parameter;
[0036] A virtual-real synchronous control instruction generation module, which is connected to the time extraction module, and is used to trigger the generation of the virtual-real synchronous control instruction when the difference between the timestamp and the generation time is less than 50 ms.
[0037] The multi-modal virtual-real fusion interaction method and system provided by this application cover core components such as a depth perception camera, a microphone array, a gesture recognition model, and a tactile execution unit. The depth perception camera is responsible for collecting three-dimensional coordinate data of the hand, and the microphone array synchronously records voice information. These data are then sent to the gesture recognition model to parse out operation intention parameters. Based on these parameters, the system obtains the real-time state of the virtual device, and accordingly generates tactile feedback parameters, which are associated with the device state data through linear mapping. Then, the tactile feedback parameters are converted into vibration waveforms and output by the tactile execution unit. At the same time, the voice data is parsed into instructions, which are spatio-temporally aligned with the operation intention parameters to jointly form a virtual-real synchronous control instruction. This method not only realizes the precise linkage between gestures and the state of virtual devices and tactile feedback, but also significantly enhances the coordination of multi-modal instructions and effectively reduces the misoperation rate. Description of the Drawings
[0038] The specification drawings forming a part of this application are used to provide a further understanding of the present invention. The schematic embodiments of the present invention and their descriptions are used to explain the present invention and do not constitute an improper limitation to the present invention. In the drawings:
[0039] Figure 1 It is a schematic flow chart of the multi-modal virtual-real fusion interaction method provided by this application;
[0040] Figure 2 It is a schematic connection diagram of the multi-modal virtual-real fusion interaction system provided by this application.
[0041] Through the above-mentioned accompanying drawings, specific embodiments of the present application have been shown, and there will be a more detailed description hereinafter. These drawings and textual descriptions are not intended to limit the scope of the concept of the present application in any way, but to illustrate the concept of the present application to those skilled in the art by referring to specific embodiments. Detailed Description of the Embodiments
[0042] To make the objectives, technical solutions, and advantages of the present application clearer, the technical solutions in the present application will be clearly and completely described below with reference to the accompanying drawings in the present application. Obviously, the described embodiments are some, but not all, of the embodiments of the present application. All other embodiments obtained by those of ordinary skill in the art based on the embodiments in the present application without creative efforts shall fall within the scope of protection of the present application.
[0043] In the description of the present invention, terms such as "first", "second", "third", "fourth", etc., if any, in the specification and claims of the present invention and the above-mentioned accompanying drawings are used to distinguish similar objects and do not necessarily have to be used to describe a specific order or sequence. It should be understood that such data may be interchanged under appropriate circumstances so that the embodiments of the present invention described herein can be implemented in an order other than those illustrated or described herein.
[0044] In the present invention, words such as "exemplary" or "for example" are used to represent examples, illustrations, or explanations. Any embodiment or design solution described as "exemplary" or "for example" in the present application should not be construed as being more preferred or having more advantages than other embodiments or design solutions. Rather, the use of words such as "exemplary" or "for example" is intended to present relevant concepts in a specific manner.
[0045] To address the above problems, the present application provides a multi-modal virtual-real fusion interaction method and system. The method uses a depth perception camera and a microphone array to collect hand motion and voice data. In combination with a gesture recognition model, the three-dimensional coordinates of the hand are parsed to output operation intention parameters. According to the operation intention, the state of the virtual device is obtained in real time, tactile feedback parameters are generated, and through linear mapping, they are associated with the device state and converted into vibration waveforms and output to the tactile execution unit. At the same time, the voice data is parsed into instructions, which are spatio-temporally aligned with the operation intention parameters to form virtual-real synchronous control instructions. This technical concept realizes low-cost, low-latency gesture and voice fusion interaction, accurately links the state of the virtual device, enhances multi-modal instruction coordination, and significantly reduces the misoperation rate.
[0046] The technical solutions of the present application and how the technical solutions of the present application solve the above technical problems will be described in detail below with specific embodiments. These specific embodiments below can be combined with each other, and the same or similar concepts or processes may not be repeated in some embodiments. The embodiments of the present application will be described below with reference to the accompanying drawings.
[0047] Figure 1 This is a schematic flowchart of the multi-modal virtual-real fusion interaction method provided for this application, which details the multi-modal virtual-real fusion interaction method. As Figure 1 shown, the multi-modal virtual-real fusion interaction method provided in this embodiment includes:
[0048] S101: Collect three-dimensional coordinate data of the hand through a depth perception camera and collect voice data through a microphone array.
[0049] Specifically, the specific implementation of step S101 includes:
[0050] I. Hardware configuration stage:
[0051] 1. Deploy the physical structures of the depth perception camera and the microphone array
[0052] Select the Intel RealSense D455 depth perception camera as the acquisition device for the three-dimensional coordinate data of the hand. This camera is built-in with an infrared laser projector and a binocular infrared sensor, and supports outputting a depth image stream of 30 frames per second at a resolution of 1280×720.
[0053] Configure the ReSpeaker 6-Mic circular microphone array as the voice data acquisition device. The 6 digital MEMS microphones in the array are installed in a circular layout on a PCB substrate with a diameter of 80mm, and the distance between each microphone is 40mm.
[0054] 2. Initialize the parameters of the depth perception camera
[0055] Call the rs-depth module in the Intel RealSense SDK to enable the stereo vision depth calculation algorithm in the camera. This algorithm is based on the binocular disparity principle and calculates the depth value of pixel points through the disparity maps collected by the left and right infrared sensors.
[0056] Set the depth perception range to 0.2m to 1.5m, and the coordinate accuracies of the X-axis (horizontal direction), Y-axis (vertical direction), and Z-axis (depth direction) in the camera coordinate system are ±1.5mm, ±1.8mm, and ±2.0mm respectively.
[0057] 3. Configure the hand coordinate data acquisition process
[0058] Use the Hand Tracking module in the OpenCV library and call the BlazePalm hand detection model under the mediapipe framework to extract the key points of the hand contour in the depth image in real time.
[0059] The two-dimensional pixel coordinates of the 21 detected hand key points (including the palm root point and each finger joint point) are converted into three-dimensional world coordinate system coordinates through the rs2_deproject_pixel_to_point function, and output as a three-dimensional coordinate data set containing [x, y, z] values in the data format of a 32-bit floating-point array.
[0060] 4. Set the voice acquisition parameters of the microphone array
[0061] Initialize the microphone array through the ReSpeaker SDK provided by Seeed Studio, set the voice sampling rate to 48 kHz, the quantization bit width to 16 bit, and use the beamforming algorithm (Minimum Variance Distortionless Response, MVDR) for sound source orientation to suppress noise in non-target directions.
[0062] Activate the Voice Activity Detection (VAD) module, and judge the effective voice segment based on the energy threshold method in the G.729 standard. Trigger data storage only when the voice signal amplitude continuously exceeds -35 dBFS and the duration is greater than 200 ms.
[0063] 5. Implement a multimodal data synchronization mechanism
[0064] Synchronize the hardware clocks of the camera and the microphone array through the Precision Time Protocol (PTP) under the Linux system, and the timestamp accuracy reaches the microsecond level.
[0065] Attach a 64-bit Unix timestamp to each frame of depth image data. Segment the voice data block by 10 ms and attach timestamps with the same time reference. Achieve hardware-level time alignment of the two types of data through a shared memory queue to ensure that the spatio-temporal synchronization error is less than ±1 ms.
[0066] 6. Data preprocessing and output interface
[0067] The three-dimensional hand coordinate data is smoothed through the Kalman Filter. The state vector includes the position (x, y, z) and velocity (vx, vy, vz), and the observation noise covariance matrix is set to diag([0.1, 0.1, 0.1]).
[0068] The output data is encapsulated in JSON format, including three fields: timestamp, the coordinates of 21 hand key points (hand_joints), the original depth image frame (depth_frame), and the audio data chunk (audio_chunk), and is sent to the data processing unit through the USB3.0 interface at a cycle of 10 ms.
[0069] II. Technical Verification Metrics
[0070] Calibration error between the depth perception camera coordinate system and the physical world coordinate system: < 2 mm for the X / Y axis and < 3 mm for the Z axis (at a measurement distance of 1 m)
[0071] Recognition latency of voice commands: The time from the sound wave reaching the microphone to the output of the data packet is < 15 ms
[0072] Synchronization accuracy of multimodal data: After PTP synchronization, the time deviation between visual and voice data is < ±1 ms
[0073] S102: Input the three-dimensional hand coordinate data into the gesture recognition model, and output the operation intention parameters through the gesture recognition model.
[0074] Specifically, inputting the three-dimensional hand coordinate data into the gesture recognition model and outputting the operation intention parameters through the gesture recognition model includes:
[0075] The gesture recognition model is a lightweight convolutional neural network model, and the input layer of the lightweight convolutional neural network model receives the three-dimensional hand coordinate data;
[0076] Output the operation intention parameters including the knob rotation angle and the pressing force through the output layer of the lightweight convolutional neural network model.
[0077] Specific implementation of step S102:
[0078] I. Gesture Recognition Model Construction Stage
[0079] 1. Define the structure of the lightweight convolutional neural network model
[0080] Adopt the MobileNetV3-Small architecture as the core network of the gesture recognition model. The input layer is designed with 63 nodes to receive the three-dimensional coordinates of 21 hand key points (each key point contains three values of x, y, and z, totaling a 63-dimensional input vector).
[0081] Input data preprocessing: Normalize the three-dimensional hand coordinate data to the interval [-1, 1]. The normalization formula is:
[0082]
[0083] Among them, xmin and xmax are respectively the minimum and maximum values of the x - coordinates in the training dataset, and the same applies to the y and z coordinates.
[0084] 2. Configure the model hierarchy and parameters
[0085] The model includes the following layer structure: input layer (63 nodes) → 1×1 convolutional layer (16 channels, ReLU activation) → depthwise separable convolutional layer (32 channels, stride 2) → SE (Squeeze - and - Excitation) attention module → global average pooling layer → fully connected layer (2 nodes).
[0086] The two nodes of the output layer respectively correspond to the rotation angle of the knob (unit: degree, range 0° - 360°) and the pressing force (unit: dimensionless, range 0 - 1). The activation functions are the linear function (for rotation angle) and the Sigmoid function (for pressing force).
[0087] 3. Model training and optimization
[0088] Dataset construction: Collect 10,000 groups of hand gesture samples. Each group contains the three - dimensional coordinates of 21 key points and the corresponding true rotation angle of the knob (measured by an encoder) and pressing force (measured by a pressure sensor).
[0089] Loss function: The Huber loss is used for the rotation angle, and the binary cross - entropy loss is used for the pressing force. The total loss is the weighted sum of the two (weight ratio 1:1).
[0090] Optimizer: Use the Adam optimizer. The initial learning rate is set to 0.001, the batch size is 32, and the number of training epochs is 100.
[0091] II. Gesture recognition inference stage
[0092] 1. Real - time data input and feature extraction
[0093] Parse the JSON - formatted hand coordinates data output in step S101 into a 63 - dimensional vector, and perform real - time normalization processing according to the normalization parameters in the training stage.
[0094] Extract spatial local features through the 1×1 convolutional layer of MobileNetV3 - Small, use the depthwise separable convolutional layer to reduce the computational amount, and the SE module dynamically adjusts the channel weights to enhance the key feature response.
[0095] 2. Calculation of operation intention parameters
[0096] The fully connected layer maps the feature after global average pooling to a 2 - dimensional output:
[0097] Knob rotation angle: directly output a floating-point value. For example, an output value of 123.45 means a rotation of 123.45°.
[0098] Pressing force: Compressed to the range of 0 - 1 through the Sigmoid function. For example, an output of 0.75 means the pressing intensity is 75% of the full scale.
[0099] 3. Output result calibration and filtering
[0100] Perform Kalman filtering on the rotation angle (the state vector includes the angle and angular velocity), set the observation noise covariance to 0.1, and the process noise covariance to 0.01 to suppress jitter.
[0101] The pressing force output is filtered by moving average (window length 5 frames) to eliminate instantaneous noise.
[0102] III. Technical verification indicators
[0103] Knob angle prediction error: Root Mean Square Error (RMSE) of the test set < 2° (at an operating distance of 1m).
[0104] Pressing force prediction error: Mean Absolute Error (MAE) of the test set < 0.05 (range 0 - 1).
[0105] Single inference latency: On NVIDIA Jetson Nano hardware, the end-to-end time from input data to output result < 5ms.
[0106] IV. Example implementation process for reference
[0107] 1. Load the pre-trained MobileNetV3-Small model weight file (in the format of TensorFlow Lite model).
[0108] 2. Initialize the TensorFlow Lite interpreter and bind the input / output tensors:
[0109] Input tensor: shape = [1, 63], dtype = float32, corresponding to the normalized 63-dimensional hand coordinates.
[0110] Output tensor: shape = [1, 2], dtype = float32, the 0th position is the rotation angle, and the 1st position is the pressing force.
[0111] 3. Real-time loop:
[0112] Read the latest hand coordinate data from the data queue in step S101.
[0113] Perform normalization → model inference → filtering process → output the operation intention parameters to step S103.
[0114] S103: Obtain the corresponding real-time status data of the virtual device according to the operation intention parameter, generate a haptic feedback parameter based on the real-time status data of the virtual device, and associate the haptic feedback parameter with the real-time status data of the virtual device through a linear mapping relationship.
[0115] Specifically, the linear mapping relationship is f = K × I, where f represents the vibration frequency in the haptic feedback parameter, K is a preset mapping coefficient, and I represents the current value in the real-time status data of the virtual device. The haptic feedback parameter is used to generate the vibration waveform.
[0116] Specifically, when the real-time status data of the virtual device contains a fault code, the haptic feedback parameter is generated according to a gradient change pattern, and the gradient change pattern is used to indicate the continuous change process of the vibration frequency linearly increasing from 10 Hz to 100 Hz.
[0117] Specific implementation of step S103:
[0118] Real-time status data acquisition stage of the virtual device:
[0119] 1. Establish a communication interface between the operation intention parameter and the virtual device
[0120] The operation intention parameters (knob rotation angle, pressing force) are transmitted to the virtual device simulation module in the Unity3D engine through the UDP protocol. The simulation module embeds a Modbus TCP protocol stack and establishes a long connection with the real-time status database of the virtual device.
[0121] The real-time status data of the virtual device includes a current value (I, unit: A), a voltage value, a temperature value, and a fault code (for example, E001 represents an overcurrent fault). The data update frequency is 100 Hz and is stored in a circular buffer through shared memory.
[0122] 2. Haptic feedback parameter generation logic
[0123] Linear mapping under normal working conditions:
[0124] Define the value of the mapping coefficient K as 50 Hz / A. When the current value I in the real-time status data of the virtual device is 0.8 A, the vibration frequency f = 50 × 0.8 = 40 Hz.
[0125] The haptic feedback parameter is encapsulated in JSON format and contains two fields: vibration frequency (frequency) and amplitude. The amplitude is fixed at 0.5 N (Newton).
[0126] Gradient change pattern under fault conditions:
[0127] When a fault code (such as E001) is detected, start the linear growth algorithm: the vibration frequency starts from 10 Hz and increases at a rate of 900 Hz per second (i.e., increases by 0.9 Hz every 1 ms) until it reaches 100 Hz and then remains constant.
[0128] The duration of the gradient change process is 100 ms. The frequency update is triggered by a timer interrupt, and the frequency value is updated every 1 ms and written into the haptic feedback parameter.
[0129] Data processing and mapping implementation:
[0130] 3. Real-time status data matching and filtering
[0131] According to the knob rotation angle (e.g., 123.45°) in the operation intention parameter, perform the nearest neighbor interpolation algorithm in the real-time status database of the virtual device to match the current value corresponding to the angle:
[0132] The current-angle correspondence table stored in the database is at 1° intervals (0°: 0 A, 1°: 0.01 A,..., 360°: 3.6 A). Round 123.45° to 123° and 124°, and calculate the weighted average value I = 0.01×(123.45 - 123)×(current corresponding to 124° - current corresponding to 123°) + current corresponding to 123°.
[0133] 4. Fault code detection and mode switching
[0134] The fault code field in the real-time status data of the virtual device uses 8-bit binary encoding (e.g., 0x01 indicates normal, 0x02 indicates E001 fault). Detect the fault status through a bit mask: if the result of the bitwise AND operation between the status word and 0x02 is non-zero, it is determined as an E001 fault and the gradient change mode is triggered.
[0135] Technical verification indicators:
[0136] Haptic feedback parameter generation delay: The time taken from receiving the operation intention parameter to outputting the haptic feedback parameter is < 2 ms.
[0137] Gradient change linearity error: The maximum deviation between the frequency change curve and the ideal straight line is < ±0.5 Hz (verified by least squares fitting).
[0138] Fault code response delay: The time taken from the occurrence of the fault to the start of haptic feedback is < 5 ms.
[0139] Example implementation process that those skilled in the art can refer to:
[0140] 1. Configure the Modbus TCP server of the virtual device in Unity 3D, set the IP address to 192.168.1.100, the port to 502, and define the data register addresses:
[0141] Current value: register 40001, data type is float32; Fault code: register 40005, data type is uint8.
[0142] 2. Initialize the haptic feedback parameter generation thread, bind it to CPU core 3, and set the thread priority to the real-time level (SCHED_FIFO).
[0143] 3. Real-time loop:
[0144] Read the operation intention parameters and the real-time status data of the virtual device from the shared memory.
[0145] If the fault code is 0x00, execute the linear mapping f = K × I; if it is 0x02, call the gradient change algorithm to generate the frequency sequence.
[0146] Write the haptic feedback parameters into the vibration waveform generation queue in step S104.
[0147] S104: Generate a vibration waveform according to the haptic feedback parameters and output it to the haptic execution unit.
[0148] Specifically, after generating the vibration waveform according to the haptic feedback parameters and outputting it to the haptic execution unit, it further includes:
[0149] The haptic execution unit includes an array of vibration motors, and the drive circuit of the array of vibration motors receives the vibration waveform and converts it into a pulse width modulation signal.
[0150] The specific implementation manner of step S104 includes:
[0151] Vibration waveform generation stage:
[0152] 1. Haptic feedback parameter parsing and waveform definition
[0153] The haptic feedback parameters include the vibration frequency (f, unit: Hz) and the amplitude (A, unit: N), and extract the fields from the JSON data output in step S103:
[0154] Frequency range: 10 Hz to 100 Hz (dynamically adjusted during the fault mode gradient change), fixed frequency under normal working conditions (such as 40 Hz).
[0155] The amplitude is fixed at 0.5 N, and the corresponding drive voltage range of the vibration motor is 0 - 3.3 V (the linear relationship between voltage and amplitude is: voltage = amplitude × 6.6 V / N).
[0156] 2. Implementation of Sine Wave Generation Algorithm
[0157] Generate a sine wave form using the look-up table method:
[0158] Pre-generate a sine wave table of 4096 points (one cycle), and each point is stored as a 16-bit signed integer, with the value range from -32768 to +32767.
[0159] Calculate the sampling interval according to the target frequency f: Phase increment Δθ = (f × 4096) / sampling rate, where the sampling rate is set to 10 kHz. For example, when f = 40 Hz, Δθ = (40 × 4096) / 10000 = 16.384 points / sampling.
[0160] Calculate the waveform data in real time: Output waveform points in the phase accumulation manner through a circular buffer to avoid phase truncation error.
[0161] Vibration waveform output and signal conversion:
[0162] 3. Drive Circuit Configuration
[0163] The tactile execution unit consists of 8 LRAs (Linear Resonant Actuators) of model C10-100 to form a 4×2 array. The working voltage of a single motor is 3.3 V, and the resonant frequency is 175 Hz ± 10 Hz.
[0164] The drive circuit uses the TIDRV2605L tactile drive chip, which supports the PWM input mode. The PWM carrier frequency is set to 25 kHz, and the dead time is 100 ns.
[0165] 4. PWM Signal Generation Process
[0166] Convert the sine waveform data into a PWM signal with a variable duty cycle:
[0167] Waveform amplitude mapping formula: Duty cycle D = 50% + (A / 0.5N) × 50%, where A is the amplitude in the tactile feedback parameter (for example, when A = 0.5 N, D = 100%).
[0168] Frequency control: Write the target frequency register (address 0x16) through the I 2 C interface of DRV2605L, and the chip automatically adjusts the PWM period to match the target frequency.
[0169] The output stage of the drive circuit adopts the H-bridge topology structure, and controls the forward and reverse rotation of the motor through four PWM signals (PH1 / PH2, PH3 / PH4) to achieve bidirectional vibration.
[0170] Special handling of the fault mode:
[0171] 5. Dynamic Generation of Gradient Change Waveforms
[0172] When a fault code is detected, start the linear frequency increment algorithm:
[0173] The initial frequency f_start = 10 Hz, the termination frequency f_end = 100 Hz, and the change rate Δf = 90 Hz / 100 ms = 0.9 Hz / ms.
[0174] Update the frequency value every 1 ms: f_current = f_start + Δf × t, where t is the time counter (0 ≤ t ≤ 100 ms).
[0175] Ensure waveform phase continuity: During the frequency change process, maintain the continuous calculation of the phase accumulator to avoid waveform jumps.
[0176] Technical verification indicators:
[0177] 1. Vibration waveform generation delay: The time from receiving tactile feedback parameters to outputting the PWM signal is < 2 ms.
[0178] 2. PWM signal accuracy: Frequency error < ±1 Hz, duty cycle error < ±5% (under the power supply condition of 3.3 V).
[0179] 3. Fault mode response consistency: During the gradient change process, the linear correlation coefficient R between the measured frequency and the target value 2 ≥ 0.99.
[0180] Example implementation process for reference:
[0181] 1. Initialize the DRV2605L driver chip:
[0182] Configure the working mode as PWM input through the I 2 C bus (SCL = 100 kHz) (write 0x03 to register 0x01).
[0183] Set the overcurrent protection threshold to 1.2 A (write 0x0C to register 0x25).
[0184] 2. Create two independent threads:
[0185] Thread 1 (high priority): Read the tactile feedback parameters from the queue in step S103, generate sine wave data and write it into the DAC buffer.
[0186] Thread 2 (real-time level): Send the DAC data to the PWM generator of the DRV2605L through the SPI interface, triggering data transmission every 0.1 ms.
[0187] 3. Monitor the fault flag bit in real time:
[0188] If a fault code is detected, immediately interrupt the current waveform generation, switch to the gradient change mode, and reset the phase accumulator.
[0189] S105: Parse the voice data to obtain a voice command, perform spatio-temporal alignment processing on the voice command and the operation intention parameter, and generate a virtual-real synchronization control command.
[0190] Specifically, the spatio-temporal alignment processing of the voice command and the operation intention parameter to generate a virtual-real synchronization control command includes:
[0191] Extract the timestamp of the voice command and the generation time of the operation intention parameter;
[0192] When the difference between the timestamp and the generation time is less than 50 ms, trigger the generation of the virtual-real synchronization control command.
[0193] Specific implementation of step S105:
[0194] The voice command parsing stage includes:
[0195] 1. Voice data preprocessing and endpoint detection
[0196] From the voice data block output in step S101 (sampling rate 48 kHz, quantization 16 bit), through the Voice Activity Detection (VAD) module in the WebRTC open source library, based on the Gaussian Mixture Model (GMM), judge the voice segment and non-voice segment, and only retain the effective voice segment with an amplitude exceeding -35 dBFS and a duration ≥ 200 ms.
[0197] Perform pre-emphasis processing on the effective voice segment (filter coefficient α = 0.97), then frame (frame length 25 ms, frame shift 10 ms) and add a Hamming window to generate a Short-Time Fourier Transform (STFT) spectrogram.
[0198] 2. Voice recognition model deployment
[0199] Adopt the Vosk open source voice recognition engine, load the pre-trained small English model (vosk-model-small-en-us-0.15), the input is Mel Frequency Cepstral Coefficients (MFCC, 26 dimensions), and the output is a phoneme sequence and the corresponding text.
[0200] Custom instruction vocabulary: includes instruction words and parameters such as "rotate", "press", "stop" (such as "rotate30degrees"), and is parsed into a structured JSON instruction through a Finite State Machine (FSM), and the fields include action type (action), parameter value (value), and timestamp (timestamp).
[0201] The spatio-temporal alignment processing stage includes:
[0202] 3. Multi-modal timestamp matching
[0203] Extract the generated timestamp t_gesture from the operation intention parameter (output of step S102), extract the start timestamp t_voice of the voice segment from the voice command, and calculate the time difference Δt = |t_gesture - t_voice|.
[0204] Timestamp synchronization mechanism: Synchronize the clocks of the depth camera, microphone array, and processing host at the microsecond level based on the NTP protocol, and the timestamp error ≤ ±1ms.
[0205] 4. Instruction fusion logic
[0206] When Δt ≤ 50ms, trigger the generation of virtual-real synchronization control instructions:
[0207] If the voice command is "rotate X degrees" and the gesture parameter is the rotation angle θ of the knob, then generate the control instruction: {"type": "synced_rotate", "angle": θ, "voice_param": X}, and verify whether the deviation between θ and X is within the range of ±5°. If it exceeds the limit, trigger the error code E201.
[0208] If the voice command is "press" and the gesture parameter pressing force ≥ 0.8, then generate the control instruction: {"type": "synced_press", "force": 0.8}.
[0209] The generation of virtual-real synchronization control instructions includes:
[0210] 5. Instruction encapsulation and transmission
[0211] The control instruction is encapsulated in the Protobuf format, and the defined fields include: instruction ID (uint64), timestamp (int64), operation type (enum), parameter value (float32 array), checksum (CRC32).
[0212] Send it to the virtual device control interface (IP: 192.168.1.200, port: 6000) through the UDP protocol. The data packet length is fixed at 128 bytes, and the sending period is 10ms.
[0213] Technical verification indicators include:
[0214] 1. Voice command recognition latency: The time taken from the end of the voice segment to generate the structured JSON instruction < 20ms (tested on NVIDIA Jetson Nano).
[0215] 2. Space-time alignment error: The standard deviation σ of the measured Δt < 2 ms (1000 sample tests).
[0216] 3. Multimodal instruction fusion accuracy: The correct rate of instruction matching in the test set (500 groups of samples) ≥ 95% (within the deviation tolerance).
[0217] The example implementation process includes:
[0218] 1. Configure the Vosk speech recognition environment: Install vosk-api 0.3.31, load the model file into memory, initialize the recognizer, and set the regular expression of the instruction vocabulary.
[0219] 2. Start the NTP time synchronization service: Run the chronyd service on the host, configure it as an NTP client, and synchronize it to the LAN NTP server (192.168.1.1).
[0220] 3. Real-time processing loop:
[0221] Thread 1: Read the valid speech segment from the speech data queue, and call the Vosk recognizer to generate text instructions.
[0222] Thread 2: Read the latest data from the operation intention parameter queue and extract the timestamp t_gesture.
[0223] Thread 3: Compare t_voice and t_gesture. If Δt ≤ 50 ms, execute the instruction fusion logic and send control instructions.
[0224] This instance provides a multimodal virtual-real fusion interaction method. This method synchronously collects hand motion and speech data through a depth camera and a microphone array, uses a lightweight convolutional neural network (MobileNetV3-Small) to real-time analyze the operation intention parameters of gestures (rotation angle, pressing force), combines the virtual device current value to generate tactile feedback parameters through linear mapping (f = K × I) or fault gradient mode (10 Hz → 100 Hz), and drives the vibration motor array to output PWM modulated tactile signals; at the same time, uses the Vosk speech engine to parse instructions and perform space-time alignment with gesture parameters (time difference ≤ 50 ms), and generates virtual-real synchronous control instructions through clock synchronization (NTP protocol) and instruction fusion verification, realizing multimodal low-latency coordination of gestures, speech, and tactile feedback, improving operation accuracy and reducing the false trigger rate.
[0225] Figure 2 It is a connection schematic diagram of the multimodal virtual-real fusion interaction system provided by this application; as Figure 2 shown, this application provides a multimodal virtual-real fusion interaction system, and the system is applied toFigure 1 The interactive method described in the embodiment, the system includes:
[0226] A data acquisition unit, which is used to collect three-dimensional hand coordinate data through a depth perception camera and collect voice data through a microphone array;
[0227] A gesture recognition unit, which is connected to the data acquisition unit. The gesture recognition unit is used to input the three-dimensional hand coordinate data into a gesture recognition model and output operation intention parameters through the gesture recognition model;
[0228] A virtual-real state tactile mapping unit, which is connected to the gesture recognition unit. The virtual-real state tactile mapping unit is used to obtain corresponding real-time state data of a virtual device according to the operation intention parameters, generate tactile feedback parameters based on the real-time state data of the virtual device, and associate the tactile feedback parameters with the real-time state data of the virtual device through a linear mapping relationship;
[0229] A tactile waveform driving unit, which is connected to the virtual-real state tactile mapping unit. The tactile waveform driving unit is used to generate a vibration waveform according to the tactile feedback parameters and output it to a tactile execution unit;
[0230] A multi-modal instruction spatio-temporal alignment unit, which is connected to the tactile waveform driving unit. The multi-modal instruction spatio-temporal alignment unit is used to parse the voice data to obtain a voice instruction, perform spatio-temporal alignment processing on the voice instruction and the operation intention parameters, and generate a virtual-real synchronous control instruction.
[0231] Specifically, the depth perception camera is an RGB-D camera, and the depth sensor of the RGB-D camera acquires the three-dimensional hand coordinate data at a sampling frequency not lower than 30Hz.
[0232] Specifically, the gesture recognition unit includes:
[0233] A lightweight convolutional neural network module, which is used to receive the three-dimensional hand coordinate data through the input layer of the lightweight convolutional neural network model;
[0234] An operation intention parameter output module, which is connected to the lightweight convolutional neural network module. The operation intention parameter output module is used to output operation intention parameters including a knob rotation angle and a pressing force through the output layer of the lightweight convolutional neural network model.
[0235] Specifically, the multi-modal instruction spatio-temporal alignment unit includes:
[0236] A time extraction module, which is used to extract the timestamp of the voice command and the generation time of the operation intention parameter;
[0237] A virtual-real synchronization control instruction generation module, which is connected to the time extraction module. The virtual-real synchronization control instruction generation module is used to trigger the generation of the virtual-real synchronization control instruction when the difference between the timestamp and the generation time is less than 50 ms.
[0238] An embodiment of the multi-modal virtual-real fusion interaction system provided in this example is specifically described as follows:
[0239] 1. Data acquisition unit
[0240] Hardware configuration and connection method:
[0241] An Intel RealSense D455 model RGB-D camera is adopted, which is connected to the main control unit through a USB 3.0 interface. It has a built-in binocular infrared sensor and a VGA resolution (640×480) depth sensor, and outputs three-dimensional coordinate data (in the format of a 32-bit floating-point array) of 21 hand key points at a sampling frequency of 30 Hz.
[0242] Configure a ReSpeaker 6-Mic circular microphone array. Six digital MEMS microphones are installed on a PCB substrate with a diameter of 80 mm in a circular layout, and the distance between each microphone is 40 mm. The voice data is transmitted to the main control unit through an I 2 S interface at a sampling rate of 48 kHz.
[0243] Data processing flow:
[0244] The RGB-D camera calls the BlazePalm hand detection model of the mediapipe framework to extract the hand key point coordinates in the depth image in real time, and converts the pixel coordinates into three-dimensional world coordinate system coordinates through the rs2_deproject_pixel_to_point function of the OpenCV library.
[0245] The microphone array suppresses environmental noise through the MVDR beamforming algorithm in the Seeed Studio SDK, and activates the voice endpoint detection module based on the energy threshold method (threshold -35 dBFS) of the G.729 standard, only retaining the valid voice segments with a duration ≥ 200 ms.
[0246] 2. Gesture recognition unit
[0247] Module composition and connection method:
[0248] Lightweight Convolutional Neural Network Module: Based on the MobileNetV3-Small architecture, the input layer is designed with 63 nodes (x / y / z coordinates of 21 key points), and the input data is normalized to the interval [-1, 1]. The normalization formula is:
[0249]
[0250] where xmin and xmax are the minimum and maximum values of the x coordinate in the training dataset, and the same applies to the y and z coordinates.
[0251] The network structure includes a 1×1 convolutional layer (16 channels), a depthwise separable convolutional layer (32 channels), an SE attention module, and a global average pooling layer.
[0252] Operation Intention Parameter Output Module: Outputs 2D parameters through a fully connected layer. The 0th bit is the knob rotation angle (linearly activated, range 0° - 360°), and the 1st bit is the pressing force (Sigmoid activated, range 0 - 1).
[0253] Data Transmission:
[0254] The output parameters are filtered by Kalman filter (observation noise covariance matrix diag([0.1, 0.1, 0.1])) and moving average filter (window length 5 frames), and then transmitted to the virtual-real state tactile mapping unit through the SPI interface.
[0255] 3. Virtual-Real State Tactile Mapping Unit
[0256] Virtual Device Communication and Data Processing:
[0257] Connects to the virtual device real-time status database through the Modbus TCP protocol, and reads register addresses 40001 (current value I, float32 type) and 40005 (fault code, uint8 type).
[0258] Normal Working Condition Mapping: Generates the vibration frequency parameter according to the formula f = 50×I (e.g., when I = 0.8A, f = 40Hz).
[0259] Fault Mode Mapping: When the fault code 0x02 is detected, start the gradient change algorithm, and the vibration frequency linearly increases from 10Hz to 100Hz at a rate of 0.9Hz / ms for 100ms.
[0260] Data Output:
[0261] The tactile feedback parameters (JSON format, including frequency and amplitude fields) are transmitted to the tactile waveform driving unit through shared memory.
[0262] 4. Tactile Waveform Driving Unit
[0263] Hardware and Signal Conversion:
[0264] Adopt the TIDRV2605L haptic drive chip to receive vibration frequency parameters through the I 2 C interface (address 0x5A), and generate a PWM signal with a 25kHz carrier wave.
[0265] Drive the C10-100 model LRA vibration motor array with a 4×2 layout. When the amplitude is fixed at 0.5N, the PWM duty cycle is set to 100%.
[0266] Signal Output:
[0267] Output a bidirectional vibration signal through the H-bridge drive circuit. The layout spacing of the motors is 20mm, covering the palm haptic perception area.
[0268] 5. Multimodal Instruction Spatiotemporal Alignment Unit
[0269] Module Function:
[0270] Time Extraction Module: Synchronize the hardware clocks of the RGB-D camera, microphone array, and host through the NTP protocol (error ≤ ±1ms), and extract the start timestamp of the voice instruction and generate a timestamp for the operation intention parameter.
[0271] Virtual-Reality Synchronization Control Instruction Generation Module: When the timestamp difference ≤ 50ms, trigger instruction fusion: The voice instruction is parsed into a structured JSON instruction (such as "rotate30°") by the Vosk engine (model vosk-model-small-en-us-0.15).
[0272] It is encapsulated with the gesture parameters into a Protobuf format data packet (fields include instruction_id, timestamp, action_type, parameters, crc32), and sent to the virtual device control interface (IP: 192.168.1.200, port 6000) through the UDP protocol at a 10ms cycle.
[0273] Technical Verification Metrics
[0274] 1. Gesture Recognition Latency: The time from data input to parameter output is < 5ms (NVIDIA Jetson Nano platform).
[0275] 2. Haptic Feedback Latency: The fault mode response time is < 5ms, and the linearity error of gradient change is < ±0.5Hz.
[0276] 3. Multimodal Instruction Alignment Accuracy: The timestamp synchronization error is < ±1ms, and the instruction matching correct rate is ≥ 95%.
[0277] Other embodiments of the present application will be readily apparent to those skilled in the art upon considering the specification and practicing the invention disclosed herein. The present application is intended to cover any variations, uses, or adaptations of the present application, which follow the general principles of the present application and include known common knowledge or conventional technical means in the technical field not disclosed in the present application. The specification and examples are only to be considered as exemplary, and the true scope and spirit of the present application are pointed out by the following claims.
[0278] It should be understood that the present application is not limited to the exact structures described above and shown in the drawings, and various modifications and changes can be made without departing from its scope. The scope of the present application is only limited by the appended claims.
Claims
1. A multi-modal virtual-real fusion interaction method, characterized in that Including: Collecting three-dimensional coordinate data of the hand through a depth perception camera and collecting voice data through a microphone array; Inputting the three-dimensional coordinate data of the hand into a gesture recognition model and outputting operation intention parameters through the gesture recognition model; Obtaining real-time status data of the corresponding virtual device according to the operation intention parameters, generating tactile feedback parameters based on the real-time status data of the virtual device, and associating the tactile feedback parameters with the real-time status data of the virtual device through a linear mapping relationship; Generating a vibration waveform according to the tactile feedback parameters and outputting it to a tactile execution unit; Parsing the voice data to obtain a voice command, performing spatio-temporal alignment processing on the voice command and the operation intention parameters, and generating a virtual-real synchronous control command.
2. The method according to claim 1, wherein The step of inputting the three-dimensional coordinate data of the hand into a gesture recognition model and outputting operation intention parameters through the gesture recognition model includes: The gesture recognition model is a lightweight convolutional neural network model, and the three-dimensional coordinate data of the hand is received through the input layer of the lightweight convolutional neural network model; Outputting operation intention parameters including a knob rotation angle and a pressing force through the output layer of the lightweight convolutional neural network model.
3. The method according to claim 1, characterized in that, The linear mapping relationship is f = K×I, where f represents the vibration frequency in the tactile feedback parameters, K is a preset mapping coefficient, and I represents the current value in the real-time status data of the virtual device. The tactile feedback parameters are used to generate the vibration waveform.
4. The method according to claim 1, characterized in that The step of performing spatio-temporal alignment processing on the voice command and the operation intention parameters to generate a virtual-real synchronous control command includes: Extracting the timestamp of the voice command and the generation time of the operation intention parameters; When the difference between the timestamp and the generation time is less than 50 ms, triggering the generation of the virtual-real synchronous control command.
5. The method according to claim 1, characterized in that, After generating the vibration waveform according to the tactile feedback parameters and outputting it to the tactile execution unit, it further includes: The tactile execution unit includes an array vibration motor, and the drive circuit of the array vibration motor receives the vibration waveform and converts it into a pulse width modulation signal.
6. The method according to claim 1, characterized in that, The step of generating tactile feedback parameters based on the real-time status data of the virtual device includes: When the real-time status data of the virtual device contains a fault code, the tactile feedback parameters are generated in a gradient change mode, and the gradient change mode is used to indicate a continuous change process in which the vibration frequency linearly increases from 10 Hz to 100 Hz.
7. A multi-modal virtual-real fusion interaction system, characterized in that, The system is applied to the method according to any one of claims 1-6. The system includes: A data acquisition unit for collecting three-dimensional coordinate data of the hand through a depth perception camera and collecting voice data through a microphone array; A gesture recognition unit connected to the data acquisition unit, and the gesture recognition unit is used to input the three-dimensional coordinate data of the hand into a gesture recognition model and output operation intention parameters through the gesture recognition model; A virtual-real state tactile mapping unit, which is connected to the gesture recognition unit. The virtual-real state tactile mapping unit is used to obtain corresponding real-time state data of the virtual device according to the operation intention parameters, generate tactile feedback parameters based on the real-time state data of the virtual device, and associate the tactile feedback parameters with the real-time state data of the virtual device through a linear mapping relationship; A tactile waveform driving unit, which is connected to the virtual-real state tactile mapping unit. The tactile waveform driving unit is used to generate a vibration waveform according to the tactile feedback parameters and output it to the tactile execution unit; A multi-modal instruction spatio-temporal alignment unit, which is connected to the tactile waveform driving unit. The multi-modal instruction spatio-temporal alignment unit is used to parse the voice data to obtain a voice instruction, perform spatio-temporal alignment processing on the voice instruction and the operation intention parameters, and generate a virtual-real synchronization control instruction.
8. The system according to claim 7, wherein The depth perception camera is an RGB-D camera, and the depth sensor of the RGB-D camera acquires the three-dimensional hand coordinate data at a sampling frequency of not less than 30 Hz.
9. The system according to claim 7, wherein The gesture recognition unit includes: A lightweight convolutional neural network module, which is used to receive the three-dimensional hand coordinate data through the input layer of the lightweight convolutional neural network model; An operation intention parameter output module, which is connected to the lightweight convolutional neural network module. The operation intention parameter output module is used to output operation intention parameters including the knob rotation angle and the pressing force through the output layer of the lightweight convolutional neural network model.
10. The system according to claim 7, characterized in that, The multi-modal instruction spatio-temporal alignment unit includes: A time extraction module, which is used to extract the timestamp of the voice instruction and the generation time of the operation intention parameters; A virtual-real synchronization control instruction generation module, which is connected to the time extraction module. The virtual-real synchronization control instruction generation module is used to trigger the generation of the virtual-real synchronization control instruction when the difference between the timestamp and the generation time is less than 50 ms.
Citation Information
Cited By
Man-machine intelligent interaction method based on multi-modal information fusion
CN121051699A
Self-adaptive man-machine interaction method and system based on multi-modal perception
CN121934723A
Adaptive human-computer interaction method and system based on multi-modal perception
CN121934723B