Multi-sensory interaction synchronization system and method based on dynamic timestamps
By adopting dynamic timestamping technology and the collaborative work of multiple modules in a multi-sensory interaction system, the problems of signal delay differences and insufficient static calibration in traditional systems are solved, and high-precision real-time synchronization and smooth interactive experience of multi-sensory signals are achieved.
Patent Information
- Application Number
- CN202510646444.0
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- Filing Date
- 2025-05-20
- Publication Date
- 2025-06-20
- Estimated Expiration
- 2045-05-20
AI Technical Summary
Traditional multi-sensory interaction systems have large differences in signal transmission and processing delays due to differences in device hardware processing cycles, communication protocols and data processing logic, resulting in a large difference in signal transmission and processing delays, resulting in 'tactile-visual asynchronous'. The existing technology relies on static calibration methods and is unable to adapt to delay fluctuations in dynamic scenarios, resulting in an expansion of synchronization errors and affecting user experience.
Using a multi-sensory interaction synchronization system based on dynamic timestamps, a unified time reference is provided for the device through a high-precision time source. The signal encoding module encodes the multi-sensory signal into a unified format and embeds the dynamic timestamp. The delay compensation module dynamically adjusts the signal transmission timing according to the sensory signal priority. The interface adaptation module supports multiple communication protocols.
It realizes high-precision real-time synchronization of multi-sensory signals, reduces device clock drift and cross-device time deviation, improves interaction fluency and user experience, and reduces system integration costs.
Smart Images

Figure CN120179078A_ABST
Abstract
Description
Technical Field
[0001] The present invention relates to the technical field of human-computer interaction, and particularly to a multi-sensory interaction synchronization system and method based on dynamic timestamps. Background Art
[0002] In immersive interaction scenarios such as the metaverse and virtual reality, the precise synchronization of multi-modal sensory signals is the core foundation for constructing a realistic experience. However, traditional multi-sensory interaction systems face severe challenges: due to significant differences in hardware processing cycles, communication protocols (Bluetooth / Wi-Fi / wired), and data processing logics among different types of devices (such as XR headsets, haptic gloves, and odor generators), the transmission and processing delays of visual, auditory, tactile, and olfactory signals can differ by more than 50 ms. For example, the Bluetooth transmission delay of haptic gloves is usually 20 - 30 ms, while the wired transmission delay of XR headsets is only 5 - 10 ms. This delay difference causes "tactile-visual asynchrony" - when a user touches an object in a virtual scene, the tactile feedback lags significantly behind the visual image, creating a strong sense of sensory disconnection.
[0003] More critically, existing technologies rely on static calibration methods (presetting fixed delay compensation values) and cannot adapt to the real-time fluctuations in delays (with a fluctuation range of up to ±20 ms) caused by network fluctuations, device load changes, or multi-person interactions in dynamic scenarios. When the scene complexity suddenly increases (such as rapid virtual environment switching or multi-device concurrent interaction), the static compensation mechanism fails, and the synchronization error further expands, resulting in discomfort such as dizziness and operation disconnection for users, severely restricting the popularization of immersive experiences. In addition, the closed interfaces and inconsistent protocols of device manufacturers lead to extremely high costs for multi-modal data fusion, further exacerbating the system integration difficulty. Summary of the Invention
[0004] The purpose of the present invention is to provide a multi-sensory interaction synchronization system and method based on dynamic timestamps to solve the problems raised in the above background art.
[0005] To achieve the above purpose, the present invention provides the following technical solution: A multi-sensory interaction synchronization system based on dynamic timestamps, comprising: A time reference module for providing a unified time reference for all devices through a high-precision time source, so that the clock synchronization error of the devices does not exceed 1 μs, and the high-precision time source includes an atomic clock or a satellite time service device; A signal encoding module for encoding visual, auditory, tactile, and olfactory signals into data frames in a unified format and embedding dynamic timestamps in the data frames, where the dynamic timestamps are used to indicate the decoding time of the signals at the terminal device; A delay compensation module, which is used to dynamically adjust the signal transmission timing based on a preset sensory signal priority, where signals with a higher priority are preferentially allocated transmission resources; An interface adaptation module, which is used to support multiple communication protocols and realize the rapid access and data interaction of devices from different manufacturers.
[0006] Preferably, the calculation formula of the dynamic timestamp is: , where is the reference time obtained through the time reference module, accurate to the nanosecond level; is the transmission delay between devices predicted based on a long short-term memory network, with the unit of millisecond; is the scene complexity coefficient, which is obtained by normalizing parameters related to the scene, and the value range is from 0 to 1; is the dynamic compensation factor, and its calculation formula is , where is the empirical coefficient, is a very small constant to prevent division by zero.
[0007] Preferably, the preset order of sensory signal priorities is: The priority of tactile signals is higher than that of visual signals, the priority of visual signals is higher than that of olfactory signals, and the priority of olfactory signals is higher than that of auditory signals; Signals with different priorities are transmitted using different transmission protocols. In the signal buffer of the terminal device, signals with a higher priority have the priority decoding permission.
[0008] Preferably, the time reference module synchronizes the clocks of all devices through the Network Time Synchronization Protocol, and the Network Time Synchronization Protocol includes the Network Time Protocol or the Precision Time Protocol; Among them, in an environment lacking a high-precision time source, the time reference module can achieve distributed time synchronization through a network composed of edge computing nodes and through the Reference Broadcast Synchronization Protocol, and the time synchronization error does not exceed 5 μs.
[0009] Preferably, the delay compensation module includes: A prediction unit, which predicts the transmission delay through a long short-term memory network model integrating an attention mechanism, and the calculation formula is , where is the historical hidden state, is the attention weight matrix at each time step, and this matrix dynamically adjusts the weight coefficient according to the signal type; or, model training is carried out through federated learning technology, and each device calculates the model gradient locally and uploads it to the server for aggregation and update to avoid the transmission of original delay data; A scheduling unit, which adjusts the signal priority in real time through a fuzzy logic algorithm, combining the scene type and the sensory signal fusion threshold. The priority adjustment formula is , where is the signal type, is the scene type, is the sensory signal fusion threshold; or, by accessing the electroencephalogram signal feedback path, the signal compensation priority is dynamically adjusted according to the electroencephalogram signal monitoring results.
[0010] Preferably, the interface adaptation module includes: A protocol conversion sub-module for implementing the conversion of multiple communication protocols, which include but are not limited to OpenXR, DCPv2.0, and MQTT.
[0011] The present invention also provides a multi-sensory interaction synchronization method based on dynamic timestamps, including the following steps: Synchronize the clocks of all devices through a high-precision time source to generate a unified reference time; Encode visual, auditory, tactile, and olfactory signals into data frames in a unified format, predict the transmission delay between devices based on a long short-term memory network, calculate a dynamic compensation factor in combination with the scene complexity coefficient, generate a dynamic timestamp and embed it into the data frame; Dynamically adjust the signal transmission timing according to the preset sensory signal priority, and preferentially allocate transmission resources to signals with high priority; The terminal device decodes the signal according to the dynamic timestamp and feeds the actual delay data back to the delay compensation module to form an accuracy synchronization closed loop.
[0012] The present invention also provides an electronic device, which is a physical device and includes: A processor and a memory, and the memory is communicatively connected to the processor; The memory is used to store at least one executable instruction executed by the processor, and the processor is used to execute the executable instruction to implement the multi-sensory interaction synchronization method based on dynamic timestamps as described above.
[0013] The present invention also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it implements the multi-sensory interaction synchronization method based on dynamic timestamps as described above.
[0014] Compared with the prior art, the beneficial effects of the present invention are: The present invention constructs a globally unified time coordinate system for multi-modal signals by adopting an atomic clock or satellite timekeeping combined with the Network Time Protocol, eliminates the clock drift of devices, ensures that the time deviation across devices is at an extremely low level, realizes high-precision real-time synchronization by uniformly encoding multi-sensory signals and embedding dynamic timestamps, meets the needs of human sensory fusion, predicts the transmission delay through an LSTM neural network, provides a reliable basis for signal synchronization, ensures the transmission of key signals and improves the interaction fluency by setting the priority of sensory signals and reasonably allocating resources, realizes device compatibility and reduces the integration cost by supporting multiple communication protocols, further optimizes the prediction accuracy and protects data privacy by integrating attention mechanisms, federated learning, etc., and enhances the system adaptability and reliability by combining dynamic priority scheduling and semantic web technology, comprehensively improving the interaction experience. BRIEF DESCRIPTION OF THE DRAWINGS
[0015] Figure 1 FIG. is a schematic structural diagram of a multi-sensory interaction synchronization system based on dynamic timestamps provided by an embodiment of the present invention; Figure 2 FIG. is an exemplary flowchart of a multi-sensory interaction synchronization method based on dynamic timestamps provided by an embodiment of the present invention; Figure 3 FIG. is a schematic structural diagram of an electronic device provided by an embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS
[0016] The technical solutions in the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings in the embodiments of the present invention. Obviously, the described embodiments are only a part of the embodiments of the present invention, rather than all the embodiments. All other embodiments obtained by those of ordinary skill in the art based on the embodiments of the present invention without creative efforts shall fall within the protection scope of the present invention.
[0017] Please refer to Figure 1 , the present invention provides a multi-sensory interaction synchronization system based on dynamic timestamps, including: A time reference module 11, configured to provide a unified time reference for all devices through a high-precision time source, so that the clock synchronization error of the devices does not exceed 1 μs, and the high-precision time source includes an atomic clock or a satellite timekeeping device; A signal encoding module 12, configured to encode visual, auditory, tactile, and olfactory signals into data frames in a unified format, and embed a dynamic timestamp in the data frame, where the dynamic timestamp is used to indicate the decoding time of the signal at the terminal device; A delay compensation module 13, configured to dynamically adjust the signal transmission timing based on a preset priority of sensory signals, where signals with higher priority are preferentially allocated transmission resources; The interface adaptation module 14 is used to support multiple communication protocols and enable the rapid access and data interaction of devices from different manufacturers.
[0018] In an optional embodiment, the time reference module 11 synchronizes the clocks of all devices through the Network Time Synchronization Protocol, which includes the Network Time Protocol or the Precision Time Protocol. Among them, in an environment lacking a high-precision time source, the time reference module 11 can achieve distributed time synchronization through a network composed of edge computing nodes using the Reference Broadcast Synchronization Protocol, with a time synchronization error not exceeding 5 μs.
[0019] Specifically, the time reference module 11 uses an atomic clock (such as a rubidium atomic clock with a time accuracy of ≤1 μs) or satellite time service (GPS / Beidou) as the reference time source to synchronize the clocks of all devices through the Network Time Protocol (NTP) or the Precision Time Protocol (PTP), ensuring that the time deviation across devices is ≤1 μs.
[0020] The system constructs a reference clock system using high-precision time sources, and the core solution includes two types of high-precision time service technologies. First, a rubidium atomic clock is selected as the local time reference, which realizes timing based on the energy level transition characteristics of rubidium atoms, and the time accuracy can reach the sub-microsecond level (≤1 μs), suitable for closed environments with low network dependence. Second, access to the Global Navigation Satellite System (such as GPS, Beidou), and realize wide-area time synchronization through the nanosecond-level accurate time signals broadcast by satellites. At the clock synchronization protocol level, it supports the hybrid deployment of the Network Time Protocol (NTP) and the Precision Time Protocol (PTP): The NTP protocol realizes millisecond-level fast synchronization through UDP packets, suitable for devices with lower real-time requirements; the PTP protocol is based on the IEEE1588 standard, uses hardware timestamp technology, and through the master-slave clock architecture and path delay compensation algorithm, can strictly control the time deviation across devices within 1 μs, ensuring the highly unified clocks of multi-sensory devices.
[0021] This time synchronization mechanism constructs a unified time coordinate system covering the entire system, fundamentally solving the problem of time inconsistency in the process of multi-modal signal acquisition and processing. The cumulative error of the device's local clock caused by factors such as crystal oscillator temperature drift and aging will cause misalignment of multi-modal data on the time axis (such as the delay asynchronization of audio and video signals). The global time reference can control the time deviation of each device below the sensory threshold through real-time calibration and dynamic compensation, ensuring that multi-modal signals such as vision, hearing, and touch are fused and processed in the same time dimension, providing a basic guarantee for the accurate analysis and collaborative response of subsequent multi-sensory interaction data.
[0022] In an offline environment without atomic clocks or satellite signals (such as underground industrial scenarios), distributed time synchronization is achieved through the networking of edge computing nodes. By adopting an improved RBS (Reference Broadcast Synchronization) protocol, the edge nodes calibrate time through multi-hop broadcasting, with an error ≤ 5 μs; it forms a'main-backup collaboration' with the atomic clock synchronization mode to enhance the robustness of the system.
[0023] In an alternative embodiment, the calculation formula for the dynamic timestamp is: , where is the reference time obtained through the time reference module 11, accurate to the nanosecond level; is the transmission delay between devices predicted based on a long short-term memory network, with the unit of millisecond; is the scene complexity coefficient, obtained by normalizing parameters related to the scene, with a value range of 0 to 1; is the dynamic compensation factor.
[0024] Specifically, in a multi-sensory interaction system, the differences between different modality signals are extremely large. To achieve precise synchronization, signals such as vision (RGB frames), audition (audio PCM data), touch (pressure / vibration signals), and olfaction (gas release instructions) need to be unified and converted into a standardized data frame format. This format mainly consists of the following two parts: Payload: The payload part carries the original sensory data, which are the core information carriers for system interaction. For visual signals, it appears as a pixel matrix, recording the color and brightness information of each frame; auditory signals exist in the form of audio sampling points, and each sampling point corresponds to the amplitude of the sound waveform at a specific moment; touch signals include the pressure values collected by pressure sensors or the drive parameters of vibration motors; while olfactory signals are reflected as the instruction parameters for controlling gas release devices, such as the type, concentration, and duration of the released gas.
[0025] Dynamic timestamp: The dynamic timestamp module is the key to achieving precise synchronization of multi-modal signals, and its core lies in embedding the decoding time expected by the signal for the terminal device.
[0026] In the above calculation formula for the dynamic timestamp: The UTC timestamp synchronized by an atomic clock, accurate to the nanosecond level, provides an absolute time reference for the entire system to ensure the consistency of the time scales between different devices; A prediction model is constructed based on the Long Short-Term Memory Network (LSTM). By learning historical transmission data (such as features like packet size, network bandwidth changes, device distance, etc.), the model dynamically predicts the transmission delay of the current signal between different devices, with the unit being milliseconds (ms). For example, when there is network congestion or the device distance is far, the value will increase accordingly; is used to dynamically adjust the synchronization accuracy according to the scene complexity, and its calculation formula is: where is an empirical coefficient, with a default value of 0.8, which can be fine-tuned according to the actual application scenario; is a very small constant to prevent division by zero, with a value of ; is the scene complexity coefficient, with a value range between 0 and 1, which is obtained by normalizing key parameters such as frame rate, number of devices, network load, etc. For example, in a high-frame-rate game scenario or a complex environment where a large number of devices work together, the value approaches 1, and the system will enhance the synchronization accuracy compensation to ensure the synchronization quality of multi-modal signals.
[0027] In an optional embodiment, the preset priority order of the sensory signals is as follows: The priority of the tactile signal is higher than that of the visual signal, the priority of the visual signal is higher than that of the olfactory signal, and the priority of the olfactory signal is higher than that of the auditory signal; Signals with different priorities are transmitted using different transmission protocols. In the signal buffer of the terminal device, signals with higher priority have the priority decoding permission.
[0028] Specifically, a hierarchical scheduling strategy is constructed based on the physiological characteristics of human senses, and the specific priority settings are as follows: Tactile signal: As the fastest perception channel of the human body, the nerve conduction speed is about 100m / s, and the human perception threshold is strictly controlled at ≤10ms. To meet the ultra-low latency requirement, it is transmitted using the UDP protocol combined with FEC (Forward Error Correction) coding, and a dedicated priority queue is configured in the terminal device to ensure that tactile data is decoded and rendered first.
[0029] Visual signal: The visual nerve conduction speed is about 40m / s, and the perception threshold is ≤50ms. The system adopts a dynamic bit rate adjustment strategy, and automatically switches between the H.264 / H.265 coding standards in combination with the network condition to balance the transmission delay while ensuring the image quality.
[0030] Olfactory signal: The olfactory conduction speed is the slowest, about 1 m / s, and the perception threshold is ≤500 ms. Given its high requirement for transmission reliability, the TCP protocol is used for transmission, and a circular buffer is set at the receiving end to achieve smooth playback through a double-buffering mechanism.
[0031] Auditory signal: Although the sound wave conduction speed reaches 340 m / s, the brain has a relatively high tolerance for audio synchronization. The system uses adaptive resampling technology to dynamically adjust the playback speed by calculating the difference in PTS (Presentation Time Stamp) between the audio and video, ensuring the final synchronization of multi-modal signals.
[0032] Based on the characteristic differences of the above various sensory signals, the system adopts a priority scheduling strategy: high-priority signals (such as tactile signals) are preferentially allocated network bandwidth and have the priority to decode in the buffer of the terminal device to ensure the immediacy of key sensory feedback; while low-priority signals (such as olfactory signals) reasonably utilize network resources on the basis of ensuring data reliability to achieve the efficient coordination and precise synchronization of multi-sensory signals.
[0033] In an optional embodiment, the delay compensation module 13 includes: A prediction unit, which predicts the transmission delay through a long short-term memory network model that integrates the attention mechanism. The calculation formula is , where is the historical hidden state, is the attention weight matrix at each time step. This matrix dynamically adjusts the weight coefficient according to the signal type. Specifically, the system assigns corresponding weight coefficients according to the characteristics of different sensory signals and their contribution degrees to the synchronization accuracy: for tactile signals with high time-delay sensitivity, the weight coefficient is set to 1.5 to enhance its priority in synchronization calculation; for auditory signals, considering its relatively low requirement for time accuracy, the weight coefficient is set to 0.8. Through this dynamic adjustment mechanism, the model can adaptively balance the fusion strategy of multi-modal data. After actual measurement and verification, the prediction error can be stably controlled within 1.5 ms, significantly improving the synchronization accuracy of the system.
[0034] Or, the model is trained through federated learning technology. Each device calculates the model gradient locally and uploads it to the server for aggregation and update to avoid the transmission of raw delay data.
[0035] Under the model training mechanism through federated learning technology, after each terminal device completes the model gradient calculation locally, only the encrypted gradient parameters are uploaded to the central server for aggregation and update, avoiding the cross-device transmission of original latency data throughout the process. This training mode of "data does not move while the model moves" not only ensures data privacy and security but also makes full use of the computing power resources of edge devices, effectively solving the problem of data leakage risk existing in traditional centralized training and achieving the dual goals of privacy protection and model performance optimization.
[0036] A scheduling unit, which adjusts the signal priority in real time through a fuzzy logic algorithm, combining the scene type and the sensory signal fusion threshold. The priority adjustment formula is , where is the signal type, is the scene type, is the sensory signal fusion threshold; or, by accessing the EEG feedback path, the signal compensation priority is dynamically adjusted according to the EEG monitoring results.
[0037] Specifically, in the multi-modal interaction scenario, the dynamic adjustment mechanism of signal priority can be precisely optimized according to the specific application scenario. For example, in the space movie experience, the system recognizes the immersive viewing demand based on the scene type parameters, and temporarily increases the priority weight coefficient of the olfactory signal from the default 0.3 to 0.8 through the dynamic adjustment algorithm, forming a synergy with the 0.85 weight of the visual signal, so that when the audience feels the interstellar explosion scene, the smoke and burnt smell can be triggered synchronously, significantly enhancing the immersive experience of the plot.
[0038] In addition, the system supports the feedback closed-loop access of electroencephalogram (EEG) signals, and the amplitude and latency of the P300 EEG component are monitored in real time through a high-precision head-mounted device. When the feature of decreased immersion is detected (such as the P300 wave amplitude is lower than 20% of the baseline value and lasts for more than 3 seconds), the adaptive compensation mechanism is triggered: first, the Bayesian inference model is used to judge the signal short board of the current scene. If the auditory signal is missing, the priority of the audio signal is increased by 30%, and the fusion thresholds of other modalities are adjusted synchronously, forming an intelligent control closed-loop of "physiological signal acquisition - state evaluation - parameter optimization - multi-modal compensation" to achieve the dynamic optimization of immersion with millisecond-level response.
[0039] In an alternative embodiment, the interface adaptation module 14 includes: A protocol conversion sub-module for implementing the conversion of multiple communication protocols, and the communication protocols include but are not limited to OpenXR, DCPv2.0, and MQTT; Device description parsing sub-module. The device description parsing sub-module adopts semantic web technology, uses ontology language to describe device attributes, and parses device attributes through a semantic reasoning engine to achieve automatic matching of device capabilities and reasonable allocation of compensation resources; or, enhances the time synchronization ability of the system in special environments through the distributed time synchronization function between edge computing nodes.
[0040] Optionally, in the device description parsing sub-module, the OWL ontology language is used to construct a device capability description framework, and the semantic expression of device features is realized through the structured definition of classes, properties, and instances. Taking a tactile feedback device as an example, by defining the subclass HapticGlove of the Device class and setting the forceFeedbackAccuracy property to 0.1mm; for an odor generating device, it is accurately characterized through the OdorGenerator class and the responseTime property (with a value of 100ms). Combining semantic reasoning engines such as Jena, based on the device capability description, a semantic matching algorithm is automatically executed, and a dynamic compensation model is established for the capability differences between heterogeneous devices. When the system detects the access of a device with high latency but high precision, the reasoning engine will dynamically allocate compensation resources such as the GPU acceleration module and low-latency network channels in the cloud computing resource pool to the target device according to preset rules to ensure the real-time performance and accuracy of multi-modal data synchronization processing.
[0041] Or, in an edge computing scenario without atomic clock deployment, the ReferenceBroadcastSynchronization (RBS) protocol is deeply optimized, and an adaptive weight adjustment mechanism and a two-way timestamp interaction strategy are introduced. By constructing an edge node topology awareness model, multi-hop path analysis is performed on the nodes participating in synchronization, and nodes with excellent link quality and fewer hops are preferentially selected as reference nodes. During the synchronization process, each node periodically broadcasts a beacon packet containing the local timestamp, and the receiving node fuses the timestamps received multiple times through the Kalman filtering algorithm to effectively suppress clock drift. Through actual measurement verification, this solution can stably control the time synchronization error within 5μs in a typical industrial Internet of Things scenario. At the same time, the system designs an atomic clock mode as the main time source, which automatically switches when an external high-precision time source is available, forming a primary-backup redundant architecture with the distributed synchronization mechanism to ensure the time synchronization reliability of the system in a complex network environment.
[0042] In this embodiment, the present invention constructs a globally unified time coordinate system for multi-modal signals by using an atomic clock or satellite timekeeping in combination with the Network Time Protocol, eliminates clock drift of devices, ensures that the time deviation across devices is at an extremely low level, realizes high-precision real-time synchronization by uniformly encoding multi-sensory signals and embedding dynamic timestamps, meets the needs of human sensory fusion, predicts transmission delays through an LSTM neural network to provide a reliable basis for signal synchronization, ensures the transmission of key signals and improves interaction fluency by setting the priority of sensory signals and reasonably allocating resources, realizes device compatibility and reduces integration costs by supporting multiple communication protocols, further optimizes prediction accuracy and protects data privacy by integrating attention mechanisms, federated learning, etc., and enhances system adaptability and reliability by combining dynamic priority scheduling and semantic web technology, comprehensively improving the interaction experience.
[0043] Based on the above embodiment, as Figure 2 shown, the present invention also provides a multi-sensory interaction synchronization method based on dynamic timestamps, including the following steps: Step 100, synchronize the clocks of all devices through a high-precision time source to generate a unified reference time.
[0044] Specifically, deploy a high-precision atomic clock or a global satellite navigation system (such as GPS, Beidou) as the time source, and use the Network Time Protocol (NTP) or the Precision Time Protocol (PTP) to synchronize and calibrate the clocks of all devices participating in the interaction. Specifically, the master device periodically sends time synchronization packets to the slave devices. After receiving the packets, the slave devices calculate the time offset and compensate it. After multiple iterations, a unified reference time is generated , ensuring that the time error within the system is controlled within the microsecond level.
[0045] Step 200, encode visual, auditory, tactile, and olfactory signals into data frames in a unified format, predict the transmission delay between devices based on a long short-term memory network, calculate a dynamic compensation factor in combination with the scene complexity coefficient, and generate a dynamic timestamp and embed it into the data frame.
[0046] Specifically, adopt cross-modal data fusion technology to encode signals such as vision (RGB video stream, depth image), audition (multi-channel audio), touch (force feedback data), and olfaction (gas concentration parameter encoding) into data frames in a unified format that conforms to the ISO / IEC23008-11 standard; Construct a transmission delay prediction model based on a long short-term memory network (LSTM), and predict the transmission delay Dpredict between devices by analyzing historical transmission data (including features such as network bandwidth and device load); Introduce the scene complexity coefficient , this coefficient is obtained by weighted calculation of parameters such as the number of dynamic objects in the scene and the density of interaction devices, and its value range is [0, 1]; According to the formula generate a dynamic timestamp, where is an adjustable compensation coefficient, and its value is determined by experimental optimization. Finally, the timestamp is embedded in the header of the data frame.
[0047] Step 300, dynamically adjust the signal transmission timing according to the preset sensory signal priorities, and preferentially allocate transmission resources to signals with higher priorities.
[0048] Specifically, establish a sensory signal priority matrix, divide the signals into emergency categories (such as tactile feedback, warning audio), important categories (such as key action visual signals), and ordinary categories (such as environmental background sounds, non-critical visual elements). Based on the dynamic programming algorithm, combined with the current network state and device resource occupancy, adopt a priority queue scheduling strategy for signals with higher priorities, and preferentially allocate transmission resources by reserving bandwidth, compressing low-priority data, etc., to ensure that the transmission delay of key signals is lower than the threshold.
[0049] Step 400, the terminal device decodes the signal according to the dynamic timestamp and feeds back the actual delay data to the delay compensation module to form a precision synchronization closed loop.
[0050] Specifically, after receiving the data frame, the terminal device extracts the dynamic timestamp , calculates the actual transmission delay in combination with the local time, takes the difference between the actual transmission delay and the predicted delay as feedback data, updates the parameters of the LSTM prediction model through the backpropagation algorithm, and dynamically adjusts the compensation factor , forms a closed-loop synchronization mechanism of "prediction - transmission - feedback - optimization", and continuously improves the synchronization accuracy of multi-sensory signals.
[0051] In this embodiment, the present invention constructs a globally unified time coordinate system for multi-modal signals by using an atomic clock or satellite timekeeping combined with the Network Time Protocol, eliminates device clock drift, ensures that the cross-device time deviation is at an extremely low level, realizes high-precision real-time synchronization by uniformly encoding multi-sensory signals and embedding dynamic timestamps, meets the needs of human sensory fusion, predicts the transmission delay through the LSTM neural network to provide a reliable basis for signal synchronization, ensures the transmission of key signals and improves the interaction fluency by setting sensory signal priorities and reasonably allocating resources, realizes device compatibility and reduces the integration cost by supporting multiple communication protocols, further optimizes the prediction accuracy and protects data privacy by integrating attention mechanisms, federated learning, etc., and enhances the system adaptability and reliability by combining dynamic priority scheduling and semantic web technology, comprehensively improving the interaction experience.
[0052] Further, the cache data processing device for intelligently analyzing the operation habits of investment users can run the above-mentioned multi-sensory interaction synchronization system and method based on dynamic timestamps. For the specific implementation, reference can be made to the method embodiments, which will not be elaborated here.
[0053] Based on the above embodiments, as Figure 3 shown, the present invention also provides an electronic device, which includes: At least one processor 22, at least one memory 21, a communication interface 23, and a communication bus 24, and the processor 22 is communicatively connected to the memory 21; In this embodiment, the memory 21 can be implemented in any suitable manner. For example, the memory 21 can be a read-only memory, a mechanical hard disk, a solid-state drive, or a USB flash drive, etc.; the memory 21 is used to store executable instructions executed by at least one of the processors. In this embodiment, the processor 22 can be implemented in any suitable manner. For example, the processor 22 can take the form of, for example, a microprocessor or a processor, and a computer-readable medium storing computer-readable program code (such as software or firmware) executable by the (micro)processor, logic gates, switches, application-specific integrated circuits (ASICs), programmable logic controllers, and embedded microcontrollers, etc.; the processor is used to execute the executable instructions to implement the multi-sensory interaction synchronization method based on dynamic timestamps as described above.
[0054] Based on the above embodiments, the present invention also provides a computer-readable storage medium, in which a computer program is stored, and when the computer program is executed by a processor, it implements the multi-sensory interaction synchronization method based on dynamic timestamps as described above.
[0055] Those of ordinary skill in the art can realize that the modules and method steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods for each specific application to implement the described functions, but such implementation should not be considered to exceed the scope of the present invention.
[0056] Those skilled in the art can clearly understand that for the convenience and conciseness of description, the specific working processes of the above-described devices, equipment, and modules can refer to the corresponding processes in the foregoing method embodiments, which will not be elaborated here.
[0057] In the several embodiments provided in the present application, it should be understood that the disclosed devices, apparatuses and methods can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the modules is only a logical function division. There may be other division methods in actual implementation, such as multiple modules or units can be combined or integrated into another device, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be an indirect coupling or communication connection through some interfaces, devices or equipment, which can be electrical, mechanical or other forms.
[0058] The modules described as separate components may or may not be physically separated, and the components shown as modules may or may not be physical units, that is, they may be located in one place or distributed on multiple network units. Some or all of the units may be selected according to actual needs to achieve the purpose of the solution of this embodiment.
[0059] In addition, each functional module in each embodiment of the present invention may be integrated into one processing module, or each module may exist physically separately, or two or more modules may be integrated into one module.
[0060] If the functions are implemented in the form of software function modules and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present invention, or the part that contributes to the prior art or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for a computer device (which can be a personal computer, a server, or a network device, etc.) to perform all or part of the steps of the methods described in the various embodiments of the present invention. The aforementioned storage media include: U disk, mobile hard disk, read-only storage server, random access storage server, disk or optical disk, and other media that can store program instructions.
[0061] In addition, it should be noted that the combination of the various technical features in this case is not limited to the combination described in the claims of this case or the combination described in the specific embodiments. All technical features described in this case can be freely combined or combined in any way unless there is a contradiction between them.
[0062] It should be noted that the above examples are only specific embodiments of the present invention, and the present invention is obviously not limited to the above examples, and there are many similar variations. All variations directly derived or associated from the contents disclosed by the technicians in this field should fall within the protection scope of the present invention.
[0063] The above are only the preferred embodiments of the present invention and are not intended to limit the protection scope of the present invention. Any modifications, equivalent replacements, improvements, etc. made within the spirit and principle of the present invention shall be included within the protection scope of the present invention.
Claims
1. A multi-sensory interactive synchronization system based on dynamic timestamps, characterized in that: include: A time reference module, used to provide a unified time reference for all devices through a high-precision time source, so that the clock synchronization error of the devices does not exceed 1 μs. The high-precision time source includes an atomic clock or a satellite timing device; A signal encoding module, used to encode visual, auditory, tactile and olfactory signals into a data frame of a unified format, and embed a dynamic timestamp in the data frame, wherein the dynamic timestamp is used to indicate the decoding time of the signal in the terminal device; A delay compensation module is used to dynamically adjust the signal transmission timing based on the preset sensory signal priority, wherein the signal with a high priority is preferentially allocated transmission resources; The interface adapter module is used to support multiple communication protocols and realize fast access and data interaction of devices from different manufacturers.
2. The multi-sensory interactive synchronization system based on dynamic timestamp according to claim 1 is characterized in that: The calculation formula of the dynamic timestamp is: ,in, The reference time obtained by the time reference module is accurate to nanosecond level; is the inter-device transmission delay predicted based on the LSTM network, in milliseconds; is the scene complexity coefficient, which is obtained by normalizing the parameters related to the scene and ranges from 0 to 1; is the dynamic compensation factor, and its calculation formula is ,in is the empirical coefficient, A very small constant to prevent division by zero.
3. The multi-sensory interactive synchronization system based on dynamic timestamp according to claim 1 is characterized in that: The preset sensory signal priority order is: Tactile signals have a higher priority than visual signals, visual signals have a higher priority than olfactory signals, and olfactory signals have a higher priority than auditory signals; Signals of different priorities are transmitted using different transmission protocols. In the signal buffer of the terminal device, signals with higher priorities have priority decoding rights.
4. The multi-sensory interactive synchronization system based on dynamic timestamp according to claim 1 is characterized in that: The time reference module realizes synchronization of all device clocks through a network time synchronization protocol, wherein the network time synchronization protocol includes a network time protocol or a precision time protocol; In an environment lacking a high-precision time source, the time reference module can achieve distributed time synchronization through a network composed of edge computing nodes by referring to the broadcast synchronization protocol, and the time synchronization error does not exceed 5μs.
5. The multi-sensory interactive synchronization system based on dynamic timestamp according to claim 2 is characterized in that: The delay compensation module comprises: The prediction unit predicts the transmission delay by integrating the long short-term memory network model with the attention mechanism. The calculation formula is: ,in, For the historical hidden state, is the attention weight matrix for each time step, which dynamically adjusts the weight coefficient according to the signal type; or, the model is trained through federated learning technology, where each device calculates the model gradient locally and uploads it to the server for aggregate update to avoid the transmission of original delayed data; The scheduling unit uses a fuzzy logic algorithm to adjust the signal priority in real time in combination with the scene type and the sensory signal fusion threshold. The priority adjustment formula is: ,in, is the signal type, is the scene type, is the sensory signal fusion threshold; or, by accessing the EEG signal feedback path, the signal compensation priority is dynamically adjusted according to the EEG signal monitoring results.
6. The multi-sensory interactive synchronization system based on dynamic timestamp according to claim 1, characterized in that: The interface adapter module comprises: A protocol conversion submodule, used to implement conversion of multiple communication protocols, including but not limited to OpenXR, DCPv2.0 and MQTT; The device description parsing submodule adopts semantic web technology, uses ontology language to describe device attributes, and parses device attributes through a semantic reasoning engine to achieve automatic matching of device capabilities and reasonable allocation of compensation resources; or, through the distributed time synchronization function between edge computing nodes, enhance the system's time synchronization capability in special environments.
7. A multi-sensory interactive synchronization method based on dynamic timestamp, characterized in that: The following steps are involved: Synchronize the clocks of all devices through a high-precision time source to generate a unified reference time; Encode visual, auditory, tactile, and olfactory signals into a unified format of data frames, predict the transmission delay between devices based on the long short-term memory network, calculate the dynamic compensation factor based on the scene complexity coefficient, generate a dynamic timestamp and embed it into the data frame; According to the preset sensory signal priority, the signal transmission timing is dynamically adjusted, and transmission resources are preferentially allocated to signals with high priority; The terminal device decodes the signal according to the dynamic timestamp and feeds back the actual delay data to the delay compensation module to form a precision synchronization closed loop.
8. An electronic device, characterized in that: The electronic device comprises: A processor and a memory, wherein the memory is communicatively connected to the processor; The memory is used to store at least one executable instruction executed by the processor, and the processor is used to execute the executable instruction to implement the multi-sensory interaction synchronization method based on dynamic timestamp as described in any one of claims 1 to 6.
9. A computer-readable storage medium, characterized in that: The computer-readable storage medium stores a computer program, and when the computer program is executed by a processor, the multi-sensory interaction synchronization method based on dynamic timestamps as described in any one of claims 1 to 6 is implemented.
Citation Information
Patent Citations
Video timestamp event identification and reasoning method based on multi-modal large model
CN119723431A
Audio associations for interactive media event triggering
US20210389868A1
Method and device for timestamping and synchronization with high-accuracy timestamps in low-power sensor systems
US20220414036A1
Timestamp generating method, device and system
WO2014173267A1
Cited By
Multi-modal data high-precision time synchronization acquisition system and method based on Syntaclos
CN120508185A
Multi-protocol conversion control method and controller for copper liquid air supply drying treatment system
CN120567948A
Multi-protocol conversion control method and controller for copper liquid air replenishment and drying processing system
CN120567948B
Tactile feedback interaction method for outer surface of vacuum cup
CN120762536A
All-media all-signal integrated management platform
CN120935094A