Visual identification and attribute virtualization-based interactive system and collaborative configuration
By combining visual recognition and attribute virtualization interaction system with multimodal perception data packets and dynamic mapping technology, the problems of non-universal hardware and static software in existing interaction systems are solved, realizing cross-scene adaptability and user experience consistency, and improving the scalability and immersion of XR applications.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- PIMAX TECH (SHANGHAI) CO LTD
- Filing Date
- 2025-12-31
- Publication Date
- 2026-04-17
AI Technical Summary
Existing interactive systems suffer from non-universal hardware design, static software configuration, and poor cross-scene scalability, resulting in poor system scalability, difficulty in cross-scene reuse, and difficulty in supporting large-scale, highly immersive, and content-varying XR application scenarios.
By using a visual recognition and attribute virtualization interaction system, visual data, inertial motion data, and grip pressure data of physical props are collected simultaneously to form a multimodal perception data package. Semantic type recognition and operability status analysis are performed, and dynamic mapping is performed by combining user status and virtual environment status to generate an interaction configuration parameter set. The hardware interface functional logic is dynamically virtualized through a general interaction module to monitor and optimize multimodal consistency and achieve seamless switching.
It improves the accuracy of semantic and operability analysis of physical props in complex environments, realizes the system's strong scalability, flexible adaptation and unified user experience of software and hardware collaborative interaction, reduces the complexity of development and maintenance, and ensures the continuity and smoothness of highly immersive interactive experience.
Smart Images

Figure CN121879569A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the fields of human-computer interaction and virtual reality technology, and in particular to an interactive system and collaborative configuration based on visual recognition and attribute virtualization. Background Technology
[0002] In the field of motion-sensing interaction, virtual reality (VR) and extended reality (XR) interaction technologies are developing towards higher immersion and more natural interaction. From early dedicated props based on single sensors, they have evolved into modular interactive devices integrating multiple sensors such as inertial measurement and optical positioning. Physical props, as key carriers connecting users and virtual content, are gradually improving in terms of intelligence and versatility. Related research focuses on improving motion capture accuracy, reducing interaction latency, and enriching haptic feedback formats, laying the foundation for immersive experiences.
[0003] Existing interactive systems generally suffer from architectural limitations. At the hardware level, props with different functions typically employ independent, non-universal, customized electronic designs, leading to high R&D and maintenance costs. At the software level, prop recognition largely relies on short-range wireless pairing or manual scanning, failing to achieve seamless, real-time binding; their functional logic (such as button mapping and feedback modes) is usually fixed at the factory or subject to limited static configuration at the application layer, lacking the ability to dynamically adapt to real-time user status, virtual environment, and interaction intent. This results in poor system scalability, difficulty in cross-scenario reuse, and an inability to support large-scale, highly immersive, and content-varying next-generation XR application scenarios. Summary of the Invention
[0004] In view of the aforementioned existing problems, the present invention is proposed.
[0005] Therefore, this invention provides a visual recognition and attribute virtualization-based interactive system and collaborative configuration to solve the problems of poor system scalability, insufficient cross-scene dynamic adaptation capability, and inconsistent user experience.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, the present invention provides an interactive system and collaborative configuration based on visual recognition and attribute virtualization, characterized by comprising the following steps: When a user holds a physical prop, visual data, inertial motion data, and grip pressure data of the physical prop are collected simultaneously to form a multimodal perception data package; Based on multimodal perception data packets, semantic type identification and operability status analysis are performed on physical props, and the confidence scores of various data sources are integrated to generate a description vector containing semantic type and operability status. By combining the real-time acquired user status and virtual environment status, the description vector is dynamically mapped to generate an interactive configuration parameter set adapted to the current context; The interactive configuration parameter set is wirelessly sent to the general interactive module in the physical prop, driving the general interactive module to dynamically virtualize the functional logic of its hardware interface according to the parameter set. During the interaction, the multimodal consistency between virtual feedback, haptic feedback and user operations is monitored, and the interaction configuration parameter set is fine-tuned in real time to optimize consistency. When a context switch is detected, dynamic mapping and parameter generation are re-executed based on the new context state, and the general interaction module is driven to update its virtualization configuration to achieve seamless switching.
[0007] As a preferred embodiment of the visual recognition and attribute virtualization interactive system and collaborative configuration described in this invention, wherein: when the user holds a physical prop, visual data, inertial motion data, and grip pressure data of the physical prop are simultaneously collected to form a multimodal perception data packet, the specific steps are as follows: Visual frame data is generated by capturing preset visual identifiers and prop outline point clouds on physical props using the depth camera of a head-mounted device. The linear acceleration and angular velocity of the prop in three-dimensional space are collected by the inertial measurement unit built into the prop, and an inertial data stream is generated. Pressure data is generated by acquiring images of the pressure distribution on the hand contact surface through a pressure sensor array integrated into the handle of the prop; Based on a unified time reference, the visual frame data, inertial data stream, and pressure data are timestamped and their spatial coordinates are unified, and then packaged to generate a multimodal sensing data package.
[0008] As a preferred embodiment of the visual recognition and attribute virtualization-based interactive system and collaborative configuration described in this invention, the steps of performing semantic type recognition and operability state analysis on entity props based on multimodal perception data packets, and fusing the confidence levels of various data sources to generate a description vector containing semantic type and operability state are as follows: Decode visual identifiers from the visual data of the multimodal perception data packet and query the local database to obtain the basic semantic type corresponding to the identifier; Synchronously analyze the inertial data stream, identify the motion pattern to determine the macroscopic motion state of the prop, and analyze the pressure data to determine the grip posture; Calculate confidence scores for the visual recognition results, action state judgment results, and grip posture analysis results respectively; The decision is weighted based on the confidence scores of each result. When the confidence score of visual recognition is lower than the threshold, the weight of the action and grip posture analysis results is increased, and finally a unified description vector is output.
[0009] As a preferred embodiment of the visual recognition and attribute virtualization-based interactive system and collaborative configuration described in this invention, the specific steps of dynamically mapping the description vector by combining the real-time acquired user state and virtual environment state to generate an interactive configuration parameter set adapted to the current context are as follows: Real-time collection of user physiological sensor data and task events in virtual applications; calculation of user excitement index and task urgency parameters. The description vector, along with parameters for user excitement and task urgency, is input into the dynamic mapping rule base. The mapping rule base adjusts the strategy based on the input matching predefined interaction metaphors, and the strategy defines how to adjust the parameters of the underlying interaction logic; Based on the matching strategy, an interaction configuration parameter set is generated, which includes function mapping relationships, trigger thresholds, feedback types and strengths.
[0010] As a preferred embodiment of the visual recognition and attribute virtualization-based interactive system and collaborative configuration described in this invention, the specific steps of wirelessly distributing the interactive configuration parameter set to the universal interactive module within the physical prop, and driving the universal interactive module to dynamically virtualize the functional logic of its hardware interface according to the parameter set, are as follows: The main controller of the general interaction module receives the set of interaction configuration parameters and parses the function mapping relationships within them. The main controller configures the specified physical input pins as specific logic function interfaces according to the function mapping relationship and loads the corresponding signal processing algorithms; The main controller configures the specified physical output pin as a specific feedback drive interface according to the function mapping relationship and loads the corresponding waveform output protocol; After configuration is complete, the module enters the working state, and its physical interface runs according to the virtual functional logic defined by the parameter set.
[0011] As a preferred embodiment of the visual recognition and attribute virtualization-based interactive system and collaborative configuration described in this invention, the following steps are taken: During the interaction process, the multimodal consistency between virtual feedback, tactile feedback, and user operations is monitored, and the interactive configuration parameter set is fine-tuned in real time to optimize consistency. When an interactive event is triggered, the virtual scene rendering event time, the general interactive module execution feedback time, and the operation event time captured by the inertial sensor are recorded synchronously. Calculate the delay difference between the virtual feedback moment and the haptic feedback moment, as well as the matching degree between the intensity of the operation event and the intensity of the feedback signal; If the delay difference exceeds the preset threshold or the intensity matching degree is not in compliance, a fine-tuning instruction containing the timing offset compensation amount and intensity adjustment coefficient is generated. Fine-tuning instructions are sent to the virtual content rendering engine and the general interaction module respectively, so as to make real-time corrections to the triggering timing and output intensity of subsequent events.
[0012] As a preferred embodiment of the visual recognition and attribute virtualization-based interactive system and collaborative configuration described in this invention, the following steps are taken: When a context switch is detected, dynamic mapping and parameter generation are re-executed based on the new context state, and the general interactive module is driven to update its virtualization configuration to achieve seamless switching. Continuously monitor scene identifiers in the virtual environment or user-initiated switching commands as trigger signals for scene switching; Once the switching signal is confirmed, the description vector is immediately maintained unchanged based on the current holding state, but new context state parameters are obtained. Using the unchanged description vector and the new context state parameters as input, the dynamic mapping process is re-executed to generate a new set of interaction configuration parameters; The new set of interactive configuration parameters is distributed to the general interactive module in an incremental update manner. The module switches the interface logic while maintaining the basic connection, so as to achieve uninterrupted interaction.
[0013] Secondly, this invention provides a unified interaction and hardware / software co-configuration system based on visual recognition and attribute virtualization, comprising, The multimodal synchronization module simultaneously collects visual data, inertial motion data, and grip pressure data of the physical prop when the user holds it, forming a multimodal perception data packet. The perception fusion module, based on multimodal perception data packets, performs semantic type recognition and operability state analysis on physical props, and integrates the confidence of various data sources to generate a description vector containing semantic type and operability state. The dynamic mapping module combines the real-time acquired user state and virtual environment state to dynamically map the description vector and generate an interactive configuration parameter set that adapts to the current context. The configuration distribution module wirelessly distributes the interactive configuration parameter set to the general interactive module in the physical prop, and drives the general interactive module to dynamically virtualize the functional logic of its hardware interface according to the parameter set. The consistency optimization module monitors the multimodal consistency between virtual feedback, haptic feedback and user operations during the interaction process, and fine-tunes the interaction configuration parameter set in real time to optimize consistency. The scenario switching module, when a scenario switch is detected, re-executes dynamic mapping and parameter generation based on the new scenario state, and drives the general interaction module to update its virtualization configuration to achieve seamless switching.
[0014] Thirdly, the present invention provides a computer device, including a memory and a processor, wherein the memory stores a computer program, wherein: when the computer program is executed by the processor, it implements any step of the visual recognition and attribute virtualization-based interactive system and collaborative configuration method as described in the first aspect of the present invention.
[0015] Fourthly, the present invention provides a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the visual recognition and attribute virtualization-based interactive system and collaborative configuration method as described in the first aspect of the present invention.
[0016] The beneficial effects of this invention are as follows: By introducing a multimodal perception fusion and confidence decision-making mechanism, the accuracy and robustness of semantic and operability analysis of physical props in complex environments are effectively improved; by combining real-time user state and virtual environment state to dynamically map and generate parameters for interaction logic, the system can deeply understand interaction intent and achieve intelligent configuration that adapts to the context; through the virtualization of hardware interface functions and wireless parameter distribution of the general interaction module, true normalization and reconfigurability are achieved at the hardware level, significantly reducing the complexity of system development and maintenance; by monitoring and optimizing the consistency of multimodal feedback and achieving seamless configuration updates when a context switch is detected, the continuity and smoothness of the highly immersive interactive experience are ensured at the system level, thus constructing a highly scalable, flexible, and user-unified software and hardware collaborative interaction system. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the following description of the embodiments will be briefly introduced. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0018] Figure 1 This is a flowchart of the visual recognition and attribute virtualization-based interactive system and collaborative configuration method in Example 1. Detailed Implementation
[0019] To make the above-mentioned objects, features and advantages of the present invention more apparent and understandable, the specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings.
[0020] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0021] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0022] Example 1, referring to Figure 1 This is the first embodiment of the present invention, which provides a visual recognition and attribute virtualization-based interactive system and collaborative configuration, characterized by including the following steps: When a user holds a physical prop, visual data, inertial motion data, and grip pressure data of the physical prop are collected simultaneously to form a multimodal perception data package; Based on multimodal perception data packets, semantic type identification and operability status analysis are performed on physical props, and the confidence scores of various data sources are integrated to generate a description vector containing semantic type and operability status. By combining the real-time acquired user status and virtual environment status, the description vector is dynamically mapped to generate an interactive configuration parameter set adapted to the current context; The interactive configuration parameter set is wirelessly sent to the general interactive module in the physical prop, driving the general interactive module to dynamically virtualize the functional logic of its hardware interface according to the parameter set. During the interaction, the multimodal consistency between virtual feedback, haptic feedback and user operations is monitored, and the interaction configuration parameter set is fine-tuned in real time to optimize consistency. When a context switch is detected, dynamic mapping and parameter generation are re-executed based on the new context state, and the general interaction module is driven to update its virtualization configuration to achieve seamless switching.
[0023] It should be noted that by using multi-source perception fusion and confidence-based decision-making, the robustness and accuracy of prop recognition in complex environments are improved; context awareness is introduced to achieve dynamic intelligent adaptation of interaction logic; hardware interface software virtualization technology is used to support multiple functions with a single set of general-purpose hardware, reducing costs; real-time multimodal feedback calibration ensures the consistency of immersion; and finally, seamless function transitions are achieved during context switching, ensuring a highly consistent and smooth cross-scene interactive experience.
[0024] Specifically, when the user holds the physical prop, visual data, inertial motion data, and grip pressure data of the physical prop are collected simultaneously to form a multimodal perception data packet. The specific steps are as follows: Visual frame data is generated by capturing preset visual identifiers and prop outline point clouds on physical props using the depth camera of a head-mounted device. The linear acceleration and angular velocity of the prop in three-dimensional space are collected by the inertial measurement unit built into the prop, and an inertial data stream is generated. Pressure data is generated by acquiring images of the pressure distribution on the hand contact surface through a pressure sensor array integrated into the handle of the prop; Based on a unified time reference, the visual frame data, inertial data stream, and pressure data are timestamped and their spatial coordinates are unified, and then packaged to generate a multimodal sensing data package.
[0025] It should be noted that the depth camera integrated into the head-mounted device actively emits infrared structured light and receives reflections, simultaneously acquiring RGB color images and depth information of the scene. Its processor processes the images in real time, identifying specific visual identifier patterns pasted on the prop surface and simultaneously calculating the 3D point cloud coordinates of the prop surface surrounding the identifier. The identifier information and corresponding 3D contour data are then encapsulated into a complete visual frame. This step, through active optical detection, simultaneously acquires the target's identity code and precise 3D geometric information, providing the system with the necessary visual features for prop recognition and spatial attitude estimation, achieving non-contact, rapid identity and attitude initialization.
[0026] A miniature inertial measurement unit (IMU), encapsulated within a general-purpose interactive module, continuously samples the raw readings of its triaxial accelerometer and triaxial gyroscope at a high frequency. The processor within this unit performs preprocessing on this raw data, including temperature compensation and bias correction, and packages the processed linear acceleration and angular velocity data in chronological order to form a continuous time-series data stream. This step, by measuring the prop's own motion dynamics parameters, provides high-frequency raw motion information independent of external vision, offering a crucial data source for subsequent real-time assessment of the prop's precise trajectory, speed changes, and subtle vibrations.
[0027] An array of flexible pressure sensors embedded within the handle housing of the prop allows each sensing unit to detect the normal pressure applied to its surface in real time. The array's drive circuit sequentially reads the resistance or capacitance changes of all sensing units at a predetermined scanning frequency and converts them into digital signals. The main controller arranges these discrete pressure values according to the physical positions of the sensing units in the array, combining them into a digital image representing the spatial distribution of pressure on the handle surface. This step enables refined perception of the user's gripping behavior, quantifying the abstract "holding" action into an analyzable pressure distribution map, providing direct tactile evidence for determining the user's hand posture, grip strength, and intentions.
[0028] The system employs a high-precision clock source to assign a unified timestamp to all data acquisition events. A data synchronization management unit receives timestamped raw data from the camera, inertial measurement unit, and pressure array. Using the visual frame's time as a reference, it aligns the timelines of the inertial data stream and pressure data through an interpolation algorithm. Simultaneously, based on pre-calibrated spatial transformation relationships between sensors, all data is uniformly transformed to the world coordinate system of the head-mounted device. Finally, the spatiotemporally aligned heterogeneous data is encapsulated into a structured data packet. This step resolves the temporal and spatial asynchrony issue of multi-source heterogeneous sensor data, generating a foundation of fused sensing data with strict spatiotemporal correlation, ensuring the accuracy and effectiveness of subsequent fusion analysis.
[0029] Specifically, the step of performing semantic type identification and operability state analysis on entity props based on multimodal perception data packets, and fusing the confidence levels of various data sources to generate a description vector containing semantic type and operability state, involves the following steps: Decode visual identifiers from the visual data of the multimodal perception data packet and query the local database to obtain the basic semantic type corresponding to the identifier; Synchronously analyze the inertial data stream, identify the motion pattern to determine the macroscopic motion state of the prop, and analyze the pressure data to determine the grip posture; Calculate confidence scores for the visual recognition results, action state judgment results, and grip posture analysis results respectively; The decision is weighted based on the confidence scores of each result. When the confidence score of visual recognition is lower than the threshold, the weight of the action and grip posture analysis results is increased, and finally a unified description vector is output.
[0030] It should be noted that the system invokes a dedicated image processing unit to locate and decode specific regions in the visual frame data, extracting the unique string or number encoded by the visual identifier. Subsequently, this number is sent to a lightweight database stored on the head-mounted device or a local edge computing node for rapid querying. The database returns the basic item semantic type pre-bound to this number, such as "longsword," "shield," or "medical kit." This step, through rapid decoding and querying, accurately converts the optical identifier of the physical item into its preset digital identity, providing a stable and reliable foundation for item identity authentication for the entire system and ensuring the initial accuracy of the interaction logic binding.
[0031] The system's inertial data processing thread continuously receives inertial data streams and compares them with a series of pre-stored typical action templates (such as rapid waving, slow movement, and stillness) in the time and frequency domains to determine the current motion mode. Simultaneously, the pressure data processing thread analyzes the shape, center of gravity position, and total pressure value of the pressure distribution image, determining the user's grip method based on a predefined posture classification model (such as single-handed grip, two-handed grip, and fingertip pinch). This step, through parallel real-time analysis of motion and pressure signals, quantifies the user's manipulation state of the prop from different physical dimensions, providing crucial behavioral contextual information for understanding the user's real-time interaction intentions.
[0032] The confidence assessment unit receives the three intermediate results and their original data features, and performs a reliability assessment on each result. For visual recognition, the confidence level is calculated based on the sharpness, completeness, and decoding matching degree of the identifier image. For action state judgment, the confidence level is calculated based on the similarity between the current inertial data features and the matching template. For grip posture analysis, the confidence level is calculated based on the quality of the pressure image and the output probability of the classification model. This step assigns a quantitative reliability measure to the output of each sensing source, unifying sensing information of different natures and reliability levels onto a comparable scale, providing an objective basis for subsequent intelligent decision fusion.
[0033] The fusion decision-maker receives three results and their confidence scores, and internally employs a set of dynamic weighting rules. When the visual recognition confidence score is higher than a set threshold, the system primarily outputs the semantic type of the visual recognition. Once the visual confidence score falls below the threshold, the decision-maker automatically reduces the weight of visual recognition and correspondingly increases the influence of action state based on inertial data and grip posture based on pressure data in the final decision. A weighted algorithm is then used to output a comprehensive and optimal descriptive vector. This step creates an adaptive and robust decision-making mechanism, enabling the system to automatically switch to a behavior- and touch-based inference mode in complex environments where vision-dominated recognition fails. This ensures the continuity and reliability of the entire interaction process under different environmental conditions.
[0034] Specifically, the step of dynamically mapping the description vector by combining the real-time acquired user state and the virtual environment state to generate an interaction configuration parameter set adapted to the current context involves the following steps: Real-time collection of user physiological sensor data and task events in virtual applications; calculation of user excitement index and task urgency parameters. The description vector, along with parameters for user excitement and task urgency, is input into the dynamic mapping rule base. The mapping rule base adjusts the strategy based on the input matching predefined interaction metaphors, and the strategy defines how to adjust the parameters of the underlying interaction logic; Based on the matching strategy, an interaction configuration parameter set is generated, which includes function mapping relationships, trigger thresholds, feedback types and strengths.
[0035] It should be noted that the system collects raw signals such as heart rate and skin conductance at a fixed frequency using a wristband-style biosensor worn by the user, and then filters and extracts features from these signals. Simultaneously, the system monitors task status change messages issued by the virtual content engine. A dedicated parameter calculation unit receives these processed biosignal features and task event identifiers, and outputs a quantified user arousal index and a parameter representing the urgency of the current virtual task through a predefined algorithm model. This step, by quantifying the user's internal physiological state and external task pressure, provides objective and calculable contextual input for the dynamic adjustment of the interactive system, enabling the configuration process to respond to the user's real-time physical and mental state and the dynamic needs of the virtual environment.
[0036] The system combines the description vector representing the item's identity and holding state with the excitement index and task urgency parameters calculated in the previous step into a complete input tuple. This tuple is passed to a dynamically mapped rule base stored in memory via a standard software interface. This rule base is a structured set of lookup tables or decision trees, where each rule defines the mapping relationship between specific input conditions and corresponding output strategies. This step completes the aggregation and formatting of multi-dimensional perceived information (item, user, environment), preparing data for subsequent rule-based intelligent decision-making and serving as a crucial bridge connecting perception and execution.
[0037] After receiving the input tuple, the dynamic mapping rule base's internal matching engine begins comparing each value in the tuple with the triggering conditions of each rule in the base. The matching process follows priority and accuracy principles; for example, it prioritizes rules with high task urgency or searches for the most suitable descriptive vector type within the given conditions. When a matching rule is found, the engine immediately locks onto and invokes the specific "interaction metaphor adjustment strategy" associated with that rule. This strategy exists in the form of machine-readable code or a configuration file. This step, through the rule matching mechanism, automatically transforms complex situational states into concrete, executable interactive logic modification instructions, replacing the tedious work of manually predefining all possibilities and achieving automated and intelligent configuration.
[0038] The system executes the invoked "interaction metaphor adjustment strategy," which contains a series of instructions to modify the basic interaction template. Based on these instructions, the generator specifically adjusts the mapping between virtual functions and physical input pins, resets sensitivity thresholds such as trigger pull, and selects the vibration waveform file and playback intensity level of the haptic feedback motor. Finally, all these adjusted parameters are organized and encapsulated into a complete, uniformly formatted digital configuration file. This step is the final output of dynamic configuration; it transforms the abstract strategy into a set of specific operational instructions required to drive the general-purpose hardware module, enabling the same physical prop to generate unique interactive behavior configurations based on the real-time context, achieving deep personalization and scenario-based adaptation.
[0039] Specifically, the steps for wirelessly transmitting the interaction configuration parameter set to the universal interaction module within the physical prop, and driving the universal interaction module to dynamically virtualize the functional logic of its hardware interface according to the parameter set, are as follows: The main controller of the general interaction module receives the set of interaction configuration parameters and parses the function mapping relationships within them. The main controller configures the specified physical input pins as specific logic function interfaces according to the function mapping relationship and loads the corresponding signal processing algorithms; The main controller configures the specified physical output pin as a specific feedback drive interface according to the function mapping relationship and loads the corresponding waveform output protocol; After configuration is complete, the module enters the working state, and its physical interface runs according to the virtual functional logic defined by the parameter set.
[0040] It should be noted that the general-purpose interactive module continuously listens for and receives wireless data packets from the host system through its integrated wireless communication chip. After confirming the integrity of the data packets, the module's main controller decodes them according to a predefined communication protocol, extracting a structured set of interactive configuration parameters. The controller further parses this parameter set, clarifying the specific functional mapping entries defined within it, such as mapping "physical pin A1" to "main trigger input". This step completes the reliable reception and accurate interpretation of remote configuration commands, providing a unique and clear instruction basis for the subsequent precise reprogramming of hardware interfaces, and is the primary step in realizing the software definition of hardware functions.
[0041] The main controller consults the parsed function mapping relationship to locate the specific physical input pin that needs to be configured. Through register write operations, the controller reconfigures the pin's operating mode from a general-purpose input to a dedicated input interface with specific interrupt triggering characteristics. Simultaneously, it loads a dedicated signal processing algorithm program corresponding to this logical function interface from the module's internal memory or the received parameter set; for example, a firmware module containing debouncing and analog threshold judgment. This step dynamically redefines the behavioral logic of the hardware input circuit through software instructions, enabling the same physical pin to carry multiple different interactive input semantics, achieving on-demand allocation and flexible reuse of hardware input resources.
[0042] Based on the functional mapping relationship, the main controller locates the designated physical output pin. The controller then configures the output drive mode and electrical characteristics of this pin, setting it as a feedback drive interface suitable for driving a specific actuator. Subsequently, the controller indexes and loads a complex waveform data file or real-time generation protocol that matches the current virtual feedback type from the parameter set or local library. This step achieves dynamic binding between the physical output channel and rich virtual feedback effects, enabling simple output pins to generate diverse haptic feedback that precisely matches the virtual content, thus enhancing interactive expressiveness.
[0043] Once all specified input and output pins have been reconfigured and driver loaded according to the interactive configuration parameter set, the main controller executes a state switching command, causing the module to officially enter "operating mode" from "configuration mode". In this mode, the module's firmware will analyze the input signals in real time according to the newly loaded logic and control the actuators according to the newly bound output protocol. Its overall behavior is completely defined by the issued parameter set and decoupled from the original physical design of the hardware. This step marks the end of a complete dynamic virtualization process. General-purpose hardware is instantly transformed into a virtual device with specific interactive functions, achieving "plug and play" and providing fundamental hardware flexibility to cope with diverse interactive scenarios.
[0044] Specifically, during the interaction process, the multimodal consistency between virtual feedback, haptic feedback, and user operations is monitored, and the interaction configuration parameter set is fine-tuned in real time to optimize consistency. The specific steps are as follows: When an interactive event is triggered, the virtual scene rendering event time, the general interactive module execution feedback time, and the operation event time captured by the inertial sensor are recorded synchronously. Calculate the delay difference between the virtual feedback moment and the haptic feedback moment, as well as the matching degree between the intensity of the operation event and the intensity of the feedback signal; If the delay difference exceeds the preset threshold or the intensity matching degree is not in compliance, a fine-tuning instruction containing the timing offset compensation amount and intensity adjustment coefficient is generated. Fine-tuning instructions are sent to the virtual content rendering engine and the general interaction module respectively, so as to make real-time corrections to the triggering timing and output intensity of subsequent events.
[0045] It should be noted that when the system determines an interactive event has occurred based on the processed input signal (such as the electrical signal of pulling a trigger), a high-precision timer is triggered. This timer captures three key time points in parallel: the GPU instruction timestamp indicating that the corresponding visual effect (such as a spark) is submitted for rendering from the graphics rendering pipeline; the confirmation signal timestamp confirming via wireless communication that the general-purpose interactive module has received the instruction and started driving the vibration motor; and the timestamp identifying the starting point of a specific impact waveform generated by user operation (such as recoil) by analyzing the inertial measurement unit data stream. This step, through high-precision timing and correlation with multi-source signals, establishes for the first time a comparable and unified spatiotemporal observation benchmark for the occurrence time of a single interactive event across different software and hardware subsystems, providing a precise data foundation for quantitatively evaluating the overall system response consistency.
[0046] The system's time analysis unit reads the three moments mentioned above. First, it calculates the absolute time difference between the visual / auditory feedback moment and the tactile feedback moment as the "cross-modal delay." Simultaneously, the signal analysis unit extracts the peak acceleration or energy of the current operation event from the inertial sensor data as the "operation intensity," and reads the peak voltage or duty cycle of the output vibration waveform from the feedback drive log of the general interaction module as the "feedback intensity." The ratio or correlation between the two is calculated as the "intensity matching degree." This step transforms the subjective sensory experience of "asynchrony" and "inconsistency in intensity" into two objective, calculable physical quantity indicators, enabling the system to accurately diagnose the specific type of consistency problem—whether it is a timing misalignment or an energy imbalance.
[0047] The system's decision-making unit compares the calculated "cross-modal latency" with a stored, human-perception-calibrated comfort latency threshold, and compares the "intensity matching degree" with an ideal ratio range. If the latency exceeds the threshold, an instruction is generated instructing the rendering engine to advance or delay the trigger time of the next similar event by a specific number of milliseconds. If the intensity matching degree does not match, another instruction is generated instructing the general interaction module to scale its output power by a specific percentage factor. This step automates the decision-making process from "problem diagnosis" to "correction solution generation." It generates specific, executable compensation parameters based on objective measurements and preset rules, rather than general adjustment suggestions.
[0048] After generating the fine-tuning instructions, the system sends the timing offset instructions to the engine module responsible for rendering and sound effects via internal inter-process communication. This module immediately adjusts its internal event response scheduling logic. Simultaneously, it sends the intensity adjustment instructions to the main controller of the general interaction module via a low-latency wireless link. The main controller dynamically updates the gain parameters of its output driver module. Thereafter, the system immediately applies the corrected parameters to newly occurring interactive events and continuously monitors the effects, forming a real-time closed loop. This step completes the full closed loop of perception-decision-execution, enabling the system to proactively and dynamically eliminate multi-sensory dissonance caused by differences in physical latency or gain mismatch between different subsystems, thereby continuously ensuring and optimizing the immersive and realistic experience of the user.
[0049] Specifically, when a scenario switch is detected, dynamic mapping and parameter generation are re-executed based on the new scenario state, and the general interaction module is driven to update its virtualization configuration to achieve seamless switching. The specific steps are as follows: Continuously monitor scene identifiers in the virtual environment or user-initiated switching commands as trigger signals for scene switching; Once the switching signal is confirmed, the description vector is immediately maintained unchanged based on the current holding state, but new context state parameters are obtained. Using the unchanged description vector and the new context state parameters as input, the dynamic mapping process is re-executed to generate a new set of interaction configuration parameters; The new set of interactive configuration parameters is distributed to the general interactive module in an incremental update manner. The module switches the interface logic while maintaining the basic connection, so as to achieve uninterrupted interaction.
[0050] It should be noted that the system runs a background monitoring service. This service continuously parses scene status messages broadcast by the virtual content engine, identifying specific scene identifier change events contained within them. Simultaneously, it monitors the user interface, capturing explicit mode-switching commands issued by the user via gestures, voice, or menu buttons. This monitoring service filters and confirms these two types of signals, transforming them into standardized context-switching trigger signals. This step establishes a proactive, dual-channel context change perception mechanism, enabling the system to promptly and accurately detect when functional reconfiguration is needed, based on changes within the virtual world or external user commands.
[0051] Once the switching trigger signal is confirmed to be valid, the system first freezes the descriptive vector generated by the perception fusion module, which describes the current holding state and type of the prop, and stores it in a temporary buffer to ensure that it will not be overwritten by subsequent real-time analysis. At the same time, the system immediately re-collects and calculates a new set of user state parameters and environmental state parameters reflecting the post-switching context from the biosensor interface and virtual engine. This step adopts a "state freeze and parameter refresh" strategy, which, while maintaining a stable basic understanding of the physical prop, quickly updates the dynamic contextual information, preserving accurate holding state information and providing new decision-making basis for function remapping.
[0052] The system combines the frozen description vector in the temporary buffer with the newly acquired contextual state parameters as a complete input, and submits it again to the rule base of the dynamic mapping module for matching calculation. Based on this new input, the mapping rule base performs a fast strategy retrieval and decision-making process, outputting a set of interaction metaphor adjustment strategies suitable for the new context, and generating a completely new set of structured interaction configuration parameters. This step reuses the core dynamic mapping capability, but the input contextual parameters have changed, thus enabling the rapid calculation of a completely different virtual function configuration scheme for the same entity prop even when the user's holding action is uninterrupted.
[0053] The newly generated set of interactive configuration parameters is sent to the general-purpose interactive module via a wireless link. This set contains only the configuration items that need to be changed. While maintaining communication with the host and power supply, the module's main controller receives and parses this incremental set of parameters, dynamically reprogramming the specified hardware interface logic online, overwriting the old configuration and activating the new one. This step enables hot updates of the configuration, allowing the general-purpose interactive module to complete function switching within milliseconds. Users can experience a smooth transition in functionality with changing contexts without needing to put down or re-identify the prop, ensuring a consistent cross-scene interactive experience.
[0054] This embodiment also provides a unified interaction and hardware / software co-configuration system based on visual recognition and attribute virtualization, including: The multimodal synchronization module simultaneously collects visual data, inertial motion data, and grip pressure data of the physical prop when the user holds it, forming a multimodal perception data packet. The perception fusion module, based on multimodal perception data packets, performs semantic type recognition and operability state analysis on physical props, and integrates the confidence of various data sources to generate a description vector containing semantic type and operability state. The dynamic mapping module combines the real-time acquired user state and virtual environment state to dynamically map the description vector and generate an interactive configuration parameter set that adapts to the current context. The configuration distribution module wirelessly distributes the interactive configuration parameter set to the general interactive module in the physical prop, and drives the general interactive module to dynamically virtualize the functional logic of its hardware interface according to the parameter set. The consistency optimization module monitors the multimodal consistency between virtual feedback, haptic feedback and user operations during the interaction process, and fine-tunes the interaction configuration parameter set in real time to optimize consistency. The scenario switching module, upon detecting a scenario switch, re-executes dynamic mapping and parameter generation based on the new scenario state and drives the general interaction module to update its virtualization configuration, achieving seamless switching. This embodiment also provides a computer device applicable to visual recognition and attribute virtualization interaction systems and collaborative configuration methods, including: a memory and a processor; the memory stores computer-executable instructions, and the processor executes the computer-executable instructions to implement the visual recognition and attribute virtualization interaction system and collaborative configuration method proposed in the above embodiments.
[0055] The computer device can be a terminal, comprising a processor, memory, communication interface, display screen, and input devices connected via a system bus. The processor provides computing and control capabilities. The memory includes non-volatile storage media and internal memory. The non-volatile storage media stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage media. The communication interface is used for wired or wireless communication with external terminals; wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen can be an LCD screen or an e-ink screen. The input devices can be a touch layer covering the display screen, buttons, a trackball, or a touchpad on the computer device's casing, or an external keyboard, touchpad, or mouse.
[0056] This embodiment also provides a storage medium storing a computer program. When executed by a processor, the program implements the visual recognition and attribute virtualization-based interactive system and collaborative configuration method as proposed in the above embodiments. The storage medium can be implemented by any type of volatile or non-volatile storage device or a combination thereof, such as Static Random Access Memory (SRAM), Electrically Erasable Programmable Read-Only Memory (EEPROM), Erasable Programmable Read Only Memory (EPROM), Programmable Red-Only Memory (PROM), Read-Only Memory (ROM), magnetic storage, flash memory, magnetic disk, or optical disk.
[0057] In summary, this invention effectively improves the accuracy and robustness of semantic and operability analysis of physical props in complex environments by introducing multimodal perception fusion and confidence decision-making mechanisms; it enables the system to deeply understand interaction intentions and achieve context-adaptive intelligent configuration by combining real-time user state and virtual environment state to dynamically map and generate parameters for interaction logic; it achieves true normalization and reconfigurability at the hardware level through the virtualization of hardware interface functions and wireless parameter distribution of general interaction modules, significantly reducing the complexity of system development and maintenance; and it ensures the continuity and smoothness of highly immersive interactive experiences at the system level by monitoring and optimizing the consistency of multimodal feedback and achieving seamless configuration updates when context switching is detected. Thus, it constructs a highly scalable, flexible, and user-unified hardware and software collaborative interaction system.
[0058] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. Based on visual recognition and attribute virtualization interaction and collaborative configuration, characterized in that, Includes the following steps: When a user holds a physical prop, visual data, inertial motion data, and grip pressure data of the physical prop are collected simultaneously to form a multimodal perception data package; Based on multimodal perception data packets, semantic type identification and operability status analysis are performed on physical props, and the confidence levels of various data sources are integrated to generate description vectors containing semantic type and operability status. By combining the real-time acquired user status and virtual environment status, the description vector is dynamically mapped to generate an interactive configuration parameter set adapted to the current context; The interactive configuration parameter set is wirelessly sent to the general interactive module in the physical prop, driving the general interactive module to dynamically virtualize the functional logic of its hardware interface according to the parameter set. During the interaction, the multimodal consistency between virtual feedback, haptic feedback and user operations is monitored, and the interaction configuration parameter set is fine-tuned in real time to optimize consistency. When a context switch is detected, dynamic mapping and parameter generation are re-executed based on the new context state, and the general interaction module is driven to update its virtualization configuration to achieve seamless switching.
2. The visual recognition based interaction and co-configuration with attribute virtualization as claimed in claim 1, wherein: When the user holds the physical prop, visual data, inertial motion data, and grip pressure data of the physical prop are collected simultaneously to form a multimodal perception data packet. The specific steps are as follows: Visual frame data is generated by capturing preset visual identifiers and prop outline point clouds on physical props using the depth camera of a head-mounted device. The linear acceleration and angular velocity of the prop in three-dimensional space are collected by the inertial measurement unit built into the prop, and an inertial data stream is generated. Pressure data is generated by acquiring images of the pressure distribution on the hand contact surface through a pressure sensor array integrated into the handle of the prop; Based on a unified time reference, the visual frame data, inertial data stream, and pressure data are timestamped and their spatial coordinates are unified, and then packaged to generate a multimodal sensing data package.
3. The visual recognition based interaction and co-configuration with attribute virtualization as claimed in claim 2, wherein: The process involves using multimodal sensing data packets to perform semantic type identification and operability state analysis on entity props, and fusing confidence levels from various data sources to generate a description vector containing semantic type and operability state. The specific steps are as follows: Decode visual identifiers from the visual data of the multimodal perception data packet and query the local database to obtain the basic semantic type corresponding to the identifier; Synchronously analyze the inertial data stream, identify the motion pattern to determine the macroscopic motion state of the prop, and analyze the pressure data to determine the grip posture; Calculate confidence scores for the visual recognition results, action state judgment results, and grip posture analysis results respectively; The decision is weighted based on the confidence scores of each result. When the confidence score of visual recognition is lower than the threshold, the weight of the action and grip posture analysis results is increased, and finally a unified description vector is output.
4. The visual recognition based interaction and co-configuration with attribute virtualization of claim 3, wherein: The step of dynamically mapping the description vector by combining the real-time acquired user state and virtual environment state to generate an interaction configuration parameter set adapted to the current context is as follows: Real-time collection of user physiological sensor data and task events in virtual applications; calculation of user excitement index and task urgency parameters. The description vector, along with parameters for user excitement and task urgency, is input into the dynamic mapping rule base. The mapping rule base adjusts the strategy based on the input matching predefined interaction metaphors, and the strategy defines how to adjust the parameters of the underlying interaction logic; Based on the matching strategy, an interaction configuration parameter set is generated, which includes function mapping relationships, trigger thresholds, feedback types and strengths.
5. The visual recognition based interaction and co-configuration with attribute virtualization as claimed in claim 4, wherein: The specific steps for wirelessly transmitting the interaction configuration parameter set to the universal interaction module within the physical prop, and driving the universal interaction module to dynamically virtualize the functional logic of its hardware interface according to the parameter set, are as follows: The main controller of the general interaction module receives the set of interaction configuration parameters and parses the function mapping relationships within them. The main controller configures the specified physical input pins as specific logic function interfaces according to the function mapping relationship and loads the corresponding signal processing algorithms; The main controller configures the specified physical output pin as a specific feedback drive interface according to the function mapping relationship and loads the corresponding waveform output protocol; After configuration is complete, the module enters the working state, and its physical interface runs according to the virtual functional logic defined by the parameter set.
6. The visual recognition based interaction and co-configuration with attribute virtualization of claim 5, wherein: During the interaction process, the multimodal consistency between virtual feedback, haptic feedback, and user operations is monitored, and the interaction configuration parameter set is fine-tuned in real time to optimize consistency. The specific steps are as follows: When an interactive event is triggered, the virtual scene rendering event time, the general interactive module execution feedback time, and the operation event time captured by the inertial sensor are recorded synchronously. Calculate the delay difference between the virtual feedback moment and the haptic feedback moment, as well as the matching degree between the intensity of the operation event and the intensity of the feedback signal; If the delay difference exceeds the preset threshold or the intensity matching degree is not in compliance, a fine-tuning instruction containing the timing offset compensation amount and intensity adjustment coefficient is generated. Fine-tuning instructions are sent to the virtual content rendering engine and the general interaction module respectively, so as to make real-time corrections to the triggering timing and output intensity of subsequent events.
7. The visual recognition based interaction and co-configuration with attribute virtualization as claimed in claim 6, wherein: When a scenario switch is detected, dynamic mapping and parameter generation are re-executed based on the new scenario state, and the general interaction module is driven to update its virtualization configuration to achieve seamless switching. The specific steps are as follows: Continuously monitor scene identifiers in the virtual environment or user-initiated switching commands as trigger signals for scene switching; Once the switching signal is confirmed, the description vector is immediately maintained unchanged based on the current holding state, but new context state parameters are obtained. Using the unchanged description vector and the new context state parameters as input, the dynamic mapping process is re-executed to generate a new set of interaction configuration parameters; The new set of interactive configuration parameters is distributed to the general interactive module in an incremental update manner. The module switches the interface logic while maintaining the basic connection, so as to achieve uninterrupted interaction.
8. A system for interaction and collaborative configuration based on visual recognition and attribute virtualization, based on the interaction and collaborative configuration based on visual recognition and attribute virtualization as described in any one of claims 1 to 7, characterized in that: include, The multimodal synchronization module simultaneously collects visual data, inertial motion data, and grip pressure data of the physical prop when the user holds it, forming a multimodal perception data packet. The perception fusion module, based on multimodal perception data packets, performs semantic type recognition and operability state analysis on physical props, and integrates the confidence of various data sources to generate a description vector containing semantic type and operability state. The dynamic mapping module combines the real-time acquired user state and virtual environment state to dynamically map the description vector and generate an interactive configuration parameter set that adapts to the current context. The configuration distribution module wirelessly distributes the interactive configuration parameter set to the general interactive module in the physical prop, and drives the general interactive module to dynamically virtualize the functional logic of its hardware interface according to the parameter set. The consistency optimization module monitors the multimodal consistency between virtual feedback, haptic feedback and user operations during the interaction process, and fine-tunes the interaction configuration parameter set in real time to optimize consistency. The scenario switching module, when a scenario switch is detected, re-executes dynamic mapping and parameter generation based on the new scenario state, and drives the general interaction module to update its virtualization configuration to achieve seamless switching.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the visual recognition and attribute virtualization-based interactive system and collaborative configuration method as described in any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the visual recognition and attribute virtualization-based interactive system and collaborative configuration method as described in any one of claims 1 to 7.