VR interactive control management system and method

By performing intra-modal representation conversion and fusion on multi-source interaction data, extracting semantic intent information, generating control instructions, and performing device channel mapping and feedback reconstruction, the problems of multimodal data fusion and cross-modal semantic understanding in existing VR interactive control technologies are solved, and dynamic response and feedback closed loops are achieved in highly complex virtual reality applications.

CN120406748BActive Publication Date: 2025-09-16HANGZHOU KAILIN CULTURE TECHNOLOGY CO LTD +1
View PDF 2 Cites 0 Cited by

Patent Information

Application Number
CN202510907593.8
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-09-16
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

Existing VR interactive control technologies lack the ability to fusion multimodal data and understand cross-modal semantics, making it difficult to support dynamic perception and response control in complex contexts. In addition, control signal generation is limited in scalability and flexibility.

Method used

By acquiring multi-source interaction data for intra-modal representation conversion and fusion, extracting semantic intent information, generating control instruction vectors, and performing device channel mapping and feedback reconstruction, closed-loop interaction between control and perception is achieved.

Benefits of technology

It realizes unified spatiotemporal structure modeling of multi-source interactive data, enhances the dynamic response capability and structural adaptability of control instruction generation, and improves the extraction accuracy and representation consistency of cross-modal semantic information.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120406748B_ABST
    Figure CN120406748B_ABST
Patent Text Reader

Abstract

The present invention discloses a VR interactive control management system and method, which relates to the field of virtual reality technology. The method includes: obtaining multi-source interaction data of users in a virtual reality environment, and performing intra-modal representation conversion and fusion to construct a unified multi-modal input tensor; extracting semantic intent representation based on a cross-modal perception structure; generating a control instruction vector by combining historical state information with current semantic intent; mapping the control instruction vector into a set of device control signals that conform to multiple VR terminal interface specifications; after the device executes the control instruction, collecting multi-dimensional feedback information and reconstructing its structure to generate a feedback representation that can flow back to the perception module, thereby realizing a closed-loop interaction between perception and control. By unifying the multi-source interaction data structure, constructing a dynamically adaptive control instruction vector, and a high-precision cross-modal semantic representation mechanism, a closed-loop interaction process with consistent structure, flexible response, and semantic alignment of perception and control in a virtual reality system is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of virtual reality technology, and in particular to a VR interactive control management system and method. Background Art

[0002] As a key form of next-generation human-computer interaction, virtual reality technology has been widely used in fields such as education and training, medical rehabilitation, remote collaboration, and intelligent manufacturing. In a virtual reality system, users interact with the virtual environment through various sensory devices. The system must promptly generate control responses based on user behavioral input and effectively perceive and process terminal feedback to achieve an immersive, highly real-time interactive experience. Existing VR interactive control technologies primarily focus on recognizing and responding to single-modal input (such as gestures and voice), lacking the ability to uniformly model and deeply understand multi-source heterogeneous data, making it difficult to support dynamic perception and responsive control requirements in complex contexts.

[0003] Furthermore, existing solutions generally employ static channel mapping or preset instruction templates, lacking adaptation mechanisms for the differences between terminal control interfaces. This significantly limits the scalability and flexibility of control signal generation. Within the control loop, feedback information is often transmitted in the form of a single physical state quantity, lacking structural and semantic descriptions. This makes it difficult to directly align and effectively utilize feedback data with perception modules, hindering the development of the system's adaptive control capabilities.

[0004] Therefore, there is an urgent need for a VR interactive control management method and system with the capabilities of multimodal data fusion, cross-modal semantic extraction, state-driven control generation, channel-level control signal mapping and structured feedback reflow to support highly complex, multi-dimensional and highly interactive virtual reality application scenarios. Summary of the Invention

[0005] Based on the above-mentioned shortcomings of the prior art, the purpose of the present invention is to provide a VR interactive control management system and method to solve the above-mentioned technical problems.

[0006] To achieve the above objectives, the present invention provides the following technical solution: a VR interactive control and management method, comprising:

[0007] S1: Obtain multi-source interaction data of users in a virtual reality environment, perform intra-modal representation conversion on each of the multi-source interaction data to form a modal vector sequence with a unified time scale and channel structure, and fuse them to generate a unified multi-modal input tensor;

[0008] S2: Based on the preset cross-modal perception structure, the user semantic intent information is extracted from the unified multimodal input tensor to form a semantic intent representation;

[0009] S3: Combine the system's internal state information at the previous moment with the current semantic intent representation to generate the corresponding control instruction vector;

[0010] S4: Perform device channel mapping on the control instruction vector to generate a set of device control signals that conform to different VR terminal control interfaces;

[0011] S5: Receives the multi-dimensional feedback information generated by the device after executing the control command, and restructures the feedback information to form a feedback representation that can flow back to S1, realizing a closed-loop interaction between control and perception.

[0012] The present invention is further configured such that the multi-source interaction data includes three-dimensional gesture trajectory data, voice signal feature data, eye movement behavior data, and physiological response data;

[0013] Independent modal representation spaces are constructed for multi-source interaction data, and each modal data is aligned to a unified time scale based on a structure-preserving resampling mechanism.

[0014] A modal vector sequence with a unified channel structure is constructed through a cross-modal channel normalization mapping strategy;

[0015] Generate a unified multimodal input tensor through modality concatenation and channel-level fusion.

[0016] The present invention is further configured such that the generation of the semantic intent representation table includes:

[0017] Split the multimodal input tensor according to the modal dimension to form a set of modal sub-tensors;

[0018] Perform bilinear interactive mapping on the channel-level coupling relationship between each modal sub-tensor to form a time-distributed modal cross tensor structure;

[0019] The modal cross tensor is structurally compressed and a unified semantic vector representation is constructed.

[0020] The present invention is further configured such that the unified semantic vector representation is generated by compressing a modal cross tensor structure constructed between a plurality of modal sub-tensors;

[0021] The modal cross tensor structure is constructed by performing channel-level bilinear interaction between the modal sub-tensors obtained by splitting the unified multimodal input tensor according to the modal dimension;

[0022] Structural compression includes multi-dimensional compression mapping of the modal cross tensor based on the channel coupling mapping kernel at a fixed time scale, and fusing the compressed representations to form a unified semantic vector representation.

[0023] The present invention is further configured such that combining the system internal state information at the previous moment with the current semantic intent representation to generate a corresponding control instruction vector includes:

[0024] Align the system's internal state information at the previous moment with the control history trajectory to construct a fusion state representation;

[0025] Perform channel matching and cross-embedding between the fusion state representation and the current semantic intent representation to construct a semantic regulation representation tensor;

[0026] The semantic control representation tensor is mapped through multi-order transformation to obtain a control instruction vector, and the instruction vector is subjected to feedback correction and scale normalization.

[0027] The present invention is further configured such that the control instruction vector is a multidimensional vector having a hierarchical channel configuration structure, comprising a state fusion representation domain, a semantic response drive domain, and a channel control weight domain, wherein the domains are arranged in an orderly manner according to a preset semantic mosaic structure;

[0028] State fusion represents the structural alignment between the historical state of the domain representation system and the semantic input;

[0029] The semantic response driving domain encodes the structurally active regions in the current semantic vector through channel index mapping;

[0030] The channel control weight domain uses a non-overlapping index method to record the cross-channel dynamic response amplitude and structure allocation priority. The subdomains maintain structural independence through cross-domain mapping rules and share a unified channel index table.

[0031] The present invention is further configured such that generating a set of device control signals conforming to different VR terminal control interfaces includes:

[0032] Perform multi-channel response structure expansion on the input control instruction vector to form a high-dimensional control response tensor;

[0033] Based on the preset device channel mapping rules, the control response tensor is mapped to the functional channels of multiple virtual reality terminals;

[0034] Perform structural strength adjustment and offset correction on the control signals of each terminal channel after mapping;

[0035] The adjusted control signals are combined to generate a device control signal set that meets the control interface specifications of each virtual reality device.

[0036] The present invention is further configured such that forming a reflowable feedback representation comprises:

[0037] Receiving multi-channel and multi-dimensional feedback data generated by the virtual reality terminal after executing the control instruction;

[0038] Perform unified encoding and mapping on multi-channel and multi-dimensional feedback data to form a high-dimensional feedback tensor;

[0039] Based on the preset multimodal tensor mapping rules, the high-dimensional feedback tensor is transformed into a complex reconstruction of the spatial and channel dimensions to generate a structured feedback representation.

[0040] The structured feedback representation is mapped to the feedback input space preset by S1 to form a feedback closed-loop representation, which is used to realize closed-loop interaction and dynamic adaptation between the control and perception modules of the virtual reality system.

[0041] The present invention is further configured such that the structured feedback representation is composed of a feedback modal channel set, a semantic mapping index structure, and a time-series synchronization identification vector;

[0042] The feedback modality channel set is generated from the original feedback data through encoding mapping, maintaining structural consistency with the channel dimension of the multimodal input tensor;

[0043] The semantic mapping index structure describes the mapping relationship between each feedback channel and the semantic intent extraction path in the cross-modal perception module in a discrete indexing manner;

[0044] The timing synchronization identification vector is constructed with a fixed-length tensor, which explicitly calibrates the starting position, update frequency, and timing accuracy of the feedback signal in the sensing cycle according to a unified time scale.

[0045] The structured feedback representation is organized as a whole in a closed-loop mapping space, maintaining an alignable configuration with the input structure of the perception module, and achieving structural-level data reflow through a channel reference index mechanism.

[0046] The present invention also provides a VR interactive control management system, the system comprising:

[0047] Multimodal perception preprocessing module: This module obtains multi-source interaction data of users in a virtual reality environment, performs intra-modal representation conversion on each of the multi-source interaction data to form a modal vector sequence with a unified time scale and channel structure, and fuses them to generate a unified multimodal input tensor.

[0048] Cross-modal semantic extraction module: Based on the preset cross-modal perception structure, it extracts user semantic intent information from the unified multimodal input tensor to form a semantic intent representation;

[0049] Semantic control instruction generation module: combines the system internal state information at the previous moment with the current semantic intent representation to generate the corresponding control instruction vector;

[0050] Terminal control signal adaptation module: This module processes the control instruction vector through device channel mapping to generate a set of device control signals that conform to different VR terminal control interfaces.

[0051] Multi-dimensional feedback structure reconstruction module: It receives the multi-dimensional feedback information generated by the device after executing the control command, and restructures the feedback information to form a feedback representation that can flow back to S1, realizing closed-loop interaction between control and perception.

[0052] The present invention provides a VR interactive control management system and method. The method obtains multi-source interaction data of a user in a virtual reality environment, performs intra-modal representation conversion processing on the multi-source interaction data respectively, forms a modal vector sequence with a unified time scale and channel structure, and fuses the modal vector sequence to generate a unified multi-modal input tensor; based on a preset cross-modal perception structure, extracts user semantic intent information from the unified multi-modal input tensor to form a semantic intent representation; combines the system internal state information at the previous moment with the current semantic intent representation to generate a corresponding control instruction vector; performs device channel mapping processing on the control instruction vector to generate a set of device control signals that conform to different VR terminal control interfaces; receives multi-dimensional feedback information generated by the device after executing the control instruction, and performs structural reconstruction processing on the feedback information to form a feedback representation that can flow back to step one, thereby realizing closed-loop interaction between control and perception. The beneficial effects produced include:

[0053] 1. Achieve unified spatiotemporal structural modeling of multi-source interaction data: By performing intra-modal representation conversion and structural alignment on multi-source heterogeneous modalities such as the user's 3D gesture trajectories, voice signal features, eye movement behavior, and physiological responses, a unified modal vector sequence is formed and fused into a multimodal input tensor, providing a structurally consistent input foundation for subsequent semantic extraction and control generation.

[0054] 2. Enhance the dynamic responsiveness and structural adaptability of control command generation: By fusing the system's historical state with the current semantic representation to construct a semantic control representation tensor, a multi-order mapping mechanism is used to generate control command vectors with a dynamic channel activation structure. Furthermore, channel state mapping tables and interface protocols are used to achieve control compatibility with multiple types of VR terminals.

[0055] 3. Improve the extraction accuracy and representation consistency of cross-modal semantic information: By introducing a cross-modal perception structure and a modal cross-tensor construction mechanism, it is possible to effectively model high-order coupling relationships between modalities and generate semantic intent representations with fixed length and unified channel ordering through structural compression, ensuring that all types of modal information are synchronously aligned and structurally consistent at the semantic level.

[0056] The above description is only an overview of the technical solution of the present application. In order to more clearly understand the technical means of the present application, it can be implemented in accordance with the contents of the specification. In order to make the above and other purposes, features and advantages of the present application more obvious and easy to understand, the specific implementation methods of the present application are listed below. BRIEF DESCRIPTION OF THE DRAWINGS

[0057] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following briefly introduces the drawings required for describing the embodiments. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without inventive efforts. In the drawings:

[0058] Figure 1 This is a flow chart of a VR interactive control management method according to an exemplary embodiment of the present invention;

[0059] Figure 2 The figure is a structural diagram of a VR interactive control management system according to an exemplary embodiment of the present invention. DETAILED DESCRIPTION

[0060] The following describes the embodiments of the present invention with reference to the accompanying drawings and preferred embodiments. Those skilled in the art will readily appreciate the other advantages and benefits of the present invention from the disclosure herein. The present invention may also be implemented or applied through various other specific embodiments, and the various details in this specification may be modified or altered based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are intended only to illustrate the present invention and are not intended to limit the scope of protection of the present invention.

[0061] It should be noted that the illustrations provided in the following embodiments are merely schematic illustrations of the basic concept of the present invention. Therefore, the illustrations only show components related to the present invention and are not drawn according to the number, shape, and size of components in actual implementation. In actual implementation, the type, quantity, and proportion of each component may be changed arbitrarily, and the component layout may also be more complex.

[0062] In the following description, numerous details are discussed to provide a more thorough explanation of the embodiments of the present invention. However, it will be apparent to those skilled in the art that the embodiments of the present invention may be practiced without these specific details. In other embodiments, well-known structures and devices are shown in block diagram form rather than in detail to avoid obscuring the embodiments of the present invention.

[0063] Example 1

[0064] A VR interactive control management method, such as Figure 1 Shown, including:

[0065] S1: Obtain multi-source interaction data of users in a virtual reality environment, perform intra-modal representation conversion on each of the multi-source interaction data to form a modal vector sequence with a unified time scale and channel structure, and fuse them to generate a unified multi-modal input tensor;

[0066] S2: Based on the preset cross-modal perception structure, the user semantic intent information is extracted from the unified multimodal input tensor to form a semantic intent representation;

[0067] S3: Combine the system's internal state information at the previous moment with the current semantic intent representation to generate the corresponding control instruction vector;

[0068] S4: Perform device channel mapping on the control instruction vector to generate a set of device control signals that conform to different VR terminal control interfaces;

[0069] S5: Receives the multi-dimensional feedback information generated by the device after executing the control command, and restructures the feedback information to form a feedback representation that can flow back to S1, realizing a closed-loop interaction between control and perception.

[0070] The present invention is further configured such that the multi-source interaction data includes three-dimensional gesture trajectory data, voice signal feature data, eye movement behavior data, and physiological response data;

[0071] We construct independent modal representation spaces for multi-source interaction data, and align each modal data to a unified time scale based on the structure-preserving resampling mechanism. Specifically, for each modality, Construct modal tensor representations separately ,in, is the three-dimensional gesture trajectory mode, is the speech signal characteristic mode, is the eye movement behavior mode, is the physiological response mode, is the modal feature dimension, is the original time frame number of the modal;

[0072] For the original modal tensor Resample to generate a uniform time scale tensor , the resampling function is defined as: , ,in, For modal No. Channel, The resampled frame comes from the original Normalized mapping weights of frames, , Original frame The corresponding local structure position index in channel order, is the inter-modal neighborhood mapping function based on structural correlation, To unify the target time frame number, all modalities have the same frame number after resampling, avoiding the timing mismatch problem caused by inconsistent time scales between modalities. The unified alignment is achieved through the structure-preserving resampling mechanism;

[0073] The modal vector sequence with a unified channel structure is constructed through a cross-modal channel normalization mapping strategy. Perform channel alignment transformation to generate a uniform dimension vector sequence ,in, , It is the modal channel mapping matrix with unit column norm, responsible for converting features of different dimensions to a unified dimension , To unify the modal channel dimension, which does not change with the modality, and improve the structural compatibility and expression consistency between modalities, the channel normalization mapping strategy ensures the comparability between modal tensors;

[0074] A unified multimodal input tensor is generated by modal splicing and channel-level fusion. All modal vector sequences are spliced ​​and fused to generate the final unified multimodal input tensor. ,in, is the channel dimension splicing operation, It is a channel coupling fusion function, which is used to introduce structural redundancy removal operations between local channels to maintain information independence.

[0075] The present invention is further configured such that the generation of the semantic intent representation table includes:

[0076] The multimodal input tensor is split according to the modal dimension to form a set of modal sub-tensors. Specifically, the unified multimodal input tensor is split according to the modal dimension into , ,in, is the gesture modality subtensor, is the speech modality subtensor, is the eye movement modality subtensor, is the physiological modality subtensor, is the modal unified channel number, is the number of time frames;

[0077] The channel-level coupling relationship between each modal sub-tensor is bilinearly interactively mapped to form a time-distributed modal cross tensor structure. , construct its cross tensor ,in, For modal and In time 2D channel index The cross tensor at , For time frame The fourth-order mapping kernel function under is used to adjust the cross-importance between different channel combinations. Mode , In time Place , The value of the channel, after the cross tensor is constructed, all combined cross tensor sets are defined as , contains a total of six three-dimensional cross tensors, which uniformly constitute the set of modal cross structures;

[0078] The modal cross tensor is structurally compressed and a unified semantic vector representation is constructed.

[0079] The present invention is further configured such that the unified semantic vector representation is generated by compressing a modal cross tensor structure constructed between a plurality of modal sub-tensors;

[0080] The modal cross tensor structure is constructed by performing channel-level bilinear interaction between the modal sub-tensors obtained by splitting the unified multimodal input tensor according to the modal dimension;

[0081] Structural compression includes multi-dimensional compression mapping of modal cross tensors based on channel coupling mapping kernel at a fixed time scale, and fusing each compression representation to form a unified semantic vector representation. Specifically, for each cross tensor , perform channel compression and time aggregation mapping: ,in, For modal The compressed representation of is the structural compression kernel, used to map from the dual-channel coupling dimension to the semantic feature dimension , all Cascade and fuse to finally build a unified semantic vector representation , It is a structural feature fusion function with channel merging and compression characteristics. is the final semantic vector dimension, satisfying , is the number of time frames.

[0082] The present invention is further configured such that combining the system internal state information at the previous moment with the current semantic intent representation to generate a corresponding control instruction vector includes:

[0083] Align the system internal state information at the previous moment with the control history trajectory to construct a fusion state representation. Specifically, use the structure alignment mapping function , construct the fusion state representation: ,in, is the fusion state representation, To align the coupling kernel, dynamically adjust the matching degree between the state channel and the historical trajectory. is the system state vector The value of the dimension, represents the first The main channel and the Position elements of the auxiliary dimension, They are the system state vector channel index, the main channel index of the control history trajectory, and the auxiliary dimension index of the control history trajectory;

[0084] Perform channel matching and cross-embedding on the fusion state representation and the current semantic intent representation to construct the semantic regulation representation tensor: ,in, represents a tensor for semantic regulation, is the cross-channel chimera coefficient kernel, 、 is the channel dimension parameter after mapping, satisfying , used for subsequent control vector expansion, is the dimension of the fusion state representation vector, is the dimension of the semantic intent vector, Represents the current semantic intention;

[0085] The semantic control representation tensor is mapped through multi-order transformation to obtain the control instruction vector , and perform feedback correction and scale normalization on the instruction vector to obtain the processed .

[0086] The present invention is further configured such that the control instruction vector is a multidimensional vector having a hierarchical channel configuration structure, comprising a state fusion representation domain, a semantic response drive domain, and a channel control weight domain, each domain being arranged in order according to a preset semantic mosaic structure. Specifically, Parsed into three ordered structural subdomains ;

[0087] State fusion representation domain Characterize the structural alignment between the system's historical state and semantic input, and map the main components in the control instructions related to the historical state coupling: , characterizes the dependence of control reasoning on the system's internal state and historical trajectory, Extract the mapping function for the structure, extract the A dominant component, To control the mapping channel mosaic kernel, the channel matching rules are defined, the type is tensor kernel matrix, It represents the fusion state;

[0088] The semantic response driven domain encodes the structurally active region in the current semantic vector through channel index mapping, and extracts the high-response channel index structure of the current semantic vector in the mapping space: ,in, is the current semantic vector, encoding the user’s current intention, It is a semantically driven mapping function that extracts the high-response component of the semantically dominated channel and explicitly encodes the control-driven mode dominated by semantic intention, which has strong timeliness.

[0089] The channel control weight domain uses a non-overlapping index method to record the cross-channel dynamic response amplitude and structural allocation priority. The subdomains maintain structural independence through cross-domain mapping rules and share a unified channel index table. The current cross-channel dynamic amplitude and priority are recorded through non-overlapping index configuration. It serves to adjust the channel activation amplitude and map channel sorting. It does not participate in semantic logic modeling, the mapping is unique, and there is no overlap.

[0090] The present invention is further configured such that generating a set of device control signals conforming to different VR terminal control interfaces includes:

[0091] Perform multi-channel response structure expansion on the input control instruction vector to form a high-dimensional control response tensor. Specifically, the control instruction vector is expanded into a structure response tensor: , ,in, is the control instruction vector, is the structural expansion weight core matrix, is the expanded channel response width, Defined as a channel structure embedding function, it expands the control channel into a set of response channels by column broadcasting, with nonlinear adjustability. This structural expansion operation encodes the low-dimensional instruction vector into a high-dimensional tensor with a control signal distribution structure to support functional differences among multiple terminal devices.

[0092] Based on the preset device channel mapping rules, the control response tensor is mapped to the functional channels of multiple virtual reality terminals, and the response tensor is mapped to the function channels of multiple virtual reality terminals. Functional channel space projected onto each virtual reality terminal: , ,in, For the Function channel mapping matrix of VR terminals, For the The number of independent functional channels supported by each VR device, is the number of VR terminals supported, After mapping, it belongs to The original control signal vector of each terminal;

[0093] Perform structural strength adjustment and offset correction on the control signals of each terminal channel after mapping to obtain the corrected control signals ;

[0094] The adjusted control signals are aggregated to generate a device control signal set that meets the control interface specifications of each virtual reality device. Specifically, the adjustment signals of all terminals are aggregated: , is the final device control signal set, For the The identification code of a device can be MAC, UUID or device type. For the The adjusted control vector of the device, To construct a collection, it is not a simple splicing, but a structural packaging.

[0095] The present invention is further configured such that forming a reflowable feedback representation comprises:

[0096] Receiving multi-channel and multi-dimensional feedback data generated by the virtual reality terminal after executing the control instruction;

[0097] The multi-channel and multi-dimensional feedback data are uniformly encoded and mapped to form a high-dimensional feedback tensor. Specifically, the original tensor of the feedback channel received from the virtual reality terminal is , after unified coding mapping function Get a high-dimensional feedback tensor , For the Each terminal feedback channel data, is the number of feedback channels, is the feedback dimension of each channel, is a multi-source channel fusion encoder, is the feedback encoding matrix, is the spatial dimension after reconstruction, To reconstruct the posterior channel dimension;

[0098] Based on the preset multimodal tensor mapping rules, the high-dimensional feedback tensor is transformed into complex reconstruction of space and channel dimensions to generate structured feedback representation. Perform structural reconstruction and define the reconstruction function as , , the specific mapping formula is: , is a structure reconstruction function with nonlinear cross expansion effect. Reconstruct the weight system for the channel space, is the spatial scale of the reconstructed structure, is the channel scale of the reconstructed structure, Ensure alignment restoration of the original structure for the modulo operation;

[0099] The structured feedback representation is mapped to the feedback input space preset by S1 to form a feedback closed-loop representation, which is used to realize closed-loop interaction and dynamic adaptation between the control and perception modules of the virtual reality system.

[0100] The present invention is further configured such that the structured feedback representation is composed of a feedback modal channel set, a semantic mapping index structure, and a time-series synchronization identification vector;

[0101] The feedback modal channel set is generated by encoding mapping from the original feedback data, maintaining the structural consistency with the multimodal input tensor channel dimension, and is expressed as: ,in, is the modal structure extraction function, is the number of structured feedback channels;

[0102] The semantic mapping index structure describes the mapping relationship between each feedback channel and the semantic intent extraction path in the cross-modal perception module in a discrete indexing manner, and the expression is: ,in, Extract the mapping for the channel index, is the channel semantic index, is the set of participating channels for mapping;

[0103] The timing synchronization identification vector is constructed with a fixed-length tensor, and the starting position, update frequency and timing accuracy of the feedback signal in the perception cycle are explicitly calibrated according to a unified time scale. The expression is: , is the starting step identifier, is the update frequency, is the timing precision factor;

[0104] The structured feedback representation is organized as a whole in a closed-loop mapping space, maintaining an alignable configuration with the input structure of the perception module, and achieving structural-level data reflow through a channel reference index mechanism.

[0105] Example 1

[0106] See also Figure 2 , the exemplary VR interactive control management system includes:

[0107] Multimodal perception preprocessing module: This module obtains multi-source interaction data of users in a virtual reality environment, performs intra-modal representation conversion on each of the multi-source interaction data to form a modal vector sequence with a unified time scale and channel structure, and fuses them to generate a unified multimodal input tensor.

[0108] Cross-modal semantic extraction module: Based on the preset cross-modal perception structure, it extracts user semantic intent information from the unified multimodal input tensor to form a semantic intent representation;

[0109] Semantic control instruction generation module: combines the system internal state information at the previous moment with the current semantic intent representation to generate the corresponding control instruction vector;

[0110] Terminal control signal adaptation module: This module processes the control instruction vector through device channel mapping to generate a set of device control signals that conform to different VR terminal control interfaces.

[0111] Multi-dimensional feedback structure reconstruction module: It receives the multi-dimensional feedback information generated by the device after executing the control command, and restructures the feedback information to form a feedback representation that can flow back to S1, realizing closed-loop interaction between control and perception.

[0112] It should be noted that the VR interactive control management system provided in the above embodiment and the VR interactive control management method provided in the above embodiment are based on the same concept, wherein the specific manner in which each module and unit performs operations has been described in detail in the method embodiment and will not be repeated here. In actual applications, the VR interactive control management system provided in the above embodiment can allocate the above functions to different functional modules as needed, that is, divide the internal structure of the system into different functional modules to complete all or part of the functions described above, and this is not limited here.

[0113] The above embodiments can be implemented in whole or in part by software, hardware, firmware or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer program are loaded or executed on a computer, the process or function described in the embodiment of the present application is generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server or data center to another website, computer, server or data center via a wired (e.g., infrared, wireless, microwave, etc.) method. The computer-readable storage medium can be any available medium that can be accessed by a computer or a data storage device such as a server or data center that contains one or more available media sets. The available medium can be a magnetic medium (e.g., a floppy disk, a hard disk, a tape), an optical medium (e.g., a DVD), or a semiconductor medium. The semiconductor medium can be a solid-state drive.

[0114] It should be understood that the term "and / or" as used herein simply describes a relationship between associated objects, indicating that three possible relationships exist. For example, "A and / or B" can represent: A alone, A and B together, or B alone. A and B can be singular or plural. Furthermore, the character " / " as used herein generally indicates an "or" relationship between the associated objects, but it may also indicate an "and / or" relationship. For specific understanding, please refer to the context.

[0115] In this application, "at least one" means one or more, and "plurality" means two or more. "At least one of the following" or similar expressions refers to any combination of these items, including any combination of single or plural items. For example, at least one of a, b, or c can mean: a, b, c, ab, ac, bc, or abc, where a, b, and c can be single or plural.

[0116] It should be understood that in the various embodiments of the present application, the size of the serial numbers of the above-mentioned processes does not mean the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present application.

[0117] Those skilled in the art will appreciate that the units and algorithm steps of each example described in conjunction with the embodiments disclosed herein can be implemented in electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are performed in hardware or software depends on the specific application and design constraints of the technical solution. Professional and technical personnel can use different methods to implement the described functions for each specific application, but such implementation should not be considered beyond the scope of this application.

[0118] Those skilled in the art will clearly understand that, for the convenience and brevity of description, the specific working processes of the systems, devices and units described above can refer to the corresponding processes in the aforementioned method embodiments and will not be repeated here.

[0119] In the several embodiments provided in this application, it should be understood that the disclosed system can be implemented in other ways. For example, the device embodiments described above are merely schematic. For example, the division of the units is merely a logical function division. In actual implementation, there may be other division methods, such as multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the mutual coupling or direct coupling or communication connection shown or discussed can be through some interfaces, indirect coupling or communication connection of devices or units, which can be electrical, mechanical or other forms.

[0120] The units described as separate components may or may not be physically separate, and the components shown as units may or may not be physical units, that is, they may be located in one place or distributed across multiple network units. Some or all of these units may be selected to achieve the purpose of this embodiment according to actual needs.

[0121] In addition, each functional unit in each embodiment of the present application may be integrated into one processing unit, or each unit may exist physically separately, or two or more units may be integrated into one unit.

[0122] If the functions are implemented in the form of software functional units and sold or used as independent products, they can be stored in a computer-readable storage medium. Based on this understanding, the technical solution of the present application, or the part that contributes to the prior art, or the part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for enabling a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the method described in each embodiment of the present application. The aforementioned storage medium includes various media that can store program codes, such as a USB flash drive, a mobile hard disk, a read-only memory (ROM), a random access memory (RAM), a magnetic disk or an optical disk.

[0123] The above description is merely a specific embodiment of the present application, but the scope of protection of the present application is not limited thereto. Any changes or substitutions that can be easily conceived by a person skilled in the art within the technical scope disclosed in this application should be included in the scope of protection of this application. Therefore, the scope of protection of this application should be based on the scope of protection of the claims.

Claims

1. A VR interactive control management method, characterized in that: include: S1: Obtain multi-source interaction data of users in a virtual reality environment, perform intra-modal representation conversion on each of the multi-source interaction data to form a modal vector sequence with a unified time scale and channel structure, and fuse them to generate a unified multi-modal input tensor; S2: Based on the preset cross-modal perception structure, the user semantic intent information is extracted from the unified multimodal input tensor to form a semantic intent representation; S3: Combine the system's internal state information at the previous moment with the current semantic intent representation to generate the corresponding control instruction vector; S4: Perform device channel mapping on the control instruction vector to generate a set of device control signals that conform to different VR terminal control interfaces; S5: Receives the multi-dimensional feedback information generated by the device after executing the control command, and restructures the feedback information to form a feedback representation that can flow back to S1, realizing a closed-loop interaction between control and perception.

2. A VR interactive control management method according to claim 1, characterized in that: Multi-source interaction data includes three-dimensional gesture trajectory data, voice signal feature data, eye movement behavior data, and physiological response data; Independent modal representation spaces are constructed for multi-source interaction data, and each modal data is aligned to a unified time scale based on a structure-preserving resampling mechanism. A modal vector sequence with a unified channel structure is constructed through a cross-modal channel normalization mapping strategy; Generate a unified multimodal input tensor through modality concatenation and channel-level fusion.

3. A VR interactive control management method according to claim 2, characterized in that: The generation of semantic intent representation table includes: Split the multimodal input tensor according to the modal dimension to form a set of modal sub-tensors; Perform bilinear interactive mapping on the channel-level coupling relationship between each modal sub-tensor to form a time-distributed modal cross tensor structure; The modal cross tensor is structurally compressed and a unified semantic vector representation is constructed.

4. A VR interactive control management method according to claim 3, characterized in that: The unified semantic vector representation is generated by structural compression of the modal cross tensor structure constructed between multiple modal sub-tensors; The modal cross tensor structure is constructed by performing channel-level bilinear interaction between the modal sub-tensors obtained by splitting the unified multimodal input tensor according to the modal dimension; Structural compression includes multi-dimensional compression mapping of the modal cross tensor based on the channel coupling mapping kernel at a fixed time scale, and fusing the compressed representations to form a unified semantic vector representation.

5. A VR interactive control management method according to claim 4, characterized in that: Combining the system's internal state information at the previous moment with the current semantic intent representation, the corresponding control instruction vector is generated, including: Align the system's internal state information at the previous moment with the control history trajectory to construct a fusion state representation; Perform channel matching and cross-embedding between the fusion state representation and the current semantic intent representation to construct a semantic regulation representation tensor; The semantic control representation tensor is mapped through multi-order transformation to obtain a control instruction vector, and the instruction vector is subjected to feedback correction and scale normalization.

6. A VR interactive control management method according to claim 5, characterized in that: The control instruction vector is a multidimensional vector with a hierarchical channel configuration structure, which includes a state fusion representation domain, a semantic response drive domain, and a channel control weight domain. Each domain is arranged in an orderly manner according to a preset semantic mosaic structure. State fusion represents the structural alignment between the historical state of the domain representation system and the semantic input; The semantic response driving domain encodes the structurally active regions in the current semantic vector through channel index mapping; The channel control weight domain uses a non-overlapping index method to record the cross-channel dynamic response amplitude and structure allocation priority. The subdomains maintain structural independence through cross-domain mapping rules and share a unified channel index table.

7. A VR interactive control management method according to claim 5, characterized in that: The device control signal set generated to comply with different VR terminal control interfaces includes: Perform multi-channel response structure expansion on the input control instruction vector to form a high-dimensional control response tensor; Based on the preset device channel mapping rules, the control response tensor is mapped to the functional channels of multiple virtual reality terminals; Perform structural strength adjustment and offset correction on the control signals of each terminal channel after mapping; The adjusted control signals are combined to generate a device control signal set that meets the control interface specifications of each virtual reality device.

8. A VR interactive control management method according to claim 7, characterized in that: Forming a reflowable feedback representation includes: Receiving multi-channel and multi-dimensional feedback data generated by the virtual reality terminal after executing the control instruction; Perform unified encoding and mapping on multi-channel and multi-dimensional feedback data to form a high-dimensional feedback tensor; Based on the preset multimodal tensor mapping rules, the high-dimensional feedback tensor is transformed into a complex reconstruction of the spatial and channel dimensions to generate a structured feedback representation. The structured feedback representation is mapped to the feedback input space preset by S1 to form a feedback closed-loop representation, which is used to realize closed-loop interaction and dynamic adaptation between the control and perception modules of the virtual reality system.

9. A VR interactive control management method according to claim 8, characterized in that: The structured feedback representation consists of a feedback modality channel set, a semantic mapping index structure, and a temporal synchronization identification vector; The feedback modality channel set is generated from the original feedback data through encoding mapping, maintaining structural consistency with the channel dimension of the multimodal input tensor; The semantic mapping index structure describes the mapping relationship between each feedback channel and the semantic intent extraction path in the cross-modal perception module in a discrete indexing manner; The timing synchronization identification vector is constructed with a fixed-length tensor, which explicitly calibrates the starting position, update frequency, and timing accuracy of the feedback signal in the sensing cycle according to a unified time scale. The structured feedback representation is organized as a whole in a closed-loop mapping space, maintaining an alignable configuration with the input structure of the perception module, and achieving structural-level data reflow through a channel reference index mechanism.

10. A VR interactive control management system, used to implement a VR interactive control management method according to any one of claims 1 to 9, characterized in that: include: Multimodal perception preprocessing module: This module obtains multi-source interaction data of users in a virtual reality environment, performs intra-modal representation conversion on each of the multi-source interaction data to form a modal vector sequence with a unified time scale and channel structure, and fuses them to generate a unified multimodal input tensor. Cross-modal semantic extraction module: Based on the preset cross-modal perception structure, it extracts user semantic intent information from the unified multimodal input tensor to form a semantic intent representation; Semantic control instruction generation module: combines the system internal state information at the previous moment with the current semantic intent representation to generate the corresponding control instruction vector; Terminal control signal adaptation module: This module processes the control instruction vector through device channel mapping to generate a set of device control signals that conform to different VR terminal control interfaces. Multi-dimensional feedback structure reconstruction module: It receives the multi-dimensional feedback information generated by the device after executing the control command, and restructures the feedback information to form a feedback representation that can flow back to S1, realizing closed-loop interaction between control and perception.

Citation Information

Patent Citations

  • Multi-modal information processing and interaction system

    CN112613534A

  • Human-computer interaction method and device and storage medium

    CN119292459A