VR interactive control management system and method

Through multimodal data fusion and closed-loop interactive processing, the multimodal data fusion and cross-modal semantic understanding problems in existing VR interaction control technology are solved, dynamic perception and response control in high-complex virtual reality applications are realized, and the system's adaptive regulation capabilities are improved.

CN120406748AActive Publication Date: 2025-08-01HANGZHOU KAILIN CULTURE TECHNOLOGY CO LTD +1

Patent Information

Application Number
CN202510907593.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-07-02
Publication Date
2025-08-01
Estimated Expiration
2045-07-02

AI Technical Summary

Technical Problem

The existing VR interactive control technology lacks multimodal data fusion and cross-modal semantic understanding capabilities, and it is difficult to support dynamic perception and response control in complex contexts. There are limitations in the scalability and flexibility of the generation of control signals. The feedback information lacks structural description, making it difficult to realize the adaptive regulation of the system.

Method used

By obtaining multi-source interactive data, the modal representation conversion and fusion are performed, a unified multi-modal input tensor is generated, semantic intention information is extracted, control instructions are generated in combination with system state, device channel mapping is performed, and feedback information is structurally reconstructed to realize closed-loop interaction between control and perception.

Benefits of technology

It realizes unified spatio-temporal structure modeling of multi-source interactive data, enhances the dynamic response ability and structural adaptability of control instructions, improves the extraction accuracy and characterization consistency of cross-modal semantic information, and supports high-complexity and multi-dimensional virtual reality applications.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120406748A_ABST
    Figure CN120406748A_ABST
Patent Text Reader

Abstract

The invention discloses a VR interactive control management system and method, and relates to the technical field of virtual reality. The method comprises the steps of obtaining multi-source interaction data of a user in a virtual reality environment, performing intra-modal representation conversion and fusion, and constructing a unified multi-modal input tensor; semantic intention representation is extracted based on the cross-modal perception structure; generating a control instruction vector in combination with the historical state information and the current semantic intention; mapping the control instruction vector into an equipment control signal set conforming to various VR terminal interface specifications; after the equipment executes the control instruction, multi-dimensional feedback information is collected, the structure of the multi-dimensional feedback information is reconstructed, feedback representation capable of flowing back to the sensing module is generated, and closed-loop interaction between sensing and control is achieved. Through unifying a multi-source interaction data structure, a dynamically adaptive control instruction vector and a high-precision cross-modal semantic representation mechanism are constructed, and a closed-loop interaction process with consistent sensing and control structures, flexible response and semantic alignment in a virtual reality system is realized.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the technical field of virtual reality, and particularly to a VR interactive control management system and method. Background Technique

[0002] As an important form of the new generation of human-computer interaction, virtual reality technology has been widely applied in many fields such as education and training, medical rehabilitation, remote collaboration, and intelligent manufacturing. In a virtual reality system, a user interacts with a virtual environment through various sensing devices. The system needs to generate control responses in a timely manner according to the user's behavioral inputs and effectively sense and process the terminal feedback, so as to achieve an immersive and highly real-time interaction experience. Existing VR interaction control technologies mainly focus on the recognition and response of single-modal inputs (such as gestures, voices), lacking the ability of unified modeling and deep semantic understanding of multi-source heterogeneous data, and it is difficult to support the dynamic sensing and response control requirements in complex contexts.

[0003] In addition, existing solutions generally adopt static channel mapping or preset instruction templates, lacking an adaptation mechanism for the differences between different terminal control interfaces, resulting in obvious limitations in the scalability and flexibility of the generation of control signals. In the control closed-loop, feedback information is mostly transmitted back in the form of a single physical state quantity, lacking structural and semantic descriptions, and it is difficult to achieve direct alignment and effective utilization between the feedback data and the sensing module, which is not conducive to building the adaptive regulation ability of the system.

[0004] Therefore, there is an urgent need for a VR interactive control management method and system with the capabilities of multi-modal data fusion, cross-modal semantic extraction, state-driven control generation, channel-level control signal mapping, and structured feedback return to support high-complexity, multi-dimensional, and strongly interactive virtual reality application scenarios. Summary of the Invention

[0005] Based on the above-mentioned disadvantages of the prior art, the purpose of the present invention is to provide a VR interactive control management system and method to solve the above technical problems.

[0006] To achieve the above purpose, the present invention provides the following technical solution: A VR interactive control management method, including: S1: Obtain multi-source interaction data of a user in a virtual reality environment, respectively perform in-modal representation conversion processing on the multi-source interaction data to form a modal vector sequence with a unified time scale and channel structure, and perform fusion processing on it to generate a unified multi-modal input tensor; S2: Based on a preset cross-modal perception structure, extract user semantic intention information from the unified multi-modal input tensor to form a semantic intention representation; S3: Combine the system internal state information of the previous moment with the current semantic intention representation to generate a corresponding control instruction vector; S4: Perform device channel mapping processing on the control instruction vector to generate a set of device control signals that conform to different VR terminal control interfaces; S5: Receive the multi-dimensional feedback information generated by the device after executing the control instruction, and perform structure reconstruction processing on the feedback information to form a feedback representation that can flow back to S1, realizing the closed-loop interaction between control and perception.

[0007] The present invention is further configured such that the multi-source interaction data includes three-dimensional gesture trajectory data, voice signal feature data, eye movement behavior data, and physiological response data; Construct independent modal representation spaces for the multi-source interaction data respectively, and align each modal data to a unified time scale based on a structure-preserving resampling mechanism; Construct a modal vector sequence with a unified channel structure through a cross-modal channel normalization mapping strategy; Generate a unified multi-modal input tensor through modal splicing and channel-level fusion.

[0008] The present invention is further configured such that the generation of the semantic intention characterization table includes: Slice the multi-modal input tensor according to the modal dimension to form a set of modal sub-tensors; Perform bilinear interaction mapping on the channel-level coupling relationship between each modal sub-tensor to form a time-distributed modal cross-tensor structure; Compress the structure of the modal cross-tensor and construct a unified semantic vector representation.

[0009] The present invention is further configured such that the unified semantic vector representation is generated by structure compression of the modal cross-tensor structure constructed between multiple modal sub-tensors; The modal cross-tensor structure is constructed by performing channel-level bilinear interaction between each modal sub-tensor obtained by slicing the unified multi-modal input tensor according to the modal dimension; Structure compression includes performing multi-dimensional compression mapping on the modal cross-tensor based on a channel coupling mapping kernel at a fixed time scale, and fusing each compressed representation to form a unified semantic vector representation.

[0010] The present invention is further configured such that the generation of the corresponding control instruction vector by combining the system internal state information of the previous moment and the current semantic intention characterization includes: Align the system internal state information of the previous moment with the control history trajectory to construct a fused state representation; Perform channel matching cross-embedding on the fused state representation and the current semantic intention characterization to construct a semantic regulation representation tensor; Obtain the control instruction vector by performing multi-order transformation mapping on the semantic regulation representation tensor, and perform feedback correction and scale normalization processing on the instruction vector.

[0011] The present invention is further configured such that the control instruction vector is a multi-dimensional vector having a hierarchical channel configuration structure, including a state fusion representation domain, a semantic response driving domain, and a channel regulation weight domain, and each domain is arranged in an orderly manner according to a preset semantic embedding structure; The state fusion representation domain represents the structural alignment relationship between the system historical state and the semantic input; The semantic response driving domain encodes the structurally active regions in the current semantic vector through a channel index mapping method; The channel regulation weight domain records the cross-channel dynamic response amplitude and the structural allocation priority by using a non-overlapping index method, and the structural independence is maintained between each sub-domain, and a unified channel index table is shared through a cross-domain mapping rule.

[0012] The present invention is further configured such that the set of device control signals that conform to different VR terminal control interfaces includes: Performing a multi-channel response structure expansion on the input control instruction vector to form a high-dimensional control response tensor; Based on a preset device channel mapping rule, mapping the control response tensor to the functional channels of multiple virtual reality terminals; Performing a structural strength adjustment and offset correction on the control signals of each terminal channel after mapping; Combining the adjusted control signals to generate a set of device control signals that meet the control interface specifications of each virtual reality device.

[0013] The present invention is further configured such that the formation of a feedback representation that can be refluxed includes: Receiving multi-channel multi-dimensional feedback data generated by a virtual reality terminal after executing a control instruction; Performing a unified coding mapping on the multi-channel multi-dimensional feedback data to form a high-dimensional feedback tensor; Based on a preset multi-modal tensor mapping rule, performing complex reconstruction conversion on the high-dimensional feedback tensor in the spatial and channel dimensions to generate a structured feedback representation; Mapping the structured feedback representation to a preset feedback input space S1 to form a feedback closed-loop representation for realizing closed-loop interaction and dynamic adaptation between the control and perception modules of the virtual reality system.

[0014] The present invention is further configured such that the structured feedback representation is composed of a feedback modal channel set, a semantic mapping index structure, and a time sequence synchronization identification vector; The feedback modal channel set is generated by encoding and mapping the original feedback data, and maintains the structural consistency with the channel dimension of the multi-modal input tensor; The semantic mapping index structure describes the mapping relationship between each feedback channel and the semantic intention extraction path in the cross-modal perception module in a discrete indexing manner; The timing synchronization identification vector is constructed with a fixed-length tensor, and explicitly calibrates the start position, update frequency, and timing accuracy of the feedback signal in the sensing period according to a unified time scale; The structured feedback representation is overall organized in the closed-loop mapping space, maintains an alignable configuration with the input structure of the sensing module, and realizes data reflux at the structural level through the channel reference indexing mechanism.

[0015] The present invention also provides a VR interactive control and management system, and the system includes: Multimodal perception preprocessing module: Obtain multi-source interaction data of the user in the virtual reality environment, perform intra-modal representation conversion processing on the multi-source interaction data respectively to form a modal vector sequence with a unified time scale and channel structure, and perform fusion processing on it to generate a unified multimodal input tensor; Cross-modal semantic extraction module: Based on a preset cross-modal perception structure, extract user semantic intention information from the unified multimodal input tensor to form a semantic intention representation; Semantic regulation instruction generation module: Combine the system internal state information at the previous moment and the current semantic intention representation to generate a corresponding control instruction vector; Terminal control signal adaptation module: Perform device channel mapping processing on the control instruction vector to generate a set of device control signals that conform to different VR terminal control interfaces; Multidimensional feedback structure reconstruction module: Receive the multidimensional feedback information generated after the device executes the control instruction, and perform structure reconstruction processing on the feedback information to form a feedback representation that can flow back to S1, realizing the closed-loop interaction between control and perception.

[0016] The present invention provides a VR interactive control and management system and method. The method obtains multi-source interaction data of the user in the virtual reality environment, performs intra-modal representation conversion processing on the multi-source interaction data respectively to form a modal vector sequence with a unified time scale and channel structure, and performs fusion processing on it to generate a unified multimodal input tensor; based on a preset cross-modal perception structure, extract user semantic intention information from the unified multimodal input tensor to form a semantic intention representation; combine the system internal state information at the previous moment and the current semantic intention representation to generate a corresponding control instruction vector; perform device channel mapping processing on the control instruction vector to generate a set of device control signals that conform to different VR terminal control interfaces; receive the multidimensional feedback information generated after the device executes the control instruction, and perform structure reconstruction processing on the feedback information to form a feedback representation that can flow back to step one, realizing the closed-loop interaction between control and perception. The beneficial effects generated include: 1. Implement unified spatio-temporal structure modeling for multi-source interaction data: By performing in-modal representation conversion and structure-preserving alignment on multi-source heterogeneous modalities such as three-dimensional gesture trajectories, speech signal features, eye movement behaviors, and physiological responses from users, a unified modal vector sequence is formed and fused into a multi-modal input tensor, providing a structurally consistent input basis for subsequent semantic extraction and control generation. 2. Enhance the dynamic response ability and structural adaptability of control instruction generation: By fusing the system historical state and the current semantic representation to construct a semantic regulation representation tensor, a multi-order mapping mechanism is used to generate a control instruction vector with a dynamic channel activation structure, and control compatibility with multiple types of VR terminals is achieved through a channel state mapping table and an interface protocol. 3. Improve the extraction accuracy and representation consistency of cross-modal semantic information: By introducing a cross-modal perception structure and a modal cross-tensor construction mechanism, the high-order coupling relationship between modalities can be effectively modeled, and a semantic intention representation with a fixed length and unified channel sorting is generated through a structure compression method, ensuring synchronous alignment and structurally consistent expression of various modal information at the semantic level.

[0017] The above description is only an overview of the technical solution of this application. In order to be able to understand the technical means of this application more clearly, it can be implemented according to the content of the specification. And in order to make the above and other purposes, features, and advantages of this application more obvious and understandable, the specific embodiments of this application are specifically given below. BRIEF DESCRIPTION OF THE DRAWINGS

[0018] In order to more clearly illustrate the technical solutions in the embodiments of the present invention, the following will briefly introduce the drawings required for the description of the embodiments. Obviously, the drawings in the following description are only some embodiments of the present invention. For those of ordinary skill in the art, without creative efforts, other drawings can be obtained based on these drawings. In the drawings: Figure 1 It is a flowchart of a VR interaction control management method shown in an exemplary embodiment of the present invention; Figure 2 It is a structural schematic diagram of a VR interaction control management system shown in an exemplary embodiment of the present invention. DETAILED DESCRIPTION OF THE EMBODIMENTS

[0019] The following will describe the embodiments of the present invention with reference to the drawings and preferred embodiments. Those skilled in the art can easily understand the other advantages and effects of the present invention from the content disclosed in this specification. The present invention can also be implemented or applied through other different specific embodiments, and various details in this specification can also be modified or changed based on different viewpoints and applications without departing from the spirit of the present invention. It should be understood that the preferred embodiments are only for explaining the present invention, rather than for limiting the protection scope of the present invention.

[0020] It should be noted that the illustrations provided in the following embodiments only illustrate the basic concept of the present invention schematically. Therefore, only the components related to the present invention are shown in the drawings, rather than being drawn according to the number, shape, and size of the components in actual implementation. The form, quantity, and proportion of each component in actual implementation can be arbitrarily changed, and the component layout form may also be more complex.

[0021] In the following description, a large number of details are explored to provide a more thorough explanation of the embodiments of the present invention. However, it is obvious to those skilled in the art that the embodiments of the present invention can be implemented without these specific details. In other embodiments, well-known structures and devices are shown in the form of block diagrams rather than in detail to avoid making the embodiments of the present invention difficult to understand.

[0022] Embodiment 1

[0023] A VR interactive control management method, as Figure 1 shown, includes: S1: Obtain multi-source interaction data of the user in the virtual reality environment, perform intra-modal representation conversion processing on the multi-source interaction data respectively to form a modal vector sequence with a unified time scale and channel structure, and perform fusion processing on it to generate a unified multi-modal input tensor; S2: Based on a preset cross-modal perception structure, extract the user's semantic intention information from the unified multi-modal input tensor to form a semantic intention representation; S3: Combine the system internal state information at the previous moment with the current semantic intention representation to generate a corresponding control instruction vector; S4: Perform device channel mapping processing on the control instruction vector to generate a set of device control signals that conform to different VR terminal control interfaces; S5: Receive the multi-dimensional feedback information generated after the device executes the control instruction, and perform structure reconstruction processing on the feedback information to form a feedback representation that can flow back to S1, realizing the closed-loop interaction between control and perception.

[0024] The present invention is further configured such that the multi-source interaction data includes three-dimensional gesture trajectory data, voice signal feature data, eye movement behavior data, and physiological response data; Respectively construct independent modal representation spaces for the multi-source interaction data, and align each modal data to a unified time scale based on a structure-preserving resampling mechanism. Specifically, for each modality respectively construct a modal tensor representation , where is the three-dimensional gesture trajectory modality, is the voice signal feature modality, is the eye movement behavior modality, is the physiological response mode, is the modal feature dimension, is the original number of modal time frames; Resample the original modal tensor to generate a unified time-scale tensor , and the resampling function is defined as: , , where is the th channel, the th resampled frame comes from the normalized mapping weight of the original th frame, , is the local structure position index corresponding to the original frame in the channel order, is the inter-modal neighborhood mapping function based on structural correlation, is the unified target number of time frames. After resampling, the number of frames of all modalities is the same, avoiding the timing mismatch problem caused by inconsistent time scales between modalities, and realizing unified alignment through the structure-preserving resampling mechanism; Construct a sequence of modal vectors with a unified channel structure through a cross-modal channel normalization mapping strategy, and perform channel alignment conversion on each resampled modal tensor to generate a sequence of unified-dimension vectors , where , is the modal channel mapping matrix with unit column norm, which is responsible for converting features of different dimensions to a unified dimension , is the unified modal channel dimension, which does not change with the modality, improving the structural compatibility and expression consistency between cross-modalities. The channel normalization mapping strategy ensures comparability between modal tensors; Generate a unified multi-modal input tensor through modal concatenation and channel-level fusion. All sequences of modal vectors are concatenated and fused to generate the final unified multi-modal input tensor , where is the channel dimension concatenation operation, is the channel coupling fusion function, which is used to introduce the operation of removing structural redundancy between local channels and maintain information independence.

[0025] This invention is further set as that the generation of the semantic intention representation table includes: Slice the multi-modal input tensor along the modal dimension to form a set of modal sub-tensors. Specifically, slice the unified multi-modal input tensor along the modal dimension into , , where is the gesture modality sub-tensor, is the speech modality sub-tensor, is the eye movement modality sub-tensor, is the physiological modality sub-tensor, is the modality unified channel number, is the number of time frames; Perform a bilinear interaction mapping on the channel-level coupling relationship between each modality sub-tensor to form a time-distributed modality cross-tensor structure. For any two modalities , construct its cross-tensor , where is the modality and at time two-dimensional channel index at the cross-tensor, is the fourth-order mapping kernel function under time frame , used to adjust the cross-importance between different channel combinations, are respectively the modalities , at time at the , channel values. After the cross-tensor construction is completed, all combined cross-tensor sets are defined as , a total of six three-dimensional cross-tensors, which together constitute the modality cross-structure set; Compress the structure of the modality cross-tensor and construct a unified semantic vector representation.

[0026] The present invention is further configured such that the unified semantic vector representation is generated by structure compression of the modality cross-tensor structure constructed between multiple modality sub-tensors; The modality cross-tensor structure is constructed by performing channel-level bilinear interaction between the modality sub-tensors obtained by slicing the unified multi-modal input tensor along the modality dimension; Structure compression includes performing multi-dimensional compression mapping on the modality cross-tensor based on the channel coupling mapping kernel at a fixed time scale, and fusing the compressed representations to form a unified semantic vector representation. Specifically, for each cross-tensor , perform channel compression and time aggregation mapping: , where is the compressed representation of the modality , is the structure compression kernel, used to map from the two-channel coupling dimension to the semantic feature dimension , cascade and fuse all to finally construct the unified semantic vector representation , is the structural feature fusion function, with channel merging and compression characteristics, is the dimension of the final semantic vector, satisfying , is the number of time frames.

[0027] The present invention is further configured such that the generation of the corresponding control instruction vector by combining the system internal state information of the previous moment with the current semantic intention representation includes: Structurally align the system internal state information of the previous moment with the control history trajectory to construct a fused state representation. Specifically, use the structural alignment mapping function to construct the fused state representation: , where is the fused state representation, is the alignment coupling kernel, dynamically adjusting the matching degree between the state channel and the history trajectory, is the value of the -th dimension of the system state vector, represents the position element of the -th main channel and the -th auxiliary dimension in the control trajectory tensor, are respectively the system state vector channel index, the main channel index of the control history trajectory, and the auxiliary dimension index of the control history trajectory; Perform channel matching cross-embedding on the fused state representation and the current semantic intention representation to construct a semantic regulation representation tensor, and construct the semantic regulation tensor: , where is the semantic regulation representation tensor, is the cross-channel embedding coefficient kernel, , are the mapped channel dimension parameters, satisfying , for subsequent control vector expansion, is the dimension of the fused state representation vector, is the dimension of the semantic intention vector, is the current semantic intention representation; Map the semantic regulation representation tensor through multi-order transformation to obtain the control instruction vector , and perform feedback correction and scale normalization processing on the instruction vector to obtain the processed .

[0028] The present invention is further configured such that the control instruction vector is a multi-dimensional vector with a hierarchical channel configuration structure, including a state fusion representation domain, a semantic response driving domain, and a channel regulation weight domain. Each domain is arranged in an orderly manner according to a preset semantic embedding structure. Specifically, is parsed into three ordered structural sub-domains ; State fusion representation domain Characterize the structural alignment relationship between the historical state of the system and the semantic input, and map the main components related to the coupling with the historical state in the control instruction: , characterize the dependence characteristics of the control inference on the internal state and historical trajectory of the system, is a structure extraction mapping function, extracting the th dominant component, is the chimeric kernel of the control mapping channel, defining the channel matching rule, and the type is a tensor kernel matrix, is the fused state representation; The semantic response driving domain encodes the structurally active area in the current semantic vector through the channel index mapping method, and extracts the high-response channel index structure of the current semantic vector in the mapping space: , where is the current semantic vector, encoding the user's current intention, is the semantic driving mapping function, extracting the high-response components of the semantic dominant channels, explicitly encoding the control driving mode dominated by semantic intentions, and having strong timeliness; The channel regulation weight domain records the cross-channel dynamic response amplitude and the structural allocation priority in a non-overlapping index manner. The sub-domains maintain structural independence through the cross-domain mapping rule and share a unified channel index table. Through the non-overlapping index configuration, the current cross-channel dynamic amplitude and priority are recorded, and the function is to adjust the channel activation amplitude and map the channel sorting, and do not participate in the semantic logic modeling, and the mapping is unique and non-overlapping.

[0029] The present invention is further configured such that the set of device control signals that conform to different VR terminal control interfaces includes: Perform multi-channel response structure expansion on the input control instruction vector to form a high-dimensional control response tensor. Specifically, expand the control instruction vector into a structural response tensor: , , where is the control instruction vector, is the structure expansion weight kernel matrix, is the expanded channel response width, is defined as the channel structure embedding function, and the control channel is broadcast by column to expand into a set of response channels, with non-linear adjustability. This structure expansion operation encodes the low-dimensional instruction vector into a high-dimensional tensor with a control signal distribution structure, which is used to support the functional differences of multi-terminal devices; Based on the preset device channel mapping rule, map the control response tensor to the functional channels of multiple virtual reality terminals, and project the response tensor into the functional channel space of each virtual reality terminal: , , where is the The functional channel mapping matrix of a VR terminal is the number of independent functional channels supported by the th VR device, is the number of supported VR terminals, is the original control signal vector belonging to the th terminal after mapping; Structurally adjust the intensity and offset correction of the control signals of each terminal channel after mapping to obtain the corrected control signals ; Combine the adjusted control signals to generate a set of device control signals that meet the control interface specifications of each virtual reality device. Specifically, aggregate the adjustment signals of all terminals: , is the final set of device control signals, is the identification code of the th device, which can be MAC, UUID or device type, is the adjusted control vector of the th device, is a constructed set, not a simple concatenation, but a structural packaging.

[0030] The present invention is further configured such that the formation of the refluxable feedback representation includes: Receiving multi-channel and multi-dimensional feedback data generated by the virtual reality terminal after executing the control instruction; Perform unified coding mapping on the multi-channel and multi-dimensional feedback data to form a high-dimensional feedback tensor. Specifically, let the original tensor of the feedback channel received from the virtual reality terminal be , and obtain the high-dimensional feedback tensor through the unified coding mapping function , is the feedback channel data of the th terminal, is the number of feedback channels, is the feedback dimension per channel, is the multi-source channel fusion encoder, is the feedback coding matrix, is the reconstructed spatial dimension, is the reconstructed channel dimension; Based on the preset multi-modal tensor mapping rules, perform complex reconstruction transformation on the spatial and channel dimensions of the high-dimensional feedback tensor to generate a structured feedback representation. Perform structural reconstruction on the high-dimensional feedback tensor , define the reconstruction function as , , and the specific mapping formula is: , is the structure reconstruction function, which has a non-linear cross-expansion effect is the channel space reconstruction right system, is the scale of the reconstructed structural space, is the scale of the structural channel after reconstruction, The modulo operation ensures the alignment and restoration of the original structure;

[0031] Map the structured feedback representation to the feedback input space preset by S1 to form a feedback closed-loop representation, which is used to realize the closed-loop interaction and dynamic adaptation between the control and perception modules of the virtual reality system.

[0032] The present invention is further configured such that the structured feedback representation is composed of a feedback modality channel set, a semantic mapping index structure, and a timing synchronization identification vector; The feedback modality channel set is generated by encoding and mapping the original feedback data, and maintains the structural consistency with the channel dimension of the multi-modal input tensor. The expression is: where, is the modality structure extraction function, is the number of structured feedback channels; The semantic mapping index structure describes the mapping relationship between each feedback channel and the semantic intention extraction path in the cross-modal perception module in a discrete indexing manner. The expression is where, is the channel index extraction mapping, is the channel semantic index, is the mapping participation channel set; The timing synchronization identification vector is constructed with a fixed-length tensor, and explicitly calibrates the start position, update frequency, and timing accuracy of the feedback signal in the perception cycle according to a unified time scale. The expression is , is the start step identification, is the update frequency, is the timing accuracy factor; The structured feedback representation is overall organized in the closed-loop mapping space, maintains an alignable configuration with the input structure of the perception module, and realizes data backflow at the structural level through the channel reference index mechanism.

[0033] Embodiment 1

[0034] Please refer to Figure 2 , an exemplary VR interactive control and management system includes: Multi-modal perception preprocessing module: Obtain multi-source interaction data of the user in the virtual reality environment, perform intra-modal representation conversion processing on the multi-source interaction data respectively to form a modal vector sequence with a unified time scale and channel structure, and perform fusion processing on it to generate a unified multi-modal input tensor; Cross-modal semantic extraction module: Based on a preset cross-modal perception structure, extract user semantic intention information from the unified multi-modal input tensor to form a semantic intention representation; Semantic regulation instruction generation module: Combine the internal state information of the system at the previous moment with the current semantic intention representation to generate a corresponding control instruction vector; Terminal control signal adaptation module: Perform device channel mapping processing on the control instruction vector to generate a set of device control signals that conform to different VR terminal control interfaces; Multi-dimensional feedback structure reconstruction module: Receive the multi-dimensional feedback information generated by the device after executing the control instruction, and perform structure reconstruction processing on the feedback information to form a feedback representation that can flow back to S1, realizing the closed-loop interaction between control and perception.

[0035] It should be noted that a VR interactive control management system provided by the above embodiment and a VR interactive control management method provided by the above embodiment belong to the same concept. The specific ways in which each module and unit perform operations have been described in detail in the method embodiment, and will not be repeated here. In practical applications, a VR interactive control management system provided by the above embodiment can, according to needs, allocate the above functions to different functional modules, that is, divide the internal structure of the system into different functional modules to complete all or part of the functions described above. This is not limited here either.

[0036] The above embodiments can be implemented in whole or in part by software, hardware, firmware, or any other combination. When implemented using software, the above embodiments can be implemented in whole or in part in the form of a computer program product. The computer program product includes one or more computer instructions or computer programs. When the computer instructions or computer programs are loaded or executed on a computer, the processes or functions described in the embodiments of the present application are generated in whole or in part. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable devices. The computer instructions can be stored in a computer-readable storage medium, or transmitted from one computer-readable storage medium to another computer-readable storage medium. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center by wire (such as infrared, wireless, microwave, etc.). The computer-readable storage medium can be any available medium that the computer can access, or a data storage device such as a server or data center that includes one or more collections of available media. The available media can be magnetic media (such as floppy disks, hard disks, magnetic tapes), optical media (such as DVDs), or semiconductor media. The semiconductor media can be a solid-state drive.

[0037] It should be understood that the term "and / or" in this text is merely a description of the association relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent three situations: A exists alone, A and B exist simultaneously, and B exists alone. Here, A and B can be singular or plural. Additionally, the character " / " in this text generally represents an "or" relationship between the preceding and following associated objects, but it may also represent an "and / or" relationship, and the specific meaning can be understood by referring to the context before and after.

[0038] In this application, "at least one" means one or more, and "a plurality" means two or more. "At least one of the following" or its similar expressions refer to any combination of these items, including any combination of single items or plural items. For example, at least one of a, b, or c can represent: a, b, c, a - b, a - c, b - c, or a - b - c, where a, b, and c can be single or multiple.

[0039] It should be understood that in various embodiments of this application, the magnitudes of the serial numbers of the above processes do not imply the sequence of execution. The execution sequence of each process should be determined by its function and internal logic, and should not impose any limitation on the implementation process of the embodiments of this application.

[0040] Those of ordinary skill in the art can realize that the units and algorithm steps of each example described in combination with the embodiments disclosed herein can be implemented by electronic hardware, or a combination of computer software and electronic hardware. Whether these functions are executed in a hardware or software manner depends on the specific application and design constraints of the technical solution. Professional technicians can use different methods to implement the described functions for each specific application, but such implementation should not be considered to exceed the scope of this application.

[0041] Those skilled in the art can clearly understand that for the convenience and brevity of description, the specific working processes of the systems, devices, and units described above can refer to the corresponding processes in the foregoing method embodiments, and will not be elaborated herein.

[0042] In several embodiments provided in this application, it should be understood that the disclosed system can be implemented in other ways. For example, the device embodiments described above are merely illustrative. For example, the division of the units is only a logical function division, and there can be other division methods in actual implementation. For example, multiple units or components can be combined or integrated into another system, or some features can be ignored or not executed. Another point is that the displayed or discussed coupling or direct coupling or communication connection between each other can be through some interfaces, and the indirect coupling or communication connection of the devices or units can be in an electrical, mechanical, or other form.

[0043] The unit described as a separation component may or may not be physically separated. The component displayed as a unit may or may not be a physical unit, that is, it may be located in one place or distributed to multiple network units. Some or all of the units can be selected according to actual needs to achieve the purpose of the solution of this embodiment.

[0044] In addition, each functional unit in various embodiments of the present application can be integrated in a processing unit, or each unit can exist physically alone, or two or more units can be integrated in one unit.

[0045] If the above function is implemented in the form of a software functional unit and sold or used as an independent product, it can be stored in a computer-readable storage medium. Based on such an understanding, the technical solution of the present application, in essence, or the part that contributes to the prior art, or a part of the technical solution, can be embodied in the form of a software product. The computer software product is stored in a storage medium and includes several instructions for causing a computer device (which can be a personal computer, a server, or a network device, etc.) to execute all or part of the steps of the methods described in various embodiments of the present application. The foregoing storage medium includes: various media such as USB flash drives, mobile hard disks, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical discs that can store program codes.

[0046] The above is only the specific implementation manner of the present application, but the protection scope of the present application is not limited thereto. Any person skilled in the art can easily think of changes or substitutions within the technical scope disclosed by the present application, and all of them should be covered by the protection scope of the present application. Therefore, the protection scope of the present application should be subject to the protection scope of the claims.

Claims

1. A VR interactive control management method, characterized in that, Including: S1: Obtain the multi-source interaction data of the user in the virtual reality environment, respectively perform intra-modal representation conversion processing on the multi-source interaction data to form a sequence of modal vectors with a unified time scale and channel structure, and perform fusion processing on them to generate a unified multi-modal input tensor; S2: Based on a preset cross-modal perception structure, extract the user's semantic intention information from the unified multi-modal input tensor to form a semantic intention representation; S3: Combine the internal state information of the system at the previous moment with the current semantic intention representation to generate a corresponding control instruction vector; S4: Perform device channel mapping processing on the control instruction vector to generate a set of device control signals that conform to different VR terminal control interfaces; S5: Receive the multi-dimensional feedback information generated by the device after executing the control instruction, and perform structure reconstruction processing on the feedback information to form a feedback representation that can flow back to S1, realizing a closed-loop interaction between control and perception.

2. The VR interactive control management method according to claim 1, wherein, The multi-source interaction data includes three-dimensional gesture trajectory data, voice signal feature data, eye movement behavior data, and physiological response data; Respectively construct independent modal representation spaces for the multi-source interaction data, and align each modal data to a unified time scale based on a structure-preserving resampling mechanism; Construct a sequence of modal vectors with a unified channel structure through a cross-modal channel normalization mapping strategy; Generate a unified multi-modal input tensor through modal splicing and channel-level fusion.

3. The VR interactive control management method according to claim 2, wherein, The generation of the semantic intention representation table includes: Slice the multi-modal input tensor according to the modal dimension to form a set of modal sub-tensors; Perform bilinear interaction mapping on the channel-level coupling relationship between each modal sub-tensor to form a time-distributed modal cross-tensor structure; Compress the structure of the modal cross-tensor and construct a unified semantic vector representation.

4. A VR interactive control and management method according to claim 3, characterized in that, The unified semantic vector representation is generated by structure compression of the modal cross-tensor structure constructed between multiple modal sub-tensors; The modal cross-tensor structure is constructed by performing channel-level bilinear interaction between each modal sub-tensor obtained by slicing the unified multi-modal input tensor according to the modal dimension; Structure compression includes performing multi-dimensional compression mapping on the modal cross-tensor based on a channel coupling mapping kernel at a fixed time scale, and fusing each compressed representation to form a unified semantic vector representation.

5. The VR interactive control management method according to claim 4, wherein, Combining the internal state information of the system at the previous moment with the current semantic intention representation to generate a corresponding control instruction vector includes: Align the internal state information of the system at the previous moment with the control history trajectory to construct a fused state representation; Perform channel matching cross-embedding on the fused state representation and the current semantic intention representation to construct a semantic regulation representation tensor; Obtain the control instruction vector by performing multi-order transformation mapping on the semantic regulation representation tensor, and perform feedback correction and scale normalization processing on the instruction vector.

6. A VR interactive control and management method according to claim 5, characterized in that, The control instruction vector is a multi-dimensional vector with a hierarchical channel configuration structure, including a state fusion representation domain, a semantic response drive domain, and a channel regulation weight domain, and each domain is arranged in an orderly manner according to a preset semantic embedding structure; The state fusion representation domain represents the structural alignment relationship between the system's historical state and the semantic input; The semantic response drive domain encodes the structurally active regions in the current semantic vector through a channel index mapping method; The channel regulation weight domain records the cross-channel dynamic response amplitude and the structure allocation priority in a non-overlapping indexing manner. The sub-domains maintain structural independence through cross-domain mapping rules and share a unified channel index table.

7. A VR interactive control and management method according to claim 5, characterized in that Generating a set of device control signals that conform to different VR terminal control interfaces includes: Performing multi-channel response structure expansion on the input control instruction vector to form a high-dimensional control response tensor; Based on the preset device channel mapping rules, mapping the control response tensor to the functional channels of multiple virtual reality terminals; Performing structural strength adjustment and offset correction on the control signals of each terminal channel after mapping; Combining the adjusted control signals to generate a set of device control signals that meet the control interface specifications of each virtual reality device.

8. A VR interaction control management method according to claim 7, characterized in that Forming a feedback representation that can be refluxed includes: Receiving multi-channel and multi-dimensional feedback data generated by the virtual reality terminal after executing the control instruction; Performing unified coding mapping on the multi-channel and multi-dimensional feedback data to form a high-dimensional feedback tensor; Based on the preset multi-modal tensor mapping rules, performing complex reconstruction transformation on the high-dimensional feedback tensor in the spatial and channel dimensions to generate a structured feedback representation; Mapping the structured feedback representation to the preset feedback input space of S1 to form a feedback closed-loop representation for realizing closed-loop interaction and dynamic adaptation between the control and perception modules of the virtual reality system.

9. The VR interaction control and management method according to claim 8, characterized in that, The structured feedback representation consists of a feedback modality channel set, a semantic mapping index structure, and a timing synchronization identification vector; The feedback modality channel set is generated by encoding and mapping the original feedback data, maintaining structural consistency with the channel dimensions of the multi-modal input tensor; The semantic mapping index structure describes the mapping relationship between each feedback channel and the semantic intention extraction path in the cross-modal perception module in a discrete indexing manner; The timing synchronization identification vector is constructed with a fixed-length tensor, explicitly calibrating the start position, update frequency, and timing accuracy of the feedback signal in the perception cycle according to a unified time scale; The structured feedback representation is overall organized in a closed-loop mapping space, maintaining an alignable configuration with the input structure of the perception module, and realizing data reflux at the structural level through the channel reference index mechanism.

10. A VR interactive control management system for implementing the VR interactive control management method according to any one of claims 1-9, characterized in that, Including: Multi-modal perception preprocessing module: Obtaining multi-source interaction data of the user in the virtual reality environment, respectively performing intra-modal representation conversion processing on the multi-source interaction data to form a sequence of modal vectors with a unified time scale and channel structure, and fusing them to generate a unified multi-modal input tensor; Cross-modal semantic extraction module: Based on the preset cross-modal perception structure, extracting user semantic intention information from the unified multi-modal input tensor to form a semantic intention representation; Semantic regulation instruction generation module: Combining the system internal state information of the previous moment with the current semantic intention representation to generate a corresponding control instruction vector; Terminal control signal adaptation module: Performing device channel mapping processing on the control instruction vector to generate a set of device control signals that conform to different VR terminal control interfaces; Multi-dimensional feedback structure reconstruction module: Receiving the multi-dimensional feedback information generated by the device after executing the control instruction, and performing structure reconstruction processing on the feedback information to form a feedback representation that can be refluxed to S1, realizing closed-loop interaction between control and perception.

Citation Information

Patent Citations

  • Multi-modal information processing and interaction system

    CN112613534A

  • Human-computer interaction method and device and storage medium

    CN119292459A

  • Multi-mode AI glasses vision-electroencephalogram cooperative control method, device and equipment

    CN120085760A

  • Transmodal input fusion for a wearable system

    US20190362557A1

  • Device control method, conflict processing method, corresponding apparatus and electronic device

    US20220317641A1

Cited By

  • AR interaction method and system for health and epidemic prevention science popularization

    CN120767006A

  • An ar interaction method and system for health and epidemic prevention popularization

    CN120767006B

  • Large language model driven VR intelligent role emotion feedback method and device

    CN121277361A