A multimodal perception immersive digital interaction system and method
By combining a multimodal perception fusion module and a contextual cognitive reasoning module with a holographic biosensing matrix and a dynamic contextual framework, the problem of insufficient multimodal perception in existing digital interaction systems is solved, enabling accurate identification and real-time response of user intentions, and improving the naturalness and immersion of the interaction.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Applications(China)
- Current Assignee / Owner
- CHINA JILIANG UNIV
- Filing Date
- 2026-01-07
- Publication Date
- 2026-06-02
Smart Images

Figure CN122133049A_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of digital interaction technology, and in particular to a multimodal perception immersive digital interaction system and method. Background Technology
[0002] With the continuous development of society, human-computer interaction technology has increasingly become one of the core driving forces for technological progress and industrial transformation. Driven by the wave of digitalization and intelligence, users have higher and higher requirements for interactive experience. Traditional single-modal interaction methods can no longer meet people's needs for naturalness, immersion and real-time. The rise of multimodal perception technology has made it possible to integrate multi-channel information such as vision, hearing, touch and even physiological signals, which has promoted the evolution of human-computer interaction towards a more intelligent direction.
[0003] Currently, digital interactive systems generally suffer from insufficient multimodal information fusion capabilities. Most systems still rely on single or limited perception channels and lack a comprehensive perception and dynamic response mechanism for user behavior, emotional state, and environmental context. This results in a mechanical interaction process, delayed feedback, and difficulty in achieving an immersive experience. This fragmented perception and interaction mode limits the system's accurate understanding of user intentions and adaptive response, seriously affecting the naturalness and fluency of the interaction. Summary of the Invention
[0004] In view of the problems existing in the existing multimodal perception immersive digital interaction systems and methods, this invention is proposed.
[0005] Therefore, the problem to be solved by the present invention is to solve the technical problems of mechanical interaction, slow response, weak immersion and inaccurate intent understanding caused by insufficient multimodal perception fusion in existing digital interaction systems.
[0006] To solve the above-mentioned technical problems, the present invention provides the following technical solution: In a first aspect, embodiments of the present invention provide a multimodal perception immersive digital interaction system, comprising: a multimodal perception fusion module, used to collect multidimensional human condition data of users through a holographic biosensor matrix and generate user state representation; a contextual cognition reasoning module, used to construct a dynamic contextual framework based on user state representation and combined with environmental contextual information; and an immersive interaction generation module, used to generate immersive interactive content according to a multi-channel feedback mechanism driving digital interaction based on the dynamic contextual framework, and to optimize the interaction strategy through the multi-channel feedback mechanism.
[0007] As a preferred embodiment of the multimodal perception immersive digital interaction system of the present invention, the multimodal perception fusion module includes a heterogeneous sensing collaboration submodule, a cross-domain feature extraction submodule, and a semantic fusion representation submodule. The heterogeneous sensing collaboration submodule is used to schedule and calibrate sensor units of different modes in the holographic biosensing matrix; The cross-domain feature extraction submodule is used to perform time-frequency-space multidimensional analysis on multimodal signals and extract semantic features of multimodal signals; The semantic fusion representation submodule is used to construct a cross-modal attention fusion network, which maps heterogeneous features to the semantic space to generate user state representations.
[0008] As a preferred embodiment of the multimodal perception immersive digital interaction system of the present invention, the contextual cognition reasoning module includes a contextual perception integration submodule, an intent graph construction submodule, and a contextual evolution deduction submodule. The context-aware integration submodule is used to dynamically access and parse context information from digital scenes to establish a quantifiable environment state descriptor. The intent graph construction submodule is used to integrate user state representation and environmental state descriptor to drive a knowledge graph-based semantic association engine and generate an intent topology structure that includes the user's potential goals and behavioral paths. The context evolution deduction submodule is used to introduce a temporal memory network to dynamically track and probabilistically infer the intention topology, and output a contextual cognitive framework that reflects the current digital interaction.
[0009] As a preferred embodiment of the multimodal perception immersive digital interaction system of the present invention, the immersive interaction generation module includes a multi-channel collaborative orchestration submodule, a real-time rendering driving submodule, and a closed-loop strategy evolution submodule. The multi-channel collaborative orchestration submodule is used to parse semantic instructions in the dynamic context framework and generate a multimodal interaction instruction set; The real-time rendering driver submodule is used to synchronously generate an immersive sensory output stream based on the multimodal interaction instruction set of the scheduled holographic biosensing matrix. The closed-loop strategy evolution submodule is used to collect user behavioral responses to immersive sensory output streams, evaluate the effectiveness of behavioral responses using reinforcement learning algorithms, dynamically adjust the multi-channel feedback mechanism, and complete the multi-channel feedback mechanism to optimize the interaction strategy.
[0010] As a preferred embodiment of the multimodal sensing immersive digital interaction system of the present invention, the heterogeneous sensing collaborative submodule includes a sensing timing alignment unit and a modal adaptive calibration unit. The sensing timing alignment unit is used to synchronize the original data streams of each modal sensor in the holographic biosensing matrix with nanosecond-level timestamps. The modal adaptive calibration unit is used to dynamically adjust the raw data stream of each modal sensor and calibrate the sensor units of different modalities in the holographic biosensing matrix. The cross-domain feature extraction submodule includes a spatiotemporal semantic coding unit and a cross-modal decoupling representation unit; The spatiotemporal semantic coding unit is used to jointly embed multimodal signals in a unified spatiotemporal coordinate system to generate a multi-channel feature tensor with position awareness capability. The cross-modal decoupling characterization unit is used to output semantic features for denoising multimodal signals by combating shared multi-channel feature tensor noise in decoupled multimodal signals. The semantic fusion representation submodule includes an attention-guided fusion unit and a dynamic representation compression unit; The attention-guided fusion unit is used to construct a hierarchical cross-modal attention mechanism, which dynamically weights the semantic contribution of different modal features based on the current interaction task. The dynamic representation compression unit is used to map high-dimensional fused features to a low-dimensional manifold space to generate a user state representation vector.
[0011] As a preferred embodiment of the multimodal perception immersive digital interaction system of the present invention, the context perception integration submodule includes a scene semantic parsing unit and an environment state quantization unit. The scene semantic parsing unit is used to perform ontology-driven semantic annotation on contextual information from digital scenes and extract spatial and event nodes of environmental states. The environmental state quantization unit is used to map the space of environmental state and event nodes to the context parameter space to generate an environmental state descriptor. The intent graph construction submodule includes a user environment semantic alignment unit and a dynamic intent topology generation unit; The user environment semantic alignment unit is used to align the user state representation vector and the environment state descriptor in the unified knowledge embedding space and identify their paths in the semantic association engine. The dynamic intent topology generation unit is used to activate concept nodes in the knowledge graph based on the recognized semantic association engine path to complete the generation of intent topology structure; The scenario evolution deduction submodule includes a temporal intent tracking unit and a probabilistic scenario deduction unit; The temporal intent tracking unit is used to update the weights of each node in the intent topology using a memory-enhanced recursive network; The probabilistic scenario inference unit is used to perform multi-step predictions based on the weights of each node in the updated intent topology, and output a contextual cognition framework.
[0012] As a preferred embodiment of the multimodal perception immersive digital interaction system of the present invention, the multi-channel collaborative orchestration submodule includes a semantic action mapping unit and a cross-modal timing scheduling unit. The semantic action mapping unit is used to map the dynamic context framework to semantic actions. The cross-modal timing scheduling unit is used to sort multimodal interaction instructions based on semantic action mapping and generate a multimodal interaction instruction set; The real-time rendering driver submodule includes a heterogeneous output engine calling unit and a sensory stream synchronous synthesis unit; The heterogeneous output engine calling unit is used to process multimodal interaction instruction sets; The sensory stream synchronization synthesis unit is used to fuse the output signals of each modality under the processing of multimodal interaction instruction set to generate an immersive sensory output stream; The closed-loop policy evolution submodule includes a response feature capture unit and a policy gradient optimization unit; The response feature capture unit is used to extract the user's sensory signals to the current immersive sensory output stream from the holographic biosensing matrix and construct a multidimensional response feature vector. The policy gradient optimization unit is used to input the multi-dimensional response feature vector as a reward signal into the reinforcement learning algorithm, iteratively update the parameter configuration of the multi-channel feedback mechanism, and complete the optimization interaction strategy of the multi-channel feedback mechanism.
[0013] Secondly, embodiments of the present invention provide a multimodal perception-based immersive digital interaction method, which includes: collecting multidimensional human condition data of users through a holographic biosensor matrix to generate user state representation; constructing a dynamic contextual framework based on user state representation and combined with environmental contextual information; generating immersive interactive content according to a multi-channel feedback mechanism that drives digital interaction based on the dynamic contextual framework, and optimizing the interaction strategy through the multi-channel feedback mechanism.
[0014] Thirdly, embodiments of the present invention provide a computer device, including a memory and a processor, wherein the memory stores a computer program, and the processor executes the computer program to implement any step of the above-described multimodal perception immersive digital interaction system.
[0015] Fourthly, embodiments of the present invention provide a computer-readable storage medium having a computer program stored thereon, wherein: when the computer program is executed by a processor, it implements any step of the above-described multimodal perception immersive digital interaction system.
[0016] The beneficial effects of this invention are as follows: By constructing a closed-loop architecture integrating perception, cognition, and interaction, this invention achieves deep fusion and dynamic understanding of the user's multidimensional human state and environmental context, significantly improving the system's accuracy in recognizing user intentions and its real-time response. The holographic biosensing matrix and multimodal signal collaborative processing mechanism effectively overcome the shortcomings of traditional systems, such as single perception channels and fragmented information, enhancing the naturalness and immersion of the interaction. Based on a dynamic contextual framework-driven multi-channel collaborative feedback and closed-loop strategy evolution mechanism, the system possesses continuous adaptive optimization capabilities, enabling it to generate highly personalized, smooth, and context-consistent immersive interactive experiences, thereby comprehensively solving key problems in existing technologies such as mechanical interaction, delayed feedback, and insufficient immersion. Attached Figure Description
[0017] To more clearly illustrate the technical solutions of the embodiments of the present invention, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort. Wherein: Figure 1 This is a schematic diagram of a multimodal perception immersive digital interaction system and method provided in an embodiment of the present invention.
[0018] Figure 2 This is a flowchart illustrating a multimodal perception immersive digital interaction system and method provided in an embodiment of the present invention.
[0019] Figure 3 This is a schematic diagram of the structure of a medium for a multimodal sensing immersive digital interaction system and method provided in an embodiment of the present invention.
[0020] Figure 4 This is a schematic diagram of the structure of a computing device for a multimodal sensing immersive digital interaction system and method provided in an embodiment of the present invention. Detailed Implementation
[0021] To make the above-mentioned objects, features, and advantages of the present invention more apparent and understandable, specific embodiments of the present invention will be described in detail below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of the present invention, and not all of them. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort should fall within the protection scope of the present invention.
[0022] Many specific details are set forth in the following description in order to provide a full understanding of the invention. However, the invention may also be practiced in other ways different from those described herein, and those skilled in the art can make similar extensions without departing from the spirit of the invention. Therefore, the invention is not limited to the specific embodiments disclosed below.
[0023] Secondly, the term "one embodiment" or "embodiment" as used herein refers to a specific feature, structure, or characteristic that may be included in at least one implementation of the present invention. The phrase "in one embodiment" appearing in different places in this specification does not necessarily refer to the same embodiment, nor is it a single or selective embodiment that is mutually exclusive with other embodiments.
[0024] This invention is described in detail with reference to the schematic diagrams. When detailing the embodiments of this invention, for ease of explanation, the cross-sectional views illustrating the device structure may be partially enlarged, not adhering to the usual scale. Furthermore, the schematic diagrams are merely examples and should not be construed as limiting the scope of protection of this invention. In actual fabrication, the three-dimensional spatial dimensions of length, width, and depth should be included.
[0025] Furthermore, in the description of this invention, it should be noted that the terms "upper," "lower," "inner," and "outer," etc., indicate the orientation or positional relationship based on the orientation or positional relationship shown in the accompanying drawings. These terms are used solely for the convenience of describing the invention and for simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation. Therefore, they should not be construed as limitations on the invention. In addition, the terms "first," "second," or "third" are used for descriptive purposes only and should not be construed as indicating or implying relative importance.
[0026] Unless otherwise explicitly specified and limited, the terms "installation," "connection," and "joining" in this invention should be interpreted broadly. For example, they can refer to fixed connections, detachable connections, or integral connections; similarly, they can refer to mechanical connections, electrical connections, or direct connections, or indirect connections through an intermediate medium, or internal connections between two components. Those skilled in the art can understand the specific meaning of the above terms in this invention based on the specific circumstances.
[0027] Example Reference Figures 1-4 This is the first embodiment of the present invention, which provides a multimodal sensing immersive digital interaction system, including: S1: Multimodal perception fusion module, used to collect multidimensional human condition data of users through holographic biosensor matrix and generate user state representation.
[0028] The multimodal perception fusion module includes a heterogeneous sensing collaboration submodule, a cross-domain feature extraction submodule, and a semantic fusion representation submodule. The heterogeneous sensing collaboration submodule is used to schedule and calibrate sensor units of different modes in the holographic biosensing matrix; The cross-domain feature extraction submodule is used to perform time-frequency-space multidimensional analysis on multimodal signals and extract semantic features of multimodal signals; The semantic fusion representation submodule is used to construct a cross-modal attention fusion network, which maps heterogeneous features to the semantic space to generate user state representations.
[0029] S2: Contextual cognitive reasoning module, used to construct a dynamic contextual framework based on user state representation and combined with environmental context information.
[0030] The contextual cognition reasoning module includes a context perception integration submodule, an intent graph construction submodule, and a contextual evolution deduction submodule. The context-aware integration submodule is used to dynamically access and parse context information from digital scenes to establish quantifiable environmental state descriptors. The intent graph construction submodule is used to integrate user state representation and environment state descriptor to drive a knowledge graph-based semantic association engine and generate an intent topology structure that includes the user's potential goals and behavioral paths. The context evolution inference submodule is used to introduce a temporal memory network to dynamically track and infer the probabilistic structure of intent topology, and output a contextual cognitive framework that reflects the current digital interaction.
[0031] S3: Immersive Interaction Generation Module, used to generate immersive interactive content based on a multi-channel feedback mechanism that drives digital interaction according to a dynamic context framework, and to optimize interaction strategies through the multi-channel feedback mechanism.
[0032] The immersive interactive generation module includes a multi-channel collaborative orchestration submodule, a real-time rendering driving submodule, and a closed-loop strategy evolution submodule. The multi-channel collaborative orchestration submodule is used to parse semantic instructions in a dynamic context framework and generate a multimodal interaction instruction set; The real-time rendering driver submodule is used to synchronously generate an immersive sensory output stream based on the multimodal interaction instruction set of the scheduled holographic biosensing matrix. The closed-loop strategy evolution submodule is used to collect user behavioral responses to immersive sensory output streams, combine reinforcement learning algorithms to evaluate the effectiveness of behavioral responses, dynamically adjust the multi-channel feedback mechanism, and complete the optimization of the interaction strategy through the multi-channel feedback mechanism.
[0033] Further refinement of S1: The heterogeneous sensing collaboration submodule includes a sensing timing alignment unit and a modal adaptive calibration unit; The sensing timing alignment unit is used to synchronize the raw data streams of each modal sensor in the holographic biosensing matrix with nanosecond-level timestamps; The modal adaptive calibration unit is used to dynamically adjust the raw data stream of each modal sensor and calibrate the sensor units of different modalities in the holographic biosensing matrix. The cross-domain feature extraction submodule includes a spatiotemporal semantic coding unit and a cross-modal decoupling representation unit; The spatiotemporal semantic coding unit is used to jointly embed multimodal signals in a unified spatiotemporal coordinate system to generate a multi-channel feature tensor with position-aware capability; The cross-modal decoupling representation unit is used to output semantic features for denoising multimodal signals by combating shared multichannel feature tensor noise in decoupled multimodal signals; The semantic fusion representation submodule includes an attention-guided fusion unit and a dynamic representation compression unit; Attention-guided fusion units are used to construct hierarchical cross-modal attention mechanisms, dynamically weighting the semantic contributions of different modal features based on the current interaction task; The dynamic representation compression unit is used to map high-dimensional fused features to a low-dimensional manifold space to generate user state representation vectors.
[0034] Further refinement for S2: The context-aware integration submodule includes a scene semantic parsing unit and an environment state quantization unit; The scene semantic parsing unit is used to perform ontology-driven semantic annotation on contextual information from digital scenes and extract spatial and event nodes of environmental states. The environmental state quantization unit is used to map the space of environmental states and event nodes to the context parameter space to generate an environmental state descriptor. The intent graph construction submodule includes a user environment semantic alignment unit and a dynamic intent topology generation unit; The user environment semantic alignment unit is used to align the user state representation vector and the environment state descriptor in the unified knowledge embedding space and identify their paths in the semantic association engine. The dynamic intent topology generation unit is used to activate concept nodes in the knowledge graph based on the recognized semantic association engine path, and complete the generation of intent topology structure; The context evolution inference submodule includes a temporal intent tracking unit and a probabilistic context inference unit; The temporal intent tracking unit is used to update the weights of each node in the intent topology using a memory-enhanced recursive network; The probabilistic contextual inference unit is used to perform multi-step predictions based on the weights of each node in the updated intent topology, and outputs a contextual cognition framework.
[0035] Further refinement for S3: The multi-channel collaborative orchestration submodule includes a semantic action mapping unit and a cross-modal timing scheduling unit; The semantic action mapping unit is used to map semantic actions to dynamic context frames; The cross-modal timing scheduling unit is used to sort multimodal interaction instructions based on semantic action mapping and generate a multimodal interaction instruction set; The real-time rendering driver submodule includes a heterogeneous output engine calling unit and a sensory stream synchronous compositing unit; The heterogeneous output engine calling unit is used to process multimodal interactive instruction sets; The sensory stream synchronous synthesis unit is used to fuse the output signals of each modality under the processing of multimodal interaction instruction set to generate an immersive sensory output stream; The closed-loop policy evolution submodule includes a response feature capture unit and a policy gradient optimization unit; The response feature capture unit is used to extract the user's sensory signals to the current immersive sensory output stream from the holographic biosensing matrix and construct a multidimensional response feature vector; The policy gradient optimization unit is used to input the multi-dimensional response feature vector as a reward signal into the reinforcement learning algorithm, iteratively update the parameter configuration of the multi-channel feedback mechanism, and complete the optimization of the interaction policy of the multi-channel feedback mechanism.
[0036] The reinforcement learning algorithm used by the policy gradient optimization unit can be described by the following policy gradient formula: in, Indicates the first The strategy network parameters of the multi-channel feedback mechanism in round iteration. Indicates the learning rate. Indicates the state Select action The strategy probability, This represents the cumulative reward signal calculated based on the multidimensional response feature vector. Indicates the first The strategy network parameters of the multi-channel feedback mechanism in round iteration. This represents a single interaction trajectory.
[0037] Furthermore, this invention uses a multimodal perception fusion module to collect and fuse multidimensional human state data of users in real time, generating a unified user state representation. This representation is input into a contextual cognition reasoning module, which combines environmental context information to construct a dynamic contextual framework, achieving a deep understanding of user intentions and interaction behavior. Subsequently, an immersive interaction generation module coordinates multi-channel feedback resources based on this contextual framework to generate immersive sensory output, and continuously collects user responses and optimizes interaction strategies through a closed-loop strategy evolution mechanism. The three core modules form a complete closed loop of perception, cognition, interaction, feedback, and optimization. Each sub-module and unit is closely linked in terms of temporal synchronization, semantic alignment, and strategy iteration, thereby systematically solving the problems of insufficient multimodal fusion, delayed response, and lack of immersion in existing technologies.
[0038] In a preferred embodiment, a multimodal perception-based immersive digital interaction method includes: collecting multidimensional human condition data of users through a holographic biosensor matrix to generate user state representation; constructing a dynamic contextual framework based on the user state representation and combined with environmental contextual information; generating immersive interactive content according to a multi-channel feedback mechanism that drives digital interaction based on the dynamic contextual framework; and optimizing the interaction strategy through the multi-channel feedback mechanism.
[0039] The above-mentioned unit modules can be embedded in the processor of the computer device in hardware form or independent of it, or they can be stored in the memory of the computer device in software form, so that the processor can call and execute the corresponding operations of the above modules.
[0040] In one embodiment, a computer device is provided, which may be a terminal. The computer device includes a processor, memory, a communication interface, a display screen, and an input device connected via a system bus. The processor of the computer device provides computing and control capabilities. The memory of the computer device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores an operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs stored in the non-volatile storage medium. The communication interface of the computer device is used for wired or wireless communication with external terminals. Wireless communication can be achieved through Wi-Fi, carrier networks, NFC (Near Field Communication), or other technologies. The display screen of the computer device may be an LCD screen or an e-ink display screen. The input device of the computer device may be a touch layer covering the display screen, or buttons, a trackball, or a touchpad located on the casing of the computer device, or an external keyboard, touchpad, or mouse, etc.
[0041] In summary, this invention, by constructing a closed-loop architecture integrating perception, cognition, and interaction, achieves deep fusion and dynamic understanding of the user's multidimensional human state and environmental context, significantly improving the system's accuracy in recognizing user intentions and its real-time response. The holographic biosensing matrix and multimodal signal collaborative processing mechanism effectively overcome the shortcomings of traditional systems, such as single perception channels and fragmented information, enhancing the naturalness and immersion of the interaction. Based on a dynamic contextual framework-driven multi-channel collaborative feedback and closed-loop strategy evolution mechanism, the system possesses continuous adaptive optimization capabilities, enabling the generation of highly personalized, smooth, and context-consistent immersive interactive experiences, thus comprehensively solving key problems in existing technologies such as mechanical interaction, delayed feedback, and insufficient immersion.
[0042] After introducing the method and system of exemplary embodiments of the present invention, the following references are made. Figure 3A computer-readable storage medium according to exemplary embodiments of the present invention will be described, please refer to... Figure 3 The computer-readable storage medium shown is an optical disc 30, on which a computer program (i.e., a program product) is stored. When the computer program is run by a processor, it implements the steps described in the above method implementation. For example, the multimodal perception fusion module is used to collect multidimensional human condition data of the user through a holographic biosensor matrix and generate a user state representation; the contextual cognitive reasoning module is used to construct a dynamic contextual framework based on the user state representation and combined with environmental context information; and the immersive interaction generation module is used to generate immersive interactive content according to the multi-channel feedback mechanism of digital interaction driven by the dynamic contextual framework, and to optimize the interaction strategy through the multi-channel feedback mechanism. The specific implementation methods of each step will not be repeated here.
[0043] It should be noted that examples of computer-readable storage media may also include, but are not limited to, phase-change memory (PRAM), static random access memory (SRAM), dynamic random access memory (DRAM), other types of random access memory (RAM), read-only memory (ROM), electrically erasable programmable read-only memory (EEPROM), flash memory or other optical and magnetic storage media, which will not be elaborated here.
[0044] After introducing the methods and media of exemplary embodiments of the present invention, the following references are made. Figure 4 A computational device for adaptive recovery of low-voltage power grid self-healing control according to an exemplary embodiment of the present invention.
[0045] Figure 4 A block diagram is shown of an exemplary computing device 40 suitable for implementing embodiments of the present invention. The computing device 40 may be a computer system or a server. Figure 4 The computing device 40 shown is merely an example and should not impose any limitation on the functionality and scope of use of the embodiments of the present invention.
[0046] like Figure 4 As shown, the components of computing device 40 may include, but are not limited to: one or more processors or processing units 401, system memory 402, and bus 403 connecting different system components (including system memory 402 and processing unit 401).
[0047] The computing device 40 typically includes a variety of computer system readable media. These media can be any available media that can be accessed by the computing device 40, including volatile and non-volatile media, and removable and non-removable media.
[0048] System memory 402 may include computer system readable media in the form of volatile memory, such as random access memory (RAM) 4021 and / or cache memory 4022. Computing device 40 may further include other removable / non-removable, volatile / non-volatile computer system storage media. By way of example only, ROM 4023 may be used to read and write non-removable, non-volatile magnetic media (…). Figure 4 (Not shown in the image, usually referred to as "hard drive"). Although not shown in... Figure 4 The diagram illustrates that disk drives for reading and writing to removable non-volatile disks (e.g., "floppy disks") and optical disc drives for reading and writing to removable non-volatile optical discs (e.g., CD-ROMs, DVD-ROMs, or other optical media) can be provided. In these cases, each drive can be connected to bus 403 via one or more data media interfaces. System memory 402 may include at least one program product having a set (e.g., at least one) of program modules configured to perform the functions of the embodiments of the present invention.
[0049] A program / utility 4025 having a set (at least one) of program modules 4024 may be stored, for example, in system memory 402, and such program modules 4024 include, but are not limited to, an operating system, one or more application programs, other program modules, and program data, each or some combination of these examples may include an implementation of a network environment. Program modules 4024 typically perform the functions and / or methods described in the embodiments of the present invention.
[0050] The computing device 40 can also communicate with one or more external devices 404 (such as a keyboard, pointing device, display, etc.). This communication can be performed via the input / output (I / O) interface 405. Furthermore, the computing device 40 can also communicate with one or more networks (e.g., local area networks (LANs), wide area networks (WANs), and / or public networks, such as the Internet) via a network adapter 406. Figure 4 As shown, network adapter 406 communicates with other modules of computing device 40 (such as processing unit 401) via bus 403. It should be understood that, although... Figure 4 As not shown, it can be used in conjunction with computing device 40 with other hardware and / or software modules.
[0051] The processing unit 401 executes various functional applications and data processing by running programs stored in the system memory 402. For example, the multimodal perception fusion module is used to collect multidimensional human condition data of users through a holographic biosensor matrix and generate user state representations; the contextual cognition reasoning module is used to construct a dynamic contextual framework based on user state representations and combined with environmental context information; and the immersive interaction generation module is used to generate immersive interactive content based on a multi-channel feedback mechanism that drives digital interaction according to the dynamic contextual framework, and optimize the interaction strategy through the multi-channel feedback mechanism.
[0052] Those skilled in the art will understand that, for the sake of convenience and brevity, the specific working processes of the systems, devices, and units described above can be referred to the corresponding processes in the foregoing method embodiments, and will not be repeated here.
[0053] In the several embodiments provided by this invention, it should be understood that the disclosed systems, apparatuses, and methods can be implemented in other ways. The apparatus embodiments described above are merely illustrative. For example, the division of units is only a logical functional division, and in actual implementation, there may be other division methods. Furthermore, multiple units or components may be combined or integrated into another system, or some features may be ignored or not executed. Additionally, the coupling or direct coupling or communication connection shown or discussed may be through some communication interface; the indirect coupling or communication connection between apparatuses or units may be electrical, mechanical, or other forms.
[0054] The units described as separate components may or may not be physically separate. The components shown as units may or may not be physical units; that is, they may be located in one place or distributed across multiple network units. Some or all of the units can be selected to achieve the purpose of this embodiment according to actual needs.
[0055] In addition, the functional units in the various embodiments of the present invention can be integrated into one processing unit, or each unit can exist physically separately, or two or more units can be integrated into one unit.
[0056] If the functionality is implemented as a software functional unit and sold or used as an independent product, it can be stored in a processor-executable, non-volatile, computer-readable storage medium. Based on this understanding, the technical solution of this invention, or the part that contributes to the prior art, or a portion of the technical solution, can be embodied in the form of a software product. This computer software product is stored in a storage medium and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute all or part of the steps of the methods of the various embodiments of this invention. The aforementioned storage medium includes various media capable of storing program code, such as USB flash drives, portable hard drives, read-only memory (ROM), random access memory (RAM), magnetic disks, or optical disks.
[0057] Finally, it should be noted that the above embodiments are merely specific implementations of the present invention, used to illustrate the technical solutions of the present invention, and not to limit it. The scope of protection of the present invention is not limited thereto. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that any person skilled in the art can still modify or easily conceive of changes to the technical solutions described in the foregoing embodiments within the technical scope disclosed in the present invention, or make equivalent substitutions for some of the technical features; and these modifications, changes, or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be covered within the scope of protection of the present invention. Therefore, the scope of protection of the present invention should be determined by the scope of the claims.
[0058] Furthermore, although the operations of the method of the present invention are described in a specific order in the accompanying drawings, this does not require or imply that these operations must be performed in that specific order, or that all the operations shown must be performed to achieve the desired result. Additionally or alternatively, certain steps may be omitted, multiple steps may be combined into one step, and / or one step may be broken down into multiple steps.
[0059] It should be noted that the above embodiments are only used to illustrate the technical solutions of the present invention and are not intended to limit it. Although the present invention has been described in detail with reference to preferred embodiments, those skilled in the art should understand that modifications or equivalent substitutions can be made to the technical solutions of the present invention without departing from the spirit and scope of the technical solutions of the present invention, and all such modifications or substitutions should be covered within the scope of the claims of the present invention.
Claims
1. A multimodal sensing immersive digital interaction system, characterized in that: include, The multimodal perception fusion module is used to collect multidimensional human condition data of users through a holographic biosensing matrix and generate user state representations. The contextual cognition reasoning module is used to construct a dynamic contextual framework based on user state representation and environmental context information. The immersive interaction generation module is used to generate immersive interactive content based on a multi-channel feedback mechanism that drives digital interaction according to a dynamic context framework, and to optimize the interaction strategy through the multi-channel feedback mechanism.
2. The multimodal sensing immersive digital interaction system as described in claim 1, characterized in that: The multimodal perception fusion module includes a heterogeneous sensing collaboration submodule, a cross-domain feature extraction submodule, and a semantic fusion representation submodule; The heterogeneous sensing collaboration submodule is used to schedule and calibrate sensor units of different modes in the holographic biosensing matrix; The cross-domain feature extraction submodule is used to perform time-frequency-space multidimensional analysis on multimodal signals and extract semantic features of multimodal signals; The semantic fusion representation submodule is used to construct a cross-modal attention fusion network, which maps heterogeneous features to the semantic space to generate user state representations.
3. The multimodal sensing immersive digital interaction system as described in claim 2, characterized in that: The contextual cognition reasoning module includes a contextual awareness integration submodule, an intent graph construction submodule, and a contextual evolution deduction submodule; The context-aware integration submodule is used to dynamically access and parse context information from digital scenes to establish a quantifiable environment state descriptor. The intent graph construction submodule is used to integrate user state representation and environmental state descriptor to drive a knowledge graph-based semantic association engine and generate an intent topology structure that includes the user's potential goals and behavioral paths. The context evolution deduction submodule is used to introduce a temporal memory network to dynamically track and probabilistically infer the intention topology, and output a contextual cognitive framework that reflects the current digital interaction.
4. The multimodal sensing immersive digital interaction system as described in claim 3, characterized in that: The immersive interactive generation module includes a multi-channel collaborative orchestration submodule, a real-time rendering driving submodule, and a closed-loop strategy evolution submodule. The multi-channel collaborative orchestration submodule is used to parse semantic instructions in the dynamic context framework and generate a multimodal interaction instruction set; The real-time rendering driver submodule is used to synchronously generate an immersive sensory output stream based on the multimodal interaction instruction set of the scheduled holographic biosensing matrix. The closed-loop strategy evolution submodule is used to collect user behavioral responses to immersive sensory output streams, evaluate the effectiveness of behavioral responses using reinforcement learning algorithms, dynamically adjust the multi-channel feedback mechanism, and complete the multi-channel feedback mechanism to optimize the interaction strategy.
5. The multimodal sensing immersive digital interaction system as described in claim 4, characterized in that: The heterogeneous sensing collaborative submodule includes a sensing timing alignment unit and a modal adaptive calibration unit. The sensing timing alignment unit is used to synchronize the original data streams of each modal sensor in the holographic biosensing matrix with nanosecond-level timestamps. The modal adaptive calibration unit is used to dynamically adjust the raw data stream of each modal sensor and calibrate the sensor units of different modalities in the holographic biosensing matrix. The cross-domain feature extraction submodule includes a spatiotemporal semantic coding unit and a cross-modal decoupling representation unit; The spatiotemporal semantic coding unit is used to jointly embed multimodal signals in a unified spatiotemporal coordinate system to generate a multi-channel feature tensor with position awareness capability. The cross-modal decoupling characterization unit is used to output semantic features for denoising multimodal signals by combating shared multi-channel feature tensor noise in decoupled multimodal signals. The semantic fusion representation submodule includes an attention-guided fusion unit and a dynamic representation compression unit; The attention-guided fusion unit is used to construct a hierarchical cross-modal attention mechanism, which dynamically weights the semantic contribution of different modal features based on the current interaction task. The dynamic representation compression unit is used to map high-dimensional fused features to a low-dimensional manifold space to generate a user state representation vector.
6. The multimodal sensing immersive digital interaction system as described in claim 5, characterized in that: The context-aware integration submodule includes a scene semantic parsing unit and an environment state quantization unit; The scene semantic parsing unit is used to perform ontology-driven semantic annotation on contextual information from digital scenes and extract spatial and event nodes of environmental states. The environmental state quantization unit is used to map the space of environmental state and event nodes to the context parameter space to generate an environmental state descriptor. The intent graph construction submodule includes a user environment semantic alignment unit and a dynamic intent topology generation unit; The user environment semantic alignment unit is used to align the user state representation vector and the environment state descriptor in the unified knowledge embedding space and identify their paths in the semantic association engine. The dynamic intent topology generation unit is used to activate concept nodes in the knowledge graph based on the recognized semantic association engine path to complete the generation of intent topology structure; The scenario evolution deduction submodule includes a temporal intent tracking unit and a probabilistic scenario deduction unit; The temporal intent tracking unit is used to update the weights of each node in the intent topology using a memory-enhanced recursive network; The probabilistic scenario inference unit is used to perform multi-step predictions based on the weights of each node in the updated intent topology, and output a contextual cognition framework.
7. The multimodal sensing immersive digital interaction system as described in claim 6, characterized in that: The multi-channel collaborative orchestration submodule includes a semantic action mapping unit and a cross-modal timing scheduling unit; The semantic action mapping unit is used to map the dynamic context framework to semantic actions. The cross-modal timing scheduling unit is used to sort multimodal interaction instructions based on semantic action mapping and generate a multimodal interaction instruction set; The real-time rendering driver submodule includes a heterogeneous output engine calling unit and a sensory stream synchronous synthesis unit; The heterogeneous output engine calling unit is used to process multimodal interaction instruction sets; The sensory stream synchronization synthesis unit is used to fuse the output signals of each modality under the processing of multimodal interaction instruction set to generate an immersive sensory output stream; The closed-loop policy evolution submodule includes a response feature capture unit and a policy gradient optimization unit; The response feature capture unit is used to extract the user's sensory signals to the current immersive sensory output stream from the holographic biosensing matrix and construct a multidimensional response feature vector. The policy gradient optimization unit is used to input the multi-dimensional response feature vector as a reward signal into the reinforcement learning algorithm, iteratively update the parameter configuration of the multi-channel feedback mechanism, and complete the optimization interaction strategy of the multi-channel feedback mechanism.
8. A multimodal sensing immersive digital interaction method, based on the multimodal sensing immersive digital interaction system according to any one of claims 1 to 7, characterized in that: include, The user's multidimensional human condition data is collected by a holographic biosensor matrix to generate a user state representation. Based on user state representation and combined with environmental context information, a dynamic contextual framework is constructed. Based on the dynamic context framework, a multi-channel feedback mechanism for driving digital interaction is used to generate immersive interactive content and optimize the interaction strategy through the multi-channel feedback mechanism.
9. A computer device comprising a memory and a processor, wherein the memory stores a computer program, characterized in that: When the processor executes the computer program, it implements the steps of the multimodal perception immersive digital interaction system according to any one of claims 1 to 7.
10. A computer-readable storage medium having a computer program stored thereon, characterized in that: When the computer program is executed by the processor, it implements the steps of the multimodal perception immersive digital interaction system according to any one of claims 1 to 7.