Multi-mode holographic light field interaction system, method and device based on AGI

Through the FusionNet network and Laplace operator adaptive smoothing processing, combined with cross-modal mapping module and light field reconstruction network, the problem of insufficient multimodal information fusion is solved, high-precision three-dimensional light field data generation and dynamic holographic image rendering are realized, and the intelligence and user experience of the interactive system are improved.

CN120541776AInactive Publication Date: 2025-08-26YIBU DISTANCE (JINAN) INTERNET OF THINGS TECH CO LTD
View PDF 0 Cites 3 Cited by

Patent Information

Application Number
CN202510648877.X
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-05-20
Publication Date
2025-08-26
Estimated Expiration
Not applicable · inactive patent

AI Technical Summary

Technical Problem

The prior art is difficult to effectively realize the precise fusion of multimodal information and the real-time rendering of dynamic holographic images. Especially when processing multiple sensory data, the system's response ability and accuracy are insufficient, making it difficult to provide a smooth and high-quality holographic visual experience.

Method used

The FusionNet network model is used to extract visual, audio and tactile features, combined with the Laplace operator for adaptive smoothing processing, and feature alignment is performed through the cross-modal mapping module. The light field reconstruction network is used to generate high-precision three-dimensional light field data, and dynamically optimize the holographic image details and rendering effects with real-time interactive feedback.

Benefits of technology

It realizes efficient processing and precise fusion of multimodal data, generates high-quality three-dimensional light field data, provides a smooth dynamic holographic image interactive experience, enhances the intelligence and adaptability of the system, and improves the personalization and dynamic response capabilities of the user experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120541776A_ABST
    Figure CN120541776A_ABST
Patent Text Reader

Abstract

The invention discloses an AGI-based multi-mode holographic light field interaction system, method and device. The method comprises the following steps: S1, collecting and preprocessing multi-mode interaction data; s2, constructing a FusionNet network model, extracting visual, audio and tactile features, and fusing to generate multi-modal feature representation; s3, adopting a Laplace operator to carry out adaptive smoothing processing on the features, retaining significant feature information and removing high-frequency noise; s4, aligning the multi-modal features through a cross-modal mapping module, and generating three-dimensional light field data by using the light field reconstruction network; s5, decomposing the three-dimensional light field data, converting the three-dimensional light field data into a dynamic holographic image through light field display equipment, and performing layered rendering; and S6, adjusting network parameters according to real-time feedback, and optimizing holographic image details and rendering effects. According to the invention, multi-modal data fusion and real-time holographic image generation can be realized, and the intellectualization and immersion of interaction experience are significantly improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to the field of light field interaction technology, and in particular to an AGI-based multimodal holographic light field interaction system, method, and device. Background Art

[0002] In the research of modern intelligent interaction technologies, with the rapid development of artificial intelligence (AI) and holographic technology, achieving a more natural, immersive, and intelligent user interaction experience has become a key issue. Although numerous technologies have been proposed to enhance the interactive experience, traditional interaction methods still have certain limitations. In particular, existing technologies often struggle to effectively achieve accurate and efficient multimodal information fusion and natural three-dimensional visual effects when combining multimodal perception with holographic light field interaction.

[0003] Traditional interactive technologies mostly rely on single-modal input methods, such as vision, touch, or audio. While these single-modal interactive systems have achieved certain application results in certain fields, their responsiveness and accuracy remain insufficient in complex environments, especially when processing multimodal data. For example, traditional systems based on vision and sound are generally unable to effectively process data from different senses, resulting in significant limitations in the user experience. These systems often cannot make immediate adjustments and adaptations based on user feedback, and they also struggle to effectively switch between multiple input methods in real time. To address this shortcoming, researchers in recent years have attempted to introduce the concept of multimodal perception, enabling systems to simultaneously process data from multiple senses, such as vision, hearing, and touch. However, even these multimodal perception systems often face a common challenge: how to extract effective features from massive amounts of sensory information and accurately integrate them for more efficient decision-making and feedback.

[0004] Especially when it comes to holographic display technology, while holographic images excel in enhancing visual effects and immersion, existing light field rendering techniques are often constrained by various factors, including processing power, computing resources, and image reconstruction accuracy. Traditional holographic technologies typically rely on complex optical devices and image reconstruction algorithms. However, due to technical limitations, these holographic displays are typically limited to static images and lack dynamic interaction. Although some systems have made progress in achieving dynamic light field displays, they still face many challenges in terms of interactive experience, image accuracy, and real-time performance.

[0005] Currently, research on image processing and multimodal data fusion using artificial intelligence (AI) technology is increasing, especially in areas such as computer vision, natural language processing, and voice recognition. However, the application of AI technology in these traditional fields is typically based on single-modal training and reasoning, and the challenges of multimodal information processing have not yet been effectively addressed. In particular, in AI-based systems, how to process and analyze data from different senses in real time, extract effective features, and make intelligent interactive decisions is an urgent problem that needs to be solved. AGI (artificial general intelligence), as an AI model with broad cognitive and reasoning capabilities, can process information across modalities and make adaptive adjustments based on different interactive scenarios. However, most current AGI systems still lack effective multimodal data fusion capabilities, and their application is particularly limited in dynamic interactive systems.

[0006] Furthermore, existing holographic display systems often rely on extensive computing resources to process the fusion of multimodal data and the rendering of three-dimensional light fields, which significantly limits the real-time and accuracy of these systems. Especially when processing high-dimensional data, existing holographic image rendering technologies often struggle to achieve efficient three-dimensional reconstruction and often fail to provide a smooth and high-quality holographic visual experience. Therefore, existing holographic display technologies still face many difficulties in dynamic interaction, especially how to accurately render complex holographic images during real-time interaction and ensure real-time updates and adaptation of the light field, which remains a significant challenge.

[0007] Therefore, how to provide a multimodal holographic light field interaction system, method and device based on AGI is a problem that technical personnel in this field urgently need to solve. Summary of the Invention

[0008] One purpose of the present invention is to propose an AGI-based multimodal holographic light field interaction system, method, and device. The present invention makes full use of AGI technologies such as holographic display technology and multimodal data fusion, and describes in detail how to achieve dynamic holographic image generation and interactive experience through steps such as multimodal feature extraction, feature alignment, and light field reconstruction. By using the FusionNet network model to extract visual, audio, and tactile features, the Laplace operator is used to adaptively smooth the features, and the cross-modal mapping module is combined for precise feature alignment, and high-precision three-dimensional light field data is further generated through the light field reconstruction network. Finally, the holographic image details and light field rendering effects are dynamically optimized in combination with real-time interactive feedback, thereby achieving a personalized interactive experience.

[0009] The AGI-based multimodal holographic light field interaction method according to an embodiment of the present invention includes the following steps:

[0010] S1. Collect multimodal interaction data, perform preprocessing, and generate a standardized data set;

[0011] S2. Based on the standardized dataset, a FusionNet network model is constructed to extract visual features, audio features, and tactile features, and fuse them to generate a multimodal feature representation;

[0012] S3. Adopting the Laplace operator to perform adaptive smoothing on the multimodal feature representation. The Laplace operator performs weighted smoothing based on local differences, retaining significant feature information and removing high-frequency noise.

[0013] S4. Aligning the multimodal feature representations through a cross-modal mapping module and generating three-dimensional light field data through a light field reconstruction network. The cross-modal mapping module uses a cross-modal alignment algorithm to accurately align features from different modalities by establishing a multi-level embedding space and performing iterative optimization.

[0014] S5. performing light field decomposition on the three-dimensional light field data, converting the decomposed light field data into a real-time dynamic holographic image through a light field display device, and performing layered rendering processing;

[0015] S6. Based on real-time interaction feedback, dynamically adjust network parameters, optimize holographic image details and light field rendering effects, and feed back to the next round of interaction, thereby achieving a personalized dynamic interactive experience.

[0016] Optionally, the multimodal interaction data includes visual images, audio signals and tactile feedback signals.

[0017] Optionally, the preprocessing includes noise removal, missing value filling, outlier removal and data standardization.

[0018] Optionally, the FusionNet network model includes a convolutional layer, a residual connection, a channel attention module and a feature fusion layer. The convolutional layer is used to extract visual, audio and tactile features at different levels. The residual connection is used to alleviate the gradient vanishing problem in information transmission. The channel attention module is used to weight the feature responses of different channels. The feature fusion layer fuses various features in multimodal data through a self-attention mechanism and a bidirectional encoding method to generate a comprehensive multimodal feature representation.

[0019] Optionally, the light field reconstruction network includes a light field projection layer and a three-dimensional feature mapping layer, the light field projection layer performs weighted projection on multimodal features according to spatial frequency, and the three-dimensional feature mapping layer performs feature mapping and three-dimensional coordinate reconstructing on the projection results to generate three-dimensional light field data.

[0020] Optionally, the S2 specifically includes:

[0021] S21. Based on the standardized data set, construct a FusionNet network model, wherein the FusionNet network model includes a convolutional layer, a residual connection, a channel attention module, and a feature fusion layer;

[0022] S22. Extracting visual features, audio features, and tactile features through a convolution layer. The convolution layer performs a layer-by-layer convolution operation on the standardized dataset using multiple convolution kernels to extract feature information at different levels. The convolution operation result is a preliminary representation of the visual features, audio features, and tactile features.

[0023] S23, alleviating the gradient vanishing problem in information transfer through residual connections, which directly connect the input and output feature maps to ensure that feature information is effectively transmitted in the network and avoid the gradient vanishing phenomenon in the FusionNet network;

[0024] S24. Weighting the feature responses of different channels by a channel attention module, wherein the channel attention module weights the features of each channel by calculating a weight coefficient of each channel;

[0025] S25. Various features in the multimodal data are fused through a feature fusion layer. The feature fusion layer adopts a self-attention mechanism to calculate the similarity between different features, weight each feature, and combine it with a bidirectional encoding method to deeply fuse the visual, audio, and tactile features to generate a comprehensive multimodal feature representation.

[0026] Optionally, the S3 specifically includes:

[0027] S31. Calculate the local difference of the multimodal feature representation using the Laplace operator:

[0028]

[0029] Where x and y represent the spatial positions of the multimodal feature representation, Δ(x,y) is the difference at position (x,y), f(x,y) represents the eigenvalue at position (x,y), f(x+i,y+j) represents the eigenvalue at position (x+i,y+j), and k is the size of the local window.

[0030] S32. Perform weighted smoothing on the Laplace operator according to the local difference, wherein the weighted smoothing operation is adjusted according to the local difference:

[0031] w(x,y)=exp(-α·Δ(x,y));

[0032] Where w(x,y) is the smoothing weight of the position (x,y), exp(·) represents the natural exponential function with e as the base, and α is the smoothing coefficient, which controls the strength of weighted smoothing.

[0033] S33. Perform weighted smoothing on the multimodal feature representation to obtain a smoothed feature representation:

[0034]

[0035] Where f′(x,y) is the smoothed eigenvalue, and w(x+i,y+j) is the smoothing weight of the position (x+i,y+j);

[0036] S34. Retain salient feature information and remove high-frequency noise. Through the weighted smoothing operation, the smoothed feature representation can effectively remove high-frequency noise while retaining important salient feature information.

[0037] Optionally, the S4 specifically includes:

[0038] S41. Align features from different modalities using a cross-modal alignment algorithm. The cross-modal alignment algorithm establishes a multi-level embedding space and performs iterative optimization to accurately align features from different modalities.

[0039]

[0040] Among them, f v (x i ),f a (x i ),f t (x i ) represent the embedded features of visual, audio and tactile features at the i-th modality position, L align is the alignment loss, N is the number of modalities, and λ is the weight for adjusting the alignment of audio and tactile features;

[0041] S42. Based on the aligned multimodal features, perform spatial frequency weighted projection through a light field reconstruction network, wherein the light field reconstruction network includes a light field projection layer and a three-dimensional feature mapping layer, and the light field projection layer performs weighted projection on the multimodal features:

[0042]

[0043] Among them, P proj (x,y,z) represents the projection result at position (x,y,z), f i (x, y, z) represents the value of the ith modal feature at position (x, y, z) after alignment, w i is the projection weight, which determines the contribution of different modal features to the projection results;

[0044] S43, performing feature mapping and three-dimensional coordinate reconstruction on the projection result through the three-dimensional feature mapping layer:

[0045]

[0046] Among them, F 3D (x, y, z) represents the value of the reconstructed three-dimensional light field data at position (x, y, z), T map (P proj (x, y, z)) is the mapping function from the projection result to the three-dimensional coordinates;

[0047] S44. Through the above-mentioned projection and mapping steps, three-dimensional light field data is obtained, and displayed through a light field display device.

[0048] The AGI-based multimodal holographic light field interaction system according to an embodiment of the present invention includes:

[0049] The data processing module is used to collect multimodal interaction data, perform preprocessing, and generate standardized data sets;

[0050] The feature fusion module is used to build the FusionNet network model, extract visual features, audio features, and tactile features, and fuse them to generate multimodal feature representations;

[0051] The feature smoothing module is used to adaptively smooth the multimodal feature representation using the Laplace operator, perform weighted smoothing based on local differences, retain significant feature information and remove high-frequency noise;

[0052] A feature alignment module is used to align multimodal feature representations through a cross-modal mapping module and generate 3D light field data through a light field reconstruction network;

[0053] The holographic rendering module is used to perform light field decomposition on 3D light field data, convert the decomposed light field data into real-time dynamic holographic images through a light field display device, and perform layered rendering processing;

[0054] The interactive feedback module is used to dynamically adjust network parameters based on real-time interactive feedback, optimize holographic image details and light field rendering effects, and feed back to the next round of interaction, thereby achieving a personalized dynamic interactive experience.

[0055] The AGI-based multimodal holographic light field interaction device according to an embodiment of the present invention includes:

[0056] memory for storing computer programs;

[0057] A processor is used to implement the AGI-based multimodal holographic light field interaction method when executing the computer program.

[0058] The beneficial effects of the present invention are:

[0059] First, this invention effectively addresses the problem of insufficient multimodal information fusion in traditional interactive systems by introducing an AGI-based multimodal holographic light field interaction method. Through the FusionNet network model, it accurately extracts features from different modalities, such as vision, audio, and touch, and generates a comprehensive multimodal feature representation through an adaptive feature fusion method. This feature fusion method, difficult to implement in traditional single-sensor modes, significantly improves the system's ability to process multimodal information in complex environments, enhancing the system's intelligence and adaptability.

[0060] Secondly, the present invention has made innovations in the generation and real-time interaction of holographic images. The Laplace operator is used to adaptively smooth multimodal features, effectively removing noise and retaining significant feature information, making the final generated holographic image more accurate and delicate. On this basis, a cross-modal mapping module is used to accurately align multimodal features, and a light field reconstruction network is used to generate high-quality three-dimensional light field data, ensuring high precision and realism of the image. In addition, the introduction of the light field decomposition and rendering module ensures that the image can be dynamically presented on the light field display device, providing a smooth interactive experience.

[0061] Finally, dynamic adjustment of network parameters based on real-time interactive feedback enables continuous optimization of holographic image details and fine-tuning of light field rendering effects, further enhancing the personalized and dynamic responsiveness of the interactive experience. As the interaction process deepens, the system intelligently adjusts based on user behavior and feedback, making each interaction more precise and tailored to user needs. This not only enhances the user experience but also makes the technology more adaptable and continuously optimized in practical applications. BRIEF DESCRIPTION OF THE DRAWINGS

[0062] The accompanying drawings are used to provide a further understanding of the present invention and constitute a part of the specification. Together with the embodiments of the present invention, they are used to explain the present invention and do not constitute a limitation of the present invention. In the accompanying drawings:

[0063] Figure 1 This is a flow chart of the AGI-based multimodal holographic light field interaction method proposed in the present invention;

[0064] Figure 2 Flowchart of feature alignment and 3D light field data generation for the AGI-based multimodal holographic light field interaction method proposed in this invention;

[0065] Figure 3 This is a module structure diagram of the AGI-based multimodal holographic light field interaction system proposed in this invention. DETAILED DESCRIPTION

[0066] The present invention will now be described in further detail with reference to the accompanying drawings, which are simplified schematic diagrams that illustrate the basic structure of the present invention in a schematic manner.

[0067] refer to Figure 1-2 The AGI-based multimodal holographic light field interaction method includes the following steps:

[0068] S1. Collect multimodal interaction data, perform preprocessing, and generate a standardized data set;

[0069] S2. Based on the standardized dataset, a FusionNet network model is constructed to extract visual features, audio features, and tactile features, and fuse them to generate a multimodal feature representation;

[0070] S3. Adopting the Laplace operator to perform adaptive smoothing on the multimodal feature representation. The Laplace operator performs weighted smoothing based on local differences, retaining significant feature information and removing high-frequency noise.

[0071] S4. Aligning the multimodal feature representations through a cross-modal mapping module and generating three-dimensional light field data through a light field reconstruction network. The cross-modal mapping module uses a cross-modal alignment algorithm to accurately align features from different modalities by establishing a multi-level embedding space and performing iterative optimization.

[0072] S5. performing light field decomposition on the three-dimensional light field data, converting the decomposed light field data into a real-time dynamic holographic image through a light field display device, and performing layered rendering processing;

[0073] S6. Based on real-time interaction feedback, dynamically adjust network parameters, optimize holographic image details and light field rendering effects, and feed back to the next round of interaction, thereby achieving a personalized dynamic interactive experience.

[0074] This method utilizes AGI technology to enhance the system's ability to process complex data by fusing multimodal data such as vision, audio, and touch. Through precise feature extraction, fusion, smoothing, and alignment, this method generates highly accurate three-dimensional light field data. It also optimizes holographic image rendering through dynamic interactive adjustment of system parameters, enabling a personalized, dynamic interactive experience.

[0075] In this embodiment, the multimodal interaction data includes visual images, audio signals and tactile feedback signals.

[0076] The present invention clarifies the specific composition of multimodal interaction data, including visual images, audio signals and tactile feedback signals, thereby ensuring the full integration of different sensory information and improving the intelligence of the interactive system and user experience.

[0077] In this embodiment, the preprocessing includes noise removal, missing value filling, outlier removal and data standardization.

[0078] The present invention adopts technologies such as noise removal, missing value filling, outlier elimination and data normalization in the data preprocessing stage to ensure the quality of input data, improve the accuracy of subsequent feature extraction and fusion, and provide a more reliable foundation for multimodal data processing.

[0079] In this embodiment, the FusionNet network model includes a convolutional layer, a residual connection, a channel attention module and a feature fusion layer. The convolutional layer is used to extract visual, audio and tactile features at different levels. The residual connection is used to alleviate the gradient vanishing problem in information transmission. The channel attention module is used to weight the feature responses of different channels. The feature fusion layer fuses various features in multimodal data through a self-attention mechanism and a bidirectional encoding method to generate a comprehensive multimodal feature representation.

[0080] The present invention effectively improves the extraction and fusion capabilities of multimodal features through the design of the FusionNet network model, including convolutional layers, residual connections, channel attention modules and feature fusion layers, especially in alleviating the gradient vanishing problem, weighting the feature responses of different channels and deeply fusing multimodal data, greatly improving the accuracy and efficiency of feature processing.

[0081] In this embodiment, the light field reconstruction network includes a light field projection layer and a three-dimensional feature mapping layer. The light field projection layer performs weighted projection on multimodal features according to spatial frequency, and the three-dimensional feature mapping layer performs feature mapping and three-dimensional coordinate reconstruction on the projection results to generate three-dimensional light field data.

[0082] The light field reconstruction network of the present invention combines the light field projection layer and the three-dimensional feature mapping layer, adopts spatial frequency weighted projection and three-dimensional coordinate reconstruction technology, can accurately generate high-quality three-dimensional light field data, and improves the rendering effect and accuracy of holographic images.

[0083] In this embodiment, S2 specifically includes:

[0084] S21. Based on the standardized data set, construct a FusionNet network model, wherein the FusionNet network model includes a convolutional layer, a residual connection, a channel attention module, and a feature fusion layer;

[0085] S22. Extracting visual features, audio features, and tactile features through a convolution layer. The convolution layer performs a layer-by-layer convolution operation on the standardized dataset using multiple convolution kernels to extract feature information at different levels. The convolution operation result is a preliminary representation of the visual features, audio features, and tactile features.

[0086] S23, alleviating the gradient vanishing problem in information transfer through residual connections, which directly connect the input and output feature maps to ensure that feature information is effectively transmitted in the network and avoid the gradient vanishing phenomenon in the FusionNet network;

[0087] S24. Weighting the feature responses of different channels by a channel attention module, wherein the channel attention module weights the features of each channel by calculating a weight coefficient of each channel;

[0088] S25. Various features in the multimodal data are fused through a feature fusion layer. The feature fusion layer adopts a self-attention mechanism to calculate the similarity between different features, weight each feature, and combine it with a bidirectional encoding method to deeply fuse the visual, audio, and tactile features to generate a comprehensive multimodal feature representation.

[0089] This paper describes in detail each step in the FusionNet network model, from feature extraction in the convolutional layer to optimization of the residual connection, channel attention, and feature fusion layers, effectively improving the depth and breadth of multimodal data fusion and ensuring that the generated multimodal feature representation is of high quality and accuracy.

[0090] In this embodiment, S3 specifically includes:

[0091] S31. Calculate the local difference of the multimodal feature representation using the Laplace operator:

[0092]

[0093] Where x and y represent the spatial positions of the multimodal feature representation, Δ(x,y) is the difference at position (x,y), f(x,y) represents the eigenvalue at position (x,y), f(x+i,y+j) represents the eigenvalue at position (x+i,y+j), and k is the size of the local window.

[0094] S32. Perform weighted smoothing on the Laplace operator according to the local difference, wherein the weighted smoothing operation is adjusted according to the local difference:

[0095] w(x,y)=exp(-α·Δ(x,y));

[0096] Where w(x,y) is the smoothing weight of the position (x,y), exp(·) represents the natural exponential function with e as the base, and α is the smoothing coefficient, which controls the strength of weighted smoothing.

[0097] S33. Perform weighted smoothing on the multimodal feature representation to obtain a smoothed feature representation:

[0098]

[0099] Where f′(x,y) is the smoothed eigenvalue, and w(x+i,y+j) is the smoothing weight of the position (x+i,y+j);

[0100] S34. Retain salient feature information and remove high-frequency noise. Through the weighted smoothing operation, the smoothed feature representation can effectively remove high-frequency noise while retaining important salient feature information.

[0101] The present invention introduces the Laplace operator to perform adaptive smoothing on multimodal features, which can effectively remove high-frequency noise and retain important significant feature information, further improving the clarity and accuracy of features and providing high-quality data support for subsequent light field generation.

[0102] In this embodiment, the S4 specifically includes:

[0103] S41. Align features from different modalities using a cross-modal alignment algorithm. The cross-modal alignment algorithm establishes a multi-level embedding space and performs iterative optimization to accurately align features from different modalities.

[0104]

[0105] Among them, f v (x i ),f a (x i ),f t (x i ) represent the embedded features of visual, audio and tactile features at the i-th modality position, L align is the alignment loss, N is the number of modalities, and λ is the weight for adjusting the alignment of audio and tactile features;

[0106] S42. Based on the aligned multimodal features, perform spatial frequency weighted projection through a light field reconstruction network, wherein the light field reconstruction network includes a light field projection layer and a three-dimensional feature mapping layer, and the light field projection layer performs weighted projection on the multimodal features:

[0107]

[0108] Among them, P proj (x,y,z) represents the projection result at position (x,y,z), f i (x, y, z) represents the value of the ith modal feature at position (x, y, z) after alignment, w i is the projection weight, which determines the contribution of different modal features to the projection results;

[0109] S43, performing feature mapping and three-dimensional coordinate reconstruction on the projection result through the three-dimensional feature mapping layer:

[0110]

[0111] Among them, F 3D (x, y, z) represents the value of the reconstructed three-dimensional light field data at position (x, y, z), T map (P proj (x, y, z)) is the mapping function from the projection result to the three-dimensional coordinates;

[0112] S44. Through the above-mentioned projection and mapping steps, three-dimensional light field data is obtained, and displayed through a light field display device.

[0113] By combining a cross-modal alignment algorithm with a light field reconstruction network, the present invention enables features from different modalities to be accurately aligned and generate three-dimensional light field data, greatly improving the accuracy of feature alignment and ensuring the matching and authenticity of the ultimately generated three-dimensional light field data with the actual environment.

[0114] refer to Figure 3 , an AGI-based multimodal holographic light field interaction system, including:

[0115] The data processing module is used to collect multimodal interaction data, perform preprocessing, and generate standardized data sets;

[0116] The feature fusion module is used to build the FusionNet network model, extract visual features, audio features, and tactile features, and fuse them to generate multimodal feature representations;

[0117] The feature smoothing module is used to adaptively smooth the multimodal feature representation using the Laplace operator, perform weighted smoothing based on local differences, retain significant feature information and remove high-frequency noise;

[0118] A feature alignment module is used to align multimodal feature representations through a cross-modal mapping module and generate 3D light field data through a light field reconstruction network;

[0119] The holographic rendering module is used to perform light field decomposition on 3D light field data, convert the decomposed light field data into real-time dynamic holographic images through a light field display device, and perform layered rendering processing;

[0120] The interactive feedback module is used to dynamically adjust network parameters based on real-time interactive feedback, optimize holographic image details and light field rendering effects, and feed back to the next round of interaction, thereby achieving a personalized dynamic interactive experience.

[0121] This invention builds a complete multimodal holographic light field interaction system through the collaborative work of a data processing module, a feature fusion module, a feature smoothing module, a feature alignment module, a holographic rendering module, and an interactive feedback module. The effective collaboration between these modules improves the system's overall performance and real-time feedback capabilities, providing users with a more fluid and personalized interactive experience.

[0122] AGI-based multimodal holographic light field interaction device, including:

[0123] memory for storing computer programs;

[0124] A processor is used to implement the AGI-based multimodal holographic light field interaction method when executing the computer program.

[0125] The AGI-based multimodal holographic light field interaction device provided by the present invention realizes the effective execution of the methods described in claims 1 to 8 through the cooperation of memory and processor, ensuring that the system can efficiently process multimodal data and generate high-quality dynamic holographic images, further improving the practicality and efficiency of the device in actual applications.

[0126] Example 1:

[0127] To verify the feasibility of the present invention in practice, the present invention was applied to a virtual reality (VR) interactive entertainment system, aiming to enhance the user's interactive experience in a holographic environment. In this system, users interact with the system through multimodal sensory inputs such as vision, audio, and touch. At the same time, the system performs intelligent analysis and processing based on AGI technology, generates high-quality three-dimensional holographic images in real time, and feeds them back to the user. The following is a detailed description of the embodiment, showing how the technical solution of the present invention can address the shortcomings of traditional interactive systems in multimodal data fusion, real-time feedback, and high-quality holographic image generation.

[0128] In a virtual reality interactive entertainment system, users wear a VR headset, haptic feedback gloves, and headphones to interact. The system collects the user's visual images, audio signals, and tactile feedback signals to generate multimodal interaction data. After preprocessing, this data enters the FusionNet network model for feature extraction and fusion. Specifically, the system first uses a camera to capture a real-time image of the user and performs image processing to extract visual features such as the user's facial expressions and hand movements. Simultaneously, ambient audio information collected by the headphones and vibration signals from the tactile gloves are also collected and processed into audio and tactile features. The preprocessing process includes steps such as noise removal, missing value filling, and data normalization to ensure the quality and consistency of the multimodal data.

[0129] The system then uses the FusionNet network model to extract visual, audio, and tactile features and deeply fuses them to generate a comprehensive multimodal feature representation. In this process, convolutional layers are used to extract features at different levels, residual connections address the vanishing gradient problem in information transfer, the channel attention module weights the feature responses of different channels, and the feature fusion layer deeply fuses multimodal data using a self-attention mechanism. This process enables the system to fully understand and integrate diverse sensory information, providing accurate data support for subsequent interaction and image generation.

[0130] After feature fusion is complete, the system enters the Laplace operator adaptive smoothing phase. The Laplace operator performs weighted smoothing on multimodal features based on local differences, preserving significant feature information while effectively removing high-frequency noise. This smoothing process enables the system to extract clearer and more accurate features, further improving image quality and interaction accuracy.

[0131] Next, the system aligns features from different modalities through a cross-modal mapping module. The cross-modal alignment algorithm employed precisely aligns features from different modalities through multi-level embedding and iterative optimization, providing more accurate input data for the light field reconstruction network. The light field reconstruction network generates 3D light field data through a light field projection layer and a 3D feature mapping layer, further enhancing the spatial expressiveness and realism of the holographic image.

[0132] Through the light field decomposition and rendering module, 3D light field data is converted into real-time dynamic holographic images, which are presented to users via a light field display device. This display device dynamically adjusts the details and rendering effects of the holographic images based on real-time user feedback, making the interaction process more fluid and immersive. During the interaction process, users can receive instant responses, whether using gestures, voice commands, or tactile feedback, significantly improving the interactive experience.

[0133] To verify the effectiveness of the method of the present invention, we conducted multiple rounds of tests on the virtual reality interactive entertainment system, recording the user's response time, system rendering effect, interaction accuracy and other performance during the interaction process. The test data is as follows:

[0134] Table 1 Comparison of experimental data before and after the test

[0135]

[0136]

[0137] The test data shows that the method of the present invention significantly improves system performance and user experience in many aspects. In particular, it has achieved significant improvements in response time, image rendering quality, and interaction accuracy. These results show that the present invention not only solves the problems of interaction delay and low image quality existing in the prior art, but also achieves a personalized and precise interaction experience through intelligent multimodal data fusion and real-time feedback optimization, greatly improving the overall performance of the virtual reality interactive entertainment system.

[0138] The embodiments of the present invention successfully address the shortcomings of traditional interactive systems in processing complex sensory input, real-time feedback, and high-quality image generation by combining advanced AGI technology, holographic light field display technology, and multimodal data processing methods, verifying the feasibility and superiority of the present invention.

[0139] The above description is only a preferred specific embodiment of the present invention, but the scope of protection of the present invention is not limited thereto. Any technician familiar with the technical field, within the technical scope disclosed by the present invention, who makes equivalent replacements or changes based on the technical solution and inventive concept of the present invention, should be covered by the scope of protection of the present invention.

Claims

1. AGI-based multimodal holographic light field interaction method, characterized by: The steps include: S1. Collect multimodal interaction data, perform preprocessing, and generate a standardized data set; S2. Based on the standardized dataset, a FusionNet network model is constructed to extract visual features, audio features, and tactile features, and fuse them to generate a multimodal feature representation; S3. Adopting the Laplace operator to perform adaptive smoothing on the multimodal feature representation. The Laplace operator performs weighted smoothing based on local differences, retaining significant feature information and removing high-frequency noise. S4. Aligning the multimodal feature representations through a cross-modal mapping module and generating three-dimensional light field data through a light field reconstruction network. The cross-modal mapping module uses a cross-modal alignment algorithm to accurately align features from different modalities by establishing a multi-level embedding space and performing iterative optimization. S5. performing light field decomposition on the three-dimensional light field data, converting the decomposed light field data into a real-time dynamic holographic image through a light field display device, and performing layered rendering processing; S6. Based on real-time interaction feedback, dynamically adjust network parameters, optimize holographic image details and light field rendering effects, and feed back to the next round of interaction, thereby achieving a personalized dynamic interactive experience.

2. The AGI-based multimodal holographic light field interaction method according to claim 1, characterized in that: The multimodal interaction data includes visual images, audio signals and tactile feedback signals.

3. The AGI-based multimodal holographic light field interaction method according to claim 1, characterized in that: The preprocessing includes noise removal, missing value filling, outlier removal and data normalization.

4. The AGI-based multimodal holographic light field interaction method according to claim 1, characterized in that: The FusionNet network model includes convolutional layers, residual connections, channel attention modules and feature fusion layers. The convolutional layers are used to extract visual, audio and tactile features at different levels. The residual connections are used to alleviate the gradient vanishing problem in information transmission. The channel attention modules are used to weight the feature responses of different channels. The feature fusion layer fuses various features in multimodal data through a self-attention mechanism and bidirectional encoding to generate a comprehensive multimodal feature representation.

5. The AGI-based multimodal holographic light field interaction method according to claim 1, characterized in that: The light field reconstruction network includes a light field projection layer and a three-dimensional feature mapping layer. The light field projection layer performs weighted projection on multimodal features according to spatial frequency, and the three-dimensional feature mapping layer performs feature mapping and three-dimensional coordinate reconstruction on the projection results to generate three-dimensional light field data.

6. The AGI-based multimodal holographic light field interaction method according to claim 1, characterized in that: The S2 specifically includes: S21. Based on the standardized data set, construct a FusionNet network model, wherein the FusionNet network model includes a convolutional layer, a residual connection, a channel attention module, and a feature fusion layer; S22. Extracting visual features, audio features, and tactile features through a convolution layer. The convolution layer performs a layer-by-layer convolution operation on the standardized dataset using multiple convolution kernels to extract feature information at different levels. The convolution operation result is a preliminary representation of the visual features, audio features, and tactile features. S23, alleviating the gradient vanishing problem in information transfer through residual connections, which directly connect the input and output feature maps to ensure that feature information is effectively transmitted in the network and avoid the gradient vanishing phenomenon in the FusionNet network; S24. Weighting the feature responses of different channels by a channel attention module, wherein the channel attention module weights the features of each channel by calculating a weight coefficient of each channel; S25. Various features in the multimodal data are fused through a feature fusion layer. The feature fusion layer adopts a self-attention mechanism to calculate the similarity between different features, weight each feature, and combine it with a bidirectional encoding method to deeply fuse the visual, audio, and tactile features to generate a comprehensive multimodal feature representation.

7. The AGI-based multimodal holographic light field interaction method according to claim 1, characterized in that: The S3 specifically includes: S31. Calculate the local difference of the multimodal feature representation using the Laplace operator: Where x and y represent the spatial positions of the multimodal feature representation, Δ(x,y) is the difference at position (x,y), f(x,y) represents the eigenvalue at position (x,y), f(x+i,y+j) represents the eigenvalue at position (x+i,y+j), and k is the size of the local window. S32. Perform weighted smoothing on the Laplace operator according to the local difference, wherein the weighted smoothing operation is adjusted according to the local difference: w(x,y)=exp(-α·Δ(x,y)); Where w(x,y) is the smoothing weight of the position (x,y), exp(·) represents the natural exponential function with e as the base, and α is the smoothing coefficient, which controls the strength of weighted smoothing. S33. Perform weighted smoothing on the multimodal feature representation to obtain a smoothed feature representation: Where f′(x,y) is the smoothed eigenvalue, and w(x+i,y+j) is the smoothing weight of the position (x+i,y+j); S34. Retain salient feature information and remove high-frequency noise. Through the weighted smoothing operation, the smoothed feature representation can effectively remove high-frequency noise while retaining important salient feature information.

8. The AGI-based multimodal holographic light field interaction method according to claim 1, characterized in that: The S4 specifically includes: S41. Align features from different modalities using a cross-modal alignment algorithm. The cross-modal alignment algorithm establishes a multi-level embedding space and performs iterative optimization to accurately align features from different modalities. Among them, f v (x i ),f a (x i ),f t (x i ) represent the embedded features of visual, audio and tactile features at the i-th modality position, L align is the alignment loss, N is the number of modalities, and λ is the weight for adjusting the alignment of audio and tactile features; S42. Based on the aligned multimodal features, perform spatial frequency weighted projection through a light field reconstruction network, wherein the light field reconstruction network includes a light field projection layer and a three-dimensional feature mapping layer, and the light field projection layer performs weighted projection on the multimodal features: Among them, P proj (x,y,z) represents the projection result at position (x,y,z), f i (x, y, z) represents the value of the ith modal feature at position (x, y, z) after alignment, w i is the projection weight, which determines the contribution of different modal features to the projection results; S43, performing feature mapping and three-dimensional coordinate reconstruction on the projection result through the three-dimensional feature mapping layer: Among them, F 3D (x, y, z) represents the value of the reconstructed three-dimensional light field data at position (x, y, z), T map (P proj (x, y, z)) is the mapping function from the projection result to the three-dimensional coordinates; S44. Through the above-mentioned projection and mapping steps, three-dimensional light field data is obtained, and displayed through a light field display device.

9. An AGI-based multimodal holographic light field interaction system, which executes the AGI-based multimodal holographic light field interaction method according to any one of claims 1 to 8, characterized in that: include: The data processing module is used to collect multimodal interaction data, perform preprocessing, and generate standardized data sets; The feature fusion module is used to build the FusionNet network model, extract visual features, audio features, and tactile features, and fuse them to generate multimodal feature representations; The feature smoothing module is used to adaptively smooth the multimodal feature representation using the Laplace operator, perform weighted smoothing based on local differences, retain significant feature information and remove high-frequency noise; A feature alignment module is used to align multimodal feature representations through a cross-modal mapping module and generate 3D light field data through a light field reconstruction network; The holographic rendering module is used to perform light field decomposition on 3D light field data, convert the decomposed light field data into real-time dynamic holographic images through a light field display device, and perform layered rendering processing; The interactive feedback module is used to dynamically adjust network parameters based on real-time interactive feedback, optimize holographic image details and light field rendering effects, and feed back to the next round of interaction, thereby achieving a personalized dynamic interactive experience.

10. AGI-based multimodal holographic light field interaction device, characterized by: include: Memory for storing computer programs; A processor, configured to implement the AGI-based multimodal holographic light field interaction method according to any one of claims 1 to 8 when executing the computer program.

Citation Information

Cited By

  • AI-supported multi-modal desktop interactive projection method and system

    CN121433504A

  • An ai-supported multi-modal tabletop interaction projection method and system

    CN121433504B

  • Intelligent interaction method and system based on generative artificial intelligence and light field display

    CN121455344A