An adaptive augmented reality content presentation method based on artificial intelligence

By acquiring multimodal data and modeling user cognitive states, combined with 3D scene understanding and real-time interactive intent recognition, the problem of insufficient dynamic response in existing augmented reality content presentation methods has been solved. This has enabled adaptive personalized content presentation and efficient interaction, improving user experience and system performance.

CN122239941APending Publication Date: 2026-06-19郑凌昊
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
郑凌昊
Filing Date
2026-03-24
Publication Date
2026-06-19

AI Technical Summary

Technical Problem

Existing augmented reality content presentation methods lack dynamic responses to user cognitive states, environmental changes, and interaction needs. They cannot effectively combine multimodal sensor data, resulting in insufficient personalized content presentation, unnatural and unsmooth interactions, and inflexible system resource scheduling, leading to a decline in performance and user experience.

Method used

By employing multimodal environmental data acquisition and dynamic modeling of user cognitive states, combined with 3D scene semantic understanding and augmented reality content generation, and through real-time interactive intent recognition and dynamic allocation of system resources, the system optimizes user attention prediction and ambient lighting adaptation to achieve adaptive content presentation.

Benefits of technology

It enables personalized content presentation, improves content quality and adaptability, enhances the naturalness and efficiency of interaction, and ensures the optimization of user experience and system performance.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122239941A_ABST
    Figure CN122239941A_ABST
Patent Text Reader

Abstract

This invention relates to the field of augmented reality technology, specifically to an artificial intelligence-based adaptive augmented reality content presentation method, comprising the following steps: S1, multimodal environment data acquisition and preprocessing; S2, dynamic modeling of user cognitive state; S3, 3D scene semantic understanding modeling; S4, augmented reality content generation optimization; S5, dynamic prediction and guidance of user attention; S6, real-time interactive intent recognition and response; S7, dynamic adaptation of ambient lighting and visual comfort; S8, dynamic allocation and optimization of system resources; and S9, continuous optimization of system performance evaluation. This invention provides personalized basis for content presentation strategies through dynamic modeling of user cognitive state; provides a spatial understanding foundation for content placement and interaction through 3D scene semantic understanding and modeling; improves content quality and adaptability through augmented reality content generation and optimization; and enhances information transmission efficiency through dynamic prediction and guidance of user attention.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of augmented reality technology, and more specifically to an adaptive augmented reality content presentation method based on artificial intelligence. Background Technology

[0002] In recent years, with the rapid development of artificial intelligence (AI) and augmented reality (AR) technologies, AI-based AR applications have been widely used in education, healthcare, industry, and many other fields. AR technology, through the fusion of virtual and reality, provides users with an interactive and immersive experience, while AI technology, through multimodal data processing, user behavior analysis, and intelligent reasoning, offers greater personalization and adaptability for the generation and presentation of AR content. Although traditional methods of presenting augmented reality content have made some progress, they still face many limitations in areas such as dynamic adaptation, user cognition, and interaction optimization.

[0003] Current augmented reality (AR) content presentation methods mostly employ static content generation and presentation strategies, lacking dynamic responses to user cognitive states, environmental changes, and interaction needs. Many systems fail to effectively integrate multimodal sensor data, such as EEG signals and eye-tracking data, to perceive the user's cognitive state in real time, resulting in a lack of personalized content presentation. Although some methods attempt to achieve environmental understanding through 3D scene modeling, they often cannot efficiently handle dynamic changes in complex scenes, leading to unnatural and unsmooth content placement and interaction. Existing interactive intent recognition relies heavily on single sensory sources, such as gestures or speech recognition, failing to fully integrate multi-channel interactive signals such as vision and hearing, limiting the richness and accuracy of interaction. Many systems have relatively simple resource scheduling and optimization, unable to flexibly allocate computing resources according to task requirements and device status, leading to a decline in system performance and user experience under high load. Therefore, we propose an AI-based adaptive augmented reality content presentation method. Summary of the Invention

[0004] In view of the above-mentioned shortcomings of the existing technology, the first objective of the present invention is to provide an adaptive augmented reality content presentation method based on artificial intelligence, thereby solving the problems in the background technology.

[0005] To achieve the above objectives, the present invention provides the following technical solution:

[0006] An AI-based adaptive augmented reality content presentation method includes the following steps:

[0007] S1. Multimodal environment data acquisition and preprocessing;

[0008] S2, Dynamic modeling of user cognitive state;

[0009] S3, 3D scene semantic understanding modeling;

[0010] S4. Augmented reality content generation optimization;

[0011] S5, Dynamic Prediction and Guidance of User Attention;

[0012] S6. Real-time interactive intent recognition and response;

[0013] S7. Dynamic adaptation to visual comfort under ambient lighting;

[0014] S8. System resource dynamic allocation optimization;

[0015] S9, System performance evaluation is continuously optimized.

[0016] The present invention is further configured such that: in step S1, multimodal environment data acquisition and preprocessing:

[0017] S1.1 Deploy a multimodal sensing network consisting of an RGB-D camera, an inertial measurement unit (IMU), an ambient light sensor, and a depth sensor, and use a timestamp alignment mechanism to achieve spatiotemporal synchronization of heterogeneous data;

[0018] S1.2 To address Gaussian noise, salt-and-pepper noise, and motion blur interference in sensor-acquired data, a cascaded processing framework of adaptive Kalman filtering and bilateral filtering is constructed.

[0019] S1.3. Employ a dimensionality reduction strategy that combines principal component analysis (PCA) with manifold learning to reduce the dimensionality of high-dimensional environmental features;

[0020] S1.4 Construct a spatiotemporal context model based on a Long Short-Term Memory (LSTM) network to perform temporal correlation modeling on continuous frame data;

[0021] S1.5 Establish a multi-index data quality assessment system, and use quantitative indicators such as information entropy, signal-to-noise ratio, and feature integrity to monitor the preprocessing results in real time.

[0022] The present invention is further configured such that: in step S2, dynamic modeling of user cognitive state:

[0023] S2.1 Deploy wearable physiological sensing devices to simultaneously collect users' electroencephalogram (EEG), eye movement trajectory, skin conductance response (GSR), and heart rate variability (HRV) data;

[0024] S2.2 Perform feature engineering processing on the collected physiological signals to extract quantitative indicators that reflect the user's cognitive state;

[0025] S2.3. A cognitive state classification model is constructed using a deep belief network (DBN) to classify users' cognitive states into five categories: focused, fatigued, confused, distracted, and relaxed.

[0026] S2.4 Construct a dynamic tracking model of cognitive state based on particle filtering to achieve real-time estimation and prediction of the user's cognitive state;

[0027] S2.5 Establish a baseline library of user cognitive features and achieve personalized calibration of the cognitive model through transfer learning;

[0028] In step S2, dynamic modeling of user cognitive state, the following features are used to fuse multimodal physiological signals:

[0029]

[0030] In the formula, K represents the query vector, derived from current cognitive state features; K represents the key vector, derived from historical physiological signal features; V represents the value vector, containing multimodal physiological features. Indicates the dimension of the key vector; The (·) function normalizes the attention score into a probability distribution. This formula calculates the similarity between the query and the key, assigns dynamic weights to the physiological features of different modalities, and achieves effective fusion of multi-source information.

[0031] The present invention is further configured such that: in step S3, the three-dimensional scene semantic understanding modeling:

[0032] S3.1. A fusion method based on structure of motion reconstruction (SfM) and multi-view stereo matching (MVS) is used to reconstruct the three-dimensional geometry of the scene;

[0033] S3.2 Construct a multi-scale semantic segmentation network based on Transformer to achieve pixel-level classification and recognition of objects in the scene;

[0034] S3.3 Based on the semantic segmentation results, construct a topological relationship graph of the scene to describe the spatial location and functional association between objects;

[0035] S3.4 Design a scene dynamic change detection algorithm based on Siamese network to monitor dynamic events in the environment in real time;

[0036] S3.5 Establish a multi-dimensional scenario understanding confidence assessment system to quantify the reliability of scenario modeling;

[0037] In step S3, the semantic understanding of the 3D scene is used to optimize the semantic segmentation model.

[0038]

[0039] In the formula For cross-entropy loss, Indicates the true label, Indicates the predicted probability; For Dice's loss; The weight coefficient (0.7) balances the contributions of the two losses. This hybrid loss function solves the class imbalance problem in semantic segmentation and improves the accuracy of the segmentation boundary.

[0040] The present invention is further configured such that: in step S4, during the optimization of augmented reality content generation:

[0041] S4.1 Construct a BERT-based context-aware requirement parsing model to accurately understand users' AR content needs;

[0042] S4.2. Based on the parsed requirements, call the corresponding content generation module to generate multimodal AR content including text, images, and 3D models;

[0043] S4.3 Establish a content optimization framework based on reinforcement learning to dynamically adjust AR content attributes according to scene characteristics and user cognitive state;

[0044] S4.4 Design an AR content spatial layout optimization algorithm based on genetic algorithm to achieve reasonable placement of multiple contents;

[0045] S4.5. Employs physically based rendering (PBR) technology and real-time global illumination algorithms to enhance the realism and immersion of AR content;

[0046] In step S4, augmented reality content generation, the content spatial layout is optimized.

[0047]

[0048] In the formula This represents the visibility score of the i-th AR content; The occlusion overlap between the i-th and j-th content is represented; S represents the total area occupied by all AR content. , , The weighting coefficients are set to 0.5, 0.3, and 0.2 respectively. This objective function comprehensively optimizes content visibility, occlusion, and space occupation to achieve the optimal layout of AR content.

[0049] The present invention is further configured such that: in step S5, the user attention dynamic prediction guidance:

[0050] S5.1 Integrating bottom-up and top-down attention mechanisms to construct a visual attention prediction model that conforms to the characteristics of AR scenarios;

[0051] S5.2 Construct a user attention shift prediction model based on Hidden Markov Model (HMM) to predict the attention trajectory in the next 1-3 seconds;

[0052] S5.3. Generate personalized attention guidance strategies based on attention prediction results;

[0053] S5.4 Establish an attention load monitoring and balancing mechanism to prevent cognitive fatigue caused by information overload;

[0054] S5.5 Build a closed-loop attention feedback system to verify the attention guidance effect through user interaction behavior and continuously optimize the model.

[0055] The present invention is further configured such that: in step S6, the real-time interactive intent recognition response:

[0056] S6.1 Deploy a multimodal interaction perception system to simultaneously collect user gestures, voice, eye movements, and head posture interaction signals;

[0057] S6.2 Construct a multimodal interaction intent recognition model based on Transformer to achieve real-time classification and prediction of user interaction intent;

[0058] S6.3. Generate an adaptive interaction response strategy based on the identified interaction intent;

[0059] S6.4 Design a multimodal interactive feedback system to provide interactive status feedback through multiple channels including vision, hearing, and touch;

[0060] S6.5 Establish an interactive performance evaluation index system, including response latency, recognition accuracy, operation efficiency and user satisfaction, and continuously optimize and improve the interactive experience.

[0061] The present invention is further configured such that: in step S7, dynamic adaptation of ambient lighting visual comfort:

[0062] S7.1. By combining ambient light sensors with image analysis, the ambient light characteristics are comprehensively extracted.

[0063] S7.2 Construct a visual comfort assessment model based on physiological indicators and subjective evaluation;

[0064] S7.3. Dynamically adjust the display parameters of AR content based on ambient lighting characteristics and visual comfort model;

[0065] S7.4 Optimize stereoscopic vision parameters to reduce visual fatigue, taking into account the stereoscopic display characteristics of AR devices;

[0066] S7.5 Build a real-time visual fatigue monitoring system to determine the user's visual fatigue state through the fusion of multiple physiological indicators;

[0067] In step S7, the adaptation of ambient lighting and visual comfort is used to predict user visual comfort.

[0068]

[0069] In the formula, C represents the visual comfort score (1-5 points); This represents the k-th visual parameter (brightness, contrast, saturation); For the intercept term, The coefficients are linear. The coefficients of the interaction terms; As the error term, this multivariate nonlinear regression model takes into account the interaction between visual parameters, and can more accurately predict user comfort under different display conditions.

[0070] The present invention is further configured such that: in step S8, system resource dynamic allocation optimization:

[0071] S8.1 Deploy a system resource monitoring module to collect real-time status parameters of CPU, GPU, memory, storage, and network;

[0072] S8.2 Establish a multi-factor-based task priority evaluation model to dynamically determine the resource allocation priority of each task in the AR system;

[0073] S8.3 Design a resource allocation optimization algorithm based on reinforcement learning to realize the dynamic allocation and scheduling of system resources;

[0074] S8.4 Construct an edge-cloud collaborative computing framework and dynamically decide on the offloading strategy for computing tasks based on task characteristics and network status;

[0075] S8.5 Establish a system performance and energy consumption balance optimization mechanism to minimize energy consumption while meeting the performance requirements of AR applications.

[0076] The present invention is further configured such that: in step S9, system performance evaluation and continuous optimization:

[0077] S9.1 Design a multi-dimensional performance evaluation system that includes technical indicators, user experience indicators, and efficiency indicators;

[0078] S9.2 Build an automated performance testing platform to achieve continuous monitoring and evaluation of various system indicators;

[0079] S9.3. Design a scientific user experience evaluation method that combines subjective evaluation with objective physiological indicators to comprehensively evaluate the system experience;

[0080] S9.4. The Bayesian optimization algorithm is adopted, with the performance evaluation index as the objective function, to dynamically adjust the system parameters; the multi-armed bandit algorithm is designed to explore the optimal configuration strategy in different scenarios; through transfer learning, the optimization strategy learned in a specific scenario is transferred to similar scenarios to accelerate the optimization process.

[0081] S9.5 Establish a systematic system iteration and version control process to ensure the orderly implementation and effect tracking of optimization measures.

[0082] Beneficial effects

[0083] Compared with known public technologies, the technical solution provided by this invention has the following beneficial effects:

[0084] This invention provides personalized guidance for content presentation strategies through dynamic modeling of user cognitive states; it offers a spatial understanding foundation for content placement and interaction through 3D scene semantic understanding and modeling; it enhances content quality and adaptability through augmented reality content generation and optimization; it improves information transmission efficiency through dynamic prediction and guidance of user attention; it achieves natural and efficient human-computer interaction through real-time interactive intent recognition and response; it ensures a balance between content visibility and viewing comfort through dynamic adaptation of ambient lighting and visual comfort; it optimizes system performance and energy consumption through dynamic allocation and optimization of system resources; and it improves the quality of AR content presentation and user experience through system performance evaluation and continuous optimization. Attached Figure Description

[0085] Figure 1 This is a flowchart illustrating the steps of an adaptive augmented reality content presentation method based on artificial intelligence according to the present invention. Detailed Implementation

[0086] To make the objectives, technical solutions, and advantages of the embodiments of the present invention clearer, the technical solutions in the embodiments of the present invention will be clearly and completely described below. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. All other embodiments obtained by those skilled in the art based on the embodiments of the present invention without creative effort are within the scope of protection of the present invention.

[0087] The present invention will be further described below with reference to embodiments.

[0088] Example

[0089] like Figure 1 As shown, the present invention provides a technical solution: an adaptive augmented reality content presentation method based on artificial intelligence, comprising the following steps:

[0090] S1. Multimodal environment data acquisition and preprocessing;

[0091] S2, Dynamic modeling of user cognitive state;

[0092] S3, 3D scene semantic understanding modeling;

[0093] S4. Augmented reality content generation optimization;

[0094] S5, Dynamic Prediction and Guidance of User Attention;

[0095] S6. Real-time interactive intent recognition and response;

[0096] S7. Dynamic adaptation to visual comfort under ambient lighting;

[0097] S8. System resource dynamic allocation optimization;

[0098] S9, System performance evaluation and continuous optimization;

[0099] In step S1, the preprocessing of multimodal environment data acquisition:

[0100] S1.1 Deploy a multimodal sensing network consisting of an RGB-D camera, an inertial measurement unit (IMU), an ambient light sensor, and a depth sensor, and use a timestamp alignment mechanism to achieve spatiotemporal synchronization of heterogeneous data;

[0101] S1.2 To address Gaussian noise, salt-and-pepper noise, and motion blur interference in sensor-acquired data, a cascaded processing framework of adaptive Kalman filtering and bilateral filtering is constructed.

[0102] S1.3. Employ a dimensionality reduction strategy that combines principal component analysis (PCA) with manifold learning to reduce the dimensionality of high-dimensional environmental features;

[0103] S1.4 Construct a spatiotemporal context model based on a Long Short-Term Memory (LSTM) network to perform temporal correlation modeling on continuous frame data;

[0104] S1.5 Establish a multi-index data quality assessment system, and use quantitative indicators such as information entropy, signal-to-noise ratio, and feature integrity to monitor the preprocessing results in real time;

[0105] In step S2, dynamic modeling of user cognitive state:

[0106] S2.1 Deploy wearable physiological sensing devices to simultaneously collect users' electroencephalogram (EEG), eye movement trajectory, skin conductance response (GSR), and heart rate variability (HRV) data;

[0107] S2.2 Perform feature engineering processing on the collected physiological signals to extract quantitative indicators that reflect the user's cognitive state;

[0108] S2.3. A cognitive state classification model is constructed using a deep belief network (DBN) to classify users' cognitive states into five categories: focused, fatigued, confused, distracted, and relaxed.

[0109] S2.4 Construct a dynamic tracking model of cognitive state based on particle filtering to achieve real-time estimation and prediction of the user's cognitive state;

[0110] S2.5 Establish a baseline library of user cognitive features and achieve personalized calibration of the cognitive model through transfer learning;

[0111] In step S2, dynamic modeling of user cognitive state, the following features are used to fuse multimodal physiological signals:

[0112]

[0113] In the formula, K represents the query vector, derived from current cognitive state features; K represents the key vector, derived from historical physiological signal features; V represents the value vector, containing multimodal physiological features. Indicates the dimension of the key vector; The (·) function normalizes the attention score into a probability distribution. This formula calculates the similarity between the query and the key, assigns dynamic weights to the physiological features of different modalities, and achieves effective fusion of multi-source information.

[0114] In step S3, the three-dimensional scene semantic understanding modeling:

[0115] S3.1. A fusion method based on structure of motion reconstruction (SfM) and multi-view stereo matching (MVS) is used to reconstruct the three-dimensional geometry of the scene;

[0116] S3.2 Construct a multi-scale semantic segmentation network based on Transformer to achieve pixel-level classification and recognition of objects in the scene;

[0117] S3.3 Based on the semantic segmentation results, construct a topological relationship graph of the scene to describe the spatial location and functional association between objects;

[0118] S3.4 Design a scene dynamic change detection algorithm based on Siamese network to monitor dynamic events in the environment in real time;

[0119] S3.5 Establish a multi-dimensional scenario understanding confidence assessment system to quantify the reliability of scenario modeling;

[0120] In step S3, the semantic understanding of the 3D scene is used to optimize the semantic segmentation model.

[0121]

[0122] In the formula For cross-entropy loss, Indicates the true label, Indicates the predicted probability; For Dice's loss; The weight coefficient (0.7) balances the contributions of the two losses. This hybrid loss function solves the class imbalance problem in semantic segmentation and improves the accuracy of the segmentation boundary.

[0123] In step S4, during the optimization of augmented reality content generation:

[0124] S4.1 Construct a BERT-based context-aware requirement parsing model to accurately understand users' AR content needs;

[0125] S4.2. Based on the parsed requirements, call the corresponding content generation module to generate multimodal AR content including text, images, and 3D models;

[0126] S4.3 Establish a content optimization framework based on reinforcement learning to dynamically adjust AR content attributes according to scene characteristics and user cognitive state;

[0127] S4.4 Design an AR content spatial layout optimization algorithm based on genetic algorithm to achieve reasonable placement of multiple contents;

[0128] S4.5. Employs physically based rendering (PBR) technology and real-time global illumination algorithms to enhance the realism and immersion of AR content;

[0129] In step S4, augmented reality content generation, the content spatial layout is optimized.

[0130]

[0131] In the formula This represents the visibility score of the i-th AR content; The occlusion overlap between the i-th and j-th content is represented; S represents the total area occupied by all AR content. , , The weighting coefficients are set to 0.5, 0.3, and 0.2 respectively. This objective function comprehensively optimizes content visibility, occlusion, and space occupation to achieve the optimal layout of AR content.

[0132] In step S5, user attention dynamic prediction guidance:

[0133] S5.1 Integrating bottom-up and top-down attention mechanisms to construct a visual attention prediction model that conforms to the characteristics of AR scenarios;

[0134] S5.2 Construct a user attention shift prediction model based on Hidden Markov Model (HMM) to predict the attention trajectory in the next 1-3 seconds;

[0135] S5.3. Generate personalized attention guidance strategies based on attention prediction results;

[0136] S5.4 Establish an attention load monitoring and balancing mechanism to prevent cognitive fatigue caused by information overload;

[0137] S5.5 Build a closed-loop attention feedback system to verify the attention guidance effect through user interaction behavior and continuously optimize the model;

[0138] In step S6, the real-time interactive intent recognition response:

[0139] S6.1 Deploy a multimodal interaction perception system to simultaneously collect user gestures, voice, eye movements, and head posture interaction signals;

[0140] S6.2 Construct a multimodal interaction intent recognition model based on Transformer to achieve real-time classification and prediction of user interaction intent;

[0141] S6.3. Generate an adaptive interaction response strategy based on the identified interaction intent;

[0142] S6.4 Design a multimodal interactive feedback system to provide interactive status feedback through multiple channels including vision, hearing, and touch;

[0143] S6.5 Establish an interactive performance evaluation index system, including response latency, recognition accuracy, operation efficiency and user satisfaction, and continuously optimize and improve the interactive experience;

[0144] In step S7, dynamic adaptation of ambient lighting visual comfort:

[0145] S7.1. By combining ambient light sensors with image analysis, the ambient light characteristics are comprehensively extracted.

[0146] S7.2 Construct a visual comfort assessment model based on physiological indicators and subjective evaluation;

[0147] S7.3. Dynamically adjust the display parameters of AR content based on ambient lighting characteristics and visual comfort model;

[0148] S7.4 Optimize stereoscopic vision parameters to reduce visual fatigue, taking into account the stereoscopic display characteristics of AR devices;

[0149] S7.5 Build a real-time visual fatigue monitoring system to determine the user's visual fatigue state through the fusion of multiple physiological indicators;

[0150] In step S7, the adaptation of ambient lighting and visual comfort is used to predict user visual comfort.

[0151]

[0152] In the formula, C represents the visual comfort score (1-5 points); This represents the k-th visual parameter (brightness, contrast, saturation); For the intercept term, The coefficients are linear. The coefficients of the interaction terms; As the error term, this multivariate nonlinear regression model takes into account the interaction between visual parameters, and can more accurately predict user comfort under different display conditions.

[0153] In step S8, the dynamic allocation and optimization of system resources:

[0154] S8.1 Deploy a system resource monitoring module to collect real-time status parameters of CPU, GPU, memory, storage, and network;

[0155] S8.2 Establish a multi-factor-based task priority evaluation model to dynamically determine the resource allocation priority of each task in the AR system;

[0156] S8.3 Design a resource allocation optimization algorithm based on reinforcement learning to realize the dynamic allocation and scheduling of system resources;

[0157] S8.4 Construct an edge-cloud collaborative computing framework and dynamically decide on the offloading strategy for computing tasks based on task characteristics and network status;

[0158] S8.5 Establish a system performance and energy consumption balance optimization mechanism to minimize energy consumption while meeting the performance requirements of AR applications;

[0159] Step S9, system performance evaluation and continuous optimization:

[0160] S9.1 Design a multi-dimensional performance evaluation system that includes technical indicators, user experience indicators, and efficiency indicators;

[0161] S9.2 Build an automated performance testing platform to achieve continuous monitoring and evaluation of various system indicators;

[0162] S9.3. Design a scientific user experience evaluation method that combines subjective evaluation with objective physiological indicators to comprehensively evaluate the system experience;

[0163] S9.4. The Bayesian optimization algorithm is adopted, with the performance evaluation index as the objective function, to dynamically adjust the system parameters; the multi-armed bandit algorithm is designed to explore the optimal configuration strategy in different scenarios; through transfer learning, the optimization strategy learned in a specific scenario is transferred to similar scenarios to accelerate the optimization process.

[0164] S9.5 Establish a systematic system iteration and version control process to ensure the orderly implementation and effect tracking of optimization measures.

[0165] In this embodiment, dynamic modeling of user cognitive states provides a personalized basis for content presentation strategies; semantic understanding and modeling of 3D scenes provide a spatial understanding foundation for content placement and interaction; augmented reality content generation and optimization improve content quality and adaptability; dynamic prediction and guidance of user attention enhances information transmission efficiency; real-time interactive intent recognition and response achieve natural and efficient human-computer interaction; dynamic adaptation of ambient lighting and visual comfort ensures a balance between content visibility and viewing comfort; dynamic allocation and optimization of system resources optimize system performance and energy consumption; and system performance evaluation and continuous optimization improve the quality of AR content presentation and user experience.

[0166] Working principle

[0167] like Figure 1 As shown, in practical use, the process unfolds from two dimensions: multimodal environmental data acquisition and user cognitive state modeling. Environmental data is collected synchronously through a multi-sensor network including RGB-D cameras and IMUs. After noise is processed by adaptive Kalman filtering and bilateral filtering, PCA and manifold learning are used for dimensionality reduction. A spatiotemporal context model is constructed through an LSTM network to provide high-quality input for scene understanding. Wearable devices are used to collect physiological signals such as EEG and eye movement of users, and 86-dimensional cognitive feature vectors are extracted. A deep belief network is used to classify cognitive states into five categories, such as focus and fatigue. Dynamic tracking is achieved by combining particle filtering. Through multimodal feature fusion and attention mechanism, dynamic weights are assigned to different physiological features to form the basis for personalized content presentation.

[0168] In the scene understanding and content generation stages, a fusion method of SfM and MVS is used to reconstruct the 3D geometric structure. Pixel-level object recognition is achieved through the Transformer semantic segmentation network. The segmentation accuracy is optimized by combining cross-entropy and Dice hybrid loss functions. After parsing user needs based on the BERT model, GPT and StableDiffusion are called to generate multimodal AR content. The spatial layout is optimized through genetic algorithms to comprehensively balance content visibility, occlusion, and space occupation. Hidden Markov models are used to predict user attention trajectories. Multimodal guidance strategies are combined to improve information transmission efficiency. Display parameters are dynamically adjusted based on ambient lighting characteristics and a multivariate nonlinear regression model for visual comfort to ensure viewing comfort.

[0169] System resource management and performance optimization form a closed-loop guarantee mechanism. By monitoring the status of resources such as CPU and GPU in real time, reinforcement learning algorithms are used to dynamically allocate computing tasks. Combined with the edge-cloud collaborative framework, task offloading is achieved. While ensuring a rendering frame rate of 60fps, energy consumption is reduced by more than 30%. The performance evaluation system covers technical, experience and efficiency indicators. Through automated testing platforms and user experience evaluation methods, continuous optimization is carried out to ultimately achieve adaptive improvement in AR content presentation quality and user experience.

[0170] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions will not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention.

Claims

1. An adaptive augmented reality content presentation method based on artificial intelligence, characterized in that, Includes the following steps: S1. Multimodal environment data acquisition and preprocessing; S2, Dynamic modeling of user cognitive state; S3, 3D scene semantic understanding modeling; S4. Augmented reality content generation optimization; S5, Dynamic Prediction and Guidance of User Attention; S6. Real-time interactive intent recognition and response; S7. Dynamic adaptation to visual comfort under ambient lighting; S8. System resource dynamic allocation optimization; S9, System performance evaluation is continuously optimized.

2. The adaptive augmented reality content presentation method based on artificial intelligence according to claim 1, characterized in that: In step S1, the preprocessing of multimodal environment data acquisition: S1.1 Deploy a multimodal sensing network consisting of an RGB-D camera, an inertial measurement unit (IMU), an ambient light sensor, and a depth sensor, and use a timestamp alignment mechanism to achieve spatiotemporal synchronization of heterogeneous data; S1.2 To address Gaussian noise, salt-and-pepper noise, and motion blur interference in sensor-acquired data, a cascaded processing framework of adaptive Kalman filtering and bilateral filtering is constructed. S1.

3. Employ a dimensionality reduction strategy that combines principal component analysis (PCA) with manifold learning to reduce the dimensionality of high-dimensional environmental features; S1.4 Construct a spatiotemporal context model based on a Long Short-Term Memory (LSTM) network to perform temporal correlation modeling on continuous frame data; S1.5 Establish a multi-index data quality assessment system, and use quantitative indicators such as information entropy, signal-to-noise ratio, and feature integrity to monitor the preprocessing results in real time.

3. The adaptive augmented reality content presentation method based on artificial intelligence according to claim 1, characterized in that: In step S2, dynamic modeling of user cognitive state: S2.1 Deploy wearable physiological sensing devices to simultaneously collect users' electroencephalogram (EEG), eye movement trajectory, skin conductance response (GSR), and heart rate variability (HRV) data; S2.2 Perform feature engineering processing on the collected physiological signals to extract quantitative indicators that reflect the user's cognitive state; S2.

3. A cognitive state classification model is constructed using a deep belief network (DBN) to classify users' cognitive states into five categories: focused, fatigued, confused, distracted, and relaxed. S2.4 Construct a dynamic tracking model of cognitive state based on particle filtering to achieve real-time estimation and prediction of the user's cognitive state; S2.5 Establish a baseline library of user cognitive features and achieve personalized calibration of the cognitive model through transfer learning; In step S2, dynamic modeling of user cognitive state, the following features are used to fuse multimodal physiological signals: In the formula, K represents the query vector, derived from current cognitive state features; K represents the key vector, derived from historical physiological signal features; V represents the value vector, containing multimodal physiological features. Indicates the dimension of the key vector; The (·) function normalizes the attention score into a probability distribution. This formula calculates the similarity between the query and the key, assigns dynamic weights to the physiological features of different modalities, and achieves effective fusion of multi-source information.

4. The adaptive augmented reality content presentation method based on artificial intelligence according to claim 1, characterized in that: In step S3, the three-dimensional scene semantic understanding modeling: S3.

1. A fusion method based on structure of motion reconstruction (SfM) and multi-view stereo matching (MVS) is used to reconstruct the three-dimensional geometry of the scene; S3.2 Construct a multi-scale semantic segmentation network based on Transformer to achieve pixel-level classification and recognition of objects in the scene; S3.3 Based on the semantic segmentation results, construct a topological relationship graph of the scene to describe the spatial location and functional association between objects; S3.4 Design a scene dynamic change detection algorithm based on Siamese network to monitor dynamic events in the environment in real time; S3.5 Establish a multi-dimensional scenario understanding confidence assessment system to quantify the reliability of scenario modeling; In step S3, the semantic understanding of the 3D scene is used to optimize the semantic segmentation model. In the formula For cross-entropy loss, Indicates the true label, Indicates the predicted probability; For Dice's loss; The weight coefficient (0.7) balances the contributions of the two losses. This hybrid loss function solves the class imbalance problem in semantic segmentation and improves the accuracy of the segmentation boundary.

5. The adaptive augmented reality content presentation method based on artificial intelligence according to claim 1, characterized in that: In step S4, during the optimization of augmented reality content generation: S4.1 Construct a BERT-based context-aware requirement parsing model to accurately understand users' AR content needs; S4.

2. Based on the parsed requirements, call the corresponding content generation module to generate multimodal AR content including text, images, and 3D models; S4.3 Establish a content optimization framework based on reinforcement learning to dynamically adjust AR content attributes according to scene characteristics and user cognitive state; S4.4 Design an AR content spatial layout optimization algorithm based on genetic algorithm to achieve reasonable placement of multiple contents; S4.

5. Employs physically based rendering (PBR) technology and real-time global illumination algorithms to enhance the realism and immersion of AR content; In step S4, augmented reality content generation, the content spatial layout is optimized. In the formula This represents the visibility score of the i-th AR content; The occlusion overlap between the i-th and j-th content is represented; S represents the total area occupied by all AR content. , , The weighting coefficients are set to 0.5, 0.3, and 0.2 respectively. This objective function comprehensively optimizes content visibility, occlusion, and space occupation to achieve the optimal layout of AR content.

6. The adaptive augmented reality content presentation method based on artificial intelligence according to claim 1, characterized in that, In step S5, user attention dynamic prediction guidance: S5.1 Integrating bottom-up and top-down attention mechanisms to construct a visual attention prediction model that conforms to the characteristics of AR scenarios; S5.2 Construct a user attention shift prediction model based on Hidden Markov Model (HMM) to predict the attention trajectory in the next 1-3 seconds; S5.

3. Generate personalized attention guidance strategies based on attention prediction results; S5.4 Establish an attention load monitoring and balancing mechanism to prevent cognitive fatigue caused by information overload; S5.5 Build a closed-loop attention feedback system to verify the attention guidance effect through user interaction behavior and continuously optimize the model.

7. The adaptive augmented reality content presentation method based on artificial intelligence according to claim 1, characterized in that, In step S6, the real-time interactive intent recognition response: S6.1 Deploy a multimodal interaction perception system to simultaneously collect user gestures, voice, eye movements, and head posture interaction signals; S6.2 Construct a multimodal interaction intent recognition model based on Transformer to achieve real-time classification and prediction of user interaction intent; S6.

3. Generate an adaptive interaction response strategy based on the identified interaction intent; S6.4 Design a multimodal interactive feedback system to provide interactive status feedback through multiple channels including vision, hearing, and touch; S6.5 Establish an interactive performance evaluation index system, including response latency, recognition accuracy, operation efficiency and user satisfaction, and continuously optimize and improve the interactive experience.

8. The adaptive augmented reality content presentation method based on artificial intelligence according to claim 1, characterized in that, In step S7, dynamic adaptation of ambient lighting visual comfort: S7.

1. By combining ambient light sensors with image analysis, the ambient light characteristics are comprehensively extracted. S7.2 Construct a visual comfort assessment model based on physiological indicators and subjective evaluation; S7.

3. Dynamically adjust the display parameters of AR content based on ambient lighting characteristics and visual comfort model; S7.4 Optimize stereoscopic vision parameters to reduce visual fatigue, taking into account the stereoscopic display characteristics of AR devices; S7.5 Build a real-time visual fatigue monitoring system to determine the user's visual fatigue state through the fusion of multiple physiological indicators; In step S7, the adaptation of ambient lighting and visual comfort is used to predict user visual comfort. In the formula, C represents the visual comfort score (1-5 points); This represents the k-th visual parameter (brightness, contrast, saturation); For the intercept term, The coefficients are linear. The coefficients of the interaction terms; As the error term, this multivariate nonlinear regression model takes into account the interaction between visual parameters, and can more accurately predict user comfort under different display conditions.

9. The adaptive augmented reality content presentation method based on artificial intelligence according to claim 1, characterized in that, In step S8, the dynamic allocation and optimization of system resources: S8.1 Deploy a system resource monitoring module to collect real-time status parameters of CPU, GPU, memory, storage, and network; S8.2 Establish a multi-factor-based task priority evaluation model to dynamically determine the resource allocation priority of each task in the AR system; S8.3 Design a resource allocation optimization algorithm based on reinforcement learning to realize the dynamic allocation and scheduling of system resources; S8.4 Construct an edge-cloud collaborative computing framework and dynamically decide on the offloading strategy for computing tasks based on task characteristics and network status; S8.5 Establish a system performance and energy consumption balance optimization mechanism to minimize energy consumption while meeting the performance requirements of AR applications.

10. The adaptive augmented reality content presentation method based on artificial intelligence according to claim 1, characterized in that, Step S9, system performance evaluation and continuous optimization: S9.1 Design a multi-dimensional performance evaluation system that includes technical indicators, user experience indicators, and efficiency indicators; S9.2 Build an automated performance testing platform to achieve continuous monitoring and evaluation of various system indicators; S9.

3. Design a scientific user experience evaluation method that combines subjective evaluation with objective physiological indicators to comprehensively evaluate the system experience; S9.

4. The Bayesian optimization algorithm is adopted, with the performance evaluation index as the objective function, to dynamically adjust the system parameters; the multi-armed bandit algorithm is designed to explore the optimal configuration strategy in different scenarios; through transfer learning, the optimization strategy learned in a specific scenario is transferred to similar scenarios to accelerate the optimization process. S9.5 Establish a systematic system iteration and version control process to ensure the orderly implementation and effect tracking of optimization measures.