Visual perception interaction system and method based on AI scene understanding

By adopting AI-driven scene understanding technology in augmented reality devices, real-time analysis and adjustment of display parameters are solved, and the user's visual perception quality and interactive experience are significantly improved.

CN120215708APending Publication Date: 2025-06-27SOUTHEAST UNIV +1
View PDF 0 Cites 2 Cited by

Patent Information

Application Number
CN202510281390.2
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2025-03-11
Publication Date
2025-06-27

AI Technical Summary

Technical Problem

Existing augmented reality devices have poor display effects in complex environments due to changes in lighting and high scene complexity, which makes users prone to visual fatigue and lack quantitative analysis and closed-loop feedback mechanisms for the overall scene complexity.

Method used

Using AI-based scenario understanding technology, the environment images are captured in real time through image sensors, image features are extracted using convolutional neural networks, entropy values ​​are calculated to quantify scene complexity, and display brightness and contrast are dynamically adjusted, and display parameters are optimized in combination with user feedback.

Benefits of technology

It significantly improves the user's visual perception quality, improves the display adaptability of AR devices in complex environments, reduces user visual fatigue, and achieves a more natural and smooth human-computer interaction experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120215708A_ABST
    Figure CN120215708A_ABST
Patent Text Reader

Abstract

The invention discloses a visual perception interaction system and method based on AI scene understanding, and belongs to the technical field of augmented reality. The system captures a user environment in real time through an image sensor, performs feature extraction and target detection in combination with a convolutional neural network (CNN), quantifies scene complexity, and dynamically adjusts brightness, contrast and display parameters of AR equipment. The core module comprises complexity analysis, adaptive brightness / contrast adjustment, user feedback optimization and global scoring. By calculating the image entropy, the standard deviation and the user score, the system optimizes the display effect in real time, and the visual definition and comfort of the user are improved. The method solves the problem of poor AR display adaptability in a complex environment, remarkably enhances the human eye perception quality, and is suitable for the fields of virtual reality, automatic driving and the like.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] The present invention relates to a human-computer interaction optimization technology. By using existing deep learning and AI technologies, it can effectively improve the perception experience of see-through AR devices and achieve the purpose of improving the human eye perception quality. Background Art

[0002] With the emergence of concepts such as the metaverse and digital twin, adaptive real-time human eye perception feedback has become increasingly indispensable in the process of improving the display quality of augmented reality devices. Especially in the applications of augmented reality (AR) and virtual reality (VR) technologies, an efficient visual perception interaction system can significantly enhance the user experience and system performance. Most current visual perception technologies focus on the recognition and processing of single targets. Although these technologies perform well in static scenarios or simple tasks, they face many challenges in complex and changing environments.

[0003] Existing augmented reality devices often directly display virtual digital signals without special design for complex see-through scenarios and dynamically changing external environments. By using deep learning technology, the ability of image processing and recognition can be significantly improved, and a large amount of visual data can be processed in real time, enabling precise analysis of the changing scenarios. During the use of augmented reality devices, the user's visual perception is affected not only by the real environment but also by the overlay of virtual content. How to achieve effective scene understanding and dynamic adjustment in such a complex multi-modal environment is one of the main challenges faced by current technologies.

[0004] In recent years, the emergence of deep learning, especially convolutional neural networks (CNNs), has greatly improved the ability of image processing and visual perception. Through an end-to-end learning model, a computer can automatically extract image features from a large amount of data to achieve more accurate object recognition, semantic segmentation, and object detection. However, with the increase in scene complexity, traditional vision systems based on object recognition can no longer meet the requirements of scene understanding. For example, in an autonomous driving scenario, the system not only needs to recognize vehicles and pedestrians but also understand their spatial layout, motion intention, and potential dangers in the scene. AI scene understanding technology combines various advanced technologies such as deep learning, spatio-temporal information fusion, and image semantic analysis, aiming to enable a machine to not only recognize individual targets in a scene but also deeply analyze the layout, dynamic changes, and interaction relationships between people and objects in the whole scene. Through this advanced scene understanding, the system can make more intelligent decisions for complex environments, improving overall safety and interactivity.

[0005] In summary, existing AR devices often have poor display effects due to light changes and high scene complexity in complex environments, and users are prone to visual fatigue. Traditional methods rely on static parameter adjustment and cannot dynamically adapt to changing scenarios. Although deep learning technology has been used for object recognition, there is a lack of quantitative analysis of the overall scene complexity and a closed-loop feedback mechanism. The present invention fills the technical gap in dynamic display optimization through AI-driven scene understanding and adaptive adjustment. Summary of the Invention

[0006] The object of the present invention is to propose an AI-based scene understanding technology for augmented reality devices, so as to improve the visual perception quality of users, form an intelligent interaction system, provide efficient multi-modal perception and scene analysis, and meet the requirements for intelligent augmented reality human-computer interaction systems in fields such as autonomous driving, smart home, and virtual reality.

[0007] To achieve the above object, the present invention adopts the following technical solutions:

[0008] A visual perception interaction system based on AI scene understanding, comprising:

[0009] An image sensor: used to capture the user's environmental image in real time;

[0010] A complexity analysis module: extracting image features through a convolutional neural network and calculating the entropy value to quantify the scene complexity;

[0011] A human-computer interaction database: used to store display parameters corresponding to different levels of scene complexity, including display brightness and contrast;

[0012] An adaptive adjustment module: retrieving the corresponding display parameters from the human-computer interaction database according to the scene complexity calculated by the complexity analysis module, and dynamically adjusting the display brightness and contrast in combination with the ambient light intensity and the image standard deviation;

[0013] A user feedback optimization module: optimizing the display parameters based on the clarity score of the display image by the user after adjusting the display brightness and contrast, and updating the human-computer interaction database.

[0014] Preferably, the image entropy value calculation formula of the complexity analysis module is:

[0015]

[0016] where p(x i ) represents the probability distribution of the i-th pixel x i in the input image, and n represents the number of pixels.

[0017] Preferably, the brightness adjustment formula of the adaptive adjustment module is:

[0018]

[0019] The contrast adjustment formula is:

[0020]

[0021] Wherein, I current is the external ambient light intensity, I reference is the reference light intensity, L base is the base brightness, L new is the display brightness of the virtual image inside the AR display device, C new is the new contrast value, std(I current ) is the standard deviation of the current image, std(I reference ) is the standard deviation of the reference image, C base is the reference contrast.

[0022] Preferably, the user feedback optimization module adjusts the parameters according to the user clarity score S clarity The formula is:

[0023] P new = P base ·(1 + λ·(S clarity - S target ))

[0024] Where P new is the adjusted parameter, S clarity is the image clarity feedback by the user, S target is the target clarity value, λ is the feedback sensitivity coefficient, P base is the base brightness L base or the reference contrast C base .

[0025] Preferably, it further includes a global scoring module for generating a system performance score by comprehensively considering clarity, response time, and user satisfaction.

[0026] Preferably, the comprehensive scoring formula of the global scoring module is:

[0027] S total = αS clarity + βS response + γS satisfaction

[0028] Where S clarity , S response and S satisfaction are the scores of clarity, response time, and satisfaction respectively, and α, β, γ are the weight coefficients of each item.

[0029] The present invention also provides a visual perception interaction method based on AI scene understanding for the above system, including the following steps:

[0030] The user wears an AR headset, and the image sensor externally placed in front of the headset captures the environmental image in the user's field of view in real time;

[0031] Use a convolutional neural network (CNN) to perform feature extraction and object detection on the input image to generate a feature map;

[0032] Calculate the image entropy value based on the probability distribution of the feature map to quantify the scene complexity;

[0033] Dynamically adjust the brightness of the display device according to the ratio of the current ambient light intensity to the reference value;

[0034] Dynamically adjust the contrast of the display device according to the ratio of the standard deviation of the input image to the reference value;

[0035] Receive the feedback score of the user on the display clarity, and optimize the display parameters in combination with the target clarity value.

[0036] Preferably, the feature extraction step includes:

[0037] Use a convolutional neural network (CNN) to perform a convolution operation on the input image to extract edge and texture features;

[0038] The formula is as follows:

[0039]

[0040] Here, K is the convolution kernel for extracting the feature map of the image I input of.

[0041] Calculate the intersection over union (IoU) of the object detection box and the ground truth box to evaluate the object overlapping area;

[0042]

[0043] Adjust the feature weights based on the IoU result to generate a multi-dimensional feature vector.

[0044] Preferably, it further includes a scene classification and parameter archiving step:

[0045] Divide the scene into low, medium, and high complexity levels based on the entropy value H;

[0046] Store the optimized parameters for different levels of scenes to form a historical database;

[0047] Preferentially call the historical parameters in similar scenes to accelerate the adaptive adjustment process.

[0048] Preferably, it further includes the steps of: generating a global performance score by synthesizing clarity, response time, and user satisfaction, and driving the system to adaptively iterate and optimize.

[0049] Preferably, the system also classifies and summarizes different scenarios for future reference and optimization of display settings. Integrating the requirements in different scenarios, the system optimizes the display device parameters

[0050]

[0051] Here, P i are the display parameters (brightness and contrast) in each scenario, and w i is the weight of the scenario.

[0052] After the above processing and optimization, the AR headset applies the adjusted display parameters to its display screen, thereby generating an optimized fused image I of the virtual and real environments output . The image seen by the user through the AR headset is the result of real-time analysis and adaptive adjustment, significantly improving the quality of visual perception and the interaction experience.

[0053] Beneficial effects: The present invention captures the user's environment in real time through an image sensor, combines a convolutional neural network (CNN) for feature extraction and target detection, quantifies the scene complexity, and dynamically adjusts the brightness, contrast, and display parameters of the AR device. The core modules include complexity analysis, adaptive brightness / contrast adjustment, user feedback optimization, and global scoring. By calculating the image entropy value, standard deviation, and user score, the system optimizes the display effect in real time, improving the user's visual clarity and comfort. The present invention solves the problem of poor AR display adaptability in complex environments, significantly enhancing the quality of human eye perception, and is applicable to fields such as virtual reality and autonomous driving.

[0054] The present invention integrates visual perception, spatio-temporal information processing, and deep learning technologies. The system can not only classify and summarize complex scenarios, but also deeply understand the human eye's response to complex scenarios, and adjust and enhance the display parameters of the display device in real time to achieve a more natural and smooth interaction experience. BRIEF DESCRIPTION OF THE DRAWINGS

[0055] Figure 1 is a system architecture diagram.

[0056] Figure 2 is a system interface diagram. DETAILED DESCRIPTION OF THE INVENTION

[0057] The following further explains and illustrates the present invention in conjunction with the drawings and embodiments.

[0058] As Figure 1 shown, the implementation steps of the present invention include:

[0059] 1. The user wears an AR headset and can enter this system from the virtual display interface. Figure 2 This is the home page of the system diagram. At the beginning, the system needs to capture the environmental images in the user's field of view in real time through the image sensor on the outer side of the front of the headset (I input ). After the external scene is obtained, it enters the complexity analysis module, records the scene name, assigns a scene number, records this image as the scene content, and the user can select scene tasks for label classification. A series of preliminary assignments are made to the brightness, color appearance, contrast, and dynamic range of the captured scene. The preliminary assignment calls historical parameters, assuming that this image has a resolution of 1920x1080 pixels and a color depth of 24 bits.

[0060] 2. Normalize the input RGB image and adjust the size to meet the input requirements of the subsequent CNN model. Perform preliminary filtering on each channel (R, G, B) to remove noise interference.

[0061] 3. Use a Convolutional Neural Network (CNN) for feature extraction. This process sharpens the image through a 3x3 convolutional kernel (K), combines each pixel with adjacent pixels, and prepares for subsequent object detection. The convolution operation follows the following formula:

[0062]

[0063] This process can effectively extract edges and other important features in the image. Through multi-layer convolution stacking (such as VGG or ResNet structures), a high-dimensional feature map (Feature Map) is generated.

[0064] 4. Based on the feature map, use an object detection algorithm (such as YOLO or Faster R-CNN) to generate prediction boxes (PD). Compare the intersection over union (IoU) of the prediction boxes with the preset ground truth boxes (GT):

[0065]

[0066] Example: GT = [100, 150, 200, 250], PD = [120, 160, 220, 270].

[0067] Upper left corner: ((120, 160)), lower right corner: ((200, 250)).

[0068] Calculate the overlapping area: Area of Overlap = 80 * 90 = 7200.

[0069] Calculate the total area: Area of Union = 10000 + 10000 - 7200 = 16800.

[0070] Calculate IoU: IoU = 7200 / 16800 = 0.428.

[0071] 5. Dynamically adjust the weights of each region in the feature map according to the IoU result. Regions with high IoU (significant target overlap) are given higher weights, and regions with low IoU (background or interference) have their weights reduced.

[0072] 6. Calculate the entropy value based on the probability distribution of the feature map to quantify the scene complexity:

[0073]

[0074] Assume the feature distribution is: p(x i ) = [0.1, 0.2, 0.25, 0.2, 0.25]. Calculate the scene complexity through the following Python code:

[0075]

[0076] Output result: Scene complexity H = 2.153. The information entropy H provides a basis for subsequent brightness and contrast adjustments.

[0077] 7. Divide the scene level according to the entropy value H:

[0078] Low complexity: H < 1.5

[0079] Medium complexity: 1.5 ≤ H ≤ 2.5

[0080] High complexity: H > 2.5

[0081] Store the optimization parameters (brightness, contrast) of the current scene in the historical human-computer interaction database for quick call in similar scenes.

[0082] 8. Pass the scene complexity level (such as "medium complexity") and the corresponding parameters to the adaptive adjustment module to trigger the dynamic optimization of brightness and contrast.

[0083] The adaptive adjustment module, based on the scene complexity calculated by the complexity analysis module, which is medium complexity in this example, retrieves the corresponding display parameters from the human-computer interaction database as reference values and optimizes the brightness and contrast based on the lighting environment of the current image. Assume the base brightness L base = 200, the current light intensity I current = 450, and the reference light intensity I reference = 300. Calculate the new brightness value through the formula:

[0084]

[0085] The contrast adjustment example is based on the following formula. If the reference contrast is C base= 1.2, the current standard deviation is 20, and the reference standard deviation is 15. Then the updated contrast ratio is:

[0086]

[0087] 9. The user feedback optimization module optimizes the display parameters and updates the human-computer interaction database based on the clarity score of the displayed image given by the user after adjusting the display brightness and contrast ratio. The formula is:

[0088]

[0089] where P new is the adjusted parameter, S clarity = 8 is the image clarity feedback by the user, S target = 7 is the target clarity value, λ = 0.1 is the feedback sensitivity coefficient, i = 1 or 2, is the reference brightness L base = 200, is the reference contrast ratio C base = 1.2. The calculation gives

[0090] 10. The user can give a real-time score to the optimized virtual display information to evaluate the system performance based on personal subjective perception. The user evaluates and scores from three aspects: clarity, response time, and user satisfaction. The larger the score, the better the virtual signal perception quality brought by the display device to the human eye. Suppose these scores are S clarity = 8, S response = 9, S satisfaction = 7, and the weights are set as follows: α = 0.5, β = 0.3, γ = 0.2. Calculate the comprehensive score S total = αS clarity + βS response + γS satisfaction = 0.5·8 + 0.3·9 + 0.2·7 = 4 + 2.7 + 1.4 = 8.1

[0091] 11. Set weights for different scenarios according to the scene complexity calculated by deep learning to optimize and enhance the virtual digital signal display parameters of the display device. Suppose the scene weights are w1 = 0.5, w2 = 0.3. Using the optimization formula, calculate the adjusted display parameter P opt :

[0092]

Claims

1. A visual perception interaction system based on AI scene understanding, characterized in that: include: Image sensor: used to capture images of the user's environment in real time; Complexity analysis module: extracts image features through convolutional neural networks and calculates entropy values ​​to quantify scene complexity; Human-computer interaction database: used to store display parameters corresponding to different levels of scene complexity, including display brightness and contrast; Adaptive adjustment module: retrieves the corresponding display parameters from the human-computer interaction database according to the scene complexity calculated by the complexity analysis module, and dynamically adjusts the display brightness and contrast based on the ambient light intensity and image standard deviation; User feedback optimization module: Based on the user's rating of the clarity of the displayed image after adjusting the display brightness and contrast, the display parameters are optimized and the human-computer interaction database is updated.

2. The system according to claim 1, characterized in that The image entropy value calculation formula of the complexity analysis module is: Among them, p(x i ) represents the i-th pixel x in the input image i The probability distribution of n is the number of pixels.

3. The system according to claim 1, characterized in that The brightness adjustment formula of the adaptive adjustment module is: The contrast adjustment formula is: Among them, I current is the external ambient light intensity, I reference is the reference light intensity, L base is the basic brightness, L new is the display brightness of the virtual image inside the AR display device, C new is the new contrast value, std(I current ) is the standard deviation of the current image, std(I reference ) is the standard deviation of the reference image, C base is the base contrast.

4. The system according to claim 1, characterized in that The user feedback optimization module is based on the user clarity score S clarity Adjust the parameters, the formula is: Where P new is the adjusted parameter, S clarity is the image clarity reported by the user, S target is the target clarity value, λ is the feedback sensitivity coefficient, i=1 or 2, The reference brightness L base , The reference contrast C base .

5. The system according to claim 1, characterized in that It also includes a global scoring module that integrates clarity, response time and user satisfaction to generate a system performance score.

6. The system according to claim 5, characterized in that The comprehensive scoring formula of the global scoring module is: S total =αS clarity +βS response +γS satisfaction Where S clarity , S response and S satisfaction are the scores of clarity, response time and satisfaction respectively, and α, β, and γ are the weight coefficients of each item.

7. A visual perception interaction method based on AI scene understanding based on the system of any one of claims 1 to 4, characterized in that: The following steps are involved: An input image of the user's environment is captured in real time through an image sensor; Using a convolutional neural network (CNN) to perform feature extraction and target detection on the input image to generate a feature map; Calculate the image entropy value based on the probability distribution of the feature map to quantify the scene complexity; Dynamically adjust the brightness of the display device according to the ratio of the current ambient light intensity to the reference value; Dynamically adjust the contrast of the display device according to the ratio of the standard deviation of the input image to the reference value; Receive user feedback on display clarity and optimize display parameters based on target clarity values.

8. The method according to claim 7, characterized in that The feature extraction step comprises: Use a 3×3 convolution kernel to perform a convolution operation on the input image to extract edge and texture features; Calculate the intersection over union (IoU) between the target detection box and the true box to evaluate the target overlap area; The feature weights are adjusted based on the IoU results to generate a multi-dimensional feature vector.

9. The method according to claim 7, characterized in that: It also includes scene classification and parameter archiving steps: Based on the entropy value H, the scenes are divided into low, medium and high complexity levels; Store the optimized parameters of different levels of scenes to form a historical human-computer interaction database; Prioritize calling historical parameters in similar scenarios to speed up the adaptive adjustment process.

10. The method according to claim 7, characterized in that The steps are also included: generating a global performance score based on comprehensive clarity, response time and user satisfaction, and driving adaptive iterative optimization of the system.

Citation Information

Cited By

  • Self-adaptive display buffer system and method for enhancing visual comfort and medium

    CN120491920A

  • Display equipment brightness adaptive control method based on multi-mode environment perception and deep reinforcement learning

    CN121034204A