A Brightness Adaptive Control Method for Display Devices Based on Multimodal Environment Perception and Deep Reinforcement Learning
By using multimodal environmental perception and deep reinforcement learning, the system collects and analyzes environmental and interactive data of display devices in real time, identifies scenes, and iteratively updates the brightness strategy. This solves the problem of inaccurate brightness control in existing technologies, and achieves intelligent and precise adaptive brightness control, thereby improving user experience and device performance.
Patent Information
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- SHENZHEN EASYQUICK TECH CO LTD
- Filing Date
- 2025-09-11
- Publication Date
- 2026-05-26
AI Technical Summary
Existing display device brightness control technologies cannot generate scene feature identification vectors that accurately reflect the current environment and user status. They also lack an effective mechanism to iteratively update the brightness decision strategy based on the comprehensive reward value, resulting in insufficient intelligence and precision in brightness control, which fails to meet the diverse needs of users in different scenarios.
Based on multimodal environment perception and deep reinforcement learning, multidimensional environment perception data and multidimensional interaction data are collected in real time to generate scene feature identification vectors. By performing cluster analysis on the historical brightness adjustment records of display devices, scene feature threshold vectors are obtained to identify the current usage scenario. The initial strategy is retrieved from the brightness decision strategy library. Combined with device and user feedback data, a comprehensive reward formula is constructed to iteratively update the brightness decision strategy.
It enables precise brightness adjustment of display devices in different scenarios, improves the user's visual experience, reduces device power consumption, takes into account the user's eye health, and dynamically optimizes the brightness adjustment strategy to meet the needs of visualization effects, power consumption and user feedback.
Smart Images

Figure CN121034204B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of display control technology, and in particular to a display device brightness adaptive control method based on multimodal environment perception and deep reinforcement learning. Background Technology
[0002] In today's digital age, display devices are widely used in various aspects of people's lives and work, such as mobile phones, computers, televisions, and various smart terminals. As people's demands for visual experience continue to increase, brightness control of display devices has become one of the key factors affecting user experience. Different environmental conditions and user needs place different demands on the brightness of display devices. For example, in well-lit outdoor environments, higher brightness is needed to ensure that the screen content is clearly visible; while at night or in low-light environments, lower brightness can meet viewing needs while protecting the eyes and saving power. A display device brightness adaptive control method based on multimodal environmental perception and deep reinforcement learning is of great significance. In the future, with the continued popularization of smart devices and people's increasing pursuit of a high-quality life, this technology is expected to be applied to more types of display devices, driving the display industry towards a more intelligent and user-friendly direction.
[0003] However, existing display device brightness control technologies cannot generate scene feature vectors that accurately reflect the current environment and user status. In scene recognition, there is a lack of effective methods for clustering and analyzing historical brightness adjustment records of display devices, making it difficult to obtain accurate scene feature threshold vectors for each usage scenario, thus hindering precise identification of the current usage scenario. Furthermore, existing technologies lack an effective mechanism for iteratively updating brightness decision strategies based on comprehensive reward values, resulting in insufficiently intelligent and precise brightness control, failing to meet the diverse brightness needs of users in different scenarios.
[0004] Therefore, this invention proposes a display device brightness adaptive control method based on multimodal environment perception and deep reinforcement learning. Summary of the Invention
[0005] This invention provides a display device brightness adaptive control method based on multimodal environmental perception and deep reinforcement learning. Multimodal environmental perception technology can collect multi-dimensional environmental perception data in real time, including light intensity, color temperature, and ambient color, as well as multi-dimensional interaction data such as user operating habits and usage time, to comprehensively understand the current environment and user status. Deep reinforcement learning, based on this rich data, can continuously learn and optimize to formulate more intelligent and precise brightness control strategies. This method not only significantly improves the user's visual experience in different scenarios but also effectively reduces device power consumption, achieving energy conservation and environmental protection, while also considering the user's eye health.
[0006] This invention provides a display device brightness adaptive control method based on multimodal environment perception and deep reinforcement learning, comprising:
[0007] It collects multi-dimensional environmental perception data and multi-dimensional interaction data in real time under the current environment, and generates scene feature identification vectors under the current environment based on the multi-dimensional environmental perception data and multi-dimensional interaction data.
[0008] By performing cluster analysis on the historical brightness adjustment records of the display device, the scene feature threshold vector of each usage scenario is obtained, and the scene feature identifier vector is matched with the scene feature threshold vectors of all usage scenarios to identify the current usage scenario;
[0009] Retrieve the initial brightness decision strategy for the current usage scenario from the brightness decision strategy library;
[0010] Control the display device to execute the initial brightness decision strategy, and collect device feedback data and user feedback data in real time after the execution of the initial brightness decision strategy;
[0011] Obtain the multi-objective standard differentiation weights of multiple control objectives in the current usage scenario, and construct the comprehensive reward formula R based on the multi-objective standard differentiation weights of multiple control objectives in the current usage scenario, the visualization reward formula R1, the power consumption reward formula R2, the eye protection reward formula R3, and the user feedback reward formula R4.
[0012] The comprehensive reward value of the initial brightness decision strategy is calculated based on device feedback data, user feedback data, and the comprehensive reward formula R.
[0013] Through a priority experience replay mechanism, the current brightness decision strategy is iteratively updated based on the comprehensive reward value of the initial brightness decision strategy. At the same time, the scene benchmark brightness / or multi-objective standard differential weights of the current use scenario are corrected based on the iterated brightness decision strategy to obtain the brightness adaptive control result driven by deep reinforcement learning.
[0014] Preferably, multi-dimensional environmental perception data and multi-dimensional interaction data of the current environment are collected in real time, including:
[0015] The ambient light intensity is collected in real time within a preset visible light band in the current environment. At the same time, the proportion of harmful blue light bands within the preset band is collected as ambient light spectral distribution data, and ambient temperature data is also collected.
[0016] Real-time data collection of user interaction distance distribution and user eye gaze duration;
[0017] Generate scene feature identification vectors for the current environment based on multi-dimensional environmental perception data and multi-dimensional interaction data;
[0018] The multidimensional environmental perception data includes ambient light intensity, ambient light spectral distribution data, and ambient temperature data.
[0019] Multidimensional interactive data includes data on the distribution of user interaction distance with display devices and data on user eye gaze duration.
[0020] Preferably, by performing cluster analysis on the historical brightness adjustment records of the display device, a scene feature threshold vector for each usage scenario is obtained, including:
[0021] The original feature sets and feature sequences of multiple historical moments are extracted from the historical brightness adjustment records of the display device. A standard feature matrix is generated based on the original feature sets of all historical moments. The feature correlation between different dimensions is obtained by calculating the Pearson correlation coefficient between feature sequences of different dimensions. A feature correlation matrix is constructed based on the feature correlation between different dimensions.
[0022] The standard feature matrix is reduced in dimensionality based on a preset dimensionality reduction algorithm and feature correlation matrix to obtain a three-dimensional feature matrix.
[0023] The high-dimensional spatial similarity matrix is calculated based on the standard feature matrix, and the low-dimensional spatial similarity matrix is calculated based on the three-dimensional feature matrix.
[0024] Based on gradient descent minimization of the loss function, the high-dimensional similarity matrix, and the low-dimensional similarity matrix, the three-dimensional feature matrix is optimized to obtain the three-dimensional optimized feature matrix.
[0025] The spatial normalization of the 3D optimized feature matrix is performed to obtain the coordinate matrix to be clustered.
[0026] The coordinate matrix to be clustered is divided into clusters to obtain the final sample coordinate clusters corresponding to each use case.
[0027] High-dimensional feature sets are extracted from the final sample coordinate clusters of each use case, and support vector machines are used to construct boundary models of each dimension. Based on the boundary models of each dimension, the initial threshold vectors of the corresponding use cases are generated.
[0028] Construct the convex hull of the final sample coordinate cluster in three-dimensional space for each use scenario, determine the vertex set of the convex hull, and reverse map the vertex set to a high-dimensional feature space to obtain high-dimensional vertices. Generate the corrected threshold vector for the corresponding use scenario based on the extreme values of each dimension of the high-dimensional vertices.
[0029] The initial threshold vector and the modified threshold vector for each use scenario are weighted to obtain the scene feature threshold vector for each use scenario.
[0030] Preferably, the coordinate matrix to be clustered is divided into clusters to obtain the final sample coordinate clusters corresponding to each use case, including:
[0031] Based on the coordinate matrix to be clustered, the coordinates of all samples are determined, and the local density and distance to high-density points of each sample coordinate are calculated. Then, the cluster center is selected based on the local density and distance to high-density points of all sample coordinates.
[0032] All sample coordinates are arbitrarily divided based on cluster centers to obtain multiple sample coordinate clusters. The feature covariance matrix of each sample coordinate cluster is calculated, and the Mahalanobis distance of each sample coordinate cluster is calculated based on the feature covariance matrix. It is determined whether the sum of the Mahalanobis distances of all the currently obtained sample coordinate clusters is less than a preset threshold. If so, the final sample coordinate cluster is obtained. Otherwise, all sample coordinates are arbitrarily re-divided until the sum of the Mahalanobis distances of the latest obtained sample coordinate clusters is less than the preset threshold. Then, the latest obtained sample coordinate clusters are taken as the final sample coordinate clusters.
[0033] Each final sample coordinate cluster is matched with the usage scenario type to obtain the final sample coordinate cluster corresponding to each usage scenario.
[0034] Preferably, the scene feature identifier vector is matched with the scene feature threshold vectors of all usage scenarios to identify the current usage scenario, including:
[0035] The ratio of the number of elements in the scene feature identifier vector that satisfy the threshold condition of the corresponding dimension in the scene feature threshold vector of a single use scenario to the total number of elements is taken as the matching degree between the scene feature identifier vector and the scene feature threshold vector of the corresponding use scenario.
[0036] Determine if there exists a maximum matching degree between the scene feature threshold vector and the scene feature identifier vector for multiple use scenarios. If so, then the use scenario for which the element value of the maximum weight dimension in the scene feature identifier vector meets the threshold condition of the corresponding dimension in the scene feature threshold vector is taken as the current use scenario.
[0037] Preferably, real-time acquisition of device feedback data and user feedback data after executing the initial brightness decision strategy includes:
[0038] The device collects real-time data on the contrast ratio of the displayed content, the brightness output power, and the proportion of harmful blue light bands within the preset wavelength range after the initial brightness decision strategy is implemented. At the same time, it collects real-time data on the brightness value manually adjusted by the user and the duration of the user's eye gaze as user feedback data.
[0039] Preferably, based on the differentiated weights of multiple control objectives in the current usage scenario, the visualization reward formula R1, the power consumption reward formula R2, the eye protection reward formula R3, and the user feedback reward formula R4, a comprehensive reward formula R is constructed, including:
[0040]
[0041] R4 = α × (1 - Standard value of user accommodation deviation rate) + β × Standard value of user gaze stability
[0042] +γ×User explicit rating standard value
[0043] The standard differential weights of multiple control objectives in the current usage scenario are modified and the boundary constraints are applied to obtain the modified differential weights of multiple control objectives in the current usage scenario.
[0044] Based on the multi-objective correction and differentiation weights of multiple control objectives in the current usage scenario, the visualization reward formula R1, the power consumption reward formula R2, the eye protection reward formula R3, and the user feedback reward formula R4, a comprehensive reward formula R is constructed:
[0045] R=ω1×R1+ω2×R2+ω3×R3+ω4×R4
[0046] In the formula, ω1, ω2, ω3, and ω4 are the multi-objective standard differentiation weights of multiple control objectives in the current usage scenario.
[0047] Preferably, the multi-objective standard differentiation weights of multiple control objectives in the current usage scenario are corrected and boundary constraints are applied to obtain the multi-objective corrected differentiation weights of multiple control objectives in the current usage scenario, including:
[0048] Based on the key dimension elements in the scene feature identifier vector, calculate the weight correction value of each weight in the multi-objective standard differential weight of multiple control objectives in the current usage scenario;
[0049] The correction values of each weight in the multi-objective standard differential weight of multiple control objectives under the current usage scenario are normalized to obtain the multi-objective corrected differential weight of multiple control objectives under the current usage scenario.
[0050] Preferably, the comprehensive reward value of the initial brightness decision strategy is calculated based on device feedback data, user feedback data, and the comprehensive reward formula R, including:
[0051] Based on user feedback data, including manually adjusted brightness values, eye fixation duration, and user ratings, we analyzed the standard values for user accommodation deviation rate, user fixation stability, and user display rating.
[0052] The display content contrast, brightness output power, proportion of harmful blue light bands in the preset band, standard value of user adjustment deviation rate, standard value of user gaze stability, and standard value of user display score in the device feedback data are substituted into the comprehensive reward formula R to obtain the comprehensive reward value of the initial brightness decision strategy.
[0053] Preferably, based on user feedback data including manually adjusted brightness values, user eye fixation duration, and user ratings, standard values for user accommodation deviation rate, user gaze stability, and user display rating are analyzed, including:
[0054] The ratio of the difference between the brightness value manually adjusted by the user and the brightness value recommended by the system to the brightness value recommended by the system is taken as the user adjustment deviation rate. The user adjustment deviation rate is then standardized to obtain the standard value of the user adjustment deviation rate.
[0055] Based on user eye fixation duration data from user feedback, the amplitude of user fixation point fluctuation was determined, and the stability of user fixation was determined based on the amplitude of fixation point fluctuation.
[0056]
[0057] The user gaze stability is standardized to obtain a standard value for user gaze stability;
[0058] Standardize the user ratings in the user feedback data to obtain the standard values for user display ratings.
[0059] The beneficial effects of this invention compared to existing technologies are as follows: Real-time acquisition of multi-dimensional environmental perception data and multi-dimensional interaction data generates scene feature identification vectors, comprehensively and accurately depicting the current environmental characteristics and providing rich information for subsequent scene recognition. By clustering analysis of historical brightness adjustment records of display devices, feature threshold vectors for each usage scenario are obtained and matched with identification vectors, enabling accurate identification of the current usage scenario and laying the foundation for brightness adjustment strategy selection. Retrieving initial strategies from a brightness decision strategy library ensures the targeted nature of the adjustment strategy. Real-time acquisition of device and user feedback data after strategy execution, combined with differentiated weights of multiple control objectives and various reward formulas, constructs a comprehensive reward formula and calculates the comprehensive reward value, providing a quantitative basis for strategy evaluation. Utilizing a priority experience replay mechanism, the brightness decision strategy is iteratively updated based on the reward value, while simultaneously correcting the scene baseline brightness or multi-objective weights, achieving dynamic optimization of the brightness adjustment strategy. This allows display devices to adaptively adjust brightness based on deep reinforcement learning, taking into account multiple needs such as visualization effects, power consumption, eye protection, and user feedback, improving user experience and optimizing device performance.
[0060] Other features and advantages of the invention will be set forth in the description which follows, and will be apparent in part from the description, or may be learned by practicing the invention. The objects and other advantages of the invention may be realized and obtained by means of the structures particularly pointed out in this application.
[0061] The technical solution of the present invention will be further described in detail below with reference to the accompanying drawings and embodiments. Attached Figure Description
[0062] The accompanying drawings are provided to further illustrate the invention and form part of the specification. They are used in conjunction with embodiments of the invention to explain the invention and do not constitute a limitation thereof. In the drawings:
[0063] Figure 1 This is a flowchart of the display device brightness adaptive control method based on multimodal environment perception and deep reinforcement learning in an embodiment of the present invention;
[0064] Figure 2 This is a flowchart illustrating the acquisition process of multidimensional environmental perception data and multidimensional interactive data in an embodiment of the present invention.
[0065] Figure 3 This is a flowchart illustrating the collection process of device feedback data and user feedback data in an embodiment of the present invention. Detailed Implementation
[0066] The preferred embodiments of the present invention will be described below with reference to the accompanying drawings. It should be understood that the preferred embodiments described herein are for illustration and explanation only and are not intended to limit the present invention.
[0067] like Figure 1 As shown, this invention provides a display device brightness adaptive control method based on multimodal environment perception and deep reinforcement learning, comprising:
[0068] It collects multi-dimensional environmental perception data and multi-dimensional interaction data in real time under the current environment, and generates scene feature identification vectors under the current environment based on the multi-dimensional environmental perception data and multi-dimensional interaction data.
[0069] By performing cluster analysis on the historical brightness adjustment records of the display device, the scene feature threshold vector of each usage scenario is obtained, and the scene feature identifier vector is matched with the scene feature threshold vectors of all usage scenarios to identify the current usage scenario;
[0070] Retrieve the initial brightness decision strategy for the current usage scenario from the brightness decision strategy library;
[0071] Control the display device to execute the initial brightness decision strategy, and collect device feedback data and user feedback data in real time after the execution of the initial brightness decision strategy;
[0072] Obtain the multi-objective standard differentiation weights of multiple control objectives in the current usage scenario, and construct the comprehensive reward formula R based on the multi-objective standard differentiation weights of multiple control objectives in the current usage scenario, the visualization reward formula R1, the power consumption reward formula R2, the eye protection reward formula R3, and the user feedback reward formula R4.
[0073] The comprehensive reward value of the initial brightness decision strategy is calculated based on device feedback data, user feedback data, and the comprehensive reward formula R.
[0074] Through a priority experience replay mechanism, the current brightness decision strategy is iteratively updated based on the comprehensive reward value of the initial brightness decision strategy. At the same time, the scene benchmark brightness / or multi-objective standard differential weights of the current use scenario are corrected based on the iterated brightness decision strategy to obtain the brightness adaptive control result driven by deep reinforcement learning.
[0075] In this embodiment, the current environment refers to the current surrounding environment of the display device.
[0076] In this embodiment, the usage scenario is a category categorized based on cluster analysis of environmental, user interaction, and other characteristics under different usage conditions of the display device. Different usage scenarios have their own unique feature threshold vectors, such as indoor eye-protection scenarios, indoor entertainment scenarios, outdoor strong light scenarios, low light eye-protection scenarios, and extreme temperature scenarios. Each scenario corresponds to specific environmental conditions, user behavior patterns, and appropriate brightness adjustment strategies.
[0077] In this embodiment, the historical brightness adjustment record of the display device records relevant information about the brightness adjustment of the display device at different points in the past, including raw feature sets such as ambient light intensity, blue light ratio, ambient temperature, user-screen distance, and user viewing time, as well as feature sequences of multiple dimensions. These records are a retention of various states and operations during the use of the display device.
[0078] In this embodiment, the brightness decision strategy library is a collection that stores brightness decision strategies corresponding to different usage scenarios. Once the current usage scenario is identified, the matching initial brightness decision strategy can be retrieved from it. It functions like a strategy repository, containing pre-defined brightness adjustment schemes for various possible usage scenarios.
[0079] In this embodiment, the initial brightness decision strategy is a preliminary brightness adjustment strategy retrieved from the brightness decision strategy library, tailored to the currently identified usage scenario. This strategy serves as the initial approach, controlling the display device to perform brightness adjustment operations. Subsequently, based on device feedback data and user feedback data collected after executing this strategy, a reward value is calculated using a comprehensive reward formula, and the strategy is iteratively updated to continuously optimize the brightness adjustment effect.
[0080] In this embodiment, the display device is controlled to execute an initial brightness decision strategy: that is, according to the initial brightness decision strategy retrieved from the brightness decision strategy library for the current usage scenario, the display device is driven to perform actual brightness adjustment operations. For example, if the initial brightness decision strategy specifies that the brightness should be adjusted to a certain value in the current indoor eye-protection scenario, the display device will adjust its own display brightness according to this strategy.
[0081] In this embodiment, the multi-objective standard differentiated weights of multiple control objectives are obtained under the current usage scenario: based on the current usage scenario, the weights corresponding to each of the multiple control objectives, such as visibility, power consumption, eye protection, and user feedback, are determined. The importance of each objective varies in different usage scenarios, so the weights will also differ. For example, in an indoor eye-protection scenario, more emphasis may be placed on eye protection and power consumption, resulting in relatively high weights for these two aspects; while in an outdoor bright light scenario, visibility will have a higher weight. These weights are used to construct a comprehensive reward formula to balance the roles of different control objectives in brightness adjustment decisions.
[0082] In this embodiment, a priority experience replay mechanism is used: In deep reinforcement learning, this mechanism prioritizes replaying samples based on their importance (e.g., samples that have a significant impact on the overall reward value). This allows the learning process to focus more on key information and accelerates convergence to a better policy.
[0083] The strategy is based on an iterative update based on the overall reward value: After the initial brightness decision strategy is executed, the overall reward value is calculated by combining device feedback data and user feedback data to adjust the current brightness decision strategy. If the overall reward value is high, it indicates that the strategy is effective, and the relevant decisions can be strengthened appropriately; if the reward value is low, the strategy is adjusted and optimized. For example, by adjusting the specific parameters of brightness adjustment in the strategy, we can continuously try to find a strategy that can obtain a higher reward value.
[0084] Correcting the scene's baseline brightness or the differentiated weights of multiple control objectives: Using an iterative brightness decision strategy, the scene's baseline brightness or the differentiated weights of multiple control objectives are corrected for the current usage scenario. For example, if it is found that the existing baseline brightness and weights are insufficient to achieve a good brightness adjustment effect in the current scenario, the baseline brightness is adjusted according to the new strategy, or the weights of each control objective are reallocated. Through continuous iterative updates to the strategy and correction of relevant parameters, adaptive brightness control based on deep reinforcement learning is ultimately achieved, enabling the display device to automatically adjust to a more suitable brightness in different scenarios.
[0085] like Figure 2As shown, in order to provide a comprehensive and accurate data foundation for subsequent scene recognition, it is further proposed to collect multi-dimensional environmental perception data and multi-dimensional interaction data of the current environment in real time, including:
[0086] The ambient light intensity is collected in real time within a preset visible light band in the current environment. At the same time, the proportion of harmful blue light bands within the preset band is collected as ambient light spectral distribution data, and ambient temperature data is also collected.
[0087] Real-time data collection of user interaction distance distribution and user eye gaze duration;
[0088] Generate scene feature identification vectors for the current environment based on multi-dimensional environmental perception data and multi-dimensional interaction data;
[0089] The multidimensional environmental perception data includes ambient light intensity, ambient light spectral distribution data, and ambient temperature data.
[0090] Multidimensional interactive data includes data on the distribution of user interaction distance with display devices and data on user eye gaze duration.
[0091] In this embodiment, the illuminance within the preset visible light band refers to the illuminance value collected by the ambient light sensor integrated into the display device within a specific set visible light band range (typically 380nm-760nm). This value reflects the brightness of the current ambient light and is one of the important indicators describing the characteristics of ambient light. For example, in a normal indoor office environment, this illuminance may be between 300 lux and 500 lux; while on a sunny outdoor day, the illuminance may far exceed 10,000 lux.
[0092] In this embodiment, the proportion of harmful blue light within a preset wavelength band is defined as follows: This preset wavelength band generally refers to 400nm-500nm, with particular emphasis on 415nm-455nm. The proportion of blue light within this wavelength band in the entire spectrum is collected using an ambient light sensor. An excessively high proportion of harmful blue light may damage the human eye and affect visual health. This proportion may vary under different environmental conditions.
[0093] In this embodiment, ambient temperature data is collected using a temperature sensor to monitor the temperature of the environment in which the display device is located. The temperature sensor is compatible with the wide operating temperature range of the display device (-20℃ to 70℃). Ambient temperature affects the performance and stability of the display device; extreme temperatures may necessitate special handling of the brightness adjustment strategy. For example, at temperatures below 0℃ or above 60℃, the maximum brightness adjustment may be limited to avoid device malfunction. Furthermore, ambient temperature may indirectly reflect the usage scenario; for instance, high temperatures may indicate outdoor scenarios or poor heat dissipation around the device.
[0094] In this embodiment, the user-display device interaction distance distribution data is obtained by measuring the vertical distance between the user and the display device screen using a distance sensor (such as an infrared ranging module), and presented as distribution data. The measurement range is 0.2m-2m, with an error not exceeding ±5cm. This data reflects the spatial relationship between the user and the display device; different distances may affect the user's viewing needs for the displayed content. Generally, at closer distances, users are more sensitive to details in the displayed content and may require higher contrast and appropriate brightness; at farther distances, higher overall brightness may be needed to ensure visibility. By analyzing this distance distribution data, the display device can better adapt its brightness to the user's viewing needs.
[0095] In this embodiment, user eye gaze duration data is obtained by using a visual sensor (such as a front-facing camera) and an image recognition algorithm to extract user eye features, and then calculating the duration of the user's gaze at the display device screen. User gaze duration reflects the user's level of attention to the displayed content and the degree of visual fatigue. Prolonged gaze can lead to eye fatigue, at which point the display device can appropriately adjust its brightness to alleviate fatigue. For example, when it detects that the user's eye gaze duration exceeds 30 minutes continuously, the brightness is automatically reduced by 5%-10% (not lower than the minimum brightness threshold for key information to be visible) to reduce eye fatigue.
[0096] In this embodiment, a scene feature identifier vector for the current environment is generated based on multi-dimensional environmental perception data and multi-dimensional interaction data:
[0097] The aforementioned multi-dimensional environmental perception data (ambient light intensity, ambient light spectral distribution data, i.e., the proportion of harmful blue light bands within the preset wavelength range, and ambient temperature data) and multi-dimensional interaction data (distribution data of the interaction distance between the user and the display device, and user eye gaze duration data) are comprehensively processed. First, the data from each dimension may undergo standardization preprocessing, mapping different dimensions to the [0,1] interval to eliminate dimensional differences. Then, these feature data are fused into a single vector, namely, the scene feature identifier vector for the current environment. This vector comprehensively reflects the characteristics of the current environment and user interaction, and is used for subsequent matching with scene feature threshold vectors for various usage scenarios to identify the current usage scenario and provide a basis for brightness adjustment.
[0098] To provide quantitative evidence for accurately identifying the current usage scenario, a further proposal is made to obtain a scenario feature threshold vector for each usage scenario by performing cluster analysis on the historical brightness adjustment records of the display device, including:
[0099] The original feature sets and multi-dimensional feature sequences for multiple historical moments are extracted from the historical brightness adjustment records of the display device. A standard feature matrix is generated based on the original feature sets of all historical moments. Assume the display device has 10,000 historical brightness adjustment records. From these records, multiple features for different historical moments are extracted. For example, at the first historical moment, the recorded features include an ambient light intensity of 300 lux, a blue light percentage of 20%, an ambient temperature of 25°C, a user-screen distance of 0.5m, and a user viewing time of 10 minutes. All features at this single moment constitute an original feature set. At the second historical moment, the ambient light intensity is 400 lux, the blue light percentage is 22%, and so on. These multiple features at different moments constitute the original feature sets. The feature values of a single dimension are sorted across all historical moments to obtain the feature sequence of that single dimension, such as the feature sequence composed of the feature values of ambient light intensity across all historical moments. The original feature sets of these 10,000 records are integrated to form a standard feature matrix.
[0100] By calculating the Pearson correlation coefficient between feature sequences of different dimensions, the feature correlation between different dimensions is obtained, and a feature correlation matrix is constructed based on the feature correlation between different dimensions. Multiple features at different historical moments are extracted from numerous historical brightness adjustment records of the display device, such as ambient light intensity, blue light ratio, ambient temperature, user-screen distance, and user gaze duration; these constitute the original feature set. The original feature sets from all historical moments are integrated to form a standard feature matrix. The degree of correlation between these features is measured by calculating the Pearson correlation coefficient between feature sequences of different dimensions. For example, the correlation coefficient between ambient light intensity and user-screen distance reveals the strength of their correlation. Based on these correlations, a feature correlation matrix is constructed. This matrix records the correlation between features of each dimension. In subsequent data processing, features with strong correlations are weighted differently during dimensionality reduction and other operations to avoid duplicate weights, making data processing more reasonable.
[0101] The standard feature matrix is dimensionality-reduced using a pre-defined dimensionality reduction algorithm and a feature correlation matrix to obtain a three-dimensional feature matrix. This is achieved by combining the pre-defined dimensionality reduction algorithm (such as the t-SNE manifold learning algorithm) with the previously constructed feature correlation matrix. Because the original standard feature matrix has a high dimensionality, which is not conducive to analysis and clustering, it is transformed into a three-dimensional feature matrix through dimensionality reduction. This process aims to reduce dimensionality while preserving as much of the inherent relationships between scenes in the original data as possible, preparing for subsequent analysis in three-dimensional space. For example, the 5-dimensional features of a record (300 lux, 20%, 25℃, 0.5m, 10 minutes) are transformed into three-dimensional coordinates (0.2, 0.4, 0.6) after processing with the t-SNE algorithm and the feature correlation matrix. This process again emphasizes the importance of preserving as much of the inherent relationships between scenes in the original data as possible, preparing for subsequent analysis in three-dimensional space. For example, in a high-dimensional space, some samples might be grouped into the same scene because of the ambient light intensity and the similar distance between the user and the screen. After dimensionality reduction, these samples should still be as close as possible in the three-dimensional space to maintain the inherent connection between these scenes.
[0102] A high-dimensional similarity matrix is calculated based on the standard feature matrix, and a low-dimensional similarity matrix is calculated based on the three-dimensional feature matrix. The high-dimensional similarity matrix, calculated based on the standard feature matrix, describes the degree of similarity between samples in the high-dimensional space. For example, assuming samples A (300 lux, 20%, 25℃, 0.5m, 10 minutes) and B (320 lux, 21%, 26℃, 0.55m, 12 minutes), by calculating their similarity in the high-dimensional space using a formula (such as converting to similarity based on Euclidean distance), a similarity value of 0.8 is obtained, indicating that they are highly similar in the high-dimensional space. Simultaneously, a low-dimensional similarity matrix is calculated based on the dimensionality-reduced three-dimensional feature matrix to represent the degree of similarity between samples in the three-dimensional space. Suppose that the coordinates of sample A after dimensionality reduction are (0.2, 0.4, 0.6) and the coordinates of sample B after dimensionality reduction are (0.22, 0.42, 0.62). Similarly, the similarity value between them in three-dimensional space is calculated to be 0.85 using a specific formula, indicating that they are also quite similar in three-dimensional space.
[0103] Based on gradient descent minimization of the loss function and the similarity matrices in both high and low dimensions, the 3D feature matrix is optimized to obtain an optimized 3D feature matrix. The core idea is to make the similarity distribution of the low-dimensional (3D) samples as close as possible to the similarity distribution of the high-dimensional samples. By iteratively adjusting the coordinate values in the 3D feature matrix, the loss function value, which measures the difference in similarity between the two, continuously decreases, ultimately resulting in the optimized 3D feature matrix. At this point, the sample distribution in the 3D space better reflects the relationship between the original high-dimensional samples. For example, initially, a sample in the 3D feature matrix might have coordinates (0.2, 0.4, 0.6). This coordinate value is iteratively adjusted. In each iteration, the similarity between this sample and other samples in both the high-dimensional and low-dimensional similarity matrices is calculated. Then, according to the gradient descent algorithm, the partial derivative of the loss function (such as KL divergence, which measures the difference between two similarity distributions) with respect to this coordinate is calculated, and the coordinate value is updated in the opposite direction of the gradient. Suppose that after one iteration, the coordinates are updated to (0.21, 0.41, 0.61), causing the loss function value, which measures the difference in similarity between the two, to continuously decrease. After multiple iterations, a three-dimensional optimized feature matrix is finally obtained, at which point the sample distribution in the three-dimensional space can better reflect the relationship between the samples in the original high-dimensional space.
[0104] Spatial normalization is performed on the 3D optimized feature matrix to obtain the coordinate matrix to be clustered: The 3D optimized feature matrix is spatially normalized, mapping its coordinate values to the [0,1] interval, generating the coordinate matrix to be clustered. For example, if a sample in the 3D optimized feature matrix has coordinates (10, 20, 30), after normalization (assuming the formula: new value = (original value - minimum value) / (maximum value - minimum value)), and assuming the minimum value is 0 and the maximum value is 100, then after normalization, it becomes (0.1, 0.2, 0.3). This is done to ensure all samples are on a uniform scale, facilitating subsequent cluster analysis, eliminating the influence of different coordinate value ranges, and making the clustering results more accurate.
[0105] The process involves clustering the coordinate matrix to be clustered to obtain the final sample coordinate clusters for each application scenario. First, the local density and high-density distance of each sample coordinate are calculated. Assuming there is a sample C (0.3, 0.5, 0.7) in the coordinate matrix to be clustered, its distance to samples within a certain range is calculated (e.g., the cutoff distance is set to the 2nd quantile of all distances between samples). The number of samples within this cutoff distance is counted. Let's assume the local density of sample C is calculated to be 15 (meaning there are 15 samples within a certain range). Then, the high-density distance is calculated, i.e., the distance to the nearest sample among those with a higher local density than sample C, let's assume it's 0.2. Based on these data, cluster centers are selected, for example, samples with both high local density and high-density distance are selected as cluster centers. Then, all sample coordinates are divided based on the cluster centers to obtain multiple sample coordinate clusters. For example, using a certain cluster center as a reference, surrounding samples are grouped into one sample coordinate cluster. The feature covariance matrix of each sample coordinate cluster is calculated, and the Mahalanobis distance of each sample coordinate cluster is calculated based on this. For example, the Mahalanobis distance of a sample coordinate cluster is calculated to be 0.5. The reasonableness of the partition is determined by whether the sum of the Mahalanobis distances of all sample coordinate clusters is less than a preset threshold (assuming the preset threshold is 10). If it is unreasonable, the partition is re-partitioned until the condition is met, resulting in the final sample coordinate clusters corresponding to each use case. Different final sample coordinate clusters represent different use case categories, and samples within the same cluster have high similarity in features.
[0106] High-dimensional feature sets are extracted from the final sample coordinate clusters of each usage scenario, and support vector machines (SVMs) are used to construct boundary models for each dimension. Based on these boundary models, initial threshold vectors for the corresponding usage scenarios are generated. For example, in a given final sample coordinate cluster, high-dimensional features such as ambient light intensity, blue light percentage, ambient temperature, user-screen distance, and user gaze duration are extracted. For each dimension, a boundary model is constructed using a support vector machine (SVM). For instance, for the ambient light intensity dimension, the ambient light intensity values in this sample coordinate cluster are used as positive samples, and the ambient light intensity values in other sample coordinate clusters are used as negative samples. A linear SVM classifier is trained to obtain the decision boundary parameters. Assuming the decision boundary for the ambient light intensity dimension after training is: when the ambient light intensity is greater than 250 lux, it falls within the feature range of this sample coordinate cluster. Based on these boundary models, an initial threshold vector is generated for the corresponding use scenario. For example, for this scenario, the initial threshold vector may be (250 lux, 15%, 20℃, 0.4m, 8 minutes). This vector initially determines the value range of each use scenario in each dimension of the feature.
[0107] The process involves constructing the convex hull of the final sample coordinate cluster in 3D space for each use scenario, determining the vertex set of the convex hull, and then mapping the vertex set back to a high-dimensional feature space to obtain high-dimensional vertices. Based on the extreme values of each dimension of the high-dimensional vertices, a corrected threshold vector for the corresponding use scenario is generated. For example, the coordinates of a certain final sample coordinate cluster in 3D space form a shape, and the vertex set of its convex hull includes vertices D (0.2, 0.3, 0.4), E (0.5, 0.6, 0.7), etc. The vertex set is then mapped back to the high-dimensional feature space to obtain high-dimensional vertices. Assume that after the reverse mapping, the ambient light intensity of vertex D in the high-dimensional feature space is 350 lux, and the blue light percentage is 25%. Based on the extreme values of each dimension of the high-dimensional vertices, a corrected threshold vector for the corresponding use scenario is generated. For example, the maximum and minimum values of the convex hull vertices of a scene sample coordinate cluster in 3D space, after being mapped to a high-dimensional feature space, may correct the threshold range of ambient light intensity in the initial threshold vector. For instance, if the maximum ambient light intensity of a high-dimensional vertex is 400 lux and the minimum is 200 lux, the threshold range of ambient light intensity in the initial threshold vector may be corrected to obtain a corrected threshold vector (200 lux, 12%, 18℃, 0.35m, 7 minutes), making the threshold more consistent with the actual scene boundary.
[0108] The initial threshold vector and the modified threshold vector for each usage scenario are weighted to obtain the scene feature threshold vector for each usage scenario. The weights are typically set based on the actual situation (e.g., high-dimensional features account for 0.7). Assuming the initial threshold vector is (250 lux, 15%, 20℃, 0.4m, 8 minutes) and the modified threshold vector is (200 lux, 12%, 18℃, 0.35m, 7 minutes), the weighted calculation is performed as follows:
[0109] The new ambient light intensity threshold = 0.7 × 250 + 0.3 × 200 = 235 lux;
[0110] The new blue light percentage threshold = 0.7 × 15% + 0.3 × 12% = 14.1%;
[0111] Similarly, the scene feature threshold vector for each usage scenario is obtained (235 lux, 14.1%, 19.4℃, 0.385m, 7.7 minutes).
[0112] This vector takes into account both the initial threshold and the threshold corrected based on the convex hull in three-dimensional space, and more accurately describes the feature boundary of each use case, providing a quantitative basis for accurately identifying the current use case.
[0113] To achieve a reasonable clustering partition of the coordinate matrix to be clustered, a further step is proposed to partition the coordinate matrix to obtain the final sample coordinate clusters corresponding to each application scenario, including:
[0114] Based on the coordinate matrix to be clustered, all sample coordinates are determined, and the local density and distance to high-density points for each sample coordinate are calculated. Cluster centers are then selected based on the local density and distance to high-density points of all sample coordinates. Assume we have a coordinate matrix to be clustered containing 100 sample coordinates. Each sample coordinate is a three-dimensional vector representing environmental feature information extracted and processed from historical brightness adjustment records of display devices, such as (0.3, 0.5, 0.2), (0.6, 0.4, 0.7), etc. These 100 sample coordinates are identified from this coordinate matrix. Taking one sample coordinate A (0.4, 0.3, 0.5) as an example, its local density is calculated. First, the cutoff distance is determined. Assuming that the 2nd quantile of the distance between samples is used, the cutoff distance is 0.1. Then, the number of other samples within a radius of 0.1 centered on sample A is counted. If there are 8 samples, then the local density of sample A is 8.
[0115] Next, we calculate the high-density distance. For sample A, among all samples with higher local density, we find the distance to the nearest sample. Assume sample B (0.5, 0.4, 0.6) has a local density of 10, which is greater than sample A, and the distance between sample B and sample A is 0.15. Among all samples with higher local density than sample A, sample B is the closest to sample A, so the high-density distance for sample A is 0.15. For the sample with the highest global density, assuming sample C has the highest local density among these 100 samples (15), its high-density distance is its maximum distance to all other samples, assumed to be 0.3.
[0116] By analyzing the local density and high-density point distance of all 100 sample coordinates, we plotted them in a two-dimensional decision graph (local density as the ordinate and high-density point distance as the abscissa). For example, we found that sample D (0.7, 0.6, 0.4) has a local density of 12 and a high-density point distance of 0.2, placing it in a relatively high position on the graph, meaning both its local density and high-density point distance are large. Therefore, we selected sample D as a cluster center. Using the same method, we might select several other similar samples as cluster centers. These cluster centers represent the core positions of different use cases in the feature space.
[0117] Based on the cluster centers, all sample coordinates are arbitrarily divided to obtain multiple sample coordinate clusters. The feature covariance matrix of each sample coordinate cluster is calculated, and the Mahalanobis distance of each sample coordinate cluster is calculated based on the feature covariance matrix. It is then determined whether the sum of the Mahalanobis distances of all currently obtained sample coordinate clusters is less than a preset threshold. If so, the final sample coordinate cluster is obtained; otherwise, all sample coordinates are arbitrarily re-divided until the sum of the Mahalanobis distances of the most recently obtained sample coordinate clusters is less than the preset threshold. At this point, the most recently obtained sample coordinate clusters are taken as the final sample coordinate clusters.
[0118] Based on the selected cluster center (e.g., sample D), the coordinates of these 100 samples are initially arbitrarily divided. Assume that after the division, three sample coordinate clusters are obtained, namely cluster 1, cluster 2, and cluster 3.
[0119] For cluster 1, which contains 30 sample coordinates, the feature covariance matrix of cluster 1 is calculated. This matrix reflects the correlation and dispersion among the features within the cluster. Assuming that the three dimensions of the sample coordinates in cluster 1 represent ambient light intensity, blue light percentage, and user-screen distance, the calculated feature covariance matrix will show, for example, the correlation between ambient light intensity and blue light percentage, as well as the dispersion of data in each dimension.
[0120] Based on this characteristic covariance matrix, the Mahalanobis distance of cluster 1 is further calculated. Mahalanobis distance takes into account the distribution characteristics of the data, measures the distance between a sample and the cluster center, and eliminates the influence of dimensions. Assume that the calculated Mahalanobis distance of cluster 1 is 2.5. The Mahalanobis distances of clusters 2 and 3 are calculated using the same method, assumed to be 3.0 and 2.8 respectively.
[0121] The Mahalanobis distances of these three sample coordinate clusters are summed, i.e., 2.5 + 3.0 + 2.8 = 8.3. Assuming the preset threshold is 10, since 8.3 is less than 10, it indicates that the differences between samples within the current sample coordinate clusters are relatively small, and the clustering effect is good. Therefore, these three sample coordinate clusters can be used as the final sample coordinate clusters.
[0122] If the sum of the calculated Mahalanobis distances is greater than a preset threshold, such as 12, it indicates that the current clustering is not ideal, and the differences between samples within the sample coordinate clusters are large. It is necessary to re-divide these 100 sample coordinates arbitrarily. After re-dividing, the Mahalanobis distance of each new sample coordinate cluster is recalculated and summed, then compared with the preset threshold. This process is repeated until the sum of the Mahalanobis distances of all the latest obtained sample coordinate clusters is less than the preset threshold. The sample coordinate clusters at this point are the final sample coordinate clusters.
[0123] Each final sample coordinate cluster is matched with the usage scenario type to obtain the final sample coordinate cluster corresponding to each usage scenario. After obtaining the final three sample coordinate clusters, they are analyzed. Taking one of the final sample coordinate clusters (e.g., cluster 1) as an example, it is found that most of the samples in it have characteristics such as high ambient light intensity (assuming that the ambient light intensity represented by the corresponding coordinate values after transformation is mostly above 800 lux), moderate blue light ratio (e.g., 25%-35%), and a relatively far distance between the user and the screen (e.g., 1.5m-2m).
[0124] Based on an understanding of the historical usage of display devices and an analysis of the environmental and user interaction characteristics represented by the samples in each sample coordinate cluster, cluster 1 was matched with the outdoor strong light scenario type because these characteristics match the features of an outdoor strong light scenario. Clusters 2 and 3 were analyzed and matched in the same way to obtain the final sample coordinate cluster corresponding to each usage scenario. For example, cluster 2 samples mostly have characteristics such as low ambient light intensity, low blue light ratio, and close user-screen distance, which may match them with an indoor eye-protection scenario; cluster 3 samples have other specific combination of characteristics, which may match them with an indoor entertainment scenario.
[0125] To accurately identify the current usage scenario, a further method is proposed to match the scenario feature identifier vector with the scenario feature threshold vectors of all usage scenarios to identify the current usage scenario, including:
[0126] The ratio of the number of elements in the scene feature identifier vector that satisfy the threshold condition of the corresponding dimension in the scene feature threshold vector of a single use scenario to the total number of elements is taken as the matching degree between the scene feature identifier vector and the scene feature threshold vector of the corresponding use scenario.
[0127] The scene feature identifier vector contains various feature information of the current environment, and each use scenario has a corresponding scene feature threshold vector, which specifies the value range of the scene in each feature dimension.
[0128] Suppose there are three usage scenarios: indoor eye protection scenario, outdoor strong light scenario, and indoor entertainment scenario. Their corresponding scenario feature threshold vectors are as follows:
[0129] Indoor eye protection scenarios: Ambient light intensity threshold range is 0-300 lux, blue light percentage threshold range is 0-30%, ambient temperature threshold range is 15-30℃, user-screen distance threshold range is 0.4-0.8m, and user gaze duration threshold range is 15-30 minutes.
[0130] Outdoor high-light scenarios: Ambient light intensity threshold range is 8000-100000 lux, blue light percentage threshold range is 30-70%, ambient temperature threshold range is 10-40℃, user-screen distance threshold range is 0.5-2m, and user gaze duration is arbitrary.
[0131] Indoor entertainment scenarios: Ambient light intensity threshold range is 0-300 lux, blue light percentage threshold range is 0-30%, ambient temperature threshold range is 15-30℃, user-screen distance threshold range is 0.2-0.6m, and user gaze duration threshold range is 10-25 minutes.
[0132] Each element of the scene feature identifier vector is compared with the corresponding threshold condition in the scene feature threshold vector of the indoor eye protection scene. Of the five elements in the scene feature identifier vector, three elements—blue light percentage (20%), ambient temperature (25℃), and user gaze duration (20 minutes)—satisfy the threshold conditions for the corresponding dimension of the indoor eye protection scene. Therefore, the matching degree between the scene feature identifier vector and the scene feature threshold vector of the indoor eye protection scene is 3 ÷ 5 = 0.6.
[0133] Similarly, compared with outdoor strong light scenes, only the ambient temperature (25℃) and the distance between the user and the screen (1m) meet the threshold conditions of their corresponding dimensions, with a matching degree of 2÷5=0.4.
[0134] Compared with indoor entertainment scenarios, the ambient light intensity (500 lux) does not meet its ambient light intensity threshold range. Only three elements are met: blue light percentage (20%), ambient temperature (25℃), and user gaze duration (20 minutes). The matching degree is 3÷5=0.6.
[0135] This matching degree reflects the extent to which the characteristics of the current environment match the characteristics of a certain use case.
[0136] Determine if there exists a maximum matching degree between the scene feature threshold vector and the scene feature identifier vector for multiple use scenarios. If so, then the use scenario for which the element value of the maximum weight dimension in the scene feature identifier vector meets the threshold condition of the corresponding dimension in the scene feature threshold vector is selected as the current use scenario.
[0137] After calculating the matching degree between the scene feature identifier vector and the scene feature threshold vectors of all usage scenarios, it will determine whether there are multiple usage scenarios whose scene feature threshold vectors and scene feature identifier vectors have the maximum matching degree.
[0138] After calculating the matching degree between the scene feature identifier vector and the scene feature threshold vector of all usage scenarios, it was found that the matching degree of indoor eye protection scene and indoor entertainment scene is 0.6, which is the maximum matching degree.
[0139] Assuming that in this scene feature vector, the ambient light intensity dimension has the highest weight according to pre-set weights, let's further examine the ambient light intensity element value (500 lux) in the scene feature vector and its threshold condition in the corresponding scene feature threshold vector for the ambient light intensity dimension in these two use cases where the matching degree is the highest. The ambient light intensity threshold range for indoor eye-protection scenes is 0-300 lux, which is not met; the ambient light intensity threshold range for indoor entertainment scenes is also 0-300 lux, which is also not met.
[0140] Let's assume another scenario: if the maximum weighted dimension becomes the distance between the user and the screen, its value in the scene feature identifier vector is 1m. The threshold range for user-screen distance in indoor eye-protection scenarios is 0.4-0.8m, which does not meet the requirement; the threshold range for user-screen distance in indoor entertainment scenarios is 0.2-0.6m, which also does not meet the requirement.
[0141] Let's re-assume the maximum weight dimension is blue light percentage, with a value of 20%. The blue light percentage threshold range for indoor eye-protection scenarios is 0-30%, which satisfies this requirement; the same applies to indoor entertainment scenarios. However, if the weighting of blue light percentage is slightly higher for indoor eye-protection scenarios than for indoor entertainment scenarios, then the indoor eye-protection scenario is ultimately determined as the current usage scenario. This method allows for a more accurate identification of the actual usage scenario when there are cases with the same match degree.
[0142] like Figure 3 As shown, in order to provide comprehensive feedback information for evaluating the initial brightness decision strategy, it is further proposed to collect device feedback data and user feedback data in real time after the initial brightness decision strategy is implemented, including:
[0143] The device collects real-time data on the contrast ratio of the displayed content, the brightness output power, and the proportion of harmful blue light bands within the preset wavelength range after the initial brightness decision strategy is implemented. At the same time, it collects real-time data on the brightness value manually adjusted by the user and the duration of the user's eye gaze as user feedback data.
[0144] In this embodiment, the display content contrast ratio refers to the ratio of the brightness of the brightest part to the darkest part of the content displayed on the device. It is an important indicator for measuring display quality and visibility. Higher contrast ratio makes displayed content such as images and text clearer and more vivid, helping users to more easily identify key information. For example, when reading text, high contrast ratio makes the boundary between text and background clear, reducing eye strain; when watching images or videos, it can present richer details and vivid colors. In the display device brightness adaptive control method based on multimodal environment perception and deep reinforcement learning, the display content contrast ratio is used as one of the device feedback data to evaluate the brightness adjustment effect. By collecting this data in real time, it can be determined whether the current brightness setting meets the visibility requirements. For example, when the display content contrast ratio is ≥5:1, the corresponding key information visibility reward in the reward function may receive a higher score. If the contrast ratio is insufficient, it may be necessary to adjust the brightness to improve visibility.
[0145] In this embodiment, brightness output power represents the power consumed by the display device to achieve a certain brightness level. It reflects the device's energy consumption and is crucial for power management. Different brightness settings correspond to different brightness output powers; generally, higher brightness results in higher power consumption. In this adaptive brightness control method, brightness output power serves as device feedback data to evaluate the impact of brightness adjustment strategies on power consumption. For example, when designing the reward function, if the brightness output power decreases compared to the previous round, a positive reward can be obtained in the power optimization reward part, and the greater the decrease, the higher the reward value. This encourages the system to generate more energy-efficient brightness control commands, achieving the control objective of "minimizing device power consumption."
[0146] In this embodiment, the brightness value manually adjusted by the user is a value set by the user based on their own visual perception and actual needs, after manually adjusting the brightness of the display device. It directly reflects the user's subjective preference for the current display brightness.
[0147] To provide a quantitative tool for accurately evaluating the effectiveness of the initial brightness decision strategy, a comprehensive reward formula R is constructed based on the multi-objective standard differential weights of multiple control objectives in the current usage scenario, the visual reward formula R1, the power consumption reward formula R2, the eye protection reward formula R3, and the user feedback reward formula R4. This formula includes:
[0148]
[0149] R4 = α × (1 - Standard value of user accommodation deviation rate) + β × Standard value of user gaze stability
[0150] +γ×User explicit rating standard value
[0151] The standard differential weights of multiple control objectives in the current usage scenario are modified and the boundary constraints are applied to obtain the modified differential weights of multiple control objectives in the current usage scenario.
[0152] Based on the multi-objective correction and differentiation weights of multiple control objectives in the current usage scenario, the visualization reward formula R1, the power consumption reward formula R2, the eye protection reward formula R3, and the user feedback reward formula R4, a comprehensive reward formula R is constructed:
[0153] R=ω1×R1+ω2×R2+ω3×R3+ω4×R4
[0154] In the formula, ω1, ω2, ω3, and ω4 are the multi-objective standard differentiation weights of multiple control objectives in the current usage scenario.
[0155] In this embodiment, the visual contrast threshold for the current usage scenario is a lower limit set for contrast to ensure the visibility of key information in the displayed content, based on the currently identified usage scenario. Different usage scenarios have different requirements for visual contrast due to varying environmental conditions and user needs. For example, in an indoor eye-protection scenario, considering the potential eye fatigue from prolonged viewing, a relatively low contrast threshold might be set to ensure clear visibility of key information such as text. In contrast, in outdoor bright light scenarios, a higher visual contrast threshold might be set to make the displayed content stand out in strong light. This threshold serves as a crucial basis for judging the visibility of the displayed content. If the current display content contrast is lower than this threshold, it means that the visibility of key information may be insufficient. The system may then adjust brightness or other methods to increase the contrast to meet the visibility requirements. Simultaneously, in the reward function, the visibility reward component may be assigned a corresponding value based on the comparison result with this threshold.
[0156] In this embodiment, the power consumption limit for the current usage scenario is a pre-set maximum power consumption value allowed for the display device in that scenario, based on the characteristics of the current usage scenario. Different usage scenarios will have different power consumption limits due to factors such as device operating requirements and energy supply conditions. For example, in scenarios using battery-powered mobile display devices, the power consumption limit may be set relatively low to ensure device battery life; while in scenarios using indoor fixed display devices connected to a power source with less energy consumption restriction, the power consumption limit may be relatively lenient. This power consumption limit is used to constrain the brightness adjustment strategy of the display device. When the brightness output power approaches or exceeds this limit, the system will take corresponding measures, such as reducing brightness or optimizing other energy consumption factors, to ensure that the device power consumption is within an acceptable range, while also helping to achieve rational energy utilization while meeting visibility and other requirements.
[0157] To obtain multi-objective modified differential weights that better reflect the actual situation of the current application scenario, a further proposal is made to modify and constrain the standard multi-objective differential weights of multiple control objectives in the current application scenario, thereby obtaining multi-objective modified differential weights of multiple control objectives in the current application scenario, including:
[0158] Based on the key dimension elements in the scene feature identifier vector, calculate the weight correction values of each weight in the multi-objective standard differential weights for multiple control objectives in the current usage scenario:
[0159] The scene feature vector comprehensively reflects the characteristics of the current environment and user interaction. Key dimensions, such as the standardized values of ambient light intensity, blue light percentage, and user gaze duration, play a crucial indicative role in adjusting the weights of multiple control objectives (visibility, power consumption, eye protection, user feedback, etc.). For example, a high standardized value for ambient light intensity in the scene feature vector indicates strong ambient light; to ensure visibility, the visibility weight may need to be increased. Correspondingly, the weights of other objectives may need to be adjusted. In the specific calculation, based on pre-defined rules and the values of these key dimensions, the correction value for the weight of each control objective is calculated.
[0160] For example, visibility weight correction: Δω1=0.1×I std ;
[0161] Among them, I std This is a standardized value for ambient light intensity; the stronger the ambient light, the greater the increase in visibility weight.
[0162] Δω2=0.08×(1-P ratio )+0.04×D std
[0163] Among them, P ratio The ratio of the current brightness output power to the brightness output power of the previous round (P) ratio = Current brightness output power / Previous brightness output power). If the current brightness output power is lower than the previous round, P ratio A value less than 1 indicates a positive value, meaning the power consumption weight will increase; conversely, a value less than 1 indicates a decrease. D std The greater the distance between the user and the display device, the more power the device may need to maintain visibility. Therefore, the larger the distance standardization value, the larger the power consumption weight correction value.
[0164] Eye protection weight correction: Δω3=0.1×B std +0.05×T std ;
[0165] Among them, B stdThis represents a standardized value indicating the proportion of blue light. A higher proportion of blue light and a longer viewing time result in a greater increase in eye protection weight. (T) std A standardized value representing the duration of a user's gaze;
[0166] User feedback weight correction: Δω4=0.05×(1-standard value of user adjustment deviation rate) 2 (The smaller the standard value of the user-adjusted deviation rate, the greater the reduction in the weight of user feedback.)
[0167] In this way, the standard differentiated weights of multiple control objectives can be modified in a targeted manner based on the specific characteristics of the current scenario.
[0168] The corrected weight values of each weight in the multi-objective standard differential weights for multiple control objectives in the current usage scenario are normalized to obtain the corrected differential weights for multiple control objectives in the current usage scenario:
[0169] After calculating the corrected weight values, since the sum of these corrected values is not necessarily 1, they cannot be directly used as new weights. Therefore, normalization is required to ensure that the sum of the corrected weights is 1, thus guaranteeing a reasonable relative proportion between the weights of each control objective and conforming to the basic definition of weights. Specifically, the multi-objective standard differential weights for multiple control objectives in the current application scenario are first added to their respective corrected weight values to obtain a new set of weight values. Then, this new set of weight values is normalized. For example, if the corrected weights are ω1+Δω1, ω2+Δω2, ω3+Δω3, and ω4+Δω4, the normalized multi-objective corrected differential weights are: ω′1=(ω1+Δω1) / ∑(ω i +Δω i ), where i = 1, 2, 3, 4. After normalization, the resulting multi-objective modified differential weights take into account the influence of the current scene features on the weights of each control objective, while ensuring that the sum of the weights is 1. This can be used to construct a comprehensive reward formula to more reasonably balance the roles of different control objectives in brightness adjustment decisions.
[0170] To provide accurate basis for strategy evaluation and optimization, a comprehensive reward value for the initial brightness decision strategy is further proposed, calculated based on device feedback data, user feedback data, and the comprehensive reward formula R. This includes:
[0171] Based on user feedback data, including manually adjusted brightness values, eye fixation duration, and user ratings, we analyzed the standard values for user accommodation deviation rate, user fixation stability, and user display rating.
[0172] The display content contrast, brightness output power, proportion of harmful blue light bands in the preset band, standard value of user adjustment deviation rate, standard value of user gaze stability, and standard value of user display score in the device feedback data are substituted into the comprehensive reward formula R to obtain the comprehensive reward value of the initial brightness decision strategy.
[0173] To extract quantifiable standard values for evaluating strategy effectiveness from user feedback data, this paper further proposes standard values for user accommodation deviation rate, user gaze stability, and user display rating based on user-manually adjusted brightness values, user eye fixation duration data, and user ratings from the user feedback data. These standard values include:
[0174] The ratio of the difference between the brightness value manually adjusted by the user and the brightness value recommended by the system to the brightness value recommended by the system is taken as the user adjustment deviation rate. The user adjustment deviation rate is then standardized to obtain the standard value of the user adjustment deviation rate.
[0175] Based on user eye fixation duration data from user feedback, the amplitude of user fixation point fluctuation was determined, and the stability of user fixation was determined based on the amplitude of fixation point fluctuation.
[0176]
[0177] The user gaze stability is standardized to obtain a standard value for user gaze stability;
[0178] Standardize the user ratings in the user feedback data to obtain the standard values for user display ratings.
[0179] In this embodiment, the user adjustment deviation rate is standardized to obtain a standard value for the user adjustment deviation rate:
[0180] The user adjustment deviation rate reflects the degree of difference between the brightness manually adjusted by the user and the system-recommended brightness. The calculation formula is (user-adjusted brightness - system-recommended brightness) / system-recommended brightness. However, this value can vary significantly depending on the specific brightness value. To better integrate it into calculations such as the comprehensive reward formula, standardization is necessary. The purpose of standardization is to constrain it within a specific range, making it comparable. A common method is to standardize the user adjustment deviation rate U... adjust Through formula U adjust,std =max(-1,min(1,U) adjust The user adjustment deviation rate standard value U is obtained by processing the data. adjust,std This is constrained to the interval [-1, 1]. For example, if the user's adjustment deviation rate U adjust A value of 2 (indicating that the user manually adjusts the brightness to 3 times the system's recommended brightness), after standardization, Uadjust,std It is 1; if U adjust A value of -0.5 (indicating that the user manually adjusted the brightness to half of the system's recommended brightness) means U adjust,std The value is -0.5. This standardized value can be directly used to calculate user feedback rewards, etc., to measure the user's acceptance of the system's recommended brightness.
[0181] In this embodiment, the fluctuation range of the user's gaze point is determined based on the user's eye fixation duration data in the user feedback data:
[0182] After acquiring user gaze duration data through visual sensors, image recognition algorithms can not only determine gaze duration but also track changes in the user's gaze point's position on the screen. The gaze point fluctuation amplitude is the range of positional change of the gaze point during the user's gaze on the screen. For example, using a fixed point on the screen as a reference, the maximum distance the user's gaze point moves around that reference point over a period of time is the gaze point fluctuation amplitude.
[0183] In this embodiment, the user gaze stability is standardized to obtain a standard value for user gaze stability:
[0184] User gaze stability is calculated using a formula, representing the relative stability of a user's gaze at the screen. However, the calculated gaze stability value will vary depending on the screen size. To facilitate standardized calculation and comparison, standardization is necessary. This is typically achieved using the formula U... fixation,std =max(0,min(1,U) fixation The value is processed and constrained to the [0,1] interval to obtain the standard value U of user gaze stability. fixation,std This standardized value for user gaze stability can be used to calculate the user feedback reward item in the comprehensive reward formula, reflecting the user's stable viewing experience.
[0185] In this embodiment, the user ratings in the user feedback data are standardized to obtain the user display rating standard value:
[0186] If the display device supports user ratings of its display quality (typically 1-5 stars), these ratings need to be standardized to ensure they are appropriately incorporated into the system's calculations. The standardization method involves converting the user rating into a UTF-8 value. rating Through formula U rating,std =(U rating The value of -1) / 4 is processed and mapped to the [0,1] interval to obtain the user display rating standard value U. rating,std For example, if a user rates it 3 stars, then U rating,std= (3-1) / 4 = 0.5. This standardized user display rating can be directly used to calculate the user feedback reward item in the comprehensive reward formula, quantifying the user's subjective evaluation of the display effect and enabling the system to better optimize brightness decision strategies based on user ratings.
[0187] Obviously, those skilled in the art can make various modifications and variations to this invention without departing from its spirit and scope. Therefore, if these modifications and variations fall within the scope of this invention and its equivalents, this invention also intends to include these modifications and variations.
Claims
1. A display device brightness adaptive control method based on multimodal environment perception and deep reinforcement learning, characterized in that, include: It collects multi-dimensional environmental perception data and multi-dimensional interaction data in real time under the current environment, and generates scene feature identification vectors under the current environment based on the multi-dimensional environmental perception data and multi-dimensional interaction data. By performing cluster analysis on the historical brightness adjustment records of the display device, the scene feature threshold vector of each usage scenario is obtained, and the scene feature identifier vector is matched with the scene feature threshold vectors of all usage scenarios to identify the current usage scenario; Retrieve the initial brightness decision strategy for the current usage scenario from the brightness decision strategy library; Control the display device to execute the initial brightness decision strategy, and collect device feedback data and user feedback data in real time after the execution of the initial brightness decision strategy; Obtain the multi-objective standard differential weights of multiple control objectives in the current usage scenario, and visualize the reward formula based on the multi-objective standard differential weights of multiple control objectives in the current usage scenario. Power consumption reward formula Eye Protection Reward Formula User feedback reward formula A comprehensive reward formula was constructed. ; Based on device feedback data, user feedback data, and a comprehensive reward formula Calculate the overall reward value of the initial brightness decision strategy; Through a priority experience replay mechanism, the current brightness decision strategy is iteratively updated based on the comprehensive reward value of the initial brightness decision strategy. At the same time, the scene benchmark brightness / or multi-objective standard differential weights of the current use scenario are corrected based on the iterated brightness decision strategy to obtain the brightness adaptive control result driven by deep reinforcement learning. Among them, the multi-objective standard differentiated weights and visualized reward formulas based on the multi-control objectives in the current usage scenario. Power consumption reward formula Eye Protection Reward Formula User feedback reward formula A comprehensive reward formula was constructed. ,include: ; ; ; ; The standard differential weights of multiple control objectives in the current usage scenario are modified and the boundary constraints are applied to obtain the modified differential weights of multiple control objectives in the current usage scenario. Based on multiple control objectives, differentiated weights are adjusted and a visual reward formula is developed for the current application scenario. Power consumption reward formula Eye Protection Reward Formula User feedback reward formula A comprehensive reward formula was constructed. : ; In the formula, , , , To differentiate the weights of multiple control objectives in the current usage scenario.
2. The display device brightness adaptive control method based on multimodal environment perception and deep reinforcement learning according to claim 1, characterized in that, Real-time acquisition of multi-dimensional environmental perception data and multi-dimensional interaction data of the current environment, including: The ambient light intensity is collected in real time within a preset visible light band in the current environment. At the same time, the proportion of harmful blue light bands within the preset band is collected as ambient light spectral distribution data, and ambient temperature data is also collected. Real-time data collection of user interaction distance distribution and user eye gaze duration; Generate scene feature identification vectors for the current environment based on multi-dimensional environmental perception data and multi-dimensional interaction data; The multidimensional environmental perception data includes ambient light intensity, ambient light spectral distribution data, and ambient temperature data. Multidimensional interactive data includes data on the distribution of user interaction distance with display devices and data on user eye gaze duration.
3. The display device brightness adaptive control method based on multimodal environment perception and deep reinforcement learning according to claim 1, characterized in that, By performing cluster analysis on the historical brightness adjustment records of display devices, a scene feature threshold vector for each usage scenario is obtained, including: The original feature sets and feature sequences of multiple historical moments are extracted from the historical brightness adjustment records of the display device. A standard feature matrix is generated based on the original feature sets of all historical moments. The feature correlation between different dimensions is obtained by calculating the Pearson correlation coefficient between feature sequences of different dimensions. A feature correlation matrix is constructed based on the feature correlation between different dimensions. The standard feature matrix is reduced in dimensionality based on a preset dimensionality reduction algorithm and feature correlation matrix to obtain a three-dimensional feature matrix. A high-dimensional spatial similarity matrix is calculated based on the standard feature matrix, and a low-dimensional spatial similarity matrix is calculated based on the three-dimensional feature matrix. Based on gradient descent minimization of the loss function, the high-dimensional similarity matrix, and the low-dimensional similarity matrix, the three-dimensional feature matrix is optimized to obtain the three-dimensional optimized feature matrix. The spatial normalization of the 3D optimized feature matrix is performed to obtain the coordinate matrix to be clustered. The coordinate matrix to be clustered is divided into clusters to obtain the final sample coordinate clusters corresponding to each use case. High-dimensional feature sets are extracted from the final sample coordinate clusters of each use case, and support vector machines are used to construct boundary models of each dimension. Based on the boundary models of each dimension, the initial threshold vectors of the corresponding use cases are generated. Construct the convex hull of the final sample coordinate cluster in three-dimensional space for each use scenario, determine the vertex set of the convex hull, and reverse map the vertex set to a high-dimensional feature space to obtain high-dimensional vertices. Generate the corrected threshold vector for the corresponding use scenario based on the extreme values of each dimension of the high-dimensional vertices. The initial threshold vector and the modified threshold vector for each use scenario are weighted to obtain the scenario feature threshold vector for each use scenario.
4. The display device brightness adaptive control method based on multimodal environment perception and deep reinforcement learning according to claim 3, characterized in that, The coordinate matrix to be clustered is divided into clusters to obtain the final sample coordinate clusters corresponding to each use case, including: Based on the coordinate matrix to be clustered, the coordinates of all samples are determined, and the local density and distance to high-density points of each sample coordinate are calculated. Then, the cluster center is selected based on the local density and distance to high-density points of all sample coordinates. All sample coordinates are arbitrarily divided based on cluster centers to obtain multiple sample coordinate clusters. The feature covariance matrix of each sample coordinate cluster is calculated, and the Mahalanobis distance of each sample coordinate cluster is calculated based on the feature covariance matrix. It is determined whether the sum of the Mahalanobis distances of all the currently obtained sample coordinate clusters is less than a preset threshold. If so, the final sample coordinate cluster is obtained. Otherwise, all sample coordinates are arbitrarily re-divided until the sum of the Mahalanobis distances of the latest obtained sample coordinate clusters is less than the preset threshold. Then, the latest obtained sample coordinate clusters are taken as the final sample coordinate clusters. Each final sample coordinate cluster is matched with the usage scenario type to obtain the final sample coordinate cluster corresponding to each usage scenario.
5. The display device brightness adaptive control method based on multimodal environment perception and deep reinforcement learning according to claim 1, characterized in that, The scene feature identifier vector is matched with the scene feature threshold vectors of all usage scenarios to identify the current usage scenario, including: The ratio of the number of elements in the scene feature identifier vector that satisfy the threshold condition of the corresponding dimension in the scene feature threshold vector of a single use scenario to the total number of elements is taken as the matching degree between the scene feature identifier vector and the scene feature threshold vector of the corresponding use scenario. Determine if there exists a maximum matching degree between the scene feature threshold vector and the scene feature identifier vector for multiple use scenarios. If so, then the use scenario for which the element value of the maximum weight dimension in the scene feature identifier vector meets the threshold condition of the corresponding dimension in the scene feature threshold vector is taken as the current use scenario.
6. The display device brightness adaptive control method based on multimodal environment perception and deep reinforcement learning according to claim 1, characterized in that, Real-time acquisition of device feedback data and user feedback data after the execution of the initial brightness decision strategy, including: The device collects real-time data on the contrast ratio of the displayed content, the brightness output power, and the proportion of harmful blue light bands within the preset wavelength range after the initial brightness decision strategy is implemented. At the same time, it collects real-time data on the brightness value manually adjusted by the user and the duration of the user's eye gaze as user feedback data.
7. The display device brightness adaptive control method based on multimodal environment perception and deep reinforcement learning according to claim 1, characterized in that, The standard differentiation weights of multiple control objectives in the current usage scenario are corrected and boundary constraints are applied to obtain the corrected differentiation weights of multiple control objectives in the current usage scenario, including: Based on the key dimension elements in the scene feature identifier vector, calculate the weight correction value of each weight in the multi-objective standard differential weight of multiple control objectives in the current usage scenario; The correction values of each weight in the multi-objective standard differential weight of multiple control objectives under the current usage scenario are normalized to obtain the multi-objective corrected differential weight of multiple control objectives under the current usage scenario.
8. The display device brightness adaptive control method based on multimodal environment perception and deep reinforcement learning according to claim 1, characterized in that, Based on device feedback data, user feedback data, and a comprehensive reward formula Calculate the overall reward value of the initial brightness decision strategy, including: Based on user feedback data, including manually adjusted brightness values, eye fixation duration, and user ratings, we analyzed the standard values for user accommodation deviation rate, user fixation stability, and user display rating. Substitute the display contrast, brightness output power, percentage of harmful blue light bands within the preset wavelength, user adjustment deviation rate standard value, user gaze stability standard value, and user display score standard value from the device feedback data into the comprehensive reward formula. The overall reward value of the initial brightness decision strategy is obtained.
9. The display device brightness adaptive control method based on multimodal environment perception and deep reinforcement learning according to claim 8, characterized in that, Based on user feedback data, including manually adjusted brightness values, eye fixation duration, and user ratings, standard values for user accommodation deviation rate, eye fixation stability, and display rating were analyzed, including: The ratio of the difference between the brightness value manually adjusted by the user and the brightness value recommended by the system to the brightness value recommended by the system is taken as the user adjustment deviation rate. The user adjustment deviation rate is then standardized to obtain the standard value of the user adjustment deviation rate. Based on user eye fixation duration data from user feedback, the amplitude of user fixation point fluctuation was determined, and the stability of user fixation was determined based on the amplitude of fixation point fluctuation. ; The user gaze stability is standardized to obtain a standard value for user gaze stability; Standardize the user ratings in the user feedback data to obtain the standard values for user display ratings.
Citation Information
Patent Citations
CN117558244A
CN119479519A
CN120358642A
CN120379117A