Display device backlight optimization method and server based on deep learning
Patent Information
- Application Number
- CN202610778465.2
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2026-06-02
- Publication Date
- 2026-08-28
- Estimated Expiration
- 2046-06-02
AI Technical Summary
然而,上述方法普遍仅从单帧或相邻帧的静态亮度统计特征出发,缺乏对显示内容叙事节奏、语义演变趋势及视觉焦点迁移路径的深层解析能力,难以在多背光分区的空间维度与连续画面的时间维度上生成与显示内容语义结构连续匹配的动态背光调控序列
[0007]与现有技术相比,本发明的有益效果是:通过构建显示内容帧序列到时序语义空间的流形演化轨迹特征,以流形几何结构完整刻画显示内容的叙事节奏与视觉焦点迁移路径,再由背光拓扑映射网络在保持流形局部邻域和全局分布结构的前提下将该轨迹特征端到端映射为多维背光调控曲线参数,使得显示内容语义的渐变与突变能够一致地反映在背光分区的亮度与色温调控上,最终生成与画面切换叙事节奏和视觉焦点迁移路径同步匹配的背光同步控制指令序列,实现了背光发光分布对显示内容语义时空演进的连续跟随与动态重构。
Smart Images

Figure CN122337148B_ABST
Abstract
Description
Technical Field
[0001] This invention relates to the field of deep learning technology, specifically to a deep learning-based method for optimizing backlight in display devices and a server. Background Technology
[0002] Backlight control of display devices is a crucial aspect of enhancing the visual presentation of terminal devices. Existing technologies utilize local dimming to divide the display panel backlight into multiple independently controllable luminous zones. Based on the brightness statistics of each zone within the image content, the driving current of each zone is adjusted, achieving a localized light control effect that brightens high-brightness areas and darkens low-brightness areas. To adapt to dynamic image changes, related methods generate backlight adjustment signals by detecting changes in the brightness distribution of preset areas in the display frame sequence and send instructions to the backlight driver according to a preset timing control strategy. However, these methods generally only consider the static brightness statistics of a single frame or adjacent frames, lacking the ability to deeply analyze the narrative rhythm, semantic evolution trends, and visual focus migration paths of the displayed content. This makes it difficult to generate a dynamic backlight control sequence that continuously matches the semantic structure of the displayed content across the spatial dimensions of multiple backlight zones and the temporal dimensions of continuous images. Summary of the Invention
[0003] The purpose of this invention is to provide a deep learning-based method and server for optimizing the backlight of display devices, in order to solve the problems mentioned in the background art.
[0004] This invention provides a deep learning-based method for optimizing the backlight of display devices, comprising: Obtain the display content frame sequence corresponding to the display content to be optimized, wherein the display content frame sequence is composed of multiple display content frame units arranged in chronological order; The display content frame sequence is input into a preset temporal semantic deconstruction network. The visual semantic encoding layer in the temporal semantic deconstruction network performs semantic parsing on each display content frame unit to generate a visual semantic embedding vector sequence. The sequence is then processed by the temporal manifold embedding layer in the temporal semantic deconstruction network to obtain the manifold evolution trajectory features of the display content frame sequence in the temporal semantic space. The backlight topology mapping network performs topology-preserving end-to-end mapping processing on the manifold evolution trajectory features to generate multidimensional backlight control curve parameters. The backlight topology mapping network learns the continuous mapping relationship from the manifold evolution trajectory feature space to the backlight partition control parameter space during the training phase. The backlight partition signal is expanded and processed according to the parameters of the multidimensional backlight control curve to obtain the backlight partition driving signal sequence. The backlight zone driving signal sequence and the display content frame sequence are synchronized and aligned according to the timestamp to generate a backlight synchronization control command sequence. The backlight synchronization control command sequence is then output to the backlight driver to control the backlight zone to perform dynamic light emission adjustment that matches the narrative rhythm of the display content and the shift of visual focus.
[0005] This invention provides a display device backlight optimization server, comprising: A processor; a storage device having a computer program stored thereon; a network interface for providing network communication functions; when the computer program is executed by the processor, the processor enables the processor to implement the aforementioned deep learning-based display device backlight optimization method.
[0006] The present invention provides a readable storage medium storing a program or instructions, which, when executed by a processor, implement the aforementioned deep learning-based display device backlight optimization method.
[0007] Compared with the prior art, the beneficial effects of the present invention are as follows: by constructing the manifold evolution trajectory features from the display content frame sequence to the temporal semantic space, the narrative rhythm and visual focus migration path of the display content are completely characterized by the manifold geometric structure. Then, the backlight topology mapping network maps the trajectory features end-to-end into multi-dimensional backlight control curve parameters while maintaining the local neighborhood and global distribution structure of the manifold. This allows the gradual changes and abrupt changes in the semantics of the display content to be consistently reflected in the brightness and color temperature control of the backlight partition. Finally, a backlight synchronization control command sequence that is synchronously matched with the narrative rhythm of the screen switching and the visual focus migration path is generated, realizing the continuous following and dynamic reconstruction of the spatiotemporal evolution of the semantics of the display content by the backlight emission distribution. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments or related technologies of this application, the accompanying drawings used in the description of the embodiments or related technologies will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 This is a flowchart illustrating a deep learning-based backlight optimization method for display devices, provided as an embodiment of this application.
[0010] Figure 2 This is a schematic diagram of the basic structure of a display device backlight optimization server provided in an embodiment of this application.
[0011] Figure 3 This is a functional block diagram of a display device backlight optimization device provided in an embodiment of this application. Detailed Implementation
[0012] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0013] Please see Figure 1 , Figure 1 This is a flowchart of a deep learning-based display device backlight optimization method provided in an embodiment of this application. The method can be executed by a display device backlight optimization server, or by a display device backlight optimization server and a server together. The method includes steps 110-150.
[0014] This invention discloses a deep learning-based method for optimizing the backlight of display devices. Applied to a display device backlight optimization server, this method generates backlight synchronization control commands that match the narrative rhythm and visual focus shifts of the displayed content frame sequences through temporal semantic parsing, topology-preserving mapping, and partition signal expansion. These commands are then output to the backlight driver to achieve dynamic adjustment of the backlight zones. For ease of understanding, this invention uses a backlight optimization scenario during streaming media playback on a terminal device display as an example to illustrate the implementation logic. However, this example does not constitute a limitation on the application scenario.
[0015] Step 110: Obtain the display content frame sequence corresponding to the display content to be optimized. The display content frame sequence is composed of multiple display content frame units arranged in chronological order.
[0016] The display content frame sequence originates from the streaming media data stream to be rendered by the display device. It continuously captures snapshots of the screen from the frame buffer of the graphics processing unit at preset sampling intervals; each snapshot is a display content frame unit. Each display content frame unit is organized in a three-dimensional data structure, including pixel height coordinates, pixel width coordinates, and color information channels. The color information channels include at least red, green, and blue component intensity channels. All display content frame units are arranged according to the order of their sampling times, forming an ordered set on the timeline. This ordered set preserves the complete temporal evolution of the displayed content, and each display content frame unit is associated with its own timestamp information to mark its precise position on the timeline.
[0017] Step 120: Input the display content frame sequence into a preset temporal semantic deconstruction network. Perform semantic parsing on each display content frame unit through the visual semantic encoding layer in the temporal semantic deconstruction network to generate a visual semantic embedding vector sequence. Then, perform spatiotemporal manifold mapping processing through the temporal manifold embedding layer in the temporal semantic deconstruction network to obtain the manifold evolution trajectory features of the display content frame sequence in the temporal semantic space.
[0018] After acquiring the sequence of display content frames, it is used as the input to a preset temporal semantic deconstruction network, which consists of a visual semantic coding layer and a temporal manifold embedding layer connected in series.
[0019] The visual semantic encoding layer independently performs semantic parsing on each display content frame unit, mapping the original spatial information containing pixel intensity distribution and color distribution to a semantic feature space with a preset dimension. Each point in the semantic feature space corresponds to a visual semantic embedding vector, representing the comprehensive expression of the corresponding frame in visual semantic aspects such as color composition, brightness distribution, and object outline.
[0020] The visual semantic embedding vectors of all display content frame units are arranged in order according to their timestamps, forming a visual semantic embedding vector sequence. This sequence is then input into a temporal manifold embedding layer, which is configured to perform manifold modeling of the evolution path of the sequence in the temporal semantic space. By capturing the offset direction and offset amount of the visual semantic embedding vectors in the temporal neighborhood and analyzing the changes in their local geometry, a manifold evolution trajectory feature containing the visual semantic offset sequence and the local curvature features of the semantic manifold is generated, which fully depicts the spatiotemporal pattern of the semantic content of the display content flowing, turning, and converging on the time axis.
[0021] In this embodiment of the invention, the processing mechanism of the temporal semantic deconstruction network for the display content frame sequence can be further implemented according to steps 121 to 126.
[0022] Step 121: Extract the pixel brightness statistical distribution information and pixel color statistical distribution information of each display content frame unit in the display content frame sequence, and generate the visual low-level statistical description vector of the display content frame unit.
[0023] For each display content frame unit, all its spatial positions are traversed first, and the luminance intensity value of each position in the luminance channel is accumulated. The luminance intensity value is obtained by weighted summation of the intensity values of the corresponding position in the red, green, and blue channels. The weighting coefficient of the red channel is a preset first luminance weight, the weighting coefficient of the green channel is a preset second luminance weight, and the weighting coefficient of the blue channel is a preset third luminance weight.
[0024] Next, the range of luminance intensity values is divided into multiple consecutive luminance intensity intervals. The proportion of spatial locations falling into each interval is calculated to form a luminance distribution histogram vector. Simultaneously, the intensity values of each spatial location in the red, green, and blue channels are normalized to eliminate the offset caused by differences in overall illumination intensity, resulting in normalized red, green, and blue channel intensity values. The ranges of these normalized values are then further divided into multiple consecutive color intervals. The proportion of spatial locations falling into each interval is calculated to form red, green, and blue component distribution histogram vectors, respectively.
[0025] Finally, the histogram vectors of brightness distribution, red component distribution, green component distribution, and blue component distribution are concatenated along the feature dimension to obtain the visual low-level statistical description vector of the display content frame unit.
[0026] Step 122: Input the visual low-level statistical description vector into the multi-level residual coding module of the visual semantic coding layer. Through multiple residual processing units connected in series in the multi-level residual coding module, the visual low-level statistical description vector is compressed at each level of spatial scale and expanded in semantic channels to generate a multi-level semantic response map.
[0027] The multi-level residual coding module consists of a first residual processing unit, a second residual processing unit, and a third residual processing unit connected sequentially according to the data processing flow. The first residual processing unit receives the visual low-level statistical description vector as input. Internally, the first residual processing unit contains a first main path and a first skip connection path. The first main path is composed of a first fully connected layer, a first batch normalization layer, a first linear rectified activation layer, a second fully connected layer, and a second batch normalization layer connected in series. The first fully connected layer expands the dimension of the visual low-level statistical description vector to a first expanded dimension. Then, the first batch normalization layer stabilizes the feature distribution, followed by the first linear rectified activation layer introducing non-linear mapping capability. Finally, the second fully connected layer compresses it to a first compressed dimension, and the second batch normalization layer stabilizes the output again. The first skip connection path contains a first skip fully connected layer, which directly maps the visual low-level statistical description vector to the first compressed dimension. The output of the first main path and the output of the first skip connection path are added element-wise. The result is processed by the second linear rectified activation layer to generate the first residual response vector. The dimension of the first residual response vector is greater than that of the visual underlying statistical description vector, and this process realizes the first expansion of the semantic channel.
[0028] The first residual response vector is fed into the second residual processing unit. The structure of the second residual processing unit is similar to that of the first residual processing unit, but its internal first fully connected layer expands the dimension to a second expanded dimension, and the second fully connected layer compresses the dimension to a second compressed dimension. Its skip connection path maps the first residual response vector to the second compressed dimension. After passing through the second linear rectified activation layer, the second residual response vector is generated, and the dimension is increased again.
[0029] The third residual processing unit processes the second residual response vector using similar logic to generate the third residual response vector. Simultaneously, to preserve spatial structural information at different abstraction levels, the first residual processing unit performs spatial folding based on the ratio of the dimension of the first residual response vector to the dimension of the visual bottom-level statistical descriptive vector, generating the first semantic response map; the second residual processing unit performs spatial folding based on the ratio of the dimension of the second residual response vector to the dimension of the first residual response vector, generating the second semantic response map; and the third residual processing unit generates the third semantic response map. These three semantic response maps together constitute a multi-level semantic response map, with the third semantic response map being the top-level semantic response map with the smallest spatial scale.
[0030] Step 123: Input the top-level semantic response map with the smallest spatial scale in the multi-level semantic response map into the nonlocal context modeling module of the visual semantic coding layer. Capture the long-range dependencies between spatial locations in the display content frame units through nonlocal operations to generate a context-enhanced semantic response map.
[0031] The nonlocal context modeling module first selects the third semantic response map from the multi-level semantic response maps as the top-level semantic response map, whose shape can be represented symbolically as ChnTop×SptH×SptW. The module pre-defines a learnable spatial coordinate encoding vector for each spatial location, with the dimension of this vector equal to ChnTop. The spatial coordinate encoding vector for each location is then element-wise added to the feature vector of the corresponding location in the top-level semantic response map to obtain the location-enhanced top-level semantic response map.
[0032] Next, channel linear mapping is performed on the position-enhanced top-level semantic response map using the first linear transformation weight matrix Wqry, the second linear transformation weight matrix Wkey, and the third linear transformation weight matrix Wval. The mapping operation is achieved by adjusting the combination of channel information, generating a query feature map Qmap, a key feature map Kmap, and a value feature map Vmap, respectively. The spatial dimensions of Qmap, Kmap, and Vmap (SptH×SptW) are flattened to SptNum, resulting in the query matrix Qmtx, the key matrix Kmtx, and the value matrix Vmtx, all with dimensions ChnTop×SptNum.
[0033] Then, the matrix product of the transpose of Qmtx and Kmtx is calculated, and the product is multiplied by a scaling factor SclFac, which is the reciprocal of the square root of the number of channels ChnTop, to obtain the original attention score matrix AtnRaw, with dimensions SptNum × SptNum. Flexible maximum normalization is performed on each column of AtnRaw to generate a normalized attention weight matrix AtnWgt. The matrix product of Vmtx and AtnWgt is calculated to generate the global context aggregation matrix CtxAgg, with dimensions ChnTop × SptNum. The spatial dimensions of CtxAgg are restored to SptH × SptW, and the channel dimensions are mapped back to ChnTop through the fourth linear transformation layer Wout to obtain the global context feature map CtxFeat.
[0034] Subsequently, a learnable gating scalar GamCtx is introduced. GamCtx × CtxFeat is then added element-wise to the top-level semantic response map to obtain the residual augmented feature map ResEnh, which is then subjected to layer normalization. The layer-normalized ResEnh is fed into multiple parallel branches, each with an independent set of linear transformation weight matrices. Within each branch, the aforementioned linear mapping, attention calculation, weighted aggregation, and residual fusion processes are repeated. The feature maps output from each branch are concatenated along the channel dimension and compressed to ChnTop via an output linear projection layer to generate a multi-head non-local context feature map MheadCtx.
[0035] Finally, global average pooling is performed along the spatial dimension on the top-level semantic response graph to obtain the channel statistics vector ChnStat. ChnStat is then passed sequentially through the first fully connected layer, the channel linear rectified activation layer, and the second fully connected layer to generate the channel attention weight vector ChnAtn. MheadCtx is then scaled channel-by-channel using ChnAtn to obtain the context-enhanced semantic response graph CtxEnh.
[0036] Step 124: Input the semantic response maps at all levels except the top-level semantic response map and the context-enhanced semantic response map into the feature pyramid aggregation module of the visual semantic coding layer. Perform multi-scale feature fusion on the semantic response maps at all levels through top-down paths and lateral connections to generate a multi-scale fused semantic feature map.
[0037] The feature pyramid aggregation module receives a first semantic response map, a second semantic response map, and a context-enhanced semantic response map. First, an upsampling operation is performed on the context-enhanced semantic response map. The upsampling factor Nsamp is determined by the ratio of the spatial dimensions of the context-enhanced semantic response map to the second semantic response map. Upsamping uses nearest-neighbor interpolation, scaling up the spatial height and spatial width of the context-enhanced semantic response map by a factor of Nsamp, resulting in the upsampled enhanced semantic response map.
[0038] Simultaneously, a pointwise convolution is performed on the second semantic response map through a first horizontal convolutional layer to align its channel count with the upsampled enhanced semantic response map. The aligned second semantic response map and the upsampled enhanced semantic response map are then added element-wise to obtain the first fused response map. Next, the first fused response map is upsampled again, with the upsampling factor Msamp determined by the ratio of the spatial dimensions of the first fused response map to the first semantic response map, generating an upsampled fused response map. A pointwise convolution is then performed on the first semantic response map through a second horizontal convolutional layer to align the channel count, and the aligned first semantic response map is then added element-wise to the upsampled fused response map to obtain the second fused response map.
[0039] Finally, a standard convolution operation with a kernel size of Ksze×Ksze is applied to the second fused response map to eliminate the grid-like aliasing artifacts generated during the upsampling process, resulting in a multi-scale fused semantic feature map. This feature map integrates cross-scale visual information from low-level color texture to high-level semantic concepts.
[0040] Step 125: Input the multi-scale fused semantic feature map into the semantic compression module of the visual semantic coding layer, and generate a visual semantic embedding vector with a set dimension through channel dimension weighting and spatial dimension global pooling compression.
[0041] The semantic compression module first performs global average pooling on the spatial dimension of the multi-scale fused semantic feature map, compressing the feature map of each channel into a spatial mean scalar to form the initial channel vector.
[0042] Simultaneously, global max pooling is performed on the spatial dimension of the multi-scale fused semantic feature map, compressing the feature map of each channel into a spatial maximum scalar, forming a maximum channel vector. The initial channel vector and the maximum channel vector are fed into a channel-gated branch consisting of a first compressed fully connected layer, a compressed linear rectified activation layer, a second compressed fully connected layer, and a compressed flexible maximum layer. The first compressed fully connected layer compresses the channel dimension to 1 / Rcmp of the original, and the second compressed fully connected layer restores the channel dimension to the original number of channels. The compressed flexible maximum layer outputs a channel-gated weight vector Wgate of the same length as the original number of channels. Wgate is multiplied channel-wise with the multi-scale fused semantic feature map to obtain a channel-weighted feature map.
[0043] Next, global average pooling is performed again on the channel-weighted feature map to scale with spatial dimensions, and then passed through an output fully connected layer. The number of output neurons in this fully connected layer is equal to the set dimension Demb, ultimately generating a visual semantic embedding vector. The visual semantic embedding vector represents the global visual semantics of the display content frame unit as a compact numerical sequence.
[0044] Step 126: Based on the visual semantic embedding vectors of all display content frame units, the vectors are processed sequentially through the temporal dependency coding unit, manifold coordinate generation unit, and manifold trajectory description unit of the temporal manifold embedding layer to obtain the visual semantic offset sequence and semantic manifold local curvature features, which are then combined into manifold evolution trajectory features.
[0045] This step completes manifold modeling based on the visual semantic embedding vector sequence, which includes three sub-processes: temporal dependency encoding, manifold coordinate generation, and manifold trajectory description.
[0046] Step 1261: Arrange the visual semantic embedding vectors of all display content frame units into a visual semantic embedding vector sequence according to the timestamp order of the display content frame sequence.
[0047] The visual semantic embedding vectors generated for each display content frame unit in step 125 are stored sequentially in a sequential container from the earliest to the latest timestamp of each display content frame unit, forming a sequence of visual semantic embedding vectors.
[0048] Step 1262: Input the visual semantic embedding vector sequence into the temporal dependency coding unit of the temporal manifold embedding layer, and perform temporal context modeling on the visual semantic embedding vector sequence through a gated temporal processing mechanism to generate a temporally enhanced visual semantic embedding vector sequence containing contextual information.
[0049] The timing-dependent coding unit is implemented using a bidirectional gated timing encoder, which includes a forward-gated timing subunit and a reverse-gated timing subunit. The gating structure of each subunit includes a reset gate and an update gate.
[0050] The forward-gated temporal subunit processes the visual semantic embedding vectors sequentially in ascending order of time steps. For time step t, the forward-gated temporal subunit uses the forward hidden state of the previous time step and the current visual semantic embedding vector, selectively forgetting some information from the previous forward hidden state through a reset gate, and controlling the fusion ratio between the previous forward hidden state and the current candidate hidden state through an update gate, thus generating the forward hidden state for the current time step. The reverse-gated temporal subunit performs symmetrical processing in descending order of time steps to generate the reverse hidden state. The forward and reverse hidden states of each time step are concatenated along the feature dimension to form the temporally enhanced visual semantic embedding vector for the current time step. The temporally enhanced visual semantic embedding vectors of all time steps constitute a sequence of temporally enhanced visual semantic embedding vectors.
[0051] Step 1263: Input the temporal augmented visual semantic embedding vector sequence into the manifold coordinate generation unit of the temporal manifold embedding layer, and calculate the manifold embedding coordinate sequence for the temporal augmented visual semantic embedding vector sequence in the temporal semantic space.
[0052] The manifold coordinate generation unit maintains a trainable manifold basis matrix, where the number of rows equals the dimension of the temporal augmented visual semantic embedding vectors, and the number of columns equals the dimension Ldim of the target manifold space. For each temporal augmented visual semantic embedding vector in the sequence, its inner product with each column of the manifold basis matrix is calculated. The resulting Ldim inner product values are combined into an Ldim-dimensional vector, which serves as the initial projection of that time step into the manifold space.
[0053] Subsequently, using the locally linear embedding assumption, the initial projection at each time step is locally regularized. For time step t, the initial projections of its Knnb neighboring time steps on the time axis are found, and a set of reconstruction weights is calculated such that the initial projection of the current time step can be linearly represented by the initial projections of its neighboring time steps with minimum error. The initial projections of each neighboring time step are weighted and summed using this set of reconstruction weights to obtain the manifold embedding coordinates of time step t. This process yields the manifold embedding coordinate sequence for all time steps.
[0054] Step 1264: In the manifold trajectory description unit of the temporal manifold embedding layer, the manifold embedding coordinates of adjacent time steps in the manifold embedding coordinate sequence are differentially divided to obtain the visual semantic offset sequence. The change direction and angle of multiple consecutive offsets in the visual semantic offset sequence are statistically analyzed to generate semantic manifold local curvature features.
[0055] For time steps t and t+1 in the manifold embedding coordinate sequence, the manifold embedding coordinates of time step t+1 are subtracted from the manifold embedding coordinates of time step t to obtain an Ldim-dimensional difference vector, named the instantaneous semantic offset vector. The instantaneous semantic offset vectors of all consecutive adjacent time step pairs are arranged in order to form the visual semantic offset sequence.
[0056] To measure the curvature of the manifold trajectory, for any three consecutive time steps t, t+1, and t+2, the instantaneous semantic offset vector Sft1 from time step t to t+1 and the instantaneous semantic offset vector Sft2 from time step t+1 to t+2 are first calculated using the method described above. Next, the magnitudes Len1 and Len2 of Sft1 and Sft2 are calculated. The inner product of Sft1 and Sft2 is calculated, and the inner product value is divided by Len1 × Len2 to obtain the cosine of the offset angle. This cosine of the offset angle is then used to obtain the offset direction rotation angle through the inverse cosine function. The offset direction rotation angle is calculated for all combinations of three adjacent time steps, and the sequence of these values is used as the local curvature feature of the semantic manifold.
[0057] Step 1265: Combine the visual semantic offset sequence and the semantic manifold local curvature features into manifold evolution trajectory features.
[0058] Wherein, the combination method comprises padding the position of the visual semantic offset sequence at the first time step with a zero vector to make its length consistent with the length of the visual semantic embedding vector sequence, then concatenating the padded visual semantic offset sequence and the local curvature features of the semantic manifold along the feature dimension, and combining them into a composite feature tensor under the premise of ensuring time step alignment, which is used as the manifold evolution trajectory feature.
[0059] Step 130: performing topology-preserving end-to-end mapping processing on the manifold evolution trajectory features based on a backlight topology mapping network to generate multi-dimensional backlight adjustment curve parameters, wherein the backlight topology mapping network learns the continuous mapping relationship from the manifold evolution trajectory feature space to the backlight partition adjustment parameter space during the training stage.
[0060] The backlight topology mapping network receives the manifold evolution trajectory features generated in step 120 as input, and performs an end-to-end forward calculation. The internal structure of the network enables the mapping process to maintain the local neighborhood relationship and global distribution structure of the manifold trajectory, that is, two adjacent trajectory points on the visual semantic manifold are also as adjacent as possible in the mapping result in the backlight adjustment parameter space, and vice versa. This topology-preserving property ensures that the gradual and abrupt changes of the display content semantics can be consistently reflected in the backlight adjustment curve parameters. The network outputs multi-dimensional backlight adjustment curve parameters, which are organized according to backlight partitions. Each backlight partition corresponds to a set of brightness time-series curve parameters and a set of color temperature time-series curve parameters, which determine the continuous curve shape of the luminous state of the partition changing with time in the subsequent time period.
[0061] Step 131: slicing the manifold evolution trajectory features according to a preset time window length to obtain a plurality of partially overlapping manifold trajectory fragment features in time sequence.
[0062] The preset time window length WinLen determines the number of time steps contained in each slice. The manifold evolution trajectory feature is taken as the input feature tensor TfIn, and its total length in the time dimension is TotStep. Starting from the position where the time step index is 0, a sub-tensor with a length of WinLen is intercepted as the first manifold trajectory fragment feature; then, the window is slid backward with a step size Stride, and Stride<WinLen to ensure that there is an overlapping area between adjacent fragments, and the next fragment is intercepted. This traverses the entire time dimension until the window covers the end of the total length, and Nfrag manifold trajectory fragment features are obtained in total. The setting of the overlapping area enables the semantic manifold information at the time boundary to be correlated and expressed in two adjacent fragments.
[0063] Step 132: inputting the manifold trajectory fragment features into the time decomposition processing layer of the backlight topology mapping network, performing time scale transformation and channel splitting operation on each manifold trajectory fragment feature, and generating a trajectory detail decomposition component and a trajectory macro trend component.
[0064] The temporal decomposition processing layer has two sets of parallel temporal convolution channels built in.
[0065] The first temporal convolution channel uses a one-dimensional causal convolution unit with a kernel size of Kdet×1. The kernel size is small and is used to capture fluctuation details in a short time range. This convolution unit performs convolution scanning on the input manifold trajectory segment features along the time dimension and outputs trajectory detail decomposition components.
[0066] The second temporal convolution channel uses a one-dimensional causal convolution unit with a kernel size of Ktrd×1, and is combined with a moving average pooling operation. The kernel size Ktrd of this convolution branch is larger than Kdet, and it is combined with a pooling operation with a larger stride to filter high-frequency fluctuations and output the macro trend component of the trajectory.
[0067] The trajectory detail decomposition component preserves the high-frequency semantic changes in the local area of the manifold trajectory, while the trajectory macro trend component reflects the overall evolution direction of the semantic manifold within the segment.
[0068] Step 133: The trajectory detail decomposition components are fed into the local topology preservation unit of the backlight topology mapping network. By calculating the local neighborhood order-preserving mapping relationship of the trajectory detail decomposition components in the topology space, a local topology preservation descriptor is generated.
[0069] The local topology-preserving unit is constructed based on the principle of local linear embedding of high-dimensional data. For the trajectory detail decomposition component DetCmp, it contains feature vectors from WinLen time steps in the time dimension. For each target time step, Kloc nearest neighbor time steps are selected from its WinLen time steps based on the feature Euclidean distance. A set of reconstruction weights Wrec is calculated, which satisfies that the sum of all Kloc weight values is 1, and minimizes the sum of squared reconstruction errors between the feature vector of the target time step and the weighted linear combination of the feature vectors of the Kloc nearest neighbor time steps. This calculation process can be completed by solving the local covariance matrix and combining it with the Lagrange multiplier method. The element value of the covariance matrix is the inner product of the differences of the feature vectors between the Kloc nearest neighbors.
[0070] After obtaining Kloc reconstruction weights, they are arranged in order, and zeros are padded at the remaining time steps to form the local neighborhood order-preserving mapping relationship for the target time step. The local neighborhood order-preserving mapping relationships of all time steps are arranged sequentially to generate the local topology-preserving descriptor.
[0071] Step 134: The trajectory macro trend component is fed into the global topology alignment unit of the backlight topology mapping network to perform global structure alignment mapping on the trajectory macro trend component and generate a global topology descriptor.
[0072] The global topology alignment unit is implemented through a self-organizing mapping mechanism. Internally, each unit maintains a two-dimensional global topology node mesh, where each node is a weight vector with the same dimension as the macro-trend component feature of the trajectory. After inputting the macro-trend component McrCmp, for each time step feature vector in McrCmp, its feature distance to all nodes in the global topology node mesh is calculated, and the node with the smallest distance is selected as the best matching unit.
[0073] Then, based on a preset neighborhood function, the weight vectors of the best matching unit and its nodes within its topological neighborhood are updated, with the update magnitude gradually decreasing as the time step and iteration count increase. After traversing all time steps in McrCmp, for the feature vector of each time step, its best matching unit is found again. The row and column indices of this best matching unit in the two-dimensional grid are converted into a fixed-length position encoding vector, which is then used as the global topological descriptor. The global topological descriptors of each time step constitute a sequence of global topological descriptors of the same length as McrCmp.
[0074] Step 135: The local topology preservation descriptor and the global topology descriptor are fused and spliced through the feature fusion layer of the backlight topology mapping network to generate a topology fusion feature vector.
[0075] The feature fusion layer performs a concatenation operation along the feature dimension on the local topology preservation descriptor LocDesc and the global topology descriptor GlbDesc at each time step, connecting the two end to end to form a longer vector, called the topology fusion feature vector. This concatenation operation preserves the independence of local and global topology information, enabling subsequent mapping layers to utilize the two topological cues in a discriminative manner.
[0076] Step 136: Input the topology fusion feature vector into the first mapping hidden layer of the backlight topology mapping network. The first mapping hidden layer performs linear transformation and nonlinear activation processing on the topology fusion feature vector to generate the first mapping hidden layer vector.
[0077] The first mapping hidden layer consists of a linear transformation part composed of the first hidden layer fully connected weight matrix Wmp1 and the first hidden layer bias vector Bmp1, followed by a linear unit activation function with leakage correction. The topology fusion feature vector ValTopo is matrix multiplied with Wmp1 and accumulated with Bmp1 to obtain a linear response. Then, a linear unit with leakage correction is applied to each element of the linear response. When the element value is greater than 0, it retains its original value. When the element value is less than or equal to 0, it is multiplied by a preset small slope coefficient NegSlp to generate the first mapping hidden layer vector Hid1. This layer maps the topology fusion features to a hidden representation space with a larger capacity.
[0078] Step 137: Input the first mapping hidden layer vector into the second mapping hidden layer of the backlight topology mapping network. The second mapping hidden layer performs dimensional expansion processing on the first mapping hidden layer vector to generate the second mapping hidden layer vector.
[0079] The second mapped hidden layer contains a fully connected weight matrix Wmp2 and a bias vector Bmp2, followed by a linear unit activation function with leakage correction. The first mapped hidden layer vector Hid1 is matrix-multiplied with Wmp2 and then summed with Bmp2, followed by nonlinear activation to generate the second mapped hidden layer vector Hid2. Dimensional expansion ensures the features have sufficient capacity to represent the rich control information of all backlight zones at different times.
[0080] Step 138: Input the second mapping hidden layer vector into the output mapping layer of the backlight topology mapping network. The output mapping layer reshapes the second mapping hidden layer vector according to the number of backlight partitions and the parameter type of the control curve, generating the brightness time series curve parameters and color temperature time series curve parameters corresponding to each backlight partition.
[0081] Let the total number of backlight zones be NumZone. The brightness time series curve of each zone is described by NumBr control point parameters, and the color temperature time series curve of each zone is described by NumCt control point parameters. The output mapping layer first uses the output fully connected weight matrix Wout to project the second mapping hidden layer vector Hid2 onto a vector of length NumZone × (NumBr + NumCt).
[0082] Subsequently, the vector is reshaped into a matrix of shape NumZone × (NumBr + NumCt), where each row of the matrix corresponds to a backlight zone. The first NumBr values in each row represent the luminance time-series curve parameters for that zone, and the last NumCt values represent the color temperature time-series curve parameters for that zone. Both the luminance and color temperature time-series curve parameters are control point sequences, and the curve segment between every two consecutive control points is defined using cubic spline interpolation.
[0083] Step 139: Organize the brightness time series curve parameters and color temperature time series curve parameters corresponding to all backlight zones according to the zone identifier to obtain multi-dimensional backlight control curve parameters.
[0084] Assign a unique partition identifier to each backlight partition. Associate each row of the matrix obtained in step 138 with the corresponding partition identifier and construct a key-value pair set. The key is the partition identifier, and the value is a combination of the brightness time series curve parameters and the color temperature time series curve parameters corresponding to that partition. This key-value pair set is the multi-dimensional backlight control curve parameter, which can be directly indexed and called in subsequent steps.
[0085] Step 140: Perform backlight partition signal expansion processing based on the multidimensional backlight control curve parameters to obtain the backlight partition drive signal sequence.
[0086] After obtaining the multi-dimensional backlight control curve parameters, the backlight zone signal expansion processing is responsible for discretizing the continuous curve parameters into a timing electrical signal representation that can be directly executed by the backlight driver. This step analyzes the control curve of each backlight zone, interpolates on the time axis with a step size matching the display frame rate, generates a discretized driving value time series, and performs smooth transition processing on the signals across zones, finally obtaining a multi-channel timing signal that can be directly used to drive the light-emitting array. This process can be achieved through steps 141 to 144.
[0087] Step 141: Analyze the multi-dimensional backlight control curve parameters and extract the brightness time series curve parameters and color temperature time series curve parameters that are associated one-to-one with each backlight zone identifier.
[0088] Iterate through each key-value pair in the multi-dimensional backlight control curve parameters, record the partition identifier, and extract the corresponding brightness time series curve parameter BrCrv and color temperature time series curve parameter CtCrv. BrCrv is a set of control point sequences of length NumBr, and CtCrv is a set of control point sequences of length NumCt.
[0089] Step 142: For each backlight zone, perform time-domain discretization interpolation on its brightness time series curve parameters to generate a brightness driving value time series. The time step of the time-domain discretization interpolation is synchronized with the timestamp interval of the display content frame sequence.
[0090] Taking the backlight zone (Zonek) as an example, obtain the brightness time series curve parameter BrCrv corresponding to this zone. The timestamp interval of the displayed content frame sequence is denoted as DeltaT. Starting from the time point tstart and using DeltaT as the step size, generate a series of sampling time points. For each sampling time point tq, perform cubic spline interpolation using three consecutive control points in BrCrv to calculate the brightness interpolation result at tq.
[0091] Since cubic spline interpolation requires the time interval between adjacent control points as a parameter, each interpolation selects an interval covering tq and its neighboring control points. The calculation process uses natural boundary conditions, i.e., the second derivative at both ends of the curve is zero. The brightness interpolation results calculated from all sampling time points are arranged in chronological order to form a brightness driving value time series. The time-domain discretization interpolation processing of the color temperature time series curve parameters is performed in the same way, with the sampling step size also being DeltaT, generating a color temperature driving value time series.
[0092] Step 143: Combine the brightness drive value and color temperature drive value belonging to the same backlight zone and the same time step into a drive value tuple, and attach the corresponding time step index and backlight zone identifier to the drive value tuple. According to the order of the time step index, the drive value tuples of all time steps in the same backlight zone are connected in series to generate a backlight zone single-channel drive signal sequence.
[0093] For the backlight zone at time step tq, the corresponding luminance value BrVal and color temperature value CtVal are extracted from the luminance driving value time series and the color temperature driving value time series, respectively, and combined into a driving value tuple. This driving value tuple is then appended with a time step index and a backlight zone identifier to form a labeled driving value tuple. Subsequently, tq is used to sequentially sample all time points, and the resulting labeled driving value tuples are concatenated in ascending order of time step index to form the single-channel driving signal sequence for the backlight zone.
[0094] Step 144: Arrange the single-channel drive signal sequences of all backlight zones in parallel to obtain the multi-channel drive signal sequences of the backlight zones. Perform signal smoothing transition processing on the multi-channel drive signal sequences of the backlight zones. Eliminate the step abrupt changes in the signal by gradually connecting the drive value tuples between adjacent time steps to obtain the smooth drive signal sequence of the backlight zones. Output the smooth drive signal sequence of the backlight zones as the backlight zone drive signal sequence.
[0095] The NumZone backlight zone single-channel drive signal sequences are arranged in order according to the zone identifier to form a multi-channel data structure. Signal smoothing processing is performed independently for each channel: for the luminance value BrVal at time step tq of that channel, the luminance values of the SmtLen time steps before and after tq are taken, and their weighted moving average is calculated. The weights are distributed using a Gaussian weight distribution centered at tq. The original luminance value is replaced with this weighted moving average, and the color temperature value is processed in the same way. After traversing all time steps of all channels, step-like abrupt changes are suppressed, and a smooth backlight zone drive signal sequence with a gradual transition is formed, which is output as the final backlight zone drive signal sequence.
[0096] Step 150: Synchronize and align the backlight partition drive signal sequence with the display content frame sequence according to the timestamp to generate a backlight synchronization control command sequence, and output the backlight synchronization control command sequence to the backlight driver to control the backlight partition to perform dynamic light emission adjustment that matches the narrative rhythm of the display content and the shift of visual focus.
[0097] Precisely pairing the backlight partition drive signal sequence with the display content frame sequence at the timestamp level is the foundation for ensuring close synchronization between backlight changes and screen switching. This step generates structured backlight synchronization control instructions, which are sent to the backlight driver for execution after conflict detection and transition processing. This enables the backlight emission distribution to follow the narrative rhythm of the displayed content and the visual focus migration path in real time. Furthermore, it can dynamically reconstruct the visual immersion boundary according to the scene's mood and the characteristics of the interface logic flow. The implementation mechanism is further realized through steps 151 to 157.
[0098] Step 151: Obtain the timestamp of each display content frame unit in the display content frame sequence, extract the time step index corresponding to each drive value tuple in the backlight partition drive signal sequence, and establish a mapping table from time step index to timestamp.
[0099] Read the original timestamp TSeig of each display content frame unit from the display content frame sequence obtained in step 110 and store it as a timestamp list. Simultaneously, read the time step index attached to each drive value tuple from the backlight partition drive signal sequence output in step 144. Convert the time step index into a virtual timestamp with the same dimensions as the original timestamp using a linear mapping function, with the conversion formula being Tidx × DeltaT + starting offset. Record each time step index and its corresponding virtual timestamp as a row in the mapping table.
[0100] Step 152: Based on the mapping table, associate each drive value tuple in the backlight partition drive signal sequence with the display content frame unit with the nearest timestamp, and generate a pairing sequence of display content frame units and backlight partition drive value tuples.
[0101] For each time step index in the mapping table, the timestamp list is queried using its virtual timestamp to find the display content frame unit with the smallest absolute difference from that virtual timestamp. The frame identifier of that display content frame unit and the driving value tuple corresponding to that time step index are packaged into a pair record. After all time step indices have been processed, the pairing sequence is obtained.
[0102] Step 153: Perform timing conflict detection on the paired sequence. When the same display content frame unit is associated with multiple driving value tuples at different time steps, select the preferred driving value tuple from the multiple driving value tuples at different time steps according to the principle of minimum time interval, and discard the remaining driving value tuples.
[0103] Traverse the pairing sequence and construct a hash map with the frame identifier as the key and the list of driving value tuples as the value. If the list length is greater than 1, it indicates that the frame is associated with multiple driving value tuples, resulting in a timing conflict. Calculate the absolute value of the difference between the virtual timestamp and the frame timestamp for each driving value tuple in the list, select the driving value tuple with the smallest difference as the first choice, and discard the remaining tuples. Update the pairing record for this frame.
[0104] Step 154: Construct backlight synchronization control instructions using the paired sequences after timing conflict detection. The backlight synchronization control instructions include instruction frame timestamps, a set of target backlight zone identifiers, and brightness adjustment values and color temperature adjustment values corresponding to each backlight zone.
[0105] For each display content frame unit, a list of driving value tuples paired with it is collected. Since a frame may be associated with multiple driving value tuples from different backlight zones, the backlight zone identifiers corresponding to these driving value tuples are extracted to form a target backlight zone identifier set. Based on the zone identifiers, the corresponding brightness adjustment values and color temperature adjustment values are parsed from the driving value tuples. The frame timestamp, the target backlight zone identifier set, and the brightness adjustment values and color temperature adjustment values of each zone are encapsulated into a backlight synchronization control instruction object.
[0106] Step 155: Organize multiple consecutive backlight synchronization control commands into a backlight synchronization control command sequence according to timestamp order, and insert transition command frames at the beginning and end of the backlight synchronization control command sequence.
[0107] All backlight synchronization control commands are sorted in ascending order of their frame timestamps to form an ordered command sequence. An initial transition command frame is inserted before the first command in the sequence, setting both its brightness and color temperature adjustment values to predefined initial backlight state values. A termination transition command frame is inserted after the last command in the sequence, setting both its brightness and color temperature adjustment values to predefined cutoff backlight state values. The initial and cutoff backlight state values can be the minimum backlight value or a neutral color temperature value set by the hardware. After inserting the transition command frames, a complete backlight synchronization control command sequence is formed.
[0108] Step 156: The backlight synchronization control command sequence is transmitted to the backlight driver. The backlight driver parses the brightness adjustment value and color temperature adjustment value in the backlight synchronization control command sequence and converts the brightness adjustment value and color temperature adjustment value into driving electrical signals for the backlight zones. After receiving the driving electrical signals, the backlight zones adjust the luminous brightness and luminous color synchronously according to the brightness adjustment value and color temperature adjustment value, so that the backlight luminous distribution changes dynamically in real time according to the narrative rhythm of the screen switching of the displayed content and the visual focus migration path.
[0109] The backlight synchronization control command sequence generated in step 155 is sent to the backlight driver through the data transmission interface. The backlight driver has a built-in command parsing unit that reads the commands one by one, parses out the frame timestamps as the synchronization time, and reads out the target backlight zone identifier set and the corresponding brightness adjustment value and color temperature adjustment value.
[0110] The analog-to-digital converter maps the brightness adjustment value to the duty cycle of a pulse width modulation signal and the color temperature adjustment value to the current ratio of different color emitting channels, generating a driving electrical signal. This driving signal acts on the emitting elements of each backlight zone. The emitting elements determine their luminous intensity according to the brightness adjustment value and adjust the mixing ratio of each color channel according to the color temperature adjustment value to present the corresponding luminous color. As the instruction sequence continues, the backlight luminous distribution changes in real time, the spatial distribution of luminous intensity moves with the shifting visual focus of the displayed content, and the luminous color temperature switches according to the changing mood and tone of the scene.
[0111] Step 157: Based on the scene emotion flow characteristics and interface logic flow characteristics of the display content frame unit, dynamically adjust the luminous intensity attenuation slope of the backlight partition within the preset visual immersion boundary area to achieve visual immersion boundary reconstruction.
[0112] During the execution of the instruction sequence, in order to enhance the visual immersion, the boundary area of the backlight partition needs to be dynamically reconstructed according to the scene emotion changes and interface logic changes of the displayed content, which is specifically achieved through steps 1571 to 1575.
[0113] Step 1571: Extract the semantic offset between adjacent display content frame units from the visual semantic embedding vector sequence, and construct the scene emotion transition characteristic curve based on the degree of directional consistency of the semantic offset. The scene emotion transition characteristic curve reflects the change pattern of the emotional tone of the display content over time.
[0114] For adjacent vectors in a visual semantic embedding vector sequence, the element-wise difference between the preceding and following vectors is calculated to obtain the semantic offset. Then, directional consistency is calculated for multiple consecutive semantic offsets, defining a metric mechanism: for every three consecutive semantic offsets, the sign of the inner product between any two is checked; if all inner products are positive, the segment is considered directionally consistent. Time periods with consistent directions are marked as high emotional stability intervals, and time periods with inconsistent directions are marked as emotional transition intervals. These intervals are connected along the time axis with line segments, and the slope at the transition point represents the intensity of emotional transition, forming a scene emotional transition characteristic curve.
[0115] Step 1572: Extract the interface element displacement field between adjacent display content frame units from the visual semantic embedding vector sequence, and construct the interface logic flow characteristic curve based on the spatial distribution density of the interface element displacement field. The interface logic flow characteristic curve reflects the evolution pattern of the operation focus and information layout of the display content.
[0116] An optical flow estimation algorithm is used to calculate the displacement vector at each pixel position between adjacent display content frame units, and all displacement vectors constitute the interface element displacement field. Threshold filtering is applied to the magnitude of the displacement field, removing vectors with magnitudes smaller than a preset silent threshold and retaining significant displacement vectors. The display area is divided into a uniform grid, and the number of significant displacement vectors falling into each grid unit is counted to generate a spatial distribution density map. Principal component analysis is performed on the spatial distribution density map to determine the principal direction and dispersion of the spatial distribution of displacement vectors. The distribution difference value of the spatial distribution density map between adjacent frames is calculated along the time axis, and the time series of the distribution difference value is used to construct the interface logic flow characteristic curve.
[0117] Step 1573: Input the scene emotion transition characteristic curve and the interface logic transition characteristic curve into the visual immersion boundary calculation unit. Analyze the transition slope between the peaks and troughs of the scene emotion transition characteristic curve and the density of inflection points of the interface logic transition characteristic curve through the visual immersion boundary calculation unit to generate a visual immersion boundary range adjustment signal.
[0118] The visual immersion boundary calculation unit employs a convolutional loop processing module with a gated fusion mechanism. For the scene emotion transition characteristic curve, its first-order difference is calculated to obtain a transition slope sequence; for the interface logic transition characteristic curve, the number of times it crosses a preset change threshold within a unit time window is counted to obtain a turning point density sequence. The transition slope sequence and the turning point density sequence are simultaneously input into the gated fusion module, which outputs a boundary expansion / contraction coefficient. This boundary expansion / contraction coefficient dynamically adjusts the width of the preset visual immersion boundary region, generating a visual immersion boundary range adjustment signal.
[0119] Step 1574: Determine the luminous intensity attenuation slope adjustment amount within the preset visual immersion boundary area based on the visual immersion boundary range adjustment signal. The luminous intensity attenuation slope adjustment amount determines the steepness of the brightness attenuation of the backlight partition from the core area of the display screen to the edge area.
[0120] A pre-defined baseline attenuation function is defined as the luminance value at the core area boundary (LumCore) minus the distance from the core area boundary (Dst) multiplied by the baseline attenuation coefficient. When the visual immersion boundary range adjustment signal indicates boundary expansion, the baseline attenuation coefficient is decreased to obtain the real-time attenuation coefficient; when the boundary contracts, the baseline attenuation coefficient is increased. This real-time attenuation coefficient multiplied by the distance yields the luminance attenuation slope adjustment, which directly affects the slope of the change in the luminance drive value of each backlight zone within the preset visual immersion boundary area.
[0121] Step 1575: Update the brightness drive value of the backlight partition in the preset visual immersion boundary area in the backlight partition drive signal sequence according to the luminous intensity attenuation slope adjustment amount, so that the brightness drive value of the backlight partition in the preset visual immersion boundary area is redistributed according to the updated luminous intensity attenuation slope, and re-synchronize and align the updated backlight partition drive signal sequence with the display content frame sequence to generate an updated backlight synchronization control command sequence, and output the updated backlight synchronization control command sequence to the backlight driver for visual immersion boundary reconstruction.
[0122] A list of backlight zone identifiers is identified whose geometric positions fall within a preset visual immersion boundary area. For each zone in this list, the brightness drive value is recalculated using an updated real-time attenuation coefficient based on its distance from the core area of the display screen, and the corresponding value in the original backlight zone drive signal sequence is replaced.
[0123] Next, the synchronization alignment and instruction construction process in step 150 is repeated to obtain an updated backlight synchronization control instruction sequence, which is then sent to the backlight driver. The backlight driver adjusts the luminous intensity of the boundary region according to the updated instructions, thereby reconstructing the visual immersion boundary.
[0124] Based on steps 110-150 above, this embodiment of the invention further includes an optimization process for correcting subsequent mapping parameters by analyzing the deviation between the actual luminescence state and the expected manifold trajectory, as detailed in steps 210-240: Step 210: Obtain the actual luminous intensity distribution sequence of the backlight partition within the preset time window, and extract the trajectory segment corresponding to the preset time window from the manifold evolution trajectory features generated by the temporal manifold embedding layer.
[0125] By using an array of photosensitive sensors installed within the backlight module, the actual luminance values of each backlight zone within a preset time window (TWinBack) are collected to form an actual luminance intensity distribution sequence. Simultaneously, based on the start and end times of TWinBack, sub-tensors corresponding to the time periods are extracted from the manifold evolution trajectory features generated in step 126, serving as trajectory segments.
[0126] Step 220: Input the actual luminous intensity distribution sequence and trajectory segment into a preset spatiotemporal deviation analysis network. The spatiotemporal deviation analysis network performs cross-modal alignment processing on the temporal evolution mode of the actual luminous intensity distribution sequence and the manifold evolution mode of the trajectory segment in the common feature subspace to generate a spatiotemporal deviation representation vector.
[0127] The spatiotemporal bias analysis network comprises two branches, one processing the actual luminescence intensity distribution sequence and the other the trajectory segment using an encoder structure, and the other engaging in cross-attention interaction within a common feature subspace. The outputs of the two branches are matched through a feature adaptation layer guided by contrastive loss, and finally, the pointwise difference vectors of the two feature sequences, aligned in the temporal dimension, are concatenated to form the spatiotemporal bias representation vector.
[0128] Step 230: Input the spatiotemporal deviation representation vector into the preset recursive constraint generation network. The spatiotemporal deviation representation vector is modeled on a temporal dependency through the temporal state encoding layer of the recursive constraint generation network to generate a deviation state context vector. The deviation state context vector is then mapped to the backlight partition control parameter correction constraint amount through the constraint response decoding layer of the recursive constraint generation network.
[0129] The temporal state encoding layer within the recursive constraint generation network employs gated recursive units to process the spatiotemporal deviation representation vector step by step, updating and outputting the hidden state sequence as the deviation state context vector. The constraint response decoding layer takes the deviation state context vector as input and maps it layer by layer through multiple fully connected layers to generate backlight partition control parameter correction constraint quantities whose dimensions are identical to the parameters of the multidimensional backlight control curve.
[0130] Step 240: Based on the backlight partition control parameter correction constraint, the parameters of the multidimensional backlight control curve are fine-tuned under topological constraints. This ensures that when the backlight topology mapping network performs the topology-preserving end-to-end mapping process in the future, the generated multidimensional backlight control curve parameters are constrained by the backlight partition control parameter correction constraint, so as to suppress the spatiotemporal semantic drift between the manifold evolution trajectory characteristics and the actual luminous intensity distribution sequence.
[0131] The backlight zoning control parameter correction constraint is used as a correction regularization term. This constraint cost term, defined by the correction constraint, is added to the objective function when calculating the loss at the output mapping layer of the backlight topology mapping network. In subsequent iterations or online fine-tuning phases, the optimization algorithm minimizes prediction errors while suppressing deviations from the correction constraint, thus reducing the semantic gap between the newly generated multidimensional backlight control curve parameters and the current actual backlight emission state.
[0132] As another embodiment, it further includes a pre-compensation mechanism based on the attenuation of the backlight adjustment effect to counteract the physical response lag generated by the backlight element during continuous dynamic adjustment, the specific steps of which are as follows: steps 310-350.
[0133] Step 310: Obtain the backlight zone control history record corresponding to the executed instruction frame in the backlight synchronization control instruction sequence, and extract the brightness adjustment value sequence and color temperature adjustment value sequence of each backlight zone in the backlight zone control history record.
[0134] The system reads the history of recently executed instruction frames from the instruction execution log module of the backlight driver. It iterates through each record, locks the backlight zone identifier, extracts the brightness adjustment value sequence and color temperature adjustment value sequence, and arranges them by timestamp.
[0135] Step 320: Input the brightness adjustment value sequence and the color temperature adjustment value sequence into the preset adjustment effect attenuation analysis network. The time difference coding layer in the adjustment effect attenuation analysis network performs time-series coding on the brightness adjustment difference sequence and the color temperature adjustment difference sequence of adjacent adjustment command steps to generate the adjustment effect attenuation feature stream.
[0136] The temporal differential coding layer first calculates the difference between the brightness adjustment values of adjacent time steps, forming a brightness adjustment difference sequence; it then calculates the difference between the color temperature adjustment values of adjacent time steps, forming a color temperature adjustment difference sequence. The two difference sequences are interleaved according to their time step positions and input into a one-dimensional convolutional neural network for temporal encoding. This network consists of stacked convolutional and pooling layers, and outputs a modulation effect attenuation feature stream of the same length as the input sequence.
[0137] Step 330: Input the modulation effect decay feature stream into the decay pattern extraction layer in the modulation effect decay analysis network, and capture the decay patterns of the modulation effect decay feature stream under different time receptive fields through multi-scale dilated convolution operation along the time dimension to generate multi-scale decay pattern feature maps.
[0138] The decay pattern extraction layer comprises multiple dilated convolutional branches arranged side-by-side, with the dilation rate of each branch increasing sequentially. For example, the dilation rates are set to Rate1, Rate2, and Rate3, corresponding to an increasing receptive field size. The moderating effect decay feature stream flows through all branches simultaneously, with each branch capturing the response decay change pattern at its corresponding time scale. The feature maps output from each branch are concatenated along the channel dimension to generate a multi-scale decay pattern feature map.
[0139] Step 340: Input the multi-scale attenuation mode feature map into the preset backlight dynamic response compensation network, and the backlight dynamic response compensation network maps the multi-scale attenuation mode feature map into the backlight driving signal pre-compensation increment for future time steps.
[0140] The backlight dynamic response compensation network adopts an encoder-decoder architecture: the encoder gradually compresses the spatial size of the multi-scale attenuation mode feature map and increases the number of channels through several layers of convolution to capture high-level attenuation features; the decoder gradually restores the time step dimension through transposed convolution, and introduces skip connections from the encoder to retain fine-grained information. Finally, it outputs a pre-compensation increment sequence with the same length as the future preset step size. Each element in this sequence is a tuple composed of the pre-compensation value of the luminance channel and the pre-compensation value of the color temperature channel.
[0141] Step 350: Based on the backlight drive signal pre-compensation increment, perform pre-correction on subsequent instructions in the backlight partition drive signal sequence that have not yet been sent to the backlight driver. Merge the backlight drive signal pre-compensation increment with the backlight partition drive signal at the corresponding time step to generate and output the compensated backlight partition drive signal sequence, so as to offset the response lag and sensitivity attenuation generated by the backlight partition during continuous dynamic adjustment.
[0142] For each future time step in the subsequent instructions that have not yet been sent, if a corresponding pre-compensation increment exists, the brightness adjustment value is added to the brightness compensation value in the pre-compensation increment, and the color temperature adjustment value is added to the color temperature compensation value in the pre-compensation increment. After the correction is completed, this part of the drive signal is reassembled into a compensated backlight zone drive signal sequence, replacing the original drive signal sequence, and continues to be output to the downstream backlight driver according to step 150. By injecting the compensation amount in advance, the actual emission curve of the subsequent backlight element will be closer to the ideal shape defined by the original control curve parameters.
[0143] In another implementation, to facilitate the reproduction of the embodiments of the present invention by those skilled in the art, the training process and parameter configuration of the deep learning neural network model, network layer, module and unit involved in the embodiments of the present invention are explained in detail below.
[0144] The training of the temporal semantic deconstruction network is independent of the backlight topology mapping network. It requires a large-scale video content sample set prepared beforehand, containing video clips of various scene types, different image complexities, and different narrative rhythms. Each video clip is divided into a continuous sequence of display content frames. The training objective is to enable the temporal semantic deconstruction network to generate manifold evolution trajectory features that accurately characterize the semantic content of the frame sequence and its manifold structure. Training employs a self-supervised learning paradigm, based on joint optimization of reconstruction loss and manifold continuity constraints.
[0145] Specifically, during training, each display content frame unit is parsed by a visual semantic encoding layer to obtain a sequence of visual semantic embedding vectors. Then, manifold embedding coordinates for a subset of time steps are randomly selected in the temporal manifold embedding layer. These coordinates are back-mapped to the visual semantic space using a multilayer perceptron decoder to reconstruct the corresponding frame's visual semantic embedding vector. The first loss term is the cosine similarity loss between the reconstructed vector and the original vector. The second loss term is the manifold smoothing loss, which requires that the variance of two adjacent steps in the manifold embedding coordinate sequence be linearly correlated with the variance of two adjacent steps in the corresponding visual semantic embedding vector sequence. The third loss term is the manifold local conformal loss, ensuring that the angular relationships in the visual semantic embedding vector space are preserved in the manifold embedding coordinate space. The three loss terms are weighted and summed according to set weights Scalpha, Sbeta, and Sgamma to form the total training loss for temporal semantic deconstruction.
[0146] The training process uses an adaptive moment estimation optimizer with an initial learning rate of LrInit, which decays to LrMin in each training epoch using a cosine annealing strategy. The batch size is set to BatchSz, and the training runs for a total of EpochNum epochs. Parameters within the visual semantic encoding layer are initialized using a truncated normal distribution with a mean of 0 and a standard deviation of StdInit. The manifold basis matrix of the temporal manifold embedding layer is randomly initialized using a Gaussian distribution with a variance of VarMfd. Batch normalization is applied after all fully connected and convolutional layers in this network.
[0147] II. The training of the backlight topology mapping network is conducted after the temporal semantic deconstruction network has been trained and its parameters are fixed. The training dataset consists of multiple labeled samples. Each labeled sample contains the manifold evolution trajectory features of a video segment output by the temporal semantic deconstruction network, as well as the ground truth values of the multidimensional backlight control curve parameters corresponding to that video segment, annotated by experts. The backlight partitions are divided into multi-row, multi-column matrices determined by the display area ratio. The brightness time-series curve parameters and color temperature time-series curve parameters of each partition are generated frame-by-frame by professional lighting technicians based on the emotional changes in the video content.
[0148] During training, manifold evolution trajectory features are input into the backlight topology mapping network, which outputs predicted multidimensional backlight control curve parameters. A composite loss function is constructed, which includes the smoothed average absolute error loss between the predicted parameters and the ground truth, and the topology-preserving regularization loss. The calculation process of the topology-preserving regularization loss is as follows: random pairs of time steps are sampled in the manifold evolution trajectory feature space. If the feature distance between two time steps is less than a preset manifold neighborhood threshold, the distance between the corresponding parameter vectors of the two time steps is calculated in the backlight control parameter space, and this distance is accumulated as a contribution to the loss value. At the same time, for pairs whose distance in the manifold space is greater than a preset manifold non-neighborhood threshold, if the corresponding distance in the parameter space is less than a certain boundary value, a penalty is also incurred.
[0149] Training employs an adaptive moment estimation optimizer with an initial learning rate of LrBtm, weight decay of WdBtm, and epochs of Btm. The kernel weights of the one-dimensional causal convolutional units in the temporal decomposition processing layer are initialized using random orthogonal matrices to preserve the geometric isometry of the initial transformation. The Kloc values in the local topology-preserving units are set to small integers, and the node grid size of the global topology-aligning units is determined empirically. The fully connected weights of the first mapping hidden layer, the second mapping hidden layer, and the output mapping layer are initialized using a truncated normal distribution.
[0150] III. The spatiotemporal bias analysis network and the recursive constraint generation network are jointly trained. The training data is constructed as follows: During the actual operation of the system, for each preset time window (TWinBack), the actual luminous intensity distribution sequence, the corresponding manifold evolution trajectory feature segment, and the optimal correction amount of the backlight control parameters in the next time period (as the optimization target) are recorded. This optimal correction amount is obtained by retrospectively analyzing the manual scoring grid of the actual luminous effect and the semantic matching degree of the content. Both branches of the spatiotemporal bias analysis network are based on a temporal convolutional network as the backbone. The input sequence is processed by multiple dilated convolutional modules to extract features, and then matched through a feature adaptation layer.
[0151] The training loss function includes the alignment loss between the actual luminescence intensity distribution sequence and the trajectory segment in the common feature subspace, which is the mean square error of the features at the corresponding time step. A smoothed average absolute error loss is calculated between the output of the recursively constrained generation network and the optimal correction amount.
[0152] During joint optimization, the total loss is a weighted sum of two parts, with weights Wst and Wrc, respectively. The optimizer used is a stochastic gradient descent optimizer with a Newtonian momentum term, with a learning rate of LrStRc, a momentum coefficient of MomStRc, and a training epoch of EpochStRc. In the temporal state encoding layer of the recursive constraint generator network, the hidden state dimension of the gated recursive units is determined by the dimension of the bias vector. The number of neurons in each layer of the constraint response decoding layer decreases layer by layer until it matches the dimension of the multidimensional backlight control curve parameters.
[0153] IV. The modulation effect attenuation analysis network and the backlight dynamic response compensation network are trained in an end-to-end pipeline. The training samples are segments extracted from the long-term operation log of the backlight driver. The input of each sample is a sequence of brightness adjustment values and a sequence of color temperature adjustment values for multiple consecutive time steps. The expected output is an attenuation reference label that measures the degree of response attenuation of the backlight element during that time period. The attenuation reference label is pre-calibrated through a step response experiment in which a drive signal of known intensity is sent to the backlight module and the actual luminous intensity is measured by a photosensitive sensor.
[0154] The temporal difference coding layer of the attenuation analysis network consists of stacked one-dimensional convolutional layers, with the kernel size increasing from Kenc1 to KencN along the depth direction and the number of channels gradually increasing. In the attenuation mode extraction layer, the number of dilated convolutional channels in parallel is ChnlDil, and the dilation rate increases geometrically. This network converts historical instructions into features reflecting the attenuation dynamics. The backlight dynamic response compensation network uses these features as input and outputs a pre-compensation increment sequence.
[0155] The matching loss between the simulated luminous intensity after pre-compensation and the corresponding historical values of the actual light sensor is calculated. This loss is defined as the normalized root mean square error of the two intensity sequences. The network is trained using an adaptive moment estimation optimizer with a learning rate of LrCmp and a weight decay coefficient of WdCmp. The number of training epochs is set according to the log data size, and early stopping is triggered when the validation loss no longer decreases after several consecutive epochs. The number of convolutional layers in the encoder of the backlight dynamic response compensation network is typically EncLyr, and the number of transposed convolutional layers in the decoder is symmetrical to that in the encoder. The number of channels is doubled during the encoding stage and halved during the decoding stage. Skip connections concatenate the outputs of each layer of the encoder with the corresponding layer of the decoder. The concatenated features are then combined along the channel dimension and input into subsequent layers.
[0156] It should be understood that, in the specific implementation of the embodiments of the present invention, those skilled in the art can combine the color science formulas commonly used in the field of display device driving with the CIE standard colorimetric methods to unify the dimensions and weighting coefficients of the luminance intensity values. The intensity values of the red, green, and blue channels can be weighted and summed using luminance weighting coefficients conforming to ITU-R BT.709 or BT.2020 recommendations to generate luminance intensity values linearly related to the photoelectric conversion characteristics of the display device. Normalization processing can refer to the gray-world assumption or perfect reflection assumption commonly used in color constancy algorithms to eliminate overall illumination shift and ensure that the color distribution histogram is comparable under different scene brightness levels.
[0157] Furthermore, those skilled in the art can also combine classic Farneback or Lucas-Kanade dense optical flow algorithms in the field of optical flow estimation to extract pixel-level displacement vector fields from the display content frame sequence. For constructing the displacement field of interface elements, a pyramid hierarchical and iterative optimization strategy can be introduced to calculate the motion vectors of adjacent frames, and noise displacements can be filtered out by setting a silence threshold. Principal component analysis can be implemented based on the eigenvalue decomposition of the displacement field covariance matrix to obtain the principal direction and dispersion of displacement, thereby quantifying the logical flow characteristics of the interface.
[0158] Furthermore, those skilled in the art can combine backlight module pulse width modulation driving and color channel calibration methods to convert brightness adjustment values and color temperature adjustment values into driving electrical signals. The brightness adjustment value can be converted into the duty cycle of the pulse width modulation signal using a linear or piecewise linear mapping table, and the color temperature adjustment value can be calculated as the current ratio of each primary color channel based on the blackbody trajectory coordinates of the corresponding color temperature on the CIE1931 chromaticity diagram. The data collected by the photosensitive sensor array can be cross-calibrated with a standard colorimeter to obtain the conversion coefficient from sensor readings to absolute brightness, supporting the accurate construction of the actual luminous intensity distribution sequence.
[0159] This invention constructs manifold evolution trajectory features from the display content frame sequence to the temporal semantic space. The manifold geometry fully characterizes the narrative rhythm and visual focus migration path of the display content. Then, the backlight topology mapping network maps these trajectory features end-to-end into multi-dimensional backlight control curve parameters while maintaining the local neighborhood and global distribution structure of the manifold. This ensures that the gradual changes and abrupt changes in the semantics of the display content are consistently reflected in the brightness and color temperature control of the backlight partitions. Finally, a backlight synchronization control command sequence that is synchronized with the narrative rhythm of the screen switching and the visual focus migration path is generated, realizing the continuous following and dynamic reconstruction of the spatiotemporal evolution of the semantics of the display content by the backlight emission distribution.
[0160] Please see Figure 2The figure is a schematic diagram of the basic structure of a display device backlight optimization server 200 provided in an embodiment of this application. The display device backlight optimization server 200 includes: a processor 201; a storage device 202 on which a computer program 2020 is stored; and a network interface 203 for providing network communication functions. When the computer program 2020 is executed by the processor 201, the processor 201 implements any of the deep learning-based display device backlight optimization methods described above.
[0161] Please see Figure 3 This application provides a functional block diagram of a display device backlight optimization device, which includes: The display content acquisition module is used to acquire the display content frame sequence corresponding to the display content to be optimized. The display content frame sequence is composed of multiple display content frame units arranged in chronological order. The temporal semantic processing module is used to input the display content frame sequence into a preset temporal semantic deconstruction network, perform semantic parsing on each display content frame unit through the visual semantic encoding layer in the temporal semantic deconstruction network, generate a visual semantic embedding vector sequence, and perform spatiotemporal manifold mapping processing through the temporal manifold embedding layer in the temporal semantic deconstruction network to obtain the manifold evolution trajectory features of the display content frame sequence in the temporal semantic space. An end-to-end mapping module is used to perform topology-preserving end-to-end mapping processing on the manifold evolution trajectory features based on a backlight topology mapping network to generate multidimensional backlight control curve parameters. The backlight topology mapping network learns a continuous mapping relationship from the manifold evolution trajectory feature space to the backlight partition control parameter space during the training phase. The backlight partition expansion module is used to perform backlight partition signal expansion processing according to the multi-dimensional backlight control curve parameters to obtain a backlight partition driving signal sequence. The control instruction generation module is used to synchronize and align the backlight partition drive signal sequence with the display content frame sequence according to the timestamp, generate a backlight synchronization control instruction sequence, and output the backlight synchronization control instruction sequence to the backlight driver to control the backlight partition to perform dynamic light emission adjustment that matches the narrative rhythm of the display content and the shift of visual focus.
[0162] Based on the above, a readable storage medium is provided, on which a program or instructions are stored, and when the program or instructions are executed by a processor, the steps of the above method are implemented.
[0163] Furthermore, it should be noted that this application also provides a computer program product, which may include a computer program that can be stored in a computer-readable storage medium. The processor of the display device backlight optimization server reads the computer program from the computer-readable storage medium, and the processor can execute the computer program, causing the display device backlight optimization server to perform the aforementioned... Figure 1 The methods described in the corresponding embodiments are already known, and therefore will not be repeated here. Furthermore, the beneficial effects of using the same method will also not be repeated. For technical details not disclosed in the computer program product embodiments related to this application, please refer to the description of the method embodiments of this application.
[0164] It should be noted that the various embodiments in this specification are described in a progressive manner, with each embodiment focusing on the differences from other embodiments. Similar or identical parts between embodiments can be referred to interchangeably. For the systems or apparatus disclosed in the embodiments, since they correspond to the methods disclosed in the embodiments, the descriptions are relatively simple, and relevant parts can be referred to the method section.
Claims
1. A backlight optimization method for display devices based on deep learning, characterized in that, include: Obtain the display content frame sequence corresponding to the display content to be optimized, wherein the display content frame sequence is composed of multiple display content frame units arranged in chronological order; The display content frame sequence is input into a preset temporal semantic deconstruction network. The visual semantic encoding layer in the temporal semantic deconstruction network performs semantic parsing on each display content frame unit to generate a visual semantic embedding vector sequence. The sequence is then processed by the temporal manifold embedding layer in the temporal semantic deconstruction network to obtain the manifold evolution trajectory features of the display content frame sequence in the temporal semantic space. The backlight topology mapping network performs topology-preserving end-to-end mapping processing on the manifold evolution trajectory features to generate multidimensional backlight control curve parameters. The backlight topology mapping network learns the continuous mapping relationship from the manifold evolution trajectory feature space to the backlight partition control parameter space during the training phase. The backlight partition signal is expanded and processed according to the parameters of the multidimensional backlight control curve to obtain the backlight partition driving signal sequence. The backlight partition drive signal sequence and the display content frame sequence are synchronized and aligned according to the timestamp to generate a backlight synchronization control command sequence, and the backlight synchronization control command sequence is output to the backlight driver to control the backlight partition to perform dynamic light emission adjustment that matches the narrative rhythm of the display content and the shift of visual focus; The end-to-end mapping process based on the backlight topology mapping network to perform topology-preserving mapping on the manifold evolution trajectory features generates multidimensional backlight control curve parameters, including: The manifold evolution trajectory features are sliced according to a preset time window length to obtain multiple temporally overlapping manifold trajectory segment features; The manifold trajectory segment features are input into the time decomposition processing layer of the backlight topology mapping network. Time scale transformation and channel splitting operations are performed on each manifold trajectory segment feature to generate trajectory detail decomposition components and trajectory macro trend components. The trajectory detail decomposition components are fed into the local topology preservation unit of the backlight topology mapping network. By calculating the local neighborhood order-preserving mapping relationship of the trajectory detail decomposition components in the topology space, a local topology preservation descriptor is generated. The trajectory macro trend component is fed into the global topology alignment unit of the backlight topology mapping network, and the trajectory macro trend component is subjected to global structure alignment mapping to generate a global topology descriptor. The feature fusion layer of the backlight topology mapping network fuses and splices the local topology preservation descriptor and the global topology descriptor to generate a topology fusion feature vector. The topology fusion feature vector is input into the first mapping hidden layer of the backlight topology mapping network. The first mapping hidden layer performs linear transformation and nonlinear activation processing on the topology fusion feature vector to generate the first mapping hidden layer vector. The first mapping hidden layer vector is input into the second mapping hidden layer of the backlight topology mapping network, and the second mapping hidden layer performs dimensional expansion processing on the first mapping hidden layer vector to generate the second mapping hidden layer vector. The second mapping hidden layer vector is input into the output mapping layer of the backlight topology mapping network. The output mapping layer reshapes the second mapping hidden layer vector according to the number of backlight partitions and the parameter type of the control curve, generating brightness time series curve parameters and color temperature time series curve parameters corresponding to each backlight partition. The brightness time series curve parameters and color temperature time series curve parameters corresponding to all backlight zones are organized according to the zone identifier to obtain the multidimensional backlight control curve parameters.
2. The method according to claim 1, characterized in that, The backlight partition signal expansion processing based on the multi-dimensional backlight control curve parameters to obtain the backlight partition drive signal sequence includes: The parameters of the multi-dimensional backlight control curve are analyzed, and the brightness time series curve parameters and color temperature time series curve parameters that are associated one-to-one with each backlight zone identifier are extracted. For each backlight zone, time-domain discretization interpolation is performed on its brightness time series curve parameters to generate a brightness driving value time series. The time step of the time-domain discretization interpolation is synchronized with the timestamp interval of the display content frame sequence. For each backlight zone, time-domain discretization interpolation is performed on its color temperature time series curve parameters to generate a color temperature driving value time series. The time step of the time-domain discretization interpolation is aligned with the time step of the brightness driving value time series. The brightness drive value and color temperature drive value belonging to the same backlight zone and the same time step are combined into a drive value tuple, and the corresponding time step index and backlight zone identifier are attached to the drive value tuple. According to the order of the time step index, the drive value tuples of all time steps in the same backlight zone are connected in series to generate a single-channel drive signal sequence for the backlight zone. The single-channel drive signal sequences of all backlight zones are arranged in parallel to obtain a multi-channel drive signal sequence of the backlight zones. The multi-channel drive signal sequence of the backlight zones is subjected to signal smoothing transition processing. The step abrupt change in the signal is eliminated by the gradual connection of the drive value tuples between adjacent time steps to obtain a smooth drive signal sequence of the backlight zones. The smooth drive signal sequence of the backlight zones is output as the drive signal sequence of the backlight zones.
3. The method according to claim 1, characterized in that, The step of synchronizing and aligning the backlight partition drive signal sequence with the display content frame sequence according to timestamps to generate a backlight synchronization control command sequence, and outputting the backlight synchronization control command sequence to the backlight driver to control the backlight partition to perform dynamic light emission adjustment that matches the narrative rhythm of the displayed content and the shift of visual focus, includes: Obtain the timestamp of each display content frame unit in the display content frame sequence, and extract the time step index corresponding to each drive value tuple in the backlight partition drive signal sequence, and establish a mapping table from time step index to timestamp; Based on the mapping table, each drive value tuple in the backlight partition drive signal sequence is associated with the display content frame unit with the nearest timestamp, generating a pairing sequence of display content frame units and backlight partition drive value tuples; Timing conflict detection is performed on the pairing sequence. When the same display content frame unit is associated with multiple driving value tuples at different time steps, the preferred driving value tuple is selected from the multiple driving value tuples at different time steps according to the principle of minimum time interval, and the remaining driving value tuples are discarded. Backlight synchronization control instructions are constructed using a pairing sequence after timing conflict detection. The backlight synchronization control instructions include instruction frame timestamps, a set of target backlight zone identifiers, and brightness adjustment values and color temperature adjustment values corresponding to each backlight zone. Multiple consecutive backlight synchronization control commands are organized into a backlight synchronization control command sequence according to timestamp order, and transition command frames are inserted at the beginning and end of the backlight synchronization control command sequence. The backlight synchronization control command sequence is transmitted to the backlight driver, which parses the brightness adjustment value and color temperature adjustment value in the backlight synchronization control command sequence and converts the brightness adjustment value and color temperature adjustment value into driving electrical signals for the backlight zones. After receiving the driving electrical signal in the backlight zone, the backlight brightness and color are adjusted synchronously according to the brightness adjustment value and color temperature adjustment value, so that the backlight distribution changes dynamically in real time to follow the narrative rhythm of the screen switching and the visual focus migration path of the displayed content. Based on the scene emotion flow characteristics and interface logic flow characteristics of the displayed content frame unit, the intensity attenuation slope of the backlight partition in the preset visual immersion boundary area is dynamically adjusted to achieve visual immersion boundary reconstruction.
4. The method according to claim 3, characterized in that, The method of dynamically adjusting the luminous intensity attenuation slope of the backlight partition within a preset visual immersion boundary area based on the scene emotion flow characteristics and interface logic flow characteristics of the displayed content frame unit to achieve visual immersion boundary reconstruction includes: The semantic offset between adjacent display content frame units is extracted from the visual semantic embedding vector sequence. A scene emotion transition characteristic curve is constructed based on the degree of directional consistency of the semantic offset. The scene emotion transition characteristic curve reflects the change pattern of the emotional tone of the display content over time. The interface element displacement field between adjacent display content frame units is extracted from the visual semantic embedding vector sequence. An interface logic flow characteristic curve is constructed based on the spatial distribution density of the interface element displacement field. The interface logic flow characteristic curve reflects the evolution pattern of the operation focus and information layout of the display content. The scene emotion transition characteristic curve and the interface logic transition characteristic curve are input into the visual immersion boundary calculation unit. The visual immersion boundary calculation unit analyzes the transition slope between the peaks and troughs of the scene emotion transition characteristic curve and the density of the turning points of the interface logic transition characteristic curve to generate a visual immersion boundary range adjustment signal. The luminous intensity attenuation slope adjustment amount within the preset visual immersion boundary area is determined based on the visual immersion boundary range adjustment signal. The luminous intensity attenuation slope adjustment amount determines the steepness of the brightness attenuation of the backlight partition from the core area of the display screen to the edge area. The brightness driving value of the backlight partition in the preset visual immersion boundary region in the backlight partition driving signal sequence is updated according to the light intensity attenuation slope adjustment amount, so that the brightness driving value of the backlight partition in the preset visual immersion boundary region is redistributed according to the updated light intensity attenuation slope. The updated backlight partition drive signal sequence is re-synchronized and aligned with the display content frame sequence to generate an updated backlight synchronization control command sequence, and the updated backlight synchronization control command sequence is output to the backlight driver for visual immersion boundary reconstruction.
5. The method according to claim 1, characterized in that, The step involves inputting the display content frame sequence into a preset temporal semantic deconstruction network, performing semantic parsing on each display content frame unit through the visual semantic encoding layer in the temporal semantic deconstruction network to generate a visual semantic embedding vector sequence, and then performing spatiotemporal manifold mapping processing through the temporal manifold embedding layer in the temporal semantic deconstruction network to obtain the manifold evolution trajectory features of the display content frame sequence in the temporal semantic space, including: Extract the pixel brightness statistical distribution information and pixel color statistical distribution information of each display content frame unit in the display content frame sequence, and generate the visual low-level statistical description vector of the display content frame unit; The visual low-level statistical description vector is input into the multi-level residual coding module of the visual semantic coding layer. The visual low-level statistical description vector is compressed at each level of spatial scale and expanded in semantic channel by multiple residual processing units connected in series in the multi-level residual coding module to generate a multi-level semantic response map. The top-level semantic response map with the smallest spatial scale in the multi-level semantic response map is input into the nonlocal context modeling module of the visual semantic coding layer. The long-range dependencies between spatial locations in the display content frame unit are captured through nonlocal operations to generate a context-enhanced semantic response map. The semantic response maps at all levels except the top-level semantic response map and the context-enhanced semantic response map are input into the feature pyramid aggregation module of the visual semantic coding layer. Multi-scale feature fusion is performed on the semantic response maps at all levels through top-down paths and lateral connections to generate a multi-scale fused semantic feature map. The multi-scale fused semantic feature map is input into the semantic compression module of the visual semantic coding layer. Through channel dimension weighting and spatial dimension global pooling compression, a visual semantic embedding vector with a set dimension is generated. Based on the visual semantic embedding vectors of all display content frame units, the vectors are processed sequentially through the temporal dependency coding unit, manifold coordinate generation unit, and manifold trajectory description unit of the temporal manifold embedding layer to obtain the visual semantic offset sequence and semantic manifold local curvature features, which are then combined into the manifold evolution trajectory features.
6. The method according to claim 5, characterized in that, The visual semantic embedding vector based on all display content frame units is processed sequentially through the temporal dependency coding unit, manifold coordinate generation unit, and manifold trajectory description unit of the temporal manifold embedding layer to obtain the visual semantic offset sequence and semantic manifold local curvature features, which are then combined into the manifold evolution trajectory features, including: According to the timestamp order of the displayed content frame sequence, the visual semantic embedding vectors of all displayed content frame units are arranged into a visual semantic embedding vector sequence; The visual semantic embedding vector sequence is input into the temporal dependent coding unit of the temporal manifold embedding layer. The temporal context modeling of the visual semantic embedding vector sequence is performed through the gated temporal processing mechanism to generate a temporally enhanced visual semantic embedding vector sequence containing contextual information. The temporal enhanced visual semantic embedding vector sequence is input into the manifold coordinate generation unit of the temporal manifold embedding layer, and the manifold embedding coordinate sequence is calculated for the temporal enhanced visual semantic embedding vector sequence in the temporal semantic space; In the manifold trajectory description unit of the temporal manifold embedding layer, the manifold embedding coordinates of adjacent time steps in the manifold embedding coordinate sequence are differentially analyzed to obtain the visual semantic offset sequence. The change direction and angle of multiple consecutive offsets in the visual semantic offset sequence are statistically analyzed to generate semantic manifold local curvature features. The visual semantic offset sequence and the semantic manifold local curvature features are combined to form the manifold evolution trajectory features.
7. The method according to claim 5 or 6, characterized in that, The step of inputting the top-level semantic response map with the smallest spatial scale from the multi-level semantic response map into the nonlocal context modeling module of the visual semantic coding layer, and capturing the long-range dependencies between spatial positions in the display content frame units through nonlocal operations to generate a context-enhanced semantic response map includes: The top-level semantic response map with the smallest spatial scale is determined from the multi-level semantic response map, and the top-level semantic response map has a preset channel dimension and spatial size; The top-level semantic response map is enhanced by position-aware embedding. A learnable spatial coordinate encoding vector is superimposed at each spatial location of the top-level semantic response map, so that the feature vector of each spatial location carries its absolute position information, thereby generating a position-enhanced top-level semantic response map. The location-enhanced top-level semantic response map is linearly mapped using multiple preset linear transformation weight matrices to obtain a query feature map, a key feature map, and a value feature map. The spatial dimensions of the query feature map, the key feature map, and the value feature map are flattened into two-dimensional matrix forms to obtain a query matrix, a key matrix, and a value matrix. Each column of the query matrix, the key matrix, and the value matrix corresponds to the feature representation of the original spatial location. Matrix multiplication is performed between the query matrix and the transpose of the key matrix, and then multiplied by a scaling factor to obtain the original attention score matrix, where each element of the original attention score matrix represents the feature affinity between corresponding spatial location pairs; Flexible maximum normalization is applied to each column of the original attention score matrix to obtain a normalized attention weight matrix. Each column of the normalized attention weight matrix describes the distribution of the dependence intensity of all spatial locations on the corresponding location in that column. Perform matrix multiplication between the value matrix and the normalized attention weight matrix to aggregate value features from all spatial locations, thereby obtaining a global context aggregation matrix; restore the global context aggregation matrix to the spatial dimensions of the original top-level semantic response map, and map the channel dimensions back to the original channel dimensions of the top-level semantic response map through a preset fourth linear transformation layer to obtain a global context feature map; The global context feature map and the top-level semantic response map are adaptively fused channel by channel. The proportion of global context information injection is controlled by a learnable gating scalar to generate a residual enhanced feature map. Layer normalization is then performed on the residual enhanced feature map to stabilize the feature distribution. The normalized residual enhancement feature map is input into the multi-head parallel processing branch. The linear mapping, attention calculation, weighted aggregation and residual fusion processes are repeatedly executed in different feature subspaces. The features obtained from each branch are concatenated along the channel dimension, and the concatenated features are compressed to the original channel dimension of the top-level semantic response map through the output linear projection layer to generate a multi-head non-local context feature map. Channel attention weight vectors are generated using the channel statistics information of the top-level semantic response map. The multi-head nonlocal context feature map is then scaled channel by channel to suppress redundant context responses and enhance the expression of effective long-range dependencies, resulting in a context-enhanced semantic response map.
8. A display device backlight optimization server, characterized in that, include: A processor; a storage device having a computer program stored thereon; a network interface for providing network communication functions; wherein when the computer program is executed by the processor, the processor implements the deep learning-based display device backlight optimization method as described in any one of claims 1-7.
9. A readable storage medium, characterized in that, The readable storage medium stores a program or instructions that, when executed by a processor, implement the deep learning-based display device backlight optimization method as described in any one of claims 1-7.
Citation Information
Patent Citations
Mini LED backlight module intelligent control system of distributed driving architecture and method of Mini LED backlight module intelligent control system
CN121260114A
HDR image encoding and decoding methods and devices
US20150201222A1