Ambient light synchronization method, apparatus, and equipment based on video content color analysis

By downsampling and pixel weight masking of video frames, combined with temporal filtering and perceptual mapping, the problems of high computational load, high latency and inaccurate synchronization in existing ambient light synchronization schemes are solved, achieving efficient, stable and precise control of lighting and image synchronization.

CN122093543APending Publication Date: 2026-05-26SHENZHEN CHOUMEI CULTURAL BROADCASTING CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-02-11
Publication Date
2026-05-26

Smart Images

  • Figure CN122093543A_ABST
    Figure CN122093543A_ABST
Patent Text Reader

Abstract

This application discloses an ambient light synchronization method, apparatus, and device based on video content color analysis, comprising: obtaining the current video frame and its display timestamp from the video stream; downsampling the video frame; generating a pixel weight mask for eliminating interference areas based on the downsampled video frame; extracting candidate primary colors of the downsampled video frame based on the downsampled video frame, the pixel weight mask, and a preset spatial weight matrix; performing temporal filtering and perceptual mapping on the candidate primary colors to obtain the luminaire-driven color; generating and sending an ambient light synchronization command to at least one target luminaire based on the luminaire-driven color and its target effective time, to instruct the target luminaire to synchronize its emission state to the luminaire-driven color at the target effective time. This application can achieve lightweight computation, interference-resistant primary color extraction, smooth flicker-free operation, color gamut adaptation, and precise time-series alignment for video-linked ambient light synchronization.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent lighting and video signal processing technology, and in particular to an ambient light synchronization method, apparatus and device based on video content color analysis. Background Technology

[0002] With the widespread adoption of smart lighting and audio-visual entertainment devices, ambient light synchronization based on video image color has become an important means to enhance the immersive experience of watching movies and playing games. Existing ambient light synchronization solutions typically perform global color statistics on the original video frames and then send the results directly to the lighting fixtures for execution. However, in practical applications, the following problems exist: First, directly using high-resolution raw frames for calculation leads to high processing volume and latency, making it difficult to meet real-time synchronization requirements. Second, video images often contain interfering areas such as subtitles, faces, and overexposed highlights. These areas are not the main colors of the image, and directly including them in the statistics will cause the extracted main color to deviate from the true atmosphere of the image, affecting the synchronization effect. Third, the extracted main color is not filtered, which can easily lead to problems such as hue jumps, flickering, and a grayish-white appearance in low-saturation images, resulting in a poor visual experience. Fourth, the color gamut of the video and the physical color gamut of the lighting fixtures usually do not match, and directly sending colors will result in undisplayed or color-distorted images. Fifth, there are processing delays, transmission delays, and clock asynchronies between the playback terminal, gateway, and lighting fixtures, which can easily lead to problems such as the content of the image and changes in lighting being out of sync and showing significant lag. Finally, the existing solution lacks a unified end-to-end timing control and instruction optimization mechanism, resulting in low synchronization accuracy and insufficient stability.

[0003] There is currently no effective solution to the above problems. Summary of the Invention

[0004] This application provides an ambient light synchronization method, apparatus, and device based on video content color analysis to solve the technical problems of existing ambient light synchronization schemes, such as large computational load, high latency, inaccurate main color due to interference, color flickering and abrupt changes, color gamut mismatch, and latency accumulation leading to asynchrony between light and image.

[0005] To address the aforementioned technical problems, in a first aspect, embodiments of this application provide an ambient light synchronization method based on video content color analysis, applied to a playback terminal, comprising: The current video frame and its display timestamp are obtained from the video stream, the video frame is downsampled, and a pixel weight mask for eliminating interference areas is generated based on the downsampled video frame. Based on the downsampled video frame, the pixel weight mask, and the preset spatial weight matrix, the candidate primary color of the downsampled video frame is extracted; The candidate primary colors are subjected to temporal filtering and perceptual mapping to obtain the lighting drive colors. Based on the luminaire driving color and the target effective time of the luminaire driving color, an ambient light synchronization command is generated and sent to at least one target luminaire to instruct the target luminaire to synchronize its luminous state to the luminaire driving color at the target effective time. The target effective time is calculated based on the display timestamp and the estimated end-to-end latency.

[0006] In some embodiments, generating a pixel weight mask for excluding interference regions based on the downsampled video frame includes: Detect at least two different types of interference regions in the downsampled video frame, the interference regions including: subtitle or text regions, face or skin color regions, and bright or overexposed regions; Based on the detection results of each interference region and the preset weight assignment strategy, a pixel weight sub-mask corresponding to each interference region is generated. The pixel weight sub-mask includes: subtitle or text mask, face or skin color mask and highlight or overexposure mask. Based on the generated pixel weight sub-masks, they are fused according to a preset fusion rule to generate a pixel weight mask for eliminating interference areas.

[0007] In some embodiments, generating a pixel weight sub-mask corresponding to each interference region based on the detection result of each interference region and a preset weight assignment strategy includes: The pixel values ​​corresponding to the detected subtitle or text region are set as the first weight value to indicate complete exclusion, and the remaining pixel values ​​are set as the second weight value to indicate normal participation in the candidate primary color extraction, so as to obtain the subtitle or text mask corresponding to the subtitle or text region. The pixel values ​​corresponding to the detected face or skin color region are set to a third weight value that represents a reduction in weight. The third weight value is between the first weight value and the second weight value. The remaining pixel values ​​are set to the second weight value to obtain the face or skin color mask corresponding to the face or skin color region. The first pixel value corresponding to the detected bright or overexposed area is set as the first weight value, and the second pixel value corresponding to the detected bright or overexposed area is set as the gradient weight value. The brightness value of the first pixel value is greater than the first brightness threshold and the saturation value is less than the saturation threshold. The brightness value of the second pixel value is between the first brightness threshold and the second brightness threshold. The gradient weight value is determined according to the first brightness threshold and the second brightness threshold, and varies between the first weight value and the second weight value.

[0008] In some embodiments, the step of fusing the generated pixel weight sub-masks according to a preset fusion rule to generate a pixel weight mask for excluding interference regions includes: For each pixel position in the downsampled video frame, obtain the mask value of all pixel weight sub-masks at each pixel position; Based on the acquired multiple mask values, a fused mask value is calculated according to the preset fusion rule, which is the minimum value rule; The pixel weight mask is generated based on the fusion mask value calculated for each pixel location.

[0009] In some embodiments, extracting candidate primary colors of the downsampled video frames based on the downsampled video frames, the pixel weight mask, and a preset spatial weight matrix includes: The downsampled video frame is converted to the first color space to obtain the hue value, saturation value and brightness value of each pixel in the downsampled video frame; Based on the pixel weight mask value of the pixel weight mask, the preset spatial weight matrix and the saturation value, calculate the contribution weight of each pixel position to the hue histogram; A weighted hue histogram is constructed using the hue values ​​and the contribution weights; Based on the weighted hue histogram, candidate hue, candidate saturation, and candidate brightness of the downsampled video frame are determined and used to form the candidate primary color.

[0010] In some embodiments, constructing a weighted hue histogram using the hue value and the contribution weight includes: The range of hue values ​​is divided into a preset number of bins, and the width of each bin is calculated. The bins are continuous and non-overlapping hue intervals. For each pixel position in the downsampled video frame, obtain the pixel weight mask value of the pixel weight mask at each pixel position. If the pixel weight mask value is greater than 0 and the saturation value is greater than a preset saturation threshold, then determine the corresponding bin index based on the hue value and the bin width. The contribution weights are accumulated into the histogram values ​​of the bins corresponding to the bin indexes to obtain the initial hue histogram. If the cumulative proportion of the peak bins in the initial hue histogram is greater than a preset proportion threshold, then the histogram values ​​of the peak bins are attenuated, and the histogram values ​​of the second peak bins are increased, to obtain the final weighted hue histogram.

[0011] In some embodiments, determining the candidate hue, candidate saturation, and candidate brightness of the downsampled video frame based on the weighted hue histogram includes: Peak detection is performed on the weighted hue histogram, and the hue corresponding to the peak with the highest percentage is determined as the main peak hue; Based on the binning position and binning width corresponding to the main peak color, the candidate hue of the downsampled video frame is determined; Based on the pixel weight mask value and the preset spatial weight matrix, the saturation and brightness values ​​of all pixels within the preset hue range are weighted and summed to determine candidate saturation and candidate brightness, respectively. The preset hue range is centered on the candidate hue.

[0012] In some embodiments, the method further includes: If the candidate saturation is less than a preset saturation threshold, the candidate primary color is corrected to a target neutral color determined based on the candidate brightness, and used as the semantically corrected candidate primary color. If the candidate saturation is not less than the preset saturation threshold, and the weighted hue histogram contains at least two peaks that satisfy the preset significance condition, and the difference in the proportion of any two peaks is less than the preset difference threshold, then: Obtain the hue of the stable color output from the previous frame as the reference hue; Calculate the shortest distance on the color wheel between the hue corresponding to each peak that meets the condition and the reference hue; The hue corresponding to the peak with the shortest distance to the reference hue is determined as the corrected candidate hue, and the corrected candidate hue is combined with the candidate saturation and the candidate brightness to form the semantically corrected candidate primary color.

[0013] In some embodiments, performing temporal filtering and perceptual mapping on the candidate primary color to obtain the lamp-driven color includes: The candidate primary color is subjected to temporal filtering to obtain a stable target color, wherein the stable target color includes stable hue, stable saturation and stable brightness. The stable target color is perceptually mapped from the video color gamut to the target luminaire color gamut to obtain the luminaire-driven color.

[0014] In some embodiments, performing temporal filtering on the candidate primary color to obtain a stable target color includes: Calculate the normalized color difference between the candidate primary color and the stable color of the previous video frame. The normalized color difference is calculated based on the shortest hue difference on the color wheel, the saturation difference, and the brightness difference. The smoothing coefficient of the exponential moving average filter is dynamically adjusted based on the normalized color difference and the motion intensity of the current video frame. Based on the adjusted smoothing coefficient, exponential sliding filtering is performed on the candidate hue, candidate saturation and candidate brightness respectively to obtain a stable target color; The smoothing coefficients of the dynamically adjusted exponential moving average filter include: If the normalized color difference exceeds a preset color difference threshold or the motion intensity exceeds a preset motion threshold, the smoothing coefficient is increased to accelerate the color response; otherwise, the basic filtering coefficient is used to maintain color stability.

[0015] In some embodiments, the perceptual mapping from the video color gamut to the target luminaire color gamut of the stable target color to obtain the luminaire-driven color includes: The stable target color is converted from the first color space to a device-independent second color space; Based on the physical luminous gamut of the target luminaire, the color coordinates in the second color space are clipped to obtain the chromaticity coordinates located within the physical luminous gamut. Based on the brightness of the stable target color and the global brightness statistics of the downsampled video frame, the mapped target brightness is calculated. The rate of change of the mapped target brightness is limited; The driving color of the luminaire is determined based on the chromaticity coordinates after gamut clipping and the target brightness after limiting the rate of change of brightness.

[0016] In some embodiments, the estimated end-to-end latency is calculated as follows: Based on the processing latency estimate, network latency estimate, luminaire response latency estimate, and buffer latency estimate, the full-link latency estimate is calculated. The network latency estimate is determined based on the measured value of the network round-trip latency for communication with the target luminaire. The buffer latency estimate is determined based on the historical statistical fluctuation of the network round-trip latency. The full-link latency estimate is the total estimated latency of the entire process from the acquisition of the video frame to the target luminaire performing illumination state synchronization. Accordingly, the target effective time is calculated based on the displayed timestamp and the estimated end-to-end latency, including: The target effective time of the lamp-driven color is obtained by adding the display timestamp and the estimated end-to-end delay.

[0017] In some embodiments, generating and sending the ambient light synchronization command includes: The perceived color difference is calculated by comparing the driving color of the lamp with the driving color of the last successfully sent lamp. Select the corresponding dynamic threshold or static threshold based on the motion intensity of the current video frame; If the perceived color difference is greater than the selected threshold, or if the maximum silent time has been exceeded since the last successful transmission, an ambient light synchronization command is generated. The ambient light synchronization command includes at least the luminaire driving color and the target effective time. Otherwise, the generation and transmission of the ambient light synchronization command are suppressed.

[0018] In some embodiments, generating and sending ambient light synchronization commands to at least one target luminaire includes: Based on the multiple different spatial partitions of the downsampled video frame, determine the primary color of each partition; Based on the preset mapping relationship between screen partitions and physical light positions, assign exclusive light-driving colors to different target lights based on the primary color of the corresponding partition; Based on the unique luminaire driver color of each target luminaire, a corresponding ambient light synchronization command is generated and sent.

[0019] Secondly, embodiments of this application provide an ambient light synchronization method based on video content color analysis, applied to a gateway, including: Receive an ambient light synchronization command from at least one playback terminal, the ambient light synchronization command including at least the lamp driving color and the target effective time of the lamp driving color; The received ambient light synchronization command is converted into a format compatible with the communication protocol supported by the target luminaire; The effective time of the target is calibrated based on a clock synchronization mechanism; The ambient light synchronization command, after format conversion and time calibration, is sent to the corresponding target luminaire.

[0020] Thirdly, embodiments of this application provide an ambient light synchronization method based on video content color analysis, applied to a target lighting fixture, including: Receive an ambient light synchronization command sent from a playback terminal or gateway, wherein the ambient light synchronization command includes at least the lamp driving color and the target effective time of the lamp driving color; When the local system time reaches the time window of the target effective time, the current luminous state is synchronized with the color driven by the lamp. The transition process of synchronizing the current luminous state with the color driven by the lamp is calculated by interpolation using the shortest path on the color wheel.

[0021] In some embodiments, the interpolation calculation using the shortest path on the color wheel includes: The current luminous state and the color driven by the lamp are converted to the first color space to obtain the current hue and the target hue, respectively. Calculate the difference between the target hue and the current hue, and adjust the difference to the shortest angular distance on the hue wheel; Based on the adjusted shortest angular distance, interpolation is performed within a preset transition time to complete the change from the current hue to the target hue.

[0022] In some embodiments, the method further includes: During the transition process or when the light is in steady state, the operating parameters of the target luminaire are monitored; If the operating parameters exceed the preset safety threshold, the light emission state of the target lamp will be automatically adjusted to a safety mode, which includes reducing brightness or switching to a preset safety color.

[0023] In some embodiments, the method further includes: If no new ambient light synchronization command is received within a preset command reception timeout period since the last valid ambient light synchronization command was received, the target luminaire will be controlled to slowly transition from its current luminous state to a preset steady-state safety state.

[0024] Fourthly, this application provides an ambient light synchronization system based on video content color analysis, including a playback terminal, a gateway, and at least one target lamp. The playback terminal is communicatively connected to the target lamp, and the gateway serves as a communication relay between the playback terminal and the target lamp. The playback terminal is configured to: obtain the current video frame and its display timestamp from the video stream; perform downsampling processing on the video frame; and generate a pixel weight mask for eliminating interference areas based on the downsampled video frame; extract candidate primary colors of the downsampled video frame based on the downsampled video frame, the pixel weight mask, and a preset spatial weight matrix; perform temporal filtering and perceptual mapping on the candidate primary colors to obtain the lamp-driven color; and generate and send an ambient light synchronization command to at least one target lamp based on the lamp-driven color and the target effective time of the lamp-driven color, to instruct the target lamp to synchronize its illumination state to the lamp-driven color at the target effective time, wherein the target effective time is calculated based on the display timestamp and the estimated end-to-end latency. The gateway is configured to: receive an ambient light synchronization instruction from at least one playback terminal, the ambient light synchronization instruction including at least the luminaire driving color and the target effective time of the luminaire driving color; convert the received ambient light synchronization instruction into a format compatible with the communication protocol supported by the target luminaire; calibrate the target effective time based on a clock synchronization mechanism; and send the format-converted and time-calibrated ambient light synchronization instruction to the corresponding target luminaire. The target luminaire is used to: receive an ambient light synchronization command sent from a playback terminal or gateway, the ambient light synchronization command including at least the luminaire driving color and the target effective time of the luminaire driving color; when the local system time reaches the time window of the target effective time, synchronize the current luminous state to the luminaire driving color, the transition process of synchronizing the current luminous state to the luminaire driving color is calculated by interpolation using the shortest path on the color wheel.

[0025] Fifthly, embodiments of this application provide an ambient light synchronization device based on video content color analysis, applied to a playback terminal, comprising: The frame processing and mask generation module is used to obtain the current video frame and its display timestamp from the video stream, perform downsampling processing on the video frame, and generate a pixel weight mask for eliminating interference areas based on the downsampled video frame. The candidate primary color extraction module is used to extract the candidate primary color of the downsampled video frame based on the downsampled video frame, the pixel weight mask, and the preset spatial weight matrix. The temporal filtering and mapping module is used to perform temporal filtering and perceptual mapping on the candidate primary color to obtain the lamp driving color. The instruction generation and sending module is used to generate and send an ambient light synchronization instruction to at least one target luminaire based on the luminaire driving color and the target effective time of the luminaire driving color, so as to instruct the target luminaire to synchronize its light emission state to the luminaire driving color at the target effective time. The target effective time is calculated based on the display timestamp and the estimated end-to-end latency.

[0026] Sixthly, embodiments of this application provide an ambient light synchronization device based on video content color analysis, applied to a gateway, comprising: The first receiving module is configured to receive an ambient light synchronization instruction from at least one playback terminal, wherein the ambient light synchronization instruction includes at least the lamp driving color and the target effective time of the lamp driving color. A conversion module is used to convert the received ambient light synchronization command into a format compatible with the communication protocol supported by the target luminaire; The calibration module is used to calibrate the effective time of the target based on a clock synchronization mechanism. The sending module is used to send the ambient light synchronization command, which has been format-converted and time-calibrated, to the corresponding target luminaire.

[0027] Seventhly, embodiments of this application provide an ambient light synchronization device based on video content color analysis, applied to a target lighting fixture, comprising: The second receiving module is used to receive an ambient light synchronization instruction sent from a playback terminal or gateway. The ambient light synchronization instruction includes at least the lamp driving color and the target effective time of the lamp driving color. The synchronization module is used to synchronize the current luminous state with the color driven by the lamp when the local system time reaches the time window where the target effective time is located. The transition process of synchronizing the current luminous state with the color driven by the lamp is calculated by interpolation using the shortest path on the color wheel.

[0028] Eighthly, embodiments of this application provide an electronic device, including: a memory and a processor, wherein the processor and the memory are communicatively connected to each other, the memory stores computer instructions, and the processor executes the computer instructions to implement the steps of the methods described in the first, second, and third aspects.

[0029] In a ninth aspect, embodiments of this application provide a computer-readable storage medium storing computer program instructions that, when executed by a processor, implement the steps of the methods described in the first, second, and third aspects.

[0030] This application provides an ambient light synchronization method, apparatus, and device based on video content color analysis. According to the above embodiments, by obtaining the current video frame and its display timestamp from the video stream, a picture timing reference can be provided for light synchronization. By downsampling the video frame, the amount of data and computational overhead of image processing can be significantly reduced, solving the problems of high computational load, high processing latency, and poor real-time performance caused by directly using high-resolution frames in existing solutions. By generating a pixel weight mask for eliminating interference areas based on the downsampled video frame, the influence of non-subject areas such as subtitles, faces, and overexposed highlights on color statistics can be suppressed, making the main color more in line with the real atmosphere of the picture, solving the problems of main color deviation, atmosphere distortion, and poor synchronization effect caused by interference areas. By extracting candidate main colors of the downsampled video frame based on the downsampled video frame, the pixel weight mask, and a preset spatial weight matrix, stable, accurate, and visually consistent main colors of the picture can be extracted with low computational load, combined with interference shielding and spatial attention, solving the problems of inaccurate main color extraction, susceptibility to noise / mixed colors, and inability to reflect the real atmosphere of the picture. By performing temporal filtering and perceptual mapping on candidate primary colors, temporal filtering can eliminate inter-frame color jumps and flickering, improving visual smoothness. Perceptual mapping can adapt video colors to colors that the lighting fixtures can render correctly, solving problems such as hue jumps, light flickering, poor visual experience, mismatch between video and lighting color gamuts, color bias, or inability to display. By calculating the target effective time based on the display timestamp and the estimated end-to-end latency, the overall latency of processing, transmission, and lighting fixture response can be compensated based on the screen display time, aligning the light triggering moment with the screen content timing. This solves the problems of asynchronous lighting and screen, significant lag, and timing errors caused by accumulated processing and transmission latency. By generating and sending ambient light synchronization commands based on the lighting fixture-driven color and its target effective time, precise color and precise timing can be bound and sent, achieving synchronization between lighting and video in content and time. This solves problems such as lack of unified timing control, low synchronization accuracy, insufficient stability, and poor immersion. Attached Figure Description

[0031] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments recorded in this application. For those skilled in the art, other drawings can be obtained from these drawings without creative effort.

[0032] Figure 1 This is a flowchart illustrating an ambient light synchronization method based on video content color analysis provided in an embodiment of this application. Figure 2This is a flowchart illustrating an ambient light synchronization method based on video content color analysis, provided in another embodiment of this application. Figure 3 This is a flowchart illustrating an ambient light synchronization method based on video content color analysis, provided in another embodiment of this application. Figure 4 This is a flowchart illustrating an ambient light synchronization method based on video content color analysis, provided in another embodiment of this application. Figure 5 This is a flowchart illustrating an ambient light synchronization method based on video content color analysis, provided in another embodiment of this application. Figure 6 This is a flowchart illustrating an ambient light synchronization method based on video content color analysis, provided in another embodiment of this application. Figure 7 This is a flowchart illustrating an ambient light synchronization method based on video content color analysis, provided in another embodiment of this application. Figure 8 This is a flowchart illustrating an ambient light synchronization method based on video content color analysis, provided in another embodiment of this application. Figure 9 This is a flowchart illustrating an ambient light synchronization method based on video content color analysis, provided in another embodiment of this application. Figure 10 This is a flowchart illustrating an ambient light synchronization method based on video content color analysis, provided in another embodiment of this application. Figure 11 This is a flowchart illustrating an ambient light synchronization method based on video content color analysis, provided in another embodiment of this application. Figure 12 This is a flowchart illustrating an ambient light synchronization method based on video content color analysis, provided in another embodiment of this application. Figure 13 This is a flowchart illustrating an ambient light synchronization method based on video content color analysis, provided in another embodiment of this application. Figure 14 This is a schematic diagram of the structural composition of an ambient light synchronization device based on video content color analysis provided in an embodiment of this application; Figure 15 This is a schematic diagram of the structural composition of an ambient light synchronization device based on video content color analysis provided in another embodiment of this application; Figure 16 This is a schematic diagram of the structural composition of an ambient light synchronization device based on video content color analysis provided in another embodiment of this application; Figure 17 This is a schematic diagram of the structural composition of an electronic device provided in an embodiment of this application. Detailed Implementation

[0033] To enable those skilled in the art to better understand the technical solutions in this application, the technical solutions in the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only a part of the embodiments of this application, and not all of the embodiments. Based on the embodiments in this application, all other embodiments obtained by those skilled in the art without creative effort should fall within the scope of protection of this application.

[0034] It should be noted that the information and data related to users involved in the embodiments of this application are all information and data authorized by the user or fully authorized by the relevant parties. Furthermore, the collection, storage, use, processing, transmission, provision, disclosure, and application of the relevant data all comply with relevant laws, regulations, and standards, and necessary confidentiality measures have been taken. They do not violate public order and good morals, and corresponding operation entry points are provided for users or relevant parties to choose to authorize or refuse.

[0035] It should also be noted that in the embodiments of this application, certain software, components, models and other existing solutions in the industry may be mentioned. These should be regarded as exemplary and are only intended to illustrate the feasibility of implementing the technical solution of this application. However, it does not mean that the applicant has used or necessarily used the solution.

[0036] The ambient light synchronization method based on video content color analysis provided in the embodiments of this application will be described below with reference to the accompanying drawings.

[0037] Figure 1 This is a flowchart illustrating an ambient light synchronization method based on video content color analysis according to an embodiment of this application. Although this specification provides method operation steps or apparatus structures as shown in the following embodiments or figures, the method or apparatus may include more or fewer operation steps or module units through conventional or non-inventive means. In steps or structures where there is no logically necessary causal relationship, the execution order of these steps or the module structure of the apparatus is not limited to the execution order or module structure shown in the embodiments or figures of this specification. When the method or module structure is applied in actual devices, servers, or terminal products, it can be executed sequentially or in parallel according to the method or module structure shown in the embodiments or figures (e.g., in a parallel processor or multi-threaded processing environment, or even in a distributed processing or server cluster implementation environment). Figure 1 As shown, this method can be applied to a playback terminal and may include: S101: Obtain the current video frame and its display timestamp from the video stream, perform downsampling processing on the video frame, and generate a pixel weight mask for eliminating interference areas based on the downsampled video frame.

[0038] Specifically, the aforementioned playback terminal can refer to electronic devices with video decoding, playback, and network communication capabilities, including but not limited to smart TVs, set-top boxes, projectors, mobile phones, tablets, and computers. The aforementioned video stream can refer to encoded and compressed video data transmitted sequentially in time, originating from local files, live web streaming, or online on-demand playback. The aforementioned display timestamp can be used to identify the moment a video frame should be displayed on the screen, serving as the reference time for video frame synchronization with lighting. The aforementioned downsampling process can be the process of reducing the resolution of video frames by a fixed ratio (e.g., reducing the width and height to 1 / 2, 1 / 4, and 1 / 8 of their original values, respectively) while maintaining the overall color distribution of the image. Downsampling of video frames yields a low-resolution image. The aforementioned interference area can refer to areas in the video frame that do not belong to the main atmosphere of the image and are prone to distortion in the extraction of the primary color, including subtitle text areas, skin tone areas, and overexposed bright areas. The aforementioned pixel weight mask can be a two-dimensional weight matrix with the same size as the downsampled video frame, where the value (0-1) of each element represents the "weight" or "confidence" of the corresponding pixel in subsequent calculations.

[0039] Specifically, the video stream can first be decoded using the decoder in the playback terminal to generate raw video frames (high-resolution video frames F) in YUV or RGB format and placed into the decoder buffer. Then, the current video frame is obtained from the decoder buffer or the screen rendering pre-buffer (decoding post-buffer or rendering pipeline front-end), and the corresponding presentation time stamp (PTS), denoted as Tdisplay, is read synchronously. Afterward, the playback terminal downsamples the current high-resolution video frame F, using methods such as GPU texture downsampling (Mipmap L3 or L4 level) or CPU-side bilinear / nearest neighbor sampling, reducing the obtained high-resolution video frame F (e.g., 1920×1080) to a low-resolution video frame (e.g., 120×68) by a fixed ratio (e.g., 1 / 16), resulting in the downsampled video frame F_small. Afterwards, the playback terminal performs interference region detection on the downsampled video frame F_small, identifying at least two different types of interference regions, such as subtitle text regions, skin color regions, and overexposed highlight regions. Based on the type of each region, the terminal assigns a corresponding weight value to each pixel position, generating a pixel weight mask Final_Mask that is consistent with the size of the downsampled video frame.

[0040] In this embodiment, a unified timing benchmark between image content and lighting control is established by acquiring video frames and display timestamps. Downsampling significantly reduces the amount of data and computational overhead in image processing, effectively reducing processing latency and improving system real-time performance. By generating pixel weight masks, the influence of interfering areas such as subtitles, faces, and overexposed highlights can be eliminated, preventing invalid or abnormal colors from participating in the primary color statistics and ensuring that the subsequently extracted primary color accurately reflects the main atmosphere of the video image.

[0041] In some embodiments, see Figure 2 As shown, the pixel weight mask for eliminating interference regions generated based on the downsampled video frame in S101 above can, in specific implementation, include: S11: Detect at least two different types of interference regions in the downsampled video frame, the interference regions including: subtitle or text regions, face or skin color regions, and bright or overexposed regions; S12: Based on the detection results of each interference region and the preset weight assignment strategy, generate a pixel weight sub-mask corresponding to each interference region. The pixel weight sub-mask includes: subtitle or text mask, face or skin color mask and highlight or overexposure mask. S13: Based on the generated pixel weight sub-masks, they are fused according to a preset fusion rule to generate a pixel weight mask for eliminating interference areas.

[0042] Specifically, firstly, the playback terminal detects at least two preset types of interference regions in parallel. In a preferred embodiment of the present invention, the detection mainly targets three types of interference that have a significant negative impact on the extraction of the main color: (a) subtitle / text regions, which are usually located in a specific area at the bottom of the screen and have stable characteristics of high brightness, high edge density and rectangular clustering, which are prone to causing meaningless flickering in ambient light; (b) face / skin color regions, which have a clear numerical distribution range in a specific color space (such as YCrCb or HSV color space) and are often the focus of the screen (1 / 3 of the area in the center of the screen). If they are completely excluded, they will affect the response of ambient light to the narrative subject, but if they are not processed, they will cause the ambient light to be biased towards skin color for a long time; (c) bright / overexposed areas, such as flash, specular reflection, etc., which are characterized by extremely high brightness (such as brightness value V > first brightness threshold 240) and extremely low saturation (such as saturation value S < saturation threshold 30). Directly participating in the calculation will instantly increase the brightness of the ambient light, causing dazzling flickering.

[0043] For each type of detected interference region, the playback terminal independently generates a corresponding pixel weight sub-mask. Each sub-mask (i.e., pixel weight sub-mask) is a two-dimensional matrix with the exact same size as the downsampled video frame F_small. The rules for generating the sub-mask can be determined by a preset weight assignment strategy. In short, the strategy is as follows: for interference regions that need to be strictly excluded (such as subtitle / text regions, highlight / overexposure regions), the corresponding pixels are marked with extremely low weights (such as 0) in the sub-mask, indicating that these pixels do not participate in the subsequent main color extraction calculation at all; for interference that needs to be weakened (such as skin color on a face), it is marked with an intermediate weight value (such as 0.3), indicating that these pixels participate in the main color extraction, but their contribution is reduced; for pixels in non-interference regions, it is marked with standard weights (such as 1), indicating that these pixels participate normally in the main color extraction, and their contribution is not suppressed; it should be noted that the first pixel value of the highlight / overexposure region can be set to extremely low weight, and the second pixel value can be set to a gradient weight value, with the gradient weight value varying between the first weight value and the second weight value. The final result is three pixel weighted sub-masks: Mask_subtitle (title / text mask), Mask_skin (face / skin color mask), and Mask_highlight (highlight / overexposure mask).

[0044] After all types of sub-masks are generated, the playback terminal combines them into a final pixel weight mask, Final_Mask, using a preset, unified fusion rule. This fusion rule operates pixel-by-pixel: for each pixel position (or each pixel) (i,j) of the downsampled video frame F_small, the playback terminal obtains the mask value (i.e., weight value) of that pixel position in the three sub-masks (Mask_subtitle(i,j), Mask_skin(i,j), Mask_highlight(i,j)), and performs a combined calculation on these three mask values ​​to obtain the fusion weight value of that pixel in the final mask, Final_Mask. The final generated Final_Mask will be used in subsequent candidate primary color extraction steps. The mask value of each pixel directly determines its contribution to the calculation of the primary color of the image (a higher mask value results in a greater contribution; a mask value of 0 indicates no contribution).

[0045] In this embodiment, by detecting three types of core interference regions in parallel, the main sources of interference in the main color extraction process can be comprehensively covered, avoiding main color distortion caused by the failure to detect a single interference region, thus improving the comprehensiveness of interference suppression. A differentiated weighting strategy (strict exclusion / weakening of influence / normal participation) is adopted to address the impact degree of different types of interference regions. This avoids main color jumps and flickering caused by subtitles and overexposed areas, while also ensuring the responsiveness of the face area to the main subject of the image, thus improving the rationality of main color extraction. Through independent generation and unified fusion of sub-masks, precise suppression of interference regions is achieved, ensuring that the final pixel weight mask can accurately distinguish between valid pixels and interference pixels. This provides a reliable weighting basis for subsequent candidate main color extraction, guaranteeing that the extracted candidate main color can truly reflect the main atmosphere of the video image.

[0046] In some embodiments, see Figure 3 As shown, in S12 above, the generation of pixel weight sub-masks corresponding to each interference region based on the detection results of each interference region and the preset weight assignment strategy can, in specific implementation, include: S121: Set the pixel value corresponding to the detected subtitle or text region to a first weight value that indicates complete exclusion, and set the remaining pixel values ​​to a second weight value that indicates normal participation in candidate primary color extraction, so as to obtain the subtitle or text mask corresponding to the subtitle or text region. S122: Set the pixel value corresponding to the detected face or skin color region to a third weight value that represents a reduction in weight. The third weight value is between the first weight value and the second weight value. Set the remaining pixel values ​​to the second weight value to obtain the face or skin color mask corresponding to the face or skin color region. S123: Set the first pixel value corresponding to the detected bright or overexposed area as the first weight value, and set the second pixel value corresponding to the detected bright or overexposed area as the gradient weight value. The brightness value of the first pixel value is greater than the first brightness threshold and the saturation value is less than the saturation threshold. The brightness value of the second pixel value is between the first brightness threshold and the second brightness threshold. The gradient weight value is determined according to the first brightness threshold and the second brightness threshold, and varies between the first weight value and the second weight value.

[0047] Specifically, a subtitle / text mask (high brightness, dense edges, rectangular area, stable position) is applied and masked out, a face / skin color mask (lightweight YCrCb / HSV range + optional face bounding box) is applied with reduced weight, and the highlighted / overexposed mask is clipped to limit the area to avoid flickering.

[0048] Specifically, the process of generating a sub-mask for each type of interference region can be as follows: Generating a subtitle / text mask (Mask_subtitle): First, subtitle / text regions are detected on the downsampled video frame F_small. In one embodiment, this is achieved by analyzing the luminance channel (Y, or Y=0.299R+0.587G+0.114B if RGB) and edge density (e.g., using the Sobel operator to calculate the gradient) of approximately 20% of the bottom height region of the image. If a region simultaneously satisfies both an average luminance higher than a threshold (e.g., Y_mean>200 (8-bit)) and an edge density higher than a threshold (e.g.,>15%), it is identified as a subtitle / text region. In the generated Mask_subtitle matrix, all pixels identified as subtitle / text regions are set to a first weight value indicating complete exclusion, such as 0 (i.e., Mask=0). All other pixels are set to a second weight value indicating normal participation, such as 1.

[0049] Generate a face / skin mask: Convert the downsampled video frame F_small to a color space suitable for face / skin detection, such as YCrCb or HSV color space. Identify face / skin pixels according to a preset face / skin range (e.g., in YCrCb color space, luminance Y∈[80,255], red chromaticity Cr∈[133,173], blue chromaticity Cb∈[77,127]). In the generated Mask_skin matrix, pixels identified as faces / skin are assigned a third weight value representing a reduction in weight, between the first and second weight values, for example, 0.3 (this value can be adjusted according to the actual situation). This means that face / skin pixels still contribute color, but their influence is significantly reduced. All non-face / skin pixels are assigned a second weight value of 1.

[0050] Generate a highlight / overexposure mask (Mask_highlight): Convert the downsampled video frame F_small to the HSV color space, identify the first pixel value with a brightness value greater than a first brightness threshold (e.g., 240) and a saturation value less than a saturation threshold (e.g., 30), and the second pixel value with a brightness value between the first brightness threshold and a second brightness threshold (e.g., 220) (i.e., V∈[220,240]), and determine them as highlight / overexposure areas. In the generated Mask_highlight matrix, the first pixel value is set to a first weight value of 0, and the second pixel value is set to a gradient weight. The gradient weight value is determined according to the first brightness threshold and the second brightness threshold, and its calculation formula can be: Gradient weight value = (first brightness threshold - brightness value) / (first brightness threshold - second brightness threshold). When the first brightness threshold is 240 and the second brightness threshold is 220, the gradient weight value = (240-V) / 20, and its value varies between the first weight value 0 and the second weight value 1.

[0051] This differentiated assignment strategy (such as assigning 0 to subtitles / highlights, 0.3 to faces / skin tones, and 1 to normal; assigning 0 to the first pixel value in the highlight / overexposure mask and assigning a gradient weight value to the second pixel value) reflects the algorithm's understanding of the image content, rather than simple binarization filtering.

[0052] In this embodiment, by employing a "complete exclusion" strategy for the subtitle / text area, the persistent interference of subtitle flicker on the extraction of the main color is effectively eliminated. By using a "reduced weight" strategy instead of complete exclusion for the face / skin color area, the algorithm demonstrates its understanding of the main subject: avoiding long-term dominance of ambient light by skin color while preserving a moderate color influence on the subject as the narrative focus. By employing a "tiered processing" strategy for highlighted / overexposed areas, distinguishing between core overexposed areas (complete exclusion) and transitional highlighted areas (gradual weighting), the continuity of brightness in real optical phenomena is simulated. Through these designs, the final lighting effect changes more closely match the perceptual characteristics of the human eye, providing users with a more comfortable and immersive surround light experience. This multi-layered, differentiated processing strategy is one of the key technological foundations for achieving high-quality video ambient light synchronization.

[0053] In some embodiments, see Figure 4 As shown, in S13 above, the generated pixel weight sub-masks are fused according to a preset fusion rule to generate a pixel weight mask for eliminating interference regions. In specific implementations, this may include: S131: For each pixel position in the downsampled video frame, obtain the mask value of all pixel weight sub-masks at each pixel position; S132: Based on the acquired multiple mask values, calculate the fused mask value according to the preset fusion rule, wherein the preset fusion rule is the minimum value rule; S133: Generate the pixel weight mask based on the fusion mask value calculated for each pixel position.

[0054] Specifically, after generating three pixel weighted submasks (Mask_subtitle, Mask_skin, Mask_highlight), the mask values ​​(i.e. weight values) of all pixel weighted submasks at each pixel position in the downsampled video frame F_small can be obtained pixel by pixel.

[0055] Specifically, for any pixel position (i,j) of the downsampled video frame F_small, the playback terminal obtains the mask value Mask_subtitle(i,j) of the subtitle / text mask at pixel position (i,j), the mask value Mask_skin(i,j) of the face / skin color mask at pixel position (i,j), and the mask value Mask_highlight(i,j) of the highlight / overexposure mask at pixel position (i,j), thus obtaining three corresponding mask values ​​(all ranging from 0 to 1).

[0056] Based on the three acquired mask values, the playback terminal calculates the fusion mask value according to a preset fusion rule (such as the minimum value rule). The specific logic of the minimum value rule is as follows: the playback terminal compares the three mask values ​​at pixel position (i,j) and selects the mask value with the smallest value as the fusion mask value Final_Mask(i,j) for pixel position (i,j). The formula for calculating the fusion mask value is as follows: Final_Mask(i,j)=min(Mask_subtitle(i,j), Mask_skin(i,j), Mask_highlight(i,j)).

[0057] For all pixel positions in the downsampled video frame F_small, the playback terminal performs the aforementioned "obtain mask value → calculate fusion mask value" operation. After processing all pixel positions, based on the fusion mask value Final_Mask(i,j) calculated for each pixel position (i,j), a two-dimensional matrix with the exact same size as the downsampled video frame F_small is generated. This final pixel weight mask, Final_Mask, is used to eliminate interference regions. This Final_Mask will be directly applied to subsequent candidate primary color extraction and can serve as the core basis for pixel contribution weights, achieving comprehensive suppression of all interference regions.

[0058] In this embodiment, sub-mask fusion can be completed quickly by using preset fusion rules, avoiding increasing the computational burden on the playback terminal and ensuring the real-time performance of the method. By using a pixel-by-pixel fusion method, it can be ensured that the fused mask value at each pixel position can accurately reflect its interference level, avoiding situations where local interference is not suppressed or effective pixels are mistakenly suppressed, thus improving the accuracy of the final pixel weight mask and providing accurate weight support for subsequent candidate primary color extraction.

[0059] S102: Based on the downsampled video frame, the pixel weight mask, and the preset spatial weight matrix, extract the candidate primary color of the downsampled video frame.

[0060] Specifically, the aforementioned preset spatial weight matrix (or simply spatial weight matrix) can be a two-dimensional weight matrix with the same size as the downsampled video frame F_small. Its value distribution follows a specific functional law (such as a Gaussian distribution), making the pixel weights in the central region of the image higher than those in the edge region. This reflects the compositional rule that the main subject of the image is usually located in the center, thereby improving the semantic accuracy of the primary color extraction. The aforementioned candidate primary color can refer to one or more colors extracted from the downsampled video frame that best represent the atmosphere of the main subject of the image. They are usually characterized by a set of parameters in the color space (such as candidate hue θ_candidate, candidate saturation S_candidate, and candidate brightness V_candidate).

[0061] Specifically, the playback terminal first generates or loads a preset spatial weight matrix W_spatial. Then, it converts each pixel (i,j) of the downsampled video frame F_small from the YUV or RGB color space to the HSV color space (which can be called the first color space, where H∈[0,360], S∈[0,1], V∈[0,1]), obtaining the hue value H(i,j), saturation value S(i,j), and brightness value V(i,j) for each pixel position (i,j). Then, for each pixel (i,j) in the downsampled video frame F_small, it calculates the contribution weight of that pixel position (i,j) to the hue histogram by combining its corresponding pixel weight mask value Final_Mask(i,j), spatial weight matrix value W_spatial(i,j), and its own saturation value S(i,j). Next, a weighted hue histogram (Histogram(θ)) is constructed based on the hue values ​​H(i,j) of all pixels and their contribution weights. The horizontal axis of the weighted hue histogram (Histogram(θ)) represents the discretized hue range (bins), and the vertical axis represents the sum of the contribution weights of all pixels falling into each hue bin. Then, the weighted hue histogram (Histogram(θ)) is analyzed; for example, by detecting the peak values ​​of the histogram, the hue corresponding to the peak with the highest proportion is determined as the main peak hue θ_main. Finally, based on the determined main peak hue θ_main, and the saturation and brightness information of the set of pixels close to this main peak hue, candidate saturation values ​​S_candidate and candidate brightness values ​​V_candidate are obtained through weighted calculations, thus collectively constituting the candidate primary color (θ_candidate, S_candidate, V_candidate).

[0062] The formula for calculating the preset spatial weight matrix can be as follows: W_spatial(i,j)=exp(-((i-H_s / 2)²+(j-W_s / 2)²) / (2σ²)) Where: (H_s / 2, W_s / 2) are the coordinates of the center of the image; σ=min(H_s,W_s) / 4 (standard deviation, controlling the decay rate); the weight of the center area is ≈1.0, and the weight of the edge area is ≈0.3~0.5.

[0063] In this embodiment, by introducing a spatial weight matrix, the influence of the central subject area on the primary color decision is enhanced, making the extracted color more consistent with the visual focus of human viewing. By constructing a weighted hue histogram and analyzing its distribution, the dominant hue of the image can be efficiently and stably summarized from a large number of pixels. By combining contribution weights and hue distribution to extract candidate primary colors, the robustness of the results is ensured, resisting interference from local noise and secondary colors, and accurately capturing the narrative primary color tone of the video image.

[0064] In some embodiments, see Figure 5 As shown, in the above-mentioned S102, the extraction of candidate primary colors of the downsampled video frame based on the downsampled video frame, the pixel weight mask, and the preset spatial weight matrix can, in specific implementation, include: S21: Convert the downsampled video frame to the first color space to obtain the hue value, saturation value and brightness value of each pixel position in the downsampled video frame; S22: Based on the pixel weight mask value of the pixel weight mask, the preset spatial weight matrix and the saturation value, calculate the contribution weight of each pixel position to the hue histogram; S23: Construct a weighted hue histogram using the hue values ​​and the contribution weights; S24: Based on the weighted hue histogram, determine the candidate hue, candidate saturation, and candidate brightness of the downsampled video frame and form the candidate primary color.

[0065] Specifically, the playback terminal can convert the downsampled video frame F_small from the original color space (such as RGB or YUV) to the HSV color space (i.e., the first color space). The HSV color space represents color as three intuitive components: hue (H), saturation (S), and value (V). After conversion, the playback terminal obtains the hue value H(i,j)∈[0,360), saturation value S(i,j)∈[0,1], and value V(i,j)∈[0,1] for each pixel position (i,j) in the video frame.

[0066] For each pixel position (i,j), the playback terminal can calculate its "contribution weight" W_contrib(i,j) to the main color decision based on three factors: the pixel weight mask value Final_Mask(i,j), which reflects whether the pixel is excluded or downweighted; the preset spatial weight matrix value W_spatial(i,j), which is usually a two-dimensional Gaussian weight with a high center and low edge, which can be used to emphasize the central area of ​​the image (a common location for visual subjects); and the pixel's own saturation value S(i,j), the higher the saturation, the greater the contribution to the main color. The specific calculation formula is as follows: W_contrib(i,j)=Final_Mask(i,j)×W_spatial(i,j)×S(i,j) , this contribution weight integrates interference exclusion, spatial importance and color vividness, and can be used to characterize the effectiveness, reliability and importance of the pixel in the statistical distribution of the overall hue of the image.

[0067] Then, a weighted hue histogram (Histogram(θ)) can be constructed using the hue values ​​H(i,j) of all pixels and their contribution weights W_contrib(i,j). This histogram divides the hue ring from 0 to 360 degrees into N consecutive bins. The value of each bin is not a simple pixel count, but the sum of the contribution weights of all pixels whose hue falls within that bin.

[0068] Next, the weighted hue histogram can be analyzed, for example, by detecting its significant peaks to determine the dominant hue range. Then, within this dominant hue range, the saturation and brightness values ​​of relevant pixels are weighted and calculated based on the pixel weight mask value Final_Mask(i,j) and the spatial weight matrix to determine the candidate saturation S_candidate and candidate brightness V_candidate, which, together with the candidate hue θ_candidate, form the candidate primary color (θ_candidate, S_candidate, V_candidate).

[0069] In this embodiment, by converting to the HSV color space, hue, saturation, and brightness are decoupled, facilitating color statistics and effective pixel selection. Instead of simple counting, a weighted hue histogram is constructed by integrating pixel weight mask values, a preset spatial weight matrix, and the pixel's own saturation weights, significantly improving the stability and anti-interference capability of the primary color. By parsing and weighting the candidate saturation and brightness from the weighted hue histogram, it is ensured that the candidate primary color is a comprehensive and accurate color description, not just a hue value. The resulting candidate primary color is stable, accurate, and flicker-resistant, providing a reliable foundation for subsequent temporal filtering and lighting drive.

[0070] In some embodiments, see Figure 6 As shown, the construction of a weighted hue histogram using the hue value and the contribution weight in step S23 above can, in specific implementation, include: S231: Divide the range of hue values ​​into a preset number of bins, calculate the width of each bin, wherein the bins are continuous and non-overlapping hue intervals; S232: For each pixel position in the downsampled video frame, obtain the pixel weight mask value of the pixel weight mask at each pixel position. If the pixel weight mask value is greater than 0 and the saturation value is greater than a preset saturation threshold, then determine the corresponding bin index based on the hue value and the bin width. S233: The contribution weights are added to the histogram values ​​of the bins corresponding to the bin indexes to obtain an initial hue histogram; S234: If the cumulative proportion of the peak bins in the initial hue histogram is greater than a preset proportion threshold, then the histogram values ​​of the peak bins are attenuated and the histogram values ​​of the secondary peak bins are increased to obtain the final weighted hue histogram.

[0071] Specifically, the playback terminal can divide the hue range [0, 360) into N (e.g., 36 or 72) consecutive and non-overlapping bins. The width of each bin is bin_width = 360 / N. An array Histogram of length N is initialized, with all elements set to zero.

[0072] Next, each pixel position (i,j) of the downsampled video frame F_small is traversed. A pixel participates in histogram construction only if its pixel weight mask value Final_Mask(i,j) is greater than 0 (i.e., not completely excluded) and its saturation value S(i,j) is greater than a preset saturation threshold S_min (e.g., 0.15). During construction, the playback terminal calculates the bin index bin_index=k=floor(H(i,j) / bin_width) based on its hue value H(i,j), and then adds the contribution weight W_contrib(i,j) of that pixel position (i,j) to the k-th bin of the histogram: Histogram[k]+=W_contrib(i,j). After traversal, the initial hue histogram is obtained.

[0073] Next, the percentage R_peak (weight of that bin / total weight) of the bin with the highest cumulative weight in the initial histogram is calculated. If R_peak exceeds a preset threshold R_max (e.g., 0.70, i.e., 70%), it is considered that the image may be excessively dominated by a single color (such as large areas of black, white, or a single background color). To prevent the resulting monotonous ambient light color, the playback terminal performs anti-monotony processing: The weights of the peak bins are attenuated, for example: Histogram[peak_bin] = Histogram[peak_bin] × (R_max / R_peak). Optionally, the weights of the second-highest peak bins can be increased, for example: Histogram[bin_second] = Histogram[bin_second] × 1.2. The processed histogram is the final weighted hue histogram.

[0074] In this embodiment, by dividing the hue value range into continuous and non-overlapping bins, the hue distribution can be discretized and normalized statistically, avoiding fluctuations in the judgment of the main color caused by single-point pixel noise, and improving the robustness and stability of hue statistics. By effectively judging the pixel weight mask value and saturation, only non-interfering and highly reliable color pixels are retained to participate in the histogram statistics, further filtering out invalid or interfering pixels such as subtitles, overexposure, and low-saturation weak colors, ensuring the accuracy of hue statistics. Weighted accumulation based on contribution weights rather than simple counting can reflect the difference in importance of different pixels in the extraction of the main color, making the constructed initial hue histogram more consistent with the main color distribution of the image, and improving the representativeness of the main color. By performing peak attenuation and secondary peak enhancement on the histogram with an excessively high proportion of the main peak, it is possible to avoid a single hue from excessively dominating the statistical results, and to prevent abrupt changes, monotony, and flickering of ambient light due to local strong colors in the image, making the hue distribution smoother and more balanced. The final weighted hue histogram obtained after equalization processing has more reasonable peaks and color distribution that better matches the overall atmosphere of the image. This provides a stable, reliable, and smooth data foundation for subsequent peak detection and candidate primary color determination, effectively improving the continuity and visual comfort of primary color extraction.

[0075] In some embodiments, see Figure 7 As shown, in S24 above, determining the candidate hue, candidate saturation, and candidate brightness of the downsampled video frame based on the weighted hue histogram can, in specific implementation, include: S241: Perform peak detection on the weighted hue histogram and determine the hue corresponding to the peak with the highest proportion as the main peak hue; S242: Determine the candidate hue of the downsampled video frame based on the binning position and binning width corresponding to the main peak color; S243: Based on the pixel weight mask value and the preset spatial weight matrix, the saturation value and brightness value of all pixels within the preset hue range are weighted and summed to determine the candidate saturation and candidate brightness respectively, wherein the preset hue range is centered on the candidate hue.

[0076] Specifically, the playback terminal smooths the final weighted hue histogram (e.g., using a three-point moving average) to suppress minor noise, then performs peak detection to obtain candidate hues θ_candidate. The mean / mode is then combined to estimate candidate saturation S_candidate and candidate brightness V_candidate.

[0077] Specifically, the weighted hue histogram Histogram(θ) / H(θ) is first smoothed (using a 3-point moving average to avoid noise spikes), then local maxima are found: H(θ_i) > H(θ_i-1) and H(θ_i) > H(θ_i+1), and then significant peaks are selected: peak height > 0.3 × max(H(θ)). If multiple peaks exist, the hue corresponding to the peak with the highest proportion is selected as the main peak hue θ_main.

[0078] Next, based on the bin position (θ_main × bin_width) and bin width (bin_width) corresponding to the main peak color, the candidate hue θ_candidate is determined according to the formula: θ_candidate = θ_main × bin_width + bin_width / 2. This means the candidate hue can be taken as the center hue value of the bin containing the main peak. This formula ensures that the candidate hue falls within the hue range corresponding to the main peak and avoids hue deviations caused by bin boundary values, making the candidate hue more accurately represent the main color of the image.

[0079] Next, a preset hue range [θ_candidate-Δθ, θ_candidate+Δθ] can be defined centered on the candidate hue θ_candidate (where the hue tolerance Δθ=15°, which can be adjusted to 10°~20° according to the actual scene). Then, the weighting coefficient Weight(i,j) is determined based on the product of the pixel weight mask value Final_Mask(i,j) and the preset spatial weight matrix W_spatial(i,j). Based on the weighting coefficient Weight(i,j), the saturation and brightness values ​​of all pixels within the preset hue range are weighted and summed to determine the candidate saturation S_candidate and candidate brightness V_candidate respectively. S_candidate=Σ(S(i,j)×Weight(i,j)) / Σ(Weight(i,j)) V_candidate=Σ(V(i,j)×Weight(i,j)) / Σ(Weight(i,j)) The weighting coefficient is Weight(i,j) = Final_Mask(i,j) × W_spatial(i,j).

[0080] Finally, the playback terminal combines the calculated θ_candidate (candidate hue), S_candidate (candidate saturation), and V_candidate (candidate luminance) to obtain the candidate primary color (θ_candidate, S_candidate, V_candidate) of the downsampled video frame.

[0081] In this embodiment, by performing peak detection on the weighted hue histogram, the hue range with the highest proportion and strongest representativeness in the image can be quickly and accurately located, ensuring that the candidate hues are taken from the main color of the image, rather than local noise or blemishes. By calculating the candidate hues based on the binning position and binning width of the main peak hue, the discretized histogram results can be restored to continuous and accurate hue values, making the candidate hues more closely match the real color distribution and improving the accuracy of the main color. By setting a preset hue range centered on the candidate hue, and only performing saturation and brightness statistics on pixels within this range, irrelevant color interference can be eliminated, ensuring that the candidate saturation, candidate brightness, and candidate hue are highly matched, making the overall main color more unified and realistic. By performing weighted summation based on pixel weight mask values ​​and spatial weight matrices, interference areas such as subtitles, overexposure, and faces can be further suppressed, while strengthening the contribution of the visually important areas in the center of the image, making the candidate saturation and candidate brightness more representative and stable. The above can provide a stable, smooth, and atmospheric base color for subsequent color correction, temporal filtering, and lighting drive output, effectively improving the visual comfort and continuity of ambient light synchronization.

[0082] In some embodiments, see Figure 8 As shown, after S24 above, in specific implementation, it may also include: S244: If the candidate saturation is less than a preset saturation threshold, the candidate primary color is corrected to a target neutral color determined based on the candidate brightness, and used as the semantically corrected candidate primary color. S245: If the candidate saturation is not less than the preset saturation threshold, and the weighted hue histogram contains at least two peaks that satisfy the preset significance condition, and the difference in the proportion of any two peaks is less than the preset difference threshold, then: Obtain the hue of the stable color output from the previous frame as the reference hue; Calculate the shortest distance on the color wheel between the hue corresponding to each peak that meets the condition and the reference hue; The hue corresponding to the peak with the shortest distance to the reference hue is determined as the corrected candidate hue, and the corrected candidate hue is combined with the candidate saturation and the candidate brightness to form the semantically corrected candidate primary color.

[0083] Specifically, the saturation of the candidate primary color, S_candidate, can be compared with a preset saturation threshold, S_threshold (e.g., 0.15, configurable range 0.10~0.25). If S_candidate < S_threshold, the current scene color is considered too pale, close to "grayish white" (below 15% saturation is considered "grayish white"), and should be corrected towards "ambient white / cool white" that conforms to human visual habits. At this time, the specific target neutral color (cool white or warm white) can be determined based on the brightness, V_candidate, of the candidate primary color. If V_candidate > 0.6 (high brightness scene), then the candidate primary color is corrected to cool white under standard daylight (simulated daylight), for example, using CIE 1931 chromaticity coordinates (x=0.313, y=0.329, corresponding to a color temperature of about 6500K).

[0084] If V_candidate≤0.6 (medium-low brightness scene), then the candidate primary color is corrected to warm white (simulating candlelight / dusk), for example, using CIE 1931 chromaticity coordinates (x=0.448, y=0.408, corresponding to a color temperature of about 2700K).

[0085] Ultimately, the determined target neutral color (cool white or warm white) can be used as the "semantically corrected candidate primary color".

[0086] If S_candidate is not less than S_threshold, and the following conditions are met simultaneously, the playback terminal determines it as a "multi-peak proximity" scenario and performs hue smoothing correction: The weighted hue histogram contains at least two peaks that meet preset significance conditions (e.g., peak percentage ≥ 10%); the percentage difference between any two valid peaks is less than a preset difference threshold (20%), meaning the difference in the percentage of candidate hues corresponding to the peaks is less than 20%. The specific correction steps are as follows: First, obtain the reference hue: The playback terminal reads the hue of the stable color output of the previous frame of video from the local buffer, denoted as θ_prev, as the reference benchmark for the color continuity between frames, i.e., the reference hue; Next, calculate the shortest hue distance: For each peak that meets the conditions, the corresponding hue θ_n (n is the peak number, such as θ1, θ2) is calculated. Considering the 360° cycle characteristic of the hue circle, the shortest angular distance between it and the reference hue θ_prev is calculated. The calculation formula is: distance_n=min(|θ_n-θ_prev|,360-|θ_n-θ_prev|).

[0087] For example: if the reference hue of the previous frame is θ_prev=20° (red-orange), and the effective peak hues of the current frame are θ1=10° (red) and θ2=200° (cyan), then: distance1=min(|θ1-θ_prev|,360-|θ1-θ_prev|)=min(|10-20|, 360-10)=10° distance2=min(|θ2-θ_prev|,360-|θ2-θ_prev|)=min(|200-20|, 360-180)=180°.

[0088] Next, select the optimal candidate hue: compare the shortest distance between all valid peak hues and the reference hue, and determine the hue corresponding to the peak with the smallest distance as the corrected candidate hue θ_candidate_fix (in the example above, θ1=10° is selected).

[0089] Finally, the output is: the corrected candidate hue θ_candidate_fix is ​​combined with the previously determined candidate saturation S_candidate and candidate brightness V_candidate to obtain the "semantically corrected candidate primary color" (θ_candidate_fix, S_candidate, V_candidate).

[0090] It should be noted that semantic correction can be performed on scenarios with no effective colors or where multi-color balance is prone to abrupt changes. For other scenarios, the candidate primary color determined in S24 above can be directly used.

[0091] In this embodiment, by semantically modifying the candidate primary color, the problems of color distortion in low-saturation images and hue jumps in multi-peak balanced images can be effectively improved, enhancing the rationality, continuity, and visual comfort of the primary color output. Specifically, for scenarios where the candidate saturation is lower than a preset saturation threshold, the candidate primary color is modified to a target neutral color determined based on the candidate brightness. This avoids outputting dirty, low-saturation, and ineffective colors when the overall image is grayish-white and lacks effective colors, making the ambient light output more in line with the image brightness atmosphere and human visual habits, ensuring stable, natural, and non-abrupt color output in low-saturation scenes. For scenarios with sufficient saturation but multiple significant peaks in the histogram and similar peak proportions, the hue of the stable color from the previous frame is introduced as a reference, and the peak hue with the shortest distance to the reference hue on the hue circle is selected as the modified candidate hue. This avoids large jumps and abrupt changes in hue between frames on the hue circle, making the color transition between adjacent frames smoother and more continuous, effectively eliminating ambient light flickering, color jumps, and sudden changes in temperature, and improving color temporal consistency. The semantic correction process only makes adaptive adjustments to the candidate hue, preserving the overall atmospheric characteristics of the original candidate saturation and brightness. This ensures the stability of the main color without destroying the brightness and saturation information of the image itself, so that the final output main color not only fits the semantics of the image but also has good inter-frame coherence.

[0092] S103: Perform temporal filtering and perceptual mapping on the candidate primary color to obtain the lamp driving color.

[0093] Specifically, the aforementioned temporal filtering process refers to smoothing the candidate primary colors of consecutive video frames over time to suppress rapid color transitions between frames and ensure the continuity and comfort of ambient light changes. The aforementioned perceptual mapping refers to the process of converting temporally stable, video-standard-compliant colors into color parameters that can be accurately reproduced on the target lighting hardware and meet human visual comfort. The aforementioned lighting drive color refers to the color parameters ultimately used to control the emission state of the target lighting fixture, and its format is adapted to the specific lighting fixture's driving protocol (such as RGB values, CIE xyY coordinates, etc.).

[0094] Specifically, the playback terminal first obtains the candidate primary color (θ_candidate, S_candidate, V_candidate) of the current video frame and the stable target color (θ_prev, S_prev, V_prev) output after processing from the previous video frame. Then, it calculates the difference or normalized color difference between the current candidate primary color and the stable color of the previous frame, and dynamically adjusts the smoothing coefficient α of an exponential moving average filter by combining the motion intensity information M of the current video frame. Using the adjusted smoothing coefficient α, the hue, saturation, and brightness components of the current candidate primary color are calculated using exponential moving average filtering to obtain the temporally stable target color (θ_stable, S_stable, V_stable). After that, it enters the perceptual mapping stage: first, the stable target color (θ_stable, S_stable, V_stable) is converted from the HSV color space to the device-independent CIE XYZ color space (i.e., the second color space), and then converted to the CIE xyY color space (i.e., the second color space) which is easier for color gamut processing, to obtain the color coordinates (x, y) and brightness. Next, based on the pre-obtained physical luminous color gamut of the target luminaire, the calculated color coordinates (x, y) are judged and clipped. If the coordinates fall outside the color gamut, they are mapped to the color gamut boundary to obtain the clipped chromaticity coordinates (x_clip, y_clip). Simultaneously, the brightness is normalized and a brightness change rate limit is applied to obtain the final driving brightness L_target. Finally, the chromaticity coordinates (x_clip, y_clip) and the driving brightness L_target are combined to form the final luminaire driving color.

[0095] By applying adaptive temporal filtering to candidate primary colors, color jitter caused by scene transitions or rapid changes in image content is effectively smoothed, avoiding the flickering effect of ambient light and improving visual comfort. Through perceptual mapping from the video color gamut to the lighting fixture color gamut, the colors derived from video content analysis are accurately and consistently reproduced on smart lighting fixtures of different brands and models, resolving the color distortion problem. Brightness normalization and speed limiting prevent drastic fluctuations in ambient light brightness, protecting both the user's visual experience and the safety of the lighting fixture hardware.

[0096] In some embodiments, see Figure 9 As shown, the candidate primary color in S103 above undergoes temporal filtering and perceptual mapping to obtain the lamp driving color. In specific implementations, this may include: S31: Perform temporal filtering on the candidate primary color to obtain a stable target color, wherein the stable target color includes stable hue, stable saturation and stable brightness; S32: Perform a perceptual mapping from the video color gamut to the target luminaire color gamut on the stable target color to obtain the luminaire-driven color.

[0097] In some embodiments, see Figure 10 As shown, the temporal filtering process performed on the candidate primary color in S31 above to obtain a stable target color can, in specific implementation, include: S311: Calculate the normalized color difference between the candidate primary color and the stable color of the previous video frame. The normalized color difference is calculated based on the shortest hue difference on the color wheel, the saturation difference, and the brightness difference. S312: Dynamically adjust the smoothing coefficient of the exponential moving average filter based on the normalized color difference and the motion intensity of the current video frame; S313: Based on the adjusted smoothing coefficient, exponential sliding filtering is performed on the candidate hue, candidate saturation and candidate brightness respectively to obtain a stable target color; In specific implementations, the smoothing coefficient of the dynamically adjusted exponential moving average filter in S312 may include: If the normalized color difference exceeds a preset color difference threshold or the motion intensity exceeds a preset motion threshold, the smoothing coefficient is increased to accelerate the color response; otherwise, the basic filtering coefficient is used to maintain color stability.

[0098] Specifically, temporal filtering achieves inter-frame color smoothing and suppresses rapid transitions through adaptive exponential sliding filtering (IIR), which may include the following steps: First, calculate the normalized color difference between the candidate primary color and the stable color of the previous frame: The playback terminal first reads the candidate primary color (θ_candidate, S_candidate, V_candidate) of the current video frame, and the stable target color (θ_prev, S_prev, V_prev) output from the previous frame, and calculates the normalized color difference according to the following rules: 1. Calculate the shortest hue difference on the color wheel: Considering the 360° cycle of the color wheel, the hue difference is: ΔH=min(|θ_candidate-θ_prev|,360-|θ_candidate-θ_prev|); 2. Calculate the saturation difference and brightness difference: ΔS=|S_candidate-S_prev|, ΔV=|V_candidate-V_prev|; 3. Calculate normalized color difference: Normalize the differences in each dimension to a uniform scale. The formula is: color_diff=sqrt((ΔH / 180)²+ΔS²+ΔV²) (After normalization, the color difference range is 0~2, and the larger the value, the more drastic the color change).

[0099] Afterwards, the playback terminal pre-configures the basic filtering parameters: α_base=0.3 (basic smoothing coefficient for slow scenes, configurable range 0.2~0.5), color_diff threshold=0.4 (scene switching judgment threshold), M_threshold=30 (motion intensity threshold, unit: pixels / frame, estimated by the average displacement of 8×8 block matching of adjacent frames), and dynamically adjusts the smoothing coefficient α in combination with the motion intensity M of the current frame (obtained by block matching / optical flow estimation). 1. If color_diff > 0.4 (indicating scene transition / drastic color change): α = min(0.8, α_base + 0.3 × color_diff) to speed up color response; 2. If M > M_threshold (indicated as high motion intensity / fast camera movement): α = min(0.6, α_base + 0.2 × M), appropriately speed up the response; For other scenes (slow scenes / static images): α = α_base, maintain high smoothness to ensure color stability.

[0100] Next, the playback terminal performs IIR filtering on the candidate hue, candidate saturation, and candidate brightness components respectively. The hue component needs to be processed for 360° loop characteristics, which can be done as follows: 1. Candidate Hue Filtering: Hue values ​​range from 0° to 360° (distributed in a ring, with 0° and 360° representing the same hue). If linear filtering is directly applied to hue values ​​crossing the 0° / 360° boundary, a "cross-ring jump" problem will occur (e.g., θ_candidate=10°, θ_prev=350°, linear calculation will yield (350+10) / 2=180°, which does not match the expected smooth transition to 0°). Therefore, it is necessary to first adjust the hue values ​​crossing the boundary before performing exponential sliding filtering. The specific logic is as follows: First, calculate the absolute difference between the current candidate hue θ_candidate and the stable hue θ_prev from the previous frame. If |θ_candidate - θ_prev| > 180°, it is determined to have crossed the 0° / 360° boundary (e.g., θ = 10°, θ_prev = 350°), and then the candidate hue θ_candidate is adjusted. If θ_candidate > θ_prev (e.g., θ_candidate = 10°, θ_prev = 350°), shift the hue value of the previous frame upward by 360° to get θ_prev_adjusted = θ_prev + 360 (i.e., 350 + 360 = 710°). If θ_candidate < θ_prev (e.g., θ_candidate = 350°, θ_prev = 10°), shift the current candidate hue value upward by 360° to obtain θ_adjusted = θ_candidate + 360 (i.e., 350 + 360 = 710°). Then, an exponential moving average is applied to the adjusted hue value using an adaptive smoothing coefficient α, with the formula θ_stable=(1-α) ×θ_prev_adjusted+α×θ_adjusted (as in the example above, θ_stable=(1-α)×710+α×10). The filtered result is modulo 360 to bring it back to the range of 0° to 360°. The formula is θ_stable=θ_stable%360 (for example, 710°%360=350° in the above example. After filtering 10° and 350°, a smooth value of about 350° is obtained, which meets the expectation of visual continuity). If |θ_candidate-θ_prev|≤180°, perform linear filtering directly, with the formula θ_stable=(1-α)×θ_prev+α×θ_candidate, ensuring a smooth hue transition.

[0101] 2. Candidate saturation filtering: The candidate saturation value ranges from 0 to 1 (without cyclic characteristics). No boundary processing is required; instead, an exponential moving average filter is directly applied based on the adaptive smoothing coefficient α. The formula is S_stable = (1-α) × S_prev + α × S_candidate. Here, (1-α) represents the weight of the stable saturation of the previous frame, and α represents the weight of the current candidate saturation. A larger α indicates a stronger influence of the current frame's saturation on the result (faster response); a smaller α indicates a higher weight of the previous frame's saturation (stronger stability). This suppresses rapid changes in saturation and ensures continuous changes in ambient light color concentration.

[0102] 3. Candidate brightness filtering: The candidate brightness value ranges from 0 to 1 (without looping characteristics). The filtering logic is consistent with saturation, and the formula is V_stable=(1-α) ×V_prev+α×V_candidate. This filtering method can avoid sudden brightness changes caused by local highlights / dark areas in the image. At the same time, combined with an adaptive α coefficient, it accelerates brightness response during scene transitions and fast camera movements, and maintains stable brightness in static images, thus both conforming to changes in image brightness and preventing flickering of ambient light.

[0103] In this embodiment, through the above filtering process, a time-domain stable target color (θ_stable, S_stable, V_stable) is finally obtained. All three components achieve smooth inter-frame transitions, which not only ensures the real-time response of ambient light to changes in video images, but also avoids visual discomfort caused by color / brightness jumps.

[0104] In some embodiments, see Figure 11 As shown, the process in S32 above, which involves perceptually mapping the stable target color from the video color gamut to the target luminaire color gamut to obtain the luminaire-driven color, can, in specific implementations, include: S321: Convert the stable target color from the first color space to a device-independent second color space; S322: Based on the physical luminous gamut of the target luminaire, the color coordinates in the second color space are clipped to obtain the chromaticity coordinates located within the physical luminous gamut; S323: Based on the brightness of the stable target color and the global brightness statistics of the downsampled video frame, calculate the mapped target brightness; S324: Limit the rate of change of the mapped target brightness; S325: Determine the driving color of the luminaire based on the chromaticity coordinates after gamut clipping and the target brightness after limiting the rate of change of brightness.

[0105] Specifically, the playback terminal can convert the time-domain stabilized target color (θ_stable, S_stable, V_stable) from the video standard color gamut to the driving color adapted to the target lighting hardware. Through steps such as color space conversion, color gamut clipping, brightness normalization, and rate limiting, it can ensure accurate color reproduction, safe brightness output, and compliance with human visual habits. The specific process is as follows: 1. Convert the stable target color to a device-independent second color space (CIE XYZ / xyY). The playback terminal first converts the stable target colors (θ_stable, S_stable, V_stable) from the first color space (HSV) to the Rec.709 RGB color gamut, then converts them to the device-independent CIE XYZ space, and finally converts them to the CIE xyY space for easier color gamut processing. (1) HSV→Rec.709 RGB conversion: RGB_709 (value 0~255); (2) Gamma Decoding (Linearization): Perform Rec.709 Gamma decoding on the RGB values. The formula is as follows: R_linear=(RGB_709.R / 255) 2.2 G_linear=(RGB_709.G / 255) 2.2 B_linear=(RGB_709.B / 255) 2.2 .

[0106] (3) RGB to CIE XYZ conversion: calculated using the standard transformation matrix: [X]=[0.4124,0.3576,0.1805][R_linear] [Y]=[0.2126,0.7152,0.0722][G_linear] [Z]=[0.0193,0.1192,0.9505][B_linear].

[0107] (4) XYZ→xyY conversion: Calculate the chromaticity coordinates and luminance using the following formula: x=X / (X+Y+Z), y=Y / (X+Y+Z), Y_luminance=Y.

[0108] 2. Perform color gamut cropping based on the physical color gamut of the target lighting fixture.

[0109] The playback terminal preloads the physical luminous color gamut range of the target lighting fixture (e.g., Philips Hue color gamut triangle: red (0.675, 0.322), green (0.409, 0.518), blue (0.167, 0.040)), and performs color gamut clipping on the chromaticity coordinates (x, y). If (x, y) falls within the luminaire's color gamut triangle, retain it directly: x_clip=x, y_clip=y; If (x, y) exceeds the color gamut, perform minimum color difference clipping (preserve hue, reduce saturation): Draw a ray from the white point (0.313, 0.329) at point D65 to the target point (x, y); Traverse each boundary of the color gamut triangle and find the first intersection point between the ray and the boundary; Use the coordinates of the intersection point as the chromaticity coordinates (x_clip, y_clip) after clipping.

[0110] 3. Calculate the mapped target brightness based on global brightness statistics.

[0111] The playback terminal combines the global brightness statistics of the downsampled video frames to normalize V_stable, avoiding issues such as excessive brightness / darkness and flickering. (1) Calculate the average brightness of the frame. The calculation formula is as follows: V_frame_avg=Σ(V(i,j)×Final_Mask(i,j)) / Σ(Final_Mask(i,j)) (Mask is the pixel weight mask); (2) Brightness adaptive adjustment: V_normalized=V_stable×(1+0.3×(0.5-V_frame_avg)) (Increase brightness if the frame is too dark, and decrease it if the frame is too bright); (3) Application user preferences: V_user = V_normalized × brightness_scale (brightness_scale ∈ [0.3, 1.5], which can be customized by the user); (4) Brightness limit in Night / Eye Protection Mode: If night mode is enabled (automatic / manual activation from 22:00 to 06:00), V_max = 0.4; otherwise, V_max = 1.0. V_clamped=min (V_user,V_max) (Limit the upper limit of brightness).

[0112] 4. Apply rate of change limitation to the mapped brightness.

[0113] To prevent ambient light flicker, the playback terminal limits the maximum brightness change per frame to ΔV_max = 0.15 (maximum change per frame 15%). If |V_clamped-V_prev|>ΔV_max: V_final=V_prev+sign (V_clamped-V_prev)×ΔV_max; Otherwise: V_final = V_clamped; Gamma encoding converted to lamp brightness: L_target=(V_final^(1 / 2.2))×L_max (L_max is the maximum brightness of the lamp, such as 1000 lumens).

[0114] 5. Color temperature preference adjustment (optional)

[0115] If the user sets a warm / cool bias (bias∈[-1,1]), the stable hue and saturation will be adjusted accordingly: Warm color preference (bia < 0): θ_final = θ_stable + bias × 30 (hue shifts towards red and orange), S_final = S_stable × (1 + 0.2 × |bias|) (increases saturation); Cool color preference (bias > 0): θ_final = θ_stable + bias × 30 (hue shifts towards blue-cyan), S_final = S_stable × (1 - 0.1 × bias) (reduces saturation); When there is no preference, θ_final = θ_stable and S_final = S_stable.

[0116] 6. Output lighting fixture driving color

[0117] The playback terminal combines the chromaticity coordinates (x_clip, y_clip) after gamut clipping with the final brightness L_target to obtain the lamp-driven color (x_clip, y_clip, L_target) in CIE xyY format, and outputs the color change flag is_significant_change (to determine whether it is a significant color change, which is used for subsequent synchronization time calculation).

[0118] In this embodiment, temporal filtering, through adaptive smoothing coefficient adjustment, achieves "fast response in dramatic scenes and high stability in static scenes," effectively suppressing inter-frame color jumps and flickering, and ensuring continuous and natural changes in ambient light. Perceptual mapping solves the problem of adapting the video color gamut to the physical color gamut of the lighting fixtures. Minimum color difference clipping ensures color reproduction, while brightness normalization and speed limiting balance visual comfort and lighting fixture hardware safety. This application supports personalized configurations such as user preferences and night mode to adapt to different usage scenarios. The final output lighting-driven color not only matches the atmosphere of the video but also conforms to human visual habits, significantly enhancing the immersive experience of ambient light synchronization.

[0119] S104: Based on the luminaire driving color and the target effective time of the luminaire driving color, generate and send an ambient light synchronization command to at least one target luminaire to instruct the target luminaire to synchronize its light emission state to the luminaire driving color at the target effective time. The target effective time is calculated based on the display timestamp and the estimated end-to-end latency.

[0120] Specifically, the aforementioned target effective time can refer to the absolute point in time when the ambient light synchronization command is scheduled to actually take effect and be executed at the target luminaire. The aforementioned end-to-end latency estimate can refer to the predicted total time consumed by the entire processing and communication chain from the acquisition of video frames from the playback terminal to the completion of the target luminaire's light-emitting state switching or synchronization. The aforementioned ambient light synchronization command can refer to a structured control data packet, the content of which may include the target luminaire identifier, target color parameters, target effective time, and sequence information used to ensure communication reliability.

[0121] Specifically, the playback terminal first calculates the estimated end-to-end latency Δpipe. This estimated latency can include at least: terminal-side processing latency estimate (covering the computation time for color analysis, filtering mapping, etc.), network transmission latency estimate (calculated based on round-trip latency measurements with the luminaire or gateway), target luminaire response latency estimate (inherent latency determined by the luminaire type or protocol), and buffer latency to cope with network jitter. Then, the display timestamp Tdisplay of the video frame is added to the estimated end-to-end latency Δpipe to obtain the target effective time Ttarget for the luminaire-driven color corresponding to that frame, i.e., Ttarget = Tdisplay + Δpipe. Afterward, the playback terminal encapsulates the luminaire-driven color, target effective time Ttarget, target luminaire identifier, and other necessary information into an ambient light synchronization command data packet. Before transmission, a judgment can be made based on energy-saving strategies, such as comparing the current luminaire-driven color with the color difference of the last transmission. If the difference does not exceed a dynamic threshold and has not reached the maximum silent time, the current transmission may be skipped. If a decision is made to send, the instruction packet is sent to the target light fixture or intermediate gateway via a low-latency network protocol (such as UDP or a vendor-specific entertainment mode protocol).

[0122] By calculating the target activation time based on display timestamps and estimated end-to-end latency, predictive synchronization is achieved. This precisely "schedules" changes in ambient light at a specific future moment, compensating for inherent delays in processing, transmission, and execution. This results in a synchronization effect between lighting changes and video display that is imperceptible to the human eye. By generating structured synchronization commands and selecting appropriate communication protocols, accurate and efficient transmission of control information is ensured. Furthermore, by introducing energy-saving transmission strategies, network bandwidth consumption and system power consumption are effectively reduced while maintaining synchronization effectiveness. Ultimately, this step bridges the gap between video content analysis and physical lighting control, making it crucial for achieving an immersive ambient light synchronization experience.

[0123] In some embodiments, the estimated end-to-end latency in S104 above is calculated as follows: Based on the processing latency estimate, network latency estimate, luminaire response latency estimate, and buffer latency estimate, the full-link latency estimate is calculated. The network latency estimate is determined based on the measured value of the network round-trip latency for communication with the target luminaire. The buffer latency estimate is determined based on the historical statistical fluctuation of the network round-trip latency. The full-link latency estimate is the total estimated latency of the entire process from the acquisition of the video frame to the target luminaire performing illumination state synchronization. Accordingly, the target effective time in S104 above is calculated based on the displayed timestamp and the estimated end-to-end latency. In specific implementation, it may include: The target effective time of the lamp-driven color is obtained by adding the display timestamp and the estimated end-to-end delay.

[0124] Specifically, the estimated end-to-end latency is calculated using the following formula: Δpipe = T_compute + T_network + T_device + T_buffer, where T_compute is the processing latency estimate, which can include the overall computation latency of video frame sampling, weight mask generation, primary color extraction, temporal filtering, and perceptual mapping. This latency is obtained by smoothing the measured computation time of the previous N frames using an exponential moving average, i.e., T_compute = EMA(measured computation time, α = 0.1); T_network is the network latency estimate, which is half of the median of the most recent network round-trip latency historical values, i.e., T_network = median(RTT_history [last 10 frames]) / 2; T_device is the lighting response latency estimate, which can be selected according to the device type of the target lighting fixture; T_buffer is the buffer latency estimate, which is obtained by multiplying the historical statistical standard deviation of the network round-trip latency by a coefficient, i.e., T_buffer = 1.5 × std_dev(RTT_history [last 10 frames]).

[0125] The target effective time Ttarget for the lighting driver color can be obtained by adding the display timestamp Tdisplay to the estimated end-to-end latency Δpipe, i.e., Ttarget = Tdisplay + Δpipe.

[0126] Furthermore, if multiple consecutive frames detect that the actual effective time is later than the target effective time, the estimated end-to-end latency Δpipe is adjusted upwards; if multiple consecutive frames detect that the actual effective time is earlier than the target effective time, the estimated end-to-end latency Δpipe is cautiously adjusted downwards, and Δpipe is limited to a preset range to achieve dynamic adaptive adjustment.

[0127] Specifically, the playback terminal can continuously monitor the deviation between the actual effective time of the lights and the target effective time, and perform precise dynamic closed-loop adjustments to the end-to-end latency estimate Δpipe: If the actual time of the lamp activation is detected to be later than the target activation time Ttarget for 5 consecutive frames and the deviation exceeds 10ms (i.e., actual time > Ttarget + 10ms), it is determined that the end-to-end latency estimate is insufficient. At this time, the end-to-end latency estimate Δpipe is increased by 5ms to increase the advance of the command activation in order to compensate for the additional latency in actual operation. If the actual activation time of the light fixture is detected to be earlier than the target activation time Ttarget for 5 consecutive frames and the deviation exceeds 10ms (i.e., actual_time < Ttarget - 10ms), it is determined that the end-to-end latency estimate is overestimated. At this time, the end-to-end latency estimate Δpipe is reduced by 3ms to carefully reduce the lead time and avoid the light changes deviating from the display rhythm of the screen too early. After adjusting upwards or downwards, the estimated end-to-end latency Δpipe must be limited to a preset range of 30ms to 100ms (i.e., Δpipe = clamp(Δpipe, 30ms, 100ms)) to prevent excessive adjustment from causing Δpipe to be too small to cover the basic latency, or too large to cause serious timing deviations in synchronization, ensuring that the adjusted Δpipe is always within a reasonable and effective range.

[0128] In this embodiment, by accurately estimating the end-to-end latency estimate separately for processing latency, network latency, lighting response latency, and buffer latency, the entire process latency from video frame acquisition to lighting execution can be fully covered, making latency prediction more accurate and reliable. By adding the display timestamp to the end-to-end latency estimate to obtain the target effective time, predictive timing control can be achieved, effectively compensating for inherent latency in processing, transmission, and device execution, ensuring precise alignment between lighting changes and video playback, and enhancing the immersive experience of ambient light synchronization. Dynamic closed-loop adjustment of the end-to-end latency estimate can address uncertainties such as network jitter and system load fluctuations, ensuring stable, drift-free, and seamless synchronization during long-term operation.

[0129] In some embodiments, the generation and transmission of the ambient light synchronization command in S104 above may, in specific implementation, include: The perceived color difference is calculated by comparing the driving color of the lamp with the driving color of the last successfully sent lamp. Select the corresponding dynamic threshold or static threshold based on the motion intensity of the current video frame; If the perceived color difference is greater than the selected threshold, or if the maximum silent time has been exceeded since the last successful transmission, an ambient light synchronization command is generated. The ambient light synchronization command includes at least the luminaire driving color and the target effective time. Otherwise, the generation and transmission of the ambient light synchronization command are suppressed.

[0130] Specifically, intelligent control of command sending can be achieved by combining color difference perception with scene-adaptive threshold strategies. The specific process is as follows: 1. Perceived Color Difference Calculation: The playback terminal can compare the lamp driving color (x_current, y_current, L_current) calculated in the current video frame with the driving color (x_last_sent, y_last_sent, L_last_sent) successfully sent to the lamp last time. The perceived color difference ΔE between the two is calculated in the CIE xyY color space using the following formula: ΔE=sqrt((x_current-x_last_sent)²+(y_current-y_last_sent)²)+0.5×|(L_current-L_last_sent) / L_max| Where Δx=x_current-x_last_sent and Δy=y_current-y_last_sent are the chromaticity coordinate differences, ΔL=L_current-L_last_sent is the luminance difference, and L_max is the maximum luminance value of the target lamp. This formula takes into account the differences in human visual perception between chromaticity and luminance, making the color difference calculation more in line with the visual experience.

[0131] Scene threshold selection: Based on the comparison between the motion intensity M of the current video frame and the preset motion threshold M_threshold, select the corresponding color difference judgment threshold. If the motion intensity M > M_threshold, it is determined to be a dynamic scene (the scene changes rapidly), and a lenient dynamic threshold threshold_dynamic = 0.05 is used. If the motion intensity M ≤ M_threshold, it is judged as a static scene (stable image), and a strict static threshold threshold_static = 0.02 is used.

[0132] Basic transmission judgment: Compare the calculated perceived color difference ΔE with the selected threshold. If ΔE > threshold (dynamic threshold or static threshold), it is determined that a command needs to be sent, send_command=true is set, and the color of the last sent frame is updated to the color of the current frame; if ΔE ≤ threshold, the command transmission is suppressed, send_command=false is set, and the command transmission of this frame is skipped.

[0133] Forced Sending at a Time: The playback terminal keeps track of the silence duration (time_since_last_sent) since the last successful command. Even if the perceived color difference does not reach the threshold, if time_since_last_sent > 500ms, send_command=true will still be forcibly set to ensure that commands are sent periodically to maintain the device communication connection.

[0134] Command generation and transmission: If send_command is true, the playback terminal encapsulates information such as the lamp driver color, target effective time Ttarget, device identifier, and auto-incrementing serial number into a structured ambient light synchronization command packet and sends it to the target lamp through a low-latency communication protocol; if send_command is false, neither the command nor the instruction is generated or sent.

[0135] In this embodiment, through the coordinated control of color difference perception judgment, scene-adaptive threshold selection, and forced transmission of maximum silence time, invalid commands caused by minute color jitter can be effectively filtered out, reducing network bandwidth consumption and system power consumption of the playback terminal and target lights. By dynamically switching between static and dynamic thresholds based on the motion intensity of video frames, ambient light stability can be ensured in static scenes, and synchronization response sensitivity can be guaranteed in dynamic scenes, balancing the stability and real-time performance of ambient light synchronization. The forced transmission mechanism of maximum silence time can periodically maintain device communication connections, avoiding problems such as light fixture offline and synchronization failure caused by prolonged silence.

[0136] In some embodiments, the generation and sending of the ambient light synchronization command to at least one target luminaire in S104 above may, in specific implementation, include: Based on the multiple different spatial partitions of the downsampled video frame, determine the primary color of each partition; Based on the preset mapping relationship between screen partitions and physical light positions, assign exclusive light-driving colors to different target lights based on the primary color of the corresponding partition; Based on the unique luminaire driver color of each target luminaire, a corresponding ambient light synchronization command is generated and sent.

[0137] Specifically, the downsampled video frames can be divided into multiple spatial partitions, which may include, but are not limited to, a four-region division (left, right, top, and bottom) or a 3×3 grid of nine regions (top left, top center, top right-left center, center, right center-bottom left, bottom center, and bottom right). The primary color extraction process is then performed independently on each spatial partition to obtain the corresponding primary color for that partition.

[0138] Then, based on the preset mapping relationship between screen partitions and physical light positions, the target lights in different positions are bound to the corresponding spatial partitions in the video frames. Each target light is assigned a unique light drive color determined based on the primary color of the corresponding partition, so that the left light corresponds to the primary color of the left side of the screen, the right light corresponds to the primary color of the right side of the screen, the top light corresponds to the primary color of the upper side of the screen, and the bottom light corresponds to the primary color of the lower side of the screen.

[0139] When the physical location of the target luminaire lies between multiple zones, the weights of each adjacent zone are calculated using bilinear interpolation based on the luminaire's normalized coordinates. The colors of the adjacent zones are then weighted and mixed to obtain a unique luminaire-driven color with a smooth transition. In scenarios with limited computing resources, the luminaire can be directly mapped to the nearest zone, and the dominant color of that zone can be directly used as the unique luminaire-driven color.

[0140] Finally, based on the exclusive luminaire driver color and the corresponding target effective time of each target luminaire, an ambient light synchronization command adapted to each target luminaire is generated and sent to the corresponding target luminaire to realize distributed ambient light synchronization control of multiple areas and multiple luminaire positions.

[0141] The four regions are divided as follows: the left region satisfies x∈[0,W_s / 2], the right region satisfies x∈[W_s / 2,W_s], the upper region satisfies y∈[0,H_s / 2], and the lower region satisfies y∈[H_s / 2,H_s], where W_s is the width of the downsampled video frame and H_s is the height of the downsampled video frame. Primary color extraction, semantic correction, temporal stabilization, and perceptual mapping are performed independently on each spatial region to obtain the corresponding primary colors (θ_L,S_L,V_L) and (θ_R,S_R,V_R).

[0142] Light fixture position mapping rules (example: 4 lights): Lamp_Left → Main color of the left side of the screen Lamp_Right → Main color of the right side of the screen Lamp_Top → The main color of the upper part of the screen Lamp_Bottom → The primary color of the lower part of the screen; When the physical location of the target luminaire is located between multiple zones, the mixed weights of adjacent zones are calculated using bilinear interpolation based on the luminaire's normalized coordinates (x_lamp, y_lamp), i.e., w_topleft=(1-x_lamp)×(1-y_lamp), w_topright=x_lamp×(1-y_lamp), w_bottomleft=(1-x_lamp)×y_lamp, w_bottomright=x_lamp ×y_lamp, and then calculated according to: R_lamp=w_toplef×R_TL+w_topright×R_TR+w_bottomleft×R_BL+w_bottomright×R_BR, G_lamp=w_topleft×G_TL+w_topright×G_TR+w_bottomleft×G_BL+w_bottomright×G_BR, B_lamp=w_topleft×B_TL+w_topright×B_TR+w_bottomleft×B_BL+w_bottomright×R_BR, The RGB color components of adjacent partitions are weighted and mixed to obtain the exclusive lighting drive color with a smooth transition for the lighting fixture.

[0143] In this embodiment, by performing multi-spatial partitioning on the downsampled video frames and assigning dedicated lighting driver colors to target lights at different physical locations, the spatial distribution of ambient light and video images can be kept consistent, significantly enhancing the immersive and enveloping experience of ambient light synchronization. Based on the preset mapping relationship between screen partitions and physical light positions, distributed synchronous control of multiple lights can be achieved, enhancing the matching degree and realism between ambient light and image content. By independently generating corresponding ambient light synchronization commands for each target light, the emission colors of lights at different locations can accurately reflect the content changes in the corresponding image areas, further enriching the expressive forms and visual effects of ambient light synchronization and improving the overall viewing and entertainment experience.

[0144] Figure 12 A flowchart illustrating an ambient light synchronization method based on video content color analysis is provided in another embodiment of this application. This method can be applied to a gateway and may include: S201: Receive an ambient light synchronization command from at least one playback terminal, the ambient light synchronization command including at least the lamp driving color and the target effective time of the lamp driving color; S202: Convert the received ambient light synchronization command into a format compatible with the communication protocol supported by the target luminaire; S203: Based on the clock synchronization mechanism, calibrate the effective time of the target; S204: Send the ambient light synchronization command, after format conversion and time calibration, to the corresponding target luminaire.

[0145] Specifically, firstly, the gateway monitors in real time for ambient light synchronization commands from at least one playback terminal. These commands are structured instruction packets generated by the playback terminal based on video color analysis, and may include at least the target luminaire identifier, the luminaire driving color (x, y, L_target), and the target activation time Ttarget. They may also include auxiliary information such as the instruction sequence number (sequence_id) and the luminaire color transition duration. The gateway performs preliminary verification on the received commands, filtering out commands with valid formats and valid luminaire identifiers, and discarding invalid or abnormal commands to ensure the accuracy of command transmission.

[0146] Because different brands and types of target lighting fixtures support different communication protocols (e.g., Hue series lighting fixtures support ZigBee, ordinary smart lights support Wi-Fi, and BLE light strips support Bluetooth), and the ambient light synchronization commands sent by the playback terminal are in a unified standardized format, they cannot be directly recognized and executed by the target lighting fixtures. Therefore, the gateway needs to determine the communication protocol type supported by each target lighting fixture based on the preset mapping relationship between lighting fixture identifiers and communication protocols. It then converts the received standardized ambient light synchronization commands into a command format compatible with the communication protocol of that target lighting fixture, ensuring that core information such as the lighting fixture drive color (x, y, L_target) and target activation time Ttarget in the command are not lost or distorted, thus achieving protocol adaptation between the playback terminal and the target lighting fixtures.

[0147] Subsequently, to resolve the clock discrepancy issue between the gateway, the playback terminal, and the target lights, and to prevent the lights' activation time from deviating from the target value and causing asynchrony between the lights and the video image due to clock asynchrony, the gateway calibrates the target activation time Ttarget in the ambient light synchronization command based on a preset clock synchronization mechanism (such as the NTP network time protocol). Specifically, the gateway synchronizes its own clock with the standard network time in real time, calculates the deviation value ΔT between the gateway clock and the playback terminal clock and the target light clock, and corrects Ttarget based on the deviation value to obtain the calibrated target activation time Ttarget_calibrated = Ttarget + ΔT, ensuring that all target lights can accurately execute color switching commands under a unified time reference.

[0148] Finally, the gateway sends the format-converted and time-calibrated ambient light synchronization command to the corresponding target luminaire via the corresponding communication module (ZigBee module, Wi-Fi module, Bluetooth module, etc.) according to the communication protocol requirements of the target luminaire. During the transmission process, the gateway records the command transmission time, luminaire identification, and command sequence number, which can be used for subsequent command packet loss detection and retransmission to ensure that each target luminaire can accurately receive and execute the ambient light synchronization command, ultimately achieving collaborative synchronous control of multiple terminals and multiple luminaires.

[0149] The gateway can monitor and receive ambient light synchronization commands from at least one playback terminal in real time through preset communication ports (such as UDP port 8888 and TCP port 2100). These commands are structured command packets (Command_Packet) generated by the playback terminal based on video color analysis. They may include at least the target luminaire identifier (device_id), luminaire driving color (x, y, L_target), and target activation time (Ttarget), and may also include auxiliary information such as command sequence number (sequence_id) and luminaire color transition duration (transition_time). The gateway can perform format verification and validity validation on the received commands. Format validation: Verify the completeness of the fields in the instruction package (e.g., whether it contains the required fields device_id and Ttarget) and the validity of the data format (e.g., whether the x / y coordinates are in the range of 0 to 1 and whether L_target is in the range of 0 to L_max). Validity verification: Match the local preset whitelist of lamps, filter out commands with valid lamp identifiers, and discard abnormal commands with non-existent identifiers, offline status, or incompatible permissions; After successful verification, the gateway records the instruction reception timestamp T_receive_gw, providing basic data for subsequent RTT measurement and time calibration, and ensuring the accuracy and effectiveness of instruction transmission.

[0150] The gateway can pre-store the mapping relationship between target luminaire identifiers and communication protocols, and match the corresponding protocol type according to the luminaire identifier in the instruction: If the target lighting fixture supports Hue Entertainment low-latency streaming mode, the standardized instructions will be converted into the HueEntertainment protocol format, based on the instruction structure adapted for DTLS long connection, where the DTLS long connection is maintain_dtls_connection(device_ip,port=2100); If the target lighting fixture supports the UDP protocol, it will be converted to the custom UDP_Custom protocol format, conforming to the packet encapsulation rules of udp_socket(device_ip,port=8888); If the target lighting fixture supports the Bluetooth protocol, it is converted to the BLE_GATT protocol format, matching the GATT service data structure of ble_connect(device_mac,service_uuid).

[0151] During the format conversion process, the gateway only adapts the header encoding and data segment arrangement of the instructions, strictly ensuring that the accuracy of the lamp driving color (x,y,L_target) is not lost and the value of the target effective time Ttarget is not distorted. At the same time, it supplements the fixed fields required by the protocol (such as the entertainment mode flag bit of the Hue protocol and the checksum field of the UDP protocol) to achieve seamless adaptation between the unified instruction format of the playback terminal and the differentiated protocols of the lamps.

[0152] The gateway can accurately calibrate the target effective time Ttarget in the ambient light synchronization command based on a dual mechanism of NTP clock synchronization and RTT dynamic compensation. Basic clock synchronization: The gateway synchronizes its own clock with the standard network time in real time through the NTP network time protocol, ensuring that the clock reference deviation between the gateway clock and the playback terminal and the target lights is controlled within 1ms; RTT Latency Compensation: Every 10 frames of instructions received by the gateway, a timestamp echo request packet (containing the gateway's sending time T_send) is sent to the playback terminal. After receiving the echo response, the round-trip time RTT = T_receive - T_send is calculated, and the RTT is stored in a sliding window history array (RTT_history) of length 50. The median of RTT_history, RTT_median, is taken, and the one-way network latency T_network = RTT_median / 2 is calculated. Time calibration calculation: The gateway combines the clock deviation value ΔT_clock (the clock difference between the gateway and the target light) and the one-way network delay T_network to correct Ttarget, and obtain the calibrated target effective time: Ttarget_calibrated=Ttarget+T_network+ΔT_clock; Range limitation: The calibrated Ttarget_calibrated is limited to the range of [Ttarget-10ms, Ttarget+50ms] to avoid abnormal calibration values ​​due to network jitter and ensure that all target lights can accurately execute color switching commands under a unified time reference.

[0153] The gateway will send the ambient light synchronization command, after format conversion and time calibration, to the corresponding target luminaire through the corresponding communication module (DTLS module, UDP module, BLE module, etc.) according to the communication protocol requirements of the target luminaire, and execute a multi-dimensional communication guarantee mechanism: Connection maintenance: The gateway monitors the connection status with the lights in real time. If no data packet is sent for 500 ms, it automatically sends a heartbeat packet (send_heartbeat) to maintain a long connection and avoid command transmission failure due to connection loss. Bandwidth latency adaptation: If RTT_median > 20ms (network congestion) or packet loss rate > 5% is detected, the gateway triggers an adaptation strategy—increasing the command sending advance by 10ms (send_advance_time = Δpipe + 10ms) and sending redundant packets; if RTT_median > 50ms, the command update frequency is further reduced from 60fps to 30fps, while the x / y coordinates are quantized to 3 decimal places and L_target is quantized to 10 levels to reduce the data packet size; Sending records and retransmissions: The gateway records the sending time, lamp identifier, instruction sequence number (sequence_id), and sending status (success / failure / pending confirmation) of each instruction; if the packet loss rate is detected to be greater than 10%, the unconfirmed instructions are retransmitted, and the retransmission interval is dynamically adjusted according to the packet loss rate (the higher the packet loss rate, the shorter the retransmission interval). Multi-lamp collaborative transmission: For multi-lamp commands that need to be executed synchronously, the gateway sends commands in batches according to the time requirement of Ttarget_calibrated, Δpipe duration in advance, to ensure that all lamps receive the commands at the same time and take effect at the target time, ultimately realizing collaborative synchronous control of multiple terminals and multiple lamps.

[0154] In this embodiment, by uniformly receiving ambient light synchronization commands from playback terminals at the gateway, centralized scheduling among multiple terminals and multiple lights can be achieved, improving the overall system synergy and scalability. By converting the ambient light synchronization commands into a communication protocol format supported by the target lights, compatibility with lighting devices of different brands and communication methods (such as low-latency entertainment mode, UDP, and Bluetooth) can be achieved, significantly improving the versatility and adaptability of the solution. By calibrating the target effective time based on a clock synchronization mechanism, network latency, clock skew, and transmission jitter can be effectively compensated, ensuring precise time alignment between the lights and video images, significantly improving the accuracy and smoothness of ambient light synchronization. The gateway completes protocol conversion and timing calibration before issuing commands, reducing the processing pressure and protocol adaptation complexity of the playback terminals, which is conducive to achieving stable, low-latency ambient light synchronization on more terminal devices. Through the gateway's centralized forwarding and reliable transmission mechanism, the command delivery rate is improved, reducing asynchrony problems caused by packet loss and latency, ensuring consistent control of multiple lights, and further enhancing the immersive viewing experience.

[0155] Figure 13 This is a flowchart illustrating an ambient light synchronization method based on video content color analysis, provided in another embodiment of this application. This method can be applied to target lighting fixtures and may include: S301: Receive an ambient light synchronization command sent from a playback terminal or gateway, wherein the ambient light synchronization command includes at least the lamp driving color and the target effective time of the lamp driving color; S302: When the local system time reaches the time window of the target effective time, the current light emission state is synchronized with the color driven by the lamp. The transition process of synchronizing the current light emission state with the color driven by the lamp is calculated by interpolation using the shortest path on the color wheel.

[0156] In some embodiments, the interpolation calculation using the shortest path on the color wheel in S302 above may, in specific implementation, include: The current luminous state and the color driven by the lamp are converted to the first color space to obtain the current hue and the target hue, respectively. Calculate the difference between the target hue and the current hue, and adjust the difference to the shortest angular distance on the hue wheel; Based on the adjusted shortest angular distance, interpolation is performed within a preset transition time to complete the change from the current hue to the target hue.

[0157] Specifically, the target luminaire uses its corresponding communication module (DTLS module, UDP module, BLE module, etc., matching the gateway / playback terminal protocol) to listen for and receive ambient light synchronization commands sent from the playback terminal or gateway in real time. The ambient light synchronization commands are structured commands after format conversion and time calibration (if from the gateway), or standardized commands sent directly by the playback terminal (if directly connected). The commands include at least the luminaire driving color (x, y, L_target) and the target effective time Ttarget_calibrated. The target effective time Ttarget_calibrated is calculated at least based on the video frame display timestamp Tdisplay and the estimated end-to-end latency Δpipe, and may also include auxiliary information such as the command sequence number sequence_id and the transition time transition_time (default 30ms). The luminaire performs preliminary verification on the received commands, verifying the legality of the command format and the range of the luminaire driving color parameters (x / y coordinates 0~1 interval, brightness L_target, 0~L_max range), filtering out valid commands and discarding abnormal and invalid commands to ensure that the commands are executable.

[0158] The luminaire synchronizes its local system time with the standard network time in real time (echoing the NTP clock synchronization mechanism on the gateway side), continuously monitoring the difference between the local system time and the target effective time Ttarget_calibrated. When the local system time reaches the time window of the target effective time (Ttarget_calibrated ± 1ms, compatible with micro-hour time deviation), the luminous state switching is initiated, synchronizing the current luminous state with the luminaire's driving color. The luminous state transition process uses interpolation calculation based on the shortest path on the color wheel, and in specific implementations, it may include: 1. Color Space Conversion: Convert the color parameters (x_current, y_current, L_current) of the current luminaire's emission state and the luminaire driving color (x_target, y_target, L_target) in the command to the HSV first color space. Extract the current hue H_current and the target hue H_target respectively, while retaining the current saturation S_current, brightness V_current and the target saturation S_target, brightness V_target.

[0159] 2. Shortest Angular Distance Calculation: Calculate the original difference ΔH between the target hue H_target and the current hue H_current, ΔH = H_target - H_current. Since the hue wheel is a 360° circular distribution, adjust the original difference ΔH to the shortest angular distance on the hue wheel. The specific adjustment rules are as follows: if ΔH > 180°, then the shortest angular distance ΔH_short = ΔH - 360°; if ΔH < -180°, then the shortest angular distance ΔH_short = ΔH + 360°; if |ΔH| ≤ 180°, then the shortest angular distance ΔH_short = ΔH. This ensures that the interpolation path is the shortest path on the hue wheel, avoiding unnecessary color transition redundancy.

[0160] 3. Interpolation Transition Execution: Based on the adjusted shortest angular distance ΔH_short, combined with the preset transition time transition_time in the instruction (default 30ms), hue interpolation is performed uniformly within the transition time, while saturation and brightness are simultaneously interpolated linearly. This achieves a smooth transition from the current luminous state to the target lamp-driven color, avoiding visual discomfort caused by sudden color changes and meeting the core requirements of lamp-end transition.

[0161] In this embodiment, the luminaire accurately receives and verifies the ambient light synchronization command, ensuring command validity and preventing incorrect or abnormal commands that could cause erratic lighting states. This guarantees the stability of ambient light synchronization and adapts to both gateway forwarding and direct terminal connection modes, enhancing the flexibility of the solution. By monitoring the matching of the local system time with the target effective time, the luminaire is ensured to initiate the lighting state switch at the precise time, echoing the timing calibration logic of the gateway / playback terminal, further ensuring the synchronization of light and video images and enhancing the immersive experience. Interpolation calculations are performed using the shortest path on the color wheel, combined with HSV color space interpolation processing, to achieve a smooth transition of luminaire lighting states. This avoids visual discomfort caused by sudden color changes, shortens transition time, reduces color redundancy, and makes light changes more in line with the rhythm of video content, meeting the user's visual comfort needs. Unifying color space conversion and interpolation logic adapts to different formats of luminaire driver color commands, improving the luminaire's ability to adapt to commands. This eliminates the need for additional interpolation processing by the playback terminal / gateway, reducing upstream equipment overhead and ensuring consistent transition effects across multiple luminaires.

[0162] In some embodiments, after S302 above, the following may also be included: During the transition process or when the light is in steady state, the operating parameters of the target luminaire are monitored; If the operating parameters exceed the preset safety threshold, the light emission state of the target lamp will be automatically adjusted to a safety mode, which includes reducing brightness or switching to a preset safety color.

[0163] Specifically, during the transition from light emission state in S302, or when the luminaire is in a steady-state light emission state (maintaining the luminaire's driving color remains unchanged), the target luminaire monitors its own core operating parameters in real time to achieve safety protection. This can specifically include: Operating parameter monitoring: The lamp has a built-in monitoring module that continuously collects its own operating parameters, including at least the lamp housing temperature, LED bead current, and brightness output value. The monitoring frequency is 100ms / time to ensure real-time detection of parameter anomalies.

[0164] Safety threshold judgment: The lamp has preset safety thresholds for each working parameter (such as housing temperature safety threshold ≤60℃, LED lamp current safety threshold ≤200mA, brightness safety threshold ≤1000lm). The real-time monitored working parameters are compared with the corresponding safety thresholds to determine whether there are any parameters exceeding the standard.

[0165] Safety Mode Switching: If any operating parameter is detected to exceed the preset safety threshold (e.g., casing temperature reaches 65℃, current reaches 220mA), the safety protection mechanism is automatically triggered, adjusting the light emission state of the target lamp to a safety mode. The safety mode includes, but is not limited to: reducing brightness (reducing the current brightness to the maximum brightness corresponding to the safety threshold, such as reducing it from 1000lm to 800lm), switching to a preset safety color (e.g., low-brightness warm white to avoid strong light stimulation and overheating), and phased shutdown (if the parameters are severely exceeded, the lamp is briefly shut off for 5 seconds and then restored to a low-brightness safety color), until the operating parameters drop to within the safety threshold range, and then automatically restored to the light emission state required by the instruction (if the instruction reception timeout period has not been exceeded), ensuring the safety of the lighting equipment and the user's safety.

[0166] In this embodiment, by monitoring the operating parameters in real time when the luminaire performs color transition and steady-state light emission, and automatically switching to a safety mode that reduces brightness or switches to a preset safety color when the parameters exceed the preset safety threshold, the safety of the luminaire equipment and the safety of the user can be effectively guaranteed, avoiding visual discomfort caused by sudden changes in light and maintaining the immersive atmosphere. At the same time, it can realize local autonomous and rapid protection of the luminaire, reduce dependence on upstream equipment, improve the safety closed loop of the entire ambient light synchronization solution, and enhance the reliability and practicality of the system.

[0167] In some embodiments, after S302 above, the following may also be included: If no new valid instruction is received within the preset instruction reception timeout period since the last valid ambient light synchronization instruction was received, the target luminaire will be controlled to slowly transition from the current luminous state to the preset steady-state safety state.

[0168] Specifically, after the light emission state transition is completed in S302, the target luminaire continuously listens for new valid ambient light synchronization commands and simultaneously starts a command reception timeout timer. In practice, this may include: Timeout start: Upon receiving the last valid ambient light synchronization command, the luminaire starts the local timing module and begins accumulating the command reception silence duration time_since_last_valid_cmd.

[0169] Timeout detection: The lamp has a local preset command reception timeout (the preset value is configurable, the default is 5000ms, which is 5s, echoing the heartbeat cycle logic of connection maintenance in the briefing). The accumulated silence time is compared with the preset timeout in real time to determine whether a command timeout has occurred.

[0170] Steady-state safety transition: If time_since_last_valid_cmd exceeds the preset instruction reception timeout and no new valid ambient light synchronization instruction is received during this period, it is determined that the communication between the playback terminal / gateway is interrupted or the instruction transmission is abnormal. At this time, the lamp automatically controls itself to slowly transition from the current light-emitting state to the preset steady-state safety state according to the preset transition speed (such as a transition time of 1000ms). The preset steady-state safety state includes low brightness constant light (such as 200lm warm white) and soft extinguishing, to avoid the lamp from suddenly being brightly lit or instantly extinguishing due to communication interruption, taking into account both user experience and lamp energy saving safety. If a valid ambient light synchronization instruction is received again in the future, the steady-state safety transition will stop immediately and the light-emitting state will be adjusted according to the new instruction requirements.

[0171] In this embodiment, by monitoring the timeout of ambient light synchronization commands and gradually transitioning the luminaires to a steady-state safety state after the timeout, abnormal situations such as communication interruptions and command loss can be effectively addressed, preventing the luminaires from being in an abnormal lighting state. Simultaneously, the smooth transition method enhances the user's visual experience and ensures more stable and reliable system operation.

[0172] An embodiment of this application provides an ambient light synchronization system based on video content color analysis, which may include a playback terminal, a gateway, and at least one target luminaire. The playback terminal is communicatively connected to the target luminaire, and the gateway serves as a communication relay between the playback terminal and the target luminaire. The playback terminal is configured to: obtain the current video frame and its display timestamp from the video stream; perform downsampling processing on the video frame; and generate a pixel weight mask for eliminating interference areas based on the downsampled video frame; extract candidate primary colors of the downsampled video frame based on the downsampled video frame, the pixel weight mask, and a preset spatial weight matrix; perform temporal filtering and perceptual mapping on the candidate primary colors to obtain the lamp-driven color; and generate and send an ambient light synchronization command to at least one target lamp based on the lamp-driven color and its target effective time, instructing the target lamp to synchronize its illumination state to the lamp-driven color at the target effective time, wherein the target effective time is calculated based on the display timestamp and the estimated end-to-end latency; and generate and send an ambient light synchronization command to at least one target lamp to instruct the target lamp to synchronize its illumination state to the lamp-driven color at the target effective time, wherein the target effective time is calculated based at least on the display timestamp. The gateway is configured to: receive an ambient light synchronization instruction from at least one playback terminal, the ambient light synchronization instruction including at least the luminaire driving color and the target effective time of the luminaire driving color; convert the received ambient light synchronization instruction into a format compatible with the communication protocol supported by the target luminaire; calibrate the target effective time based on a clock synchronization mechanism; and send the format-converted and time-calibrated ambient light synchronization instruction to the corresponding target luminaire. The target luminaire is used to: receive an ambient light synchronization command sent from a playback terminal or gateway, the ambient light synchronization command including at least the luminaire driving color and the target effective time of the luminaire driving color; when the local system time reaches the time window of the target effective time, synchronize the current luminous state to the luminaire driving color, the transition process of synchronizing the current luminous state to the luminaire driving color is calculated by interpolation using the shortest path on the color wheel.

[0173] Specifically, the system adopts a distributed architecture of "playback terminal - gateway - target luminaire". The playback terminal, as the instruction generation core, is responsible for analyzing video content to extract the luminaire driving color and generating synchronization instructions. The gateway, as the relay adaptation core, is responsible for protocol conversion, time calibration, and reliable instruction forwarding. The target luminaire, as the execution core, is responsible for accurately receiving instructions and smoothly switching the illumination state. The three establish collaboration through a preset communication protocol to achieve a closed-loop control of "video color analysis - instruction generation - instruction relay - state switching". It can adapt to flexible deployment scenarios where a single playback terminal corresponds to multiple target luminaires or multiple playback terminals correspond to multiple target luminaires, taking into account both versatility and scalability.

[0174] For details regarding the implementation of the playback terminal, gateway, and target lighting fixtures, please refer to the preceding descriptions; this manual will not elaborate further.

[0175] In this embodiment, the system achieves a closed-loop synchronization of the entire process—from video color analysis, primary color extraction, instruction generation, protocol transfer, timing calibration, and smooth color change of the lights—through the coordinated operation of the playback terminal, gateway, and target lights. On the terminal side, downsampling, mask weighting, and temporal filtering are used to extract stable, human-perceived light-driving colors. On the gateway side, multi-protocol compatibility, clock synchronization, and low-latency reliable forwarding are achieved, resolving cross-device and cross-protocol adaptation and timing deviation issues. On the light fixture side, the shortest hue path interpolation is used to achieve a smooth transition, and local security protection and instruction timeout exception handling capabilities are provided. It is compatible with various smart lights and communication methods, maintaining a high degree of synchronization between lighting and video images even in complex network environments, significantly enhancing the immersive viewing and entertainment experience while ensuring device operational safety and user experience.

[0176] In a specific implementation scenario, the playback terminal or terminal side of the above system may include the following modules: Module A: Video Decoding and Rendering Module Function: Decode the current frame F from the video stream, obtain its display timestamp Tdisplay (PTS), and prepare to pass the frame data to the color extraction module.

[0177] Output: Raw YUV / RGB frame F + display timestamp Tdisplay.

[0178] Logical association: Pass the output frames F and Tdisplay to module B for color analysis.

[0179] Module B: Color Extraction Module (GPU / NEON / SIMD Accelerated)

[0180] Function: Receive frame F output by module A, perform low-resolution sampling (1 / 8~1 / 32), generate a mask (masking subtitles / skin color / highlight areas), calculate the HSV hue histogram, and extract candidate primary colors.

[0181] Input: Raw frame F (from module A).

[0182] Output: Candidate primary color (hue θ, saturation S_raw, brightness V_raw) + inter-frame motion intensity M.

[0183] Logical association: Pass the candidate primary color and motion intensity M to module C for temporal stabilization.

[0184] Module C: Color Stabilization and Mapping Module

[0185] Function: Receives candidate primary colors from module B, applies temporal filtering (IIR) to smooth color transitions, performs perceptual mapping (color gamut conversion, Gamma correction, brightness normalization, and user preference adjustment), and outputs the final target color for the lamp.

[0186] Input: Candidate primary color (θ, S_raw, V_raw) + motion intensity M (from module B) + previous frame stable color (θ_prev, S_prev, V_prev).

[0187] Output: Stable target color for the luminaire (HRGB or CIE xyY format + brightness L_target).

[0188] Logical association: Pass the stable target color to module D for time synchronization scheduling.

[0189] Module D: Synchronization and Scheduling Module

[0190] Function: Receives the target color output by module C and Tdisplay of module A, calculates the target effective time Ttarget=Tdisplay+Δpipe, generates multi-lamp bit instructions (global color or region mapping), and decides whether to send (based on color difference threshold and frame rate adaptation).

[0191] Input: Target color (from module C) + Tdisplay (from module A) + network RTT estimate (from module E).

[0192] Output: Multi-lamp instruction packet {lamp ID, target color, Ttarget, sequence_id}.

[0193] Logical association: Pass the instruction packet to module E for transmission over the network.

[0194] Module E: Communication Module

[0195] Function: Receive instruction packets generated by module D and send them to the gateway or lighting fixtures via low-latency protocols (Hue Entertainment / UDP / BLEGATT) to maintain long-term connections, measure RTT, and implement bandwidth degradation strategies.

[0196] Input: Multi-lamp instruction package (from module D).

[0197] Output: Network RTT estimate (feedback to module D) + confirmation sent.

[0198] Logical connection: Establish bidirectional communication with the gateway / lighting fixture side, and feed back RTT data to module D to optimize Δpipe.

[0199] The gateway / bridge side of the above system (optionally, such as Philips Hue Bridge) may include the following modules: Module F: Protocol Adaptation and Clock Alignment Function: Receives instruction packets from terminals, converts them into the native protocols of various lighting fixtures (Zigbee / BLE / Wi-Fi), synchronizes clocks (NTP / PTP), and aggregates instructions from multiple terminals (to avoid conflicts).

[0200] Input: Terminal command packet {lamp ID, target color, Ttarget, sequence_id}.

[0201] Output: Adapted lighting instructions + clock offset correction.

[0202] Logical association: The adapted instructions are sent to each lamp.

[0203] The lighting fixture side (smart bulb / strip) of the above system may include the following modules: Module G: Color Interpolation and Transition Function: Receives instruction packets, interpolates between the current color and the target color using the shortest hue path, and smoothly transitions within the [Ttarget-ε, Ttarget+ε] time window (20-80ms).

[0204] Input: target color + Ttarget (from gateway or terminal).

[0205] Output: Real-time drive current (PWM / constant current) controls the LED.

[0206] Logical association: Execute the color change and simultaneously feed back the status to module H for safety monitoring.

[0207] Module H: Gamma / Power Limiting and Fault Fallback

[0208] Function: Monitors lamp temperature and power, automatically reduces brightness after prolonged high brightness, detects abnormalities, disconnects and returns to a safe color (such as warm white).

[0209] Input: Current brightness L + temperature T + connection status (from module G)

[0210] Output: Limited brightness L_safe + fault flag

[0211] Logical correlation: Implement safety constraints on the output of module G to ensure that the hardware is not overloaded.

[0212] The data flow between the above modules can be as follows: Video decoding and rendering module A (frame + timestamp) → Color extraction module B (primary color extraction) → Color stabilization and mapping module C (stabilization mapping) → Synchronization and scheduling module D (synchronization scheduling) → Communication module E (communication transmission) → Protocol adaptation and clock alignment module F (protocol adaptation) → Color interpolation and transition module G (color transition) → Power limiting and fault fallback module H (safety limiting).

[0213] Based on the above embodiments, this invention intervenes in the "early stages" of the video processing chain (post-decoding buffer or rendering pipeline front-end), and systematically solves five major problems of existing technologies through a three-layer mechanism of "lightweight color extraction + temporal stabilization + predictive synchronization": 1. To address the latency issue: Frame data is directly acquired on the post-decoding buffer side (bypassing screen rendering), and low-resolution sampling (1 / 8~1 / 32) is accelerated by GPU / NEON hardware, which compresses the color calculation time to 2-6ms; combined with predictive timestamp (Ttarget=Tdisplay+Δpipe) and the pre-soft buffering mechanism of the lamps, the end-to-end latency is controlled within 40-60ms, reaching the threshold of imperceptibility to the human eye.

[0214] 2. To address the issue of color instability: Construct a triple stabilization mechanism of "mask + spatial weighting + temporal filtering": eliminate interference areas through subtitle / skin tone / highlight masks, adopt spatial weighting with priority given to the center area of ​​the image, and then suppress high-frequency jitter through exponential sliding filter (IIR) to ensure that the extracted main color reflects the main color of the image narrative rather than noise.

[0215] 3. To address cross-device synchronization issues: Design a unified timestamp alignment protocol (based on NTP / PTP clock synchronization). All lamps use "absolute time Ttarget" as the execution benchmark, rather than relying on the device's own clock. Issue instructions to slow-responding devices in advance and extend the transition time for fast devices to achieve a multi-lamp consistency error of <15ms.

[0216] 4. Regarding the computational energy consumption issue: a heuristic hue histogram (36-72 bins) is used to replace the complete k-means clustering, reducing the computational complexity from O(n·k·iter) to O(n / s²+N) (s is the sampling ratio, N is the number of bins); dynamic frame rate rollback (60→30→15Hz) and differential updates (sent only when the color difference is greater than the threshold) are supported, reducing CPU / network load.

[0217] 5. Regarding the color gamut mapping issue: Establish a segmented mapping chain of "video color gamut → CIE XYZ → lighting color gamut": First, convert BT.709 / BT.2020 to the device-independent CIE XYZ space, then perform minimum error clipping based on the color gamut triangle of each lighting fixture (prioritizing hue and sacrificing saturation), and apply Gamma correction and brightness limiting to ensure visual comfort.

[0218] Based on the above-described ambient light synchronization method based on video content color analysis, one or more embodiments of this specification also provide an ambient light synchronization device based on video content color analysis. The device may include apparatus (including distributed systems), software (applications), modules, plug-ins, servers, clients, etc., using the methods described in the embodiments of this specification, combined with necessary implementation hardware. Based on the same innovative concept, the devices in one or more embodiments provided in this specification are as described in the following embodiments. Since the implementation schemes and methods for solving the problem are similar, the implementation of specific devices in the embodiments of this specification can refer to the implementation of the foregoing methods, and repeated details will not be elaborated further. As used below, the terms "unit" or "module" can refer to a combination of software and / or hardware that implements a predetermined function. Although the devices described in the following embodiments are preferably implemented in software, hardware implementation, or a combination of software and hardware, is also possible and contemplated. Figure 14 This is a schematic diagram of the structural composition of an ambient light synchronization device based on video content color analysis according to an embodiment of this application, as shown below. Figure 14As shown, this ambient light synchronization device based on video content color analysis, applied to a playback terminal, may include: The frame processing and mask generation module 1401 can be used to obtain the current video frame and its display timestamp from the video stream, perform downsampling processing on the video frame, and generate a pixel weight mask for excluding interference areas based on the downsampled video frame. The candidate primary color extraction module 1402 can be used to extract the candidate primary color of the downsampled video frame based on the downsampled video frame, the pixel weight mask and the preset spatial weight matrix; The temporal filtering and mapping module 1403 can be used to perform temporal filtering and perception mapping on the candidate primary color to obtain the lamp driving color. The instruction generation and sending module 1404 can be used to generate and send an ambient light synchronization instruction to at least one target luminaire based on the luminaire driving color and the target effective time of the luminaire driving color, so as to instruct the target luminaire to synchronize its light emission state to the luminaire driving color at the target effective time. The target effective time is calculated based on the display timestamp and the estimated end-to-end delay.

[0219] Figure 15 This is a schematic diagram of the structural composition of an ambient light synchronization device based on video content color analysis according to an embodiment of this application, as shown below. Figure 15 As shown, this ambient light synchronization device based on video content color analysis, applied to a gateway, may include: The first receiving module 1501 can be used to receive an ambient light synchronization instruction from at least one playback terminal, wherein the ambient light synchronization instruction includes at least the lamp driving color and the target effective time of the lamp driving color. The conversion module 1502 can be used to convert the received ambient light synchronization command into a format compatible with the communication protocol supported by the target luminaire; The calibration module 1503 can be used to calibrate the effective time of the target based on a clock synchronization mechanism; The sending module 1504 can be used to send the ambient light synchronization command, which has been format-converted and time-calibrated, to the corresponding target luminaire.

[0220] Figure 16 This is a schematic diagram of the structural composition of an ambient light synchronization device based on video content color analysis according to an embodiment of this application, as shown below. Figure 16 As shown, this ambient light synchronization device based on video content color analysis, applied to target lighting fixtures, may include: The second receiving module 1601 can be used to receive an ambient light synchronization instruction sent from a playback terminal or gateway. The ambient light synchronization instruction includes at least the lamp driving color and the target effective time of the lamp driving color. The synchronization module 1602 can be used to synchronize the current luminous state with the color driven by the lamp when the local system time reaches the time window where the target effective time is located. The transition process of synchronizing the current luminous state with the color driven by the lamp is calculated by interpolation using the shortest path on the color wheel.

[0221] The descriptions and functions of the above modules can be found in the section on ambient light synchronization methods based on video content color analysis, and will not be repeated here.

[0222] This application also provides an electronic device, such as... Figure 17 As shown, the electronic device may include a processor 1701 and a memory 1702, wherein the processor 1701 and the memory 1702 may be connected via a bus or other means. Figure 17 Taking the example of a connection between China and Israel via a bus.

[0223] Processor 1701 can be a central processing unit (CPU). Processor 1701 can also be other general-purpose processors, digital signal processors (DSPs), application-specific integrated circuits (ASICs), field-programmable gate arrays (FPGAs), or other programmable logic devices, discrete gate or transistor logic devices, discrete hardware components, or combinations of the above types of chips.

[0224] The memory 1702, as a non-transitory computer-readable storage medium, can be used to store non-transitory software programs, non-transitory computer-executable programs, and modules, such as the program instructions / modules corresponding to the audio modulation method based on physiological signals in the embodiments of the present invention. The processor 1701 executes various functional applications and data processing of the processor by running the non-transitory software programs, instructions, and modules stored in the memory 1702, thereby implementing the audio modulation method based on physiological signals in the above method embodiments.

[0225] Memory 1702 may include a program storage area and a data storage area. The program storage area may store the operating system and applications required for at least one function; the data storage area may store data created by processor 1701, etc. Furthermore, memory 1702 may include high-speed random access memory and may also include non-transitory memory, such as at least one disk storage device, flash memory device, or other non-transitory solid-state storage device. In some embodiments, memory 1702 may optionally include memory remotely located relative to processor 1701, and these remote memories may be connected to processor 1701 via a network. Examples of such networks include, but are not limited to, the Internet, intranets, local area networks, mobile communication networks, and combinations thereof.

[0226] The one or more modules are stored in the memory 1702 and are executed by the processor 1701. Figure 1 , Figure 12 , Figure 13 The ambient light synchronization method based on video content color analysis described in the embodiments.

[0227] The specific details of the aforementioned electronic device can be understood by referring to the relevant descriptions and effects in the above method embodiments, and will not be repeated here.

[0228] This specification also provides a computer storage medium storing computer program instructions, which, when executed, implement the steps of the above-described ambient light synchronization method based on video content color analysis.

[0229] This specification also provides a computer program product, which includes a computer program that, when executed, implements the steps of the above-described ambient light synchronization method based on video content color analysis.

[0230] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The program can be stored in a computer-readable storage medium, and when executed, it can include the processes of the embodiments of the above methods. The storage medium can be a magnetic disk, optical disk, read-only memory (ROM), random access memory (RAM), flash memory, hard disk drive (HDD), or solid-state drive (SSD), etc.; the storage medium can also include combinations of the above types of memory.

[0231] The various embodiments in this specification are described in a progressive manner. For the same or similar parts between the various embodiments, please refer to each other. The focus of each embodiment is to describe the differences from other embodiments.

[0232] The systems, devices, modules, or units described in the above embodiments can be implemented by computer chips or entities, or by products with certain functions.

[0233] For ease of description, the above devices are described separately by function as various units. Of course, in implementing this application, the functions of each unit can be implemented in one or more software and / or hardware.

[0234] As can be seen from the above description of the embodiments, those skilled in the art can clearly understand that this application can be implemented by means of software plus necessary general-purpose hardware platforms. Based on this understanding, the technical solution of this application, in essence, or the part that contributes to the prior art, can be embodied in the form of a software product. This computer software product can be stored in a storage medium, such as ROM / RAM, magnetic disk, optical disk, etc., and includes several instructions to cause a computer device (which may be a personal computer, server, or network device, etc.) to execute certain parts of the methods of various embodiments of this application.

[0235] This application can be used in a wide variety of general-purpose or special-purpose computer system environments or configurations. For example: personal computers, server computers, handheld or portable devices, tablet devices, multiprocessor systems, microprocessor-based systems, set-top boxes, programmable consumer electronics devices, network PCs, minicomputers, mainframe computers, distributed computing environments including any of the above systems or devices, etc.

[0236] This application can be described in the general context of computer-executable instructions, such as program modules, that are executed by a computer. Generally, program modules include routines, programs, objects, components, data structures, etc., that perform a specific task or implement a specific abstract data type. This application can also be practiced in distributed computing environments where tasks are performed by remote processing devices connected via a communication network. In distributed computing environments, program modules can reside in local and remote computer storage media, including storage devices.

[0237] The above description is merely a preferred embodiment of this specification and is not intended to limit this specification. Various modifications and variations can be made to the embodiments described herein by those skilled in the art. Any modifications, equivalent substitutions, improvements, etc., made within the spirit and principles of this specification should be included within the scope of protection of this specification.

Claims

1. An ambient light synchronization method based on video content color analysis, characterized in that, Applied to playback terminals, including: The current video frame and its display timestamp are obtained from the video stream, the video frame is downsampled, and a pixel weight mask for eliminating interference areas is generated based on the downsampled video frame. Based on the downsampled video frame, the pixel weight mask, and the preset spatial weight matrix, the candidate primary color of the downsampled video frame is extracted; The candidate primary colors are subjected to temporal filtering and perceptual mapping to obtain the lighting drive colors. Based on the luminaire driving color and the target effective time of the luminaire driving color, an ambient light synchronization command is generated and sent to at least one target luminaire to instruct the target luminaire to synchronize its luminous state to the luminaire driving color at the target effective time. The target effective time is calculated based on the display timestamp and the estimated end-to-end latency.

2. The method according to claim 1, characterized in that, The step of generating a pixel weight mask for excluding interference regions based on the downsampled video frame includes: Detect at least two different types of interference regions in the downsampled video frame, the interference regions including: subtitle or text regions, face or skin color regions, and bright or overexposed regions; Based on the detection results of each interference region and the preset weight assignment strategy, a pixel weight sub-mask corresponding to each interference region is generated. The pixel weight sub-mask includes: subtitle or text mask, face or skin color mask and highlight or overexposure mask. Based on the generated pixel weight sub-masks, they are fused according to a preset fusion rule to generate a pixel weight mask for eliminating interference areas.

3. The method according to claim 2, characterized in that, The step of generating a pixel weight sub-mask corresponding to each interference region based on the detection results of each interference region and a preset weight assignment strategy includes: The pixel values ​​corresponding to the detected subtitle or text region are set as the first weight value to indicate complete exclusion, and the remaining pixel values ​​are set as the second weight value to indicate normal participation in the candidate primary color extraction, so as to obtain the subtitle or text mask corresponding to the subtitle or text region. The pixel values ​​corresponding to the detected face or skin color region are set to a third weight value that represents a reduction in weight. The third weight value is between the first weight value and the second weight value. The remaining pixel values ​​are set to the second weight value to obtain the face or skin color mask corresponding to the face or skin color region. The first pixel value corresponding to the detected bright or overexposed area is set as the first weight value, and the second pixel value corresponding to the detected bright or overexposed area is set as the gradient weight value. The brightness value of the first pixel value is greater than the first brightness threshold and the saturation value is less than the saturation threshold. The brightness value of the second pixel value is between the first brightness threshold and the second brightness threshold. The gradient weight value is determined according to the first brightness threshold and the second brightness threshold, and varies between the first weight value and the second weight value.

4. The method according to claim 2, characterized in that, The generated pixel weight sub-masks are fused according to a preset fusion rule to generate a pixel weight mask for eliminating interference regions, including: For each pixel position in the downsampled video frame, obtain the mask value of all pixel weight sub-masks at each pixel position; Based on the acquired multiple mask values, a fused mask value is calculated according to the preset fusion rule, which is the minimum value rule; The pixel weight mask is generated based on the fusion mask value calculated for each pixel location.

5. The method according to claim 1, characterized in that, The step of extracting candidate primary colors from the downsampled video frames based on the downsampled video frames, the pixel weight mask, and a preset spatial weight matrix includes: The downsampled video frame is converted to the first color space to obtain the hue value, saturation value and brightness value of each pixel in the downsampled video frame; Based on the pixel weight mask value of the pixel weight mask, the preset spatial weight matrix and the saturation value, calculate the contribution weight of each pixel position to the hue histogram; A weighted hue histogram is constructed using the hue values ​​and the contribution weights; Based on the weighted hue histogram, candidate hue, candidate saturation, and candidate brightness of the downsampled video frame are determined and used to form the candidate primary color.

6. The method according to claim 5, characterized in that, The step of constructing a weighted hue histogram using the hue value and the contribution weight includes: The range of hue values ​​is divided into a preset number of bins, and the width of each bin is calculated. The bins are continuous and non-overlapping hue intervals. For each pixel position in the downsampled video frame, obtain the pixel weight mask value of the pixel weight mask at each pixel position. If the pixel weight mask value is greater than 0 and the saturation value is greater than a preset saturation threshold, then determine the corresponding bin index based on the hue value and the bin width. The contribution weights are accumulated into the histogram values ​​of the bins corresponding to the bin indexes to obtain the initial hue histogram. If the cumulative proportion of the peak bins in the initial hue histogram is greater than a preset proportion threshold, then the histogram values ​​of the peak bins are attenuated, and the histogram values ​​of the second peak bins are increased, to obtain the final weighted hue histogram.

7. The method according to claim 5, characterized in that, The step of determining the candidate hue, candidate saturation, and candidate brightness of the downsampled video frame based on the weighted hue histogram includes: Peak detection is performed on the weighted hue histogram, and the hue corresponding to the peak with the highest percentage is determined as the main peak hue; Based on the binning position and binning width corresponding to the main peak color, the candidate hue of the downsampled video frame is determined; Based on the pixel weight mask value and the preset spatial weight matrix, the saturation and brightness values ​​of all pixels within the preset hue range are weighted and summed to determine candidate saturation and candidate brightness, respectively. The preset hue range is centered on the candidate hue.

8. The method according to any one of claims 5-7, characterized in that, The method further includes: If the candidate saturation is less than a preset saturation threshold, the candidate primary color is corrected to a target neutral color determined based on the candidate brightness, and used as the semantically corrected candidate primary color. If the candidate saturation is not less than the preset saturation threshold, and the weighted hue histogram contains at least two peaks that satisfy the preset significance condition, and the difference in the proportion of any two peaks is less than the preset difference threshold, then: Obtain the hue of the stable color output from the previous frame as the reference hue; Calculate the shortest distance on the color wheel between the hue corresponding to each peak that meets the condition and the reference hue; The hue corresponding to the peak with the shortest distance to the reference hue is determined as the corrected candidate hue, and the corrected candidate hue is combined with the candidate saturation and the candidate brightness to form the semantically corrected candidate primary color.

9. The method according to claim 1, characterized in that, The step of performing temporal filtering and perceptual mapping on the candidate primary colors to obtain the lighting drive colors includes: The candidate primary color is subjected to temporal filtering to obtain a stable target color, wherein the stable target color includes stable hue, stable saturation and stable brightness. The stable target color is perceptually mapped from the video color gamut to the target luminaire color gamut to obtain the luminaire-driven color.

10. The method according to claim 9, characterized in that, The step of performing temporal filtering on the candidate primary color to obtain a stable target color includes: Calculate the normalized color difference between the candidate primary color and the stable color of the previous video frame. The normalized color difference is calculated based on the shortest hue difference on the color wheel, the saturation difference, and the brightness difference. The smoothing coefficient of the exponential moving average filter is dynamically adjusted based on the normalized color difference and the motion intensity of the current video frame. Based on the adjusted smoothing coefficient, exponential sliding filtering is performed on the candidate hue, candidate saturation and candidate brightness respectively to obtain a stable target color; The smoothing coefficients of the dynamically adjusted exponential moving average filter include: If the normalized color difference exceeds a preset color difference threshold or the motion intensity exceeds a preset motion threshold, the smoothing coefficient is increased to accelerate the color response; otherwise, the basic filtering coefficient is used to maintain color stability.

11. The method according to claim 9, characterized in that, The step of performing a perceptual mapping from the video color gamut to the target lighting color gamut on the stable target color to obtain the lighting-driven color includes: The stable target color is converted from the first color space to a device-independent second color space; Based on the physical luminous gamut of the target luminaire, the color coordinates in the second color space are clipped to obtain the chromaticity coordinates located within the physical luminous gamut. Based on the brightness of the stable target color and the global brightness statistics of the downsampled video frame, the mapped target brightness is calculated. The rate of change of the mapped target brightness is limited; The driving color of the luminaire is determined based on the chromaticity coordinates after gamut clipping and the target brightness after limiting the rate of change of brightness.

12. The method according to claim 1, characterized in that, The estimated end-to-end latency is calculated as follows: Based on the processing latency estimate, network latency estimate, luminaire response latency estimate, and buffer latency estimate, the full-link latency estimate is calculated. The network latency estimate is determined based on the measured value of the network round-trip latency for communication with the target luminaire. The buffer latency estimate is determined based on the historical statistical fluctuation of the network round-trip latency. The full-link latency estimate is the total estimated latency of the entire process from the acquisition of the video frame to the target luminaire performing illumination state synchronization. Accordingly, the target effective time is calculated based on the displayed timestamp and the estimated end-to-end latency, including: The target effective time of the lamp-driven color is obtained by adding the display timestamp and the estimated end-to-end delay.

13. The method according to claim 1, characterized in that, The generation and transmission of the ambient light synchronization command includes: The perceived color difference is calculated by comparing the driving color of the lamp with the driving color of the last successfully sent lamp. Select the corresponding dynamic threshold or static threshold based on the motion intensity of the current video frame; If the perceived color difference is greater than the selected threshold, or if the maximum silent time has been exceeded since the last successful transmission, an ambient light synchronization command is generated. The ambient light synchronization command includes at least the luminaire driving color and the target effective time. Otherwise, the generation and transmission of the ambient light synchronization command are suppressed.

14. The method according to claim 1, characterized in that, The generation and sending of ambient light synchronization commands to at least one target luminaire includes: Based on the multiple different spatial partitions of the downsampled video frame, determine the primary color of each partition; Based on the preset mapping relationship between screen partitions and physical light positions, assign exclusive light-driving colors to different target lights based on the primary color of the corresponding partition; Based on the unique luminaire driver color of each target luminaire, a corresponding ambient light synchronization command is generated and sent.

15. An ambient light synchronization method based on video content color analysis, characterized in that, Applied to gateways, including: Receive an ambient light synchronization command from at least one playback terminal, the ambient light synchronization command including at least the lamp driving color and the target effective time of the lamp driving color; The received ambient light synchronization command is converted into a format compatible with the communication protocol supported by the target luminaire; The effective time of the target is calibrated based on a clock synchronization mechanism; The ambient light synchronization command, after format conversion and time calibration, is sent to the corresponding target luminaire.

16. An ambient light synchronization method based on video content color analysis, characterized in that, Applied to target lighting fixtures, including: Receive an ambient light synchronization command sent from a playback terminal or gateway, wherein the ambient light synchronization command includes at least the lamp driving color and the target effective time of the lamp driving color; When the local system time reaches the time window of the target effective time, the current luminous state is synchronized with the color driven by the lamp. The transition process of synchronizing the current luminous state with the color driven by the lamp is calculated by interpolation using the shortest path on the color wheel.

17. The method according to claim 16, characterized in that, The method of interpolating using the shortest path on the color wheel includes: The current luminous state and the color driven by the lamp are converted to the first color space to obtain the current hue and the target hue, respectively. Calculate the difference between the target hue and the current hue, and adjust the difference to the shortest angular distance on the hue wheel; Based on the adjusted shortest angular distance, interpolation is performed within a preset transition time to complete the change from the current hue to the target hue.

18. The method according to claim 16, characterized in that, The method further includes: During the transition process or when the light is in steady state, the operating parameters of the target luminaire are monitored; If the operating parameters exceed the preset safety threshold, the light emission state of the target lamp will be automatically adjusted to a safety mode, which includes reducing brightness or switching to a preset safety color.

19. The method according to claim 16, characterized in that, The method further includes: If no new ambient light synchronization command is received within a preset command reception timeout period since the last valid ambient light synchronization command was received, the target luminaire will be controlled to slowly transition from its current luminous state to a preset steady-state safety state.

20. An ambient light synchronization system based on video content color analysis, characterized in that, It includes a playback terminal, a gateway, and at least one target light fixture, wherein the playback terminal is communicatively connected to the target light fixture, and the gateway serves as a communication relay between the playback terminal and the target light fixture; The playback terminal is used to: obtain the current video frame and its display timestamp from the video stream, perform downsampling processing on the video frame, and generate a pixel weight mask for eliminating interference areas based on the downsampled video frame; Based on the downsampled video frame, the pixel weight mask, and the preset spatial weight matrix, the candidate primary color of the downsampled video frame is extracted; The candidate primary colors are subjected to temporal filtering and perceptual mapping to obtain the lighting drive colors. Based on the luminaire driving color and the target effective time of the luminaire driving color, an ambient light synchronization command is generated and sent to at least one target luminaire to instruct the target luminaire to synchronize its light emission state to the luminaire driving color at the target effective time. The target effective time is calculated based on the display timestamp and the estimated end-to-end latency. The gateway is configured to: receive an ambient light synchronization instruction from at least one playback terminal, the ambient light synchronization instruction including at least the luminaire driving color and the target effective time of the luminaire driving color; convert the received ambient light synchronization instruction into a format compatible with the communication protocol supported by the target luminaire; calibrate the target effective time based on a clock synchronization mechanism; and send the format-converted and time-calibrated ambient light synchronization instruction to the corresponding target luminaire. The target luminaire is used to: receive an ambient light synchronization command sent from a playback terminal or gateway, wherein the ambient light synchronization command includes at least the luminaire driving color and the target effective time of the luminaire driving color; When the local system time reaches the time window of the target effective time, the current luminous state is synchronized with the color driven by the lamp. The transition process of synchronizing the current luminous state with the color driven by the lamp is calculated by interpolation using the shortest path on the color wheel.

21. An ambient light synchronization device based on video content color analysis, characterized in that, Applied to playback terminals, including: The frame processing and mask generation module is used to obtain the current video frame and its display timestamp from the video stream, perform downsampling processing on the video frame, and generate a pixel weight mask for eliminating interference areas based on the downsampled video frame. The candidate primary color extraction module is used to extract the candidate primary color of the downsampled video frame based on the downsampled video frame, the pixel weight mask, and the preset spatial weight matrix. The temporal filtering and mapping module is used to perform temporal filtering and perceptual mapping on the candidate primary color to obtain the lamp driving color. The instruction generation and sending module is used to generate and send an ambient light synchronization instruction to at least one target luminaire based on the luminaire driving color and the target effective time of the luminaire driving color, so as to instruct the target luminaire to synchronize its light emission state to the luminaire driving color at the target effective time. The target effective time is calculated based on the display timestamp and the estimated end-to-end latency.

22. An ambient light synchronization device based on video content color analysis, characterized in that, Applied to gateways, including: The first receiving module is configured to receive an ambient light synchronization instruction from at least one playback terminal, wherein the ambient light synchronization instruction includes at least the lamp driving color and the target effective time of the lamp driving color. A conversion module is used to convert the received ambient light synchronization command into a format compatible with the communication protocol supported by the target luminaire; The calibration module is used to calibrate the effective time of the target based on a clock synchronization mechanism. The sending module is used to send the ambient light synchronization command, which has been format-converted and time-calibrated, to the corresponding target luminaire.

23. An ambient light synchronization device based on video content color analysis, characterized in that, Applied to target lighting fixtures, including: The second receiving module is used to receive an ambient light synchronization instruction sent from a playback terminal or gateway. The ambient light synchronization instruction includes at least the lamp driving color and the target effective time of the lamp driving color. The synchronization module is used to synchronize the current luminous state with the color driven by the lamp when the local system time reaches the time window where the target effective time is located. The transition process of synchronizing the current luminous state with the color driven by the lamp is calculated by interpolation using the shortest path on the color wheel.

24. An electronic device, characterized in that, include: A memory and a processor, the processor and the memory being communicatively connected to each other, the memory storing computer instructions, the processor executing the computer instructions to implement the steps of the method according to any one of claims 1 to 19.

25. A computer-readable storage medium, characterized in that, The computer-readable storage medium stores computer program instructions that, when executed by a processor, implement the steps of the method according to any one of claims 1 to 19.