A set-top box control method, device, equipment and medium

CN122601893APending Publication Date: 2026-08-18SHENZHEN SKYWORTH DIGITAL TECH CO LTD
View PDF 0 Cites 0 Cited by

Patent Information

Application Number
CN202610599461.8
Authority / Receiving Office
CN · China
Patent Type
Applications(China)
Current Assignee / Owner
Filing Date
2026-04-30
Publication Date
2026-08-18

AI Technical Summary

Technical Problem

[0004]本发明提供一种机顶盒控制方法、装置、计算机设备及存储介质,以解决现有技术中机顶盒智能化不足、无法根据观看对象状态自适应调节的问题

Benefits of technology

[0015] This application provides a set-top box control method, apparatus, device, and medium, comprising: acquiring image data of a viewed object; determining the state information of the viewed object based on the image data; and adjusting the device operating parameters of the set-top box based on the state information of the viewed object. This application acquires image data of the viewed object and determines its state information accordingly, thereby adaptively adjusting the device operating parameters of the set-top box. This enables the set-top box to achieve intelligent automatic control based on the user's actual viewing state, allowing parameter adjustment to be completed without the user manually operating the remote control, effectively improving the convenience and smoothness of human-computer interaction.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN122601893A_ABST
    Figure CN122601893A_ABST
Patent Text Reader

Abstract

The application relates to the field of set top box control, and discloses a set top box control method, device, equipment and medium, which comprises the following steps: acquiring image data of a watching object; determining state information of the watching object based on the image data; and adjusting a device running parameter of a set top box based on the state information of the watching object. The application solves the problems of insufficient intelligence of a set top box in the prior art and the inability to adaptively adjust according to the state of a watching object.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This invention relates to the field of set-top box control, and more particularly to a set-top box control method, apparatus, equipment, and medium. Background Technology

[0002] With the widespread adoption of smart TVs and set-top boxes, home entertainment scenarios are becoming increasingly diverse, and user viewing behavior has broken through the limitations of the traditional seated mode. In the home environment, users often adopt relaxed postures such as lying down, side-lying, or prone, and viewing distances vary significantly, ranging from close (within one meter) to far (over three meters), and include both solo viewing and group viewing scenarios. However, the current human-computer interaction system of set-top boxes still revolves around remote control button operation. While some high-end devices integrate voice control, the overall interaction logic remains static and fixed. Regardless of the user's posture or distance, the size of user interface elements, audio output intensity, and operating mode remain uniformly configured, leading to multi-dimensional adaptation imbalances. When users operate in a lying position, the angular deviation between the remote control and the receiving device significantly reduces button recognition accuracy; in far-distance viewing scenarios, the default user interface icons and text sizes are too small, making it difficult for users to clearly identify content; and when viewing at close range, the fixed high volume output can easily cause auditory discomfort. Furthermore, users frequently need to manually navigate to the system settings menu to adjust interface scaling, volume parameters, and interaction mode switching, resulting in a lengthy and costly learning process. Existing technologies, such as image recognition-based smart channel switching systems for set-top boxes, are only applicable to gesture recognition for channel switching. Their application is limited to replacing remote control button operations with single commands, supporting only a passive response mode where users initiate commands. This leaves the set-top box stuck in a rudimentary interaction stage, failing to achieve intelligent adaptation.

[0003] To address the aforementioned issues, existing technologies urgently need improvement. Summary of the Invention

[0004] This invention provides a set-top box control method, device, computer equipment, and storage medium to solve the problems of insufficient intelligence in existing set-top boxes and their inability to adaptively adjust according to the viewing object's status.

[0005] Firstly, a set-top box control method is provided, including: Acquire image data of the object being viewed; Determine the status information of the viewed object based on image data; Adjust the set-top box's operating parameters based on the status information of the viewed object.

[0006] Optionally, the set-top box's operating parameters can be adjusted based on the viewing object's status information, including: The state information of the viewed object is fused to obtain the scene feature vector; Based on scene feature vectors, the corresponding scene strategy is obtained by matching using a pre-set scene strategy library; Adjust the device operating parameters of the set-top box based on scenario-based strategies.

[0007] Optionally, the state information includes attitude information; Determining the state information of the viewed object based on image data includes: Based on image data, obtain the key coordinates of the human body of the viewed object; Based on key human body coordinates, calculate trunk tilt angle, body compression ratio, and hip support status; The posture information of the viewed object is determined based on the torso tilt angle, body compression ratio, and hip support status.

[0008] Optionally, the status information includes distance information; Determining the state information of the viewed object based on image data includes: Identify facial regions in image data; Calculate the proportion of the face region in the image data; Based on preset mapping rules, the percentage is mapped to distance information.

[0009] Optionally, the status information includes distance information; Determining the state information of the viewed object based on image data includes: Based on image data, obtain the key coordinates of the human body of the viewed object; Distance information of the viewed object is calculated based on key human body coordinates.

[0010] Optionally, the status information may also include information on the number of viewers, their age group, or their identity.

[0011] Optionally, adjust the set-top box's operating parameters based on scenario-based strategies, including: If the number of times the same scenario strategy is matched consecutively reaches a preset threshold within a preset time range, the device operating parameters of the set-top box will be adjusted based on the scenario strategy.

[0012] Secondly, a set-top box control device is provided, comprising: The acquisition module is used to acquire image data of the object being viewed; The determination module is used to determine the status information of the viewed object based on image data; The adjustment module is used to adjust the set-top box's operating parameters based on the status information of the viewed object.

[0013] Thirdly, a computer device is provided, including a memory, a processor, and a computer program stored in the memory and executable on the processor, wherein the processor executes the computer program to implement the steps of the set-top box control method described above.

[0014] Fourthly, a computer-readable storage medium is provided, which stores a computer program that, when executed by a processor, implements the steps of the set-top box control method described above.

[0015] This application provides a set-top box control method, apparatus, device, and medium, comprising: acquiring image data of a viewed object; determining the state information of the viewed object based on the image data; and adjusting the device operating parameters of the set-top box based on the state information of the viewed object. This application acquires image data of the viewed object and determines its state information accordingly, thereby adaptively adjusting the device operating parameters of the set-top box. This enables the set-top box to achieve intelligent automatic control based on the user's actual viewing state, allowing parameter adjustment to be completed without the user manually operating the remote control, effectively improving the convenience and smoothness of human-computer interaction. Attached Figure Description

[0016] To more clearly illustrate the technical solutions of the embodiments of the present invention, the drawings used in the description of the embodiments of the present invention will be briefly introduced below. Obviously, the drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0017] Figure 1 This is a flowchart illustrating a set-top box control method provided in an embodiment of the present invention; Figure 2 This is a schematic diagram of the set-top box control device provided in an embodiment of the present invention. Detailed Implementation

[0018] The technical solutions of the embodiments of the present invention will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some, not all, of the embodiments of the present invention. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0019] It should be understood that, when used in this specification and the appended claims, the term "comprising" indicates the presence of the described features, integrals, steps, operations, elements and / or components, but does not exclude the presence or addition of one or more other features, integrals, steps, operations, elements, components and / or collections thereof.

[0020] It should also be understood that the term “and / or” as used in this specification and the appended claims refers to any combination of one or more of the associated listed items and all possible combinations, and includes such combinations.

[0021] As used in this specification and the appended claims, the term "if" may be interpreted, depending on the context, as "when," "once," or "in response to determination." Similarly, the phrase "if determined" or "if matched to [described condition or event]" may be interpreted, depending on the context, as meaning "once determined," "in response to determination," "once matched to [described condition or event]," or "in response to matched to [described condition or event]."

[0022] Furthermore, in the description of this invention and the appended claims, the terms "first," "second," "third," etc., are used only to distinguish descriptions and should not be construed as indicating or implying relative importance.

[0023] References to "one embodiment" or "some embodiments" as described in this specification mean that one or more embodiments of the invention include a specific feature, structure, or characteristic described in connection with that embodiment. Therefore, the phrases "in one embodiment," "in some embodiments," "in other embodiments," "in still other embodiments," etc., appearing in different parts of this specification do not necessarily refer to the same embodiment, but rather mean "one or more, but not all, embodiments," unless otherwise specifically emphasized. The terms "comprising," "including," "having," and variations thereof mean "including but not limited to," unless otherwise specifically emphasized.

[0024] It should be understood that the sequence number of each step in the following embodiments does not imply the order of execution. The execution order of each process should be determined by its function and internal logic, and should not constitute any limitation on the implementation process of the embodiments of the present invention.

[0025] The data processing methods, apparatus, devices, and media provided in the embodiments of this application will be described in detail below.

[0026] like Figure 1 The diagram shown is a flowchart of a set-top box control method provided in an embodiment of this application. Specifically, the set-top box control method may include the following steps: S101. Obtain the image data of the object being viewed.

[0027] A set-top box typically refers to a device that connects a television set to an external signal source. Its main function is to receive, decode, and process digital television signals or online streaming media content, and then display it on the television set. In this embodiment, the set-top box also has image acquisition and processing capabilities, and can adaptively adjust device parameters according to the user's status.

[0028] The audience refers to the individual or group that is using or watching the content displayed on the set-top box. This audience can be a single user or multiple users.

[0029] Image data refers to digital data containing visual information about the viewed object, acquired through image acquisition devices (such as cameras). This data typically exists in the form of consecutive frames and can be used for subsequent image processing and analysis.

[0030] There are several ways to acquire image data of the viewed object. For example, the set-top box can be configured with a built-in camera, which is set to periodically capture still images of the viewing area at a fixed frame rate (e.g., 10 frames per second) and transmit these image data to the processing module. Alternatively, the set-top box can connect to a separate image acquisition device via an external interface. This device, upon receiving a trigger command from the set-top box, acquires a single image and transmits it back. Furthermore, a miniature camera integrated into the set-top box's remote control can be used to acquire an image when the user presses a specific button and send it to the set-top box for processing.

[0031] S102. Determine the status information of the viewed object based on image data.

[0032] Status information refers to various types of information describing the current situation of the viewed object, such as the viewer's posture, distance from the screen, and the number of people present. This information forms the basis for the set-top box's intelligent decision-making.

[0033] Various image analysis techniques can be employed to determine the state information of the viewed object based on the image data. For example, brightness variation analysis of the image data can be used to determine whether a moving object exists within the viewing area, thus preliminarily determining whether the viewed object is present. Alternatively, edge detection processing can be performed on the image data to identify the presence of a preliminary human silhouette, thereby determining the presence of the viewed object. Furthermore, analyzing changes in the area of ​​specific color regions in the image can infer changes in the position or size of the viewed object within the image, thereby indirectly obtaining some state information.

[0034] S103. Adjust the device operating parameters of the set-top box based on the status information of the viewing object.

[0035] Device operating parameters refer to various configuration parameters that can be adjusted during the operation of the set-top box, such as the display ratio of the user interface, volume level, and interaction methods (such as enabling or disabling voice control). Adjusting these parameters aims to optimize the viewing experience.

[0036] Regarding adjusting the set-top box's operating parameters based on the viewing object's status information, corresponding operations can be performed according to the determined status information. For example, when the viewing object's presence is detected, the set-top box's volume can be adjusted to a preset default value. Alternatively, when the viewing object is detected leaving, the set-top box's screen brightness can be reduced or it can enter standby mode. Furthermore, the user interface's display mode can be directly switched based on the presence or absence of a viewing object; for example, a standard interface can be displayed when a viewing object is detected, and a saver interface can be displayed when no viewing object is detected.

[0037] The following example will provide a more detailed explanation of the above technical solution: User A is watching television at home, and the set-top box continuously acquires image data from the viewing area through its built-in image acquisition device. Initially, User A is sitting on the sofa, about two meters away from the television screen. At this time, the set-top box's image processing module receives continuous image data.

[0038] First, the set-top box performs the step of acquiring image data of the viewed object. The built-in image acquisition device continuously captures video images of the area where user A is located at a frequency of approximately 15 frames per second, and transmits these image data to the set-top box's image processing unit.

[0039] Next, the set-top box determines the status information of the viewed object based on the image data. The image processing unit analyzes the received image data. For example, by initially identifying the pixel distribution and brightness changes in the image, the system can determine that there is a preliminary human-shaped area in the image, thus confirming that the viewed object, user A, is present. Furthermore, by analyzing the position and size of this human-shaped area in the image, the system can preliminarily infer that user A is currently in a state of "present and relatively stationary".

[0040] Subsequently, the set-top box adjusts its operating parameters based on the viewing object's status information. According to the determined "present and relatively stationary" status information, the set-top box's control module queries preset simple rules. For example, if the rule is set to "when a viewing object is detected, set the volume to the default value and display the user interface in standard mode," the set-top box will perform the corresponding operation. Specifically, the set-top box's audio output module will be adjusted to a preset volume level, and its display rendering engine will ensure that the user interface is presented at a standard size and layout.

[0041] Through the above process, this set-top box control method can perceive the viewing object and adaptively adjust the device's operating parameters based on this perception. For example, when user A enters the viewing area from outside the room, the system can sense its presence and automatically turn on or adjust to a basic viewing mode without requiring user A to manually operate the remote control to start the set-top box or adjust the volume.

[0042] Based on the above examples, the set-top box control method provided in this embodiment demonstrates an improvement over the prior art. In the prior art, the interaction mode of a set-top box is usually static and fixed; regardless of the user's state, the device's operating parameters, such as the user interface size and volume output, remain unchanged. For example, when user A enters the viewing area, the set-top box in the prior art still requires the user to manually operate the remote control to turn on the device or adjust the volume to a suitable level.

[0043] This application provides a set-top box control method, comprising: acquiring image data of a viewed object; determining the state information of the viewed object based on the image data; and adjusting the device operating parameters of the set-top box based on the state information of the viewed object. This application acquires image data of the viewed object and determines its state information accordingly, thereby adaptively adjusting the device operating parameters of the set-top box. This enables the set-top box to achieve intelligent automatic control based on the user's actual viewing state, allowing parameter adjustment to be completed without the user manually operating the remote control, effectively improving the convenience and smoothness of human-computer interaction.

[0044] In some embodiments, this application proposes a set-top box control method. This method acquires image data of a viewing object, determines the viewing object's state information based on the image data, and then adjusts the set-top box's operating parameters based on the viewing object's state information. However, in practical applications, directly adjusting device parameters based on raw, discrete viewing object state information may lead to overly frequent adjustments, a lack of holistic consideration, or an inability to accurately capture the viewing object's true intentions and the surrounding environment, thereby affecting the consistency and comfort of the user experience.

[0045] In response, this application further proposes the above-mentioned method of adjusting the device operating parameters of the set-top box based on the state information of the viewing object, including: fusing the state information of the viewing object to obtain a scene feature vector; matching the corresponding scene strategy based on the scene feature vector using a preset scene strategy library; and adjusting the device operating parameters of the set-top box based on the scene strategy.

[0046] The fusion processing of the viewing object's state information integrates state information from different dimensions to form a more representative and generalized data representation. This approach avoids making decisions using raw, discrete state information directly, thereby improving the accuracy and robustness of the decisions and better reflecting the overall scene of the viewing object. For example, statistical methods such as weighted average and principal component analysis (PCA) can be used to reduce and integrate multidimensional state information, or neural networks (such as fully connected layers and recurrent neural networks, RNNs) can be used to learn and extract features from sequential or multimodal state information to generate a high-dimensional scene feature vector. The scene feature vector is the result of the fusion processing; it is a multidimensional numerical vector that abstractly represents the overall scene or state of the currently viewed object.

[0047] Based on scene feature vectors and utilizing a pre-defined scene policy library, corresponding scene policies are matched to obtain the desired outcome. The aim is to map abstract scene features to specific, executable device adjustment schemes, achieving intelligent and scenario-based device control. The scene policy library is a knowledge base storing various predefined scenes and their corresponding control policies. The matching process can be rule-based, where the scene policy library stores a series of rules; when a scene feature vector meets the conditions of a certain rule, the corresponding scene policy is matched. Alternatively, machine learning classification algorithms, such as Support Vector Machines (SVM), decision trees, or neural networks, can be used to pre-train a model that can classify scene feature vectors into predefined scene categories, with each category corresponding to a scene policy. A scene policy is a set of pre-defined device operating parameter adjustment schemes for a specific viewing scenario, such as adjusting volume, brightness, playback content, and pausing.

[0048] Adjusting the set-top box's operating parameters based on scene strategies refers to modifying various operating parameters of the set-top box according to specific instructions or parameter values ​​defined in the matched scene strategy. This allows abstract control strategies to be translated into actual device operations, achieving refined and automated control of the set-top box. The set-top box's control module can receive instructions from the scene strategy and directly modify parameters such as volume, brightness, playback mode, and content recommendation list through internal APIs or hardware interfaces. Scene strategies can contain a series of operation instructions, such as "pause playback," "switch to children's mode," "reduce volume to X," and "recommend category Y programs." The set-top box's control system parses these instructions and executes the corresponding operations.

[0049] This application's solution fuses state information such as the viewing object's posture and distance to generate a scene feature vector that comprehensively reflects the current viewing context. This scene feature vector is then used to match the most suitable scene strategy from a pre-set scene strategy library. Ultimately, the set-top box's operating parameters are adjusted according to the matched scene strategy. The set-top box control is no longer simply a passive response to single state information, but rather a comprehensive assessment of the viewing context, considering multiple factors and executing corresponding device adjustment schemes. For example, when the viewing object exhibits a focused, close-range viewing posture, the system may recognize this as an "immersive viewing" scene and adjust screen brightness, volume, and content recommendations accordingly to enhance the viewing experience. This context-based intelligent adjustment significantly improves the intelligence level of set-top box control and the comfort of the user experience.

[0050] The following is a concrete example to illustrate this. Assume the set-top box system acquires the viewer's posture information (e.g., leaning forward, leaning back, sitting upright) and distance information (e.g., near, medium, far from the screen) and the number of viewers (e.g., single, two, multiple viewers) from image data. The system first fuses this state information. For example, if it detects that the viewer is leaning forward, close to the screen, and is watching alone, the fusion processing module might generate a scene feature vector representing "focused learning / working." This scene feature vector is then input into a scene policy library for matching. The scene policy library pre-sets various scenes and their corresponding policies. For example, the policy corresponding to the "focused learning / working" scene might include: adjusting the screen brightness to medium-high, adjusting the volume to a low level, and recommending educational or documentary content. Once the policy is matched, the set-top box's control module will execute the corresponding parameter adjustments, thereby providing the viewer with a viewing environment consistent with their current situation.

[0051] Through the above technical solution, this application achieves comprehensive utilization of the viewing object's state information, avoiding frequent and unstable operations that may result from direct adjustments based on single or discrete state information. By matching the fused scene feature vector with a preset scene strategy library, the set-top box can achieve more intelligent and contextualized adjustments to its operating parameters, making the set-top box control more precise, stable, and in line with the user's actual needs. This significantly improves the consistency and comfort of the user experience, reduces the frequency of manual adjustments by the user, and allows the set-top box to better adapt to different viewing scenarios.

[0052] In some embodiments, this application proposes a set-top box control method, which includes acquiring image data of a viewing object, determining the viewing object's state information based on the image data, and adjusting the set-top box's device operating parameters based on the viewing object's state information. However, in practical applications, if the state information is not specific enough or cannot accurately reflect the viewing object's actual physiological state, such as its posture, the set-top box's device operating parameters may not be adjusted precisely enough, failing to effectively improve the viewing experience or meet the viewing object's comfort needs.

[0053] In response, this application further proposes that the aforementioned state information includes posture information, and that the state information of the viewing object is determined based on image data, including obtaining the key human body coordinates of the viewing object based on image data; calculating the torso tilt angle, body compression ratio, and hip support state based on the key human body coordinates; and determining the posture information of the viewing object based on the torso tilt angle, body compression ratio, and hip support state.

[0054] Posture information refers to data related to the viewer's body posture and physique, reflecting their comfort level, fatigue level, or concentration. Posture information can be obtained in various ways, such as by analyzing the relative positions, angles, or deformations of different body parts. Obtaining the viewer's key coordinates involves identifying and locating the two-dimensional or three-dimensional position information of the human skeleton or joints in the image. This step is fundamental to posture analysis, providing precise positions of various body parts for subsequent posture calculations. Implementation methods can include using deep learning models, such as OpenPose and AlphaPose, to estimate human posture from image data and directly output the coordinates of key points; or using traditional image processing techniques, such as edge detection and feature point matching, combined with a human skeleton model for key point localization. Calculating the trunk tilt angle, body compression ratio, and hip support state are specific indicators used to identify the viewer's posture. The trunk tilt angle reflects the degree of tilt of the viewer's upper body relative to the vertical direction, indicating whether they are leaning forward, backward, or sideways. Body compression ratio reflects the degree of extension or contraction of the viewer's body, for example, by measuring the ratio of the distance from the shoulder to the hip to a standard height. Hip support status reflects whether the viewer is sitting stably, in a semi-reclined or standing position, and can be determined by the relative position and contact area between the hip key points and the seat surface or ground. Determining the viewer's posture information involves combining the above quantitative indicators to form a final judgment of the viewer's posture. This determination process can employ methods such as rule engines, machine learning classifiers, or fuzzy logic.

[0055] This application's solution first acquires image data of the viewer, and then, going beyond generalized state information, delves deeper to extract posture information reflecting the viewer's physiological state. Specifically, the system accurately identifies and obtains key human coordinates from the image data; these coordinates form the basis for human posture analysis. Subsequently, based on these key coordinates, the system can calculate a series of specific posture indicators, such as trunk tilt angle, body compression ratio, and hip support status. These indicators characterize the viewer's body posture from different dimensions; for example, the trunk tilt angle reflects the degree of upper body tilt, the body compression ratio reveals the body's extended or hunched state, and the hip support status indicates the stability of the sitting posture. Finally, by comprehensively analyzing these quantitative indicators, the system can accurately determine the viewer's posture information. This detailed posture information, as more specific state information, provides a more accurate basis for adjusting the subsequent operating parameters of the set-top box device, effectively solving the problem of inaccurate adjustments caused by insufficient specific state information and significantly improving the viewing experience.

[0056] In one specific implementation, a camera mounted above a television or set-top box can be used to acquire a real-time video stream of the viewed object as image data. This video stream is then fed into a built-in image processing module. This module can integrate a deep learning-based pose estimation algorithm, such as the OpenPose or AlphaPose model, which can identify and output the key coordinates of the viewed object's body in real time from video frames, such as the two-dimensional or three-dimensional coordinates of 17 or more key points including the head, shoulders, elbows, wrists, hips, knees, and ankles. After acquiring these key coordinates, the system can further calculate pose indices. For example, the torso tilt angle can be obtained by calculating the angle between the vector formed by connecting the midpoints of the left and right shoulders and the midpoints of the left and right hips and the vertical axis of the image. The body compression ratio can be obtained by comparing the vertical distance between the shoulder key points and the hip key points with a preset average sitting height. The hip support state can be determined by analyzing the vertical position of the hip key points in the image and their relative positional relationship with the knee key points to determine whether the viewed object is fully seated, semi-reclined, or standing. Finally, the calculated data, such as trunk tilt angle, body compression ratio, and hip support status, can be input into a preset posture judgment model. This model can be a rule-based expert system or a trained classifier, thereby outputting specific posture information of the viewed object, such as "comfortable sitting posture," "forward-leaning focused posture," or "fatigued semi-reclining posture."

[0057] Through the aforementioned technical solution, this application can extract accurate posture information from the image data of the viewing object. This quantitative analysis of key posture indicators such as torso tilt angle, body compression ratio, and hip support status enables the system to more accurately determine the physiological state and comfort of the viewing object. Therefore, adjusting the set-top box's operating parameters based on this precise posture information allows for more personalized and intelligent control. For example, it can automatically adjust screen brightness, volume, or playback content when the viewing object's posture is unfavorable, thereby significantly improving the comfort and health of the viewing experience and effectively solving the problem of inaccurate device parameter adjustments due to insufficient specificity of status information.

[0058] In some other embodiments, this application proposes a set-top box control method. This method first acquires image data of the viewed object, determines the state information of the viewed object based on the image data, and then adjusts the device operating parameters of the set-top box based on the state information of the viewed object. However, in practical applications, how to efficiently and accurately obtain the distance information between the viewed object and the set-top box from two-dimensional image data, and further guide the adjustment of device operating parameters, is a technical problem that needs to be solved.

[0059] In response, this application further proposes that the status information in the aforementioned set-top box control method includes distance information. Determining the status information of the viewed object based on image data includes: identifying facial regions in the image data; calculating the proportion of the facial regions in the image data; and mapping the proportion to distance information based on a preset mapping rule.

[0060] Distance information refers to the spatial distance between the viewer and the set-top box. This information is crucial for assessing the viewer's viewing habits and health (e.g., whether they spend extended periods watching at close range), as well as optimizing the viewing experience (e.g., automatically adjusting screen brightness, volume, or display mode based on distance). Besides image saturation mapping, distance information can be obtained in other ways, such as using a binocular vision system to calculate depth based on parallax, or employing a time-of-flight (ToF) sensor to directly measure the round-trip time of light to acquire distance data.

[0061] Face region recognition in image data refers to locating and selecting the specific position of the viewer's face within a captured image frame using image processing and pattern recognition techniques. This step is crucial for obtaining distance information. Various techniques can be employed to achieve face region recognition. For example, deep learning-based convolutional neural network (CNN) models, such as YOLO (YouOnly Only Look Once) or SSD (Single Shot MultiBox Detector), can detect faces in images in real time and accurately. Alternatively, traditional machine learning methods, such as cascaded classifiers based on Haar features, can be used to scan and match predefined facial features in images using a trained model.

[0062] Calculating the proportion of a face region in image data refers to quantifying the percentage of the face region within the entire image frame after it has been identified. This proportion is inversely proportional to the distance between the viewer and the image acquisition device: the closer the distance, the larger the face appears in the image, and the higher its proportion; conversely, the farther the distance, the lower the proportion. By calculating this proportion, a quantitative indicator directly related to distance changes can be obtained. Specifically, after obtaining the pixel coordinates of the face region (e.g., the coordinates of the top-left corner of the bounding box and its width and height), the pixel area of ​​the face region can be calculated and divided by the total pixel area of ​​the entire image frame to obtain a dimensionless proportion. Another approach is to calculate the ratio of the width or height of the face region to the total width or height of the image, as an alternative indicator of proportion.

[0063] Mapping facial proportions to distance information based on pre-established mapping rules refers to converting calculated facial region proportions into actual physical distance values ​​using a pre-established correspondence. This pre-established mapping rule can be established in various ways. For example, images can be acquired at different known distances, and the corresponding facial proportions recorded. A lookup table can then be constructed, and during actual operation, the calculated proportion value is matched against the lookup table to obtain the distance. Alternatively, a mathematical model can be established using regression analysis, such as a linear or non-linear function, that describes the quantitative relationship between facial proportions and actual distances. Distance information can be calculated by substituting the proportion value into this function.

[0064] This application's solution acquires image data of the viewed object and, based on this, determines the object's distance information by first identifying the facial region within the image data. This allows the system to focus on the portion of the image most relevant to the distance to the viewed object. Subsequently, the proportion of the facial region in the image data is calculated. This proportion directly reflects the relative size of the viewed object's face in the image, and this relative size is intrinsically correlated with the distance between the viewed object's face and the set-top box. Finally, based on a preset mapping rule, the proportion is mapped to distance information, thus transforming the two-dimensional information in the image into distance information with actual physical meaning. This distance information can be used by the set-top box for subsequent adjustments to device operating parameters. For example, when the distance is too close, the screen brightness can be reduced or the user can be prompted to maintain an appropriate distance to protect their eyesight. This method of indirectly obtaining distance through image proportion avoids reliance on complex depth sensors, reduces hardware costs and computational complexity, and provides sufficiently accurate distance judgment criteria, enabling the set-top box to respond more intelligently to the actual viewing state of the viewed object.

[0065] The following example illustrates this. Assume the set-top box has a built-in ordinary RGB camera to capture image data of the viewed object. After the camera captures the image data, the set-top box's processor can run a pre-trained deep learning-based face detection model, such as a lightweight MobileNet-SSD model. This model can detect and locate face regions in the image in real time and return a bounding box for the face region. After identifying the bounding box of the face region, the system can calculate the pixel area of ​​the bounding box and divide it by the total pixel area of ​​the entire image frame to obtain the proportion of the face region in the image. For example, if the image resolution is 1920x1080 pixels and the detected face region is 300x400 pixels, then the proportion is (300*400) / (1920*1080)≈0.057. Subsequently, the set-top box can use a preset mapping rule to convert this proportion value into distance information. The mapping rule can be a lookup table established through calibration before the system leaves the factory. For example: when the percentage is greater than 0.1, the distance is less than 0.8 meters; when the percentage is between 0.05 and 0.1, the distance is between 0.8 meters and 1.5 meters; when the percentage is less than 0.05, the distance is greater than 1.5 meters. By querying this lookup table, the system can quickly obtain the distance information between the viewing object and the set-top box.

[0066] Through the aforementioned technical solution, the set-top box can accurately obtain distance information between the viewer and the set-top box using the acquired image data. By identifying facial regions in the image and calculating their proportions, combined with preset mapping rules, it can effectively convert two-dimensional image information into meaningful distance data. This allows the set-top box to more precisely perceive the viewer's viewing status, such as determining whether the viewing distance is too close or too far, thereby adjusting device operating parameters such as screen brightness, volume, and display mode, thus improving the user's viewing experience and helping to protect the user's eye health.

[0067] In other embodiments, this application proposes a set-top box control method. This method acquires image data of a viewed object, determines the object's state information based on the image data, and then adjusts the set-top box's operating parameters based on the object's state information. However, in practical applications, accurately determining certain state information, especially distance information, directly from image data may face challenges. For example, changes in ambient light, occlusion, or insufficient image resolution can all affect the accuracy of distance estimation, resulting in inaccurate adjustments to the set-top box's operating parameters.

[0068] In response, this application further proposes that the status information includes distance information, and that the status information of the viewing object is determined based on image data, including obtaining the key human coordinates of the viewing object based on image data; and calculating the distance information of the viewing object based on the key human coordinates.

[0069] Distance information refers to the spatial distance between the viewer and the set-top box. This information is crucial for assessing the viewer's viewing habits, comfort, and potential health risks (such as prolonged close-range viewing). Distance information can be expressed as absolute distance values, such as in centimeters or meters, or as relative distance levels, such as "near," "medium," and "far." Human body key coordinates refer to the positional information of specific points on the viewer's body in the image, such as joints or feature points like the head, shoulders, elbows, wrists, hips, knees, and ankles. These coordinate points are the foundational data for human posture analysis, size estimation, and spatial positioning. Human body key coordinates can be obtained by analyzing and identifying image data using deep learning models, such as OpenPose or AlphaPose, or by combining depth maps obtained from depth sensors. Calculating the distance information of the viewer based on human body key coordinates utilizes the geometric characteristics of these coordinates and image processing techniques to indirectly and relatively accurately deduce the viewer's distance. For example, known average human body dimensions or user-specific dimensions obtained through calibration can be used, combined with the pixel distances between key coordinate points in the image, to deduce the actual distance using perspective projection principles. In addition, binocular or multi-view vision systems can be used to calculate distances by using the parallax of key coordinate points from different perspectives.

[0070] The proposed solution uses the image data of the viewed object as input to first accurately identify and extract the key coordinates of the object's human body. These key coordinates provide information about the object's spatial position and relative size within the image. Subsequently, the system uses these extracted key coordinates, combined with a pre-defined geometric model or depth estimation algorithm, to calculate the distance between the viewed object and the image acquisition device. For example, the actual distance can be deduced using perspective principles based on the relative positions of the key coordinates in the image and known human proportions. In this way, abstract image data is transformed into concrete distance information, providing a more accurate and reliable basis for adjusting the set-top box's operating parameters based on the viewed object's status information. This method avoids complex and environmentally sensitive distance estimations directly from the image, instead utilizing the stability of human structural features, making distance information acquisition more robust and accurate, thereby improving the intelligence of set-top box operating parameter adjustments and enhancing the user experience.

[0071] The following example illustrates this: After acquiring image data of the viewed object, the system can utilize a pre-trained convolutional neural network model. This model, trained on a large dataset of human poses, can accurately detect and output the key coordinates of the viewed object's body from the input image. These coordinates include the pixel coordinates of 17 or more key points such as the head center, left and right shoulders, left and right elbows, left and right wrists, left and right hips, left and right knees, and left and right ankles. Once these key coordinates are obtained, the system can further use them to calculate distance information. For example, the pixel distance between key points on the viewed object's head and hips can be selected and combined with pre-calibrated camera parameters (such as focal length and sensor size) and known average human height proportions. Using simple trigonometric relationships or perspective projection formulas, the pixel distance can be converted into actual physical distance. Furthermore, the distance information of the viewed object can be estimated by calculating the number of pixels occupied by the width or height of the face region in the image and comparing it with the preset actual width or height of the face.

[0072] Through the above technical solution, this application can more accurately calculate the distance information of the viewed object from image data by utilizing the geometric characteristics of key human body coordinates. This method avoids the complex and environmentally susceptible distance estimation directly from the image, instead utilizing the stability of human body structural features to make the acquisition of distance information more accurate. Therefore, when adjusting the set-top box's operating parameters based on the distance information of the viewed object, it can provide a more accurate basis, thereby achieving more intelligent and user-friendly adjustments to the device's operating parameters, such as automatically adjusting screen brightness, volume, or recommended content based on the viewing distance, effectively improving the comfort and health of the user's viewing experience.

[0073] In other embodiments, this application proposes a set-top box control method. This method acquires image data of a viewing object, determines the viewing object's state information based on the image data, and then adjusts the set-top box's operating parameters based on the viewing object's state information. Furthermore, the method fuses the viewing object's state information to obtain a scene feature vector; based on the scene feature vector, it matches a corresponding scene strategy using a pre-defined scene strategy library; and adjusts the set-top box's operating parameters based on the scene strategy. However, relying solely on the viewing object's basic state information (such as posture and distance) may not fully capture the complexity of the viewing scene, especially in scenarios involving multiple viewers or requiring personalized services. This could lead to inaccurate set-top box parameter adjustments or failure to meet the needs of specific users.

[0074] In this regard, this application further proposes that the status information also includes information on the number of viewers, their age group, or their identity.

[0075] The quantity information refers to the number of viewing objects (i.e., people) within the viewing area. This information enables the system to identify whether the current scene is a single person or multiple people watching, thus providing a more precise basis for adjusting the set-top box's operating parameters. For example, the number of viewing objects can be obtained by performing object detection on the acquired image data, identifying and counting faces or human body contours in the image. Another approach is to use a deep learning model to perform semantic segmentation on the image, distinguishing the areas containing people, and then calculating the number of people using methods such as connected component analysis.

[0076] Age group information refers to information used to characterize the age group category to which the viewer belongs. This age group information generally includes at least one of children, young adults, and the elderly. It is not a specific age, but rather a group label used to distinguish whether someone is a child, young adult, or elderly person, so that the data can be adjusted accordingly. In one implementation, the face region of the viewer can be located and extracted from image data. Facial features in the face region are extracted, including but not limited to facial contours, skin texture, facial proportions, and the distribution of facial key points. The extracted facial features are then input into a pre-trained age group classification model. The model outputs a classification result for the corresponding age group. Based on the classification result, the viewer is classified as a child, young adult, or elderly person.

[0077] Identity information refers to the specific identification of the viewer. This information enables the system to recognize a specific viewer, thereby providing more personalized services. For example, facial recognition technology can be used to compare faces detected in image data with a pre-stored user face database to identify the viewer. Furthermore, biometric identification technologies such as gait recognition and body posture analysis can be used to extract unique features of the viewer from image data and match them with known identity information to determine the viewer's identity.

[0078] This application's solution acquires image data of the viewing object, determines the viewing object's state information based on this image data, and then performs fusion processing on the state information to obtain a scene feature vector. Based on this scene feature vector, a scene strategy is matched to adjust the set-top box's operating parameters. Furthermore, this application incorporates the number, age group, or identity information of the viewing objects into the state information. This means that when determining the viewing object's state information, in addition to traditional information such as posture and distance, the system also identifies how many viewing objects are currently present, or who the specific viewing object is. This richer and more specific contextual information is integrated into the viewing object's state information, enabling subsequent fusion processing to generate a more comprehensive and refined scene feature vector. This scene feature vector not only reflects the viewing object's behavior and location but also includes the number or identity attributes of the viewing objects. Therefore, when using a preset scene strategy library for matching, the system can match more accurate and personalized scene strategies. For example, when a specific user or multiple users are identified, a strategy specifically designed for that user or multi-user scenario can be triggered, thereby achieving a more intelligent and practical adjustment of the set-top box's operating parameters.

[0079] The following example illustrates this: The set-top box system continuously acquires image data of the viewing area through its built-in camera. The system first processes the image data, identifying key human coordinates of the viewed object and calculating its torso tilt angle, body compression ratio, and hip support state to determine its posture information (e.g., identified as a "semi-reclining posture"). Simultaneously, the system identifies facial regions in the image data, calculates their proportion within the image, and maps them to distance information using a pre-defined mapping rule (e.g., identified as "far from the screen"). Building on this, the system further analyzes the image data, using a face detection algorithm to identify two faces in the image, thus determining the number of viewed objects to be "2". Simultaneously, a face recognition algorithm identifies the viewed objects' identities from pre-stored face data. The system identifies one face as "Mr. Zhang" and the other as "Ms. Zhang," thus confirming the viewed objects' identities. This posture information, distance information, number information, and identity information together constitute the complete state information of the viewed objects. The system fuses this state information to generate a scene feature vector containing characteristics such as "semi-reclining posture," "relatively far distance," "two people watching," and "Mr. Zhang and Ms. Zhang." Subsequently, the system uses a pre-set scene strategy library to match the corresponding scene strategy, such as "Couple's Cinema Mode." Based on this "Couple's Cinema Mode" strategy, the set-top box's operating parameters are adjusted; for example, the volume is moderately increased, the screen brightness is adjusted to a soft mode, and movie content suitable for couples to watch together is recommended.

[0080] The following example illustrates this further. The set-top box system continuously acquires image data of the viewing area through its built-in camera. The system first processes the image data to identify the age group of the viewer. If the viewer is identified as a child, a scene strategy tailored to the child's vision is applied, adjusting the corresponding device operating parameters, such as enlarging the UI, based on the principle of improving the elderly's audiovisual experience.

[0081] By incorporating information such as the number of viewers, their age group, or their identity into the status information, the set-top box can gain a more comprehensive and in-depth understanding of the viewing scenario. This solves the problem of insufficient precision in parameter adjustments or inability to meet personalized needs when relying solely on basic status information. Specifically, the system can distinguish between single-person and multi-person viewing scenarios and identify specific viewers, thus incorporating richer information when generating scene feature vectors. The matched scene strategies can more accurately reflect actual viewing needs, thereby enabling personalized and intelligent adjustments to the set-top box's operating parameters. For example, it can automatically adjust volume and screen ratio based on the number of viewers, or load personalized viewing preference settings and content recommendations based on the identified user's identity, greatly enhancing the personalization and comfort of the user experience.

[0082] In other embodiments, this application proposes a set-top box control method. This method first acquires image data of a viewed object and determines the object's state information based on the image data. Then, it adjusts the set-top box's operating parameters based on the object's state information. Specifically, the adjustment process includes fusing the object's state information to obtain a scene feature vector; matching the scene feature vector with a pre-defined scene strategy library to obtain a corresponding scene strategy; and adjusting the set-top box's operating parameters based on the scene strategy.

[0083] In some embodiments of this application, the set-top box's operating parameters are adjusted based on the matched scene strategy. However, in practical applications, due to the potential for momentary fluctuations in the viewing object's state information or interference from factors such as ambient light, the system may frequently match different scene strategies within a short period. This could lead to frequent adjustments to the set-top box's operating parameters, affecting not only the continuity of the viewing experience but also increasing the device's operational burden and instability. To address this, this application further proposes adjusting the set-top box's operating parameters based on scene strategies, including: if the number of consecutive matches of the same scene strategy within a preset time range reaches a preset threshold, then adjusting the set-top box's operating parameters based on that scene strategy.

[0084] The preset duration range refers to a pre-defined time interval used to limit the statistical period for the number of consecutive matches of a scene strategy. This duration range can be a fixed time value, such as 5 seconds, 10 seconds, or 30 seconds, to ensure that the stability of the scene strategy is observed within a sufficient time window. Furthermore, this duration range can be dynamically adjusted based on the viewing object's behavior pattern, the rate of environmental change, or the set-top box's response characteristics. For example, when the system detects that the viewing object is in a relatively stable state, the duration range can be appropriately extended to reduce unnecessary strategy switching; when the viewing object's behavior is active, the duration range can be shortened to improve the system's response sensitivity. The number of consecutive matches of the same scene strategy refers to the number of times the system identifies and confirms the consecutive occurrence of the same scene strategy within the preset duration range. This can be achieved by maintaining a counter, which increments whenever the system consecutively matches the same scene strategy; once a different scene strategy is matched, the counter is reset or cleared. Another implementation method is to use a sliding time window mechanism, which counts the frequency or duration of consecutive matches of the same scene strategy within each time step to determine its continuity. A preset threshold is a pre-defined value used to determine whether the conditions for triggering adjustments to the set-top box's operating parameters are met. This threshold can be an integer, such as 2, 3, or 5 times, indicating that within a preset time range, the same scenario strategy needs to match consecutively a certain number of times before adjustments are triggered. This threshold can be configured according to the needs of the actual application scenario. For example, a higher threshold can be set for scenarios requiring high stability, while a lower threshold can be set for scenarios requiring fast response.

[0085] The solution proposed in this application effectively solves the problem of frequent adjustments to set-top box operating parameters due to instantaneous fluctuations in scene strategies. Specifically, after the system determines the state information of the viewed object based on its image data, further fuses and processes it to obtain a scene feature vector, and then uses a preset scene strategy library to match the corresponding scene strategy, it does not immediately adjust the set-top box's operating parameters. Instead, it continuously monitors whether the same scene strategy is continuously matched within a preset time range. Only when the number of such continuous matches reaches a preset threshold will the system adjust the set-top box's operating parameters based on that scene strategy. This ensures that parameter adjustments are only triggered after a certain scene strategy has been stably identified and persisted for a period of time, thereby avoiding frequent switching of device parameters due to short-lived, non-continuous scene changes and improving the consistency of the user experience.

[0086] The following is a specific example to illustrate this. As a concrete implementation, assume the preset duration is set to 5 seconds and the preset threshold is set to 3 times. During the operation of the set-top box control method, the system continuously acquires image data of the viewed object and determines its status information, thereby matching a scene strategy. If, within a 5-second time window, the system matches the scene strategy of "viewed object viewing at close range" 3 times (or more) consecutively—for example, matching this strategy in the 1st, 2nd, and 3rd seconds—then the system will trigger adjustments to the set-top box's operating parameters, such as lowering the screen brightness or volume. However, if within that 5-second window, the system first matches "viewed object viewing at close range," then momentarily matches "viewed object leaving," and then matches "viewed object viewing at close range" again, since the consecutive matching count of "viewed object viewing at close range" does not reach 3 times, the system will not trigger parameter adjustments, thus avoiding invalid adjustments due to brief scene changes.

[0087] Through the above technical solution, the operating parameters of the set-top box can be adjusted based on continuous and stable judgment of the scene strategy, which significantly enhances the stability of the set-top box control system and effectively avoids frequent switching of equipment parameters caused by short-term fluctuations in the state of the viewing object or environmental interference. This provides the viewing object with a more stable and consistent viewing experience, reduces the interference caused to users by unnecessary parameter adjustments, and reduces the operating burden of the equipment caused by frequent adjustments.

[0088] In one embodiment, this application also provides a set-top box control device. For example... Figure 2 As shown, it includes an acquisition module 21, a determination module 22, and an adjustment module 23. Detailed descriptions of each functional module are as follows: Acquisition module 21 is used to acquire image data of the viewed object; The determination module 22 is used to determine the status information of the viewed object based on the image data; Adjustment module 23 is used to adjust the device operating parameters of the set-top box based on the status information of the viewing object.

[0089] This invention also provides an electronic device, including a memory, a processor, and a program stored in the memory and executable on the processor. When the processor executes the program, it implements the aforementioned set-top box control method; to avoid repetition, this will not be described again here. Alternatively, the electronic device can implement the functions of each module in this embodiment of the set-top box control device; this will also not be described again here.

[0090] This invention also provides a readable storage medium storing a program. When the program is executed by a processor, it implements the aforementioned set-top box control method. To avoid repetition, this will not be described again here. Alternatively, when the program is executed by a processor, it implements the functions of each module in this embodiment of the set-top box control device, which will also not be described again here.

[0091] The above embodiments are only used to illustrate the technical solutions of the present invention, and are not intended to limit it. Although the present invention has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some of the technical features. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the spirit and scope of the technical solutions of the embodiments of the present invention, and should all be included within the protection scope of the present invention.

[0092] Those skilled in the art will understand that all or part of the processes in the methods of the above embodiments can be implemented by a computer program instructing related hardware. The computer program can be stored in a non-volatile computer-readable storage medium. When executed, the computer program can include the processes of the embodiments of the above methods. Any references to memory, storage, databases, or other media used in the embodiments provided in this application can include non-volatile and / or volatile memory. Non-volatile memory can include read-only memory (ROM), programmable ROM (PROM), electrically programmable ROM (EPROM), electrically erasable programmable ROM (EEPROM), or flash memory. Volatile memory can include random access memory (RAM) or external cache memory. By way of illustration and not limitation, RAM is available in various forms, such as static RAM (SRAM), dynamic RAM (DRAM), synchronous DRAM (SDRAM), dual data rate SDRAM (DDRSDRAM), enhanced SDRAM (ESDRAM), synchronous link DRAM (SLDRAM), Rambus direct RAM (RDRAM), direct memory bus dynamic RAM (DRDRAM), and memory bus dynamic RAM (RDRAM), etc.

[0093] Those skilled in the art will clearly understand that, for the sake of convenience and brevity, the above-described division of functional units and modules is used as an example. In practical applications, the above functions can be assigned to different functional units and modules as needed, that is, the internal structure of the device can be divided into different functional units or modules to complete all or part of the functions described above.

Claims

1. A set-top box control method, characterized in that, include: Acquire image data of the object being viewed; The status information of the viewed object is determined based on the image data; Adjust the set-top box's operating parameters based on the status information of the viewed object.

2. The set-top box control method according to claim 1, characterized in that, Adjusting the set-top box's operating parameters based on the viewing object's status information includes: The state information of the viewed object is fused to obtain a scene feature vector; Based on the scene feature vector, the corresponding scene strategy is obtained by matching using a preset scene strategy library; Adjust the set-top box's operating parameters based on the scenario strategy.

3. The set-top box control method according to claim 1, characterized in that, The status information includes attitude information; Determining the state information of the viewing object based on the image data includes: Based on the image data, obtain the key coordinates of the human body of the object being viewed; Based on the aforementioned key human body coordinates, calculate the torso tilt angle, body compression ratio, and hip support status; The posture information of the viewing object is determined based on the torso tilt angle, body compression ratio, and hip support state.

4. The set-top box control method according to claim 1, characterized in that, The status information includes distance information; Determining the state information of the viewing object based on the image data includes: Identify the facial regions in the image data; Calculate the proportion of the face region in the image data; Based on preset mapping rules, the percentage is mapped to the distance information.

5. The set-top box control method according to claim 1, characterized in that, The status information includes distance information; Determining the state information of the viewing object based on the image data includes: Based on the image data, obtain the key coordinates of the human body of the object being viewed; The distance information of the viewed object is calculated based on the key coordinates of the human body.

6. The set-top box control method according to claim 1, characterized in that, The status information also includes information on the number of viewers, their age group, or their identity.

7. The set-top box control method according to claim 2, characterized in that, The adjustment of the set-top box's operating parameters based on the scenario strategy includes: If the number of times the same scenario strategy is matched consecutively reaches a preset threshold within a preset time range, the device operating parameters of the set-top box will be adjusted based on the scenario strategy.

8. A set-top box control device, characterized in that, include: The acquisition module is used to acquire image data of the object being viewed; The determination module is used to determine the status information of the viewed object based on the image data; The adjustment module is used to adjust the device operating parameters of the set-top box based on the status information of the viewed object.

9. A computer device comprising a memory, a processor, and a computer program stored in the memory and executable on the processor, characterized in that, When the processor executes the computer program, it implements the steps of the set-top box control method as described in any one of claims 1 to 7.

10. A computer-readable storage medium storing a computer program, characterized in that, When the computer program is executed by the processor, it implements the steps of the set-top box control method as described in any one of claims 1 to 7.