Display device and control method thereof
By using detection models with varying computational loads to recognize gestures and body information in display devices, and combining these gestures and body information to determine control commands, the problem of single control commands in existing technologies is solved, thereby improving the level of intelligence and user experience.
Patent Information
- Application Number
- CN202111302345.9
- Authority / Receiving Office
- CN · China
- Patent Type
- Patents(China)
- Current Assignee / Owner
- Filing Date
- 2021-11-04
- Publication Date
- 2026-02-03
- Estimated Expiration
- 2041-11-04
AI Technical Summary
Existing display devices rely on simple control commands determined by gesture information, resulting in low intelligence and a poor user experience.
The image to be detected is extracted at a preset time interval. The first detection model with less computational load is used to determine whether it includes gesture information. If gesture information is present, the second detection model with greater computational load is used to identify gesture and limb information. The control command is determined by combining gesture and limb information.
It increases the number of control commands that users can use through gestures and body language, improves the intelligence of the display device and the user experience, and reduces the computational load and power consumption caused by invalid recognition.
Smart Images

Figure CN116069280B_ABST
Abstract
Description
Technical Field
[0001] This application relates to the field of electronic technology, and in particular to a display device and its control method. Background Technology
[0002] With the continuous development of electronic technology, display devices such as televisions are able to perform more and more functions. For example, the display device can capture images of the user through its video capture device, and the processor can recognize the user's gesture information in the image and then execute the command corresponding to the gesture information.
[0003] However, with existing technology, the control commands determined by gesture information in display devices are relatively simple, resulting in low intelligence and poor user experience. Summary of the Invention
[0004] This application provides a display device and its control method to solve the problems of low intelligence level and poor user experience in display devices.
[0005] This application provides a display device, comprising: a display screen configured to display images; a video acquisition device configured to acquire video data; and a controller configured to extract a frame of image to be detected from a series of consecutive frames of video data acquired by the video acquisition device at preset time intervals; use a first detection model to determine whether the image to be detected includes human gesture information; if so, continue to extract a preset number of images to be detected from the video data at the preset time interval and a preset number, and use a second detection model to identify human gesture information and limb information in the preset number of images to be detected respectively; wherein the amount of data used for calculation by the first detection model is less than the amount of data used for calculation by the second detection model; and execute control commands corresponding to the gesture information and limb information in the preset number of images to be detected.
[0006] A second aspect of this application provides a control method for a display device, comprising: extracting a frame of an image to be detected from a series of consecutive frames of video data acquired by a video acquisition device of the display device at a preset time interval; determining, using a first detection model, whether the image to be detected includes human gesture information; if so, continuing to extract a preset number of images to be detected from the video data at the preset time interval and a preset number, and using a second detection model to identify human gesture information and limb information in the preset number of images to be detected respectively; wherein the amount of data used for calculation by the first detection model is less than the amount of data used for calculation by the second detection model; and executing control commands corresponding to the gesture information and limb information in the preset number of images to be detected.
[0007] In summary, the display device and control method provided in this application enable the display device to determine different control commands based on gesture and limb information in the image to be detected. This expands the number of control commands that users can issue to the display device using this interactive method, thereby improving the intelligence level of the display device and the user experience. Furthermore, a first detection model with lower computational requirements is used to identify whether the image to be detected contains gesture information. Only after the first detection model determines that gesture information is included is a second detection model with higher computational requirements used to identify the gesture and limb information. This reduces the computational load and power consumption caused by invalid recognition and improves the processor's computational efficiency. Attached Figure Description
[0008] To more clearly illustrate the technical solutions in the embodiments of this application or the prior art, the drawings used in the description of the embodiments or the prior art will be briefly introduced below. Obviously, the drawings described below are only some embodiments of this application. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.
[0009] Figure 1 This is a schematic diagram illustrating the operating scenario of the display device used in this application;
[0010] Figure 2 This is a schematic diagram of the hardware structure of a hardware system in a display device.
[0011] Figure 3 This is a schematic diagram of an embodiment of a control method for a display device;
[0012] Figure 4 This is a schematic diagram of another embodiment of a control method for a display device;
[0013] Figure 5 A schematic diagram showing the coordinates of key hand points provided in this application;
[0014] Figure 6 A schematic diagram illustrating different extension and contraction states of key hand points provided in this application;
[0015] Figure 7 A schematic diagram illustrating an application scenario of the control method for the display device provided in this application;
[0016] Figure 8 A schematic diagram illustrating the use of gesture information and body information to jointly determine control commands, as provided in this application;
[0017] Figure 9 A schematic flowchart of an embodiment of the control method for the display device provided in this application;
[0018] Figure 10A schematic diagram of an embodiment of the mapping relationship provided in this application;
[0019] Figure 11 A schematic diagram of another embodiment of the mapping relationship provided in this application;
[0020] Figure 12 A schematic diagram of gesture and limb information in an image provided in this application;
[0021] Figure 13 A schematic diagram of an embodiment of the movement position of the target control provided in this application;
[0022] Figure 14 A schematic diagram of another embodiment of the movement position of the target control provided in this application;
[0023] Figure 15 A schematic flowchart of an embodiment of the control method for the display device provided in this application;
[0024] Figure 16 A schematic flowchart of an embodiment of the control method for the display device provided in this application;
[0025] Figure 17 A schematic diagram of the virtual frame provided in this application;
[0026] Figure 18 A schematic diagram illustrating the correspondence between the virtual frame and the display screen provided in this application;
[0027] Figure 19 A schematic diagram illustrating the movement of the target control provided in this application;
[0028] Figure 20 A schematic diagram of the area of the virtual frame provided in this application;
[0029] Figure 21 A schematic diagram of the edge region provided in this application;
[0030] Figure 22 A schematic diagram illustrating the state of the gesture information provided in this application;
[0031] Figure 23 A schematic diagram of an embodiment of the re-established virtual frame provided in this application;
[0032] Figure 24 A schematic diagram of another embodiment of the re-established virtual frame provided in this application;
[0033] Figure 25 A schematic diagram of one embodiment of the movement of the target control provided in this application;
[0034] Figure 26A schematic diagram of another embodiment of the movement of the target control provided in this application;
[0035] Figure 27 This is a schematic flowchart of an embodiment of the control method for the display device provided in this application. Detailed Implementation
[0036] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of this application, and not all embodiments. Based on the embodiments of this application, all other embodiments obtained by those of ordinary skill in the art without creative effort are within the scope of protection of this application.
[0037] The terms “first,” “second,” “third,” “fourth,” etc. (if present) in the specification, claims, and accompanying drawings of this application are used to distinguish similar objects and are not necessarily used to describe a particular order or sequence. It should be understood that such data can be interchanged where appropriate so that the embodiments of this application described herein can be implemented, for example, in orders other than those illustrated or described herein. Furthermore, the terms “comprising” and “having,” and any variations thereof, are intended to cover a non-exclusive inclusion; for example, a process, method, system, product, or apparatus that comprises a series of steps or units is not necessarily limited to those steps or units explicitly listed, but may include other steps or units not explicitly listed or inherent to such processes, methods, products, or apparatus.
[0038] The concepts involved in this application will first be described with reference to the accompanying drawings. It should be noted that the following description of various concepts is only to make the content of this application easier to understand and does not imply a limitation on the scope of protection of this application. Among them, the terms "module", "unit", "component", etc. used in various embodiments of this application can refer to any known or later developed hardware, software, firmware, artificial intelligence, fuzzy logic, or combination of hardware and / or software code capable of performing the functions associated with the element.
[0039] Figure 1 This is a schematic diagram illustrating the operating scenario of the display device used in this application. Figure 1As shown, the user can operate the display device 200 through the control device 100. Alternatively, the video acquisition device 201, such as a camera, installed on the display device 200 can also acquire video data including the user's body and respond to the user's gestures, body information, etc., based on the images in the video data, and then execute corresponding control commands based on the user's action information. This allows the user to control the display device 200 without the need for the remote control 100, thereby enriching the functionality of the display device 200 and improving the user experience.
[0040] like Figure 1 As shown, the display device 200 can also communicate with the server 300 via various communication methods. In various embodiments of this application, the display device 200 may establish a wired or wireless communication connection with the server 300 via a local area network, a wireless local area network, or other networks. The server 300 may provide the display device 200 with various content and interactive features.
[0041] For example, display device 200 interacts by sending and receiving information, as well as with an Electronic Program Guide (EPG), receiving software updates, or accessing a remotely stored digital media library. Server 300 can be a group or multiple groups, and can be one or more types of servers. Other network services, such as video-on-demand and advertising services, are provided through server 300.
[0042] The display device 200 can be, in one sense, a liquid crystal display, an OLED (Organic Light Emitting Diode) display, or a projection display device; in another sense, it can be a smart TV or a display system consisting of a monitor and a set-top box. The specific type, size, and resolution of the display device are not limited, but those skilled in the art will understand that the display device 200 can be modified in terms of performance and configuration as needed.
[0043] In addition to providing broadcast television reception functionality, the display device 200 may also include smart network television functionality that provides computer support. Examples include IPTV, smart TV, Internet Protocol TV (IPTV), etc. In some embodiments, the display device may not have broadcast television reception functionality.
[0044] In other examples, the display device 200 may have additional functions or fewer of the aforementioned functions. This application does not specifically limit the implementation of the display device 200; for example, the display device 200 may be any electronic device such as a television set.
[0045] For example, Figure 2 This is a schematic diagram of the hardware structure of a hardware system in a display device. For example... Figure 2 It shows Figure 1 The display device 200 may specifically include: a panel 1, a backlight assembly 2, a motherboard 3, a power board 4, a back cover 5, and a base 6. The panel 1 is used to display the image to the user; the backlight assembly 2, located below the panel 1, is typically an optical component used to provide sufficient brightness and a uniformly distributed light source, enabling the panel 1 to display images normally. The backlight assembly 2 also includes a back plate 20, on which the motherboard 3 and power board 4 are mounted. Typically, some protruding structures are stamped into the back plate 20, and the motherboard 3 and power board 4 are fixed to the protrusions by screws or hooks; the back cover 5 covers the panel 1 to conceal the backlight assembly 2, motherboard 3, and power board 4, achieving an aesthetically pleasing effect; the base 6 supports the display device. Optionally, Figure 2 The device also includes a keypad, which can be mounted on the back panel of the display device; this application does not limit this.
[0046] In addition, the display device 200 may also include a sound reproduction device (not shown in the figure), such as an audio component, including an I2S interface with a power amplifier (AMP) and a speaker, for the purpose of sound reproduction. Typically, the audio component can achieve at least two channels of sound output; to achieve a panoramic surround sound effect, multiple audio components are required to output multiple channels of sound, which will not be described in detail here.
[0047] It should be noted that the display device 200 can be implemented using specific forms such as an OLED display screen. Thus, for example... Figure 2 The template included in the display device 200 shown has been changed accordingly, which will not be described in detail here. This application does not limit the specific internal structure of the display device 200.
[0048] The control method of the display device provided in this application will be described below with specific embodiments. These specific embodiments can be combined with each other, and the same or similar concepts or processes may not be described again in some embodiments.
[0049] In some embodiments, the execution subject of the control method for the display device provided in this application can be the display device itself, specifically a controller or control unit, processor, or processing unit such as a CPU, MCU, or SOC in the display device. In subsequent embodiments of this application, a controller is used as an example of the execution subject. When the controller acquires video data through the video acquisition device of the display device, it performs gesture recognition based on multiple consecutive frames of the video data, and then executes the corresponding action based on the recognized gesture information.
[0050] In some embodiments, Figure 3This is a schematic diagram of an embodiment of a control method for a display device, wherein when the controller acquires video data from the video acquisition device... Figure 3 The image to be detected on the right is used to identify gesture A within that image. A gesture recognition algorithm identifies the gesture information in the image, including an "OK" shape, its position, and size. Subsequently, the controller, based on the cursor position on the "OK" control displayed on the screen, determines the corresponding control command for the "OK" gesture as "click the OK control," and then executes this command.
[0051] In other embodiments, Figure 4 This is a schematic diagram of another embodiment of a control method for a display device. In this method, after the controller recognizes gestures in each frame of video data from a video acquisition device, it compares two frames of the image to be detected and determines that the user's gesture B in the image to be detected has moved from the left side of the previous frame to the right side of the next frame, indicating that the user's gesture B has moved. Subsequently, the controller can determine that the control command corresponding to the gesture information is "move the cursor to the right" based on the currently displayed content on the screen, which is a moving cursor C. The distance moved can be related to the movement distance corresponding to the gesture information in the image to be detected. Subsequent embodiments of this application will provide a method for calculating the correlation between the gesture movement distance in the image to be detected and the cursor movement distance on the display screen.
[0052] From the above Figure 3 and Figure 4 As can be seen from the illustrated embodiment, when the controller in the display device can determine the user's gesture information through video data collected by the video acquisition device, and then execute the control commands expressed by the user through gestures, the user can control the display device without relying on remote control, mobile phone or other control devices. This enriches the functions of the display device, adds fun to controlling the display device, and can greatly improve the user experience of the display device.
[0053] This application does not limit the specific method by which the controller determines the gesture information in a frame of an image to be detected. For example, a machine learning model can be used to identify the gesture information in the image to be detected based on image recognition.
[0054] In some embodiments, this application also provides a method for recognizing gesture information, which can determine the gesture information of the hand by defining the coordinates of key points of the human hand in the image to be detected, and can be better applied to the scenario of display devices. For example... Figure 5 This is a schematic diagram of the coordinates of key hand points provided in this application, as shown in... Figure 5In the example shown, the human hand is marked with 21 key points numbered 1-21 according to the positions of the fingers, joints, and palm.
[0055] Figure 6 This is a schematic diagram illustrating different extension and contraction states of key hand points provided in this application. When the controller recognizes gesture information in the image to be detected, it first determines the orientation of the hand in the image using image recognition algorithms. If the image includes key points on the palm side, it continues to identify all key points and determine the position of each key point. For example, Figure 6 In the leftmost image, the distances between key points 9-12 corresponding to the middle finger are relatively sparse and dispersed, indicating that the middle finger is in an extended state. Figure 6 In the middle image, the key points between 9 and 12 corresponding to the middle finger of the hand are more concentrated in the upper part and more dispersed in the lower part, indicating that the middle finger is in a semi-bent state; Figure 6 In the image on the right, the key points 9-12 corresponding to the middle finger are close together and clustered, indicating that the middle finger is in a fully curled state. Therefore, different distances and distribution ratios between key points can be defined to... Figure 6 To distinguish between different states, then according to Figure 6 The same method can be used for Figure 5 After recognizing each key point corresponding to the five fingers, the gesture information in the image to be detected is obtained.
[0056] In some embodiments, this application also provides a control method for a display device, wherein the controller can recognize gesture information and limb information in an image to be detected, and determine and execute control commands based on these two types of information. For example, Figure 7 A schematic diagram illustrating an application scenario of the control method for the display device provided in this application. Figure 7 In the scene shown, the specific structure of the display device 200 is similar to... Figures 1-2 As shown in the diagram, at this time, the user of the display device 200 can use gestures and limbs to express control commands. Subsequently, after the display device 200 acquires video data through its video acquisition device, the controller in the display device 200 identifies the image to be detected in the multi-frame images, and at the same time identifies the user's gesture information and limb information in the image to be detected.
[0057] For example, Figure 8 This is a schematic diagram illustrating the use of gesture information and body information to jointly determine control commands, as provided in this application, wherein it is assumed that... Figure 8 If the gesture information F on the left side is the "OK" gesture, and the body information G is the elbow pointing to the upper left corner, then the control command that can be determined based on the gesture information F and the body information G is to click the control displayed on the left side of the screen. Figure 8The gesture information H on the right side is the "OK" gesture, and the body information I is an elbow pointing to the upper right corner. Therefore, the control command that can be determined based on the gesture information H and the body information I is to click the control displayed on the right side of the screen.
[0058] As can be seen from the above embodiments, the control method for the display device provided in this embodiment enables the controller to determine different control commands based on gesture information and limb information in the image to be detected. This increases the number of control commands that users can issue to the display device using this interactive method, further improving the intelligence level of the display device and the user experience.
[0059] In some embodiments, if the computing power of the display device's controller supports it, the controller can perform gesture and limb information recognition on each frame of the image to be detected extracted from the video data. However, since the computational load required for common gesture and limb recognition is large, it greatly increases the computational load required by the controller. Furthermore, users are not constantly controlling the display device most of the time. Therefore, the display device provided in this application is equipped with at least two detection models, denoted as the first detection model and the second detection model. The second detection model is used to recognize gesture and limb information in the image to be detected, while the first detection model has a smaller computational load and data volume than the second detection model and can be used to identify whether the image to be detected includes gesture information. The following describes... Figure 9 The control method of the display device provided in this application will be described in detail.
[0060] Figure 9 A schematic flowchart of an embodiment of the control method for the display device provided in this application is shown below. Figure 9 The control methods shown include:
[0061] S101: Extract one frame of the image to be detected from a series of consecutive frames of video data acquired by the video acquisition device of the display device at a preset time interval.
[0062] This application can be applied to, for example, Figure 7 In the scenario shown, the process is executed by the controller within the display device. When the display device is in operation, its video acquisition device collects video data in its orientation direction. The controller, acting as the execution entity, then extracts a frame of the image to be detected from the video data at preset time intervals. For example, if the frame rate of the video data collected by the video acquisition device is 60 frames per second, the controller can sample at a frame rate of 30 frames per second, extracting one frame of the image to be detected at each frame interval for subsequent processing. In this case, the preset time interval is 1 / 30 of a second.
[0063] S102: Use the first detection model to determine whether the image to be detected contains human gesture information.
[0064] Specifically, targeting Figure 7 In the application scenario, when a user needs to control the display device, they can stand in the direction the video capture device is facing and make corresponding gestures and body movements according to the control commands they want to receive from the display device. At this time, the video capture device will capture images that include the user's gesture and body information. When the user does not need to control the display device, the video images captured by the video capture device within its capture range will not include the user's gesture and body information.
[0065] Therefore, if the image to be detected before S102 does not contain gesture information and the second detection model is not used to process the image to be detected, the controller will use the first detection model with less computational load to process the image to be detected in S102, and determine whether the image to be detected contains gesture information through the first detection model.
[0066] In some embodiments, the controller uses a gesture category detection model as the first detection model to implement a global perception algorithm, thereby determining whether the image to be detected contains gesture information. The global perception algorithm refers to an algorithm that the controller can start by default and maintain running after power-on. It has the characteristics of low computational load and simple detection types, and can be used only to obtain specific information and for other non-global functions such as starting a second detection model for detection.
[0067] In some embodiments, the first detection model is trained using multiple training images, each of which includes different gesture information to be trained. The controller then uses the first detection model to compare the learned gesture information with the image to be detected, thereby determining whether the image to be detected contains gesture information. However, the first detection model may not be used to specifically identify gesture information, while the second detection model may be used to determine gesture information through specific joint recognition algorithms.
[0068] S103: If it is determined in S102 that the image to be detected includes human gesture information, then it is determined that the user wants to control the display device. The controller then continues to acquire the image to be detected and uses the second detection model to recognize the gesture information and limb information in the image to be detected.
[0069] In some embodiments, after detecting that the image to be detected includes human gesture information, the controller can continue to extract the image to be detected from multiple frames of images captured by the video acquisition device at preset time intervals, and use a second detection model instead of the first detection model to process the subsequently extracted images to be detected, thereby identifying the gesture information and limb information of each frame of the image to be detected. Alternatively, the controller can reduce the preset time interval to extract the image to be detected in fewer time intervals.
[0070] In some embodiments, the controller may process the image to be detected that includes human gesture information in S102 using the second detection model, and then continue to process subsequent detection images using the second detection model.
[0071] S104: Determine the corresponding control command based on the gesture and limb information in the image to be detected from the preset number of frames determined in S103, and execute the control command.
[0072] In some embodiments, to improve the accuracy of recognition, the controller can continuously acquire multiple frames of images for processing. For example, when it is determined in S102 that the image to be detected includes human gesture information, in S103, after acquiring a preset number (e.g., 3) of images to be detected at a preset time interval, gesture information recognition and limb information recognition are performed on these 3 images to be detected respectively. Finally, when the gesture information and limb information in these 3 images to be detected are the same, it is determined that subsequent calculations will be performed based on these identical gesture information and limb information, which can prevent inaccurate recognition caused by occasional errors due to other factors.
[0073] When the gesture and limb information in the aforementioned preset number of images to be detected are all identical (or partially identical, and the proportion of partially identical information is greater than a threshold (e.g., the threshold could be 80%), the controller then determines the control command corresponding to that gesture and limb information based on the mapping relationship. For example, Figure 10 This is a schematic diagram of an embodiment of the mapping relationship provided in this application. The mapping relationship includes multiple control commands (control command 1, control command 2, etc.) and a correspondence between each control command and corresponding gesture information and limb information. For example, control command 1 corresponds to gesture information 1 and limb information 1, control command 2 corresponds to gesture information 2 and limb information 2, and so on. The specific implementation can be found in [reference needed]. Figure 8 Different combinations of gesture and body information can correspond to different control commands.
[0074] In some embodiments, the above mapping relationship can be preset or specified by the user of the display device, and can be stored in the controller in advance, so that the controller can determine the corresponding control command from the mapping relationship and continue to execute it based on the gesture information and limb information it determines.
[0075] In other embodiments, Figure 11 This is a schematic diagram of another embodiment of the mapping relationship provided in this application. Figure 11In the mapping relationship shown, gesture information and body information are each associated with a control command. At this time, the controller can determine a control command based on the gesture information or body information, and then use the other information to verify the determined control command, thereby improving the accuracy of the obtained control command. When the control commands determined by the two pieces of information are different, it indicates that there is an error in recognition. In this case, the control command can be not executed or re-recognition can be performed to prevent the execution of incorrect control commands.
[0076] In some other embodiments, the mapping relationship provided in this application may also include control commands corresponding to "not executing any command", for example, Figure 12 This application provides a schematic diagram of gesture and body information in an image, wherein the user in the image has their back to the display device, and their hands are facing the display device. Although the user does not intend to control the display device, they are... Figure 9 The illustrated process involves a first detection model determining that the current image to be detected contains gesture information. Subsequently, a second detection model identifies this gesture information and limb information. The controller can then determine, based on the mapping relationship, that it should not execute any commands regarding the current gesture and limb information. This mapping relationship could include, for example, a gesture indicating an open palm and limb information indicating an elbow pointing diagonally downwards.
[0077] In summary, the control method for the display device provided in this embodiment enables the controller to determine different control commands based on gesture and limb information in the image to be detected. This expands the number of control commands that users can issue to the display device using this interactive method, further improving the intelligence of the display device and the user experience. Furthermore, this embodiment uses a first detection model with lower computational complexity to identify whether the image to be detected includes gesture information. Only after the first detection model determines that gesture information is included is a second detection model with higher computational complexity used to identify the gesture and limb information. This reduces the computational load and power consumption caused by invalid recognition and improves the controller's computational efficiency.
[0078] In combination with the above Figure 9 In the specific implementation of steps S101-S104, when the control command is a one-time control operation such as clicking a control displayed on the screen, returning to the homepage, or modifying the volume, as shown in Figure S104, after executing the control command, the process ends, the second detection model stops recognizing gesture and limb information, and the process returns to S101 to continue extracting the image to be detected, and the first detection model is used again to recognize gesture information, thereby re-executing the process as described above. Figure 9 The entire process is shown below.
[0079] In another specific implementation, when the control command is a movement command that moves the target control such as the mouse on the display screen to the position corresponding to the gesture information, after the movement command is executed in S104, it should return to S103 and repeat the process of S103-S104 to detect the user's continuous movement actions and realize the continuous movement of the target control on the display screen.
[0080] In some embodiments, during the repeated execution of S103-S104, if the human gesture and limb information in the currently acquired preset number of images to be detected corresponds to a stop command, or if the second detection model determines that the preset number of images to be detected do not contain human gesture and limb information, the process can be terminated, the use of the second detection model to identify gesture and limb information can be stopped, and the process can return to S101 to continue extracting images to be detected, and the first detection model can be used again to identify gesture information, thereby re-executing the process as described above. Figure 9 The entire process is shown below.
[0081] In some embodiments, when the control command is a movement command that controls the target control such as the mouse on the display screen to move to the position corresponding to the gesture information, and the controller is repeatedly executing S103-S104, it is understood that the user's gesture should be in a continuous movement state. If the movement is too fast, the controller may fail to detect the gesture information and limb information in multiple frames of the image to be detected during a certain detection process. In this case, the controller may not stop executing the process immediately, but can predict the gesture information and limb information that may appear at the moment based on the previous or multiple detection results, and execute the subsequent movement command based on the predicted gesture information and limb information.
[0082] For example, Figure 13This is a schematic diagram of an embodiment of the target control movement position provided in this application. After the controller executes S103 for the first time and detects gesture information K and limb information L in the image to be detected, in S104, it executes a movement command to move the target control to position ① on the display screen. After the controller executes S103 for the second time and detects gesture information K and limb information L in the image to be detected, in S104, it executes a movement command to move the target control to position ② on the display screen. However, assuming that the user moves too fast after the second detection, the controller fails to recognize the gesture information and limb information in the image to be detected when it executes S103 for the third time, and fails to move the target control on the display screen. Subsequently, when the controller executes S103 for the fourth time and is able to detect the gesture information K and limb information L in the image to be detected, in S104, when it executes the movement command to move the target control to position ④ on the display screen, the change in moving the target control directly from position ② to position ④ on the display screen is too large, causing the user to experience a paused or stuttering viewing effect, which greatly affects the user experience.
[0083] Therefore, in this embodiment, when the controller executes S103 for the third time and fails to recognize gesture information and limb information in the image to be detected, since the target control is still being moved on the display screen, the controller can predict the gesture information K and limb information L that may appear in the image to be detected in the third time based on the movement comfort and movement direction of the gesture information K and limb information L recognized in the first and second times. Then, based on the predicted position corresponding to the predicted gesture information and limb information, the controller executes a movement command to move the target control to position ③ on the display screen.
[0084] final, Figure 14 This is a schematic diagram of another embodiment of the movement position of the target control provided in this application. When the above prediction method is used, for the hand gesture information and limb information changing according to ①-②-③-④ in the image to be detected collected at the same time interval, although the hand gesture information and limb information are not recognized in the image to be detected during the third execution of S103, the position ③ on the display screen is still predicted based on the predicted hand gesture information and limb information. This ensures that the target control on the display screen will change its position uniformly according to ①-②-③-④ throughout the process, avoiding... Figure 13 The pause and lag when the target control moves directly from position ② to position ④ greatly improves the display effect, making the operation of the display device by the user through gestures and body movements more smooth and fluid, and further improving the user experience.
[0085] To achieve the above process, in some embodiments, after each execution of S103, the controller stores and records the gesture and limb information obtained in this execution of S103, so as to make predictions if gesture and limb information are not detected in subsequent executions. In some embodiments, if no gesture and limb information is detected when the process in S103 is executed multiple times consecutively (e.g., 3 times), prediction is no longer performed, and the current process is stopped, and execution restarts from S101.
[0086] Based on the above embodiments, in the specific implementation process, the controller can maintain a gesture movement speed v and movement direction α according to the recognition results of the second detection model. The gesture movement speed v and movement direction α can be obtained based on the frame rate and the movement distance between multiple frames (generally three frames). When a gesture cannot be detected but the limb can be detected, multi-frame action prediction (generally three frames) will be added to prevent situations such as focus reset and mouse lag that affect the user experience due to the sudden loss of gesture detection. The predicted gesture position for the next frame can be obtained based on the movement speed v and direction α. Of course, a speed threshold β is required. If the gesture movement speed exceeds the threshold β, it will be fixed at speed β. This is to prevent the speed of the gesture from being too fast and affecting the user experience.
[0087] In some embodiments, when using the second detection model to identify gesture and limb information in the image to be detected in the above example, the recognition result of a single frame of the image to be detected is not taken as the standard. Instead, a preset number of images to be detected are extracted over a preset time interval, and gesture and limb information is detected in all of these images before the control commands corresponding to the same gesture and limb information are executed. In the specific implementation process, the controller of the display device can dynamically adjust the preset time interval according to the operating parameters of the display device. For example, the controller determines the preset time interval to be 100ms when the current load is light, that is, one frame of the image to be detected is extracted every 100ms. Assuming the preset number is 8, the preset number of images to be detected corresponds to a time range of 800ms. If the controller detects gesture and limb information in all 8 frames of the image to be detected within this time range, it indicates that the gesture and limb information is real and valid, and the control commands corresponding to the same gesture and limb information can be executed. When the controller determines that the load is heavy based on the current load exceeding a threshold, it sets a preset time interval of 200ms, meaning that one frame of the image to be detected is extracted every 200ms. At this point, the controller can adjust the preset number to 4, thus ensuring the authenticity and validity of gesture and limb information within the 800ms time range corresponding to the 4 frames of the image to be detected. Therefore, in the control method provided in this embodiment, the controller can dynamically adjust the preset number according to the preset time interval, and the two are inversely proportional. This reduces the computational load on the controller under heavy load and prevents the recognition time from being prolonged due to a large preset number when the preset time interval is long. Ultimately, it achieves a certain level of recognition efficiency while ensuring recognition accuracy.
[0088] In some embodiments, Figure 15 A flowchart illustrating an embodiment of the control method for the display device provided in this application can be used as... Figure 9 The control method shown is a specific implementation method, and its specific implementation method and principle are similar to... Figure 9 The same applies as shown, so I will not repeat it again.
[0089] In some embodiments, the controller uses a second detection model to identify human gesture information in the image to be detected, while the first detection model is also trained on images including gesture information. Therefore, after each execution, the controller... Figure 9 After the entire process shown, the gesture information identified by the second detection model during this execution can be used to train and update the first detection model, thereby enabling a more effective update of the first detection model based on the currently detected gesture information, thus improving the real-time performance and applicability of the first detection model.
[0090] In the specific implementation of the foregoing embodiments of this application, although the display screen can be controlled based on gesture and limb information in the image to be detected, the human body may only be located in a small part of the image to be detected captured by the video acquisition device of the display device. This results in the human body's gesture information moving a long distance when the user performs a long-distance movement operation to control the controls on the display screen, causing inconvenience to the user. Therefore, this application also provides a control method for a display device. By establishing a mapping relationship between a "virtual frame" in the image to be detected and the display screen, the user can control the display device by simply moving their gestures within the virtual frame to instruct the target control to move on the display screen, greatly reducing the user's range of motion and improving the user experience. The "virtual frame" and related applications provided in this application will be described below with reference to specific embodiments. The virtual frame is only an exemplary name and may also be called a mapping frame, recognition area, mapping area, etc. This application does not limit its name.
[0091] For example, Figure 16 A schematic flowchart of an embodiment of the control method for the display device provided in this application is shown below. Figure 16 The method shown can be applied to Figure 7 In the scenario shown, the method is executed by a controller in the display device and is used to recognize a user's movement command to move the control via gesture information when the display device displays controls such as a mouse. Specifically, the method includes:
[0092] S201: When the display device is in operation, its video acquisition device will collect video data in its orientation direction. The controller, acting as the execution entity, then acquires the video data and extracts a frame of the image to be detected from the video data at preset time intervals. It then identifies the human gesture information in the image to be detected.
[0093] The specific implementation of S201 can be referenced from S101-S103. For example, the controller can use the first detection model to determine whether each extracted image to be detected includes gesture information, and use the second detection model to identify the gesture information and limb information in the image to be detected that includes gesture information. The specific implementation and principle will not be elaborated here. Alternatively, in S201, when the target control is displayed on the display device or when an application that needs to display the target control is running, it can be indicated that the target control may need to be moved. Therefore, after each acquisition of the image to be detected, the second detection model is directly used to identify the gesture information and / or limb information in the image to be detected. The identified gesture information and / or limb information can be used to determine the movement command in the future.
[0094] S202: After the first image to be detected is extracted and recognized in S201, the controller determines that the first image to be detected includes gesture information. Then, the controller establishes a virtual frame based on the gesture information in the first image to be detected, and establishes a positional mapping relationship between the virtual frame and the display screen of the display device. The controller can display the target control at a preset first display position, wherein the first display position can be the center position of the display screen.
[0095] For example, Figure 17 This is a schematic diagram of the virtual frame provided in this application. When the first image to be detected includes gesture information K and limb information L, and the gesture information and limb information are an outstretched palm corresponding to a command to move a target control displayed on the screen, then the controller establishes a virtual frame centered on the first focal position P where the gesture information K is located, and displays the target control at the center position of the screen. In some embodiments, the shape of the virtual frame can be rectangular, and the aspect ratio of the rectangle is the same as the aspect ratio of the screen, but the area of the virtual frame can be different from the area of the screen. Figure 17 As shown, the positional mapping relationship between the virtual frame and the display screen is represented by the dotted line in the figure. In this positional mapping relationship, the focal point P of the virtual frame corresponds to the midpoint Q of the display screen, and the four vertices of the rectangular virtual frame correspond to the four vertices of the rectangular display screen. Since the aspect ratio of the virtual frame is the same as the aspect ratio of the display screen, a focal point position inside the rectangular virtual frame can correspond to a display position on the display screen. This allows the display position on the display screen to change accordingly when the focal point position inside the rectangular virtual frame changes.
[0096] In some embodiments, the above mapping relationship can be represented by the relative distance between the focus position in the virtual frame and a target position within the virtual frame, and the relative distance between the display position on the display screen and the same target position on the display screen. For example, assuming the vertex P0 at the lower left corner of the virtual frame is the origin, the coordinates of point P can be represented as (x, y); assuming the vertex Q0 at the lower left corner of the display screen is the origin, the coordinates of point Q can be represented as (X, Y). Then the mapping relationship can be represented as: X / x along the long side of the rectangle and Y / y along the wide side of the rectangle.
[0097] In S201-S202 above, the controller completes the establishment of the rectangular virtual frame and the mapping relationship. Subsequently, in S203-S204, the virtual frame and the mapping relationship can be applied so that the movement of the focus position corresponding to the gesture information can correspond to the movement of the target control position on the display screen.
[0098] S203: When the second image to be detected includes gesture information, and the second focus position corresponding to the gesture information is in the rectangular virtual box, determine the second display position on the display screen according to the second focus position and the mapping relationship.
[0099] S204: Move the target control on the display screen to the second display position determined in S203.
[0100] Specifically, Figure 18 This application provides a schematic diagram illustrating the correspondence between a virtual frame and a display screen. Assuming that a virtual frame is established in the first image to be detected at the first focal point P of the human gesture information, a target control "mouse" can be simultaneously displayed at the first display position Q in the center of the display screen. Subsequently, when the second focal point P' of the human gesture information within the virtual frame in the second image to be detected moves relative to the first image towards the upper right, the controller can determine the second relative distance between the corresponding second display position Q' on the display screen and the lower left target position on the display screen based on the first relative distance between the second focal point and the target position within the virtual frame, combined with the proportion in the mapping relationship. Finally, the controller can calculate the actual position of the second display position Q' on the display screen based on the second relative distance and the coordinates of the lower left target position, and display the target control at the second display position Q'.
[0101] Figure 19 A schematic diagram of the movement of the target control provided in this application is shown, wherein, as shown... Figure 18 During the process shown, when the human gesture information between the first image to be detected and the second image to be detected moves from the first focal position P to the second focal position P', the controller can display the target control on the first display position Q and the second display position Q' on the display screen according to the change of the focal position within the virtual frame. In this process, the user perceives that the target control displayed on the display screen moves accordingly following the movement of its gesture information.
[0102] It is understandable that the above S203-S204 process can be executed repeatedly, and the display position can be determined for the focus position corresponding to the gesture information in each detected image, and the target control can be repeatedly and continuously controlled to move on the display screen.
[0103] In this embodiment, the location of the gesture information is taken as the focus position, for example, a key point in the gesture information is taken as the focus position. In other embodiments, the key point of the limb information can also be taken as the focus position, etc. The implementation method is the same, and will not be described again.
[0104] Furthermore, it should be noted that the above examples use the first and second images to be detected as single frames, such as... Figure 16 It can also be with, for example Figure 9 The methods shown are combined, and the image to be detected includes multiple frames of images to be detected, thereby determining the corresponding focus position based on the gesture information identified in the multiple frames of images to be detected.
[0105] In summary, the control method for the display device provided in this embodiment can establish a mapping relationship between a "virtual frame" in the image to be detected and the display screen, so that when the user controls the display device, he or she can move the target control on the display screen simply by moving his or her hand within the virtual frame, which greatly reduces the range of user movements and improves the user experience.
[0106] In the specific implementation of the above embodiments, when the controller creates a virtual frame, the size of the virtual frame can be related to the distance between the human body and the video acquisition device. For example, Figure 20 This is a schematic diagram of the area of the virtual frame provided in this application. When the distance between the human body and the video acquisition device is far, the area corresponding to the gesture information in the image to be detected is small, so a smaller virtual frame can be set; when the distance between the human body and the video acquisition device is close, the area corresponding to the gesture information in the image to be detected is large, so a larger virtual frame can be set. The area of the virtual frame can have a linear relationship proportional to the distance, or it can be divided into multiple levels of mapping relationship based on the distance (i.e., a certain frame size corresponds to a certain distance). The specific mapping relationship can be adjusted according to the actual situation. In some embodiments, the controller can determine the distance between the human body and the display device (the video acquisition device is set on the display device) based on the infrared or other arbitrary form of ranging unit set on the display device. Alternatively, the controller can also determine the corresponding distance based on the area corresponding to the gesture information in the image to be detected, and then determine the area of the virtual frame based on the area of the gesture information, etc.
[0107] In some embodiments, when the established virtual bounding box is close to the edge of the image to be detected, the accuracy of recognizing gesture information may be reduced due to limitations such as image recognition processing algorithms. Therefore, the controller can also establish an optimal control range for the edge region surrounding the edge of the image to be detected. For example, Figure 21 The schematic diagram of the edge region provided in this application shows that the edge region refers to the area within the image to be detected, outside the optimal control range, where the distance between it and a boundary of the image to be detected is less than a preset distance. Figure 21In the image to be detected above, assuming the virtual bounding box constructed based on the gesture information in the first image to be detected is completely outside the edge region and within the optimal control range, subsequent calculations can continue. However, if a portion of the virtual bounding box constructed by the controller based on the gesture information in the first image to be detected is within the edge region, then... Figure 21 In the image to be detected below, if the left side of the virtual frame is located within the edge area, the controller can compress the virtual frame horizontally, resulting in a horizontally compressed virtual frame. It can be understood that a positional mapping relationship can then be established between the compressed virtual frame and the display screen. At this point, the movement distance of the focus position corresponding to the gesture information will correspond to a larger change in display position on the screen. Although this results in a faster horizontal movement of the target control for the user, it avoids the controller recognizing gesture information from the edge area of the image to be detected, thus improving the accuracy of gesture recognition and the overall control process.
[0108] In the above embodiments, a virtual frame in the image to be detected is provided, allowing the user to control the movement of a target control on the display screen by moving gestures within the virtual frame. However, in some cases, due to large movements or overall body movement, the user's gestures may move outside the virtual frame, resulting in unrecognizable gestures and affecting the control effect. For example, Figure 22 This is a schematic diagram of the state of gesture information provided in this application. In state S1, the second image to be detected includes gesture information, and the second focus position corresponding to the gesture information can be inside the established virtual frame K1. At this time, the control method in the aforementioned embodiment can be executed normally, and the display position of the target control is determined by the focus position of the gesture information in the virtual frame. Figure 22 In state S2, the second image to be detected includes gesture information, and the second focus position corresponding to the gesture information may appear outside the virtual box K1 in the image to be detected. At this time, it will be impossible to determine the display position of the target control through the focus position of the gesture information in the virtual box.
[0109] Therefore, when the controller recognizes that the second focus corresponding to the gesture information in the second image to be detected is located at point P2 outside the virtual frame, it can re-establish the virtual frame K2 with the point P2 where the second focus is located at this time as the center, and establish the mapping relationship between the virtual frame K2 and the display screen. Figure 23 A schematic diagram of an embodiment of the re-established virtual frame provided in this application shows that, Figure 23Within the newly re-established virtual frame K2, the second focus position P2 is located at the center of the virtual frame K2. Therefore, the controller still needs to control the target control to be displayed in the center position on the display screen according to the second focus position P2, so as to give the user the viewing effect of resetting the target control, thereby avoiding the problem of being unable to control the target control due to the removal of the virtual frame by gesture information.
[0110] Figure 24 A schematic diagram of another embodiment of the re-established virtual frame provided in this application, wherein, in this manner, when... Figure 22 When the gesture information shown in state S2 appears outside the virtual frame K1 in the image to be detected, the controller resets the virtual frame. At this time, assuming the controller displays the target control at the first relative position Q1 on the display screen based on the position information of the gesture information in the virtual frame K1 in the previous image, the virtual frame K2 is re-established based on the relative position of the first relative position Q1 on the entire display screen, ensuring that the relative position of the second focus position P2 within the virtual frame K2 is the same as the relative position of the first relative position Q1 on the display screen. Therefore, the controller can continue to display the target control at the first relative position Q2, completing the reset of the virtual frame K2 without the target control jumping to the center position of the display screen. In subsequent images to be detected, when the gesture information changes within the virtual frame K2, the controller can determine the display position of the target control based on the focus position of the gesture information within the virtual frame K2, thus achieving focus reset without the user's knowledge. This avoids the problem of being unable to control the target control after the gesture information is removed from the virtual frame, making the entire process smoother and further improving the user experience.
[0111] In some embodiments, after the controller performs the above process and re-establishes the virtual frame, it can display relevant prompts on the display screen to inform the user that the rectangle has been re-established and to display information about the re-established rectangle. For example, the controller can display text, images, or other information at the edge of the display screen to indicate that the virtual frame has been rebuilt. Alternatively, after the controller determines that the virtual frame needs to be re-established during the above process, it can also display a prompt to update the virtual frame on the display screen. After receiving confirmation from the user, it can then execute the process of rebuilding the virtual frame. This makes the entire process user-controllable and ensures that the rebuild is performed according to the user's intention, preventing invalid rebuilds due to user withdrawal.
[0112] In some embodiments, during the movement of the target control described above, if the controller fails to recognize gesture information in a preset number of consecutive images to be detected during the control process, it can stop displaying the target control on the display screen, thereby ending the process. Figure 16 The process is illustrated below. Alternatively, if the controller does not display any gesture information in the image to be detected within a certain preset time period, it can also stop displaying the target control on the screen and end the process. Or, if the controller recognizes that the gesture information in the image to be detected corresponds to a stop command during the control process, it can also stop displaying the target control on the screen and end the process.
[0113] In some embodiments, such as Figure 16 During the execution of the method shown, the controller determines the display position on the screen based on the focus position of the gesture information within the virtual frame in each frame of the image to be detected, and displays the target control at that position. In one specific implementation, Figure 25 This is a schematic diagram of an embodiment of the movement of the target control provided in this application, from... Figure 25 As can be seen, assuming the controller determines that the gesture information in image 1 is located at the focal position P1 within the virtual frame, and thus controls the display screen to show the target control at position Q1; the gesture information in image 2 is located at the focal position P2 within the virtual frame, and thus controls the display screen to show the target control at position Q2; and the gesture information in image 3 is located at the focal position P3 within the virtual frame, and thus controls the display screen to show the target control at position Q3. However, in the above process, because the user may move too quickly during the P1-P2 transition when making the gesture, the target control displayed on the screen will move between Q1 and Q2, giving the user the impression of uneven movement speed and abrupt changes in the target control.
[0114] Therefore, once the controller determines the position of the second focus, the processing performed by the controller can be referred to... Figure 26 The state changes in, among which, Figure 26 This is a schematic diagram of another embodiment of the movement of the target control provided in this application. (See diagram below.) Figure 26 As shown, after the controller determines the first focus position P1 and the second focus position P2 in the virtual frame, it also compares the distance between the second focus position and the first focus position with a preset time interval. If the ratio of the distance between P1 and P2 to the preset time interval (i.e., the interval between extracting the images to be detected where the first and second focus positions are located) is greater than a preset threshold, it indicates that the movement speed of the gesture information is too fast. If the second display position of the target control is determined based on the second focus position and the target control is displayed, it may lead to problems such as... Figure 25The display effect is shown. Therefore, the controller determines a third focal position P2' between the first focal position and the second focal position, wherein the distance between the third focal position P2' and the first focal position P1 and the preset time interval is not greater than a preset threshold, and the third focal position P2' can be a point located on the connecting line between P1 and P2, with P1, P2', and P2 being linearly connected. Subsequently, the controller can determine a second display position Q2' on the display screen based on the third focal position P2' and the mapping relationship, and control the target control to move from the first display position Q1 to the second display position Q2'.
[0115] During the aforementioned movement, since the gesture information moves to the second focal position P2, the target displayed on the screen does not move to the display position Q2 corresponding to the second focal position, but instead moves to the second display position Q2' corresponding to the third focal position P2'. Therefore, when the controller processes the third image to be detected after the second image to be detected, if the third image to be detected includes gesture information and the fourth focal position P3 corresponding to the gesture information is located in the rectangular virtual frame, and the distance between the fourth focal position P3 and the third focal position P2' is not greater than a preset threshold, the third display position Q3 corresponding to the fourth focal position can be determined according to the mapping relationship, and the target control on the display screen can be controlled to move from the second display position Q2' to the third display position Q3.
[0116] Ultimately, throughout the entire process described above, when the gesture information moves too quickly between positions P1 and P2, the target control displayed on the screen can reduce its movement length. When the gesture information moves less quickly between positions P2 and P3, the distance "reduced" during the movement between P1 and P2 can be made up. From the user's perspective, when the gesture information moves from position P1 on the left side of the virtual frame to position P3 on the right side, the target control on the screen will also move from position Q1 on the left side of the screen to position Q3 on the right side. Thus, even when the user's gesture information moves too quickly between P1 and P2, the overall movement speed of the target control displayed on the screen between P1 and P3 can be kept relatively constant, giving the user a sense of uniform movement speed and continuous change of the target control.
[0117] Figure 27 A schematic flowchart of an embodiment of the control method for the display device provided in this application is shown below. Figure 27In one specific implementation, the display device's controller first performs gesture detection. If the gesture is normal, it maps the cursor position on the TV interface based on the hand's position within the virtual frame, performing gesture movement control, gesture click detection, and gesture return detection. If the gesture disappears, it performs multi-frame (usually three-frame) motion prediction. If the gesture is detected again during this process, the focus is reset. If the distance is close, movement continues; if the distance is far, the focus is reset to the center of the TV screen. During focus reset, the virtual frame needs to be regenerated. Furthermore, if no gesture is detected multiple times, the mouse cursor on the TV interface is cleared first. If no gesture is detected for an extended period, the gesture recognition process exits, and a global gesture detection scheme is entered until a focused gesture is detected.
[0118] In the foregoing embodiments, the control method for the display device provided in the embodiments of this application has been described. To implement the functions of the methods provided in the embodiments of this application, the display device, as the executing entity, may include a hardware structure and / or software modules, implementing the above functions in the form of a hardware structure, software modules, or a combination of hardware and software modules. Whether a particular function is executed in the form of a hardware structure, a software module, or a combination of hardware and software modules depends on the specific application and design constraints of the technical solution.
[0119] It should be noted that the above division of the display device into various modules is merely a logical functional division. In actual implementation, all or part of these modules can be integrated into a single physical entity, or they can be physically separated. These modules can be implemented entirely in software via processing elements; entirely in hardware; or some modules can be implemented in software via processing elements, while others are implemented in hardware. Processing elements can be separate entities or integrated into a chip within the device. Alternatively, they can be stored as program code in the device's memory, invoked and executed by a processing element. The implementation of other modules is similar. Furthermore, these modules can be integrated together or implemented independently. The processing element described here can be an integrated circuit with signal processing capabilities. In implementation, each step of the above method or each of the above modules can be completed through integrated logic circuits in the processor element or through software instructions.
[0120] For example, these modules can be one or more integrated circuits configured to implement the above methods, such as one or more application-specific integrated circuits (ASICs), one or more digital signal processors (DSPs), or one or more field-programmable gate arrays (FPGAs). As another example, when a module is implemented using processing element scheduler code, the processing element can be a general-purpose processor, such as a central processing unit (CPU) or other processor capable of calling program code. Furthermore, these modules can be integrated together to implement a system-on-a-chip (SOC).
[0121] In the above embodiments, implementation can be achieved entirely or partially through software, hardware, firmware, or any combination thereof. When implemented using software, it can be implemented entirely or partially in the form of a computer program product. The computer program product includes one or more computer instructions. When the computer program instructions are loaded and executed on a computer, all or part of the processes or functions described in the embodiments of this application are generated. The computer can be a general-purpose computer, a special-purpose computer, a computer network, or other programmable device. The computer instructions can be stored in a computer-readable storage medium or transmitted from one computer-readable storage medium to another. For example, the computer instructions can be transmitted from one website, computer, server, or data center to another website, computer, server, or data center via wired (e.g., coaxial cable, fiber optic, digital subscriber line (DSL)) or wireless (e.g., infrared, wireless, microwave, etc.) means. The computer-readable storage medium can be any available medium that a computer can access or a data storage device such as a server or data center that integrates one or more available media. The available medium can be a magnetic medium (e.g., floppy disk, hard disk, magnetic tape), an optical medium (e.g., DVD), or a semiconductor medium (e.g., solid-state disk (SSD)).
[0122] This application also provides an electronic device, including: a processor and a memory; wherein the memory stores a computer program, and when the processor executes the computer program, the processor can be used to execute a control method for a display device as described in any of the foregoing embodiments of this application.
[0123] This application also provides a computer-readable storage medium storing a computer program, which, when executed, can be used to perform a control method for a display device as described in any of the foregoing embodiments of this application.
[0124] This application also provides a chip for executing instructions, the chip being used to execute a control method for a display device executed by an electronic device as described in any of the foregoing embodiments of this application.
[0125] Those skilled in the art will understand that all or part of the steps of the above-described method embodiments can be implemented by hardware related to program instructions. The aforementioned program can be stored in a computer-readable storage medium. When executed, the program performs the steps of the above-described method embodiments; and the aforementioned storage medium includes various media capable of storing program code, such as ROM, RAM, magnetic disks, or optical disks.
[0126] Finally, it should be noted that the above embodiments are only used to illustrate the technical solutions of this application, and are not intended to limit them. Although this application has been described in detail with reference to the foregoing embodiments, those skilled in the art should understand that modifications can still be made to the technical solutions described in the foregoing embodiments, or equivalent substitutions can be made to some or all of the technical features therein. Such modifications or substitutions do not cause the essence of the corresponding technical solutions to deviate from the scope of the technical solutions of the embodiments of this application.
Claims
1. A display device, characterized in that, include: The display screen is configured to display images; The video capture device is configured to capture video data; The controller is configured to extract one frame of image to be detected from a series of consecutive frames of video data acquired by the video acquisition device at preset time intervals; use a first detection model to determine whether the image to be detected contains human gesture information; if so, continue to extract a preset number of images to be detected from the video data at the preset time interval and a preset number, and use a second detection model to identify human gesture information and limb information in the preset number of images to be detected respectively; wherein the amount of data calculated by the first detection model is less than the amount of data calculated by the second detection model; execute control commands corresponding to the gesture information and limb information in the preset number of images to be detected, wherein the limb information includes at least one of elbow pointing to the upper left corner and elbow pointing to the upper right corner; The controller is also configured to: When the control command is a movement command that controls the target control on the display screen to move to the position corresponding to the gesture information and the limb information, if the gesture information and the limb information in multiple frames of images to be detected cannot be detected, the current gesture information and the limb information are predicted based on the previous or multiple detection results, and subsequent movement commands are executed based on the predicted gesture information and limb information. The controller is further specifically configured as follows: When all or part of the gesture information and limb information in the preset number of images to be detected are the same, the control command corresponding to the all or part of the same gesture information and limb information is determined by a mapping relationship; wherein, the mapping relationship includes: multiple control commands, and the correspondence between each control command and the gesture information and limb information; the control command is executed; or, The gesture information and the limb information each correspond to a control command. After a control command is determined based on the gesture information or the limb information, the determined control command is verified using another piece of information. When the control commands determined by the two pieces of information are different, the control command is not executed or is re-identified. The controller is further configured to update the first detection model using a preset number of human gesture information in images to be detected obtained from the second detection model.
2. The display device according to claim 1, characterized in that, The control command is a movement command that controls the target control on the display screen to move to the position corresponding to the gesture information; The controller is further configured to: repeatedly extract a preset number of images to be detected from the video data according to the preset time interval and preset number, and use a second detection model to identify the gesture information and limb information of the human body in the preset number of images to be detected, and execute a control command to move the target control to the position corresponding to the gesture information in each preset number of images to be detected.
3. The display device according to claim 2, characterized in that, The controller is further configured to: when the preset number of images to be detected does not include the gesture information, determine the predicted position of the gesture information in the preset number of images to be detected based on the movement speed and movement direction of the gesture information in the multi-frame images to be detected corresponding to the control command executed last time; and execute a movement command to control the target control to move to the predicted position.
4. The display device according to claim 3, characterized in that, The controller is also configured to store the gesture information and limb information in the preset number of images to be detected.
5. The display device according to any one of claims 1-4, characterized in that, The controller is further configured to: determine the preset quantity based on the preset time interval; wherein the length of the preset time interval is inversely proportional to the value of the preset quantity.
6. The display device according to any one of claims 1-4, characterized in that, The controller is also configured to: after executing the control command, stop using the second detection model to identify the gesture information and limb information of the human body in the preset number of images to be detected; or, When the gesture and limb information of the human body in the preset number of images to be detected is identified as corresponding to a stop command, the second detection model is stopped from identifying the gesture and limb information of the human body in the preset number of images to be detected. or, When the preset number of images to be detected do not contain human gesture information and limb information, the second detection model is stopped from recognizing human gesture information and limb information in the preset number of images to be detected.
7. The display device according to any one of claims 1-4, characterized in that, The controller is also configured to determine a preset time interval corresponding to the operating parameters of the display device.
8. A control method for a display device, characterized in that, include: According to a preset time interval, one frame of the image to be detected is extracted from a series of consecutive frames of video data acquired by the video acquisition device of the display device. The first detection model is used to determine whether the image to be detected contains human gesture information; If so, according to the preset time interval and preset number, a preset number of images to be detected are extracted from the video data, and the second detection model is used to identify the human gesture information and limb information in the preset number of images to be detected respectively; wherein, the amount of data calculated by the first detection model is less than the amount of data calculated by the second detection model, and the limb information includes at least one of elbow pointing to the upper left corner and elbow pointing to the upper right corner; Execute the control commands corresponding to the gesture information and limb information in the preset number of images to be detected; The method further includes: When the control command is a movement command that controls the target control on the display screen to move to the position corresponding to the gesture information and the limb information, if there are gesture information and limb information in multiple frames of images to be detected that cannot be detected, the current gesture information and limb information are predicted based on the previous or multiple detection results, and subsequent movement commands are executed based on the predicted gesture information and limb information. The method further includes: When all or part of the gesture information and limb information in the preset number of images to be detected are the same, the control command corresponding to the all or part of the same gesture information and limb information is determined by a mapping relationship; wherein, the mapping relationship includes: multiple control commands, and the correspondence between each control command and the gesture information and limb information; the control command is executed; or, The gesture information and the limb information each correspond to a control command. After a control command is determined based on the gesture information or the limb information, the determined control command is verified using another piece of information. When the control commands determined by the two pieces of information are different, the control command is not executed or is re-identified. The method further includes: The first detection model is updated using the human gesture information from a preset number of images to be detected obtained from the second detection model.
Citation Information
Patent Citations
Robot-based target image display processing method and system
CN107742303A
Attribute detecting method, device and system and storage medium
CN108875538A
Gesture recognition method and related device
CN111881862A
Gesture recognition method and device
CN112364799A