Display device, human posture detection method and application

By identifying key points and skeletal lines of the user in the display device, and combining facial features and limb IDs, the accuracy problem of user posture detection under long-distance interaction is solved, improving the continuity of user interaction and the gaming experience.

CN116114250BActive Publication Date: 2026-02-13HISENSE VISUAL TECH CO LTD
View PDF 4 Cites 0 Cited by

Patent Information

Application Number
CN202180057495.X
Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
Priority Date
2021-08-20
Filing Date
2021-09-10
Publication Date
2026-02-13
Estimated Expiration
2041-09-10

AI Technical Summary

Technical Problem

Existing display devices struggle to accurately track and recognize user gestures, especially in long-distance interaction scenarios, leading to issues such as recognition loss and discontinuous tracking.

Method used

By capturing images through a camera, identifying key points and skeletal lines of the user, establishing correlations, dynamically adjusting the camera angle to maintain the user's proper position in the frame, and establishing correlations through facial features and body IDs in multi-person interaction scenarios to ensure the accuracy of locking onto and tracking the person.

Benefits of technology

It enables accurate detection and tracking of user posture in long-distance interaction scenarios, reduces recognition loss, and improves the continuity of user interaction and gaming experience.

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN116114250B_ABST
    Figure CN116114250B_ABST
Patent Text Reader

Abstract

Some embodiments of the present application disclose a display device, a human posture detection method and application. The display device comprises a camera configured to capture images; a display configured to display a user interface; an external input device configured to input a current video played on the user interface; and a controller configured to determine that a user is changing channels based on the captured images obtained by the camera, obtain a screenshot of the user interface, and configure optimized image parameters and optimized sound parameters for the current video played on the user interface according to the screenshot.
Need to check novelty before this filing date? Find Prior Art

Description

[0001] This application claims priority to Chinese Patent Application No. 202011262476.4, filed November 12, 2020; No. 202011267208.1, filed November 13, 2020; No. 202110960378.6, filed August 20, 2021, the contents of which are incorporated herein by reference in their entirety. TECHNICAL FIELD

[0002] The present application relates to display device technology, in particular to a display device, a human body posture detection method and application. BACKGROUND

[0003] Display devices are becoming more and more functional, and more display devices are equipped with image capture devices such as cameras. The user image is obtained through the camera, and the "limb movement" application program is used to enable the display device to display the user's body image in real time. When the user's limb movement changes, the application program will also display the changed image, and through the detection program, the posture of the limb movement is detected and corrected to achieve the effect of movement guidance. SUMMARY

[0004] Some embodiments of the present application provide a display device, comprising: a camera configured to capture a target user in a dodgeball game and a background image in a detection area; a display configured to display a dodgeball game user interface containing the target user and the background image; and a controller configured to: control the user interface to display a first identifier at a first position in the user interface according to a first size of the target user in the background image and the first position of the target user in the user interface, wherein the first identifier is used to trigger dodgeball launch, and the size of the first identifier corresponds to the first size; and control the user interface to stop displaying the first identifier when the upper body of the target user is not completely located in the capture area of the camera and the target user is leaving the capture area.

[0005] Some embodiments of the present application provide an application of human body posture detection, comprising: controlling a dodgeball game user interface to display a first identifier at a first position according to a first size of a target user in a background image of the user interface and the first position of the target user in the user interface, wherein the first identifier is used to trigger dodgeball launch, and the size of the first identifier corresponds to the first size; and controlling the user interface to stop displaying the first identifier when the upper body of the target user is not completely located in the capture area of the camera and the target user is leaving the capture area.

[0006] Some embodiments of the present application provide a display device, comprising: a camera configured to capture images; a display configured to display a user interface; an external input device configured to input a current video played on the user interface; a controller configured to: determine, based on the captured images obtained by the camera, that a user is changing channels, obtain a screenshot of the user interface; and configure, according to the screenshot, an optimized image parameter and an optimized sound parameter for the current video played on the user interface.

[0007] Some embodiments of the present application provide an application of human posture detection, comprising: determining, based on captured images obtained, that a user is changing channels, obtaining a screenshot of a user interface; and configuring, according to the screenshot, an optimized image parameter and an optimized sound parameter for a current video played on the user interface.

[0008] Some embodiments of the present application provide a display device, comprising: a display configured to display an image frame; and a controller connected to the display, the controller being configured to: perform human posture detection on a current frame image in a video to be detected, determine a human detection box in the current frame image; determine whether a size and / or position of a target human detection box in the current frame image changes compared to a target human detection box in a previous frame image of the current frame image; if the size and / or position of the target human detection box in the current frame image changes compared to the target human detection box in the previous frame image, and the change amplitude is less than a preset change threshold, adjust the size and / or position of the target human detection box in the current frame image based on the target human detection box in the previous frame image, and send the adjusted current frame image to the display for display.

[0009] Some embodiments of the present application provide a human posture detection method, comprising: performing human posture detection on a current frame image in a video to be detected, determining a human detection box in the current frame image; determining whether a size and / or position of a target human detection box in the current frame image changes compared to a target human detection box in a previous frame image of the current frame image; if the size and / or position of the target human detection box in the current frame image changes compared to the target human detection box in the previous frame image, and the change amplitude is less than a preset change threshold, adjusting the size and / or position of the target human detection box in the current frame image based on the target human detection box in the previous frame image, and displaying the adjusted current frame image. BRIEF DESCRIPTION OF DRAWINGS

[0010] Figure 1 Fig. 1 is a schematic diagram of an operating scenario between a display device and a control device according to an embodiment of the present application;

[0011] Figure 2Hardware configuration block diagram of the display device 200 according to one or more embodiments of the present application;

[0012] Figure 3 Hardware configuration block diagram of the control device 100 according to one or more embodiments of the present application;

[0013] Figure 4 Software configuration diagram of the display device 200 according to one or more embodiments of the present application;

[0014] Figure 5 Icon control interface display diagram of the application program in the display device 200 according to one or more embodiments of the present application;

[0015] Figure 6 Portrait tracking diagram according to one or more embodiments of the present application;

[0016] Figures 7-8 AI fitness scene diagram according to one or more embodiments of the present application;

[0017] Figure 9 Application program UI interface according to one or more embodiments of the present application;

[0018] Figure 10 Display device application program interface diagram according to one or more embodiments of the present application;

[0019] Figure 11 Display device dodge ball game user interface diagram according to one or more embodiments of the present application;

[0020] Figure 12 Display device dodge ball game user interface diagram according to one or more embodiments of the present application;

[0021] Figures 13-14 Display device recognizing user limb diagram according to one or more embodiments of the present application;

[0022] Figures 15A-15B Display device recognizing user channel switching action diagram according to one or more embodiments of the present application;

[0023] Figures 16-18 Display interface diagram of the display device according to one or more embodiments of the present application. DETAILED DESCRIPTION

[0024] In order to make the purposes, embodiments and advantages of the present application clearer, the following will combine the drawings in the exemplary embodiments of the present application to clearly and completely describe the exemplary embodiments of the present application. Obviously, the described exemplary embodiments are only a part of the embodiments of the present application, but not all the embodiments.

[0025] All other embodiments obtained by those of ordinary skill in the art based on the example embodiments described in the present application without creative effort are within the scope of the claims of the present application. In addition, although the disclosure in the present application is introduced according to one or more examples, it should be understood that each aspect of the disclosure can also constitute a complete embodiment by itself. It should be noted that the brief description of the terms in the present application is only for the convenience of understanding the embodiments described below, and is not intended to limit the embodiments of the present application. Unless otherwise specified, these terms should be understood according to their ordinary and general meanings.

[0026] Figure 1 For the operation scenario between the display device and the control device according to one or more embodiments of the present application, as shown in Figure 1 The user can operate the display device 200 through the mobile terminal 300 and the control device 100. The control device 100 can be a remote controller, and the communication between the remote controller and the display device includes infrared protocol communication, Bluetooth protocol communication, wireless or other wired ways to control the display device 200. The user can input user instructions through the keys on the remote controller, voice input, control panel input, etc. to control the display device 200. In some embodiments, mobile terminals, tablets, computers, laptops, and other smart devices can also be used to control the display device 200.

[0027] In some embodiments, the mobile terminal 300 can install a software application with the display device 200, implement connection communication through a network communication protocol, and achieve the purpose of one-to-one control operation and data communication. The audio and video content displayed on the mobile terminal 300 can also be transmitted to the display device 200 to achieve synchronous display function. The display device 200 also communicates data with the server 400 through various communication methods. The display device 200 can be connected through a local area network (LAN), a wireless local area network (WLAN), and other networks. The server 400 can provide various content and interaction to the display device 200. The display device 200 can be a liquid crystal display, an OLED display, or a projection display device. In addition to providing a broadcast receiving television function, the display device 200 can also provide a smart network television function with computer support function.

[0028] Figure 2 An example is shown in the configuration block diagram of the control device 100 according to the example embodiment. As shown in Figure 2As shown, the control device 100 includes a controller 110, a communication interface 130, a user input / output interface 140, a memory, and a power supply. The control device 100 can receive user input commands and convert them into commands that the display device 200 can recognize and respond to, acting as an intermediary for interaction between the user and the display device 200. The communication interface 130 is used for external communication and includes at least one of a Wi-Fi chip, a Bluetooth module, NFC, or a replacement module. The user input / output interface 140 includes at least one of a microphone, a touchpad, a sensor, buttons, or a replacement module.

[0029] Figure 3 A hardware configuration block diagram of a display device 200 according to an exemplary embodiment is shown. For example... Figure 3 The display device 200 shown includes at least one of the following: a tuner / demodulator 210, a communicator 220, a detector 230, an external device interface 240, a controller 250, a display 260, an audio output interface 270, a memory, a power supply, and a user interface 280. The controller includes a central processing unit, a video processor, an audio processor, a graphics processor, RAM, ROM, and a first to nth interface for input / output. The display 260 can be at least one of a liquid crystal display, an OLED display, a touch display, and a projection display, and can also be a projection device and a projection screen. The tuner / demodulator 210 receives broadcast television signals via wired or wireless reception and demodulates audio and video signals, such as EPG data signals, from multiple wireless or wired broadcast television signals. The detector 230 is used to collect signals from the external environment or signals interacting with the external environment. The controller 250 and the tuner / demodulator 210 can be located in different separate devices; that is, the tuner / demodulator 210 can also be located in an external device of the main device where the controller 250 is located, such as an external set-top box.

[0030] In some embodiments, the controller 250 controls the operation of the display device and responds to user operations through various software control programs stored in memory. The controller 250 controls the overall operation of the display device 200. The user can input user commands through a graphical user interface (GUI) displayed on the monitor 260, and the user input interface receives the user input commands through the GUI. Alternatively, the user can input user commands by inputting specific sounds or gestures, and the user input interface receives the user input commands by recognizing the sounds or gestures through sensors.

[0031] In some embodiments, a "user interface" is a medium interface between an application program or an operating system and a user for interaction and information exchange, which realizes the conversion between the internal form of information and the form that the user can accept. The commonly used form of user interface is a graphic user interface (GUI), which refers to a user interface related to computer operation displayed in a graphical manner. It can be an icon, window, control, etc. interface element displayed in the display screen of an electronic device, wherein the control can include at least one of the following visual interface elements: icon, button, menu, tab, text box, dialog box, status bar, navigation bar, Widget, etc.

[0032] Figure 4 For the software configuration diagram of the display device 200 according to one or more embodiments of the present application, as shown in Figure 4 the system is divided into four layers from top to bottom, namely, an application layer (referred to as "application layer" for short), an application framework layer (referred to as "framework layer" for short), an Android runtime and system library layer (referred to as "system runtime library layer" for short), and a kernel layer. The kernel layer at least includes at least one of the following drivers: audio driver, display driver, Bluetooth driver, camera driver, WIFI driver, USB driver, HDMI driver, sensor driver (such as fingerprint sensor, temperature sensor, pressure sensor, etc.), and power supply driver, etc.

[0033] Figure 5 For the icon control interface display diagram of the application program in the display device 200 according to one or more embodiments of the present application, as shown in Figure 5 the application layer includes at least one application program, which can display corresponding icon controls in the display, such as: live TV application program icon control, video on demand application program icon control, media center application program icon control, application center icon control, game application icon control, etc. The live TV application program can provide live TV through different signal sources. The video on demand application program can provide videos from different storage sources. Unlike the live TV application program, the video on demand provides video display from certain storage sources. The media center application program can provide various multimedia content playing application programs. The application center can provide storage of various application programs.

[0034] <Body detection technology applied to portrait tracking>

[0035] The content of the application filed on August 21, 2020, application number: 202010849806.3, name "a portrait positioning method of display device"; and the application filed on December 31, 2020, application number: 202011620179.2, name "a display device and a portrait positioning method" of the present application are hereby incorporated into the present application.

[0036] In the embodiments of the present application, Figure 6 As shown in the portrait tracking schematic diagram according to one or more embodiments of the present application, Figure 6 The camera 231 can be built-in or externally connected to the display device 200 as a kind of detector 230, and can detect image data after starting to run. The camera 231 can be connected to the controller 250 through the interface component, so as to send the detected image data to the controller 250 for processing. In order to detect the image, the camera 231 can include a lens assembly and a holder assembly, and the lens assembly is arranged on the holder assembly. The holder assembly can drive the lens assembly to rotate in order to change the orientation of the lens assembly. With different orientations of the lens assembly, the lens assembly can take video of the user located at different positions, so as to obtain the user image data. Obviously, different orientations correspond to image acquisition of different areas. When the user is located to the left of the front position relative to the display 275, the first rotating shaft on the holder assembly can drive the fixed part and the lens assembly to rotate to the left, so that the user portrait position in the captured image is located in the center area of the picture. When the user's body image position is below, the second rotating shaft in the holder assembly can drive the lens assembly to rotate upward, so as to raise the shooting angle, and the user portrait position is located in the center area of the picture.

[0037] Based on the above-mentioned camera 231, an automatic control program can be set in the display device 200, so as to adjust the orientation of the lens assembly in the camera 231 by detecting the position of the user, and repeat the detection process of the portrait position at a certain frequency, so as to realize the tracking of the portrait position, which specifically includes the following steps:

[0038] Detecting the portrait position. After the camera 231 is started, multiple frames of images are captured in real time, and the captured images are sent to the controller 250 of the display device 200. After the camera 231 is started, the controller 250 can perform image processing according to the started application program on one hand, for example, controlling the display 275 to display the image; on the other hand, the controller 250 can analyze the image by calling the detection program, so as to determine the position of the user. The detection of the portrait position can be completed by the image processing program. That is, the image captured by the camera 231 is captured in real time, and the body information is detected. The body information can include key points and a body frame. The position of the key points and the body frame in the image is determined by the position information of the detected key points and the body frame, so as to determine the portrait position. The key points can be a series of points in the human body image that can represent the characteristics of the human body. For example, eyes, ears, noses, necks, shoulders, elbows, wrists, waists, knee joints, and ankle joints.

[0039] The determination of the key points can be obtained by image recognition, that is, by analyzing the characteristic shape in the picture and matching with the preset template, the image corresponding to the key points is determined, and the position corresponding to the image is obtained, so as to obtain the position corresponding to each key point. The position can be represented by the number of pixel points in the image. According to the resolution and the visual angle of the camera 231, a plane rectangular coordinate system can be constructed with the upper left corner of the image as the origin and the right and downward directions as the positive directions. Each pixel point in the image can be represented by the rectangular coordinate system.

[0040] When the position of the user changes or the posture changes, the positions of some key points will change. With the change, the relative position of the human body in the image captured by the camera 231 will also change. For example, when the human body moves to the left, the position of the human body in the image captured by the camera 231 will be to the left, which is not convenient for image analysis and real-time display.

[0041] Therefore, after detecting the portrait position, the portrait position is compared with the preset area in the image, so as to determine whether the current portrait position is in the preset area. In some embodiments, the portrait position can be represented by the center position of the body frame, and the center position of the body frame can be calculated by the coordinates of the detected key points. For example, the x-axis coordinate of the center position x0=(x1+x2) / 2 is calculated by obtaining the x-axis coordinates of the key points on the left and right sides of the body frame.

[0042] Figures 7-8For the AI fitness scene according to one or more embodiments of the present application, but in actual application, due to the influence of user posture, and the different needs in different application programs, the method of using the center position as the portrait position judgment cannot obtain better display, detection, tracking effect in some application scenarios. In order to obtain more accurate portrait position judgment, in some embodiments, taking the AI fitness scene as an example, as shown in Figure 7 , Figure 8 After identifying the plurality of key points, a skeleton line diagram can be established according to the identified key points, so as to further determine the position of the portrait according to the skeleton line diagram. The skeleton line can be determined by connecting a plurality of key points. In different postures of the user, the shape of the skeleton line is also different.

[0043] It should be noted that the skeleton line drawn can also dynamically adjust the shooting position of the camera according to the motion change rule of the skeleton line. For example, if it is judged that the motion state change process of the skeleton line is from a squatting posture to a standing posture, the angle of view of the camera can be raised so that the portrait in the standing posture can also be in the appropriate area in the image, that is, from Figure 7 to Figure 8 The effect shown in the transition. If it is judged that the motion state change process of the skeleton line is from a standing posture to a squatting posture, the angle of view of the camera can be lowered so that the portrait in the squatting posture can also be in the appropriate area in the image, that is, from Figure 8 to Figure 7 The effect shown in the transition.

[0044] <Application of limb detection technology to multi-person tracking>

[0045] The contents of the present application of the present application, the inventor on August 21, 2020, application number: 202010847711.8, name "a face feature value creation method, character locking tracking method and display device"; and the present application of the present application, the inventor on February 4, 2021, application number: 202110155019.3, name "a face feature value creation method, character locking tracking method and display device" are added to the present application.

[0046] In some embodiments, taking the AI fitness scenario as an example, in the fitness follow-up mode, a certain person needs to be tracked and the action needs to be marked to generate follow-up data and count the follow-up result. If in the mobile phone scenario, the face or the body is close to the screen, the face or the body occupies a large proportion in the screen, and the relative moving distance of the detected person image in each frame image is small, generally no loss (out of screen) phenomenon will occur. However, the display device is different from the mobile phone scenario. When a person interacts with the display device, the distance between the person and the display device is generally far, the screen occupation of the face or the body is small, and the moving distance in the front and back frames of images will be large. For example, the person quickly moves in front of the screen, which will easily cause the recognition loss of the person, that is, the out-of-screen phenomenon. Since the current fitness function is mainly based on the body information for body following, the camera realizes the portrait following function by recognizing the face or the body as the recognition basis. No matter it is the body or the face information, an ID will be generated each time the person is recognized as the identification of the body or the face. However, when the recognition is lost and then re-recognized, that is, the person is out of the screen and then enters the screen again, a new ID information will be generated, causing the inconsistency of the IDs of the same person before and after. Therefore, when the fitness or the camera tracks a specific person, the loss will cause irreversibility, and it is impossible to effectively track the same person.

[0047] The display device provided by the embodiment of the application is configured to perform the following steps:

[0048] S11, acquiring person image information collected by a camera.

[0049] Since the person who can interact with the display device at the same time can be one or more, when at least one person interacts with the display device, for example, when the at least one person uses the display device for video call, AI fitness or camera portrait tracking, etc., the camera performs real-time image collection. The person image information collected by the camera includes image information of the at least one person, from which the body action and the facial feature information of the person can be read.

[0050] S12, recognizing the person image information, determining a locked tracking person, and creating facial feature information and specific body ID information of the locked tracking person.

[0051] When multiple persons interact with the display device, or when only one person initially interacts with the display device but other persons appear in the camera shooting area during the interaction, the display device cannot determine which instruction of which person as the control instruction for response, that is, cannot determine which person as the specific person for tracking, in the multi-person interaction scenario. Therefore, it is necessary to determine the locked tracking person when interacting. The locked tracking person is one of the persons interacting with the display device, and only the instruction generated by the locked tracking person is responded to in subsequent interaction. To realize the locked tracking of the same person, one of the multiple persons interacting with the display device is selected as the locked tracking person. If only one person interacts with the display device, the locked tracking person is the person. When determining the locked tracking person, whether each person makes a specific action can be used for judgment, and the action recognition of the person can be determined according to the body key point information of the person.

[0052] S13, establishing an association relationship between the specific body ID information and the face feature information of the locked tracking person, to obtain the face feature value of the locked tracking person.

[0053] Since each person has his / her own face feature information, and the face feature information of different persons is different. Therefore, each person can be identified by the face feature information, and if the same or similar face feature information is identified, it can be identified as the same person.

[0054] Since in the normal case, when the locked tracking person walks out of the picture, the corresponding specific body ID information is lost, that is, the corresponding specific body ID information of the person is deleted after the person walks out of the picture. If the person walks into the picture again, the corresponding body ID information is generated again, and it is easy to identify the same person as two persons.

[0055] Therefore, in some embodiments, the specific body ID information and the face feature information of the locked tracking person are associated, the specific body ID information and the face feature information associated with each other are used as the face feature value of the locked tracking person, to identify the locked tracking person, and the face feature information will not be deleted because the person walks out of the picture, and will be saved in the controller.

[0056] In some embodiments, in the AI fitness scenario, the controller performs the locked tracking of the locked tracking person based on the face feature value, and is further configured to perform the following steps:

[0057] Step 1311, when the camera application is an AI fitness application, determining that the locked tracking person is a fitness person.

[0058] Step 1312, based on the face feature value of the fitness person, continuously collecting follow-up action information of the fitness person based on the demonstration video presented in the user interface.

[0059] Step 1313, generating a follow-up picture based on the follow-up action information, and displaying the follow-up picture in the user interface. The follow-up picture is displayed on one side of the picture where the demonstration video is located.

[0060] The above embodiment is implemented in the AI fitness scene. At the same time, the character of the AI fitness function configured by the display device can be one or more. If the character is one, the character of the fitness is the locked tracking character. The user interface of the display presents the demonstration video, which facilitates the follow-up of the fitness person. At this time, the camera is applied to the AI fitness application. The AI fitness application calls the camera to always collect the follow-up action of the locked tracking character, and displays it in the user interface of the display.

[0061] <Body detection technology applied to dodgeball game>

[0062] The display device can also apply human body posture detection to game applications. In the following, the display device based on body recognition to predict motion trends dodgeball display technical solution and user interface will be taken as an example to describe the display device and the dodgeball display method based on body recognition.

[0063] In some embodiments, Figure 9 For the application UI interface according to one or more embodiments of the application, for example, the application UI interface includes 4 applications installed on the television, which are news headlines, cinema on demand, AR dodgeball, karaoke, etc. By using a remote control, voice, etc. Method to move the focus on the display screen, you can select different application programs or other function buttons.

[0064] In some embodiments, the television display screen is configured to display other interactive elements while displaying the application UI interface. The interactive elements can include, for example, television home control, search control, message button control, mailbox control, browser control, favorites control, signal bar control, etc.

[0065] To improve the convenience and image of the television UI, in some embodiments, the controller of the display device in the embodiment of the application controls the UI of the television in response to the operation of the interactive element. For example, the user clicks the search control through the controller such as a remote control, which can display the search UI on top of other UIs. That is, the UI of the application component mapped by the control interactive element can be enlarged or run and displayed full screen.

[0066] In some embodiments, the interactive element can also be operated by a sensor, which can be, but is not limited to, an acoustic input sensor, such as a microphone, which can detect a voice command including a desired interactive element indication. For example, a user can use a "dodgeball" or any other suitable identifier to identify a desired interactive element, such as a search control, and can also describe a desired action to be performed in relation to the desired interactive element. The controller can identify the voice command and submit data characterizing the interaction to the UI or its processing components or engines.

[0067] In some embodiments, Figure 10 A schematic diagram of a display device application interface according to one or more embodiments of the present application. Referring to Figure 10 , a user can control the focus of the display screen through a remote control, select an AR dodgeball application so that its icon is displayed in high brightness in the user interface of the display screen, and then click the high brightness icon to open the application mapped by the icon.

[0068] In some embodiments, Figure 11 A schematic diagram of a display device dodgeball game user interface according to one or more embodiments of the present application. Referring to Figure 11 , the display device provided by the present application includes a camera, a display, and a controller. The camera is configured to capture a target user and a background image in a dodgeball game in a detection area. The display is configured to display a dodgeball game user interface including the target user and the background image. The controller is configured to control the user interface to display a first identifier at a first position in the user interface according to a first size of the target user in the background image and a first position of the target user in the user interface, and the first identifier is used to trigger a dodgeball launch, and the size of the first identifier corresponds to the first size, as shown in Figure 11 .

[0069] In some embodiments, when the upper torso of the target user is not completely located in the capture area and the target user is leaving the capture area, the controller controls the user interface to no longer display the first identifier.

[0070] In some embodiments, the camera captures a target user image and a background image of the target user in its monitoring area, and the target user and the background image are both displayed in a dodgeball game user interface. The controller controls the user interface to display a first identifier at a first position in the user interface according to a first size of the target user in the background image and a first position of the target user in the game user interface.

[0071] The controller determines the size of the first mark according to the size of the target user in the background image, i.e. the user who is playing dodgeball game, i.e. the larger the size of the user in the background image, the larger the first mark; the smaller the size of the user in the background image, the smaller the first mark.

[0072] When the user moves from the first position to the second position displayed in the user interface, the controller will determine the size of the second mark according to the second size of the user at the second position, and display the second mark in the game user interface.

[0073] In some embodiments, the controller controls the user interface to display the first mark in the first position, and the size of the first mark corresponds to the first size, specifically including: when the target user moves from far to near relative to the camera, the first size gradually increases, and the size of the first mark corresponds to the increase of the first mark; when the target user moves from near to far relative to the camera, the first size gradually decreases, and the size of the first mark corresponds to the decrease of the first mark.

[0074] For example, when the user moves from the first position to the second position, the controller can control the second mark to be enlarged or reduced compared with the first mark according to the change of the first size and the second size of the target user in the background image.

[0075] In some embodiments, the controller controls the first mark to be displayed on the target user in the first position. For example, the first mark can be displayed as an approximate rectangular frame, and when the controller locates the first position of the target user in the game user interface, an approximate rectangular frame corresponding to the first size of the target user at that time, i.e. the first mark, is displayed on the target user.

[0076] In some embodiments, Figure 12 A schematic diagram of a display device for a dodgeball game user interface according to one or more embodiments of the present application. Reference is made to Figure 12 When the upper body of the target user is not completely located in the collection area, and the target user is leaving the collection area, the controller controls the user interface to stop displaying the first mark.

[0077] It can be understood that the target user leaving the game includes: the target user walking towards the camera capture area but part of the upper body of the user is in the capture area, the target user is crossing the edge of the camera capture area and part of the upper body of the user is in the capture area, and the target user has crossed the edge of the camera capture area and the upper body of the user is completely out of the capture area. The display device provided in the present application can identify that the target user is crossing the edge of the camera capture area and part of the upper body of the user is in the capture area, and control the user interface to no longer be in the first position and trigger the game to continue to launch the dodgeball. As shown in the dashed approximate rectangular frame in the figure, the dashed approximate rectangular frame is used for convenience of understanding and is not displayed in the game user interface of the display device, and the dashed approximate rectangular frame represents part of the upper body of the target user recognized by the controller.

[0078] In some embodiments, Figures 13-14 A schematic diagram of a display device for recognizing a user's limbs according to one or more embodiments of the present application. Referring to Figure 13 The controller controls the user interface to display a first mark in the first position, and the size of the first mark corresponds to the first size, specifically including that the controller: identifies the left elbow of the target user in the camera captured image as a first positioning point, and identifies the right elbow as a second positioning point; determines the size of the first mark displayed in the first position according to a first interval distance between the first positioning point and the second positioning point.

[0079] The controller decomposes the upper body of the target user captured by the camera into points according to the image recognition model, judges the motion trend of the user according to the position changes of the points of the upper limbs, and infers the overall position from part of the limbs, so as to change the position and size of the first mark, the position being the first position provided in the present application, and the size being the size of the first mark provided in the present application.

[0080] For example, the controller identifies the left elbow of the target user as a first positioning point and the right elbow as a second positioning point, and then controls the user interface to display a first mark with a width of 10 cm at the first position according to a first interval distance between the first positioning point and the second positioning point, for example, the first interval distance is 10 cm. The first mark can be implemented as an approximate rectangular frame.

[0081] In some embodiments, the controller identifies the target user in the image captured by the camera by the following steps. First, a single frame image is extracted from the captured video every preset number of frames; human body detection is performed on the single frame image to determine whether the single frame image contains a human body; if the single frame image contains a human body, face detection is performed on the single frame image within the human body frame of the human body to determine whether the human body frame contains a face; if the human body frame contains a face, feature extraction is performed on the face to obtain the face feature of the user in the captured video; the face feature of the user in the captured video is compared with a preset family member face feature library to determine whether the user in the captured video is a family member and a game user.

[0082] In some embodiments, the controller applies a face detection algorithm (for example, a deformable part model) to the human body in the image captured by the camera, performs face detection on the single frame image within the detected human body frame to determine whether the human body frame contains a face. If the human body frame contains a face, feature extraction is performed on the face to obtain the face feature of the user. The face feature of the user is compared with a family member face feature library to determine whether the user is a family member and a game user.

[0083] In some embodiments, with reference to Figure 14 , the controller is further configured to identify the left elbow of the target user in the image captured by the camera as a first positioning point, the right elbow as a second positioning point, and other joints as third positioning points, the third positioning points being located at the left hand, and / or the left shoulder, and / or the neck, and / or the right shoulder, and / or the right hand, and / or the left waist, and / or the right waist; according to the position changes of the first positioning point, the second positioning point, and the third positioning points, it is determined whether the upper body of the target user is completely within the capture area and whether the target user is leaving the capture area.

[0084] For example, the left hand of the user is identified as point 1, the left elbow is identified as point 2, the left shoulder is identified as point 3, the neck is identified as point 4, the right shoulder is identified as point 5, the right elbow is identified as point 6, the right hand is identified as point 7, the left waist is identified as point 8, and the right waist is identified as point 9.

[0085] The first positioning point is implemented as point 2, the second positioning point is implemented as point 6, and the third positioning point can be implemented as one or a combination of multiple remaining points. The controller identifies the above-mentioned multiple points in the collected image, tracks the multiple points, and calculates the overall position of the user according to the multiple points, for example, according to the nine points, to obtain a first position. The interval distance between point 2 and point 6 determines the size of the first identification display frame. The position change of the multiple points determines the movement direction of the target user, whether the upper torso of the target user is completely in the collection area, and whether the target user is leaving the collection area.

[0086] In some embodiments, the controller extracts a single frame image from the collected video every preset frame number (for example, every 90 frames) from the collected image received from the camera. A human body detection algorithm in image recognition technology (for example, a Convolutional Pose Machine that detects each joint of a human body and determines the range formed by these joints as a human body frame of the human body) is applied to perform human body detection on the single frame image. The position change of the multiple points determines the movement direction of the target user, whether the upper torso of the target user is completely in the collection area, and whether the target user is leaving the collection area.

[0087] In some embodiments, the body features of each family member and the corresponding family member identification are stored in the family member feature library. The body features of the family member store the coordinates of each body feature point of the family member. For example, the coordinates of the right hand of family member 03 are (10, 0), the coordinates of the left hand are (-10, 0), the coordinates of the right shoulder are (5, -10), and the coordinates of the left shoulder are (-5, -10).

[0088] In some embodiments, according to a predetermined connection rule (for example, connecting the left shoulder and the right shoulder, then connecting the right shoulder and the right arm elbow, and then connecting the right arm elbow and the right hand), the upper torso feature points of the user in the collected video are sequentially connected to obtain the upper torso edge graph of the user. Similarly, the upper torso edge graph of each family member is obtained.

[0089] For a family member, the similarity of each line in the upper torso edge graph of the user to the corresponding line in the upper torso edge graph of the family member is determined by the angle of the line. For example, for family member A, the angle of the line formed by the left hand, the left elbow, and the left shoulder in the upper torso edge graph of the user is 27 degrees. The angle of the line formed by the left hand, the left elbow, and the left shoulder in the upper torso edge graph of the target user is 30 degrees. Therefore, the similarity of the line in the upper torso edge graph of the target user to the line in the upper torso edge graph of family member A is 1-(|27-30| / 30) = 0.9.

[0090] After the similarity of each line of the upper body edge graph of the target user and the corresponding line of the upper body edge graph of the family member is determined, the average of the similarity of each line of the upper body edge graph of the target user and the corresponding line of the upper body edge graph of the family member is determined as the similarity of the upper body feature of the target user and the upper body feature of the family member.

[0091] Based on the above description of the display device based on the limb recognition to predict the motion direction of the dodgeball display scheme, the application also provides a dodgeball display method based on limb recognition, which comprises: according to the first size of the target user in the dodgeball game user interface background image and the first position of the target user in the user interface, controlling the user interface to display a first identifier at the first position, wherein the first identifier is used to trigger the dodgeball launch, and the size of the first identifier corresponds to the first size; when the upper body trunk of the target user is not completely located in the collection area of the camera, and the target user is leaving the collection area, controlling the user interface to no longer display the first identifier. The specific operation and steps of the dodgeball display method based on limb recognition have been described in detail in the above display device implementation scheme, and will not be repeated here.

[0092] In some embodiments, controlling the user interface to display a first identifier at the first position, and the size of the first identifier corresponds to the first size, specifically includes: when the target user moves from far to near relative to the camera, the first size gradually increases, and the size of the first identifier corresponds to the increase of the first identifier; when the target user moves from near to far relative to the camera, the first size gradually decreases, and the size of the first identifier corresponds to the decrease of the first identifier. The specific operation and steps of the dodgeball display method based on limb recognition have been described in detail in the above display device implementation scheme, and will not be repeated here.

[0093] In some embodiments, controlling the user interface to display a first identifier at the first position, and the size of the first identifier corresponds to the first size, specifically includes: identifying the left arm elbow of the target user in the collection image as a first positioning point, and identifying the right arm elbow as a second positioning point; according to the first interval distance between the first positioning point and the second positioning point, determining the size of the first identifier displayed at the first position. The specific operation and steps of the dodgeball display method based on limb recognition have been described in detail in the above display device implementation scheme, and will not be repeated here.

[0094] In some embodiments, the method further comprises: identifying the left arm elbow of the target user in the captured image as a first positioning point, the right arm elbow as a second positioning point, and other joints as third positioning points, the third positioning points being located at the left hand, and / or the left shoulder, and / or the neck, and / or the right shoulder, and / or the right hand, and / or the left waist, and / or the right waist; determining whether the upper body of the target user is completely in the capture area and whether the target user is leaving the capture area according to tracking the position changes of the first positioning point, the second positioning point, and the third positioning points. The specific operations and steps of the body recognition-based dodgeball display method have been described in detail above in the display device implementation scheme, and will not be repeated here.

[0095] In some embodiments, the control of the user interface to display the first identifier at the first position specifically comprises the controller: displaying the first identifier at the first position of the target user. The specific operations and steps of the body recognition-based dodgeball display method have been described in detail above in the display device implementation scheme, and will not be repeated here.

[0096] The embodiments of the present application have the beneficial effects that by constructing the first size, the size of the first identifier can be controlled; by further constructing the first position, the position of the first identifier can be obtained; by further controlling the first identifier not to be displayed, the dodgeball can not be emitted when the user leaves the game; by constructing the first positioning point, the second positioning point, and the third positioning point, the game user's body can be recognized, the game user's motion trend can be predicted, the size of the first identifier can be adjusted according to the size of the user displayed on the screen, and the user can not play the game when the user is not in the screen display range.

[0097] <Body detection technology applied to channel switching>

[0098] In some embodiments, the controller determines that the user is switching channels based on the captured image obtained by the camera, specifically comprising:

[0099] The controller analyzes the captured image obtained by the camera at a first frequency to identify the presence of the user. Specifically, after the display device is turned on, the controller controls the camera to capture images and obtains image previews thereof, and uses image recognition algorithms on the preview images to detect whether a user is watching the display device. Before the controller detects the user, the controller analyzes the captured image at a first frequency, for example, the controller analyzes the image captured by the camera at an interval of every 10 frames to detect whether a user is present in the detection range of the camera. Through the setting of the first frequency, the computing resources of the controller can be effectively saved.

[0100] After recognizing the user, the controller analyzes the captured image at a second frequency to identify the user's upper limb movement, and when the user's elbow position and hand position change and the height difference between the hand position and the elbow position is less than a height threshold, the controller recognizes that the user is performing a channel changing operation, and the second frequency is greater than the first frequency. Figures 15A-15B For the display device to recognize the user's channel changing operation according to one or more embodiments of the present application, in some embodiments, after the camera captures the image of the user, the controller sends the captured data frame to the image recognition model for analysis to determine whether the user is performing a channel changing operation. The image recognition model identifies the user's upper limbs, i.e., arms, torso, as lines and nodes, for example, the elbow as point 1 and the hand as point 2, and their relative positions before the user performs a channel changing operation are as shown in Figure 15A In some embodiments, when the user uses the remote controller to change the channel or switch the first piece source to the second piece source, the user usually has an action of raising the arm and pointing at the display device; the controller analyzes the data frame captured by the camera and uses the image recognition model to identify the position of point 1 representing the elbow and point 2 representing the hand, and when the positions of point 1 and point 2 change, the position of point 2 rises, and the positions of point 1 and point 2 are close to parallel, i.e., the height difference between the position of point 2 representing the hand and the position of point 1 representing the elbow is less than a height threshold, for example, the height threshold is implemented as 5 cm, the controller determines that the user's action is very likely to be the user picking up the remote controller to switch the first piece source to the second piece source, i.e., the user performs a channel changing operation.

[0101] At this time, the controller sends a first instruction containing a screenshot of the second piece source to the server, and the first instruction is used to make the server identify the second piece source according to the screenshot to determine whether the optimized image parameters and the optimized sound parameters can be provided; when the server can provide the optimized image parameters and the optimized sound parameters, the controller controls the user interface to play the second piece source configured with the optimized image parameters and the optimized sound parameters, as shown in Figure 15B

[0102] In some embodiments, the display device identifies the user's body performing a channel changing operation as follows:

[0103] ​In step 901, the display device detects the user to speed up the preview analysis frequency. The display device controller determines that the user is channel switching based on the acquisition image obtained by the camera, specifically including that the controller determines whether the display device audio output is momentarily silent within a first time length after identifying the user's channel switching action based on the acquisition image; if so, it is determined that the user is channel switching; otherwise, it is determined that the user is not channel switching. For example, before the controller detects the user after the display device is turned on, the controller analyzes the acquisition image at a first frequency to identify whether a user is watching the display device. For example, the controller analyzes the image at an interval of every 10 frames to detect whether a user exists in the detection range of the camera. Through the setting of the first frequency, the computing resources of the controller can be effectively saved.

[0104] In step 902, the display device performs user behavior analysis. After identifying the user, the controller analyzes the acquisition image at a second frequency to perform user behavior analysis, for example, which can be implemented to identify the user's upper limb movement.

[0105] In step 903, a mark is made when the user's hand-raising action is detected. By analyzing the acquisition image, the controller identifies the change in the user's elbow and hand position to make a mark. For example, when the height difference between the hand position and the elbow position is less than a height threshold, the controller determines that the user has performed a suspected channel switching operation.

[0106] In step 904, the display device determines whether the sound change is received and the interval time is less than 30 seconds. The controller determines whether the display device audio output is momentarily silent within a first time length, i.e., 30 seconds, after identifying that the user has performed a suspected channel switching operation. The momentary silence usually occurs when the display device is channel switching or switching the video source. If the controller determines that the display device audio output is momentarily silent within the first time length, i.e., 30 seconds, it is determined that the user switches the first video source to the second video source. At this time, the controller sends a first instruction containing a screenshot of the second video source to the server, which is used to identify the second video source based on the screenshot to determine whether the optimized image parameters and the optimized sound parameters can be provided. When the server can provide the optimized image parameters and the optimized sound parameters, the controller controls the user interface to play the second video source configured with the optimized image parameters and the optimized sound parameters. It should be noted that the current playing video can also be identified based on the acquired screenshot locally on the display device.

[0107] If the controller determines that the display device audio output is not momentarily silent within the first time length, i.e., 30 seconds, it is determined that the user does not switch the second video source, and the analyzed data is discarded, as shown in step 904-1.

[0108] In some embodiments, the control configures the optimized image parameters and the optimized sound parameters for the current video played by the user interface according to the screenshot, specifically including that the controller: sends a first instruction containing the screenshot to the server, the first instruction being used to make the server identify the current video according to the screenshot to determine the optimized image parameters and the optimized sound parameters that can be provided; receives a second instruction containing the optimized image parameters and the optimized sound parameters sent from the server; configures the optimized image parameters and the optimized sound parameters for the current video played by the user interface according to the second instruction. For example, when the controller of the display device identifies and determines that the user picks up the remote controller to switch the first piece of source played by the user interface to the second piece of source, the display device takes a screenshot of the second piece of source played by the current user interface; then the controller sends the screenshot to the server through the first instruction.

[0109] In some embodiments, the screenshot is a screenshot of the display device logo at the corner position of the user interface and / or a screenshot of the text at the edge position of the user interface. The screenshot is a screenshot of the display device logo at the corner position of the user interface and / or a screenshot of the text at the edge position of the user interface. For example, when the user is watching a live channel, the display device logo is usually located at the upper left corner of the display device screen, and the controller can improve the efficiency of image recognition and reduce the amount of data transmission by taking a screenshot of the display device logo for the server to identify. When the user is watching a film and television variety show, the name of the show is usually located at the upper edge, lower edge or side edge of the display device screen, and the controller can improve the efficiency of image recognition and reduce the amount of data transmission by taking a screenshot of the text information at the edge position of the screen for the server to identify, as shown in step 904-2.

[0110] Based on the above description of the display device automatic configuration video parameter scheme, the application further provides a display device end automatic configuration video parameter method, which comprises: determining that the user is changing the channel based on the acquired capture image, acquiring a screenshot of the user interface; configuring the optimized image parameters and the optimized sound parameters for the current video played by the user interface according to the screenshot. The specific operations and steps of the automatic configuration video parameter method have been described in detail in the above display device implementation scheme, and will not be repeated here.

[0111] In some embodiments, determining that the user is changing the channel based on the acquired capture image specifically includes: after identifying the user's channel changing action based on the capture image, determining whether the audio output is momentarily muted within a first time length; if yes, determining that the user is changing the channel; otherwise, determining that the user is not changing the channel. The specific operations and steps of the automatic configuration video parameter method have been described in detail in the above display device implementation scheme, and will not be repeated here.

[0112] In some embodiments, determining whether the user is changing channel based on the captured image specifically includes: analyzing the captured image at a first frequency to identify the presence of the user; and analyzing the captured image at a second frequency to identify the movement of the upper limbs of the user after identifying the presence of the user, wherein the second frequency is greater than the first frequency, and wherein the movement of the upper limbs of the user is identified when the position of the elbow and the position of the hand of the user change and the height difference between the position of the elbow and the position of the hand is less than a height threshold. The specific operations and steps of the method for automatically configuring video parameters are described in detail above in the implementation scheme of the display device, and thus will not be described here again.

[0113] In some embodiments, configuring the optimized image parameters and the optimized sound parameters for the current video played by the user interface according to the screenshot specifically includes: sending a first instruction containing the screenshot to the server, wherein the first instruction is used to make the server identify the current video according to the screenshot to determine the optimized image parameters and the optimized sound parameters that can be provided; receiving a second instruction containing the optimized image parameters and the optimized sound parameters sent by the server; and configuring the optimized image parameters and the optimized sound parameters for the current video played by the user interface according to the second instruction. The specific operations and steps of the method for automatically configuring video parameters are described in detail above in the implementation scheme of the display device, and thus will not be described here again.

[0114] In some embodiments, the screenshot is a display device logo screenshot of the current video at the corner position of the user interface and / or a text screenshot of the current video at the edge position of the user interface. The specific operations and steps of the method for automatically configuring video parameters are described in detail above in the implementation scheme of the display device, and thus will not be described here again.

[0115] Based on the above description of the implementation scheme of the server for automatically configuring video parameters, the present application further provides a method for automatically configuring video parameters on the server side, which specifically includes: receiving a first instruction containing a second piece of source screenshot sent by a display device; identifying the name of the second piece of source based on the screenshot contained in the first instruction to determine whether the optimized image parameters and the optimized sound parameters corresponding to the second piece of source can be provided; and sending a second instruction containing the optimized image parameters and the optimized sound parameters to the display device when the optimized image parameters and the optimized sound parameters can be provided. The specific operations and steps of the implementation scheme of the server for automatically configuring video parameters are described in detail above, and thus will not be described here again.

[0116] The beneficial effects of the embodiments of the present application are that the user image can be collected by the camera to realize instant detection of the user's limb action of changing the channel; further, the current playing video can be recognized by acquiring the user interface screenshot; further, the optimized sound image parameters are configured for the recognized video, and when the display device uses the set-top box as a display through HDMI connection, the accuracy of recognizing the user's channel changing, recognizing the program information, and automatically configuring the playing parameters after the channel changing can be improved.

[0117] <Body detection precision improvement algorithm>

[0118] In some embodiments, taking the AI fitness scene as an example, some embodiments can improve the detection precision of human posture. Figures 16-18 The display interface of the display device according to one or more embodiments of the present application is schematically shown in Figure 16 The left area is a coach action display image, and the right area is a user body image collected by the camera in real time. When the user appears in the screen, the area where the user is first detected is the area enclosed by the solid line, and then the user's action is detected. By comparing whether the user's action is consistent with the coach's action, the accuracy of the user is obtained and displayed on the current display interface.

[0119] In human posture detection, any slight difference in the previous and subsequent two frames of pictures, such as light, frame rate, background, etc., will cause differences in the detection results, and the human detection box will exist in the visual display process. The jitter will affect the user experience.

[0120] Considering the above technical problems, the display device provided in the embodiments of the present application can prevent the human detection box in the human posture detection result from appearing jitter due to the slight difference between the previous and subsequent two adjacent frames of pictures, thereby improving the visual display effect. The following detailed embodiments are described in detail.

[0121] In some embodiments of the present application, the display device 200 can collect the motion video of the user through the camera 201, and send the collected motion video as a to-be-detected video to the controller 250 for processing.

[0122] The controller 250 determines the information related to the person in each frame of image in the to-be-detected video by using the human posture detection network model after receiving the to-be-detected video, which includes the position information of the person, including the center point position (xc, yc), width (wc), and height (hc) of the human detection box; and the human posture information, including the human joint position information (xloc, yloc).

[0123] The center point position (xc, yc) of the human body detection frame can be used to determine the position of the human body detection frame, and the width (wc) and height (hc) of the human body detection frame can be used to determine the size of the human body detection frame.

[0124] The controller 250, after obtaining the human body posture detection result of each frame of image, determines whether there is a change in the position of the target human body detection frame in the current frame of image compared with the target human body detection frame in the previous frame of image based on the center point position of the target human body detection frame in the current frame of image and the center point position of the target human body detection frame in the previous frame of image, and / or determines whether there is a change in the size of the target human body detection frame in the current frame of image compared with the target human body detection frame in the previous frame of image based on the width and height of the target human body detection frame in the current frame of image and the width and height of the target human body detection frame in the previous frame of image.

[0125] If the size and / or position of the target human body detection frame in the current frame of image changes compared with the target human body detection frame in the previous frame of image, it is determined whether the change amplitude of the change is less than a preset change threshold.

[0126] If the change amplitude of the change is less than the preset change threshold, it can be considered that the change is not caused by the posture change of the user, but caused by some small differences between the two adjacent frames of images. In this case, the size and / or position of the target human body detection frame in the current frame of image can be adjusted based on the target human body detection frame in the previous frame of image, so that the size and / or position of the target human body detection frame in the current frame of image is the same as that of the target human body detection frame in the previous frame of image, and then the adjusted current frame of image is sent to the display for display. Since the size and / or position of the target human body detection frame in the adjusted current frame of image is the same as that of the target human body detection frame in the previous frame of image, when the current frame of image is visually displayed, the human body detection frame in the human body posture detection result can be effectively prevented from shaking due to the small differences between the two adjacent frames of images, and the visual display effect is improved.

[0127] If the change amplitude of the change is greater than or equal to the preset change threshold, it can be considered that the change is caused by the posture change of the user, and in this case, the current frame of image can be directly sent to the display for display. At this time, the user can observe that the human body detection frame displayed on the display screen changes with the posture change of the user, thereby ensuring the accuracy of the visual display of the human body posture detection result.

[0128] The display device provided in the embodiments of the present application can determine whether the size and / or position of the target human body bounding box in the current frame image changes by comparing the size and / or position of the target human body bounding box in the current frame image with the size and / or position of the target human body bounding box in the previous frame image. When the size and / or position of the target human body bounding box in the current frame image changes and the change amplitude is less than a preset change threshold, the size and / or position of the target human body bounding box in the current frame image is adjusted based on the target human body bounding box in the previous frame image, and the adjusted current frame image is sent to the display for display. Thus, the human body bounding box in the human body posture detection result can be prevented from shaking due to the slight difference between the two adjacent frame images, and the visual display effect is improved.

[0129] Based on the content described in the above embodiments, in some embodiments of the present application, when multiple people exist in the detection scene at the same time and the number of people is more than the required number of people in the detection scene (for example, one person is required and only one person needs to be detected when doing bodybuilding), the target human body bounding box will be switched between multiple people, showing the phenomenon of human body bounding box drift. In order to prevent the target human body bounding box from drifting, in the embodiments of the present application, a human body bounding box surplus strategy can be used to establish more human body bounding boxes than the required number of people in the human body posture detection process.

[0130] In the current frame image, the IOU (Intersection-over-Union, intersection-over-union) of each human body bounding box in the current frame image and the target human body bounding box in the previous frame image is determined respectively; the human body bounding box in the current frame image with the maximum IOU with the target human body bounding box in the previous frame image is determined as the target human body bounding box in the current frame image.

[0131] In a feasible implementation, the center point position (xc, yc) and the width and height information (wc, hc) of each human body bounding box in the current frame image and the center point position (xc, yc) and the width and height information (wc, hc) of the target human body bounding box in the previous frame image can be used to calculate the IOU of each human body bounding box in the current frame image and the target human body bounding box in the previous frame image respectively.

[0132] It can be understood that in the visual display, the frame rate factor mainly affects the accuracy of the display of the target human body bounding box. Therefore, in the embodiments of the present application, an IOU threshold related to the frame rate can be set to reduce the influence of the frame rate on the target human body bounding box, so as to adapt to different image capture devices and display devices.

[0133] In a feasible implementation, the IOU threshold can be calculated by formula 1,

[0134]

[0135] wherein k is an adjustable constant, and θ is the frame rate of the video to be detected.

[0136] In some embodiments,

[0137] In some embodiments, after determining the target human body bounding box in the current frame image, it is determined whether the intersection over union (IOU) of the target human body bounding box in the current frame image and the target human body bounding box in the previous frame image is greater than a preset IOU threshold. If the IOU of the target human body bounding box in the current frame image and the target human body bounding box in the previous frame image is greater than the preset IOU threshold, it is determined whether the size and / or position of the target human body bounding box in the current frame image changes compared to the target human body bounding box in the previous frame image. If the IOU of the target human body bounding box in the current frame image and the target human body bounding box in the previous frame image is less than or equal to the preset IOU threshold, a prompt information is outputted, which is used to prompt the user that the human body in the current frame image has moved out.

[0138] In some embodiments, a change threshold can be set to determine whether the target human body bounding box is shaking.

[0139] wherein, when determining whether the size and / or position of the target human body bounding box in the current frame image changes compared to the target human body bounding box in the previous frame image, if the size and / or position of the target human body bounding box in the current frame image changes compared to the target human body bounding box in the previous frame image, and the change amplitude is less than a preset change threshold, it is determined that the target human body bounding box is shaking. At this time, the size and / or position of the target human body bounding box in the current frame image is adjusted based on the target human body bounding box in the previous frame image, and the adjusted current frame image is sent to the display for display. If the size and / or position of the target human body bounding box in the current frame image does not change compared to the target human body bounding box in the previous frame image, or the size and / or position of the target human body bounding box in the current frame image changes compared to the target human body bounding box in the previous frame image, and the change amplitude is greater than or equal to the preset change threshold, it is determined that the target human body bounding box is not shaking. At this time, the current frame image is sent to the display for display.

[0140] wherein the change amplitude δ of the size and / or position of the target human body bounding box in the current frame image compared to the target human body bounding box in the previous frame image can be determined by formula 2:

[0141]

[0142] wherein I tsize and / or position information of the target human body bounding box in the previous frame image, I t-1 size and / or position information of the target human body bounding box in the previous frame image.

[0143] It can be understood that, due to different frame rates of image acquisition by different camera acquisition devices, and different frame rates supported by different display devices. When the frame rate is large, the image display is smooth, the interval time between the upper and lower frames is short, and the difference between the two frames of images is small, so the above change amplitude is small; and when the frame rate is small, the image display is not smooth, the interval time between the upper and lower frames is long, and the difference between the two frames of images is large, so the above change amplitude is large. In these two cases, in order to adjust the change threshold to avoid misjudgment, a change threshold related to the frame rate can be set to determine whether the target human body bounding box exists jitter.

[0144] In a possible implementation, the preset change threshold is set according to formula 3;

[0145]

[0146] Wherein, α is an adjustable constant, and θ is the frame rate of the video to be detected.

[0147] In some embodiments,

[0148] In some embodiments, referring to Figure 17 , it is assumed that Figure 17 the solid line box in is the target human body bounding box in the current frame image, and the dashed line box is the target human body bounding box in the previous frame image, then since the size and / or position of the target human body bounding box in the current frame image changes compared with the target human body bounding box in the previous frame image, therefore, when the above change amplitude is less than the preset change threshold, the intuitive feeling feedback to the user is that the target human body bounding box in the display interface exists jitter.

[0149] In order to avoid the above jitter, in the embodiments of the present application, if the size and / or position of the target human body bounding box in the current frame image changes compared with the target human body bounding box in the previous frame image, and the change amplitude is less than the preset change threshold, then based on the target human body bounding box in the previous frame image, the size and / or position of the target human body bounding box in the current frame image is adjusted, so that the size and / or position of the target human body bounding box in the current frame image is the same as the target human body bounding box in the previous frame image. Referring to Figure 18 , since the size and / or position of the adjusted target human body bounding box in the current frame image is the same as the target human body bounding box in the previous frame image, the intuitive feeling feedback to the user is that the target human body bounding box in the display interface remains stationary.

[0150] The display device provided in the embodiments of the present application sets the IOU threshold and the change threshold based on the frame rate of the video to be detected, which can effectively reduce the influence of the frame rate on the judgment result when judging whether the human body detection frame is shaking, and can prevent the human body detection frame from shaking and guarantee the accuracy of the visual display of the human body detection frame.

[0151] Based on the content described in the above embodiments, in some embodiments, due to the instability of the human posture detection model and the light factor, the problem of joint point shaking may also exist in the visual display process.

[0152] In some embodiments, in order to prevent the problem of joint point shaking in the visual display process, a floating point method can be used to calibrate the human joint position coordinates, and the stability of the human joint display can be improved based on the light parameter.

[0153] In a feasible implementation, the human joint position (xloc, yloc) in the current frame image can be converted into the floating point human joint position (xf, yf) by the following formula, as formula 4:

[0154]

[0155] wherein i = 0, 1, …, s-1, s, represents the size of the heat map obtained after human posture detection of the current frame image; vi,j represents the value at the coordinate (xi, yj) in the heat map; γ ∈ (0, 1), represents the light parameter.

[0156] wherein the value of γ is different under different light conditions, when the light is sufficient, vi,j is generally large, at this time, γ should be set to be large, and when the light is insufficient or too strong, γ should be set to be small, such as 0.05.

[0157] In the present embodiment, after the floating point human joint position (xf, yf) is determined, the human joint to be displayed is labeled in the current frame image according to the floating point human joint position (xf, yf), and the current frame image labeled with the human joint to be displayed is sent to the display for display.

[0158] The display device provided in the embodiments of the present application performs floating point processing on the human joint information in the current frame image based on the light parameter, which can effectively prevent the problem of shaking of the human joint in the visual display of the human posture detection structure, and improve the visual display effect.

[0159] Some embodiments of the present application also provide a human posture detection method, which comprises:

[0160] S601, human posture detection is performed on a current frame image in a video to be detected to determine a human detection box in the current frame image.

[0161] S602, it is determined whether a size and / or position of a target human detection box in the current frame image changes compared with a target human detection box in a previous frame image of the current frame image.

[0162] S603, if the size and / or position of the target human detection box in the current frame image changes compared with the target human detection box in the previous frame image, and a change amplitude is less than a preset change threshold, the size and / or position of the target human detection box in the current frame image is adjusted based on the target human detection box in the previous frame image, and the adjusted current frame image is displayed.

[0163] The human posture detection method provided in the embodiments of the present application can determine whether the size and / or position of the target human detection box in the current frame image changes by comparing the size and / or position of the target human detection box in the current frame image with the target human detection box in the previous frame image. When the size and / or position of the target human detection box in the current frame image changes, and the change amplitude is less than the preset change threshold, the size and / or position of the target human detection box in the current frame image is adjusted based on the target human detection box in the previous frame image, and the adjusted current frame image is displayed. Thus, the human detection box in the human posture detection result can be prevented from shaking due to the slight difference between the two adjacent frames of pictures, and the visual display effect is improved.

[0164] Based on the above embodiments, in some embodiments of the present application, when the current frame image includes two or more human detection boxes, the IOU of each human detection box in the current frame image with the target human detection box in the previous frame image is determined: the human detection box with the maximum IOU of the target human detection box in the previous frame image in the current frame image is determined as the target human detection box in the current frame image.

[0165] In the embodiments of the present application, it is assumed that the current frame image includes two human detection boxes, which are human detection box 1 and human detection box 2. Then, IOU1 of the human detection box 1 with the target human detection box in the previous frame image and IOU2 of the human detection box 2 with the target human detection box in the previous frame image are determined. If IOU1>IOU2, the human detection box 1 is taken as the target human detection box in the current frame image; if IOU2>IOU1, the human detection box 2 is taken as the target human detection box in the current frame image.

[0166] In a feasible implementation, after the target human detection box in the current frame image is determined, it is judged whether the human has moved out of the current frame image.

[0167] wherein it is determined whether the intersection over union (IOU) of the target human body bounding box in the current frame image and the target human body bounding box in the previous frame image is greater than a preset IOU threshold value; if the IOU of the target human body bounding box in the current frame image and the target human body bounding box in the previous frame image is greater than the preset IOU threshold value, it is determined that the human body in the current frame image has not moved out; if the IOU of the target human body bounding box in the current frame image and the target human body bounding box in the previous frame image is less than or equal to the preset IOU threshold value, it is determined that the human body in the current frame image has moved out.

[0168] In some embodiments, the preset IOU threshold value is set according to formula 5,

[0169]

[0170] wherein k is an adjustable constant, and θ is the frame rate of the video to be detected.

[0171] In some embodiments, if it is determined that the human body in the current frame image has moved out, a prompt information is outputted, which is used to prompt the user that the human body in the current frame image has moved out; if it is determined that the human body in the current frame image has not moved out, it is continued to determine whether the target human body bounding box in the current frame image has changed.

[0172] In a feasible implementation, the bounding box information of the target human body bounding box in the current frame image and the bounding box information of the target human body bounding box in the previous frame image are determined respectively, and the bounding box information includes the center point coordinate information, the width information and the length information of the human body bounding box.

[0173] The difference value between the bounding box information of the target human body bounding box in the current frame image and the bounding box information of the target human body bounding box in the previous frame image is determined; when the difference value is zero, it is determined that the size and / or position of the target human body bounding box in the current frame image has no change compared with the target human body bounding box in the previous frame image; when the difference value is not zero, it is determined that the size and / or position of the target human body bounding box in the current frame image has changed compared with the target human body bounding box in the previous frame image.

[0174] In a feasible implementation, when it is determined that the size and / or position of the target human body bounding box in the current frame image has changed compared with the target human body bounding box in the previous frame image, the change amplitude of the size and / or position of the target human body bounding box in the current frame image is determined according to the difference value. If the change amplitude is less than a preset change threshold value, the size and / or position of the target human body bounding box in the current frame image is adjusted based on the target human body bounding box in the previous frame image, and the adjusted current frame image is sent to the display for display; otherwise, the current frame image is sent to the display for display.

[0175] In a feasible implementation, when it is determined that the size and / or position of the target human body bounding box in the current frame image has no change compared with the target human body bounding box in the previous frame image, or the change amplitude of the size and / or position of the target human body bounding box in the current frame image is greater than or equal to the preset change threshold, the current frame image is sent to the display for display.

[0176] In some embodiments, the preset change threshold is set according to formula 6

[0177]

[0178] wherein, a is an adjustable constant, and θ is the frame rate of the video to be detected.

[0179] The human posture detection method provided in the embodiments of the present application sets the IOU threshold and the change threshold based on the frame rate of the video to be detected, so that when judging whether the human body bounding box is jittered, the influence of the frame rate on the judgment result can be effectively reduced, the human body bounding box can be prevented from jittering, and the accuracy of visual display of the human body bounding box can be ensured.

[0180] Based on the content described in the above embodiments, in some embodiments, before the current frame image is displayed, the human body joint position (xloc, yloc) in the current frame image can be converted into a floating-point human body joint position (xf, yf) by formula 7 as follows:

[0181]

[0182] wherein, i = 0, 1, …, s-1, s, represents the size of the heat map obtained after human posture detection of the current frame image; vi,j represents the value at the coordinate (xi, yj) in the heat map; γ ∈ (0, 1), represents the illumination parameter.

[0183] wherein, the value of γ is different under different illumination conditions, when the light is sufficient, vi,j is generally large, at this time, γ should be set to be large, and when the illumination is insufficient or too strong, γ should be set to be small.

[0184] The human posture detection method provided in the embodiments of the present application performs floating-point processing on the human body joint information in the current frame image based on the illumination parameter, which can effectively prevent the human body joint from jittering when the human posture detection structure is visually displayed, and improve the visual display effect.

[0185] For the sake of explanation, the foregoing descriptions have been presented in terms of specific embodiments. However, it is to be appreciated that specific embodiments described herein are not intended to limit the scope of the present application, which is defined with reference to the following claims. Various modifications and changes can be made thereto by those skilled in the art which fall within the scope of the present application as defined by the following claims. The embodiments were chosen and described in order to explain the principles of the application and the practical application and to enable others skilled in the art to understand for implementing various embodiments and with various modifications as are suited to the particular use contemplated.

Claims

1. A display device, characterized by comprising: The application comprises: a display for displaying a user interface; an image acquisition interface for acquiring images within a photographable range; a controller configured to: in response to a user inputted control instruction for gesture detection, control the image acquisition interface to acquire user images entering the photographable range at a first frequency; control the image acquisition interface to acquire a height difference between a user hand position and an elbow position at a second frequency; the second frequency is greater than the first frequency; if the height difference is less than a height threshold value, and the display device outputs mute audio within a first time length, acquire a screenshot of the user interface; configure image parameters and sound parameters of a current video played by the user interface according to the screenshot.

2. The display device of claim 1, wherein, The controller performing the control of the image acquisition interface to acquire a height difference between a user hand position and an elbow position at a second frequency is further configured to: input the user images into an image recognition model to mark the hand position and the elbow position respectively by the image recognition model, to obtain a hand position mark and an elbow position mark; when the hand position mark and the elbow position mark corresponding positions change, recalculate the height difference according to the current positions of the hand position and the elbow position.

3. The display device of claim 1, wherein, The controller performing the configuration of image parameters and sound parameters of a current video played by the user interface according to the screenshot is further configured to: generate a first instruction based on the screenshot; send the first instruction to a server and receive a second instruction fed back by the server; the second instruction includes the image parameters and sound parameters of the current video identified by the server according to the screenshot; in response to the second instruction, perform parameter adjustment on the current video according to the image parameters and sound parameters.

4. The display device of claim 1, wherein, The controller is further configured to: in response to a user inputted game start instruction, control the display to display a user interface containing a user image and a background image; compare user face features extracted from the user image with a preset family member face feature library to determine the user in the user image; mark the left arm elbow of the user as a first positioning point, mark the right arm elbow of the user as a second positioning point, and mark other joints of the user as third positioning points; the other joints include left hand, left shoulder, neck, right shoulder, right hand, left waist or right waist; according to a spacing distance between the first positioning point and the second positioning point, display a first mark according to a first size of the user at a first position of the user interface; the first mark is a rectangular frame with a width same as the spacing distance and a height same as the height of the user; when the user moves from the first position to a second position, track the position change of the positioning points; generate a second mark according to the position change and a second size of the user at the second position, and display the second mark on the user interface.

5. The display device of claim 4, wherein, The controller is further configured to: after tracking the position change of the positioning points, determine the torso position of the user in the photographable range; if the user is not completely located in the photographable range and the user is in a state of leaving the user interface, the first identifier and the second identifier are revoked; the state of leaving the user interface is that the user is crossing the edge of the photographable range and the torso position of the user is completely located in the photographable range, or the user has crossed the edge of the photographable range and the torso position of the user is not completely located in the photographable range.

6. The display device of claim 4, wherein, The controller configured to perform the comparison of the user facial features extracted from the user image with the preset family member facial feature library is further configured to: perform human body detection on the user image by using a human body detection algorithm to obtain the user joint points; connect the joint points according to a predetermined connection rule to obtain a user upper body edge graph; calculate the similarity of the user upper body edge graph and the family member upper body edge graph to obtain a comparison result according to the similarity.

7. The display device of claim 6, wherein, The controller configured to calculate the similarity of the user upper body edge graph and the family member upper body edge graph is further configured to: obtain the body features of the family member input by the user in advance; the body features include the coordinates of body feature points; connect the feature points according to the predetermined connection rule and the coordinates of the body feature points to obtain a family member upper body edge graph; calculate a first connection angle of the user upper body edge graph and a second connection angle of the family member upper body edge graph, respectively; calculate the similarity according to the first connection angle and the second connection angle.

8. The display device of claim 1, wherein, The controller is further configured to: obtain a to-be-detected video in response to a detection instruction input by the user; the to-be-detected video is a video collected when at least one user performs a motion; perform human body posture detection on each frame of image in the to-be-detected video by using a human body posture detection network model to obtain human body detection information; the human body detection information includes human body position information and human body posture information; the human body position information includes the center point position, width and height of a human body detection frame; the human body posture information includes human body joint point position information; determine a change amplitude between a first human body detection frame of a current frame of image and a second human body detection frame of a previous frame of image according to the center point position, width and height of the first human body detection frame and the center point position, width and height of the second human body detection frame; if the change amplitude is less than a preset change threshold, adjust the first human body detection frame according to the center point position, width and height of the second human body detection frame, and control the display to display the current frame of image after the adjustment; if the change amplitude is greater than or equal to the preset change threshold, control the display to display the current frame of image without adjustment.

9. The display device of claim 8, wherein, The controller is further configured to: if the current frame of image includes two or more first human body detection frames, calculate the intersection over union of each first human body detection frame and the second human body detection frame in the previous frame of image, respectively; mark the first human body detection frame with the largest intersection over union as a target human body detection frame.

10. The display device of claim 9, wherein, The controller, after marking the first human body detection frame with the maximum intersection over union as a target human body detection frame, is further configured to: set an intersection over union threshold according to a frame rate of the video to be detected; if the intersection over union is greater than the intersection over union threshold, determine a change amplitude between the first human body detection frame of the current frame image and the second human body detection frame of the previous frame image according to a center point position, a width and a height of the first human body detection frame and a center point position, a width and a height of the second human body detection frame; and if the intersection over union is less than or equal to the intersection over union threshold, control the display to display prompt information; the prompt information is used to represent that the user has moved out of the current frame image.

11. The display device of claim 1, wherein, The controller is further configured to: invoke a detection program to analyze the user image to determine a user portrait position; if the user portrait position or user posture changes, compare the user portrait position with a preset region in the calibration image to adjust a shooting position of the image acquisition interface according to the preset region.

12. The display device of claim 11, wherein, The controller, after invoking the detection program to analyze the user image, is further configured to: detect limb information in the user image using an image processing program; the limb information includes key points and a limb box; determine key point coordinates of the key points in a preset planar rectangular coordinate system; calculate a center position of the limb box according to the key point coordinates.

13. The display device of claim 12, wherein, The controller is further configured to: calculate a user portrait position obtained according to a corresponding motion state of a plurality of skeleton lines determined by the key points; calculate a key point distance between at least two key points detected in the user image; calculate a distance between the user and the display device according to the key point distance, and match a preset adjustment step length according to the distance; adjust a shooting angle of the image acquisition interface according to the user portrait position and the distance, so that the user portrait position is located in a preset judgment region; the shooting angle is determined according to the preset adjustment step length. 14.A human posture detection method applied to a display device, the display device comprising a display, an image capturing interface, and a controller, the method comprising: The method comprises: in response to a control instruction for posture detection input by a user, control the image acquisition interface to acquire a user image entering a photographable range at a first frequency; control the image acquisition interface to acquire a height difference between a hand position and an elbow position at a second frequency; the second frequency is greater than the first frequency; if the height difference is less than a height threshold value, and the display device outputs a mute audio within a first time length, acquire a screenshot of a user interface after channel switching; configure an optimized image parameter and an optimized sound parameter for a current video played by the user interface according to the screenshot.

15. The human pose detection method of claim 14, wherein, The method further comprises: in response to a game start instruction input by a user, control the display to display a user interface containing a user image and a background image; compare a user face feature extracted from the user image with a preset family member face feature library to determine a user in the user image. The user's left elbow is marked as the first positioning point, the user's right elbow is marked as the second positioning point, and the user's other joints are marked as the third positioning point; the other joints include the left hand, left shoulder, neck, right shoulder, right hand, left waist or right waist; Based on the interval between the first positioning point and the second positioning point, a first identifier is displayed at a first position on the user interface according to the user's first size; the first identifier is a rectangular frame with a width equal to the interval and a height equal to the user's height. When the user moves from the first location to the second location, the positional change of the positioning point is tracked; A second identifier is generated based on the position change and the second size of the user at the second position, and the second identifier is displayed on the user interface.

16. The human pose detection method of claim 14, wherein, In the step of controlling the image acquisition interface to acquire the height difference between the user's hand position and elbow position at a second frequency, the method further includes: The user image is input into an image recognition model, and the hand position and the elbow position are marked by the image recognition model to obtain hand position identifiers and elbow position identifiers respectively; When the corresponding positions of the hand position marker and the elbow position marker change, the height difference is recalculated based on the current positions of the hand position and the elbow position.

17. The human pose detection method of claim 14, wherein, The method further includes: In response to a detection command input by the user, a video to be detected is acquired; the video to be detected is a video captured when at least one user is moving. Human pose detection is performed on each frame of the video to be detected using a human pose detection network model to obtain human detection information. The human detection information includes human position information and human pose information. The human position information includes the center point position, width, and height of the human detection box. The human pose information includes the position information of human joint points. The change amplitude between the first human detection box and the second human detection box is determined based on the center point position, width, and height of the first human detection box in the current frame image and the center point position, width, and height of the second human detection box in the previous frame image. If the change amplitude is less than a preset change threshold, the first human body detection box is adjusted according to the center point position, width and height of the second human body detection box, and the display is controlled to display the adjusted current frame image. If the change magnitude is greater than or equal to a preset change threshold, the display is controlled to show the current frame image without adjustment.

18. The human pose detection method of claim 17, wherein, The method further includes: After tracking the positional change of the location point, it is determined that the user is located in the torso position within the shooting range; If the user is not fully within the camera range and is in the process of leaving the user interface, then the first and second identifiers are revoked; the process of leaving the user interface means that the user is crossing the edge of the camera range and the user's torso is completely within the camera range, or that the user has crossed the edge of the camera range and the user's torso is not within the camera range at all.

19. The human pose detection method of claim 14, wherein, The method further comprises: calling a detection program to analyze the user image to determine a user portrait position; if the user portrait position or user posture changes, comparing the user portrait position with a preset region in the calibration image to adjust the shooting position of the image acquisition interface according to the preset region.

20. The human pose detection method of claim 19, wherein, In the step of calling the detection program to analyze the user image, the method further comprises: detecting limb information in the user image using an image processing program; the limb information comprises key points and a limb box; determining key point coordinates of the key points in a preset planar rectangular coordinate system; calculating a center position of the limb box according to the key point coordinates.

Citation Information

Patent Citations

  • A display device and a human face positioning method

    CN112672062B

  • Face feature value creating method, figure locking and tracking method and display equipment

    CN112862859A

  • Floor mat system and associated, computer medium and computer-implemented methods for monitoring and improving health and productivity of employees

    CN103781408A

  • Method of performing function of device and device for performing the method

    CN109284001A