Screen interaction method and device, display device and storage medium

CN120872219BActive Publication Date: 2026-08-07GUANGZHOU SHIYUAN ELECTRONICS CO LTD +1
View PDF 3 Cites 0 Cited by

Patent Information

Authority / Receiving Office
CN · China
Patent Type
Patents(China)
Current Assignee / Owner
GUANGZHOU SHIYUAN ELECTRONICS CO LTD
Filing Date
2024-04-30
Publication Date
2026-08-07

AI Technical Summary

Technical Problem

[0003]本申请提供一种屏幕交互方法、装置、显示设备及存储介质,旨在解决现有技术中智能交互性能低的技术问题

Benefits of technology

[0047]一个或多个应用程序,其中所述一个或多个应用程序被存储于所述存储器中,并配置为由所述处理器执行以实现任一项所述的屏幕交互方法中的步骤。

✦ Generated by Eureka AI based on patent content.

Smart Images

  • Figure CN120872219B_ABST
    Figure CN120872219B_ABST
Patent Text Reader

Abstract

The application provides a screen interaction method and device, a display device and a storage medium. The hand key points of a master controlling a target screen are obtained by user positioning and skeleton tracking on a target image containing the master; target pose information of the hand limbs of the master is determined according to three-dimensional position information of each hand key point, the target pose information including at least one of a bending degree, a pointing direction and a motion trend; a target gesture of the master is determined according to the target pose information of the hand limbs and standard pose information of the hand limbs in each preset gesture; and the target screen is interactively controlled in response to an operation event corresponding to the target gesture. No additional interactive device such as a mouse is needed, the device configuration cost is reduced, the pointing direction of the hand is analyzed based on a three-dimensional space to determine a target interaction position, the accuracy of the interaction is ensured, and the interaction performance is improved.
Need to check novelty before this filing date? Find Prior Art

Description

Technical Field

[0001] This application relates to the field of intelligent interaction technology, specifically to a screen interaction method, device, display equipment, and storage medium. Background Technology

[0002] With the development of intelligent interaction technology, the application of interaction scenarios has become increasingly diversified. When users interact with the display screen, they generally control the interaction by changing the screen cursor or screen focus with a mouse, or by controlling the interaction through touch screen. However, these interaction methods limit users to operating the mouse only from where it is placed, or to standing in front of the screen and interacting through touch screen, which restricts intelligent interaction and reduces its performance. Summary of the Invention

[0003] This application provides a screen interaction method, apparatus, display device, and storage medium, aiming to solve the technical problem of low intelligent interaction performance in the prior art.

[0004] In a first aspect, this application provides a screen interaction method, including:

[0005] The key hand points of the controller who controls the target screen are obtained, and the key hand points are determined by user localization and skeleton tracking of the target image containing the controller.

[0006] Based on the three-dimensional position information of each of the key hand points, the target pose information of the controller's hand limbs is determined, and the target pose information includes at least one of bending degree, pointing and movement trend;

[0007] Based on the target posture information of the hand limbs and the standard posture information of the hand limbs in each preset gesture, the target gesture of the controller is determined;

[0008] Interactive control of the target screen in response to the operation event corresponding to the target gesture.

[0009] This solution determines at least one target pose information from three-dimensional space, including the bending degree, pointing direction, and movement trend of the hand limbs. Further, based on the target pose information and the standard pose information of the hand limbs in each preset gesture, it determines the target gesture of the controller to achieve gesture-based interactive control. This avoids the limitations of two-dimensional space, ensures the flexibility of the controller's position movement during interaction, and eliminates the need for additional interactive devices such as mice, reducing equipment configuration costs. Furthermore, by analyzing the pointing direction of the hand in three-dimensional space to determine the target interaction position, it ensures the accuracy of the interaction and improves interaction performance. In addition, the implementation of this application uses standard pose information to represent the hand limbs in the preset gestures, resulting in greater algorithm flexibility. Moreover, it eliminates the need to retrain the algorithm model when adding or adjusting the preset gesture library, improving the algorithm's versatility and expanding its applicability.

[0010] In one possible implementation of this application, determining the target pose information of the controller's hand limbs based on the three-dimensional position information of each of the key hand points includes:

[0011] Based on the three-dimensional position information of the target hand key points, the curvature and orientation of the hand limb corresponding to the target hand key points are determined, and the target hand key points are at least two hand key points on the same hand limb;

[0012] The movement trend of the hand limb is determined based on the curvature and orientation of the hand limb in at least two target images;

[0013] The bending degree of the hand limb, the pointing direction, and the movement trend are set as the target pose information of the hand limb.

[0014] In one possible implementation of this application, determining the curvature and orientation of the hand limb corresponding to the target hand key points based on the three-dimensional position information of the target hand key points includes:

[0015] The bending angle of each segment is calculated based on the three-dimensional position information of every two adjacent key points of the same hand limb.

[0016] The bending angle of each segment in each hand limb is counted to obtain the bending degree of the hand limb;

[0017] If the curvature of the hand limb is less than the preset curvature of the hand limb, then the direction of the hand limb is calculated based on the three-dimensional position information of the main control key point and the three-dimensional position information of the auxiliary key point. The auxiliary key point includes at least one of the wrist key point, interphalangeal key point and palmodigital key point corresponding to the hand limb.

[0018] This solution determines the degree of curvature of each hand limb based on the relationship between the curvature of each limb and the preset curvature corresponding to each degree of curvature. It further sets the curvature of the hand limb as the curvature, that is, it classifies the curvature of the hand limb and calculates and determines the direction, thereby improving the fault tolerance of the curvature confirmation of the hand limb and avoiding gesture recognition failure due to differences in the gesture habits of different people.

[0019] In one possible implementation of this application, determining the target gesture of the controller based on the target pose information of the hand limbs and the standard pose information of the hand limbs in each preset gesture includes:

[0020] The first evaluation information is determined based on the curvature of the hand limbs and the standard curvature of the hand limbs in each preset gesture.

[0021] Based on the direction of the hand limbs and the standard direction of the hand limbs in each preset gesture, determine the second evaluation information; or,

[0022] The third evaluation information is determined based on the movement trends of the hand limbs and the standard movement trends of the hand limbs in each preset gesture.

[0023] The target gesture of the controller is determined based on at least one of the first evaluation information, the second evaluation information, and the third evaluation information.

[0024] This solution can enhance the flexibility and versatility of gesture recognition.

[0025] In one possible implementation of this application, before determining the controller's target gesture based on the target pose information of the hand limbs and the standard pose information of the hand limbs in each preset gesture, the following steps are included:

[0026] Obtain the gesture type of each preset gesture, and the standard pose information corresponding to the gesture type;

[0027] The standard pose information includes at least one of standard curvature, standard orientation, and standard motion tendency.

[0028] In one possible implementation of this application, after obtaining the key points of the hand of the controller who controls the target screen, the method includes:

[0029] The hand joint model is used to predict the three-dimensional information of the key points of the hand to obtain the initial three-dimensional position information of the key points of the hand.

[0030] Based on the camera intrinsic parameters corresponding to the target image, the initial three-dimensional position information of the hand key points is mapped in two dimensions to obtain the predicted key points.

[0031] If the distance between the predicted key point and the hand key point is less than a preset distance threshold, then the initial three-dimensional position information is set as the three-dimensional position information of the hand key point.

[0032] This solution improves the accuracy of three-dimensional position information by enhancing the model precision of the hand joint model.

[0033] In one possible implementation of this application, obtaining the key hand points of the controller controlling the target screen includes:

[0034] Obtain the target image corresponding to the target screen;

[0035] Human detection is performed on the target image using a human detection model to identify the main controller in the target image and the hand area of ​​the main controller;

[0036] The key points of the hand region are detected by the hand key point detection model to determine the two-dimensional position information of the hand joints of the controller, as well as the confidence level of the two-dimensional position information;

[0037] Based on the two-dimensional position information and the confidence level corresponding to each two-dimensional position information, key points of the hand are determined.

[0038] This solution uses a model to detect the hand region and key points, improving the efficiency of key point detection. At the same time, it determines the key points of the hand based on the confidence level corresponding to the two-dimensional position information, avoiding the limitations of image key point detection and ensuring the rationality and accuracy of the key points of the hand.

[0039] Secondly, this application also provides a screen interaction device, the device comprising:

[0040] The acquisition module is used to acquire the key hand points of the master controller who controls the target screen. The key hand points are determined by user positioning and skeleton tracking of the target image containing the master controller.

[0041] The pose determination module is used to determine the target pose information of the controller's hand limbs based on the three-dimensional position information of each of the key hand points. The target pose information includes at least one of bending degree, pointing and movement trend.

[0042] The gesture determination module is used to determine the target gesture of the controller based on the target pose information of the hand limbs and the standard pose information of the hand limbs in each preset gesture.

[0043] An interactive control module is used to interactively control the target screen in response to an operation event corresponding to the target gesture.

[0044] Thirdly, this application also provides a display device, the display device comprising:

[0045] One or more processors;

[0046] Memory; and

[0047] One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the steps in any of the screen interaction methods described above.

[0048] Fourthly, this application also provides a computer-readable storage medium having a computer program stored thereon, the computer program being loaded by a processor to perform the steps in any of the screen interaction methods described herein.

[0049] This application provides a screen interaction method, apparatus, display device, and storage medium. The method involves acquiring key hand points of a controller who is controlling a target screen. These key hand points are determined through user localization and skeletal tracking of a target image containing the controller. Based on the three-dimensional position information of each key hand point, target pose information of the controller's hand limbs is determined. This target pose information includes at least one of curvature, pointing direction, and movement trend. Based on the target pose information of the hand limbs and the standard pose information of the hand limbs in each preset gesture, a target gesture of the controller is determined. The target screen is then interactively controlled in response to an operation event corresponding to the target gesture. This solution determines at least one target pose information from the bending degree, pointing direction, and movement trend of the hand limbs in three-dimensional space. Further, based on the target pose information and the standard pose information of the hand limbs in each preset gesture, it determines the target gesture of the controller to achieve gesture-based interactive control. This avoids the limitations of two-dimensional space, ensures the flexibility of the controller's position movement during interaction, and eliminates the need for additional interactive devices such as mice, reducing equipment configuration costs. Furthermore, by analyzing the pointing direction of the hand in three-dimensional space to determine the target interaction position, it ensures the accuracy of the interaction and thus improves interaction performance. Attached Figure Description

[0050] To more clearly illustrate the technical solutions in the embodiments of this application, the accompanying drawings used in the description of the embodiments will be briefly introduced below. Obviously, the accompanying drawings described below are only some embodiments of the present invention. For those skilled in the art, other drawings can be obtained based on these drawings without creative effort.

[0051] Figure 1 This is a schematic diagram of a scenario for the screen interaction method provided in an embodiment of this application;

[0052] Figure 2This is a schematic flowchart of an embodiment of the screen interaction method provided in this application.

[0053] Figure 3 A schematic flowchart of one implementation scheme for determining target pose information in the screen interaction method provided in this application embodiment;

[0054] Figure 4 A schematic diagram of constraint information between three-dimensional hands included in the hand joint model provided for the implementation scheme of this application;

[0055] Figure 5 A schematic flowchart of one embodiment of the screen interaction method provided in this application for determining the target gesture of the main controller;

[0056] Figure 6 A schematic flowchart of one embodiment of the screen interaction method provided in this application for determining key hand points;

[0057] Figure 7 Flowchart of another embodiment of the screen interaction method provided in this application;

[0058] Figure 8 This is a schematic diagram of an embodiment of the screen interaction device provided in this application.

[0059] Figure 9 This is a schematic diagram of an embodiment of the display device provided in this application. Detailed Implementation

[0060] The technical solutions of the embodiments of this application will be clearly and completely described below with reference to the accompanying drawings. Obviously, the described embodiments are only some embodiments of the present invention, and not all embodiments. Based on the embodiments of the present invention, all other embodiments obtained by those skilled in the art without creative effort are within the scope of protection of the present invention.

[0061] In the description of this invention, it should be understood that the terms "center," "longitudinal," "lateral," "length," "width," "thickness," "upper," "lower," "front," "rear," "left," "right," "vertical," "horizontal," "top," "bottom," "inner," and "outer," etc., indicating orientation or positional relationships based on the orientation or positional relationships shown in the accompanying drawings, are only for the convenience of describing the invention and simplifying the description, and do not indicate or imply that the device or element referred to must have a specific orientation, or be constructed and operated in a specific orientation, and therefore should not be construed as a limitation of the invention. Furthermore, the terms "first" and "second" are used for descriptive purposes only and should not be construed as indicating or implying relative importance or implicitly specifying the number of indicated technical features. Thus, features defined with "first" and "second" may explicitly or implicitly include one or more of the stated features. In the description of this invention, "a plurality of" means two or more, unless otherwise explicitly specified.

[0062] In this embodiment, "and / or" describes the relationship between associated objects, indicating that three relationships can exist. For example, A and / or B can represent three cases: A alone, A and B simultaneously, and B alone. Additionally, the character " / ", unless otherwise specified, generally indicates that the preceding and following associated objects have an "or" relationship.

[0063] In this application, the term "exemplary" is used to mean "serving as an example, illustration, or description." Any embodiment described as "exemplary" in this application is not necessarily to be construed as being more preferred or advantageous than other embodiments. The following description is provided to enable any person skilled in the art to make and use the invention. Details are set forth in the following description for purposes of explanation. It should be understood that those skilled in the art will recognize that the invention can be made without using these specific details. In other instances, well-known structures and processes will not be described in detail to avoid obscuring the description of the invention with unnecessary detail. Therefore, the invention is not intended to be limited to the embodiments shown, but is consistent with the broadest scope of the principles and features disclosed in this application.

[0064] Gesture recognition is a technology that enables interaction by analyzing human hand gestures. It converts human gestures into computer-understandable instructions, thus enabling interaction with computers or other devices. Gesture recognition technology is applied in multiple fields, including human-computer interaction, virtual reality, augmented reality, smart homes, and healthcare, offering advantages such as naturalness, convenience, and intuitiveness in interaction design.

[0065] In existing solutions, large screens primarily use touchscreens and remote controls for interaction. Touchscreens offer precise positioning but require close-range operation, and the large screen size makes the user experience less than ideal. Remote controls support interaction from medium to long distances but generally do not support actual cursor positioning. Remote controls require additional hardware, increasing sales costs and incurring additional maintenance costs (charging, preventing loss, etc.), especially for public facilities.

[0066] Therefore, embodiments of this application provide a screen interaction method, apparatus, device, and computer-readable storage medium (hereinafter referred to as storage medium). By acquiring a target image of a target screen for remote control, the method analyzes the key points of the controller's hand, determines the target pose information of the hand limbs based on the three-dimensional position information of the key points, and further determines the target gesture based on the target pose information of the hand limbs and the standard pose information of the hand limbs in a preset gesture, thereby realizing remote interactive control, improving the flexibility and accuracy of remote interactive control, and avoiding the problem of increased costs associated with using remote control devices for remote interaction. These will be described in detail below.

[0067] This application provides a screen interaction method, apparatus, display device, and computer-readable storage medium, which will be described in detail below.

[0068] The screen interaction method in this embodiment of the invention is applied to a screen interaction device, which is set in a display device. The display device is provided with one or more processors, a memory, and one or more applications, wherein one or more applications are stored in the memory and configured to be executed by the processor to implement the screen interaction method. The display device can be a terminal, such as a mobile phone or a tablet computer, or it can be a server or a service cluster composed of multiple servers.

[0069] like Figure 1 As shown, Figure 1 This is a schematic diagram of a screen interaction method according to an embodiment of the present application. The screen interaction scenario in this embodiment includes a display device 100 (the display device 100 integrates a screen interaction device), and a computer-readable storage medium corresponding to the screen interaction is run in the display device 100 to perform the screen interaction steps.

[0070] Understandable, Figure 1The display device in the scenario of the screen interaction method shown, or the device included in the display device, does not constitute a limitation on the embodiments of the present invention. That is, the number or type of device included in the scenario of the screen interaction method, or the number or type of device included in each device, does not affect the overall implementation of the technical solution in the embodiments of the present invention, and can all be considered as equivalent substitutions or derivatives of the technical solutions claimed in the embodiments of the present invention.

[0071] In this embodiment of the invention, the display device 100 is mainly used for: acquiring key hand points of a controller who controls a target screen, wherein the key hand points are determined by user positioning and skeletal tracking of a target image containing the controller; determining target pose information of the controller's hand limbs based on the three-dimensional position information of each key hand point, wherein the target pose information includes at least one of curvature, pointing, and movement trend; determining the controller's target gesture based on the target pose information of the hand limbs and the standard pose information of the hand limbs in each preset gesture; and interactively controlling the target screen in response to an operation event corresponding to the target gesture.

[0072] In this embodiment of the invention, the display device 100 can be an independent display device, or a network of display devices or a cluster of display devices. For example, the display device 100 described in this embodiment of the invention includes, but is not limited to, a computer, a network host, a single network display device, a set of multiple network display devices, or a cloud display device composed of multiple display devices. The cloud display device is composed of a large number of computers or network display devices based on cloud computing.

[0073] Those skilled in the art will understand that Figure 1 The application environment shown is merely one application scenario of the solution in this application and does not constitute a limitation on the application scenario of the solution in this application. Other application environments may include those that are more specific to this application. Figure 1 The number of display devices shown, or the network connectivity of display devices, for example... Figure 1 Only one display device is shown in the diagram. It is understood that the scenario of this screen interaction method may also include one or more other display devices, which are not specifically limited here. The display device 100 may also include a memory for storing data, such as storing image information acquired by shooting.

[0074] Furthermore, in the screen interaction method scenario of this application, the display device 100 is equipped with a display device 200, which serves as the target large screen. Alternatively, the display device 100 may not have a display device 200 communicating with an external display device. The display device 200 is used to output the results of the screen interaction method executed within the display device. The display device 100 can access a background database 300 (the background database can be located in the local storage of the display device or it can be located in the cloud). The background database 300 stores information related to screen interaction.

[0075] It should be noted that, Figure 1 The schematic diagram of the screen interaction method shown is merely an example. The scenarios of the screen interaction method described in the embodiments of the present invention are for the purpose of more clearly illustrating the technical solutions of the embodiments of the present invention, and do not constitute a limitation on the technical solutions provided in the embodiments of the present invention.

[0076] Based on the scenarios described above for screen interaction methods, embodiments of screen interaction methods are proposed.

[0077] like Figure 2 The diagram shown is a flowchart of an embodiment of the screen interaction method in this application. The screen interaction method includes steps S201-S204:

[0078] S201. Obtain the key points of the hands of the master controller who controls the target screen.

[0079] The key hand points are determined by user localization and skeletal tracking of a target image containing the controller.

[0080] The target image is an image of the environment corresponding to the target screen, and the target image includes at least one controller who controls the target screen.

[0081] It is understood that the target image can be captured by the camera device of the target screen itself, or by the camera device that transmits data with the target screen; this application does not make any specific limitations.

[0082] The controller is the user who interacts with the target screen using gestures, and the target screen adjusts its display according to the controller's gesture changes.

[0083] Among them, the key points of the hand are the coordinate points used to represent the form of the hand movements of the controller. For example, they can be key points representing the position of each hand limb. For example, they include key points representing the thumb, index finger, middle finger, ring finger and little finger. The key points representing a hand limb can include: fingertip key points, wrist key points, interphalangeal key points, palm and finger key points, etc.

[0084] In some embodiments of this application, hand key points may include identification information, which is used to identify which hand key point is represented by the corresponding hand key point, such as representing a fingertip key point, a wrist key point, etc.

[0085] In some other embodiments of this application, there are hand constraint relationships among the hand key points to characterize which hand key point corresponds to which position of the hand key point. For example, the hand constraint relationship can be a mapping relationship between the hand key point and a preset hand key point model or a set of hand key point positions.

[0086] The user positioning, that is, identifying the user's location in the target image, can be understood to be achieved through methods such as user face recognition and human body recognition.

[0087] The skeletal tracking refers to locating and tracking the human skeleton of each user to identify skeletal changes in each user. Based on these changes, the system further determines the controller and key points of the controller's hands. It is understood that skeletal tracking can be implemented using a model. For example, user images are input into a preset model for skeletal prediction to determine the key points of the human skeleton corresponding to the target image, thereby determining the key points of the controller's hands. It is understood that the preset model can be trained to output key points of the hands, i.e., it only performs skeletal tracking on the controller's hands to obtain the key points of the hands.

[0088] In one embodiment of this application, the screen interaction method is applied to a display device, which includes the target screen, a camera corresponding to the target screen, and a processor. For example, the display device can be a smart TV, a conference large screen device, etc. In the application scenario, the user can trigger remote screen control and send it to the processor through screen interaction controls displayed on the target screen or screen interaction controls displayed on a mobile device connected to the target screen. The processor responds to the remote screen control operation, controls the camera to start, and performs target image acquisition. After obtaining the target image, the processor inputs the target image into a preset model for user positioning and skeleton tracking, and then outputs the key points of the controller's hand.

[0089] S202. Based on the three-dimensional position information of each of the key hand points, determine the target pose information of the controller's hand limbs.

[0090] The target pose information includes at least one of curvature, orientation, and motion trend.

[0091] The three-dimensional position information, namely the position coordinates of the hand key points in three-dimensional space, can be understood as follows: the three-dimensional space can be, but is not limited to, the world coordinate system corresponding to the camera that acquires the target image, or a preset three-dimensional coordinate system that has a positional mapping relationship with the hand key points. It can be understood that, in the embodiments of this application, the target screen has preset pose information in the three-dimensional space, and the pose information of the target screen in the three-dimensional space can be obtained through camera calibration.

[0092] The hand limbs include, but are not limited to: thumb, index finger, middle finger, ring finger, and little finger. For example, the key points of the hand include: fingertip joints, wrist joints, interphalangeal joints, and metacarpophalangeal joints, which characterize the index finger.

[0093] The curvature includes the curvature of each hand limb, for example, the curvature of the thumb, index finger, middle finger, ring finger, and little finger.

[0094] In one embodiment of this application, the curvature can be determined by calculating the relative offset between the three-dimensional position information corresponding to the key points of the hand using a preset offset calculation formula. For example, the relative offset can be the angular offset, position offset, etc. between the key points of the hand.

[0095] In other embodiments of this application, the curvature can also be calculated by a preset model. For example, the three-dimensional position information corresponding to the key points of the hand is input into the preset model, and the curvature of the hand limb is output through model analysis. It is understood that the preset model can be obtained through training.

[0096] The direction referred to here, that is, the direction of each hand limb, can be exemplarily defined as up, down, left, right, front, back, etc., and may also be further limited to include upper left, lower left, etc., but this application does not make specific limitations.

[0097] It is understood that the direction is defined according to the three-dimensional position information corresponding to the three-dimensional space. For example, in the three-dimensional space, the x-axis coordinate corresponds to left and right direction, the y-axis coordinate corresponds to up and down direction, and the z-axis coordinate corresponds to front and back direction. Furthermore, the direction of the hand limb can be determined according to the relative positional relationship between the three-dimensional positional information corresponding to the key hand points representing the same hand limb. For example, the direction of the index finger can be determined according to the relative positional relationship between the three-dimensional positional information corresponding to the fingertip joint and the three-dimensional positional information corresponding to the interphalangeal joint in the key hand points representing the index finger.

[0098] The movement trend is used to characterize the direction of hand movement, such as sliding to the left or sliding to the right.

[0099] Specifically, in one embodiment of this application, the motion trend can be determined by combining the three-dimensional position information of the key hand points in the corresponding frames of the target image and the position change vector between the key hand points in the target image.

[0100] In another embodiment of this application, the motion trend can also be predicted by analyzing the pixel motion direction of the corresponding hand region in the target image using an optical flow algorithm. This application does not make any specific limitations on this.

[0101] Specifically, in one embodiment of this application, after determining the key points of the hand, the display device further performs three-dimensional position information mapping through the model to obtain the three-dimensional position information corresponding to each key point of the hand. Furthermore, based on the relative positional relationship between the three-dimensional position information of each key point of the hand, as well as the positional movement direction, the target pose information of the controller's hand limbs is determined.

[0102] S203. Determine the target gesture of the controller based on the target posture information of the hand limbs and the standard posture information of the hand limbs in each preset gesture.

[0103] The target gesture is a gesture used to instruct the target screen to interact. It is understood that the target gesture can be any type of gesture, including dynamic gestures, static gestures, and static gestures carrying a running trend. This application does not make any specific limitations.

[0104] It is understood that the preset gesture corresponds to the gesture type of the target gesture, which may include dynamic gestures, static gestures, and static gestures carrying a motion trend.

[0105] In one embodiment of this application, after obtaining the target limb posture of the hand limb, the display device compares the target posture with the standard limb posture of each preset gesture, and selects the preset gesture that matches the standard limb posture with the target limb posture as the target gesture. The definition of "matching" can be that the similarity is greater than a preset similarity threshold, or the posture error is greater than a preset error threshold, and a matching judgment is made.

[0106] S204. Perform interactive control on the target screen in response to the operation event corresponding to the target gesture.

[0107] Specifically, after receiving the target gesture, the display device further obtains the interactive control command corresponding to the target gesture, and controls the target screen to make adjustments according to the interactive control command.

[0108] The interactive control instructions corresponding to the target gesture are not specifically limited in this application. For example, the target gesture triggered by the main controller includes two categories: regular gesture actions and swipe gesture actions (i.e., dynamic gestures, or static gestures with a motion trend). The corresponding operation instructions may include screenshot, return to the previous menu, confirm, freeze / unfreeze the screen, return to the desktop, open the menu, mute, and turn pages (previous page, next page), etc.

[0109] In this embodiment, key hand points of the controller corresponding to the target screen are acquired by capturing target images. Furthermore, based on the three-dimensional position information of the key hand points, at least one target pose information, including the bending degree, pointing direction, and movement trend of the hand limbs, is determined from a three-dimensional spatial perspective. Further, based on the target pose information and the standard pose information of the hand limbs in each preset gesture, the target gesture of the controller is determined to achieve gesture-based interactive control. This avoids the limitations of two-dimensional space, ensures the flexibility of the controller's position movement during interaction, and eliminates the need for additional interactive devices such as mice, reducing equipment configuration costs. Moreover, the hand pointing analysis based on three-dimensional space determines the target interaction position, ensuring the accuracy of the interaction and thus improving interaction performance. In this embodiment, the standard pose information is used to represent the hand limbs in the preset gestures, resulting in higher algorithm flexibility. Furthermore, the algorithm model does not need to be retrained when adding or adjusting the preset gesture library, improving the algorithm's versatility and expanding its applicability.

[0110] Furthermore, based on the above implementation scheme, this application provides an implementation scheme for determining target pose information, see [link to relevant documentation]. Figure 3 , Figure 3 This is a flowchart illustrating one implementation of the target pose information determination method in the screen interaction method provided in this application, specifically including steps S301-S303:

[0111] S301. Based on the three-dimensional position information of the target hand key points, determine the curvature and orientation of the hand limb corresponding to the target hand key points, wherein the target hand key points are at least two hand key points on the same hand limb.

[0112] Specifically, in the embodiments of this application, a three-dimensional information prediction model is used to predict the three-dimensional position information of key points of the target hand in the target image, thereby obtaining the three-dimensional position information of key points of the target hand.

[0113] For example, in the embodiments of this application, the three-dimensional information prediction model is a hand joint model. The three-dimensional information of the key hand points is predicted by using the hand joint model to obtain the three-dimensional position information of the key hand points. The hand joint model is constructed based on kinematic structure, including constraint information between the three-dimensional hand parts. The type of joint model is not limited; it can be a parametric template model (e.g., MANO, Model for Articulated Hands with Object Interaction, a model for hand pose estimation. It is a deep learning-based method designed to accurately estimate the three-dimensional pose and gestures of the hand), or an optimized constraint set of finger lengths learned based on statistical information. For example, see [link to relevant documentation]. Figure 4 , Figure 4 This diagram illustrates the constraint information between the three-dimensional hand joints in the hand joint model provided for the implementation of this application. The constraint information includes 21 key hand points, forming a tree structure with the root node at the bottom and child nodes at the top. The wrist key point is the root node. First-level child nodes are the five palmodigital key points connected to the wrist key point. Second-level child nodes are the points corresponding to the lower-middle joints of the fingers closest to the fingertip key point, connected to each palmodigital key point. Fourth-level nodes are the fingertip key points of each finger (not shown in the diagram, representing the fingertip tip). Third-level child nodes are the points corresponding to the upper-middle joints of the fingers located between the palmodigital key points and the fingertip key points, connected to each palmodigital key point. Bending the parent node affects the 3D position of the child nodes. Due to the limitations of human anatomy, the direction and range of bending for each node are restricted, thus reducing the degrees of freedom.

[0114] In this scheme, each point on the fingers of the hand joint model has 1-2 degrees of freedom in bending along two axes, and the wrist itself contains 6 degrees of freedom (position in three dimensions and rotation in three dimensions), thus forming constraint information between the three-dimensional hand parts. If there are no constraints on the 3D points, 21*3 variables need to be solved. In this implementation scheme, after kinematic model constraints, with the finger joint length fixed (which can be understood as, after obtaining the key points of the hand, the corresponding finger length can be calculated based on the key points of the hand), only 23 variables need to be solved, significantly reducing the difficulty of solving and the number of degrees of freedom. Therefore, the process of predicting the three-dimensional information of the key points of the hand by using the hand joint model to perform three-dimensional information prediction can reduce the amount of data processing while improving the calculation accuracy of the three-dimensional position information.

[0115] In this embodiment, a hand keypoint detection model is used to detect keypoints in the hand region, outputting two-dimensional position information and confidence levels for each keypoint. In other embodiments, the hand keypoint detection model also outputs relative / estimated depth, relative distance between finger joints, and predicted values ​​of finger joint rotation angles. Further, based on the preliminary results output by the network (two-dimensional position information, and the confidence levels of the two-dimensional position information, or including relative / estimated depth, relative distance between finger joints, and predicted values ​​of finger joint rotation angles), a hand joint model is used to further solve for a set of hand 3D keypoint positions that are most consistent with (with the smallest error) the actual output values ​​from all viewpoints (this process is the step mentioned in the above embodiment of determining hand keypoints based on the two-dimensional position information and the confidence levels corresponding to each two-dimensional position information). This position is the final determined three-dimensional position information of the hand keypoints, which also satisfies some constraints on hand length (based on statistical information from offline data).

[0116] S302. Determine the movement trend of the hand limb based on the bending degree and direction of the hand limb in at least two target images.

[0117] For example, in this embodiment of the application, the movement direction corresponding to each hand key point is determined based on the positional change information between the three-dimensional position of each key point and the corresponding historical key point three-dimensional position, and the movement direction is taken as the hand movement trend. The three-dimensional position of the historical key point is determined based on the historical target image corresponding to the target screen. The acquisition time of the historical target image is separated from the acquisition time of the target image by a preset time interval, which is used to limit the historical target image to the previous frame or several frames of the target image.

[0118] S303. Set the bending degree of the hand limb, the pointing direction, and the movement trend as the target pose information of the hand limb.

[0119] Specifically, in the embodiments of this application, after calculating the curvature, the direction, and the movement trend, the obtained curvature, direction, and movement trend of the hand limb can be set as the target pose information of the hand limb.

[0120] It is understood that in some other embodiments of this application, the curvature, orientation, and motion trend may be calculated as at least one of the target pose information.

[0121] This solution combines curvature, direction, and motion trend as target pose information to determine the target gesture, ensuring the accuracy of gesture determination while improving the flexibility and diversity of gesture configuration. For example, a gesture can correspond to different directions and / or different motion trends, which can correspond to different gesture commands. From the user's perspective, users have more configuration space for the same gesture, and do not need to remember many cumbersome correspondences between gestures and finger commands, thus reducing the difficulty of gesture interaction.

[0122] Furthermore, based on the above implementation scheme, this application also provides another implementation scheme with determined curvature and orientation, specifically including the following steps:

[0123] (1) Based on the three-dimensional position information of each two adjacent target hand key points on the same hand limb, the bending angle of each limb segment is calculated;

[0124] (2) Calculate the bending angle of each segment in each hand limb to obtain the bending degree of the hand limb;

[0125] (3) If the curvature of the hand limb is less than the preset curvature of the hand limb, the direction of the hand limb is calculated based on the three-dimensional position information of the main control key point and the three-dimensional position information of the auxiliary key point of the hand limb.

[0126] Among them, the main control key point is the distal key point of the hand limb, for example, the distal key point is the fingertip key point; the auxiliary key points include at least one of the wrist key point, interphalangeal key point, and palmodigital key point corresponding to the hand limb.

[0127] For example, auxiliary key points are determined based on the directional offset between at least two of the wrist key points, interphalangeal key points, and palmodigital key points and the main control key point. For instance, if the target key points include the wrist key points and interphalangeal key points, the first directional offset between the wrist key points and the main control key point, and the second directional offset between the interphalangeal key points and the main control key point are calculated respectively. If the relative distance between the first directional offset and the second directional offset is less than a preset relative distance, any one of the wrist key points and interphalangeal key points is set as an auxiliary key point. If the relative distance between the first directional offset and the second directional offset is greater than or equal to the preset relative distance, the midpoint between the wrist key points and interphalangeal key points is set as an auxiliary key point. The specific implementation method is not specifically limited in this application.

[0128] Specifically, in one embodiment of this application, the finger joint bending angle is calculated based on the three-dimensional position information of every two adjacent target hand key points on the same hand limb. This angle can be obtained by accumulating the angles between the connecting bones of adjacent key points on each finger. This angle is used to determine the degree of finger bending. In this embodiment, the degree of bending can correspond to different levels of bending, including three levels: straightening, partial bending, and full bending. Each level corresponds to a different preset degree of bending. The degree of bending of each hand limb is divided by the preset degree of bending. For example, straightening corresponds to the first preset degree of bending; partial bending corresponds to the second preset degree of bending; and full bending corresponds to the third preset degree of bending. The first preset degree of bending is less than the second preset degree of bending, and the second preset degree of bending is less than the third preset degree of bending. Further, the degree of bending of the hand limb is determined based on the relationship between the degree of bending of each hand limb and the preset degree of bending corresponding to each degree of bending. The degree of bending is further set as the degree of bending of the hand limb. That is, the degree of bending of the hand limb is classified to improve the fault tolerance of the degree of bending of the hand limb and avoid the failure of gesture recognition due to differences in the gesture habits of different people.

[0129] Furthermore, in the embodiments of this application, the preset curvature of each hand limb can be general or individually preset. That is, the thresholds for different fingers may be different. For example, for the index finger, <= 30 degrees indicates straightness, and >= 120 degrees indicates full flexion; for the little finger, <= 40 degrees indicates straightness, and >= 100 degrees indicates full flexion. This threshold is obtained offline based on training data statistics and remains fixed during actual operation.

[0130] It is understood that in some other embodiments of this application, the gesture recognition can be based on more or less preset curvature settings to set more or fewer levels of curvature, depending on different accuracy recognition requirements.

[0131] Specifically, in the embodiments of this application, when the curvature of the hand limb is less than the first preset curvature of the hand limb, it indicates that the hand limb is in a straight state. At this time, the direction of the hand limb is calculated based on the three-dimensional position information of the fingertip key points and the three-dimensional position information of the auxiliary key points corresponding to the hand limb.

[0132] Specifically, in the implementation scheme of this application, the direction corresponding to each hand limb is determined. Specifically, the direction of the hand limb can be determined based on the relative positional relationship between the three-dimensional positional information corresponding to the key points of the hand representing the same hand limb. For example, the direction of the index finger is determined based on the relative positional relationship between the three-dimensional positional information corresponding to the fingertip key point and the three-dimensional positional information corresponding to the interphalangeal key point among the key points of the hand representing the index finger.

[0133] It is understandable that in some application scenarios of this application, the hand limbs that need to be judged for pointing can be limited. In this case, it is only necessary to perform pointing calculation on the limited hand limbs. The specific implementation of pointing calculation will not be elaborated. For example, the index finger is limited to be judged for pointing. In this case, it is only necessary to perform the pointing calculation on the index finger.

[0134] Furthermore, based on any of the above embodiments, this application also provides a specific implementation method for determining the target gesture of the controller, specifically, see [link to relevant documentation]. Figure 5 , Figure 5 A flowchart illustrating one embodiment of the screen interaction method provided in this application for determining the target gesture of the controller is shown, specifically including steps S501-S504:

[0135] S501. Determine the first evaluation information based on the curvature of the hand limb and the standard curvature of the hand limb in each preset gesture.

[0136] The first evaluation information corresponds to the initial gesture set of the hand limbs. It can be understood that an initial gesture set may correspond to different directions and movement trends, which correspond to different preset gestures. This solution determines the initial gesture set first based on the curvature of the hand limbs and the standard curvature of the hand limbs in each preset gesture.

[0137] Specifically, in the embodiments of this application, the curvature of each hand limb is determined to be straight, bent, or semi-bent, and matched with the standard curvature of each hand limb in the preset gestures to determine whether it is straight, bent, or semi-bent, and then an initial set of gestures with the same curvature as the hand limbs is selected from the preset gestures.

[0138] S502. Determine the second evaluation information based on the direction of the hand limbs and the standard direction of the hand limbs in each preset gesture.

[0139] The second evaluation information is the target standard pointing that has the highest similarity to the pointing of the hand limb.

[0140] S503. Based on the movement trend of the hand limbs and the standard movement trend of the hand limbs in each preset gesture, determine the third evaluation information.

[0141] The third evaluation information is the target standard movement trend that has the highest similarity to the movement trend of the hand limb.

[0142] S504. Determine the target gesture of the controller based on at least one of the first evaluation information, the second evaluation information, and the third evaluation information.

[0143] Specifically, based on the above implementation scheme, before determining the target gesture of the controller according to the target pose information of the hand limbs and the standard pose information of the hand limbs in each preset gesture, the controller obtains the gesture type of each preset gesture, the standard pose information corresponding to the gesture type, and the standard evaluation information corresponding to the gesture type. The standard pose information includes at least one of standard curvature, standard pointing, and standard movement trend. For example, for the action of pointing to the left, the determination condition is: the thumb is straight, the other four fingers are bent, and the thumb points to the left. For the action of spreading five fingers, the determination condition is: all fingers are straight, and the direction is arbitrary.

[0144] In this scheme, preset gestures are characterized by at least one of standard curvature, standard pointing, and standard motion trend, ensuring the flexibility of preset gesture storage and the convenience of modification.

[0145] Specifically, in one embodiment of this application, a target gesture is selected from the first preset gestures by comparing the similarity between the direction of the hand limbs in the target pose information and the standard direction of the hand limbs in the initial gesture set. For example, a target gesture with the thumb extended, the other four fingers bent, and the thumb extended to the left is selected.

[0146] Specifically, in some other embodiments of this application, a target gesture is selected from the first preset gestures by comparing the similarity between the pointing and movement trends of the hand limbs in the target pose information and the standard pointing and movement trends of the hand limbs in the initial gesture set. For example, a target gesture is selected where the thumb is straight, the other four fingers are bent, and the thumb is straightened to the left and slid to the left.

[0147] Specifically, in some other embodiments of this application, target gestures are selected from the first preset gestures by comparing the similarity between the movement trend of the hand limbs in the target pose information and the standard movement trend of the hand limbs in the initial gesture set. For example, target gestures with the thumb extended, the index finger straightened, the other three fingers bent, and the thumb and index finger gradually moving away from or towards each other (e.g., the movement trend of increasing distance between the key points of the thumb and index finger tips) can be selected. The finger commands of this trend can be used to control page zooming, etc., further enhancing the flexibility and diversity of gesture recognition.

[0148] Furthermore, in one embodiment of this application, determining the three-dimensional position information of the key hand points specifically includes the following steps:

[0149] (1) The hand key points are predicted in three dimensions using a hand joint model to obtain the initial three-dimensional position information of the hand key points;

[0150] (2) Based on the camera intrinsic parameters corresponding to the target image, the initial three-dimensional position information of the hand key points is mapped in two dimensions to obtain the predicted key points;

[0151] (3) If the distance between the predicted key point and the hand key point is less than a preset distance threshold, then the initial three-dimensional position information is set as the three-dimensional position information of the hand key point.

[0152] Specifically, in this embodiment, after predicting the 3D information of the hand key points using a hand joint model, the 3D position P (initial 3D position information) of each hand key point in the camera coordinate system is calculated. Combined with the camera's intrinsic parameters, this position is projected onto the 2D coordinates in the image plane. The 2D Euclidean distance between this projected 2D coordinate and the model's 2D detection result is calculated as the reprojection error. If the distance between the predicted key point and the hand key point is greater than a preset distance threshold, the projection error is minimized for model optimization. Based on the optimized hand joint model, the 3D information prediction of the hand key points is performed again until the distance between the predicted key point and the hand key point is less than the preset distance threshold. Then, the initial 3D position information is set as the 3D position information of the hand key point.

[0153] This application also provides a specific implementation scheme for determining key points of the hand, see [link to relevant documentation]. Figure 6 , Figure 6 This is a schematic flowchart of one embodiment of the screen interaction method provided in this application for determining hand key points, specifically including steps S601-S604:

[0154] S601. Obtain the target image for controlling the target screen.

[0155] Specifically, in this embodiment of the application, a target image is acquired by a camera device installed on the target screen, and the target image includes at least one user.

[0156] S602. Perform human detection on the target image using a human detection model to determine the main controller in the target image and the hand area of ​​the main controller.

[0157] Specifically, in this embodiment, a human detection model is used to detect human bodies in the target image. This model uses a pre-trained deep learning model to detect the location regions of all human bodies and the location regions of each person's hands from the target image in the video stream. Here, the location region is defined as a rectangular area on the target image, typically represented by the top left corner plus width and height, the top left corner plus the top right corner, or the center point plus width and height. This application does not require a specific representation.

[0158] S603. The key points of the hand region are detected by the hand key point detection model to determine the two-dimensional position information of the hand joints of the controller and the confidence level of the two-dimensional position information.

[0159] Among them, the confidence level of the two-dimensional position information is used to evaluate the accuracy of the hand joints. It is understandable that due to the occlusion of the hand in the image and the incompleteness of the image information, the accuracy of the two-dimensional position information of the hand joints may be biased. Therefore, the corresponding confidence level is output as auxiliary information to determine the key points of the hand.

[0160] Specifically, in the implementation scheme of this application, the hand keypoint model uses a pre-trained deep learning model to extract hand keypoints from the hand location region in the original video stream based on the results of the human detection model. Here, hand keypoints refer to the positional information of the corresponding joints of the hand, typically consisting of 21 or 26 points, with each keypoint corresponding to a fixed joint. The positional information of the keypoints can be 2D coordinates or may include relative / estimated depth information (i.e., 2.5D coordinates).

[0161] It should be noted that this application does not restrict the deep learning model. It can be the classic OpenPose network (OpenPose is a deep learning-based multi-person pose estimation system designed to detect and track key points of the human body from images or videos. Its core is a deep neural network capable of end-to-end pose estimation), the lightweight AlphaPose (AlphaPose is a deep learning-based multi-person pose estimation system designed to detect and track key points of the human body from images or videos, similar to OpenPose), the high-precision HRNet (HRNet: High-Resolution Network, a deep learning network for human pose estimation that offers better performance while preserving higher resolution information compared to traditional pose estimation methods), and other similar networks. The network must contain at least the 2D coordinate information of all or some of the key points, and may also provide additional information such as confidence level, relative / estimated depth, relative distance between finger joints, and predicted values ​​of finger joint rotation angles.

[0162] That is, it can be understood that, in the implementation scheme of this application, the two-dimensional position information of the hand joints of the controller determined by the hand key point detection model can be 2D coordinate information or 2.5D coordinate information.

[0163] Understandably, this application does not impose restrictions on the deep learning model. It can be a single model simultaneously detecting the person and corresponding hand region (one-stage), or multiple models detecting the person and hand region separately (multi-stage). In multi-stage models, hand region detection can use the entire image input or be performed individually for each person based on the detected human body region. There are also no restrictions on the network structure of the deep learning model. It can be a lightweight YOLO ("YOLO" is a popular object detection algorithm, short for "You Only Look Once." It is a real-time object detection algorithm that can quickly and accurately detect multiple objects in an image or video, providing bounding boxes and class labels for each object.) series of networks, or EfficientDet (EfficientDet is an object detection model based on the EfficientNet architecture, designed to achieve higher efficiency and accuracy in object detection tasks. It employs a series of innovative technologies, including network structure design, Feature Pyramid Network (FPN), and Bidirectional Feature Network (BiFPN), to achieve excellent performance in object detection tasks), or Transformer-based ViTDet or DETR, etc.

[0164] ViTDet is an object detection model based on the Vision Transformer (ViT), applying the Transformer architecture to object detection tasks. Unlike traditional Convolutional Neural Networks (CNNs), ViTDet uses a global self-attention mechanism to capture global contextual information in images. ViTDet uses several variations to adapt to object detection tasks; for example, it adds positional encoding information to the input image to preserve the spatial structure of the image. By introducing object detection-related designs into the Transformer architecture, ViTDet achieves performance comparable to or even better than traditional object detection models on some datasets. DETR is an end-to-end object detection model that does not require traditional anchor boxes, candidate boxes, etc., but directly outputs the objects in the image and their locations through the Transformer architecture. DETR uses an encoder-decoder architecture, where the encoder extracts features from the input image, and the decoder predicts the object's category and location. DETR uses a self-attention mechanism to globally model the features of the input image, thus effectively detecting objects at different scales and resolutions. DETR demonstrates high performance on some commonly used object detection datasets and has end-to-end advantages, making the training and inference processes simpler and more efficient.

[0165] S604. Determine the key points of the hand based on the two-dimensional position information and the confidence level corresponding to each two-dimensional position information.

[0166] Furthermore, in the embodiments of this application, the joint information can be further corrected by the model based on the confidence level of the two-dimensional position information and the hand constraint relationship between the two-dimensional position information, thereby determining the corrected key points of the hand.

[0167] In this implementation scheme, the hand region and key points of the hand are detected by a model, which improves the efficiency of key point detection. At the same time, the key points of the hand are determined based on the confidence level corresponding to the two-dimensional position information, avoiding the limitations of image key point detection and ensuring the rationality and accuracy of the key points of the hand.

[0168] Furthermore, based on the above implementation scheme, this application also provides an implementation scheme for a screen interaction method, as detailed below. Figure 7 , Figure 7 The flowchart of another implementation scheme of the screen interaction method provided in this application is as follows: Specifically, in the screen interaction method, a human detection model is used to detect all people (users) in the scene corresponding to the target image, obtaining the position areas of all people in the scene and the position areas of each person's hands. If no controller is currently selected, the controller detection process begins; if a controller already exists, the step of obtaining the controller's hand key points for controlling the target screen is executed, entering the controller tracking process, which only requires hand key point detection and gesture recognition for the controller. That is, if a controller is detected, a hand key point model is used to determine the hand key points, and then the target gesture is recognized based on the hand key points, and the controller information is recorded (the target gesture recognition step is executed repeatedly). If a gesture is triggered, the target screen is interactively controlled based on the target gesture.

[0169] Specifically, the master controller detection process includes: treating each user as a candidate master controller, using a hand keypoint detection model to detect hand keypoints in the hand area of ​​each candidate master controller, and sequentially calling a hand keypoint deep learning module to recognize hand movements based on the 3D position information of each hand keypoint for both hands of each candidate master controller, thereby calculating hand gestures. If the calculated gesture matches a predefined master controller gesture, then that candidate is selected as the master controller. In subsequent algorithm processes, only the interaction actions of this master controller will be processed, ignoring the interaction actions of other people (if any) in the video frame.

[0170] Currently, the commercial display industry generally does not incorporate air gesture control technology. Supporting gesture control via built-in or external cameras will provide a significant competitive advantage without incurring excessive additional costs. The screen interaction method provided in this application acquires key hand points of the controller controlling the target screen. These key hand points are determined through user localization and skeletal tracking of a target image containing the controller. Based on the three-dimensional position information of each key hand point, the target pose information of the controller's hand limbs is determined, including at least one of curvature, pointing, and movement trend. Based on the target pose information of the hand limbs and the standard pose information of the hand limbs in various preset gestures, the target gesture of the controller is determined. Interactive control of the target screen is performed in response to the operation event corresponding to the target gesture. The algorithm of this solution is based solely on a monocular RGB camera. Compared to other products on the market that require infrared, depth, and other sensors, this sensor is simpler and lower in cost, improving product competitiveness and achieving a leading domestic level in technology. Furthermore, this solution uses trigger gesture recognition to identify the controller, making the algorithm's application scenarios not limited to single-person scenarios. This approach improves usability in multi-user scenarios, avoiding accidental operations caused by multiple users triggering gestures. It maintains operational continuity by continuously tracking the same controller, preventing interference from others entering the camera's field of view. Furthermore, this solution uses 3D keypoint information to match preset gestures, offering greater algorithmic flexibility compared to methods using neural networks for gesture classification. Adding or adjusting the preset gesture library does not require retraining the algorithm model, improving its versatility and expanding its applicability.

[0171] To better implement the screen interaction method in the embodiments of this application, a screen interaction device is also provided in the embodiments of this application, such as... Figure 8 As shown, the screen interaction device includes modules 801-804:

[0172] The acquisition module 801 is used to acquire the key hand points of the main controller controlling the target screen. The key hand points are determined by user positioning and skeleton tracking of the target image containing the main controller.

[0173] The pose determination module 802 is used to determine the target pose information of the hand limbs of the master controller based on the three-dimensional position information of each of the key hand points. The target pose information includes at least one of bending degree, pointing and movement trend.

[0174] The gesture determination module 803 is used to determine the target gesture of the main controller based on the target pose information of the hand limbs and the standard pose information of the hand limbs in each preset gesture.

[0175] The interactive control module 804 is used to interactively control the target screen in response to the operation event corresponding to the target gesture.

[0176] This solution determines at least one target pose information from three-dimensional space, including the bending degree, pointing direction, and movement trend of the hand limbs. Further, based on the target pose information and the standard pose information of the hand limbs in each preset gesture, it determines the target gesture of the controller to achieve gesture-based interactive control. This avoids the limitations of two-dimensional space, ensures the flexibility of the controller's position movement during interaction, and eliminates the need for additional interactive devices such as mice, reducing equipment configuration costs. Furthermore, by analyzing the pointing direction of the hand in three-dimensional space to determine the target interaction position, it ensures the accuracy of the interaction and improves interaction performance. In addition, the implementation of this application uses standard pose information to represent the hand limbs in the preset gestures, resulting in greater algorithm flexibility. Moreover, it eliminates the need to retrain the algorithm model when adding or adjusting the preset gesture library, improving the algorithm's versatility and expanding its applicability.

[0177] In one possible implementation of this application, the pose determination module 802 is used to determine the target pose information of the controller's hand limbs based on the three-dimensional position information of each of the hand key points, including:

[0178] Based on the three-dimensional position information of the target hand key points, the curvature and orientation of the hand limb corresponding to the target hand key points are determined, and the target hand key points are at least two hand key points on the same hand limb;

[0179] The movement trend of the hand limb is determined based on the curvature and orientation of the hand limb in at least two target images;

[0180] The bending degree of the hand limb, the pointing direction, and the movement trend are set as the target pose information of the hand limb.

[0181] In one possible implementation of this application, the pose determination module 802 is used to determine the curvature and orientation of the hand limb corresponding to the target hand key points based on the three-dimensional position information of the target hand key points, including:

[0182] The bending angle of each segment is calculated based on the three-dimensional position information of every two adjacent key points of the same hand limb.

[0183] The bending angle of each segment in each hand limb is counted to obtain the bending degree of the hand limb;

[0184] If the curvature of the hand limb is less than the preset curvature of the hand limb, then the direction of the hand limb is calculated based on the three-dimensional position information of the main control key point and the three-dimensional position information of the auxiliary key point. The auxiliary key point includes at least one of the wrist key point, interphalangeal key point and palmodigital key point corresponding to the hand limb.

[0185] This solution determines the degree of curvature of each hand limb based on the relationship between the curvature of each limb and the preset curvature corresponding to each degree of curvature. It further sets the curvature of the hand limb as the curvature, that is, it classifies the curvature of the hand limb and calculates and determines the direction, thereby improving the fault tolerance of the curvature confirmation of the hand limb and avoiding gesture recognition failure due to differences in the gesture habits of different people.

[0186] In one possible implementation of this application, the gesture determination module 803 is used to determine the target gesture of the controller based on the target pose information of the hand limbs and the standard pose information of the hand limbs in each preset gesture, including:

[0187] The first evaluation information is determined based on the curvature of the hand limbs and the standard curvature of the hand limbs in each preset gesture.

[0188] Based on the direction of the hand limbs and the standard direction of the hand limbs in each preset gesture, determine the second evaluation information; or,

[0189] The third evaluation information is determined based on the movement trends of the hand limbs and the standard movement trends of the hand limbs in each preset gesture.

[0190] The target gesture of the controller is determined based on at least one of the first evaluation information, the second evaluation information, and the third evaluation information.

[0191] This solution can enhance the flexibility and versatility of gesture recognition.

[0192] In one possible implementation of this application, before the gesture determination module determines the target gesture of the controller based on the target pose information of the hand limbs and the standard pose information of the hand limbs in each preset gesture, it further includes a function for:

[0193] Obtain the gesture type of each preset gesture, and the standard pose information corresponding to the gesture type;

[0194] The standard pose information includes at least one of standard curvature, standard orientation, and standard motion tendency.

[0195] In one possible implementation of this application, after the acquisition module 801 acquires the key points of the hand of the controller of the target screen, it further includes a module for:

[0196] The hand joint model is used to predict the three-dimensional information of the key points of the hand to obtain the initial three-dimensional position information of the key points of the hand.

[0197] Based on the camera intrinsic parameters corresponding to the target image, the initial three-dimensional position information of the hand key points is mapped in two dimensions to obtain the predicted key points.

[0198] If the distance between the predicted key point and the hand key point is less than a preset distance threshold, then the initial three-dimensional position information is set as the three-dimensional position information of the hand key point.

[0199] This solution improves the accuracy of three-dimensional position information by enhancing the model precision of the hand joint model.

[0200] In one possible implementation of this application, the acquisition module 801 is used to acquire key hand points of the controller controlling the target screen, including:

[0201] Obtain the target image corresponding to the target screen;

[0202] Human detection is performed on the target image using a human detection model to identify the main controller in the target image and the hand area of ​​the main controller;

[0203] The key points of the hand region are detected by the hand key point detection model to determine the two-dimensional position information of the hand joints of the controller, as well as the confidence level of the two-dimensional position information;

[0204] Based on the two-dimensional position information and the confidence level corresponding to each two-dimensional position information, key points of the hand are determined.

[0205] This solution uses a model to detect the hand region and key points, improving the efficiency of key point detection. At the same time, it determines the key points of the hand based on the confidence level corresponding to the two-dimensional position information, avoiding the limitations of image key point detection and ensuring the rationality and accuracy of the key points of the hand.

[0206] The screen interaction device provided in this application includes an acquisition module for acquiring key hand points of the controller who controls the target screen. These key hand points are determined through user localization and skeletal tracking of a target image containing the controller. A pose determination module is used to determine the target pose information of the controller's hand limbs based on the three-dimensional position information of each key hand point. The target pose information includes at least one of curvature, pointing direction, and movement trend. A gesture determination module is used to determine the controller's target gesture based on the target pose information of the hand limbs and the standard pose information of the hand limbs in each preset gesture. An interaction control module is used to interactively control the target screen in response to an operation event corresponding to the target gesture. This solution determines at least one target pose information from the bending degree, pointing direction, and movement trend of the hand limbs in three-dimensional space. Further, based on the target pose information and the standard pose information of the hand limbs in each preset gesture, it determines the target gesture of the controller to achieve gesture-based interactive control. This avoids the limitations of two-dimensional space, ensures the flexibility of the controller's position movement during interaction, and eliminates the need for additional interactive devices such as mice, reducing equipment configuration costs. Furthermore, by analyzing the pointing direction of the hand in three-dimensional space to determine the target interaction position, it ensures the accuracy of the interaction and thus improves interaction performance.

[0207] Furthermore, it is understood that in some other embodiments of this application, a display device is also provided, which integrates any of the screen interaction devices provided in the embodiments of the present invention, the display device comprising:

[0208] One or more processors;

[0209] Memory; and

[0210] One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor of the steps in the screen interaction method described in any of the above embodiments of the screen interaction method.

[0211] As can be seen from the above embodiments, in some embodiments of this application, the processors and memory in the display device are integrated on the circuit board body included in the display device, and the circuit board body is disposed in the display device.

[0212] It is understood that in some other embodiments of this application, the processor and memory in the display device may not be integrated into the display device, that is, the processor and memory are respectively disposed in the display device as components of the display device.

[0213] like Figure 9 As shown, Figure 9This is a schematic diagram of an embodiment of the display device provided in this application.

[0214] Specifically, a display device may include components such as a processor 1001 with one or more processing cores, a memory 1002 with one or more computer-readable storage media, a power supply 1003, and an input unit 1004. Those skilled in the art will understand that... Figure 9 The display device structure shown does not constitute a limitation on the display device and may include more or fewer components than shown, or combine certain components, or have different component arrangements. Wherein:

[0215] The processor 1001 is the image processing center, connecting various parts of the display device via various interfaces and lines. It executes software programs and / or modules stored in the memory 1002, and calls data stored in the memory 1002, to perform various functions and process data of the display device, thereby providing overall monitoring of the display device. It is understood that the processor 1001 communicates with the controller via signal transmission. Optionally, the processor 1001 may include one or more processing cores; preferably, the processor 1001 may integrate an application processor and a modem processor, wherein the application processor mainly handles the operating system, user interface, and applications, while the modem processor mainly handles wireless communication. It is understood that the aforementioned modem processor may also not be integrated into the processor 1001.

[0216] The memory 1002 can be used to store software programs and modules. The processor 1001 executes various functional applications and data processing by running the software programs and modules stored in the memory 1002. The memory 1002 may mainly include a program storage area and a data storage area. The program storage area may store the operating system, application programs required for at least one function (such as sound playback function, image playback function, etc.), etc.; the data storage area may store data created according to the use of the display device, etc. In addition, the memory 1002 may include high-speed random access memory, and may also include non-volatile memory, such as at least one disk storage device, flash memory device, or other volatile solid-state storage device. Accordingly, the memory 1002 may also include a memory controller to provide the processor 1001 with access to the memory 1002.

[0217] In some embodiments of this application, the screen interaction device can be implemented as a computer program, and the computer program can be implemented as follows: Figure 9 The device operates on the display device shown. The display device's memory can store the various program modules that make up the screen interaction device, for example, Figure 8The diagram shows a response acquisition module 801, a pose determination module 802, a gesture determination module 803, and an interaction control module 804. The computer program comprised of these modules causes the processor to execute the steps of the screen interaction methods in the various embodiments of this application described in this specification.

[0218] For example, Figure 9 The display device shown can be used as follows Figure 8 The response acquisition module 801 in the screen interaction device shown executes step S201. The display device can execute step S202 through the pose determination module 802. The display device can execute step S203 through the gesture determination module 803. The display device can execute step S204 through the interaction control module 804. The display device includes a processor, memory, and network interface connected via a system bus. The processor of the display device provides computing and control capabilities. The memory of the display device includes a non-volatile storage medium and internal memory. The non-volatile storage medium stores the operating system and computer programs. The internal memory provides an environment for the operation of the operating system and computer programs in the non-volatile storage medium. The network interface of the display device is used to communicate with external display devices via a network connection. When the computer program is executed by the processor, it implements a screen interaction method.

[0219] The display device also includes a power supply 1003 that supplies power to the various components. Preferably, the power supply 1003 can be logically connected to the processor 1001 through a power management system, thereby enabling functions such as charging, discharging, and power consumption management through the power management system. The power supply 1003 may also include one or more DC or AC power supplies, recharging systems, power fault detection circuits, power converters or inverters, power status indicators, and other arbitrary components.

[0220] The display device may also include an input unit 1004, which can be used to receive input digital or character information, and generate keyboard, mouse, joystick, optical or trackball signal inputs related to user settings and function control.

[0221] Although not shown, the display device may also include display units, etc., which will not be described in detail here. Specifically, in this embodiment, the processor 1001 in the display device loads the executable files corresponding to the processes of one or more application programs into the memory 1002 according to the following instructions, and the processor 1001 runs the application programs stored in the memory 1002 to realize various functions, as follows:

[0222] The key hand points of the controller who controls the target screen are obtained, and the key hand points are determined by user localization and skeleton tracking of the target image containing the controller.

[0223] Based on the three-dimensional position information of each of the key hand points, the target pose information of the controller's hand limbs is determined, and the target pose information includes at least one of bending degree, pointing and movement trend;

[0224] Based on the target posture information of the hand limbs and the standard posture information of the hand limbs in each preset gesture, the target gesture of the controller is determined;

[0225] The target screen is interactively controlled in response to the operation event corresponding to the target gesture.

[0226] Those skilled in the art will understand that all or part of the steps in the various methods of the above embodiments can be performed by instructions, or by instructions controlling related hardware. These instructions can be stored in a computer-readable storage medium and loaded and executed by a processor.

[0227] Therefore, embodiments of the present invention provide a computer-readable storage medium, which may include: read-only memory (ROM), random access memory (RAM), magnetic disk, or optical disk, etc. A computer program is stored thereon, which is loaded by a processor to execute the steps in any of the screen interaction methods provided in the embodiments of the present invention. For example, the computer program loaded by the processor can execute the following steps:

[0228] The key hand points of the controller who controls the target screen are obtained, and the key hand points are determined by user localization and skeleton tracking of the target image containing the controller.

[0229] Based on the three-dimensional position information of each of the key hand points, the target pose information of the controller's hand limbs is determined, and the target pose information includes at least one of bending degree, pointing and movement trend;

[0230] Based on the target posture information of the hand limbs and the standard posture information of the hand limbs in each preset gesture, the target gesture of the controller is determined;

[0231] The target screen is interactively controlled in response to the operation event corresponding to the target gesture.

[0232] In the above embodiments, the descriptions of each embodiment have different focuses. For parts not described in detail in a certain embodiment, please refer to the detailed descriptions of other embodiments above, which will not be repeated here.

[0233] In practice, each of the above units or structures can be implemented as an independent entity or can be arbitrarily combined to be implemented as the same or several entities. For the specific implementation of each of the above units or structures, please refer to the previous method embodiments, which will not be repeated here.

[0234] For details on the implementation of each of the above operations, please refer to the previous examples, which will not be repeated here.

[0235] The foregoing has provided a detailed description of a screen interaction method, apparatus, display device, and storage medium provided in the embodiments of this application. Specific examples have been used to illustrate the principles and implementation methods of the present invention. The descriptions of the above embodiments are only for the purpose of helping to understand the method and core ideas of the present invention. At the same time, for those skilled in the art, there will be changes in the specific implementation methods and application scope based on the ideas of the present invention. Therefore, the content of this specification should not be construed as a limitation of the present invention.

Claims

1. A screen interaction method, characterized in that, include: The key hand points of the controller who controls the target screen are obtained, and the key hand points are determined by user localization and skeleton tracking of the target image containing the controller. The bending angle of each segment is calculated based on the three-dimensional position information of every two adjacent target hand key points on the same hand limb. The target hand key points are at least two hand key points on the same hand limb. The bending angle of each segment in each hand limb is counted to obtain the bending degree of the hand limb; If the curvature of the hand limb is less than the preset curvature of the hand limb, then the direction of the hand limb is calculated based on the three-dimensional position information of the main control key point and the three-dimensional position information of the auxiliary key point. The auxiliary key point includes at least one of the wrist key point, interphalangeal key point and palmodigital key point corresponding to the hand limb. The movement trend of the hand limb is determined based on the curvature and orientation of the hand limb in at least two target images; The bending degree of the hand limb, the pointing direction, and the movement trend are set as the target pose information of the hand limb; Based on the target posture information of the hand limbs and the standard posture information of the hand limbs in each preset gesture, the target gesture of the controller is determined; The target screen is interactively controlled in response to the operation event corresponding to the target gesture.

2. The screen interaction method according to claim 1, characterized in that, The step of determining the target gesture of the controller based on the target pose information of the hand limbs and the standard pose information of the hand limbs in each preset gesture includes: The first evaluation information is determined based on the curvature of the hand limbs and the standard curvature of the hand limbs in each preset gesture. Based on the direction of the hand limbs and the standard direction of the hand limbs in each preset gesture, determine the second evaluation information; or, The third evaluation information is determined based on the movement trends of the hand limbs and the standard movement trends of the hand limbs in each preset gesture. The target gesture of the controller is determined based on at least one of the first evaluation information, the second evaluation information, and the third evaluation information.

3. The screen interaction method according to claim 2, characterized in that, Before determining the target gesture of the controller based on the target pose information of the hand limbs and the standard pose information of the hand limbs in each preset gesture, the process includes: Obtain the gesture type of each preset gesture, and the standard pose information corresponding to the gesture type; The standard pose information includes at least one of standard curvature, standard orientation, and standard motion tendency.

4. The screen interaction method according to claim 1, characterized in that, After acquiring the key hand points of the controller who controls the target screen, the method includes: The hand joint model is used to predict the three-dimensional information of the key points of the hand to obtain the initial three-dimensional position information of the key points of the hand. Based on the camera intrinsic parameters corresponding to the target image, the initial three-dimensional position information of the hand key points is mapped in two dimensions to obtain the predicted key points. If the distance between the predicted key point and the hand key point is less than a preset distance threshold, then the initial three-dimensional position information is set as the three-dimensional position information of the hand key point.

5. The screen interaction method according to any one of claims 1-4, characterized in that, The acquisition of the key hand points of the controller who controls the target screen includes: Obtain the target image corresponding to the target screen; Human detection is performed on the target image using a human detection model to identify the main controller in the target image and the hand area of ​​the main controller; The key points of the hand region are detected by the hand key point detection model to determine the two-dimensional position information of the hand joints of the controller, as well as the confidence level of the two-dimensional position information; Based on the two-dimensional position information and the confidence level corresponding to each two-dimensional position information, key points of the hand are determined.

6. A screen interaction device, characterized in that, The device includes: The acquisition module is used to acquire the key hand points of the master controller who controls the target screen. The key hand points are determined by user positioning and skeleton tracking of the target image containing the master controller. The pose determination module is used to calculate the bending angle of each segment based on the three-dimensional position information of every two adjacent target hand key points on the same hand limb. The target hand key points are at least two hand key points on the same hand limb. The bending angle of each segment in each hand limb is counted to obtain the bending degree of the hand limb. If the bending degree of the hand limb is less than the preset bending degree of the hand limb, the pointing of the hand limb is calculated based on the three-dimensional position information of the main control key point and the three-dimensional position information of the auxiliary key points of the hand limb. The auxiliary key points include at least one of the wrist key point, interphalangeal key point, and palmodigital key point corresponding to the hand limb. The movement trend of the hand limb is determined based on the bending degree and pointing of the hand limb in at least two target images. The bending degree, pointing, and movement trend of the hand limb are set as the target pose information of the hand limb. The gesture determination module is used to determine the target gesture of the controller based on the target pose information of the hand limbs and the standard pose information of the hand limbs in each preset gesture. An interactive control module is used to interactively control the target screen in response to an operation event corresponding to the target gesture.

7. A display device, characterized in that, The display device includes: One or more processors; Memory; and One or more applications, wherein the one or more applications are stored in the memory and configured to be executed by the processor to implement the steps of the screen interaction method according to any one of claims 1 to 5.

8. A computer-readable storage medium, characterized in that, It stores a computer program, which is loaded by a processor to perform the steps of the screen interaction method according to any one of claims 1 to 5.

Citation Information

Patent Citations

  • Method for automatically generating annotation data of hand and method for calculating skeleton length

    CN112767300A

  • Gesture recognition method and device, equipment and storage medium

    CN116597473A

  • Dynamic gesture recognition method and device, related equipment and handwriting recognition method

    CN116863541A